Skip to main content Inicio Creadores synthetic-sciences openscience long-context
long-context Extend context windows of transformer models using RoPE, YaRN, ALiBi, and position interpolation techniques. Use when processing long documents (32k-128k+ tokens), extending pre-trained models beyond original context limits, or implementing efficient positional encodings. Covers rotary embeddings, attention biases, interpolation methods, and extrapolation strategies for LLMs.
Ir a la instalación Skills Marketplace Descubre y explora habilidades de IA creadas por la comunidad.
Instalar con Codex o Claude Copia este prompt, pégalo en Codex, Claude u otro asistente, y deja que revise la página de la skill y la instale por ti.
Copiar promptMostrar detalles del prompt Un comando directo omite el prompt de revisión. Revisa el origen antes de ejecutarlo.
npx skills add https://github.com/synthetic-sciences/openscience --skill long-contextEl comando permanece en una sola línea. Desplázate horizontalmente para revisarlo antes de copiarlo.
¿Prefieres una copia local? Descarga los archivos que SkillsMP tiene disponibles ahora.
Descargar Zip Descargando... Más de este repositorio Diffusion-based molecular docking. Predict protein-ligand binding poses from PDB/SMILES, confidence scores, virtual screening, for structure-based drug design. Not for affinity prediction.
Fast inference and fine-tuning platform with serverless and on-demand GPU deployments. OpenAI-compatible API for chat completions, embeddings, function calling, vision, and structured output. Supports SFT, DPO, and RL fine-tuning. SOC2 + HIPAA compliant.
Serverless inference, fine-tuning, embeddings, image generation, and batch processing on 200+ open-source models via an OpenAI-compatible API. Use when you need fast, cost-effective access to open-source LLMs without managing infrastructure.
Ocupaciones relacionadas SOC
Basado en la clasificación ocupacional SOC
Explorador de archivos
4 archivos name long-context description Extend context windows of transformer models using RoPE, YaRN, ALiBi, and position interpolation techniques. Use when processing long documents (32k-128k+ tokens), extending pre-trained models beyond original context limits, or implementing efficient positional encodings. Covers rotary embeddings, attention biases, interpolation methods, and extrapolation strategies for LLMs. category llm-tools version 1.0.0 author Synthetic Sciences license MIT tags ["Emerging Techniques","Long Context","RoPE","YaRN","ALiBi","Position Interpolation","Extended Context","Rotary Embeddings","Attention Bias","Context Extension","Positional Encoding"] dependencies ["transformers","torch","flash-attn"]
Long Context: Extending Transformer Context Windows
When to Use This Skill
Use Long Context techniques when you need to:
Process long documents (32k, 64k, 128k+ tokens) with transformer models
Extend context windows of pre-trained models (LLaMA, Mistral, etc.)
Implement efficient positional encodings (RoPE, ALiBi)
Train models with length extrapolation capabilities
Deploy models that handle variable-length inputs efficiently
Fine-tune existing models for longer contexts with minimal compute
Key Techniques : RoPE (Rotary Position Embeddings), YaRN, ALiBi (Attention with Linear Biases), Position Interpolation
Papers : RoFormer (arXiv 2104.09864), YaRN (arXiv 2309.00071), ALiBi (arXiv 2108.12409), Position Interpolation (arXiv 2306.15595)
Installation
pip install transformers torch
pip install einops
pip install rotary-embedding-torch
pip install flash-attn --no-build-isolation
Quick Start
RoPE (Rotary Position Embeddings)
import torch
import torch.nn as nn
class RotaryEmbedding (nn.Module):
"""Rotary Position Embeddings (RoPE)."""
def __init__ (self, dim, max_seq_len=8192 , base=10000 ):
super ().__init__()
inv_freq = 1.0 / (base ** (torch.arange(0 , dim, 2 ).float () / dim))
self .register_buffer( , inv_freq)
.max_seq_len = max_seq_len
( ):
t = torch.arange(seq_len, device=device).type_as( .inv_freq)
freqs = torch.outer(t, .inv_freq)
emb = torch.cat((freqs, freqs), dim=- )
emb.cos(), emb.sin()
( ):
x1, x2 = x.chunk( , dim=- )
torch.cat((-x2, x1), dim=- )
( ):
q_embed = (q * cos) + (rotate_half(q) * sin)
k_embed = (k * cos) + (rotate_half(k) * sin)
q_embed, k_embed
rope = RotaryEmbedding(dim= , max_seq_len= )
cos, sin = rope(seq_len= , device= )
q_rotated, k_rotated = apply_rotary_pos_emb(query, key, cos, sin)
"inv_freq"
self
def
forward
self, seq_len, device
self
self
1
return
def
rotate_half
x
"""Rotate half the hidden dimensions."""
2
1
return
1
def
apply_rotary_pos_emb
q, k, cos, sin
"""Apply rotary embeddings to queries and keys."""
return
64
8192
2048
'cuda'
ALiBi (Attention with Linear Biases) def get_alibi_slopes (num_heads ):
"""Get ALiBi slope values for each attention head."""
def get_slopes_power_of_2 (n ):
start = 2 ** (-(2 ** -(math.log2(n) - 3 )))
ratio = start
return [start * (ratio ** i) for i in range (n)]
if math.log2(num_heads).is_integer():
return get_slopes_power_of_2(num_heads)
else :
closest_power = 2 ** math.floor(math.log2(num_heads))
slopes = get_slopes_power_of_2(closest_power)
extra = get_slopes_power_of_2(2 * closest_power)
slopes.extend(extra[0 ::2 ][:num_heads - closest_power])
return slopes
def create_alibi_bias (seq_len, num_heads ):
"""Create ALiBi attention bias."""
context_position = torch.arange(seq_len)
memory_position = torch.arange(seq_len)
relative_position = memory_position[None , :] - context_position[:, None ]
slopes = torch.tensor(get_alibi_slopes(num_heads))
alibi = slopes[:, None , None ] * relative_position[None , :, :]
return alibi
num_heads = 8
seq_len = 2048
alibi_bias = create_alibi_bias(seq_len, num_heads).to('cuda' )
attn_scores = attn_scores + alibi_bias
attn_weights = torch.softmax(attn_scores, dim=-1 )
Position Interpolation for LLaMA from transformers import LlamaForCausalLM, LlamaTokenizer
model = LlamaForCausalLM.from_pretrained("meta-llama/Llama-2-7b-hf" )
model.config.rope_scaling = {
"type" : "linear" ,
"factor" : 16.0
}
model.config.rope_scaling = {
"type" : "dynamic" ,
"factor" : 16.0
}
Core Concepts
1. RoPE (Rotary Position Embeddings)
Encodes absolute position via rotation matrix
Provides relative position dependency in attention
Enables length extrapolation
Mathematical formulation:
q_m = (W_q * x_m) * e^(imθ)
k_n = (W_k * x_n) * e^(inθ)
where θ_j = base^(-2j/d) for j ∈ [0, d/2)
Decaying inter-token dependency with distance
Compatible with linear attention
Better extrapolation than absolute position encodings
2. YaRN (Yet another RoPE extensioN)
NTK-aware interpolation (Neural Tangent Kernel)
Attention temperature scaling
Efficient context extension (10× less tokens vs baselines)
yarn_config = {
"scale" : 16 ,
"original_max_position" : 2048 ,
"extrapolation_factor" : 1.0 ,
"attn_factor" : 1.0 ,
"beta_fast" : 32 ,
"beta_slow" : 1 ,
}
Extends LLaMA to 128k tokens
2.5× less training steps than baselines
State-of-the-art context window extension
3. ALiBi (Attention with Linear Biases)
No positional embeddings added to tokens
Apply distance penalty directly to attention scores
Bias proportional to key-query distance
attention_bias[i, j] = -m * |i - j|
where m = slope for each attention head
11% faster training vs sinusoidal embeddings
11% less memory usage
Strong length extrapolation (train 1k, test 2k+)
Inductive bias towards recency
4. Position Interpolation
Linearly down-scale position indices
Interpolate within trained range (vs extrapolate beyond)
Minimal fine-tuning required
# Original: position indices [0, 1, 2, ..., L]
# Extended: position indices [0, 0.5, 1.0, ..., L/2]
# (for 2× extension)
scaled_position[i] = i / extension_factor
LLaMA 7B-65B extended to 32k tokens
1000 fine-tuning steps sufficient
600× better stability than extrapolation
Method Comparison Method Max Context Training Needed Memory Extrapolation Best For RoPE 8k-32k Full pre-training Moderate Good New models YaRN 32k-128k Minimal (10× efficient) Moderate Excellent Extending existing models ALiBi Unlimited Full pre-training Low (-11%) Excellent Training from scratch Position Interpolation 32k+ Minimal (1k steps) Moderate Poor (by design) Quick extension
Implementation Patterns
HuggingFace Transformers Integration from transformers import AutoModelForCausalLM, AutoConfig
config = AutoConfig.from_pretrained("mistralai/Mistral-7B-v0.1" )
config.rope_scaling = {
"type" : "yarn" ,
"factor" : 8.0 ,
"original_max_position_embeddings" : 8192 ,
"attention_factor" : 1.0
}
model = AutoModelForCausalLM.from_config(config)
config.rope_scaling = {
"type" : "linear" ,
"factor" : 4.0
}
config.rope_scaling = {
"type" : "dynamic" ,
"factor" : 8.0
}
Custom RoPE Implementation class LongContextAttention (nn.Module):
"""Multi-head attention with RoPE."""
def __init__ (self, hidden_size, num_heads, max_seq_len=32768 ):
super ().__init__()
self .num_heads = num_heads
self .head_dim = hidden_size // num_heads
self .q_proj = nn.Linear(hidden_size, hidden_size)
self .k_proj = nn.Linear(hidden_size, hidden_size)
self .v_proj = nn.Linear(hidden_size, hidden_size)
self .o_proj = nn.Linear(hidden_size, hidden_size)
self .rotary_emb = RotaryEmbedding(
dim=self .head_dim,
max_seq_len=max_seq_len
)
def forward (self, hidden_states ):
batch_size, seq_len, _ = hidden_states.shape
q = self .q_proj(hidden_states)
k = self .k_proj(hidden_states)
v = self .v_proj(hidden_states)
q = q.view(batch_size, seq_len, self .num_heads, self .head_dim).transpose(1 , 2 )
k = k.view(batch_size, seq_len, self .num_heads, self .head_dim).transpose(1 , 2 )
v = v.view(batch_size, seq_len, self .num_heads, self .head_dim).transpose(1 , 2 )
cos, sin = self .rotary_emb(seq_len, device=hidden_states.device)
q, k = apply_rotary_pos_emb(q, k, cos, sin)
attn_output = F.scaled_dot_product_attention(q, k, v)
attn_output = attn_output.transpose(1 , 2 ).contiguous()
attn_output = attn_output.view(batch_size, seq_len, -1 )
output = self .o_proj(attn_output)
return output
Fine-tuning for Long Context
Minimal Fine-tuning (Position Interpolation) from transformers import Trainer, TrainingArguments
model.config.max_position_embeddings = 32768
model.config.rope_scaling = {"type" : "linear" , "factor" : 16.0 }
training_args = TrainingArguments(
output_dir="./llama-32k" ,
num_train_epochs=1 ,
max_steps=1000 ,
per_device_train_batch_size=1 ,
gradient_accumulation_steps=16 ,
learning_rate=2e-5 ,
warmup_steps=100 ,
logging_steps=10 ,
save_steps=500 ,
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=long_document_dataset,
)
trainer.train()
YaRN Fine-tuning
git clone https://github.com/jquesnelle/yarn
cd yarn
python scripts/train.py \
--model meta-llama/Llama-2-7b-hf \
--scale 16 \
--rope_theta 10000 \
--max_length 32768 \
--batch_size 1 \
--gradient_accumulation 16 \
--steps 400 \
--learning_rate 2e-5
Best Practices
1. Choose the Right Method
use_method = "ALiBi"
use_method = "YaRN"
use_method = "Position Interpolation"
use_method = "Linear RoPE Scaling"
2. Scaling Factor Selection
scaling_factor = 2.0
scaling_factor = 4.0
scaling_factor = 8.0
scaling_factor = 16.0
steps_needed = 100 * scaling_factor
3. Fine-tuning Data
train_data = [
{"text" : long_doc_32k_tokens},
{"text" : long_doc_24k_tokens},
{"text" : long_doc_16k_tokens},
]
train_data = [
{"text" : short_doc_2k_tokens},
]
4. Avoid Common Pitfalls
model.config.rope_scaling = {"type" : "linear" , "factor" : 16.0 }
model.config.rope_scaling = {"type" : "linear" , "factor" : 16.0 }
fine_tune(model, long_documents, steps=1000 )
scale_to_1M_tokens()
Production Deployment
Inference with Long Context from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"togethercomputer/LLaMA-2-7B-32K" ,
torch_dtype=torch.float16,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("togethercomputer/LLaMA-2-7B-32K" )
long_text = "..." * 30000
inputs = tokenizer(long_text, return_tensors="pt" , truncation=False ).to('cuda' )
outputs = model.generate(
**inputs,
max_new_tokens=512 ,
temperature=0.7 ,
)
response = tokenizer.decode(outputs[0 ], skip_special_tokens=True )
Memory Optimization
model.gradient_checkpointing_enable()
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-2-7b-hf" ,
attn_implementation="flash_attention_2" ,
torch_dtype=torch.float16
)
from vllm import LLM
llm = LLM(
model="togethercomputer/LLaMA-2-7B-32K" ,
max_model_len=32768 ,
gpu_memory_utilization=0.9
)
Resources
See Also
references/rope.md - Detailed RoPE implementation and theory
references/extension_methods.md - YaRN, ALiBi, Position Interpolation comparisons
references/fine_tuning.md - Complete fine-tuning guide for context extension