Skip to main content 홈 크리에이터 synthetic-sciences openscience long-context
long-context Extend context windows of transformer models using RoPE, YaRN, ALiBi, and position interpolation techniques. Use when processing long documents (32k-128k+ tokens), extending pre-trained models beyond original context limits, or implementing efficient positional encodings. Covers rotary embeddings, attention biases, interpolation methods, and extrapolation strategies for LLMs.
설치로 이동 Skills Marketplace 커뮤니티가 만든 AI 스킬을 발견하고 탐색하세요.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
직접 명령은 검토 Prompt를 거치지 않습니다. 실행하기 전에 소스를 확인하세요.
npx skills add https://github.com/synthetic-sciences/openscience --skill long-context명령은 한 줄로 유지됩니다. 복사하기 전에 가로로 스크롤해 전체 내용을 확인하세요.
로컬 사본을 원하시나요? SkillsMP에서 현재 제공할 수 있는 파일을 다운로드하세요.
Zip 다운로드 다운로드 중... name long-context description Extend context windows of transformer models using RoPE, YaRN, ALiBi, and position interpolation techniques. Use when processing long documents (32k-128k+ tokens), extending pre-trained models beyond original context limits, or implementing efficient positional encodings. Covers rotary embeddings, attention biases, interpolation methods, and extrapolation strategies for LLMs. category llm-tools version 1.0.0 author Synthetic Sciences license MIT tags ["Emerging Techniques","Long Context","RoPE","YaRN","ALiBi","Position Interpolation","Extended Context","Rotary Embeddings","Attention Bias","Context Extension","Positional Encoding"] dependencies ["transformers","torch","flash-attn"]
Long Context: Extending Transformer Context Windows
When to Use This Skill
Use Long Context techniques when you need to:
Process long documents (32k, 64k, 128k+ tokens) with transformer models
Extend context windows of pre-trained models (LLaMA, Mistral, etc.)
Implement efficient positional encodings (RoPE, ALiBi)
Train models with length extrapolation capabilities
Deploy models that handle variable-length inputs efficiently
Fine-tune existing models for longer contexts with minimal compute
Key Techniques : RoPE (Rotary Position Embeddings), YaRN, ALiBi (Attention with Linear Biases), Position Interpolation
Papers : RoFormer (arXiv 2104.09864), YaRN (arXiv 2309.00071), ALiBi (arXiv 2108.12409), Position Interpolation (arXiv 2306.15595)
Installation
pip install transformers torch
pip install einops
pip install rotary-embedding-torch
pip install flash-attn --no-build-isolation
Quick Start
RoPE (Rotary Position Embeddings)
import torch
import torch.nn as nn
class RotaryEmbedding (nn.Module):
"""Rotary Position Embeddings (RoPE)."""
def __init__ (self, dim, max_seq_len=8192 , base=10000 ):
super ().__init__()
inv_freq = 1.0 / (base ** (torch.arange(0 , dim, 2 ).float () / dim))
self .register_buffer( , inv_freq)
.max_seq_len = max_seq_len
( ):
t = torch.arange(seq_len, device=device).type_as( .inv_freq)
freqs = torch.outer(t, .inv_freq)
emb = torch.cat((freqs, freqs), dim=- )
emb.cos(), emb.sin()
( ):
x1, x2 = x.chunk( , dim=- )
torch.cat((-x2, x1), dim=- )
( ):
q_embed = (q * cos) + (rotate_half(q) * sin)
k_embed = (k * cos) + (rotate_half(k) * sin)
q_embed, k_embed
rope = RotaryEmbedding(dim= , max_seq_len= )
cos, sin = rope(seq_len= , device= )
q_rotated, k_rotated = apply_rotary_pos_emb(query, key, cos, sin)
"inv_freq"
self
def
forward
self, seq_len, device
self
self
1
return
def
rotate_half
x
"""Rotate half the hidden dimensions."""
2
1
return
1
def
apply_rotary_pos_emb
q, k, cos, sin
"""Apply rotary embeddings to queries and keys."""
return
64
8192
2048
'cuda'
ALiBi (Attention with Linear Biases) def get_alibi_slopes (num_heads ):
"""Get ALiBi slope values for each attention head."""
def get_slopes_power_of_2 (n ):
start = 2 ** (-(2 ** -(math.log2(n) - 3 )))
ratio = start
return [start * (ratio ** i) for i in range (n)]
if math.log2(num_heads).is_integer():
return get_slopes_power_of_2(num_heads)
else :
closest_power = 2 ** math.floor(math.log2(num_heads))
slopes = get_slopes_power_of_2(closest_power)
extra = get_slopes_power_of_2(2 * closest_power)
slopes.extend(extra[0 ::2 ][:num_heads - closest_power])
return slopes
def create_alibi_bias (seq_len, num_heads ):
"""Create ALiBi attention bias."""
context_position = torch.arange(seq_len)
memory_position = torch.arange(seq_len)
relative_position = memory_position[None , :] - context_position[:, None ]
slopes = torch.tensor(get_alibi_slopes(num_heads))
alibi = slopes[:, None , None ] * relative_position[None , :, :]
return alibi
num_heads = 8
seq_len = 2048
alibi_bias = create_alibi_bias(seq_len, num_heads).to('cuda' )
attn_scores = attn_scores + alibi_bias
attn_weights = torch.softmax(attn_scores, dim=-1 )
Position Interpolation for LLaMA from transformers import LlamaForCausalLM, LlamaTokenizer
model = LlamaForCausalLM.from_pretrained("meta-llama/Llama-2-7b-hf" )
model.config.rope_scaling = {
"type" : "linear" ,
"factor" : 16.0
}
model.config.rope_scaling = {
"type" : "dynamic" ,
"factor" : 16.0
}
Core Concepts
1. RoPE (Rotary Position Embeddings)
Encodes absolute position via rotation matrix
Provides relative position dependency in attention
Enables length extrapolation
Mathematical formulation:
q_m = (W_q * x_m) * e^(imθ)
k_n = (W_k * x_n) * e^(inθ)
where θ_j = base^(-2j/d) for j ∈ [0, d/2)
Decaying inter-token dependency with distance
Compatible with linear attention
Better extrapolation than absolute position encodings
2. YaRN (Yet another RoPE extensioN)
NTK-aware interpolation (Neural Tangent Kernel)
Attention temperature scaling
Efficient context extension (10× less tokens vs baselines)
yarn_config = {
"scale" : 16 ,
"original_max_position" : 2048 ,
"extrapolation_factor" : 1.0 ,
"attn_factor" : 1.0 ,
"beta_fast" : 32 ,
"beta_slow" : 1 ,
}
Extends LLaMA to 128k tokens
2.5× less training steps than baselines
State-of-the-art context window extension
3. ALiBi (Attention with Linear Biases)
No positional embeddings added to tokens
Apply distance penalty directly to attention scores
Bias proportional to key-query distance
attention_bias[i, j] = -m * |i - j|
where m = slope for each attention head
11% faster training vs sinusoidal embeddings
11% less memory usage
Strong length extrapolation (train 1k, test 2k+)
Inductive bias towards recency
4. Position Interpolation
Linearly down-scale position indices
Interpolate within trained range (vs extrapolate beyond)
Minimal fine-tuning required
# Original: position indices [0, 1, 2, ..., L]
# Extended: position indices [0, 0.5, 1.0, ..., L/2]
# (for 2× extension)
scaled_position[i] = i / extension_factor
LLaMA 7B-65B extended to 32k tokens
1000 fine-tuning steps sufficient
600× better stability than extrapolation
Method Comparison Method Max Context Training Needed Memory Extrapolation Best For RoPE 8k-32k Full pre-training Moderate Good New models YaRN 32k-128k Minimal (10× efficient) Moderate Excellent Extending existing models ALiBi Unlimited Full pre-training Low (-11%) Excellent Training from scratch Position Interpolation 32k+ Minimal (1k steps) Moderate Poor (by design) Quick extension
Implementation Patterns
HuggingFace Transformers Integration from transformers import AutoModelForCausalLM, AutoConfig
config = AutoConfig.from_pretrained("mistralai/Mistral-7B-v0.1" )
config.rope_scaling = {
"type" : "yarn" ,
"factor" : 8.0 ,
"original_max_position_embeddings" : 8192 ,
"attention_factor" : 1.0
}
model = AutoModelForCausalLM.from_config(config)
config.rope_scaling = {
"type" : "linear" ,
"factor" : 4.0
}
config.rope_scaling = {
"type" : "dynamic" ,
"factor" : 8.0
}
Custom RoPE Implementation class LongContextAttention (nn.Module):
"""Multi-head attention with RoPE."""
def __init__ (self, hidden_size, num_heads, max_seq_len=32768 ):
super ().__init__()
self .num_heads = num_heads
self .head_dim = hidden_size // num_heads
self .q_proj = nn.Linear(hidden_size, hidden_size)
self .k_proj = nn.Linear(hidden_size, hidden_size)
self .v_proj = nn.Linear(hidden_size, hidden_size)
self .o_proj = nn.Linear(hidden_size, hidden_size)
self .rotary_emb = RotaryEmbedding(
dim=self .head_dim,
max_seq_len=max_seq_len
)
def forward (self, hidden_states ):
batch_size, seq_len, _ = hidden_states.shape
q = self .q_proj(hidden_states)
k = self .k_proj(hidden_states)
v = self .v_proj(hidden_states)
q = q.view(batch_size, seq_len, self .num_heads, self .head_dim).transpose(1 , 2 )
k = k.view(batch_size, seq_len, self .num_heads, self .head_dim).transpose(1 , 2 )
v = v.view(batch_size, seq_len, self .num_heads, self .head_dim).transpose(1 , 2 )
cos, sin = self .rotary_emb(seq_len, device=hidden_states.device)
q, k = apply_rotary_pos_emb(q, k, cos, sin)
attn_output = F.scaled_dot_product_attention(q, k, v)
attn_output = attn_output.transpose(1 , 2 ).contiguous()
attn_output = attn_output.view(batch_size, seq_len, -1 )
output = self .o_proj(attn_output)
return output
Fine-tuning for Long Context
Minimal Fine-tuning (Position Interpolation) from transformers import Trainer, TrainingArguments
model.config.max_position_embeddings = 32768
model.config.rope_scaling = {"type" : "linear" , "factor" : 16.0 }
training_args = TrainingArguments(
output_dir="./llama-32k" ,
num_train_epochs=1 ,
max_steps=1000 ,
per_device_train_batch_size=1 ,
gradient_accumulation_steps=16 ,
learning_rate=2e-5 ,
warmup_steps=100 ,
logging_steps=10 ,
save_steps=500 ,
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=long_document_dataset,
)
trainer.train()
YaRN Fine-tuning
git clone https://github.com/jquesnelle/yarn
cd yarn
python scripts/train.py \
--model meta-llama/Llama-2-7b-hf \
--scale 16 \
--rope_theta 10000 \
--max_length 32768 \
--batch_size 1 \
--gradient_accumulation 16 \
--steps 400 \
--learning_rate 2e-5
Best Practices
1. Choose the Right Method
use_method = "ALiBi"
use_method = "YaRN"
use_method = "Position Interpolation"
use_method = "Linear RoPE Scaling"
2. Scaling Factor Selection
scaling_factor = 2.0
scaling_factor = 4.0
scaling_factor = 8.0
scaling_factor = 16.0
steps_needed = 100 * scaling_factor
3. Fine-tuning Data
train_data = [
{"text" : long_doc_32k_tokens},
{"text" : long_doc_24k_tokens},
{"text" : long_doc_16k_tokens},
]
train_data = [
{"text" : short_doc_2k_tokens},
]
4. Avoid Common Pitfalls
model.config.rope_scaling = {"type" : "linear" , "factor" : 16.0 }
model.config.rope_scaling = {"type" : "linear" , "factor" : 16.0 }
fine_tune(model, long_documents, steps=1000 )
scale_to_1M_tokens()
Production Deployment
Inference with Long Context from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"togethercomputer/LLaMA-2-7B-32K" ,
torch_dtype=torch.float16,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("togethercomputer/LLaMA-2-7B-32K" )
long_text = "..." * 30000
inputs = tokenizer(long_text, return_tensors="pt" , truncation=False ).to('cuda' )
outputs = model.generate(
**inputs,
max_new_tokens=512 ,
temperature=0.7 ,
)
response = tokenizer.decode(outputs[0 ], skip_special_tokens=True )
Memory Optimization
model.gradient_checkpointing_enable()
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-2-7b-hf" ,
attn_implementation="flash_attention_2" ,
torch_dtype=torch.float16
)
from vllm import LLM
llm = LLM(
model="togethercomputer/LLaMA-2-7B-32K" ,
max_model_len=32768 ,
gpu_memory_utilization=0.9
)
Resources
See Also
references/rope.md - Detailed RoPE implementation and theory
references/extension_methods.md - YaRN, ALiBi, Position Interpolation comparisons
references/fine_tuning.md - Complete fine-tuning guide for context extension