Skip to main content 首页 创作者 synthetic-sciences openscience long-context
long-context Extend context windows of transformer models using RoPE, YaRN, ALiBi, and position interpolation techniques. Use when processing long documents (32k-128k+ tokens), extending pre-trained models beyond original context limits, or implementing efficient positional encodings. Covers rotary embeddings, attention biases, interpolation methods, and extrapolation strategies for LLMs.
跳到安装 Skills Marketplace 发现并探索由社区构建的 Agent Skills
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/synthetic-sciences/openscience --skill long-context命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
下载 Zip 下载中... 同仓库更多 Skills Diffusion-based molecular docking. Predict protein-ligand binding poses from PDB/SMILES, confidence scores, virtual screening, for structure-based drug design. Not for affinity prediction.
Fast inference and fine-tuning platform with serverless and on-demand GPU deployments. OpenAI-compatible API for chat completions, embeddings, function calling, vision, and structured output. Supports SFT, DPO, and RL fine-tuning. SOC2 + HIPAA compliant.
Serverless inference, fine-tuning, embeddings, image generation, and batch processing on 200+ open-source models via an OpenAI-compatible API. Use when you need fast, cost-effective access to open-source LLMs without managing infrastructure.
synthetic-sciences
synthetic-sciences/openscience
打开 GitHub 仓库 name long-context description Extend context windows of transformer models using RoPE, YaRN, ALiBi, and position interpolation techniques. Use when processing long documents (32k-128k+ tokens), extending pre-trained models beyond original context limits, or implementing efficient positional encodings. Covers rotary embeddings, attention biases, interpolation methods, and extrapolation strategies for LLMs. category llm-tools version 1.0.0 author Synthetic Sciences license MIT tags ["Emerging Techniques","Long Context","RoPE","YaRN","ALiBi","Position Interpolation","Extended Context","Rotary Embeddings","Attention Bias","Context Extension","Positional Encoding"] dependencies ["transformers","torch","flash-attn"]
Long Context: Extending Transformer Context Windows
When to Use This Skill
Use Long Context techniques when you need to:
Process long documents (32k, 64k, 128k+ tokens) with transformer models
Extend context windows of pre-trained models (LLaMA, Mistral, etc.)
Implement efficient positional encodings (RoPE, ALiBi)
Train models with length extrapolation capabilities
Deploy models that handle variable-length inputs efficiently
Fine-tune existing models for longer contexts with minimal compute
Key Techniques : RoPE (Rotary Position Embeddings), YaRN, ALiBi (Attention with Linear Biases), Position Interpolation
Papers : RoFormer (arXiv 2104.09864), YaRN (arXiv 2309.00071), ALiBi (arXiv 2108.12409), Position Interpolation (arXiv 2306.15595)
Installation
pip install transformers torch
pip install einops
pip install rotary-embedding-torch
pip install flash-attn --no-build-isolation
Quick Start
RoPE (Rotary Position Embeddings)
import torch
import torch.nn as nn
class RotaryEmbedding (nn.Module):
"""Rotary Position Embeddings (RoPE)."""
def __init__ (self, dim, max_seq_len=8192 , base=10000 ):
super ().__init__()
inv_freq = 1.0 / (base ** (torch.arange(0 , dim, 2 ).float () / dim))
self .register_buffer( , inv_freq)
.max_seq_len = max_seq_len
( ):
t = torch.arange(seq_len, device=device).type_as( .inv_freq)
freqs = torch.outer(t, .inv_freq)
emb = torch.cat((freqs, freqs), dim=- )
emb.cos(), emb.sin()
( ):
x1, x2 = x.chunk( , dim=- )
torch.cat((-x2, x1), dim=- )
( ):
q_embed = (q * cos) + (rotate_half(q) * sin)
k_embed = (k * cos) + (rotate_half(k) * sin)
q_embed, k_embed
rope = RotaryEmbedding(dim= , max_seq_len= )
cos, sin = rope(seq_len= , device= )
q_rotated, k_rotated = apply_rotary_pos_emb(query, key, cos, sin)
"inv_freq"
self
def
forward
self, seq_len, device
self
self
1
return
def
rotate_half
x
"""Rotate half the hidden dimensions."""
2
1
return
1
def
apply_rotary_pos_emb
q, k, cos, sin
"""Apply rotary embeddings to queries and keys."""
return
64
8192
2048
'cuda'
ALiBi (Attention with Linear Biases) def get_alibi_slopes (num_heads ):
"""Get ALiBi slope values for each attention head."""
def get_slopes_power_of_2 (n ):
start = 2 ** (-(2 ** -(math.log2(n) - 3 )))
ratio = start
return [start * (ratio ** i) for i in range (n)]
if math.log2(num_heads).is_integer():
return get_slopes_power_of_2(num_heads)
else :
closest_power = 2 ** math.floor(math.log2(num_heads))
slopes = get_slopes_power_of_2(closest_power)
extra = get_slopes_power_of_2(2 * closest_power)
slopes.extend(extra[0 ::2 ][:num_heads - closest_power])
return slopes
def create_alibi_bias (seq_len, num_heads ):
"""Create ALiBi attention bias."""
context_position = torch.arange(seq_len)
memory_position = torch.arange(seq_len)
relative_position = memory_position[None , :] - context_position[:, None ]
slopes = torch.tensor(get_alibi_slopes(num_heads))
alibi = slopes[:, None , None ] * relative_position[None , :, :]
return alibi
num_heads = 8
seq_len = 2048
alibi_bias = create_alibi_bias(seq_len, num_heads).to('cuda' )
attn_scores = attn_scores + alibi_bias
attn_weights = torch.softmax(attn_scores, dim=-1 )
Position Interpolation for LLaMA from transformers import LlamaForCausalLM, LlamaTokenizer
model = LlamaForCausalLM.from_pretrained("meta-llama/Llama-2-7b-hf" )
model.config.rope_scaling = {
"type" : "linear" ,
"factor" : 16.0
}
model.config.rope_scaling = {
"type" : "dynamic" ,
"factor" : 16.0
}
Core Concepts
1. RoPE (Rotary Position Embeddings)
Encodes absolute position via rotation matrix
Provides relative position dependency in attention
Enables length extrapolation
Mathematical formulation:
q_m = (W_q * x_m) * e^(imθ)
k_n = (W_k * x_n) * e^(inθ)
where θ_j = base^(-2j/d) for j ∈ [0, d/2)
Decaying inter-token dependency with distance
Compatible with linear attention
Better extrapolation than absolute position encodings
2. YaRN (Yet another RoPE extensioN)
NTK-aware interpolation (Neural Tangent Kernel)
Attention temperature scaling
Efficient context extension (10× less tokens vs baselines)
yarn_config = {
"scale" : 16 ,
"original_max_position" : 2048 ,
"extrapolation_factor" : 1.0 ,
"attn_factor" : 1.0 ,
"beta_fast" : 32 ,
"beta_slow" : 1 ,
}
Extends LLaMA to 128k tokens
2.5× less training steps than baselines
State-of-the-art context window extension
3. ALiBi (Attention with Linear Biases)
No positional embeddings added to tokens
Apply distance penalty directly to attention scores
Bias proportional to key-query distance
attention_bias[i, j] = -m * |i - j|
where m = slope for each attention head
11% faster training vs sinusoidal embeddings
11% less memory usage
Strong length extrapolation (train 1k, test 2k+)
Inductive bias towards recency
4. Position Interpolation
Linearly down-scale position indices
Interpolate within trained range (vs extrapolate beyond)
Minimal fine-tuning required
# Original: position indices [0, 1, 2, ..., L]
# Extended: position indices [0, 0.5, 1.0, ..., L/2]
# (for 2× extension)
scaled_position[i] = i / extension_factor
LLaMA 7B-65B extended to 32k tokens
1000 fine-tuning steps sufficient
600× better stability than extrapolation
Method Comparison Method Max Context Training Needed Memory Extrapolation Best For RoPE 8k-32k Full pre-training Moderate Good New models YaRN 32k-128k Minimal (10× efficient) Moderate Excellent Extending existing models ALiBi Unlimited Full pre-training Low (-11%) Excellent Training from scratch Position Interpolation 32k+ Minimal (1k steps) Moderate Poor (by design) Quick extension
Implementation Patterns
HuggingFace Transformers Integration from transformers import AutoModelForCausalLM, AutoConfig
config = AutoConfig.from_pretrained("mistralai/Mistral-7B-v0.1" )
config.rope_scaling = {
"type" : "yarn" ,
"factor" : 8.0 ,
"original_max_position_embeddings" : 8192 ,
"attention_factor" : 1.0
}
model = AutoModelForCausalLM.from_config(config)
config.rope_scaling = {
"type" : "linear" ,
"factor" : 4.0
}
config.rope_scaling = {
"type" : "dynamic" ,
"factor" : 8.0
}
Custom RoPE Implementation class LongContextAttention (nn.Module):
"""Multi-head attention with RoPE."""
def __init__ (self, hidden_size, num_heads, max_seq_len=32768 ):
super ().__init__()
self .num_heads = num_heads
self .head_dim = hidden_size // num_heads
self .q_proj = nn.Linear(hidden_size, hidden_size)
self .k_proj = nn.Linear(hidden_size, hidden_size)
self .v_proj = nn.Linear(hidden_size, hidden_size)
self .o_proj = nn.Linear(hidden_size, hidden_size)
self .rotary_emb = RotaryEmbedding(
dim=self .head_dim,
max_seq_len=max_seq_len
)
def forward (self, hidden_states ):
batch_size, seq_len, _ = hidden_states.shape
q = self .q_proj(hidden_states)
k = self .k_proj(hidden_states)
v = self .v_proj(hidden_states)
q = q.view(batch_size, seq_len, self .num_heads, self .head_dim).transpose(1 , 2 )
k = k.view(batch_size, seq_len, self .num_heads, self .head_dim).transpose(1 , 2 )
v = v.view(batch_size, seq_len, self .num_heads, self .head_dim).transpose(1 , 2 )
cos, sin = self .rotary_emb(seq_len, device=hidden_states.device)
q, k = apply_rotary_pos_emb(q, k, cos, sin)
attn_output = F.scaled_dot_product_attention(q, k, v)
attn_output = attn_output.transpose(1 , 2 ).contiguous()
attn_output = attn_output.view(batch_size, seq_len, -1 )
output = self .o_proj(attn_output)
return output
Fine-tuning for Long Context
Minimal Fine-tuning (Position Interpolation) from transformers import Trainer, TrainingArguments
model.config.max_position_embeddings = 32768
model.config.rope_scaling = {"type" : "linear" , "factor" : 16.0 }
training_args = TrainingArguments(
output_dir="./llama-32k" ,
num_train_epochs=1 ,
max_steps=1000 ,
per_device_train_batch_size=1 ,
gradient_accumulation_steps=16 ,
learning_rate=2e-5 ,
warmup_steps=100 ,
logging_steps=10 ,
save_steps=500 ,
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=long_document_dataset,
)
trainer.train()
YaRN Fine-tuning
git clone https://github.com/jquesnelle/yarn
cd yarn
python scripts/train.py \
--model meta-llama/Llama-2-7b-hf \
--scale 16 \
--rope_theta 10000 \
--max_length 32768 \
--batch_size 1 \
--gradient_accumulation 16 \
--steps 400 \
--learning_rate 2e-5
Best Practices
1. Choose the Right Method
use_method = "ALiBi"
use_method = "YaRN"
use_method = "Position Interpolation"
use_method = "Linear RoPE Scaling"
2. Scaling Factor Selection
scaling_factor = 2.0
scaling_factor = 4.0
scaling_factor = 8.0
scaling_factor = 16.0
steps_needed = 100 * scaling_factor
3. Fine-tuning Data
train_data = [
{"text" : long_doc_32k_tokens},
{"text" : long_doc_24k_tokens},
{"text" : long_doc_16k_tokens},
]
train_data = [
{"text" : short_doc_2k_tokens},
]
4. Avoid Common Pitfalls
model.config.rope_scaling = {"type" : "linear" , "factor" : 16.0 }
model.config.rope_scaling = {"type" : "linear" , "factor" : 16.0 }
fine_tune(model, long_documents, steps=1000 )
scale_to_1M_tokens()
Production Deployment
Inference with Long Context from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"togethercomputer/LLaMA-2-7B-32K" ,
torch_dtype=torch.float16,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("togethercomputer/LLaMA-2-7B-32K" )
long_text = "..." * 30000
inputs = tokenizer(long_text, return_tensors="pt" , truncation=False ).to('cuda' )
outputs = model.generate(
**inputs,
max_new_tokens=512 ,
temperature=0.7 ,
)
response = tokenizer.decode(outputs[0 ], skip_special_tokens=True )
Memory Optimization
model.gradient_checkpointing_enable()
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-2-7b-hf" ,
attn_implementation="flash_attention_2" ,
torch_dtype=torch.float16
)
from vllm import LLM
llm = LLM(
model="togethercomputer/LLaMA-2-7B-32K" ,
max_model_len=32768 ,
gpu_memory_utilization=0.9
)
Resources
See Also
references/rope.md - Detailed RoPE implementation and theory
references/extension_methods.md - YaRN, ALiBi, Position Interpolation comparisons
references/fine_tuning.md - Complete fine-tuning guide for context extension