Skip to main content

more-than-quick-glance

Implement LASER-KV-style KV-cache compression for LLM inference pipelines using block-wise accumulative budgeting and hybrid exact-attention/LSH token selection. Use when: 'optimize KV cache for long context', 'compress KV cache without losing accuracy', 'implement LASER-KV', 'reduce LLM memory for 128k context', 'block-wise cache eviction strategy', 'fix attention-score greedy bias in cache pruning'.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
ndpvt-web/arxiv-claude-skills
آخر نشاط في المصدر
١٣ فبراير ٢٠٢٦ في ١٣:٣٥
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
١٤
التفرعات
٣

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.