Abhi
@abhiram_ai_guy
Anúncio de SkillOtimizar a inferência de modelos de linguagem
Resumo da publicação · inglês
A skill for tuning self-hosted LLM inference around the model, hardware, API requirements, and serving workload, including parallelism, KV tiers, and prefill/decode separation.