Abhi
@abhiram_ai_guy
Anuncio de un SkillOptimizar la inferencia de modelos de lenguaje
Resumen de la publicación · inglés
A skill for tuning self-hosted LLM inference around the model, hardware, API requirements, and serving workload, including parallelism, KV tiers, and prefill/decode separation.