Abhi
@abhiram_ai_guy
Annonce d’un SkillOptimiser l’inférence des modèles de langage
Résumé de la publication · anglais
A skill for tuning self-hosted LLM inference around the model, hardware, API requirements, and serving workload, including parallelism, KV tiers, and prefill/decode separation.