Abhi
@abhiram_ai_guy
Skill release postTune LLM inference
Post summary
A skill for tuning self-hosted LLM inference around the model, hardware, API requirements, and serving workload, including parallelism, KV tiers, and prefill/decode separation.
Menu
Explore reviewed X posts about Skills, the people sharing them, and the projects they discuss.
Counts include collected and reviewed Skill posts only. Empty days mean nothing was collected here. Dates use UTC.
1 collected posts
A skill for tuning self-hosted LLM inference around the model, hardware, API requirements, and serving workload, including parallelism, KV tiers, and prefill/decode separation.