06:56 UTCExpand X postSkill release postTune LLM inferencePost summaryA skill for tuning self-hosted LLM inference around the model, hardware, API requirements, and serving workload, including parallelism, KV tiers, and prefill/decode separation.Original postView Skill details · inference-god-mode/codex-skills-inference-optimizer (abhiram1809/inference-god-mode/codex-skills-inference-optimizer)View repository · inference-god-mode (abhiram1809/inference-god-mode)