Abhi
@abhiram_ai_guy
Skill 공개 소식언어 모델 추론 최적화
게시물 요약 · 영어
A skill for tuning self-hosted LLM inference around the model, hardware, API requirements, and serving workload, including parallelism, KV tiers, and prefill/decode separation.
메뉴
SkillsMP는 abhiram1809/inference-god-mode에서 1개의 skill을 수집했습니다. skill을 열어 소스와 세부 정보를 확인하세요.
이 프롬프트를 복사해 사용 중인 AI 어시스턴트에 입력하세요.
Follow https://skillsmp.com/skill-install/prompt.md to install Agent Skills from https://github.com/abhiram1809/inference-god-mode.A skill for tuning self-hosted LLM inference around the model, hardware, API requirements, and serving workload, including parallelism, KV tiers, and prefill/decode separation.
수집된 skill 1개 중 1개를 표시합니다.
Plan, deploy, benchmark, and tune self-hosted open-weight LLM inference on one machine or Kubernetes GPU clusters, including quantization, full-context capacity, OpenAI or Anthropic APIs, cache-aware routing, and prefill/decode disaggregation.
원문 언어: 영어