미분류
Plan, deploy, benchmark, and tune self-hosted open-weight LLM inference on one machine or Kubernetes GPU clusters, including quantization, full-context capacity, OpenAI or Anthropic APIs, cache-aware routing, and prefill/decode disaggregation.
2026년 10월 2일