المهنة
مطوّرو البرمجيات
الوصف
Optimize LLM costs and latency through KV caching and prompt caching. Use when (1) structuring prompts for cache hits, (2) configuring API cache_control for Anthropic/Cohere/OpenAI/Gemini, (3) setting up self-hosted inference with vLLM/SGLang/Ollama, (4)…
لغة النص الأصلي: الإنجليزية
آخر تحديث