Enforces latency/throughput/cost rules for LLM-calling code. Use when code calls Anthropic/OpenAI/Gemini/Groq, defines system prompts, tool schemas, structured output schemas, streaming handlers, or agent loops — or when user mentions latency, TTFT, TPS, token budgets, cache hit rate, thinking tokens, or "make AI faster". Provider-impartial. Framework-agnostic.
2026-04-25