Fast structured generation and serving for LLMs with RadixAttention prefix caching. Use for JSON/regex outputs, constrained decoding, agentic workflows with tool calls, or when you want faster inference than vLLM on workloads with heavy prefix sharing. The…
原文の言語: 英語