Benchmark a hybrid-attention model across the LMCache performance ladder (vLLM no-hybrid-allocator → hybrid allocator + prefix caching → hybrid allocator + LMCache) and produce a decode-throughput / TTFT / cache-hit-rate comparison. Use when asked to…
Review a GitHub pull request against the LMCache coding standards
Create a GitHub pull request from the current branch to the upstream repo, using the repo's PR template
Check local changes against the LMCache coding standards and fix issues before creating a PR