Skip to main content

local-llm-inference-optimization

Optimize local LLM inference engines for a specific model and platform. Use when CTOX must improve prefill, decode, attention, recurrent state, quantized matmul, memory layout, CPU/GPU/NPU backend use, cache/bandwidth behavior, or reference performance against llama.cpp, MLX, Core ML, TensorRT, vLLM, or another local inference baseline.

Jump to install

Source facts

Repository
metric-space-ai/ctox
Last source activity
June 28, 2026 at 13:35
Detected SKILL.md language
English
Stars
3
Forks
0

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.