Skip to main content
Run any Skill in Manus
with one click

llama-build

Stars0
Forks0
UpdatedJune 18, 2026 at 20:25

Build llama.cpp from source, run local LLM inference with GPU acceleration, and configure coding agents to use local models. Use this skill whenever the user wants to: compile llama.cpp, set up local AI inference, run GGUF models locally, configure llama-server or llama-cli, create model launcher scripts, work with Unsloth quantized models, or connect coding tools (Claude Code, opencode, pi.dev, Continue, Cursor) to a local llama.cpp server. Triggers on mentions of llama.cpp, GGUF, local LLM serving, Metal/ROCm/HIP/CUDA GPU backends, model quantization (Q4, Q3, IQ3, Q8, etc), MTP/Multi-Token Prediction, or running models from Hugging Face locally. Also triggers when the user asks about using local models with Claude Code, opencode, pi.dev, or any OpenAI-compatible API client.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
2 files
SKILL.md
readonly