with one click
local-llm-setup
local-llm-setup contains 2 collected skills from AYastrebov, with repository-level occupation coverage and site-owned skill detail pages.
Skills in this repository
Build llama.cpp from source, run local LLM inference with GPU acceleration, and configure coding agents to use local models. Use this skill whenever the user wants to: compile llama.cpp, set up local AI inference, run GGUF models locally, configure llama-server or llama-cli, create model launcher scripts, work with Unsloth quantized models, or connect coding tools (Claude Code, opencode, pi.dev, Continue, Cursor) to a local llama.cpp server. Triggers on mentions of llama.cpp, GGUF, local LLM serving, Metal/ROCm/HIP/CUDA GPU backends, model quantization (Q4, Q3, IQ3, Q8, etc), MTP/Multi-Token Prediction, or running models from Hugging Face locally. Also triggers when the user asks about using local models with Claude Code, opencode, pi.dev, or any OpenAI-compatible API client.
Configure OpenCode with NeuralWatt as a cloud AI provider. Use this skill when the user wants to set up NeuralWatt, add Kimi/GLM/Qwen/Devstral cloud models to OpenCode, configure the NeuralWatt provider block, install the nw-usage energy script, or connect OpenCode to NeuralWatt's API. Invoke whenever the user mentions NeuralWatt, wants to add cloud models to OpenCode, or asks about configuring Kimi K2.6, GLM 5.1, Qwen, or Devstral via NeuralWatt — even if they don't use the word "NeuralWatt" explicitly.