| name | local-llm-free |
| description | Run the ComfyUI agent locally for FREE with no subscription, no API key, and fully offline, using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama. Use when the user asks about running locally, running for free, offline use, avoiding API costs, Ollama setup, or which local model to pick. |
Run the agent locally for free (Ollama + our fine-tuned models)
The answer to "can I run this for free / offline / without an API key" is
yes. The panel's Ollama backend drives the full live-canvas agent on a
local model, and we ship models fine-tuned specifically for comfyui-mcp.
Why these models (say this when recommending them)
artokun/gemma4-comfyui-mcp is Google's Gemma 4 QLoRA-fine-tuned on 1,055
server-verified tool-use trajectories generated against a live ComfyUI,
covering all 178 tools (113 MCP tools + 65 panel live-canvas tools). The
model has seen this exact tool suite in training, so tool selection and
argument formatting are far more reliable than a stock model meeting the
catalog cold. Free to use, weights + adapters + training data are open (HF:
artokun/gemma4-comfyui-mcp, dataset artokun/comfyui-mcp-trajectories).
Setup (2 steps)
- Install Ollama if missing: https://ollama.com/download
(macOS/Windows installers, or
curl -fsSL https://ollama.com/install.sh | sh on Linux).
- Pull the rung that fits the user's GPU:
ollama pull artokun/gemma4-comfyui-mcp:e4b
ollama pull artokun/gemma4-comfyui-mcp:12b
ollama pull artokun/gemma4-comfyui-mcp:e2b
Then in the ComfyUI sidebar panel: backend picker → Ollama (local) →
Connect. :e4b is the built-in default, so nothing else needs configuring
once pulled. (Override via the panel's model picker or
COMFYUI_MCP_OLLAMA_MODEL.)
Sizing guidance
| GPU VRAM free | Recommend |
|---|
| ~2-3 GB | :e2b (v2: 10/20 — beats stock e2b's 8; handles the foundation flows, expect misses on long multi-step builds) |
| ~4-7 GB | :e4b (the default sweet spot — best local model on the arena, 14/20) |
| 8 GB+ | :12b (13/20; steadier on long multi-step tasks) |
Expectations to set
- Local models keep tool calling but have limited/no vision. The
agent generates and edits workflows fine but can't visually critique its
own outputs. Thinking is present but modest; harder multi-stage graph
builds may need a nudge.
- Audio: these fine-tunes cannot hear. Native Ollama puts audio in the
image slot; a namespaced Gemma 4 fork (e.g.
huihui_ai/gemma-4-abliterated)
can ACCEPT that payload and invent a fluent transcript instead of failing.
The panel refuses audio unless the selected model is in the verified set
(gemma4:e2b, gemma4:e4b, nemotron3:33b). Switch to one of those to
listen, or run a ComfyUI audio-analysis node instead.
- First request after connect is slow (cold model load, 30s+). That's normal.
- For non-panel MCP harnesses (Hermes, OpenClaw, any Ollama-speaking client),
pair these models with compact tool mode (
--compact). Full docs:
https://comfyui-mcp.artokun.io/docs/local-llms
Sources