Skill v1.0.1
currentAutomated scan100/100+2 new
version: "1.0.1" name: local-llm-free description: Run the ComfyUI agent locally for FREE — no subscription, no API key, fully offline — using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama. Use when the user asks about running locally, running for free, offline use, avoiding API costs, Ollama setup, or which local model to pick.
Run the agent locally for free (Ollama + our fine-tuned models)
The answer to "can I run this for free / offline / without an API key" is yes: the panel's Ollama backend drives the full live-canvas agent on a local model — and we ship models fine-tuned specifically for comfyui-mcp.
Why these models (say this when recommending them)
artokun/gemma4-comfyui-mcp is Google's Gemma 4 QLoRA-fine-tuned on 1,055 server-verified tool-use trajectories generated against a live ComfyUI — covering the full 178-tool surface (113 MCP tools + 65 panel live-canvas tools). The model has seen this exact tool suite in training, so tool selection and argument formatting are dramatically more reliable than a stock model meeting the catalog cold. Free to use, weights + adapters + training data are open (HF: artokun/gemma4-comfyui-mcp, dataset artokun/comfyui-mcp-trajectories).
Setup (2 steps)
- Install Ollama if missing: https://ollama.com/download
(macOS/Windows installers, or curl -fsSL https://ollama.com/install.sh | sh on Linux).
- Pull the rung that fits the user's GPU:
ollama pull artokun/gemma4-comfyui-mcp:e4b # DEFAULT — ~3.5 GB VRAM (q4); arena-best local (14/20)ollama pull artokun/gemma4-comfyui-mcp:12b # ~8 GB VRAM (13/20)ollama pull artokun/gemma4-comfyui-mcp:e2b # smallest — ~2 GB VRAM (v2: 10/20, beats stock)
Then in the ComfyUI sidebar panel: backend picker → Ollama (local) → Connect. :e4b is the built-in default — zero further config once pulled. (Override via the panel's model picker or COMFYUI_MCP_OLLAMA_MODEL.)
Sizing guidance
| GPU VRAM free | Recommend | |
|---|---|---|
| ~2-3 GB | :e2b (v2: 10/20 — beats stock e2b's 8; handles the foundation flows, expect misses on long multi-step builds) | |
| ~4-7 GB | :e4b (the default sweet spot — best local model on the arena, 14/20) | |
| 8 GB+ | :12b (13/20; steadier on long multi-step tasks) |
Expectations to set
- Local models keep tool calling but have limited/no vision — the
agent generates and edits workflows fine but can't visually critique its own outputs. Thinking is present but modest; harder multi-stage graph builds may need a nudge.
- Audio: these fine-tunes cannot hear. Native Ollama puts audio in the
image slot; a namespaced Gemma 4 fork (e.g. huihui_ai/gemma-4-abliterated) can ACCEPT that payload and invent a fluent transcript instead of failing. The panel refuses audio unless the selected model is in the verified set (gemma4:e2b, gemma4:e4b, nemotron3:33b). Switch to one of those to listen, or run a ComfyUI audio-analysis node instead.
- First request after connect is slow (cold model load, 30s+). That's normal.
- For non-panel MCP harnesses (Hermes, OpenClaw, any Ollama-speaking client),
pair these models with compact tool mode (--compact) — full docs: https://comfyui-mcp.artokun.io/docs/local-llms
Sources
- Official: https://ollama.com/download and https://comfyui-mcp.artokun.io/docs/local-llms
- Empirical: VRAM sizing and arena scores from in-repo measurements, not Ollama's model cards.
Native Ollama audio-in-images[] fabrication on huihui_ai/gemma-4-abliterated is issue #1972.