waymakerone local
Run Claude Code against a model on your own Mac, through a small gateway that the CLI starts and stops for you. waymakerone local and waymaker local are the same command (CLI 2.17.0 and later).
There are two ways to use it:
| Command | What happens | Use it when |
|---|---|---|
waymakerone local --offline | Every model request goes to the local model. No model requests leave your Mac; tools you approve (WebFetch, Bash, MCP servers you add) can still use the network. | You have no internet or no credits, and need to finish a task or do basic work. |
waymakerone local | Hybrid (the default). Your frontier model keeps the session, under your normal Claude sign-in. Claude Code's small/fast model slot runs locally. | You want to offset some frontier usage without giving up frontier quality on the main work. |
Local answers are lower quality than the cloud. The command says so before it starts, and --offline says so explicitly.
Before you start
- Claude Code on your PATH (
claude). - Ollama installed (
brew install ollama). If Ollama is already running, the command uses it. If nothing is answering, it startsollama serveand stops it again when you exit. It never stops an Ollama it did not start. - At least 24 GB of memory. Below that the command refuses to run.
--codingneeds 48 GB. - The model. The default is
gpt-oss:20b(about 13 GB). If it is not on your Mac, the command prints its size and the exactollama pullcommand, checks you have about 1.5× that in free disk, and asks first. It never downloads silently. Pass--yesto agree in advance.
Unlike the rest of the CLI, local needs no WAYMAKER_API_KEY.
Usage
waymakerone local [--offline] [--coding] [--yes] [--force] [--dry-run] [claude args...]
waymakerone local stats [--days N]
| Flag | What it does |
|---|---|
--offline | Everything runs locally. Every model slot points at the local model, the context is set to 64k, and your MCP servers are not loaded unless you pass your own --mcp-config. |
--coding | Use qwen3-coder:30b-a3b instead of gpt-oss:20b (48 GB or more). It is for Claude Code only: in our test it did a multi-file rename in Claude Code but failed the same rename in another harness, and its MCP tool calls were unreliable. |
--yes | Allow downloading a missing model. |
--force | Run even while another local-model job holds the machine-wide lock. |
--dry-run | Print the mode, model, endpoints and environment variable names, and start nothing. |
Anything else is passed to claude. If one of claude's own flags clashes with these, put it after --:
waymakerone local --offline -p "summarise README.md"
waymakerone local -- --help
What it measures, and what it does not
waymakerone local stats reads ~/.waymaker/local-usage.jsonl: one line per request with the time, local or upstream, the model name, token counts, duration and status. It never holds your prompts, the replies or any credentials. The file is readable by you alone (mode 0600), and moves to local-usage.jsonl.1 once it reaches 5 MB. Nothing in it is sent anywhere.
waymaker local stats --days 7
How hybrid works
Claude Code chooses a model per request by name, but has one base URL. The gateway listens on 127.0.0.1. A request for a local- model goes to Ollama. Everything else goes to api.anthropic.com unchanged, with your sign-in headers forwarded exactly as Claude Code sent them. Your credentials are never sent to Ollama.
Hybrid needs your normal Claude sign-in to work through a custom base URL. We tested this on 2026-10-11 with a subscription (OAuth) sign-in and no API key set, and it worked. Claude Code decides which requests use the small/fast slot. A trivial -p prompt used none, so how much a real session offloads has not been measured yet.
What has been tested
- Measured on one Mac only: Apple M5 Pro with 48 GB, Ollama 0.40.1, Claude Code 2.1.296, one run each (n=1).
--offline -p "reply with the single word ready"repliedreadywith the model warm. The prompt was 17,130 tokens.- In hybrid, a frontier request went upstream and returned 200 with the subscription sign-in. A small/fast-slot request ran locally: 21,430 prompt tokens in 21.6 s.
- Not tested: 24 GB Macs. The 24 GB tier is inferred, not measured. Intel Macs are not supported.
Limits
- 64k context. The gateway refuses a local request it estimates is larger than the model's context, with the error Claude Code uses to compact the conversation. Without that refusal, Ollama would silently drop the start of the conversation. The estimate errs high, so compaction can start a little early.
- No tool profile yet.
--offlineloads no MCP servers by default, because a full catalogue does not fit in 64k. A curated local tool set comes in a later release. - One local model at a time. The command takes the same machine-wide lock as other local-model jobs (
~/.fleet/locks/ollama). In--offlineit refuses while another job holds the lock. In hybrid it still starts, sends everything to your frontier model, and tells you so. - Writes still ask. Nothing is auto-approved. Claude Code's own permission prompts are the check before any write.