Waymaker
WaymakerDocs
Developer Documentation
Documentation

waymakerone local

Run Claude Code against a model on your own Mac, through a small gateway that the CLI starts and stops for you. waymakerone local and waymaker local are the same command (CLI 2.17.0 and later).

There are two ways to use it:

CommandWhat happensUse it when
waymakerone local --offlineEvery model request goes to the local model. No model requests leave your Mac; tools you approve (WebFetch, Bash, MCP servers you add) can still use the network.You have no internet or no credits, and need to finish a task or do basic work.
waymakerone localHybrid (the default). Your frontier model keeps the session, under your normal Claude sign-in. Claude Code's small/fast model slot runs locally.You want to offset some frontier usage without giving up frontier quality on the main work.

Local answers are lower quality than the cloud. The command says so before it starts, and --offline says so explicitly.

Before you start

  • Claude Code on your PATH (claude).
  • Ollama installed (brew install ollama). If Ollama is already running, the command uses it. If nothing is answering, it starts ollama serve and stops it again when you exit. It never stops an Ollama it did not start.
  • At least 24 GB of memory. Below that the command refuses to run. --coding needs 48 GB.
  • The model. The default is gpt-oss:20b (about 13 GB). If it is not on your Mac, the command prints its size and the exact ollama pull command, checks you have about 1.5× that in free disk, and asks first. It never downloads silently. Pass --yes to agree in advance.

Unlike the rest of the CLI, local needs no WAYMAKER_API_KEY.

Usage

waymakerone local [--offline] [--coding] [--yes] [--force] [--dry-run] [claude args...]
waymakerone local stats [--days N]
FlagWhat it does
--offlineEverything runs locally. Every model slot points at the local model, the context is set to 64k, and your MCP servers are not loaded unless you pass your own --mcp-config.
--codingUse qwen3-coder:30b-a3b instead of gpt-oss:20b (48 GB or more). It is for Claude Code only: in our test it did a multi-file rename in Claude Code but failed the same rename in another harness, and its MCP tool calls were unreliable.
--yesAllow downloading a missing model.
--forceRun even while another local-model job holds the machine-wide lock.
--dry-runPrint the mode, model, endpoints and environment variable names, and start nothing.

Anything else is passed to claude. If one of claude's own flags clashes with these, put it after --:

waymakerone local --offline -p "summarise README.md"
waymakerone local -- --help

What it measures, and what it does not

waymakerone local stats reads ~/.waymaker/local-usage.jsonl: one line per request with the time, local or upstream, the model name, token counts, duration and status. It never holds your prompts, the replies or any credentials. The file is readable by you alone (mode 0600), and moves to local-usage.jsonl.1 once it reaches 5 MB. Nothing in it is sent anywhere.

waymaker local stats --days 7

How hybrid works

Claude Code chooses a model per request by name, but has one base URL. The gateway listens on 127.0.0.1. A request for a local- model goes to Ollama. Everything else goes to api.anthropic.com unchanged, with your sign-in headers forwarded exactly as Claude Code sent them. Your credentials are never sent to Ollama.

Hybrid needs your normal Claude sign-in to work through a custom base URL. We tested this on 2026-10-11 with a subscription (OAuth) sign-in and no API key set, and it worked. Claude Code decides which requests use the small/fast slot. A trivial -p prompt used none, so how much a real session offloads has not been measured yet.

What has been tested

  • Measured on one Mac only: Apple M5 Pro with 48 GB, Ollama 0.40.1, Claude Code 2.1.296, one run each (n=1).
    • --offline -p "reply with the single word ready" replied ready with the model warm. The prompt was 17,130 tokens.
    • In hybrid, a frontier request went upstream and returned 200 with the subscription sign-in. A small/fast-slot request ran locally: 21,430 prompt tokens in 21.6 s.
  • Not tested: 24 GB Macs. The 24 GB tier is inferred, not measured. Intel Macs are not supported.

Limits

  • 64k context. The gateway refuses a local request it estimates is larger than the model's context, with the error Claude Code uses to compact the conversation. Without that refusal, Ollama would silently drop the start of the conversation. The estimate errs high, so compaction can start a little early.
  • No tool profile yet. --offline loads no MCP servers by default, because a full catalogue does not fit in 64k. A curated local tool set comes in a later release.
  • One local model at a time. The command takes the same machine-wide lock as other local-model jobs (~/.fleet/locks/ollama). In --offline it refuses while another job holds the lock. In hybrid it still starts, sends everything to your frontier model, and tells you so.
  • Writes still ask. Nothing is auto-approved. Claude Code's own permission prompts are the check before any write.