Calling AI from a Host app
Preview — HQ only. The AI gateway is running for Waymaker's own Solutions while billing is built. It isn't available to other organisations yet, and the details on this page can change before it is.
Your app on Host can call Claude, GPT and Gemini through one endpoint, with one key per Solution. You don't open a provider account and you don't put a provider key in your app. What the app spends is recorded against its Solution.
This is raw model access for software you build. It is separate from One, the AI inside Waymaker.
The endpoint
https://ai.waymakerapp.com/v1
| Route | Compatible with |
|---|---|
POST /v1/chat/completions | OpenAI Chat Completions |
POST /v1/responses | OpenAI Responses |
POST /v1/messages | Anthropic Messages |
GET /v1/models | the models your key may call |
Streaming works on all three POST routes.
The key
A key starts with wmai_ and belongs to exactly one Solution. It is shown once, when it's
created. We keep only a fingerprint of it, so a lost key can't be shown again: revoke it and create
another. During the preview, keys are issued by Waymaker on request.
Send it as Authorization: Bearer wmai_… (OpenAI SDKs) or x-api-key: wmai_… (Anthropic SDKs).
Call from a server, never from a browser. Use your app's backend or an Ambassador, and keep the key in the environment, not in client code. The endpoint doesn't answer browser cross-origin requests.
A revoked key stops working within about a minute, everywhere.
OpenAI SDK
import OpenAI from 'openai'
const openai = new OpenAI({
apiKey: process.env.WAYMAKER_AI_KEY,
baseURL: 'https://ai.waymakerapp.com/v1',
})
const stream = await openai.chat.completions.create({
model: 'anthropic/claude-sonnet-4-5',
messages: [{ role: 'user', content: 'Summarise this ticket in one line.' }],
stream: true,
})
for await (const chunk of stream) process.stdout.write(chunk.choices[0]?.delta?.content ?? '')
Anthropic SDK
import Anthropic from '@anthropic-ai/sdk'
const anthropic = new Anthropic({
apiKey: process.env.WAYMAKER_AI_KEY,
baseURL: 'https://ai.waymakerapp.com/v1',
})
const message = await anthropic.messages.create({
model: 'anthropic/claude-sonnet-4-5',
max_tokens: 512,
messages: [{ role: 'user', content: 'Summarise this ticket in one line.' }],
})
The Anthropic SDK adds /v1/messages to the base URL itself, so the base URL above is the same for
both SDKs.
Models
Model names are provider/model. A key may call only the models on its allow-list; the default is:
anthropic/claude-sonnet-4-5anthropic/claude-haiku-4-5openai/gpt-4.1openai/gpt-4.1-minigoogle/gemini-2.5-flash
GET /v1/models returns your key's list. Asking for any other model returns 400 with the allowed
models named in the message.
Limits
Each key has two per-minute limits:
| Limit | Default |
|---|---|
| Requests per minute | 60 |
| Tokens per minute (input + output) | 100,000 |
Over either limit you get 429 with a Retry-After header. Tokens are counted when a response
finishes, so one large request can take a key over its token limit, and the next request that
minute is refused.
Errors
Errors come back in the shape your SDK expects:
| Status | Meaning |
|---|---|
401 | Missing, malformed or revoked key |
400 | Model not on the key's allow-list, or the body isn't JSON |
413 | Request body over 10 MB |
429 | Over a per-key limit; see Retry-After |
502 / 503 | The model provider or the gateway couldn't be reached. Retry with backoff |
Every response carries an x-request-id. Quote it when you ask for help.
We don't retry a failed model call for you. A retry is a second call, and it would be billed twice.
Privacy
Prompts and completions aren't stored by the gateway. We record usage metadata only: the model, token counts, status and latency. Responses are never cached or shared between Solutions.