Waymaker
WaymakerDocs
Developer Documentation
Documentation

Calling AI from a Host app

Preview — HQ only. The AI gateway is running for Waymaker's own Solutions while billing is built. It isn't available to other organisations yet, and the details on this page can change before it is.

Your app on Host can call Claude, GPT and Gemini through one endpoint, with one key per Solution. You don't open a provider account and you don't put a provider key in your app. What the app spends is recorded against its Solution.

This is raw model access for software you build. It is separate from One, the AI inside Waymaker.

The endpoint

https://ai.waymakerapp.com/v1
RouteCompatible with
POST /v1/chat/completionsOpenAI Chat Completions
POST /v1/responsesOpenAI Responses
POST /v1/messagesAnthropic Messages
GET /v1/modelsthe models your key may call

Streaming works on all three POST routes.

The key

A key starts with wmai_ and belongs to exactly one Solution. It is shown once, when it's created. We keep only a fingerprint of it, so a lost key can't be shown again: revoke it and create another. During the preview, keys are issued by Waymaker on request.

Send it as Authorization: Bearer wmai_… (OpenAI SDKs) or x-api-key: wmai_… (Anthropic SDKs).

Call from a server, never from a browser. Use your app's backend or an Ambassador, and keep the key in the environment, not in client code. The endpoint doesn't answer browser cross-origin requests.

A revoked key stops working within about a minute, everywhere.

OpenAI SDK

import OpenAI from 'openai'

const openai = new OpenAI({
  apiKey: process.env.WAYMAKER_AI_KEY,
  baseURL: 'https://ai.waymakerapp.com/v1',
})

const stream = await openai.chat.completions.create({
  model: 'anthropic/claude-sonnet-4-5',
  messages: [{ role: 'user', content: 'Summarise this ticket in one line.' }],
  stream: true,
})
for await (const chunk of stream) process.stdout.write(chunk.choices[0]?.delta?.content ?? '')

Anthropic SDK

import Anthropic from '@anthropic-ai/sdk'

const anthropic = new Anthropic({
  apiKey: process.env.WAYMAKER_AI_KEY,
  baseURL: 'https://ai.waymakerapp.com/v1',
})

const message = await anthropic.messages.create({
  model: 'anthropic/claude-sonnet-4-5',
  max_tokens: 512,
  messages: [{ role: 'user', content: 'Summarise this ticket in one line.' }],
})

The Anthropic SDK adds /v1/messages to the base URL itself, so the base URL above is the same for both SDKs.

Models

Model names are provider/model. A key may call only the models on its allow-list; the default is:

  • anthropic/claude-sonnet-4-5
  • anthropic/claude-haiku-4-5
  • openai/gpt-4.1
  • openai/gpt-4.1-mini
  • google/gemini-2.5-flash

GET /v1/models returns your key's list. Asking for any other model returns 400 with the allowed models named in the message.

Limits

Each key has two per-minute limits:

LimitDefault
Requests per minute60
Tokens per minute (input + output)100,000

Over either limit you get 429 with a Retry-After header. Tokens are counted when a response finishes, so one large request can take a key over its token limit, and the next request that minute is refused.

Errors

Errors come back in the shape your SDK expects:

StatusMeaning
401Missing, malformed or revoked key
400Model not on the key's allow-list, or the body isn't JSON
413Request body over 10 MB
429Over a per-key limit; see Retry-After
502 / 503The model provider or the gateway couldn't be reached. Retry with backoff

Every response carries an x-request-id. Quote it when you ask for help.

We don't retry a failed model call for you. A retry is a second call, and it would be billed twice.

Privacy

Prompts and completions aren't stored by the gateway. We record usage metadata only: the model, token counts, status and latency. Responses are never cached or shared between Solutions.