How routing works
graze/auto tries llama-4-maverick first and moves to gpt-6-luna only when the first answer is unusable. You can also name any model directly.
flowchart LR A[Your agent] --> G[Graze] G -->|first| L[llama-4-maverick] G -.->|if needed| P[gpt-6-luna] L --> R[Answer + receipt] P --> R
When Graze moves to the next model
- The model returns an error, for example a provider outage.
- The answer is empty.
- You asked for JSON (
response_format) and the answer is not valid JSON. - You required a tool call (
tool_choice: "required") and none came back.
Graze does not judge whether an answer is good, only whether it is usable.
Billing across attempts
A failed attempt that returned an error costs nothing. An attempt that produced tokens is billed, even if Graze then moved on, so the receipt always shows the true cost.
Streaming
Streamed calls can switch models only before the first token arrives. After that the stream is not re-checked.
Reasoning models
Some models think before answering, and the thinking counts toward max_tokens. Graze gives them room: gpt-6-luna gets at least 400 output tokens, deepseek-v4-pro 600 and kimi-k3 800. You pay only for tokens actually used.
Available models
| Model | Notes |
|---|---|
graze/auto |
Recommended. llama-4-maverick, then gpt-6-luna |
llama-4-maverick |
Fast, no reasoning step |
gpt-6-luna |
Reasons briefly, low cost |
deepseek-v4-pro |
Reasoning model |
kimi-k3 |
Reasoning model, close to Claude Sonnet in price |