$GRAZE CA: only trust the address shown on usegraze.world and posted by @usegraze. Claude, OpenAI and Kimi credits · 28 Sep
Graze docs

How routing works

graze/auto tries llama-4-maverick first and moves to gpt-6-luna only when the first answer is unusable. You can also name any model directly.

flowchart LR
  A[Your agent] --> G[Graze]
  G -->|first| L[llama-4-maverick]
  G -.->|if needed| P[gpt-6-luna]
  L --> R[Answer + receipt]
  P --> R

When Graze moves to the next model

  • The model returns an error, for example a provider outage.
  • The answer is empty.
  • You asked for JSON (response_format) and the answer is not valid JSON.
  • You required a tool call (tool_choice: "required") and none came back.

Graze does not judge whether an answer is good, only whether it is usable.

Billing across attempts

A failed attempt that returned an error costs nothing. An attempt that produced tokens is billed, even if Graze then moved on, so the receipt always shows the true cost.

Streaming

Streamed calls can switch models only before the first token arrives. After that the stream is not re-checked.

Reasoning models

Some models think before answering, and the thinking counts toward max_tokens. Graze gives them room: gpt-6-luna gets at least 400 output tokens, deepseek-v4-pro 600 and kimi-k3 800. You pay only for tokens actually used.

Available models

Model Notes
graze/auto Recommended. llama-4-maverick, then gpt-6-luna
llama-4-maverick Fast, no reasoning step
gpt-6-luna Reasons briefly, low cost
deepseek-v4-pro Reasoning model
kimi-k3 Reasoning model, close to Claude Sonnet in price