Routing
Decide which offer handles a request, then understand fallback.
Open Routing, choose a strategy and restrictions, and select Save routing. Routing first filters offers for the exact requested model, then orders the eligible offers. It does not substitute a different model.
Choose a strategy
| Strategy | First choice |
|---|---|
| Cheapest only | Lowest estimated input cost plus the requested output budget. |
| Fastest | Lowest recent time to first token (TTFT), with price breaking speed ties. |
| Balanced | Cheapest offer within 1.5× the fastest estimated TTFT. |
Prices compare the estimated uncached input and requested output counts. A high output budget can change the winning offer even if you expect a short answer. Final billing uses actual usage.
Worked example
Suppose your request estimates 1,000 input tokens and allows 500 output tokens. These illustrative offers have equal input and output rates:
| Offer | Input / 1M | Output / 1M | Estimated request cost | TTFT |
|---|---|---|---|---|
| A | $1 | $1 | $0.0015 | 1,000 ms |
| B | $2 | $2 | $0.0030 | 600 ms |
| C | $3 | $3 | $0.0045 | 500 ms |
Cheapest only picks A. Fastest picks C. Balanced sets a 750 ms threshold (500 × 1.5) and picks B, the cheaper of B and C. All figures are hypothetical, not current Marketplace prices.
Understand TTFT
TTFT measures the time until a generated text token, reasoning token, or tool start. Headers, keep-alives, and empty blocks do not count. Only streaming requests provide routing TTFT samples; total response latency is a different measurement.
Each offer's estimate combines 80% of its previous value with 20% of the new sample. It expires after 24 hours without a sample. An unmeasured offer uses the median measured TTFT; when none are measured, offers have equal assumed speed and price breaks ties.
Apply a price cap and model lists
Price cap applies separately to both input and output USD rates per million tokens. For example, a $2 cap rejects an offer with $1 input and $3 output. Leave the field empty for no cap. An offer exactly at the cap is eligible.
Blacklist excludes checked models. Whitelist permits only checked models. The key's Selected models restriction also applies: the model must pass both settings. A model restriction returns 403; offers excluded only by price can leave you with 503.
Context/output limits, available inventory, active status, and offer health also affect eligibility. A seller's own offers cannot serve their buyer requests.
Understand fallback and failures
Fallback means trying the next eligible offer when an attempt fails before response bytes reach you. The gateway selects at most three offers in strategy order. Retryable upstream failures and inventory contention can move the request to the next offer.
Two counted upstream failures within one minute temporarily remove an offer for 60 seconds. Authentication failures can pause the affected offer. These rules do not guarantee that another eligible offer exists.
Once response bytes have been sent, the gateway does not retry the request on another offer. A later failure is sent as a stream error when the connection permits it; a disconnected client cannot receive that event. Inspect stream events as well as the initial HTTP status.
Explain a route you observed
- Save
x-request-idandx-unused-listingfrom the response. - Check the model, status, cost, and total latency in Analytics → Requests.
- Compare the model's offers and TTFT in Marketplace.
- Check Routing and the key's settings for restrictions.
- Use Analytics → Audit Log → Routing to identify settings changes.
The Requests page does not expose per-attempt fallback history or offer TTFT. Current Marketplace values can differ from the values at request time. Quote both response IDs when the visible information cannot explain a route.