Troubleshooting and FAQ
Find the request, interpret its failure, and choose the next action.
Start with the request ID
Save x-request-id, then search for it in Analytics → Requests. Keep x-unused-listing too: on a failed call it can identify the last attempted offer, not a successfully completed response. Requests rejected before authentication may not appear in your account history.
Record the client, model, time, HTTP status, error body, and whether the stream had started. Do not include your API key, provider credentials, or prompt content when they are not needed to reproduce the problem.
Match the symptom to a fix
The reason labels below come from request history. A status stored for an interrupted stream is not a replacement for its already-sent HTTP status.
| Status / reason | Likely cause to check | Action |
|---|---|---|
400 · Invalid request | Unsupported option, malformed message/tool history, or context/output limit. | Read the error's parameter and message. Remove unsupported fields; reduce history or output budget. |
401 · Authentication failed | Missing, invalid, or revoked buyer key; an upstream credential can also fail. | Check the full unused.market key. If it works elsewhere, include the request and listing IDs in a report. |
402 · Insufficient credits | Credit hold does not fit the balance or key's monthly budget. | Add credits, review the key limit, or lower the output budget. The error message distinguishes the monthly limit. |
403 · Access denied | Key or routing model restrictions; possible provider denial. | Check Selected models, Blacklist, and Whitelist. If those permit the model, retain the IDs. |
404 · Model not found | Unknown model ID. | Copy the exact ID from Marketplace or the model list. |
408 · Client stopped reading | Stream consumer did not keep reading response data. | Drain the stream promptly and check client backpressure. |
429 · Rate limit exceeded | Key rate limit or provider throttling. | Reduce concurrency and request rate. Respect Retry-After when present. |
499 · Client disconnected | Client cancelled or closed the connection. | Check your cancellation logic, client timeout, and network. The disconnected client cannot receive a 499 response. |
500 · Server error (500) | Server/provider failure; historical rows do not establish a more precise cause. | Retry an appropriate request later and retain its ID if repeated. |
502 · Invalid or interrupted provider response | Malformed response or stream ended without a valid completion. | Treat partial output as incomplete. Retry only after handling any partial work. |
503 · Temporarily unavailable | No usable offer, price cap, temporary offer recovery, or gateway saturation. | Check Marketplace and Routing. Respect Retry-After when returned. |
504 · Request timed out | Upstream idle or total request deadline elapsed. | Retry later or request a shorter response; inspect whether partial output arrived. |
529 · Provider overloaded | Provider could not accept the request. | Back off and retry later. |
See API reference for the two client error envelopes. For example, 402 uses insufficient_credits in the OpenAI envelope and billing_error in the Anthropic envelope.
Why did this request use a different offer?
The strategy orders currently eligible offers for the same model. Key restrictions, the price cap, context/output limits, inventory, health, and fallback affect the result. Compare Routing with the worked example, and keep the response's listing ID.
Why is my model absent from the model list?
GET /v1/models returns models usable by your key under current routing restrictions. A public Marketplace entry can still be excluded by your allowed models, price cap, temporarily recovering offers, or ownership. It is not a list of every model ever registered.
Why did a failed request cost money?
Once output starts, a later disconnect, timeout, or provider error can still leave billable delivered usage. For an interrupted stream, the gateway settles reported input when available, otherwise estimated input, and estimated delivered output. Missing cache usage is not guessed. Error therefore does not mean “no tokens and no charge.”
Why is available spending lower than expected?
In-progress requests reserve uncached input plus their requested output budgets. The final charge can be much smaller, and unused reservations are released after settlement. Check concurrent work and unusually large output limits before adding another top-up. See credit holds.
Why is there no cache saving?
The provider may not report cache reads, the offer's effective cache-read rate may equal input, or the prompt/provider path may not be cache-compatible. Free cache also depends on reported cache categories. Check recorded cached input instead of assuming that a repeated prompt hit cache.
Why does the stream fail after HTTP 200?
The HTTP response has already started. Read the error event and do not treat a partial stream as a completed answer. The gateway never switches to another offer once response bytes reach the client.
Why does a coding agent ask for an unsupported endpoint?
Only Chat Completions and Messages are public inference APIs. /v1/responses is not implemented as a public route. Use the tool's supported compatibility mode if one exists; check the guide's visible Tested or Untested status before relying on it.