Pricing, credits and cache
Calculate a charge and understand temporary credit holds.
Add and inspect credits
Open Credits → Add credits and choose an available payment method. Card credits arrive after payment confirmation. USDC on Base uses the deposit information shown in the panel; follow its network and token instructions and wait for confirmation.
Your credits pay for buyer requests. Seller earnings and Payouts are separate balances.
Calculate the request cost
An offer prices input and output separately in USD per million tokens. You pay for the tokens used, not in million-token blocks. For normal cache billing:
charge_usd = (
(input_tokens - cached_input_tokens) * input_rate
+ cached_input_tokens * cache_read_rate
+ output_tokens * output_rate
) / 1_000_000Here, input_tokens includes cache reads and writes. Cache writes use the regular input rate unless the offer enables free cache. Rates are USD per million tokens. The total is rounded upward once, to the next $0.000001, after combining the categories.
Worked example with cache reads
Suppose an offer charges $2 input, $0.20 cache reads, and $8 output per million tokens. Your completed request reports 10,000 input tokens, including 6,000 cache reads, and 500 output tokens.
| Category | Tokens | USD / 1M | Cost |
|---|---|---|---|
| Input without cache reads | 4,000 | $2 | $0.0080 |
| Cache reads | 6,000 | $0.20 | $0.0012 |
| Output | 500 | $8 | $0.0040 |
| Total | $0.0132 |
At the regular input rate, the same usage would cost $0.024. The cache saving is $0.0108. These are illustrative rates, not current offer prices.
Understand the hold and release
A hold is a temporary credit reservation. Before sending the provider request, the gateway reserves estimated input at the uncached rate plus your requested output budget. It cannot count on a cache hit before the provider reports one.
In the example above, if estimated input is 10,000 and max_tokens is 2,000, the hold is (10,000 × $2 + 2,000 × $8) / 1,000,000 = $0.036. After charging $0.0132, the unused $0.0228 is released. Concurrent requests can therefore reduce your available spending capacity before their final charges are known.
The hold must fit both your credits and the key's remaining monthly budget. A large output budget can cause 402 even when you expect a much smaller final bill. Use a realistic output limit. Offer rates and reference prices are captured when the request is admitted; later price edits do not change that request's agreed rates.
Understand cache prices
| Setting | Billing behavior |
|---|---|
| Seller's explicit cache-read rate | Uses that rate, capped at the regular input rate. |
| Automatic cache-read price | Uses the seller input rate multiplied by the reference catalog's cache/input ratio, rounded upward at the rate's precision. |
| Missing or invalid reference cache ratio | Falls back to the regular input rate. |
| Cache write | Uses the regular input rate under normal billing. |
| Free cache | Reported cache reads and writes cost zero; other input and output remain billable. |
Free cache subtracts both reported cache categories from regular input. With cache_write_tokens available, its formula is:
charge_usd = (
(input_tokens - cached_input_tokens - cache_write_tokens) * input_rate
+ output_tokens * output_rate
) / 1_000_000The same final rounding rule applies. The offer's free-cache policy is captured for the request with its rates.
What “reported cache” means
The provider must return cache usage counters. The gateway does not infer a hit because a prompt resembles an earlier prompt. If cache usage is missing, it does not invent a discount. OpenAI-shaped usage exposes reported reads under prompt_tokens_details.cached_tokens; Anthropic-shaped usage distinguishes cache reads and creation.
In Analytics → Requests, cached input appears under Tokens in and the cost can show cache savings. Cached input is already part of total input; do not add it again when calculating total tokens.
Preserve cache-compatible prompts
Keep reusable instructions and tool definitions stable when testing cache behavior, and compare the provider-reported counters. Anthropic cache_control is accepted for supported text blocks, tools, and requests, but upstream support determines its effect. Streaming supplies TTFT measurements; enabling it is not evidence of a cache hit.
Claude Code can place plain role: "system" messages inside messages. The gateway folds these instructions into the system prompt; system messages carrying per-message output_config.effort use a separate preserved path. Folding does not preserve those inline prompt-cache positions. Changing client format or reaching a different offer can also change the request seen by the provider; do not assume a previous hit carries over.
TODO(verify): Provider-specific cache eligibility, minimum prompt size, retention, and cache affinity across different seller credentials are not established by the gateway code. Confirm them for the selected provider before relying on repeated hits.
Read reference-price savings
A reference price is the catalog comparison rate captured for the request, not a promise about a provider invoice. Analytics → Savings compares your paid cost with the same recorded token mix at reference rates. If required rates are missing, that usage is shown separately as unpriced. Negative savings mean you paid above the comparison cost.
A failed stream can still have a charge for delivered usage. See Troubleshooting before assuming every failure is free.