Verification checklist
- ✓On the usage page, group Responses or Chat Completions by service tier. A model total hides the split.
- ✓Reconcile
service_tieron the response, not the request.defaultmeans standard rates even if the request saidfast. - ✓Check Project Service Tier. Omitting
service_tieruses that project default, not Standard. - ✓A Flex
429 Resource Unavailableshould not appear on the bill. If the retry dropsservice_tier, the next call may have left Flex pricing. - ✓On older models, Fast usage can show up as
priority. Filtering the bill for the stringfastdrops half of that tier.
1. The price gap is a latency tier on the same call
This is easy to mix up with the Batch API. Batch is a different path: upload JSONL, wait up to 24 hours, discount attached to that endpoint. Flex and Fast are not that path. They are still synchronous Responses or Chat Completions calls, plus a service_tier.
OpenAI documents three layers. Omit the parameter, or send auto, and the call uses the project setting. An untouched project uses default: standard price, standard latency. flex prices tokens at Batch API rates, and prompt-cache discounts can stack, in exchange for slower responses and occasional lack of capacity. fast (formerly priority; both values are the same lane) buys lower, steadier latency at a per-token premium.
The model name can be identical. A finance rollup of “what gpt cost this month” blends half-price offline work with double-price chat, then someone swaps the model. The lever is the tier.
2. Flex is cheap because a miss can be free
The Flex guide is explicit: evaluations, enrichment, async jobs, not a user waiting on a chat turn. Price matches Batch, but the call still waits inside this HTTP request. Official SDKs default to a 10-minute timeout; Flex samples raise it to 15. On a 408, the SDK retries twice before throwing.
When capacity is short, the API returns 429 Resource Unavailable and does not charge. That is the line between Flex and “a slow standard call.” A slow standard call still bills standard. A Flex call that never got capacity should not be on the invoice.
The retry policy is where the discount disappears. The docs offer two paths: exponential backoff, stay on Flex, wait for capacity, keep the price; or set service_tier to auto, or delete the parameter, and take the project default. The second path is not “try again at half price.” If someone set the project to Fast, the “cheap retry” leaves on the premium tier. A log that only says “Flex failed, retry succeeded” will not match the bill.
3. A successful Fast call is not proof you paid Fast
The Fast mode guide says Priority was renamed Fast on 30 July 2026. priority and fast select the same lane. The service_tier on the response is the tier that actually served the call. For GPT-5.6 and earlier, the response returns priority whether you sent fast or priority. Group usage by service tier and those rows say priority too.
The premium is not a universal “2× coupon.” The documented example is GPT-5.6 Sol, where Fast costs twice the matching Standard rate: $8 / $40 per million input / output tokens on short context, $16 / $60 on long context, with promotional pricing at least through 21 November 2026. Other models use the pricing page. Cached-input discounts still apply. Fine-tuned models and embeddings are outside Fast.
The quiet part is the downgrade. If traffic ramps too fast, some Fast requests are served at standard speed and billed at standard rates, and the response says service_tier: "default". The published rule of thumb: once you are at 1 million input tokens per minute, don’t grow more than 50% every 15 minutes. The exact point moves with model and load. The ramp limit is organization-wide, not per project. Switching models or snapshots, or dumping an overnight ETL into Fast, is how teams hit it.
Downgraded calls succeed. An alert that only watches 4xx and 5xx stays quiet. Users say Fast felt ordinary today. Finance sees two unit prices on one model. Neither starts by reading the response field.
4. Project default, rate limits, and Scale are three ledgers
Project settings can set Project Service Tier to Fast. Requests that omit service_tier then move to Fast gradually, not in one cutover. There is no code diff, and the bill gets more expensive over hours. The other direction is a hard reject. Error codes say a tier the project disallows — including a tier that auto or an omitted parameter resolves to — returns 400 with param service_tier. In that policy, fast is evaluated as priority. Scale Tier sits outside it.
Fast and Standard share the model’s rate limit. There is no separate pool. That is the opposite of Batch’s separate rate-limit pool. Fast does not buy extra TPM. Scale Tier is a third ledger: Fast billing is separate and does not draw purchased Scale TPM, and Scale spillover does not automatically become Fast.
Fast mode for GPT-6 Astra has no latency SLA. On GPT-5.6 and earlier, Fast and Scale share the same SLA treatment, and an enterprise agreement may include service credits when a target is missed. Paying the premium without a latency promise is documented for the newest model, not a support improvisation.
5. Split the tier. Keep seat renewals off the API balance
Keep three columns: the request parameter, service_tier on the response, and usage grouped by service tier. When they disagree, trust the response and the usage page. fast in, default out, is a downgrade, not a lost request. priority on the dashboard and fast in code is an alias, not a second product.
API prepaid credit and a ChatGPT web seat are still two bills. The tier discount lives on service_tier, not on the card BIN. Isolate web-seat renewals on a capped virtual card such as RDVCC so those auto-renewals don’t compete with API top-ups. A virtual card does not stop a Flex 429, and it does not stop Fast from being billed as standard.