Million-token API pricing: verified rates from OpenAI, Google and xAI

Share





Among the vendor pricing pages we could verify, OpenAI’s gpt-5.6-luna is the cheapest standard short-context text model in this comparison at $0.20 input, $0.02 cached input and $1.20 output per 1M tokens, while Google’s gemini-3.8-flash lists $0.75 input, $0.075 context caching and $3.75 output through 31 December 2026, and xAI’s grok-build-0.1 lists $1.00 input, $0.20 cached input and $2.00 output below its 200k-token long-context threshold.

That is the verified answer to the cheap-million-token question from the source the public record. It is not the whole bill. The same pages split pricing by input, cached input, output, context threshold, service tier, regional uplift and tool calls.

Verified headline rates from vendor pages

OpenAI’s pricing page lists gpt-5.6-luna on Standard short context at $0.20 input, $0.02 cached input, $0.25 cache writes and $1.20 output per 1M tokens, and on Standard long context at $0.40 input, $0.04 cached input, $0.50 cache writes and $1.80 output per 1M tokens. The captured page shows short-context and long-context tables, but it does not state the threshold that moves a request from one to the other.

Google’s pricing page lists gemini-3.8-flash on the Paid Standard tier at $0.75 input, $3.75 output and $0.075 context caching per 1M tokens through 31 December 2026, then $1.50 input, $7.50 output and $0.15 context caching from 1 January 2027. The same page lists context-cache storage at $0.50 per 1,000,000 tokens per hour through 31 December 2026 and $1.00 from 1 January 2027.

xAI’s pricing page lists grok-build-0.1 at $1.00 input, $0.20 cached input and $2.00 output per 1M tokens below 200k prompt tokens, then $2.00, $0.40 and $4.00 once the prompt reaches 200k tokens. The same page lists grok-4.3 at $1.25, $0.20 and $2.50 below 200k prompt tokens, then $2.50, $0.40 and $5.00 above that threshold.

Why “1M tokens” is at least three prices?

OpenAI, Google and xAI all separate input from output, and all three publish some form of discounted reuse pricing. On gpt-5.6-luna, output costs six times the Standard short-context input rate, according to OpenAI’s pricing page. On gemini-3.8-flash, output costs five times the Paid Standard input rate through 31 December 2026, according to Google’s pricing page.

The reuse line also differs by vendor. OpenAI publishes cached-input and cache-write rates. Google publishes context-caching token rates plus a separate hourly storage charge. xAI publishes cached-input prices and bills long-context rates for all tokens in a request once its prompt reaches the model’s long-context threshold.

How tier rules change the effective rate?

On OpenAI’s pricing page, gpt-5.6-luna drops from $0.20 input and $1.20 output on Standard short context to $0.10 input and $0.60 output on Batch and on Flex, while Fast mode rises to $0.40 input and $2.40 output. The same page says regional processing endpoints add a 10% uplift for eligible models released on or after 5 March 2026.

On Google’s pricing page, gemini-3.8-flash Batch and Flex both cut the Paid Standard rate in half to $0.375 input, $1.875 output and $0.0375 context caching through 31 December 2026, while Priority raises those figures to $1.35 input, $6.75 output and $0.135 context caching through 31 December 2026.

On xAI’s pricing page, Batch discounts vary by model: grok-4.3 and the grok-4.20-0309 variants get 20% off, while models not listed get no Batch discount. The same page says Priority Processing carries a 2x premium over standard rates, and the US regional endpoint bills token usage at 1.1x global rates.

xAI also publishes a separate fast path: Grok 4.7 Fast costs $4.00 input, $1.00 cached input and $12.00 output below 200k prompt tokens, and $6.00, $1.50 and $18.00 above 200k. The page also says Grok 4.7 Fast is only available through Cursor and Grok Build, not on the public xAI API.

Why tool fees make token-only comparisons incomplete?

OpenAI’s pricing page adds tool charges on top of token billing, including web search at $10.00 per 1,000 calls, file-search tool calls at $2.50 per 1,000 calls and container sessions from $0.03 for 1 GB to $1.92 for 64 GB per 20-minute session. The same page says search content tokens are billed at model rates.

xAI’s pricing page lists web_search at $5 per 1,000 calls, code_execution at $5 per 1,000 calls, attachment_search at $10 per 1,000 calls and collections_search or file_search at $2.50 per 1,000 calls. The same page says X Search changes on 21 September 2026 at 12:00 PM PT from $5 per 1,000 calls to $5 per 1,000 posts fetched and $10 per 1,000 user profiles fetched.

Google’s pricing page puts grounding with Google Search and Google Maps on separate meters: each gets 5,000 free requests or prompts per month on the Paid tier for Gemini 3.x, then costs $14 per 1,000 requests or queries.

Buyer’s decision table: verified per-1M text rates

Figures below are in USD per 1M tokens unless noted. Where a vendor publishes a date change or threshold, it is shown in the same row.

OptionVendorInputCached inputOutputContextBest fitEvidence tier
gpt-5.6-luna (Standard)OpenAI$0.20$0.02$1.20Short context on the page; long-context Standard is $0.40 / $0.04 / $1.80. OpenAI also lists cache writes at $0.25 short and $0.50 long.Lowest verified standard rate in this comparisonVendor page
gpt-5.6-luna (Batch or Flex)OpenAI$0.10$0.01$0.60Short context; long-context Batch and Flex are $0.20 / $0.02 / $0.90. Cache writes are $0.125 short and $0.25 long.Backfills and lower-priority workloadsVendor page
gemini-3.8-flash (Paid Standard)Google$0.75 through 31 December 2026; $1.50 from 1 January 2027$0.075 through 31 December 2026; $0.15 from 1 January 2027$3.75 through 31 December 2026; $7.50 from 1 January 2027Paid tier; context-cache storage is $0.50 per 1,000,000 tokens per hour through 31 December 2026, then $1.00 from 1 January 2027Production traffic if you can live with the published 2027 step-upVendor page
gemini-3.8-flash (Paid Batch or Flex)Google$0.375 through 31 December 2026; $0.75 from 1 January 2027$0.0375 through 31 December 2026; $0.075 from 1 January 2027$1.875 through 31 December 2026; $3.75 from 1 January 2027Batch and Flex both halve the Paid Standard rate on the pageDeferred work and price-sensitive bulk trafficVendor page
grok-build-0.1xAI$1.00 below 200k; $2.00 at or above 200k$0.20 below 200k; $0.40 at or above 200k$2.00 below 200k; $4.00 at or above 200kText API; long-context rates apply to all tokens once the prompt reaches 200k tokensCheapest verified xAI text rate in the source the public recordVendor page
grok-4.3xAI$1.25 below 200k; $2.50 at or above 200k$0.20 below 200k; $0.40 at or above 200k$2.50 below 200k; $5.00 at or above 200k1M context; the page also says grok-4.3 gets a 20% Batch discountLong-context xAI workloads with a published Batch discountVendor page
grok-4.7 FastxAI$4.00 below 200k; $6.00 above 200k$1.00 below 200k; $1.50 above 200k$12.00 below 200k; $18.00 above 200kAvailable only through Cursor and Grok Build; not on the public xAI APITeams already buying the faster managed pathVendor page

Source: vendor pricing pages from OpenAI, Google and xAI.

The cheapest verified standard row in this sample is still gpt-5.6-luna. The cheapest verified xAI text row in the source the public record is grok-build-0.1, not grok-4.7. Google’s gemini-3.8-flash sits between them on headline token rates, but its published date change on 1 January 2027 matters if you are budgeting beyond this year.

What we could not verify?

We could not verify public vendor-page pricing for Xiaomi or DeepSeek from the source the public record, so we removed those rows instead of repeating tracker figures. The third-party sources below, including Price Per Token and llm-prices, are useful for market scanning, but the captured the public record text does not include current Xiaomi or DeepSeek model-level prices that we could check line by line against a vendor page.

We also could not verify the exact token threshold that separates short-context and long-context billing on OpenAI’s pricing page. The page publishes both schedules, but the captured text does not state where the switch happens.

The next date that changes the math is 1 January 2027

Google’s pricing page says gemini-3.8-flash Paid Standard moves from $0.75 input, $0.075 context caching and $3.75 output to $1.50, $0.15 and $7.50 on 1 January 2027. If your budget spans 2026 and 2027, that is the date to model now.

Sources



Maya Chen
Maya Chen
Maya Chen covers AI agents, orchestration frameworks, tool-use, and evaluation. She focuses on what actually works in production—failure modes, safety boundaries, and measurable performance—without the hype.

Read more

Local News