Million-Token Models Under One Dollar: What the Published Price Cards Actually Prove

Share




Z.ai lists GLM-5.3-Flash at $0.15 per 1M input tokens and $0.50 per 1M output tokens, putting both token columns below $1. But its published pricing page does not state that model’s context window, so it cannot by itself prove a million-token offering. The clearest rule is to treat “under one dollar” as applying separately to input and output, then require a published context limit alongside the rate card.

Z.ai publishes sub-$1 rates for several models, including GLM-4.7-FlashX and GLM-5.3-Flash. Google publishes an input rate below $1 for Gemini 3.8 Flash through 31 December 2026, but its output rate exceeds the threshold. Neither of the cited first-party pricing pages pairs those prices with a stated million-token context window.

At a glance

What does “under one dollar” mean for an API model?

For API buyers, “under one dollar per million tokens” is incomplete unless it says which tokens. Providers commonly publish separate rates for input tokens and output tokens. Input is the material sent to the model, while output is the text, code, tool trace or other response returned by it.

A model qualifies under the strict reading used here only if its published input price is below $1 per 1M tokens and its published output price is also below $1 per 1M tokens. A model with cheap input but expensive output is not an under-$1 model on this definition. This matters because many useful workloads generate far more output than a short chat reply.

A second test is separate. A “million-token model” should have a published context window of at least 1M tokens. The context window is the maximum material a model can consider in one request, not the amount of information a search tool might retrieve.

That distinction prevents a misleading shortlist. A vendor can offer low token rates without a million-token window. A vendor can also offer a very large window while charging more than $1 for one or both token columns.

The strict test therefore has three parts:

  • A stated context window of at least 1M tokens.
  • A published input rate below $1 per 1M tokens.
  • A published output rate below $1 per 1M tokens.

Where a price page proves only the rate, this article says so. Where a listing claims both pricing and context, the evidence is identified as a marketplace listing rather than a vendor document. That may still be useful for exploration, but it is weaker evidence for a purchase decision.

Which published price cards meet the under-$1 rate test?

Z.ai’s published price table shows GLM-5.3-Flash at $0.15 input and $0.50 output per 1M tokens. On price alone, it passes the strict both-columns test. The same table lists cached input at $0.03 per 1M tokens and describes cached-input storage as limited-time free.

Z.ai also lists GLM-4.7-FlashX at $0.07 input and $0.40 output per 1M tokens. That also passes the rate test. Its listed cached-input price is $0.01 per 1M tokens.

GLM-4-32B-0414-128K is listed at $0.10 for input and $0.10 for output. It is cheap in both columns. Its name, however, identifies it as a 128K model rather than a million-token model.

Z.ai lists GLM-4.6V at $0.30 input and $0.90 output. That puts the vision model below the threshold on both published token columns. Again, the cited price page does not state a context limit.

GLM-4.6V-FlashX is listed at $0.04 input and $0.40 output per 1M tokens. It is another rate-card match. It should not be described as a verified million-token model until a first-party model document supplies that context specification.

Z.ai lists GLM-OCR at $0.03 input and $0.03 output per 1M tokens. That is the lowest stated pair in the cited table. It is an OCR model, so buyers should not assume it serves as a general-purpose replacement for a text reasoning or coding model.

ProviderModelPublished contextInput $/1MOutput $/1MDoes the price pass?What the cited page proves
Z.aiGLM-5.3-FlashNot stated on cited pricing page$0.15$0.50Yes, both columnsFirst-party pricing; context not shown
Z.aiGLM-4.7-FlashXNot stated on cited pricing page$0.07$0.40Yes, both columnsFirst-party pricing; context not shown
Z.aiGLM-4-32B-0414-128K128K in model name$0.10$0.10Yes, both columnsNot a million-token model on the published identifier
Z.aiGLM-4.6VNot stated on cited pricing page$0.30$0.90Yes, both columnsFirst-party vision-model pricing; context not shown
Z.aiGLM-4.6V-FlashXNot stated on cited pricing page$0.04$0.40Yes, both columnsFirst-party vision-model pricing; context not shown
Z.aiGLM-5.3-FlashXNot stated on cited pricing page$0.37$1.25No, output exceeds $1First-party pricing; context not shown
GoogleGemini 3.8 Flash StandardNot stated on cited pricing page$0.75$3.75No, output exceeds $1First-party pricing through 31 December 2026
GoogleGemini 3.8 Flash BatchNot stated on cited pricing page$0.375$1.875No, output exceeds $1First-party batch pricing through 31 December 2026
XiaomiMiMo-V2.6-FlashNot verified from first-party documentation hereNot verified from first-party documentation hereNot verified from first-party documentation hereNot assessedRequires current Xiaomi documentation
XiaomiMiMo-V2.6-ProNot verified from first-party documentation hereNot verified from first-party documentation hereNot verified from first-party documentation hereNot assessedRequires current Xiaomi documentation

Source: Z.ai and Google pricing pages linked in each applicable cell. A price card is not, by itself, proof of a context-window limit.

Why does Z.ai make the shortlist incomplete rather than confirmed?

Z.ai provides a detailed first-party token price table. That table distinguishes input, cached input, cached-input storage and output. It is stronger evidence for pricing than a reseller’s model catalogue because the vendor itself publishes the figures.

The gap is context. The cited page does not list context windows for GLM-5.3-Flash, GLM-4.7-FlashX, GLM-4.6V or GLM-4.6V-FlashX. A price card cannot establish that these are million-token models when it does not state their maximum prompt size.

That does not make the low prices unhelpful. It makes them conditional. A buyer who already knows a model’s context limit from a current contractual schedule or model card can use the listed token rates as a starting point.

For public comparison, though, the safer label is “rate-qualified, context unverified.” This avoids turning a low price into a broader claim that the cited evidence does not support. It also prevents readers from mistaking a model family name for a technical specification.

Z.ai lists GLM-4.7-Flash and GLM-4.5-Flash as free for input, cached input, storage and output. Free access can be useful for evaluation. The price page alone does not explain the operational limits, available capacity, context window or terms that would determine whether free access suits production traffic.

The same caution applies to limited-time offers. Z.ai marks cached-input storage as “Limited-time Free” for several models. That phrase is not a permanent pricing commitment. Teams building a cost forecast should record the date of the rate card and ask what charge applies when the promotional period ends.

Why does Google fail the strict test despite cheap input?

Google lists Gemini 3.8 Flash Standard input at $0.75 per 1M tokens through 31 December 2026. That is below the threshold. Google lists its Standard output price at $3.75 per 1M tokens for the same period.

That output price decides the strict result. Gemini 3.8 Flash is not under $1 per 1M tokens when the phrase applies to both directions. It is an under-$1 input option, not an under-$1 input-and-output option.

Google’s Batch tier lists $0.375 input and $1.875 output per 1M tokens through 31 December 2026. Flex has the same published token prices. Both reduce input costs, but neither moves output below $1.

The distinction matters most in output-heavy work. Code generation, long reports, agent traces and structured extraction can produce substantial output. A purchasing comparison that mentions only Gemini’s discounted input rate leaves out the more expensive half of the bill.

Google says its paid tier offers higher rate limits, context caching and Batch API access. These are access and operational features, not evidence that a model meets an under-$1 output requirement. They may still make a higher-priced SKU suitable for a particular workload.

Google includes thinking tokens in Gemini 3.8 Flash’s output price. That makes the output column particularly important for buyers using tasks that cause the model to spend tokens on internal reasoning. A low input price cannot offset unlimited or unexpectedly high output use by itself.

What changes on 1 January 2027?

Google says Gemini 3.8 Flash Standard input rises to $1.50 per 1M tokens on 1 January 2027. Its Standard output price rises to $7.50 on that date.

Google lists Batch input at $0.75 from 1 January 2027. Batch output becomes $3.75 per 1M tokens from the same date. Flex follows the same posted schedule.

This is a concrete example of why model comparisons need an effective date. A rate that qualifies on input today can stop qualifying when a posted price change takes effect. Long-lived product budgets should retain the date, model identifier, service tier and region or contract terms used in the calculation.

Google’s other Flash entries in the cited page follow the same structure. Gemini 3.7 Flash has the same listed Standard input and output rates through 31 December 2026. Gemini 3.6 Flash also shows $0.75 Standard input and $3.75 Standard output through that date.

How do caching and storage change the real bill?

Context caching stores reusable prompt material so later requests can refer to it without paying the full ordinary input rate again. It is useful when many requests share a large document set, codebase instruction, policy pack or system prompt. It is not a substitute for checking the model’s context-window limit.

Google lists Gemini 3.8 Flash Standard context caching at $0.075 per 1M tokens through 31 December 2026. Google also lists storage at $0.50 per 1M tokens per hour in that period.

Google lists Batch and Flex cache writes at $0.0375 per 1M tokens through 31 December 2026. Their storage charge remains $0.50 per 1M tokens per hour. Lower cache-write pricing can help repeated workloads, but stored context continues to create a time-based cost.

Z.ai lists GLM-5.3-Flash cached input at $0.03 per 1M tokens. Its ordinary input rate for that model is $0.15 per 1M tokens. The difference shows why cache-read and ordinary-input prices should not be blended without saying which traffic is cached.

GLM-4.7-FlashX has a published cached-input rate of $0.01 per 1M tokens. Its ordinary input rate is $0.07 per 1M tokens. A workload with a repeated long prompt may therefore have a different effective input cost from a workload that sends new material every time.

There are limits to this saving. Caching usually helps when requests reuse the same material. It does not reduce generated output, web-search use, image generation or other separately priced features.

Which extra charges can overwhelm a low token rate?

Token prices are only one line in an API budget. Tool calls, search, image generation, video generation and storage may carry their own charges. These services can be valuable, but they should be costed separately from the language model’s input and output tokens.

Z.ai prices built-in Web Search at $0.01 per use. A workflow that searches on every request has an additional fixed cost, regardless of whether its model token rates are below $1 per 1M tokens.

Z.ai lists GLM-Image at $0.015 per image. It lists CogView-4 at $0.01 per image. Image work therefore needs a request-level budget as well as a text-token budget.

Z.ai lists CogVideoX-3 at $0.2 per video. That is a different unit from tokens. Comparing it directly with an input price per 1M tokens would be meaningless.

Google gives paid Gemini users 5,000 free Google Search grounding requests per month across Gemini 3.x models. Google then lists $14 per 1,000 search requests. Search-heavy applications should include that threshold in their forecast.

Google also lists 5,000 free Maps prompts per month for Gemini 3. It then charges $14 per 1,000 search queries. The exact workload matters: a text-only assistant and a grounded location assistant can have very different unit economics.

How should buyers compare costs without being misled?

Start with a workload, not a headline price. Estimate ordinary input, cached input, output and any tool calls for one completed task. Then multiply by expected volume and test a high-usage case, not just an average day.

Keep input and output separate. A retrieval system may send long documents and produce short answers, making input rates important. A coding agent may produce long patches and repeated explanations, making output rates more important.

Separate cache traffic from uncached traffic. If a shared system prompt is reused thousands of times, a cached-input price may dominate its cost. If each request contains a different customer document, ordinary input pricing matters more.

Check what the provider includes in output. Google explicitly says Gemini 3.8 Flash output pricing includes thinking tokens. This makes a direct token-price comparison harder when another vendor describes reasoning consumption differently or does not expose it in the same way.

Do not infer capacity from a free tier. Google describes its free tier as limited access to certain models. The paid tier is the relevant price card for production applications needing higher rate limits, according to Google.

Ask for the exact model identifier. Names such as “Flash,” “Standard,” “Batch” and “Flex” can refer to different service options. A rate attached to one tier should not be assumed to apply to another tier.

Record the price date. Google’s Gemini prices have a stated change date of 1 January 2027. A spreadsheet without an effective date can become wrong even if every formula is correct.

What should a procurement checklist include?

A useful shortlist should be based on documents that answer both the technical and commercial questions. A pricing page answers the commercial part. A model card, API reference or contract schedule should answer the context and access part.

  • Exact model name and API identifier.
  • Maximum context window, stated in tokens.
  • Input, cached-input and output price per 1M tokens.
  • Whether reasoning or thinking tokens count as output.
  • Cache storage price and expiry rules.
  • Batch, Flex, priority or other service-tier restrictions.
  • Tool, grounding, image, audio and video charges.
  • Free-tier limits and paid-tier access conditions.
  • Effective date, currency and applicable contract terms.

This checklist is deliberately dull. It is also what keeps a low headline rate from becoming an unexpected invoice. The right model is not necessarily the one with the smallest input figure; it is the one whose capability, context limit, access conditions and total workload cost match the job.

What can the available documentation not verify?

The cited Z.ai pricing page verifies token prices, cached-input prices and several tool prices. It does not state context windows for GLM-5.3-Flash, GLM-4.7-FlashX, GLM-4.6V or GLM-4.6V-FlashX. Those models cannot be confirmed as million-token options from that page alone.

The cited Google pricing pages verify Gemini 3.8 Flash rates, tier differences, caching charges and the 1 January 2027 price changes. They do not state context windows alongside the relevant pricing entries. They therefore cannot independently verify a million-token context claim for the listed Gemini models.

No Xiaomi first-party model card, API documentation or pricing page appears in the material available here. The Xiaomi MiMo entries are therefore not assessed as verified options in this shortlist. A marketplace listing may be useful for discovery, but a buyer should obtain Xiaomi documentation or a contractual SKU schedule before treating pricing or context claims as procurement evidence.

Likewise, a model name alone is not enough to establish access restrictions. Buyers need to know whether a rate applies to general API access, a limited promotion, a specific region, a reseller, a batch-only service or an enterprise agreement. That information can alter both cost and availability.

What should buyers check next?

Google’s next published Gemini 3.8 Flash price change is 1 January 2027. Before then, the practical next step is to obtain model documentation that pairs each candidate’s context limit with its exact API price and access terms.

Sources


Maya Chen
Maya Chen
Maya Chen covers AI agents, orchestration frameworks, tool-use, and evaluation. She focuses on what actually works in production—failure modes, safety boundaries, and measurable performance—without the hype.

Read more

Local News