Google Gemini’s free tier lists text, image, audio and video inputs, making it the clearest low-cost starting point in public sources. However, no source here provides a first-party, like-for-like price, latency and quality test across the cheap multimodal models. That means there is no defensible single “cheapest” paid multimodal API recommendation.
At a glance
- Google Gemini is listed with a no-card free tier and text, image, audio and video support, but Google’s unpaid-service terms have data-use implications.
- Google says it can use prompts, files and responses submitted through unpaid Gemini services to improve products.
- Mistral provides API access in Free mode without a credit card, although its supplied first-party pages do not establish which free-mode models support every modality.
- PricePerToken lists low token rates for several models labelled vision or multimodal elsewhere, but token prices alone do not establish image, audio or video costs.
- A test published on 20 September 2026 found that image token counts can reverse a text-price ranking.
What is the best cheap multimodal API?
The best-supported answer is Google Gemini for low-risk prototyping, not because it is proven cheapest, but because the cited material identifies a no-card API tier and lists multimodal inputs. The free-tier directory lists Gemini 3.6 Flash with text, image, audio and video support, a 1M-token context window, and 15 requests per minute.
That is a practical starting point for a developer testing image descriptions, video summaries or audio analysis. It is not a recommendation to place confidential production data on the free tier. Google says unpaid Gemini API quota is an unpaid service, with different data terms from paid use.
For paid work, the available evidence does not support naming Ling 3.0 Flash VL, DeepSeek V4.1 Flash or Z.ai GLM 5.3 FlashX as the winner. Public listings show token prices, but they do not establish the complete cost of an image, audio clip or video file. They also do not establish direct vendor availability, supported outputs, regional access, rate limits, latency or measured quality for the same multimodal task.
Why is a token-price ranking not enough?
Token price is only one part of the bill. A multimodal provider must convert an image, audio clip or video into billable units before a model can process it. Different models can turn the same image into very different token counts.
ClawRouters sent an identical 1024×768 PNG chart directly to five OpenAI-family models on 20 September 2026. The publisher reported that GPT-4o-mini used 25,522 prompt tokens, while GPT-4o used 786. Those figures concern OpenAI models only, but they show why a generic input-token price can mislead buyers.
At the prices used in that test, GPT-4o-mini cost 95% more per image than GPT-4o. The cheaper text-token rate lost because the image used far more tokens. A price comparison without image accounting cannot answer which API is cheapest for vision work.
The same test also found that request settings can change the result. Setting GPT-4o-mini image detail to “low” reduced the measured image count to 2,854 tokens. The publisher reported that the chart remained readable in its test, but warned that lower detail is not safe for dense text or fine-grained OCR.
That does not prove anything about Gemini, Mistral, Ling, DeepSeek or GLM models. It does set a useful buying rule: measure a representative file on the provider’s own endpoint before committing to a model. Do that for every input type your product accepts.
Which low-cost APIs have direct API access documented?
Mistral has the clearest first-party setup path in public sources. Mistral says Free mode enables API access by default without a credit card. Its quickstart shows a request to the https://api.mistral.ai/v1/chat/completions endpoint.
The same guide says Free mode has limited usage and rate limits. It does not give the numeric limits in the cited first-party page. A buyer should therefore treat Free mode as a test environment until the organisation’s dashboard shows the current allowance and limits.
Mistral says new accounts default to Free mode. It also says plan access applies across Studio, Vibe and API usage, which matters if several teams share an organisation account.
Google’s terms also establish that Gemini API access exists through a Cloud project. Google says paid Gemini API access requires a Cloud Project with an active billing account. The terms distinguish that route from unpaid quota and from direct interactions in Google AI Studio.
Google publishes a list of countries and territories where Gemini API and Google AI Studio are available. The page also says users outside listed regions should try Gemini through the Gemini Enterprise Agent Platform. Availability is therefore an early procurement check, not an afterthought.
What do the available listings say about models and costs?
The table keeps the published prices from the earlier comparison, but changes what they mean. They are text-token listings, not verified end-to-end multimodal quotes. They should be used to identify candidates for testing, rather than to declare a cheapest paid vision, audio or video API.
| Option | Published input / output price | Published context | What the available evidence supports | What remains unverified |
|---|---|---|---|---|
| InclusionAI Ling 3.0 Flash VL | $0.06 / $0.18 per 1M tokens | 131K tokens | Third-party directory listing of model name, context and token rates. | First-party modality support, direct API route, regions, image accounting, rate limits, latency and quality. |
| DeepSeek V4.1 Flash | $0.120 / $0.420 per 1M tokens | 1.0M tokens | Third-party directory listing of token rates and context. | Whether this specific model supports the required multimodal inputs, plus first-party image, audio and video accounting. |
| Z.ai GLM 5.3 Flash | $0.000 / $0.000 listed input and output | 1.0M tokens | Third-party directory listing only. | The terms, availability and scope of the listed zero-price entry, plus modality support and direct API conditions. |
| Google Gemini 3.6 Flash | Not provided in the supplied primary pricing material. | 1M tokens | Directory lists text, image, audio and video inputs; 15 RPM and 1,500 requests per day. | Like-for-like paid pricing, image tokenisation and comparative latency for a defined workload. |
| Mistral Free mode | Not provided in the supplied first-party pricing material. | Varies by model. | First-party documentation confirms no-card API access in Free mode. | Current numeric free allowance, specific multimodal model access and a first-party multimedia price schedule. |
Source: PricePerToken, Awesome Free LLM APIs, Mistral documentation, and Google Gemini API terms.
The published prices above are not interchangeable. They come from different services and evidence tiers, and they may include distinct routing, caching or access conditions.
What can free tiers actually support?
Free tiers are useful for proving that a workflow works. They are less useful for proving a production cost model. Limits can be low, they can change, and an app may need a paid plan before it can serve users in certain regions.
The directory lists Gemini 3.5 Flash-Lite at 30 requests per minute and 1,500 requests per day. It lists the same multimodal input categories as Gemini 3.6 Flash. This makes the tier suitable for small prototypes, but the directory is not Google’s price or capability contract.
The directory also lists Cohere Command A+ and Command A Vision as text-plus-image models on a trial key. It says the trial includes 1,000 API calls a month, with non-commercial use only. That is enough for a narrow evaluation, not sustained production traffic.
The same directory lists Z AI’s GLM-4.6V-Flash as multimodal with one concurrent request. This is a separate model from GLM 5.3 Flash. It should not be used as evidence that GLM 5.3 Flash has the same capabilities or pricing.
For Mistral, the official documentation is more useful on billing behaviour than on the exact free allowance. Mistral says each plan’s included monthly usage is shared across Studio, API and Vibe Code. A team could exhaust an allowance through testing in another Mistral product without noticing.
Mistral says usage can stop after the included allowance runs out if pay-as-you-go is not enabled. That protects against surprise spending. It can also stop an unattended workflow, so teams should alert on usage before launching scheduled jobs.
How should privacy and regional rules affect the choice?
Google gives a clear reason not to treat “free” as a generic answer. Google says it uses content and generated responses from unpaid services to provide, improve and develop products and services. The terms say human reviewers may read, annotate and process API input and output for quality purposes.
Google explicitly tells users not to submit sensitive, confidential or personal information to unpaid services. This includes files such as images, videos and documents. A team testing customer support screenshots, medical records or employee documents should not treat the free tier as a neutral default.
Paid access changes that boundary. Google says it does not use prompts, files or responses from Paid Services to improve its products. Google still says it logs prompts and responses for a limited period for safety, security and legal purposes.
Region is equally important. Google says API clients made available to users in the European Economic Area, Switzerland or the United Kingdom may use only Paid Services. A free-tier proof of concept cannot simply become a customer-facing launch in those markets.
Google’s availability page says API access is limited to its published countries and territories. Buyers should verify the location of their development team and their users before they build around an endpoint.
How should a team test a cheap multimodal API?
The right comparison is a small test based on the files and prompts that the product will actually use. Test each candidate through its direct endpoint where possible. Record the provider-returned usage fields, response time, error rate and answer quality.
- Use at least one representative image, document, audio clip or video clip for each planned feature.
- Keep the prompt, file and evaluation criteria identical across models.
- Record billable input and output units returned by the API, rather than estimating from file size.
- Measure the result at the default setting and at any lower-cost detail setting.
- Check whether rate limits permit the expected burst traffic.
- Repeat the test after a model or pricing change.
ClawRouters says it repeated its image-token measurement across three consecutive runs with identical results. That is a useful model for a procurement test: measure the provider response, document the method and limit conclusions to the models actually tested.
Quality needs its own score. A model that is cheap but misses figures in a chart, fails OCR or hallucinates facts from a video can cost more in manual review. public sources does not provide a dated benchmark that compares quality and latency for the same multimodal workload across the recommended candidates.
Which option fits each buying situation?
Choose Gemini’s free tier for a non-sensitive prototype when you need one API that the available directory lists as accepting text, images, audio and video. The directory lists Gemini 3.6 Flash with 15 RPM and 1,500 requests per day. Move to paid access before sending confidential data or serving users in the EEA, Switzerland or the UK.
Choose Mistral Free mode when no-card API setup is the immediate requirement and the team can validate the available models in its account. Mistral’s official quickstart says an API key can be created in Free mode. Confirm modality, rate limits and available usage before choosing it for a multimodal product.
Use Ling 3.0 Flash VL, DeepSeek V4.1 Flash or GLM 5.3 Flash only as test candidates until their vendors provide the facts needed for a real multimodal comparison. The directory’s low published token rates make Ling and DeepSeek worth putting on a benchmark shortlist. They do not prove that either model is a cheaper vision, audio or video API.
What we could not verify?
We could not verify first-party documentation for the supported inputs and outputs of InclusionAI Ling 3.0 Flash VL, DeepSeek V4.1 Flash or Z.ai GLM 5.3 Flash. We also could not verify their direct API availability, supported regions, current rate limits, image-token accounting, audio pricing, video pricing or first-party paid rate cards.
We could not verify a dated, like-for-like quality or latency comparison for these specific models on the same image, audio or video workload. The PricePerToken entries are directory listings, not a substitute for vendor documentation or a controlled test. They should not support a claim that any named model is the cheapest paid multimodal API.
For Gemini, the supplied first-party terms document data handling, paid access and regional restrictions, but they do not provide an exact model-by-model multimedia rate card. For Mistral, the supplied first-party pages document API setup and billing controls, but do not establish the free-mode multimodal model list or a current image, audio and video price schedule.
What should buyers watch next?
Watch for a provider-published multimedia pricing page that states how images, audio and video are metered, and test it against your own files. Also track model retirement schedules before building a production dependency. Google lists Gemini 2.5 Flash Image for shutdown on 2 October 2026, while Gemini 3.8 Flash has no shutdown date announced.
Sources
- https://openrouter.ai/models
- https://docs.mistral.ai/admin/billing-usage/subscriptions
- https://docs.mistral.ai/getting-started/quickstarts/studio/activate-and-generate-api-key
- https://ai.google.dev/gemini-api/docs/deprecations
- https://ai.google.dev/gemini-api/terms
- https://ai.google.dev/gemini-api/docs/available-regions
- https://www.siliconflow.com/articles/the-cheapest-llm-api-provider
- https://www.clawrouters.com/blog/cheapest-vision-multimodal-llm-api-2026
- https://pricepertoken.com/cheapest
- https://github.com/mnfst/awesome-free-llm-apis
- https://docs.bigmodel.cn/cn/faq/registration-login.md

