Google Gemini API pricing (2026): the per-token cost table

Short answer: Per million tokens, the Gemini API costs $0.10 in / $0.40 out for Gemini 2.5 Flash-Lite (the cheapest model from any major provider), $0.30 / $2.50 for Gemini 2.5 Flash (same rate as the newer 3.5 Flash-Lite), and $1.25 / $10 for Gemini 2.5 Pro (up to 200K tokens; $2.50 / $15 above that). Newest: Gemini 3.8 Flash at $0.75 / $3.75 (released Sept 2, 2026; introductory pricing through Dec 31, 2026). Note that Google now limits the 2.5 models to developers who have already used them — new projects are steered to 3.5 Flash-Lite or 3.8 Flash. Other current options: Gemini 3.5 Flash $1.50 / $9 and Gemini 3.1 Pro Preview $2 / $12. The Batch API takes 50% off and context caching cuts cached input ~90%. Verified against Google's official pricing page.

Gemini API model pricing (per 1M tokens)

ModelInputOutputBest for
Gemini 2.5 Flash-Lite (cheapest)$0.10$0.40High-volume classification, routing, extraction
Gemini 2.5 Flash$0.30$2.50Best value workhorse for most tasks
Gemini 3.1 Flash-Lite$0.25$1.50Newer lite tier, stronger reasoning
Gemini 3.5 Flash-Lite$0.30$2.50Latest lite tier; same price as 2.5 Flash, better quality
Gemini 3.5 Flash$1.50$9.00Higher quality, native thinking
Gemini 3.8 Flash (newest)$0.75 (intro)$3.75 (intro)Released Sep 2, 2026; Google's recommended Flash for new projects
Gemini 3.7 Flash / 3.6 Flash$0.75 (intro)$3.75 (intro)Earlier 2026 Flash releases at the same intro rate
Gemini 2.5 Pro (≤200K)$1.25$10.00Flagship reasoning at low cost
Gemini 2.5 Pro (>200K)$2.50$15.00Long-context flagship
Gemini 3.1 Pro Preview (≤200K)$2.00$12.00Newest Pro preview; hardest reasoning

Prices are USD per million tokens (MTok), standard paid tier. Flash and Flash-Lite use a single flat rate; Gemini 2.5 Pro and 3.1 Pro Preview charge more above 200K tokens. The $0.75 / $3.75 rate on Gemini 3.8, 3.7 and 3.6 Flash is introductory pricing through December 31, 2026, after which it rises to $1.50 / $7.50. Access to the Gemini 2.5 models (2.5 Flash-Lite, 2.5 Flash, 2.5 Pro) is now limited to developers who have actively used them; Google recommends new projects start on 3.5 Flash-Lite or 3.8 Flash. Google also offers Gemini 3.5 Flash Cyber, a specialized cybersecurity-focused model available through limited access programs with pricing disclosed on request — it is not part of the standard self-serve API pricing above. Note: Gemini 2.0 Flash and 2.0 Flash-Lite were shut down on June 1, 2026 — migrate to the 2.5 line or newer.

What is the cheapest Gemini model?

Gemini 2.5 Flash-Lite, at $0.10 / $0.40 per million tokens, is the cheapest Gemini model — and the cheapest production model from any major LLM provider. Three stacking levers cut it further:

For how Gemini stacks up against the other providers, see cheapest LLM API: OpenAI vs Anthropic vs Gemini.

Batch API and context caching discounts

FeatureDiscountApplies to
Batch API50% off in & outAsync jobs (24-hour SLA), all models
Context caching~90% off cached inputRepeated context, plus per-hour storage fee
Free tierNo charge (rate-limited)Testing and low-volume use

Worked example: cost of a typical Gemini 2.5 Flash request

A request that sends 20,000 input tokens and generates 2,000 output tokens on Gemini 2.5 Flash:

The same request on Gemini 2.5 Flash-Lite costs (20,000 × $0.10 + 2,000 × $0.40) / 1,000,000 = $0.0028 — about 4x cheaper, which is why Flash-Lite is the right pick for high-volume, low-complexity work.

How Gemini API pricing compares

Gemini is the price leader across the budget and mid tiers in 2026. Gemini 2.5 Flash-Lite ($0.10/$0.40) still edges OpenAI's new GPT-6 Luna ($0.10/$0.50) on output and undercuts Anthropic's Claude Haiku 4.5 ($1/$5) — though new projects may not get 2.5 access. Gemini 3.5 Flash-Lite ($0.30/$2.50) beats GPT-6.1 Sol ($2/$10) on price. Even Gemini 3.1 Pro Preview ($2 input) is cheaper than GPT-6 Astra and Claude Fable 5.1 ($10 input each). Full breakdown: Cheapest LLM API, OpenAI API pricing, and Anthropic API pricing.