Gemini 2.5 Flash-Lite pricing

Google's cheapest model, keeping the full million-token window and multimodal input. Priced to compete directly with GPT-5 nano on high-volume classification work.

List prices checked

By Google · Small / fast tier

Token pricing

Per million tokens, from the official price list.

Token typePrice per 1MNotes
Input$0.100Everything you send, including the system prompt
Output$0.400Everything the model generates, reasoning tokens included
Cached input$0.025Repeated prefixes served from cache
Blended (3:1)$0.175One number for ranking, at three input tokens per output token

Source: Google pricing (May 2026). Published list prices only — committed-use discounts, regional variation, and per-request charges are not included. Not yet independently re-checked against the vendor’s page.

Capabilities

PropertyValue
Context window1M tokens
Max output66K tokens
Inputs acceptedtext, image, audio, video
Cost to fill the context once$0.10

What it costs per month

Token prices applied to three workload shapes. Input-heavy and output-heavy workloads rank models very differently.

In-app chat assistant

A support or help assistant handling a few thousand conversations a month, with short prompts and short answers.

$4.00/mo

  • 20M input tokens$2.00
  • 5M output tokens$2.00
  • If fully cached$2.50

RAG over documents

Retrieval-augmented answering where each request stuffs several retrieved passages into the prompt, so input dominates.

$36/mo

  • 300M input tokens$30
  • 15M output tokens$6.00
  • If fully cached$14

Coding agent

An agentic loop that reads files and writes patches, generating heavy output and re-sending large context on every turn.

$80/mo

  • 400M input tokens$40
  • 100M output tokens$40
  • If fully cached$50

Gemini 2.5 Flash-Lite vs other small / fast models

Same capability tier, ranked by blended token price.

ModelInputOutputContextCompare
Mistral Small$0.100$0.300131KHead to head
GPT-4o mini$0.150$0.600128KHead to head
Cohere Command R$0.150$0.600128KHead to head
GPT-5 nano$0.050$0.400400KHead to head
GPT-4.1 mini$0.400$1.601MHead to head

Questions about pricing

How much does the Gemini 2.5 Flash-Lite API cost in 2026?
$0.100 per million input tokens and $0.400 per million output tokens, from Google's published pricing. The table above converts that into a monthly figure for three common workloads.
What is the Gemini 2.5 Flash-Lite context window?
1M tokens, with up to 66K tokens of output per request. Filling the window costs whatever that many input tokens cost — at $0.100 per million, a single full-context request runs about $0.10.
Does Gemini 2.5 Flash-Lite support prompt caching?
Yes. Cached input is billed at $0.025 per million tokens, roughly 4× cheaper than uncached input. For an agent that resends the same system prompt and file context on every turn, that discount usually matters more than the headline rate does.
Is Gemini 2.5 Flash-Lite cheaper than the alternatives?
On a blended 3:1 input-to-output basis it works out at $0.175 per million tokens. The comparison table below ranks it against the other small / fast models in the dataset, and the head-to-head pages model both against identical workloads.

Keep comparing