Small and fast AI models compared

Cheap, low-latency models for classification, routing, extraction, and cleanup. At these rates per-token cost stops being the constraint and throughput starts being one.

List prices checked

Token prices

ModelProviderInputOutputCachedBlendedContext
GPT-5 nanoOpenAI$0.050$0.400$0.0050$0.138400K
Gemini 2.5 Flash-LiteGoogle$0.100$0.400$0.025$0.1751M
Mistral SmallMistral AI$0.100$0.300$0.150131K
GPT-4o miniOpenAI$0.150$0.600$0.075$0.262128K
Cohere Command RCohere$0.150$0.600$0.262128K
GPT-4.1 miniOpenAI$0.400$1.60$0.100$0.7001M

Ranked for: In-app chat assistant

A support or help assistant handling a few thousand conversations a month, with short prompts and short answers.

#ModelInput costOutput costMonthly
1GPT-5 nano$1.00$2.00$3.00
2Mistral Small$2.00$1.50$3.50
3Gemini 2.5 Flash-Lite$2.00$2.00$4.00
4GPT-4o mini$3.00$3.00$6.00
5Cohere Command R$3.00$3.00$6.00
6GPT-4.1 mini$8.00$8.00$16

Ranked for: RAG over documents

Retrieval-augmented answering where each request stuffs several retrieved passages into the prompt, so input dominates.

#ModelInput costOutput costMonthly
1GPT-5 nano$15$6.00$21
2Mistral Small$30$4.50$35
3Gemini 2.5 Flash-Lite$30$6.00$36
4GPT-4o mini$45$9.00$54
5Cohere Command R$45$9.00$54
6GPT-4.1 mini$120$24$144

Ranked for: Coding agent

An agentic loop that reads files and writes patches, generating heavy output and re-sending large context on every turn.

#ModelInput costOutput costMonthly
1GPT-5 nano$20$40$60
2Mistral Small$40$30$70
3Gemini 2.5 Flash-Lite$40$40$80
4GPT-4o mini$60$60$120
5Cohere Command R$60$60$120
6GPT-4.1 mini$160$160$320

Head-to-head in this tier