- AI models
AI model API pricing (2026)
Token prices for 25 models from 7 providers, taken from each provider's own price list. Grouped by capability tier, because the cheapest model overall is rarely the answer to the question you are actually asking.
List prices checked
Frontier models
The most capable tier. Price spread inside this group is wide enough that the choice is usually a budget decision before it is a capability one.
| Model | Provider | Input /1M | Output /1M | Context | In-app chat assistant |
|---|---|---|---|---|---|
| DeepSeek-V4-Pro | DeepSeek | $0.435 | $0.870 | 1M | $13/mo |
| Mistral Large | Mistral AI | $2.00 | $6.00 | 131K | $70/mo |
| GPT-5 | OpenAI | $1.25 | $10.00 | 400K | $75/mo |
| Gemini 2.5 Pro | $1.25 | $10.00 | 1M | $75/mo | |
| OpenAI o3 | OpenAI | $2.00 | $8.00 | 200K | $80/mo |
| Cohere Command A | Cohere | $2.50 | $10.00 | 256K | $100/mo |
| Claude Sonnet 5 | Anthropic | $3.00 | $15.00 | 200K | $135/mo |
| Grok 4 | xAI | $3.00 | $15.00 | 256K | $135/mo |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | 200K | $225/mo |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | 200K | $450/mo |
Balanced models
Strong general-purpose models at a fraction of frontier pricing. Where most production traffic should land.
| Model | Provider | Input /1M | Output /1M | Context | In-app chat assistant |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash | DeepSeek | $0.140 | $0.280 | 1M | $4.20/mo |
| Grok 3 mini | xAI | $0.300 | $0.500 | 131K | $8.50/mo |
| Codestral | Mistral AI | $0.300 | $0.900 | 262K | $11/mo |
| GPT-5 mini | OpenAI | $0.250 | $2.00 | 400K | $15/mo |
| Gemini 2.5 Flash | $0.300 | $2.50 | 1M | $19/mo | |
| OpenAI o4-mini | OpenAI | $1.10 | $4.40 | 200K | $44/mo |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | 200K | $45/mo |
| GPT-4.1 | OpenAI | $2.00 | $8.00 | 1M | $80/mo |
| GPT-4o | OpenAI | $2.50 | $10.00 | 128K | $100/mo |
Small / fast models
Cheap, fast models for classification, routing, extraction, and cleanup passes at volume.
| Model | Provider | Input /1M | Output /1M | Context | In-app chat assistant |
|---|---|---|---|---|---|
| GPT-5 nano | OpenAI | $0.050 | $0.400 | 400K | $3.00/mo |
| Mistral Small | Mistral AI | $0.100 | $0.300 | 131K | $3.50/mo |
| Gemini 2.5 Flash-Lite | $0.100 | $0.400 | 1M | $4.00/mo | |
| GPT-4o mini | OpenAI | $0.150 | $0.600 | 128K | $6.00/mo |
| Cohere Command R | Cohere | $0.150 | $0.600 | 128K | $6.00/mo |
| GPT-4.1 mini | OpenAI | $0.400 | $1.60 | 1M | $16/mo |
Questions
- Which AI API is cheapest in 2026?
- On a blended 3:1 input-to-output basis, GPT-5 nano at $0.138 per million tokens — against Claude Fable 5 at $20.00, a spread of well over an order of magnitude. Cheapest is only the right question once you have fixed the capability tier, which is why the table is grouped by tier.
- Why compare on a blended rate instead of input price?
- Because input price alone flatters reasoning models. Their thinking tokens bill as output, and on an agentic workload output can exceed input several times over. The blended figure assumes three input tokens per output token, which is roughly where chat and retrieval workloads sit.
- Do these prices include batch or cached discounts?
- No. These are standard pay-as-you-go list prices. Cached input is shown as a separate column where the provider publishes it, because for agents resending the same context every turn it often matters more than the headline rate.