- AI models
- Small / fast
- Mistral Small
Mistral Small pricing
An open-weight model available both as a hosted API and as a download, which makes it the cheap tier you can also run yourself if the economics change.
List prices checked
Token pricing
Per million tokens, from the official price list.
| Token type | Price per 1M | Notes |
|---|---|---|
| Input | $0.100 | Everything you send, including the system prompt |
| Output | $0.300 | Everything the model generates, reasoning tokens included |
| Blended (3:1) | $0.150 | One number for ranking, at three input tokens per output token |
Source: Mistral AI pricing (May 2026). Published list prices only — committed-use discounts, regional variation, and per-request charges are not included. Not yet independently re-checked against the vendor’s page.
Capabilities
| Property | Value |
|---|---|
| Context window | 131K tokens |
| Inputs accepted | text, image |
| Cost to fill the context once | $0.01 |
What it costs per month
Token prices applied to three workload shapes. Input-heavy and output-heavy workloads rank models very differently.
In-app chat assistant
A support or help assistant handling a few thousand conversations a month, with short prompts and short answers.
$3.50/mo
- 20M input tokens$2.00
- 5M output tokens$1.50
RAG over documents
Retrieval-augmented answering where each request stuffs several retrieved passages into the prompt, so input dominates.
$35/mo
- 300M input tokens$30
- 15M output tokens$4.50
Coding agent
An agentic loop that reads files and writes patches, generating heavy output and re-sending large context on every turn.
$70/mo
- 400M input tokens$40
- 100M output tokens$30
Mistral Small vs other small / fast models
Same capability tier, ranked by blended token price.
| Model | Input | Output | Context | Compare |
|---|---|---|---|---|
| Gemini 2.5 Flash-Lite | $0.100 | $0.400 | 1M | Head to head |
| GPT-4o mini | $0.150 | $0.600 | 128K | Head to head |
| Cohere Command R | $0.150 | $0.600 | 128K | Head to head |
| GPT-5 nano | $0.050 | $0.400 | 400K | Head to head |
| GPT-4.1 mini | $0.400 | $1.60 | 1M | Head to head |
Questions about pricing
- How much does the Mistral Small API cost in 2026?
- $0.100 per million input tokens and $0.300 per million output tokens, from Mistral AI's published pricing. The table above converts that into a monthly figure for three common workloads.
- What is the Mistral Small context window?
- 131K tokens. Filling the window costs whatever that many input tokens cost — at $0.100 per million, a single full-context request runs about $0.01.
- Is Mistral Small cheaper than the alternatives?
- On a blended 3:1 input-to-output basis it works out at $0.150 per million tokens. The comparison table below ranks it against the other small / fast models in the dataset, and the head-to-head pages model both against identical workloads.