Cheapest AI model for classification and tagging

High-volume, low-stakes calls that return a label or a few words.

List prices checked

How this is scored

Scored at 500M input and 5M output tokens. Output is almost irrelevant at this ratio, so this is effectively a ranking on input price — and at these volumes the gap between the cheap tier and the frontier tier is thousands of dollars a month.

All 25, ranked for this workload

#ModelProviderCost for this workloadNotes
1GPT-5 nanoOpenAI$27/moThe cheapest model OpenAI publishes, intended for classification, routing, and cleanup passes where the task is simple and the volume is enormous.
2Mistral SmallMistral AI$52/moAn open-weight model available both as a hosted API and as a download, which makes it the cheap tier you can also run yourself if the economics change.
3Gemini 2.5 Flash-LiteGoogle$52/moGoogle's cheapest model, keeping the full million-token window and multimodal input. Priced to compete directly with GPT-5 nano on high-volume classification work.
4DeepSeek-V4-FlashDeepSeek$71/moA million-token context window at a fraction of any Western model's price, with cache-hit input billed at a fiftieth of the base rate. DeepSeek has published notice of a significant price rise, so treat these rates as a floor rather than a plan.
5GPT-4o miniOpenAI$78/moThe model that made per-token cost stop mattering for most simple tasks, and the benchmark every subsequent cheap model has been priced against.
6Cohere Command RCohere$78/moThe cheap end of Cohere's range, aimed at retrieval pipelines where the model's job is to ground an answer in supplied documents rather than to reason from scratch.
7GPT-5 miniOpenAI$135/moA fifth of GPT-5's price with the same 400k context window, which makes it the cheapest way to process very long documents without dropping to a small model.
8Grok 3 minixAI$153/moAn unusually cheap reasoning model, with output priced below most competitors' input rates — which matters because reasoning models are output-heavy by nature.
9CodestralMistral AI$155/moA code-specialised model tuned for fill-in-the-middle completion, priced for the high request volume that inline editor autocomplete generates.
10Gemini 2.5 FlashGoogle$163/moThe best price-to-context ratio on the market: a million tokens of multimodal input for a fraction of what frontier models charge, with optional thinking budgets.
11GPT-4.1 miniOpenAI$208/moThe million-token context window at a fifth of the price, which is an unusual combination and the main reason this model is still in production stacks.
12DeepSeek-V4-ProDeepSeek$222/moFrontier-class reasoning at roughly a fifth of Western frontier pricing, with a million-token window. The trade-offs are Chinese data residency, a 500-request concurrency cap, and a price rise the vendor has already announced.
13Claude Haiku 4.5Anthropic$525/moThe fast, cheap end of the Claude family, priced at a third of Sonnet while keeping the same context window. Suited to classification, extraction, and high-volume routing.
14OpenAI o4-miniOpenAI$572/moCheap reasoning, strong on maths and code relative to its price. Same caveat as every reasoning model: budget for the thinking tokens, not just the visible answer.
15GPT-5OpenAI$675/moOpenAI's flagship, priced aggressively below the other frontier models on input and with a 90% cached-input discount that makes long system prompts nearly free to reuse.
16Gemini 2.5 ProGoogle$675/moA million-token context window with native video and audio input, priced at frontier-model rates below 200k tokens and at a higher tier above it — a detail that surprises people modelling long-context costs.
17Mistral LargeMistral AI$1,030/moThe French lab's flagship, priced below every US frontier model and hosted in the EU, which is often the deciding factor rather than the benchmark scores.
18GPT-4.1OpenAI$1,040/moA million-token context window at a mid-tier price, which remains the reason to reach for it over newer models when the whole codebase has to fit in one request.
19OpenAI o3OpenAI$1,040/moA reasoning model that spends tokens thinking before answering. The list price understates the real cost, because reasoning tokens bill as output and can dominate the total.
20GPT-4oOpenAI$1,300/moThe natively multimodal model that set the mid-market price for over a year. Now more expensive than newer models that outperform it, but still the default in a great deal of shipped code.
21Cohere Command ACohere$1,300/moCohere's enterprise flagship, built around retrieval-augmented generation and multilingual work, and sold heavily as a private deployment inside customer infrastructure.
22Claude Sonnet 5Anthropic$1,575/moThe workhorse of the Claude family and the default for production work at volume. Introductory pricing of $2/$10 per million tokens runs to 31 August 2026; the $3/$15 standard rate shown here is what applies after that.
23Grok 4xAI$1,575/moxAI's flagship, priced to match Claude Sonnet exactly, with a large context window and live access to posts on X as a differentiator.
24Claude Opus 5Anthropic$2,625/moAnthropic's model for complex agentic coding and enterprise work. Prompt caching cuts repeated input to a tenth of the base rate, which for an agent resending the same context every turn matters more than the headline price.
25Claude Fable 5Anthropic$5,250/moAnthropic's top model for long-running agentic work, priced at twice Opus 5. Worth it only where a longer autonomous run genuinely replaces human supervision, since the token cost compounds across a multi-hour session.

Published list prices only. Rows that do not publish a rate for an axis this workload depends on are excluded rather than shown as cheap.

Questions

Cheapest AI model for classification and tagging in 2026?
GPT-5 nano from OpenAI — $27/mo. Mistral Small is second at $52/mo. This ranking is specific to the workload described above; a different usage shape reorders it.
How is this ranking weighted?
Scored at 500M input and 5M output tokens. Output is almost irrelevant at this ratio, so this is effectively a ranking on input price — and at these volumes the gap between the cheap tier and the frontier tier is thousands of dollars a month.
Why does this differ from the general cheapest list?
Because a single price axis never describes a real workload. Ranking by headline rate answers "who is cheapest per unit"; this page answers "who is cheapest for this job", and the two orders are often very different — which is the whole reason it exists as a separate table.
What is not accounted for?
Committed-use discounts, regional price variation, per-request charges, support plans, and minimum retention. These are published list prices applied to one stated workload — a shortlist to price properly with the vendor, not a quotation.

Other workloads