Price board · verified 24 June 2026
What intelligence costs today.
Every model API on one board, priced per million tokens. No estimates and no affiliate rankings — the published number, and a link to where we read it.
10×
between the cheapest and dearest output token on this board — Claude Haiku 4.5 at $5.00 against Claude Fable 5 at $50.00.
| Relative cost | ||||
|---|---|---|---|---|
| Claude Opus 5Anthropic | $5.00 | $25.00 | 1M | |
| Claude Opus 4.8Anthropic | $5.00 | $25.00 | 1M | |
| Claude Fable 5Anthropic | $10.00 | $50.00 | 1M | |
| Claude Sonnet 5Anthropic · Introductory pricing to 2026-08-31 | $2.00 | $10.00 | 1M | |
| Claude Sonnet 4.6Anthropic | $3.00 | $15.00 | 1M | |
| Claude Haiku 4.5Anthropic | $1.00 | $5.00 | 200K |
Click any column to re-rank the board.
Cost calculator
Price your actual workload.
A sticker price per million tokens tells you very little on its own. Put in the shape of your traffic and see the monthly bill — including what prompt caching takes off it.
Estimated monthly spend
Prices are list rates before any committed-use discount.
- Claude Haiku 4.5cheapest$172.56
- Claude Sonnet 5$347.28
- Claude Sonnet 4.6$517.68
- Claude Opus 5$862.80
- Claude Opus 4.8$862.80
- Claude Fable 5$1,726
Tool index
The rest of the stack.
The model is one line item. These are the categories you end up choosing from around it, and what each one is actually for.
Gateways & routers
One endpoint in front of many models. Useful for failover, spend controls, and swapping models without touching application code.
4 listed
Coding agents
Agents that read, write, and run code in a real repository rather than suggesting snippets.
5 listed
Agent frameworks
Libraries for building the loop yourself: tool calling, state, orchestration, evaluation hooks.
6 listed
Vector & retrieval
Where your embeddings live, and what queries them when a model needs grounding.
6 listed
Evals & observability
Tracing, cost attribution, and regression testing. The category most teams skip and later wish they had not.
6 listed
Local inference
Run open-weight models on your own hardware — laptop, workstation, or a GPU box you own.
5 listed
Inference hosting
Someone else runs the GPUs and bills you per token or per second for open-weight models.
6 listed
Coverage
6 models verified. 7 providers queued.
We publish a price only after reading it off the provider's own page and recording the date. That is slower than scraping everything, and it is the entire reason to trust the board. The gaps are listed openly rather than filled in with guesses.