Last Updated: October 4, 2026

Claude API pricing ranges from $1 to $50 per million tokens, depending on which model you call and whether it's input or output — Haiku 4.5 is the cheapest at $1/$5 per million, while Fable 5.1 tops out at $10/$50. Most teams building on Claude use Sonnet 5 as the default workhorse at $2/$10 per million tokens, reserving Opus 5.5 for tasks that need the extra reasoning quality.
Here's exactly what each model costs, how prompt caching and the Batch API cut that bill by up to 90%, and what a realistic monthly cost actually looks like for common use cases.
Claude API Pricing by Model
Model | Input (per 1M tokens) | Output (per 1M tokens) | Best for |
|---|---|---|---|
Haiku 4.5 | $1.00 | $5.00 | High-volume, latency-sensitive tasks |
Sonnet 5 | $2.00 | $10.00 | General-purpose default — coding, writing, agents |
Opus 5.5 | $4.00 | $20.00 | Complex reasoning, high-stakes accuracy |
Fable 5.1 | $10.00 | $50.00 | Specialized/frontier workloads |
These are standard API rates, billed per million tokens processed — input tokens are what you send (your prompt, context, documents), output tokens are what the model generates back. Output almost always costs 5x the input rate across every model in the lineup, which matters more than it looks once you're generating long responses instead of just asking short questions.
How Each Model's Pricing Actually Breaks Down
Haiku 4.5 is the budget tier, built for tasks where speed and volume matter more than depth — classification, simple extraction, chat responses that don't need deep reasoning. At $1 input / $5 output per million tokens, it's cheap enough to run at genuine scale without the bill becoming the bottleneck.
Sonnet 5 is what most production applications actually run on. At $2/$10 per million tokens, it's priced as the default — capable enough for coding, agentic workflows, and long-form writing, without paying Opus-level rates for tasks that don't need them. If you're not sure which model to start with, this is the one.
Opus 5.5 costs double Sonnet's rate ($4/$20) and is built for the harder end of the task spectrum — multi-step reasoning, high-stakes accuracy, work where a wrong answer is expensive enough that the extra cost per call is worth it.
Fable 5.1 sits at the top of the pricing table ($10/$50) as the most specialized model in the current lineup, priced for workloads where its specific capabilities justify a 5x premium over Opus.
A detail that catches people off guard: newer model generations sometimes use a different tokenizer than older ones, which can mean the same piece of text breaks into more tokens than it used to — so a lower sticker price per token doesn't always translate to a lower total bill, a nuance confirmed across multiple current rate breakdowns tracking Anthropic's pricing changes through 2026. Always test actual token counts on your real prompts rather than assuming the rate card tells the whole story.
Cutting Your Actual Bill: Caching and Batch Processing
Two features meaningfully change what you actually pay, and most people underuse both.
Prompt caching can cut costs by up to 90% on repeated context. If your application sends the same system prompt, document, or instructions on every call, caching stores that content so subsequent calls only pay roughly a tenth of the standard input rate to reuse it. There are two cache tiers — a 5-minute cache and a longer 1-hour cache — with cache writes costing a premium (roughly 1.25x the base rate for the short cache, 2x for the long one) that pays for itself after just one or two reuses. For any application that sends a large, unchanging context block repeatedly — a long system prompt, a reference document, a codebase snapshot — caching is close to mandatory for cost control, not optional.
The Batch API applies a flat 50% discount on both input and output pricing for any workload that doesn't need an instant response — bulk classification jobs, overnight data processing, large-scale content generation. Results typically return within 24 hours. This discount stacks with prompt-cache pricing, so a cached, batched call can end up costing a fraction of the standard rate.
Together, a workload that uses both intelligently can realistically run at 20-30% of its "sticker price" cost — which is the gap between what a rate card suggests a project costs and what it actually costs once it's built properly.

Real-World Cost Examples
To make the per-token numbers concrete:
A customer support chatbot on Sonnet 5, handling a typical 500-token question with 300-token context and a 200-token response, costs roughly $0.003 per conversation — meaning 10,000 conversations a month runs about $30 in model costs before caching discounts.
A coding agent on Sonnet 5 working through a large codebase with heavy context reuse can cut costs substantially with prompt caching, since the codebase context — the expensive part — gets reused across many calls rather than resent in full each time.
A bulk document-summarization job on Haiku 4.5, run through the Batch API overnight, costs a fraction of what the same job would cost on Opus run synchronously during the day — the right model-plus-feature combination for the task can be the difference between a $50 job and a $500 one.
The pattern across all three: the model you pick and whether you use caching/batch matters more to your final bill than the headline per-token rate does.
How to Estimate Your Own Monthly Cost
Before committing to a model, run the math on your actual expected volume rather than guessing from the rate card. The formula is straightforward:
(Average input tokens per call × input rate) + (average output tokens per call × output rate) × expected monthly call volume = estimated monthly cost
A worked example: say you're building a feature that summarizes support tickets on Sonnet 5, averaging 800 input tokens (the ticket plus instructions) and 150 output tokens (the summary) per call, running 20,000 times a month. That's (800 × $0.000002) + (150 × $0.00001) = $0.0016 + $0.0015 = $0.0031 per call, or roughly $62/month at that volume — before any caching benefit, since each ticket is unique and there's limited repeated context to cache here.
Compare that to a different workload — an internal tool that answers questions against the same 50-page policy document all day, on the same model. There, the document itself (a large, unchanging block of context) is exactly what prompt caching is built for: pay the cache-write premium once, then pay roughly a tenth of the standard input rate on every subsequent call that reuses it. The ticket-summarizer example gets little benefit from caching; the policy-document example gets most of its savings from it. Knowing which category your actual workload falls into, before you build, is what separates an accurate cost estimate from a surprising bill. Our guide to how LLMs actually work covers the token-based mechanics behind this pricing model in more depth if the underlying concept is new.
Common Claude API Pricing Mistakes
A few patterns account for most unexpectedly high bills:
Defaulting to Opus for everything. Opus 5.5 costs double Sonnet 5's rate, and for a large share of tasks — drafting, classification, straightforward Q&A — Sonnet performs close enough that the extra cost buys little. Reserve Opus for the subset of calls where accuracy genuinely justifies it, not the whole pipeline by default.
Ignoring output length. Because output tokens cost roughly 5x input tokens across the board, an unconstrained prompt that lets the model ramble costs meaningfully more than one with a clear length instruction or max-token cap. This is one of the cheapest optimizations available and one of the most commonly skipped.
Not using the Batch API for anything that can wait. If a workload doesn't need a response in seconds — overnight processing, bulk analysis, non-interactive jobs — running it synchronously instead of through the Batch API means leaving a flat 50% discount on the table for no real benefit.
Resending the same large context on every call. Without caching, a workflow that repeatedly sends the same system prompt, reference document, or codebase pays full input price on that content every single time. This is usually the single biggest lever available once a workload is already live. For a broader look at where AI projects typically overspend, see our AI pricing guide.
Claude API vs. the Claude Pro/Max Subscription
These solve different problems and the pricing isn't really comparable directly. The Claude Pro ($20/month) and Max ($100-$200/month) subscriptions are for using Claude yourself, through the chat interface or Claude Code, with usage limits rather than per-token billing. The API is for building Claude into your own product or workflow, where you pay per token processed with no subscription fee and no usage cap beyond your own budget.
If you're an individual using Claude for your own work, the subscription is almost always cheaper and simpler. If you're building an application that other people or systems will call programmatically, the API is the only option — subscriptions aren't built for that. Some teams end up paying for both: a Max subscription for the people doing the work directly, and API credits for whatever they're building that needs to call Claude automatically.

Frequently Asked Questions
How much does the Claude API cost per month?
There's no fixed monthly cost — the API bills per million tokens processed, so your bill depends entirely on usage volume and which model you call. A low-volume app on Sonnet 5 might run $10-50/month; a high-volume production system can run into thousands. Use the per-model rates above and your expected token volume to estimate your actual cost.
Is Claude API pricing cheaper than ChatGPT's API?
It depends on the specific models compared and your usage pattern — both providers price their budget, mid, and frontier tiers differently, and discounts like caching and batch processing shift the real-world comparison further. Our Claude vs. ChatGPT comparison covers how the two platforms differ beyond price alone.
What's the difference between input and output token pricing?
Input tokens are what you send to the model — your prompt, any context, documents, or instructions — and output tokens are what the model generates in response. Output tokens cost roughly 5x the input rate across every Claude model, which is why long, verbose responses affect your bill more than long prompts do.
Does prompt caching actually save significant money?
Yes, substantially, but only if your application repeatedly sends the same context. Cached reads cost roughly a tenth of the standard input rate, so any workflow that reuses a large system prompt, document, or codebase across many calls sees real savings. A one-off query with no repeated context gets no benefit from caching at all.
Do I need a Claude Pro subscription to use the API?
No, they're entirely separate. API access is billed per token through Anthropic's developer console with no subscription required, while Pro and Max are consumer plans for using Claude directly through chat or Claude Code. For how the consumer plans work, see our complete guide to Claude AI.
Conclusion
Claude API pricing isn't one number — it's four models spanning $1 to $10 per million input tokens, with output consistently priced around 5x higher, and two features (prompt caching and the Batch API) that can cut a real-world bill by up to 90% when used correctly. The model choice matters less than most people assume in isolation; what actually determines your monthly cost is whether you're reusing context efficiently and routing non-urgent work through batch processing. Start with Sonnet 5 as your default, measure actual token usage on your real workload, and layer in caching before you worry about which model's sticker price looks cheapest.
By Sameer Khan
AI-assisted. Researched, reviewed, and edited by Sameer Khan before publication.

