Claude API Cost Calculator

Enter your usage once and see what it costs per month on every current Claude model, with GPT-5.6 and Gemini rows for comparison. Prices come from the vendors' own docs, verified July 2026.

ModelCost per requestCost per month
Claude Haiku 4.5
$0.00450$45.00
Claude Sonnet 5
Introductory pricing through Aug 31, 2026; $3 input / $15 output per MTok from Sep 1, 2026.
$0.00900$90.00
Claude Sonnet 4.6
$0.0135$135.00
Claude Sonnet 4.5
$0.0135$135.00
Claude Opus 4.8
$0.0225$225.00
Claude Opus 4.7
$0.0225$225.00
Claude Opus 4.6
$0.0225$225.00
Claude Opus 4.5
$0.0225$225.00
Claude Fable 5
$0.0450$450.00
Claude Opus 4.1
Retires Aug 5, 2026. Migrate to Opus 4.8.
$0.0675$675.00
For comparison
GPT-5.6 Sol
OpenAI
$0.0250$250.00
GPT-5.6 Terra
OpenAI
$0.0125$125.00
Gemini 2.5 Pro
Google
Prompts over 200K tokens bill at $2.50 input / $15 output per MTok.
$0.00750$75.00
Gemini 2.5 Flash
Google
Text input; audio input bills at $1 per MTok.
$0.00185$18.50
DeepSeek V4 Flash
DeepSeek
Cache-miss input; cache-hit input is $0.0028 per MTok.
$0.00042$4.20
DeepSeek V4 Pro
DeepSeek
Cache-miss input; cache-hit input is $0.003625 per MTok.
$0.00130$13.05
GLM-5.2
Zhipu
Zhipu's Z.ai international API pricing (already quoted in USD).
$0.00500$50.00
Grok 4.5
xAI
$0.00700$70.00
Models from Opus 4.7 onward use a new tokenizer that produces roughly 30% more tokens for the same text than Sonnet 4.6 and earlier. When comparing per-token prices across that line, remember the newer models also count more tokens per document.

Frequently asked questions

Where do these prices come from?

Straight from the vendors' own documentation: Anthropic's pricing page for Claude, OpenAI's pricing page for GPT, and Google's Gemini API pricing page. Last verified July 12, 2026.

Does this include prompt caching or the Batch API?

No. These are the standard per-token API rates. Prompt caching and batch processing can reduce real costs substantially; check each vendor's pricing page for those rates.

What is the difference between input and output tokens?

Input tokens are everything you send the model: the system prompt, conversation history, and documents. Output tokens are what the model writes back. Output tokens usually cost around five times more, so long answers dominate the bill faster than long prompts.

Why does the same text cost more on newer Claude models?

Opus 4.7 and later use a new tokenizer that produces roughly 30% more tokens for the same text. Even at an identical per-token price, a document costs more to process on a new-tokenizer model than the raw rate suggests.

Tired of re-explaining yourself to every AI?

MemoryPlugin gives Claude, ChatGPT, Gemini, and 20+ other AI tools one shared long-term memory and a searchable chat history. Tell one AI something once, and the rest have it from your next conversation.

Try MemoryPlugin free