The reader who types claude api pricing into Google in 2026 is not window-shopping — they are halfway through a build, watching a token meter tick, and trying to decide whether to swap tiers before the next commit. This article is the honest rate card: every live Standard Global per-million-token price for Claude Opus 5.5, Opus 5, Opus 4.x, Sonnet 5.5, Sonnet 5, Sonnet 4.x, Haiku 4.5, Fable 5.1, and the retired SKUs still billable through Bedrock and Google Cloud — all verified today, 2026-10-06, against Anthropic's own pricing docs at docs.claude.com. Then it stacks that card against the dual-agent Planner-plus-Executor split that WizardGenie ships for game dev, because the real cost cut is not picking the cheapest Claude SKU — it's running one expensive reasoner for 500 planning tokens and one cheap typer for every diff.
What people typing claude api pricing actually want
DataForSEO lists claude api pricing at 5,400 searches a month, keyword difficulty 18, competition 0.03, intent commercial (verified 2026-10-06 against tools/research-pricing-article.md). The adjacent long-tails tell the same story: anthropic claude api pricing (720/mo, KD 7), claude code api pricing (320/mo, KD 12), claude ai api pricing (260/mo, KD 30). The intent is uniform — someone wants a rate card they can plug into a spreadsheet. The reason it's a commercial query, not an informational one, is that readers already know what Claude is; what they want is the per-million-token number and the fine print on cache and batch lanes.
Three facts frame the whole answer. One: Claude's rate schedule changed again in 2026 — Opus 5.5 landed at a lower base price than Opus 5 did, which is unusual (Anthropic has historically held frontier prices flat and shipped the step-down one tier below). Two: the retired Opus 4 and Opus 4.1 are still billable at the pre-5-series rate ($15 input / $75 output per MTok) on AWS Bedrock and Google Cloud only — the first-party Claude API has moved on. Three: Claude Haiku 4.5 at $1 input / $5 output per MTok is the cheapest frontier-adjacent SKU any major lab publishes, which makes it a serious Executor candidate for dual-agent coding setups.
Everything below cites the live source for every number. If a figure is quoted without a verification date, it does not belong in a pricing article — this cluster rotates too fast for memory.
The 2026 Claude API rate card (Standard Global, verified 2026-10-06)
Prices are per million tokens, Standard tier, Global scope. US-only inference carries a roughly 10 percent uplift on most SKUs; Bedrock cross-region matches the Standard Global column almost line for line. Source: Anthropic's own pricing docs page at docs.claude.com (verified 2026-10-06), cross-checked against the Claude Model Pricing — All Platforms PDF and the 2026-08-31 List Prices PDF, both hosted under www-cdn.anthropic.com.
| Model | Input / MTok | Output / MTok | 5-min cache write | 1-hour cache write | Cache read hit | Context |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 | $10.00 | $50.00 | $12.50 | $20.00 | $0.25 | 1M tokens |
| Claude Mythos 5.1 (limited) | $10.00 | $50.00 | $12.50 | $20.00 | $0.25 | 1M tokens |
| Claude Opus 5.5 | $4.00 | $20.00 | $5.00 | $8.00 | $0.20 | 1M tokens |
| Claude Opus 5 | $5.00 | $25.00 | $6.25 | $10.00 | $0.50 | 1M tokens |
| Claude Opus 4.8 / 4.7 / 4.6 / 4.5 | $5.00 | $25.00 | $6.25 | $10.00 | $0.50 | 200K tokens |
| Claude Opus 4.1 (retired, Bedrock + GCP) | $15.00 | $75.00 | $18.75 | $30.00 | $1.50 | 200K tokens |
| Claude Opus 4 (retired, GCP only) | $15.00 | $75.00 | $18.75 | $30.00 | $1.50 | 200K tokens |
| Claude Sonnet 5.5 | $2.00 | $10.00 | $2.50 | $4.00 | $0.20 | 1M tokens |
| Claude Sonnet 5 | $2.00 | $10.00 | $2.50 | $4.00 | $0.20 | 1M tokens |
| Claude Sonnet 4.6 / 4.5 | $3.00 | $15.00 | $3.75 | $6.00 | $0.30 | 200K tokens |
| Claude Sonnet 4 (retired, Bedrock + GCP) | $3.00 | $15.00 | $3.75 | $6.00 | $0.30 | 200K tokens |
| Claude Haiku 4.5 | $1.00 | $5.00 | $1.25 | $2.00 | $0.10 | 200K tokens |
| Claude Haiku 3.5 (retired, Bedrock + GCP) | $0.80 | $4.00 | $1.00 | $1.60 | $0.08 | 200K tokens |
Three honest observations from the table. First, output is five times input across every tier — a long agentic session that ships thousands of lines of generated code bills more on the output pass than on the context, which is why cache math matters less for coding than for retrieval. Second, Opus 5.5 is a genuine step-down from Opus 5 (same tier, 20 percent cheaper on both input and output) rather than a brand-new premium SKU — the premium SKU is now Fable 5.1 at $10 / $50. Third, Haiku 4.5 is the cheapest frontier-adjacent model on the market: $1 input and $5 output beats any Claude Haiku generation before it except the retired 3.5, which is 20 cents cheaper on input but materially weaker on the coding and reasoning benchmarks Anthropic publishes in the Models Overview.
Context windows and max output: tokens are only half the cost
Verified 2026-10-06 against the Claude Models Overview page at platform.claude.com: Opus 5.5, Sonnet 5.5, Sonnet 5, and Fable 5.1 all expose a 1 million token context window. Opus 4.x and Sonnet 4.x stay at 200K. Haiku 4.5 is 200K. Max synchronous output on the Messages API is 128K tokens for Opus 5.5, Sonnet 5.5, and Fable 5.1; Haiku 4.5 is 64K. On the Message Batches API, Opus 5.5 / Sonnet 5.5 / Opus 5 / Opus 4.8 / Opus 4.7 / Opus 4.6 / Sonnet 4.6 can be run with up to 300K output tokens via the output-300k-2026-03-24 beta header.
Why context size matters for the bill: at Opus 5.5's $4-per-MTok input rate, a 1M-token load is $4. Load it twice in a session (which is normal for long coding sessions) and you have paid $8 before the model has written a single token. For the same 1M load on Haiku 4.5 at $1 input, you pay $1 twice — $2 total. Context is where tier choice bites first on long sessions.
Max output caps also set the upper bound on how much code a single call can emit. For a dual-agent setup where the Planner is Opus 5.5 and the Executor is Haiku 4.5, the 64K cap on Haiku is almost always enough — a single diff is rarely bigger than 10K tokens and a whole-file rewrite rarely bigger than 40K. The 128K ceiling on Opus / Sonnet / Fable only comes into play when the Planner is also doing the writing, which is a pattern you want to avoid anyway (see dual-agent section below).
Prompt caching: the 10%, 5%, and 2.5% lanes
Prompt caching is the single biggest cost lever on long sessions. Verified 2026-10-06 on docs.claude.com: cache read hits bill at 10 percent of base input on most SKUs, 5 percent on Opus 5.5, and 2.5 percent on Fable 5.1 and Mythos 5.1 per the footnotes on the pricing page. That means a 500K-token system prompt that would cost $2 to re-send cold on Opus 5.5 costs $0.10 to re-send warm from the 1-hour cache. For an agentic session that re-sends the same big system prompt dozens of times, the cache lane turns the first request's full price into a one-time setup fee.
Cache write pricing is where the trade-off lives. 5-minute cache writes cost 1.25x base input on most SKUs. 1-hour cache writes cost 2x base input on most SKUs (except Opus 5.5, where the 1-hour write is 2x and the 5-min write is a mild 1.25x). The math only pays off if you expect to hit the cache more than two or three times inside the window — a Planner that fires one prompt and never comes back does not benefit; a long coding session with persistent system context benefits a lot.
For game dev specifically, the big cache win is the Mem Palace long-term memory that WizardGenie exposes (src/app/wizard-genie/page.tsx, verified 2026-10-06). The Mem Palace holds project conventions, past decisions, and codebase shape across sessions; feeding that same context to Claude every session via a 1-hour cache write means every session after the first pays 10 percent (or 5 percent on Opus 5.5) of the base input rate for the Mem Palace block, not 100 percent. Over a weekend game jam that opens the editor eight times, that is roughly 90 percent off the context cost on every session past the first.