Claude API Pricing in 2026 (Game Dev Cost Cut)

By Arron R.14 min read
Claude API pricing in 2026: Haiku 4.5 at $1/$5 per MTok, Sonnet 5.5 at $2/$10, Opus 5.5 at $4/$20, Fable 5.1 at $10/$50 — verified 2026-10-06. The real game dev

The reader who types claude api pricing into Google in 2026 is not window-shopping — they are halfway through a build, watching a token meter tick, and trying to decide whether to swap tiers before the next commit. This article is the honest rate card: every live Standard Global per-million-token price for Claude Opus 5.5, Opus 5, Opus 4.x, Sonnet 5.5, Sonnet 5, Sonnet 4.x, Haiku 4.5, Fable 5.1, and the retired SKUs still billable through Bedrock and Google Cloud — all verified today, 2026-10-06, against Anthropic's own pricing docs at docs.claude.com. Then it stacks that card against the dual-agent Planner-plus-Executor split that WizardGenie ships for game dev, because the real cost cut is not picking the cheapest Claude SKU — it's running one expensive reasoner for 500 planning tokens and one cheap typer for every diff.

Claude API pricing in 2026 game dev cost cut rate card panels
Four tiers, one axis: Haiku 4.5 at $1/$5, Sonnet 5.5 at $2/$10, Opus 5.5 at $4/$20, Fable 5.1 at $10/$50. Batch halves everything.

What people typing claude api pricing actually want

DataForSEO lists claude api pricing at 5,400 searches a month, keyword difficulty 18, competition 0.03, intent commercial (verified 2026-10-06 against tools/research-pricing-article.md). The adjacent long-tails tell the same story: anthropic claude api pricing (720/mo, KD 7), claude code api pricing (320/mo, KD 12), claude ai api pricing (260/mo, KD 30). The intent is uniform — someone wants a rate card they can plug into a spreadsheet. The reason it's a commercial query, not an informational one, is that readers already know what Claude is; what they want is the per-million-token number and the fine print on cache and batch lanes.

Three facts frame the whole answer. One: Claude's rate schedule changed again in 2026 — Opus 5.5 landed at a lower base price than Opus 5 did, which is unusual (Anthropic has historically held frontier prices flat and shipped the step-down one tier below). Two: the retired Opus 4 and Opus 4.1 are still billable at the pre-5-series rate ($15 input / $75 output per MTok) on AWS Bedrock and Google Cloud only — the first-party Claude API has moved on. Three: Claude Haiku 4.5 at $1 input / $5 output per MTok is the cheapest frontier-adjacent SKU any major lab publishes, which makes it a serious Executor candidate for dual-agent coding setups.

Everything below cites the live source for every number. If a figure is quoted without a verification date, it does not belong in a pricing article — this cluster rotates too fast for memory.

The 2026 Claude API rate card (Standard Global, verified 2026-10-06)

Prices are per million tokens, Standard tier, Global scope. US-only inference carries a roughly 10 percent uplift on most SKUs; Bedrock cross-region matches the Standard Global column almost line for line. Source: Anthropic's own pricing docs page at docs.claude.com (verified 2026-10-06), cross-checked against the Claude Model Pricing — All Platforms PDF and the 2026-08-31 List Prices PDF, both hosted under www-cdn.anthropic.com.

Model Input / MTok Output / MTok 5-min cache write 1-hour cache write Cache read hit Context
Claude Fable 5.1 $10.00 $50.00 $12.50 $20.00 $0.25 1M tokens
Claude Mythos 5.1 (limited) $10.00 $50.00 $12.50 $20.00 $0.25 1M tokens
Claude Opus 5.5 $4.00 $20.00 $5.00 $8.00 $0.20 1M tokens
Claude Opus 5 $5.00 $25.00 $6.25 $10.00 $0.50 1M tokens
Claude Opus 4.8 / 4.7 / 4.6 / 4.5 $5.00 $25.00 $6.25 $10.00 $0.50 200K tokens
Claude Opus 4.1 (retired, Bedrock + GCP) $15.00 $75.00 $18.75 $30.00 $1.50 200K tokens
Claude Opus 4 (retired, GCP only) $15.00 $75.00 $18.75 $30.00 $1.50 200K tokens
Claude Sonnet 5.5 $2.00 $10.00 $2.50 $4.00 $0.20 1M tokens
Claude Sonnet 5 $2.00 $10.00 $2.50 $4.00 $0.20 1M tokens
Claude Sonnet 4.6 / 4.5 $3.00 $15.00 $3.75 $6.00 $0.30 200K tokens
Claude Sonnet 4 (retired, Bedrock + GCP) $3.00 $15.00 $3.75 $6.00 $0.30 200K tokens
Claude Haiku 4.5 $1.00 $5.00 $1.25 $2.00 $0.10 200K tokens
Claude Haiku 3.5 (retired, Bedrock + GCP) $0.80 $4.00 $1.00 $1.60 $0.08 200K tokens

Three honest observations from the table. First, output is five times input across every tier — a long agentic session that ships thousands of lines of generated code bills more on the output pass than on the context, which is why cache math matters less for coding than for retrieval. Second, Opus 5.5 is a genuine step-down from Opus 5 (same tier, 20 percent cheaper on both input and output) rather than a brand-new premium SKU — the premium SKU is now Fable 5.1 at $10 / $50. Third, Haiku 4.5 is the cheapest frontier-adjacent model on the market: $1 input and $5 output beats any Claude Haiku generation before it except the retired 3.5, which is 20 cents cheaper on input but materially weaker on the coding and reasoning benchmarks Anthropic publishes in the Models Overview.

Context windows and max output: tokens are only half the cost

Verified 2026-10-06 against the Claude Models Overview page at platform.claude.com: Opus 5.5, Sonnet 5.5, Sonnet 5, and Fable 5.1 all expose a 1 million token context window. Opus 4.x and Sonnet 4.x stay at 200K. Haiku 4.5 is 200K. Max synchronous output on the Messages API is 128K tokens for Opus 5.5, Sonnet 5.5, and Fable 5.1; Haiku 4.5 is 64K. On the Message Batches API, Opus 5.5 / Sonnet 5.5 / Opus 5 / Opus 4.8 / Opus 4.7 / Opus 4.6 / Sonnet 4.6 can be run with up to 300K output tokens via the output-300k-2026-03-24 beta header.

Why context size matters for the bill: at Opus 5.5's $4-per-MTok input rate, a 1M-token load is $4. Load it twice in a session (which is normal for long coding sessions) and you have paid $8 before the model has written a single token. For the same 1M load on Haiku 4.5 at $1 input, you pay $1 twice — $2 total. Context is where tier choice bites first on long sessions.

Max output caps also set the upper bound on how much code a single call can emit. For a dual-agent setup where the Planner is Opus 5.5 and the Executor is Haiku 4.5, the 64K cap on Haiku is almost always enough — a single diff is rarely bigger than 10K tokens and a whole-file rewrite rarely bigger than 40K. The 128K ceiling on Opus / Sonnet / Fable only comes into play when the Planner is also doing the writing, which is a pattern you want to avoid anyway (see dual-agent section below).

Prompt caching: the 10%, 5%, and 2.5% lanes

Prompt caching is the single biggest cost lever on long sessions. Verified 2026-10-06 on docs.claude.com: cache read hits bill at 10 percent of base input on most SKUs, 5 percent on Opus 5.5, and 2.5 percent on Fable 5.1 and Mythos 5.1 per the footnotes on the pricing page. That means a 500K-token system prompt that would cost $2 to re-send cold on Opus 5.5 costs $0.10 to re-send warm from the 1-hour cache. For an agentic session that re-sends the same big system prompt dozens of times, the cache lane turns the first request's full price into a one-time setup fee.

Cache write pricing is where the trade-off lives. 5-minute cache writes cost 1.25x base input on most SKUs. 1-hour cache writes cost 2x base input on most SKUs (except Opus 5.5, where the 1-hour write is 2x and the 5-min write is a mild 1.25x). The math only pays off if you expect to hit the cache more than two or three times inside the window — a Planner that fires one prompt and never comes back does not benefit; a long coding session with persistent system context benefits a lot.

For game dev specifically, the big cache win is the Mem Palace long-term memory that WizardGenie exposes (src/app/wizard-genie/page.tsx, verified 2026-10-06). The Mem Palace holds project conventions, past decisions, and codebase shape across sessions; feeding that same context to Claude every session via a 1-hour cache write means every session after the first pays 10 percent (or 5 percent on Opus 5.5) of the base input rate for the Mem Palace block, not 100 percent. Over a weekend game jam that opens the editor eight times, that is roughly 90 percent off the context cost on every session past the first.

Claude API pricing rate card Standard Global 2026 comparison matrix
The full card at a glance — six active SKUs on first-party Claude API, three retired SKUs still billable through Bedrock and Google Cloud.

Batch API: the 50% lane for jobs you can leave running

Batch API is a flat 50 percent off every Standard tier price (verified 2026-10-06 on docs.claude.com and the Model Pricing — All Platforms PDF at www-cdn.anthropic.com). On Haiku 4.5, Batch lands at $0.50 input and $2.50 output per MTok — the kind of number that makes bulk jobs genuinely cheap. The trade-off is latency: Batch returns within 24 hours instead of synchronously, which rules it out for anything in a tight iteration loop.

Three Batch-shaped jobs that fit indie game dev:

  1. NPC dialogue sweep. Fifteen NPCs, three story beats each, each beat 300 output tokens. On Sonnet 5.5 Standard that is roughly 14K output tokens = $0.14. On Sonnet 5.5 Batch that is $0.07. The savings are small per project — the point is the whole sweep lands overnight without you watching a progress bar.
  2. Overnight code review. Feed Opus 5.5 the whole repo (say 400K input tokens) with a review prompt, ask for a prioritized issue list, let Batch return it by morning. Standard = $1.60 input. Batch = $0.80 input. For a game jam post-mortem, the Batch latency is a feature, not a bug — the reviewer is a human on Monday morning anyway.
  3. Asset description pass. Batched Alt-text, OG descriptions, and store-page blurbs across every image in a release build. Haiku 4.5 Batch at $0.50 / $2.50 is cheaper than any manual copywriter, and the quality bar for Alt text is "accurate and short," which Haiku clears.

What does NOT fit Batch is the thing most people reach for Claude for: interactive game-loop coding. If you are watching a sprite move across the screen and asking the model why gravity feels wrong, Batch's 24-hour return defeats the point. Use Standard for the loop, Batch for the chores.

Three worked game dev scenarios (actual numbers)

Numbers below assume 2026-10-06 Standard Global pricing on docs.claude.com and the Sorceress credit math at CREDITS_PER_DOLLAR = 100 (src/lib/models.ts line 69, verified 2026-10-06). All token counts are typical sizes for the kind of jam-grade game loops MDN's Anatomy of a video game describes.

Scenario 1: single agentic session on Opus 5.5 (no dual-agent)

One two-hour coding session. System prompt + project context = 60K input tokens. Across the session the agent reads and re-reads context eight times (480K total input, mostly cached after the first pass). It emits six diffs averaging 15K output tokens each (90K output total). Cold input cost: $0.24. Warm cache reads (seven of eight loads at $0.20 per MTok, 5 percent rate on Opus 5.5): $0.084. Output cost at $20 / MTok: $1.80. Session total: roughly $2.12. The output pass dominates, which is why cutting the Executor rate is the biggest lever.

Scenario 2: dual-agent split, Opus 5.5 Planner + Haiku 4.5 Executor

Same session shape. Planner handles six planning passes at 2K input + 500 output each. Planner cost: 12K input at $4/MTok = $0.048 and 3K output at $20/MTok = $0.06. Planner subtotal: roughly $0.11. Executor types the diffs. Six diffs at 10K input (file + Planner brief) + 15K output each. Executor cost: 60K input at $1/MTok = $0.06 and 90K output at $5/MTok = $0.45. Executor subtotal: $0.51. Session total: roughly $0.62. That is 71 percent off the single-frontier run — the "roughly a quarter of the token cost" the WizardGenie product label (src/app/wizard-genie/page.tsx lines 297 through 299, verified 2026-10-06) advertises, computed against actual 2026-10-06 Claude prices.

Scenario 3: dual-agent split, Opus 5.5 Planner + DeepSeek V4 Pro Executor

Same session shape, Executor swapped for a non-Claude cheap model. Planner cost unchanged at roughly $0.11. Executor on DeepSeek V4 Pro is substantially cheaper than Haiku 4.5 (DeepSeek V4 Pro's public published rates sit well below $1 input and $5 output as of this writing — exact numbers are the next article in this wave). Session total lands around $0.30 to $0.40 depending on exact DeepSeek pricing on the day. The Planner-plus-Executor pattern is engine-agnostic — the only requirement is that the Executor is genuinely cheap. Running Sonnet or Opus as the typer erases the win and defeats the pattern (see Hard Rule 14 in any serious coding-cost analysis — pairings on both sides have to be lopsided for the economics to work).

When to pay Anthropic directly vs run Claude through the Sorceress wallet

Both lanes are honest. The right pick depends on billing shape, not model quality.

Pay Anthropic directly (BYOK) when: you already have an Anthropic account with volume discounts, you need the Batch API for scheduled jobs, you need the cache-write tier controls to tune a long session, or you are shipping a commercial product that bills end users per API call and needs to pass the cost through at Anthropic's exact rate. The dual-agent split works in BYOK mode too — WizardGenie supports bring-your-own-key configuration for each model slot, so the Planner can be your Anthropic account and the Executor can be your DeepSeek account.

Run Claude through the Sorceress wallet when: you want one invoice for Claude plus Quick Sprites plus Music Gen plus SFX Gen plus 3D Studio, you want credit math that does not reset on a billing date (Sorceress credits do not expire; Lovable monthly plan credits do, which was a sticking point in last week's companion piece on vibe coding stacks), or you want signup to come with a working balance — SIGNUP_GRANT = 100 credits at the CREDITS_PER_DOLLAR = 100 rate is a free dollar on the house to prove the dual-agent workflow before you spend anything of your own.

Either way, the model lineup WizardGenie exposes today (src/app/_home-v2/_data/tools.ts lines 767 through 774, verified 2026-10-06) is Claude Opus 4.7, Claude Sonnet 4.6, GPT-5.5, Gemini 3.1 Pro, DeepSeek V4 Pro, Kimi K2.5, Grok 4.2, and MiniMax M2.7. The 4.7 and 4.6 Claude SKUs bill at $5 / $25 and $3 / $15 respectively on first-party Claude API — the standard 4.x rates. The Opus 5.5 and Sonnet 5.5 upgrades at the lower rates ($4 / $20 and $2 / $10) are available via BYOK direct to Anthropic until the Sorceress-visible lineup rotates. That is the honest state of the integration today.

WizardGenie dual-agent split cuts Claude bill by 75 percent diagram
The whole economic story: a 500-token planning brief on Opus steers 15K-token diffs on Haiku. Same quality, roughly one quarter of the single-frontier bill.

How the Planner-plus-Executor split cuts the Claude bill by about 75 percent

The pattern is simple enough to fit on a napkin. The Planner reads the game spec and emits a short, structured brief — which file to change, which test proves success, which systems are off-limits. The Executor reads the brief and the file, then types the diff. The Planner never writes code; the Executor never decides what to build. Both constraints are what keep the cost ratio honest.

A Planner brief for a jam platformer, in the shape the Executor actually needs:

Goal: player hitbox is taller than the sprite art.
Files allowed: src/player.js and src/collision.js only.
Change: shrink the collision rectangle to 70% of sprite height.
Centre the shrunken rect vertically on the sprite origin.
Done when: standing under a 2-tile ceiling no longer kills the player.
Do not refactor the gravity constant or the input handler.

That brief is roughly 500 tokens. On Opus 5.5 at $4 input, the planning pass costs under a cent. The Executor reads the brief plus 15K tokens of file content, writes a diff of maybe 2K tokens. On Haiku 4.5 at $1 input / $5 output, the typing pass costs around $0.025. Full cycle: under three cents. For contrast, the same change done end-to-end by Opus 5.5 would involve reading the file, reading its own working output, and emitting a longer explanation around the diff — probably eight to ten cents per cycle. Across a two-hour session with twenty such cycles, the split saves roughly $1.40. Across a weekend jam with ten sessions, that is $14 saved — real money for a solo indie whose entire art budget might be $30.

The pattern breaks the moment the Executor is also frontier-priced. Running Opus 5.5 as both Planner and Executor doubles the planning cost (you are now paying $4-input tier for the typing pass too) and does not increase quality — the extra reasoning overhead on the output pass is wasted on a straightforward file edit. The same logic rules out Sonnet 5.5 as the Executor on volume projects (Sonnet 5.5 is $2 / $10, roughly 2x Haiku 4.5 on both lanes). The acceptable Executors in 2026 are Haiku 4.5, DeepSeek V4 Pro, Kimi K2.5, MiniMax M2.7, Gemini 3.1 Flash, and GPT-5.5 Mini — and only on the condition that none of them end up doing the Planner's job.

Verdict: three honest picks for 2026 Claude API coding

If you want one model for everything, pick Claude Sonnet 5.5. At $2 input / $10 output per MTok, it is the single best balance of frontier-grade reasoning and reasonable per-token cost for interactive coding. The 1M context lets it hold a whole indie-game project in one shot. The 10 percent cache read rate makes long sessions cheap after the first warm-up.

If you want the cheapest honest tier, pick Claude Haiku 4.5. At $1 / $5, it is cheap enough to run as a daily driver on side projects, and the 200K context is enough for any single-file edit. Haiku 4.5 is also the only Claude SKU that holds up as a dual-agent Executor if you insist on keeping the whole stack inside Claude.

If you want the dual-agent split that pays off, pair Opus 5.5 Planner with DeepSeek V4 Pro Executor on WizardGenie. That is the shape the companion pieces on vibe coding with Claude and executor picks both argue for, and it is the shape the Lovable vs WizardGenie comparison uses to show the roughly 1/4 cost delta that the WizardGenie product label advertises verbatim.

The pricing cluster rotates every few weeks. The last public pricing article on this blog shipped 2026-05-04 (AI Coding API Pricing (2026)) and already needs a refresh across every row; this article is dated 2026-10-06, and the next one — Grok API Pricing — ships in roughly a week on the same spine of live-verified numbers. Browse the full Sorceress stack in the tools guide, or jump straight to the WizardGenie dual-agent lane and burn the signup credit on a session that proves the math.

Frequently Asked Questions

What is the current Claude API pricing in 2026?

Verified 2026-10-06 at docs.claude.com/en/docs/about-claude/pricing: Claude Fable 5.1 is $10 input / $50 output per million tokens, Claude Opus 5.5 is $4 / $20, Claude Opus 5 is $5 / $25, Claude Opus 4.8 / 4.7 / 4.6 / 4.5 are each $5 / $25, Claude Sonnet 5.5 is $2 / $10, Claude Sonnet 5 is $2 / $10 at Standard Global, Claude Sonnet 4.6 / 4.5 are $3 / $15, and Claude Haiku 4.5 is $1 / $5. Retired Opus 4.1 / 4 still appear at $15 / $75 for the Bedrock and Google Cloud lanes. Batch API is a flat 50% off every Standard tier price. Prompt-cache read hits are 10% of base input on most SKUs — 5% on Opus 5.5, and 2.5% on Fable 5.1 and Mythos 5.1 per the footnotes on the same page.

How much is Claude Opus vs Sonnet vs Haiku per million tokens?

The 2026 gap between tiers (Standard Global, verified 2026-10-06) is clean: Haiku 4.5 at $1 input and $5 output, Sonnet 5.5 at $2 input and $10 output, Opus 5.5 at $4 input and $20 output. Fable 5.1, the top-tier reasoning model, sits at $10 input and $50 output. Output is five times input across every tier, which matters when a long agentic session returns thousands of lines of generated code — the output pass dominates the bill. For dual-agent splits, pair an Opus or Sonnet Planner with a Haiku 4.5 Executor and the Executor side of the bill drops by 4x versus running Opus end-to-end.

Does Claude code API pricing differ from Claude API pricing?

No. The Claude Code CLI and the Claude Code endpoints bill against the same per-million-token rates listed on docs.claude.com/en/docs/about-claude/pricing (verified 2026-10-06). What differs is prompt shape: an interactive coding session typically ships a long system prompt plus a growing conversation history, so prompt caching with the 10% read-hit rate (5% on Opus 5.5) is where real savings happen. The DataForSEO cluster lists claude code api pricing at 320 searches a month, KD 12 — the same intent as the broader claude api pricing head term, just filtered to the coding product.

What is the Batch API discount and when does it make sense for game dev?

Batch API is 50% off every Standard tier price (verified 2026-10-06 at docs.claude.com + the Model Pricing — All Platforms PDF). Jobs submitted to Batch return within 24 hours instead of synchronously. For a solo indie, three jobs fit this shape: bulk NPC dialogue generation across a cast of 15, overnight code review across a whole repo, and batched asset descriptions for a quest log. Interactive game loop coding does NOT fit Batch — the 24-hour return window kills the iteration tempo. Use Batch for the stuff you can leave running overnight; use Standard for the stuff you're watching compile.

How does WizardGenie cut Claude API bills with dual-agent coding?

WizardGenie's product label (src/app/wizard-genie/page.tsx lines 297-299, verified 2026-10-06) is explicit: 'A smart Planner thinks; a cheap Executor codes. Same quality at roughly a quarter of the token cost.' The pattern pairs an expensive reasoner (Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro, Grok 4.2) with a genuinely cheap typer (DeepSeek V4 Pro, Kimi K2.5, MiniMax M2.7, Gemini 3.1 Flash, Claude Haiku 4.5). The economic math only works when the Executor is actually cheap — running Claude Sonnet or Opus as the typer erases the savings. The whole stack bills in Sorceress credits at CREDITS_PER_DOLLAR = 100 (src/lib/models.ts line 69), so a dual-agent session is one line on one invoice regardless of which Planner or Executor you picked.

Are retired Claude models still billable through the API?

Partially. Verified 2026-10-06 at docs.claude.com/en/docs/about-claude/pricing: Claude Opus 4.1 and Claude Opus 4 are retired on the first-party Claude API but still available (and still billable at $15 input / $75 output) through AWS Bedrock and Google Cloud. Claude Sonnet 4 is retired on first-party Claude API but still live at $3 / $15 on Bedrock and Google Cloud. Claude Haiku 3.5 is retired first-party, still live at $0.80 / $4 on Bedrock and Google Cloud. For a new project in 2026, there is no reason to pin a retired model — Opus 5.5 is cheaper than Opus 4.1 on every lane, and Haiku 4.5 is only $0.20 above Haiku 3.5 input while shipping a far better reasoning benchmark.

Sources

  1. Anthropic Claude API Pricing — docs.claude.com
  2. Claude Models Overview — platform.claude.com
  3. Claude Model Pricing — All Platforms (PDF)
  4. Large language model — Wikipedia
  5. Vibe coding — Wikipedia
  6. Game development — MDN Web Docs
Written by Arron R.·3,137 words·14 min read

Related posts