Tooling

Every Major AI Model in 2026: Claude, GPT, Gemini & Grok Pricing, Benchmarks, and How to Choose

Token rates, benchmarks, context windows, and which model fits which job and which professional, across Claude, GPT, Gemini, and Grok. One reference, kept current.

Every Major AI Model in 2026 — AISEOShift

Pricing for frontier AI models changes almost monthly, and the official pages are scattered across API docs, plan pages, and help-center rate cards. This is a single, current reference for the four stacks most builders are comparing right now: Anthropic’s Claude, OpenAI’s GPT and Codex, Google’s Gemini, and xAI’s Grok.

Last updated: June 10, 2026. Everything below is accurate as of that date. Because these numbers move fast, treat this as a working snapshot and confirm against the official pricing pages (linked at the end) before you budget anything important.

A note on how to read token pricing: rates are quoted per one million tokens, split into input (what you send) and output (what the model generates). Output almost always costs several times more than input, so output volume is usually what drives your bill.

The Claude model lineup (June 2026)

Anthropic now runs a four-tier lineup, from the new Mythos-class flagship down to the fast, cheap workhorse.

ModelInput ($/M)Output ($/M)ContextMax outputBest for
Claude Fable 5$10$501M128kHardest agentic, coding, research
Claude Opus 4.8$5$251M-Flagship reasoning, complex work
Claude Sonnet 4.6$3$151M-Balanced daily driver
Claude Haiku 4.5$1$5--High-volume, fast, cheap tasks

A few things worth unpacking.

Claude Fable 5 (released June 9, 2026) is the first publicly available model from Anthropic’s Mythos class. It supports a 1M-token context window by default and up to 128k output tokens per request. Adaptive thinking is always on, and it ships with effort and task budgets, memory, context editing, compaction, and vision input. In unguarded domains it sits a clear tier above Opus 4.8, reportedly scoring 95.0% versus 88.6% on SWE-bench Verified and 80.0% versus 69.2% on SWE-bench Pro. One important caveat: Fable 5 is the safeguarded public deployment of the same weights as the restricted Claude Mythos 5, and requests touching cybersecurity, biology, chemistry, or distillation fall back to Opus 4.8. It is also a Covered Model with 30-day data retention and no zero-data-retention option.

Claude Opus 4.8 (launched May 28, 2026) is the flagship for complex reasoning and the default for most heavy knowledge work. At $5/$25 it costs a fraction of older Opus generations, and it offers a Fast Mode at $10/$50 for faster output without dropping to a smaller model. It supports the 1M-token context window.

Claude Sonnet 4.6 at $3/$15 is the balanced daily driver: most of the capability for routine work at a meaningfully lower price, also with 1M context.

Claude Haiku 4.5 at $1/$5 is the speed-and-volume tier for classification, extraction, routing, and other high-throughput tasks where latency and cost matter more than maximum reasoning.

The Claude cost levers that actually matter

The headline token rate is only the starting point. Three levers move your real bill more than which model you pick:

Prompt caching cuts cached input by about 90%. If you reuse the same context across turns (a system prompt, a codebase, a long document), caching is the single biggest optimization available. In agentic and chat workloads where context repeats constantly, this is the difference between a sane bill and a scary one.

Batch processing is 50% cheaper. For anything asynchronous (bulk generation, evaluations, offline pipelines), the Batch API halves the cost across models.

Output is 5x input across the lineup. Because output is the expensive half, the cheapest optimization is often just asking for less: tighter outputs, structured formats, and not regenerating what you already have.

The 1M context window is a flat rate on the newer models. Opus 4.8, Opus 4.7, and Sonnet 4.6 support 1M-token context with no surcharge, so you can use long context without a pricing penalty.

Claude subscription plans

If you are using Claude through the app or Claude Code rather than the raw API, you are on a subscription, not token billing.

PlanPriceWhat you get
Free$0Basic access
Pro$20/mo ($17/mo annual)5x usage, flagship Opus access, Claude Code in terminal/web/desktop
Max 5x$100/mo~5x Pro usage (~88k tokens / 5-hr window)
Max 20x$200/mo~20x Pro usage (~220k tokens / 5-hr window)
Team Standard$25/seat/mo ($20 annual)Collaboration, min 5 seats
Team Premium$125/seat/mo ($100 annual)Adds Claude Code
EnterpriseCustom500k context, SCIM, audit logs, Compliance API

For Claude Code specifically, Anthropic’s own data puts the average user at about $6 per developer per day, with 90% of users staying under $12/day. At full-time usage that lands around $100-$200 per developer per month, which is exactly where the Max plan sits. So the practical rule: heavy Claude Code users usually come out ahead on Max 5x or 20x rather than paying per token.

OpenAI Codex: what it actually is

Codex is easy to misunderstand because the name has been reused. In 2026, Codex is OpenAI’s cloud-based autonomous coding agent, powered by the GPT-5 family. It runs multi-step software tasks in isolated sandboxes and can work in parallel across many projects. It is the closest direct competitor to Anthropic’s Claude Code, not a single model.

The billing model changed meaningfully this year. On April 2, 2026, OpenAI switched Codex from per-message billing to token-based credits across all plans. Credits are now consumed as (input tokens x input rate) + (cached input tokens x cached rate) + (output tokens x output rate), with rates varying by the underlying model.

Codex models and API pricing

ModelInput ($/M)Output ($/M)Notes
GPT-5.3-Codex$1.75$14Default for many agent tasks and code review
GPT-5.4-Minilower-cost tierlower-cost tierDesigned for lighter tasks, lower credit rate
GPT-5.5$5$30Higher capability, uses fewer tokens; $0.50 cached input

GPT-5.3-Codex (released February 24, 2026) is the workhorse: strong agentic performance at $1.75/$14, and the default for code review and many agent tasks.

GPT-5.4-Mini is the lighter, cheaper option for simpler tasks where you do not need full agentic muscle.

GPT-5.5 is the high-capability tier at $5/$30, notable because it is designed to reach comparable results using significantly fewer tokens than GPT-5.4, which can offset its higher per-token rate. Cached input is around $0.50/M.

As with Claude, cached input is the big lever: repeated context costs roughly 10% of the regular input rate, the most important single cost optimization in agentic workflows where the same project context is sent turn after turn.

Codex subscription plans

PlanPriceWhat you get
Plus$20/mo10-60 cloud tasks per 5-hr window, full GPT-5.3-Codex access
Pro (5x)$100/mo~5x Plus capacity; the sweet spot for active developers

In practice, Codex runs roughly $100-$200 per developer per month, with wide variance depending on which model you use, how many parallel instances you run, how much you automate, and whether you lean on fast mode.

Codex vs Claude Code: how to think about it

If your real question is “which coding agent,” here is the honest framing.

Both are cloud-capable autonomous agents that run multi-step tasks, work in sandboxes, and bill by tokens or subscription. The differences that matter day to day:

Claude Code leans on Anthropic’s lineup, where Opus 4.8 and now Fable 5 post the highest published SWE-bench scores, and the 1M-token context is a flat rate. If you work in large codebases where long context without a surcharge matters, that is a real advantage.

Codex leans on the GPT-5 family, with GPT-5.3-Codex as an efficient default and GPT-5.5 as a token-efficient high-end option. Its credit system and parallel cloud tasks are well suited to running many automated jobs at once.

For most teams the deciding factors are not the spec sheet. They are which model your developers trust on your actual code, which context limits fit your repo, and which billing model matches your usage pattern. Run both on a real task from your own codebase before committing. Benchmarks are a starting point, not a verdict.

Output price per 1M tokens across flagship AI models

The full 2026 pricing comparison: Claude vs GPT vs Gemini vs Grok

Claude and Codex are the focus above, but you are almost certainly weighing them against Google and xAI. Here is the whole field on one chart, by API token rate.

ProviderModelInput ($/M)Output ($/M)ContextNotes
AnthropicClaude Fable 5$10$501MMythos-class; guarded-domain fallback to Opus 4.8
AnthropicClaude Opus 4.8$5$251MFlagship; Fast Mode $10/$50
AnthropicClaude Sonnet 4.6$3$151MBalanced daily driver
AnthropicClaude Haiku 4.5$1$5-Budget / high-volume
OpenAIGPT-5.5$5$30largeToken-efficient flagship; ~$0.50 cached input
OpenAIGPT-5.2$1.75$14largeMainline general model
OpenAIGPT-5.3-Codex$1.75$14largeDefault model inside Codex
GoogleGemini 3.1 Pro$2$12200k$4/$18 above 200k tokens
GoogleGemini 3 Flash$0.50$3largeBudget price-performance pick
xAIGrok 4.3$1.25$2.501MCheap flagship
xAIGrok 4.1 Fast$0.20$0.502MCheapest rate, largest context

A dash means the figure was not confirmed at publication; check the provider’s docs. The short read: Grok leads on raw cost-efficiency, Gemini Flash is the budget price-performance pick, OpenAI sits in the competitive middle, and Claude runs a clean tier from Haiku at the bottom to Fable 5 at the top.

On the consumer side the market has standardized: a roughly $20/month tier gets you flagship access from every provider, the premium tier is $200/month (ChatGPT Pro, Claude Max 20x), and Google AI Ultra tops out around $250/month for the full tool and video suite.

Benchmarks: how the models actually compare

Price only matters next to capability. SWE-bench Verified (real-world coding) is the most-watched number; here is roughly where the field sits in mid-2026.

ModelSWE-bench VerifiedOther notes
Claude Fable 5~95.0% (unguarded)SWE-bench Pro ~80.0%; GDPval ~1932 Elo
GPT-5.5~88.7%Reported #1 on the public leaderboard
Claude Opus 4.8~88.6%Current flagship
Claude Opus 4.7~87.6%
GPT-5.3-Codex~85.0%Codex default
Gemini 3.1 Pro~80.6%GPQA ~94.3% (science)
Grok 4~72-75% self-reported~58.6% in independent testing

Two honest caveats. OpenAI has flagged training-data contamination concerns across all frontier models on SWE-bench Verified, and SWE-bench Pro (multi-language, standardized scaffold) is emerging as the more reliable successor. And scaffolding choices move scores a lot, which is why Grok’s self-reported and independently tested numbers differ so much. Treat these as a guide, not gospel, and benchmark on your own work.

Capability and feature matrix

FamilyContextVision inputPrompt cachingNotable
Claude 4.x + Fable 5up to 1MYes~90% off cachedExtended thinking, memory, context editing, batch 50% off
OpenAI GPT-5 / CodexlargeYes~10% of input rateToken-credit billing, parallel cloud agents
Google Gemini 3.x200k (Pro)YesAvailableDeep Google Workspace and tool integration
xAI Grok 4.xup to 2MYesAvailableLargest context windows, lowest prices

Feature availability and discount levels shift often, so verify the specifics (especially caching and batch terms for Gemini and Grok) against each provider’s current docs.

How to choose, and how to not overpay

A few rules that hold across both stacks:

Match the model to the task. Do not run a flagship on work a mid-tier model handles fine. Haiku, Sonnet, and GPT-5.4-Mini exist precisely so you do not pay Opus or GPT-5.5 rates for routine jobs.

Cache aggressively. On both platforms, cached input is the single biggest saving (about 90% off on Claude, roughly 10% of input rate on Codex). Any workflow that resends the same context should be caching it.

Use batch for anything asynchronous. Claude’s Batch API is 50% off for non-interactive workloads.

Pick subscription over API once you are a heavy daily user. If you are coding full time, Claude Max or Codex Pro almost always beats per-token billing. If your usage is spiky or light, the API or a lower tier wins.

Watch output, not just input. Output is where the cost concentrates. Tighter prompts and structured outputs cut bills faster than switching models.

Which model for which job

Most people overpay by defaulting to the flagship for everything, or underperform by using a cheap model on work that needs reasoning. Here is the task-to-model mapping that keeps both cost and quality in the right place.

Type of workBest fitWhy
Hard, multi-step coding and agentsFable 5, or GPT-5.3-Codex in CodexTop SWE-bench scores; built for long agentic tasks
Everyday coding, refactors, reviewsSonnet 4.6 or GPT-5.3-CodexStrong enough for most code at a much lower rate
Complex reasoning, analysis, strategyOpus 4.8Flagship reasoning without Fable 5 pricing
Long-document work (contracts, reports)Opus 4.8 or Sonnet 4.61M context at flat rate handles big inputs
High-volume writing and contentSonnet 4.6Best quality-per-dollar for routine generation
Classification, extraction, tagging, routingHaiku 4.5 or GPT-5.4-MiniFast and cheap; reasoning is not the bottleneck
Vision tasks (screenshots, diagrams, docs)Fable 5, Opus 4.8, Sonnet 4.6Vision input across the modern lineup
Scientific or frontier researchFable 5Highest capability; note safeguard fallbacks
Customer-facing chatbots at scaleHaiku 4.5, step up to Sonnet 4.6Latency and cost matter more than peak IQ
Regulated / YMYL topics (legal, medical, finance)Opus 4.8Reliable, and avoids Fable 5’s guarded-domain fallbacks

The general pattern: start one tier lower than your instinct says, test on your real task, and only move up if quality actually requires it. The cheap tiers exist because most work does not need a flagship.

Which model for which kind of professional

The same logic, mapped to who you are and what your day actually looks like.

Role / fieldDay-to-day driverReach for when…
Software engineerSonnet 4.6 / GPT-5.3-CodexGnarly bug or large refactor: Fable 5 or GPT-5.5
Founder / solopreneurPro plan with Sonnet 4.6Big strategic or financial call: Opus 4.8
Content marketer / SEOSonnet 4.6Cornerstone research piece: Opus 4.8
Data analyst / scientistOpus 4.8Bulk parsing and labeling: Haiku 4.5
Researcher / academicOpus 4.8, Fable 5 for the hardest problemsNote guarded-domain fallbacks on Fable 5
Customer support / opsHaiku 4.5Nuanced escalations and summaries: Sonnet 4.6
Lawyer / compliance / financeOpus 4.8Long contracts and filings: 1M context on Opus
Healthcare / life sciencesOpus 4.8Avoid Fable 5 for biology/chemistry (it falls back anyway)
Designer / productSonnet 4.6 with visionReviewing dense mockups or flows: Opus 4.8
Student / learnerFree or Pro with Sonnet 4.6Heavy project weeks: Max 5x
Enterprise / regulated orgOpus 4.8 on Enterprise (500k context, audit logs)Governance and compliance needs

Two cross-cutting notes. If your field touches cybersecurity, biology, chemistry, or distillation, remember that Fable 5 quietly falls back to Opus 4.8 in those domains, so you may as well plan around Opus there. And if you are a heavy daily user in any of these roles, a Max or Codex Pro subscription almost always beats per-token API billing.

What it actually costs: three worked examples

Token rates feel abstract until you run a real job through them. Here are three common workloads, with the assumptions stated so you can adjust. All figures are before caching and batch discounts, which can cut them dramatically.

1. A batch of 100 articles (about 2k input + 2k output tokens each, so 0.2M input and 0.2M output total):

ModelCostWith Batch (50% off)
Haiku 4.5$1.20$0.60
Sonnet 4.6$3.60$1.80
Gemini 3 Flash$0.70-
Opus 4.8$6.00$3.00

2. A RAG support chatbot (10,000 messages/month, ~1k input + 300 output tokens each, so 10M input and 3M output):

ModelMonthly cost (no caching)Why caching matters
Haiku 4.5~$25Most input is a repeated system prompt and context
Gemini 3 Flash~$14so caching can cut the input line by ~90%
Sonnet 4.6~$75turning these into a fraction of the listed cost

3. A full day of AI-assisted coding (heavy context reuse: ~5M input, mostly cached, plus ~500k output):

Without caching, 5M input + 0.5M output on Sonnet 4.6 is about $22.50; on Opus 4.8 about $37.50. With prompt caching on the repeated codebase context, the input line drops by roughly 90%, which is exactly why Anthropic’s own data puts the average Claude Code user near $6/day rather than $30+. The lesson repeats: caching, not model choice, is usually the biggest lever on a real bill.

Frequently asked questions

What is the cheapest AI model in 2026? Among the major frontier models, Grok 4.1 Fast is the cheapest at about $0.20/$0.50 per million tokens. On Anthropic’s lineup, Haiku 4.5 ($1/$5) is the budget tier, and Google’s Gemini 3 Flash ($0.50/$3) is competitive.

Which AI model is best for coding? The top SWE-bench Verified scores cluster around GPT-5.5 (~88.7%) and Claude Opus 4.8 (~88.6%), with Fable 5 higher (~95%) in unguarded domains. For most teams the real choice is Claude Code or GPT-5.3-Codex; test both on your own codebase.

How much does Claude cost per million tokens? As of June 2026: Fable 5 $10/$50, Opus 4.8 $5/$25, Sonnet 4.6 $3/$15, Haiku 4.5 $1/$5. Caching cuts cached input ~90%; batch is 50% off.

Is Claude Fable 5 worth it over Opus 4.8? For the hardest agentic, coding, and research work, yes. But it costs double and falls back to Opus 4.8 in guarded domains (cybersecurity, biology, chemistry), so for those topics or routine work, Opus 4.8 is better value.

Codex or Claude Code? Both are autonomous coding agents. Claude Code runs on Anthropic’s models with flat-rate 1M context; Codex runs on the GPT-5 family with token credits and strong parallel cloud tasks. Trial both on real code.

Subscription or pay-per-token? Full-time users almost always save with a subscription (Claude Max or Codex Pro, ~$100-$200/mo). Light or spiky usage is cheaper on the API or a lower tier.

Deep dives by vendor and use case

This page is the hub. For the full per-provider breakdowns and the two questions everyone asks, go deeper here:

The bottom line

As of June 2026, Claude’s lineup runs from Fable 5 at $10/$50 down to Haiku at $1/$5, with a flagship Opus 4.8 at $5/$25 and a 1M-token context that is flat-rate on the newer models. OpenAI’s Codex runs on the GPT-5 family, from an efficient GPT-5.3-Codex at $1.75/$14 to a token-efficient GPT-5.5 at $5/$30, billed through a token-credit system since April.

The right choice is rarely the cheapest sticker price or the highest benchmark. It is the model your team trusts on your real work, at a context limit that fits your codebase, billed in a way that matches how you actually use it. Pick on that, then use caching, batching, and right-sizing to keep the bill sane.

For the strategic read on what cheaper, more capable models mean for visibility and getting cited, see our piece on Claude Fable 5 and AI search and the broader future of SEO in the AI era. For optimizing specifically for Anthropic’s models, see our Claude SEO guide.

Sources and official pricing to verify against: Claude API pricing, Claude plans, Fable 5 and Mythos 5 docs, OpenAI Codex pricing, Gemini pricing, and Grok / xAI pricing.


Changelog, June 10, 2026: Expanded from Claude + Codex to a full cross-vendor comparison (added Google Gemini and xAI Grok), plus benchmark and feature matrices, worked cost examples, and an FAQ. This is a living reference and is refreshed as pricing changes.