Commercial guide - Last reviewed 2026-08-07
LLM API Pricing Comparison 2026: OpenAI vs Claude vs Gemini
Current OpenAI, Anthropic Claude, and Google Gemini token prices per million tokens, what the tiers mean, and why the price sheet rarely predicts your bill.
Direct answer for LLM API pricing comparison
The short answer
As of August 2026, frontier-tier list prices cluster at $5 per million input tokens and $25-$30 per million output tokens (OpenAI GPT-5.6 Sol, Claude Opus 5), premium reasoning tiers run higher (Claude Fable 5 at $10/$50), mid tiers at roughly $1.50-$2 in / $7.50-$12 out (GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.6 Flash), and high-volume tiers at $0.20-$1 in / $1.20-$5 out (GPT-5.6 Luna, Claude Haiku 4.5). But per-token price explains a minority of real spend: retries, agent loops, and retrieval typically multiply raw usage, so model routing and workflow design move bills more than provider choice.
Use frontier tiers (Sol, Opus 5, Fable 5, Gemini Pro) only where reasoning quality is the product — and route everything else down-tier.
Mid tiers (Terra, Sonnet 5, Gemini Flash) now handle most production workloads; this is the default tier to price against.
High-volume tiers (Luna, Haiku 4.5) plus batch and caching discounts can cut classification, extraction, and summarization costs 10-50x versus frontier pricing.
Comparison table
| Factor | Frontier tier | Efficient tier |
|---|---|---|
| List price range (per M tokens) | $5-$10 input / $25-$50 output (Sol, Opus 5, Fable 5). | $0.20-$2 input / $1.20-$12 output (Luna, Haiku 4.5, Terra, Gemini Flash). |
| Best fit | Agent planning steps, complex reasoning, code generation where quality failures are expensive. | Classification, extraction, summarization, RAG answer synthesis, high-volume chat. |
| Cost drivers to watch | Long contexts and chain-of-thought multiply output tokens — the expensive side of the ratio. | Volume itself: cheap unit prices invite 50-500x agentic usage growth (the Jevons trap). |
| Discount levers | Prompt caching (cache hits ~10% of input price on Claude), batch APIs (~50% off). | Same levers apply — batch + caching on an efficient tier is the cheapest managed option. |
Worked example
Published list prices — August 2026 (per million tokens)
- List prices as published on Aug 7, 2026
- Batch (~50% off) and cache discounts excluded
- Claude Sonnet 5 promo ($2/$10) ends Aug 31, 2026 → $3/$15
- GPT-5.6 Terra & Luna cut July 30, 2026
| Model | Input $/M | Output $/M |
|---|---|---|
| Claude Fable 5 (Anthropic) | $10.00 | $50.00 |
| GPT-5.6 Sol (OpenAI) | $5.00 | $30.00 |
| Claude Opus 5 (Anthropic) | $5.00 | $25.00 |
| Gemini 3.1 Pro (Google, preview) | $2.00 | $12.00 |
| GPT-5.6 Terra (OpenAI) | $2.00 | $12.00 |
| Claude Sonnet 5 (promo to Aug 31) | $2.00 | $10.00 |
| Gemini 3.6 Flash (Google) | $1.50 | $7.50 |
| Claude Haiku 4.5 (Anthropic) | $1.00 | $5.00 |
| GPT-5.6 Luna (OpenAI) | $0.20 | $1.20 |
| Self-hosted 70B INT8 (NavyaAI benchmark, blended) | ~$0.47 | + ops floor |
Unit prices keep falling — Terra and Luna were cut again in July 2026 — yet bills keep rising, because workflows multiply usage faster than prices drop. Route by task tier, cache aggressively, batch what can wait, and measure the workflow multiplier before blaming the price sheet.
Frequently asked questions
How much does the OpenAI API cost per million tokens?
As of August 2026: GPT-5.6 Sol at $5 input / $30 output per million tokens, GPT-5.6 Terra at $2/$12, and GPT-5.6 Luna at $0.20/$1.20 after the July 30 price cut. Batch processing roughly halves these rates.
How much does the Claude API cost per million tokens?
As of August 2026: Claude Fable 5 at $10 input / $50 output per million tokens, Opus 5 at $5/$25, Sonnet 5 at $2/$10 (promotional until Aug 31, 2026, then $3/$15), and Haiku 4.5 at $1/$5. Cache hits cost about 10% of the input price; the Batch API cuts prices ~50%.
Which LLM API is cheapest in 2026?
On list price, OpenAI's GPT-5.6 Luna ($0.20/$1.20 per million tokens) is the cheapest mainstream tier, followed by Claude Haiku 4.5 ($1/$5) and Gemini Flash tiers. But cheapest-per-token rarely means cheapest-per-task: routing the right tier per task and using batch and caching discounts matters more than picking one provider.
Why is my AI bill rising when token prices keep falling?
Because usage grows faster than prices fall. Agentic workflows multiply token consumption 50-500x per task through loops, tool calls, retries, and retrieval — and roughly 72% of production AI cost sits outside the model invoice entirely. That is the core finding of our AI Cost Report.
Do batch and caching discounts really change the economics?
Yes, materially. Prompt caching prices cache hits at ~10% of normal input cost — transformative for long shared system prompts and RAG contexts. Batch APIs cut both sides ~50% for anything that tolerates delayed responses. Combined, an efficient-tier model with caching and batch can run 10-50x cheaper than naive frontier-tier usage.
References & related
Apply this to your stack
Request a free AI inference audit before changing providers or buying GPUs.
Share your monthly spend, token volume, model stack, RAG or agent pattern, and latency target. NavyaAI will identify the first cost levers to inspect.
Request Free Audit