Commercial guide - Last reviewed 2026-09-17
LLM API Pricing September 2026: OpenAI vs Claude vs Gemini
OpenAI, Claude, and Gemini API token prices verified September 2026 — including the long-context and promotional rates the headline price sheet leaves out.
Direct answer for openai api pricing
The short answer
As of September 2026, frontier-tier prices cluster at $4-$5 per million input tokens and $20-$25 per million output tokens (OpenAI GPT-5.6 Sol at its promotional $4/$20, Claude Opus 5 at $5/$25), premium reasoning tiers run higher (Claude Fable 5 at $10/$50), mid tiers at $2 in / $10-$12 out (GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.1 Pro), and high-volume tiers at $0.20-$1 in / $1.20-$5 out (GPT-5.6 Luna, Gemini 3.7 Flash, Claude Haiku 4.5). Two rates the headline sheet hides move bills more than the tier you pick: long-context requests reprice upward — Sol goes from $4/$20 to $8/$30 past its threshold, and Gemini 3.1 Pro doubles its input rate above 200K context — and several current rates are promotional with published expiry dates. Beyond that, per-token price explains a minority of real spend: retries, agent loops, and retrieval multiply raw usage, so routing and workflow design matter more than provider choice.
Use frontier tiers (Sol, Opus 5, Fable 5, Gemini 3.1 Pro) only where reasoning quality is the product — and route everything else down-tier.
Mid tiers (Terra, Sonnet 5) handle most production workloads; this is the default tier to price against.
High-volume tiers (Luna at $0.20/$1.20, Gemini 3.7 Flash at $0.75/$3.75, Haiku 4.5 at $1/$5) plus batch and caching discounts can cut classification, extraction, and summarization costs 10-50x versus frontier pricing.
Check the expiry before you budget: Gemini Flash rates double on Jan 1 2027, and GPT-5.6 Sol's $4/$20 is promotional through at least Nov 21 2026. One went the other way: Claude Sonnet 5's $2/$10 launched as introductory pricing and is now the standard price — Anthropic cancelled the planned Sept 1 increase to $3/$15.
Comparison table
| Factor | Frontier tier | Efficient tier |
|---|---|---|
| List price range (per M tokens) | $4-$10 input / $20-$50 output (Sol at its promo rate, Opus 5, Fable 5). | $0.20-$2 input / $1.20-$12 output (Luna, Gemini Flash, Haiku 4.5, Terra). |
| Long-context repricing | Sol $4/$20 to $8/$30 past its long-context threshold; Gemini 3.1 Pro goes from $2/$12 to $4/$18 above 200K. | Terra $2/$12 to $4/$18 and Luna $0.20/$1.20 to $0.40/$1.80 — the cheap tier is not cheap once contexts get long. |
| Best fit | Agent planning steps, complex reasoning, code generation where quality failures are expensive. | Classification, extraction, summarization, RAG answer synthesis, high-volume chat. |
| Cost drivers to watch | Long contexts and chain-of-thought multiply output tokens — the expensive side of the ratio. | Volume itself: cheap unit prices invite 50-500x agentic usage growth (the Jevons trap). |
| Discount levers | Prompt caching (cache hits ~10% of input price on Claude), batch APIs (~50% off). | Same levers apply — batch + caching on an efficient tier is the cheapest managed option. |
Worked example
Published list prices — September 2026 (per million tokens)
- Standard-context prices, re-verified Sept 17, 2026 against each provider's own pricing page
- Batch (~50% off) and cache discounts excluded
- Long-context tiers cost more — see the comparison table above
- Claude Sonnet 5's $2/$10 is now standard — the planned Sept 1 increase to $3/$15 was cancelled
- Both Gemini Flash intro rates double on Jan 1, 2027
- GPT-5.6 Sol has a promo at $4/$20 through at least Nov 21, 2026
| Model | Input $/M | Output $/M |
|---|---|---|
| Claude Fable 5 (Anthropic) | $10.00 | $50.00 |
| GPT-5.6 Sol (OpenAI, list after promo — last seen Aug 23) | $5.00 | $30.00 |
| Claude Opus 5 (Anthropic) | $5.00 | $25.00 |
| GPT-5.6 Sol (OpenAI, promo to Nov 21) | $4.00 | $20.00 |
| Gemini 3.1 Pro (Google, to 200K context) | $2.00 | $12.00 |
| GPT-5.6 Terra (OpenAI) | $2.00 | $12.00 |
| Claude Sonnet 5 (Anthropic) | $2.00 | $10.00 |
| Claude Haiku 4.5 (Anthropic) | $1.00 | $5.00 |
| Gemini 3.8 Flash (Google, intro) | $0.75 | $3.75 |
| Gemini 3.7 Flash (Google, intro) | $0.75 | $3.75 |
| Gemini 3.6 Flash (Google, intro) | $0.75 | $3.75 |
| GPT-5.6 Luna (OpenAI) | $0.20 | $1.20 |
| Self-hosted 70B INT8 (NavyaAI benchmark, blended) | ~$0.47 | + ops floor |
Unit prices keep falling — Terra and Luna were cut in July 2026, and Gemini 3.7 Flash landed in August at half the prior Flash rate — yet bills keep rising, because workflows multiply usage faster than prices drop. Three things distort the headline number: long-context repricing, promo rates with expiry dates, and the workflow multiplier. Route by task tier, cache aggressively, batch what can wait, and measure the multiplier before blaming the price sheet.
Frequently asked questions
How much does the OpenAI API cost per million tokens?
As of September 2026: GPT-5.6 Sol at a promotional $4 input / $20 output per million tokens through at least Nov 21, 2026, GPT-5.6 Terra at $2/$12, and GPT-5.6 Luna at $0.20/$1.20. Long-context requests are billed higher — Sol at $8/$30, Terra at $4/$18, Luna at $0.40/$1.80. Batch processing roughly halves the standard rates.
How much does the Claude API cost per million tokens?
As of September 2026: Claude Fable 5 at $10 input / $50 output per million tokens, Opus 5 at $5/$25, Sonnet 5 at $2/$10, and Haiku 4.5 at $1/$5. Sonnet 5's price launched as introductory through Aug 31, 2026; Anthropic has since made it the standard price and cancelled the planned increase to $3/$15. Cache hits cost about 10% of the input price; the Batch API cuts prices ~50%.
Which LLM API is cheapest in 2026?
On list price, OpenAI's GPT-5.6 Luna ($0.20/$1.20 per million tokens) is the cheapest mainstream tier, followed by Gemini 3.8, 3.7 and 3.6 Flash at their $0.75/$3.75 introductory rate, then Claude Haiku 4.5 ($1/$5). Two caveats: the Gemini Flash rates double on Jan 1, 2027, and every one of these tiers reprices upward on long-context requests. Cheapest-per-token also rarely means cheapest-per-task — routing the right tier per task and using batch and caching discounts matters more than picking one provider.
Why is my AI bill rising when token prices keep falling?
Because usage grows faster than prices fall. Agentic workflows multiply token consumption 50-500x per task through loops, tool calls, retries, and retrieval — and roughly 72% of production AI cost sits outside the model invoice entirely. That is the core finding of our AI Cost Report.
Do batch and caching discounts really change the economics?
Yes, materially. Prompt caching prices cache hits at ~10% of normal input cost — transformative for long shared system prompts and RAG contexts. Batch APIs cut both sides ~50% for anything that tolerates delayed responses. Combined, an efficient-tier model with caching and batch can run 10-50x cheaper than naive frontier-tier usage.
References & related
Apply this to your stack
Request a free AI inference audit before changing providers or buying GPUs.
Share your monthly spend, token volume, model stack, RAG or agent pattern, and latency target. NavyaAI will identify the first cost levers to inspect.
Request Free Audit