Commercial guide - Last reviewed 2026-08-23
LLM API Pricing August 2026: OpenAI vs Claude vs Gemini
OpenAI, Claude, and Gemini API token prices verified August 2026 — including the long-context and promotional rates the headline price sheet leaves out.
Direct answer for openai api pricing
The short answer
As of August 2026, frontier-tier list prices cluster at $5 per million input tokens and $25-$30 per million output tokens (OpenAI GPT-5.6 Sol, Claude Opus 5), premium reasoning tiers run higher (Claude Fable 5 at $10/$50), mid tiers at roughly $2-$3 in / $10-$15 out (GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.1 Pro), and high-volume tiers at $0.20-$1 in / $1.20-$5 out (GPT-5.6 Luna, Gemini 3.7 Flash, Claude Haiku 4.5). Two rates the headline sheet hides move bills more than the tier you pick: long-context requests reprice upward — Sol goes from $5/$30 to $10/$45 past its threshold, and Gemini 3.1 Pro doubles its input rate above 200K context — and several current rates are promotional with published expiry dates. Beyond that, per-token price explains a minority of real spend: retries, agent loops, and retrieval multiply raw usage, so routing and workflow design matter more than provider choice.
Use frontier tiers (Sol, Opus 5, Fable 5, Gemini 3.1 Pro) only where reasoning quality is the product — and route everything else down-tier.
Mid tiers (Terra, Sonnet 5) handle most production workloads; this is the default tier to price against.
High-volume tiers (Luna at $0.20/$1.20, Gemini 3.7 Flash at $0.75/$3.75, Haiku 4.5 at $1/$5) plus batch and caching discounts can cut classification, extraction, and summarization costs 10-50x versus frontier pricing.
Check the expiry before you budget: Claude Sonnet 5's $2/$10 intro ends Aug 31 2026, and both Gemini Flash rates double on Jan 1 2027.
Comparison table
| Factor | Frontier tier | Efficient tier |
|---|---|---|
| List price range (per M tokens) | $5-$10 input / $25-$50 output (Sol, Opus 5, Fable 5). | $0.20-$2 input / $1.20-$12 output (Luna, Gemini Flash, Haiku 4.5, Terra). |
| Long-context repricing | Sol $5/$30 to $10/$45 past its long-context threshold; Gemini 3.1 Pro input doubles to $4 above 200K. | Terra $2/$12 to $4/$18 and Luna $0.20/$1.20 to $0.40/$1.80 — the cheap tier is not cheap once contexts get long. |
| Best fit | Agent planning steps, complex reasoning, code generation where quality failures are expensive. | Classification, extraction, summarization, RAG answer synthesis, high-volume chat. |
| Cost drivers to watch | Long contexts and chain-of-thought multiply output tokens — the expensive side of the ratio. | Volume itself: cheap unit prices invite 50-500x agentic usage growth (the Jevons trap). |
| Discount levers | Prompt caching (cache hits ~10% of input price on Claude), batch APIs (~50% off). | Same levers apply — batch + caching on an efficient tier is the cheapest managed option. |
Worked example
Published list prices — August 2026 (per million tokens)
- Standard-context list prices, re-verified Aug 23, 2026
- Batch (~50% off) and cache discounts excluded
- Long-context tiers cost more — see the comparison table above
- Claude Sonnet 5 intro ($2/$10) ends Aug 31, 2026 → $3/$15
- Both Gemini Flash intro rates double on Jan 1, 2027
- GPT-5.6 Sol has a promo at $4/$20 through at least Nov 21, 2026
| Model | Input $/M | Output $/M |
|---|---|---|
| Claude Fable 5 (Anthropic) | $10.00 | $50.00 |
| GPT-5.6 Sol (OpenAI) | $5.00 | $30.00 |
| Claude Opus 5 (Anthropic) | $5.00 | $25.00 |
| GPT-5.6 Sol (OpenAI, promo to Nov 21) | $4.00 | $20.00 |
| Gemini 3.1 Pro (Google, to 200K context) | $2.00 | $12.00 |
| GPT-5.6 Terra (OpenAI) | $2.00 | $12.00 |
| Claude Sonnet 5 (intro to Aug 31) | $2.00 | $10.00 |
| Claude Haiku 4.5 (Anthropic) | $1.00 | $5.00 |
| Gemini 3.7 Flash (Google, intro) | $0.75 | $3.75 |
| Gemini 3.6 Flash (Google, intro) | $0.75 | $3.75 |
| GPT-5.6 Luna (OpenAI) | $0.20 | $1.20 |
| Self-hosted 70B INT8 (NavyaAI benchmark, blended) | ~$0.47 | + ops floor |
Unit prices keep falling — Terra and Luna were cut in July 2026, and Gemini 3.7 Flash landed in August at half the prior Flash rate — yet bills keep rising, because workflows multiply usage faster than prices drop. Three things distort the headline number: long-context repricing, promo rates with expiry dates, and the workflow multiplier. Route by task tier, cache aggressively, batch what can wait, and measure the multiplier before blaming the price sheet.
Frequently asked questions
How much does the OpenAI API cost per million tokens?
As of August 2026: GPT-5.6 Sol at $5 input / $30 output per million tokens, GPT-5.6 Terra at $2/$12, and GPT-5.6 Luna at $0.20/$1.20 after the July 30 price cut. Sol also has a promotional rate of $4/$20 running through at least Nov 21, 2026. Long-context requests are billed higher — Sol at $10/$45, Terra at $4/$18, Luna at $0.40/$1.80. Batch processing roughly halves the standard rates.
How much does the Claude API cost per million tokens?
As of August 2026: Claude Fable 5 at $10 input / $50 output per million tokens, Opus 5 at $5/$25, Sonnet 5 at $2/$10 (promotional until Aug 31, 2026, then $3/$15), and Haiku 4.5 at $1/$5. Cache hits cost about 10% of the input price; the Batch API cuts prices ~50%.
Which LLM API is cheapest in 2026?
On list price, OpenAI's GPT-5.6 Luna ($0.20/$1.20 per million tokens) is the cheapest mainstream tier, followed by Gemini 3.7 and 3.6 Flash at their $0.75/$3.75 introductory rate, then Claude Haiku 4.5 ($1/$5). Two caveats: the Gemini Flash rates double on Jan 1, 2027, and every one of these tiers reprices upward on long-context requests. Cheapest-per-token also rarely means cheapest-per-task — routing the right tier per task and using batch and caching discounts matters more than picking one provider.
Why is my AI bill rising when token prices keep falling?
Because usage grows faster than prices fall. Agentic workflows multiply token consumption 50-500x per task through loops, tool calls, retries, and retrieval — and roughly 72% of production AI cost sits outside the model invoice entirely. That is the core finding of our AI Cost Report.
Do batch and caching discounts really change the economics?
Yes, materially. Prompt caching prices cache hits at ~10% of normal input cost — transformative for long shared system prompts and RAG contexts. Batch APIs cut both sides ~50% for anything that tolerates delayed responses. Combined, an efficient-tier model with caching and batch can run 10-50x cheaper than naive frontier-tier usage.
References & related
Apply this to your stack
Request a free AI inference audit before changing providers or buying GPUs.
Share your monthly spend, token volume, model stack, RAG or agent pattern, and latency target. NavyaAI will identify the first cost levers to inspect.
Request Free Audit