← Blog AI Cost Management

LLM API Pricing 2026: Cost per 1M Tokens Compared

By FavTray · · Updated

Written by the team behind FavTray, a Mac developer dashboard for the menu bar. We check every app we cover against its own site, docs or a hands-on test, and each post says which.

Short answer: in LLM API pricing for September 2026, the cheapest capable models are Gemini 2.5 Flash-Lite ($0.10/$0.40 per million input/output tokens), GPT-6 Luna ($0.10/$0.50) and DeepSeek Flash ($0.15–0.30/$0.60–1.20). The mid tier has converged on $2 input and $10–12 output (Claude Sonnet 5, GPT-6 Sol, Gemini 3.1 Pro Preview). Top models cost $10/$50. Batch jobs are 50% off at Anthropic, OpenAI and Google.

Every price below comes from the vendor’s own pricing page, checked on September 26, 2026: Anthropic, OpenAI, Google and DeepSeek. Prices change often, so check again before you commit a budget.

LLM API pricing comparison: current models per 1M tokens

Anthropic (Claude)

ModelInputOutputCached input (read)Context windowBest for
Claude Fable 5.1$10.00$50.00$0.251MHardest reasoning, long agentic runs
Claude Opus 5.5$4.00$20.00$0.201MAgentic coding, complex work
Claude Sonnet 5$2.00$10.00$0.201MDaily coding, analysis
Claude Haiku 4.5$1.00$5.00$0.10200KFast, lightweight tasks

No long-context surcharge: Claude 4.6 and later models bill the full 1M window at these rates.

OpenAI

ModelInputOutputCached inputContext windowBest for
GPT-6 Astra$10.00$50.00$1.001.05MHardest end-to-end work
GPT-6 Sol$2.00$10.00$0.201.05MCoding, agentic workflows
GPT-6 Luna$0.10$0.50$0.011.05MHigh-volume, focused tasks

Prompts over 272K input tokens cost 2× input and 1.5× output for the whole request. Older models are still sold, from GPT-5.6 Sol ($4/$20 on promotion until at least November 21, 2026) down to gpt-5-nano ($0.05/$0.40).

Google (Gemini)

ModelInputOutputCached inputContext windowFree tier
Gemini 3.1 Pro Preview$2.00 ($4 over 200K)$12.00 ($18 over 200K)$0.201MNo
Gemini 3.8 Flash$0.75*$3.75*$0.0751MYes
Gemini 3.5 Flash$1.50$9.00$0.151MYes
Gemini 3.5 Flash-Lite$0.30$2.50$0.031MYes
Gemini 3.1 Flash-Lite$0.25$1.50$0.0251MYes
Gemini 2.5 Pro$1.25 ($2.50 over 200K)$10.00 ($15 over 200K)$0.1251MYes
Gemini 2.5 Flash$0.30$2.50$0.031MYes
Gemini 2.5 Flash-Lite$0.10$0.40$0.011MYes

*Launch price through December 31, 2026; $1.50 input and $7.50 output from January 1, 2027. Gemini also charges an hourly storage fee for cached tokens.

DeepSeek

ModelInput (cache miss)OutputCached inputContext window
DeepSeek Flash (V4.1-Flash)$0.30 peak / $0.15 off-peak$1.20 / $0.60$0.006 / $0.0031M
DeepSeek V4 Pro$1.32 / $0.66$3.96 / $1.98$0.044 / $0.0221M

DeepSeek’s peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays; every other hour is billed at the half-price off-peak rate.

Not listed: Mistral and hosted open-source providers such as Groq and Together. We couldn’t read current per-token prices from their own pricing pages on the day we checked, and we’d rather leave them out than print numbers we can’t source.

To see what you actually spend, check each vendor’s console. If you use subscription tools rather than raw API calls, FavTray’s free AI Usage Tracker shows your Claude, Codex, Gemini CLI, Cursor and Copilot limits in the Mac menu bar.

What does each model cost per task?

The table assumes a typical call of 2,000 input and 500 output tokens, at standard list prices with no caching or batch discount. Tokenizers differ between vendors, so treat these as estimates.

ModelCost per callCost per 1,000 calls
Gemini 2.5 Flash-Lite$0.0004$0.40
GPT-6 Luna$0.00045$0.45
DeepSeek Flash (peak)$0.0012$1.20
Gemini 3.1 Flash-Lite$0.00125$1.25
Gemini 3.8 Flash (launch price)$0.0034$3.38
Claude Haiku 4.5$0.0045$4.50
DeepSeek V4 Pro (peak)$0.0046$4.62
Claude Sonnet 5$0.009$9.00
GPT-6 Sol$0.009$9.00
Gemini 3.1 Pro Preview$0.010$10.00
Claude Opus 5.5$0.018$18.00
Claude Fable 5.1 / GPT-6 Astra$0.045$45.00

The spread from cheapest to most expensive is more than 100×. The quality gap is real but much smaller than that for routine work (formatting, extraction, short summaries), and it matters most for multi-step reasoning, where a model that gets it right first time saves a second paid attempt.

Which provider is cheapest for high-volume use?

Monthly cost at different volumes, assuming a 50/50 split between input and output tokens:

Monthly volumeClaude Sonnet 5GPT-6 SolGemini 3.1 ProClaude Haiku 4.5GPT-6 LunaGemini 2.5 Flash-Lite
1M tokens$6$6$7$3$0.30$0.25
10M tokens$60$60$70$30$3$2.50
100M tokens$600$600$700$300$30$25
1B tokens$6,000$6,000$7,000$3,000$300$250

At a billion tokens a month, the difference between Claude Sonnet 5 and Gemini 2.5 Flash-Lite is $5,750. That’s why large applications route simple queries to small models and keep the expensive ones for hard tasks. Batch processing halves every column at Anthropic, OpenAI and Google if you can wait for results.

For an individual developer spending on 5–20 million tokens a month, the gap between the mid-tier models is a few dollars; output quality and tooling matter more than the rate card.

How do batch, caching and long-context rules compare?

AnthropicOpenAIGoogleDeepSeek
Batch discount50%50% (plus Flex at 50%)50%—
Time-of-day discount———50% off-peak
Cached input10% of input (5% on Opus 5.5, 2.5% on Fable 5.1)10% of input10% of input, plus hourly storageAbout 2–3% of input
Long-prompt surchargeNone up to 1M (Claude 4.6+)Over 272K: 2× input, 1.5× outputPro models over 200K: about 2× input, 1.5× outputNone listed
Free tierSmall credit for new users—Yes, on most Flash models and 2.5 Pro—

How should you choose between quality tiers?

Use mid- and top-tier models (Sonnet 5, Opus 5.5, GPT-6 Sol, Gemini 3.1 Pro) where accuracy saves iteration: debugging, architecture, complex code generation. Use small models (Haiku 4.5, GPT-6 Luna, Gemini Flash-Lite, DeepSeek Flash) for high-volume work where “good enough” is fine: formatting, boilerplate, extraction, simple refactors.

Task characteristicRecommended tierWhy
Requires multi-step reasoningMid or topSmall models make errors that cascade
Single-turn, well-defined outputSmallExtra reasoning is wasted
Code that runs in productionMid or topBugs cost more than the savings
Internal documentationSmallMinor quality differences don’t matter
Security-sensitive code reviewMid or topMissing a vulnerability is expensive
Data transformation, formattingSmallPattern-following, not reasoning
Very long prompts (300K+ tokens)Claude, or DeepSeekNo long-context surcharge

Routing by tier is the single biggest lever on an API bill; our guide to reducing AI API costs covers it alongside caching and prompt trimming. For head-to-heads, see is Gemini cheaper than Claude?, Claude vs OpenAI API pricing and, if you’re weighing a subscription against the API, whether Claude Max is worth it.

Prices aren’t simply falling. The top tier got much cheaper, the middle got a little cheaper, and some newer small and mid models cost more than the ones they replaced. These are each vendor’s own list prices, old and new, as shown on the same pricing pages on September 26, 2026:

Vendor and tierEarlier model (input / output)Current model (input / output)Change
Anthropic topClaude Opus 4 and 4.1 (2025): $15 / $75Claude Opus 5.5: $4 / $20−73%
Anthropic midClaude Sonnet 4 to 4.6: $3 / $15Claude Sonnet 5: $2 / $10−33%
Anthropic smallClaude Haiku 3.5: $0.80 / $4Claude Haiku 4.5: $1 / $5+25%
OpenAI midGPT-4o (2024): $2.50 / $10GPT-6 Sol: $2 / $10−20% input, same output
OpenAI smallGPT-4o mini (2024): $0.15 / $0.60GPT-6 Luna: $0.10 / $0.50−33% / −17%
Google ProGemini 2.5 Pro (2025): $1.25 / $10Gemini 3.1 Pro Preview: $2 / $12+60% / +20%
Google FlashGemini 2.5 Flash: $0.30 / $2.50Gemini 3.8 Flash: $0.75 / $3.75 (launch), $1.50 / $7.50 from 2027+150% / +50%

Two things follow. First, “wait for prices to drop” isn’t a strategy: the newest model in a tier can cost more than the last one, and older models often stay on sale at their old price. Second, the cheapest option for a task is often last year’s model, so re-check your routing whenever a vendor updates its lineup. Watching your total AI spending across tools is what turns a price change into an actual saving.

Frequently Asked Questions

What is the cheapest LLM API in 2026?

Among current models from the major vendors, Gemini 2.5 Flash-Lite ($0.10 input, $0.40 output per million tokens), GPT-6 Luna ($0.10/$0.50) and DeepSeek Flash ($0.15–0.30 input, $0.60–1.20 output, cheaper off-peak) are the cheapest capable options. Prices from each vendor's pricing page, September 26, 2026.

How much does GPT-6 cost per million tokens?

GPT-6 Astra costs $10 input and $50 output per million tokens, GPT-6 Sol $2 and $10, and GPT-6 Luna $0.10 and $0.50. Prompts over 272K input tokens are billed at 2× input and 1.5× output for the whole request. Batch and Flex processing take 50% off.

Is Claude or GPT cheaper for coding?

At list price they're the same: Claude Sonnet 5 and GPT-6 Sol both cost $2 input and $10 output per million tokens, with the same $0.20 cache-read price. Claude is cheaper for prompts over 272K tokens because it doesn't add a long-context surcharge. Which one costs less on your work depends on how many tokens and retries each needs.

Why are output tokens more expensive than input tokens?

Output tokens typically cost 4–8× more than input tokens because the model generates them one at a time, each needing its own pass through the network, while input tokens are processed in parallel. Claude Sonnet 5 and GPT-6 Sol charge 5× more for output; Gemini 3.1 Pro Preview charges 6×.

How do Gemini API prices compare to OpenAI and Claude?

Gemini's cheapest models are among the cheapest anywhere (2.5 Flash-Lite at $0.10/$0.40), and Google has the most generous free tier. Its top model, Gemini 3.1 Pro Preview, costs $2/$12 for prompts up to 200K tokens, slightly more on output than Claude Sonnet 5 or GPT-6 Sol at $2/$10, and $4/$18 above 200K.

Something new: My Dock Buddies