On this page
The best-known charts of falling LLM prices are essays. a16z's "LLMflation" piece from November 2024 and Epoch AI's analysis of inference price trends are both careful and both widely linked, and both describe the market on the day they were written. If you want the number for this month, you are on your own.
This page is that measurement kept running. Every day a script reads three public sources, works out the cheapest price for three fixed levels of model capability, and rebuilds the tables below. The history goes back to October 2023, rebuilt month by month from the git history of the open-source LiteLLM price file.
I run Gravity, where every agent run is paid for in tokens, so I check these prices more than most people. Writing the dates down was the only way to answer the question I kept getting asked: how fast is this actually falling?
What the data says on 4 October 2026
- GPT-4-class tokens cost 99.8% less than in October 2023. The cheapest model rated at or above the original GPT-4 now costs $0.0575 per million tokens blended (gpt-oss-20b on DeepInfra: $0.03 input, $0.14 output), against $37.50 for gpt-4-0314 on 1 October 2023. That is about 652 times cheaper in three years.
- GPT-4o class prices fell 36% in 12 months, from $0.11 to $0.0703 per million tokens blended between 1 October 2025 and 4 October 2026. The cheapest model in the tier today is gpt-oss-120b on DeepInfra.
- Gemini 2.5 Pro class prices fell 94% in 12 months, from $2.41 to $0.152 per million tokens blended between 1 October 2025 and 4 October 2026. The cheapest model in the tier today is gemma-4-31b on DeepInfra.
- The median new model from 11 major labs lists at $0.75 input and $3.75 output per million tokens. That is the median across the 55 newest paid models of OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral, Alibaba, Moonshot, Z.ai, MiniMax and Cohere on 4 October 2026: what the labs are shipping now, not the cheapest option.
Updated . Sources for this update: LiteLLM price file commit a66adb4, OpenRouter models API (466 models listed), LMArena leaderboard of 2 October 2026.
Cite this page: "LLM API pricing history", Gravity, updated , https://gravity.fast/data/llm-api-price-history/. The data is licensed CC BY 4.0: copy it, chart it, publish it, and link back here. Downloads: data.csv (tier floors by month) and current-prices.csv (today's list prices by lab).
Price per million tokens by capability tier
Each line below holds capability fixed and asks what the cheapest way to buy it costs. A model counts toward a tier when its LMArena rating is at or above the tier's floor model: the original GPT-4 from March 2023, GPT-4o from May 2024, or Gemini 2.5 Pro from 2025. The price is the cheapest blended rate among those models, counting the labs' own APIs and six large inference hosts.
Month by month
The first snapshot of every month, newest first, with the model that set the floor and who served it.
| Snapshot | GPT-4 class | GPT-4o class | Gemini 2.5 Pro class |
|---|---|---|---|
| $0.0575 gpt-oss-20b, DeepInfra | $0.0703 gpt-oss-120b, DeepInfra | $0.152 gemma-4-31b, DeepInfra | |
| $0.0575 gpt-oss-20b, DeepInfra | $0.0703 gpt-oss-120b, DeepInfra | $0.152 gemma-4-31b, DeepInfra | |
| $0.0575 gpt-oss-20b, DeepInfra | $0.0703 gpt-oss-120b, DeepInfra | $0.152 gemma-4-31b, DeepInfra | |
| $0.05 gemma-3-4b-it, DeepInfra | $0.107 gemma-3-27b-it, DeepInfra | $0.45 gpt-5.6-luna-xhigh, OpenAI | |
| $0.05 gemma-3-4b-it, DeepInfra | $0.107 gemma-3-27b-it, DeepInfra | $0.544 deepseek-v4-pro, DeepSeek | |
| $0.05 gemma-3-4b-it, DeepInfra | $0.107 gemma-3-27b-it, DeepInfra | $1.07 kimi-k2.5-thinking, Together AI | |
| $0.05 gemma-3-4b-it, DeepInfra | $0.107 gemma-3-27b-it, DeepInfra | $1.07 kimi-k2.5-thinking, Together AI | |
| $0.05 gemma-3-4b-it, DeepInfra | $0.107 gemma-3-27b-it, DeepInfra | $1.07 kimi-k2.5-thinking, Together AI | |
| $0.05 gemma-3-4b-it, DeepInfra | $0.107 gemma-3-27b-it, DeepInfra | $1.07 kimi-k2.5-thinking, Together AI | |
| $0.05 gemma-3-4b-it, DeepInfra | $0.107 gemma-3-27b-it, DeepInfra | $1.13 gemini-3-flash, Google | |
| $0.05 gemma-3-4b-it, DeepInfra | $0.107 gemma-3-27b-it, DeepInfra | $1.13 gemini-3-flash, Google | |
| $0.05 gemma-3-4b-it, DeepInfra | $0.107 gemma-3-27b-it, DeepInfra | $3.44 gpt-5.1-high, OpenAI | |
| $0.05 gemma-3-4b-it, DeepInfra | $0.107 gemma-3-27b-it, DeepInfra | $3.44 gemini-2.5-pro, DeepInfra | |
| $0.05 gemma-3-4b-it, DeepInfra | $0.11 gemma-3-27b-it, DeepInfra | $2.41 gemini-2.5-pro, DeepInfra | |
| $0.025 gemma-3-4b-it, DeepInfra | $0.11 gemma-3-27b-it, DeepInfra | $2.41 gemini-2.5-pro, DeepInfra | |
| $0.131 gemini-2.0-flash-lite-preview-02-05, Google | $0.131 gemini-2.0-flash-lite-preview-02-05, Google | $3.44 gemini-2.5-pro, Google | |
| $0.131 gemini-2.0-flash-lite-preview-02-05, Google | $0.131 gemini-2.0-flash-lite-preview-02-05, Google | $3.44 gemini-2.5-pro, Google | |
| $0.131 gemini-2.0-flash-lite-preview-02-05, Google | $0.131 gemini-2.0-flash-lite-preview-02-05, Google | No priced model yet | |
| $0.131 gemini-2.0-flash-lite-preview-02-05, Google | $0.131 gemini-2.0-flash-lite-preview-02-05, Google | No priced model yet | |
| $0.131 gemini-2.0-flash-lite-preview-02-05, Google | $0.131 gemini-2.0-flash-lite-preview-02-05, Google | No priced model yet | |
| $0.131 gemini-2.0-flash-lite-preview-02-05, Google | $0.131 gemini-2.0-flash-lite-preview-02-05, Google | No priced model yet | |
| $0.131 gemini-1.5-flash-002, Google | $0.90 deepseek-v3, Fireworks AI | No priced model yet | |
| $0.131 gemini-1.5-flash-002, Google | $5.25 gemini-1.5-pro-002, Google | No priced model yet | |
| $0.131 gemini-1.5-flash-002, Google | $5.25 gemini-1.5-pro-002, Google | No priced model yet | |
| $0.131 gemini-1.5-flash-002, Google | $5.25 gemini-1.5-pro-002, Google | No priced model yet | |
| $0.131 gemini-1.5-flash-002, Google | $5.25 gemini-1.5-pro-002, Google | No priced model yet | |
| $0.263 gpt-4o-mini-2024-07-18, OpenAI | $7.50 gpt-4o-2024-05-13, OpenAI | No priced model yet | |
| $0.263 gpt-4o-mini-2024-07-18, OpenAI | $7.50 gpt-4o-2024-05-13, OpenAI | No priced model yet | |
| $6.00 claude-3-5-sonnet-20240620, Anthropic | $7.50 gpt-4o-2024-05-13, OpenAI | No priced model yet | |
| $7.50 gpt-4o-2024-05-13, OpenAI | $7.50 gpt-4o-2024-05-13, OpenAI | No priced model yet | |
| $15.00 gpt-4-turbo-2024-04-09, OpenAI | No priced model yet | No priced model yet | |
| $15.00 gpt-4-0125-preview, OpenAI | No priced model yet | No priced model yet | |
| $15.00 gpt-4-0125-preview, OpenAI | No priced model yet | No priced model yet | |
| $15.00 gpt-4-0125-preview, OpenAI | No priced model yet | No priced model yet | |
| $15.00 gpt-4-1106-preview, OpenAI | No priced model yet | No priced model yet | |
| $15.00 gpt-4-1106-preview, OpenAI | No priced model yet | No priced model yet | |
| $37.50 gpt-4-0314, OpenAI | No priced model yet | No priced model yet | |
| $37.50 gpt-4-0314, OpenAI | No priced model yet | No priced model yet |
Blended price = (3 × input + output) ÷ 4, in US dollars per million tokens. Monthly rows read the LiteLLM price file as it stood at 00:00 UTC on the 1st of the month (from its git history); the top row reads today's file. Tier membership uses the LMArena leaderboard of 2 October 2026.
The moves that mattered
The GPT-4 class line is the longest, and its steps are worth reading one at a time, because each one is a different kind of price drop.
- October and November 2023: $37.50 blended. GPT-4 itself was the only model in its class, at $30 input and $60 output per million tokens.
- December 2023: $15. GPT-4 Turbo (gpt-4-1106-preview) arrived at $10 input and $30 output. The same lab, a newer model, less than half the price.
- June 2024: $7.50. GPT-4o at $5 and $15. It also set the floor for the GPT-4o class, which started that month.
- August 2024: $0.26. gpt-4o-mini at $0.15 input and $0.60 output. The floor fell by more than 95% in one month, because a small model cleared a bar that only frontier models had cleared before.
- October 2024 onward: about $0.13. Google's Gemini 1.5 Flash-002 and then Gemini 2.0 Flash-Lite, at $0.075 input and $0.30 output.
- September 2025 onward: a few cents. Small open-weight models (Gemma 3 4B, then gpt-oss-20b) served by DeepInfra. From here the floor belongs to hosts competing on open models, not to the labs.
The Gemini 2.5 Pro class shows the same pattern, faster. It started in July 2025 at $3.44 blended (Gemini 2.5 Pro itself, $1.25 input and $10 output) and by September 2026 an open-weight model on a host had taken the floor at about 15 cents. The frontier gets cheaper slowly; the level the frontier reached a year ago gets cheaper very fast.
If you are pricing an AI product, that gap is the useful part. I wrote up what it means for agents specifically in what it costs to run an AI agent around the clock and AI agent cost optimization.
Current API prices from 11 labs
The five newest paid text models from each lab, with input, output and cached-input prices per million tokens and the context window. This is the table to check for OpenAI API pricing, Claude API pricing, Gemini API pricing or DeepSeek API pricing on a given day. Prices are the lab's own list price wherever the model matches the lab's own entry in the LiteLLM file.
| Provider | Model | Input / 1M | Output / 1M | Cache read / 1M | Context | Last list-price change seen |
|---|---|---|---|---|---|---|
| OpenAI | GPT-6.1 Sol Pro (OpenRouter price) | $2.00 | $10.00 | $0.10 | 1.05M | between 1 August 2026 and 1 September 2026: gpt-5.6-sol output $30.00 to $20.00 |
| GPT-6.1 Sol | $2.00 | $10.00 | $0.10 | 1.05M | ||
| GPT-6 Luna Pro (OpenRouter price) | $0.10 | $0.50 | $0.01 | 1.05M | ||
| GPT-6 Luna | $0.10 | $0.50 | $0.01 | 1.05M | ||
| GPT-6 Sol Pro (OpenRouter price) | $2.00 | $10.00 | $0.20 | 1.05M | ||
| Anthropic | Claude Sonnet 5.5 | $2.00 | $10.00 | $0.20 | 1M | between 1 July 2026 and 1 August 2026: claude-sonnet-5 output $15.00 to $10.00 |
| Claude Opus 5.5 | $4.00 | $20.00 | $0.20 | 1M | ||
| Claude Fable 5.1 | $10.00 | $50.00 | $0.25 | 1M | ||
| Claude Opus 5 | $5.00 | $25.00 | $0.50 | 1M | ||
| Claude Sonnet 5 | $2.00 | $10.00 | $0.20 | 1M | ||
| Gemini 3.8 Flash | $0.75 | $3.75 | $0.075 | 1.05M | between 1 August 2026 and 1 September 2026: gemini/gemini-3.6-flash output $7.50 to $3.75 | |
| Gemini 3.7 Flash | $0.75 | $3.75 | $0.075 | 1.05M | ||
| Gemini 3.6 Flash | $0.75 | $3.75 | $0.075 | 1.05M | ||
| Gemini 3.5 Flash Lite | $0.30 | $2.50 | $0.03 | 1.05M | ||
| Gemini 3.5 Flash | $1.50 | $9.00 | $0.15 | 1.05M | ||
| xAI | Grok 4.7 | $2.00 | $6.00 | $0.50 | 500K | between 1 August 2026 and 1 September 2026: xai/grok-code-fast-1-0825 output $1.50 to $2.00 |
| Grok 4.6 | $2.00 | $6.00 | $0.50 | 500K | ||
| Grok 4.5 | $2.00 | $6.00 | $0.30 | 500K | ||
| Grok Build 0.1 | $1.00 | $2.00 | $0.20 | 256K | ||
| Grok 4.3 | $1.25 | $2.50 | $0.20 | 1M | ||
| DeepSeek | DeepSeek V4.1 Flash (OpenRouter price) | $0.003 | $2.40 | $0.003 | 1.05M | between 1 September 2026 and 1 October 2026: deepseek-v4-flash-vision-exp output $1.32 to $1.20 |
| DeepSeek V4 Flash Vision Exp | $0.30 | $1.20 | $0.006 | 1.05M | ||
| DeepSeek V4 Pro 0813 (OpenRouter price) | $0.55 | $4.20 | $0.45 | 1.05M | ||
| DeepSeek V4 Flash 0731 (OpenRouter price) | $0.0152 | $1.28 | $0.0152 | 1.05M | ||
| DeepSeek V4 Pro 0423 | $1.32 | $3.96 | $0.044 | 1.05M | ||
| Mistral | Mistral Medium 3.5 | $1.50 | $7.50 | $0.15 | 262K | between 1 September 2026 and 1 October 2026: mistral/mistral-medium output $8.10 to $7.50 |
| Mistral Small 4 | $0.15 | $0.60 | $0.015 | 262K | ||
| Devstral 2 2512 (OpenRouter price) | $0.40 | $2.00 | $0.04 | 262K | ||
| Ministral 3 14B 2512 | $0.20 | $0.20 | $0.02 | 262K | ||
| Ministral 3 8B 2512 | $0.15 | $0.15 | $0.015 | 262K | ||
| Alibaba (Qwen) | Qwen3.8 Max Prime (OpenRouter price) | $4.00 | $12.00 | $0.50 | 1M | None recorded since October 2023 |
| Qwen3.8 Omni Flash | $0.15 | $0.47 | $0.016 | 1M | ||
| Qwen3.8 Max (0902) (OpenRouter price) | $2.00 | $6.00 | $0.25 | 1M | ||
| Qwen3.8 Flash | $0.15 | $0.47 | $0.016 | 1M | ||
| Qwen3.8 27B (OpenRouter price) | $0.425 | $2.55 | $0.085 | 1M | ||
| Moonshot AI | Kimi K3 | $3.00 | $15.00 | $0.30 | 1.05M | between 1 December 2025 and 1 January 2026: moonshot/kimi-thinking-preview output $30.00 to $2.50 |
| Kimi K2.7 Code | $0.95 | $4.00 | $0.19 | 262K | ||
| Kimi K2.6 | $0.95 | $4.00 | $0.16 | 262K | ||
| Kimi K2.5 | $0.60 | $3.00 | $0.10 | 262K | ||
| Kimi K2 Thinking (OpenRouter price) | $0.60 | $2.50 | n/a | 262K | ||
| Z.ai | GLM 5.3 Prime (OpenRouter price) | $2.80 | $8.80 | $0.56 | 1M | None recorded since October 2023 |
| GLM 5.3 FlashX (OpenRouter price) | $0.37 | $1.25 | $0.09 | 1.05M | ||
| GLM 5.3 Flash | $0.15 | $0.50 | $0.03 | 1.05M | ||
| GLM 5.3 | $1.40 | $4.40 | $0.26 | 1.05M | ||
| GLM 5.2 | $1.40 | $4.40 | $0.26 | 1.05M | ||
| MiniMax | MiniMax M3 | $0.30 | $1.20 | $0.06 | 1.05M | None recorded since October 2023 |
| MiniMax M2.7 (OpenRouter price) | $0.21 | $0.84 | $0.042 | 205K | ||
| MiniMax M2.5 | $0.30 | $1.20 | $0.03 | 205K | ||
| MiniMax M2-her (OpenRouter price) | $0.30 | $1.20 | $0.03 | 66K | ||
| MiniMax M2.1 | $0.30 | $1.20 | $0.03 | 205K | ||
| Cohere | Command A+ (OpenRouter price) | $0.30 | $1.50 | $0.15 | 192K | between 1 June 2026 and 1 July 2026: command-r7b-12-2024 output $0.0375 to $0.15 |
| Command A (OpenRouter price) | $2.50 | $10.00 | n/a | 256K | ||
| Command R7B (12-2024) | $0.0375 | $0.15 | n/a | 128K | ||
| Command R (08-2024) | $0.15 | $0.60 | n/a | 128K | ||
| Command R+ (08-2024) | $2.50 | $10.00 | n/a | 128K |
The 5 newest paid text models per lab, by the date OpenRouter listed them, read on 4 October 2026. Prices in US dollars per million tokens are the lab's own list price from the LiteLLM file where the model matches one of the lab's own entries; rows marked "OpenRouter price" carry OpenRouter's listed price instead, which for open-weight models can be a third-party host's. Context is OpenRouter's figure. "Last list-price change seen" compares the lab's own entries in the LiteLLM file between snapshots: month by month before 4 October 2026, day by day after.
List price is the starting point, not the bill. Batch APIs, prompt caching, long-context surcharges and reasoning tokens all move the real number, and the AI agent pricing index shows how platforms built on these models turn token prices into plan prices.
Price change log
Dated list-price changes on the labs' own models, newest first. Monthly windows before 4 October 2026 come from the git history; from that date on, each line is a single day.
- : Mistral,
mistral/mistral-medium, input $2.70 to $1.50, output $8.10 to $7.50 per million tokens (LiteLLM price file) - : DeepSeek,
deepseek-v4-flash, input $0.44 to $0.30, output $1.32 to $1.20 per million tokens (LiteLLM price file) - : DeepSeek,
deepseek-v4-flash-vision-exp, input $0.44 to $0.30, output $1.32 to $1.20 per million tokens (LiteLLM price file) - : Google,
gemini/gemini-3.6-flash, input $1.50 to $0.75, output $7.50 to $3.75 per million tokens (LiteLLM price file) - : OpenAI,
gpt-5.6, input $5.00 to $4.00, output $30.00 to $20.00 per million tokens (LiteLLM price file) - : OpenAI,
gpt-5.6-sol, input $5.00 to $4.00, output $30.00 to $20.00 per million tokens (LiteLLM price file) - : xAI,
xai/grok-3, input $3.00 to $1.25, output $15.00 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-3-beta, input $3.00 to $1.25, output $15.00 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-3-fast-beta, input $5.00 to $1.25, output $25.00 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-3-mini, input $0.30 to $1.25, output $0.50 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-3-mini-beta, input $0.30 to $1.25, output $0.50 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-3-mini-fast, input $0.60 to $1.25, output $4.00 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-3-mini-fast-beta, input $0.60 to $1.25, output $4.00 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-4, input $3.00 to $1.25, output $15.00 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-4-fast-reasoning, input $0.20 to $1.25, output $0.50 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-4-fast-non-reasoning, input $0.20 to $1.25, output $0.50 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-4-0709, input $3.00 to $1.25, output $15.00 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-4-1-fast, input $0.20 to $1.25, output $0.50 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-4-1-fast-reasoning, input $0.20 to $1.25, output $0.50 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-4-1-fast-non-reasoning, input $0.20 to $1.25, output $0.50 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-4.20-beta-0309-reasoning, input $2.00 to $1.25, output $6.00 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-4.20-0309-reasoning, input $2.00 to $1.25, output $6.00 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-4.20-beta-0309-non-reasoning, input $2.00 to $1.25, output $6.00 to $2.50 per million tokens (LiteLLM price file) - : xAI,
xai/grok-code-fast, input $0.20 to $1.00, output $1.50 to $2.00 per million tokens (LiteLLM price file) - : xAI,
xai/grok-code-fast-1, input $0.20 to $1.00, output $1.50 to $2.00 per million tokens (LiteLLM price file) - : xAI,
xai/grok-code-fast-1-0825, input $0.20 to $1.00, output $1.50 to $2.00 per million tokens (LiteLLM price file) - : DeepSeek,
deepseek-v4-flash, input $0.14 to $0.44, output $0.28 to $1.32 per million tokens (LiteLLM price file) - : DeepSeek,
deepseek-v4-pro, input $0.435 to $1.32, output $0.87 to $3.96 per million tokens (LiteLLM price file) - : Anthropic,
claude-sonnet-5, input $3.00 to $2.00, output $15.00 to $10.00 per million tokens (LiteLLM price file) - : Cohere,
command-r7b-12-2024, input $0.15 to $0.0375, output $0.0375 to $0.15 per million tokens (LiteLLM price file) - : DeepSeek,
deepseek-chat, input $0.60 to $0.28, output $1.70 to $0.42 per million tokens (LiteLLM price file) - : DeepSeek,
deepseek-reasoner, input $0.60 to $0.28, output $1.70 to $0.42 per million tokens (LiteLLM price file) - : DeepSeek,
deepseek/deepseek-chat, input $0.27 to $0.28, output $1.10 to $0.42 per million tokens (LiteLLM price file) - : DeepSeek,
deepseek/deepseek-reasoner, input $0.55 to $0.28, output $2.19 to $0.42 per million tokens (LiteLLM price file) - : Moonshot AI,
moonshot/kimi-thinking-preview, output $30.00 to $2.50 per million tokens (LiteLLM price file) - : OpenAI,
gpt-3.5-turbo, input $1.50 to $0.50, output $2.00 to $1.50 per million tokens (LiteLLM price file) - : OpenAI,
o3, input $10.00 to $2.00, output $40.00 to $8.00 per million tokens (LiteLLM price file) - : OpenAI,
o3-2025-04-16, input $10.00 to $2.00, output $40.00 to $8.00 per million tokens (LiteLLM price file) - : Google,
gemini/gemini-2.5-flash-preview-05-20, input $0.15 to $0.30, output $0.60 to $2.50 per million tokens (LiteLLM price file) - : Mistral,
mistral/mistral-small, input $1.00 to $0.10, output $3.00 to $0.30 per million tokens (LiteLLM price file)
The 40 most recent list-price changes on the 11 labs' own models, newest first. A change in the LiteLLM file can also be a correction to that file, which is why each line names its source. New models are not price changes and do not appear here.
Methodology
Sources. Prices come from LiteLLM's model_prices_and_context_window.json, an MIT-licensed file that the LiteLLM project keeps in step with provider pricing pages. Its git history supplies one snapshot per month: the file as it stood at 00:00 UTC on the 1st, from October 2023 (the first month boundary after the file appeared) to October 2026. Today's prices come from the same file on its main branch. The per-lab table lists models from the OpenRouter models API, which also supplies the context window. Capability ratings come from the LMArena leaderboard dataset (CC BY 4.0).
Blended price. One input-to-output weighting, three parts input to one part output, turns two prices into one comparable number. It is the convention Artificial Analysis uses, and it suits chat-style workloads; output-heavy work such as long generation or reasoning will cost more than the blend suggests.
Who counts. The tier floors count the 11 labs' own APIs (OpenAI, Anthropic, Google's Gemini API, xAI, DeepSeek, Mistral, Alibaba Cloud, Moonshot AI, Z.ai, MiniMax, Cohere) and six inference hosts with public per-token prices: Together AI, Fireworks AI, Groq, DeepInfra, Cerebras and SambaNova. Resellers, cloud marketplaces (Bedrock, Vertex, Azure) and entries priced at zero are left out. Standard pay-as-you-go rates only: no batch, priority or long-context tiers.
How the tiers are defined
A tier is a rating band on LMArena's text leaderboard (style control, overall category). A model is in the GPT-4 class when its rating is at or above gpt-4-0314's, in the GPT-4o class at or above gpt-4o-2024-05-13's, and in the Gemini 2.5 Pro class at or above gemini-2.5-pro's. LMArena ratings come from human preference votes, so a tier measures what people preferred in blind comparisons, which is not the same as coding or reasoning benchmarks. I chose it because it is public, covers hundreds of models from every major lab, and is published under an open licence.
LMArena model names are matched to LiteLLM entries by name, after removing provider prefixes and suffixes such as "-instruct" or a reasoning-effort label. Unmatched models simply do not count, which can only make a floor higher than the true cheapest price, never lower. Ratings refresh monthly; a refresh can move models in or out of a tier for later snapshots, and earlier snapshots keep the result they were computed with.
Known gaps
- Google before October 2024. Until then the LiteLLM file carried no per-token price for Gemini API models (zero or per-character Vertex rates), so Google's models enter the series in October 2024. Gemini 1.5 Flash was available earlier, so the GPT-4 class floor for mid-2024 may be overstated.
- Corrections look like price changes. When the LiteLLM maintainers fix a wrong price, the change log records it like a provider move. Changes larger than 20 times in either direction are treated as unit fixes and skipped.
- OpenRouter prices for open models. Where a model has no first-party entry in LiteLLM, the per-lab table shows OpenRouter's listed price, which for open-weight models can come from a third-party host.
- Time-of-day pricing. DeepSeek's own pricing page lists peak and off-peak rates, with off-peak at half the peak price as read on 4 October 2026. The LiteLLM file, and so this page, carries the peak rate.
- Quality is not identical inside a tier. A model at the tier floor matched the anchor on one leaderboard. Hosts also differ on quantization, speed, rate limits and data retention.
Free usage is a separate question with its own tracker: free LLM API tiers, verified daily. For what agent platforms charge on top of these tokens, see the cheapest AI agent platforms and AI agent pricing trends in 2026. Gravity's own plans are on the pricing page: Basic is free with one agent, and Pro is $5 a month during the alpha.
FAQ
1. How much have LLM API prices fallen since 2023?
For a fixed level of capability, by more than 99%. The cheapest model rated at or above the original GPT-4 on LMArena cost $37.50 per million tokens blended on 1 October 2023 and costs $0.0575 on 4 October 2026 (gpt-oss-20b on DeepInfra). Prices for the newest frontier models have not fallen that fast: the median new model from 11 labs lists at $0.75 input and $3.75 output per million tokens.
2. How much does the OpenAI API cost per million tokens?
On 4 October 2026, OpenAI's newest paid models list at $0.10 to $2.00 input and $0.50 to $10.00 output per million tokens. The newest: GPT-6.1 Sol at $2.00 input and $10.00 output; GPT-6 Luna at $0.10 input and $0.50 output. Cached input, where offered, costs less; the table above has the cache-read price and context window for each model.
3. How much does the Claude API cost?
On 4 October 2026, Anthropic's newest paid models list at $2.00 to $10.00 input and $10.00 to $50.00 output per million tokens. The newest: Claude Sonnet 5.5 at $2.00 input and $10.00 output; Claude Opus 5.5 at $4.00 input and $20.00 output; Claude Fable 5.1 at $10.00 input and $50.00 output. Cached input, where offered, costs less; the table above has the cache-read price and context window for each model.
4. How much does the Gemini API cost?
On 4 October 2026, Google's newest paid Gemini models list at $0.30 to $1.50 input and $2.50 to $9.00 output per million tokens. The newest: Gemini 3.8 Flash at $0.75 input and $3.75 output; Gemini 3.7 Flash at $0.75 input and $3.75 output; Gemini 3.6 Flash at $0.75 input and $3.75 output. Cached input, where offered, costs less; the table above has the cache-read price and context window for each model.
5. How much does the DeepSeek API cost?
On 4 October 2026, DeepSeek's newest paid models list at $0.30 to $1.32 input and $1.20 to $3.96 output per million tokens. The newest: DeepSeek V4 Flash Vision Exp at $0.30 input and $1.20 output; DeepSeek V4 Pro 0423 at $1.32 input and $3.96 output. Cached input, where offered, costs less; the table above has the cache-read price and context window for each model.
6. What does "GPT-4 class" mean on this page?
A model is in the GPT-4 class when its LMArena rating (text, style control, overall) is at or above the rating of gpt-4-0314, the original GPT-4 from March 2023. The GPT-4o class uses gpt-4o-2024-05-13 as its floor and the Gemini 2.5 Pro class uses gemini-2.5-pro. Ratings come from the leaderboard published 2 October 2026, so the tiers measure what a model can do, not what its maker calls it.
7. Why is the cheapest model in a tier often served by a host, not the lab that made it?
Open-weight models such as gpt-oss, Gemma, Llama, Qwen and DeepSeek can be served by anyone, and inference hosts like DeepInfra, Together AI and Fireworks AI compete on price. The tier floors count the labs' own APIs and six large hosts with public per-token prices. Hosts can differ on speed, rate limits, quantization and data terms, so check those before switching for price alone.
8. How often is this page updated, and can I reuse the data?
The current prices and the tier floors refresh every day; the monthly history goes back to October 2023. The data is CC BY 4.0: copy it, chart it, or publish it with a link back to this page. Both CSV files are linked in the "Cite this page" box.
Sources
- BerriAI, LiteLLM, model_prices_and_context_window.json and its commit history, MIT licence, read daily since 4 October 2026 and backfilled to October 2023.
- OpenRouter, models API, read daily since 4 October 2026.
- LMArena, leaderboard dataset on Hugging Face, text style control, overall category, leaderboard of 2 October 2026, CC BY 4.0.
- Guido Appenzeller, a16z, Welcome to LLMflation: LLM inference cost is going down fast, 12 November 2024.
- Epoch AI, LLM inference prices have fallen rapidly but unequally across tasks, accessed 4 October 2026.