LLM API Pricing Comparison (2026)
Last verified 2026-08-03 ยท Prices per 1M tokens (USD) ยท Standard API tier unless noted
Compare input and output API costs for major frontier and budget models. Each row links to the official pricing page used for verification. For subscription-style coding assistants, see the coding plan comparison.
Also compare: Coding plan pricing ยท Image generation comparison ยท Assistant comparison ไธญๆ
Verdict matrix
Choose an API tier by workload economics
Choose a budget model if...
- โข You run classification, extraction, routing, or high-volume simple tasks.
- โข Latency and unit cost matter more than maximum reasoning quality.
- โข You can add evaluation, fallback, or routing logic around the model.
Choose a balanced model if...
- โข You need stronger reasoning without paying frontier-model rates for every request.
- โข The workload mixes coding, analysis, structured generation, and agent steps.
- โข You want one default model before adding specialized routing.
Choose a premium model if...
- โข Task failure is more expensive than token cost.
- โข You need difficult reasoning, long-context work, complex coding, or high-value agent decisions.
- โข You will measure quality gain against output length, latency, and review cost.
| Provider | Model | Input $/MTok | Output $/MTok | Context | Source |
|---|---|---|---|---|---|
| ๐ฐ Budget Tier (< $1/MTok input) | |||||
| OpenAI | GPT-5 nano | $0.05 | $0.4 | 400K | openai.com |
| OpenAI | GPT-4.1 nano | $0.1 | $0.4 | 1M | openai.com |
| Gemini 2.0 Flash | $0.1 | $0.4 | 1M | ai.google.dev | |
| Gemini 2.5 Flash Lite | $0.1 | $0.4 | 1M | ai.google.dev | |
| DeepSeek | Deepseek V4 Flash | $0.14 | $0.28 | 1M | api-docs.deepseek.com |
| OpenAI | GPT-4o mini | $0.15 | $0.6 | 128K | openai.com |
| OpenAI | GPT-5.4 nano | $0.2 | $1.25 | 400K | openai.com |
| OpenAI | GPT-5 mini | $0.25 | $2 | 400K | openai.com |
| Gemini 3.1 Flash Lite | $0.25 | $1.5 | 1M | ai.google.dev | |
| Gemini 2.5 Flash | $0.3 | $2.5 | 1M | ai.google.dev | |
| OpenAI | GPT-4.1 mini | $0.4 | $1.6 | 1M | openai.com |
| DeepSeek | Deepseek V4 Pro | $0.43 | $0.87 | 1M | api-docs.deepseek.com |
| OpenAI | GPT-5.4 mini | $0.75 | $4.5 | 400K | openai.com |
| ๐ Mid Tier ($1โ3/MTok input) | |||||
| Anthropic | Claude Haiku 4.5 | $1 | $5 | 1M | claude.com |
| OpenAI | GPT-5.6 Luna | $1 | $6 | 1M | openai.com |
| OpenAI | o4-mini | $1.1 | $4.4 | 200K | openai.com |
| OpenAI | GPT-5 | $1.25 | $10 | 1M | openai.com |
| OpenAI | GPT-5.1 | $1.25 | $10 | 1M | openai.com |
| Gemini 2.5 Pro | $1.25 | $10 | 1M | ai.google.dev | |
| Gemini 3.5 Flash | $1.5 | $9 | 1M | ai.google.dev | |
| OpenAI | GPT-5.2 | $1.75 | $14 | 1M | openai.com |
| OpenAI | GPT-4.1 | $2 | $8 | 1M | openai.com |
| OpenAI | o3 | $2 | $8 | 200K | openai.com |
| OpenAI | GPT-4o | $2.5 | $10 | 128K | openai.com |
| OpenAI | GPT-5.4 | $2.5 | $15 | 1M | openai.com |
| OpenAI | GPT-5.6 Terra | $2.5 | $15 | 1M | openai.com |
| โญ Premium Tier ($3+/MTok input) | |||||
| Anthropic | Claude Sonnet 4.5 | $3 | $15 | 1M | claude.com |
| Anthropic | Claude Sonnet 4.6 | $3 | $15 | 1M | claude.com |
| Anthropic | Claude Opus 4.5 | $5 | $25 | 1M | claude.com |
| Anthropic | Claude Opus 4.6 | $5 | $25 | 1M | claude.com |
| Anthropic | Claude Opus 4.7 | $5 | $25 | 1M | claude.com |
| Anthropic | Claude Opus 4.8 | $5 | $25 | 1M | claude.com |
| Anthropic | Claude Opus 5 | $5 | $25 | 1M | claude.com |
| OpenAI | GPT-5.5 | $5 | $30 | 1M | openai.com |
| OpenAI | GPT-5.6 Sol | $5 | $30 | 1M | openai.com |
| Anthropic | Claude Opus 4.1 | $15 | $75 | 1M | claude.com |
| OpenAI | GPT-5 pro | $15 | $120 | 1M | openai.com |
| OpenAI | GPT-5.4 pro | $30 | $180 | 1M | openai.com |
| OpenAI | GPT-5.5 pro | $30 | $180 | 1M | openai.com |
| ๐ Open-Weight (via API hosts; not vendor list price) | |||||
| Meta | Llama 4 Maverick (hosted) | $0.05โ0.90 | $0.05โ0.90 | 1M | groq.com ยท docs.fireworks.ai |
| DeepSeek | DeepSeek R1 (hosted) | See official pricing | See official pricing | 128K | api-docs.deepseek.com |
| Alibaba | Qwen3 Max (DashScope) | See official pricing | See official pricing | 128K | help.aliyun.com |
* Range rows: price depends on hosting provider or SKU; see linked official pages. OpenAI rows use Standard tier unless noted.
๐ Output Cost Visual Comparison ($/MTok)
๐ก Key Insights
Output often costs more than input
Most providers charge higher $/MTok for output than input. Shorter prompts and concise output formats reduce monthly spend.
Route by workload instead of defaulting to the flagship
GPT-5.6 Luna and Terra create lower-cost routing options for many workloads; reserve Sol or other premium models for tasks that justify the higher output-token price.
Open-weight hosting varies
Llama and similar models are priced by the inference host (Groq, Fireworks, etc.), not a single vendor list priceโalways check your provider's page.
Promotions change quickly
DeepSeek and others run time-limited discounts. Re-check official pages before locking budgets; this table shows last_verified date in the header.
๐งฎ Quick Cost Estimator
Estimated Monthly Cost
โ
๐ Official Pricing Pages
- OpenAI API Pricing โ openai.com
- Anthropic API Pricing โ claude.com
- Google Gemini API Pricing โ ai.google.dev
- DeepSeek API Pricing โ api-docs.deepseek.com
- Groq Pricing โ groq.com
- Fireworks AI Serverless Pricing โ docs.fireworks.ai
- Alibaba Model Studio (Qwen) Pricing โ help.aliyun.com
Prices reflect official list rates as of last verified date above. Providers may change pricing without notice. Always confirm on the linked official pages before budgeting.