LLM API Pricing Comparison (2026)
Last verified 2026-10-02 · Prices per 1M tokens (USD) · Standard API tier unless noted
Compare input and output API costs for major frontier and budget models. Each row links to the official pricing page used for verification. For subscription-style coding assistants, see the coding plan comparison.
Also compare: Coding plan pricing · Image generation comparison · Assistant comparison 中文
Verdict matrix
Choose an API tier by workload economics
Choose a budget model if...
- • You run classification, extraction, routing, or high-volume simple tasks.
- • Latency and unit cost matter more than maximum reasoning quality.
- • You can add evaluation, fallback, or routing logic around the model.
Choose a balanced model if...
- • You need stronger reasoning without paying frontier-model rates for every request.
- • The workload mixes coding, analysis, structured generation, and agent steps.
- • You want one default model before adding specialized routing.
Choose a premium model if...
- • Task failure is more expensive than token cost.
- • You need difficult reasoning, long-context work, complex coding, or high-value agent decisions.
- • You will measure quality gain against output length, latency, and review cost.
| Provider | Model | Input $/MTok | Output $/MTok | Context | Source |
|---|---|---|---|---|---|
| 💰 Budget Tier (< $1/MTok input) | |||||
| OpenAI | GPT-5 nano | $0.05 | $0.4 | 400K | openai.com |
| OpenAI | GPT-4.1 nano | $0.1 | $0.4 | 1M | openai.com |
| Gemini 2.0 Flash | $0.1 | $0.4 | 1M | ai.google.dev | |
| Gemini 2.5 Flash Lite | $0.1 | $0.4 | 1M | ai.google.dev | |
| DeepSeek | Deepseek V4 Flash | $0.14 | $0.28 | 1M | api-docs.deepseek.com |
| OpenAI | GPT-4o mini | $0.15 | $0.6 | 128K | openai.com |
| OpenAI | GPT-5 mini | $0.25 | $2 | 400K | openai.com |
| Gemini 3.1 Flash Lite | $0.25 | $1.5 | 1M | ai.google.dev | |
| Gemini 2.5 Flash | $0.3 | $2.5 | 1M | ai.google.dev | |
| Gemini 3.5 Flash Lite | $0.3 | $2.5 | 1M | ai.google.dev | |
| OpenAI | GPT-4.1 mini | $0.4 | $1.6 | 1M | openai.com |
| DeepSeek | Deepseek V4 Pro | $0.43 | $0.87 | 1M | api-docs.deepseek.com |
| Gemini 3.6 Flash | $0.75 | $3.75 | 1M | ai.google.dev | |
| Gemini 3.7 Flash | $0.75 | $3.75 | 1M | ai.google.dev | |
| Z AI | GLM-5.3-Flash | $0.15 | $0.5 | artificialanalysis.ai · artificialanalysis.ai | |
| MiniMax | MiniMax-M3 | $0.3 | $1.2 | artificialanalysis.ai · artificialanalysis.ai | |
| MiniMax | MiniMax-M2.5 | $0.3 | $1.2 | artificialanalysis.ai · artificialanalysis.ai | |
| Xiaomi | MiMo-V2.6-Flash | $0.14 | $0.28 | artificialanalysis.ai · artificialanalysis.ai | |
| StepFun | Step 3.5 Flash | $0.1 | $0.3 | artificialanalysis.ai · artificialanalysis.ai | |
| Apodex | Apodex 1.1 | $0.3 | $3 | artificialanalysis.ai · artificialanalysis.ai | |
| InclusionAI | Ling-3.0-flash-VL | $0.07 | $0.22 | artificialanalysis.ai · artificialanalysis.ai | |
| InclusionAI | Ling-2.6-1T | $0.3 | $2.5 | artificialanalysis.ai · artificialanalysis.ai | |
| IBM | Granite 4.2 30B | $0.16 | $0.65 | artificialanalysis.ai · artificialanalysis.ai | |
| 📊 Mid Tier ($1–3/MTok input) | |||||
| Anthropic | Claude Haiku 4.5 | $1 | $5 | 1M | claude.com |
| OpenAI | o4-mini | $1.1 | $4.4 | 200K | openai.com |
| OpenAI | GPT-5 | $1.25 | $10 | 1M | openai.com |
| OpenAI | GPT-5.1 | $1.25 | $10 | 1M | openai.com |
| Gemini 2.5 Pro | $1.25 | $10 | 1M | ai.google.dev | |
| Gemini 3.5 Flash | $1.5 | $9 | 1M | ai.google.dev | |
| OpenAI | GPT-5.2 | $1.75 | $14 | 1M | openai.com |
| Anthropic | Claude Sonnet 5 | $2 | $10 | 1M | claude.com |
| OpenAI | GPT-4.1 | $2 | $8 | 1M | openai.com |
| OpenAI | o3 | $2 | $8 | 200K | openai.com |
| OpenAI | GPT-4o | $2.5 | $10 | 128K | openai.com |
| Z AI | GLM-5.3 | $1.4 | $4.4 | artificialanalysis.ai · artificialanalysis.ai | |
| Alibaba | Qwen3.8 27B | $0.5 | $3 | artificialanalysis.ai · artificialanalysis.ai | |
| Alibaba | Qwen3.8 Max | $2 | $6 | artificialanalysis.ai · artificialanalysis.ai | |
| SpaceXAI | Grok 4.7 | $2 | $6 | artificialanalysis.ai · artificialanalysis.ai | |
| SpaceXAI | Grok 4.6 | $2 | $6 | artificialanalysis.ai · artificialanalysis.ai | |
| Kimi | Kimi K3 | $3 | $15 | artificialanalysis.ai · artificialanalysis.ai | |
| Kimi | Kimi K2.6 | $0.95 | $4 | artificialanalysis.ai · artificialanalysis.ai | |
| Xiaomi | MiMo-V2.6-Pro | $0.43 | $0.87 | artificialanalysis.ai · artificialanalysis.ai | |
| StepFun | Step 5 Preview | $1 | $2.7 | artificialanalysis.ai · artificialanalysis.ai | |
| Thinking Machines | Inkling | $1 | $4.05 | artificialanalysis.ai · artificialanalysis.ai | |
| ⭐ Premium Tier ($3+/MTok input) | |||||
| Anthropic | Claude Sonnet 4.5 | $3 | $15 | 1M | claude.com |
| Anthropic | Claude Sonnet 4.6 | $3 | $15 | 1M | claude.com |
| Anthropic | Claude Opus 4.5 | $5 | $25 | 1M | claude.com |
| Anthropic | Claude Opus 4.6 | $5 | $25 | 1M | claude.com |
| Anthropic | Claude Opus 4.7 | $5 | $25 | 1M | claude.com |
| Anthropic | Claude Opus 4.8 | $5 | $25 | 1M | claude.com |
| Anthropic | Claude Opus 5 | $5 | $25 | 1M | claude.com |
| Anthropic | Claude Opus 4.1 | $15 | $75 | 1M | claude.com |
| OpenAI | GPT-5 pro | $15 | $120 | 1M | openai.com |
| 🔓 Open-Weight (via API hosts; not vendor list price) | |||||
| Meta | Llama 4 Maverick (hosted) | $0.05–0.90 | $0.05–0.90 | 1M | groq.com · docs.fireworks.ai |
| DeepSeek | DeepSeek R1 (hosted) | See official pricing | See official pricing | 128K | api-docs.deepseek.com |
| Alibaba | Qwen3 Max (DashScope) | See official pricing | See official pricing | 128K | help.aliyun.com |
| NVIDIA | Nemotron 3 Ultra | $0.6 | $2.5 | artificialanalysis.ai · artificialanalysis.ai | |
| NVIDIA | Nemotron 3.5 Lightning | $0.06 | $0.2 | artificialanalysis.ai · artificialanalysis.ai | |
| Mistral | Mistral Medium 3.5 | $1.5 | $7.5 | artificialanalysis.ai · artificialanalysis.ai | |
| Mistral | Mistral Small 4 | $0.15 | $0.6 | artificialanalysis.ai · artificialanalysis.ai | |
| Upstage | Solar Pro 4 | $0.3 | $1.2 | artificialanalysis.ai · artificialanalysis.ai | |
* Range rows: price depends on hosting provider or SKU; see linked official pages. OpenAI rows use Standard tier unless noted.
📊 Output Cost Visual Comparison ($/MTok)
💡 Key Insights
Output often costs more than input
Most providers charge higher $/MTok for output than input. Shorter prompts and concise output formats reduce monthly spend.
Route by workload instead of defaulting to the flagship
GPT-5.6 Luna and Terra create lower-cost routing options for many workloads; reserve Sol or other premium models for tasks that justify the higher output-token price.
Open-weight hosting varies
Llama and similar models are priced by the inference host (Groq, Fireworks, etc.), not a single vendor list price—always check your provider's page.
Promotions change quickly
DeepSeek and others run time-limited discounts. Re-check official pages before locking budgets; this table shows last_verified date in the header.
🧮 Quick Cost Estimator
Estimated Monthly Cost
—
📋 Official Pricing Pages
- OpenAI API Pricing — openai.com
- Anthropic API Pricing — claude.com
- Google Gemini API Pricing — ai.google.dev
- DeepSeek API Pricing — api-docs.deepseek.com
- Groq Pricing — groq.com
- Fireworks AI Serverless Pricing — docs.fireworks.ai
- Alibaba Model Studio (Qwen) Pricing — help.aliyun.com
Prices reflect official list rates as of last verified date above. Providers may change pricing without notice. Always confirm on the linked official pages before budgeting.
Mistral Small 4