Skip to main content
AI Toolset
AI Picker — decision engine, not directory

Choose the right AI tool before you pay.

· 中文

Start with your task, then compare the few tools that actually fit your workflow, budget, and risk level.

LLM intelligence leaderboard

Artificial Analysis Intelligence Index

Which frontier model is actually smartest? Artificial Analysis aggregates 9 independent evaluations — including GPQA Diamond, Humanity’s Last Exam, and Terminal-Bench — into a single intelligence score. Higher is better.

Top 20 models by Artificial Analysis Intelligence Index. Parentheses show reasoning effort; bars are scaled to the leading score from a baseline of 20.1, not zero; labels show actual scores. Data snapshot: · Index version: v4.3.2. Methodology on Artificial Analysis · See our full model rankings

Component evaluations and operating costs

Thirteen leaderboards from Artificial Analysis

The nine evaluations behind the Intelligence Index, plus per-task cost, API pricing, and output speed — the tradeoffs that decide real-world value. Higher is better except the two cost charts.

GDPval-AA v2

Quantitative analysis on real-world spreadsheets and documents.

Claude Opus 5.567.3%
Claude Fable 5.161.7%
Grok 4.759.8%
Muse Spark 1.358.7%
MiMo-V2.6-Pro58.7%
Qwen3.8 Max (0902)58.4%
GLM-5.357.3%
GLM 5.3 Flash57.0%
Grok 4.655.3%
DeepSeek V4.1 Flash55.0%

τ³-Banking

Agentic tool use on realistic banking workflows.

Grok 4.650.7%
Muse Spark 1.350.5%
GLM-5.350.3%
Qwen3.8 27B48.0%
Qwen3.8 Max (0902)47.8%
Claude Fable 5.147.2%
GLM 5.3 Flash47.2%
Kimi K346.0%
Gemini 3.8 Flash44.9%
GPT-6 Astra41.4%

SciCode

Scientific computing and research code problems.

Claude Opus 5.566.9%
Claude Fable 5.163.1%
MiMo-V2.6-Pro60.9%
Kimi K359.5%
GLM-5.359.0%
Step 5 Preview58.9%
Muse Spark 1.358.8%
GPT-6 Sol57.6%
Grok 4.757.4%
Gemini 3.8 Flash56.6%

Humanity's Last Exam

Extremely hard academic questions across dozens of domains.

Claude Opus 5.561.4%
Claude Fable 5.159.1%
GPT-6 Astra54.7%
MiMo-V2.6-Pro49.4%
Muse Spark 1.348.7%
GPT-6 Sol47.9%
Gemini 3.8 Flash47.8%
Kimi K346.9%
Step 5 Preview46.5%
Grok 4.743.1%

GPQA Diamond

Graduate-level science questions (diamond subset).

GPT-6 Astra96.1%
Gemini 3.8 Flash95.3%
Grok 4.694.9%
Claude Fable 5.193.7%
Muse Spark 1.393.5%
Kimi K393.5%
MiniMax-M392.9%
Qwen3.8 Max (0902)92.8%
GLM-5.391.7%
GLM 5.3 Flash91.2%

CritPt

Critical-point reasoning under ambiguous instructions.

Claude Opus 5.531.7%
GPT-6 Astra31.7%
GPT-6 Sol30.9%
GPT-5.5 Pro30.6%
Claude Fable 5.129.7%
MiMo-V2.6-Pro26.6%
Muse Spark 1.324.9%
Kimi K323.4%
Step 5 Preview20.9%
GPT-6 Luna19.4%

AA-LCR

Long-context reasoning across extended documents.

Kimi K388.7%
Step 5 Preview88.3%
MiMo-V2.6-Pro86.3%
Claude Fable 5.185.3%
Claude Opus 5.584.7%
DeepSeek V4.1 Flash84.0%
GPT-6 Sol83.7%
GPT-6 Luna83.3%
Muse Glimmer83.3%
Muse Spark 1.383.0%

Cost per Task

Blended USD cost per Intelligence Index task, split by token type.

Muse Glimmer (high)$0.055
GPT-6 Luna (max)$0.068
Gemini 3.5 Flash-Lite$0.124
MiMo-V2.6-Pro$0.133
GLM-5.3-Flash$0.253
DeepSeek V4.1 Flash (max)$0.265
Mistral Medium 3.5$0.437
MiniMax-M3$0.508
Nemotron 3 Ultra$0.549
Inkling$0.607

Output Speed (tok/s)

Median output tokens per second from the model API.

Gemini 3.5 Flash-Lite345
Gemini 3.8 Flash (high)288
DeepSeek V4.1 Flash (max)232
Muse Spark 1.3 (max)222
Inkling163
Mistral Medium 3.5149
MiniMax-M3145
GPT-6 Luna (max)132
Nemotron 3 Ultra121
GPT-6 Sol (max)104

All charts: top 10 models per chart, measured independently by Artificial Analysis. Bar lengths are truncated to highlight differences; labels show actual scores. Data snapshot: . Methodology on Artificial Analysis

LLM API pricing detail

API pricing: cache hit, input, and output

Per-model list prices in USD per million tokens, sorted from most to least expensive by the sum of all three. Three separate bars per model; all bars share one scale, so prices are directly comparable across models.

Kimi K3 (max)
$0.300
$3.00
$15.00
GPT-6 Sol (max)
$0.200
$2.00
$10.00
Mistral Medium 3.5
$0.150
$1.50
$7.50
Grok 4.7 (xhigh)
$0.500
$2.00
$6.00
Grok 4.6 (high)
$0.500
$2.00
$6.00
Qwen3.8 Max (0902)
$0.250
$2.00
$6.00
GLM-5.3 (max)
$0.260
$1.40
$4.40
Muse Spark 1.3 (max)
$0.150
$1.25
$4.25
Inkling
$0.170
$1.00
$4.05
Gemini 3.8 Flash (high)
$0.075
$0.750
$3.75
Step 5 Preview
$0.050
$1.00
$2.70
Qwen3.8 27B (xhigh)
$0.100
$0.500
$3.00
Nemotron 3 Ultra
$0.160
$0.600
$2.50
Gemini 3.5 Flash-Lite
$0.030
$0.300
$2.50
Muse Glimmer (high)
$0.040
$0.350
$1.50
MiniMax-M3
$0.060
$0.300
$1.20
DeepSeek V4.1 Flash (max)
$0.006
$0.300
$1.20
MiMo-V2.6-Pro
$0.004
$0.435
$0.870
GLM-5.3-Flash
$0.026
$0.150
$0.500
GPT-6 Luna (max)
$0.010
$0.100
$0.500

List prices from model APIs, measured and published by Artificial Analysis. Cache-hit price is missing for models that do not offer prompt caching. Data snapshot: .

Interactive AI Picker

Choose by job, not by hype

Select a task to get a practical recommendation path: what to compare first, which tools are worth opening, and what risk or cost check to do before paying.

Writing workflow

Start with Claude or ChatGPT, then check writing-specific tools only if you need workflow features.

For most writing jobs, the first decision is not “which writing app?” but “which assistant gives the best draft, edit, and reasoning loop for your budget?”

Coding workflow

Pick the interface first: editor, extension, or terminal agent.

Cursor is strongest when the editor is the workflow. Claude Code fits terminal-heavy repo work. Plan limits and team controls matter more than raw model names.

Research workflow

Use citation-aware tools for discovery, then source-grounded tools for synthesis.

Perplexity is useful for sourced exploration; NotebookLM is better when you already have documents and need grounded synthesis.

Image workflow

Choose by output quality, control, and production constraints.

Midjourney is a strong creative benchmark, Flux is strong for model/API workflows, and Stable Diffusion-style stacks matter when control and ownership are important.

Video workflow

Start with the workflow: ideation, short clips, editing, or production assets.

Video tools are expensive and fast-changing, so use them for concrete creative jobs instead of subscribing before you know the output format you need.

API cost workflow

Estimate cost before you prototype at scale.

Check input/output token pricing, free tiers, latency expectations, and fallback models before committing to one API provider.

Team buying workflow

Before buying seats, check data policy, admin controls, audit needs, and actual usage.

The right AI tool for a team is often the one with acceptable risk controls, predictable billing, and workflows people will actually use.

Start by task

What are you trying to do?

These guide pages are the main entry points. Each one narrows the market to a practical shortlist.

Compare before you pay

Core comparisons with clear tradeoffs

Use these pages when you already know the category but need help picking the winner.

Canonical tool profiles

Ten tools we currently treat as core

These profiles support the guide and comparison pages. The site intentionally avoids pushing every long-tail tool equally.

Pricing and risk checks

Before you choose, check cost and safety

Good AI tool decisions are not just feature decisions. Cost, limits, privacy, and workflow fit matter.

Open-source trends

Track AI projects gaining momentum

Keep the picker focused on decisions, but keep a clear path to GitHub AI momentum when you want to spot fast-rising open-source projects.