๐Ÿค– AI Toolset

📅 2026-07-09 ยท 8 min read

Grok 4.5 in Cursor: Availability, Pricing, Benchmarks, and Composer 2.5 Differences

Grok 4.5 expands Cursor's first-party model pool beyond coding-specialist work. This guide separates official availability, pricing, benchmark context, and the practical difference from Composer 2.5.

Benchmark snapshot

Grok 4.5 at a glance

Provider-reported launch scores. Compare only within the same benchmark and harness.

See full comparisons โ†’
62.0%
DeepSWE 1.0
53%
DeepSWE 1.1
29.0%
SWE Marathon
83.3%
Terminal Bench 2.1
64.7%
SWE Bench Pro
15,954 avg. output tokens/task on SWE Bench Proโ‰ˆ4.2ร— fewer than Opus 4.8 Max in the official comparison

What launched

Cursor and SpaceXAI released Grok 4.5 on July 8, 2026. Cursor describes it as a mixture-of-experts model jointly trained with SpaceXAI and as its first model built for more than software engineering. The official positioning covers long-running tool use across software engineering, data science, finance, legal work, and other computer-based knowledge tasks.

This matters for Cursor users because the model is not presented as a simple Composer replacement. Cursor explicitly says Grok 4.5 and Composer 2.5 are different model weight classes and that Composer 2.5 will remain available.

Availability

According to Cursor, Grok 4.5 is available across desktop, web, iOS, CLI, and the Cursor SDK. Individual and team plans include usage from the first-party model pool, and Cursor announced double usage for the first week after launch. SpaceXAI also says the model is available through Grok Build and its API console.

Availability is not the same as unlimited usage. Check your current Cursor plan and model documentation before budgeting around a launch promotion or assuming every surface has identical limits.

Official model pricing

Cursor and SpaceXAI list the Grok 4.5 base model at $2 per million input tokens and $6 per million output tokens. Cursor also lists a fast variant at $4 per million input tokens and $18 per million output tokens.

These are model usage rates, not a substitute for checking Cursor subscription pricing, included usage, account-specific limits, or regional availability. Pricing and promotions can change.

Benchmark context

SpaceXAI's official launch page compares Grok 4.5 with other frontier models across five engineering benchmarks. The charts below reproduce the published scores as accessible HTML bars so you can compare model ordering without relying on a screenshot.

Decision matrix

Where Grok 4.5 leads โ€” and where it does not

Use this as a reading guide for the benchmark charts below, not as a universal model ranking.

BenchmarkGrok 4.5 resultPosition in official comparisonDecision takeaway
DeepSWE 1.062.0%Behind 2 modelsStrong, but not the top choice on this provider-harness evaluation.
DeepSWE 1.153%Behind 3 modelsThe harness change materially alters ordering; do not generalize from DeepSWE 1.0.
SWE Marathon29.0%Leads comparisonBest headline result in the official engineering comparison set.
Terminal Bench 2.183.3%Near the topVery competitive for terminal-heavy agent work, but not uniquely dominant.
SWE Bench Pro64.7%Behind 2 modelsCompetitive result; token efficiency may be as important as raw resolve rate for cost-sensitive workflows.
DeepSWE 1.0

Pass@1 ยท DataCurve eval, each provider's harness, run by Artificial Analysis

DeepSWE 1.1

mini-swe-agent harness ยท run by DataCurve

SWE Marathon

Resolution rate ยท pass@1

Terminal Bench 2.1
SWE Bench Pro

Resolve rate

Token-efficiency context on SWE Bench Pro

15,954
Grok 4.5 avg. output tokens/task
67,020
Opus 4.8 Max avg. output tokens/task

SpaceXAI reports Grok 4.5 using about 4.2ร— fewer output tokens on this task set. This is a provider-reported efficiency comparison, not a universal cost guarantee for every workload.

Source: SpaceXAI's official Grok 4.5 launch page. Competitor figures are reproduced from the comparison data published there. Benchmark methods differ, so compare scores only within the same benchmark and harness context.

A useful example is DeepSWE: the provider-harness-oriented 1.0 table and the mini-swe-agent 1.1 table produce materially different scores and ordering. Treat benchmark method, harness, and who ran the evaluation as part of the result.

Cursor also disclosed that an earlier snapshot of the Cursor codebase was accidentally included in training, giving Grok 4.5 an advantage on CursorBench. Cursor excluded that result from the launch comparison and says the data has been removed for future models.

Grok 4.5 vs Composer 2.5

Composer 2.5 remains Cursor's coding-specialist model for sustained multi-step coding and instruction following. Grok 4.5 uses a broader training mix and is positioned for coding plus general knowledge work, data analysis, finance, legal work, and other tool-heavy computer tasks.

DecisionFirst model to test
Long-running repo codingTest Composer 2.5 and Grok 4.5 on the same repo task
Terminal-heavy agent workStart with Grok 4.5, then compare completion cost and review burden
Data analysis or broader knowledge workStart with Grok 4.5
Stable existing Composer workflowDo not switch only because a new model launched; run a controlled pilot

Who should try it first

A launch-week benchmark table is not enough reason to standardize. The practical test is whether the model completes your real task with fewer retries and acceptable review burden at a cost your team can predict.

FAQ

Did Grok 4.5 replace Composer 2.5 in Cursor?

No. Cursor says they are different model weight classes and that Composer 2.5 remains available.

Is Grok 4.5 only for coding?

No. Cursor and SpaceXAI position it for coding, agentic tasks, and broader knowledge work.

Should I trust one benchmark number?

No. Compare benchmark task, harness, evaluation operator, and contamination caveats before using a score for a buying decision.