Cursor

Composer 2.5

Version 2.5 benchmark runs across the shared Superbash visual prompts.

Canonical model record

Current identity, limits, and pricing

Provider source · checked 2026-08-14 ↗
Status
Current
API model ID
Not publicly verified
Context
Not publicly verified
Max output
Not publicly verified
API price / 1M tokens
$0.5 input · $2.5 output

Standard mode; Cursor also publishes a faster $3 input / $15 output mode.

Visual prompt runs

Benchmark runs

Open each generated scene, or compare the same prompt across models.

9 runs

Official benchmark profile

How Composer 2.5 performs on coding benchmarks.

Cursor continued training Kimi K2.5 for Composer 2.5 and reports gains over Composer 2 across agentic, multilingual, and harder coding tasks. The release table publishes three scores.

Cursor sourceMay 2026Source report →
Agentic coding

Terminal-Bench 2.0

69.3%
Multilingual coding

SWE-bench Multilingual

79.8%
Hard agentic coding

CursorBench v3.1 harder tasks

63.2%
Full official benchmark table3 rows with source settings and peer charts
BenchmarkAreaScoreSetting / comparison
Terminal-Bench 2.0Agentic coding69.3%Cursor release-table result; no Composer effort label or detailed harness setting is published.
GPT-5.582.7%
Claude Opus 4.769.4%
Composer 2.569.3%
Composer 261.7%
SWE-bench MultilingualMultilingual coding79.8%Cursor release-table result; no Composer effort label is published.
Claude Opus 4.780.5%
Composer 2.579.8%
GPT-5.577.8%
Composer 273.7%
CursorBench v3.1 harder tasksHard agentic coding63.2%Cursor's own harder-task benchmark; no Composer effort label is published.Cursor marks the comparison figures for Opus 4.7 and GPT-5.5 as self-reported.
Claude Opus 4.7 max64.8%
GPT-5.5 xhigh64.3%
Composer 2.563.2%
Composer 252.2%