Xai

Grok 4.5

Version 4.5 benchmark runs across the shared Superbash visual prompts.

Canonical model record

Current identity, limits, and pricing

Provider source · checked 2026-08-14 ↗
Status
Current
API model ID
grok-4.5
Context
500K
Max output
Not publicly verified
API price / 1M tokens
$2 input · $6 output

Prompts of 200K tokens or more use the provider's higher long-context rate.

Visual prompt runs

Benchmark runs

Open each generated scene, or compare the same prompt across models.

7 runs

Official benchmark profile

How Grok 4.5 scores beyond our visual tests.

xAI describes Grok 4.5 as a coding, agentic-task, and knowledge-work model. xAI publishes different benchmark rows than OpenAI/Anthropic, so only direct source comparisons are shown.

SpaceXAI sourceJuly 2026Source report →
Coding

DeepSWE 1.0

62.0%
Coding

DeepSWE 1.1

53.0%
Long-horizon coding

SWE Marathon

29.0%
Full official benchmark table6 rows with source settings and peer charts
BenchmarkAreaScoreSetting / comparison
DeepSWE 1.0Coding62.0%xAI reported software-engineering benchmark.
DeepSWE 1.1Coding53.0%xAI reported software-engineering benchmark.
SWE MarathonLong-horizon coding29.0%xAI reported long-horizon coding benchmark.
SWE Bench Pro output tokensEfficiency15,954 avg.Average output tokens on SWE Bench Pro tasks.
Output-token efficiencyEfficiency4.2x fewer than Opus 4.8xAI comparison on SWE Bench Pro tasks.xAI compares Grok 4.5 against Opus 4.8 max on SWE Bench Pro tasks.
Serving speedLatency80 TPSReported generation speed.