Xai
Grok 4.6
S tier for fast agents; closer to A tier when raw coding intelligence matters most.
Canonical model record
Current identity, limits, and pricing
- Status
- Current
- API model ID
grok-4.6- Context
- 500K
- Max output
- Not publicly verified
- API price / 1M tokens
- $2 input · $6 output
Prompts of 200K tokens or more use the provider's higher long-context rate.
Visual prompt runs
Benchmark runs
Open each generated scene, or compare the same prompt across models.

City Scroll Journey
Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Ember Glider
Sunset gliding journey testing flight energy management, checkpoint flow, and atmospheric scene craft.

Helm's Deep
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Jabberwock
Dark fantasy encounter testing creature design, forest mood, and narrative staging.

Low-Poly Tower Defense
Diorama tower defense testing economy balance, wave design, placement rules, and combat readability.

Low Poly World
Stylized island build testing composition, color, and low-poly worldbuilding.

Mechanical Watch Simulator
Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Neon Drift
Synthwave time-trial racing testing drift physics, lap timing, ghost replay, and unlock progression.

Office Life
Workplace vignette testing everyday scene logic, objects, and believable office detail.

Petri Dish
Microscopic ecosystem testing organic forms, scientific clarity, and cellular detail.

Starfall Arena
Neon arena survival testing wave escalation, upgrade builds, particle feedback, and boss design.

Stormwind Trebuchet Simulator
Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Universe Simulator
Cosmic system testing orbital structure, glowing bodies, scale, and simulation readability.

Vice City
Neon coastal city testing vehicles, architecture, atmosphere, and dense urban layout.

Yingzao Fashi Assembly
Timber assembly scene testing structure, joinery, construction order, and material clarity.
Official benchmark profile
Grok 4.6’s published agentic coding and knowledge-work results.
xAI reports Grok 4.6 on a mix of composite, knowledge-work, and agentic-coding evaluations. The rows below retain xAI’s exact benchmark labels and only compare the named configurations in the same launch table.
AA Intelligence Index
61GDPVal-AA v2
1753CursorBench v3.2
69.9%Full official benchmark table10 rows with source settings and peer charts
| Benchmark | Area | Score | Setting / comparison |
|---|---|---|---|
| AA Intelligence Index | Composite capability | 61 | xAI launch-table result; the index aggregates nine benchmarks. |
| GDPVal-AA v2 | Knowledge work | 1753 | xAI launch-table rating; higher is better. |
| CursorBench v3.2 | Agentic coding | 69.9% | xAI launch-table result for the named CursorBench configuration. |
| DeepSWE v1.1 | Software engineering | 65.9% | xAI launch-table result for its DeepSWE v1.1 setup. |
| FrontierCode v1.1 (Extended) | Long-horizon coding | 61.3% | xAI launch-table result for the Extended variant. |
| Terminal-Bench v3.0 | Command-line agents | 26.0% | xAI launch-table result. This newer benchmark is not interchangeable with Terminal-Bench 2.x results elsewhere on the site. |
| APEX-Agents | Agent reliability | 57.5% | xAI launch-table result for the named APEX-Agents evaluation. |
| APEX-SWE | Software engineering | 56.4% | xAI launch-table result for the named APEX-SWE evaluation. |
| AA-Briefcase | Knowledge work | 1577 | xAI launch-table rating; higher is better. |
| Harvey LAB (Vals) | Legal work | 15.8% | xAI launch-table result for the Harvey LAB Vals evaluation. |
Superbash commentary
Our take
Grok 4.6 takes the third S-tier slot because its speed changes what is practical in agent workflows. The team found it dramatically more useful than previous Grok versions and now runs it in Hermes, while still seeing some mid-task drift and placing its coding intelligence below the very best.
Best for
Watch out
Why it is ranked here
Evidence and commentary
Superbash editorial model ranking
Takeaway: Grok 4.6 is currently placed in Tier S.
The August 2026 editorial roster places Grok 4.6 at rank 3.
Open source →Introducing Grok 4.6
Takeaway: Grok 4.6 is placed in S tier for the current visual-build and agentic-coding evidence set.
SpaceXAI reports the launch table; Superbash keeps the runs and source settings visible rather than flattening them into one score.
Open source →Keep learning