OpenAI
GPT-5.5
A very good delivery model that has been overtaken by cheaper Terra.
Canonical model record
Current identity, limits, and pricing
- Status
- Current
- API model ID
gpt-5.5- Context
- 1.05M
- Max output
- 128K
- API price / 1M tokens
- $5 input · $30 output
Visual prompt runs
Benchmark runs
Open each generated scene, or compare the same prompt across models.

Helm's Deep
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Jabberwock
Dark fantasy encounter testing creature design, forest mood, and narrative staging.

Low Poly World
Stylized island build testing composition, color, and low-poly worldbuilding.

Office Life
Workplace vignette testing everyday scene logic, objects, and believable office detail.

Petri Dish
Microscopic ecosystem testing organic forms, scientific clarity, and cellular detail.

Universe Simulator
Cosmic system testing orbital structure, glowing bodies, scale, and simulation readability.

Vice City
Neon coastal city testing vehicles, architecture, atmosphere, and dense urban layout.

Yingzao Fashi Assembly
Timber assembly scene testing structure, joinery, construction order, and material clarity.
Official benchmark profile
How GPT-5.5 scores beyond our visual tests.
GPT-5.5 rows combine OpenAI’s launch benchmarks with shared GPT-5.6 comparison rows where those make the score directly comparable to newer models.
SWE-bench Pro
58.6%Terminal-Bench 2.1
85.6%OSWorld 2.0
47.5%Full official benchmark table12 rows with source settings and peer charts
| Benchmark | Area | Score | Setting / comparison |
|---|---|---|---|
| SWE-bench Pro | Coding | 58.6% | OpenAI GPT-5.5 launch score; the GPT-5.6 shared comparison table reports 59.4% on its comparison run. |
| Terminal-Bench 2.1 | Agentic coding | 85.6% | Shared GPT-5.6 comparison table. |
| OSWorld 2.0 | Computer use | 47.5% | Shared GPT-5.6 comparison table. |
| OSWorld-Verified | Computer use | 78.7% | OpenAI GPT-5.5 launch score. |
| BrowseComp | Tool use | 84.4% | Shared GPT-5.6 comparison table and GPT-5.5 launch result. |
| BenchCAD | Computer-aided design | 44.4% | Vision2Code score without Python tool from the shared GPT-5.6 table. |
| BenchCAD with Python tool | Tool use | 59.4% | Vision2Code score with Python tool from the shared GPT-5.6 table. |
| GPQA Diamond | Academic reasoning | 93.6% | OpenAI science reasoning result and shared GPT-5.6 table. |
| FrontierMath Tier 1-3 v2 | Math | 85.3% | Shared GPT-5.6 comparison table. |
| FrontierMath Tier 4 v2 | Math | 72.5% | Shared GPT-5.6 comparison table. |
| GDPval | Professional work | 84.9% | OpenAI GPT-5.5 launch score. |
| CyberGym | Cybersecurity | 81.8% | OpenAI GPT-5.5 launch score. |
Superbash commentary
Our take
GPT-5.5 still writes clean code and fits OpenAI tooling well, but the August recording treats it as a relic of the previous generation: Terra is roughly half the API price in the current catalog and is now the team’s routine daily driver.
Best for
Watch out
Why it is ranked here
Evidence and commentary
Superbash editorial model ranking
Takeaway: GPT-5.5 is currently placed in Tier B.
The August 2026 editorial roster places GPT-5.5 at rank 12.
Open source →NEW AI Model Tier List for Vibe Coding!
Takeaway: Use GPT-5.5 in plan mode rather than as a long-horizon coding default.
The source video says it is strong and clean, but overthinks and spends too much context on extended builds.
Open source →Keep learning