Thinking Machines Lab
Inkling
Version 1.0 benchmark runs across the shared Superbash visual prompts.
Canonical model record
Current identity, limits, and pricing
- Status
- Current
- API model ID
- Not publicly verified
- Context
- 1M
- Max output
- Not publicly verified
- API price / 1M tokens
- Not publicly verified
Visual prompt runs
Benchmark runs
Open each generated scene, or compare the same prompt across models.

Helm's Deep
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Jabberwock
Dark fantasy encounter testing creature design, forest mood, and narrative staging.

Low Poly World
Stylized island build testing composition, color, and low-poly worldbuilding.

Office Life
Workplace vignette testing everyday scene logic, objects, and believable office detail.

Petri Dish
Microscopic ecosystem testing organic forms, scientific clarity, and cellular detail.

Universe Simulator
Cosmic system testing orbital structure, glowing bodies, scale, and simulation readability.

Vice City
Neon coastal city testing vehicles, architecture, atmosphere, and dense urban layout.

Yingzao Fashi Assembly
Timber assembly scene testing structure, joinery, construction order, and material clarity.
Official benchmark profile
How Inkling scores beyond our visual tests.
Thinking Machines presents Inkling as a customizable open-weights multimodal model rather than the strongest model in every category. These rows use the current website model card at effort 0.99 and temperature 1.0.
AIME 2026
97.1%GPQA Diamond
87.2%SWE-bench Verified
77.6%Full official benchmark table12 rows with source settings and peer charts
| Benchmark | Area | Score | Setting / comparison |
|---|---|---|---|
| AIME 2026 | Math | 97.1% | Effort 0.99 and temperature 1.0. |
| GPQA Diamond | Academic reasoning | 87.2% | Externally reported score reproduced in the Inkling model card. |
| SWE-bench Verified | Agentic coding | 77.6% | Effort 0.99, temperature 1.0, and a 256K-token coding trajectory cap. |
| SWE-bench Pro Public | Agentic coding | 54.3% | Effort 0.99, temperature 1.0, and a 256K-token coding trajectory cap. |
| Terminal-Bench 2.1 Best Harness | Terminal agents | 63.8% | Effort 0.99 and temperature 1.0; web-search-contaminated rollouts score zero. |
| MCP Atlas | Tool use | 76.0% | Current Thinking Machines website model-card value.The older Hugging Face README reports 74.1%; the website model card is used here. |
| BrowseComp with context management | Agentic research | 77.1% | Model-card run with context management enabled. |
| IFBench | Instruction following | 79.8% | Effort 0.99 and temperature 1.0. |
| Global-MMLU-Lite | Multilingual knowledge | 88.7% | Effort 0.99 and temperature 1.0. |
| MMMU-Pro Standard 10 | Multimodal reasoning | 73.5% | Externally reported score; long image edge resized under the model-card protocol. |
| VoiceBench | Audio | 91.4% | Effort 0.99 and temperature 1.0. |
| FORTRESS Adversarial | Safety | 78.0% | Adversarial refusal evaluation from the model card. |