Thinking Machines Lab

Inkling

Version 1.0 benchmark runs across the shared Superbash visual prompts.

Canonical model record

Current identity, limits, and pricing

Provider source · checked 2026-08-14 ↗
Status
Current
API model ID
Not publicly verified
Context
1M
Max output
Not publicly verified
API price / 1M tokens
Not publicly verified

Visual prompt runs

Benchmark runs

Open each generated scene, or compare the same prompt across models.

9 runs

Official benchmark profile

How Inkling scores beyond our visual tests.

Thinking Machines presents Inkling as a customizable open-weights multimodal model rather than the strongest model in every category. These rows use the current website model card at effort 0.99 and temperature 1.0.

Thinking Machines sourceJuly 2026Source report →
Math

AIME 2026

97.1%
Academic reasoning

GPQA Diamond

87.2%
Agentic coding

SWE-bench Verified

77.6%
Full official benchmark table12 rows with source settings and peer charts
BenchmarkAreaScoreSetting / comparison
AIME 2026Math97.1%Effort 0.99 and temperature 1.0.
Claude Fable 599.9%
GPT-5.6 Sol99.9%
GLM 5.299.2%
Inkling97.1%
GPQA DiamondAcademic reasoning87.2%Externally reported score reproduced in the Inkling model card.
GPT-5.6 Sol94.1%
Gemini 3.1 Pro94.1%
Claude Fable 592.6%
Kimi K2.691.1%
Inkling87.2%
SWE-bench VerifiedAgentic coding77.6%Effort 0.99, temperature 1.0, and a 256K-token coding trajectory cap.
Claude Fable 595.0%
GPT-5.6 Sol82.2%
DeepSeek V4 Pro80.6%
GLM 5.280.0%
Inkling77.6%
SWE-bench Pro PublicAgentic coding54.3%Effort 0.99, temperature 1.0, and a 256K-token coding trajectory cap.
Claude Fable 580.0%
GPT-5.6 Sol64.6%
GLM 5.262.1%
Inkling54.3%
Terminal-Bench 2.1 Best HarnessTerminal agents63.8%Effort 0.99 and temperature 1.0; web-search-contaminated rollouts score zero.
GPT-5.6 Sol89.5%
Claude Fable 584.6%
GLM 5.282.7%
Inkling63.8%
MCP AtlasTool use76.0%Current Thinking Machines website model-card value.The older Hugging Face README reports 74.1%; the website model card is used here.
Claude Fable 583.3%
GPT-5.6 Sol81.8%
GLM 5.277.8%
Inkling76.0%
BrowseComp with context managementAgentic research77.1%Model-card run with context management enabled.
GPT-5.6 Sol90.8%
Claude Fable 588.0%
Gemini 3.1 Pro85.9%
DeepSeek V4 Pro83.4%
Inkling77.1%
IFBenchInstruction following79.8%Effort 0.99 and temperature 1.0.
Nemotron 3 Ultra81.4%
Inkling79.8%
Gemini 3.1 Pro77.1%
DeepSeek V4 Pro76.5%
Global-MMLU-LiteMultilingual knowledge88.7%Effort 0.99 and temperature 1.0.
Claude Fable 593.3%
Gemini 3.1 Pro92.7%
GPT-5.6 Sol91.8%
GLM 5.289.2%
Inkling88.7%
MMMU-Pro Standard 10Multimodal reasoning73.5%Externally reported score; long image edge resized under the model-card protocol.
Claude Fable 584.2%
GPT-5.6 Sol83.0%
Gemini 3.1 Pro82.0%
Inkling73.5%
VoiceBenchAudio91.4%Effort 0.99 and temperature 1.0.
Gemini 3.1 Pro94.3%
Inkling91.4%
FORTRESS AdversarialSafety78.0%Adversarial refusal evaluation from the model card.
Claude Fable 596.0%
GPT-5.6 Sol82.4%
Inkling78.0%
Nemotron 3 Ultra77.6%