MiniMax
MiniMax M3
A former daily burner that has fallen behind the August field.
Canonical model record
Current identity, limits, and pricing
- Status
- Current
- API model ID
MiniMax-M3- Context
- 1M
- Max output
- Not publicly verified
- API price / 1M tokens
- $0.3 input · $1.2 output
Rates shown apply through 512K input tokens; longer prompts use the provider's higher tier.
Suggested for this guide
MiniMax
Use the same MiniMax family tested on this page. Check the current plan and model access before subscribing.
Best for: Long-context work and cost-efficient execution
Check MiniMax plansPartner link. It supports Superbash Learn at no extra cost to you.Visual prompt runs
Benchmark runs
Open each generated scene, or compare the same prompt across models.

Helm's Deep
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Jabberwock
Dark fantasy encounter testing creature design, forest mood, and narrative staging.

Low Poly World
Stylized island build testing composition, color, and low-poly worldbuilding.

Mechanical Watch Simulator
Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Office Life
Workplace vignette testing everyday scene logic, objects, and believable office detail.

Petri Dish
Microscopic ecosystem testing organic forms, scientific clarity, and cellular detail.

Stormwind Trebuchet Simulator
Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Universe Simulator
Cosmic system testing orbital structure, glowing bodies, scale, and simulation readability.

Vice City
Neon coastal city testing vehicles, architecture, atmosphere, and dense urban layout.

Yingzao Fashi Assembly
Timber assembly scene testing structure, joinery, construction order, and material clarity.
Official benchmark profile
How MiniMax M3 scores beyond our visual tests.
MiniMax publishes M3 results across coding, browser work, tool use, spreadsheets, and computer control. Its strongest relative result in the official comparison chart is SVG-Bench; terminal and OS-control scores remain below the largest closed models.
SWE-bench Pro
59.0%Terminal-Bench 2.1
66.0%VIBE V2
50.1%Full official benchmark table11 rows with source settings and peer charts
| Benchmark | Area | Score | Setting / comparison |
|---|---|---|---|
| SWE-bench Pro | Agentic coding | 59.0% | MiniMax infrastructure with Claude Code scaffolding and official-aligned evaluation logic. |
| Terminal-Bench 2.1 | Terminal agents | 66.0% | 8C16G sandbox, two-hour timeout, 128K output cap, Terminus 2. |
| VIBE V2 | Full-stack coding | 50.1% | Internal build-from-scratch benchmark with Claude Code and a three-run average. |
| SVG-Bench | Visual coding | 63.7% | Internal text/image build and edit tasks with VLM render verification; three-run average. |
| KernelBench Hard | GPU kernels | 28.8% | Claude Code on NVIDIA Blackwell sm_120; average submitted TFLOPs over theoretical peak across nine questions. |
| BrowseComp | Web research | 83.5% | WebExplorer agent framework with history discarded above 64K tokens. |
| GDPval rubrics | Professional work | 74.7% | Public GDPval cases and rubrics with pointwise scoring aligned to GDPval-AA. |
| BankerToolBench | Finance tools | 76.1% | Public dataset; Claude Code except GPT uses Codex; MiniMax M2.7 judge. |
| MCP Atlas | Tool use | 74.2% | Official public set and codebase with Gemini 2.5 Pro as judge. |
| OSWorld-Verified | Computer use | 75.2% | 361 samples, official codebase, 1920x1080, relative coordinates, and a 200-step cap. |
| SWE-fficiency | Coding efficiency | 34.8% | MiniMax open-source dataset and workflow with a two-hour timeout in Claude Code. |

Superbash commentary
Our take
MiniMax M3 now sits in C tier. It was one of the team’s favorite cheap daily burners only a few months ago, but newer models widened the gap: M3 now needs more bug checking, can confidently claim work is fixed when it is not, and is neither as cheap as DeepSeek nor as capable as Kimi.
Best for
Watch out
Why it is ranked here
Evidence and commentary
Superbash editorial model ranking
Takeaway: MiniMax M3 is currently placed in Tier C.
The August 2026 editorial roster places MiniMax M3 at rank 14.
Open source →NEW AI Model Tier List for Vibe Coding!
Takeaway: MiniMax M3 needs a strong harness and supervision; the August roster no longer treats it as a current workhorse.
The older source records why the team adopted M3, while the new recording explains why newer peers have overtaken it.
Open source →Keep learning