Independent measurements lead. Vendor results stay in separate, source-linked lanes.
Independent comparison
One harness across 9 models.
Artificial Analysis Index v4.1 combines nine evaluations across coding, agentic work, science, knowledge, physics, and long-context reasoning. Higher is better.
- Grok 4.6High61AA source ↗
- Claude Fable 5Max, Opus 4.8 fallback60AA source ↗
- Claude Opus 5Adaptive reasoning, xhigh60AA source ↗
- GPT-5.6 SolMax59AA source ↗
- Kimi K3Reasoning57AA source ↗
- 55AA source ↗
- GPT-5.5xhigh55AA source ↗
- Grok 4.5High54AA source ↗
- GPT-5.6 LunaHigh46AA source ↗
Configuration matters. Fable 5 includes its published Opus 4.8 fallback. GPT and Claude scores use the effort level shown beside each model.
Cost explorer
Price your actual workload.
Set the input and output tokens for one task. The explorer applies current API list prices and recalculates every model in the same scenario.
| Model | AA Index | Input / 1M | Output / 1M | Speed | Your task |
|---|---|---|---|---|---|
| Claude Fable 5Anthropic · Max, Opus 4.8 fallback | 60Index | $10.00 | $50.00 | 69.7tokens/sec | AA data ↗Price ↗ |
| Claude Opus 5Anthropic · Adaptive reasoning, xhigh | 60Index | $5.00 | $25.00 | 53.1tokens/sec | AA data ↗Price ↗ |
| GPT-5.6 SolOpenAI · Max | 59Index | $5.00 | $30.00 | 77.1tokens/sec | AA data ↗Price ↗ |
| Kimi K3Moonshot AI · Reasoning | 57Index | $3.00 | $15.00 | 32.0tokens/sec | AA data ↗Price ↗ |
| GPT-5.6 TerraOpenAI · Max | 55Index | $2.50 | $15.00 | 134.5tokens/sec | AA data ↗Price ↗ |
| GPT-5.5OpenAI · xhigh | 55Index | $5.00 | $30.00 | 72.5tokens/sec | AA data ↗Price ↗ |
| Grok 4.5Xai · High | 54Index | $2.00 | $6.00 | 67.1tokens/sec | AA data ↗Price ↗ |
| Grok 4.6Xai · High | 61Index | $2.00 | $6.00 | 65.1tokens/sec | AA data ↗Price ↗ |
| GPT-5.6 LunaOpenAI · High | 46Index | $1.00 | $6.00 | 178.0tokens/sec | AA data ↗Price ↗ |
What is included: standard input and generated output at API prices verified 2026-08-14.
What is not included: cache writes, tool calls, storage, batch discounts, provider markups, or the extra tokens a model may consume to finish the same job.
Blended reference: Artificial Analysis also publishes a 7:2:1 cache-hit/input/output price. The lowest in this set is GPT-5.6 Luna at $0.87 per 1M blended tokens.
Methodology
What this page will not do.
Artificial Analysis Index results come from one independent harness. The exact model configuration remains attached to every score.
Token cost scenarios multiply published input and output prices by the workload you enter. They exclude tools, storage, cache writes, and provider discounts.
A shared-table mark means the source published the compared models in one table. A same-eval mark means the benchmark name matches, but the source or configuration differs.
Missing data is not a low score. Models without a compatible public score remain unranked until a reproducible result is published.
Our Superbash visual runs are a separate, same-prompt evidence layer. Use the live run comparison to inspect build quality directly.