23 models across 5 tiers
13 have benchmark runs

S

A

B

C

D

Updated

Evidence and tools

Check the benchmarks.

Compare independent scores, API costs, and the same prompts built by different models.