13 have benchmark runs
S
A
B
C
D
Updated
Evidence and tools
Check the benchmarks.
Compare independent scores, API costs, and the same prompts built by different models.
Updated
Evidence and tools
Compare independent scores, API costs, and the same prompts built by different models.