Reference / 02
Who's actually
smarter.
Arena snapshots Aug 6–15, 2026 · 6 leaderboards
Blind human preference votes across six arena boards — Overall, WebDev, Coding, Hard Prompts, Math and Vision — side by side with the benchmarks vendors love to quote: SWE-bench Pro, Terminal-Bench, GPQA, HLE. Every number links to its source; vendor-run scores are flagged.
The leaderboard
pick a slice · sort by any column · open a row for detailsModels ranked
08
▲ of 13 tracked, top-20 cut
Top score
1506
▲ Claude Fable 5
Votes
7.78M
▲ 392 models on full board
Snapshot
Aug 13, 2026
▲ arena.ai leaderboard
— not measured / not published · * score has a caveat — see footnotes below · P preliminary rating (low vote count) · Value * = arena score per $1 of blended price (3:1 input/output mix) · sort by any column · open a row for full arena & benchmark data.
Caveats *
- Qwen3.8-Max: every benchmark score is vendor-run (Alibaba official release table) — no independent measurements yet. WebDev arena rating is preliminary.
- DeepSeek V4 Pro: benchmarks are Artificial Analysis data via a secondary source — the official model card was not directly read. Listed prices are off-peak; peak windows bill 2×.
- Muse Spark 1.2: Terminal-Bench 2.1 and DeepSWE are Meta-reported; an independent Vals AI run scored 14/50 on a common scaffold — large discrepancy, treat vendor numbers with care.
- Mistral Medium 3.5: both benchmark scores come from aggregators, unverified against an official model card.
- GPT-5.6 Sol: GPQA, HLE and SWE-bench Pro figures come from Alibaba's vendor-run Qwen3.8-Max release table (cross-vendor); METR flagged high reward-hacking on the Terminal-Bench run.
- Claude Opus 5 Terminal-Bench 2.1: figure from Meta's vendor table — unverified for Anthropic.
Off the boards
- GLM-5.3 released Aug 14, 2026 — too fresh for any arena board; open weights promised ~2 weeks post-launch after a safety review. No verified benchmarks published yet — GLM-5.2 anchors (vendor claims): SWE-bench Verified 84.2, SWE-bench Pro 62.1, GPQA 91.2, Terminal-Bench 2.1 81.0–82.7 (harness-dependent).
- Mistral Medium 3.5 not in the top-20 of any tracked arena slice.
- Kimi K3 has no Vision board entry.
- GPT-5.6 Sol is absent from Vision and Math.
- Grok 4.6 appears only on WebDev (preliminary).
- Gemini 3.1 Pro skips WebDev and Coding.
- Gemini 3.6 Flash skips WebDev and Coding.
Field test · 2026-09-03
RomaBenchmark
Six frontier models behind one gateway — CEHWA (词花) — OpenAI-compatible proxy (https://router.cehwa.com/v1) — run through a gauntlet of short but nasty tasks. Every answer verified against script-computed ground truth; every coding solution executed, not eyeballed.
| Model | R1 | R2 | R3a | R3b | T1 | T2 | T3 | Score |
|---|---|---|---|---|---|---|---|---|
| MiniMax-M3 | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 8/8 |
| Kimi-K3 | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 8/8 |
| GLM-5.3 | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 8/8 |
| DeepSeek-V4-Pro | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 8/8 |
| DeepSeek-V4-Flash | ✓ | ✓ | ✓ | ✓ | — | — | — | 4/4 math |
| GLM-5.2 | ✓ | ✓ | ✗ | ✗ | ✓ | ✓ | ✓ | 6/8 |
Tasks & ground truth
- R1 Smallest n with n≡1 (mod 2,3,4,5,6) and n≡0 (mod 7)301 — naive LCM+1 gives 61, which is not divisible by 7
- R2 Trailing zeros of 2026! in base 7335 — requires Legendre's formula, not the base-10 shortcut
- R3a Count integers 1..1000 divisible by 3 or 5 but not by 7401 — double inclusion–exclusion
- R3b Count primes in [2000, 2100]14 — error-prone manual sieving of 101 candidates
- T1 Predict output: x=[[]]*3; x[0].append(1); lambdas closing over the loop variable[[1],[1],[1]] [2,2,2] — list aliasing + late binding
- T2 Predict output: print(0.1+0.2 == 0.3, round(2.675, 2))False 2.67 — binary float representation
- T3 Write calc(s): arithmetic parser with + - * /, parentheses, unary minus — no eval/execPass = 32/32 automated tests
Latency & tokens (single run, via gateway)
| Model | R3 math | R4 code |
|---|---|---|
| MiniMax-M3 | 16.1s · 1341 tok | 10.0s · 1633 tok |
| Kimi-K3 | 46.4s · 7023 tok | 102.4s · 4419 tok |
| GLM-5.3 | 68.5s · 4789 tok | 18.2s · 2057 tok |
| DeepSeek-V4-Pro | 99.5s · 7233 tok | 43.4s · 2269 tok |
| DeepSeek-V4-Flash | 75.6s · 8640 tok | — |
| GLM-5.2 | 4.1s · 653 tok | 62.9s · 3352 tok |
- MiniMax-M3. Best default. Flawless on both tracks while being the fastest and the most token-efficient by a wide margin.
- Kimi-K3. Best presentation — tabulated, fully checkable solutions. The slowest and most verbose of the pack.
- GLM-5.3. Solid all-rounder; the only model that explained why the float traps trigger (2.675 is stored as 2.67499…).
- DeepSeek-V4-Pro. Deepest work shown — full composite sieve with correct factorizations (2021=43·47, 2023=7·17²). Slowest math round; its gateway channel died right after.
- DeepSeek-V4-Flash. Math track flawless, but the gateway channel died before the coding round and stayed down — treat as unstable.
- GLM-5.2. Code-capable (32/32) but unreliable at multi-step arithmetic: answered in 4 s with confident hallucinations. Do not trust it for exact counting.
Where models broke
- GLM-5.2 r3a: Answered 380 — derived 467 and 66 correctly, then wrote 467−66=380
- GLM-5.2 r3b: Answered 3 — hallucinated that all but three candidates are composite
⚠ Gateway stability
- 2026-09-03: both DeepSeek channels (Flash, Pro) returned “no available channel” for 5+ minutes mid-benchmark. The gateway is fine for experiments — never depend on a single model without a fallback.
Methodology
- Six frontier models behind one gateway, identical prompts, one run per model per task.
- Ground truth for every task was computed independently by script before querying the models.
- Coding task T3: the code from each answer was extracted and executed against 32 automated tests (25 randomized expressions diffed against the reference evaluator).
- Latency includes gateway and network overhead — indicative only, not a precise speed metric.
- Trial key window: 2026-09-02 → 2026-09-09; measured spend during the whole benchmark: under $0.09.