Skip to content

Reference / 02

Who's actually
smarter.

Arena snapshots Aug 6–15, 2026 · 6 leaderboards

Blind human preference votes across six arena boards — Overall, WebDev, Coding, Hard Prompts, Math and Vision — side by side with the benchmarks vendors love to quote: SWE-bench Pro, Terminal-Bench, GPQA, HLE. Every number links to its source; vendor-run scores are flagged.

The leaderboard

pick a slice · sort by any column · open a row for details

Models ranked

08

▲ of 13 tracked, top-20 cut

Top score

1506

▲ Claude Fable 5

Votes

7.78M

▲ 392 models on full board

Snapshot

Aug 13, 2026

▲ arena.ai leaderboard

— not measured / not published · * score has a caveat — see footnotes below · P preliminary rating (low vote count) · Value * = arena score per $1 of blended price (3:1 input/output mix) · sort by any column · open a row for full arena & benchmark data.

Arena

  • Text Overall1506±5

    rank #1 · 21,533 votes · claude-fable-5

  • WebDev1627±8

    rank #6 · 7,867 votes · claude-fable-5

  • Coding1554±9

    rank #1 · 5,817 votes · claude-fable-5

  • Hard Prompts1532±6

    rank #2 · 14,125 votes · claude-fable-5

  • Math1527±18

    rank #2 · 1,055 votes · claude-fable-5

  • Vision1315±9

    rank #1 · 8,333 votes · claude-fable-5

Snapshot Aug 13, 2026 · 7.78M votes · 392 models on board · leaderboard ↗

Benchmarks

  • SWE-bench Pro80.0Source ↗

    tested Jun 9, 2026

    vendor-reported

  • Terminal-Bench 2.184.6Source ↗

    tested Jun 9, 2026

    Terminus-2 harness

  • Humanity's Last Exam53.3Source ↗

    tested Aug 3, 2026

    best among flagships in Alibaba's Aug 2026 table

  • GPQA Diamond92.6Source ↗

    tested Aug 3, 2026

    vendor-run

Also: FrontierSWE 86.6 (Jul 2026, frontierswe.com)

Provider AnthropicContext 1MPrice in/out $10.00/$50.00Compare →Full model page →

Arena

  • Text Overall1498±10

    rank #4 · 3,280 votes · muse-spark-1.2 (xHigh)

  • Coding1533±20

    rank #8 · 947 votes · muse-spark-1.2 (xHigh)

  • Hard Prompts1511±13

    rank #11 · 2,148 votes · muse-spark-1.2 (xHigh)

  • Vision1290±18

    rank #10 · 1,282 votes · muse-spark-1.2 (xHigh)

Snapshot Aug 13, 2026 · 7.78M votes · 392 models on board · leaderboard ↗

Benchmarks

  • Terminal-Bench 2.182.9Source ↗

    tested Aug 5, 2026

    Meta-reported; independent Vals AI run scored 14/50 on a common scaffold — large discrepancy

Also: DeepSWE v1.1 59.3 (Aug 2026, Meta-reported)

Provider MetaContext 1.05MPrice in/out $1.25/$4.25Compare →Full model page →

Arena

  • Text Overall1493±5

    rank #7 · 20,030 votes · claude-opus-5-high

    max variant: 1489, rank 10

  • WebDev1692±9

    rank #1 · 6,448 votes · claude-opus-5-max

    high variant: 1663, rank 4

  • Coding1531±9

    rank #9 · 5,433 votes · claude-opus-5-high

    max variant: 1530, rank 12

  • Hard Prompts1520±6

    rank #5 · 13,390 votes · claude-opus-5-high

    max variant: 1516, rank 9

  • Math1553±28

    rank #1 · 437 votes · claude-opus-5-max

    high variant: 1524, rank 3; small sample

  • Vision1297±11

    rank #6 · 3,821 votes · claude-opus-5-high

Snapshot Aug 13, 2026 · 7.78M votes · 392 models on board · leaderboard ↗

Benchmarks

  • Terminal-Bench 2.186.7Source ↗

    tested Aug 5, 2026

    UNVERIFIED for Anthropic — figure from Meta's vendor table

Provider AnthropicContext 1MPrice in/out $5.00/$25.00Compare →Full model page →

Arena

  • Text Overall1491±8

    rank #8 · 7,004 votes · qwen3.8-max

  • WebDev1667±13P

    rank #3 · 2,855 votes · qwen3.8-max

  • Coding1529±13

    rank #13 · 2,146 votes · qwen3.8-max

  • Hard Prompts1516±9

    rank #8 · 4,807 votes · qwen3.8-max

  • Math1513±31

    rank #6 · 353 votes · qwen3.8-max

    small sample

  • Vision1301±9

    rank #2 · 5,344 votes · qwen3.8-max

Snapshot Aug 13, 2026 · 7.78M votes · 392 models on board · leaderboard ↗

Benchmarks

  • Terminal-Bench 2.186.6Source ↗

    tested Aug 3, 2026

    vendor-run (Alibaba official release table); no independent measurements yet

  • SWE-bench Pro67.7Source ↗

    tested Aug 3, 2026

    vendor-run (Alibaba official release table); no independent measurements yet

  • GPQA Diamond92.6Source ↗

    tested Aug 3, 2026

    vendor-run (Alibaba official release table); no independent measurements yet

  • Humanity's Last Exam43.6Source ↗

    tested Aug 3, 2026

    vendor-run (Alibaba official release table); no independent measurements yet

Also: OSWorld-Verified 86.1 (Aug 2026, vendor-run, Alibaba official release table)

Provider Alibaba (Qwen)Context 1MPrice in/out $2.00/$6.00Compare →Full model page →

Arena

  • Text Overall1489±6

    rank #12 · 11,969 votes · kimi-k3-max

  • WebDev1674±11

    rank #2 · 4,546 votes · kimi-k3-max

  • Coding1544±11

    rank #6 · 3,176 votes · kimi-k3-max

  • Hard Prompts1516±7

    rank #7 · 7,835 votes · kimi-k3-max

  • Math1493±26

    rank #13 · 515 votes · kimi-k3-max

Snapshot Aug 13, 2026 · 7.78M votes · 392 models on board · leaderboard ↗

Benchmarks

  • Terminal-Bench 2.188.3Source ↗

    tested Jul 16, 2026

    Kimi Code harness, max effort; official tech blog

  • SWE-bench Verified93.4Source ↗

    tested Aug 1, 2026

    independent measurement by Vals AI

Also: DeepSWE v1.1 67.5 (Jul 2026, vendor, Kimi Code harness); Program Bench 77.8 (Jul 2026, vendor); BrowseComp 91.2 (Jul 2026, vendor)

Provider Moonshot AIContext 1.05MPrice in/out $3.00/$15.00Compare →Full model page →

Arena

  • Text Overall1486±3

    rank #14 · 95,107 votes · gemini-3.1-pro-preview

  • Hard Prompts1507±4

    rank #14 · 61,475 votes · gemini-3.1-pro-preview

  • Math1491±9

    rank #16 · 5,072 votes · gemini-3.1-pro-preview

Snapshot Aug 13, 2026 · 7.78M votes · 392 models on board · leaderboard ↗

Benchmarks

  • SWE-bench Pro54.2Source ↗

    tested Jul 21, 2026

    official Google model card

  • Terminal-Bench 2.173.8Source ↗

    tested Jul 21, 2026

    official Google model card

  • SWE-bench Verified80.6Source ↗

    tested Aug 1, 2026

    secondary source, unverified; early preview eval was 72.1

Provider Google DeepMindContext 2MPrice in/out $2.00/$12.00Compare →Full model page →

Arena

  • Text Overall1484±6

    rank #16 · 13,815 votes · gemini-3.6-flash-high

  • Hard Prompts1504±7

    rank #18 · 9,406 votes · gemini-3.6-flash-high

  • Math1512±23

    rank #7 · 667 votes · gemini-3.6-flash-high

  • Vision1295±18

    rank #7 · 1,205 votes · gemini-3.6-flash-high

Snapshot Aug 13, 2026 · 7.78M votes · 392 models on board · leaderboard ↗

Benchmarks

  • SWE-bench Pro58.7Source ↗

    tested Jul 21, 2026

    official Google model card

  • Terminal-Bench 2.178.0Source ↗

    tested Jul 21, 2026

    official Google model card

Also: OSWorld-Verified 83.0 (Jul 2026, official Google model card)

Provider Google DeepMindContext 1MPrice in/out $1.50/$7.50Compare →Full model page →

Arena

  • Text Overall1481±6

    rank #19 · 15,558 votes · gpt-5.6-sol-xhigh

  • WebDev1622±8

    rank #7 · 8,595 votes · gpt-5.6-sol-xhigh (codex-harness)

  • Coding1526±9

    rank #15 · 4,436 votes · gpt-5.6-sol-xhigh

  • Hard Prompts1504±7

    rank #16 · 10,306 votes · gpt-5.6-sol-xhigh

Snapshot Aug 13, 2026 · 7.78M votes · 392 models on board · leaderboard ↗

Benchmarks

  • Terminal-Bench 2.188.8Source ↗

    tested Jul 9, 2026

    base, Codex harness; Ultra config: 91.9; vendor-reported, METR flagged high reward-hacking

  • GPQA Diamond94.1Source ↗

    tested Aug 3, 2026

    from Alibaba's official Qwen3.8-Max release table (cross-vendor)

  • Humanity's Last Exam47.2Source ↗

    tested Aug 3, 2026

    from Alibaba's official Qwen3.8-Max release table (cross-vendor)

  • SWE-bench Pro64.6Source ↗

    tested Aug 3, 2026

    OpenAI has not published SWE-bench Pro; figure from Alibaba's vendor-run table

Provider OpenAIContext 1.05MPrice in/out $5.00/$30.00Compare →Full model page →

Caveats *

  • Qwen3.8-Max: every benchmark score is vendor-run (Alibaba official release table) — no independent measurements yet. WebDev arena rating is preliminary.
  • DeepSeek V4 Pro: benchmarks are Artificial Analysis data via a secondary source — the official model card was not directly read. Listed prices are off-peak; peak windows bill 2×.
  • Muse Spark 1.2: Terminal-Bench 2.1 and DeepSWE are Meta-reported; an independent Vals AI run scored 14/50 on a common scaffold — large discrepancy, treat vendor numbers with care.
  • Mistral Medium 3.5: both benchmark scores come from aggregators, unverified against an official model card.
  • GPT-5.6 Sol: GPQA, HLE and SWE-bench Pro figures come from Alibaba's vendor-run Qwen3.8-Max release table (cross-vendor); METR flagged high reward-hacking on the Terminal-Bench run.
  • Claude Opus 5 Terminal-Bench 2.1: figure from Meta's vendor table — unverified for Anthropic.

Off the boards

  • GLM-5.3 released Aug 14, 2026 — too fresh for any arena board; open weights promised ~2 weeks post-launch after a safety review. No verified benchmarks published yet — GLM-5.2 anchors (vendor claims): SWE-bench Verified 84.2, SWE-bench Pro 62.1, GPQA 91.2, Terminal-Bench 2.1 81.0–82.7 (harness-dependent).
  • Mistral Medium 3.5 not in the top-20 of any tracked arena slice.
  • Kimi K3 has no Vision board entry.
  • GPT-5.6 Sol is absent from Vision and Math.
  • Grok 4.6 appears only on WebDev (preliminary).
  • Gemini 3.1 Pro skips WebDev and Coding.
  • Gemini 3.6 Flash skips WebDev and Coding.

Field test · 2026-09-03

RomaBenchmark

Six frontier models behind one gateway — CEHWA (词花) — OpenAI-compatible proxy (https://router.cehwa.com/v1) — run through a gauntlet of short but nasty tasks. Every answer verified against script-computed ground truth; every coding solution executed, not eyeballed.

ModelR1R2R3aR3bT1T2T3Score
MiniMax-M3✓✓✓✓✓✓✓8/8
Kimi-K3✓✓✓✓✓✓✓8/8
GLM-5.3✓✓✓✓✓✓✓8/8
DeepSeek-V4-Pro✓✓✓✓✓✓✓8/8
DeepSeek-V4-Flash✓✓✓✓———4/4 math
GLM-5.2✓✓✗✗✓✓✓6/8

Tasks & ground truth

  • R1 Smallest n with n≡1 (mod 2,3,4,5,6) and n≡0 (mod 7)
    301 — naive LCM+1 gives 61, which is not divisible by 7
  • R2 Trailing zeros of 2026! in base 7
    335 — requires Legendre's formula, not the base-10 shortcut
  • R3a Count integers 1..1000 divisible by 3 or 5 but not by 7
    401 — double inclusion–exclusion
  • R3b Count primes in [2000, 2100]
    14 — error-prone manual sieving of 101 candidates
  • T1 Predict output: x=[[]]*3; x[0].append(1); lambdas closing over the loop variable
    [[1],[1],[1]] [2,2,2] — list aliasing + late binding
  • T2 Predict output: print(0.1+0.2 == 0.3, round(2.675, 2))
    False 2.67 — binary float representation
  • T3 Write calc(s): arithmetic parser with + - * /, parentheses, unary minus — no eval/exec
    Pass = 32/32 automated tests

Latency & tokens (single run, via gateway)

ModelR3 mathR4 code
MiniMax-M316.1s · 1341 tok10.0s · 1633 tok
Kimi-K346.4s · 7023 tok102.4s · 4419 tok
GLM-5.368.5s · 4789 tok18.2s · 2057 tok
DeepSeek-V4-Pro99.5s · 7233 tok43.4s · 2269 tok
DeepSeek-V4-Flash75.6s · 8640 tok—
GLM-5.24.1s · 653 tok62.9s · 3352 tok
  • MiniMax-M3. Best default. Flawless on both tracks while being the fastest and the most token-efficient by a wide margin.
  • Kimi-K3. Best presentation — tabulated, fully checkable solutions. The slowest and most verbose of the pack.
  • GLM-5.3. Solid all-rounder; the only model that explained why the float traps trigger (2.675 is stored as 2.67499…).
  • DeepSeek-V4-Pro. Deepest work shown — full composite sieve with correct factorizations (2021=43·47, 2023=7·17²). Slowest math round; its gateway channel died right after.
  • DeepSeek-V4-Flash. Math track flawless, but the gateway channel died before the coding round and stayed down — treat as unstable.
  • GLM-5.2. Code-capable (32/32) but unreliable at multi-step arithmetic: answered in 4 s with confident hallucinations. Do not trust it for exact counting.

Where models broke

  • GLM-5.2 r3a: Answered 380 — derived 467 and 66 correctly, then wrote 467−66=380
  • GLM-5.2 r3b: Answered 3 — hallucinated that all but three candidates are composite

⚠ Gateway stability

  • 2026-09-03: both DeepSeek channels (Flash, Pro) returned “no available channel” for 5+ minutes mid-benchmark. The gateway is fine for experiments — never depend on a single model without a fallback.

Methodology

  • Six frontier models behind one gateway, identical prompts, one run per model per task.
  • Ground truth for every task was computed independently by script before querying the models.
  • Coding task T3: the code from each answer was extracted and executed against 32 automated tests (25 randomized expressions diffed against the reference evaluator).
  • Latency includes gateway and network overhead — indicative only, not a precise speed metric.
  • Trial key window: 2026-09-02 → 2026-09-09; measured spend during the whole benchmark: under $0.09.
Snapshots Aug 6–15, 2026. Every score links to its official source on the model page; vendor-run figures are flagged *. How we collect and verify the data:Methodology & sources →
Benchmarks don't pay the bills — prices do. Compare what these models cost per 1M tokens.Browse API pricing →
LLM Benchmarks & Arena Ratings: Claude Fable 5 tops Text Arena at 1506, Claude Opus 5 leads WebDev at 1692 — no-way.dev