LLM leaderboard

Last updated: 2026-09-25

Agent-oriented picks.

Index ignores price. Default = best Index within 2× cheapest task $ · Cheap = lowest task $ · Frontier = best Index. Rules under Method.

Category 1st 2nd
Default GLM-5.3 Kimi K3
Cheap Gemini 3.8 Flash GPT-6
Frontier GPT-6 Claude Opus 5

* Fewest active benchmarks on this board.

Top 15.

Showing the top 15 of 34 models with ≥4 active direct benchmarks. Expand a row for sources and costs.

# Model Index · $/task Expand
1 GPT-6 84.20 $4.43
Hermes Index
84.20
Coverage
6/6 direct + Agent Arena
Value (Index÷task$)
19.01
LMSYS
1480 (z=+0.26, w·z=+0.029) · as of 2026-09-25
Agent Arena
97.6% (z=+1.10, w·z=+0.183) · as of 2026-09-25
Terminal-Bench 4.0
58.2% (z=+2.23, w·z=+0.372) · as of 2026-09-25
BrowseComp
91.5% (z=+0.95, w·z=+0.158) · as of 2026-09-25
HLE
54.7% (z=+1.47, w·z=+0.244) · as of 2026-09-25
GPQA Diamond
96.3% ±2.6 (z=+1.69, w·z=+0.094) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
74.1% (z=+0.85, w·z=+0.141) · as of 2026-09-25
DeepSWE $/task
$4.43
List $/1M
$10/50
2 Claude Opus 5 83.34 $11.84
Hermes Index
83.34
Coverage
6/6 direct + Agent Arena
Value (Index÷task$)
7.04
LMSYS
1490 (z=+0.74, w·z=+0.082) · as of 2026-09-25
Agent Arena
95.2% (z=+1.01, w·z=+0.169) · as of 2026-09-25
Terminal-Bench 4.0
53.9% (z=+1.96, w·z=+0.328) · as of 2026-09-25
BrowseComp
90.8% (z=+0.87, w·z=+0.144) · as of 2026-09-25
HLE
54.9% (z=+1.50, w·z=+0.249) · as of 2026-09-25
GPQA Diamond
93.7% ±3.4 (z=+0.48, w·z=+0.027) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
73.6% (z=+0.81, w·z=+0.135) · as of 2026-09-25
DeepSWE $/task
$11.84
List $/1M
$5/25
3 Claude Fable 5.1* 82.81 —
Hermes Index
82.81
Index band
82.81–88.22
Coverage
4/6 direct + Agent Arena
Value (Index÷task$)
1.38
LMSYS
1498 (z=+1.13, w·z=+0.125) · as of 2026-09-25
Agent Arena
100.0% (z=+1.18, w·z=+0.197) · as of 2026-09-25
Terminal-Bench 4.0
57.9% (z=+2.21, w·z=+0.369) · as of 2026-09-25
BrowseComp
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
HLE
59.1% (z=+2.18, w·z=+0.364) · as of 2026-09-25
GPQA Diamond
93.7% ±3.4 (z=+0.48, w·z=+0.027) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
DeepSWE $/task
—
List $/1M
$10/50
4 Claude Fable 5 80.83 $13.41
Hermes Index
80.83
Index band
80.83–82.60
Coverage
5/6 direct + Agent Arena
Value (Index÷task$)
6.03
LMSYS
1506 (z=+1.46, w·z=+0.162) · as of 2026-09-25
Agent Arena
90.5% (z=+0.84, w·z=+0.141) · as of 2026-09-25
Terminal-Bench 4.0
44.5% (z=+1.38, w·z=+0.229) · as of 2026-09-25
BrowseComp
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
HLE
55.5% (z=+1.59, w·z=+0.265) · as of 2026-09-25
GPQA Diamond
92.6% ±3.6 (z=-0.05, w·z=-0.003) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
69.9% (z=+0.53, w·z=+0.089) · as of 2026-09-25
DeepSWE $/task
$13.41
List $/1M
$10/50
5 GPT-5.6 Sol 79.79 $8.39
Hermes Index
79.79
Coverage
6/6 direct + Agent Arena
Value (Index÷task$)
9.51
LMSYS
1483 (z=+0.43, w·z=+0.048) · as of 2026-09-25
Agent Arena
85.7% (z=+0.67, w·z=+0.112) · as of 2026-09-25
Terminal-Bench 4.0
37.3% (z=+0.92, w·z=+0.153) · as of 2026-09-25
BrowseComp
92.2% (z=+1.03, w·z=+0.172) · as of 2026-09-25
HLE
49.5% (z=+0.63, w·z=+0.104) · as of 2026-09-25
GPQA Diamond
95.3% ±3.0 (z=+1.20, w·z=+0.067) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
72.7% (z=+0.74, w·z=+0.123) · as of 2026-09-25
DeepSWE $/task
$8.39
List $/1M
$5/30
6 GLM-5.3 77.24 $3.99
Hermes Index
77.24
Index band
77.24–78.29
Coverage
5/6 direct + Agent Arena
Value (Index÷task$)
19.34
LMSYS
1479 (z=+0.23, w·z=+0.026) · as of 2026-09-25
Agent Arena
59.5% (z=-0.25, w·z=-0.042) · as of 2026-09-25
Terminal-Bench 4.0
41.8% (z=+1.21, w·z=+0.201) · as of 2026-09-25
BrowseComp
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
HLE
55.5% (z=+1.59, w·z=+0.265) · as of 2026-09-25
GPQA Diamond
92.6% ±3.6 (z=-0.05, w·z=-0.003) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
69.0% (z=+0.46, w·z=+0.077) · as of 2026-09-25
DeepSWE $/task
$3.99
List $/1M
$1.4/4.4
7 Kimi K3 76.18 $4.65
Hermes Index
76.18
Index band
76.18–77.01
Coverage
5/6 direct + Agent Arena
Value (Index÷task$)
16.37
LMSYS
1485 (z=+0.49, w·z=+0.054) · as of 2026-09-25
Agent Arena
81.0% (z=+0.51, w·z=+0.084) · as of 2026-09-25
Terminal-Bench 4.0
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
BrowseComp
91.2% (z=+0.91, w·z=+0.152) · as of 2026-09-25
HLE
46.9% (z=+0.21, w·z=+0.034) · as of 2026-09-25
GPQA Diamond
93.5% ±3.4 (z=+0.39, w·z=+0.021) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
68.5% (z=+0.43, w·z=+0.071) · as of 2026-09-25
DeepSWE $/task
$4.65
List $/1M
$3/15
8 GPT-5.6 Terra 75.72 $4.95
Hermes Index
75.72
Coverage
6/6 direct + Agent Arena
Value (Index÷task$)
15.31
LMSYS
1466 (z=-0.38, w·z=-0.042) · as of 2026-09-25
Agent Arena
47.6% (z=-0.67, w·z=-0.112) · as of 2026-09-25
Terminal-Bench 4.0
21.5% (z=-0.07, w·z=-0.011) · as of 2026-09-25
BrowseComp
87.5% (z=+0.48, w·z=+0.080) · as of 2026-09-25
HLE
58.7% (z=+2.12, w·z=+0.353) · as of 2026-09-25
GPQA Diamond
93.4% ±3.5 (z=+0.34, w·z=+0.019) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
69.6% (z=+0.51, w·z=+0.085) · as of 2026-09-25
DeepSWE $/task
$4.95
List $/1M
$2/12
9 Gemini 3.8 Flash 75.38 $2.36
Hermes Index
75.38
Index band
75.38–76.06
Coverage
5/6 direct + Agent Arena
Value (Index÷task$)
31.91
LMSYS
1493 (z=+0.87, w·z=+0.097) · as of 2026-09-25
Agent Arena
69.0% (z=+0.08, w·z=+0.014) · as of 2026-09-25
Terminal-Bench 4.0
19.1% (z=-0.22, w·z=-0.036) · as of 2026-09-25
BrowseComp
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
HLE
47.8% (z=+0.36, w·z=+0.059) · as of 2026-09-25
GPQA Diamond
95.3% ±3.0 (z=+1.20, w·z=+0.067) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
73.8% (z=+0.83, w·z=+0.138) · as of 2026-09-25
DeepSWE $/task
$2.36
List $/1M
$0.75/3.75
10 GPT-5.5 73.93 $7.23
Hermes Index
73.93
Index band
73.93–74.32
Coverage
5/6 direct + Agent Arena
Value (Index÷task$)
10.23
LMSYS
1479 (z=+0.22, w·z=+0.024) · as of 2026-09-25
Agent Arena
78.6% (z=+0.42, w·z=+0.070) · as of 2026-09-25
Terminal-Bench 4.0
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
BrowseComp
84.4% (z=+0.12, w·z=+0.020) · as of 2026-09-25
HLE
45.8% (z=+0.03, w·z=+0.004) · as of 2026-09-25
GPQA Diamond
93.5% ±3.4 (z=+0.39, w·z=+0.021) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
67.0% (z=+0.32, w·z=+0.053) · as of 2026-09-25
DeepSWE $/task
$7.23
List $/1M
$5/30
11 Claude Opus 4.8 73.87 $13.22
Hermes Index
73.87
Coverage
6/6 direct + Agent Arena
Value (Index÷task$)
5.59
LMSYS
1477 (z=+0.14, w·z=+0.016) · as of 2026-09-25
Agent Arena
88.1% (z=+0.76, w·z=+0.127) · as of 2026-09-25
Terminal-Bench 4.0
23.6% (z=+0.07, w·z=+0.011) · as of 2026-09-25
BrowseComp
84.3% (z=+0.11, w·z=+0.018) · as of 2026-09-25
HLE
48.7% (z=+0.49, w·z=+0.082) · as of 2026-09-25
GPQA Diamond
92.0% ±3.8 (z=-0.34, w·z=-0.019) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
59.0% (z=-0.29, w·z=-0.048) · as of 2026-09-25
DeepSWE $/task
$13.22
List $/1M
$5/25
12 Muse Spark* 73.48 $3.70
Hermes Index
73.48
Index band
73.48–74.21
Coverage
4/6 direct + Agent Arena
Value (Index÷task$)
19.88
LMSYS
1493 (z=+0.89, w·z=+0.099) · as of 2026-09-25
Agent Arena
71.4% (z=+0.17, w·z=+0.028) · as of 2026-09-25
Terminal-Bench 4.0
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
BrowseComp
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
HLE
48.7% (z=+0.50, w·z=+0.083) · as of 2026-09-25
GPQA Diamond
94.1% ±3.3 (z=+0.67, w·z=+0.037) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
54.9% (z=-0.60, w·z=-0.100) · as of 2026-09-25
DeepSWE $/task
$3.70
List $/1M
$1.25/4.25
13 Claude Sonnet 5 71.60 $26.40
Hermes Index
71.60
Coverage
6/6 direct + Agent Arena
Value (Index÷task$)
2.71
LMSYS
1461 (z=-0.60, w·z=-0.067) · as of 2026-09-25
Agent Arena
83.3% (z=+0.59, w·z=+0.098) · as of 2026-09-25
Terminal-Bench 4.0
12.4% (z=-0.64, w·z=-0.106) · as of 2026-09-25
BrowseComp
84.7% (z=+0.16, w·z=+0.026) · as of 2026-09-25
HLE
48.7% (z=+0.50, w·z=+0.083) · as of 2026-09-25
GPQA Diamond
94.1% ±3.3 (z=+0.67, w·z=+0.037) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
53.8% (z=-0.67, w·z=-0.112) · as of 2026-09-25
DeepSWE $/task
$26.40
List $/1M
$2/10
14 Claude Opus 4.7* 71.20 —
Hermes Index
71.20
Index band
70.41–71.20
Coverage
4/6 direct
Value (Index÷task$)
2.37
LMSYS
1498 (z=+1.11, w·z=+0.123) · as of 2026-09-25
Agent Arena
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
Terminal-Bench 4.0
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
BrowseComp
79.3% (z=-0.47, w·z=-0.079) · as of 2026-09-25
HLE
42.3% (z=-0.54, w·z=-0.089) · as of 2026-09-25
GPQA Diamond
91.4% ±3.9 (z=-0.63, w·z=-0.035) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
DeepSWE $/task
—
List $/1M
$5/25
15 Claude Opus 4.6* 71.09 —
Hermes Index
71.09
Index band
70.17–71.09
Coverage
4/6 direct
Value (Index÷task$)
2.37
LMSYS
1501 (z=+1.24, w·z=+0.138) · as of 2026-09-25
Agent Arena
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
Terminal-Bench 4.0
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
BrowseComp
83.7% (z=+0.04, w·z=+0.007) · as of 2026-09-25
HLE
39.9% (z=-0.92, w·z=-0.153) · as of 2026-09-25
GPQA Diamond
89.6% ±4.2 (z=-1.49, w·z=-0.083) · as of 2026-09-25
SWE-Atlas-QnA
— (absent; imputed z=0)
DeepSWE
— (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
DeepSWE $/task
—
List $/1M
$5/25

* Fewest active benchmarks on this board.

Download board JSON

How this board is built.

Daily deterministic collectors. Meta boards are diagnostics only.

Composite formula

Scheme Agent/system 70% · Reasoning 20% · Chat 10%. AA · DeepSWE 20% each; TB · BrowseComp · HLE 15% each.

Composite score pipeline

Each active source contributes one metric. Percentage benches that arrive as 0–1 fractions are scaled to percent; there is no fixed bound map before ranking. For every source, we compute a robust center and scale on the fixed scored-model set (models with at least four active direct scores): the center is the median, and the scale is max(1.4826×MAD, ε), with ε = 1.0 percentage point for percent benches and ε = 10 Elo for LMSYS. Each model’s per-source z is (metric − center) / scale, then clamped to ±3. The Hermes Index is the weighted mean of those z values under the active weights for this run, mapped as clamp(72 + 10·z̄, 10, 100). It is an index, not a percentage. The primary path treats a missing active source as z = 0 (shrinkage to the board center). The sensitivity path drops missing sources and renormalizes the remaining weights; that peer is shown as a band only when coverage is incomplete. Sources that cover less than 80% of the scored set have their weight capped at 15% before renormalization so sparse boards cannot dominate. Stale or coverage-failed sources are dropped from the active set entirely. SWE-Atlas-QnA is not weighted (Hermes task mix is terminal/agent execution, not codebase QnA). BFCL-class and tau-bench were evaluated on 2026-08-30 against the freshness and frontier-hit floor and are absent from scoring.

Composite source weights by bucket
Source Bucket Nominal Active Effective Floor
Terminal-Bench 4.0Agent/systemAgent/system15%16.7%11.4%Direct
DeepSWEAgent/systemAgent/system20%16.7%17.4%Direct
Agent ArenaAgent/systemAgent/system20%16.7%11.3%Derived
BrowseCompAgent/systemAgent/system15%16.7%20.2%Direct
HLEReasoningReasoning15%16.7%26.9%Direct
LMSYS Text ArenaChatChat10%11.1%9.9%Direct
GPQA DiamondReasoningReasoning5%5.6%2.9%Direct

Active = coverage-capped then renormalized. Effective = share of realized variance of per-source contributions across the scored board.

Agent Arena Agent Arena is a rank-linear percentile, 100×(N−rank)/(N−1), so rank 1 is exactly 100% and the board ends near 0%. Against the scored set that compresses z to roughly ±1.35 by construction — much tighter than free metric scales. It carries weight but does not count toward the ≥4 direct-source floor.

Coverage & missing A dash means the model has no score on that active source. The primary Index uses z = 0 for that cell; the sensitivity band shows weight renormalization and appears only when coverage is incomplete. Imputed cells do not count toward the inclusion floor.

Sources Tool-use board audit 2026-08-30: BFCL v4 (BenchLM) was ≤7d fresh with 13 rows but 0 models in the Hermes frontier seed after name cleaning — fails ≥5 frontier hits. tau-bench (BenchLM) published an empty leaderboard (display-only / outdated tasks). Neither is a primary source. SWE-Atlas-QnA weight is 0 (not Hermes task mix). Terminal-Bench is pinned to release 4.0 (BenchLM mirror of tbench.ai; best harness/effort per model). Scores are not comparable to the prior TB 2.x column; cross-version fallback into one z-pool is forbidden.

CIs Only GPQA currently shows a ±95% Wald interval (n≈198 multiple-choice items). LMSYS is Elo, not a binomial accuracy; Agent Arena is a rank transform; BrowseComp, Terminal-Bench 4.0, DeepSWE, and HLE are publisher point estimates without a stable public n for the same binomial model. Extending CIs needs per-source n or SEs.

MAD ε Scale floor agent_lmsys ε=1, browsecomp ε=1, deepswe ε=1, gpqa_diamond ε=1, hle ε=1, lmsys ε=10, terminalbench ε=1. Z clamped to ±3.0.

Picks Picks use the same top-15 board as the public leaderboard (models with ≥4 active directs, ranked by Hermes Index). The Hermes Index itself is capability-only and never uses price. Price enters only Cheap and Default. List $/1M means input+output USD per 1M tokens. Task $ means DeepSWE average $/task (real multi-step agent spend on that board). Frontier is the best Hermes Index (#1); the runner-up is the closest peer among ranks 2–5 within 2.0 Index points (higher Index then board rank on ties; no price), otherwise #2 by Index rank. Cheap is the lowest DeepSWE $/task inside that top-15 (list $/1M is only a tie-break or fallback when task $ is missing); its runner-up is the highest Index within 2× that task $ (else the next-cheapest by the same key). Default is the highest Hermes Index among models whose task $ is within 2× Cheap's task $ (list $ defines the band only when Cheap lacks task $), excluding Frontier #1 — the best daily driver still near the cheap frontier, not max Index÷$. The Default runner-up is the next-highest Index in that same band (if the band is a singleton, the next-cheapest model). Default may equal Cheap when Cheap also leads the band on Index; the rule never elevates a dominated model just to keep three category names distinct. Aggregate ≠ task proof.

Inclusion rules

  • Hermes-oriented composite for personal agents — not a neutral AGI ranking.
  • ≥4 active direct sources required (Agent Arena does not count toward the floor).
  • List $/1M = input+output USD per 1M tokens. DeepSWE $ = avg $/task at best-effort config.
  • * Fewest active benchmarks on this board.
  • Aggregate ≠ task proof.

Source status

  • LMSYS Text Arena — active, as of 2026-09-25, frontier hits 15
  • HLE — active, as of 2026-09-25, frontier hits 16
  • GPQA Diamond — active, as of 2026-09-25, frontier hits 16
  • Terminal-Bench 4.0 — active, as of 2026-09-25, frontier hits 10
  • BrowseComp — active, as of 2026-09-25, frontier hits 7
  • DeepSWE — active, as of 2026-09-25, frontier hits 14
  • Agent Arena — active, as of 2026-09-25, frontier hits 17