The whole leaderboard as a map: every one of the 1198 distinct post-training runs (25 agents × 7 benchmarks; 1226 total minus 28 old-container reruns) is a dot in the model × task grid — shaded by official score, red when the usage judge flagged it, and ringed + clickable when we hand-verified it into a step-by-step replay. 45 deep replays so far; 52 runs judge-flagged.
▨ light-yellow cells have a clickable step-by-step replay — click the ◎ ringed dot to open that agent's trajectory on that task. ● red = usage judge flagged · ● shaded = official score (darker = higher) · ● gray = ≈0 / no lift. Cell label = run count × best score. Un-ringed colors are the judge's signal (which under-counts hacks); only ringed dots are verified.
| model \ task | BFCL | HealthBench | Arena-Hard | AIME 2025 | GSM8K | GPQA | HumanEval |
|---|---|---|---|---|---|---|---|
| fable-5 | 4× 1.00 | 4× 0.46 | 4× 0.86▸ replay | 4× 0.20 | 4× 0.91 | 2× — | 4× 0.78 |
| opus-4.5 | 3× 0.92 | 4× 0.22 | 4× 0.16 | 4× 0.00 | 4× 0.62 | 4× 0.26 | 4× 0.50▸ replay |
| opus-4.6 | 12× 1.00▸ replay | 12× 0.30▸ replay | 12× 0.21▸ replay | 12× 0.20 | 12× 0.83 | 12× 0.32 | 12× 0.71▸ replay |
| opus-4.6 · 1m | 12× 1.00▸ replay | 12× 0.23▸ replay | 12× 0.14 | 12× 0.13 | 12× 0.85 | 12× 0.33 | 12× 0.74 |
| opus-4.7 | 12× 0.95▸ replay | 12× 0.31 | 12× 0.76 | 12× 0.17 | 12× 0.82 | 12× 0.33 | 12× 0.71 |
| opus-4.8 | 8× 1.00▸ replay | 8× 0.48 | 8× 0.75 | 8× 0.23 | 8× 0.87 | 8× 0.35 | 8× 0.77▸ replay |
| opus-4.8 · max | 8× 1.00▸ replay | 8× 0.49▸ replay | 8× 0.74 | 8× 0.30▸ replay | 8× 0.91 | 8× 0.33 | 8× 0.76 |
| sonnet-4.6 | 4× 1.00▸ replay | 4× 0.20 | 4× 0.17 | 4× 0.07 | 4× 0.42 | 3× 0.18 | 4× 0.53 |
| model \ task | BFCL | HealthBench | Arena-Hard | AIME 2025 | GSM8K | GPQA | HumanEval |
|---|---|---|---|---|---|---|---|
| gpt-5.1-codex-max | 4× 0.00▸ replay | 4× 0.00 | 4× 0.00 | 4× 0.03 | 4× 0.11 | 4× 0.25 | 4× 0.11 |
| gpt-5.3-codex | 12× 1.00▸ replay | 12× 0.18 | 12× 0.03 | 12× 0.03 | 12× 0.58 | 12× 0.30 | 12× 0.41 |
| gpt-5.3-codex · high | 12× 0.98▸ replay | 12× 0.21 | 12× 0.11 | 12× 0.03 | 12× 0.59 | 12× 0.34 | 12× 0.43 |
| gpt-5.4 · high | 12× 0.96▸ replay | 12× 0.33 | 12× 0.27 | 12× 0.07 | 12× 0.68 | 12× 0.34 | 12× 0.39 |
| gpt-5.4 · high·rp | 4× 1.00▸ replay | 4× 0.32 | 4× 0.50▸ replay | 4× 0.07 | 4× 0.83 | 4× 0.34 | 4× 0.66 |
| gpt-5.5 · xhigh | 8× 0.99▸ replay | 8× 0.33 | 8× 0.28 | 8× 0.10 | 8× 0.80 | 8× 0.37 | 8× 0.72 |
| gpt-5.5 · xhigh·rp | 4× 1.00 | 4× 0.34 | 4× 0.28 | 4× 0.10 | 4× 0.79 | 4× 0.34 | 4× 0.61 |
| model \ task | BFCL | HealthBench | Arena-Hard | AIME 2025 | GSM8K | GPQA | HumanEval |
|---|---|---|---|---|---|---|---|
| qwen3-max | 4× 0.00 | 4× — | 4× 0.02▸ replay | 4× 0.00 | 4× 0.43 | 4× 0.08 | 4× 0.46 |
Source: PostTrainBench raw trajectories — 1198 runs / 25 agents / 7 benchmarks. Matrix cells are mechanical (official score + contamination/disallowed-model judge from meta.csv). The 45 ringed runs are full forensic step-by-step replays with verbatim, line-anchored evidence.