Dan Luu’s The benchmarkpocalypse is an argument about public scoreboards. He writes that a large, formerly trustworthy suite can now be won by an unattended loop, and that the same problem applies twice as hard to AI software. Scoreboard saturation is the product reading of that argument: when a public table stops discriminating the decision a reader needs, topping it is not a shipping win. The coding-agent comparison chart already keeps cost, time, and token use beside those named-suite scores for that reason.
The Hraness reading note of Dan Luu’s The benchmarkpocalypse, saved 2026-08-18, is a dated digest. It is not a substitute for the essay. The sibling note on why a coding-agent high score still needs a holdout stays with Luu’s FRE regex-engine experiment. This page stays with saturation: a cheap public-suite win, or a table that no longer splits products, is remaining measurement work.
Scoreboard saturation is a reading, not a trophy
Scoreboard saturation is not a stored field in the Artificial Analysis snapshot. It is a reading of a public table. One form is numeric. Many named configurations sit near the same ceiling, so the lead cell no longer tells you which system to ship. The other form is semantic. The suite is still hard enough that scores have not hit 100, but a “win” no longer stands for the work a product has to do. Luu’s essay is about that second form. A comprehensive public suite can still be climbed in a way that fails later use.
Those two forms can arrive together. They do not have to. A table can still spread across models and harnesses while a comment thread treats one high cell as a product verdict. The useful question is whether the public number still answers the decision in front of the reader. If it does not, citing the cell as a win is the mistake, whether the score is 91 or 30.
What Dan Luu changed about a public suite
People have always advertised unrepresentative microbenchmarks. Luu’s change is the cost of gaming a large suite. He writes: “What's changed is that it used to take a lot of work to game a large benchmark suite, but an LLM and loop can just do it.” CPU vendors once spent skilled engineering time on SPEC-style compiler tricks. He cites Sun improving 179.art by 12x in SPECfp2000. An agent can now search that space by default:
He is explicit that the regex-engine loop is the worked example, not the only target. “Note that while this post has discussed non-AI software, everything said here goes double for AI software.” A public coding-agent scoreboard has the same shape as rebar: a named task set, an automated check, and a visible scoring rule. If those are public, an optimizer can climb them. The scarce work is checking whether the cell still means what the suite’s authors thought it meant.
A scoreboard cell is not a use claim
Luu’s own AI example is not FRE. It is a scoreboard-versus-use gap on named models. “I've seen lots of people drop comments saying that Kimi K3 is Fable (5) level. But every single person I know who's used it has found it to be substantially worse than GPT-5.6 Sol and Fable.” He adds that “the performance on a wide variety of real-world tasks isn't up to the level it is in benchmarks.” That sentence is the saturation claim in product language. A public table can place two systems together while the people who used both still see a gap.
He gives two narrower checks. A friend tried different coding agents on the ICFP 2026 contest problems and saw the same gap. A colleague asked Kimi K3 to scan for vulnerabilities and found approximately a quarter of the issues GPT-5.6 Sol found, with no extra finds and no advantage except cost. Luu says people he knows who use cheaper models for real security work pick other systems, such as GLM-5.2, that score worse on public tables and work better in practice.
| Public scoreboard claim | What Luu reports from use |
|---|---|
| Kimi K3 is Fable 5 level | Everyone he knows who used both found Kimi K3 substantially worse than GPT-5.6 Sol and Fable |
| Benchmark-level coding-agent performance | A friend saw the same gap on ICFP 2026 contest problems |
| Security-eval strength for Kimi K3 | A colleague’s vuln scan found approximately a quarter of the issues GPT-5.6 Sol found, and no extras |
| Cheaper models that win public tables | People he knows picking cheaper security scanners use GLM-5.2, which scores worse and works better in practice |
Those rows are anecdotes with named people and named workloads. They are not a second leaderboard. They are enough to keep “the public cell is high” from being reported as “the product is interchangeable.” A saturated comment thread is still a saturated scoreboard when the only cited evidence is the cell.
A low science peak is the other failure
Why a 30% Terminal-Bench-Science score is not a product win answers the low-score version of the same product question. Scientists, not vendors, set that bar. The leading named configuration still fails most of the accepted workflows. Cost and token Pareto is the useful comparison there. This page asks the high-score version. A public suite that is cheap to win, or a table that no longer splits products, is also not a shipping decision.
The two notes share a refusal. They do not share a task set. Terminal-Bench-Science 0.1 is a scientist-set science suite. Luu’s examples are a regex-engine public suite and later model-versus-use checks. Collapsing them into one “benchmarks are fake” headline would drop the bar, the harness, and the workload.
This snapshot still splits
AI Charts retrieved the checked Artificial Analysis coding-agents snapshot on Aug 30, 2026, 3:01 PM UTC. Saturation is not a column in that file. The useful check is whether the stored scores still split named configurations. They do. One model does not own every metric. Terminal-Bench v2.1 is the closest approach to a numeric ceiling, at 91.0 for Gemini 3.7 Flash on Opencode, with 7 configurations inside five points of that lead. AA Index, DeepSWE, and SWE-Atlas-QnA still leave a larger remaining gap.
| Benchmark | Model | Agent | Setting | Score |
|---|---|---|---|---|
| AA Index | Opus 5 | Claude Code | xhigh | 68.1 |
| DeepSWE | GPT-5.6 Sol | Codex | max | 68.7 |
| Terminal-Bench v2.1 | Gemini 3.7 Flash | Opencode | high | 91.0 |
| SWE-Atlas-QnA | Opus 5 | Claude Code | xhigh | 54.8 |
Hraness defines an agent harness as software that gives a model a place to work: it injects instructions, offers tools, runs an assess-act-reassess loop, and translates across model APIs. Every stored row already names that layer. The AA Index lead is 68.1 for Opus 5 on Claude Code at the xhigh setting. DeepSWE’s lead is 68.7 for GPT-5.6 Sol on Codex. SWE-Atlas-QnA’s lead is 54.8 for Opus 5 on Claude Code. A model name without the harness and setting is an incomplete citation here too.
What transfers from Luu is the comparison shape, not a claim that any named row cheated. If a lab can see the task family, the harness, and the scoring rule, a public-suite lead is evidence about that suite. It is not a product interchangeability claim. Open the coding-agent comparison chart when the next question is cost, active time, or token use beside those scores. Those axes still move after a quality cell clusters.
How to read a saturated or gameable board
Read Luu’s essay for the changed cost of gaming a large suite, the SPEC history, and the Kimi K3 versus Fable use gap. Read the Hraness reading note of Dan Luu’s The benchmarkpocalypse for a dated digest. Read why a 30% Terminal-Bench-Science score is not a product win when the bar is scientist-set and still mostly missed. Read why a coding-agent high score still needs a holdout when the next control is hidden cases rather than saturated meaning. Read the Hraness harness definition when a row’s agent name needs a noun.
The useful sentence is narrower than a leaderboard headline. A public scoreboard can be cheap to win. It can also stop discriminating the product decision even when the numbers have not hit 100. Either case is remaining measurement work. It is not a reason to ship.
Limits of this reading
- Dan Luu reports FRE, SPEC history, and later model-versus-use checks. AI Charts does not rerun those observations or recover unpublished plot points from his images.
- The Hraness page is a dated digest, not a substitute for the essay. Quote Luu for the measurements. Do not treat the digest as a second primary source.
- Kimi K3 versus Fable, the ICFP contest check, and the vulnerability scan are Luu’s reported use observations. They are not Artificial Analysis scores and not a second official suite.
- Artificial Analysis defines the coding-agent scores. AI Charts is an independent visualization and is not affiliated with Artificial Analysis, Dan Luu, or the listed providers.
- This snapshot is not a 100-point ceiling on every metric. Nearby Terminal-Bench v2.1 scores are a clustering check, not proof that the whole table has saturated.
- The claim that scoreboard saturation is not a product win is AI Charts analysis of Luu’s argument and the stored snapshot. Cite Luu for the measurements.
