Skip to benchmark notes

Coding-agent scores still need expertise

A stored coding-agent score names a model, harness, and setting on a published suite. Faye and Goedecke show why that number still needs a person who can specify and audit the work.

Lars Faye’s AI Coding will Prevent Expertise, published July 22, 2026, argues that coding assistants demand the expertise they also prevent novices from forming. Sean Goedecke’s LLMs reward expertise, published July 24, 2026, argues that the same models amplify domain knowledge rather than flatten it. This note places those two claims next to the current coding-agent comparison.

The checked snapshot answers a narrow measurement question. It records that a named model, on a named agent harness, at a named effort setting, scored a stored value on a named suite. Faye and Goedecke answer a different question: who can specify the work that produced a plausible result, and who can audit that result once it exists. A high cell is not a credential for the person reading it.

Faye’s expert-novice bind

Faye names a skilled-orchestrator paradox. The skills that make an assistant useful (taste, review, and the ability to reject a fluent wrong turn) are the skills that unrestricted generation can skip. Veterans already have a history of failing, tracing, and rewriting. Newcomers are told to accelerate with tools that remove that friction.

He cites The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers, a study of novice programmers that JetBrains highlighted. Heavy assistance produced confident finishes without the planning those novices would have done alone. People who ignored bad suggestions wrote the code they already intended. Faye quotes the study’s verdict: “Finished with an 'illusion of competence' rather than true understanding.

The learning setup is inverted. The student has to steer the mentor first, so domain knowledge is what makes an answer checkable. Without it, the model confirms the direction already implied by the prompt. Faye treats generation as a leaky abstraction and keeps Joel Spolsky’s limit attached: “the abstractions save us time working, but they don’t save us time learning.” He also quotes François Chollet: “You cannot interpolate your way through a completely unique system failure.

A Hraness reading note of Faye’s essay, saved 2026-08-24, records that bind as a dated digest. The note is a companion citation. The argument lives in Faye’s essay.

Goedecke’s claim that models reward expertise

Goedecke starts from the opposite public story: if everyone can ask the same model for sort-of-okay CSS, then prompting looks like a skill with no content. He calls that reading wrong. He writes: “The most important skill in prompting is expertise in the domain you're prompting for.

His illustration is Terence Tao’s conversation with ChatGPT about a counterexample to the Jacobian Conjecture. Tao’s messages stay short, push back when an answer looks too complex, and propose formulations the model did not choose. Goedecke is clear that those habits are not a transferable prompt recipe. They work because Tao understands the mathematics well enough to pull a useful fragment out of a long reply and to notice what looks strange.

The same constraint appears in software. Goedecke writes that “system design problems are dominated by concrete specifics, not generic principles.” Familiarity with a codebase supports questions a generic design lecture cannot: whether a simpler path already exists, whether the system already does the work, and which local terms would make the problem smaller. Novices can still get something from the same model. Experts can steer it much harder.

He treats that gap as durable. For many tasks, “the human is the bottleneck, not the model” because the hard work is communicating the desired solution and judging whether the output is that solution. A Hraness reading note of Goedecke’s essay, saved 2026-08-05, keeps that constraint as a dated digest.

What the snapshot names on each row

AI Charts retrieved the checked snapshot on Aug 18, 2026, 11:03 AM UTC. The dataset contains 59 model-agent configurations across 28 models, 9 agent harnesses, and 10 providers. The dataset page defines each metric and lists the highest stored score for that metric. This note copies those named fields. It does not invent an operator, a rank, or a claim about who can read the table.

Highest stored score by benchmark in the Artificial Analysis snapshot retrieved Aug 18, 2026, 11:03 AM UTC
BenchmarkModelAgentSettingScore
AA IndexOpus 5Claude Codexhigh66.7
DeepSWEGPT-5.6 SolCodexmax68.7
Terminal-Bench v2GPT-5.6 SolCodexmax87.7
SWE-Atlas-QnAOpus 5Claude Codexhigh54.8

AA Index is the snapshot’s overall 0–100 score across code changes, terminal work, and repository understanding. DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA stay separate. The highest stored AA Index is 66.7 for Opus 5 on Claude Code at the xhigh setting. That sentence is complete only when the model, harness, setting, and suite stay attached to the number.

Hraness defines an agent harness as software that gives a model a place to work: it injects instructions, offers tools, runs an assess-act-reassess loop, and translates across model APIs. The snapshot already treats that layer as part of the observation. Two rows that share a model and differ only in harness or setting are different measurements.

Highest stored AA Index configurations in the Artificial Analysis snapshot retrieved Aug 18, 2026, 11:03 AM UTC
ModelAgentSettingAA Index
Opus 5Claude Codexhigh66.7
GPT-5.6 SolCodexmax66.6
Fable 5 (with fallback)Claude Codemax65.8
Opus 5Claude Codemax65.5
GPT-5.6 SolCodexxhigh65.1

Artificial Analysis publishes the coding-agent comparison that this snapshot copies. AI Charts does not recalculate those scores. The public page is the source for the numbers. Faye and Goedecke are the sources for why a person still has to stand next to them.

A named row still needs a reader

The snapshot can tell you that a configuration did well on a published suite at the retrieval date. It cannot tell you whether the person citing that cell can write a specification the harness will follow, or whether that person can notice when the output is fluent and wrong. Those jobs sit outside the table.

Faye’s bind is about how that reader is formed. If generation removes the friction that builds taste, the next cohort can inherit a scoreboard it is poorly equipped to audit. Goedecke’s claim is about how that reader is used. The same named harness yields more when the operator already knows the domain well enough to reject a plausible dead end.

Those two essays agree on the scarce resource. The model can produce a large space of possible programs. Selecting, specifying, and checking a result still belongs to someone who can recognize a good one. A stored AA Index does not record that someone.

Specify and audit stay human work

Goedecke’s Tao example is a specification problem. The useful turn is not a longer prompt. It is a person who already knows which objection, simplification, or local term would change the next step. Faye’s prescription keeps the same work on the person after the model answers: documentation, exercises, and checks against sources other than the model.

A coding-agent row makes that split concrete. The harness is the loop that offers tools and retries. The setting is the effort the run was allowed. The score is the suite’s verdict on that loop. None of those fields says whether the human goal was stated well, whether a wrong file was accepted, or whether the change should have been smaller. Those checks are the expertise the essays describe.

The snapshot is still the right place to look up the named configuration. Open the comparison chart to pin a model or change axes. Open the dataset page for metric definitions and the full configuration table. Keep Faye and Goedecke next to that lookup when the decision is whether the person in the loop can stand behind the result.

A different question from holdouts

Why a high score still needs a holdout asks whether a public-suite win survives cases the optimizer did not see. This page asks whether the person citing the win can specify and audit the work. Both questions attach to the same leaders table. They fail for different reasons.

A holdout can falsify a score when the suite was visible and the hidden cases were not. Missing expertise can leave a true suite score standing while the production change is still wrong, incomplete, or impossible to review. The first failure is about the task set. The second is about the reader.

How to read this scoreboard

Read Faye’s essay for the expert-novice bind, the novice-programmer study, and the leaky-abstraction limit. Read the Hraness reading note of that essay for a dated digest. Read Goedecke’s essay for the claim that prompting skill is domain expertise, and the Hraness reading note of that essay for its digest. Read the Hraness harness definition when a row’s agent name needs a noun. Read the dataset page when you need the current metric definitions. Read why a high score still needs a holdout when the next question is unseen tasks rather than an unseen reader.

The useful sentence is narrower than a leaderboard headline. A high coding-agent score means the named model, harness, and setting did well on the visible suite at the retrieval date. Faye and Goedecke add the missing clause: that number still needs a person who can say what the work should be and check whether the output is that work.

Limits of this reading

  • Lars Faye and Sean Goedecke own their essays. AI Charts does not rerun the novice-programmer study, Tao’s conversation, or any other example they cite.
  • The Hraness pages are dated digests, not substitutes for the essays. Quote Faye and Goedecke for the arguments and the Hraness notes only for their own digest sentences.
  • Artificial Analysis defines the coding-agent scores. AI Charts is an independent visualization and is not affiliated with Artificial Analysis, Lars Faye, Sean Goedecke, or the listed providers.
  • Highest stored scores are observations of named configurations in this snapshot. They are not general ranks, operator credentials, or production guarantees.
  • The snapshot does not record who specified a run, who reviewed the output, or whether that person could detect a fluent error.
  • This is a checked snapshot, not a live mirror. Cite the retrieval timestamp when quoting a value.

Sources

  1. AI Coding will Prevent ExpertiseLars Faye, 2026. The July 22, 2026 essay argues that coding assistants demand the expertise they also prevent novices from forming, and treats generation as a leaky abstraction.
  2. Hraness reading note: AI Coding will Prevent ExpertiseHraness, 2026. The Hraness reading note is a dated digest of Lars Faye’s essay, used here as a crawlable companion citation rather than a substitute for the original.
  3. LLMs reward expertiseSean Goedecke, 2026. The July 24, 2026 essay argues that LLMs amplify domain expertise rather than flatten it, and that specifying and judging a result remain human constraints.
  4. Hraness reading note: LLMs reward expertiseHraness, 2026. The Hraness reading note is a dated digest of Sean Goedecke’s essay, used here as a crawlable companion citation rather than a substitute for the original.
  5. Coding AgentsArtificial Analysis, 2026. The public coding-agents comparison is the upstream source of the checked AI Charts snapshot. Model names, agent harnesses, settings, AA Index scores, and mean API costs are Artificial Analysis measurements.

Results describe the named model, harness, task set, budget, and evaluation version. They do not establish performance on every production repository.