Calvin French-Owen writes that small models have arrived, in an essay published August 26, 2026. He reports gpt-5.6-luna near 100 tokens per second, research-thread API bills in the tens of cents, and a personalized news eval near $0.10 versus about $1 on Sonnet-class models. He treats inference cost as the reason consumer AI companies were scarce. He still reaches for frontier models when the work is hard coding. This note keeps that split.
AI Charts publishes named coding-agent rows on a Pareto frontier and a scoreboard. Each stored row names a model, an agent harness, and an effort setting. A consumer-cheap model that wins a news-eval dollar chart is not a reason to collapse those rows into one cheap-won column. Closing a scoreboard is a different event from closing a consumer bill. That sibling note asks whether a SemiAnalysis era composite should erase the named rows. This page asks whether a ten-cent news eval should.
This page is not the Hraness reading digest of French-Owen’s essay. The digest, saved 2026-08-27, is a dated companion citation. Quote French-Owen for the measurements. Quote the digest only for its own sentences.
French-Owen measures a consumer bill
The essay’s evidence is a personal cost chart, not a public coding-agent suite. French-Owen has been running gpt-5.6-luna through a codebase, email, and a knowledge base. The speed claim is about 100 tokens per second. The bill claim is that fairly complicated research threads stay in the tens of cents, including searches across thousands of emails.
His pet eval is a daily personalized news site. The prompt asks a model to research him, then build a micro-site of stories from Hacker News, Reddit, and Twitter. On Sonnet-class models he spent about $1 to get anywhere. On luna he reports decent results at an average of about $0.10. He writes: “But looking at luna, the results are pretty decent, and the average cost is ~$0.10.”
| Observation | Quoted figure | What it measures |
|---|---|---|
| gpt-5.6-luna speed | About 100 tokens per second | Interactive throughput on his runs |
| Research-thread API bill | tens of cents | Complicated personal research, including large email search |
| Personalized news eval | $0.10 versus $1 on Sonnet-class models | His daily news-site prompt, not a coding-agent suite |
Those figures belong to French-Owen’s runs and his news prompt. They are not AA Index cells, DeepSWE cells, or mean API costs from Artificial Analysis’s coding-agents comparison. A ten-cent news eval can be a faithful citation of the essay and still leave the coding-agent scoreboard untouched.
Token cost, not taste, blocked consumer AI
Investors asked him why consumer AI companies have been scarce. His answer is a cost structure, not a missing product idea. “There's a straightforward answer: token costs.” The pre-AI consumer playbook assumed a cheap-to-run website, virality, then an ads marketplace. Per-request inference broke that capital path. A product that spends a dollar to assemble one personalized edition cannot charge a newspaper subscription and survive.
A Hraness reading note of the essay records the same limit as a dated digest: “Calvin French-Owen argues that small, fast models have crossed a cost-quality threshold that unlocks consumer AI.” The digest’s cost sentence is the one this page keeps in view: “Inference cost, not product taste, is why consumer AI companies have been scarce.”
That claim is about whether a consumer loop can pay for itself. It is not a claim about which named coding-agent configuration leads a public suite. Crossing a news-eval dollar threshold can unlock a class of products and still leave the scoreboard’s columns in place.
He still keeps frontier models for hard coding
The same essay refuses to retire the expensive models. For coding work French-Owen almost always reaches for Fable 5 and GPT-5.6 Sol. He says that habit made the small-model progress easy to miss. He also expects demand for frontier-level models to keep compounding in engineering, hard science, and model training.
He reports a second demand curve from Peter Reinhardt. About 95% of that operator work is “token spewer” responsiveness: hopping on calls, nudging people, and blocking and tackling. The remaining slice is the novel-breakthrough work he still assigns to an expensive model. Cheap models, in this telling, can start to cover the responsive slice. They do not replace the slice that still needs a frontier stack.
French-Owen also names GLM 5.3 as a new option on a general Pareto frontier. That sentence is his. The checked coding-agent snapshot does not store a GLM-5.3 row, and this page does not mint one. A general cost-quality remark is not a stored AA Index cell.
A news-eval win is not a coding-agent cell
AI Charts retrieved the checked snapshot on Aug 28, 2026, 3:04 PM UTC. The dataset contains 57 model-agent configurations across 27 models, 10 agent harnesses, and 11 providers. The snapshot has no consumer-bill field and no news-eval dollar column. Each row is already a product citation: a model name, a harness name, and an effort setting, with scores copied from Artificial Analysis.
The snapshot does not store a gpt-5.6-luna coding-agent row. That absence is part of the argument. A consumer-eval dollar chart does not mint a named AA Index cell. Inventing a luna scoreboard line from a ten-cent news prompt would collapse two jobs into one column.
The snapshot does store named Fable 5 and GPT-5.6 Sol configurations, the models French-Owen still reaches for when the work is hard coding. Those rows stay attached to a harness and a setting. They are coding-agent observations on the retrieval date. They are not consumer-bill observations.
| Model | Agent | Setting | AA Index |
|---|---|---|---|
| Fable 5 (with fallback) | Claude Code | max | 67.2 |
| GPT-5.6 Sol | Codex | max | 65.1 |
Those two lines are already on the live coding-agent comparison. They remain two named products. A cheap-model headline that erased the harness or the setting would drop the only fields that make a stored row citeable.
The harness still names the row
Hraness defines an agent harness as software that gives a model a place to work: it injects instructions, offers tools, runs an assess-act-reassess loop, and translates across model APIs. French-Owen’s own close already points at that layer. He says fast, cheap, good-enough models still need new harnesses, prompt-injection safety, roles, and permissions before they can run a business.
A cheaper engine does not delete the place the engine works. If the consumer job is a news loop, the missing work is a harness that can run that loop at a ten-cent bill. If the coding job is a named scoreboard cell, the row still has to say which harness and which setting produced the score. Keeping those jobs in different columns is how a later cheap-model row can appear without rewriting the frontier rows already stored.
Closing a bill is not closing a scoreboard
Open models can close a scoreboard and still lose the product asks whether a SemiAnalysis era composite should collapse named coding-agent rows into an open-won headline. This page asks the adjacent question for a consumer bill. Both notes keep the snapshot as named configurations. They do not share a source suite.
| Source | Question it answers | What a win there means |
|---|---|---|
| French-Owen news eval | Can a small fast model assemble a personalized edition at a consumer-viable bill? | A dollar-chart observation on his prompt |
| French-Owen coding habit | Which models does he still use for hard coding? | A product preference, not a second news eval |
| AI Charts coding-agent snapshot | What did this named model, harness, and setting score on the stored metrics? | A configuration observation on the retrieval date |
The useful failure mode is a headline that treats the first row as a substitute for the third. “Small models have arrived” can be a faithful citation of French-Owen and still be the wrong instruction for this snapshot. The snapshot would have to invent a luna coding-agent cell, or drop harness and setting from the Fable and Sol rows, to print a cheap-won rank. The essay does not ask for that drop. It keeps frontier models for hard coding while it celebrates the cheaper consumer loop.
How to read this page
Read the French-Owen essay for the luna speed, the research-thread bills, the news-eval dollar comparison, the consumer-capital argument, and the coding-versus-responsiveness split. Read the Hraness reading note for a dated digest of those claims. This page is not that digest. Read the coding-agent comparison when you need the live named rows. Read why closing a scoreboard is not closing a consumer bill when the adjacent source is SemiAnalysis rather than a news-eval dollar chart. Read the Hraness harness definition when a later cheap-model row still needs a noun for the software around the weights.
The useful sentence is narrower than the essay title. Small fast models can close a consumer bill on a news eval. French-Owen still picks frontier models for hard coding. AI Charts keeps the live coding-agent snapshot as named model, harness, and setting rows for the same reason.
Limits of this comparison
- French-Owen reports his own luna runs, news-eval bills, and coding preference. AI Charts does not rerun that prompt or recover unpublished cost traces.
- The Hraness page is a dated digest, not a substitute for the essay. Quote French-Owen for the measurements and the Hraness note only for its own digest sentences. This page is not that digest.
- Artificial Analysis defines the coding-agent scores. AI Charts is an independent visualization and is not affiliated with Artificial Analysis, French-Owen, or the listed providers.
- This page does not add a gpt-5.6-luna coding-agent cell. A consumer-eval dollar chart is not a stored AA Index observation.
- Named Fable 5 and GPT-5.6 Sol scores belong to the stored model, harness, setting, task set, and evaluation version on the retrieval date. They do not establish results for every repository or production workflow.
- This is a checked snapshot, not a live mirror. Cite the retrieval timestamp when quoting a value.