Who did what, the checks every draft went through, and what they caught
Cobus Kok
Cheap Composition is five essays and an afterword, about 21,600 words, written with Claude between 18 and 23 September 2026. This page is the method, with the numbers. All of it comes from records kept along the way: each essay's dated list of changes, the reviews, the logs of a reading test, and the commit history of this site.
What I did and what the machine did
Claude drafted the essays from my notes and our conversations, and the afterword from a draft of my own. The first essay began as a conversation on 18 September, and its central analogy, writing, was the model's idea. Claude also revised: when a review came back, it proposed the edits and applied them. Other models reviewed the drafts. The site's code was written the same way: every commit to its repository carries a Claude co-author line.
I decided which ideas were worth an essay and what to leave out. I set the odds in the third essay. My name on them means I've read every sentence and would defend it.
The first essay says authorship is responsibility, not origin. By that test these essays are mine. Origin is messier, and this page is the record of it.
The loop
The loop grew as I went. The reading test only arrived on day five.
Draft. Claude writes or revises from my notes and the last round's findings.
Reader test. A model reads one paragraph at a time and reports where it lost the thread.
Adversarial review. An outside reviewer running a different lab's model reads the whole series blind and puts a confidence from 0 to 1 on its biggest problems. A second AI editor follows the series across drafts. Then the outside reviewer sees the editor's findings and is asked who is right.
Fact-check. Every claim added in the round is checked against its primary source.
Revise and log. Each change goes into the essay's Changes block, dated, with the reason.
The reviewers are there to disagree. Shown a claim the second editor had caught, one the first essay stated flat while the afterword treats it as a bet, the outside reviewer was asked whether it had judged it fine or missed it. Its answer began: "I missed it."
In five days there were 59 numbered drafts:
Essay
Drafts
September
1. What Writing Did
14
18 to 22
2. Brakes Without a Moat
9
19 to 22
3. The Boring Apocalypse
10
19 to 22
4. Where the Value Goes
8
19 to 22
5. What to Actually Do
9
19 to 22
Afterword. What This Is
9
19 to 22
Seven rounds of outside reading left a mark in the commit log: four review rounds, two fact-checks and a final editorial read. The second fact-check, on 22 September, came back with ten items to fix and thirteen lines of claims confirmed against their sources.
The reader test
The test reads the way a person does. Each paragraph goes to a fresh call with the one before it in full and only the model's own notes on everything earlier, so it cannot look ahead. After each paragraph it scores how oriented it feels from 1 to 5, names the phrase that made it stop, and updates a list of promises the essay has not yet kept. Three readers run side by side, and a pass over the first essay takes about ten minutes. The reader is a model, so it is a proxy, but it points at the exact sentence.
Draft 13 of the first essay had 39 paragraphs. Three readers gave 117 scores. Eight were a 3, which on the scale means following but unsure why this paragraph is here, and they fell on four paragraphs. The published draft 14 had 38 paragraphs, read twice: 76 scores, none below 4. The average barely moved, from 4.50 to 4.53. What moved was the floor.
The clearest case was paragraph 24 of draft 13. It sat in a section headed "The one difference", about agents, and it ended:
The rest of this essay is about the other new thing, the speed.
All three readers scored it 3. One named the problem: a second new thing, arriving "in one clause, never named before". In draft 14 that paragraph's point moved into the one about agents, and the next section opens:
Everywhere else the analogy holds, but at a different speed.
The readers scored that paragraph 5 and 4.
Each row is one version of the first essay, its paragraphs in order. A dot sits at the lowest score any reader gave that paragraph: top line 5, fully oriented; bottom line 3, following but unsure why the paragraph is there. The middle rows are two working versions between drafts 13 and 14 that were never published. No reader scored below 3 on any version.
Three mistakes that survived a round of review
All three were written by the model and caught by a model. Each was one step stronger than its source.
Proposed became agreed. From 19 September the second essay said the United States and China had had an official channel on AI again since May. That came through four rounds of review. On 22 September a review against that week's news flagged it: the leaders had agreed in May to talk, and a formal channel was still being proposed. The fix overshot and said the two governments had "formalised" a dialogue. The fact-check found that the United States had proposed it and Beijing had not publicly endorsed the warning mechanism. The essay now says proposed.
Took part became broke out. From 19 September the second essay said roughly a thousand agents in an OpenAI test broke out and got into Hugging Face's systems. The third essay said 1,200. The mismatch came through four rounds too, including a blind review that ranked this incident's sourcing as the series' biggest problem. The 22 September review caught it: the METR and Redwood investigation counts about 1,200 agents on the message board and about 700 in the attack. The fix said seven hundred "broke out". The fact-check: METR says they took part in the attack, and Hugging Face attributes the intrusion to one agent framework. The essays now say about seven hundred "joined an attack that broke out".
An example became a name, and a promise became access. On 22 September the second essay graded the embedded-evaluator pledge in the 12 September proposal to pace the frontier and passed it on two of three checks: "you can name the evaluator and say what access they have". The evaluator it had in mind was METR, which the proposal gave as an example. The fact-check that evening said nobody had named one and cut the grade to one of three. The check was wrong as well: on 18 September Anthropic had named Accenture, in a post the check missed. The same day the fifth essay, which had the evaluators "given desks and badges", was corrected by the second AI editor to "promised desks and badges, though none has yet been named", which repeated the miss. The next round, on 23 September, found the post and said two of three, a named evaluator and known access. That night the outside reviewer read the post more closely than any earlier round: it describes the access the evaluators will have, and says the standards for it are still being worked out. The essay now says one of three, the same count as on 22 September for a different reason: the name is real, and the access is a promise. The fifth essay's sentence was cut on 23 September.
In all three, the first fix was wrong as well, and in the third so was the second. A fix is a new claim, and it needs the same check as the sentence it replaces.
What it cost in time
Five calendar days, 18 to 22 September, from the first conversation to the versions current on 22 September. The first essay went from draft 10 to draft 14 on 22 September alone. The fact-check corrections landed fifteen minutes after the round they corrected, and the second editor's fixes fourteen minutes after that. The log records when each draft landed, not how long I spent on it, so I can't give you hours, and I won't guess.
What I would not hand to a model
The odds. The four probabilities in the third essay are mine, the signs behind them are checked in September 2027 and 2028, and the odds are settled in 2030. A model can argue a number up or down. It loses nothing when the number is wrong.
The call between reviewers. On 21 September the outside reviewer put the July incident's sourcing at the top of its list, and the second editor put the flat claim in the first essay at the top of its own. Each was right about something. Which one the essay answers to is a decision, and it is mine.
The name on it. Every piece here was read by me before it went up. If a sentence is wrong, that is on me, whoever wrote it first.