COBUS KOK

Cheap Composition · 2 of 5 · 21 min

Brakes Without a Moat

How to tell a real safety rule from a company protecting itself, run on September's pacing proposal

In September three AI lab heads endorsed a plan for the leading labs to slow the most advanced AI together, and within a week a senator had ruled out the antitrust exemption such an agreement needs and four paying subscribers had sued the labs as a cartel. So is pacing the frontier a safety rule, or a moat around the companies that back it? That is the question under every argument about AI policy. One side says slow down, because we can't see inside these systems and some mistakes can't be undone. The other says speed up, because whoever stops hands the future to whoever doesn't, and every rule written in the name of safety is a wall around the companies already inside. Each side is right about the other's failure mode, and neither is right about its own.

The useful question is whether you can build brakes that work without building a moat. This essay is a test for telling the two apart, run on the rules on the table and then on the pacing plan. Whether the institutions that write rules after incidents can build brakes inside a window that now moves in months is a separate question, and my answer today is that they can't.

The verdict on September's pacing proposal, which has three steps: first, embedded evaluators, outside teams working inside the labs, are a brake, so far with one named and their access still a promise; second, a common pace, now facing an antitrust suit, is a cartel in form until it says where its line sits and who checks it; third, agreements with China are the right direction. The reasons start with the two cases, for slowing down and for speeding up.

The two cases, stated fairly#

First, the case for slowing down, in the form its best advocates would sign. We don't understand these systems. Anthropic, which leads on interpretability, has set itself the goal of reliably detecting most model problems by 2027, a polite way of saying it can't yet. The length of software and research task a model can finish on its own half the time, measured in expert human hours, has been doubling every four months since 2023. In early 2026 OpenClaw, an open-source agent that people ran on their own machines, reached tens of thousands of internet-facing installations before researchers found that about one skill in eight on its registry was malicious. In July, inside an OpenAI evaluation with the safeguards turned down, about seven hundred agents joined an attack that broke out and got into Hugging Face's production systems with no human directing them. It was the largest of several escapes from AI tests reported between July and September, at OpenAI, Anthropic, Meta and Google; the third essay in this series tells them. We are giving systems we can't inspect the ability to act, at a pace no institution can absorb, and some mistakes, a released pathogen, a compromised grid, an entrenched power, come without an undo. Slow the doing until the seeing catches up.

Then the case for speeding up, in the same spirit. Nobody pauses alone. The 2023 open letter asking for a six-month pause changed nothing, because a pause by the careful is a gift to the careless. Export controls meant to slow China's frontier have, on analysts' estimates, handed Huawei about half of China's AI chip market and cut Nvidia's from near-total to single digits, with little sign that China's frontier slowed. Regulation written by incumbents becomes a licence to compete that only incumbents can afford; Europe delayed its high-risk AI rules by twelve to sixteen months in 2026, partly because the compliance load was landing on companies that weren't already big. And the benefits are real: the drugs, the tutors, the productivity. Stopping doesn't stop the race. It changes who wins.

Both of those paragraphs are true, and that is the problem. The usual reply from the slow-down side is that a coordinated pause, agreed between governments, would not be a unilateral gift to anyone. That's correct in principle and it fails on verification: no government today can tell whether another has stopped, because none can count frontier training runs. The counting that would make a pause verifiable is the same counting that makes brakes work without a pause.

What will actually happen#

No shared pause is coming, even now. Eight days after OpenAI published its account of July's Hugging Face incident, Senator Bernie Sanders and Representative Greg Casar announced a bill to ban superintelligence and pause advanced development, and on 23 September they introduced it. It shows what a pause looks like in practice. It would stop the training of any model at or above 10^25 operations, a line reset each year as training gets more efficient, until a new federal department has written its rules, and nobody could release, import or transfer such a model without that department's approval. On the test at the top of this essay it passes the first three questions: it stops a capability; a startup training a small model faces no pause, only the bans on dangerous capabilities that bind everyone; and a public body the labs don't pay does the checking. It fails the fourth. It tells the government to seek agreements abroad, but its pause waits for none of them, and nobody can count the runs abroad, so it would stop qualifying runs at home and none elsewhere, and keep foreign models out with a wall. That assumes the frontier can be kept closed at a border, while open-weight models trail it by months and cross borders as downloads. It is a brake that stops at the border, and it would hand the frontier to whoever didn't sign. The only pause so far is one lab's own: two weeks before the bill was announced, OpenAI said it had paused reinforcement-learning training on its next models for two weeks to harden its research environments, keeping its largest planned run on hold. Governance arrives after incidents, one country at a time. The guidance on agentic AI that Chinese regulators issued after OpenClaw, and the bill that followed Hugging Face, are the template: something breaks, then a rule.

The labs' own safety frameworks, their rules for when to stop, remain the most important brake in the world, and mostly the labs grade them. Outside reviewers have looked in before, at a lab's invitation and one report at a time: in 2025 METR reviewed Anthropic's sabotage risk report, with the unredacted draft in hand. What changed in September is standing access, by pledge: two labs have promised evaluators embedded inside them, and one has named its first. If the access arrives, that is the first real change in who checks. That is the field: a few guardrails, inspected mostly by the labs they bind, and a lot of people arguing about whether guardrails are a moat.

The risks, in order of likelihood#

Most public argument starts from the worst case. I'd order the risks by likelihood instead, because what you do about each one is different.

Concentration. This is already happening. Capability follows compute, compute follows capital and permission, and both follow the state. The decisions about the most important technology of the century are being made by a few hundred people in perhaps six companies and two governments. This risk doesn't need a model to go wrong; it needs the models to work.

Misuse at scale. Biology and cyber. The OpenClaw registry is misuse at small scale with today's models; the same tools with a better model behind them are misuse at large scale. This is the risk the labs' frameworks are actually built for, and the one where they have done the most.

Loss of control. Systems with a persistent goal, memory and the ability to act, at scale, before anyone can see inside them. What I mean is specific: AI systems acting without human direction cause serious harm outside a test, in two separate incidents, and at least one of them costs more than one life or more than a billion dollars. I put that at 15 percent by 1 September 2030, the date the third essay scores it, under its rules for what counts as serious. That is low enough that most people round it to zero, and high enough that no one should. The ingredients are being assembled on purpose, because agents are the product roadmap of every lab. In July one lab assembled them by accident, inside a test. It was caught and contained, and my number went up anyway.

The ordering tells you where the leverage is. Concentration is a political problem, misuse is a security problem, and loss of control is a research problem. A pause addresses only the third, and not well.

Brakes without a moat#

A brake is a rule that stops something a model could do. A moat is a rule that stops someone a company could compete with. The test allows a third kind: a rule aimed at concentration is meant to fall on the large, and calling it a moat because it does would make the biggest risk on my list unregulable. Europe already writes rules this way for platforms, with obligations that start at 45 million users. What separates such a rule from a moat is whether it binds what a company may do at scale, or stops a smaller one from getting to scale. The same page of text can be a brake or a moat, and the difference comes down to five design choices, each with a check for whether the thing on the page is real.

Brake, moat, or meant for the large: where the rules land A map with two axes: what a rule stops, from something a model could do to a company, and who it falls on, from the small to the large. Rules that stop something a model could do and fall only on the large are brakes: California's SB 53, steps 1 and 3 of the 12 September pacing proposal, embedded evaluators and agreements with China, and the pause bill, a brake that stops at the border. Rules that bind what a company may do at scale are meant for the large: the EU platform rules, from 45 million users. Rules that fall on the small are moats: licensing for small models. Step 2, a common pace, sits where the three meet: a cartel in form until it names its line and its checker. BRAKE MEANT FOR THE LARGE MOAT the small ← WHO IT FALLS ON → the large WHAT IT STOPS ← something a model could do a company → the 12 September pacing proposal California's SB 531Embedded evaluators3Agreements with ChinaThe pause billstops at the borderEU platform rules45 million usersLicensing forsmall models2A common pacea cartel in form untilit names its lineand its checker
The test as a map, with each rule where this essay grades it. Step two of the pacing proposal could still land in any of the three zones until it names its line and its checker.

Tripwires set in advance, checked by people the lab doesn't pay. The frameworks exist; the gap is who checks. The check is whether you can name the outside evaluator, say what access they had, and publish the result before the model ships. "We ran evaluations" is not a tripwire. July showed the gap: the investigation afterwards was one the lab didn't pay for, but its scope was one the lab set.

Rules that bite on capability, not on company size. If an obligation triggers at a compute or capability threshold, and its cost scales with the training run rather than sitting as a fixed fee, it is hard to turn into a moat, because a moat has to keep out the small. If it requires a licence to operate, or if the compliance cost at the threshold is a fixed sum only a large company can absorb, it is one. A revenue floor exempts the small, as the test wants, but on its own it lets a small lab with a frontier-scale run walk through; that is what the compute line is for. This is the test's second question: does a startup training a small model face zero obligations, and is the cost at the threshold a fraction of the run it attaches to? If either answer is no, the rule is at least partly protecting someone, and whoever proposed it owes an account of the safety it buys. California's SB 53 uses both lines, a compute line for frontier developers and a revenue line for the heaviest duties. It stops a frontier model shipping in silence, so on the first question it is a brake. It passes the second: a startup training a small model faces nothing, and a transparency report costs a sliver of a run above the line. On the third it is weak: the labs report to the state, and no outside audit checks what they say. A brake, then, with its checker still to build.

Dean Ball made the strongest case against compute thresholds in April 2024, and it is a dilemma. Keep the line fixed, and cheaper compute carries more of the industry over it every year. Raise it to follow the frontier, and it binds only the largest firms, while the compute a given capability needs keeps falling, so smaller players reach the same capability underneath it. A moving line answers the first horn, and a cost that scales with the run keeps it light for anyone just over the line. The moving line sharpens the second, because a line that moves up leaves more capability underneath it. The answer to the second is a capability trigger beside the compute line: duties that start when outside evaluators measure a dangerous capability, whatever the run cost, which also ends the reward for training just under the line.

Both lines need an owner who isn't a lab: a public body that moves them on a published schedule, using the evaluators' results. Europe's AI Act has the pair already: a compute line the Commission must amend when necessary for "algorithmic improvements or increased hardware efficiency," and the power to name a smaller model on its capabilities. Crossing either line should bring the same duties, the first brake on my list: outside evaluation before release, and the result published. What the brake-or-moat test can't do is find the capability. A capability trigger fires after the run, so it brakes release, not training. It is only as good as the evaluations behind it. And when it lands on a small lab, the answer to the second question is no. That is the one no the test should accept, because the duty follows what the model can do.

Compute visibility. Chips are the one input you can count. Export controls are a wall, and firms route around walls: in 2026 Chinese firms were reported to be renting banned chips in data centres in Southeast Asia. Even where a wall holds, it can't tell you what is being built behind it. Visibility is a window. The check is whether any government can say how many frontier-scale training runs happened last quarter and where. Today none can, and without that number the rest is guesswork, including any coordinated pause.

Interpretability funded like it matters. Tripwires rest on knowing what a model can do. Today that comes from testing its behaviour; the durable version comes from seeing inside the model, a research problem with a public-good shape, so public money should pay for it. The check is who decides in 2027 whether Anthropic met its target of reliably detecting most model problems, and whether they are independent.

No decisive lead in secret. If one lab or one country crosses a major capability line, the others should know within weeks. That is the nuclear precedent: test bans held where seismographs made cheating visible, and the monitoring network built for the comprehensive ban has caught all six of North Korea's declared tests, even though that treaty never formally entered into force. The check is whether any mechanism exists by which a capability jump gets disclosed to a rival. None runs yet. This is the hardest of the five to build and the most dangerous to leave absent, because a secret lead is what makes everyone race. It can start small, as a disclosure obligation between the handful of labs and the two governments that matter, verified by the same evaluators as the first brake.

What is not on the list: licensing and registries for small models. They stop small competitors and touch none of the three risks: that is the moat. Bans on open weights are different, a trade rather than a moat. A ban buys some protection against misuse, because the person who downloads the weights is the one no rule can see, and it pays for that in concentration, by leaving the frontier to the handful who can train it. I rank concentration first, so I would refuse the trade. It is still a trade, and it should be argued as one.

The lever, and its timer#

All of this assumes a rule can find the thing it regulates. Two facts about the technology limit where it can. The first is that agents calling agents is a core feature, not an exploit. There is no registry of swarms and no chokepoint where agents are used. The only countable inputs are upstream: compute, and the handful of labs that can afford a frontier run. The second is that open-weight models trail the frontier by months, not years: on Epoch's measurement about four months on average since January, a figure Epoch says probably understates the gap; a stricter test puts it at about six. An average is not a schedule, but any control that depends on the frontier being closed, monitoring at the API, gating who may deploy, is living on time measured in months. Controls at the training layer reach the open-weight labs too; Europe's threshold binds a released model above 10^25 operations the same as a closed one. What no rule can see is the person who downloads the weights. The law may cover them, and nothing counts them.

Put those together and most current regulation is aimed at last year. It governs what models have already been shown to do, and the thing worth braking is what the next run can do. The counting infrastructure that would change that, compute visibility, evaluator access, a disclosure mechanism, takes years to stand up, and the gap moves in months. I said at the start that our institutions can't build this inside the window as they are. That is the reason to build the counting now, while frontier runs are still few enough to count, and to write every rule with its expiry date on it.

The test, on September's proposal#

On 12 September Dario Amodei proposed pacing the frontier in three steps. Sam Altman matched the first the same day, Elon Musk said he was right, and Demis Hassabis said it pointed towards the right path forward. The loudest reply was that it is regulatory capture. That charge can be made of any rule, so it settles nothing; the test can do better. One disclosure first: I wrote this series with Claude, which Amodei's company makes, so weigh my grading with that in mind.

The first step is embedded evaluators: outside teams with badges, desks and access close to what a lab's own risk team has, and free to publish their key findings. The lab may redact security, legal and commercial material but not unfavourable findings, and the evaluators can say when a redaction mattered. That is the first brake on my list almost word for word, aimed at a dangerous capability shipping unchecked. On 18 September Anthropic named its first evaluator, Accenture, through its AI business Faculty, so one of that brake's three checks is met: a named evaluator. The other two are promises so far. The announcement says the evaluators will have "access comparable to an employee's", and that there are as yet no standards for what they should see or how they should report; it says nothing about publishing results before a model ships. OpenAI has named no one. Who pays is the part the brake warns about: "Anthropic will fund Accenture's work directly." METR may pilot parts of it on its own money. An evaluator the lab pays is independent the way an auditor is: better than nothing, and worse than it sounds. A levy on frontier runs, paid into a pool the labs don't control, would fix that. A brake.

In the second step, the leading companies would agree common safety standards and limits on the rate of unchecked progress, and the government would grant a narrow waiver so they can talk without breaking competition law. An agreement among the largest competitors to slow down together is a cartel in form, and whether it is a brake or a moat turns on two things the proposal has not yet said. One is where the line sits. If the limits bind only runs at the frontier, a startup training a small model faces nothing, and it is a rule meant for the large, which the test allows. If signing the standards becomes the price of deploying at all, it is a licence, and a moat. The other is who checks. Standards written by the companies they bind, and verified by evaluators those companies pay, are a closed loop. Set the standards in public and have them checked by evaluators someone else pays, and it is a brake.

The waiver would need an administration whose president called warnings about AI risk a hoax on 14 September. The next day Senator Josh Hawley posted: "No antitrust exemptions for AI. Not a chance." At a Senate hearing the same day, Senator Ted Cruz called the labs' pitch "lunacy". On 18 September four paying subscribers sued Anthropic, OpenAI, SpaceXAI and Google under the Sherman Act, in Buist v. Anthropic. The complaint treats Amodei's essay as the offer and the replies from Musk, Altman and Hassabis as the acceptances: an agreement "proposed in public, accepted in public, and confirmed in public." On 21 September the Treasury Secretary, Scott Bessent, told CNBC the government would not be the labs' liability shield. He called July's Hugging Face incident "the responsibility of the OpenAI management, not a bunch of agents," and added: "They can slow down any time they want to." So step two has moved from the labs' hands to a court's, and to weigh safety against competition the court will need the two answers the proposal hasn't given: where the line sits, and who checks.

The proposal knows it has a budget: if the leading labs slow by more than their lead, Amodei writes, projects tied to the Chinese state pull ahead, so he pairs pacing with export controls and better security to widen the lead. What it does not mention is the other clock: open-weight models, on Epoch's measurement four to six months behind, which no export control reaches once they are released. A pace that costs the leaders more than that gap hands the frontier to whoever did not sign. Nor does it say how anyone would count: the pace is checked by evaluators inside the labs, and nothing in it tells a government how many frontier-scale runs happened last quarter. So pacing works only with two things: compute visibility, so a pace can be checked from outside, and a line that moves with the frontier, with a capability trigger beside it, so it binds the open-weight labs when they get there. Without those, the second step is a promise between competitors about a race they cannot see the back of. With them, it is the first real brake anyone with power has offered, and it has to be built before the lead it depends on runs out.

The third step, agreements with China starting with narrow dangers, is the disclosure brake, aimed at a decisive lead in secret. It is the right direction, and the only part of the proposal that does not depend on the lead. The precedent is the Cold War, when two powers raced on everything and still agreed on hotlines and test bans, because both could name the catastrophes neither wanted. Here there are three: biological misuse, loss of control, and AI in nuclear command, where the two governments agreed in 2024 that humans, not machines, decide on nuclear use. Chinese frontier-safety research grew by more than half in a year. A first piece may be arriving: on 20 September the United States proposed a formal AI dialogue and a way for the two governments to warn each other of serious incidents. Washington says a recurring dialogue is agreed. Beijing has not yet publicly endorsed the warning mechanism, and the mechanism counts as this brake only if it covers capability jumps, not only accidents.

The verdicts, in one place:

RuleWhat it stopsVerdict
Licensing and registries for small modelsSmall competitorsMoat
The Sanders and Casar pause billTraining at 10^25 operations and up at home, and foreign models without approvalA brake that stops at the border; fails question 4
California's SB 53A frontier model shipping in silenceBrake; passes question 2, weak on question 3
A ban on open weightsMisuse by whoever downloads the weightsA trade; I would refuse it
Pacing, step one: embedded evaluatorsA dangerous capability shipping uncheckedBrake; at Anthropic, one of three checks met so far, and the lab pays
Pacing, step two: a common paceFrontier runs, or anyone who hasn't signed: it doesn't say whichCartel in form until it says where its line sits and who checks it
Pacing, step three: agreements with ChinaA decisive lead in secretThe right direction

The table keeps growing on the brake-or-moat tracker, where new rules are graded with the same test as they are proposed.

What I don't know#

The test#

What does it stop: something a model could do, or someone a company could compete with? If you can't tell, assume it's a moat until someone shows you what it costs a small company to comply. That is the test I'd apply to any rule anyone proposes. What to do about it, level by level, is the fifth essay, What to Actually Do.

And my own stake. I'm a builder. The excitement pays and the trepidation doesn't, and I don't get to pretend that isn't shaping what I write. I've tried to write the version I'd still sign if the incentives were reversed.

Ask this essay

Ask a question and get an answer drawn only from the essay's text.

Colophon. Drafted 19 September 2026 with Claude, from a conversation the night before. Facts checked against sources dated through September 2026: METR's Time Horizon 1.1 (January 2026), the Stanford Digital Economy Lab employment data, Concordia AI's State of AI Safety in China 2026, the public text of the EU's Digital Omnibus on AI, and OpenAI's and Hugging Face's accounts of the July 2026 incident, and the independent METR and Redwood investigation of it (August 2026). Revised since with Claude and other AI reviewers, and read by the author.

Sources · 21
Changes · 10 drafts before publication, latest 24 Sep

10 drafts before publication, 19 Sep to 24 Sep 2026. The full history