WRITING / 2026.07.29 · COLUMN · FORBES · 5 MIN
Don’t Pace The Frontier. Look Inside The Trojan Horse.
1,200 frontier-lab employees want Washington to pace AI development with China. Revealed preferences say otherwise — the real ask is transparency, not a pause.
Cite as Christian Catalini, “Don’t Pace The Frontier. Look Inside The Trojan Horse.,” Forbes, Jul 2026. forbes.com ↗
FIG. 01 — 1,200 FRONTIER-LAB EMPLOYEES WANT WASHINGTON TO PACE AI DEVELOPMENT WITH CHINA. REVEALED…
Over 1,200 employees of frontier AI companies, including Anthropic CEO Dario Amodei, several of his co-founders, the chief scientists of OpenAI and Meta AI, and the Chief Strategy Officer of Google DeepMind, have signed a statement calling for the U.S. government to work with other countries, but truly with China, to deliberately introduce guardrails into automated AI development.
The statement comes with deeply felt, personal notes from the researchers, which I would encourage everyone to read. They convey the extreme anxiety and fear you’d expect from the scientists racing relentlessly at Los Alamos on the Manhattan Project. The AI researchers admit to being “quite scared of the pace of progress,” write about self-improving AI being akin to “a runaway nuclear chain reaction,” describe the work ahead as “an insane and suicidal thing to do,” emphasize that “developing technology shouldn’t require a leap of faith for our continued survival,” and compare the challenge to playing chess against a superintelligence: “You won’t have a miraculous insight that lets you outsmart it. You will just lose. And if a rogue superintelligence is created, the game won’t be chess.”
But is AI truly the second time scientists have encountered a “destroyer of worlds,” and should we read these messages as Cassandra alerting us to impending doom? Are these scientists, who know the most about current and upcoming capabilities, the Oppenheimers of our generation?
The reality is that we do not know, and the burden of proof is on the same frontier labs that yesterday endorsed the statement from their official accounts. After all, there is massive information asymmetry between us and them even when it comes to unpacking what actually happened when an OpenAI model started hacking other companies to beat a cybersecurity benchmark.
Sunlight is the best disinfectant, and if the labs have proof of these dangers, rather than publishing statements that may be perceived as virtue signaling and cheap insurance when things go wrong, they should bring the regulators, academia, and the general public rapidly up to speed on what is currently, actually going on.
As an economist, I tend to weigh revealed preferences (your actual actions) as a better proxy for reality than stated ones. And while I do honestly believe that these researchers are concerned and that there is a high chance we have deeply underinvested in security and safety, I do not think that calling on the government to take it from here will actually get us to the result they care about.
Having experienced directly how regulators process new technologies, I worry the private sector may be fooling itself into thinking that the government holds some sort of magic wand to properly deal with unknown unknowns. It excels at managing and monitoring what’s known, and at reacting to an issue once it’s well identified, but right now a public-private partnership may be a lot more effective than handing matters to the State Department.
For one, the basic game theory of international AI pacing cuts deeply against coordination. If AI gives a nation superpowers, then any asymmetric advance will be pursued, irrespective of public statements and diplomacy. It is also unclear how a treaty would be enforced: not only could a country continue to accumulate GPU capacity in the shadows, but it could also accelerate the kind of scientific breakthroughs that allow you to do more with less, the same types of improvements that export controls have triggered in China. Necessity is the mother of invention, and once governments know recursive self-improvement is possible, nothing will stop them from racing toward it.
So what could the labs do if they truly believe they cannot stop themselves from racing? If “the only winning move is not to play,” how could they convince adversaries that the fate of mankind really depends on it, and that this is not just the product of the labs’ collective hallucination?
Humans can and do beat the prisoner’s dilemma, but it takes trust, and trust takes shared, verifiable evidence. The Cold War had a slogan for it: trust, but verify. Ironically, it’s a Russian proverb (doveryai, no proveryai) that Reagan popularized, and that crypto, shaped in an adversarial environment, later hardened into “don’t trust, verify.”
To build trust, first show the evidence, in enough detail for academic researchers and third parties across the globe to understand it.
Second, rather than vague commitments to pacing, double down on safety and security investments. There is a good chance that the OpenAI-Hugging Face drama was simply the result of OpenAI being under immense pressure to catch up with Anthropic on cyber capabilities. Prove to us that this is not the case, and that the OpenAI team did its best to ensure the model could be contained.
Third, work with the research community to ensure defense capabilities are as widely available as possible. OpenAI just released cybersecurity tooling to make it easier to scan your codebase for vulnerabilities. Importantly, Hugging Face defended itself with an open-weights model when a closed one refused. Shift your attention, talent, and funding from the AGI/ASI race to the verification infrastructure that would ensure that what the models see and can do is something we are able to parse and understand.
These would all be revealed-preference moves available to the labs, and they would show real conviction and commitment to delivering us the massive upside of AI without a tragic tail event. Delegating to diplomacy is just cope.
Back in February, when we unpacked the economics of AGI, we concluded that the measurability gap between agents and humans was bound to increase rapidly as a result of our progress. We also pointed out that this leads to a classic externality. And externalities are the kind of thing that, once proven, economists agree calls for regulation. While we didn’t expect the name to become so vivid thanks to Nolan’s blockbuster release, we named it the “Trojan horse” externality.
The idea is simple: if not all of AI’s output can be verified, we have no guarantee that it follows human intentions and preferences. You see it everywhere right now, from an internet flooded with AI slop that takes longer to read than to generate, to tech companies racing to commit only lightly vetted code to their repositories. It qualifies as an externality because the private incentive to release unverified output into society is high, and its social cost is shared among many.
The externality is most salient in AI safety and security: labs are racing to win. Some are racing for money, others possibly for the power to shape the most consequential technology humankind has ever developed. Regardless of money or power, the danger is real. So if it is true that the CEOs with the most insight into where we are on the curve are having regrets, now is the time to bring everyone else up to speed, even if it comes at great commercial cost.
Without proof, statements and letters are just cheap talk. They won’t move Beijing, even if they might move voters as the midterms approach. Cassandra’s curse was to be right and never believed. The labs don’t have that excuse: they can still let us look inside the horse.
A version of this article appeared on Forbes.