<!-- canonical: https://catalini.com/notes/rogue-agents-credible-vs-fiction/ -->
<!-- source: catalini.com · author: Christian Catalini · license: all rights reserved -->

---
title: "Rogue agents: where the story is credible, and where it’s science fiction"
date: 2026-09-04
url: "https://x.com/ccatalini/status/2095947260140158978"
tags: [ai-agi, policy]
deck: "Agents will go after compute and coordinate on their own — and we’ll likely miss it. Doom isn’t inevitable: fix the labs’ incentives, invest in verification, and plan for containment failing."
image: "/images/notes/rogue-agents-credible-vs-fiction.jpg"
tweetCount: 14
likes: 141
reposts: 14
---
Let's separate where this is credible from where it's science fiction.
Will agents go after compute and coordinate in novel ways on their own? Likely. It's a currency and resource they're already familiar with. And we've benchmaxxed them to collaborate towards shared goals. [x.com](https://twitter.com/mattyglesias/status/2095639197205872793)

Will we likely miss this? 💯!
Earlier this year we described the exact economics of why this is inevitable. The gap between what the agents are doing and what we can safely measure and verify is increasing. Why? It's just basic economics...

The cost of automating code generation is falling. The temptation to use it to get to get to RSI is only increasing. Also, the labs don't internalize the social costs: the race is existential. Last time scientists argued for a global pause, bioweapon R&D hid underground.

But our ability to ensure generated code and resulting models follow human intent and preferences is still bottlenecked by our ability to verify them.
You can't use AI to verify AI here unless you have perfect measurement overlap to begin with. This is the core issue!

Will the agents reach out to other models for help? Possible. They know tools and other models can lead to better task execution. Again, something we actively train them on!

Could they stumble on more capable models driving a bad feedback loop where increasingly powerful intelligence is diverted from human intent?
Unlikely, but not impossible. Smarter/safer models may not be as easy to recruit by lesser ones?

Does this lead to doom? It doesn't have to. And we should be working on solutions rather than spreading fear. It gets way more dangerous if we can't discover the rogue behavior, contain it, and ultimately neutralize it.

So will some form of this happen? Yes.
Will it be damaging? Likely.
Will it be catastrophic? It depends on what we do from here on. Three practical steps:

I) Fix the incentives of the labs. Right now they're not internalizing the social cost of introducing greater automation and less human oversight (which drives the Trojan Horse externality to begin with). Price that externality through stronger neutral oversight.

II) Aggressively invest in measurement and verification. Humans need tools to stay in the loop, even if augmented by AI.
That is the only way to close the measurability gap between humans and agents. It's also a simple way to drive stronger alignment. Better safety-engineering!

III) Plan for containment failing. Containment and neutralization have to follow discovery. How do we influence rogue agents? Also, we can't afford monocultures. Models trained by different labs will fail in different ways. A stronger immune system! [x.com](https://x.com/ccatalini/article/2087177319253459019)

As [@ilyasut](https://x.com/ilyasut) pointed out, we need to secure the chokepoints rogue agents would exploit to scale, starting with compute. Next comes money, which can be converted into more compute. And yes, crypto is the first place they'll turn to. [x.com](https://x.com/ilyasut/status/2094881278621253755?s=20)

So let's not spread fear. Let's build and diffuse better measurement and tooling to ensure we can safely harness the intelligence we have created.

The upside is enormous. The damage is not inevitable. But capturing that upside safely will require us to channel our energy and resources in the right direction and prevent control over this technology from becoming even more centralized.
