NOTES / 2026.09.14 · 7-post thread · 16 likes
For safety, recursive self-improvement is a red herring
AI R&D is already being automated, and even an ASI has to wait for reality on anything unmeasured. The danger arrives earlier and is already all over the economy: you automate more than you should because you don’t bear the full cost when it goes wrong. Neutral inspectors, not faith.
FIG. 01 — FROM THE ORIGINAL THREAD
For the safety discussion, recursive self-improvement is a red herring.
We already have increasing automation of AI R&D.
Any remaining bottleneck where you still need to verify it’s doing what it’s supposed to slows you down.
Even ASI has to wait for reality on anything that hasn’t been measured yet.
But the danger arrives earlier, and it’s already all over the economy, starting inside the labs: you automate more than you should because you don’t bear the full cost if it goes wrong. x.com
It’s the developer who gets lazy reviewing their code or agent swarm. It’s OpenAI & Anthropic racing while underinvesting in cyber and security engineering. It’s the lawyer who saves hours by using unverified AI-generated.
At the extreme, it’s human as a “meat proxy.”
It’s a classic externality: you collect the short-term upside while others pay for your underinvestment in verification. Society accumulates tech debt, and systemic risk. x.com
So yes, I want neutral inspectors and evaluators in.
Then we’ll know whether the tech is dangerous (wherever we are on the continuum from basic automation to RSI), or whether people were just recklessly racing and underinvesting in safety. x.com
We’ve solved this in other domains!
As @wu_jane’s research shows, measurement and metrics are the key.
Stop asking us to take safety on faith. Start showing us the evidence.
Originally published as a thread on X.