WRITING / 2026.10.02 · ESSAY · HBR · 14 MIN
AI Is Making Verification the Bottleneck for Companies
Can your organization properly steer AI-generated output—and stand behind the results?
Cite as Christian Catalini, “AI Is Making Verification the Bottleneck for Companies,” Harvard Business Review, October 2026. hbr.org ↗
FIG. 01 — CAN YOUR ORGANIZATION PROPERLY STEER AI-GENERATED OUTPUT—AND STAND BEHIND THE RESULTS?
Somewhere in your company today, an expert rejected a perfectly plausible answer. A controller caught a revenue figure that counted deferred bookings. A security engineer blocked a release that passed every test but would have handed an AI agent write access to the production database.
That override, escalation, or reversal is the most valuable data your firm produces. If you record enough of those moments and use them, alongside everything else your firm knows and measures, to train an AI model that learns from each correction, you could, in theory, eliminate the need for managers to route information and coordinate action. In doing so, execution gets turbocharged by automation. You might even be able to eliminate hierarchy altogether.
That, at least, is the pitch. Microsoft CEO Satya Nadella calls that feedback system the critical “learning loop” of the AI era. He also insists that enterprises need to own it rather than cede it to the frontier AI labs. Palantir’s CEO Alex Karp has likewise argued that companies and governments will need sovereign AI to preserve their independence. Jack Dorsey and Roelof Botha went further, offering a blueprint for an ambitious “company world model”: an omnipresent and continuously updated model of the entire business that could coordinate work autonomously.
The vision is admittedly seductive. AI makes context extremely easy to retrieve, digest, and continuously process the moment new information arises. What once took an army of managers and multiple systems of record can theoretically be handed over to a single well-tuned AI model.
The solution also draws on decades of economic thinking about why firms exist and how they organize work. In the 1930s, economist Ronald Coase argued that firms arise to reduce the costs of relying on markets: discovering prices, negotiating separate contracts, and adapting agreements as circumstances change. Firms bring activities inside their boundaries when coordinating them internally costs less than transacting through the market. In the 1970s, business historian Alfred Chandler documented how the “visible hand” of professional management outperformed market coordination in large-scale production and distribution. Their shared insight is that internal coordination can confer an economic advantage.
If AI can perform that coordinating work more cheaply and reliably than layers of human managers, the economic case for those layers weakens. Dissolving the hierarchy becomes the next logical step. With AI, leaders see a chance to preserve the benefits of internal coordination while stripping out much of the friction of human-to-human management.
The problem is that execution is only half of the story. The other half—and the source of some of that apparent friction—is the work of verification. Hierarchy does more than route information; it also verifies: deciding what information means, what deserves the firm’s attention, which outputs can be trusted, and which are worth generating at all. As the ability to generate text, media, and code becomes increasingly commoditized and abundant, the calculation as to which half creates the most value for companies evolves too. When execution is cheap, verification becomes more valuable, and firms turn into verification factories: institutions capable of properly steering AI-generated output and standing behind the results.
This explains why AI has progressed unevenly across domains. Enterprise AI has advanced fastest where verification is relatively easy. In coding, outputs can often be tested rapidly and cheaply, helping explain why it has so quickly become a success story of AI deployment.
Leaders under pressure to define an AI strategy need to recognize what its success hinges on: a coherent plan to solve the verification problem. They also need to recognize how they may be giving away their advantage in this emerging landscape by letting expertise erode and building on technology stacks they don’t own. Using the economics of verification as a framework, this piece explains what belongs in a firm’s verification loop, why leaders must retain control of it, and how doing so enables reliable autonomy without giving away the tacit knowledge and expertise that are key to long-term competitive advantage.
The Verification Bottleneck
Increasingly, what can be measured can be automated. Once the relevant information is codified and success can be measured, an AI model can use that feedback to hill-climb toward the desired outcome. But there is a catch: Any divergence between the context a human brings to a problem and the context available to the machine is a potential source of error.
Think about the last time you wrote a prompt, only to realize midway through that you had omitted a critical dimension, constraint, or desired outcome. More often than not, the missing detail was something you understood tacitly and took for granted. This is a key reason machine intelligence can feel so “jagged:” astonishingly capable along measurable dimensions, but surprisingly clumsy everywhere else.
Verification is the act of closing the gap between the weights that power an AI model and the neural paths and “weights” that power your own decision-making on the same problem. As AI makes execution increasingly cheap, the ability to properly steer and verify AI outputs better than your competitors will be the differentiator.
World-class verification factories rely on two assets: unique ground truth and talent.
For some firms, unique ground truth is simply a byproduct of scale and incumbency. If you process most of the world’s transactions in a domain, you observe that slice of the economy with a precision nobody can replicate. For others, it comes from extreme focus. If you have underwritten the same narrow risk for decades, you hold ground truth nobody else has seen. Neither is the prerequisite; unique measurement is. The best firms use measurement to turn market uncertainty into manageable risks, repeatable processes, and consistently better products.
But ground truth alone is inert. Data only records what happened; when something falls outside the norm, it takes an expert to decide what it means and how to act on it. When domain experts set objectives and verify AI output, they draw on accumulated experience: past errors, corrections, and lessons learned. That experience is itself a form of measurement, one the AI does not have access to.
Some call what remains unmeasured taste, others judgment. But these terms are vague, and their boundaries keep shifting: as model outputs improve, more of what once required human judgment is automated. A more concrete, practical distinction starts with a question: What context lives only inside a human, and what is available to the machine? Everything else follows from this distinction. Whatever isn’t yet, or cannot be, measured becomes a bottleneck to safe deployment.
Unfortunately, many firms are adopting AI in ways that erode both of these critical assets, data and talent, at once. Tools that summarize documents, interactions, and context risk flattening divergent views into premature consensus, weakening the independence that makes expert talent valuable. Building on technology stacks the firm doesn’t own and letting its experts’ traces accumulate elsewhere erodes the exclusivity of its proprietary data. Together, these choices weaken the very capability that becomes more valuable as AI improves: the ability to tell whether its output is right.
In a recent paper on the economics of verification, my coauthors and I mapped tasks along two dimensions: how cheap they are to automate, and how cheap they are to verify. When a task is cheap to automate and cheap to verify, AI can be trusted to execute on its own without a human in the loop. We called this the “safe industrial” zone. But when automation is cheap while verification depends on human expertise, things get dangerous rather quickly. Here, we expect firms to deploy AI even when its outputs are likely to fall short.
This is where a firm either defends its role as a verification factory or quietly gives it up. Catching the flaw in an output that is almost right takes experts with the independence and motivation to challenge the model, and ground truth that exists nowhere else.
Right now, AI labs are doing everything they can to bring down the cost of verification across domains by aggressively purchasing specialized data and expert evaluations—even going so far as to buy the internal records of bankrupt companies. Each time they acquire richer digital traces of an economically valuable task, they turn something their models once missed into a capability they can sell.
This expansion would be less threatening if firms were actively defending their position as verification factories. Most are doing the opposite: handing the labs the traces that make their tasks verifiable while letting the experts who could still catch a bad answer drift out of the loop.
The sequence from there is predictable. One task at a time, the labs learn to verify what only in-house experts could verify before. The firm’s edge shrinks to whatever the labs still cannot measure. It becomes a thin wrapper around their intelligence, paying per token for capabilities it once owned while the labs capture a growing share of the actual value.
Don’t Automate Away Your Expertise
As firms increasingly buy inference from the same AI labs, their advantage will lie in the difference between the answers those labs provide to everyone and the conclusions they reach using proprietary data and expertise. Without that independent contribution, research suggests firms could rapidly converge toward a monoculture.
Think about what managers really do. Good managers are verifiers. Yes, they route information, but in doing so they also determine what it means, what deserves others’ attention, and what meets their quality bar. In making these calls, they often retain a significant advantage even over frontier AI. Anyone with deep experience is a fast simulator of a particular slice of reality: a “world model” made of a person, fine-tuned through years of friction with a specific technology, market, customer base, or regulator. Much of the information behind that model is tacit and unrecorded, learned as much from what failed as from what worked. And in deciding what deserves attention, managers also wield real power within the organization.
Now consider ceding many of those responsibilities to AI. An agent could turn sales meetings, support tickets, and internal discussions into a ranked list of product priorities, then use that list to assign work. An experienced engineer hesitates over one feature because it reminds them of a past failure, but they have not yet pinned down the connection. All they write is, “Not sure I’m fully comfortable with this.” The agent records strong customer demand and no concrete technical objection, leaving the hesitation out of its conclusions. A manager may still approve the budget, but the model has already shaped what they see and what they know to question.
Handing these decisions to AI is a transfer of power, not simple automation, and it can destroy what makes a firm unique. Lose your company’s edge in what it senses and prioritizes, and everything else falls apart rather quickly.
Even an AI system trained on company data is constrained by what has been measured and made available to it. It therefore risks prioritizing what the firm has already seen and codified. This can deliver measurable short-run gains by quickly routing decision-makers to existing solutions, while weakening the firm’s ability to identify and adapt to new, unstructured problems. When AI tools struggle to distinguish an emerging signal from a distraction, they can amplify what is already known at the expense of what is still taking shape, pushing the firm toward premature consensus around the wrong answer.
Deployed across a firm, these tools can degrade the very expertise they draw on. Slop, but for decisions. If companies cut experts off from the friction that keeps them learning, they may not recognize the damage until the next big discontinuity, when the models are confidently wrong and no one inside the firm is still used to overruling them.
Companies can deploy AI without flattening the independent thinking of their managers and experts. But doing so requires systems that preserve independent judgment, expose disagreements, and help employees test competing views against new data. Rather than compressing everything into a single company-wide memory, AI should help each employee build and continually refine their own version of it. Three design principles follow.
Learning loops should test decisions against real-world outcomes and feed the insights back to both models and experts. A healthy loop starts with a clear definition of success. It records the information the AI used, the actions it took, and why experts accepted, corrected, or overrode its recommendations. The firm then tests both the AI’s recommendations and the experts’ corrections against business outcomes and new evidence from the market. Those results should guide changes to evaluation criteria, knowledge, instructions, and workflows, including what the system can approve, reject, or escalate. Proposed improvements should be tested, and the lessons incorporated into the firm’s institutional memory and shared with experts so their judgment improves too. Fine-tuning the model is one possible ingredient in this broader learning system. The loop should also flag when generation outpaces meaningful verification, a warning that the firm may be shipping output it cannot yet trust. The pattern is already visible in software: as AI adoption has accelerated, developer throughput has risen, but so have bugs and production incidents.
Agents should surface relevant knowledge, not flatten it. AI agents in a firm’s communication channels should help people articulate tacit knowledge with as little friction as possible. Because agents often lack that context, they should preserve disagreements and unresolved concerns rather than compress them into premature consensus. Experts’ objections can reveal a gap between their “weights” and the firm’s telemetry, bringing previously unrecorded experience into the system. Eliciting those objections helps the firm detect new developments and identify what still needs to be measured. Preserving them also creates an invaluable record for a postmortem, when the firm must trace a failure back to the assumptions and decisions that produced it.
World models should be supporting infrastructure, not decision-makers. Perhaps most importantly, an effective company “world model” should continuously spot hidden assumptions and identify where accurate data is still missing, then encourage employees to seek real-world friction to surface more robust evidence. Only through repeated contact with reality can the AI model be genuinely up to date. Sometimes, this can be as simple as connecting people who hold different pieces of relevant context so they can compare notes and make explicit what would otherwise remain uncodified.
These countermeasures can help ensure that, as AI becomes embedded in every company interaction, it does not inadvertently force everyone to operate from a narrow, model-compressed version of reality. By amplifying what makes the company and its experts unique, a truly customized model can become one of the firm’s most valuable strategic assets. The goal is to keep the firm’s internal market of ideas alive: treat competing views as hypotheses, prompt individuals to test them against real-world outcomes, and feed the resulting evidence and lessons back into the firm’s institutional memories and expertise. Each turn of the loop makes the firm a better verification factory.
But the ultimate design test is how AI supports consequential decisions under fundamental uncertainty. A company-wide model that decides what information matters, resolves conflicting interpretations, and allocates attention and resources risks acting like a central planner and run into Friedrich Hayek’s central critique of planned economies: the tacit knowledge it needs is dispersed among people, tied to particular circumstances, and often never written down. The model cannot reliably outperform people if it lacks critical knowledge they hold. It should help them make that knowledge explicit, test their assumptions, and learn from the results. Otherwise, it risks eroding the expertise it depends on.
Owning the Weights of Production
A firm can get every component of its verification factory exactly right and still fail to capture the value it generates. Capturing value requires owning the traces and feedback generated by the loop.
Today, leading AI labs are courting organizations with tools that reach deep into their document repositories, communication channels, and codebases. They also tempt employees with a Faustian bargain: Let the tools record how you use your computer while you work, and in exchange we will automate your work for you.
Companies may be reassured by AI labs’ “no-training” commitments for their prompts and outputs. But they also need to know who can retain and reuse the operational traces surrounding that work: which tools an agent used and in what order, how it failed, and when a human overruled it. This form of “ambient AI” captures the distinctive ways employees handle exceptions, make decisions, and prioritize information. Taken together, those traces can provide a blueprint for replicating parts of a firm’s verification engine. If providers can reuse them to build their own offerings, firms risk handing over the knowledge that gives them an edge. The stakes rise as agents become more autonomous: broader permissions and deeper context put more of the firm’s secret sauce within the provider’s reach.
Companies can protect that edge in a few ways.
Own the loop. Protecting your firm from suddenly finding itself competing with the AI labs’ models starts with owning the learning loop and controlling how its data is used. The fine-grained, proprietary measurements it generates capture how your experts verify outputs and correct mistakes. Letting the labs reuse those measurements helps them learn to perform that verification themselves. Start by checking exactly what your provider can retain and reuse, and what its no-training and zero-data-retention commitments do and do not cover.
Embrace open-weight models. Open-weight models offer a practical way to retain greater control. They make their underlying parameters available, allowing firms to download, run, and adapt the models themselves. Firms can deploy them on their own infrastructure or through a partner whose business is providing infrastructure rather than building AI products that could compete with the firm. In either case, the arrangement should keep proprietary information, operational traces, and improvements under the firm’s control. A growing coalition is working to build a thriving AI ecosystem that gives firms an alternative to relying on the closed labs. By August, more than 120 organizations had joined the Open Secure AI Alliance to advance open models and open-source tools for AI security. Open-weight models initially attracted enterprise interest for their potential to lower costs. Their strategic appeal now extends beyond those savings to control: ensuring that the data, feedback, and improvements generated through the firm’s use of AI remain assets of the firm.
Find the right partner. Owning the loop does not require building everything yourself. Working with Thinking Machines, hedge fund Bridgewater used its proprietary data to customize an open-weight model that outperforms frontier closed models on the financial tasks tested. The arrangement serves both parties. Bridgewater strengthens its own capabilities without feeding its intellectual property back into the lab’s base models. Thinking Machines can offer training infrastructure without asking customers to surrender the knowledge that gives them an edge. The partner supplies the machine intelligence infrastructure; the firm retains the learning.
…
Firms have always been verification factories, even when verification was inextricably bundled with execution. Now that AI has unbundled the two, the companies that thrive will not only build best-in-class verification systems but also keep every trace that improves them inside their boundaries.
As you prepare for this shift, ask yourself: when one of your experts inevitably overrules the AI, is that correction recorded? And do you own the system that captures it? If the answer to the first is no, you do not have a verification factory yet. If the answer to the second is no, you are building someone else’s.
A version of this article appeared on Harvard Business Review.