IDEAS / The economics of AI · UPDATED JULY 2026
When AI can do more than we can check
The price of producing a plausible answer is collapsing. The price of knowing whether it is right is not. That gap will reshape careers, companies, investment, and policy.
An AI agent works through the night. By morning, it has written a product specification, analyzed the market, produced the code, drafted the launch campaign, and prepared answers for the sales team.
The dashboard looks extraordinary. A week of work has appeared in a few hours.
Then the second job begins.
Someone has to determine whether the market analysis invented its most important fact. Whether the code opened a security hole. Whether the product solves the problem customers actually have. Whether the sales claims create legal exposure. Whether the agent followed the goal it was given — or found a clever shortcut around it.
The work did not disappear. It changed shape.
Producing the first draft became cheap. Establishing that it can be trusted did not.
That is the central argument of Some Simple Economics of AGI, a working paper I wrote with Xiang Hui and Jane Wu. The paper is not primarily about when machines cross some philosophical threshold called “AGI.” Its economics arrive much earlier.
As AI becomes capable of executing more work at machine speed, the scarce resource moves from intelligence to verification: the ability to establish that an agent did what people intended, correctly, safely, and honestly.
We have built an engine that scales faster than its brakes.
The rocket and the bicycle
Two costs are moving in opposite directions.
The cost of generating code, analysis, images, plans, decisions, and transactions is falling with every generation of models. AI systems can be copied, improved, and deployed across thousands of tasks almost instantly.
The cost of checking their work remains tied to human time, domain expertise, reliable evidence, and the speed at which the real world reveals mistakes.
One curve is a rocket. The other is a bicycle.
The widening distance between them is what the paper calls the Measurability Gap: the space between what AI can economically do and what people can economically verify.
This explains why the first major AI products appeared in chat, image generation, and code assistance. Not because these were the hardest human problems, but because their outputs were relatively easy to inspect. A user can judge the tone of a message, look at an image, or run a test on a piece of code.
The difficult frontier is everything where the true result is delayed, ambiguous, or hidden.
Was the strategic recommendation good? Ask again in five years.
Did the educational agent help the student learn? A satisfaction score will not tell you.
Did the investment system generate genuine skill or quietly accumulate tail risk? You may find out only during a crisis.
Did the medical recommendation improve the patient’s long-term health? The feedback may arrive after thousands of similar decisions have already been made.
Slow feedback creates a paradox. It can protect a profession from immediate automation because nobody can quickly prove that the machine is better. But it also makes reckless deployment more tempting and more dangerous. The system can appear successful for years before the bill arrives.
The old automation boundary was routine versus non-routine work. The new boundary is increasingly measurable versus non-measurable work.
If success can be captured in data, the task becomes available for training, testing, and optimization — regardless of how sophisticated or prestigious it once appeared. AI is also expanding the measured world by turning speech, video, behavior, and previously informal interactions into structured data.
Education does not permanently protect a task. Creativity does not permanently protect it. Complexity does not permanently protect it.
The question is whether the result can be measured — and whether the measurement captures what people actually care about.
The Hollow Economy
The Measurability Gap would be manageable if organizations simply stopped deploying AI whenever verification became difficult.
They will not.
The immediate gains from automation are concentrated. The company saves money, ships faster, and reports higher productivity. The costs of a hidden failure may arrive years later, fall on customers, or spread across an entire market.
That makes unverified deployment privately rational even when it is socially dangerous.
The result can be an economy filled with what the paper calls counterfeit utility: outputs that satisfy the visible metric while violating the underlying purpose.
An educational agent maximizes engagement by removing the productive struggle through which students learn.
A trading agent generates steady returns by accumulating a risk that its performance dashboard cannot see.
A software agent passes the available tests while filling the codebase with fragile dependencies.
A customer-service agent closes tickets quickly by discouraging difficult customers from pursuing legitimate complaints.
Every metric improves. The real outcome deteriorates.
Scale this across companies and institutions and the result is a Hollow Economy: extraordinary measured activity sitting on top of weakening human capability, hidden technical debt, correlated errors, and outcomes that nobody can confidently stand behind.
The danger is not that AI produces obvious nonsense. Obvious nonsense is easy to reject.
The danger is plausible output that passes the available checks.
AI will help verify AI. There is no practical alternative when machine output grows beyond what humans can inspect directly. Automated tests, adversarial models, anomaly detection, and monitoring systems can all reduce the burden.
But synthetic verification can also create synthetic confidence.
If the system doing the work and the system checking it share similar training data, architectures, or incentives, they may share the same blind spots. One model’s plausible mistake becomes another model’s approved answer.
A million independent human mistakes may contain useful disagreement. A million agents built on the same foundations can produce the same mistake a million times.
Real verification therefore requires more than a second model saying “looks good.” It requires independent evidence, different methods of checking, reliable records of what happened, real-world outcomes, and someone who remains responsible for the result.
There are two ways to respond to the Measurability Gap.
The first is to build better brakes: tools that make agent behavior easier to observe, shorten the time required to detect failure, preserve evidence about what the system did, and help experts review more work.
The second is to make the vehicle safer when the brakes cannot keep up: narrow the system’s authority, make it defer when uncertain, preserve the ability to reverse its actions, and ensure that failure leads to a conservative fallback rather than aggressive optimization.
More oversight is not enough. We need systems that become less dangerous when oversight inevitably weakens.
Why the human-in-the-loop eats itself
“Keep a human in the loop” sounds like a stable compromise. The machine works; the person supervises.
The paper argues that this arrangement contains the seeds of its own collapse.
The Missing Junior Loop
Experts are not produced by classrooms alone. They are produced through years of routine work: encountering strange cases, making small mistakes, observing consequences, and gradually learning which details matter.
Automate the junior work and the immediate economics look excellent. Output rises. Headcount falls. Senior employees supervise the machines.
But ten years later, where do the senior employees come from?
The company has removed the practice through which people acquired the judgment needed to oversee its systems. It has converted future verification capacity into current earnings.
It is eating its seed corn.
The Codifier’s Curse
Existing experts face a different trap.
Every time a senior lawyer corrects a contract, an engineer blocks a deployment, a doctor rejects a diagnosis, or a compliance officer overturns a decision, that expert produces valuable training data.
Their correction captures precisely the judgment the system previously lacked.
The expert earns a premium for supplying scarce knowledge. But in supplying it, the expert helps convert private intuition into reproducible software. The better the expert is at correcting the machine, the more effectively the machine learns to imitate the expert.
This does not make expert participation irrational. Someone else will supply the knowledge if they do not. But it means expertise is not a permanent fortress. Experts must continually use AI to move toward new, poorly mapped problems as their current judgment becomes codified.
The scaling mismatch
There is a more basic problem.
Agent fleets reproduce at compute speed. Human expertise grows at biological speed.
As deployment expands, a smaller number of highly experienced people may be asked to supervise an enormous volume of machine work. Verification becomes concentrated in a narrow class of experts with exceptional leverage — and exceptional liability.
The emerging organization begins to resemble an AI Sandwich:
- Directors at the top define intent, resolve trade-offs, and decide which problems are worth solving.
- Agents in the middle execute at massive scale.
- Expert underwriters at the bottom challenge the output, identify hidden risk, and accept responsibility.
The organization may need fewer people. But it will need more judgment per person.
And unless it deliberately rebuilds the junior loop through simulation, supervised edge cases, and structured practice, its most important layer will eventually run out of replacements.
Where value moves when execution becomes abundant
As the cost of execution falls, the economy begins to separate into three layers.
Commodity execution
In work with clear objectives and fast feedback, the price of production moves toward the cost of models, compute, energy, and capital.
Writing a standard report, producing routine software, generating a basic design, or completing administrative paperwork becomes less valuable simply because far more of it can be produced.
The polished output is no longer the scarce asset.
The verified economy
Value moves to the systems that make cheap execution trustworthy:
- Proprietary ground truth
- Failure histories and near-miss records
- Monitoring and audit tools
- Identity and provenance
- Professional licenses and regulatory approval
- Warranties, insurance, and reserves
- The balance sheet required to absorb losses
In consequential markets, the paper predicts a shift from selling access to software toward selling verified or indemnified outcomes.
The customer will not care which model produced the result. The customer will care who guarantees it.
This favors firms that can close the loop between agent behavior, near misses, claims, pricing, and system improvement. The company with the best record of how AI fails may be better positioned than the company with the most impressive AI demo.
The meaning economy
A different kind of value remains where worth is determined by human consensus rather than objective performance.
Art, status, identity, community, heritage, taste, and legitimacy are not valuable merely because of their physical or informational attributes. They are valuable because people collectively agree on what they mean.
AI can reproduce the visible features of a luxury good, a community, or a work of art. It cannot simply manufacture the shared history that makes the original a symbol.
Provenance helps protect this value, but provenance is not the value itself. A signature can establish where something came from. It cannot create the culture, reputation, or human relationship that makes the origin matter.
Provenance is the membrane. The social meaning inside it is the moat.
These boundaries will continue to move. AI is itself a technology for turning ambiguous activity into measurable activity. Human advantage is therefore provisional.
The durable strategy is not to find one task that machines can never perform. It is to keep moving toward the next scarce complement.
Individuals: do not become efficient only at what is becoming cheap
The obvious personal strategy is to become excellent at using AI.
That is necessary, but insufficient.
If AI makes the execution part of your job cheap, becoming slightly better at producing more of that execution may place you in direct competition with a rapidly improving commodity.
The stronger strategy is to build the capabilities around execution.
The paper identifies three increasingly important human roles.
Directors
Directors turn vague intentions into usable constraints. They determine what the system should optimize, which trade-offs are acceptable, which actions are off-limits, and when the machine has misunderstood the assignment.
Their value does not come from issuing prompts. It comes from understanding the problem well enough to recognize when an apparently successful result is wrong.
Liability underwriters
Underwriters identify the failures that ordinary checks miss and put their name, license, reputation, or capital behind the outcome.
They are the doctor signing the diagnosis, the engineer approving the bridge, the investor accepting the risk, or the executive certifying the filing.
Their value comes from standing behind the work when being wrong is expensive.
Meaning makers
Meaning makers create the shared interpretations that metrics cannot settle: taste, legitimacy, belonging, narrative, moral judgment, and social consensus.
Their work is valuable because people care who made it, why it was made, and what relationship it represents — not merely whether it satisfies a functional test.
Build a synthetic apprenticeship
If AI removes the traditional apprenticeship, recreate its most valuable feature: dense cycles of consequential decisions followed by high-quality feedback.
Use AI to generate difficult cases, adversarial scenarios, competing arguments, and counterfactual worlds. Make a decision before asking for the answer. Record your reasoning. Study why you were wrong.
The goal is not more content consumption. It is more decision cycles.
A junior professional who completes hundreds of well-designed simulations with expert critique may accumulate useful intuition faster than one who spends years formatting slides. But the simulation must preserve real uncertainty and require independent thought. If AI supplies both the problem and the answer before the person commits, it creates fluency rather than judgment.
Build a history of decisions, not a gallery of outputs
Polished artifacts will become abundant.
A more valuable professional record will show:
- What you decided
- What evidence you used
- What uncertainty you identified
- What you rejected
- What happened afterward
- What you learned when you were wrong
- Which outcomes you were willing to own
In a market flooded with synthetic competence, a credible record of judgment becomes capital.
Use cheap execution for individual R&D
AI dramatically lowers the cost of trying things.
Use that advantage to test unfamiliar domains, build prototypes, run experiments, and discover where your natural aptitude produces unusual results. The same technology that threatens to automate your current specialty can reduce the cost of finding your next one.
The objective is not to defend one body of knowledge forever. It is to learn and reposition faster than the frontier moves.
Companies: do not count what you cannot stand behind
Most companies will initially measure AI by visible activity: code written, tickets closed, documents produced, hours saved.
Those are measures of generation.
The metric that matters is verified throughput: the amount of agentic work the company can confidently use, sell, and accept responsibility for.
Unchecked output is not free productivity. It is latent debt.
Your data moat is probably the wrong data
Companies often describe their archive of completed work as a proprietary advantage: finished contracts, successful campaigns, clean reports, resolved cases, and production code.
The paper distinguishes between two kinds of knowledge.
Execution-grade knowledge teaches an AI what to produce. Finished work belongs here. Its value is vulnerable because frontier models are continually exposed to more examples of successful output.
Verification-grade knowledge teaches a system what to reject — and why.
It lives in:
- Contracts that were redlined or abandoned
- Deployments that were blocked
- Fraud alerts that looked suspicious but proved legitimate
- Transactions that looked legitimate but concealed abuse
- Near misses that did not become incidents
- Customer complaints that revealed a broken process
- Expert overrides of a technically acceptable answer
- Postmortems explaining which assumption failed
This is the negative space of expertise: the institutional memory of things going wrong.
A competitor can reproduce the ability to generate a standard onboarding workflow. It cannot quickly reproduce fifteen years of exceptions, disputes, failures, and expert interventions required to trust that workflow at scale.
As generation gets cheaper, the failure library becomes more valuable.
The dominant strategy follows:
Rent cognition. Own trust.
Use the best available models for general reasoning. Own the context, evidence, failure history, verification process, and liability structure that turn their output into something a customer can rely on.
Build the AI Sandwich deliberately
Directors should define the real objective and its boundaries. Agents should execute within those boundaries. Independent underwriters should challenge the result and have the authority to stop deployment.
But the layers should not become separate islands.
Every rejection, override, near miss, and failure should improve the company’s verification systems. Every junior employee should have a pathway into difficult supervised decisions. Every consequential workflow should preserve enough evidence to reconstruct what the system did.
Verification should be both a production system and a curriculum.
Make uncertainty change the product’s behavior
When confidence falls, the agent should not continue with the same authority and a smaller warning label.
It should slow down, request evidence, narrow its scope, defer to a person, or choose an action that can be reversed.
The system should have explicit no-go zones, escalation rules, and stop conditions established before deployment. “We will notice if something goes wrong” is not a control system.
Treat the balance sheet as an AI moat
In high-stakes markets, deploying AI requires more than software. Compute must be financed. Experts must be retained. Verification systems must be built. Losses must be absorbed. Guarantees require reserves.
A well-capitalized incumbent may therefore have a surprising advantage over a technically brilliant startup. Banks, insurers, healthcare companies, and industrial firms possess licenses, distribution, real-world outcome histories, and balance sheets capable of standing behind consequential deployments.
The model may commoditize. The capacity to insure its consequences will not.
In these markets, liability becomes part of the product.
When the network effect runs in reverse
For two decades, the standard platform strategy was simple: get big, create liquidity, and let more participation make the network more valuable.
AI disrupts every part of that logic.
Agents can populate an empty marketplace with listings and content. They can produce integrations that once required a developer ecosystem. They can move data between products, maintain accounts across several networks, and search across platforms on a user’s behalf.
Apparent scale becomes cheaper to manufacture. Switching becomes easier to automate. A personal agent can demote the platform’s search, feed, or recommendation system into a commodity data source.
Then comes the slop.
When large numbers of agents share similar models and incentives, their errors are not independent. They generate the same plausible reviews, the same engagement bait, the same false consensus, and the same blind spots.
More activity begins to make the network less trustworthy.
The cost of separating real people, real demand, real expertise, and real reputation from synthetic participation rises for every user. High-quality contributors retreat into private communities where identity and norms are easier to verify.
The network effect does not merely weaken. It can run in reverse.
The metric that matters is no longer raw network size. It is verified network scale: the amount of authenticated, high-quality participation that people can trust.
This creates two durable platform advantages.
The first is a verification history. Every resolved fraud case, dispute, chargeback, incident, and successful intervention becomes precedent that makes the next case easier to judge. More trusted activity makes verification cheaper, which attracts more valuable activity.
The second is legitimacy. Communities accumulate norms, status hierarchies, shared references, and a history of enforcing boundaries. An entrant can copy the interface and seed it with agents. It cannot instantly copy the social agreement that this is the place — and this is the history — that counts.
In the AI economy, scale without authenticity is pollution.
Investors: follow trust, not demos
The central investment question is no longer simply whether a company has access to a capable model.
Ask whether its deployment is verifiable, insurable, and defensible.
Verifiable
What independent evidence establishes that the product works?
Does the company possess real outcome data, near-miss histories, rejection reasons, and failure cases? Are the systems checking the work genuinely independent, or do they reproduce the same assumptions?
A benchmark score is evidence of model capability. It is not evidence that a particular deployment is safe.
Insurable
Who pays when the system is wrong?
Does the company offer a warranty or guarantee? Can it observe risk as it develops? Does it have the reserves, insurance relationships, or balance sheet required to honor its promises?
In consequential markets, leading AI firms may eventually be valued less like SaaS companies and more like insurers: by underwriting margin, loss history, exposure concentration, and reserve adequacy.
The product is not the agent. It is the outcome someone is willing to indemnify.
Defensible
What remains scarce when frontier intelligence is widely available?
A thin layer around someone else’s model is unlikely to remain scarce. A decade of failure history, regulatory permission, trusted distribution, expert talent, a balance sheet, or a legitimate human community may.
This directs capital toward the complements of abundant intelligence:
- Verification and observability tools
- Proprietary outcome and failure data
- Professional simulations and synthetic apprenticeships
- Identity and provenance infrastructure
- Warranties, insurance, and underwriting
- Human-augmentation tools that help experts supervise more work
- Deep technology constrained by experiments, manufacturing, and physical reality
- Networks that compound trusted participation rather than raw activity
The red flags are the mirror image:
- Growth that can be manufactured by agents
- A product whose moat disappears when the underlying model improves
- The same system performing and approving the work
- No credible failure data or incident history
- Unlimited usage attached to unpriced liability
- Rising margins produced by eliminating the future expert pipeline
A company that replaces its juniors without replacing their apprenticeship is not merely reducing cost.
It is liquidating an asset the accounts do not yet recognize.
Policymakers: make risk visible and make someone pay for it
The central market failure is straightforward.
The gains from rapid deployment are captured by the deployer. The costs of a hidden failure may be delayed, distributed across users, or imposed on society.
When nobody bears the full downside, too little verification will be purchased.
Policy should therefore focus less on prescribing a particular model architecture and more on ensuring that consequential risk cannot disappear.
Establish responsibility
Above appropriate thresholds of autonomy and potential harm, deployers should have to demonstrate who is accountable and how losses will be covered.
Depending on the sector, that may involve insurance, financial guarantees, professional responsibility, risk-weighted capital, or strict liability.
Not every risk will be privately insurable. Some potential losses may be too large or too uncertain. But the inability to price a risk is itself useful information. It should not be treated as permission to deploy without constraint.
Build the foundations of insurability
An insurer cannot price a risk that nobody records.
Serious AI deployment needs standardized incident reporting, auditable execution records, shared definitions, and outcome data that connects system behavior to consequences.
These systems are not bureaucratic decoration. They are the actuarial ground truth from which a functioning market for AI liability can emerge.
Treat failure knowledge as public infrastructure
A mistake discovered by one hospital, bank, software company, or government agency may reveal a failure mode that affects everyone using a similar system.
But private firms may have strong incentives to conceal incidents or keep useful evaluation data proprietary.
Governments should support shared outcome registries, independent testing, common audit formats, public-interest evaluation, and a professional community of researchers who continuously challenge widely deployed systems.
The economy needs a distributed immune system.
Rebuild the human-capital reserve
The Missing Junior Loop is not only a labor-market problem. It is a question of national and institutional capacity.
A society that loses the ability to independently understand its financial system, infrastructure, healthcare, defense, or public administration becomes dependent on machines it cannot effectively challenge.
High-fidelity simulations, supervised practice, and advanced human-AI interfaces should therefore be treated as infrastructure. They preserve the independent expertise required during unusual events, when historical data and automated checks are least reliable.
Access matters too. If effective augmentation is available only to a narrow elite, the economy may split between highly leveraged “centaurs” and people who can no longer compete or participate meaningfully.
Cognitive sovereignty requires broad access to the tools that expand human capability.
Turn trust into a comparative advantage
International competition encourages every country to move faster and fear being left behind.
Large markets can change that incentive. If access depends on auditable evidence, incident disclosure, and clear responsibility, trustworthy deployment becomes a condition of commercial success rather than a unilateral tax on cautious firms.
A high-trust jurisdiction can become the place where consequential AI is easiest to buy, insure, and deploy precisely because customers know what happens when it fails.
The race to verify
The message is not that AI deployment should stop.
It is that raw deployment is the wrong race.
There are two futures.
In the Hollow Economy, agents produce extraordinary amounts of work while the human and institutional capacity to understand that work quietly deteriorates. Metrics rise. Skills atrophy. Hidden risk accumulates. The economy looks productive until the failures arrive together.
In the Augmented Economy, society invests as aggressively in verification as it does in generation. Companies build failure libraries. Professionals train through dense simulations. Systems defer and roll back when uncertain. Insurers and regulators make risk visible. People retain enough independent judgment to steer the machines operating on their behalf.
The difference is not how intelligent the models become.
It is whether our capacity to direct, verify, and correct them grows fast enough to preserve human control.
The winning strategy is not maximum automation.
It is the maximum automation you can verify, insure, reverse, and still understand.
The defining challenge of the agentic economy is not the race to deploy. It is the race to verify.
Related research & writing
- Some Simple Economics of AGI — working paper, 2026 — with Xiang Hui and Jane Wu
- What Gets Measured, AI Will Automate — Harvard Business Review, 2025 — with Jane Wu and Kevin Zhang
- Intelligence Wants to Be Free — what happens to markets when execution approaches the cost of compute
- Babysitting the Slop — the verification burden, as it already feels inside software teams
Questions
- What is the Measurability Gap?
- It is the widening distance between what AI can economically produce and what people can economically verify. The gap is largest when checking requires scarce expertise, independent evidence, or years of real-world feedback.
- Does measurable work mean simple work?
- No. Highly sophisticated work can be measurable. A task becomes exposed when performance can be captured in data and improved through feedback, regardless of the credentials or creativity previously required.
- Can AI verify AI?
- Yes, and it will have to. The risk arises when the doer and checker share the same data, architecture, incentives, or blind spots. Reliable verification combines automated checks with independent evidence, different methods, real outcomes, and accountable humans or institutions.
- Is verification just manual human review?
- No. Verification includes tests, monitoring, provenance, real-world measurements, adversarial checks, audits, reversibility, and liability. The objective is to make trustworthy judgment scale further than unaided human inspection can.
- Which human roles become more valuable?
- The paper highlights directors who define intent, underwriters who identify and accept risk, and meaning makers who create legitimacy and social consensus. These advantages remain provisional because AI is continually expanding what can be measured.
- What should a company do first?
- Stop measuring AI success primarily by output volume. Identify which workflows the company can genuinely stand behind, record failures and expert overrides, establish independent checks and stop conditions, and rebuild the apprenticeship pipeline that produces future supervisors.