AI Scales Production. Institutions Must Scale Verification.
Note | CT Innovates | September 2026
I recently listened to a Chris Hayes podcast that was ostensibly about whether the AI boom is about to collapse. His guest, Ed Zitron, is one of the technology industry’s more aggressive AI skeptics, and much of their conversation concerned the familiar questions surrounding the boom: extraordinary infrastructure spending, the cost of computing, uncertain business models, and whether the enormous investment flowing into artificial intelligence can produce returns commensurate with the expectations now attached to it.
All of that was interesting. But the exchange I kept returning to afterward had almost nothing to do with Nvidia’s valuation or the possibility of an AI bubble. It concerned something much more ordinary: reading documents.
Hayes made the reasonable case that generative AI differs from some previous technology hype cycles because it already does useful things. Give a capable system a large collection of documents and it can search them, organize them, compare them, and produce a synthesis in a fraction of the time a person might require.
Zitron raised the uncomfortable question that follows. How does someone know whether the synthesis is good without understanding the material from which it was produced?
Hayes immediately saw the problem: quality control.
That brief exchange points toward an institutional constraint that may become increasingly important as AI improves. The technology is acquiring the ability to produce cognitive work at a rate that human beings and organizations may have difficulty evaluating. Production capacity can rise dramatically while verification capacity improves much more slowly.
And once those two capacities begin to diverge, the conventional productivity story becomes incomplete.
The constraint is moving
Much of the current discussion of AI productivity is organized around output. How many reports can an analyst produce? How much code can a developer generate? How many documents can a system examine? How much faster can an organization draft emails, presentations, summaries, analyses, recommendations, and responses?
These measures capture something real. AI can make many forms of cognitive production dramatically cheaper.
Organizations, however, have little use for output simply because it exists. A summary becomes valuable when someone can rely on it. An analysis becomes valuable when its reasoning survives scrutiny. A recommendation becomes useful when a person with the appropriate knowledge and authority is prepared to act on it.
Getting from generated work to trusted work often requires its own demanding sequence of activities. Someone may have to inspect the underlying evidence, notice omitted context, test whether a conclusion follows from its sources, recognize something implausible, understand the authority carried by a particular document, evaluate how uncertainty has been represented, and determine whether the resulting work is adequate for the decision at hand.
AI may compress the time required to produce the first draft of that work enormously. The burden of judgment does not necessarily shrink at the same rate.
Consider an organization that once had the capacity to produce five substantive analyses a week. With AI, it can suddenly generate twenty. Calling that a fourfold productivity increase seems reasonable until we discover that the organization has enough knowledgeable human attention to review only eight.
The institution has acquired twenty analyses’ worth of production capacity. But its usable capacity is much closer to eight.
This suggests a distinction worth making. Gross AI capacity describes how much additional work the technology allows an organization to produce. Net institutional capacity describes how much additional work the organization can responsibly understand, verify, approve, and put to use.
For institutions, the second number is ultimately the consequential one.
The distinction becomes especially important in civic and public settings, where the cost of error varies enormously. A weak first draft of an internal memo can be corrected with little consequence. An inaccurate explanation of a public decision, a mistaken benefits determination, a flawed personnel assessment, or an unsupported claim communicated to residents occupies a very different category.
As the consequence of the work increases, the institution needs more capacity for judgment.
When the human in the loop cannot keep up
This also complicates one of the most familiar assurances in responsible-AI discussions: there will be a human in the loop.
The phrase describes an important principle, but its meaning depends heavily on the conditions surrounding that human.
Imagine an employee reviewing ten AI-generated recommendations. The person may have time to inspect the evidence, question unusual conclusions, compare cases, and reject weak work. Increase the volume to a thousand recommendations and the formal workflow may remain unchanged: the employee is still technically responsible for review. In practice, something very different has happened.
Attention becomes rationed. Review becomes faster and more superficial. The likelihood that generated conclusions will be accepted because they appear polished begins to rise. Eventually the human reviewer can become a ceremonial layer of accountability attached to a machine-scale production system.
This is why meaningful human authority requires more than placing a person somewhere in the process. The CT Innovates’ AI for the Public Good framework argues that people need sufficient knowledge, access to evidence, discretion, time, and manageable scale to evaluate consequential machine-produced work. Human authority becomes fragile when the quantity or complexity of output exceeds the reviewer’s practical ability to understand it.
The design question for an AI-enabled workflow therefore includes an issue that conventional productivity measures rarely capture: How much additional trustworthy work can this institution actually absorb?
That question may prove more important than the amount of material the system can generate.
Judgment may become the scarce resource
Organizations have historically operated under many forms of production constraint. Research required hours. Reading required hours. Drafting required hours. Comparing documents, organizing evidence, searching archives, and preparing analyses all consumed scarce human time.
AI can loosen many of those constraints simultaneously. That is one reason the technology can create genuine productivity gains.
Yet removing one constraint tends to reveal another.
If analysis becomes cheap while evaluating analysis remains difficult, expert judgment grows more valuable. If drafting becomes nearly instantaneous while fact-checking remains laborious, verification absorbs a larger share of the workflow. If an organization can generate more recommendations than its leaders can thoughtfully consider, attention begins to determine the real ceiling on productivity.
The bottleneck has moved downstream.
A more useful model of AI productivity might therefore look something like this: AI Production Capacity → Human Verification → Trusted Work → Decision or Action
Every stage matters because the useful capacity of the system is ultimately constrained by the stage least able to scale.
This has practical implications for institutional design. Organizations investing heavily in AI tools may also need to invest in expertise, source access, review processes, auditability, clearer approval rights, better information architecture, and explicit limits on how much AI-produced work any one person can meaningfully supervise.
Otherwise, a strange form of abundance becomes possible: organizations surrounded by more analysis, more summaries, more recommendations, and more information than ever before while remaining unable to determine confidently what deserves their trust.
The productivity measure that matters
None of this diminishes AI’s productive potential. It makes the definition of productivity more demanding.
The useful question is not simply how much additional material AI allows an institution to produce. Leaders need to understand how much reliable, usable, human-owned capacity the technology creates once verification and judgment are included.
That distinction is embedded in a principle we use throughout Civic Intelligence: AI prepares the work. Accountable humans verify, approve, decide, and own it.
Technological progress is making the first half of that sentence easier with remarkable speed. The harder institutional challenge may increasingly reside in the second.
Organizations that recognize this early will design verification as a capacity in its own right. They will treat knowledgeable attention as infrastructure, build review requirements into workflows, and measure productivity closer to the point where information becomes trusted action.
AI can scale production extraordinarily well. The institutions that capture the greatest value from it may be those that learn how to scale judgment alongside it.
Source note: This Civic Intelligence Note was inspired by Chris Hayes’s conversation with Ed Zitron on Why Is This Happening?, particularly their discussion of document analysis, hallucination risk, coding, and the continuing quality-control problem created when AI-produced work still requires knowledgeable human review. The broader argument develops the Civic Intelligence principle that accountable human review is an operating stage of AI-assisted work rather than a disclaimer attached after the fact.