Cognitive Outsourcing at Work: Why AI Productivity Gains Are Outrunning Workplace Governance
Authored On
Modified
AI use is already widespread Better output does not prove learning Assessment must verify unaided skill

Ninety-four percent of software leaders say AI has produced a meaningful productivity gain. Just 37 percent say their organization has anything resembling formal governance over that AI-assisted work. The gap is the clearest evidence yet that a problem people mostly associate with students and chatbots has already moved well past the classroom. Cognitive outsourcing, handing off not just information retrieval but the actual labor of forming a judgment, now shows up at every level of the corporate ladder, from the newest analyst to the chief executive's desk. Most commentary treats this as a discipline problem, something better habits or more skepticism could fix. It isn't. It's structural: organizations have adopted a genuinely powerful capability without building anything that would make careful verification the default rather than a private act of willpower. The result is productivity that looks excellent on a quarterly dashboard, right up until the judgment underneath it gives out and someone finally notices.
The Scale of the Problem and Why It Matters Now
None of this rests on speculation. An analysis of roughly one hundred thousand real conversations on Anthropic's Claude platform found that the tasks people bring to the model would typically take about ninety minutes to do unassisted and that the model cut completion time by nearly 80 percent. Extrapolate that across the economy and current-generation systems could add nearly two percentage points to annual United States labor productivity growth over the next decade, almost double the recent trend. Software development makes the case most concretely: engineers contribute more to AI-linked productivity gains than any other occupation and a mid-2026 survey of applications and engineering leaders found 94 percent now reporting meaningful gains, with more than 4 in 5 also noting a measurable drop in shipped defects.
That same survey turned up something that complicates the good news. Among developers actually using AI in the build phase, only 37 percent called their organization's governance of that use formal or better and 2 in 3 agreed that AI-written code needs more testing than code a person wrote alone. Security worries, output-quality concerns and basic skills gaps topped the list of adoption barriers. This is what commentators have started calling cognitive outsourcing; a term the technology writer Enrique Dans frames as a natural extension of a decades-long habit that began with search engines replacing memory and GPS replacing spatial reasoning. MIT Sloan's Eric So has a name for the pull behind it: "AI gravity," the mounting, largely unconscious pressure to route thinking through a system that simply does it faster. Three recent findings make the case that this is urgent rather than merely interesting. The International Labour Organization's 2026 review found the striking task-level gains researchers keep documenting have not yet shown up in aggregate output. Stanford researchers tracking payroll data found a roughly 16 percent relative decline in employment for workers aged 22 to 25 in AI-exposed occupations, even as employment for older workers in the same fields held steady. And a widely discussed Brookings analysis argues today's visible gains are disproportionately produced by a generation of experts who built their judgment before these tools existed, with borrowed expertise not being renewed at anywhere near the rate it's being drawn down.

Beneficial Offloading Versus Harmful Outsourcing
The behavioral evidence backs this up well beyond anecdote at this point. A preliminary MIT Media Lab study found that 83 percent of people who had written an essay with a chatbot's help couldn't quote a single sentence from their own work minutes after submitting it. A separate survey of knowledge workers by Microsoft Research and Carnegie Mellon University found something even more telling: confidence in an AI system tracked negatively with self-reported critical-evaluation effort across five of the six cognitive activities researchers tracked and the effect was strongest right at the evaluative stage, the moment someone decides whether an answer is actually correct. The researchers flagged a self-reinforcing loop here: people tend to skip scrutiny exactly when they lack the expertise to apply it, which is also exactly when scrutiny would matter most.

It's worth taking the optimistic counterargument seriously, because it isn't a weak one. Skeptics point out that the calculator, the search engine and Wikipedia all drew similar warnings and that AI genuinely does democratize certain kinds of work. A National Bureau of Economic Research study of more than five thousand customer service agents found AI assistance raised productivity for novice agents by roughly a third while barely moving output for the most experienced ones. Fair enough. But the calculator comparison breaks down once applied to a real workplace. A calculator is reliably correct inside its domain; a language model can be confidently, fluently wrong and only a user's own domain knowledge will ever catch that. A calculator carries out an operation someone has already decided on; a language model often decides the framing itself, quietly doing the part of the job that used to signal expertise. Consultants at Boston Consulting Group who used AI on a task just outside its capability frontier saw performance drop roughly nineteen percentage points relative to working unaided, because nothing in the tool's confident tone signaled it had wandered into territory it couldn't handle. That's the distinction this argument turns on, drawn from education research on cognitive offloading: beneficial offloading frees capacity for genuine understanding, while outsourcing hands off the intrinsic cognitive work that was the actual point of the exercise. On a payroll, that's the gap between asking AI to format a report and asking it to decide what the report should conclude; easy to state as a rule, nearly impossible to see from the inside an hour before the deadline.
How the Pattern Compounds Across the Organization
Blur that line and the consequences build up quietly, at every level. A junior analyst who produces a memo with heavy, unverified AI help and one who produces the same memo unaided turn in output that looks identical, because whatever measure the firm uses treats the memo as the unit of work rather than the judgment behind it. What never gets built, in the first case, is the slow accumulation of pattern recognition: the instinct that a number looks off, or a contract clause reads strangely, the kind of expertise the firm will need from that analyst ten years on. Once skipped, that capacity is expensive to rebuild, mostly because nobody ever tracked it as worth protecting.
Higher up, the same mechanism runs with less friction and considerably higher stakes. The newsletter The Data Ecosystem describes leaders who skim a five-line AI summary and start treating the model's inferred take as their own judgment. For routine calls, maybe 4 out of 5, this costs almost nothing. Trouble concentrates in the remaining, high-stakes decisions, where a leader's own independently held context determines how good the call actually is. Senior leaders are supposed to stay close enough to catch what's wrong while keeping strategic distance; using AI this way collapses that balance instead of protecting it, producing something closer to micromanagement, where more ground gets nominally covered and less real judgment gets exercised. There's a labor-market version of the same story already showing up in the data. Research covering résumé and job-posting records for sixty-five million workers across two hundred and eighty thousand firms found a clear pattern of seniority-biased hiring: firms are quietly cutting entry-level recruitment because AI now handles the routine work that used to train juniors. Those juniors were the internal pipeline for the senior judgment a firm will need in fifteen years and that judgment can't be bought back later, since it's tied up with years of firm-specific context. What looks efficient for one firm becomes corrosive once every firm in the industry does the same thing.
Building Governance and the Policy Response
The real diagnosis here is organizational, not personal. The International Labour Organization's 2026 research brief calls it an "aggregation paradox": task-level productivity gains of 10 to 70 percent, concentrated among less experienced workers doing well-defined tasks, simply haven't shown up yet in firm-level or national productivity statistics. The brief attributes this to the same organizational lag that delayed payoffs from electrification and earlier waves of information technology by roughly a generation. The OECD gets to a similar place from different data: firm-level AI adoption more than doubled between 2023 and 2025, but adoption still clusters around large, digitally mature firms, with skills shortages cited as the main obstacle everywhere else. What's actually missing isn't sharper prompting technique. It's a system in the plainest organizational sense: clarity on who's responsible for an AI-assisted output before it counts as finished, a defined process for checking that output against domain knowledge and a governance layer that decides, on purpose, which categories of work need that scrutiny and which don't.
The government's role here is narrower than either banning the technology outright or leaving everything to individual discretion, because the line between beneficial offloading and harmful outsourcing gets drawn at the level of how a specific task is designed, not something a legislature can regulate directly. Policy's legitimate job is fixing coordination failures no single firm can solve on its own, the seniority-biased hiring pattern being the clearest case, through licensing standards and procurement rules that keep the incentive for real developmental training alive. A second role is disclosure: requiring firms to report the gap between claimed productivity gains and actual governance maturity, similar to existing reporting regimes for cybersecurity incidents, so capital markets and insurers can start pricing governance quality the way they price other operational risk. A third and probably least controversial is funding the transition itself, since training is a textbook case of market underprovision. The World Economic Forum's 2025 survey of more than a thousand global employers found fifty-nine out of every hundred workers will need meaningful retraining by 2030 and eleven of those fifty-nine are unlikely to get it under current arrangements. Regulation can also assign clear liability for harm caused by unverified AI output in regulated decisions, making the absence of a governance system expensive enough that firms have real reason to build one.

The Divide Between AI Users and AI Governors
The productivity gains here are real; nothing in this argument disputes that AI accelerates work. The risk is that the acceleration is being paid for by drawing down a form of capital, verified judgment, institutional memory, accumulated expertise, faster than any current system replenishes it. That imbalance shows up first as individual habits, then widens into a gap between what firms report and what they actually govern and eventually hardens into a policy vacuum that neither banning the technology nor lecturing people about discipline can fill. Fixing it responsibly means firms building real verification systems that run from the individual desk through middle management to company governance, schools rethinking how expertise gets built in the first place and governments sticking to the coordination failures, disclosure gaps and training externalities that only they can address. The labor market taking shape now isn't sorting along the old lines of seniority or credentials. It's splitting between AI users, who take fluent output on faith and AI governors, who've earned the authority to direct it, check it and answer for it. That sorting is already underway.
This article is based on an original research article published by The Economy Research. For the original version, please refer to The Missing System: Cognitive Outsourcing and the Policy Architecture of AI in the Workplace.
The views expressed in this article are those of the author(s) and do not necessarily reflect the official position of The Economy or its affiliates.
References
Agrawal, A. (2026) ‘AI is rewriting the economics of outsourcing’, Harvard Business Review, 5 June.
Anderson, D. (2026) ‘Leaders are starting to outsource their thinking to AI’, The Data Ecosystem, 28 May.
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, O. and Mariman, R. (2025) ‘Generative AI without guardrails can harm learning: Evidence from high school mathematics’, Proceedings of the National Academy of Sciences, 122(26), e2422633122.
Bracha, A. and Tang, J. (2025) Shaping the Future of Work: Workers’ Optimism and Pessimism about AI. Current Policy Perspectives 25-16. Boston, MA: Federal Reserve Bank of Boston.
Brynjolfsson, E., Chandar, B. and Chen, R. (2025) Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence. Stanford, CA: Stanford Digital Economy Lab.
Brynjolfsson, E., Li, D. and Raymond, L.R. (2025) ‘Generative AI at work’, The Quarterly Journal of Economics, 140(2), pp. 889–942.
Chan, C.Y.C. and Shedania, K. (2026) The Aggregation Paradox of AI: Why Do Micro-Economic Productivity Gains from AI Disappear at Scale? Geneva: International Labour Organization.
Dell’Acqua, F., McFowland III, E., Mollick, E.R., Lifshitz-Assaf, H., Kellogg, K.C., Rajendran, S., Krayer, L., Candelon, F. and Lakhani, K.R. (2026) ‘Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality’, Organization Science, 37(2), pp. 403–423.
Hosseini Maasoum, S.M. and Lichtinger, G. (2025) Generative AI as Seniority-Biased Technological Change: Evidence from U.S. Résumé and Job Posting Data. SSRN Working Paper 5425555, 31 August, revised 6 June 2026.
Info-Tech Research Group (2026) AI Adoption and Impact Study: AI in Software Development, June 2026 Top 10 Insights. Arlington, VA: Info-Tech Research Group.
Instructure (2026) New Instructure Research Shows the Current State of AI in Education: Formal Training and Support for Educators. Salt Lake City, UT: Instructure.
Kosmyna, N., Hauptmann, E., Yuan, Y.T., Situ, J., Liao, X.H., Beresnitzky, A.V., Braunstein, I. and Maes, P. (2025) Your Brain on ChatGPT: Accumulation of Cognitive Debt When Using an AI Assistant for Essay Writing Task. arXiv preprint, 10 June.
Lee, H.P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R. and Wilson, N. (2025) ‘The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers’, Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Article 1121, pp. 1–22.
Lee, K. (2026) ‘AI labor and human labor: Why replacement is slower than the hype, but more serious than workers think’, Swiss Institute of Artificial Intelligence, 16 July.
Stackpole, B. (2026) ‘“AI gravity” is pulling you toward dependency. Here’s how to push back’, MIT Sloan School of Management, 2 June.
Stephenson, R. and Armstrong, C. (2026) Student Generative Artificial Intelligence Survey 2026. HEPI Report 199. Oxford: Higher Education Policy Institute, in partnership with Kortext.
Tamkin, A. and McCrory, P. (2025) ‘Estimating AI productivity gains from Claude conversations’, Anthropic Economic Index, 25 November.
World Economic Forum (2025) The Future of Jobs Report 2025. Geneva: World Economic Forum.
Yaraghi, N. (2026a) ‘Borrowed expertise: Why AI’s productivity boom may not survive the generation that built it’, Brookings Institution, 10 July.
Yaraghi, N. (2026b) ‘Repaying the inheritance: How education and research policy can address AI’s borrowed expertise’, Brookings Institution, 20 July.