AI Training for Employees and the Future of the Middle Tier
Authored On
Modified
AI lifts weaker workers' output, but not their unaided judgment Training mid-tier workers in checking AI output pays off Top workers thrive, while the bottom tier cannot be saved

In a study of 5,172 customer support agents, researchers from Stanford and MIT recorded a 15 percent average gain in productivity and a gain of about 30 percent for less experienced and lower-skilled agents, while for the experienced and highly skilled, the impact of the AI tool was minimal. The finding is usually interpreted as good news for equality and in part it is, since a single piece of software compressed the distance between the competent and the mediocre employee. It also hides a question that businesses rarely ask explicitly, because the measurement was about performance with the tool in hand, while the ability of the worker without it was left out of the picture. AI training for employees thus acquires a weight that it did not have when it was simply a matter of access to the software and this burden falls almost entirely on the middle tier of the labor market, where the difference between a good and a mediocre use of the tool determines who will keep a position.
The Law Firm Experiment That Split Junior Lawyers in Two
Researchers at Google and MIT, led by economist David Autor, organized an experiment in eleven intellectual property law firms, where 133 lawyers were randomly assigned to two groups. One received accounts on InFlow, an unreleased Google Labs patent drafting product, along with training on the tool, while the other served as the control group. The texts were scored by patent attorneys from an independent firm, who did not know which group the author belonged to, with a standardized grid of criteria covering enforceability, accuracy, completeness, clarity and strategic ambiguity, which refers to the tactical scoping of claims. The quality of the drafts went up by 0.34 standard deviations after ten days and by 0.38 after ninety, so the benefit was maintained over the longer period.
The gain came from the bottom of the distribution. Weak drafts became fewer, medium-quality ones multiplied and the number of excellent ones remained unchanged. Younger lawyers, with less than seven years of experience, benefited most in quality and saved the most time, since they completed the ten-day task 18 minutes faster than their colleagues without access to the tool. The compression of differences was therefore real and measurable, with the ceiling of the distribution staying where it was from the beginning.
The most instructive part of the experiment came with a second assignment, in which lawyers had to note the errors in a flawed patent application without any help. The group that had used InFlow outperformed the control group by 0.32 standard deviations, but the advantage was concentrated entirely in experienced lawyers with seven or more years of experience, who outperformed those in the control group by 0.45. For the younger ones, no average gain was identified. Their scores were split into two poles, with more very low scores, fewer mediocre and more good and the number of excellent scores did not grow. The researchers described the phenomenon as a springboard for learning for some and as a hammock for others, with reservations about the small sample and the duration of three months.

An Average Worker with a Supercomputer Outruns a Genius without One
A genius without a computer and an average worker with a supercomputer don't play the same game and the analogy explains why the range of tasks where the former wins is constantly narrowing. In what requires searching, calculating, synthesizing documents, or testing hundreds of alternatives, the supercomputer is ahead by a wide margin and the advantage is lost only when its operator does not know what to ask or how to check the answer. The average worker does not need to become a genius to achieve superhuman performance. He needs to learn to operate the machine and correctly judge what it returns to him, two skills that training can build.
The most eloquent quantitative picture comes from Harvard Business School, which in collaboration with Boston Consulting Group, followed 758 consultants in real consulting tasks. Those who used GPT-4 completed an average of 12.2 percent more tasks, 25.1 percent faster, with 40 percent higher result quality. Consultants in the bottom half of the distribution improved their performance by 43 percent, the already strong ones by 17 percent. A similar pattern was recorded by MIT in an experiment with professional writing assignments, published in Science in 2023, where participants with the most limited skills benefited the most from the use of ChatGPT. The customer support study found the same pattern, with gains of about 30% for novice agents.
There are, of course, naturally talented workers who perform excellently without paid training and most do not belong to this group. For the rest, whether the supported average worker outperforms the strongest colleague working without a tool depends on the initial size of the gap and the type of work and the BCG study's percentage improvements show how quickly a gap can be closed. The same study included a task deliberately placed beyond the capabilities of the model, where consultants with access to GPT-4 were less likely to come up with a correct solution than those who worked without it. The supercomputer, even when it makes a mistake, does so in a convincing way.
AI Training for Employees Starts with Picking Who Can Spot Errors
Interviews with the experiment's experienced lawyers showed what good use looks like. The experienced ones described the tool's effect as a logic auditor, which weakened their attachment to their own text and forced them to articulate the why and how of each structural change. They spent more time on the second task, restructured the draft to address its legal vulnerabilities and accompanied the corrections with comments explaining the principle behind each. The younger ones worked successively on stylistic and administrative corrections, repeatedly identified serious flaws without correcting them and many of their notes read like instructions to an assistant who wasn't there, a pattern equally common among juniors in the control group, so it cannot be blamed on the tool.

The treated lawyers received training on the tool itself, not a program built to develop judgment, so the directions below are inferences from the pattern of the findings. HR managers targeting the middle level have a reason to select candidates with the test used in the experiment, i.e. by annotating defective text without assistance, since the ability to diagnose errors is what separated senior from junior work in the experiment. Training should teach outcome control as a stand-alone skill, with exercises where the employee has to justify each change in writing and department managers should connect younger people with experienced auditors for the duration of the apprenticeship.
A common objection is that if the tool flattens out the differences, costly training is unnecessary. The data point the other way. In assisted writing the differences were indeed compressed, but in unaided work the benefit remained with the already experienced and in BCG's work outside the limits of the model the use of the tool reduced the share of correct answers. The worker who has not learned to question the output of the system easily produces adequate texts and these are precisely the ones he risks not recognizing when they are incorrect.
Training Can Save the Middle Tier but Not the Bottom One
The top tier, consisting of employees with natural talent and strong judgment, takes advantage of the tool with little extra support and the study of law firms is consistent with this: experienced lawyers gained 0.45 standard deviations in work done without help. The study sorts lawyers by experience, not by skill tier, so reading seniors as the top tier and juniors as the lower tiers is an inference. The assumption that this tier now produces ten times the output is used here as a work scenario and not as a measured size, since none of the studies cited records such a multiplier. If the scenario is even approximate, businesses need fewer people for the same project and the demand for labor is no longer evenly distributed.
At the bottom, the picture is more unpleasant. The Stanford and MIT study found that the system disseminates the best practices of the most capable employees, which means that the contribution of a weak worker, when limited to adequate text production, is already embedded in the software. In the law firm experiment, adequacy improved and excellence stagnated. An employer pays for what the software cannot offer and the group of younger lawyers who used the tool produced more very low scores, a finding that the researchers are wary of because of the small sample. Training builds on a minimal foundation of judgment and where this is lacking, the tool gives the impression of competence without creating it.
The middle tier therefore remains the only one where the cost of training can pay off. Budget holders would do well to concentrate resources there, rather than spreading them to workers without an initial basis for judgment and bodies planning retraining programs have a reason to first check their ability to evaluate the outcome. What will happen to the lower tier, however, is not answered by any of the available experiments and the question goes beyond the limits of business, since it concerns social policy.
Three Months of Data Cannot Settle the Question
The gain for novice agents in the customer support study is real, measured in issues resolved per hour and nowhere does it prove that novices learned something that they will keep when the tool is not in front of them. The law firm experiment covered just three months, a short time compared to the years it takes for patent expertise to mature and was based on the capabilities of the 2024 and 2025 models. The models have since advanced and the study's authors admit that the effects on learning may have shifted in either direction.
The supercomputer and the genius rarely compete directly, since businesses hire people rather than tools and the question of which of the two will win the next competition concerns the tasks left to the human when the machine has taken as much as it can. In the eleven-firm experiment, the gap between the 0.45 standard deviations of the experienced and the zero average gain of the younger remains the most specific measure of what the absence of prior judgment costs.
This article reflects the analytical judgment of The Economy Editorial Board and does not constitute policy advice or the official position of any affiliated institution.
References
Autor, D. et al. (2026) Does AI assistance enhance or erode expertise? NBER Working Paper 35720.
Brynjolfsson, E., Li, D. and Raymond, L.R. (2025) 'Generative AI at work', Quarterly Journal of Economics, 140(2), pp. 889-942.
Dell'Acqua, F. et al. (2026) 'Navigating the jagged technological frontier', Organization Science, 37(2), pp. 403-423.
Noy, S. and Zhang, W. (2023) 'Experimental evidence on the productivity effects of generative artificial intelligence', Science, 381(6654), pp. 187-192.