🔍 Read the full analysis: The Referee Shortage: How AI Made Doing Cheap And Checking Expensive on ThorstenMeyerAI.com
Get movie-night favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A source analysis describes a widening gap between AI’s ability to produce work and the human capacity to check whether that work is correct and fit for use. It cites examples in mathematics, software and contract work, while noting that several software metrics come from companies that sell review tools and should be read with care.
An analysis published this week argues that AI-generated work is expanding faster than the human capacity to verify it, drawing on examples from mathematics, software development and contract workflows. The gap matters because organisations may be able to produce more material than experts can responsibly review, leaving correctness, accountability and training of future reviewers as unresolved constraints.
The analysis says OpenAI published 722 mathematical manuscripts produced from roughly 4,000 problems, with an average result taking about three hours of compute. It contrasts that volume with the careful verification of an earlier result from the same programme: a proposed counterexample to an Erdős conjecture was examined by five leading mathematicians. OpenAI has said some results that are not formally checked could have issues. The source describes the wider problem as “verification abundance, adjudication scarcity”: software can check whether a proof follows from its stated assumptions, but human experts still have to judge whether the result is meaningful and whether it addresses the right question.
Software metrics cited in the analysis point to a similar review bottleneck. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The analysis also cites a peer-reviewed 2026 study in which 61% of AI-agent pull requests received no human review before being merged or closed.
In professional services, the source points to OpenAI’s partnership with contract-software company Ironclad. It says GPT-6 Astra was trained on real contracting workflows and met 55% of evaluation criteria on average across 11 tasks. That is described as a substantial improvement over the previous model, but the remaining criteria still need to be found and assessed before the work can be relied on.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Human Review Is the New Bottleneck
The immediate implication is not that AI output is necessarily wrong, but that more output does not automatically mean more usable work. A result, code change or contract draft may need expert scrutiny before an organisation can act on it. If review capacity does not grow alongside production, work can be delayed, accepted with insufficient scrutiny, or filtered by the people who produced it.
The analysis identifies a potential economic shift: when generation becomes abundant but judgement remains scarce, organisations may place greater value on people who can assess work and take responsibility for approving it. Senior engineers, specialist lawyers, auditors and scientific reviewers could become constraints on how much AI-generated material can be safely used. This is the analysis’s interpretation, not a measured forecast of wages or job growth.
There is also a workforce concern. The source argues that junior staff traditionally learn judgement by doing the work that AI may now draft: writing code, preparing contracts or developing proofs. If those entry-level tasks shrink without replacement training, the future pool of experienced reviewers could narrow even as demand for review rises.
As an affiliate, we earn on qualifying purchases.
Evidence Across Three Workflows
The examples differ in what “checking” means. In mathematics, formal proof tools can verify that a proof establishes a stated theorem, but they do not decide whether the theorem is useful or whether it captures the intended claim. In software, tests can check specified behaviour without proving that the tests reflect all user needs or that a change is safe in the broader system. In contracts, a model can draft or analyse language, while people remain responsible for identifying errors such as a missed approval rule or an unsuitable jurisdiction clause.
The software figures need qualification. Faros AI and LinearB sell code-review products, as the source notes, so their findings should be interpreted with awareness of that commercial interest. The analysis says the direction is consistent across the cited sources, but the data does not establish that AI alone caused every change in review time, acceptance rates or review practices.
Technical verification also does not settle institutional responsibility. The source notes that contracts are signed by people, engineers stamp designs and researchers answer for published work. Organisations may use automated checks, but decisions with legal, safety or professional consequences still involve accountable individuals and institutions.
mathematical proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Much Review Is Being Lost
The available figures do not establish a single, comparable measure of review quality across mathematics, software and contracts. The source does not provide the full methodology or comparison windows for every statistic, and the software metrics come partly from vendors with commercial interests in review tools. The figures therefore show reported patterns, not a definitive estimate of AI’s causal effect across the industry.
It is also unclear how much of verification can be automated as models and formal tools improve, or whether organisations will add reviewers, change workflows or accept more risk. The source raises the possibility that AI reduces opportunities for junior staff to develop expertise, but it does not provide workforce data demonstrating that this is already happening at scale.
contract review automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Whether Review Capacity Catches Up
The next test is whether organisations build review processes that can handle higher output without relying on rubber-stamping or blanket suspicion of AI-generated work. That may involve stronger automated checks, clearer review standards and explicit human sign-off for high-consequence decisions. None of the cited material establishes which approach will prove most effective.
For now, the analysis points to a practical question for employers and professional fields: who will verify AI output, and how will they gain the experience to do it well? Further evidence on review outcomes, error rates and training pathways will be needed to determine whether the gap is temporary or a lasting limit on the use of AI-generated work.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described?
The analysis says AI can produce mathematical, software and contract-related work faster than people can verify it, creating a potential review-capacity bottleneck.
Did OpenAI publish 722 verified mathematical results?
The source says OpenAI published 722 mathematical manuscripts drawn from roughly 4,000 problems. It also reports that some results were not formally checked and quotes OpenAI as warning that unformalized results could have issues.
What do the software figures show?
The cited reports describe more pull requests being merged alongside longer waits for review and lower acceptance rates for AI-generated changes in one dataset. These are reported findings, not proof that AI alone caused the changes; some sources sell code-review tools.
Does the analysis say AI cannot check its own work?
No. It says automated tools can verify some properties, such as whether a proof follows from its stated theorem or code passes specified tests. It argues that people may still need to judge whether the specification is right, the work matters and someone can take responsibility for its use.
What remains unknown?
The scale of the review gap, how much automated checking will reduce it and whether reduced junior-level work will weaken future reviewer supply are not established by the material cited.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
