🔍 Read the full analysis: When Creation Scales Faster Than Verification on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A source article describes a widening gap between AI’s ability to produce work and the human capacity to verify it, using examples from mathematics, software and contract work. The figures suggest review can become a bottleneck, but several software metrics come from companies that sell code-review tools, and the source does not provide enough detail to independently assess every study.
A report published this week says an OpenAI mathematical research effort produced 722 manuscripts from roughly 4,000 problems, while a counterexample from the same programme drew careful verification by five leading mathematicians. The contrast illustrates a broader concern across research, software and contract work: AI can increase the supply of drafts and results faster than people can check whether they are sound and fit for use.
The report says OpenAI’s model generated the 722 manuscripts across 372 families of problems, with an average result taking about three hours of compute. Some results have been checked in the Lean proof assistant. OpenAI cautioned that some results not formalized in Lean could have issues. The report does not establish that all manuscripts are correct, novel or independently reviewed.
In software, the source cites figures from several industry analyses. Faros AI reported that teams merged 98% more pull requests during periods of high AI adoption, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. These are company-reported findings, not a single controlled comparison; the source notes that some cited firms sell code-review tools.
The source also describes a partnership between OpenAI and contract-software company Ironclad involving GPT-6 Astra, trained on real contracting workflows. It says Astra met an average of 55% of evaluation criteria across 11 tasks, an improvement on a previous model. That result leaves substantial room for review before contract work can be relied on. The supplied material gives no detailed evaluation protocol or independent assessment.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Sets the Pace
If AI makes it inexpensive to produce drafts, code or research, the limiting resource may shift to people able to determine whether the work is correct, relevant and safe to use. A growing queue of unchecked output can delay useful work; rushed review can also allow errors through. The source describes both risks in software: some changes may receive little or no human review, while reviewers may deprioritise AI-generated work because they expect it to need extra scrutiny.
The effects extend beyond immediate productivity. Expert review carries professional and institutional responsibility: engineers approve designs, lawyers sign or advise on contracts, and researchers answer for published claims. Models can assist those processes, but the source argues that organizations still need accountable people to make decisions. If review capacity is scarce, its cost and influence could rise as AI-generated work becomes more common.
There is also a workforce concern. The source argues that junior employees often build expertise by doing the work now being automated: drafting, coding and proving. If they mainly check machine-generated output, they may get fewer chances to develop the judgment needed for senior review. That is a plausible risk, not a measured outcome established by the figures presented.
As an affiliate, we earn on qualifying purchases.
Three Fields, One Bottleneck
In mathematics, formal verification can show that a proof follows from its stated assumptions and theorem. It cannot by itself show that the theorem addresses the right question or that the result is important. The source uses the phrase “verification abundance, adjudication scarcity” to distinguish mechanical checking from human assessment of meaning and significance. It also refers to a disputed mathematical counterexample, but supplies no date or detailed account of that episode.
In software, automated tests and code-review tools can catch particular defects, but their checks depend on what has been specified and tested. A change may pass tests that fail to cover the intended requirements. In contracting, evaluation against task criteria can measure performance on those tasks, but it does not show that every real contract risk has been captured. Across all three fields, the source’s central distinction is between producing a result and establishing that it is suitable for a particular use.
The evidence has different levels of independence. The source cites company analyses for pull-request metrics and refers to a peer-reviewed 2026 study reporting that 61% of AI-agent pull requests received no human review before being merged or closed. It does not name that study or provide its methods in the supplied material. The OpenAI and Ironclad figures are presented as results from a company partnership. These measures should not be treated as directly comparable or as a complete picture of all workplaces.
As an affiliate, we earn on qualifying purchases.
How Much Review Is Enough?
The supplied material does not identify all the underlying studies or provide their methods, publication dates and data definitions. That limits comparisons across the reported metrics. For example, a longer wait before review begins, a lower acceptance rate and a missing review are different outcomes; none alone establishes whether AI-generated work is more error-prone overall.
It is also unclear how representative the cited organisations and tasks are, how much human editing occurred before submission, and whether review practices change as teams gain experience. The source does not quantify the total cost of checking, measure the long-term supply of experienced reviewers, or establish whether AI tools can reduce that cost. Its “referee premium” is an interpretation of the evidence, not a measured market outcome.
As an affiliate, we earn on qualifying purchases.
Track Review and Training
The next useful evidence will come from transparent, independently assessable evaluations that report what was reviewed, how errors were defined, how much time review took and what happened after deployment. Organisations adopting AI can also track review queues, defects, rework and the share of work receiving substantive human assessment, rather than relying on output volume alone.
For professional and research settings, the key question is whether review standards and accountability keep pace with production. The source does not announce a new policy or a scheduled follow-up study. Whether the proposed reviewer shortage develops into a broader constraint remains an open question, and depends partly on how employers train junior staff while adopting AI tools.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described?
The report says AI-generated work is increasing across mathematics, software and contract tasks, while human verification remains time-consuming. Its central example is an OpenAI programme that produced 722 mathematical manuscripts from roughly 4,000 problems.
Does the report say all 722 manuscripts are correct?
No. The source says some results were checked in Lean and quotes OpenAI warning that some unformalized results could have issues. It does not say all manuscripts have been independently verified.
What do the software figures show?
The cited company analyses report higher pull-request volume alongside longer review waits and lower acceptance rates for AI-generated changes. The figures come from different analyses and should not be read as a single controlled study; some sources sell code-review products.
Why might human review remain necessary?
Automated checks can test a proof, code change or contract against defined criteria, but they may not show that the criteria capture the real need. People also remain accountable for many professional decisions.
Is there evidence that AI will cause a shortage of expert reviewers?
The source raises that possibility, arguing that routine work helps train future experts. It does not provide a measured forecast of a reviewer shortage, so the scale and timing of that risk remain uncertain.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
