🔍 Read the full analysis: How To Handle The Rising Cost Of Checking AI Work on ThorstenMeyerAI.com
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A source article describes a widening gap between AI-generated work and the human capacity to verify it, drawing on examples in mathematics, software and contract review. The figures indicate a potential bottleneck, but their scope and reliability vary, and some cited data comes from companies that sell review tools.
AI systems are producing more mathematical manuscripts, software changes and contract drafts, while the time and expertise needed to check that work remain limited, according to an analysis published this week by ThorstenMeyerAI.com. The article points to a potential review-capacity bottleneck: organizations may be able to generate work faster than qualified people can establish whether it is correct and fit for use.
The analysis opens with OpenAI’s reported mathematics work: a model was given about 4,000 problems and produced 722 manuscripts, grouped into 372 families. It says the average result took about three hours of compute. Some manuscripts have been checked in Lean, a proof assistant, while OpenAI has cautioned that unformalized results could contain problems. The article contrasts that volume with the careful human verification of a counterexample to an earlier Erdős conjecture, though it does not provide a detailed accounting of the review effort.
Software figures cited in the article suggest a similar imbalance. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB said its analysis of 8.1 million pull requests across 4,800 organizations found AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A 2026 peer-reviewed study, as described in the source, found 61% of AI-agent pull requests received no human review before being merged or closed.
The article also cites OpenAI’s partnership with contract-software company Ironclad. In an evaluation of 11 tasks, OpenAI’s GPT-6 Astra met an average of 55% of the criteria, according to the source. That result is presented as an improvement over a previous model, but the source does not specify the prior score or explain how the criteria were weighted. It says the remaining gaps still require human attention before contract work can be relied on.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Sets the AI Limit
If generation becomes cheap while review remains labor-intensive, the amount of AI output an organization can safely use may be limited by the people able to check it. That could shift demand toward senior engineers, auditors, specialist lawyers and scientific reviewers, whose work includes judgment and responsibility, not just spotting visible errors.
The risk is not only that errors slip through. Slow review can also delay useful work, while reviewers who distrust AI output may push it to the back of the queue. The source cites LinearB’s finding that 38% of reviewers deliberately deprioritize AI-generated changes. That is a company-reported figure, not a universal measure, but it illustrates how confidence in output affects whether productivity gains translate into usable work.
There is also a workforce concern. The analysis argues that routine work has traditionally helped junior staff develop the experience needed to become reliable reviewers. If AI absorbs much of that work, employers may need explicit training paths in drafting, testing, auditing and adjudication. The source presents this as a longer-term risk, not as an established outcome across workplaces.
As an affiliate, we earn on qualifying purchases.
Three Fields, One Verification Gap
The examples in the article cover different kinds of checking. In mathematics, formal tools such as Lean can verify that a proof follows specified rules, but human experts still need to judge whether the theorem is the right one and whether the result matters. In software, tests can confirm behavior covered by those tests, but do not establish that the tests capture the intended requirements. Contract review involves another layer: people must assess whether clauses fit the parties, jurisdiction and approval rules.
The source describes the pattern as “verification abundance, adjudication scarcity”: systems can help check defined properties, while broader decisions about relevance, intent and accountability remain harder to automate. These examples should not be treated as directly comparable measurements. The mathematics count, software metrics and contract evaluation use different methods and come from different sources.
Some software figures cited come from Faros AI and LinearB, companies that sell products related to software development and review. Their commercial interests are a reason to examine methods and definitions carefully. The source says the findings point in the same direction, but the article does not provide enough methodological detail to independently compare all the datasets.
“Verification abundance, adjudication scarcity.”
— The title of a recent paper cited by ThorstenMeyerAI.com
As an affiliate, we earn on qualifying purchases.
How Broad Are These Findings?
The figures do not establish that AI has caused review delays or lower acceptance rates in every organization. The source provides selected metrics, but not full study methods, comparison periods or definitions for every measure. In particular, a longer wait before review begins does not by itself show whether a change is incorrect or whether review teams are understaffed.
It is also unclear how much of the mathematics output has been independently verified, how GPT-6 Astra’s contract evaluation was designed, and whether the cited results represent routine production use. The source does not quantify the overall cost of checking AI work or show how that cost compares with the savings from faster generation. Its conclusion that review capacity is becoming a constraint is an interpretation of the cited examples, not a finding established by a single cross-industry study.
As an affiliate, we earn on qualifying purchases.
Track Review Alongside Output
Organizations adopting AI can monitor review time, the share of work that receives human scrutiny, correction rates and the number of changes returned or rejected—not just how much output is produced. Comparisons should identify the time period, baseline and type of work so that raw counts are not mistaken for evidence of improved quality or productivity.
For now, the source points to no single policy or new study that will resolve the gap. The next useful evidence would include independently evaluated results, transparent methods and longer-term tracking of whether junior workers still gain the experience needed for expert review. Until then, the extent of the bottleneck remains a question for employers and researchers to measure rather than assume.
mathematical proof verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described?
The source article argues that AI is increasing the amount of work produced faster than organizations can review it. It draws examples from mathematics, software development and contract workflows.
Does the cited evidence prove that AI work is less reliable?
No. The reported acceptance and review figures are observations from particular datasets and evaluations. They do not establish that AI-generated work is less reliable in every field or organization.
Why can’t automated checks handle all verification?
Automated tools can test defined properties, such as whether a proof follows formal rules or code passes specified tests. They may not determine whether the underlying question, requirements or evaluation criteria are appropriate.
Are the software metrics independent?
Not all of the cited figures are independent academic findings. Faros AI and LinearB sell software-related products, so their analyses should be read with attention to methods and possible commercial interests. The source also cites a peer-reviewed 2026 study.
What should organizations measure?
Alongside output, they can track review wait times, how much work receives human review, and what is corrected or rejected. Any comparison should specify its time window and baseline.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
