🔍 Read the full analysis: The Oversight Challenge Created By Cheap AI on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A report on AI-generated work describes a widening gap between the low cost of producing results and the time required to verify them. Data cited from mathematics, software and contract workflows suggests review capacity may constrain how much AI output organisations can safely use, though several figures come from vendors and require care.
A report published this week argues that cheap AI-generated work is outpacing human review, drawing on examples from mathematics, software development and contract workflows. The gap matters because organisations may be able to generate more material than qualified people can verify, leaving quality control and accountability as constraints on AI use.
The report says OpenAI posed about 4,000 mathematics problems to a model and produced 722 manuscripts across 372 families; the average result took about three hours of compute. Some results were formally checked in Lean, while OpenAI warned that some unformalized work “could have issues.” The report contrasts this scale with the careful verification of an earlier result from the same programme: a counterexample to an old Erdős conjecture was examined by five leading mathematicians.
In software, the report cites several datasets that point to heavier review pressure. Faros AI reported that teams in high-AI-adoption periods merged 98% more pull requests, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, found AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found 61% of AI-agent pull requests received no human review before they were merged or closed.
The report also points to OpenAI’s partnership with contract-software company Ironclad. In an evaluation across 11 tasks, OpenAI’s GPT-6 Astra met 55% of evaluation criteria on average, an improvement over its predecessor, according to the source material. The remaining criteria still require review before a draft can be used safely. The report notes that some cited software metrics come from companies selling code-review products, and says they should be read with care.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Sets the AI Limit
If AI makes drafting, coding and mathematical exploration cheaper, the practical value of that output depends on whether people can check it accurately and quickly. A queue of unreviewed work can slow deployments; weak review can allow errors through; and blanket suspicion of AI-generated work can delay sound contributions. Each outcome affects the productivity gains organisations can actually realise.
The issue also concerns responsibility. A person or institution may have to stand behind a contract, engineering design, software release or published result. AI systems can assist with production and some forms of checking, but the report argues that they do not replace the human or institutional role of deciding whether the work addresses the right problem and who is accountable if it fails.
The report’s labour-market argument is that experienced reviewers may become more valuable as AI output grows. That is an interpretation, not a measured forecast: the cited material does not establish how wages or staffing will change. But it identifies a concrete operational risk: if organisations expand generation without resourcing review, the volume of usable work may be determined by scarce expert attention.
As an affiliate, we earn on qualifying purchases.
Three Fields, Similar Pressure
In mathematics, formal proof systems can confirm that a proof follows from its stated assumptions and conclusion. They cannot by themselves establish that the theorem is the one researchers intended to prove, or that the result is important. The report describes this as a distinction between verification and adjudication: technical checking can scale, while interpreting a result still calls for expertise.
Software testing has a similar limit. Tests can show whether code passes specified checks, but they cannot establish that the tests cover the real need. The report says AI-generated changes may be harder to assess because reviewers lack the author’s reasoning and cannot readily tell where mistakes are likely to be. The cited statistics describe different populations and measures, so they should not be combined into a single industry-wide rate.
In contract work, a model’s score against evaluation criteria indicates performance on that assessment, not that every contract is ready for use. Clauses, approval rules and jurisdictional requirements still need attention from people qualified to judge the particular agreement. Across these examples, the common development is not that AI cannot help with review, but that producing more material does not automatically resolve questions of correctness, relevance or responsibility.
““Verification abundance, adjudication scarcity.””
— The report, citing a recent paper’s title
As an affiliate, we earn on qualifying purchases.
How Much Review Is Enough?
The cited figures do not provide a single, comparable measure of review quality across mathematics, software and contracting. The source material does not specify the time periods behind every vendor metric, and it flags that Faros AI and LinearB sell code-review tools. Their figures are useful as reported observations, but they do not alone establish that AI adoption caused every measured change.
It is also unclear how many unreviewed changes contained consequential errors, how often human review missed defects, or whether teams later corrected issues. The contract evaluation’s 55% average does not identify which criteria were missed or how the tasks were weighted. The report’s broader claims about future staffing, reviewer value and lost apprenticeship opportunities are concerns and projections, not confirmed outcomes across the workforce.
software pull request review tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Building Review Alongside Generation
The immediate question for organisations adopting AI is whether their review processes can keep pace with the work being generated. That means tracking review queues, time to review, acceptance and defect rates, and whether high-risk work receives qualified human attention. The cited material does not identify one agreed standard or policy that organisations are expected to adopt next.
Longer term, the report argues that employers and professional institutions will need to preserve ways for junior workers to learn the underlying craft, rather than relying only on AI-produced drafts. Whether that approach becomes common remains uncertain. As AI systems and their evaluations change, new evidence on error rates, review practices and accountability will help show whether the gap is narrowing or becoming a more persistent limit on deployment.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described in the report?
The report says AI can produce work faster and more cheaply than people can verify it, citing examples from mathematics, software development and contract workflows. It frames review capacity as a possible limit on how much generated work organisations can use.
Did OpenAI publish 722 mathematically verified results?
The source says OpenAI published 722 manuscripts arising from about 4,000 problems. Some results were formally checked in Lean, but OpenAI cautioned that some unformalized results could have issues. The source does not say all 722 were formally verified.
What did the software data find?
The report cites vendor data showing more pull requests being merged in high-AI-adoption periods alongside longer review times, as well as a separate analysis finding AI-generated changes waited longer to receive review. The figures cover different datasets and should not be treated as one universal rate.
Can AI review AI-generated work?
AI can assist with checks, but the report argues that automated verification does not necessarily confirm that the right question was asked, that tests capture the real requirement or that a result is fit for use. Human judgment and accountability may still be needed, especially for consequential work.
What remains uncertain?
The sources do not establish how often weak review leads to serious errors, whether AI adoption caused all the reported changes, or how the demand for expert reviewers will affect employment and pay. Some software statistics also come from companies that sell review tools.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
