The Oversight Challenge Created By Cheap AI
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Oversight Challenge Created By Cheap AI on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A report on AI-generated work describes a widening gap between the low cost of producing results and the time required to verify them. Data cited from mathematics, software and contract workflows suggests review capacity may constrain how much AI output organisations can safely use, though several figures come from vendors and require care.

A report published this week argues that cheap AI-generated work is outpacing human review, drawing on examples from mathematics, software development and contract workflows. The gap matters because organisations may be able to generate more material than qualified people can verify, leaving quality control and accountability as constraints on AI use.

The report says OpenAI posed about 4,000 mathematics problems to a model and produced 722 manuscripts across 372 families; the average result took about three hours of compute. Some results were formally checked in Lean, while OpenAI warned that some unformalized work “could have issues.” The report contrasts this scale with the careful verification of an earlier result from the same programme: a counterexample to an old Erdős conjecture was examined by five leading mathematicians.

In software, the report cites several datasets that point to heavier review pressure. Faros AI reported that teams in high-AI-adoption periods merged 98% more pull requests, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, found AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found 61% of AI-agent pull requests received no human review before they were merged or closed.

The report also points to OpenAI’s partnership with contract-software company Ironclad. In an evaluation across 11 tasks, OpenAI’s GPT-6 Astra met 55% of evaluation criteria on average, an improvement over its predecessor, according to the source material. The remaining criteria still require review before a draft can be used safely. The report notes that some cited software metrics come from companies selling code-review products, and says they should be read with care.

At a glance
reportWhen: Report published this week; cited studi…
The developmentA report argues that AI is making work cheaper to produce faster than experts can check it, with examples from mathematics, software and contract review.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Sets the AI Limit

If AI makes drafting, coding and mathematical exploration cheaper, the practical value of that output depends on whether people can check it accurately and quickly. A queue of unreviewed work can slow deployments; weak review can allow errors through; and blanket suspicion of AI-generated work can delay sound contributions. Each outcome affects the productivity gains organisations can actually realise.

The issue also concerns responsibility. A person or institution may have to stand behind a contract, engineering design, software release or published result. AI systems can assist with production and some forms of checking, but the report argues that they do not replace the human or institutional role of deciding whether the work addresses the right problem and who is accountable if it fails.

The report’s labour-market argument is that experienced reviewers may become more valuable as AI output grows. That is an interpretation, not a measured forecast: the cited material does not establish how wages or staffing will change. But it identifies a concrete operational risk: if organisations expand generation without resourcing review, the volume of usable work may be determined by scarce expert attention.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, Similar Pressure

In mathematics, formal proof systems can confirm that a proof follows from its stated assumptions and conclusion. They cannot by themselves establish that the theorem is the one researchers intended to prove, or that the result is important. The report describes this as a distinction between verification and adjudication: technical checking can scale, while interpreting a result still calls for expertise.

Software testing has a similar limit. Tests can show whether code passes specified checks, but they cannot establish that the tests cover the real need. The report says AI-generated changes may be harder to assess because reviewers lack the author’s reasoning and cannot readily tell where mistakes are likely to be. The cited statistics describe different populations and measures, so they should not be combined into a single industry-wide rate.

In contract work, a model’s score against evaluation criteria indicates performance on that assessment, not that every contract is ready for use. Clauses, approval rules and jurisdictional requirements still need attention from people qualified to judge the particular agreement. Across these examples, the common development is not that AI cannot help with review, but that producing more material does not automatically resolve questions of correctness, relevance or responsibility.

““Verification abundance, adjudication scarcity.””

— The report, citing a recent paper’s title

Amazon

mathematics verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Enough?

The cited figures do not provide a single, comparable measure of review quality across mathematics, software and contracting. The source material does not specify the time periods behind every vendor metric, and it flags that Faros AI and LinearB sell code-review tools. Their figures are useful as reported observations, but they do not alone establish that AI adoption caused every measured change.

It is also unclear how many unreviewed changes contained consequential errors, how often human review missed defects, or whether teams later corrected issues. The contract evaluation’s 55% average does not identify which criteria were missed or how the tasks were weighted. The report’s broader claims about future staffing, reviewer value and lost apprenticeship opportunities are concerns and projections, not confirmed outcomes across the workforce.

Amazon

software pull request review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Building Review Alongside Generation

The immediate question for organisations adopting AI is whether their review processes can keep pace with the work being generated. That means tracking review queues, time to review, acceptance and defect rates, and whether high-risk work receives qualified human attention. The cited material does not identify one agreed standard or policy that organisations are expected to adopt next.

Longer term, the report argues that employers and professional institutions will need to preserve ways for junior workers to learn the underlying craft, rather than relying only on AI-produced drafts. Whether that approach becomes common remains uncertain. As AI systems and their evaluations change, new evidence on error rates, review practices and accountability will help show whether the gap is narrowing or becoming a more persistent limit on deployment.

Amazon

AI contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development described in the report?

The report says AI can produce work faster and more cheaply than people can verify it, citing examples from mathematics, software development and contract workflows. It frames review capacity as a possible limit on how much generated work organisations can use.

Did OpenAI publish 722 mathematically verified results?

The source says OpenAI published 722 manuscripts arising from about 4,000 problems. Some results were formally checked in Lean, but OpenAI cautioned that some unformalized results could have issues. The source does not say all 722 were formally verified.

What did the software data find?

The report cites vendor data showing more pull requests being merged in high-AI-adoption periods alongside longer review times, as well as a separate analysis finding AI-generated changes waited longer to receive review. The figures cover different datasets and should not be treated as one universal rate.

Can AI review AI-generated work?

AI can assist with checks, but the report argues that automated verification does not necessarily confirm that the right question was asked, that tests capture the real requirement or that a result is fit for use. Human judgment and accountability may still be needed, especially for consequential work.

What remains uncertain?

The sources do not establish how often weak review leads to serious errors, whether AI adoption caused all the reported changes, or how the demand for expert reviewers will affect employment and pay. Some software statistics also come from companies that sell review tools.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Could The Gaming Signal Monitor Reveal Valve’s Hidden Plans?

A new gaming signal monitor hints at Valve’s potential development of a barebones Steam Machine, raising questions about their future hardware strategy.

A Texas Drainage District Walked Its Ditch on a Routine Inspection. They Found a Pipe They Didn’t Recognize Discharging Black Liquid From Tesla’s $1 Billion Lithium Refinery

Routine inspection revealed Tesla’s wastewater pipe discharging dark liquid into a Texas ditch, raising environmental concerns amid regulatory gaps.

Drafting Legal Documents? Grammarly Makes It Easier For Small Businesses

New AI-powered platform helps small businesses and self-represented litigants draft court-ready legal documents with citation verification and proper formatting.

Community volunteer action tracker for local boards

A new volunteer action tracker is being tested to improve follow-up on community board decisions, aiming for a streamlined workflow for local civic groups.