Why Checking AI Work Doesn’t Get Cheaper With More Output
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why Checking AI Work Doesn’t Get Cheaper With More Output on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A source article argues that AI is increasing the volume of mathematical, software and professional work faster than people can check it. The examples point to a growing review bottleneck, though several software metrics come from companies that sell review tools and should be read with care.

A new analysis argues that AI is producing work faster than people can verify it, citing 722 mathematical manuscripts published by OpenAI this week and software data showing review delays rising alongside AI adoption. The development matters because review capacity can limit how much AI-generated work organisations can safely use; the figures do not establish that every AI result is unreliable or that review costs have risen in every field.

OpenAI’s mathematical programme was given about 4,000 problems and produced 722 manuscripts grouped into 372 families, according to the source material. The average result took about three hours of compute. Some results have formal checks in Lean, a proof assistant, while OpenAI cautioned that unformalized results could have issues. The source contrasts that volume with the careful verification by five leading mathematicians of an earlier result from the programme: a proposed counterexample to an Erdős conjecture.

In software, the analysis cites several measures of rising review pressure. Faros AI reported that teams merged 98% more pull requests across low- and high-AI-adoption periods, while review time increased 91%. LinearB said its analysis of 8.1 million pull requests across 4,800 organisations found AI-generated changes waited 4.6 times longer for review to begin. It reported an acceptance rate of 32.7% for AI-generated changes, compared with 84.4% for human-written ones. The source notes that some cited providers sell code-review products, so their findings warrant careful interpretation.

The analysis also points to professional contract work. It says OpenAI and contract-software company Ironclad evaluated GPT-6 Astra on 11 tasks, where the model met an average of 55% of evaluation criteria. That result is described as a substantial improvement over the previous model, but still leaves criteria for a person to check before the work is usable. The source gives no further details on the evaluation design or the consequences of missed criteria.

At a glance
analysisWhen: Published this week, according to the s…
The developmentA newly published analysis uses examples from mathematics, software and contract work to argue that AI is making production faster without making human verification proportionally cheaper.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Limits AI Use

If the analysis’s pattern holds, the constraint on adopting AI may shift from generating drafts, code or proofs to deciding whether the outputs are correct and fit for purpose. Organisations can produce more work without being able to approve more of it at the same pace. That can mean longer queues, delayed releases or added risk if review is skipped.

The issue is not simply whether a model can check another model. Automated tests and formal proof tools can verify particular properties, but they cannot by themselves establish that the right problem was posed or that a result meets the organisation’s needs. In regulated or consequential work, a human may also remain responsible for signing off. The source calls the resulting scarcity a potential “referee premium”: greater demand for people able to assess and take responsibility for work.

There is also a workforce concern. Junior staff often gain the experience needed for expert judgment by doing the drafting, coding and analysis that AI can now assist with. If those tasks disappear without replacement training, organisations could weaken the future supply of reviewers even as their need for review grows. That is a risk raised by the analysis, not an outcome established by the cited figures.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evidence Across Three Fields

The examples span mathematics, software and contract workflows, but they are not directly comparable. Mathematical manuscripts, pull requests and contract evaluation criteria involve different standards and consequences. Taken together, they illustrate the analysis’s central argument: generation can scale quickly, while expert judgment depends on time, domain knowledge and accountability.

In mathematics, formal verification can establish that a proof follows from its stated assumptions. It does not necessarily show that the theorem is important, correctly framed or useful. In software, passing tests shows that code meets the tests that were written, not that the tests cover every requirement. The source also cites a peer-reviewed 2026 study reporting that 61% of AI-agent pull requests received no human review before being merged or closed, and Faros’s report of a 31.3% rise in merges with zero review during high-adoption periods.

These results describe particular studies and tracked organisations, not the entire economy. The source says OpenAI selected which of its 722 manuscripts to publish from the larger set of problems posed. That makes the programme’s own selection process relevant to how readers interpret the published output, but the material does not provide enough detail to assess the criteria or selection rate.

Amazon

software review automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Available Evidence

The source material does not provide links, full methods or date ranges for the Faros AI and LinearB analyses, making it difficult to compare their findings or determine how representative the organisations are. It also notes that several sources sell code-review tools. That commercial interest does not invalidate their data, but it is a reason to inspect definitions and methods before treating the numbers as broad industry measures.

Details are also limited for OpenAI’s mathematics programme and the Ironclad contract evaluation. The material does not explain how the average compute time was calculated, how manuscript families were grouped, or how the 11 contract tasks and evaluation criteria were selected. Nor does it quantify whether human review time has increased across professions overall. The evidence supports a concern about a possible bottleneck; it does not settle its size or long-term effects.

It remains unclear how quickly tools for automated testing, formal verification and review will improve, or whether those tools will reduce the human workload without shifting it to new forms of oversight. The analysis’s concerns about rubber-stamping, reviewer triage and lost junior training are plausible mechanisms, but the cited statistics alone do not establish how often each occurs across organisations.

Amazon

mathematical proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Track Review and Training

The next useful evidence will be independent, methodologically transparent measurement of review times, review quality and failure rates across teams with different levels of AI use. For the mathematics and contract examples, more detail on validation methods and task selection would help readers judge what the reported outputs demonstrate.

Organisations adopting AI will also need to decide how work is selected for review, who is accountable for approval and how junior staff can build the experience that expert review requires. The source article argues for protecting those training paths while using AI. Whether that approach becomes common—and whether the expected reviewer shortage materialises—remains to be seen.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main claim of the analysis?

It argues that AI can increase the volume of work faster than human review capacity grows, leaving verification and expert judgment as potential bottlenecks.

Did OpenAI publish 722 mathematical manuscripts?

The source material says OpenAI published 722 manuscripts in 372 families after its model was posed about 4,000 problems. It also says some results have formal checks, while OpenAI cautioned that unformalized results could have issues.

Do the software figures prove AI-generated code is worse?

No. The cited figures report longer review waits and lower acceptance rates for AI-generated changes in particular analyses. They do not, on their own, establish why the differences occurred or how representative the samples are.

Can AI tools verify other AI output?

They can check defined properties, such as whether code passes specified tests or a proof follows formal rules. Those checks do not necessarily establish that the requirements, tests or theorem address the right question.

What remains uncertain?

The available material does not provide full methods for several cited studies or show how widespread the reported patterns are. It is also unclear whether improved verification tools will substantially reduce human review needs.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Anthropic Connects Claude To A Marketplace Of 2,000+ AI Plugins

A BleepingComputer headline reports more than 2,000 Claude plugins and connectors, but launch details and the count’s meaning remain unverified.

Behind The Microduck: The Open Stack AI That Powers It

Hugging Face’s Microduck is an affordable, open-source robotic platform demonstrating embodied reinforcement learning, signaling a shift in accessible robotics.

An Update On Orion For Linux And Windows

Kagi says it will stop developing Orion for Linux and Windows and release both projects’ source code for community stewardship.

Pentium II At 600Mhz With Voodoo 3 Emulated On 86Box With M6 Mac Mini

Search interest surges in emulating Pentium II 600MHz with Voodoo 3 via 86Box on an M6 Mac Mini. Trigger unconfirmed.