🔍 Read the full analysis: AI Made Production Fast. Verification Still Takes Judgment on ThorstenMeyerAI.com
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
OpenAI reported producing 722 mathematical manuscripts from about 4,000 problems, while software-industry data cited in the source material shows review struggling to keep pace with AI-generated code. The figures point to an emerging constraint: producing drafts and results is getting faster, but checking their correctness, relevance and safety still requires time and accountable human judgment.
OpenAI published 722 mathematical manuscripts this week after posing about 4,000 problems to an AI model, according to source material from ThorstenMeyerAI.com. The report sets that output against the human work needed to assess whether results are correct and useful, arguing that verification—not generation—may become a constraint as AI-produced work spreads across mathematics, software and professional services.
The source says the manuscripts fall into 372 problem families and that the average result took about three hours of compute to produce. Some results have been checked formally using Lean, a proof-assistant system. OpenAI cautioned that some manuscripts without formal verification “could have issues,” according to the source. The material does not establish that all 722 papers have been independently reviewed or that the average compute time captures the full human effort involved.
The report compares this volume with the response to an earlier result from the same programme: a proposed counterexample to an old Erdős conjecture. It says five leading mathematicians carefully verified that result. The contrast illustrates the report’s central concern: machine systems can produce many candidate results, but expert review remains limited and each claim may require separate judgment.
Software figures cited in the report point to a similar pressure, though the datasets use different methods. Faros AI found teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The report also cites a peer-reviewed 2026 study finding 61% of AI-agent pull requests received no human review before they were merged or closed.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Sets the Pace
If AI makes it easier to generate code, research or contract drafts, organisations may not be able to use that output at the same rate. Review time becomes a practical limit: work can accumulate in queues, receive less scrutiny or be rejected because reviewers cannot establish its quality quickly enough. The report’s figures are a warning about workflow pressure, not proof that every AI-assisted team has the same problem.
The consequences depend on the stakes. A weak draft may cost time; an unchecked contract term, software defect or mathematical claim can carry broader consequences. Formal tools can verify specified properties, but they do not automatically decide whether the specification reflects the real need. People still have to judge relevance, consequences and acceptable risk—and often bear responsibility for the final decision.
The source also raises a workforce concern: reviewers often develop through doing the work they later assess. If junior staff mainly supervise machine-generated drafts rather than learning to write, prove or draft independently, employers could reduce one route by which experienced judgment is built. That outcome is not established by the cited data, but it is a risk organisations may need to track as they redesign entry-level roles.
As an affiliate, we earn on qualifying purchases.
Three Fields, Similar Bottleneck
The source frames mathematics, software and professional workflows as examples of the same production-and-review gap. In mathematics, Lean can check whether a formal proof follows specified rules, but it cannot by itself establish that the theorem is important or that it addresses the right question. The source’s phrase for this tension is “verification abundance, adjudication scarcity”: automated checks may scale while expert interpretation does not.
In software, the reported findings come from industry analyses as well as a peer-reviewed study. Faros AI and LinearB sell code-review tools, a commercial interest that readers should bear in mind when evaluating their measurements. The figures are not directly interchangeable: they cover different samples and outcomes. The report says their direction is consistent, but the available material does not provide enough methodology to independently compare the studies or establish a single industry-wide rate.
The professional-workflow example is an OpenAI partnership with contract-software company Ironclad. The source says GPT-6 Astra met an average of 55% of evaluation criteria across 11 tasks, describing that as a large improvement over the prior model. This is a reported evaluation result, not evidence that the system can independently handle contracts in practice. The source does not identify the evaluation protocol or provide the prior model’s score.
“Verification abundance, adjudication scarcity.”
— ThorstenMeyerAI.com
proof assistant software for mathematicians
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Much Review Is Enough?
The source does not establish how many of the 722 manuscripts have received independent expert review, how many are formally verified, or how many contain errors. It also does not provide enough detail to assess the mathematical programme’s selection process or compare the reported three-hour compute average with total research and review time.
The software findings do not prove that AI-generated changes are inherently less reliable in every setting. Acceptance rates and review delays can reflect project complexity, team policies and how “AI-generated” is defined, among other factors. The source notes that some cited companies sell review tools, and its summary does not include the full study methods or confidence ranges.
It is also unclear whether the Ironclad evaluation predicts performance on live contracts, how evaluators weighted the 11 tasks, or whether the reported 55% represents the model’s current production performance. More broadly, the report presents a plausible concern about training future reviewers, but does not quantify whether AI adoption is already reducing the number of people acquiring those skills.
As an affiliate, we earn on qualifying purchases.
Track Review and Training Outcomes
The next useful evidence will come from fuller methods and follow-up measurements: how mathematical results fare under independent review, how software review time and defect rates change across comparable teams, and whether organisations preserve meaningful human oversight as output grows. Review queues, rejection reasons and post-release defects can help distinguish a genuine quality problem from a temporary adjustment in workflow.
For professional services, details of the contract evaluation and performance on real tasks would clarify what the reported score means. Employers adopting AI can also monitor whether junior staff still get hands-on practice creating work before they are asked to approve machine-generated output. The source offers no announced deadline for such results; the central question remains whether verification capacity and training can keep pace with production.
AI project review and validation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI publish?
According to the source material, OpenAI published 722 mathematical manuscripts spanning 372 problem families after posing about 4,000 problems to a model. The source says some results were formally checked in Lean and cautions that unformalized results could have issues.
Does the report say all 722 manuscripts are verified?
No. The source says some results were checked formally, but it does not state that every manuscript received formal verification or independent expert review.
What do the software statistics show?
The cited studies report increased pull-request volume alongside longer review delays and, in one study, many AI-agent pull requests receiving no human review. Their samples and methods differ, so the figures should not be treated as one universal rate.
Why can’t AI simply verify AI-generated work?
Automated systems can check defined properties, such as whether a proof follows formal rules or code passes specified tests. Those checks do not necessarily show that the right question was asked, the tests cover the real need, or the result is safe and useful. Human judgment and accountability remain relevant.
What remains unknown?
The source does not give an independent review rate for the mathematics manuscripts, full methods for the cited software studies, or enough information to judge how well the contract-software evaluation predicts real-world performance.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
