📊 Full opportunity report: Engineering Is Automated. Research Is the Residual. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Recent evidence indicates AI systems have achieved near-complete automation of core engineering tasks in AI R&D. However, research activities remain less automatable, leaving a residual human role. This shift could accelerate AI development but raises questions about innovation and originality.
Recent benchmark data confirms that AI systems can now fully automate core engineering tasks in AI research and development, marking a significant shift in the landscape of AI innovation. While engineering automation approaches saturation, research activities remain less automatable, leaving a residual human role. This development could reshape the pace and nature of AI progress.
According to Thorsten Meyer’s analysis of Jack Clark’s recent essay, six key benchmarks measuring AI capabilities in core AI R&D skills show rapid progress toward saturation. For example, the CORE-Bench, which tests research reproduction, has improved from 21.5% to 95.5% in just over 15 months, with some authors declaring it ‘solved.’ Similarly, the MLE-Bench, assessing Kaggle competition performance, has advanced from 16.9% to 64.4% over 16 months, approaching competitive levels with mid-tier human practitioners.
These benchmarks, which evaluate tasks like reproducing research, optimizing kernels, and solving complex ML problems, indicate that AI can now handle the majority of engineering tasks involved in AI R&D at a near-human or superhuman level. The progress suggests that the bottleneck in AI development is shifting from engineering to research, which remains less automatable due to its creative and exploratory nature. Clark’s analysis implies that the residual research component may itself be a form of large-scale engineering, potentially closing faster than previously expected.
Engineering is automated.
Research is the residual.
Six skill benchmarks. Edison’s framing. The question Clark leaves open is whether research is just engineering at scale.
Jack Clark’s Import AI #455 catalogs six benchmarks measuring AI capability on AI R&D tasks and concludes “AI can today automate vast swatches, perhaps the entirety, of AI engineering.” The residual question is research. The structural read on the residual: it may not be a permanent moat.
Six skills. One trajectory.
Clark catalogs six benchmarks measuring AI capability on AI R&D-relevant tasks. Each individual benchmark could be noise. Six benchmarks moving together is a curve. The pattern is the cascade observed across the broader Clark series — visible here in the specific R&D-skill domain.

The AI Marketing Canvas, Second Edition: A Five-Step AI Plan for Marketers
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three data points. Mixed signal.
Clark provides three data points on the creative-spark question. Yes-evidence: Erdős-1051, centaur math discovery, sporadic Move-37-style moments. No-evidence: low yield, framing dependence, absence of acceleration. The mixed signal is the honest read.
The data supports two readings. Pessimistic: rare moments suggest creative insight is qualitatively distinct from engineering work. Optimistic: rare moments are an artifact of low-volume exploration; more shots on goal yields more discoveries. Both readings are consistent with Clark’s “vast swatches, perhaps the entirety” claim. They differ on the residual.

The Agentic AI Bible: The Complete and Up-to-Date Guide to Design, Develop, and Scale Goal-Driven, LLM-Powered Agents that Think, Execute and Evolve
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five dimensions Clark gestures at but leaves underdeveloped.
Clark’s section is rigorous on the empirical evidence. Five strategic dimensions matter for the institutional response that the Clark series synthesis argues is structurally inadequate.

Trust Me I Am Data Scientist Machine Learning Research: Notebook Planner – 6×9 inch Daily Planner Journal, To Do List Notebook, Daily Organizer, 114 Pages
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Two readings. Different equilibria.
The structural question Clark leaves open: is research a permanent moat that bounds automated AI R&D, or is it engineering at scale that dissolves with more shots on goal? Both readings are consistent with the current data. They differ by orders of magnitude in consequences.
Productivity multiplier years
Recursive loop operational

Computational Visual Media: 13th International Conference, CVM 2025, Hong Kong SAR, China, April 19–21, 2025, Proceedings, Part II (Lecture Notes in Computer Science)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five audiences. Asymmetric cost of being wrong.
The institutional response should not bet on inspiration being a permanent moat. If the distinction holds, capacity built is still useful. If it closes, capacity is necessary. Asymmetric cost-of-being-wrong points toward building now.
IN INDUSTRY
IN ACADEMIA
POLICYMAKERS
INVESTORS
EVERYONE ELSE
Engineering is automated. The residual is the question. The institutional response should not bet on inspiration being a permanent moat.
Implications for AI Development Speed and Innovation
This rapid automation of engineering tasks could dramatically accelerate AI research cycles, reducing costs and time-to-market for new models and capabilities. However, the residual research activities—such as hypothesis generation, novel algorithm design, and scientific discovery—may continue to require human insight, at least for now. The shift suggests that future AI progress might depend more on automating research itself, challenging traditional notions of scientific innovation.
Progress in AI R&D Skills Over Recent Months
Over the past year and a half, multiple independent benchmarks have demonstrated AI’s rapid progress in core R&D skills. The CORE-Bench, which measures research reproduction capabilities, has seen a fourfold improvement, with systems now capable of handling dependencies, code execution, and output analysis at near-human reliability. The MLE-Bench, assessing Kaggle competition performance, has similarly advanced, indicating AI’s growing proficiency in practical, real-world tasks. Parallel developments in kernel design, such as automated GPU kernel optimization, further illustrate the shift toward production-ready automation in AI infrastructure.
This pattern of rapid progress across different domains suggests that the engineering aspect of AI research is approaching full automation, while the more creative research tasks lag behind, still requiring human input.
“Clark’s conclusion is correct and possibly understated for engineering. The residual research question is real but may be less binding than the framing suggests.”
— Thorsten Meyer
Uncertain Extent of Research Automation
While engineering tasks are nearing full automation, it remains unclear how much of AI research—such as hypothesis formulation, novel algorithm development, and scientific exploration—can be automated in the near term. Clark leaves open whether research itself is reducible to large-scale engineering, which could accelerate its automation.
Next Steps in AI R&D Automation and Human Role
The next 32 months are likely to see continued progress in automating research tasks, potentially leading to a new phase where human input is primarily strategic or creative. Researchers and institutions may need to adapt to an environment where engineering is fully automated, focusing on guiding AI-driven research directions and ensuring ethical standards. Monitoring the development of new benchmarks and capabilities will be essential to assess the remaining gaps.
Key Questions
What specific tasks in AI research are still not automatable?
Tasks such as hypothesis generation, designing novel algorithms, and scientific exploration are less automatable, as they involve creative insight and abstract reasoning that current AI systems are not fully capable of replicating.
How does automation of engineering tasks affect AI research timelines?
Automation of core engineering tasks could significantly speed up research cycles, reducing the time and cost required to develop and test new AI models, potentially leading to faster innovation.
Could this trend lead to a reduction in human researchers?
While automation may reduce the need for humans in routine engineering tasks, human researchers will likely remain essential for strategic, creative, and ethical decision-making in AI development.
What are the risks associated with fully automating AI engineering?
Risks include over-reliance on automated systems, potential loss of scientific diversity, and challenges in ensuring AI-generated research aligns with ethical and safety standards.
Source: ThorstenMeyerAI.com