📊 Full opportunity report: Engineering Is Automated. Research Is the Residual. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
AI systems have achieved near-complete automation of core engineering tasks in AI R&D, with benchmarks indicating saturation. However, research activities still rely heavily on human insight. This development could accelerate AI progress but raises questions about the future role of human researchers.
Recent empirical data and expert analysis confirm that AI systems are now capable of automating the majority of core engineering tasks in AI research, while the research process itself remains less automated. This shift could significantly accelerate AI development timelines and reshape institutional strategies, making engineering a largely automated process.
Thorsten Meyer’s review of recent benchmarks reveals that AI models have reached near-saturation levels in core engineering skills, such as reproducing research experiments and competing in Kaggle-like challenges. For example, the CORE-Bench, which measures research reproduction ability, improved from 21.5% in September 2024 to 95.5% in December 2025, with one author declaring it ‘solved.’ Similarly, the MLE-Bench, assessing performance in Kaggle competitions, rose from 16.9% to 64.4% over sixteen months, surpassing mid-tier human performance.
Experts attribute these advances to improvements in AI model capabilities, including better kernel design and automated code conversion, which are now integrated into production-grade systems. The pattern across multiple benchmarks indicates a rapid approach toward saturation, suggesting that the engineering aspect of AI R&D is becoming fully automated. However, the research process—generating new hypotheses, designing experiments, and interpreting results—remains less automated, with the structural question of whether research can be automated at scale still open.
Engineering is automated.
Research is the residual.
Six skill benchmarks. Edison’s framing. The question Clark leaves open is whether research is just engineering at scale.
Jack Clark’s Import AI #455 catalogs six benchmarks measuring AI capability on AI R&D tasks and concludes “AI can today automate vast swatches, perhaps the entirety, of AI engineering.” The residual question is research. The structural read on the residual: it may not be a permanent moat.
Six skills. One trajectory.
Clark catalogs six benchmarks measuring AI capability on AI R&D-relevant tasks. Each individual benchmark could be noise. Six benchmarks moving together is a curve. The pattern is the cascade observed across the broader Clark series — visible here in the specific R&D-skill domain.

The No-BS Guide to AI for Trading & Market Research: How to Use ChatGPT, Claude & AI Tools for Market Analysis, Stock Research & Data-Driven Trading … … Required (The No-BS AI Playbooks Book 3)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three data points. Mixed signal.
Clark provides three data points on the creative-spark question. Yes-evidence: Erdős-1051, centaur math discovery, sporadic Move-37-style moments. No-evidence: low yield, framing dependence, absence of acceleration. The mixed signal is the honest read.
The data supports two readings. Pessimistic: rare moments suggest creative insight is qualitatively distinct from engineering work. Optimistic: rare moments are an artifact of low-volume exploration; more shots on goal yields more discoveries. Both readings are consistent with Clark’s “vast swatches, perhaps the entirety” claim. They differ on the residual.

AI-Powered Real Estate Investing: The 2026 Guide to AI Tools, Prompt Engineering & Automated Systems for Building a Million-Dollar Property Portfolio
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five dimensions Clark gestures at but leaves underdeveloped.
Clark’s section is rigorous on the empirical evidence. Five strategic dimensions matter for the institutional response that the Clark series synthesis argues is structurally inadequate.

AT-AP Multi-purpose Processing Assistance Platform For Model Tool
Brand Name:NoEnName_Null
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Two readings. Different equilibria.
The structural question Clark leaves open: is research a permanent moat that bounds automated AI R&D, or is it engineering at scale that dissolves with more shots on goal? Both readings are consistent with the current data. They differ by orders of magnitude in consequences.
Productivity multiplier years
Recursive loop operational

The 90-Day AI Plan for Business Leaders: A Step-by-Step Action Plan to Launch AI Projects That Work Without Technical Skills, Big Budgets, or Time-Wasting Experiments (Generative AI made Practical)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five audiences. Asymmetric cost of being wrong.
The institutional response should not bet on inspiration being a permanent moat. If the distinction holds, capacity built is still useful. If it closes, capacity is necessary. Asymmetric cost-of-being-wrong points toward building now.
IN INDUSTRY
IN ACADEMIA
POLICYMAKERS
INVESTORS
EVERYONE ELSE
Engineering is automated. The residual is the question. The institutional response should not bet on inspiration being a permanent moat.
Impact of Automation on AI Development Speed
The near-complete automation of core engineering tasks in AI research implies that the bottleneck in AI development may shift from engineering to research itself. This could lead to faster iteration cycles and more rapid deployment of AI systems. However, it also raises questions about the role of human researchers and the potential need for new institutional responses to manage increasingly autonomous AI-driven research processes.Recent Advances in AI Engineering Capabilities
Over the past two years, multiple benchmarks and research efforts have documented rapid progress in AI’s ability to perform engineering tasks. The CORE-Bench, MLE-Bench, and kernel design research demonstrate that AI models are now handling tasks traditionally performed by human engineers, such as reproducing scientific experiments, competing in data science competitions, and designing optimized hardware kernels. These developments suggest that engineering in AI R&D is approaching full automation, a trend highlighted by Thorsten Meyer’s analysis of Clark’s empirical work.“The pattern across benchmarks indicates a rapid approach toward saturation, suggesting that the engineering aspect of AI R&D is becoming fully automated.”
— Thorsten Meyer
Unresolved Questions About AI Research Automation
While engineering tasks are rapidly becoming automated, it remains unclear how much of the research process—such as hypothesis generation, experimental design, and interpretation—can be automated at scale. The structural question Clark leaves open is whether research itself is a form of engineering at scale, which could mean the residual gap closes faster than anticipated. Additionally, the institutional and strategic implications of fully automating engineering are still under exploration.
Next Steps in AI Automation and Research Development
Researchers and institutions will likely focus on developing methods to automate research activities further, including hypothesis generation and experimental design. Monitoring the evolution of benchmarks and real-world applications will be crucial to assess whether research automation progresses at a similar pace. Policymakers and industry leaders should prepare for potential shifts in research workflows and the role of human researchers in AI R&D.
Key Questions
How close are AI systems to fully automating AI research?
While engineering tasks are nearing full automation, the automation of research activities such as hypothesis creation and interpretation remains uncertain and is an active area of investigation.
What are the implications for human researchers?
If engineering is fully automated, human researchers may need to focus more on strategic oversight, hypothesis development, and ethical considerations, rather than routine experimentation.
Could this accelerate AI development timelines?
Yes, automating core engineering tasks could significantly reduce the time required to develop and deploy new AI systems, but the pace depends on progress in automating research activities.
Are there risks associated with fully automated AI research?
Potential risks include loss of human oversight, difficulty in understanding AI-generated research, and ethical concerns about autonomous scientific discovery. These issues are still under discussion.
Source: ThorstenMeyerAI.com