A new study found that advanced AI agents can perform many engineering tasks for AI research but struggle to generate original scientific contributions. Researchers tested AI agents on unpublished research questions, and while the agents completed tasks like literature reviews and running experiments, their resulting papers were rejected by reviewers.

This study highlights the current limitations of AI in performing genuine scientific discovery, indicating that while AI can automate research tasks, human creativity and original thought remain crucial for advancing scientific knowledge.
A recent study has revealed that while advanced AI agents can perform many of the engineering tasks associated with scientific research, they currently fall short of producing original work publishable at top machine learning conferences. The research, involving institutions like Princeton University, the UK AI Security Institute, and Stanford University, aimed to rigorously test the independent research capabilities of frontier AI agents.
To ensure the AI agents could not simply retrieve answers from their training data or the internet, researchers provided them with central research questions from two unpublished NeurIPS 2026 papers. Each AI agent was given six days, thousands of dollars in API credits, GPU resources, internet access, and a virtual machine to produce a conference-quality paper. The resulting papers were then reviewed by the original authors of the unpublished research.
While the AI agents successfully completed a range of engineering tasks, including conducting literature reviews, debugging software, running experiments, and managing GPU resources without human intervention, their output was ultimately deemed insufficient for publication. Reviewers concluded that the systems failed to generate original scientific contributions worthy of acceptance at a leading AI conference.
The study's authors noted that their evaluation method is a better measure of scientific reasoning than previous benchmarks, as it focuses on open-ended research problems rather than predefined tasks. However, they also acknowledged limitations, including the small sample size of two research projects and the fact that the original researchers evaluated the AI-generated papers. The findings suggest that current frontier AI agents are adept at automating many research-related engineering processes but still struggle with the core aspect of generating novel scientific insights.