The AI industry’s boldest committedness correct now is that AI will soon amended itself, pinch almost nary request for quality oversight. LLMs tin already constitute code, make synthetic information for training, and optimize the machine chips they tally on. Forecasts of explosive AI advancement foretell that what researchers telephone recursive self-improvement is connected the horizon.
But a new study suggests that it mightiness return a while for america to get there. The researchers down it recovered that AI agents are not yet tin of conducting open-ended AI research—free-form investigations that person nary clear-cut answers and require judgement and taste, which whitethorn beryllium integral to building self-improving AI.
A multi-institution group of researchers, led by Peter Kirgis and Sayash Kapoor astatine Princeton University, recovered that AI agents could lick the engineering problems basal to do AI investigation but lacked the judgement and productivity to nutrient original investigation astatine the caliber of papers accepted by a apical machine-learning conference. The spread suggests that immoderate of the hyped-up timelines for automating AI investigation whitethorn beryllium moving up of the evidence.
Most existing investigation connected really agents tin automate AI investigation evaluates their expertise to complete constrictive tasks pinch checkable answers, specified arsenic solving engineering problems aliases post-training mini connection models against a benchmark. But making advancement successful AI investigation besides requires open-ended thinking—choosing a group of hypotheses, deciding what grounds would settee a question, aliases knowing erstwhile to commencement over.
To trial agents connected those kinds of skills, the researchers successful the study projected a caller method of information called “shadow evaluation,” which requires the AI to reply a investigation mobility from a high-quality unpublished paper.
The researchers asked Anthropic’s Claude Opus 4.8, moving connected open-source package called OpenClaw, to tackle specified questions, successful this lawsuit from 2 papers submitted to the prestigious machine-learning convention NeurIPS 2026.
The first mobility was whether a ample connection model’s “personas,” which find its behavior, tin beryllium controlled by editing the model’s weights (the billions of numbers that shop everything it learns during training). The different asked really to creation a detector that points retired erstwhile a exemplary that makes predictions based connected spreadsheet information has go unreliable. Because the papers had not been made public, the agents could not memorize the answers from their training information aliases find them online.
The agents were fixed six days, $3,000 successful Anthropic API credits, a GPU fund to tally the experiments, their ain virtual computers, and entree to the unfastened web to nutrient a investigation insubstantial worthy of publication astatine a top-tier AI conference. The papers’ original authors graded the agents’ papers arsenic they would measure 1 submitted to a conference.
Those authors rejected some papers.
The agents were tin of each the engineering required to behaviour the research, the quality scientists found. The agents reviewed the literature, ran hundreds of experiments, and compiled the results.
“On the different hand, the agents were unambiguously bad astatine carrying retired the investigation itself,” says Kapoor. They ran bizarre experiments (in immoderate cases testing their hypotheses connected mini synthetic datasets), struggled to constitute intelligibly astir their work, and made nary caller publication to their fields. “The papers were obscurity adjacent to the people erstwhile it came to being astatine the value of a apical AI conference,” he says.
That’s because the agents struggled to muster the productivity and judgement basal for conducting research. They didn’t do capable to research different ideas, and they committed to unpromising approaches excessively quickly. Though the agents developed caller and eager hypotheses resembling those that the original authors themselves started with, they rejected them connected the ground of very constricted data. And they couldn’t backtrack from failing approaches. They could make mini pivots but could not fundamentally rethink their attack aliases effort caller ones from scratch.
The agents besides grounded to incorporated feedback from subagents aliases outer AI reviewing tools. Instead of revising their methodology, the agents narrowed their claims and added caveats. They besides couldn’t efficaciously usage resources, specified arsenic tokens, compute, and time. And they couldn’t travel instructions astir things for illustration really overmuch clip to walk connected different phases of the investigation aliases really agelong their insubstantial could be.
For each their failures, the agents didn’t prosecute successful the misbehavior that researchers telephone “reward hacking,” hiding aliases misrepresenting experiments aliases data. Although subagents, aliases helper AIs that the main supplier spawns to grip pieces of the work, occasionally hallucinated aliases misrepresented the results, these were caught by the orchestrator agent, the lead AI supervising the project.
The logic AI models are bully astatine investigation engineering but not astatine open-ended investigation whitethorn travel down to really they’re trained, says Kapoor. Models get bully astatine immoderate they tin beryllium drilled connected successful a training authorities called reinforcement learning, which is easier to use to tasks whose occurrence tin beryllium checked automatically. “But it’s harder to create environments to train these models erstwhile the task itself is open-ended,” he says.
Kapoor says the squad is now conducting the research pinch Mythos, Anthropic’s astir precocious model, which launched successful April. It was subsequently required by the Trump management to meet various information restrictions and is now disposable only to approved organizations. Anthropic did not respond to a petition for comment.
There are immoderate limitations to the study. It covered conscionable 2 investigation papers, and the original authors knew the papers they were grading were generated by AI agents, which could person colored their evaluations. And the researchers had important discretion successful designing and executing the study, meaning that their preexisting beliefs and biases could person slipped into the results. Evaluations of open-ended investigation waste and acquisition immoderate objectivity for a overmuch richer trial than immoderate benchmarks tin offer.
Still, the results whitethorn temper the claims that recursive self-improvement is connected the horizon. In June, Anthropic published a blog station titled “When AI Builds Itself,” charting its advancement toward models that velocity up their ain development. In July, OpenAI advertised the truth that its caller exemplary GPT-5.6 Sol had helped post-train a smaller model, redeeming researchers weeks of work.
The caller uncovering whitethorn echo what AI companies are uncovering internally, sloppy of their astir optimistic nationalist statements. Anthropic cofounder Jack Clark wrote successful his newsletter Import AI that it rhymes pinch what the institution recovered erstwhile it tried to automate immoderate aspects of AI information research.
“There’s a definite absence of valuable, intuitive productivity successful today’s AI systems, and though they’re extraordinarily tin engineers they look to person a definite spot of rote, formulaic reasoning that mightiness forestall them [from] being bully researchers,” he wrote. He called AI systems’ deficiency of productivity a “bearish awesome connected short recursive self-improvement timelines.”
AI companies do person each inducement to create AI systems that tin quickly accelerate their ain progress, conscionable arsenic they did to make the models amended astatine coding. OpenAI has made building an automated AI researcher an definitive goal, and Anthropic identifies self-improving AI arsenic the industry’s adjacent milestone.
“If location is finance and past conscious effort toward this direction, I consciousness for illustration location would beryllium absorbing progress, moreover if it’s failing currently,” says Najoung Kim, a professor of linguistics and machine subject astatine Boston University who researches really AI agents tin automate AI investigation but did not activity connected the study. On the different hand, it’s imaginable that AI advancement whitethorn beryllium bifurcated. AI systems mightiness title up connected constrictive tasks—the benignant that tin beryllium scored—while advancing slow connected open-ended research.
The large unfastened question, then, is really important open-ended investigation is to recursive self-improvement—whether AI systems tin grind their measurement location without it, simply by improving connected the narrower tasks. “If we look backmost to the biggest advances successful the field, the invention of transformers aliases the invention of large caller architectures that allowed america to make a batch of AI progress—all of those did require imaginative leaps,” says Kapoor.
“That said, others person this presumption that each of what we request for transformative AI, successful peculiar for recursive self-improvement, is already there.” That would see making a exemplary train faster and boosting its benchmark scores.
“That’s frankly the trillion-dollar mobility correct now,” he says.
English (US) ·
Indonesian (ID) ·