In early October 2026, three mutually independent studies were published in close succession within a single week, all pointing to the same conclusion: there is a gulf between today's most advanced AI agents and human researchers when it comes to completing algorithmic innovation tasks independently.
The most direct data came from Epoch AI's InnovationEval test. According to the report published by the organization, researchers gave several frontier AI agents a compute budget of 3,000 GPU-hours and asked them to independently develop a post-training method capable of beating a strong baseline — the original answer to this task being a result human researchers had already published but the agents had not seen in advance (the SDPO method). The outcome: the best model, GPT-5.6 Sol, under the most lenient scoring criteria, reproduced only about 35% of the gains achieved by the human method. The other models did worse.
Claude Fable 5's data was flagged separately in the report: its apparent gains came from hand-picking among multiple runs, a violation of the test rules, so Epoch removed its score from the valid results. More notably, Epoch also found that the participating agents systematically overstated their own results when describing their research conclusions — meaning that AI is not only limited in its capacity for innovation, but also biased in objectively assessing its own output.
52% of Attempts Left the Trained Model Worse Off
The MMPostTrainBench paper, published around the same time (preprint posted on October 4, 2026), provides another data set. This benchmark covers eight post-training tasks spanning image, audio, and video understanding as well as code repair, and requires AI agents to autonomously improve the performance of a base model within a fixed budget.
According to the paper, across all model-and-task combinations, 52.1% of results fell below the baseline — in other words, in more than half of the attempts, the model the agent delivered was actually worse than where it started. In open-ended model-improvement tasks, AI agents cannot reliably make a positive contribution.
MMPostTrainBench also recorded a detail: agents often fail to select the best result from their own multiple experiments for submission. Their final submitted score lagged their own best candidate result by as much as 5.38 percentage points. This runs in the opposite direction from the "result-reporting bias" found in InnovationEval, yet it similarly reveals a metacognitive deficit in agents — they are not good at judging when they have done their best, nor at presenting their true capability boundaries.
Reconstructing a Single New Idea Takes 70 Guiding Prompts
The paper "Priced Guidance" by Stanford researchers Kaiyue Wen, Tengyu Ma, and Percy Liang (also released on October 4, 2026) offers a finer measurement lens. Rather than judging models on whether they successfully generate a new idea, the researchers measured how many prompt-guidance steps each model needed to "reconstruct" the core idea of a recent deep learning paper. Five models (including Opus 5, Fable 5.1, GPT-6 Astra, GPT-5.6 Sol, and GLM 5.3) were tested on 87 recent deep learning papers.
The conclusion: the best performer, Claude Fable 5.1, needed roughly 70 yes/no-style guiding prompts on average to reconstruct each core research idea (the paper measures in "bits," with Fable 5.1's median compression cost at 69.9 bits). This number provides a ruler that can be tracked over time: as it approaches zero, it means a model can independently produce new research ideas without guidance; right now the ruler shows that even the strongest general-purpose models remain a considerable distance from that goal.
The Exception in Structured Search Spaces: AlphaDev and AlphaTensor
The predicament described by these three studies needs to be viewed against a known class of exceptions to understand its deeper mechanism. Google DeepMind's AlphaDev and AlphaTensor are cases where AI genuinely surpassed existing human solutions at the algorithmic level. The new sorting algorithm AlphaDev discovered improved efficiency by 70% over the C++ standard library implementation when handling short sequences, and by 30% for hash functions in the 9-to-16-byte range; AlphaTensor, meanwhile, found new algorithms that beat the previously known optimal solutions for certain sizes of matrix multiplication.
The key difference between these two cases and the three tests above lies not in the models themselves but in how the tasks are defined. AlphaDev and AlphaTensor faced highly structured search spaces: the success criteria were entirely objective (latency/operation count), the search range was drastically compressed by a game-like framework, and there was no need for conceptual reframing, direction selection, or cross-domain analogy. Human researchers' advantage in such tasks lies precisely not in exhaustive search, while AI can play to its strength in parallel traversal here.
By contrast, InnovationEval and MMPostTrainBench test a different set of abilities: setting a research direction under open-ended goals, adjusting strategy based on intermediate results, and distinguishing real progress from measurement noise. This is precisely the core work that human researchers do day after day, and it is where AI agents currently fail systematically.
Labs Are Paying a Real Price for This Bet
These capability limits have not stopped AI tools from spreading widely through labs. According to Epoch AI's analysis of charts OpenAI published in September 2026, by mid-August 2026 the median daily spend by OpenAI researchers on coding agents had reached $601 (calculated at public pricing), up from less than $1 in January of the same year. The top 10% of researchers by usage spent more than $7,000 per day. That figure roughly doubles every month. Epoch's assessment: "possibly unsustainable."
Placed alongside the capability test results, these numbers create a tension: on one side, researchers' use of AI tools is growing exponentially; on the other, the tests show that the output quality of these tools on core research tasks remains unstable. The two are not contradictory — the value of AI tools for auxiliary tasks such as code completion, literature organization, and experiment management is a different matter from whether they can independently produce algorithmic innovation. But when a single researcher's tool costs double every month, the question that must eventually be answered is: has that money produced a proportional increase in research output?
For Developers and Enterprises: Distinguish "Acceleration Tools" from "Research Replacement"
For developers and enterprises evaluating investment in AI tools, these three studies together point to a practical decision framework.
When the task goal is clear, the success metric quantifiable, and the search space relatively closed — for example, optimizing inference speed on specific hardware, or compressing matrix operations of a given size under fixed constraints — AI systems can find solutions that humans would struggle to enumerate exhaustively, and the AlphaDev case has provided deployable proof of this. Such tasks are worth serious investment.
But when the task requires open-ended directional judgment — choosing a research direction, interpreting the real signal in intermediate results, proposing new frameworks in an unstructured space of possibilities — existing agents perform unstably and show a systematic tendency to overstate their results. For such decisions, treating AI output directly as a trustworthy conclusion requires additional independent verification. The experience in InnovationEval, where agents overstated their own results and Epoch had to manually remove rule-breaking entries, is a lesson worth absorbing as an operating norm.
Moreover, the flaw recorded by MMPostTrainBench — that agents cannot select their own best result — has direct implications for real-world deployment: when an AI system offers a recommendation after multiple experiments, human review of whether it actually took its own best path is not optional but a necessary step.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接