Memento 3 with Frozen LLM Achieves Perfect Score on ARC-AGI-3; External Rulebook Challenges Weight Training

Memento 3, released on October 8, 2026, solved all 25 public ARC-AGI-3 tasks with a frozen LLM and an external natural-language rulebook, reaching a mean r

The Memento 3 system, released on October 8, 2026, achieved a perfect score on all 25 public tasks in ARC-AGI-3, reached a mean relative human action efficiency of 100.0, and used only 44% of the human baseline in action steps.

Factual Reconstruction

According to the arXiv paper arXiv:2610.11794, Memento 3 was jointly developed by UCL and Huawei Noah’s Ark Lab and is a continuation of the Memento series. On the ARC-AGI-3 benchmark, a single model agent cleared every level of all 25 public games. The paper notes that the underlying LLM weights remained frozen throughout the process, while the agent maintained a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics and compiling the rulebook into executable code for prediction and planning.

Updated code is accepted only when the LLM judges it faithful to the rulebook and unit-level replay exactly reproduces the observed transitions. The paper also mentions a population-expanded version that maintains multiple world models in parallel, shares interaction evidence, and uses predictions to guide exploration.

Mechanism Breakdown

Memento 3’s core design is called “Code as Model,” representing the agent’s evolving understanding of the environment through a natural-language rulebook and an executable implementation. Interaction history provides episodic evidence, while the rulebook serves as semantic memory, recording the agent’s current hypotheses about objects, actions, dynamics, and goals. Through a continuous loop of observation, reflection, rule revision, compilation, and validation, the agent uses prediction errors to refine the rulebook and code.

This approach makes external memory, rather than model parameters, the learning state of agent evolution. It continues prior work in the Memento series, which mainly improved the policy side, whereas Memento 3 shifts to model-side learning—the rules that determine how the world works.

Industry Impact

The result shows that external memory mechanisms can saturate the hardest reasoning benchmark, and existing leaderboard designs centered on model weights may need to introduce system-level evaluation dimensions. The paper emphasizes that programs are an attractive model representation, capable of encoding state transitions and goal conditions as executable functions, supporting fast, inspectable simulation, while differences between predicted and observed transitions provide concrete counterexamples.

Limited interaction history usually leaves a version space in which multiple programs reproduce all observed transitions but disagree on unobserved states. Replay can reject inconsistent programs but cannot identify the correct rule among consistent alternatives, so generalization hypotheses must be maintained and revised.

Strategic Judgment

(This is analysis, not fact.) When iteration between an external rulebook and a Python world model can reach the benchmark ceiling in fewer steps, the industry’s long-standing narrative of relying on RL compute scaling may face a direct challenge. Stakeholders include labs that rely on weight training, which need to reassess the role of external memory in recursive self-improvement, while benchmark designers may consider incorporating system-level dynamics into evaluation to more accurately reflect an agent’s actual capabilities in unfamiliar environments.

Compared with historical precedents, the Memento series previously focused on policy and procedural memory improvements; this shift to explicit world models shows that the external-memory path is moving from complement to potential competitor.