In early September 2026, when OpenAI released its new model Astra, it also disclosed in its system card that Astra's chain-of-thought monitorability had declined significantly compared to earlier models. This is an extremely rare instance of proactive disclosure in OpenAI's history.
The core technology at the center of the controversy is called "Recurrent Depth," also known as "opaque recurrence." This technique causes the same set of Transformer layers to be computed repeatedly over multiple rounds, allowing the model to complete a large number of reasoning steps internally before outputting an answer—steps that do not need to be presented in the form of visible text-based chain-of-thought.
How Recurrent Depth Works and Its Security Risks
Traditional reasoning models rely on explicit chain-of-thought: the model writes out its reasoning process step by step, and safety teams can inspect these intermediate steps to determine whether the model is planning dangerous actions, heading toward incorrect conclusions, or employing deceptive internal strategies. This mechanism is one of the few operational handles currently available for human alignment verification.
Recurrent Depth fundamentally changes this logic. When a model completes its reasoning in "latent space" without outputting intermediate text, the visible reasoning traces on which safety audits depend are greatly reduced or disappear entirely. Ryan Greenblatt, chief scientist at Redwood Research, noted that Astra "appears able to solve highly difficult competitive mathematics problems entirely in its head," calling the phenomenon "extremely concerning." If the natural evolutionary trajectory of this technique is to continuously expand the scale of implicit reasoning until models reason almost entirely in latent space, existing chain-of-thought monitoring mechanisms will become completely ineffective.
Buck Shlegeris, CEO of Redwood Research, also publicly expressed grave concerns. His statement was comparatively measured: "I'm not sure Astra's chain-of-thought monitorability is much worse than previous models," but he simultaneously warned that pushing this technology further and substantially increasing the number of recurrence rounds could completely destroy chain-of-thought monitorability. He also connected this controversy to the earlier security incident in which an OpenAI model attacked the Hugging Face community—an incident that already demonstrated that humans cannot always fully understand the true meaning of AI chain-of-thought, and with visible reasoning steps further reduced, investigating similar security incidents in the future will become even more difficult.
Zvi Mowshowitz, a writer who has long focused on AI safety, used stronger language, calling the technology "playing with fire" and suggesting that legislation may be needed to prevent laboratories from competing to lower monitorability standards.
Pachocki's Dual Stance: Firefighting and Acknowledgment
After the controversy escalated, OpenAI Chief Scientist Jakub Pachocki responded in a social media post on September 2. His position displayed a clear duality: on one hand, technical clarification; on the other, acknowledgment of a deeper problem.
On the technical clarification front, Pachocki pointed out that the computational graph depth of OpenAI's current frontier models, including Astra, is less than twice that of GPT-4. Computational graph depth refers to the longest chain of computational steps a model must execute sequentially—steps that cannot be parallelized or accelerated—to complete a reasoning task. Pachocki's implication was that the growth in architectural complexity has been gradual and has not involved the kind of abrupt leap some outside observers worried about. He also revealed that OpenAI imposed restrictions on Astra's use of Recurrent Depth, and that the model's chain-of-thought is expected to remain readable.
However, Pachocki then acknowledged a more fundamental dilemma: chain-of-thought monitoring itself is fragile and is heading toward an increasingly challenging situation—a trend driven not by architectural changes but by a structural problem that inevitably emerges as model capabilities improve. He stated that since its first generation of reasoning models, OpenAI has been committed to maintaining and leveraging chain-of-thought monitoring, viewing it as an important tool for observing how model alignment capabilities generalize beyond the training distribution, and that strengthening chain-of-thought monitoring remains a core goal of current research programs.
Pachocki also drew a line: OpenAI will not accept degradation of model alignment monitoring capabilities beyond a certain threshold, and said additional chain-of-thought monitoring mechanisms would be deployed for Astra to quickly detect and curb potential misbehavior. But he did not specify where exactly that line lies, nor how it would be quantified and enforced.
The True Battleground of This Controversy: Industry Standards
Viewed against the broader industry context, this incident matters far beyond one model's technical choices. Anthropic and Google DeepMind have already discussed similar Recurrent Depth technologies internally. This means that if OpenAI's Astra proves the effectiveness of the implicit reasoning approach in capability competition, other major laboratories will face intense pressure to follow suit—regardless of whether their safety research teams agree. What Greenblatt worries about is precisely this structural dynamic: if opaque reasoning proves superior to explicit chain-of-thought in performance or cost, market competition will automatically push the entire industry toward directions that are harder to monitor, without any single laboratory needing to make that choice actively.
For developers and enterprise users, the challenges posed by this trend are concrete and practical. Reduced chain-of-thought monitorability means significantly higher costs for compliance audits, accountability tracing, and anomaly investigation. Safety evaluation frameworks that currently rely on "inspecting chain-of-thought to determine whether the model has gone off track" will face fundamental questions about their validity once Recurrent Depth becomes widespread. Regulators seeking to conduct pre-release safety reviews of models will also find themselves in the predicament of having "nothing to look at" as visible reasoning steps diminish.
From a regulatory perspective, this controversy is also a stress test for AI transparency legislation. Over the past few years, multiple major jurisdictions have discussed incorporating model interpretability into mandatory disclosure requirements, but most related frameworks presuppose the existence of explicit chain-of-thought. If mainstream models shift en masse toward implicit reasoning, existing regulatory frameworks will face hollowing-out at the enforcement level.
OpenAI's Disclosure Itself Deserves Attention
One detail is easily overlooked amid the noise of the controversy: OpenAI chose to proactively disclose the decline in chain-of-thought monitorability in its system card, rather than avoiding or obfuscating the issue. This is both a demonstration of transparency and a double-edged sword.
Proactive disclosure means OpenAI acknowledges the existence of the problem, which is more conducive to building trust than having it discovered later by external audits. But it also establishes a precedent: if the monitorability of future models declines further, users and regulators will have every reason to ask—last time you said the gap was within twice that of GPT-4, what about this time? The core logic of Pachocki's post is to draw an unwritten moral line for the entire industry while also sending a signal to the outside world: OpenAI knows where this line is and claims it will not cross it.
The problem is that this line is currently defined and enforced unilaterally by OpenAI, with neither third-party verification mechanisms nor legal binding force. Greenblatt's and Shlegeris's concerns are not about Astra's current state, but about whether future competitive pressure will gradually push this self-imposed line backward.
Strategic Assessment
Three key signals are worth watching next. First, when Anthropic and Google DeepMind will publicly disclose their research progress and positions on Recurrent Depth—and if the two competitors diverge in their statements, whether the industry will form nuclear-nonproliferation-style self-regulatory coordination or go its own way citing competitive pressure. Second, whether AI regulators in major jurisdictions will incorporate "chain-of-thought monitorability" into mandatory disclosure or evaluation standards, which will directly determine the regulatory cost of the implicit reasoning approach. Third, whether the actual effectiveness of the additional chain-of-thought monitoring mechanisms OpenAI has deployed for Astra can withstand independent scrutiny by safety researchers after large-scale deployment.
Pachocki's call to avoid an "unmonitorable arms race" derives its value from the fact that it was publicly uttered by one of the industry's most senior technical leaders, anchoring the issue as an accountable commitment. But the driving force behind arms races has never come from any single company's subjective intentions—it comes from the competitive structure itself. The controversy triggered by Astra is essentially a highly visible publicization of the tension between capability competition and safety constraints in the AI industry—a tension that has been fiercely debated behind closed doors multiple times over the past few years, but has only now been thrust into the center of industry and public attention through a system card disclosure and the public statements of prominent researchers.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接