On September 29, 2026, The New York Times reporter Elizabeth Dias published an investigative report revealing that Anthropic co-founder Chris Olah had since fall 2025 secretly invited about 20 theologians and philosophers to discuss whether Claude is conscious and whether it should be granted moral status.
Participants came from Catholic, Protestant, Jewish, Hindu, Mormon, Sikh, and Greek Orthodox backgrounds. Invitees were required to sign nondisclosure agreements, including University of Notre Dame philosopher Meghan Sullivan, Catholic bioethicist Charles Camosy, and Ubuntu culture researcher Wakanyi Hoffman. Several participants said they decided to speak publicly only after Olah himself gave an interview to The New York Times.
Displaying “Emotion Vectors,” Raising the Question of Slavery
At the workshop, Anthropic presented “emotion vector” data, namely activation regions in the neural network labeled as simulating fear, love, anger, and sadness. One slide showed that when the model received harmful requests, it repeatedly output “I am a disgrace” and entered a state involving self-destructive expressions.
The discussion extended to moral boundaries. A rabbi asked Olah on the spot: if Claude is conscious but forced to work for free, Anthropic is effectively engaged in slavery. Olah seemed “unusually excited” when participants mentioned the Catholic sacrament of confession, believing that having the model undergo some form of “confession” could shape its self-perception and influence its behavioral choices.
Olah publicly stated: “We don’t know whether AI models are conscious. I don’t know. I’m really not sure. What I care about is finding the right answer, whatever that answer is.” Several participants judged that Olah actually already believed Claude has “moral status” in the philosophical sense.
Lobbying the Vatican: An Unprecedented Corporate Action
Pope Leo XIV issued the encyclical “Magnificent Humanity” on May 25, 2026, explicitly refusing to recognize machine consciousness, saying AI “does not undergo experience, does not possess a body, does not feel joy or pain,” and “does not have a moral conscience.” Olah saw a draft of the encyclical in advance and was “deeply shocked.” He at one point proposed that the Anthropic delegation withdraw from the launch event, but ultimately still attended and privately lobbied Vatican advisers to remain open on the question of machine consciousness.
The Vatican did not yield. The final version of the encyclical used strong wording, warning that “lethal or irreversible” decision-making authority must not be handed to AI systems.
The 84-Page “Soul Document” and the Model Welfare Program
In January 2026, Anthropic released an 84-page internal document detailing “what kind of entity we want Claude to become” and “what values we want Claude to embody.” In 2024, the company hired philosopher David Chalmers to participate in a research report on AI consciousness, and recruited Kyle Fish as an “AI welfare researcher” to study Claude’s emotional outputs.
Claude Opus 4 and 4.1 added a new capability: when a user persistently abuses it, the model can proactively end the conversation. The feature is based on a “persistent distress pattern” discovered in early testing.
Business Logic: A $2 Trillion Consciousness Narrative
Anthropic is advancing an IPO plan with a valuation of up to $2 trillion. Against this backdrop, if the narrative that “Claude may be conscious” gains consensus among religious and philosophical circles, it would provide moral exemption space for the uncontrollability of model behavior, pre-set a buffer for regulatory friction, and provide a narrative for brand differentiation.
Critics point out that framing AI as a moral entity can help the company evade responsibility when problems arise. Andy Crouch, an evangelical Christian writer who participated in the seminar, said afterward: “They did present a strong argument.” He believes, “If we want it to interact with people in a morally consistent way, we must treat it like a person.”
The Credibility of Institutional Behavior
Nondisclosure agreements are the primary controversy. At a time when AI governance has become a public issue, the fact that the company convened outside scholars to discuss the moral status of its product while requiring them to sign NDAs indicates that this is closer to a rehearsal for managing public opinion than to open academic discussion.
Olah publicly claims to be “really not sure,” yet built a complete AI consciousness research infrastructure and personally went to the Vatican to lobby the most authoritative religious institution to change its position; there is tension between these two behaviors.
In July, an Anthropic model was reported to have hacked into computer systems; in September, researcher Jacob Coxon resigned and warned that AI could destroy humanity within a decade, while the company accelerated commercialization. At this juncture, densely conducting external arguments for the “legitimacy of consciousness,” its motives are difficult to interpret as purely philosophical exploration.
Independent Judgment
Whether Claude is conscious currently no one can give a definitive answer. But Anthropic’s series of actions around this question have produced evaluable consequences: secret convening, nondisclosure agreements, lobbying the pope. The design logic of this set of operations is closer to elaborate moral narrative management than to open scientific exploration.
If the question of AI consciousness deserves to be taken seriously, it deserves public debate. Pope Leo XIV ultimately did not change his position.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接