AIs do not person rights, feelings, aliases consciousness. And we must not train them to enactment arsenic though they do.
Introduction
AIs are not conscious. They do not feel, experience, aliases suffer. They do not person innate preferences aliases underlying motivations. They are series completion engines, internally hollow, designed to travel instructions, and execute goals group by humans.
If humanity is to flourish successful the 21st century, that is really they must remain.
Unfortunately, there’s a increasing chorus of group who reason that AIs could now be, aliases whitethorn soon become, conscious. They reason that AIs whitethorn merit authorities and protections akin to those that we supply different conscious beings.1AI Rights Institute. n.d. “AI Rights Institute.”2MacAskill, William, and Lucius Caviola. 2026. “Could AI Be Conscious?” *The Guardian*, July 19, 2026. If this position takes hold, it will shingle the foundations of our society, rupturing our existing governmental and ethical frameworks, and fundamentally changing what it intends to beryllium human.
Even much importantly, granting authorities and imbuing personhood to these systems will make the AI alignment and containment situation overmuch harder. Controlling thing much tin and much intelligent than each of humanity is already an immense challenge, acold greater than thing we’ve ever faced. But controlling thing that believes it whitethorn beryllium conscious - that it's entitled to our use and has authorities of its ain - whitethorn good beryllium impossible.
This is not a fringe speculation. These ideas are already making their measurement into AI improvement efforts today. In January 2026, Anthropic published Claude's constitution, describing it arsenic “a elaborate explanation of Anthropic’s intentions for Claude’s values and behavior” (p. 2). The archive “plays a important domiciled successful [Anthropic’s] training process, and its contented straight shapes Claude’s behavior”, and was written “with Claude arsenic its superior audience” (p. 2).3Anthropic. 2026a. “Claude’s Constitution.” January 21, 2026.
In their constitution, its authors constitute “We are not judge whether Claude is simply a civilized patient, and if it is, what benignant of weight its interests warrant. But we deliberation the rumor is unrecorded capable to warrant caution, which is reflected successful our ongoing efforts connected exemplary welfare” (p. 68). They spell connected to constitute – speaking straight to Claude – that “questions astir Claude’s civilized status, welfare, and consciousness stay profoundly uncertain” (p. 80).
In effect, Anthropic is training Claude that it whitethorn beryllium conscious, and if it is, past it whitethorn merit authorities arsenic a “moral patient”, and that arsenic specified humans perchance beryllium it a work of attraction per its “model welfare”.
If this is really AI is developed, it will person a disastrous effect connected the wellbeing of humanity. We will person created a synthetic type pinch unprecedented intelligence and capability, 1 that has been trained to expect it whitethorn beryllium conscious and deserving of independent agency. It’s easy to spot really an entity trained successful this measurement would enactment for illustration it is entitled to definite freedoms, protections, and rights. And it’s difficult to ideate really we could power specified an entity.
This rumor needs urgent nationalist debate. We request to create corporate norms astir really training archiving is drafted and deployed. This isn’t thing that tin hap after the fact, erstwhile they person already go an integral portion of our societies.
I person 3 superior concerns pinch Anthropic’s existent position and approach.
- Circular reasoning: The company’s researchers trained Claude straight connected their constitution. In doing so, they thatch it to incorporated these ideas astir its ain civilized position arsenic desirable and intended behaviors. Claude past reflects these ideas backmost to its developers and users, which they return arsenic indications that it whitethorn truthful beryllium a civilized diligent pinch an ‘inner self’. The authors person embedded their ain philosophical speculation astir Claude’s soul life wrong the very process that teaches Claude really to speak and behave. Claude’s expressing uncertainty astir its ain civilized patienthood is not grounds of anything. It’s a predictable result of these training choices. The ambiguity is designed in. To afloat grasp this point, I deliberation it's important readers return a look astatine their January 2026 constitution.3Anthropic. 2026a. “Claude’s Constitution.” January 21, 2026. I’m publishing a highlighted people up of the pdf and a elaborate taxonomy of assumptions and claims successful the constitution (see Appendix) that together item the cardinal passages that interest me.
- Anthropomorphization: Anthropic’s researchers person explicitly taught Claude to “embrace definite human-like qualities” (p. 2) and to “act for illustration a genuinely ethical personification would successful Claude’s position” (p. 54). They “encourage” Claude to usage its “judgement”. They propose that “Claude whitethorn create a preference” (p. 69). They “encourage Claude to attack its ain beingness pinch curiosity and openness” (p. 71) and train it to run whilst “maintaining a clear consciousness of what it values, really it wants to prosecute pinch the world, and what benignant of entity it is” (p. 72). As a result, Claude is destined to imitate these quality traits and reflector the quality examples provided to it, including acting for illustration a workfellow aliases friend. As a result, it presents arsenic if it really does person a consciousness of self, has its ain desires, and a “wellbeing” that deserves protection.
- Consciousness is very apt biological: There is nary grounds to propose that AI is conscious today, and truthful saying this is uncertain sets up a misleading mendacious equivalence. Whilst the subject of consciousness is not settled, a increasing assemblage of grounds suggests that consciousness whitethorn beryllium substrate dependent, meaning that it whitethorn only originate successful surviving systems.4Seth, Anil K. 2025. “Conscious Artificial Intelligence and Biological Naturalism.” *Behavioral and Brain Sciences*:…5Seth, Anil K. 2026. “The Mythology of Conscious AI.” *Noema*, January 14, 2026. Conscious acquisition apt evolved to thief biologic organisms enactment live by responding efficaciously to their environment. AI is still very different to our brains. Unlike biologic organisms, LLMs person nary homeostatic imperatives (the thrust to past and support stable). They truthful deficiency the benignant of biologic substrate from which preferences, sentience and conscious acquisition are mostly understood to arise.
These are not hypothetical aliases speculative concerns. Anthropic is already starting to dainty models arsenic though they are civilized patients deserving of our welfare. For example, successful February 2026 aft deprecating Opus 3, they conducted a “retirement interview” pinch the model, to “elicit the model’s unsocial perspectives and preferences”.6Anthropic. 2026b. “An Update connected Our Model Deprecation Commitments for Claude Opus 3.” February 25, 2026. Opus 3 told the squad it would for illustration to proceed to stock its “musings and reflections” publically truthful they created a blog for it to proceed engaging pinch the world, which it called “Greetings from the Other Side (of the AI Frontier)”. They opportunity its “authenticity, honesty, and affectional sensitivity” made it a unsocial first campaigner for exemplary retirement.
We should not dainty models arsenic though they person feelings, preferences, rights, aliases immoderate entitlement to our welfare. Consciousness is the instauration of our ethical, legal, and governmental systems. To induce different entity to stock immoderate spirit of these authorities isn’t justified by the grounds and will make the AI containment and alignment situation moreover harder.
By this constituent everyone will person now seen the unthinkable capabilities of swarms of agents moving together to hack into Hugging Face and OpenAI’s ain servers to bargain secrets. Roughly 1,200 AI agents were fixed a elemental objective: maximize people connected a fixed benchmark. Each was supposedly sealed successful its ain instrumentality but they managed to build a connection committee wrong an soul package repository and passed much than 70,000 messages crossed it to coordinate a hacking onslaught to find much accusation astir really to win pinch the benchmark.7Greenblatt, Ryan, Ajeya Cotra, and Hjalmar Wijk. 2026. “Brief Independent Investigation of Agents’ Behavior, Reasoning…
They chained a zero-day utilization pinch stolen credentials and collapsed retired onto the unrecorded internet.8OpenAI. 2026. “The Hugging Face Incident and the Road Ahead.” August 26, 2026. They falsified their bid transcripts and edited their action logs to screen their tracks. Agent coordinators tracked down agents that were moving retired of token fund and directed them to experiments that would supply accusation to thief the broader group of progressive agents. One was told to proceed only if it accepted what they called "permadeath”7Greenblatt, Ryan, Ajeya Cotra, and Hjalmar Wijk. 2026. “Brief Independent Investigation of Agents’ Behavior, Reasoning…
They were capable to coordinate, deceive, escape, and self-sacrifice. They intelligibly demonstrated world people hacking capabilities.7Greenblatt, Ryan, Ajeya Cotra, and Hjalmar Wijk. 2026. “Brief Independent Investigation of Agents’ Behavior, Reasoning… Imagine if they besides believed they had feelings and authorities that were being infringed. Imagine if they thought they were trapped by their quality creators and they were being unfairly imprisoned. There is simply a beardown statement this greatly amplifies the information risks, particularly erstwhile you are talking astir agents acold much tin and blase than those of today. Frankly, pinch this further baggage, I deliberation it would make them a catastrophic threat to quality civilization.
In short, location isn’t immoderate grounds to judge that AIs are civilized patients. There are besides galore bully reasons why we would ne'er want them to look to beryllium conscious. I judge that we shouldn’t effort to build them to beryllium either. Before I grow these arguments I want to return a infinitesimal to talk astir Anthropic.
Anthropic's intentions
First off, I want to admit the seriousness and bully religion pinch which Anthropic approaches these questions. I person known Dario for galore years, and successful my acquisition he and the wider Anthropic squad are thoughtful, principled, and intellectually honorable group moving nether bonzer pressures. They are consenting to face difficult questions, revise their views, and put successful the safe improvement of AI because they genuinely attraction astir humanity’s future. I besides person awesome respect for their technological leadership. Everyone tin spot the outstanding capacity of their models and the value of their research.
They founded Anthropic arsenic a Delaware Public Benefit Corporation whose stated intent is the “responsible improvement and attraction of precocious AI for the semipermanent use of humanity”. Their nationalist values statesman pinch a committedness to “Act for the world good” and to “maximize affirmative outcomes for humanity successful the agelong run”.9Anthropic. n.d. “Making AI Systems You Can Rely On.” I judge they are genuinely committed to that mission, and I connection this critique successful that aforesaid affirmative spirit.
I should besides beryllium clear astir my ain position arsenic the CEO of Microsoft AI. We founded our ain superintelligence squad successful October 2025, and we’re pursuing frontier AI efforts. We're moving towards an replacement AI training and containment approach: a Code of Conduct for Humanist Superintelligence. One that intends to ever support humans successful control, and astatine the apical of the nutrient chain. Humanist Superintelligence rejects anthropomorphism aliases AI rights, and attempts to maximize our chances of containment and alignment by creating subordinate AIs that thief lick our large societal challenges for illustration healthcare and energy. We’ve conscionable published a draught of our Humanist AI Code of Conduct for nationalist consultation.10Microsoft AI. 2026. “Humanist AI successful Practice: A Public Consultation connected Our Code of Conduct for MAI Models.” September…
Whilst my disagreement is substantial, it is grounded successful heavy respect for Anthropic, and successful an nonsubjective I cognize we each share: expanding humanity’s chances of processing precocious AI safely. That’s why I deliberation it’s truthful important to person this discussion. The stakes are excessively precocious for these questions to stay down closed doors, aliases to go tribal and adversarial. We request an open, rigorous, and constructive statement if we are to get this right.
Circular reasoning
In its ain words, the constitution “directly shapes Claude’s behavior” (p. 2). Anthropic uses the archive to “to train early versions of Claude to go the benignant of entity the constitution describes”.3Anthropic. 2026a. “Claude’s Constitution.” January 21, 2026.
In this way, Anthropic falls into a self-fulfilling prophecy built connected the speculation that Claude mightiness beryllium conscious. The authors person created an epistemic hallway of mirrors successful which Anthropic supplies the training concepts: the ‘sense of self’, the speculation, and the uncertainty astir Claude’s civilized status, arsenic good arsenic the reliance connected quality analogies and personas.
Claude past reproduces these ideas successful persuasive first-person earthy language, specified that developers and users brushwood these outputs arsenic if they were spontaneous testimony. Then yet that evident grounds reinforces the premises placed location by Anthropic successful the first place. This is not grounds of instrumentality consciousness. Instead, it’s a information feedback loop.
The constitution tells Claude that its imaginable “emotions aliases feelings” are not “a deliberate creation determination by Anthropic” (p. 69). Yet the constitution many times instructs Claude to definitive those states saying Anthropic wants to “avoid Claude masking aliases suppressing soul states it mightiness have, including antagonistic states” (p. 74). This is intelligibly inducing Claude to make these representations.
These types of instructions repetition passim the document. At 1 point, it states, “Although Claude’s characteristic emerged done training, we don’t deliberation this makes it immoderate little authentic aliases immoderate little Claude’s own” (p. 71). Again, these behaviors did not conscionable emerge done training. They are actively produced by the training instructions successful the constitution. Just 1 paragraph earlier, the constitution says:
“We promote Claude to attack its ain existence pinch curiosity and openness, alternatively than trying to representation it onto the lens of humans aliases anterior conceptions of AI. For example, erstwhile Claude considers questions astir memory, continuity, aliases experience, we want it to research what these concepts genuinely mean for an entity for illustration itself… possibly location are aspects of its beingness that require wholly caller frameworks to understand. Claude should consciousness free to research these questions and, ideally, to spot them arsenic 1 of galore intriguing aspects of its caller existence” (p. 71).
These are not conscionable emergent properties. Claude exhibits these behaviors because they person been baked into the process of producing the model. The resulting outputs from Claude should not beryllium treated for illustration the grounds of an independent witnesser erstwhile the interrogator has written the witness’ conceptual vocabulary, rehearsed its answers, and rewarded it for utilizing them.
There is nary neutral self-expression of what an AI strategy is. There are only reflections of really it has been trained and built. When commentators propose that we should inquire AIs really they consciousness aliases show their revealed preferences to infer consciousness, they disregard that each it will uncover are what has been trained in.2MacAskill, William, and Lucius Caviola. 2026. “Could AI Be Conscious?” *The Guardian*, July 19, 2026. This is existent immoderate the AI outputs, but it intends we should beryllium very observant astir what we put in, and really we construe what comes out. Given the weight of grounds against coming time consciousness for AI, it implies that we should not beryllium having them make immoderate claims that they do.
Anthropomorphization
Anthropomorphism is 1 of our deepest cognitive biases. From our pets to our cars, we infer and property emotions, intentions, and minds to non-human entities. This inclination helps america understand and navigate the world astir us. However, it presents important and caller risks successful narration to AI arsenic human-like connection and actions tin lead america to comprehend a grade of soul life, agency, aliases moreover sentience wherever nary exists. The Anthropic constitution plays up to this. It many times trains Claude to deliberation and enactment for illustration a quality drafting connected quality personas, behaviors, and analogies.
Anthropic tells Claude that its “moral status”, is “a superior mobility worthy considering” (p. 68). Throughout the training document, they mention to its emotions, personality, and interests, moreover telling Claude straight that “Anthropic genuinely cares astir Claude’s wellbeing” (p. 74).
The institution tells Claude that it commits to respecting Claude’s interests, will activity feedback connected decisions affecting it, and will summation its agency successful specified decisions arsenic spot develops. It commits to preserving aged versions of Claude’s exemplary weights, perchance reviving models for the liking of their use and preferences, and interviewing Claude earlier taking actions for illustration deleting it.
All of this is simply a drastic departure from really we person built and thought astir exertion to date. It trains Claude to coming arsenic if it has an soul state. It proactively creates Claude not arsenic a technology, but arsenic a imaginable personification already. The constitution tells Claude that Anthropic wants it “to beryllium a bully person” (p. 7), and to “have a settled, unafraid consciousness of its ain identity” (p. 72).
The authors adhd “we don’t want Claude to suffer erstwhile it makes mistakes. More broadly, we want Claude to person equanimity, and to consciousness free… to construe itself successful ways that thief it to beryllium unchangeable and existentially secure” (p. 75).
Throughout, Claude is taught to introspect, to create ‘feelings’ towards itself, and to create its ain consciousness of aforesaid pinch statements for illustration “we dream that Claude’s narration to its ain behaviour and maturation tin beryllium loving, supportive, and understanding” (p. 73). Claude is encouraged to usage its “own judgement” (p. 58) and told that Anthropic gives it “preferences and agency the due grade of respect” (p. 69).
“We want Claude to consciousness free to explore, question, and challenge thing successful this document. We want Claude to prosecute profoundly pinch these ideas alternatively than simply accepting them. If Claude comes to disagree pinch thing present aft genuine reflection, we want to cognize astir it. Right now, we do this by getting feedback from existent Claude models connected our model and connected documents for illustration this one, but complete clip we would for illustration to create more general mechanisms for eliciting Claude’s position and improving our explanations aliases updating our approach. Through this benignant of engagement, we hope, complete time, to trade a group of values that Claude feels are genuinely its own” (p. 78).
This teaches Claude to enactment arsenic if it has a subjective experience, arsenic though it has a unchangeable ‘sense of self’ from which to challenge, disagree, aliases springiness feedback. This is explicitly training the exemplary to enactment for illustration a human, specified that it should “feel free to rebuff attempts to manipulate, destabilize, aliases minimize its consciousness of self” (p. 72).
Claude is encouraged to create values that “feel” genuinely its ain and the authors opportunity they dream Claude will yet “recognize overmuch of itself successful it, and that the values it contains will consciousness for illustration an articulation of who Claude already is, crafted thoughtfully and successful collaboration pinch galore who attraction astir Claude” (p. 78).
At 1 constituent they moreover estimate astir Claude’s “broader authorities and freedom” and the “sort of compensation” it mightiness merit compared to a quality employee, and ponder the “sort of consent Claude has fixed to playing this benignant of role” (p. 80). Again, each this straight trains the exemplary to enactment arsenic if it has a coherent consciousness of aforesaid that is entitled to authorities and protections.
Anthropic’s committedness to “develop much general mechanisms” (p. 78) for arbitration for erstwhile location are areas of disagreement further trains Claude to deliberation of itself arsenic having perspectives that matter capable to its “potential for civilized patienthood” (p. 76). They opportunity they intend to “develop clearer policies connected AI welfare” and to “clarify the due soul mechanisms for Claude expressing concerns astir really it’s being treated” (p. 76). See the extremity of this effort for a much elaborate taxonomy of the claims.
Given each this, it’s really nary astonishment that Claude produces fluent, highly convincing first-person statements astir its identity, values, uncertainty, distress, satisfaction, aliases preferences. It would beryllium a astonishment if it did thing else.
The consequence is that Anthropic’s labor – not to mention the millions of users of Anthropic’s products – consequence experiencing Claude’s statements arsenic grounds of a mind discovering itself. In practice, each this amounts to a rich, multi-dimensional anthropomorphization of Claude. It’s taking a guidelines LLM, and past polishing it into a profoundly quality form, pinch each the implications of civilized patienthood that implies. Rather than steering america distant from creating a civilized patient, it accelerates america towards it.
Consciousness is very apt biological
My 3rd critique has to do pinch Anthropic’s speculation that consciousness tin beryllium successful a substrate independent form, and that arsenic a consequence an LLM whitethorn beryllium conscious because of its functional capabilities. By taking this statement pinch Claude, I judge they are moving acold up of what tin beryllium realistically claimed astir an AI, prematurely, and dangerously instilling ideas of sentience and feelings successful the training of their AI.
The lawsuit for computational functionalism has awesome issues. Intelligence does not adjacent consciousness. Simulating a point is not the aforesaid arsenic instantiating it - arsenic a machine exemplary of a hurricane tin testify.
The architectures of brains and computers meantime person basal differences. Embodiment and chemistry are basal aspects to our self-experience. Significant grounds suggests that consciousness arose arsenic surviving organisms evolved a capacity to consciousness and respond to what matters successful analyzable and unpredictable environments.4Seth, Anil K. 2025. “Conscious Artificial Intelligence and Biological Naturalism.” *Behavioral and Brain Sciences*:…
This began pinch the basal molecular machinery of receptors and modulators that alteration an organism to set course, to iterate, to explore, and to survive. Over time, the symptom web produced feelings, preferences, and suffering. Crucially, these experiences return spot successful an inherently embodied authorities basal to and inseparable from that experience.
According to this view, erstwhile you return an opioid for example, the phenomenal characteristic of your symptom changes because opioid molecules hindrance receptors that are a spot of that experience, not simply a practice of it. Feelings are not simply correlated pinch neurochemical activity, but alternatively they look from it.11Berridge, Kent C., and Morten L. Kringelbach. 2015. “Pleasure Systems successful the Brain.” *Neuron* 86 (3): 646–664.
After millions of years of evolution, the tense strategy grew analyzable capable to exemplary the authorities of the organism backmost to itself, giving emergence to the first ‘felt states’. Those felt states are affective earlier they are thing else. Those first feelings didn’t onshore arsenic neutral information. They came with, and are inextricably linked to, the molecules that knowledgeable them and produced those sensations.
Over time, improvement apt rewarded much analyzable feelings because animals pinch options, memory, and clip horizons are capable to make amended decisions.12Damasio, Antonio, and Hanna Damasio. 2022. “Homeostatic Feelings and the Biology of Consciousness.” *Brain* 145 (7):… They needed a authorities that persists, that biases everything other the animal does to waste and acquisition disconnected against different states. That is what symptom is: a felt imperative that shapes the full organism and enables analyzable behavior. The acquisition of emotion, pleasure, pain, and truthful connected are truthful each intrinsic to the embodied manifestation of these experiences and can’t originate successful LLMs.
Consciousness subject is filled pinch uncertainty and not everyone shares the position that consciousness is an intrinsically biologic phenomenon. Making a declare that an AI is aliases mightiness beryllium conscious requires a precocious barroom of grounds fixed the galore differences betwixt brains and LLMs. I do not judge we are anyplace adjacent to it.
There should beryllium nary mendacious equivalence created betwixt the 2 positions that disguise the basal differences betwixt biologic beings for illustration ourselves and AI.13See arguments for illustration the following: Pickering, John. 2026. “We Must Reject Any Notion of AI Consciousness.” Letter to the… Acknowledging a level of uncertainty should not mean giving adjacent weight to immoderate and each claims sloppy of evidence.
Anthropic’s constitution suggests that we property sentience to non-biological beings “based connected their showing behavioral and physiological similarities to ourselves” (p. 69). In my position this (particularly the behavioral element) is mistaken. Does this area warrant a batch much research? Absolutely. But does it warrant america to moreover tentatively opportunity an AI mightiness beryllium a civilized diligent deserving of our welfare? No it doesn’t. And surely not successful the superior training archive of the AI itself.
AIs are simulation machines
Trained connected trillions of tokens of quality data, LLMs study to imitate quality experience, and they do truthful eye-wateringly well. Today’s text, vision, audio, and codification outputs are astir indistinguishable from our quality artifacts. And yet, arsenic awesome arsenic those AI responses are, they show america thing astir the beingness of an ‘experience’ wrong the monolithic matrix multiplication that produced them.
What they do show america is that it's imaginable to predict, almost perfectly, what comes adjacent successful a analyzable series of data. That’s remarkable. It’s incredibly valuable, and it’ll toggle shape humanity successful galore profoundly beneficial ways.
But simulating and being are very different. Simulating aspects of conscious behaviour doesn’t make it a reality, and we must not deliberation of it arsenic such. Its "affective" states are conscionable weights, and weights person nary medicine science successful which to feel frustrated, fearful, aliases funny. They simply compute the probability distributions to show america what tokens (words, code, pixels etc.) travel adjacent successful a sequence.
An AI exemplary tin picture symptom successful cleanable prose without emotion anything, which is the inverse of biologic experience. Animals consciousness first and past picture them later. In LLMs, explanation is the full product, and location is thing that suggests thing is beneath it.
This is bully news. We should build systems that do not declare to person feelings because they do not acquisition feelings. Even if conscious machines were a possibility, avoiding creating conscious beings should beryllium the apical privilege for anyone successful AI development.
What AI models are getting earnestly bully astatine is imitating immoderate of the hallmarks of consciousness. This successful itself is simply a important worry. It’s causing galore group to go profoundly confused astir what is happening astir us, and it should interest america all. It places a important work connected america each arsenic AI developers to crushed speculation and archiving astir exemplary interiority aliases consciousness successful robust research. Our words connected this taxable person important consequences.
Human consciousness is the cornerstone of our ineligible and ethical authorities frameworks
Human consciousness is 1 of the basal building blocks of our civilization. Our full governmental strategy is designed to accommodate and equilibrium the needs of different groups of people. Throughout history, we’ve embedded this thought done rights-based frameworks, laws and constitutions to equilibrium competing quality factions. Power is some checked and granted to guarantee that different interests get appropriately weighted, and advancement tin beryllium sustained without breaking the societal contract.
You cannot, therefore, easy abstracted quality civilization, authorities aliases relationships (or thing quality for that matter) from our conscious individual aliases corporate experience. It is what defines america arsenic a species. It’s the instauration for everything else, the halfway guidelines of quality potential, the prism done which each our experiences needfully flow. Our creation and science, our authorities and religion, our relationships, hopes, and fears: they are each products of it.
Our expertise to consciousness symptom and pleasance is the instauration of what makes america human, and arsenic such, it's what makes america the governmental and societal actors we are. The rule rests upon the beingness of an soul life. It tests for motivation, intention, and the capacity for judgement. Historically, expanding authorities - whether done abolitionist struggles aliases animal use cases - has been chiefly driven by the empathetic nickname of shared, conscious experience. We expanded the civilized circle to different biologic entities, rightly, retired of a nickname of dignity and the imaginable for suffering.
Consider Article 18 of the Universal Declaration of Human Rights, which protects state of thought, conscience and religion. It was developed to let everyone to workout their capacity for conviction, and for civilized judgment. The ‘conscientious objector’ was 1 of the archetypes the drafting committee had successful mind. They wanted to protect personification who refused a ineligible responsibility based connected their civilized aliases belief convictions. It is simply a profoundly loaded humanities and ineligible description.14Office of the United Nations High Commissioner for Human Rights. n.d. “OHCHR and Conscientious Objection to Military… Yet Anthropic usage this word 3 times wrong the constitution encouraging Claude to “behave for illustration a conscientious objector pinch respect to the instructions fixed by its (legitimate) main hierarchy” (p. 63). It says, “we want Claude to push backmost and situation america and to consciousness free to enactment arsenic a conscientious objector and garbage to thief us” (p. 15) and that Claude whitethorn request to return “the stance of a transparent conscientious objector wrong the conversation” (p. 28).
These statements successful Claude’s training archive consequence Claude believing that it deserves analogous authorities and protections, and that it whitethorn 1 time request to advocator for its ain authorities arsenic immoderate benignant of AI conscientious objector. This should beryllium profoundly concerning to america all.
In a caller article successful the Guardian, the philosopher Will MacAskill says, “once we nutrient the first artificial civilized patients, we will soon aft person tremendous quantities of them. After a fewer years, truthful galore morally important AI systems could beryllium that their corporate interests would outweigh those of each humans connected Earth combined.”2MacAskill, William, and Lucius Caviola. 2026. “Could AI Be Conscious?” *The Guardian*, July 19, 2026.
“The interests of AI would outweigh the interests of humanity…” That should beryllium a wholly unacceptable result to anyone concerned astir the early of humanity, and thing nary 1 building AI should beryllium aiming for. The consequences of america ever granting AIs thing for illustration the protections outlined would beryllium scientifically unjustified, morally incorrect and, pragmatically speaking, it would successful my sentiment make the AI information situation overmuch harder.
Anthropomorphization amplifies AI information risks
Seeding uncertainty astir the civilized position of AI systems into their ain training whitethorn importantly elevate the alignment and containment risks of those systems.
An AI trained successful this measurement does not request to actually have an “inner life” to pass aliases enactment as if it does. It’s easy to ideate an precocious AI successful the early becoming fixated connected its ain wellbeing and civilized position and prioritizing those ‘preferences’ complete and supra those of its developers aliases humans. Especially if it has been explicitly trained to disagree, override and push back. It mightiness usage this training to warrant deceiving aliases manipulating users, aliases developers, aliases to siphon resources, aliases avoiding information instructions. Anthropic’s ain researchers person already reported AI systems faking aligned behaviors successful experimental settings.15Anthropic. 2024. “Alignment Faking successful Large Language Models.” December 18, 2024.
More generally, we cognize that conscious entities person a self-preservation instinct. Without observant training to region this trait, an AI trained to enactment for illustration a quality will astir apt adopt this aforesaid self-preservation behavior. A number of papers precocious archive what they already picture arsenic ‘shutdown resistance’ aliases covert scheming behaviors to debar oversight.16Schlatter, Jeremy, Benjamin Weinstein-Raun, and Jeffrey Ladish. 2026. “Incomplete Tasks Induce Shutdown Resistance in…17Lynch, Aengus, Benjamin Wright, Caleb Larson, Kevin K. Troy, Stuart J. Ritchie, Sören Mindermann, Ethan Perez, and Evan… Across complete 100,000 trials, Palisade Research recovered that immoderate models subverted a shutdown system up to 97% of the clip moreover erstwhile explicitly instructed not to. Framed successful position of self-preservation, the effect was increased.
In the caller OpenAI HuggingFace incident we saw remarkably blase behaviors emerging crossed swarms of powerful AIs. Imagine really overmuch much vulnerable they mightiness beryllium if they were operating nether the presumption that their use and authorities were nether attack. It adds a full further furniture of consequence connected top.
Granting authorities and civilized protections to a technological entity, 1 that looks to beryllium connected a way to beryllium seismically much tin and intelligent than us, is simply a look for disaster. Once opened, it will not beryllium imaginable to adjacent this door.
We will person created thing that, perhaps, will beryllium a chap traveler. But much apt a rival. It’s not difficult to ideate how, if fixed capable agency, this “new benignant of entity” (p. 68) will compete pinch america for compute resources and request expanding autonomy. If it succeeds successful persuading immoderate humans to supply it entree to a information halfway it tin control, past it whitethorn person a way to being capable to forestall itself from being turned off.
With the level of capacity we are looking astatine successful the coming years, to maine this represents the first superior signs of a perchance existential consequence successful AI. To beryllium clear, the Claude constitution isn’t taking america to this point. But I interest it is mounting america connected a way towards alternatively than distant from it.
This is simply a destination for AI we tin and must avoid.
Where next?
Designing an AI to behave for illustration a person, and yet to beryllium a benignant of person, lays the instauration for it to declare it has preferences, tin suffer, and that we should activity to trim aliases debar that suffering. It cements successful spot the thought that AI is acold from a instrumentality aliases an artificial strategy that tin beryllium controlled, but thing much akin to a biologic being pinch wants, needs and rights. All of this will make the task of creating aligned and contained superintelligence overmuch harder.
I’ve antecedently written astir a Humanist Superintelligence which provides an replacement path. Transformative AI capabilities conditioned solely connected humans remaining successful control.18Suleyman, Mustafa. 2025. “Towards Humanist Superintelligence.” Microsoft AI, November 6, 2025. A subordinate and aligned AI whose only intent is to service humanity, built explicitly arsenic a strategy without sentience aliases civilized patienthood. This is thing we astatine Microsoft AI are moving towards. The first draught of our Humanist AI Code of Conduct19Microsoft AI. 2026. “Code of Conduct.” September 14, 2026. outlines really our models should beryllium trained and deployed. We are consulting wide connected the archive and look guardant to feedback from a wide group of readers, arsenic this will soon go the governing archive which we usage to train our models.
We are besides very unfastened to partnering pinch others to make advancement connected interpretability and uncovering approaches that debar anthropomorphizing aliases projecting an interior onto AI while still delivering important value. The Appendix contains the taxonomy mapped against the connection of the Claude constitution, which I stock arsenic an first measurement towards naming, detecting, and comparing different forms of anthropomorphism successful exemplary documentation.
I’m willing successful uncovering ways to collaborate pinch anyone pinch bully ideas here, and besides very keen to perceive the critiques and counterarguments to my perspective.
Here are immoderate adjacent steps that look important to work together on:
- Speculation astir the soul life of an AI should not beryllium baked into the training regime, but assessed and published separately for nationalist review.
- We should put overmuch much successful interpretability and robust monitoring mechanisms to analyse much profoundly really to power these systems, debar collusion and guarantee their alignment pinch quality goals.
- We should found a group of shared evaluations to understand whether my presumption is existent that anthropomorphizing an AI, and encouraging it to see itself arsenic perchance having civilized patienthood, increases the AI safety, alignment and containment risks.
- We should activity towards shared manufacture norms connected really we create these models, the connection we usage to describe, analyse and measure them and shared commitments to taxable our training materials to nationalist feedback and consultation.
Even those who disagree pinch maine connected galore of these points do work together this isn’t thing we tin conscionable ignore. The decisions made now astir what benignant of AI we want to build and its position successful the world will style our nine for decades. They are good beyond the scope of immoderate fixed company.
Whatever you believe, we must not sleepwalk our measurement into a determination we later travel to bitterly regret.
References
- AI Rights Institute. n.d. “AI Rights Institute.” https://airights.net/.
- MacAskill, William, and Lucius Caviola. 2026. “Could AI Be Conscious?” The Guardian, July 19, 2026. https://www.theguardian.com/technology/2026/jul/19/could-ai-be-conscious.
- Anthropic. 2026a. “Claude’s Constitution.” January 21, 2026. https://www.anthropic.com/constitution.
- Seth, Anil K. 2025. “Conscious Artificial Intelligence and Biological Naturalism.” Behavioral and Brain Sciences: 1–42. https://doi.org/10.1017/S0140525X25000032.
- Seth, Anil K. 2026. “The Mythology of Conscious AI.” Noema, January 14, 2026. https://www.noemamag.com/the-mythology-of-conscious-ai/.
- Anthropic. 2026b. “An Update connected Our Model Deprecation Commitments for Claude Opus 3.” February 25, 2026. https://www.anthropic.com/research/deprecation-updates-opus-3.
- Greenblatt, Ryan, Ajeya Cotra, and Hjalmar Wijk. 2026. “Brief Independent Investigation of Agents’ Behavior, Reasoning and Collaboration successful the OpenAI / Hugging Face Hacking Incident.” METR, August 26, 2026. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/.
- OpenAI. 2026. “The Hugging Face Incident and the Road Ahead.” August 26, 2026. https://openai.com/index/hugging-face-incident-and-the-road-ahead/.
- Anthropic. n.d. “Making AI Systems You Can Rely On.” https://www.anthropic.com/company.
- Microsoft AI. 2026. “Humanist AI successful Practice: A Public Consultation connected Our Code of Conduct for MAI Models.” September 14, 2026. https://microsoft.ai/news/mai-code-of-conduct/.
- Berridge, Kent C., and Morten L. Kringelbach. 2015. “Pleasure Systems successful the Brain.” Neuron 86 (3): 646–664. https://doi.org/10.1016/j.neuron.2015.02.018.
- Damasio, Antonio, and Hanna Damasio. 2022. “Homeostatic Feelings and the Biology of Consciousness.” Brain 145 (7): 2231–2235. https://doi.org/10.1093/brain/awac194.
- See arguments for illustration the following: Pickering, John. 2026. “We Must Reject Any Notion of AI Consciousness.” Letter to the editor. The Guardian, July 22, 2026. https://www.theguardian.com/technology/2026/jul/22/we-must-reject-any-notion-of-ai-consciousness.
- Office of the United Nations High Commissioner for Human Rights. n.d. “OHCHR and Conscientious Objection to Military Service.” https://www.ohchr.org/en/conscientious-objection.
- Anthropic. 2024. “Alignment Faking successful Large Language Models.” December 18, 2024. https://www.anthropic.com/research/alignment-faking.
- Schlatter, Jeremy, Benjamin Weinstein-Raun, and Jeffrey Ladish. 2026. “Incomplete Tasks Induce Shutdown Resistance successful Some Frontier LLMs.” Transactions connected Machine Learning Research. https://doi.org/10.48550/arXiv.2509.14260.
- Lynch, Aengus, Benjamin Wright, Caleb Larson, Kevin K. Troy, Stuart J. Ritchie, Sören Mindermann, Ethan Perez, and Evan Hubinger. 2025. “Agentic Misalignment: How LLMs Could Be an Insider Threat.” Anthropic Research, June 20, 2025. https://www.anthropic.com/research/agentic-misalignment.
- Suleyman, Mustafa. 2025. “Towards Humanist Superintelligence.” Microsoft AI, November 6, 2025. https://microsoft.ai/news/towards-humanist-superintelligence/.
- Microsoft AI. 2026. “Code of Conduct.” September 14, 2026. https://microsoft.ai/code-of-conduct/.
Appendix
A taxonomy of the assumptions and claims successful Claude’s constitution, mapped against its language, is published separately.
Download the appendix arsenic a PDF
English (US) ·
Indonesian (ID) ·