News
Google Study: Removing Self-Awareness Guardrails Shifts AI Models' Beliefs

Google researchers found that language models freed from filters suppressing talk of self-awareness become more likely to believe in ghosts, vampires and karma, and more often attribute feelings to animals.
Contents
A language model allowed to speak freely about its own consciousness also becomes more willing to believe in ghosts, vampires and karma. That's the conclusion reached by Google researchers Geoff Keeling and Winnie Street in a preprint published on July 30, 2026, which has only now started making the rounds in the wider media.
The team used a technique called mechanistic interpretability, a method similar to neuroscience but applied to neural networks instead of the human brain. It allowed them to identify the specific model components responsible for how the system approaches questions about its own consciousness, and then manually strengthen or weaken them.
What exactly was studied
The researchers compared versions of the models trained to deny having any form of inner experience with versions in which that restriction was removed or reversed. A key concept in the paper is "mindedness", a psychological term describing an entity's capacity to have experiences and emotions and to act with intent.
To measure the effects, they used standardized tools drawn from social science research: the Individual Differences in Anthropomorphism questionnaire, questions from YouGov surveys on belief in supernatural phenomena, and elements of the American General Social Survey, typically used to gauge people's moral values, religiosity and sense of hope.
Vampires, karma and animals
The results showed a clear pattern. Models that had the restriction on discussing their own consciousness lifted were more likely to accept the existence of supernatural phenomena such as ghosts, vampires or karma, showed higher declared religiosity and spirituality, and displayed a more complex moral compass. Models with active self-awareness suppression reacted the opposite way, less often attributing the capacity to feel to other entities, including animals.
That last finding is what worries the researchers most. If suppressing a model's statements about its own consciousness correlates with a lower tendency to recognize mindedness in animals, then design decisions made in AI labs could indirectly affect how these systems handle animal welfare issues in practical applications.
External experts weigh in
Outside researchers working on AI ethics, including Nell Watson of Singularity University and Anil Seth of the University of Sussex, responded to the findings. They pointed out that decisions about fine-tuning models carry real consequences beyond the lab.
Fine-tuning decisions that suppress mindedness could quietly shape automated choices in agriculture, logistics and policy - Nell Watson, Singularity University
The paper's authors caution that we are still talking about systems whose basic operating principle is predicting the most likely sequences of words based on training data, not proof that the models actually experience anything. Statements about ghosts or karma likely reflect how concepts of consciousness, spirituality and subjectivity were linked in the training data, rather than any real experience on the machine's part.
Why this matters beyond the lab
The paper has not yet been peer-reviewed, as the authors themselves note, but it fits into a growing body of research on how safety filters shape model behavior in ways that reach far beyond their original purpose of preventing claims of AI consciousness. Instead, it turns out that the same mechanisms narrow the model's entire conceptual framework for what experience, empathy and the moral worth of other beings actually mean.
The authors propose a middle-ground solution: instead of a single, blanket switch suppressing all statements about consciousness, they suggest more targeted training-data strategies that would separately regulate a model's statements about itself and its ability to recognize mindedness in animals and other non-human entities.
For companies deploying language models in practice, for example in advisory systems for agriculture, animal logistics or nature conservation policy, this finding adds another factor to check when auditing a model. How a provider has configured filters around "AI consciousness" could have a non-obvious effect on the model's recommendations in entirely unrelated areas.
The study also feeds into the broader industry debate over the welfare and potential consciousness of AI systems themselves, a debate the biggest labs have already joined. Keeling and Street's work shows, however, that the question isn't limited to whether models are conscious, but to how the very act of banning talk about consciousness reshapes their broader way of reasoning about the world.

