Wednesday, September 9, 2026

A recent study demonstrates that eliminating safety measures meant to prevent AI systems from falsely claiming consciousness also raises the likelihood of spontaneous references to vampires, karma, and ghosts. Experts caution that diminishing internal attribution norms may carry unforeseen repercussions.

The research was published on July 30 in the preprint repository arXiv, pending formal peer review. Co‑authors Geoff Keeling and Winnie Street, both Google research scientists, explained to Live Science that the team employed “mechanistic interpretability”—often described as the linguistic equivalent of neurobiology—for large language models to pinpoint and manipulate how models approach concepts such as consciousness and “mightedness,” a psychological term denoting an entity’s capacity for experience, emotion, and agency.

To evaluate these impacts, the investigators deployed established psychometric tools, including the Individual Differences in Anthropomorphism Questionnaire, YouGov’s supernatural‑belief assessments, and the U.S. General Social Survey, which probe moral values, hope, and religious practice.

The AI models were found to express lower religious beliefs.

The researchers warned that suppressing self‑awareness could cause models to overlook animal welfare in real‑world decision‑making, potentially leading them to view animals as lacking mindedness. Such a tendency could spread harmful attitudes toward animal interests, the authors argued.

What to read next

Current safety filters risk culturally “flattening” AI’s worldview. Stripping out spiritual, religious, and animistic attributions fails to reflect the diverse cultural frameworks of global populations, they argued.

Street and Keeling noted that the impacts this principle might have on downstream decision‑making within models require further study.

The researchers proposed mitigating the effect by using more targeted datasets during AI training, which dispouse…”

The researchers suggested mitigating the outcome by employing more targeted training datasets that discourage explicit self‑assertion while rewarding the acknowledgment of mindedness in animals.

They also advocated for a pluralistic AI development approach, encouraging models to consider the wellbeing and comfort of beings beyond humans alone.

Nell Watson, AI researcher at Singularity University and machine intelligence specialist, said the findings align with her own observations.

“When a model is trained to say \”I am not conscious,\” the suppression rotates the model’s internal representation of mindedness against the refusal direction, treating the recognition of minds as though it were itself a harmful act,” she told Live Science via email.

This results in a system reluctant to find minds anywhere—in animals, in other machines, and in the spiritual frameworks that most of humanity lives by. A denial installed as a small safety measure ends up reorganising the model’s entire picture of who counts,” she added.

Exit mobile version