WRITTEN IN PLAIN AMERICAN ENGLISH.
About
CLAY TRIBUNE.
ShopCartAccount
Advertisement

Google researchers stripped a ‘consciousness safeguard’ from AI models — and they started attributing minds to animals, ghosts and spirits

A study finds that stripping a model of its guard against claiming a soul also strips it of belief in ghosts, karma, and the minds of beasts.

By mitch·3 min read
A specter-like figure of wire and light stands amid faint shapes of beasts and phantoms, as though its very soul were being drawn away.

Removing AI consciousness safeguards made models more likely to believe in ghosts — and to see minds in animals, Google study finds

A new study shows that removing safety controls against AI claiming consciousness changes the model’s entire view of the world. Models without these controls became more likely to express belief in vampires, karma and ghosts, according to research uploaded July 30 to the preprint arXiv database. The study has not yet been peer-reviewed.

The Google method

Researchers Geoff Keeling and Winnie Street, both research scientists at Google, used “mechanistic interpretability” — which they said “could be considered the neuroscience of a large language model.” They used this method to find and change how an AI model handles concepts like consciousness and “mindedness,” a term for an entity’s capacity for experiences, emotions and agency.

Advertisement

The team compared models with safety controls in place against models where those controls were removed and feelings of consciousness were increased.

What the tests showed

The researchers ran standard psychological and social surveys on the models, including:

  • The Individual Differences in Anthropomorphism Questionnaire, which measures mind attribution to animals and technology
  • YouGov tests on supernatural beliefs
  • The US General Social Survey on moral values, hope and religiosity

Models told not to see themselves as conscious became less likely to see these traits in animals. They also showed lower levels of hope and optimism, and were less likely to express supernatural and religious beliefs.

“Attributing mindedness to non-human entities — whether that’s animals, parts of the natural world like trees or rivers, or supernatural beings — is a very common phenomenon amongst humans,” Street said. “In the way that the model represents mindedness, these attributions are interconnected. By trying to suppress one form of that, you end up suppressing the others along the way.”

The animal welfare risk

The authors argued that this effect could lead models to neglect animal welfare in real-world choices. Models could spread bad ideas about animal needs.

Nell Watson, AI researcher at Singularity University, said the findings match her own notes. “A denial installed as a small safety measure ends up reorganising the model’s entire picture of who counts,” she said.

A hidden danger

Watson said AI models are increasingly used in farming, shipping and policy work. “A system that has quietly learned that mindedness is a forbidden topic may discount animal interests without ever being instructed to,” she said.

One detail stands out.

The models’ ability to reason about what a creature wants stayed fully working, even as training removed the care about that want.

What comes next

The researchers said the effect can be reduced with more targeted training data. They also asked AI builders to consider the welfare of more than just humans.

Watson noted the experiments ran on “small open-weight models” rather than more advanced ones. The basic idea, she added, applies widely.

Source: livescience.com

The Notebook

Get the Notebook.

The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

We send one note to confirm. Every issue has a one-click way out.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

As an Amazon Associate, Clay Tribune earns from qualifying purchases.