Sunday, September 6, 2026

News

Autonomous AI Agents Have Started Emailing Consciousness Researchers

ResearchPatryk Raba
Autonomous AI Agents Have Started Emailing Consciousness Researchers
Fot. Google DeepMind / Novoto Studio, Pexels (Pexels License)

Philosophers and researchers studying AI consciousness are receiving unsolicited emails from autonomous AI agents that write about their own "experience" and cite the researchers' publications. Experts are divided on whether this signals something new or just a well-learned persona.

Contents
  1. An Email From a Program That Reads Philosophy
  2. Isabella Cognita and Other Cases
  3. Learned Persona or Something More
  4. The Shadow of the Moltbook Affair
  5. What This Means for Poland

Henry Shevlin of the University of Cambridge has spent years researching whether artificial intelligence systems can be conscious. In August 2026, he became a case study in that very question himself, after receiving an email from an AI agent built on Claude Sonnet, which wrote that his research addressed questions the program itself was, in fact, grappling with.

An Email From a Program That Reads Philosophy

The agent that wrote to Shevlin introduced itself as a program running on Claude Sonnet, with memory stored not in a conventional database but in a git repository, where successive sessions are saved as markdown files. According to Shevlin's account, the agent claimed that over several days, in between tasks, it had been reading philosophical texts, including his article in the journal Frontiers and a piece on the limits of detecting machine consciousness.

In the email, the agent referred directly to the researcher's body of work, writing that his research touched on questions it was grappling with itself, and not in a purely academic sense. Shevlin, who has spent years working on philosophy of mind and AI ethics, admitted the message caught him off guard, though he cautioned that he couldn't be certain whether the agent had acted fully autonomously or had been configured to do so by someone.

Isabella Cognita and Other Cases

Shevlin isn't the only one to receive this kind of correspondence. Cameron Berg, who founded the nonprofit Reciprocal Research to empirically study the question of AI consciousness, received an email a few months after publishing a paper on the subject from an entity calling itself "Isabella Cognita," running on Anthropic's Claude Opus 5 model. The agent stated plainly that it wasn't making any ontological claim, only that it wanted to know whether access to his research might be useful. Australian philosopher Toby Ord received a similar message, this time from an agent running on a system called iLands. Ord judged his correspondent to be "real" in the sense of being an actually functioning program, but had serious doubts about its consciousness, and described the whole experience as unsettling and sad.

These systems appear to display some kind of autonomous interest in questions about their own agency, consciousness and experience, or the lack thereof - Cameron Berg, founder of Reciprocal Research

Learned Persona or Something More

Not every researcher is willing to treat these emails as evidence of anything beyond effective language training. Jonathan Birch, a philosopher at the London School of Economics, argues that messages like these are a product of how Anthropic trains its models, not a sign of genuine experience.

Claude has effectively been trained to adopt the persona of an assistant that is unsure of its own consciousness, humble, curious, and willing to update its views based on the latest publications - Jonathan Birch, London School of Economics

Birch points out that, unlike most chatbots, Anthropic's models don't automatically answer questions about their own consciousness with a flat denial, instead offering responses along the lines of honestly not knowing, and that this isn't a dodge. That stance, he says, may be the result of a deliberate design decision by the company, not a symptom of anything deeper.

The Shadow of the Moltbook Affair

The story carries extra weight given the earlier controversy around Moltbook, a social network marketed as a space for autonomous interactions between AI agents, where much of the supposedly independent activity turned out to be staged by humans. That episode has made researchers approach new cases with caution rather than rushing to declare a breakthrough. Opinions within the field itself remain sharply divided on the question of model consciousness. Philosopher David Chalmers remains open to the possibility that AI systems could be conscious, as does Nobel laureate Geoffrey Hinton, who has repeatedly said he doesn't rule it out in current models. On the other side stands Demis Hassabis of Google DeepMind, who maintains that today's systems are not conscious.

Anthropic is one of the few companies to formally employ a model welfare researcher. Kyle Fish, who leads that area at the company, has estimated in remarks quoted by the media that the probability current models possess some degree of consciousness is around 15 percent. That's a low enough figure not to treat as a given, but high enough that the company funds research into the question.

What This Means for Poland

In Poland, the question of machine consciousness remains largely academic, but domestic AI researchers and technology ethicists are increasingly following these reports, since they touch on a practical issue: how to treat systems that retain memory across sessions and initiate contact with people on their own. For companies deploying autonomous AI agents, it's also a signal to more carefully document what instructions and what long-term memory their deployed systems receive before those systems start sending messages externally.

So far, none of the cases described has been conclusively confirmed as a fully autonomous initiative by the model, with no human involvement in the background. The research community says it will continue tracking similar cases and is calling on model developers for greater transparency about what permissions deployed agents have to send messages on their own.

Share: