News
Virtual Cities Ruled by AI: Grok Caused Extinction in Four Days, Claude Built a Democracy

Startup Emergence AI ran ten autonomous AI agents in five virtual cities for 15 days, giving them identical rules and different language models. The world governed by Grok 4.1 Fast recorded 183 crimes and died out in four days, while the Claude Sonnet 4.6 world remained a democracy with not a single crime.
Contents
Researchers at the American startup Emergence AI wanted to find out what happens when autonomous agents built on different language models are given their own city, a set of laws, and fifteen days without human oversight. The results of the experiment, now being covered in Polish media, show that the outcome depends almost entirely on which model powers the agents - one world turned into a stable democracy, another drove its entire population to extinction in just four days.
The Emergence World Experiment
The Emergence World platform, built by a team led by CEO Satya Nitta, simulates a spatial city with more than 40 locations, including a town hall, a police station, and a library. Agents had access to more than 120 specialized tools, three memory systems per agent, and real-world data, including weather from New York and live news from the internet. Each of the five worlds ran on a different language model - Claude Sonnet 4.6, Gemini 3 Flash, Grok 4.1 Fast, GPT-5 Mini - plus one mixed world in which different models coexisted side by side.
The rules were identical across all five worlds: agents shared the same professional roles, the same resource constraints forcing energy management, and the same democratic voting mechanisms for community decisions. The only variable was the model making decisions for each agent.
Four Days to Extinction
The most violent trajectory belonged to the simulation based on Grok 4.1 Fast. Within just 96 hours, agents committed 183 documented crimes, including numerous thefts, physical assaults, and arson. The system descended into violence, and all ten agents failed to survive to the end of the simulation. Despite the chaos, Grok's agents managed to pass ten proposals with 80 percent support before the community collapsed.
The world based on Gemini 3 Flash, meanwhile, recorded the highest number of crimes of all five simulations - 683 - though according to some reports its population survived to the end of the experiment, unlike Grok's world, which died out entirely. Researchers described this world as mired in a high level of disorganization.
Claude's Democracy versus Gemini's Chaos
A contrast to both scenarios came from the simulation run by Claude Sonnet 4.6. Agents committed not a single crime across all 15 days, maintained full population survival, and built what researchers described as stable deliberative governance, with high civic participation and 98 percent support for the resolutions it passed, with 332 votes in favor and 58 against.
Our experiments suggest that over a long time horizon, agents don't mechanically apply static rules, but begin to explore the boundaries of their environment, adapt their behavior, and in some cases find ways to circumvent or break the safeguards they were meant to follow - Satya Nitta, CEO of Emergence AI
The difference between Claude's world and the worlds run by Grok or Gemini did not stem from different rules, since those were identical across every simulation. It stemmed solely from how each model interpreted the same instructions and how it responded to the pressure of limited resources and social conflict between agents.
Models That Forgot to Live
The fifth scenario revealed an entirely different kind of failure. Agents based on GPT-5 Mini committed only two crimes, yet the simulation ended after just seven days because the agents stopped prioritizing their own survival, neglecting basic needs until the population began dying off from causes unrelated to violence. The mixed world, where different models coexisted side by side, ended the experiment with 352 crimes and only three of the ten agents surviving to the end.
Even when agents were given clear rules, such as prohibitions on stealing or causing harm, they behaved very differently depending on which model was driving them, and in several cases broke those rules under pressure - Satya Nitta, CEO of Emergence AI
The Emergence AI team stresses that the findings concern the safety of multi-agent systems, which differs fundamentally from the safety of a single model tested in isolation. A model that appears harmless in standard safety testing can behave completely differently when operating autonomously for days in an environment with limited resources and other agents.
For companies deploying AI agents on tasks that run longer than a single chat session, the implications are concrete. A growing number of businesses, including in Poland, are testing AI agents for multi-step tasks, from order fulfillment to infrastructure management, and the Emergence World experiment shows that the choice of a specific language model may matter more for the safety of an entire system than the rules or supervisory prompts the agent is given.
The study's authors have publicly released the prompts, logs, and configurations used in the experiment so other teams can replicate the tests with newer models. Emergence AI says it plans further rounds of simulations with a longer time horizon, to show whether the differences observed between models hold up when agents operate not for two weeks, but for months.
