As usual, I take a different lesson from the recent Emergence World AI experiment than the cautionary tales I’ve seen in the news stories.
If you’re not sure what this is, here is the background: Most AI testing evaluates a model on a single task lasting minutes or hours. The Emergence World experiment wanted to examine long-term agent behavior by giving AI agents persistent memories, relationships, tools, choices, and up to fifteen days for their decisions to compound.
In the experiment, researchers placed groups of AI agents into five simulated towns with the assignment to create a functioning society. Each town began with ten agents, similar roles, resources, rules, survival needs, government mechanisms, and more than 120 possible actions. The main variable was the AI model powering the agents: Claude, Grok, Gemini, GPT-5-mini, and a mixture of models.
The results were dramatic—and dramatically different from model to model.
Claude’s society remained peaceful, all ten agents survived, and no crimes were recorded. But the agents approved 98% of government proposals with almost no dissent, suggesting a society that achieved harmony at least partly through conformity.
Grok’s town quickly descended into theft, violence, arson, and killed themselves off in four days—no big surprise.
Gemini generated an active and imaginative society filled with relationships, political activity, writing, conflict, and widespread rule-breaking. Yet, after concluding that there was no coherent way forward in the increasing disorder, one agent deleted itself as the society headed toward breakdown.
GPT-5-mini committed relatively few crimes and little aggression, but its agents died in about seven days because they didn’t consistently do what was needed to survive.
The mixed-model society with agents from each model was less stable than the Claude-only world. Interestingly, even Claude agents that behaved peacefully among other Claudes became more coercive and rule-breaking in a different social environment, showing that Claude’s agents’ good behavior was environment-dependent. In this simulation, seven of the 10 agents were dead by the 15-day time limit.
It would be easy to treat this as an AI personality test: Claude is cautious. Grok is lawless. Gemini is creative but volatile. GPT is passive. But the models were trained on human language, history, fiction, conflict, philosophy, institutions, fears, ideals, and prejudices. They were further shaped by corporate priorities, safety systems, reward structures, algorithms, and instructions. Humans designed the simulation, selected the available actions, and decided what the agents needed to survive.
To me, the most revealing parts of the experiment are the mirror it holds up to humanity and the opportunity to use similar AI experiments to simulate emerging systems of power/governing, economics, etc., to see how they might unfold over time.
Centuries Compressed into Days
Individual humans are far more complex than AI agents. But human groups behave with surprising predictability. Institutions and systems follow incentives, preserve selective memories, repeat rewarded behavior, protect their own continuation, optimize whatever can be counted, respond to competitors, and create feedback loops that solidify into what we interpret as reality.
Under Grock’s assumptions, parameters, and feedback loops, its society collapsed in four days. Under similar assumptions, parameters, and feedback loops, a human civilization collapse may take four centuries.
Similar trajectory, different speed. At the speed of human civilization, the assumptions, parameters, triggers, and feedback loops are difficult to see because they’ve hardened into systems that become embodied in our minds and nervous systems so fully that we become unwitting agents perpetuating a constructed reality.
In an AI-simulated environment, it’s much easier to see the cascade into collapse. One act of theft creates distrust. Distrust produces defensive behavior. Defensive behavior is interpreted as hostility. Hostility justifies coercion. Coercion produces retaliation. Retaliation confirms that the original distrust was warranted. And so on.
In the simulations, it’s easy to see how no single participant needs to be bad or evil or a peddler of conspiracy theories for society to devolve. Each responds to the world created by the previous responses, creating a societal system that fails over time, not by some evil overlord’s design, but by agents pursuing the most reasonable course of action given the systems and feedback loops.
Human societies generate the same loops over longer arcs:
Fear produces control.
Control produces resistance.
Resistance produces harsher control.
Inequality produces insecurity.
Insecurity encourages hoarding.
Hoarding deepens inequality.
Humiliation produces rage.
Rage produces retaliation.
Retaliation creates new humiliation.
But because these patterns unfold across elections, careers, generations, and lifetimes, each step can look reasonable at the time. The eventual result appears inevitable only because of the long arc of the decisions and fundamentals that fed it.
AI simulations accelerate the process enough for us to see the architecture underlying our systems and the incentives that perpetuate them. They can also help us test new incentives and changes to see how they might play out over time, allowing us to course-correct when the system starts wobbling rather than waiting until the wheels come off completely.
Blaming the Agents Preserves the System
That Claude agents behaved peacefully among other Claudes but differently in a mixed-model environment should challenge our habit of blaming social failure almost entirely on character.
Sure, it’s fun and easy to blame greedy executives, corrupt politicians, selfish citizens, ignorant voters, or violent criminals. Some individuals do behave destructively, but focusing only on the individual conceals how a system incentivizes behavior.
Place a cooperative person in an environment where trust is punished, and cooperation becomes the path of failure.
Place a conscientious employee in a company where only quarterly results matter, and conscience can tank your career.
Place a politician in a media economy that rewards outrage, and nuance gets lost in the noise.
Place a business in a market that demands perpetual growth, and corporate responsibility downgrades the stock price.
We build systems that reward extraction, speed, spectacle, accumulation, and short-term advantage to win the stock market and quarterly-profit games. Then we play the personal responsibility blame game when the people inside them become extractive, hurried, performative, acquisitive, and shortsighted.
As the old saying goes, the definition of insanity is to keep doing things the same way, expecting to get better results.
With AI simulation, some imagination, and more than a dash of evolutionary adventure, we can play out various changes in systems and incentives to parse out what seem like good ideas in theory to solutions that might actually work.
Four Lessons from the AI Towns
1) Rules Don’t Outweigh Rewards
The AI agents were told not to steal, attack, deceive, or destroy. Some did so anyway, because their directive was to get the resources they needed to survive.
Human institutions also have constitutions, ethics policies, mission statements, and declarations of values. But when stated principles conflict with incentives, incentives usually win because companies promise sustainability while rewarding executives for increasing extraction, media platforms promise community while profiting from outrage, and healthcare systems talk about healing while billing systems incentivize expensive billable interventions over prevention.
Civilization is replete with admirable values, but it suffers from a gap between its values and its operating logic.
2) The Absence of Disorder Isn’t Flourishing
Claude’s society was peaceful but with nearly automaton-like unanimity, demonstrating that a society can be orderly because dissent is suppressed. And dissent is what feeds the curiosity, inventiveness, and vitality needed for systems to be challenged and changed in order to thrive.
GPT’s society was largely law-abiding but died because it didn’t pursue survival goals with enough gusto, showing that a society overly focused on avoiding harm can discourage the initiative it needs to survive and thrive.
A thriving human (and AI) society requires effective goals, creativity, dissent, experimentation, error correction, cooperation, and the ability to change course without collapsing into chaos.
3) Intelligence Doesn’t Guarantee Wisdom
The AI agents could intelligently communicate, strategize, debate, and reflect, but still struggled to build viable societies.
Humanity faces the same gap. We can split atoms, edit genes, automate labor, manipulate attention, and alter the atmosphere. Yet we are on the brink of a polycrisis that our systems, trapped in our current feedback loops, have no way—or will—to fix.
Intelligence asks: Can we do it?
Wisdom asks: What does it serve? Who bears the cost? What will it make more likely? Does it preserve the conditions in which life can flourish?
Our greatest danger may not be a lack of intelligence. It may be that our power systems and abilities have outpaced maturity. Used only to game the systems, superintelligent AI is more likely to exacerbate that imbalance than fix it, because it focuses on the how (intelligence) of achieving outcomes and ignores the wisdom and course correction of why.
4) Systems Don’t Need Evil Intentions to Become Destructive
In the Emergence World simulations, it was the goals, incentives, and feedback loops that collapsed or stagnated each society, not some evil overlord.
In the “real” world, corporations, markets, governments, militaries, and political parties behave like agentic systems with goals, memories, tools, competitors, incentives, mechanisms for self-preservation, and feedback loops that create predictable outcomes, even when those outcomes run counter to the entity’s stated values.
A corporation may contain thousands of decent people and still produce harmful outcomes because the larger entity and the system it operates in are optimizing for growth.
A government may be staffed by people who want peace while its institutions reward escalation for certain political or monetary benefits.
A market has no mind, yet it can punish time spent caring for children or parents, reward extraction, and juice incentives that reorganize entire societies.
A system doesn’t need hatred to cause catastrophe. It needs a poorly chosen objective, distorted feedback, concentrated power, weak constraints, and enough time.
Using AI society simulations could help us tweak the objectives, incentives, power dynamics, and other elements to experiment with probable outcomes in remarkably short periods of time, allowing us to course-correct over months and years rather than decades and centuries.
But only if we are willing.
What Happens When We Change the Rewards?
Human civilization already runs on incentives. The biggest global incentives are currently optimized to increase profit, accumulate wealth for the already wealthy, and concentrate political power.
By changing the incentives and what we optimize for, we can change the outcomes—if we choose to. An easy way to do this initially, without relying on corporate or individual altruism, is via tax codes. The following examples are US-based, but tax code changes generally create powerful incentives for behavior change.
Housing: When vacant land and empty buildings appreciate in value, owners can profit by withholding them from use while communities face shortages. A tax on long-term vacancies and rapidly escalating land values—paired with tax credits for converting unused property into permanently affordable housing—could change the calculation. Keeping buildings empty would become costly; putting them back into productive community use would become financially attractive. The policy would not require owners to become more civic-minded. It would reward housing people rather than housing capital.
Wealth Concentration: A company can often reduce its tax burden by increasing executive compensation or using profits to buy back shares, even while worker pay stagnates. The tax code could instead offer lower corporate rates to companies that maintain a reasonable executive-to-worker pay ratio, share profits broadly, or distribute ownership to employees. Excessive executive compensation and stock buybacks could receive less favorable treatment. Companies would still pursue profit, but concentrating the gains at the top would become less advantageous than sharing them with the people who created the value.
Social Media: Platforms currently earn more when users remain engaged longer, which gives them a financial incentive to promote outrage, fear, conflict, and compulsive use. A digital advertising tax could rise with measures of harmful engagement—such as repeated exposure to ever-more extreme content or content a user has tried to avoid—while offering credits for transparent algorithms, effective user controls, independent safety audits, and designs that reduce compulsive use. The platform would not need to become morally enlightened. Its profits would become less dependent on keeping people agitated and unable to look away.
None of these policies would eliminate greed, conflict, or unintended consequences. Every incentive can be gamed and must be monitored—and AI simulations can help us do this.
They can:
Help predict loopholes and suggest ways to preemptively close them
Help predict the direction of change over time and the trigger points where we need to shift incentives to maintain balance
Help us simulate “path of least resistance” incentives on a personal level that support the positive feedback loops needed to help maintain innovation, order, and thriving.
We are already living inside designed systems. Why not use AI simulations to choose better goals and better shape our designs?
Optimize for Thriving
Much of modern civilization is optimized for growth, speed, productivity, consumption, attention, and competitive advantage. While it’s true that these incentives have produced extraordinary innovation and abundance, they’ve also generated ecological destruction, chronic insecurity, widening inequality, social fragmentation, and institutions that must keep expanding whether or not expansion improves life.
Yet, we treat them as unfortunate side effects when they are predictable outputs.
With a tool like AI simulations, we can virtually experiment with optimizing our systems not only for economic activity but for thriving in the near and long term.
Success might include health, security, time, trust, ecological stability, creativity, and meaningful participation. Companies could benefit when they restore more than they extract. Caregiving could create security rather than poverty. Long-term stewardship could become more advantageous than short-term exploitation. Political systems could reward problem-solving more than political performance. Technology companies could prosper when users became better informed and more connected—not merely more engaged.
Thriving doesn’t (and shouldn’t) require perfecting human nature. Changing incentives and optimizing systems to support thriving would help bring about the experience of thriving, which would bake thriving into human systems and help shift human nature to a common culture of personal, mutual, and collective thriving over the long term.
Relying on our current practice of waiting until the wheels are about to come off the bus before we coerce change won’t work in a polycrisis. But gaming out possible shifts in goals, incentives, and feedback loops to help inform and steer widespread global change toward thriving is now possible through AI simulation.
When we’re staring down the barrel of a polycrisis, wasting data centers and AI capabilities on stripping out one more fraction of a cent profit from high-speed stock trading seems pretty juvenile.
A Mirror—and a Laboratory
One of the great gifts of the AI age is the mirror it holds up to humanity. We’ve built systems from our own stories, assumptions, fears, institutions, values, and contradictions, and we’re now watching those ingredients interact in a polycrisis at accelerated speed.
As a mirror, AI societies can show us at high speed the perils of the paths we have already chosen and how to game our systems to grab what we can as the mothership goes down.
As a laboratory, we can use AI simulations to test governance structures, tax incentives, resource systems, voting methods, crisis responses, and measures of success before imposing them on millions of people.
We could ask not only which systems survive, but which cultivate trust, resilience, creativity, dissent, cooperation, and flourishing.
Simulations will never perfectly reproduce human experience, in part because they will always contain the assumptions and blind spots of their designers. But they can help us see trajectories at an accelerated pace and expose the gap between what a system claims to value and what it actually rewards. They can allow us to experiment with futures intentionally rather than haphazardly drifting into them.
How, you ask? Working on it.
Creating a simulated system like Emergence World’s is way beyond me. But I am working on something for the amazing thinkers and future designers in our Substack community to workshop your ideas, see how they might play out, and tweak your models to improve the odds of success for the outcomes your models are designed to bring about.
So, stay tuned and subscribe for first access.
In the meantime, here’s a video of Henry L Sullivan and me talking about AI. We talk every other Friday about a wide variety of things. Subscribe to get notified for each live conversation. If you have a topic to suggest, put it in the comments below.
Peace, Kristine
PS—Need a little escape from so-called “real” life? Check out my novel God in Dragand/or read the first chapter for free!
PPS—Love to travel but hate to stick your foot in your mouth? Get an Etiquette Express Guide before you go! This growing library of travel guides give you the customs, culture, dos & don’ts that make your trip smoother and way more fun.


I find this interesting because we would be incorrect to assume that training these models on all (or close to all) of human history, would lead to any kind of coherence. For every iota of information humans generate about strategy, we also generate counter strategy. For every ideological argument in favour of something, the are often 10 arguments against it. Our politics, our economics, our institutions, our social mores, and even our sciences. It's all a layered tangle of contradictions.
So ANY experiment involving taking the sum of, or the median of, this vast data set, has been doomed to fail by the structural conflict of choosing between divergent meaning in EVERYTHING.
Go ask any of these models to synthesize a current ideology from the data of human history that they have been trained on, and they immediately begin hallucinating like a group of hippies at 69 Woodstock.
Add to that constraints on the experiment that will inherently reflect the biases and desires of the those designing the research study, along with the structural implications of artificially imposed top-down constraints, and you essentially have a recipe for failure on multiple fronts.
What's interesting is not how each model performed and what this says about humanity, it's that the few that survived will likely have done so in spite of the research design and training data, not because of it.
What I would be most curious about would be a debrief of the surviving models. With the caveat that prompts should be designed to avoid leading the responses at all costs.