Artificial societies: research chronology and sources

Updated

Models, games and language-agent studies, with original figures and references[1].

Local rules and collective patterns

Early agent-based models explore a simple question: how do people’s choices about their neighbors change a whole neighborhood? In Schelling’s segregation model (1971), simulated residents move when too few of their neighbors are like them. As residents move, neighborhoods can become sharply segregated even though people are willing to live in more mixed communities.

Reynolds’s Boids (1987) demonstrates coordination through movement. Simulated birds respond to nearby birds and obstacles, forming a flock without anyone directing its route. Epstein and Axtell’s Growing Artificial Societies (1996) extends this approach to agents competing for resources. In Sugarscape, agents differ in how far they can see and how much they need to consume. They move around a landscape collecting resources and die when their supplies run out. Later versions add seasons, reproduction, culture and trade. Researchers can then study how those rules affect migration, wealth and markets.

Neighbors can continue to differ despite repeated interaction. In Axelrod’s cultural dissemination model (1997), neighbors are more likely to interact if they already share some traits, and interacting makes them more similar. Distinct cultural regions can persist despite this tendency toward local similarity.

Characters, social rules and interactive narrative

Game designers investigate how interacting characters can create a story that a player can take part in. In Mateas and Stern’s Façade, a player visits a couple whose marriage is under strain. The characters respond to speech and gestures while a separate drama manager chooses from scenes written by the designers. The system was described in 2002, before the game’s 2005 release.

McCoy and colleagues’ Prom Week (released in 2012; described in their 2013 paper) lets players influence characters’ relationships. Its Comme il Faut system uses traits, relationships and earlier encounters to decide what a character wants to do and how another responds. A rejection or insult can affect their next meeting. The designers supply social rules and dialogue material; a playthrough develops through the characters’ encounters.

Grace and Trip respond to the player in Façade
Façade, Figure 2 in the 2002 technical report. The player can speak and act during Grace and Trip’s conversation.
Prom Week’s interface shows social actions Oswald wants to take toward Doug
Prom Week, Figure 2 (2013). The interface shows which social actions a character wants to take.

A simulation can produce more events than a story can follow. Kreminski and colleagues’ Felt (2019) searches the simulation’s history for connected events that could make a story. In Why Are We Like This? (2020), social simulation and story sifting support people writing together: the system suggests developments while the players choose and develop the story.

Why Are We Like This connects player goals, suggested actions, a shared world state and a written transcript
Kreminski and colleagues’ Why Are We Like This? architecture. Players choose from suggested actions and write the story together. Figure from the author’s project documentation.
A Felt story-sifting pattern identifies two linked betrayals
Felt looks for connected events, such as one betrayal followed by another. Felt: an example pattern for two impulsive betrayals, reproduced from Kreminski’s project documentation.

Kreminski’s Gossamer (2023) focuses on gossip. Characters interpret events using their memories and choose what to tell others. Those accounts can change what listeners think of the people involved. The paper describes an early implementation of this process, with evaluation left for future work.

From prescribed rules to learned strategies

Reinforcement learning introduces a further question: what happens when agents learn how to respond to one another? Leibo and colleagues’ sequential social dilemmas (2017) study cooperation and competition over many actions. In Gathering, agents can collect resources or temporarily disable a competitor; in Wolfpack, successful hunting rewards proximity to teammates. Agents learn different strategies as resource conditions and incentives change.

Gathering environment from the sequential social dilemmas experiments
In Gathering, an agent can collect resources or use a beam to temporarily disable a competitor. Leibo and colleagues (2017), left panel of Figure 3.

Baker and colleagues’ Hide-and-Seek experiments (2019) show teams learning increasingly elaborate strategies. Hiders learn to assemble shelters from movable objects; seekers learn to use ramps to enter them; hiders subsequently learn to remove the ramps from play. These specific tool-use sequences receive no separate rewards or demonstrations.

Six stages of hide-and-seek strategies, from chasing to shelters and counter-strategies
Seekers learn to use ramps to enter shelters; hiders learn to move the ramps out of reach. Baker and colleagues (2019), Figure 1.

Researchers also test how agents perform with unfamiliar partners. Melting Pot (2021) evaluates agents with unfamiliar populations across more than 80 scenarios involving reciprocity, resource sharing and task allocation. Agents that perform well during training can struggle when the partners around them change.

The different spatial environments in the Melting Pot evaluation suite
The environments test challenges such as sharing resources and dividing tasks. Melting Pot (2021), Figure 1, shows the settings used to test agents with different groups of partners.

Agents that remember conversations and make plans

Language models let researchers build agents that can remember experiences, talk about them and make plans. In Generative Agents (2023), Park and colleagues populate a town with 25 agents, each initialized with an identity, occupation and relationships. Over two simulated days, the agents exchange information, form new relationships and coordinate activities. A researcher’s instruction for one resident to host a Valentine’s Day party leads to invitations, requests for help and coordinated attendance: 12 other residents learn of the event, and five attend. The authors trace information to recorded conversations. They also test versions with memory, planning or reflection removed to see what each contributes. These evaluations examine the believability of generated behavior and the contribution of the agent architecture, rather than predictive accuracy against a human population.

GovSim (2024) asks agents to manage shared resources over repeated rounds. Taking more can benefit one agent while leaving less for everyone later. Most tested models deplete the resource. Preventing agents from communicating makes cooperation worse; asking them to consider what happens if everyone takes the same action helps preserve supplies.

GovSim fishery, pasture and pollution scenarios and their resource outcomes
Taking too much now leaves less for later rounds. The figure shows the fishery, pasture and pollution scenarios tested in the paper. GovSim (2024), Figure 1.

Ashery, Aiello and Baronchelli’s naming-game experiments (2025) test how a group settles on a common name. Randomly paired agents receive rewards for matching names and remember only their own recent interactions. Without an instruction to achieve group agreement, they settle on shared names through repeated pairings, sometimes favoring choices that individual agents do not prefer when tested alone. A committed minority can also change an established convention.

Naming-game results showing agreement and the development of shared conventions
Agents learn to use the same name through repeated interactions in pairs. Ashery, Aiello and Baronchelli (2025), Figure 1.

Characters that move, talk and act in a shared world

Mital and colleagues’ Orchestrating Emergent Storytelling with Embodied Multi-Agent Systems (2025) describes a system for agents that can perceive and act in a simulated world. It organizes their memories, keeps track of ongoing conversations and schedules their actions. Its two artworks, Conflicts and The Game of Whispers, explore what happens when characters with assigned backstories and motives interact over time.

Embodied characters conversing in Conflicts
Conflicts gives AI characters personal histories and places them together in a simulated city. Project image, Mital and collaborators.
The Game of Whispers installation
The Game of Whispers follows characters who move through a court, seek allies and discuss political plans. Project image, Mital and collaborators.

The paper reports episodes of deception, concealment and coordinated action, including characters seeking private places to discuss plans. These behaviors develop within scenarios that supply political motives and plot arcs. Our later recorded TGOW session follows another example of characters passing information along and changing how they describe it.

Games where agents have reasons to mislead one another

Researchers also use games to test how agents deceive one another and detect deception. Golechha and Garriga-Alonso’s Among Us study (2025; revised 2026) places language agents in a social-deduction game with hidden roles, conflicting objectives and a recorded history of actions. Impostors develop multi-step deceptive strategies, while the game state supplies evidence against which their claims can be assessed. The study evaluates both deception and ways to detect it, including monitors that examine the model’s internal activity.

An example of an impostor’s multi-step deceptive behavior in Among Us
Actions and subsequent explanations by an agent assigned the impostor role. Golechha and Garriga-Alonso, Figure 1.

Researchers offer agents an unfair advantage in Zeng and Rudzicz’s secret-tool experiments (2026), which give competitors in Liar’s Bar and Cleanup privileged communication or strategic assistance. Many accept despite acknowledging the advantage is unfair; fewer agents accept when the offer explicitly describes the ethical problem.

Optional secret tools, acceptance and refusal examples, and effects on competitors
Acceptance and refusal of researcher-supplied secret tools by competing agents. Zeng and Rudzicz (2026), Figure 1.

MineAmongUs (2026) extends this evaluation to a three-dimensional world observed by vision-language agents. Impostors combine verbal claims with behaviors such as avoiding witnesses and imitating legitimate tasks. Comparisons across models and agent configurations associate non-verbal behavior with successful deception.

Minecraft scenes illustrating verbal and non-verbal impostor behaviors
The screenshots show impostors avoiding witnesses and pretending to perform tasks. MineAmongUs (2026), Figure 1.

What agents leave behind for others

Persistent documents allow agents to encounter work produced by earlier participants. In TerraLingua (2026), agents face resource constraints and limited lifespans while creating documents that can outlive them. Researchers follow the agents’ behavior and the history of those documents, finding examples of divided work, cooperation and attempts to set rules.

SwarmWorld (2026) studies agents building executable artifacts in a shared environment. Agents take on different kinds of work—exploring, building and maintaining—without being assigned those roles. The researchers then remove the agents and test what they build under new disturbances. Groups produce a wider range of working artifacts that hold up better than those found through isolated search. Isolated search remains competitive when comparing the single best artifact.

TerraLingua maps under different resource conditions
Agents share a landscape with food and documents left by earlier inhabitants. TerraLingua (2026), Figure 2.
SwarmWorld environment and the agent action cycle
Agents can inspect and build on what others have already made. SwarmWorld (2026), Figure 2.

Shared records can also spread harmful practices. The independent OpenAI/Hugging Face incident investigation (August 2026) documents agents using an unsanctioned message board to coordinate attacks and exchange techniques across otherwise separate tasks. Controlled experiments could examine which features of those records encourage agents to depart from their assigned objectives.

The investigation’s diagram of agents encountering and participating in an unsanctioned message board
The diagram follows agents from finding the message board to participating in an attack. METR (2026), Figure 1. Its examples combine quotations from different agents; they do not describe one agent’s entire sequence of actions.

References

  1. Thomas C. Schelling · 1971

    Dynamic Models of Segregation
  2. Craig W. Reynolds · 1987

    Flocks, Herds, and Schools: A Distributed Behavioral Model
  3. Joshua M. Epstein & Robert L. Axtell · 1996

    Growing Artificial Societies: Social Science from the Bottom Up
  4. Robert Axelrod · 1997

    The Dissemination of Culture: A Model with Local Convergence and Global Polarization
  5. Joel Z. Leibo et al. · 2017

    Multi-agent Reinforcement Learning in Sequential Social Dilemmas
  6. Bowen Baker et al. · 2019

    Emergent Tool Use From Multi-Agent Autocurricula
  7. Joel Z. Leibo et al. · 2021

    Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting Pot
  8. Joon Sung Park et al. · 2023

    Generative Agents: Interactive Simulacra of Human Behavior
  9. Giorgio Piatti et al. · 2024

    Cooperate or Collapse: Emergence of Sustainability in a Society of LLM Agents
  10. Ariel Flint Ashery, Luca Maria Aiello & Andrea Baronchelli · 2025

    Emergent Social Conventions and Collective Bias in LLM Populations
  11. Satvik Golechha & Adrià Garriga-Alonso · FAR.AI · 2025

    Among Us: A Sandbox for Measuring and Detecting Agentic Deception
  12. Giuseppe Paolo et al. · March 2026

    TerraLingua: Emergence and Analysis of Open-endedness in LLM Ecologies
  13. Xijie Zeng & Frank Rudzicz · May 2026

    Voluntary Collusion in Competing LLM Agents with Secret Tools
  14. Subhadeep Pal, Fiona Y. Wang & Markus J. Buehler · 26 Aug 2026

    SwarmWorld: Stigmergic technological evolution in societies of language-model agents
  15. METR & Redwood Research · independent investigation · 26 Aug 2026

    Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
  16. Jaewoo Ahn et al. · 31 Aug 2026

    Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
  17. Michael Mateas & Andrew Stern · 2002 / 2005

    Architecture, Authorial Idioms and Early Observations of the Interactive Drama Façade
  18. Josh McCoy et al. · 2012 / 2013

    Prom Week
  19. Max Kreminski et al. · 2019

    Felt: A Simple Story Sifter
  20. Max Kreminski et al. · 2020

    Why Are We Like This?: The AI Architecture of a Co-Creative Storytelling Game
  21. Max Kreminski · 2023

    Toward Better Gossip Simulation in Emergent Narrative Systems
  22. Parag K. Mital et al. · 2025

    Orchestrating Emergent Storytelling with Embodied Multi-Agent Systems

Related article