← Latest papers
🤖 AI

The Silicon Society\textit{Silicon Society} Cookbook: Design Space of LLM-based Social Simulations

This paper systematically analyzes the design space of LLM-based social simulations, revealing that the choice of the base large language model is the most critical factor influencing simulation outcomes and highlighting the complex, non-trivial interactions between various design parameters.

Original authors: Aurélien Bück-Kaeffer (McGill University, Mila - Quebec Artificial Intelligence Institute, Ubisoft La Forge), Sneheel Sarangi (McGill University, Mila - Quebec Artificial Intelligence Institute), Maxi
Published 2026-05-04
📖 5 min read🧠 Deep dive

Original authors: Aurélien Bück-Kaeffer (McGill University, Mila - Quebec Artificial Intelligence Institute, Ubisoft La Forge), Sneheel Sarangi (McGill University, Mila - Quebec Artificial Intelligence Institute), Maximilian Puelma Touzel (McGill University, Université de Montréal), Reihaneh Rabbany (McGill University, Mila - Quebec Artificial Intelligence Institute), Zachary Yang (McGill University, Mila - Quebec Artificial Intelligence Institute, Ubisoft La Forge), Jean-François Godbout (Mila - Quebec Artificial Intelligence Institute, Université de Montréal)

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a director trying to film a movie about a bustling city, but instead of hiring thousands of human actors, you are using a fleet of advanced robots. Your goal is to make the robots interact, argue, gossip, and form opinions just like real people. This is what the authors of this paper call a "Silicon Society."

The paper is essentially a "behind-the-scenes" guide to figuring out which knobs and dials on your robot set actually change the story, and which ones are just decoration. The authors ran 595 different versions of these robot cities to see what happens when they tweak the settings.

Here is a breakdown of their findings using simple analogies:

1. The Most Important Choice: The "Actor"

The single biggest factor in how the movie turns out isn't the script or the camera; it's which robot model you choose to play the role.

  • The Analogy: Imagine you are casting a play. If you hire a robot that naturally sounds like a cheerful, agreeable customer service bot, your play will be boring and everyone will agree with each other. If you hire a robot trained on messy, real-world social media arguments, your play will be chaotic, with people changing their minds and fighting.
  • The Finding: The specific AI model used as the "brain" for the agents matters more than the network structure (who follows whom) or the size of the crowd. Some models naturally create more drama and opinion changes than others.

2. The "Makeup" Matters: Fine-Tuning

The authors tested what happens if they take a standard robot and "teach" it how to speak like a real human by feeding it millions of real social media posts (a dataset called BluePrint).

  • The Analogy: A standard robot speaks with a perfect, robotic accent that gives it away immediately. Giving it "social media makeup" (fine-tuning) is like teaching it slang, sarcasm, and how to be stubborn.
  • The Finding:
    • Realism: The "makeup" made the robots much harder for a computer detector to identify as fake. They sounded more human.
    • Drama: Without the makeup, the robots were too polite and boring; they rarely changed their opinions. With the makeup, they started arguing, changing their minds, and forming groups, just like real people on Twitter or Facebook.

3. The "Stage" vs. The "Script"

The team tested various ways to set up the scene:

  • Network Topology: Does it matter if the robots are connected randomly (like a party) or in a "hub-and-spoke" pattern (like a celebrity with many fans)?
    • Result: Not as much as you'd think. The type of robot (the actor) mattered way more than the layout of the room.
  • Biased News: What if they added one robot that only posts fake, biased news?
    • Result: Surprisingly, it didn't change anything. One noisy voice in a crowd of 1,000 wasn't enough to shift the whole group's opinion.
  • Homophily (Birds of a feather): What if they forced robots with similar opinions to sit next to each other?
    • Result: This did create "echo chambers" (groups where everyone agrees), but only if the robots were already the type to argue.

4. The "Surprise" Factor: Complex Interactions

The authors expected that if they turned up the "volume" on one setting, the result would just get louder (additive). Instead, they found the settings interact in weird, unpredictable ways.

  • The Analogy: Think of baking a cake. You might think adding more sugar makes it sweeter, and adding more eggs makes it fluffier. But in this "Silicon Society" kitchen, adding sugar to this specific type of flour might make the cake collapse, while adding it to that flour makes it perfect.
  • The Finding: You can't predict the outcome just by looking at one setting in isolation. For example, letting the robots see their own survey answers made them more conformist, but only if they were using a specific type of robot model. On other models, it didn't change anything.

5. The "Detective" Test

To see if their simulations were realistic, they used a "detective" (a BERT classifier) to try to spot which threads were written by real humans and which were written by their robots.

  • The Result: Robots that were not fine-tuned were caught almost 100% of the time because they sounded too perfect. Robots that were fine-tuned on social media data were much harder to catch, though the "detective" was still very good at spotting them.

The Bottom Line

The paper concludes that building a realistic simulation of human society isn't about finding the perfect network map or the biggest crowd. It's about choosing the right "brain" (the base AI model) and teaching it to speak like a real, messy human (fine-tuning).

If you get the brain and the voice right, the rest of the simulation tends to fall into place. If you get them wrong, no amount of complex networking will make the robots act like real people.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →