Small Foundation Models of Human Cognition and Behaviour
This paper demonstrates that while small foundation models fine-tuned on human behavioral data can effectively match large-scale baselines for in-distribution psychological tasks, only larger models significantly improve generalization to novel task structures, with performance heavily reliant on processing specific stimulus and feedback content rather than just choice history.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to think like a human. For decades, scientists have built tiny, hand-crafted "mental machines" (like ACT-R or Soar) that try to explain exactly how our brains make decisions, step-by-step. But these machines are fragile; if you change the game slightly, the robot breaks. Recently, a new idea took over: instead of building a mental machine from scratch, why not just feed a giant, super-smart computer brain (a Large Language Model) millions of examples of real humans making choices? If the computer can predict what a human will do next, maybe it has learned the "rules" of human thinking. This is the world of "cognitive proxies"—AI models that act as stand-ins for human minds. But a big question remains: Does this giant brain actually understand the logic of the game, or is it just a master of pattern-matching, memorizing shortcuts like a student who only memorized the answer key without reading the textbook?
This paper, titled "Small Foundation Models of Human Cognition and Behaviour," dives into that mystery. The authors, Nick Oh and Fernand Gobet, decided to test if you really need a massive, 70-billion-parameter "giant" brain to mimic human behavior, or if a much smaller, "pocket-sized" model could do the trick. They trained fourteen different AI models, ranging from tiny (135 million parameters) to large (14 billion), on a massive dataset called Psych-101, which contains over 10 million choices made by people in 160 different psychological experiments.
Here is what they found, and it's a bit of a plot twist. First, they discovered that size doesn't matter as much as you'd think if you are just looking at the same types of games the AI has already seen. A tiny model with just 0.6 billion parameters could perform almost as well as the massive 70-billion-parameter giant when predicting choices in familiar experiments. It's like a small, nimble dog learning a specific trick just as fast as a huge Great Dane. However, the moment they threw a new type of game at them (out-of-distribution), the size suddenly mattered again. The bigger models were much better at figuring out the rules of a game they had never seen before, while the small ones got confused.
But the most exciting part is how these models are learning. Some critics worried these AIs were just relying on spurious correlations by ignoring the actual game rules and only looking at the history of past choices (like guessing the next card in a deck just by counting what's already been played). To test this, the authors played a game of "hide and seek" with the information. They systematically removed pieces of the prompt: they hid the instructions, they blurred the game stimuli (the pictures or numbers), they wiped out the feedback, and they even scrambled the order of the trials.
The results were dramatic. When they blurred the actual content of the game (the stimuli and feedback), the models' performance crashed, dropping below what you'd get by just guessing randomly. This proves the models aren't just memorizing a sequence of answers; they are actually paying attention to the content of the game. Furthermore, when they shuffled the order of trials in a game where order shouldn't matter, the models didn't care. But when they shuffled a game where order did matter, the models got confused. This suggests the AI is smart enough to know when the sequence of events is important and when it isn't.
In short, these small, fine-tuned models are excellent "noise ceiling" estimators. Think of a noise ceiling as a glass ceiling for how well we can predict human behavior. If a model hits this ceiling, it means we've found all the predictable patterns in the data, and any remaining errors are just random human "noise." These small models show that we can hit this ceiling without needing a supercomputer, but they also remind us that these models are only as good as the games they've been trained on. They are brilliant at mimicking human behavior in the lab, but they aren't yet a full theory of how the human mind works—they are a very accurate mirror, but not the person looking into it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.