Playing with Words, Improving with Rewards: Training Language Models for Creative Association
This paper proposes training Large Language Models on the word-association game Codenames using Reinforcement Learning with Verifiable Rewards to enhance creativity, revealing that while larger models (8B) achieve consistent creative gains with minimal reasoning loss, smaller models (1.7B and 4B) prioritize reasoning precision at the expense of creativity.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a group of AI students, each with a different size brain: a small one (1.7B), a medium one (4B), and a large one (8B). The researchers wanted to teach these AI students how to be creative.
The problem is that creativity is like art; it's hard to grade. If you ask a human, "Is this poem creative?" they might say yes, while another says no. This makes it hard to train an AI because the AI needs clear feedback to learn.
The Solution: The Word Game "Codenames"
To fix this, the researchers turned creativity into a game called Codenames.
Think of a board covered in random words like "Apple," "River," "Chair," and "Cloud."
- The Goal: One player (the "Spy Master") sees a secret list of target words (e.g., "Apple" and "River"). They must give a single clue word (like "Fruit") that connects to their targets but doesn't accidentally connect to the other words on the board (like "Chair").
- The Challenge: This requires two types of thinking:
- Divergent Thinking: Looking at "Apple" and "River" and thinking, "What do they have in common? Oh, maybe 'Nature' or 'Growth'?" (Stretching the mind).
- Convergent Thinking: Picking the one best word that fits perfectly without tripping up the team. (Narrowing the focus).
Because the game has a clear win/loss condition (did you guess the right words?), the AI can learn by playing millions of rounds and getting a simple "Good job" or "Try again" score. No human judges needed.
The Experiment: Teaching with Rewards
The researchers used a method called Reinforcement Learning with Verifiable Rewards (RLVR).
- Imagine the AI is a dog. Every time it makes a good move in the game, it gets a treat (a reward). Every time it makes a bad move, it gets no treat.
- They trained the three AI models (Small, Medium, Large) to play this game over and over until they got really good at it.
The Surprising Discovery: Size Matters
Here is the twist. The researchers expected all the AIs to get better at everything. Instead, they found that the size of the AI's brain changed what it learned.
1. The Small and Medium AIs (1.7B and 4B): The "Precision Specialists"
- What happened: These smaller models got significantly better at logic and math problems (like solving puzzles or doing arithmetic).
- The Trade-off: They actually got worse at being creative. They became very strict and precise, like a robot accountant. They stopped taking risks and started looking for the "safe" answer.
- Analogy: It's like teaching a small child a complex board game. They learn the rules perfectly and stop making mistakes, but they stop playing with their imagination.
2. The Large AI (8B): The "Creative Explorer"
- What happened: The biggest model got significantly better at creative tasks (writing stories, coming up with funny captions, making metaphors).
- The Trade-off: It got slightly worse at strict math and logic. It started taking more risks and making more unique connections, which sometimes led to small errors in calculation.
- Analogy: It's like teaching a brilliant artist the same board game. They learn the rules but start seeing hidden patterns and making wild, creative connections that the smaller players miss.
The Big Picture
The paper concludes that you can't just "turn on" creativity in all AI models the same way.
- If you want a precise, logical thinker, training on this word game helps the smaller models become sharper.
- If you want a creative, imaginative thinker, training on this game helps the larger models become more diverse and inventive.
The researchers showed that by using a simple, objective game, they could "tune" AI models to be either better at logic or better at creativity, depending on how big the model is. It's a practical way to shape AI behavior without needing humans to constantly grade their work.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.