Reward function compression facilitates goal-dependent reinforcement learning
This paper proposes and experimentally validates that humans enhance reinforcement learning efficiency by initially relying on working memory to compress complex, abstract goals into simplified, stable reward functions that transfer to long-term memory, thereby freeing cognitive resources for automatic evaluation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
The Big Idea: Learning with a "Mental Backpack"
Imagine your brain has a backpack called Working Memory. This backpack is great for holding things you need right now, like a phone number you just heard or a rule you just read. But it has a strict limit: it can only hold about 3 to 4 items at once. If you try to stuff too much in, things fall out, or you get so overwhelmed you can't think clearly.
The paper argues that when humans try to learn new things based on abstract goals (like "win this game" or "collect the purple card"), we initially have to carry the rules for those goals inside this small backpack. This is cognitively expensive. It's like trying to solve a math problem while simultaneously juggling three heavy bowling balls. You can do it, but you'll be slower and make more mistakes.
The authors propose a clever solution: Reward Function Compression.
Think of this as folding a map.
- The Uncompressed Map: Imagine a map of a city where every single street, building, and tree is drawn in high detail. It's accurate, but it's huge and hard to carry in your backpack.
- The Compressed Map: After you've visited the city a few times, you realize you don't need every detail. You can fold the map down to just the main highways and the location of the park. It's tiny, fits easily in your pocket, and you can pull it out instantly without thinking.
The paper suggests that our brains do this with goals. At first, we remember every specific detail of what a "win" looks like. But with practice, we compress that information into a simple, automatic rule (e.g., "All purple things are good"). Once this rule is "compressed" and moved to Long-Term Memory (a giant warehouse outside the backpack), we don't have to carry it in our backpack anymore. This frees up mental space to actually learn the task faster.
The Experiments: How They Tested This
The researchers ran six experiments using a computer game to prove this idea. Here is how they did it:
1. The Game Setup
Participants had to learn which key to press for which picture.
- The Easy Way (Points): They got standard feedback like "+1" or "+0". This is like getting a score.
- The Hard Way (Goals): Instead of numbers, they saw abstract fractal images labeled "Goal" (Good) or "Nongoal" (Bad). They had to figure out which picture was the "Goal" and press the right key to get it.
2. Experiment 1: The "Too Many Goals" Test
The Setup: In some rounds, there was only one "Goal" image to look for. In other rounds, there were four different "Goal" images that changed.
The Result: When there was only one goal, people learned quickly. When there were four different goals, their learning slowed down significantly.
The Takeaway: This proves that our "mental backpack" gets overloaded. Trying to keep track of four different rules at once uses up all the space needed for learning.
3. Experiment 2: The "Compressible vs. Uncompressible" Test
This was the most important experiment. The researchers kept the number of goals the same (so the backpack load was the same), but they changed the structure of the goals.
- Compressible Goals: Imagine the "Goal" images were all Red Squares, and the "Nongoal" images were all Blue Triangles. Even though there were many images, the rule was simple: "If it's Red, it's a Goal." This is easy to compress.
- Uncompressible Goals: Imagine the "Goal" images were a weird mix: a Red Square, a Blue Circle, a Green Triangle, etc. There was no simple rule. You had to memorize every single specific image.
The Result: People learned much faster in the Compressible condition. Even though they had to remember the same number of images, they could "fold the map" into a simple rule ("Red = Good"). In the Uncompressible condition, they couldn't fold the map, so they had to carry the whole heavy thing in their backpack, slowing them down.
4. The Speed Connection
The researchers also measured how fast people pressed a button to "collect" the good outcome.
- Finding: The faster someone was at recognizing the "Goal" image, the better they were at learning the game.
- Why? This suggests that when the brain processes the reward automatically (because the rule is compressed), it frees up energy for the actual learning. If you are still struggling to figure out "Is this a goal?", you have less brainpower left to learn the rules of the game.
What This Means (According to the Paper)
The paper concludes that human learning isn't just about memorizing facts. It's about efficiency.
- Flexibility has a cost: We can set any goal we want (even abstract ones), but it takes a lot of mental energy at first.
- Compression is the key: To become efficient, our brains must turn complex, high-dimensional goals into simple, low-dimensional rules.
- Automaticity: Once that rule is compressed and stored in long-term memory, we stop "thinking" about the goal and start "feeling" it automatically. This frees up our working memory to learn new things faster.
In short: We learn best when we can turn a complex list of "what to do" into a simple, automatic habit. If the rules are too messy to simplify, our brains get tired, and learning suffers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.