← Latest papers
🧬 biology

Cortex and subcortex play distinct roles over learning when cortical memory is limited

This paper proposes a theoretical framework demonstrating that when cortical memory is limited, optimal learning emerges from a functional dissociation where the cortex captures the environment's general structure while subcortical circuits specialize in reward-based learning, a hypothesis that can be tested with experimental data.

Original authors: Matthew Farrell, Taro Toyoizumi

Published 2026-06-02
📖 5 min read🧠 Deep dive

Original authors: Matthew Farrell, Taro Toyoizumi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine your brain is a high-tech office with two distinct departments working together to help you make decisions: the Cortex (the "Executive Office") and the Subcortex (the "Fast-Action Crew").

This paper explores a simple but powerful idea: The Executive Office is incredibly smart but has very limited desk space (memory), while the Fast-Action Crew is less flexible but has endless storage.

Here is how the paper breaks down their relationship using a simple game.

The Game: A Tree of Choices

The researchers set up a mental game that looks like a tree.

  • You start at the trunk (the root).
  • You have to make a series of choices to move down the branches.
  • At the very end of the branches (the leaves), there might be a reward (like a cookie).
  • The problem? The tree is huge, and the "Executive Office" (the Cortex) can only remember a few specific paths at a time. It can't memorize every single branch.

The Two Departments

  1. The Fast-Action Crew (Model-Free/Subcortex): This team learns by trial and error. "If I go left, I get a cookie. If I go right, I get nothing." They are great at repeating what works, but they are slow to learn if the rules change. They need to eat the cookie many times to be sure it's there.
  2. The Executive Office (Model-Based/Cortex): This team tries to build a map of the whole tree. They don't just wait for cookies; they try to understand how the tree works. "If I go left, there's a 70% chance I'll hit a dead end, but a 30% chance I'll find a path to the cookie." This is very powerful, but because their "desk space" (memory) is limited, they can't draw the whole map. They have to choose which parts of the tree to draw.

The Big Discovery: How to Use Limited Desk Space

The paper asks: If the Executive Office can only remember a few branches, which ones should it draw?

The researchers tested two different strategies for filling that limited desk space:

Strategy A: "The Cookie Hunter" (MAXREWARD)

The Executive Office only draws the branches that lead directly to the cookie right now.

  • Pros: It's very efficient at getting the current reward.
  • Cons: If the cookie suddenly moves to a different part of the tree, this team is stuck. They have to erase their old map and start drawing a new one from scratch. They are slow to adapt.

Strategy B: "The Explorer" (MAXREACH)

The Executive Office ignores where the cookie is for a moment. Instead, it draws the branches closest to the trunk (the start of the tree).

  • Pros: It builds a general understanding of the "structure" of the tree. It knows how to get from the start to any part of the tree.
  • Cons: It might waste time drawing paths that don't currently have cookies.
  • The Magic: When the cookie moves to a new spot, the Explorer team is ready. Because they already drew the main paths near the start, they can instantly figure out the new route to the cookie without starting over.

The Verdict: When to Use Which?

The paper found that Strategy B (The Explorer) is often better in a changing world.

If the environment is stable (the cookie stays in the same spot), the "Cookie Hunter" wins. But if the cookie moves around often (which happens a lot in real life), the "Explorer" wins because it focuses on learning the general structure of the environment rather than just the specific reward.

What This Means for the Brain

The authors suggest this explains why our brain is built the way it is:

  • The Cortex (Executive Office) is specialized for learning the general structure of the world. It builds the map. It doesn't need a cookie on every branch to keep drawing the map; it just needs to know how the tree grows.
  • The Subcortex (Fast-Action Crew) is specialized for reward-based learning. It focuses on the specific paths that lead to the cookie.

This separation allows the brain to be efficient. The Cortex doesn't waste its limited memory trying to memorize every single reward location. Instead, it builds a flexible map, while the Subcortex handles the specific "where is the treat?" details.

How to Test This

The paper also suggests a way to test this in real people. If you watch someone playing this tree game:

  • If they are a "Cookie Hunter," they will struggle immediately when the reward moves.
  • If they are an "Explorer," they will adapt quickly to the new reward location because they were already paying attention to the main paths near the start of the game.

In short: To survive in a changing world, it's better to learn the map of the territory than to just memorize where the treasure is today.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →