← Latest papers
💻 computer science

Learning Long-Term Educational Investment Policies under Residential Sorting

This paper proposes a dynamic multi-agent framework using reinforcement learning to optimize long-term public-school investment policies that account for residential sorting and housing market feedback, demonstrating superior effectiveness and equity compared to existing approaches.

Original authors: Honglei Guo, Shuo Chen, Mingjie Bi, Zeyang Sun, Xiaoxi Wang, Yuhan Zhao

Published 2026-08-10
📖 3 min read☕ Coffee break read

Original authors: Honglei Guo, Shuo Chen, Mingjie Bi, Zeyang Sun, Xiaoxi Wang, Yuhan Zhao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a giant, invisible game of musical chairs played across an entire city, but instead of music stopping, the game is driven by how much money families have and how much they want their kids to go to a specific school. This is the world of residential sorting, a concept where families naturally drift toward neighborhoods that offer the best schools, much like moths to a light. But here's the twist: when a school gets better, the neighborhood around it becomes more expensive, pushing out the very families who might need that improvement the most. This creates a tricky loop: the government tries to fix a school, the school gets better, prices go up, and the people who needed the help the most get locked out. Scientists call this the "education-housing feedback loop." It's a complex puzzle because you can't just look at the school, the house, or the family in isolation; they are all dancing together, and if you change one step, the whole choreography shifts.

This paper steps into that dance floor to see if a computer can learn the best moves. The authors built a digital simulation—a virtual city with 12 neighborhoods and 4 schools—where families, housing prices, and schools all react to each other in real-time. They taught a computer agent (using a method called Reinforcement Learning, which is like teaching a dog tricks by giving it treats for good behavior) to act as the city planner. The computer's job was to decide how to split a yearly budget of money among the schools. The goal wasn't just to make the schools "good" on average, but to make sure the money was shared fairly so that rich and poor families could both access quality education without the housing market ruining the plan.

The results of this digital experiment were quite revealing. The computer, after learning from thousands of simulated years, found a strategy that outperformed traditional ways of handing out money. While other methods (like giving every school the exact same amount or just giving money to the schools with the most students) created big gaps between rich and poor, the computer's strategy achieved a high level of access for everyone (an average score of 0.4780) while keeping the inequality between families incredibly low (a Gini coefficient of just 0.0164). The simulation suggests that by explicitly accounting for the fact that better schools make houses more expensive, the government can actually design investment policies that improve schools for everyone without accidentally pricing the poor out of the neighborhood. However, the authors are careful to note that this is a simulation, not a real-world policy yet. They found that simply building more schools didn't automatically fix the problem; in fact, having too many schools in a large area sometimes made the sorting worse. The key takeaway is that to fix education, you have to understand the house next door, too.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →