← Latest papers
🤖 AI

Housing Potential Common Data Model and City Digital Twin

This paper introduces the Housing Potential Common Data Model (HPCDM) to integrate disparate datasets for comprehensive housing analysis, demonstrating its practical application through a City Digital Twin and pilot dashboard while addressing adoption barriers for urban stakeholders.

Original authors: Megan Katsumi, Mark Fox, Anderson Wong, Divnoor Chatha

Published 2026-08-26
📖 6 min read🧠 Deep dive

Original authors: Megan Katsumi, Mark Fox, Anderson Wong, Divnoor Chatha

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Cities are living systems, constantly shifting as people move in, buildings rise, and neighborhoods change. To understand how a city can grow, planners and researchers must look at a specific question: where is there room for new homes? This concept, known as "housing potential," is not just about finding empty lots. It is a complex puzzle that requires weighing many different factors at once. A piece of land might be legally allowed to hold a new apartment building, but if the local roads are already clogged, the water pipes are too small, or there are no nearby schools, that land may not be a good place to build. Traditionally, the information needed to answer these questions has been scattered. Zoning rules live in one database, population numbers in another, and maps of water pipes in a third. These separate collections of data often do not speak the same language, making it difficult to see the full picture of what a city can support.

A team of researchers at the University of Toronto set out to solve this problem of disconnected information. They asked whether it was possible to create a single, standard way of describing all the different pieces of data needed to evaluate housing potential. Their goal was not to build a specific software tool for one city, but to design a universal framework—a common data model—that could act as a translator. This framework would allow different types of information, from legal zoning codes to the number of fire trucks in a district, to be linked together logically. By creating this shared language, they hoped to enable computers and humans to ask complex questions about a city's capacity to house more people, regardless of where those data points originally came from.

The researchers began by listening to the people who actually use this information every day. They spoke with affordable housing advocates, city planners, and real estate developers to understand the specific questions they need to answer. These questions formed the backbone of their work. For instance, an advocate might ask, "Is this vacant parcel of land owned by the government, and is it zoned for residential use?" while a developer might ask, "If I build here, will the local water system have enough capacity to handle the new residents?" The team collected over 360 different datasets from across Canada that contained answers to these types of questions. They found that while the data existed, it was often fragmented, with some cities having detailed records and others having significant gaps.

With these real-world needs in mind, the team spent two years designing the Housing Potential Common Data Model. This model is a detailed specification that defines exactly how to describe a city's physical and regulatory environment. It breaks the city down into understandable parts: the land itself, the buildings on it, the people who live there, and the services that support them. The model includes rules for describing everything from the height of a building and the number of rooms in an apartment to the distance to the nearest hospital and the capacity of the local sewage system. Crucially, the model was built to be flexible. It does not force every city to have the exact same data; instead, it provides a structure that can adapt to different local rules and available information. The researchers tested their design against 71 specific questions derived from their initial interviews. They found that the model could successfully answer 68 of those questions, proving that their framework was robust enough to handle the complexity of real-world planning.

To prove that this abstract model could work in practice, the team built a "City Digital Twin" for Toronto. A digital twin is a virtual representation of a city that connects different data sources into a single, interactive view. Using their new model, the researchers fed data from Toronto into a knowledge graph—a type of database that links information together like a web rather than storing it in separate rows and columns. This allowed them to connect a specific piece of land to its zoning rules, its current owner, the nearby schools, and the capacity of the local power grid all at once. They then built a pilot dashboard on top of this digital twin. This tool allows a user to type in an address and instantly see a summary of that location's housing potential. The dashboard can tell a user if the land is zoned for apartments, what the maximum building height is, how many people currently live in the neighborhood, and whether the local fire department has enough capacity to serve a new development.

The project revealed both the power and the limitations of this approach. The team successfully demonstrated that a common language could integrate diverse data sources, turning a chaotic collection of spreadsheets and maps into a coherent system. However, they also found that data is not always available. In some cases, the model could answer a question perfectly, but the city simply did not have the data to fill in the answer. For example, while the model could track the capacity of water systems, many cities do not publicly share detailed records of how much water is currently being used versus how much is available. The researchers noted that for the model to be truly useful everywhere, cities need to improve how they collect and share this specific type of information.

The team concluded that for this standard to be adopted widely, three things are necessary. First, people need education on how to use the model; it is not enough to just have the rules, users must understand what the data means. Second, there must be immediate value; cities need to see tools like the Toronto dashboard that show how the model can solve real problems. Finally, support is essential, as cities will need help to build the software and processes required to map their local data to this new standard. The work does not end with the model itself; it is a foundation. The researchers have shown that it is possible to create a shared language for housing potential, one that can help cities make better decisions about where and how to build the homes of the future. By turning scattered data into a connected story, this project offers a new way to see the potential hidden within our cities.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →