A multi-view contrastive learning framework for spatial embeddings in risk modelling
This paper proposes a novel multi-view contrastive learning framework that generates low-dimensional spatial embeddings from satellite imagery and OpenStreetMap data to enhance the predictive accuracy and interpretability of risk models in insurance and real estate applications.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine trying to understand the character of a place just by looking at a dot on a map. For decades, insurance companies and real estate analysts have done exactly this, treating a location simply as a pair of numbers: latitude and longitude. They hoped these coordinates would tell them everything they needed to know about risk or property value. But a dot on a map is silent; it does not know if it sits in a bustling city center, a quiet rural valley, or a flood-prone riverbank. It cannot see the nearby shops, the type of buildings, or the texture of the landscape. When analysts rely only on these raw numbers, their models often miss the subtle, real-world patterns that define how people live and how nature behaves. They end up drawing artificial, boxy lines around risks that flow smoothly across the land, much like trying to describe a winding river by drawing a series of straight, rectangular steps.
To solve this, researchers have begun using a technique called "embedding." Think of an embedding as a digital fingerprint for a place. Instead of just a pair of numbers, a location is represented by a short list of values that capture its essence: what it looks like, what surrounds it, and how it fits into the wider world. This paper introduces a new way to create these fingerprints by teaching a computer to look at a place from multiple angles at once. The researchers combined satellite images, which show the physical shape of the land, with digital maps of human activity, such as the location of schools, shops, and parks. By training a computer to match these visual and contextual clues with the exact geographic coordinates, they created a system that can generate a rich, detailed description of any location using only its latitude and longitude.
The team, led by Freek Holvoet and colleagues, built a massive dataset of nearly 96,000 locations across Europe. For each spot, they gathered a satellite photo and a detailed count of nearby amenities from a global map project. They then trained a computer model to learn the connection between the visual and contextual data and the specific coordinates of that spot. The model learned to say, "This pattern of buildings and green spaces belongs to these specific coordinates." Once the model learned this relationship, it could forget the images and maps entirely. From that point on, it could take just a pair of coordinates for any new location and instantly generate a complex digital fingerprint that contained all the learned information about that place's surroundings.
The researchers tested this system in two very different real-world scenarios. First, they looked at house prices in France. They compared standard models that used only raw coordinates against models that used the new digital fingerprints. The results were clear: the models using the fingerprints predicted prices much more accurately. More importantly, the new models could "see" the landscape in a way the old ones could not. When the researchers visualized the predictions, the old models produced blocky, unnatural maps where prices jumped abruptly at invisible lines. The new models produced smooth, realistic maps that correctly identified high-value areas around Paris and in the Alps, while recognizing the lower values in rural regions. They even showed that the system could make sensible predictions for areas where it had never seen a single house sale, simply because it understood the geographic character of those regions.
In a second test, the team applied the method to flood insurance in Belgium. They analyzed thousands of flood claims to see if the digital fingerprints could help sort areas by risk. Again, the models using the fingerprints outperformed those using raw coordinates. They were better at distinguishing between high-risk and low-risk neighborhoods, capturing the true shape of the danger zones along river valleys rather than forcing them into artificial boxes. The study also showed that these digital fingerprints could be generated in less than a millisecond, making them fast enough to use in everyday insurance pricing.
This work suggests that the future of risk analysis lies in giving computers a richer sense of place. By teaching machines to understand that a location is not just a point in space, but a place with a specific visual and social context, analysts can build models that are more accurate and more fair. The researchers demonstrated that these tools can be trained on standard computers and adapted to different regions, offering a practical way to improve how we price insurance and manage risk without needing to access sensitive or proprietary data for every single calculation. The result is a system that sees the world more like a human does, understanding that where you are matters, but what is around you matters just as much.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.