From Forecasts to Auditable Reports: Evidence Contracts for LLM-Assisted Housing-Guarantee Risk Monitoring
This paper introduces an evidence-constrained reporting pipeline that integrates predictive modeling with structured evidence contracts and verification to transform sparse housing-guarantee risk forecasts into auditable, high-quality operational reports validated by domain professionals.
Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Detective's Dilemma: When AI Predicts the Future
Imagine you are a detective trying to solve a mystery before it even happens. You have a crystal ball (an Artificial Intelligence) that can guess what might go wrong next month, like a house falling into disrepair or a tenant running away with a deposit. But here's the catch: the crystal ball is a bit of a show-off. It gives you a simple answer like "Bad things might happen!" but refuses to show you why or where it got that idea. In the real world, especially with money and housing, you can't just trust a hunch. You need a paper trail. You need to be able to look at the evidence, check the math, and prove that the warning is real before you call the police or freeze assets. This is the world of "risk monitoring," where the goal isn't just to be right, but to be provable. If an AI makes a mistake, it shouldn't just be a wrong guess; it shouldn't be a made-up story that hides the truth. This paper dives into a tricky corner of science where computer scientists try to teach AI to stop guessing and start building a case file, ensuring that every warning comes with a stack of receipts that a human can actually read and verify.
From Crystal Balls to Case Files
In this study, researchers from South Korea tackled a very specific, high-stakes problem: predicting when a "jeonse" housing guarantee might fail. In the jeonse system, tenants pay a massive lump sum instead of monthly rent, and a public institution guarantees they get that money back if the landlord goes bust. The problem is that these disasters are rare (like finding a needle in a haystack), and the data is secret. If you ask a standard AI to write a report about a potential disaster, it might just make up a story to sound smart, or it might miss the rare disaster entirely because it's too focused on being "average."
The authors built a new system called an "Evidence Contract" pipeline to fix this. Think of it like a strict courtroom procedure for AI. Instead of letting the AI wander around and write a free-form essay, the system forces the AI to follow a rigid script based on hard facts.
Here is how the magic happens:
- The Prediction: First, a special AI model (a Temporal Fusion Transformer) looks at the data and predicts the risk for next month. But unlike normal models that just want to be "close" to the average, this one is trained to be extra careful about missing the big, scary disasters. It's like a smoke detector that is tuned to scream loudly even if there's just a tiny wisp of smoke, rather than waiting for a full-blown fire.
- The Evidence Hunt: Once the AI spots a potential risk, it doesn't just guess why. It goes back into the history books to find similar past situations. But it doesn't just look for similar numbers; it looks for similar patterns of thinking. It asks, "Did the computer care about the same things in the past as it does now?" This is done using a clever math trick called Centered Kernel Alignment, which matches the "reasoning" of the AI rather than just the raw numbers.
- The Contract: This is the most important part. Before the AI is allowed to write a single word of the final report, a "contract" is created. This contract is a typed list of facts: the predicted number, the direction of the risk (up or down), the key reasons, and the historical examples. The AI is strictly forbidden from making up new numbers or inventing new reasons. It can only take the facts in the contract and turn them into a readable story.
- The Audit: Finally, a strict checklist (an audit) runs over the AI's draft. It checks: "Did you say the number 23.71%? Good. Did you say the direction is 'Down'? Good. Did you mention the lease-price index? Good." If the AI tries to sneak in a made-up fact or change a number, the system catches it immediately.
What They Found
The researchers tested this system using data from South Korea's housing market from September 2015 to December 2025. They didn't just look at how accurate the predictions were; they looked at how well the system caught the rare, dangerous events (the "upper-tail" risks).
They found that their special "Regret-Sensitive" model was much better at spotting these rare disasters than standard models. While a standard model might miss 9 out of 10 big risks just to look good on average, their model caught 14 out of 25 of the worst-case scenarios. It did cost a tiny bit of accuracy on the average days (the error went up slightly), but the researchers argue this is a fair trade-off because missing a disaster is much worse than having a false alarm.
When they tested how well different AI models (8 of them!) could write the reports, the results were clear: Structure wins. When the AI was given the "Evidence Contract" (the list of allowed facts) and a strict template, the reports were much better. They were more accurate, less likely to lie, and easier for humans to trust. In fact, when they showed these reports to 51 real-world analysts and experts, most of them said the reports were useful and would help them make better decisions.
The Bottom Line
The paper suggests that we can't just let AI write reports on its own, especially when money and safety are on the line. The AI is great at crunching numbers, but it needs a human-like "case file" to work from. By forcing the AI to stick to a pre-approved list of evidence and checking its work with a strict audit, we can turn a black-box guess into an auditable, trustworthy report. The researchers didn't solve every problem in the world, but they showed a clear path: if you want AI to be a helpful partner in high-stakes decisions, you have to give it a contract to follow, not just a blank page to fill.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.