← Latest papers
💰 quantitative finance

Counterfactual Analysis via Large Language Models

This paper demonstrates that prompt-engineered GPT-3.5 models can effectively perform counterfactual analysis in online lending to predict returns on investment under hypothetical interest rate scenarios, achieving performance comparable to traditional gradient-boosted regression while exhibiting logical causal reasoning.

Original authors: Zonghao Yang

Published 2026-08-07
📖 4 min read☕ Coffee break read

Original authors: Zonghao Yang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a time traveler with a magical "What If?" machine. In the world of science, this is called counterfactual analysis. It's the art of predicting what would have happened if you had made a different choice yesterday. Did that student fail because the class was too big, or would they have aced it in a smaller room? Did that email get ignored because the subject line was boring, or would a personalized one have saved the day? Scientists love this stuff because it helps us understand cause and effect, not just guess. Usually, to answer these questions, researchers need to run expensive experiments or use complex math models that look at past data to guess the future. But recently, a new kind of digital brain has arrived: the Large Language Model (LLM). Think of these as super-smart robots that have read almost everything on the internet. They are great at writing stories and chatting, but can they also play the role of a time traveler to figure out "what if" scenarios? That is the big question this paper asks.

The author of this study decided to test these digital brains in the high-stakes world of online lending. Imagine a bank that lends money to people. They have to decide what interest rate to charge. If they charge too much, the borrower might get overwhelmed and stop paying. If they charge too little, the bank loses money. The bank knows what happened with the actual interest rate they chose, but they can never know for sure what would have happened if they had picked a different one. This missing piece of the puzzle is the "counterfactual." The researchers wanted to see if an AI, specifically a model called GPT-3.5, could look at a loan application and say, "If we had charged 10% instead of 12%, the borrower would have paid us back faster."

To make this work, the team didn't just ask the AI a simple question. They realized that talking to a robot is a lot like talking to a human; you have to ask the right way. They used a technique called prompt engineering, which is basically the art of giving the AI the perfect instructions. They tried asking the AI to pretend to be the borrower, then they tried asking it to pretend to be a credit expert. They even made the AI simulate a panel of four experts arguing with each other to reach a consensus, a method called tree-of-thought. They also let the AI peek at the predictions made by traditional computer algorithms to see if it could improve upon them.

The results were surprisingly promising. When the researchers first asked the AI to guess loan outcomes without any special instructions, it was pretty bad, getting the math right only about 2% of the time. But once they used the "expert panel" trick and gave it the right context, its performance jumped up to nearly 3%. While that might sound small, it is actually a huge deal in this field, bringing the AI almost as close to the accuracy of the best traditional math models. More importantly, the AI didn't just spit out a number; it explained its reasoning. It looked at a borrower's history, considered the interest rate, and logically argued why a lower rate might help them pay off the loan sooner.

The study suggests that these large language models are not just chatbots; they are becoming capable tools for figuring out "what if" scenarios. The AI showed it could spot when a traditional computer model might be too optimistic and adjust its prediction based on its own "knowledge" of how people behave. For instance, when the traditional model predicted a borrower would pay off a loan in 8 months with a lower rate, the AI thought, "That seems too fast given their credit history," and adjusted the prediction to 22 months. This shows the AI is doing more than just copying the math; it's actually thinking through the logic.

However, the author is careful not to call this a perfect solution. They note that while the AI is getting better, it's still being tested in a simulated environment. They found that the AI's predictions were logically consistent with the rules of the real world (like, if you lower the interest rate, people are less likely to default), but they didn't claim the AI has solved the mystery of human behavior entirely. Instead, they suggest that these models are a powerful new tool that can help lenders and scientists make better decisions by simulating different futures. It's like having a very smart, very well-read consultant who can help you visualize the consequences of your choices before you actually make them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →