← Latest papers
💰 quantitative finance

Synthetic Data for Portfolios: A Throw of the Dice Will Never Abolish Chance

This paper critiques the limitations of current generative models in portfolio management, highlighting issues like the pitfalls of excessive data generation and misalignment with portfolio construction needs, while proposing a robust pipeline for synthesizing multivariate returns and introducing a "regurgitative training" method to better evaluate model identifiability and suitability for specific financial applications.

Original authors: Adil Rengim Cetingoz, Charles-Albert Lehalle

Published 2026-07-21
📖 5 min read🧠 Deep dive

Original authors: Adil Rengim Cetingoz, Charles-Albert Lehalle

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to predict the future of a chaotic, noisy world. You have a small notebook filled with observations of how things move, but you want to know what happens next. In the world of finance, this notebook contains the daily price changes of thousands of stocks. For decades, experts have tried to build "generative models"—sophisticated digital machines that learn from this notebook and then spit out endless new, fake scenarios. It's like a chef tasting a single spoonful of soup and then trying to cook a banquet that tastes exactly the same. The hope is that by feeding these fake scenarios into a computer, we can test investment strategies without risking real money. But there's a catch: if the chef only tasted a tiny spoonful, no matter how many bowls of soup they cook, the banquet will still taste wrong. This paper dives into the messy, high-stakes kitchen of financial markets to see if these digital chefs can actually cook up a meal that helps investors, or if they are just serving up a fancy illusion.

The authors, Adil Rengim Cetingoz and Charles-Albert Lehalle, start by dropping a heavy truth bomb: throwing more dice doesn't fix a bad roll. They argue that if you train a model on a small amount of real data (like 30 years of daily stock prices), generating millions of fake data points won't magically make your predictions more accurate. In fact, it might make things worse. The model learns the "flaws" of the small sample size and then confidently repeats those flaws over and over again. It's like a student who memorizes the answers to a tiny practice test; if you give them a million copies of that test, they'll get a perfect score, but they still won't know how to solve a new problem. The paper suggests that in finance, you can't just "scale up" the data like you can with images or text; the initial sample size is a hard limit that you can't cheat.

Next, the authors tackle a weird mismatch between how these models learn and how investors actually think. Most AI models are trained to get the "big picture" right—they focus on the most obvious, loud patterns in the data, like the general rise and fall of the entire market. But for building a smart investment portfolio (especially one that bets on some stocks going up and others going down), the most important clues are often the quiet, subtle, low-risk patterns that the AI tends to ignore. It's as if the AI is a student who only studies the loud, dramatic chapters of a history book and skips the quiet footnotes, yet the teacher (the investor) needs the footnotes to pass the exam. The paper shows that standard AI models are terrible at capturing these subtle, low-risk details, which are exactly what you need to build a safe portfolio.

To fix this, the authors propose a new, custom-built pipeline for generating synthetic stock data. Instead of letting the AI guess everything at once, they break the problem down into two parts. First, they separate the "loud" market movements (the factors) from the "quiet" random noise (the residuals). They use a special type of AI called a Generative Adversarial Network (GAN) to learn the complex, time-based patterns of the loud factors, but they group similar factors together so the AI doesn't get overwhelmed. For the quiet noise, they use a simpler, more flexible math trick (a mixture of Student-t distributions) to capture the weird, heavy-tailed spikes in prices that standard models miss. They tested this new system on 433 stocks from the S&P 500. The results were promising: the fake data looked surprisingly real, capturing the same strange behaviors (like volatility clustering and heavy tails) as the real market.

However, the authors don't just stop at "it looks good." They introduce a clever new way to test if the model is actually smart or just a parrot. They call it "regurgitative training." Imagine teaching a student using a textbook you wrote, then asking them to write a new textbook based on what they learned, and finally checking if that new textbook can teach the original student. If the model can't learn from its own fake data to solve the original problem, it's not a good model. Using this test, they found that many standard models fail because they can't identify the underlying rules of the market, even when those rules are hidden in the data they generated themselves.

In short, this paper suggests that while we can't just generate infinite fake data to solve our problems, we can build better, more specialized tools if we stop treating all data the same. By respecting the limits of our small real-world samples and focusing on the quiet, subtle details that matter for investing, we can create synthetic data that is actually useful for navigating the chaotic, dice-rolling world of finance. The authors suggest that the future isn't about bigger models, but smarter, more tailored ones that understand the specific quirks of the market.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →