Identifying Model Quality Effects on User Engagement: A Within-Version Causal Estimator with Synthetic Data Validation
This paper introduces a novel causal estimation framework that leverages non-uniform quality improvements across different capabilities within a single LLM version to isolate the impact of model updates on user engagement, a method validated on synthetic data to significantly outperform standard approaches and accurately recover true causal effects after correcting for measurement error.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of artificial intelligence, companies build massive language models that can write code, draft stories, and solve complex problems. To know if these models are getting better, engineers run them through strict, offline tests where the answers are known and the scoring is precise. They can say with certainty that a new version is twelve percent better at coding than the old one. At the same time, they watch their user numbers. When a new model launches, often accompanied by a wave of press coverage and marketing campaigns, they see more people using the product. But a difficult question remains: did the improvement in the model's actual intelligence cause more people to use it, or was the surge in users simply because everyone heard about the launch?
This is a classic problem of cause and effect in a noisy environment. When a new version arrives, everything changes at once. The model gets smarter, but the marketing team runs ads, the news writes stories, and seasonal trends shift. In this storm of simultaneous changes, it is nearly impossible to isolate the specific contribution of the model's quality using standard experiments. You cannot easily keep some users on the old, worse model while others get the new one, because that creates a poor experience for the first group and is operationally difficult to maintain. Without a way to separate the model's quality from the surrounding noise, companies struggle to know if their engineering investments are truly driving user engagement.
A researcher at Google, John Tribbia, has proposed a different way to solve this puzzle by looking at how different people actually use the technology. The core idea relies on a simple observation: models do not improve evenly across every skill. A new version might get significantly better at writing computer code while only improving slightly at creative writing. This creates a natural experiment within the same group of users. Two people might be using the exact same new version of the model at the same time, but if one person spends their day writing code and the other spends it drafting fiction, they are experiencing very different levels of quality improvement. The coder sees a massive leap in performance, while the fiction writer sees a modest one. Because this difference in experience is driven by the users' own pre-existing habits and job roles, rather than a reaction to the new model, it offers a clean way to measure the true impact of quality.
To test this idea, the researcher built a simulated world where the answer was already known. He created a dataset of one hundred thousand virtual users over twenty-six weeks, introducing specific, known improvements to the model's quality in different categories. He knew exactly how much the "coding" skill improved and how much the "writing" skill improved, and he knew the true effect these improvements were supposed to have on user activity. He then applied his new method to this data to see if it could find the truth. The method worked remarkably well. By comparing users on the same version who focused on different tasks, the estimator recovered between eighty-one and eighty-eight percent of the true effect. The small gap that remained was not a failure of the method, but a result of the fact that the researchers were using past usage patterns as a proxy for current habits, which introduced a small amount of statistical noise.
The researchers addressed this noise directly. They calculated how stable the users' habits were over time and used that information to adjust their final numbers. Once this correction was applied, the method recovered the full effect, landing right on the true value. This success stood in sharp contrast to other common approaches. When the researchers tried simpler methods that ignored the specific version of the model or used real-time usage data that could be influenced by the quality changes themselves, the results were far worse. Those simpler methods only recovered about sixty-two to sixty-three percent of the true effect, missing the mark significantly. The new approach also proved it was not just finding random patterns; when the researchers shuffled the data to remove any real connection, the method correctly found no effect, confirming it was measuring something real.
The study also explored how the length of time used to measure user habits affected the results. The researchers found that using a longer history of past behavior made the measurements more precise. When they looked at a seven-week history of user habits before the new model arrived, the results were much clearer than when they looked at only three weeks. This suggests a practical tradeoff: the more historical data a company has, the more accurately they can measure the impact of their model updates. The method also successfully distinguished between different types of user outcomes. It correctly identified that quality improvements increased the number of days users were active, but it did not falsely claim that quality increased the sheer volume of messages sent, showing that the model could tell the difference between genuine engagement and mere activity.
This approach offers a powerful tool for product teams who cannot run traditional experiments. It allows them to use the natural variation in how their customers work to understand what drives engagement. By freezing the measurement of user habits before a new model arrives and then watching how different groups respond to the specific quality changes they experience, companies can isolate the true value of their engineering work. The research shows that even without a controlled experiment, it is possible to determine if a model's improvement is the reason people are using the product more, provided one accounts for the noise of the launch and the specific ways different users interact with the technology. The findings suggest that with the right statistical adjustments, the signal of quality can be heard clearly, even in the busiest of product launches.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.