A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications
This paper presents a framework demonstrating that while LLMs can generate consistent synthetic responses for security surveys, they exhibit systematic biases and homogenization that make them suitable for piloting and hypothesis generation but insufficient for replacing actual expert elicitation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the high-stakes world of cybersecurity, understanding how human experts think is crucial. When a security breach occurs, the decisions made by analysts in a Security Operations Centre (SOC) can determine whether a threat is contained or spreads across a network. To understand these decision-making processes, researchers often rely on surveys and interviews with these professionals. However, finding enough experts to answer these surveys is notoriously difficult. These specialists are often overworked, dealing with constant alerts and the threat of burnout, and many work under strict confidentiality rules that prevent them from sharing details about their daily work. As a result, studies often suffer from small sample sizes, leaving researchers with incomplete pictures of the industry.
Recently, a new tool has emerged that promises to solve this shortage: large language models. These are advanced computer programs trained on vast amounts of human writing, capable of generating text that sounds remarkably like a person. Some researchers have begun to wonder if these artificial intelligences could act as stand-ins for real experts, filling in the gaps of survey data by generating thousands of synthetic responses. The idea is appealing: if a computer can mimic an expert, it could allow scientists to test their questions, explore different scenarios, and gather data without needing to track down busy humans. But a critical question remains: can a machine truly think like a security professional, or does it merely produce a convincing imitation that misses the nuance of real human experience?
A team of researchers set out to answer this question by building a rigorous framework to test whether these artificial models could serve as reliable substitutes for human experts in security surveys. They did not simply ask the models a few questions and see what happened; instead, they designed a comprehensive experiment to see if the computer-generated answers held up under scrutiny. The researchers gathered real survey data from six cybersecurity professionals working in different sectors, ranging from healthcare to finance. These experts had varying levels of experience and held different roles, such as managers, analysts, and specialists. The researchers then asked six different large language models to simulate these specific experts, creating digital personas that matched the real people's backgrounds and job titles.
The experiment was conducted in three distinct ways to test the models from every angle. First, the researchers asked the models to act as individual experts, trying to replicate the specific opinions of the real humans. Second, they asked the models to generate the overall patterns of a whole group, simulating the aggregate results of a large survey. Finally, they tested the models over time, asking them to predict how expert opinions had shifted across several years of real-world data. Throughout this process, the researchers ran each simulation multiple times to see if the models gave the same answers when asked the same question repeatedly, checking for consistency and stability.
The results revealed a clear and significant gap between the artificial simulations and human reality. While the large language models were remarkably consistent with themselves—often giving the same answer when asked the same question multiple times—they failed to capture the true diversity of human thought. Real security experts showed a wide range of opinions; some were confident in their choices, while others were uncertain, and their answers varied depending on their specific experiences and the unique pressures of their workplaces. In contrast, the models tended to converge on a single, safe, middle-ground answer. They smoothed over the disagreements and uncertainties that are natural to human experts, producing responses that were too uniform and too polite.
When the researchers compared the distribution of answers from the models against the real data, the models appeared to be getting the general shape of the data right, but they were missing the details. They tended to avoid extreme or rare opinions, clustering their responses around the average. This "central tendency bias" meant that while the models could generate a plausible-looking survey result, they were effectively erasing the very variations that make human expertise valuable. For instance, when asked about the challenges of adopting new tools, real experts highlighted specific, gritty operational hurdles, while the models offered generic, textbook-style difficulties that lacked the texture of actual practice.
The study also examined whether these models could track changes in the industry over time. The researchers found that the models were largely time-insensitive. Even when asked to simulate responses from different years, the models produced the same static patterns, failing to reflect the evolving nature of cybersecurity threats and practices. This suggests that the models are not truly "learning" or adapting to new information in the way a human expert would; instead, they are recalling a fixed set of concepts that do not shift with the times. This lack of temporal awareness means they cannot be trusted to predict how expert opinions might change in the future or to accurately reflect historical trends.
Ultimately, the researchers concluded that while these large language models are powerful tools, they are not ready to replace human experts in security research. The models are excellent for the early stages of a study, such as helping researchers draft survey questions, testing whether a question is clear, or generating hypothetical scenarios to explore. They can serve as a useful sounding board for ideas. However, they should never be used to replace the actual collection of expert opinions. The unique, varied, and sometimes contradictory nature of human judgment is something that cannot be synthesized by an algorithm. To understand the real world of cybersecurity, researchers must still rely on the messy, complex, and invaluable insights of the humans who live and work within it. The study serves as a vital guide, ensuring that as the field embraces new technologies, it does not lose sight of the human expertise that remains the foundation of security.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.