← Latest papers
💻 computer science

Toward a Threat Actor Profiling Taxonomy for Pre-Release Risk Management of Open-Weight Frontier Models

This paper proposes a six-attribute, empirically grounded taxonomy for explicitly characterizing threat actors to standardize pre-release risk management and improve the interpretability and faithfulness of evaluations for open-weight frontier AI models.

Original authors: James Zhang

Published 2026-08-27
📖 6 min read🧠 Deep dive

Original authors: James Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the rapidly evolving world of artificial intelligence, a specific type of software has emerged that functions like a powerful, universal tool. These systems, known as frontier models, can write code, analyze complex data, and solve problems that once required human experts. For years, the most advanced versions of these tools were kept behind closed doors, accessible only through secure internet connections where developers could monitor how they were used. However, a new wave of these models is being released with their underlying "weights"—the core mathematical instructions that make them work—freely available for anyone to download. This shift is comparable to handing out the blueprints and engine parts of a high-performance car to the public; while it allows for incredible innovation and research, it also means that once the parts are out, they cannot be taken back. If someone downloads these files, they can modify the software, remove safety features, or use it for harmful purposes without the original creators ever knowing.

The central challenge for the people who build these systems is figuring out how dangerous they might be before they are released. Currently, safety teams run tests to see what a model can do, often asking it to solve difficult problems or trying to trick it into revealing harmful information. But a significant gap exists in how these tests are designed. Safety researchers often assume a generic "bad actor" might try to misuse the technology, but they rarely define exactly who that person is. Is it a lone teenager with a laptop? A well-funded criminal gang? A state-sponsored team of experts? Without a clear picture of the adversary, the tests might be too easy for a sophisticated attacker or too hard for a novice, leaving developers with a false sense of security or unnecessary fear. The question is not just what the software can do, but who is likely to use it and how much effort they would need to succeed.

A recent paper submitted to Tsinghua University addresses this uncertainty by proposing a new way to describe potential threats. The author, James Q Zhang, argues that before any safety test is run, developers must explicitly define the characteristics of the person or group they are trying to protect against. To do this, they created a structured system, or taxonomy, that breaks down a threat actor into six specific categories. Instead of vague labels like "skilled" or "well-resourced," the system asks for concrete details: How technically advanced is the actor? Do they already know about biology or cyber warfare? How many people are in their group? What kind of equipment do they have? How much money can they spend? And how much time are they willing to invest?

The researchers developed this framework by looking at how experts in other fields, such as cybersecurity and terrorism studies, already analyze threats. They found that while these fields have detailed ways to describe attackers, the artificial intelligence community had not yet adopted a similar standard. The new system organizes these six attributes into a grid with different levels of capability, ranging from a basic user with minimal resources to a highly sophisticated, well-funded institution. For example, under "financial capacity," the levels range from having less than one thousand dollars to having over ten million. Under "time horizon," the levels span from impulsive actions lasting less than a day to enduring campaigns that last for six months or more.

The paper demonstrates the value of this approach by creating two distinct profiles to show how different threats require different safety tests. The first profile describes a "radicalized graduate researcher." This individual has deep knowledge of biology and has been working on a harmful goal for several months, but they are working alone with very little money and only basic computer skills. Because this person already knows the science, a safety test for them should focus on whether the AI can help them plan the logistics of an attack, rather than teaching them the science itself. The second profile describes a "small, financially motivated criminal group." These actors have moderate computer skills and some money to rent equipment, but they lack specific knowledge about their target. For them, the most dangerous risk is that the AI could fill in their knowledge gaps, acting as a guide to help them find vulnerabilities they wouldn't have seen on their own.

By using this structured system, developers can stop guessing and start designing tests that match the real-world risks they are worried about. If a company knows they are worried about a well-funded group with a long-term plan, they can set up a safety test that gives their testers a large budget and plenty of time, simulating a serious, sustained effort. If they are worried about a lone individual with limited resources, the test can be adjusted to reflect those constraints. The author suggests that this method should become a standard part of the release process, similar to how medical researchers must write down their plans before starting a clinical trial. This would ensure that safety evaluations are not just random checks, but precise measurements of risk against specific, well-defined adversaries.

The paper emphasizes that this is particularly urgent for open-weight models, where the software is released to the public and cannot be recalled. Once these weights are out, the developers lose control, making the pre-release decision the only chance to prevent harm. The author does not claim that this system solves all safety problems or that it can predict the future with certainty. Instead, they offer a tool to make the assumptions behind safety tests visible and comparable. By forcing developers to fill in the details of who they are protecting against, the system makes it harder to overlook dangerous scenarios and easier to compare the safety of different models. The ultimate goal is to move the industry from vague warnings to clear, actionable risk management, ensuring that the powerful capabilities of these new technologies are understood in the context of the people who might try to misuse them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →