← Latest papers
💻 computer science

Training Versus Hiring in Security Operations Centers: Agent-Based and Reinforcement Learning Simulations of Training Duration Effects on CyberTeam Performance

This study utilizes agent-based modeling and reinforcement learning to demonstrate that training duration in Security Operations Centers yields nonlinear performance gains, with optimal investment occurring between 500 and 1,000 episodes and highly trained teams achieving efficiency advantages equivalent to recruiting 20–25 additional personnel.

Original authors: Max N. Yaw, Ferdinand Kpieleh, Aos Mulahuwaish, Basheer Qolomany, Jacques Bou Abdo

Published 2026-06-27
📖 4 min read☕ Coffee break read

Original authors: Max N. Yaw, Ferdinand Kpieleh, Aos Mulahuwaish, Basheer Qolomany, Jacques Bou Abdo

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a Security Operations Center (SOC) is like a high-stakes video game tournament. The organization has a budget and a critical choice to make: Should they hire a huge army of new, inexperienced players, or should they take their small, existing team and put them through an intense, long-term training camp?

This paper uses a computer simulation to answer that question. Instead of watching real people for years, the researchers built a digital world where "agents" (computer programs representing humans) play a game of cyber attack and defense. They used a special type of AI called Reinforcement Learning to act as a "training simulator." Just like a human gamer gets better the more they play, these AI agents learned strategies by playing thousands of rounds, getting rewards for winning and penalties for losing.

Here is what they discovered, broken down into simple concepts:

1. The "Three-Act" Training Story

The paper found that training doesn't make you better at a steady, straight line. It happens in three distinct phases, like the acts of a play:

  • Act 1: The Clumsy Beginner Phase (Episodes 1–500):
    Imagine a new player who just picked up the controller. They are exploring, trying random moves, and mostly failing. In the simulation, teams in this phase won almost nothing (0% to 15% win rate). They were just figuring out the rules.
  • Act 2: The "Aha!" Moment (Episodes 500–2,000):
    This is the sweet spot. The team stops guessing and starts understanding. They figure out how to work together, and their performance skyrockets. Their win rate jumps from about 15% to over 50%. This is where the most value is gained.
  • Act 3: The Master Polish (Episodes 2,000–5,000):
    The team is now a pro. They are winning almost all the time (93% to 95%). However, every extra hour of training here adds only a tiny bit more skill. It's like a master chef spending extra time to make a dish 1% tastier; it's good, but the massive gains are already done.

2. The "Sweet Spot" for Your Wallet

If you are the boss deciding how long to train your team, the paper suggests a specific strategy: Don't stop too early, but don't train forever.

  • The Danger Zone: If you stop training after the first 500 "games," you are throwing away the most valuable part of the process. You're still in the "clumsy" phase.
  • The Golden Zone: The best return on your investment happens between 500 and 1,000 games. This is where the team transforms from "novices" to "competent pros" with the biggest jump in skill.
  • The Diminishing Returns: After 2,000 games, you are still getting better, but it costs a lot of time for very small improvements.

3. The "Magic Multiplier": Small Teams vs. Big Armies

This is the most surprising finding. The researchers compared a small, highly trained team against large, untrained teams (teams that just follow basic rules and never learn).

  • The Result: A team of just 5 trained attackers performed just as well as a team of 25 to 35 untrained attackers.
  • The Analogy: Think of it like a group of 5 elite special forces soldiers versus a mob of 30 untrained civilians. The 5 trained soldiers, because they know exactly what to do and how to coordinate, can beat the much larger group.
  • The Math: The paper claims that one trained attacker is roughly equal to 3 to 7 untrained attackers.

The Bottom Line for Organizations

The paper concludes that if you have to choose between hiring 20 new people or training your current 5 people for a long time, training is the smarter move.

  • Hiring gives you a big team that is slow, clumsy, and needs a lot of people to get the job done.
  • Training turns a small team into a highly efficient machine that can do the work of a much larger group.

The study suggests that organizations should design training programs that last long enough to get past the "clumsy" phase and reach the "expert" phase (around 1,000 to 2,000 simulated training sessions). Doing so creates a team that is not only better but also much more cost-effective than simply buying more bodies.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →