← Latest papers
🤖 AI

CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations

This paper introduces CTBench, a public benchmark designed to evaluate AI agents' troubleshooting capabilities in realistic telecom network operations, revealing that while current agents excel at path restoration, they struggle with root cause analysis and often fail to provide the evidence-grounded diagnoses required in operational practice.

Original authors: Xingyu Yan, Tingting Dai, Antonio De Domenico, Mohamed Sana, Nicola Piovesan, Changchang Li, Bowen Liu, Kun Jiang, Mengjie Zhang, Dingcheng Shan, Jing-Cheng Pang, Chenwei Wu, Sijie Wu, Lianying Chao
Published 2026-08-13
📖 4 min read☕ Coffee break read

Original authors: Xingyu Yan, Tingting Dai, Antonio De Domenico, Mohamed Sana, Nicola Piovesan, Changchang Li, Bowen Liu, Kun Jiang, Mengjie Zhang, Dingcheng Shan, Jing-Cheng Pang, Chenwei Wu, Sijie Wu, Lianying Chao, Haoran Cai, Jiantao Ye, Xubin Li, Simon Mark Lucas, Xin Chen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the captain of a massive, floating city made entirely of invisible wires and glowing signals. This city is the internet, and it's run by a team of human engineers who act like detectives. When a part of the city goes dark or a message gets lost, these engineers have to figure out exactly where the break is, why it happened, and how to fix it before anyone notices. They have to look at thousands of different types of gadgets from different manufacturers, each speaking a slightly different language, all while the city keeps moving around them.

Recently, scientists have been trying to build "AI detectives"—computer programs that can think, ask questions, and solve problems just like humans do. These AI agents are like super-smart robots that can read manuals, type commands, and connect the dots faster than any human. But here's the big question: Can these robots actually handle the messy, confusing reality of fixing a real city, or are they just good at solving clean, textbook puzzles? We need to know if they can be trusted to fix our digital world without making things worse.

This is exactly what a team of researchers from Huawei and Queen Mary University of London set out to find with a new project called CTBench. Think of CTBench as a giant, high-stakes "escape room" designed specifically for AI agents. Instead of a spooky mansion, the room is a realistic simulation of a telecom network, filled with routers, switches, and firewalls from different brands like Huawei, Cisco, and H3C. The researchers created 234 tricky scenarios based on real-life problems that human engineers face every day. Some tasks ask the AI to find the "root cause" of a problem (like finding the exact broken wire in a wall), while others ask it to rebuild a path for data to travel (like rerouting traffic after a bridge collapses).

The researchers didn't just ask the AI, "Did you get the right answer?" Instead, they watched how the AI got there. They gave the AI a "gold standard" checklist of the exact steps a human expert would take to solve the problem. If the AI guessed the right answer but skipped the important detective work, or if it looked in the wrong place, it didn't get full credit. It's like grading a math test where you lose points if you write the right number but didn't show your work.

The results were a mix of impressive skills and some serious stumbling blocks. The AI agents were surprisingly good at finding the start and end points of a broken path, almost like they had a perfect map. However, when it came to figuring out why something broke, they struggled. The AI often got confused by the different languages of the various devices or missed clues that were hidden behind partial information. In fact, even when the AI gave the correct final answer, it frequently failed to provide the "evidence" or the step-by-step reasoning that a human engineer would need to trust the fix.

The study suggests that while these AI agents are getting smarter, they aren't quite ready to replace human network engineers just yet. They tend to guess the right answer without doing the deep detective work required to be safe and reliable. The researchers found that the AI performed worse when the network was complex, with many different types of devices, or when the clues were hard to find. It turns out that just because an AI can solve a puzzle doesn't mean it understands the story behind the puzzle. Until these digital detectives can show their work as clearly as a human expert, we might still need to keep our human engineers on the case to make sure the city stays connected.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →