OTel: Building Domain-Specialized Telecom LLM Foundations for Intelligent Networks
The paper introduces Open Telco (OTel), a comprehensive open-source resource featuring derived datasets and 30 post-trained baselines that significantly advance domain-specialized AI performance for embedding, reranking, and language modeling tasks in the telecommunications sector.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Modern life runs on a vast, invisible nervous system of cables, towers, and satellites that carries our calls, messages, and data across the globe. Keeping this system running requires a deep understanding of thousands of technical manuals, international agreements, and constantly changing rules that govern how networks behave. These documents are dense, interdependent, and written in a specialized language that even seasoned engineers find difficult to navigate. For years, artificial intelligence has struggled to make sense of this specific world. While general AI models can write poetry or summarize news, they often fail when asked to interpret a complex network specification or troubleshoot a technical error, because they lack the precise, domain-specific training data needed to understand the unique logic of telecommunications.
A new initiative called Open Telco, or OTel, aims to bridge this gap by providing the first open, shared foundation for building intelligent systems that truly understand telecommunications. The project brings together more than one hundred experts from universities, industry, and global organizations to create a massive library of cleaned, high-quality data and pre-trained computer models. Instead of forcing every researcher to start from scratch—scrambling to find documents, cleaning up messy text, and training models alone—OTel offers a complete toolkit. It includes datasets specifically designed to teach computers how to find the right information, how to rank that information by importance, and how to answer questions based strictly on what they find, all while knowing when to admit they do not know the answer.
The team behind OTel gathered roughly 1.1 million raw pieces of information from public sources, including international standards documents, technical white papers, and academic research. However, raw data is often too noisy to be useful for training smart systems. The researchers applied a rigorous, multi-stage filtering process to clean this information, removing errors, duplicates, and low-quality examples. This intensive work reduced the collection to about 327,000 high-confidence examples that are ready for use. These examples were then organized into four distinct types of training materials: one set to teach models how to retrieve relevant documents, another to teach them how to rank those documents, a third to teach them how to generate answers based on the retrieved text, and a fourth to teach them when to remain silent if the information is insufficient or off-topic.
Using this curated data, the researchers trained and released thirty different computer models of varying sizes, ranging from small, fast systems to large, powerful ones. These models were tested on their ability to handle telecommunications tasks, and the results showed a clear improvement over previous attempts. The models trained on OTel data became significantly better at finding the right documents, with their ranking accuracy jumping from roughly 35 percent to over 93 percent in some cases. They also became much better at ranking search results and answering questions correctly, with the largest model achieving an 88 percent success rate on technical questions. Crucially, the training taught the models to avoid guessing when they lacked enough information, a safety feature that prevents them from making up technical details.
The response from the global community has been immediate and substantial. Within months of release, these models were downloaded more than 16 million times, and the project received widespread attention from industry and media outlets around the world. This level of engagement suggests that the telecommunications industry has a deep, unmet need for open, reliable tools that can help manage its complex infrastructure. By providing a shared starting point, OTel allows researchers and engineers to focus on solving specific problems rather than rebuilding the basics. The project is not presented as a finished solution but as a reproducible foundation that the community can expand, improve, and adapt to make telecommunications networks safer, more efficient, and easier to manage.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.