← Latest papers
💬 NLP

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains

Relay-Bench is a holistic, text-only benchmark designed to evaluate large language models on multi-domain reasoning chains by presenting unsaturated composite problems that require integrating diverse skills like coding, math, and web search within a single prompt, where the leading model GPT-5.5 (xHigh) achieves a score of 43.3%.

Original authors: Liam Swayne

Published 2026-07-22
📖 4 min read☕ Coffee break read

Original authors: Liam Swayne

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where your favorite smart assistant isn't just a trivia bot or a code-cruncher, but a true multitasking wizard. In the rapidly evolving field of Artificial Intelligence, researchers are constantly building "benchmarks"—which are essentially standardized tests—to see how smart these digital brains really are. For a long time, these tests were like separate subjects in school: one exam for math, another for history, and a third for coding. But real life doesn't work that way. When you ask a human to plan a trip, they don't just calculate distances; they check the weather, read reviews, book a hotel, and maybe even write a funny email to a friend, all in one go. This paper dives into a specific corner of AI research called "reasoning chains," where the goal is to see if a model can juggle many different types of tasks at once without dropping the ball. The big question everyone cares about is: Can these AI models handle the messy, complicated, multi-step problems of the real world, or do they get confused and give up when things get too crowded?

Enter Relay-Bench, a new, super-challenging test designed to see if AI models can run a "relay race" of the mind. Instead of asking a model to solve one math problem or write one poem, the creators of this test string together a whole chain of different tasks—like solving a math puzzle, finding a specific fact on the web, writing a bit of code, and decoding a secret message—all inside a single, massive prompt. Think of it like a video game level where you have to jump over a pit, solve a riddle to open a door, and then fight a boss, all without hitting the "reset" button. The researchers built 31 of these complex challenges, some of which are so long and cluttered with fake information that they feel like reading a novel just to find a single number. They also added a "secret code" layer, where every word in the instructions is replaced with a strange three-letter code, forcing the AI to decode the message before it can even start solving the problems.

When they ran the top three AI models of the time (GPT-5.5, Gemini 3.1 Pro, and Claude Opus 4.7) through this gauntlet, the results were a bit of a shock. Even the smartest model, GPT-5.5, only got about 43.3% of the answers right. That's less than half! The other models scored even lower, with one of them only managing 16.7%. The paper suggests that while these AI models are getting very good at single tasks, they struggle significantly when asked to chain multiple different skills together, especially when the instructions are confusing or the task is very long. In fact, the test was so hard that many models simply gave up or got stuck in loops, unable to finish the race.

The researchers also found that these models sometimes "hallucinate," which is a fancy way of saying they make up answers when they don't know the solution. One model, Claude Opus 4.7, was very cautious; it refused to answer many of the encoded questions entirely, likely because its safety filters were too strict. Another model, Gemini 3.1, tried to guess more often but ended up with more wrong answers. The study highlights that current AI models are still far from being the all-around experts we hope they will be. They are like brilliant students who can ace a math test but might panic if you ask them to do math while simultaneously translating a poem and searching for a recipe. The authors suggest that it might take another year or two before AI models get good enough to consistently pass this kind of "relay race" test, and they warn that as models get smarter, these tests will need to get even harder to keep up.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →