Mosaic: Modular Orchestration of Specialised AI Components
The paper proposes MOSAIC, a modular architecture that replaces expensive frontier language models with a strict hierarchy of hardcoded orchestration logic and specialized smaller components to achieve comparable accuracy at a significantly reduced inference cost through mechanisms like model-scale substitution, early termination, and cumulative partial resolution.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The world of artificial intelligence has recently been dominated by a single, powerful idea: make the models bigger. By adding more and more digital "neurons," researchers have created systems capable of answering almost any question, from writing poetry to solving complex physics problems. These large systems, often called frontier models, are incredibly capable, but they come with a heavy price tag. Running them requires massive amounts of computer power and energy, making them expensive to use for every single task, no matter how simple. This has led scientists to ask a fundamental question: is it always necessary to use the most powerful engine available, or could a collection of smaller, specialized tools, working together with a smart plan, do the same job for a fraction of the cost?
A researcher named Nabil Islam has proposed a new way to answer this question, outlined in a report titled "MOSAIC." Instead of relying on a single giant brain to solve every problem, MOSAIC suggests building a system where a query is broken down into smaller pieces and handed off to different, smaller experts. The core idea is that a large, general-purpose model is like a Swiss Army knife: it can do everything, but it is heavy and inefficient if you only need to cut a piece of string. MOSAIC proposes using a dedicated pair of scissors for that specific task, but only if a smart manager first decides that scissors are the right tool. The report does not present a finished product that has already been tested in the real world. Instead, it offers a detailed blueprint and a set of hypotheses about how such a system could work, arguing that it would be far more efficient than current methods while maintaining high accuracy.
The MOSAIC system is built on a strict hierarchy, which is a fancy way of saying it has a clear chain of command. At the very top sits a "Frontier Model," a large, general-purpose AI that acts as the initial contact point for the user. When a user asks a question, this top model does not try to solve it alone. Instead, it passes the request to a layer of "Assistants." These assistants are not AI models themselves; they are fixed, unchangeable rules written by humans. Their only job is to look at the request and decide how to break it down. If a question is broad, the top assistant might split it into two or three smaller, more specific questions. It then passes these smaller tasks down to the next level of assistants.
This process continues down the line. The second level of assistants takes a specific domain, such as mathematics or history, and breaks the task down even further until it reaches the third level. At this bottom level, the task is so specific and clear that it can be handed directly to a "Specialised Model." These are small AI models trained only on one specific subject, like pure mathematics or a particular type of coding. Because they are small and focused, they are much cheaper and faster to run than the giant models. Once the specialized model gives an answer, the result travels back up the chain, where the assistants check if the job is done. If the answer is good enough, the system stops. If not, it tries again, but only within strict limits to ensure the process never gets stuck in an endless loop.
One of the most important features of this design is that it avoids the common mistake of letting the AI decide which tool to use. In many other systems, an AI model acts as the manager, guessing which smaller model to call. The MOSAIC report argues that this is risky because the manager AI can make mistakes or get confused. In MOSAIC, the manager is a set of fixed rules. It does not "think" or "guess"; it simply follows a pre-written script to decompose the task. This ensures that the system never gets stuck in a circle, a problem that can happen when AI agents talk to each other without clear boundaries. The report also emphasizes that the system stops as soon as a task is solved at a higher level. If the top assistant can answer a simple question, it does not waste time or money calling the deeper, more expensive layers.
The researchers propose three main ways this system saves money. First, it uses "model-scale substitution," meaning it swaps a giant, expensive model for a small, cheap one whenever the task allows. Second, it uses "early termination," which means the system stops as soon as it finds a good answer, rather than running the full process every time. Third, it uses "cumulative partial resolution," where each level of the hierarchy solves a piece of the puzzle, so the next level only has to solve what is left, rather than starting over from the beginning. The report suggests that if these three things work together, the system could match the accuracy of the giant models while using a tiny fraction of the computing power.
However, the author is very clear about what they have not yet done. This report is a proposal, not a report of results. No experiments have been run to prove that the system works as described. The researchers state that the entire idea rests on a hypothesis: that a system of small, specialized models managed by fixed rules can be just as accurate as a single giant model. They admit that this depends on the specialized models being very good at their specific jobs and the fixed rules being perfectly designed. If the rules for breaking down tasks are wrong, the system will fail. They also note that they do not yet have a way to explain why the small models would be as accurate as the big ones; they are simply suggesting that if the small models are trained well, they might be.
The report also addresses a common concern in artificial intelligence: how to know if the system is confident in its answer. Many systems try to ask the AI, "How sure are you?" The MOSAIC authors argue that this is a bad idea because an AI can be very confident and still be completely wrong. Instead, their system relies on structure. It checks if the task has been fully broken down and if the specialized model has actually performed the calculation. It does not rely on a feeling of confidence. This approach is designed to prevent the system from giving a wrong answer with a false sense of certainty.
The researchers conclude by outlining a plan to test their idea. They propose running a series of tests where the MOSAIC system is compared directly against a standard, giant model. They would measure not just how often the system gets the right answer, but also how often it has to call the expensive giant model, how deep the chain of assistants had to go, and how much computing power was used. They warn that if the system is tuned to look efficient but ends up giving wrong answers, the test will fail. The ultimate goal is to prove that intelligence does not have to be expensive, and that a well-organized team of small, focused tools can do the work of a giant, all-purpose machine. Until those tests are run, the MOSAIC system remains a compelling theory, a detailed map of a potential future where AI is cheaper, faster, and more efficient, but not yet a proven reality.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.