Coverage You Can Steer: Online Conformal Calibration for RL-Driven Hardware-Aware NAS
This paper proposes an online conformal calibration framework using adaptive feedback control to restore distribution-free coverage guarantees in reinforcement learning-driven hardware-aware neural architecture search, enabling the efficient pruning of 25–50% of candidate architectures without sacrificing accuracy or violating the target error rate despite the non-exchangeable nature of the search process.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where the tiny computers inside your smartwatch, your car's braking system, or a medical sensor could run the same powerful artificial intelligence that currently lives only in massive data centers. This is the promise of edge AI: intelligent systems that work locally, instantly, and without needing to send private data back to the cloud. But there is a catch. These edge devices have severe limits on how much memory they hold and how quickly they can process information. Designing a neural network—a type of computer program that learns from data—to fit these tight constraints is like trying to build a skyscraper that must also fit inside a shoebox. The number of possible designs is so vast, involving millions of combinations of layers, connections, and settings, that a human designer cannot possibly test them all.
To solve this, researchers use a method called neural architecture search, where a computer program automatically tries out different designs to find the best one. However, testing a single design is incredibly expensive and slow; it requires training the model on data to see how well it works. If a computer has to train thousands of designs just to find one good one, the process becomes too costly to be practical. The challenge, then, is to figure out which designs are worth the time and money to train, and which ones should be discarded immediately, without ever running the full test.
A team of researchers at Vicomtech in Spain has developed a new way to make this filtering process reliable. They tackled a problem that had been quietly undermining previous attempts: the computer's own learning process was changing the rules of the game while it was playing. In these search systems, a "controller" suggests new designs based on what it has learned so far. As the controller gets smarter, the designs it suggests change. Older methods for filtering out bad designs assumed that the designs suggested at the beginning of the search were statistically similar to those suggested at the end. But because the controller is learning and improving, this assumption is false. The designs keep shifting, and the old filters, which were calibrated on the early designs, began to fail. They would either waste time training terrible designs or, worse, accidentally throw away the single best design before it could be properly tested.
The researchers replaced this static, one-time filter with a system that learns and adjusts in real time. Instead of setting a rule once and hoping it holds, their new method uses a feedback loop to constantly check its own performance. After every design is tested, the system asks: "Did I correctly predict that this design would be good or bad?" If the system made a mistake, it slightly adjusts its internal threshold for what counts as a "good" design. This adjustment happens continuously, allowing the system to track the changing nature of the search. The result is a filter that can be dialed to a specific level of safety. If a researcher asks for a 90% guarantee that no good design will be missed, the system delivers exactly that, no matter how the search evolves. If they ask for 95%, it delivers that instead. This control is precise, reproducible, and works even when the designs being suggested are shifting rapidly.
The team tested this approach across three different types of network architectures and on various datasets, simulating the search for models that could run on microcontrollers with very limited memory. They found that their adaptive system could safely discard between 25% and 50% of the proposed designs without ever training them, saving a massive amount of computing time. Crucially, this pruning did not hurt the final quality of the solution. In fact, in one specific test case where the best design was a rare, isolated peak in a vast landscape of mediocre options, the old methods failed repeatedly, throwing away the best design. The new adaptive method found it every time.
The researchers also discovered that the "smart" controller often used to generate these designs was not actually the most efficient way to find the best architecture. When they compared a sophisticated learning controller against a simple random search, the random search performed just as well, if not better, in many cases. The controller was good at finding average designs, but it tended to get stuck in local areas and miss the rare, perfect ones. The true value of the new method, they found, lay not in the controller itself, but in the calibrated filter that managed the budget. By using the adaptive filter as a guide for where to look next, rather than just a gatekeeper, the researchers could steer the search directly toward the best solutions. This "calibrated optimism" allowed the system to explore promising but unproven designs with a known level of risk, outperforming random guessing and even more complex planning algorithms in constrained environments.
The study also looked at how these filters behave when the search builds a design layer by layer, rather than all at once. They found that a single, global rule for filtering often failed to account for the specific difficulties of different stages in the construction. For example, a rule that worked well for the early, shallow parts of a network might be too loose or too strict for the deep, complex parts. By applying their adaptive method separately to each stage of the construction, they could ensure that the safety guarantee held true at every step, regardless of how deep the network had grown. This granular control prevented the system from making systematic errors that would have gone unnoticed with a one-size-fits-all approach.
Ultimately, this work demonstrates that in the high-stakes game of designing artificial intelligence for limited devices, the most important tool is not necessarily a smarter predictor, but a more honest one. The old methods relied on assumptions about how the data would behave, assumptions that were broken by the very act of learning. The new method abandons those assumptions. Instead, it relies on a continuous, distribution-free check against reality, adjusting its confidence level moment by moment. It proves that you can have a search process that is both efficient and safe, one that knows exactly how much risk it is taking and can be tuned to take more or less, simply by turning a dial. This shift from static rules to dynamic, self-correcting control offers a robust path forward for building the next generation of intelligent, resource-constrained devices.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.