Predatory MoE: Serving Mixture-of-Experts with a Sequence-Local Expert Roster
This paper proposes a sequence-local expert roster strategy that fixes a small set of experts per sequence to drastically reduce memory footprint and expert transfers during MoE model serving, achieving near-full-residence throughput with significantly fewer resident weights by applying this constraint consistently during both training and inference.