Data-Shard-Driven Expert Differentiation in Sparse MoE: A Three-Component System with FrozenPath Anchoring and Dual-Loop Refinement
This paper proposes a distillation-free, auxiliary-loss-free three-component system combining FrozenPath anchoring, Data Shard binding, and Dual-Loop refinement to effectively eliminate expert homogenization and catastrophic language drift in sparse Mixture-of-Experts models during incremental training.