Decoupled Alignment for Robust Plug-and-Play Adaptation
This paper introduces DAPA, a training-free, plug-and-play safety enhancement method that leverages knowledge distillation and model fusion to inject alignment signals from well-aligned models into shadow-aligned ones, significantly improving defense success rates against harmful inputs without compromising performance.