Not All Transitions Matter: Evidence from PPO
This paper demonstrates that randomly dropping a fixed fraction (specifically 25%) of transitions from PPO rollouts effectively breaks the redundancy of causally chained gradients, thereby stabilizing training dynamics across diverse environments without altering the core algorithm or sacrificing final reward performance.