Xiaofeng Wang, Shaopu Wang, Jipeng Zhang, Yue Ma, Hierarchical multi-policy adversarial reinforcement learning, Engineering Applications of Artificial Intelligence, Volume 181, Part 1, 2026, 10.1016/j.engappai.2026.115297.
In complex and dynamic autonomous driving environments, deep reinforcement learning policies often struggle with generalization due to distribution shifts between training and deployment. While conventional approaches like domain randomization and adversarial training offer partial solutions, they face notable limitations. Specifically, domain randomization relies on manually defined distributions that poorly capture real-world perturbations, while adversarial training suffers from high computational costs, unstable optimization, and behavioral homogenization. To address these issues, we propose a hierarchical multi-policy adversarial reinforcement learning framework that enhances policy generalization under complex perturbations. It incorporates three key components: a multi-policy generation architecture with a shared backbone and independent heads for diverse adversary modeling; a dual-perspective policy expansion regularization mechanism; and an adaptive curriculum-based adversarial training framework. The engineering value translates to a reliable, high-performance solution for autonomous driving, as validated on high-fidelity driving benchmarks. Experimentally, our approach demonstrates superior convergence, stability, and out-of-distribution generalization, achieving performance improvements ranging from 7.36% to 50.06% over established baseline methods while maintaining computational efficiency. These results, obtained in a high-fidelity simulator, demonstrate promise for real-world autonomous driving applications and lay the groundwork for future real-world deployment.
