ETRI-Knowledge Sharing Plaform

KOREAN
논문 검색
Type SCI
Year ~ Keyword

Detail

Conference Paper Residual Policy Learning을 활용한 수중 ROV 6자유도 자세 안정화 제어
Cited - time in scopus Share share facebook twitter linkedin kakaostory
Authors
최승혁, 김용현, 조현우, 정종대
Issue Date
2026-05
Citation
한국해양과학기술협의회 공동학술대회 2026, pp.1-1
Publisher
한국해양과학기술협의회
Language
Korean
Type
Conference Paper
Abstract
Autonomous 6-DOF control of ROVs is challenging due to nonlinear hydrodynamics and compound disturbances including ocean currents, thruster degradation, and payload variation. Pure reinforcement learning (RL) fails to learn basic attitude stabilization — achieving 100% flip rate even after 3M training steps. This work applies Residual Policy Learning to ROV control, blending a PID expert with a learned policy so that the expert guarantees attitude stability while RL provides residual correction. Training was conducted in the Stonefish physics simulator across 32 parallel environments, enabled by a custom stepped simulation mode via ROS2 C++ modification. The BlueROV2 Heavy (8 thrusters, 11.5 kg) was trained with domain randomization and a three-stage curriculum over 10M steps. Comparative evaluations demonstrate that the proposed Residual RL consistently maintains a 0% flip rate across all disturbance scenarios, effectively inheriting the safety of the PID expert while achieving superior tracking performance where Pure RL fails entirely. These results validate that our approach provides both the robustness of classical control and the adaptive correction of RL. Future work will focus on sim-to-real transfer on a physical BlueROV2 Heavy.
Keyword
Residual policy learning, Underwater ROV, Attitude stabilization, Domain randomization, Physics simulator
KSP Keywords
Adaptive correction, DOF control, Nonlinear hydrodynamics, Ocean currents, Policy learning, ROV control, Six degrees of freedom(6-DoF), Three-stage, Tracking Performance, Underwater ROV, attitude stabilization