ORIGINAL RESEARCH

Aerosp. Res. Commun.

Corridor-Constrained Incremental TD3 for Nacelle Tilt Scheduling in Quad-Tiltrotor UAV Level Transition

  • School of Aeronautics and Astronautics, Graduate School, Zhejiang University, Hangzhou, China

The final, formatted version of the article will be published soon.

Abstract

Level-flight transition of quad-tiltrotor unmanned aerial vehicles requires a nacelle tilt scheduler that adapts to closed-loop transition states while satisfying actuator-rate limits and a speed-dependent admissible transition corridor. Offline tilt schedules provide useful nominal references, but their feedback adaptability is limited under disturbances, measurement noise, and variations in low-level closed-loop response. In contrast, absolute-action reinforcement learning (RL) policies treat consecutive nacelle commands as independent setpoints, which may induce abrupt inter-sample variations and excessive nacelle-rate demand. This paper proposes an energy-aware smooth incremental tilt strategy based on a corridor-constrained incremental Twin Delayed Deep Deterministic Policy Gradient (TD3) framework. Instead of learning the absolute nacelle angle, the actor outputs a bounded nacelle-angle increment. The command is then generated through bounded integration, rate saturation, and projection onto a tightened transition corridor. This action generation mechanism converts the policy from an independent setpoint generator into a constrained nacelle-evolution generator, embedding command continuity and actuator-rate compatibility while enforcing corridor feasibility. The scheduling objective combines altitude regulation, forward-speed buildup, longitudinal smoothness, and a rotor-speed-based effort proxy under a gain-scheduled low-level control architecture. The feasibility layer provides corridor admissibility of the sampled command, and a post-training analysis characterizes the conditional closed-loop boundedness. The simulations show that the proposed scheduler generates smoother nacelle trajectories than absolute-action TD3 and act-penalty-based TD3, reducing nacelle-rate demand and longitudinal jerk, while maintaining corridor feasibility, bounded altitude response, terminal-speed buildup, and comparable rotor-speed-based effort.

Summary

Keywords

Nacelle tilt scheduling, quad-tiltrotor UAV, Reinforcement learning, transition corridor, transition trajectory optimization

Received

24 May 2026

Accepted

24 July 2026

Copyright

© 2026 Yang, Du, Yu, Fang and Du. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.

*Correspondence: Changping Du

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Share article