Abstract
Level-flight transition of quad-tiltrotor unmanned aerial vehicles requires a nacelle tilt scheduler that adapts to closed-loop transition states while satisfying actuator-rate limits and a speed-dependent admissible transition corridor. Offline tilt schedules provide useful nominal references, but their feedback adaptability is limited under disturbances, measurement noise, and variations in low-level closed-loop response. In contrast, absolute-action reinforcement learning (RL) policies treat consecutive nacelle commands as independent setpoints, which may induce abrupt inter-sample variations and excessive nacelle-rate demand. This paper proposes an energy-aware mooth incremental tilt strategy based on a corridor-constrained incremental Twin Delayed Deep Deterministic Policy Gradient (TD3) framework. Instead of learning the absolute nacelle angle, the actor outputs a bounded nacelle-angle increment. The command is then generated through bounded integration, rate saturation, and projection onto a tightened transition corridor. This action generation mechanism converts the policy from an independent setpoint generator into a constrained nacelle-evolution generator, embedding command continuity and actuator-rate compatibility while enforcing corridor feasibility. The scheduling objective combines altitude regulation, forward-speed buildup, longitudinal smoothness, and a rotor-speed-based effort proxy under a gain-scheduled low-level control architecture. The feasibility layer provides corridor admissibility of the sampled command, and a post-training analysis characterizes the conditional closed-loop boundedness. The simulations show that the proposed scheduler generates smoother nacelle trajectories than absolute-action TD3 and act-penalty-based TD3, reducing nacelle-rate demand and longitudinal jerk, while maintaining corridor feasibility, bounded altitude response, terminal-speed buildup, and comparable rotor-speed-based effort.
Introduction
Tiltrotor vehicles combine the vertical take-off and landing capability of rotorcraft with the efficient cruise performance of fixed-wing aircraft, making them attractive for space-constrained deployment, long-endurance missions, and multi-regime flight operations [, ]. During a complete mission, the vehicle must pass through rotor-borne flight, transition flight, and wing-borne flight. Among these phases, transition flight is particularly challenging because the primary lift source gradually shifts from rotor thrust to wing aerodynamic lift, while the propulsion direction, aerodynamic characteristics, and control effectiveness vary with the nacelle angle and forward speed.
As the nacelles rotate from the vertical-lift configuration toward the forward-propulsion configuration, rotor thrust is redistributed between vertical support and forward acceleration. Meanwhile, rotor wake interacts with the wing and fuselage, resulting in significant variations in lift, drag, and pitching moment characteristics [–]. These aerodynamic variations are coupled with thrust-vector evolution, actuator dynamics, control allocation, and longitudinal flight-state changes. Consequently, transition dynamics exhibit strong nonlinearities, time-varying control effectiveness, and rate-limited actuation characteristics [–]. The transition performance is therefore highly sensitive to the nacelle tilt schedule, which directly affects altitude regulation, forward-speed buildup, actuator usage, transition smoothness, and energy-related performance.
For level-flight transition, the aircraft is required to maintain constant altitude while accelerating from the rotor-borne condition to the wing-borne condition. In this process, the nacelle angle and forward speed must satisfy a feasible matching relationship to maintain sufficient lift, power margin, and control authority. This admissible relationship is commonly represented by a transition corridor in the nacelle-angle--speed plane [–]. Recent studies further enrich this corridor by incorporating aerodynamic variables, angle-of-attack constraints, power limits, and attitude-safety indices [, ]. These studies indicate that the transition corridor provides not only a feasibility boundary, but also a natural framework for formulating nacelle tilt scheduling as a high-level transition-trajectory design problem. Accordingly, the key scheduling task is to generate a nacelle command that evolves with forward speed while satisfying corridor feasibility and actuator-rate compatibility.
Existing studies on transition trajectory design can be broadly categorized into corridor-based planning, trim-based optimization, direct collocation or pseudospectral optimization, model-predictive control, and heuristic search methods. Corridor construction and trim analysis have been used to establish feasible nacelle-angle--speed envelopes and equilibrium Refs. [, ]. Direct collocation, pseudospectral methods, and model-based optimization have been applied to generate constrained transition trajectories under aerodynamic, power, and attitude limitations [, ]. Heuristic search methods, such as ant colony optimization, have also been employed to obtain feasible transition paths by considering safety, power, and control-effort indices []. In addition, energy-oriented transition scheduling has received increasing attention, since different acceleration and tilt profiles may lead to distinct power demands, transition durations, and mission-level energy performance [, ].
Although these methods provide important insights into transition-corridor construction and nominal trajectory optimization, many of them rely on prior aerodynamic models, trim computations, lookup tables, or predefined trajectory parameterizations. The resulting schedules can be highly effective under nominal conditions, but they are usually applied in an open-loop or weakly feedback-driven manner. Their adaptability may therefore degrade when model uncertainty, wind disturbances, measurement noise, or variations in the low-level closed-loop response are present. This motivates a feedback-capable nacelle scheduler that can adapt the tilt evolution according to the measured transition state while preserving the physical feasibility of the nacelle command.
Reinforcement learning (RL) offers a promising tool for adaptive decision-making in nonlinear and uncertain systems. In UAV and eVTOL control, RL has been applied to trajectory tracking, attitude control, controller-parameter tuning, energy optimization, and disturbance-aware control [–]. In particular, the twin delayed deep deterministic policy gradient (TD3) algorithm is well suited for continuous-control tasks because its clipped double critics, delayed policy updates, and target policy smoothing reduce overestimation bias and improve training stability []. Recent work has shown that TD3 can be used to generate energy-oriented pitch-angle references for tilt-rotor quadrotors in cooperation with robust low-level controllers []. These results indicate the potential of RL for improving transition-related performance under uncertain operating conditions.
However, most existing RL-based studies on tiltrotor vehicles focus on pitch-angle reference generation, low-level controller synthesis, or adaptive controller tuning. The nacelle transition schedule itself has received limited attention as a constrained, smooth, and multi-objective high-level scheduling variable. This distinction is important because nacelle tilt scheduling differs from generic continuous-action control. The nacelle angle is a slow physical transition variable: adjacent commands are coupled by actuator-rate limits, its admissible range depends on forward speed through the transition corridor, and its variation directly changes thrust-vector distribution, wing-lift buildup, rotor-speed demand, and longitudinal acceleration. Therefore, a learning-based nacelle scheduler should not only improve transition performance, but also generate a physically realizable command evolution that is smooth, rate-compatible, and feasible within the admissible corridor.
This requirement renders the action representation of the learning policy a central design issue. If an RL policy directly outputs the absolute nacelle angle, consecutive commands are treated as independent setpoints. Under exploration noise, critic approximation errors, or policy updates, such an absolute-action representation may generate abrupt inter-sample variations, excessive nacelle-rate demand, or degraded longitudinal smoothness. Recent studies on action-space design in deep RL have shown that action-space shaping, parameterized actions, lifted or relative action representations, and smooth-policy regularization can significantly affect learning efficiency, trajectory smoothness, and physical compatibility [–]. In robotic and autonomous driving tasks, lifted or relative action spaces have been used to align learned commands with the physical evolution of the controlled system [], while smooth-policy approaches indicate that action regularity can be enforced through policy structure rather than solely through external filtering [26]. Nevertheless, these ideas have not been tailored to quad-tiltrotor nacelle scheduling, where the learned action should reflect the incremental and rate-limited evolution of the nacelle actuator.
Action-level smoothness alone, however, is insufficient. The learned command must also satisfy hard transition-feasibility constraints. Safety filtering and projection-based safe RL have been investigated as modular ways to enforce constraints after policy output. Predictive safety filters, robust MPC filters, constrained projection layers, and optimization-based output layers have been used to modify potentially unsafe actions while preserving the learning policy as the performance optimizer [27–30]. These studies motivate a separation between performance-oriented policy learning and hard constraint enforcement. For tiltrotor level-flight transition, however, the critical feasibility constraint is not merely a fixed input bound or a generic state constraint. It is a speed-dependent nacelle-angle corridor determined by aerodynamic lift capability, power margin, and transition feasibility. Therefore, a suitable learning-based nacelle scheduler should couple action representation with feasibility enforcement, so that the learned command follows the rate-limited evolution of the nacelle actuator while the executed command remains inside the speed-dependent transition corridor.
The above discussion shows that existing transition-scheduling studies and RL-based control methods have not yet provided an integrated solution for feedback-capable nacelle scheduling, physically consistent action evolution, and hard feasibility enforcement. To address this issue, this paper proposes an energy-aware smooth incremental tilt (ESIT) strategy for the level-flight transition of a quad-tiltrotor unmanned aerial vehicle. ESIT is developed under a hierarchical transition-control architecture, where the scheduled low-level controller maintains attitude and altitude tracking, and the high-level scheduler determines the nacelle evolution from measurable transition states. A Beetle Swarm (BS) optimized reference schedule is used to warm-start the actor and provide a physically reasonable initial tilt trend. Based on this warm start, the TD3 scheduler is reformulated with an incremental action representation, so that the actor outputs a bounded nacelle-angle increment rather than an independent absolute nacelle command. The resulting command is passed through a feasibility layer that performs bounded integration, actuator-rate saturation, and projection onto a tightened speed-dependent transition corridor. In this way, ESIT combines feedback adaptability, rotor-speed-based effort awareness, actuator-compatible command smoothness, and corridor-feasible transition scheduling within a unified framework. After training, the learned actor is further analyzed as a bounded nonlinear scheduler to characterize sampled-command corridor feasibility and conditional closed-loop boundedness under bounded-disturbance and local low-level stability assumptions.
The main contributions of this paper are summarized as follows.
A corridor-aware feedback nacelle-scheduling formulation is developed for quad-tiltrotor level-flight transition. The formulation treats the nacelle angle as a state-dependent physical transition variable constrained by actuator-rate limits and a speed-dependent admissible corridor. It integrates forward-speed buildup, altitude regulation, rotor-effort redistribution, and longitudinal smoothness into a unified high-level scheduling problem.
A physics-structured incremental TD3 strategy is proposed for nacelle command generation. By learning the nacelle-angle increment rather than the absolute nacelle angle, the proposed scheduler converts the RL policy from an independent setpoint generator into a constrained nacelle-evolution generator. The resulting bounded integration, rate saturation, and tightened-corridor projection provide a direct mechanism for producing smooth, rate-compatible, and corridor-feasible tilt schedules.
A mechanism-oriented analysis and validation framework is established for the proposed scheduler. The feasibility layer is characterized as a rate-compatible and corridor-aware command-generation mechanism, and the trained actor is analyzed as a bounded nonlinear scheduler, leading to a conditional boundedness characterization under bounded disturbances and a common Lyapunov condition for the scheduled low-level error dynamics. Simulations, including action-parameterization ablation, corridor-feasibility evaluation, Monte Carlo tests, and observation-noise evaluation, demonstrate that the proposed incremental action structure reduces nacelle-rate demand and longitudinal jerk while maintaining corridor feasibility, bounded altitude response, terminal-speed buildup, and comparable rotor-speed-based effort.
The remainder of this paper is organized as follows. Section II formulates the constrained high-level nacelle tilt-scheduling problem for quad-tiltrotor level-flight transition. Section III presents the proposed ESIT framework, including the incremental TD3 scheduler, rate-and-corridor feasibility layer, reward design, BS-guided reference generation, TD3 optimization procedure, and post-training boundedness characterization. Section IV presents comparative simulations, ablation studies, corridor evaluation, Monte Carlo tests, and observation-noise evaluation. Section V concludes this paper.
Problem formulation
This section formulates the constrained high-level nacelle tilt-scheduling problem for the level-flight transition of a quad-tiltrotor UAV. The problem is considered under a scheduled low-level flight-control architecture. Therefore, the nacelle scheduler is treated as a high-level transition-command generator, rather than as a replacement for the stabilizing attitude and altitude controller.
Longitudinal transition model
A longitudinal force-balance model [31, 32] is introduced to describe the dominant transition dynamics. The nacelle tilt angle is denoted by , where corresponds to the rotor-borne configuration and corresponds to the fully wing-borne configuration.
Let the longitudinal state vector be defined aswhere denotes the longitudinal position, the forward velocity, the altitude, the vertical velocity, the pitch angle, and the pitch rate. The resultant longitudinal-plane airspeed used only in the aerodynamic lift and drag calculations is defined as
The longitudinal state and resultant airspeed are defined in Equations 1, 2. The low-level control input ul is defined in Equation 3.where is the rotational speed of the -th rotor and is the elevator deflection or an equivalent longitudinal control-surface input.
The dominant longitudinal force and moment balances are given by Equations 4 -9.where is the vehicle mass, is the pitch-axis moment of inertia, and are the longitudinal and vertical components of the rotor thrust components, and are the aerodynamic lift and drag, is the pitch moment, and are bounded disturbance terms.
The rotor thrust and its longitudinal and vertical components are modeled aswhere is the air density, is the rotor diameter, and is the rotor thrust coefficient.
The aerodynamic lift and drag are expressed aswhere is the dynamic pressure, is the wing reference area, and and are the lift and drag coefficients, respectively.
The pitch moment is written aswhere is the longitudinal moment arm from the -th rotor to the aircraft center of gravity, and denotes the aerodynamic pitching-moment term under the adopted simplified longitudinal model. The rotor thrust components, aerodynamic forces, and pitch moment are expressed in Equations 10-12.
The level-transition process is dominated by the coupling among rotor thrust vectoring, wing lift buildup, forward acceleration, and pitch stabilization. The longitudinal transition dynamics can be compactly written aswhere represents bounded longitudinal disturbances and unmodeled dynamics, including aerodynamic uncertainty, wind disturbance, and actuator-related perturbations. In this formulation, the low-level controller generates for altitude and attitude regulation, while the high-level scheduler determines the nacelle command . This separation allows different nacelle tilt schedules to be compared under the same stabilizing low-level control architecture.
Nacelle command evolution and transition corridor
Let denote the decision interval of the high-level scheduler and . The sampled nacelle command sequence is denoted by
For a forward level-transition task, the boundary condition is specified as
The sampled nacelle command sequence and boundary conditions are given by Equations 14, 15. The nacelle angle is a physical transition variable rather than an arbitrary independent command. Its evolution affects the distribution of rotor thrust between vertical support and forward propulsion. Therefore, an overly aggressive tilt command at low forward speed may reduce the vertical rotor-thrust component before sufficient wing lift has been established. Conversely, an overly conservative tilt command at high forward speed may delay transition completion and increase rotor-speed demand. These characteristics make the nacelle schedule a key high-level variable for level-transition performance.
During level-flight transition, the forward speed and nacelle angle must satisfy a feasible coupling relationship. This relationship is represented by a transition corridor in the plane:where and are the initial and target forward speeds, respectively. The functions and denote the lower and upper admissible nacelle-angle boundaries determined by aerodynamic lift capability, power margin, actuator limitations, and attitude-safety requirements.
The above formulation shows that the high-level nacelle scheduler must address two coupled requirements. First, the nacelle command should evolve as a physically realizable transition variable rather than as a sequence of independent absolute setpoints. Second, the generated command should remain compatible with the speed-dependent admissible corridor under the scheduled low-level flight-control architecture. Therefore, the problem considered in this paper is to design a feedback-capable high-level nacelle scheduler that maps measurable transition states to a smooth, actuator-compatible, and corridor-feasible nacelle command, without replacing the stabilizing low-level attitude and altitude controller.
Proposed corridor-constrained incremental TD3 tilt strategy
This section presents the proposed energy-aware smooth incremental tilt (ESIT) strategy for solving the high-level nacelle scheduling problem formulated in Section Problem formulation. The method is developed under the same scheduled low-level flight-control architecture, so that the learned policy modifies only the nacelle-command evolution rather than the stabilizing attitude and altitude controller.
The central idea of ESIT is to formulate nacelle tilting as a constrained command-evolution process. Instead of directly learning the absolute nacelle angle at each decision instant, the actor learns a bounded nacelle-angle increment. The executable command is then obtained through bounded integration, rate saturation, and projection onto a tightened speed-dependent transition corridor. This action-generation mechanism embeds nacelle-command continuity and actuator-rate compatibility into the policy structure, while the projection layer enforces hard transition-corridor constraint.
The proposed framework is organized around five main elements: the hierarchical transition-control architecture, the scheduling state and incremental action parameterization, the rate-and-corridor feasibility layer, the multi-objective reward design, and the BS-guided reference schedule used for benchmarking and actor warm-starting. The overall framework is shown in Figure 1.
FIGURE 1
The feasibility layer is applied after the actor output and before the low-level controller. Therefore, the actor searches within the performance landscape induced by the closed-loop transition dynamics, while the final command remains compatible with the physical rate and corridor constraints. The corridor penalty in the reward is evaluated before projection, which encourages the actor to learn feasible candidate commands rather than relying entirely on the projection layer. The projection layer nevertheless remains the hard feasibility mechanism. The BS warm-start is used only to improve the initial policy quality and training efficiency. The final learned policy is evaluated as a closed-loop feedback scheduler and is not forced to track the BS reference. This distinction is essential because the objective of ESIT is not to reproduce an offline trajectory, but to learn a state-dependent nacelle-evolution policy under physical transition constraints.
Hierarchical transition-control architecture
The proposed scheduler is implemented under a hierarchical transition-control architecture. The high-level scheduler determines the evolution of the nacelle tilt angle, whereas the low-level controller tracks altitude and attitude references and allocates actuator commands to rotors and aerodynamic control surfaces. The scheduled low-level control architecture is kept unchanged for all compared tilt schedules, so that the performance differences can be attributed mainly to the high-level nacelle scheduling mechanism.
At the decision instant , the high-level scheduler receives measurable closed-loop transition information and generates the next nacelle command. The final command is then sent to the low-level controller, which produces the low-level actuator input . The closed-loop information flow is summarized in Equation 17 as
This hierarchical separation is important for the scope of this work. The RL policy is not required to stabilize the fast attitude and altitude dynamics directly. Instead, it adapts the slow nacelle transition process according to the measured closed-loop scheduling vector. The stabilizing controller provides local attitude and altitude regulation for all tested schedules, while the proposed scheduler shapes the transition behavior through the nacelle command.
Remark 1The proposed framework evaluates nacelle tilt scheduling rather than low-level stabilization. Therefore, the same scheduled low-level control architecture, actuator model, and simulation environment are used for all compared schedules. This setting isolates the effect of action representation and feasibility enforcement at the high-level scheduling layer.
Scheduling state and incremental action parameterization
The high-level scheduler uses measurable closed-loop transition information to determine the nacelle command evolution. The physical scheduling vector is defined aswhere is the altitude tracking error, is the forward speed, is the total rotor-speed index, and is the current nacelle angle. The altitude error characterizes the vertical tracking margin, the forward speed represents the progress of wing-lift buildup and determines the corridor interval, the rotor-speed index provides an effort-related signal, and the nacelle state is included because the scheduling decision depends on the current nacelle configuration.
For neural-network training and policy evaluation, the physical scheduling vector is normalized aswhere
The physical and normalized scheduling vectors are defined in Equations 18-21. The normalization does not introduce a new physical scheduling state; it only maps the measurable transition variables into a numerically suitable input range for the actor and critic networks.
Rapid longitudinal acceleration is not used directly as a policy observation. This choice avoids amplifying measurement noise in the high-level scheduler. Instead, acceleration and jerk are evaluated in the reward function to guide the policy toward smooth transition behavior.
The action representation is the key difference between the proposed scheduler and an absolute-action RL scheduler. A conventional absolute-action policy directly maps the scheduling state to the complete nacelle command:
Although this representation is simple, it treats the nacelle command at each decision instant as an independent setpoint. Therefore, the physical evolution between adjacent commands is not encoded in the action itself. Exploration noise, critic approximation errors, or policy updates may consequently lead to abrupt inter-sample command variations.
In contrast, the proposed policy outputs a normalized incremental action
The action represents the intended change of the nacelle angle rather than the absolute nacelle command. It is then converted into a nacelle-angle increment and implemented through the discrete-time command-evolution law
Thus, the RL scheduler is converted from an absolute setpoint generator into a constrained nacelle-evolution generator. This distinction is fundamental: a smoothness penalty can discourage oscillatory actions, but it does not change the underlying action variable. By contrast, the incremental formulation directly modifies the command-generation structure itself, because the learned action is the nacelle-angle increment and the final command is produced by a constrained update.
The high-level scheduling policy is denoted by
After action scaling, rate saturation, and corridor projection, this policy induces a physical nacelle command sequence
The absolute-action and incremental-action parameterizations are given by Equations 22-26. A feasible scheduler should therefore generate a physically consistent nacelle evolution that respects both actuator-rate compatibility and the speed-dependent transition corridor, rather than a sequence of unrelated absolute setpoints.
Rate-and-corridor feasibility layer
The feasibility layer converts the normalized actor output into an executable nacelle command by enforcing two physical implementation requirements: nacelle-rate compatibility and speed-dependent corridor feasibility. This subsection describes the operational command-generation mechanism, whereas the post-training boundedness characterization later analyzes the resulting closed-loop scheduler under explicit assumptions.
The normalized incremental action is first scaled into a raw nacelle-angle increment:where is the high-level decision interval and is the prescribed maximum nacelle tilt rate. For the forward level-transition task from the rotor-borne configuration to the wing-borne configuration, a monotonic rate-limited increment is adopted:
Equivalently, the unidirectional saturation is expressed by Equation 29.
The unidirectional saturation prevents reverse tilting and ensures that the command evolves consistently with the prescribed forward transition direction.
The pre-projection nacelle command is generated by bounded integration:
This step embeds the inter-sample evolution of the nacelle command into the action-generation process.
For implementation, a tightened transition corridor is introduced aswhere is a corridor margin. The margin separates the commanded trajectory from the nominal corridor boundary and provides room for sampling-induced speed variation, modeling mismatch, and low-level tracking transients.
At the current forward speed , the tightened corridor induces the admissible nacelle-angle interval
For the monotonic forward transition considered in this paper, the rate-compatible interval is
The final admissible update interval is the intersection
When is nonempty, the final command is obtained by scalar projection:where denotes the Euclidean projection onto the closed interval .
Equivalently, ifwiththen the projection can be implemented as
The tightened corridor, admissible intervals, and projected update are defined in Equations 31-34, 36-39. Thus, the actor output is not directly applied to the plant. Instead, it is first interpreted as an intended nacelle-angle increment and then processed by rate saturation and tightened-corridor projection. This design separates reward-level performance optimization from hard feasibility enforcement: the TD3 actor searches for a desirable transition behavior, while the feasibility layer enforces actuator-rate compatibility and practical corridor admissibility.
The transition corridor used in the feasibility layer should be interpreted as a scheduling corridor for high-level command generation, rather than as an exact representation of the full physical flight envelope. In practical tiltrotor transition, the admissible boundary may be affected by stall limits, power constraints, rotor--wing interference, attitude-safety requirements, actuator limits, and control-allocation capability. As a result, the raw transition envelope obtained from trimming, aerodynamic analysis, or offline simulation may contain piecewise segments, corner points, nonsmooth portions, or locally conservative regions.
The margin is selected to compensate for one-step boundary movement and to avoid operating exactly on the nominal corridor boundary. Letbe an estimated gradient bound of the smoothed corridor boundaries over . In a tabulated implementation, can be estimated from finite differences of the stored corridor table as
If the forward-speed variation over one high-level sampling interval is bounded bythen a practical margin-selection rule is
A practical corridor-margin selection rule is given by Equations 40-43. The lower bound reserves angular margin for one-step corridor-boundary movement caused by forward-speed variation, whereas the upper bound prevents the tightened corridor from becoming empty or excessively conservative. The one-step speed bound can be estimated either from an acceleration bound as , or directly from the maximum one-step forward-speed variation observed in offline envelope-generation simulations.
If becomes empty in implementation, no command can simultaneously satisfy the tightened-corridor interval and the monotonic rate interval at that decision instant. This event is recorded as a feasibility-margin event, indicating that the corridor smoothing, the margin , the high-level sampling time , the nacelle-rate limit , or the nominal corridor table should be redesigned. In the simulations of this paper, these design elements are selected so that the admissible update interval remains nonempty along the evaluated transition trajectories.
This feasibility layer should be distinguished from a continuous-time safety proof for detailed nacelle-servo dynamics. It specifies how the sampled high-level command is generated to be rate-compatible and corridor-aware. Detailed servo lag, overshoot, actuator faults, and online corridor adaptation require additional actuator-level analysis or experimental validation, which are beyond the scope of the present high-level scheduling design.
Multi-objective reward design
The level-transition scheduler must balance several competing objectives. Increasing the nacelle angle promotes forward acceleration and transition completion, but it reduces the vertical component of rotor thrust and may increase altitude deviation. Conversely, delaying nacelle rotation preserves vertical thrust margin but may increase transition duration and rotor-speed demand. In addition, abrupt nacelle variations may induce excessive longitudinal acceleration and jerk, which are undesirable for actuator compatibility and transition smoothness.
The reward function is designed as a normalized maximization-oriented surrogate of the scheduling trade-off considered in this work. It combines mission progress, actuator-effort regulation, and handling-qualities penalties, while leaving hard rate and corridor admissibility to the feasibility layer. The altitude-error, rotor-speed, acceleration, jerk, and increment terms correspond to the penalty terms in (Equation 56) with opposite signs, whereas the forward-speed term corresponds to the negative speed term in the cost. Additional shaping terms are introduced to improve learning efficiency and to reduce unnecessary reliance on the projection layer.
The feasibility layer enforces the actuator-rate and corridor constraints independently of the reward. Therefore, the reward function primarily shapes transition performance, while hard admissibility is imposed by the command-generation mechanism in Subsection Rate-and-corridor feasibility layer.
As shown in the hierarchical decomposition, the instantaneous reward is explicitly divided into three categories: mission goal reward, actuator effort penalty, and flight safety and quality penalty, which correspond to transition completion & velocity promotion, actuator operation smoothness constraint, and flight state safety maintenance, respectively. This yields the unified structural expression , where denotes the mission-oriented reward item, represents the penalty for consumption actuator operating effort and fluctuation, and stands for the comprehensive penalty related to flight safety and dynamic quality.
At each decision instant, the categorized reward components are specifically defined in Equations 44-47.in which and are non-negative weighting coefficients. The mission goal reward encourages forward-speed buildup and accelerates the overall transition progress, where the mixed coupling term facilitates synchronous optimization of nacelle tilt angle and flight velocity. The actuator effort penalty regularizes the incremental nacelle command within the admissible rate interval (acting as a soft constraint rather than enforcing the hard rate limit, which is already handled by the feasibility layer), suppresses excessive rotor-speed demand, and guides the policy to generate candidate commands close to the feasible corridor region in advance. The safety and quality penalty mainly suppresses altitude tracking deviation, excessive longitudinal acceleration and severe jerk, so as to guarantee stable and comfortable transition flight.
The acceleration and jerk penalties are defined aswhere and are prescribed smoothness thresholds. These terms suppress aggressive longitudinal transients during the lift-redistribution process.
The corridor-related penalty is evaluated using the pre-projection command:where is the Euclidean distance to the tightened corridor interval. This term encourages the actor to generate commands that are naturally close to the admissible region before projection, thereby reducing reliance on the feasibility layer during training.
The terminal completion reward is given bywhere is the target transition-completion speed, is a small terminal nacelle-angle tolerance, and is the positive terminal-completion bonus. The acceleration, jerk, corridor, and terminal reward terms are given by Equations 48-51.
Remark 2The rotor-speed-related term in the reward function serves as an operation effort proxy. It reflects the variation tendency of rotor rotational speed demand during aerodynamic lift redistribution and transition scheduling, rather than establishing a high-fidelity calibrated model of electrical energy consumption, motor operating efficiency or battery power loss.
BS-guided reference schedule
A Beetle Swarm (BS) Optimization based reference schedule is used to provide a physically feasible benchmark and actor warm-start source. The BS optimization is performed offline over a finite-dimensional parameterization of the nacelle schedule. This reference is not imposed as a tracking command during final TD3 policy evaluation; it is used only to initialize the actor toward a feasible transition trend and to provide an interpretable non-learning baseline.
The nacelle schedule is parameterized by a normalized PCHIP interpolation curve. Letdenote the free interpolation-node vector. Together with the boundary conditionsthe parameter vector determines a smooth baseline nacelle profile .
The offline optimization problem is written aswhere is the feasible parameter domain and (55) is a multi-objective transition-performance index including altitude error, speed buildup, rotor-speed-related effort, acceleration, jerk, and terminal completion terms. The BS schedule parameterization and offline optimization are formulated in Equations 52-54.
For a finite transition horizon , the transition-performance index is defined in Equations 55-57.where is the stage cost and is the terminal cost.
are weighting coefficients. The terms in (56) penalize altitude error, rotor-speed-based effort, excessive longitudinal acceleration, jerk, and aggressive nacelle increments, while rewarding forward-speed buildup through the negative term. The functions and are nonnegative penalty functions used to suppress acceleration and jerk beyond desired smoothness levels. The acceleration and jerk are performance-evaluation variables rather than direct components of the scheduling state . . This term encourages terminal speed buildup, completion of nacelle transition, and bounded terminal altitude error.
The criterion in (Equation 55) specifies the desired scheduling trade-off. The offline BS criterion follows the same scheduling trade-off as the reward design, but is expressed as a deterministic finite-horizon cost for reference generation. Compared to the reinforcement-learning reward, which acts as its normalized maximization-oriented surrogate with additional shaping terms for training efficiency and feasibility, this criterion eliminates these unnecessary implementation-related terms.
The optimized reference is used in two ways. First, it serves as a deterministic benchmark in the simulation section. Second, it provides supervised samples for actor warm-starting. Specifically, state--action pairs are extracted along the BS schedule:where
The actor can then be initialized by minimizing
Supervised warm-start samples and the pretraining loss are given by Equations 58–60.
Remark 3The BS reference provides an initial feasible tilt trend, but it does not constrain the final learned TD3 policy to reproduce the BS trajectory. After warm-starting, the actor is further optimized through closed-loop interaction with the environment, and the final policy is selected according to the fixed evaluation criterion described in Section Simulation results and performance evaluation.
TD3-based policy optimization
The proposed scheduler is trained using the twin delayed deep deterministic policy gradient algorithm. TD3 is adopted because it is suitable for continuous-action control and mitigates the overestimation bias of deterministic actor--critic learning through clipped double critics, delayed actor updates, and target policy smoothing.
The actor network is denoted by , and the two critic networks are denoted by and . Their target networks are denoted by , , and . At each training step, a transition tupleis stored in the replay buffer .
For a mini-batch sampled from , the target action is computed as
The target value iswhere is the discount factor. Each critic is updated by minimizingwhere is the mini-batch size.
The actor is updated less frequently according to the deterministic policy gradient:
The target networks are updated by soft update:where is the soft-update coefficient.
During environment interaction, the exploratory action is
The sampled action is then passed through the same scaling, rate-saturation, and corridor-projection mechanism described in Subsection Rate-and-corridor feasibility layer. Therefore, both training and evaluation use the same physical command-generation structure.
The training procedure is summarized in Algorithm 1.
Algorithm 1
1: Initialize actor , critics , and target networks.
2: Initialize replay buffer .
3: Optionally warm-start using the BS reference dataset .
4: For each episode do
5: Reset the transition environment and obtain ; compute .
6: For do
7: Generate exploratory incremental action .
8: Scale to using (Equation 27).
9: Apply rate saturation using (Equation 28).
10: Generate using (Equation 30).
11: Project onto using (Equation 35).
12: Apply to the low-level controller and simulate the plant.
13: Obtain the next physical scheduling vector and compute .
14: Compute reward using (Equation 47).
15: Store in .
16: Update critics using (Equations 63, 64).
17: Delayed update of the actor using (Equation 65).
18: Soft-update target networks using (Equation 66).
19: End For
20: End For
The preceding subsections define the command-generation and training procedure of the proposed scheduler. Before characterizing the post-training boundedness property, the scheduled low-level closed-loop representation used in the analysis is introduced below.
Low-level controller formulation
The role of the high-level scheduler is therefore to generate a feasible nacelle-angle evolution while the scheduled low-level controller family maintains altitude and attitude tracking over the admissible transition corridor.
For each admissible nacelle command, the low-level controller generates to track the altitude and attitude references. Since the control effectiveness and aerodynamic characteristics vary significantly with the forward speed and nacelle angle, the low-level controller is represented as a scheduled controller family rather than a single fixed-parameter feedback law.
Let the scheduling variable be defined aswhere denotes the admissible transition corridor. In implementation, this admissible corridor is represented by a smoothed scheduling corridor constructed from the raw offline envelope, as detailed in Subsection Rate-and-corridor feasibility layer.
The low-level control law is written aswhere denotes the low-level reference signal, and the subscript indicates that the controller parameters are scheduled according to the current transition operating condition. This notation covers gain-scheduled, interpolated, or locally parameterized low-level controllers without requiring the high-level scheduler to redesign the stabilizing controller online.
Substituting (Equation 69) into (Equation 13) yields the scheduled closed-loop longitudinal dynamics by Equation 70.where
With this scheduled closed-loop representation, the trained actor can now be treated as a bounded nonlinear scheduler, and the local boundedness of the resulting tracking-error dynamics can be characterized.
Post-training boundedness characterization
The preceding subsections define how the actor output is converted into an executable nacelle command through incremental action scaling, rate saturation, bounded integration, and corridor projection. This subsection provides a post-training boundedness characterization of the resulting closed-loop scheduler. The purpose is not to prove TD3 convergence, global optimality, or global nonlinear stability. Instead, the trained actor is treated as a bounded nonlinear scheduler, and the local boundedness of the scheduled low-level tracking error is characterized under explicit assumptions.
The analysis is conducted for the implemented closed-loop system after training. Suppose that the feasibility layer and the corridor-margin design described in Subsection Rate-and-corridor feasibility layer generate an implemented nacelle schedule satisfyingover the considered level-transition interval. This condition means that the high-level scheduling trajectory remains inside the admissible speed–nacelle-angle corridor used for the scheduled low-level control design. Although the transition task is evaluated over a finite horizon, the boundedness result below is stated in the standard UUB form. It applies when the implemented schedule remains in during transition and then stays in an admissible terminal operating region after transition completion. Over the finite transition horizon, the same differential inequality directly provides an explicit local tracking-error bound. In the simulations, corridor membership is verified by evaluating the generated transition trajectories in the forward-speed–nacelle-angle plane.
The high-level scheduler determines the nacelle-angle evolution, whereas the scheduled low-level controller family stabilizes the attitude and altitude loops over the admissible transition corridor. Let the longitudinal tracking error be defined aswhere
For the level-transition task considered in this paper, the desired altitude is constant and the nominal desired pitch motion is selected as
From the scheduled closed-loop longitudinal dynamics (Equation 71), the corresponding tracking-error dynamics can be written locally aswhere is the error-system vector field induced by the scheduled low-level closed-loop dynamics. The nominal zero-error manifold is assumed to be locally consistent with the scheduled low-level controller, namely
The tracking-error state, reference conditions, local error dynamics, and zero-error consistency condition are defined in Equations 73 -77. Any residual trim mismatch or unmodeled bias is included in the bounded disturbance term . Around the nominal zero-error manifold, the scheduled tracking-error dynamics are represented by the local expansionwhere denotes the Jacobian matrix of the scheduled low-level closed-loop error dynamics, is the disturbance input matrix, and collects higher-order nonlinear residuals.
In (Equation 78), the nacelle angle affects the tracking-error dynamics through the scheduling variable and the scheduled closed-loop vector field . Therefore, the nacelle-rate bound established by the incremental command-generation mechanism is not introduced as an independent low-level input. Instead, rate-compatible nacelle evolution supports the validity of the local scheduled model by preventing abrupt movement among operating conditions.
Assumption 1(Uniform local stability of the scheduled low-level controller). For all , there exist symmetric positive definite matrices and , independent of , such that
Assumption 2(Bounded disturbance and local residual). The lumped longitudinal disturbance arising from aerodynamic uncertainty, external wind interference and unmodeled dynamics is bounded as
Herein, denotes the predefined upper amplitude bound of the overall lumped system disturbance. Moreover, is uniformly bounded over the admissible transition corridor:where represents the maximum spectral norm bound of the disturbance input matrix varying with scheduling states. The nonlinear residual satisfiesfor all and all inside a compact neighborhood
The common-Lyapunov and local bounding assumptions are stated in Equations 79 -83. In particular, is the quadratic Lipschitz constant used to quantify the growth rate of high-order unmodeled nonlinear terms, and stands for the maximum allowable radius of the local error convergence domain where the above nonlinear bounding condition holds.where . Suppose that the implemented schedule satisfies (Equation 72) during transition and remains in an admissible terminal operating region after completion. Suppose further that Assumptions 1 and 2 hold. Let
The eigenvalue bounds used in the theorem are defined in Equation 85.
Remark 4Assumption 1 is a uniform local stability condition for the scheduled low-level controller family over the considered transition corridor. It is not an assumption on the convergence or stability of the TD3 learning algorithm. The trained actor provides a bounded high-level scheduling command, while the scheduled low-level controller family is responsible for stabilizing the altitude and attitude tracking dynamics. The use of a common is conservative, but it avoids additional terms associated with parameter-dependent Lyapunov functions or switching among local Lyapunov functions. This condition is used only to obtain a local boundedness characterization of the scheduled low-level error dynamics.
(Conditional uniform ultimate boundedness). Consider the scheduled tracking-error dynamics.
Choose a local radius such that
Then, for any solution of (Equation 84) that remains inthe tracking error is uniformly ultimately bounded. Specifically, defineand
Then the Lyapunov functionsatisfies
Consequently,
Moreover, if there exists a constant satisfyingthen the sublevel setis positively invariant, and the bound (Equation 92) holds for all .
Proof. Choose the Lyapunov function candidate
Since , it satisfies
Taking the derivative of along (Equation 84) gives
Using Assumption 1, one obtains
For , condition (Equation 86) implies
Substituting (Equation 99) into (Equation 98) yields
By Young’s inequality,
Therefore,
Using , one obtains
From the upper quadratic bound in (Equation 96),
The intermediate Lyapunov inequalities are given by Equation 100 -105. Hence,
Solving the scalar differential inequality (Equation 106) giveswhich proves (Equation 91). Taking the limit superior on both sides givesone obtainswhich proves (92).
It remains to verify the sufficient local invariance condition. Suppose that (Equation 93) holds. On the boundary (Equation 106), gives
Therefore, the trajectory cannot leave the sublevel set . Since , every satisfies
The remaining proof steps are given by Equations 107 -112. Thus, , and the residual bound remains valid for all . This completes the proof.
Remark 5Theorem 1 establishes a local and conditional boundedness result for the scheduled closed-loop tracking-error system. The conclusion depends on three conditions: the implemented nacelle schedule remains inside the admissible transition corridor, the scheduled low-level controller family satisfies the common Lyapunov condition over the corridor, and the disturbance and nonlinear residual remain bounded in the local operating region. The result does not prove TD3 convergence, global stability, or global optimality of the learned scheduler.
Remark 6The nacelle-rate bound does not appear as an independent input term in (Equation 84) because the compact closed-loop model uses as a scheduling variable and represents the low-level controller as a -scheduled controller family. The role of the incremental action parameterization is to keep the evolution of inside the admissible corridor and compatible with the nacelle actuator rate. This prevents abrupt movement among scheduled operating conditions and supports the validity of the uniform local stability condition in Assumption 1.
In summary, the boundedness characterization complements the feasibility layer. The feasibility layer defines how the actor output is converted into a corridor-aware and rate-compatible sampled command, whereas Theorem 1 characterizes the local tracking-error behavior once the implemented schedule evolves inside the admissible corridor. This result supports the mechanism of ESIT without assigning stabilizing responsibility to the TD3 actor itself.
Simulation results and performance evaluation
This section evaluates the proposed ESIT scheduler using a nonlinear quad-tiltrotor UAV model in MATLAB/Simulink. Baseline schedules are first used to examine the sensitivity of transition response to nacelle timing, followed by TD3 action-parameterization ablation, corridor-feasibility verification, Monte Carlo smoothness evaluation, closed-loop performance comparison, and observation-noise assessment.
Simulation setup and evaluation protocol
The simulations were conducted in MATLAB/Simulink using the nonlinear quad-tiltrotor UAV model. All tilt-scheduling strategies were evaluated with the same scheduled low-level attitude/altitude control architecture, so that performance differences mainly reflect the high-level nacelle scheduling strategy.
The task is a forward level-flight transition from the rotor-borne configuration to the wing-borne configuration. The nacelle angle starts from and reaches at the terminal configuration. The height command is constant, and the desired attitude is set to zero during transition. The recorded variables include height error, forward velocity, nacelle angle, nacelle rate, longitudinal acceleration, longitudinal jerk, and rotor-speed-related quantities.
For deterministic policy evaluation, exploration noise was disabled. For the reinforcement-learning-based strategies, the saved agent with the highest training return was used as the evaluation policy. For TD3-incremental, the BS-generated reference was used only for actor warm-starting and did not constrain the final learned policy.
The transition corridor was constructed from offline envelope analysis in the forward-speed--nacelle-angle plane. One boundary was obtained from trim-based admissibility evaluation in Simulink, and the other was determined by the rotor-thrust limitation. The same corridor table was used for all safety-filtered policies and for corridor-feasibility verification.
The BS reference optimization and TD3 training used a horizon, whereas all final deterministic evaluations, baseline comparisons, TD3-variant comparisons, corridor evaluations, Monte Carlo tests, and observation-noise tests used a unified simulation duration. Unless otherwise stated, all reported performance indices correspond to this final evaluation horizon.
The main simulation settings are summarized in Table 1.
TABLE 1
| Item | Setting |
|---|---|
| Simulation platform | MATLAB/Simulink |
| Solver | ode4 |
| Step size | |
| Final evaluation duration | |
| Initial nacelle angle | |
| Terminal nacelle angle | |
| Target terminal forward speed | |
| Completion-speed threshold | |
| Settling time for completion | |
| Height command | Constant |
| Desired attitude | |
| Scheduled low-level control architecture | Identical for all strategies |
| Policy evaluation mode | Deterministic, without exploration noise |
| Agent selection rule | Best-return checkpoint |
| Monte Carlo runs | 50 |
| Randomized variables | Wind parameters, initial Euler angles, initial forward speed |
| Fairness condition | Same 50 scenarios reused for TD3 variants |
Simulation setup and evaluation protocol.
TD3 training was conducted in two stages: a critic burn-in stage with an almost frozen actor, followed by actor fine-tuning with a small learning rate. The main training parameters are listed in Table 2.
TABLE 2
| Parameter | Stage 1 | Stage 2 |
|---|---|---|
| Sampling time | ||
| Training horizon | ||
| Maximum episodes | 150 | 325 |
| Discount factor | 0.98 | 0.98 |
| Replay buffer length | Inherited | |
| Mini-batch size | 256 | 256 |
| Warm-start steps | 10,000 | 5,000 |
| Critic learning rate | ||
| Actor learning rate | ||
| Target-network soft-update factor | ||
| Policy update frequency | 4 | 2 |
| Target update frequency | 2 | 2 |
| Exploration noise std. | 0.008 | 0.005 |
| Minimum exploration std. | 0.002 | 0.001 |
| Exploration decay rate | ||
| Target policy noise std. | 0.010 | 0.008 |
| Target policy noise bound |
TD3 training parameters used in the implementation.
For each simulation, the following performance indices were calculated:where is the height tracking error, is the forward velocity, is the longitudinal acceleration, is the longitudinal jerk, and is the rotor-speed-based effort proxy. The performance indices used in the simulations are defined in Equations 113 -115.
In addition, , , and were recorded.
The transition completion time is defined as the first time when reaches and remains above this threshold for at least . This avoids declaring completion from a transient velocity crossing.
Baseline tilt schedules and low-level tracking boundedness
Representative non-learning tilt schedules are first evaluated under the same scheduled low-level control architecture. This comparison serves two purposes: it shows the sensitivity of level-flight transition to nacelle timing, and it verifies that the shared low-level controller maintains bounded pitch and altitude responses under feasible nacelle schedules.
The evaluated schedules include a linear schedule, a slow–fast diagnostic schedule, a fast–slow diagnostic schedule, and a BS-optimized reference schedule. The predefined schedules were generated by normalized PCHIP interpolation with five free nodes between and , where . The command was held at the terminal value after . The BS-optimized reference was generated using the optimization horizon, and all schedules were evaluated under the unified deterministic horizon.
The normalized free-node vectors are as followed.
The linear baseline used
The slow--fast diagnostic schedule used
The fast--slow diagnostic schedule used
The BS-optimized baseline used
Figure 2 shows the baseline tilt schedules and the corresponding closed-loop responses. The fast--slow schedule advances most of the nacelle rotation to the early transition stage, whereas the slow--fast schedule delays the main tilting process. These different timing patterns produce different altitude-error, forward-velocity, and total-rotor-speed responses. The result confirms that the nacelle schedule is an active high-level variable that affects altitude regulation, speed buildup, and rotor-effort redistribution during level-flight transition.
FIGURE 2
The low-level pitch and height tracking bounds under different schedules are summarized in Table 3. The maximum absolute pitch angle remains below for all evaluated strategies. For the TD3-based strategies, the maximum absolute pitch angle is below , and the RMS pitch angle remains small. These results indicate that the shared low-level controller remains well behaved during the evaluated transitions.
TABLE 3
| Strategy | Max [deg] | RMS [deg] | Max [m] |
|---|---|---|---|
| Linear baseline | 1.0981 | 0.3378 | 0.3908 |
| Optimized baseline | 0.8753 | 0.2633 | 0.3479 |
| Slow–fast baseline | 0.6977 | 0.3925 | 0.1793 |
| Fast–slow baseline | 0.6023 | 0.2456 | 0.3036 |
| TD3 | 0.4526 | 0.1434 | 1.1480 |
| TD3 + penalty | 0.4526 | 0.1506 | 1.4060 |
| TD3 + penalty + noise | 0.4526 | 0.1422 | 1.0349 |
| TD3-incremental | 0.4526 | 0.1303 | 0.7738 |
| TD3-incremental + noise | 0.4526 | 0.1314 | 0.8198 |
Low-level tracking boundedness under different tilt schedules.
Table 3 is not intended to rank all schedules by tracking accuracy or to prove global nonlinear stability. It only verifies that the subsequent TD3 comparisons are not caused by low-level instability. The repeated maximum pitch value observed for several TD3-based cases mainly reflects a common low-level transient under the shared stabilization loop; the corresponding height-error bound, nacelle-rate demand, jerk response, and corridor trajectory still differ across high-level schedulers.
Action-parameterization ablation of TD3 tilt scheduling
This subsection isolates the effect of action parameterization on TD3-based nacelle scheduling. Three policies are compared under the same observation variables, scheduled low-level control architecture, reward components, and deterministic evaluation protocol: absolute-action TD3, penalty-based absolute-action TD3, and the proposed incremental-action TD3.
The absolute-action policy directly outputs the complete nacelle angle, so adjacent commands are not structurally coupled. The penalty-based policy retains this absolute-action representation but adds an oscillation penalty. In contrast, the proposed policy outputs a nacelle-angle increment before command filtering. Therefore, the comparison separates action-level reparameterization from reward-level smoothness shaping.
Figure 3 compares the nacelle-angle responses of the three TD3 variants. All policies reach the terminal nacelle configuration, but their transient command evolution is different. The absolute-action TD3 policy exhibits local oscillations during the middle-to-late transition stage, especially when the aircraft shifts from rotor-dominant support to wing-lift-supported flight. TD3 + penalty attenuates this fluctuation, but local irregularities remain. By contrast, TD3-incremental produces a smoother and nearly monotonic nacelle-angle trajectory.
FIGURE 3
The nacelle-rate responses in Figure 4 further highlight the difference among the action representations. Absolute-action TD3 generates large and rapidly alternating rate peaks. TD3 + penalty reduces the peak level but does not remove local rate fluctuations. TD3-incremental keeps the nacelle-rate response within a narrow range throughout the transition, indicating better compatibility with rate-limited nacelle actuation.
FIGURE 4
These results indicate that the oscillatory behavior of absolute-action TD3 is mainly associated with the mismatch between independent absolute-angle commands and the physical continuity requirement of nacelle rotation. Reward-level oscillation penalties can mitigate this behavior, but the improvement remains tuning-dependent. The incremental action representation modifies the command structure itself and therefore provides a more direct mechanism for smooth nacelle evolution.
Corridor feasibility evaluation
Command smoothness alone does not guarantee transition feasibility. A nacelle trajectory may be smooth but still leave the admissible speed--nacelle-angle corridor. Conversely, a corridor-feasible trajectory may still impose undesirable nacelle-rate demand. Therefore, the generated schedules are examined in the forward-speed--nacelle-angle plane.
Figure 5 shows the transition trajectories of the predefined baselines, the BS-optimized baseline, and the TD3-based schedulers. To match the plotting convention of the generated corridor map, the vertical axis is displayed as the converted nacelle angle . This conversion is used only for visualization and does not change the feasibility criterion defined in (Equation 16). All compared strategies remain inside the prescribed transition corridor during the evaluated level-flight transition.
FIGURE 5
The corridor trajectories show that feasibility and smoothness are distinct properties. The absolute-action TD3 trajectory satisfies the corridor constraint but still contains local irregularities, consistent with the oscillatory nacelle response in Figure 3. The penalty-based variant reduces these irregularities, whereas TD3-incremental produces the smoothest feasible path among the TD3-based schedulers. Thus, the smoothness advantage of the proposed scheduler is achieved within the same admissible speed--nacelle-angle region, rather than by relaxing or leaving the prescribed transition corridor.
Monte Carlo smoothness evaluation
Monte Carlo simulations are conducted to examine the consistency of the smoothness improvement under randomized transition conditions. Each TD3-based strategy was evaluated over 50 scenarios, and the same random scenarios were reused for absolute-action TD3, TD3 + penalty, and TD3-incremental. The randomized quantities include wind amplitude, frequency, phase, initial Euler angles, and initial forward speed.
The evaluation focuses on nacelle-rate demand, longitudinal acceleration, longitudinal jerk, altitude response, terminal speed, and rotor-speed-based effort. Figure 6 shows the smoothness statistics, and Table 4 summarizes the numerical results. Absolute-action TD3 exhibits the largest nacelle-rate demand and dispersion. TD3 + penalty reduces the mean oscillation level, but its maximum rate and jerk remain more sensitive to scenario variations. TD3-incremental yields the smallest nacelle-rate demand and the most consistent smoothness metrics among the three TD3 variants.
FIGURE 6
TABLE 4
| Metric | TD3 | TD3 + penalty | TD3-incremental |
|---|---|---|---|
| [deg/s] | |||
| [deg/s] | |||
| [m/s ] | |||
| [m/s ] | |||
| [m] | |||
| [m/s] | |||
Monte Carlo smoothness and transition-performance statistics.
Boldface denotes the best-performing value among the three TD3 variants for each metric.
Compared with absolute-action TD3, TD3-incremental reduces from to , and reduces from to . Compared with TD3 + penalty, it further reduces the maximum nacelle-rate demand from to . The standard deviations of both nacelle-rate metrics are also smaller for TD3-incremental, indicating that the smoother command behavior is less sensitive to the tested random variations.
The same trend appears in the longitudinal smoothness metrics. TD3-incremental reduces from for absolute-action TD3 and for TD3 + penalty to . Meanwhile, all three TD3 variants maintain terminal forward speeds close to the target value. Thus, the smoother response of TD3-incremental is not obtained by failing to complete the transition, but by changing the nacelle command evolution mechanism.
These Monte Carlo results do not constitute a global robustness guarantee. They show that, under the tested randomized scenarios, the proposed incremental action structure provides more consistent nacelle-rate and jerk reduction than absolute-action TD3 and reward-penalized absolute-action TD3.
Closed-loop transition performance and observation-noise evaluation
This subsection compares the overall closed-loop transition performance of representative strategies and further evaluates the sensitivity of TD3-incremental to observation noise.
Closed-loop transition performance
The compared strategies include the linear baseline, the BS-optimized baseline, TD3 + penalty, and TD3-incremental. They represent a simple predefined schedule, an offline optimized schedule, a reward-smoothed absolute-action RL policy, and the proposed incremental-action RL scheduler, respectively.
Figure 7 shows the height-error, forward-velocity, longitudinal-acceleration, and total-rotor-speed responses. All compared strategies complete the level-flight transition without low-level instability. The linear baseline provides a simple reference response, whereas the BS-optimized baseline achieves a strong nominal profile through offline optimization. TD3 + penalty improves the smoothness of absolute-action RL, but still inherits the limitations of absolute command generation. TD3-incremental produces a smoother nacelle-driven transition while maintaining bounded height response and terminal forward-speed buildup.
FIGURE 7
The quantitative indices are summarized in Table 5. TD3-incremental reduces the RMS jerk from for TD3 + penalty to , and lowers the maximum height error from to . Its terminal forward speed remains close to the target value, indicating that the transition is completed. The main trade-off is a longer completion time, which reflects the more conservative nacelle evolution induced by the incremental action structure.
TABLE 5
| Strategy | RMSE [m] | Max [m] | [m/s] | RMS [m/s ] | RMS [m/s ] | Completion time [s] | |
|---|---|---|---|---|---|---|---|
| Linear baseline | 0.1682 | 0.3908 | 24.165 | 1.1168 | 2.0975 | 26.90 | |
| Optimized baseline | 0.1363 | 0.3479 | 24.344 | 0.9637 | 2.4089 | 22.35 | |
| TD3 + penalty | 0.5425 | 1.4060 | 24.173 | 1.0729 | 0.3150 | 23.10 | |
| TD3-incremental | 0.3881 | 0.7738 | 24.052 | 0.9685 | 0.1774 | 28.35 |
Closed-loop transition performance.
These results show that TD3-incremental should not be interpreted as minimizing every scalar metric independently. Its advantage lies in producing a rate-compatible and low-jerk nacelle command while maintaining bounded altitude response, corridor feasibility, terminal speed buildup, and comparable rotor-speed-based effort. The BS-optimized baseline remains a high-quality offline reference, but it is not a feedback policy. In contrast, TD3-incremental generates the nacelle increment from measured transition states and therefore provides a feedback-capable implementation of corridor-aware nacelle scheduling.
Observation-noise evaluation
Observation noise is injected into the policy inputs to examine whether perturbed high-level measurements induce undesirable nacelle-command oscillations. The noise is added to altitude error, forward speed, and total rotor-speed index before normalization, while the low-level control architecture and evaluation protocol are kept unchanged. This test evaluates input-side sensitivity of the learned scheduler rather than observer or sensor-fusion design.
Figure 8 compares the noise-free and noisy responses of TD3-incremental. The observation noise is modeled as zero-mean Gaussian perturbations with variances of for , for , and for . The generated nacelle command remains smooth under noisy inputs. The height-error response remains bounded, the forward-speed buildup is preserved, and the total rotor-speed response remains comparable to the noise-free case.
FIGURE 8
The bounded response under observation noise is consistent with the incremental command-generation mechanism. Since the actor output is interpreted as a bounded increment and then processed by rate saturation and corridor projection, measurement perturbations cannot directly produce arbitrarily large changes in the executable nacelle command. The low-level tracking statistics in Table 3 further show that the maximum height error increases only moderately under observation noise, while the RMS pitch angle remains almost unchanged.
Discussion
The simulation results support the main mechanism of the proposed ESIT framework. The baseline comparison confirms that nacelle timing is a dominant high-level factor affecting altitude response, speed buildup, rotor-speed demand, and longitudinal smoothness. The TD3 ablation shows that the incremental action representation provides a structural mechanism for smooth command generation beyond reward-level oscillation penalties. The corridor evaluation verifies that this smoothness improvement is achieved inside the prescribed speed--nacelle-angle corridor. The Monte Carlo and observation-noise evaluations further show that the smoothness advantage remains consistent under the tested randomized scenarios and perturbed policy inputs.
These findings are consistent with the command-generation mechanism developed in Section Proposed corridor-constrained incremental TD3 tilt strategy. By learning a bounded nacelle-angle increment instead of an absolute nacelle command, the scheduler limits inter-sample command variation before the command reaches the low-level controller. The feasibility layer then enforces rate compatibility and sampled-command corridor admissibility. As a result, the learned scheduler improves nacelle-command smoothness without requiring changes to the scheduled low-level attitude/altitude control architecture.
The results should nevertheless be interpreted within the simulation scope of this paper. The rotor-speed-based index is used as an effort proxy rather than as a calibrated propulsion-power or battery-energy model. The transition corridor is constructed from the considered vehicle model and offline envelope analysis, so the feasibility conclusion depends on the fidelity of the corridor table. In addition, the Monte Carlo and observation-noise tests evaluate representative variations in wind parameters, initial conditions, and policy inputs, but they do not constitute a global robustness guarantee.
Conclusion
This paper investigated high-level nacelle tilt scheduling for level-flight transition of a quad-tiltrotor UAV. The main conclusion is that the smoothness and feasibility of reinforcement-learning-based transition scheduling depend not only on reward design, but also on the physical structure of the action representation. By treating the nacelle angle as a rate-limited transition variable rather than as an independently generated absolute setpoint, the proposed ESIT framework reformulated the TD3 action as a bounded nacelle-angle increment. The executable command was then generated through bounded integration, rate saturation, and projection onto a tightened speed-dependent transition corridor.
The proposed framework provides a mechanism for combining feedback learning with physical transition constraints. The incremental action representation embeds nacelle-command continuity and actuator-rate compatibility into the command-generation process, while the feasibility layer enforces sampled-command corridor admissibility. A BS-optimized reference schedule was used to provide a physically feasible warm-start and a non-learning benchmark, rather than as a globally optimal trajectory. After training, the learned actor was interpreted as a bounded nonlinear scheduler, and the post-training analysis characterized corridor feasibility and conditional closed-loop boundedness under bounded disturbances and local scheduled low-level stability assumptions. These results clarify the implemented scheduler’s feasibility and boundedness properties without relying on a convergence or global-optimality claim for the TD3 learning process.
The simulation results support the proposed mechanism. Compared with absolute-action TD3 and penalty-based absolute-action TD3, the TD3-incremental scheduler generated smoother nacelle trajectories, reduced nacelle-rate demand and longitudinal jerk, and maintained bounded altitude response, terminal forward-speed buildup, and comparable rotor-speed-based effort. The corridor evaluation verified that the smoothness improvement was obtained inside the admissible speed--nacelle-angle region. The Monte Carlo and observation-noise tests further showed that the generated nacelle commands remained smooth under the tested randomized transition scenarios and perturbed policy inputs. These results indicate that physically structured action parameterization can improve the implementation compatibility of learning-based nacelle scheduling.
The proposed framework is useful for transition tasks in which a nominal corridor or reference schedule is available, but the final nacelle command must remain feedback-capable, smooth, rate-compatible, and corridor-feasible. Its current scope is nevertheless limited to monotonic forward level-flight transition under the considered vehicle model, corridor table, and scheduled low-level control architecture. In addition, the energy-related objective is represented by a rotor-speed-based effort proxy rather than by a calibrated propulsion-power or battery-energy model. Future work will therefore focus on incorporating calibrated propulsion-energy models, updating the transition corridor using higher-fidelity aerodynamic or experimental data, and extending the incremental action-generation framework to reconversion, three-dimensional transition, and coupled tilt-scheduling/control-allocation problems.
Statements
Data availability statement
The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.
Author contributions
RY contributed to the conceptualization, methodology, investigation, formal analysis, and writing of the original draft. JY, TF, and ZD contributed to investigation, validation, data curation, and visualization. CD provided overall supervision, project administration, resources, and guidance for the research group. All authors contributed to the article and approved the submitted version.
Funding
The author(s) declared that financial support was received for this work and/or its publication. This work was supported by an internal research project of the authors’ research group, under the project title “Frontier Exploratory Research on Tilt-Rotor Unmanned Aerial Vehicles”.
Acknowledgments
The authors would like to acknowledge the support provided by their research group throughout the development of this work.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
References
1.
XiYWangYSunSLuY. A review of research on fight control and transition strategies fortilt-rotor unmanned aerial vehicles. In: 2025 2nd International Conference on Electrical Technology Andautomation Engineering (ETAE) (IEEE) (2025). p. 382–90.
2.
SuJHuangHZhangHWangYWangFY. Evtol performance analysis: a review from control perspectives. IEEE Trans Intell Vehicles (2024) 9:4877–89. 10.1109/tiv.2024.3387405
3.
Hernandez-GarciaRGRodriguez-CortesH. Transition flight control of a cyclic tiltrotor uav based on the gain-scheduling strategy. In: 2015 International Conference on Unmanned Aircraft Systems (ICUAS). IEEE (2015). p. 951–6.
4.
WuZLiCCaoY. Numerical simulation of rotor-wing transient interaction for a tiltrotor in the transition mode. Mathematics (2019) 7:116. 10.3390/math7020116
5.
LiuNWangYZhaoJChengXHeXWuK. A new inclining flight mode and oblique takeoff technique for a tiltrotor aircraft. Aerospace Sci Technology (2023) 139:108370. 10.1016/j.ast.2023.108370
6.
YuksekBVuruskanAOzdemirUYukselenMInalhanG. Transition flight modeling of a fixed-wing vtol uav. J Intell and Robotic Syst (2016) 84:83–105. 10.1007/s10846-015-0325-9
7.
ZhaoGCuiZZhaoQChenXLiP. Investigations on trimming strategy and unsteady aerodynamic characteristics of tiltrotor in conversion procedure. Aerospace (2024) 11:632. 10.3390/aerospace11080632
8.
FloresGLugoILozanoR. 6-dof hovering controller design of the quad tiltrotor aircraft: simulations and experiments. In: 53rd IEEE Conference on Decision and Control (IEEE) (2014). p. 6123–8.
9.
MaTWangXFuJHaoSXueP. Three-dimensional flight envelope for V/stol aircraft with multiple flight modes. Aerospace (2022) 9:691. 10.3390/aerospace9110691
10.
AppletonW. AFilipponeBojdoN.“Interaction effects on the conversion corridor of tiltrotor aircraft.” The Aeronautical Journal. 125.1294 (2021):2065–2086. 10.1017/aer.2021.33
11.
FanYWangXHuZZhangK. Nonlinear modeling and transition corridor calculation of a tiltrotor without cyclic pitch. MATEC Web Conf., MATEC Web of Conferences (Les Ulis, France: EDP Sciences) (2022) 355, 01022. 10.1051/matecconf/202235501022
12.
ZhangZTianJChenHShengS. Optimal transition trajectory search for tilt-rotor aircraft based on ant colony algorithm. In: 2023 2nd International Conference on Artificial Intelligence and Computer Information Technology (AICIT) (IEEE) (2023). p. 1–8.
13.
XueLZhangJZhangNChenQ. Optimization Method for Tilting Transition Trajectory of Tiltrotor International Conference on Guidance, Navigation and Control. Springer (2024). p. 465–73.
14.
ZhaoHWangBShenYZhangYLiNGaoZ. Development of multimode flight transition strategy for tilt-rotor vtol uavs. Drones (2023) 7:580. 10.3390/drones7090580
15.
EskandarpourAMehrandezhMGuptaKRamirez-SerranoASoltanshahM. A constrained robust switching mpc structure for tilt-rotor uav trajectory tracking problem. Nonlinear Dyn (2023) 111:17247–75. 10.1007/s11071-023-08787-y
16.
KangNLuLWhidborneJ. Energy optimization strategies for automatic tiltrotor electric vertical takeoff and landing aircraft. J Guidance, Control Dyn (2025) 48:1196–200. 10.2514/1.g008466
17.
BohnECoatesEMMoeSJohansenTA. Deep reinforcement learning attitude control of fixed-wing uavs using proximal policy optimization. In: 2019 International Conference on Unmanned Aircraft Systems (ICUAS). IEEE (2019). p. 523–33.
18.
RomeroAAljalboutESongYScaramuzzaD. Actor-critic model predictive control: differentiable optimization meets reinforcement learning for agile flight. IEEE Trans Robotics (2025) 42:673–92. 10.1109/tro.2025.3644945
19.
ElikerKLeTNBessaadN. Reinforcement learning-based optimal pitch-angle generation for energy-efficient tilt-rotor quadrotors. Control Eng Pract (2026) 171:106853. 10.1016/j.conengprac.2026.106853
20.
FujimotoSHoofHMegerD. Addressing function approximation error in actor-critic methods. In: International Conference on Machine Learning (PMLR) (2018). p. 1587–96.
21.
HausknechtMStoneP. Deep reinforcement learning in parameterized action space. arXiv Preprint arXiv:1511.04143 (2015).
22.
KanervistoASchellerCHautamakiV. Action space shaping in deep reinforcement learning. In: 2020 IEEE Conference on Games (CoG) (IEEE) (2020). p. 479–86.
23.
ZhuJWuFZhaoJ. An overview of the action space for deep reinforcement learning. In: Proceedings of the 2021 4th International Conference on Algorithms, Computing and Artificial Intelligence (2021). p. 1–10.
24.
EßerJMargolisGBUrbannOKernerSAgrawalP. Action space design in reinforcement learning for robot motor skills. In: 8th Annual Conference on Robot Learning (2024).
25.
ChisariELinigerARupenyanAVan GoolLLygerosJ. Learning from simulation, racing in reality. In: 2021 IEEE International Conference on Robotics and Automation (ICRA) (IEEE) (2021). p. 8046–52.
26.
ChenZHeXWangYJLiaoQZeYLiZet alLearning smooth humanoid locomotion through Lipschitz-constrained policies. In: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE (2025). p. 4743–50.
27.
WabersichKPZeilingerMN. Safe exploration of nonlinear dynamical systems: a predictive safety filter for reinforcement learning. arXiv preprint arXiv:1812.05506. (2018)
28.
ZanonMGrosS. Safe reinforcement learning using robust MPC. IEEE Transactions on Automatic Control (2020) 66.8: 3638-3652.
29.
PhamTHDe MagistrisGTachibanaR. Optlayer-practical constrained optimization for deep reinforcement learning in the real world. In: 2018 IEEE International Conference on Robotics and Automation (ICRA) (IEEE) (2018). p. 6236–43.
30.
LinSWangHChenZKanZ. Projection-based fast and safe policy optimization for reinforcement learning. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE (2024). p. 7426–32.
31.
PapachristosCAlexisKTzesA. Model predictive hovering-translation control of an unmanned tri-tiltrotor. In: 2013 IEEE International Conference on Robotics and Automation (EEE) (2013). p. 5425–32.
32.
PapachristosCAlexisKTzesA. Hybrid model predictive flight mode conversion control of unmanned quad-tiltrotors. In: 2013 European Control Conference (ECC). IEEE (2013). p. 1793–8.
Summary
Keywords
nacelle tilt scheduling, quad-tiltrotor UAV, reinforcement learning, transition corridor, transition trajectory optimization
Citation
Yang R, Du C, Yu J, Fang T and Du Z (2026) Corridor-constrained incremental TD3 for nacelle tilt scheduling in quad-tiltrotor UAV level transition. Aerosp. Res. Commun. 4:17000. doi: 10.3389/arc.2026.17000
Received
24 May 2026
Revised
02 June 2026
Accepted
24 July 2026
Published
17 August 2026
Volume
4 - 2026
Updates
Copyright
© 2026 Yang, Du, Yu, Fang and Du.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Changping Du, duchangping@zju.edu.cn
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.