Abstract
Quadruped locomotion requires coordinating body motion, contact forces, foothold placement, and gait timing. We present a hierarchical controller in which reinforcement learning adapts these structured locomotion references while a gait scheduler, reactive footstep planner, and whole-body impulse controller retain responsibility for high-frequency execution. The model-based backbone can sustain a nominal trot independently, allowing the policy to focus on when and how to modify an already functional locomotion controller. We train the complete hierarchy in a vectorized controller-in-the-loop environment using a six-direction disturbance curriculum. Under a common simulator and evaluation protocol, the proposed controller achieves 77.29% mean survival across the tested disturbance grid, compared with 50.35% for the model-based reference controller. Training-time channel ablations identify adaptive foothold placement as the largest marginal contributor to grid survival and reveal direction-dependent effects of gait-rate adaptation.
- Grid survival
- 77.29%
- ± 1.15 over three seeds
- MPC+WBIC reference
- 50.35%
- same simulator and protocol
- Full survival up to
- 2.0 m/s
- every seed, every direction
- Largest ablation effect
- −10.83 pp
- without foothold residuals
Method
Model-based quadruped controllers separate gait scheduling, footstep planning, force generation, and whole-body control. This gives reliable, interpretable execution, but the behaviour depends on how the intermediate references are produced: gait cadence is usually fixed and footholds come from engineered rules. Joint-space reinforcement learning avoids hand-designed references, but the policy must learn joint coordination from scratch and its outputs do not enforce dynamic or contact feasibility.
We keep the model-based stack and replace the reference generator. One policy has supervisory authority over several coupled levels of a controller that can already trot on its own. Learning can then concentrate on adapting nominal locomotion.
Policy interface
The policy runs at 100 Hz; each action is held for five 2 ms updates of the lower layers. It reads a 54-dimensional observation of the velocity command, body velocity and orientation, joint states, scheduled contacts, foot positions, swing phase, and body height. It uses no history buffer and no privileged disturbance information. The 25-dimensional action has five groups:
| Group | Dimensions | Range |
|---|---|---|
| Body-velocity and yaw-rate residuals | 3 | ±0.5 m/s, ±0.3 rad/s |
| Stance-force targets, per leg (fx, fy, fz) | 12 | ±50, ±50, 0–250 N |
| Foothold offsets, per leg (x, y) | 8 | ±0.5 m |
| Body-height residual | 1 | ±0.05 m |
| Gait-phase increment | 1 | 0–0.01 per update |
Structured controller
Gait scheduler. A normalized phase defines a diagonal trot. The policy's phase increment can speed up or pause the gait cycle without changing the contact order; a neutral action gives a 0.4 s cycle.
Footstep planner. Footholds follow the Raibert-style rule of Kim et al., with a bounded velocity-feedback correction. The policy's foothold residual is added after that bound, so it can move the touchdown point beyond the heuristic's range. Bézier swing trajectories with 0.10 m clearance connect lift-off to touchdown.
Whole-body impulse control. WBIC adjusts the nominal accelerations and the policy's desired forces subject to floating-base dynamics, unilateral contact, and a friction pyramid (µ = 0.45), so the applied forces can differ from the policy's targets. Inverse-dynamics feedforward and joint PD feedback produce the torque command, with actuator limits enforced.
The reward is a sum of bounded task-space terms of the form wi exp(−pi), covering forward-velocity tracking, height, orientation, straightness, smoothness, and phase progression. A fall ends the episode with reward −10.
Training and evaluation
Controller-in-the-loop training
We built a C++ EnvPool environment in which every instance runs the complete closed loop: MuJoCo simulates the robot, and the embedded locomotion stack performs state processing, gait and foothold generation, whole-body control, and joint-command execution. EnvPool parallelizes full copies of this loop, so the policy trains against the real execution path instead of a reduced model.
Three policies were trained from scratch with PPO using different seeds, for 85.1 million transitions each, with 256 parallel environments. Training disturbances are sudden velocity changes in six directions; their ceiling rises from 1 to 2 to 3 m/s at 8.1 and 24.3 million transitions. All three runs completed without collapse, with final timeout fractions of 92.9–98.4%.
Evaluation protocol
All experiments use a 15 kg Unitree Go2 model on flat ground with a 2 ms time step and observation noise on every measured signal. Each disturbance is an instantaneous change in body velocity, following Rudin et al. The grid covers six directions (forward, backward, left, right, up, down) and ten magnitudes from 0.5 to 5.0 m/s, with 100 episodes per cell and policy. An episode survives if the body does not touch the ground and roll and pitch stay within 0.8 and 1.0 rad for 5.5 s after the disturbance.
Two reference controllers are reimplemented in the same MuJoCo and EnvPool framework and share the plant, physics, actuator limits, noise, disturbance generator, and evaluation code:
- MPC+WBIC, the Mini Cheetah architecture of Kim et al.: convex MPC force planning, reactive foot placement, and WBIC. One fixed controller.
- Joint-space RL, a Go2 recipe adapted from Rudin et al. and implemented with rsl_rl: a 50 Hz policy that outputs 12 joint-position offsets. Three policies, 196.6 million transitions each.
Each keeps its own observations, actions, objectives, and update rate, so the comparison is between complete configured systems.
Results
Every controller completed every undisturbed episode, so the differences below reflect disturbance recovery. The proposed controller survived every episode in every direction up to 2.0 m/s for all three seeds.
| Controller | Grid survival (%) |
|---|---|
| MPC+WBIC | 50.35 |
| Joint-space RL (source recipe) | 54.01 ± 5.55 |
| Proposed | 77.29 ± 1.15 |
The largest gains over MPC+WBIC are lateral: c50 rises from 2.26 to 4.19 m/s to the left and from 1.96 to 3.80 m/s to the right. Backward survival is still 70.3% at 5.0 m/s, beyond the tested range. Seed variation is concentrated near the edge between survival and failure; the largest per-cell spread is 14.7 percentage points at 4.0 m/s leftward, while the per-seed thresholds there span only 3.95–4.26 m/s.
The joint-space result describes that configured policy. A joint-space policy trained specifically for this disturbance protocol could perform differently, so the comparison does not isolate the effect of the action representation.
Action-channel ablations
To measure what each action group contributes after the controller has adapted to its absence, we trained five more policies from scratch, each with one group held at its neutral value throughout training and evaluation. The other groups can compensate during training.
- Footholds matter most. Without foothold residuals, grid survival falls from 77.43% to 66.60%, with the largest lateral losses (−1.19 m/s left, −0.93 m/s right). Downward tolerance improves by 0.31 m/s, so the effect depends on direction.
- Cadence shifts robustness between directions. Fixing the gait cadence lowers the pooled score by only 1.12 points but reduces upward c50 by 0.62 m/s, and it is the only ablation that reduces backward c50.
- Smaller effects change sign. The body-velocity, force, and height ablations cost 2.37, 1.50, and 0.57 points, with gains in some directions and losses in others. This points to compensation between channels.
The fixed-cadence controller uses less mechanical power with similar pooled survival, so cadence adaptation adds directional flexibility in these tests without a demonstrated energy benefit. The no-foothold controller has the lowest torque RMS together with the largest survival loss, which suggests the extra actuation from learned foothold correction does useful work.
Limitations
All evidence comes from flat-ground simulation with noisy state. Survival does not measure return to the commanded motion. An impulse-matched 0.2 s force-pulse check gave qualitatively consistent results for the proposed controller but was not run for the reference controllers, and other push durations and rotational disturbances were not tested. The single-channel ablations do not resolve interactions between action groups, and the joint-space comparison is not training-matched. WBIC has been demonstrated on hardware, but transfer of the complete hierarchy has not been tested.
Citation
@misc{banus2026structured,
title = {Learning Structured Locomotion References for
Optimization-Based Whole-Body Control},
author = {Banu{\c{s}}, {\"O}mer Tar{\i}k and Saranl{\i}, Ulu{\c{c}}
and Ankaral{\i}, M. Mert},
note = {Submitted to the 2027 IEEE International Conference
on Robotics and Automation (ICRA)},
year = {2026}
}