Ömer Tarık Banuş / ICRA 2027 submission

Learning Structured Locomotion References for Optimization-Based Whole-Body Control

Ömer Tarık Banuş · Uluç Saranlı · M. Mert Ankaralı

Supplementary video. The controller recovers from lateral and sagittal disturbances; the overlays show gait cadence, foothold, and body-state references.

Abstract

Quadruped locomotion requires coordinating body motion, contact forces, foothold placement, and gait timing. We present a hierarchical controller in which reinforcement learning adapts these structured locomotion references while a gait scheduler, reactive footstep planner, and whole-body impulse controller retain responsibility for high-frequency execution. The model-based backbone can sustain a nominal trot independently, allowing the policy to focus on when and how to modify an already functional locomotion controller. We train the complete hierarchy in a vectorized controller-in-the-loop environment using a six-direction disturbance curriculum. Under a common simulator and evaluation protocol, the proposed controller achieves 77.29% mean survival across the tested disturbance grid, compared with 50.35% for the model-based reference controller. Training-time channel ablations identify adaptive foothold placement as the largest marginal contributor to grid survival and reveal direction-dependent effects of gait-rate adaptation.

Grid survival
77.29%
± 1.15 over three seeds
MPC+WBIC reference
50.35%
same simulator and protocol
Full survival up to
2.0 m/s
every seed, every direction
Largest ablation effect
−10.83 pp
without foothold residuals

Method

Model-based quadruped controllers separate gait scheduling, footstep planning, force generation, and whole-body control. This gives reliable, interpretable execution, but the behaviour depends on how the intermediate references are produced: gait cadence is usually fixed and footholds come from engineered rules. Joint-space reinforcement learning avoids hand-designed references, but the policy must learn joint coordination from scratch and its outputs do not enforce dynamic or contact feasibility.

We keep the model-based stack and replace the reference generator. One policy has supervisory authority over several coupled levels of a controller that can already trot on its own. Learning can then concentrate on adapting nominal locomotion.

Controller hierarchy: a 100 Hz RL policy maps a 54-dimensional observation to a 25-dimensional action of structured references for a 500 Hz controller made of a continuous gait scheduler, Raibert-style footstep planner, and whole-body impulse control, acting on a MuJoCo robot model.
Fig. 1. Controller hierarchy. A 100 Hz policy supplies structured references to a gait scheduler, footstep planner, and WBIC running at 500 Hz.

Policy interface

The policy runs at 100 Hz; each action is held for five 2 ms updates of the lower layers. It reads a 54-dimensional observation of the velocity command, body velocity and orientation, joint states, scheduled contacts, foot positions, swing phase, and body height. It uses no history buffer and no privileged disturbance information. The 25-dimensional action has five groups:

Action components (paper Table I). Force targets are absolute; the other groups are residuals on the nominal references.
Group Dimensions Range
Body-velocity and yaw-rate residuals 3 ±0.5 m/s, ±0.3 rad/s
Stance-force targets, per leg (fx, fy, fz) 12 ±50, ±50, 0–250 N
Foothold offsets, per leg (x, y) 8 ±0.5 m
Body-height residual 1 ±0.05 m
Gait-phase increment 1 0–0.01 per update

Structured controller

Gait scheduler. A normalized phase defines a diagonal trot. The policy's phase increment can speed up or pause the gait cycle without changing the contact order; a neutral action gives a 0.4 s cycle.

Footstep planner. Footholds follow the Raibert-style rule of Kim et al., with a bounded velocity-feedback correction. The policy's foothold residual is added after that bound, so it can move the touchdown point beyond the heuristic's range. Bézier swing trajectories with 0.10 m clearance connect lift-off to touchdown.

Whole-body impulse control. WBIC adjusts the nominal accelerations and the policy's desired forces subject to floating-base dynamics, unilateral contact, and a friction pyramid (µ = 0.45), so the applied forces can differ from the policy's targets. Inverse-dynamics feedforward and joint PD feedback produce the torque command, with actuator limits enforced.

The reward is a sum of bounded task-space terms of the form wi exp(−pi), covering forward-velocity tracking, height, orientation, straightness, smoothness, and phase progression. A fall ends the episode with reward −10.

Training and evaluation

Controller-in-the-loop training

We built a C++ EnvPool environment in which every instance runs the complete closed loop: MuJoCo simulates the robot, and the embedded locomotion stack performs state processing, gait and foothold generation, whole-body control, and joint-command execution. EnvPool parallelizes full copies of this loop, so the policy trains against the real execution path instead of a reduced model.

Three policies were trained from scratch with PPO using different seeds, for 85.1 million transitions each, with 256 parallel environments. Training disturbances are sudden velocity changes in six directions; their ceiling rises from 1 to 2 to 3 m/s at 8.1 and 24.3 million transitions. All three runs completed without collapse, with final timeout fractions of 92.9–98.4%.

Go2 quadruped model in MuJoCo, with an arrow indicating the disturbance direction.
Fig. 2. Go2 simulation environment. The arrow shows the disturbance direction.
Episode length and timeout fraction over 85.1 million transitions for three seeds, plateauing near the end of training.
Fig. 3. Episode length and timeout fraction for three seeds. Shading marks the 1, 2, and 3 m/s curriculum stages. Curves are unsmoothed.

Evaluation protocol

All experiments use a 15 kg Unitree Go2 model on flat ground with a 2 ms time step and observation noise on every measured signal. Each disturbance is an instantaneous change in body velocity, following Rudin et al. The grid covers six directions (forward, backward, left, right, up, down) and ten magnitudes from 0.5 to 5.0 m/s, with 100 episodes per cell and policy. An episode survives if the body does not touch the ground and roll and pitch stay within 0.8 and 1.0 rad for 5.5 s after the disturbance.

Two reference controllers are reimplemented in the same MuJoCo and EnvPool framework and share the plant, physics, actuator limits, noise, disturbance generator, and evaluation code:

  • MPC+WBIC, the Mini Cheetah architecture of Kim et al.: convex MPC force planning, reactive foot placement, and WBIC. One fixed controller.
  • Joint-space RL, a Go2 recipe adapted from Rudin et al. and implemented with rsl_rl: a 50 Hz policy that outputs 12 joint-position offsets. Three policies, 196.6 million transitions each.

Each keeps its own observations, actions, objectives, and update rate, so the comparison is between complete configured systems.

Results

Every controller completed every undisturbed episode, so the differences below reflect disturbance recovery. The proposed controller survived every episode in every direction up to 2.0 m/s for all three seeds.

Mean survival over the 60 disturbance conditions (paper Table IV). Learned controllers report mean ± sample SD across three seeds.
Controller Grid survival (%)
MPC+WBIC 50.35
Joint-space RL (source recipe) 54.01 ± 5.55
Proposed 77.29 ± 1.15
Six survival curves by disturbance direction. The proposed controller's 50% thresholds are 4.46 m/s forward, above 5.0 backward, 4.19 left, 3.80 right, 3.42 up, and 3.67 down.
Fig. 4. Survival under velocity disturbances, 100 episodes per cell and policy. Bands show ±1 SD across three seeds. Panel titles give the proposed controller's 50% survival threshold c50.

The largest gains over MPC+WBIC are lateral: c50 rises from 2.26 to 4.19 m/s to the left and from 1.96 to 3.80 m/s to the right. Backward survival is still 70.3% at 5.0 m/s, beyond the tested range. Seed variation is concentrated near the edge between survival and failure; the largest per-cell spread is 14.7 percentage points at 4.0 m/s leftward, while the per-seed thresholds there span only 3.95–4.26 m/s.

The joint-space result describes that configured policy. A joint-space policy trained specifically for this disturbance protocol could perform differently, so the comparison does not isolate the effect of the action representation.

Action-channel ablations

To measure what each action group contributes after the controller has adapted to its absence, we trained five more policies from scratch, each with one group held at its neutral value throughout training and evaluation. The other groups can compensate during training.

Ablation results. Change in grid survival: no foothold residual −10.83, no body-velocity residual −2.37, fixed force target −1.50, fixed cadence −1.12, fixed height −0.57 percentage points, with signed directional changes in c50.
Fig. 5. Change in grid survival (left) and in directional c50 (right; negative is worse), relative to an unablated reference at 77.43%.
  • Footholds matter most. Without foothold residuals, grid survival falls from 77.43% to 66.60%, with the largest lateral losses (−1.19 m/s left, −0.93 m/s right). Downward tolerance improves by 0.31 m/s, so the effect depends on direction.
  • Cadence shifts robustness between directions. Fixing the gait cadence lowers the pooled score by only 1.12 points but reduces upward c50 by 0.62 m/s, and it is the only ablation that reduces backward c50.
  • Smaller effects change sign. The body-velocity, force, and height ablations cost 2.37, 1.50, and 0.57 points, with gains in some directions and losses in others. This points to compensation between channels.
Nominal mechanical cost of transport, forward speed, mechanical power, and joint torque RMS for each retrained ablation controller.
Fig. 6. Undisturbed walking for each retrained controller over five episodes. Circles are episodes; diamonds are means.

The fixed-cadence controller uses less mechanical power with similar pooled survival, so cadence adaptation adds directional flexibility in these tests without a demonstrated energy benefit. The no-foothold controller has the lowest torque RMS together with the largest survival loss, which suggests the extra actuation from learned foothold correction does useful work.

Limitations

All evidence comes from flat-ground simulation with noisy state. Survival does not measure return to the commanded motion. An impulse-matched 0.2 s force-pulse check gave qualitatively consistent results for the proposed controller but was not run for the reference controllers, and other push durations and rotational disturbances were not tested. The single-channel ablations do not resolve interactions between action groups, and the joint-space comparison is not training-matched. WBIC has been demonstrated on hardware, but transfer of the complete hierarchy has not been tested.

Citation

@misc{banus2026structured,
  title  = {Learning Structured Locomotion References for
            Optimization-Based Whole-Body Control},
  author = {Banu{\c{s}}, {\"O}mer Tar{\i}k and Saranl{\i}, Ulu{\c{c}}
            and Ankaral{\i}, M. Mert},
  note   = {Submitted to the 2027 IEEE International Conference
            on Robotics and Automation (ICRA)},
  year   = {2026}
}