Spectral Skills: A Command Space for Humanoid Control
Our new paper, Learning Expressive and Compositional Motion Representation via Spectral Skills, is now on arXiv. It is joint work with Chenxiao Gao, Chen Yang, Ye Zhao, Bo Dai, and Anqi Wu.
A Unitree G1 walking with its right arm raised. The arm-up behavior is not in the base walk: it is added to the skill command by steering along one spectral direction.
The question
A humanoid stack usually has two layers. A high-level model decides what should happen, and a low-level whole-body controller makes it happen at 50 Hz. The awkward part is the message between them. If it is a joint or Cartesian trajectory, the planner has to predict every detail of the robot’s dynamics, and its errors accumulate in closed loop. If it is a latent code, we get something compact, but we usually do not know what the code means, whether the planner can predict it, or whether it can be edited.
The paper asks one thing: what command representation should a planner use to instruct a whole-body controller? We wanted three properties from the answer. It should be expressive enough to track a wide range of motions accurately, efficient to learn and to predict, and compositional, so that a walk plus a wave does not need data for every walk-with-a-wave.
Learning skills by predicting what comes next
Most latent interfaces are learned by reconstruction: encode a motion segment, decode it back. That rewards a code for remembering the segment, including details that do not matter for what happens after. We instead train the encoder by asking the decoder to predict the motion that follows the segment.
Concretely, a 0.2 second segment is encoded into a 64-dimensional skill . The decoder is a factorized diffusion model whose transition kernel has the form . The one design choice that matters most is that the skill enters only through the affine map . We call the result a spectral skill because the factorization follows the spectral view of transition operators used in spectral representation learning for RL.
Once the encoder is trained and frozen, we train a single PPO controller on top of it that tracks reference motions conditioned on skills.
Composition is addition
The affine form buys something concrete. For a fixed context, the Jacobian of the decoder’s denoised estimate with respect to the skill does not depend on the skill itself (Proposition 3.1 in the paper). Moving the skill by moves the predicted future by exactly , whatever the base skill is.
That makes it possible to ask which directions change the predicted motion most across many contexts. Averaging the squared Jacobian response over sampled contexts gives a response Gram matrix, and its eigenvectors are what we call spectral directions. They come from the trained model alone, with no labels. When we run them through the controller, many turn out to change one recognizable thing: one arm rises, the feet lift higher, the robot turns.
Composing a new skill is then
where is the encoded base motion, collects the chosen directions, and is a time-varying amplitude that turns each direction on, off, up, or down while the base skill keeps going. No new data and no retraining of the controller.
What it does
Tracking
Conditioned on spectral skills, the controller tracks more accurately than the released SONIC tracker under identical settings. On the 4,096-motion evaluation set, mean per-joint error in the world frame drops from 187.9 mm to 70.5 mm, and in the robot’s local frame from 26.7 mm to 20.5 mm. Success rate is 98.3% versus 98.9%, so SONIC is 0.6 points better there and we give up a little on the hardest motions. On the smaller 124-motion capability set both methods succeed on every clip, and our global error is 65.1 mm versus 173.9 mm. That is the 62% in the abstract. The controller was trained with about 12 times fewer simulated frames.
Almost all of the gap is drift. SONIC reproduces body configurations accurately but wanders from the commanded trajectory, and root position accounts for 98% of its squared global error. Our reading is that a representation trained to predict how motion progresses over long horizons carries the information needed to stay on the path. Matched-budget ablations point the same way: an encoder with the same architecture trained by reconstruction has about 2.8 times the global error, with local accuracy roughly unchanged.
On the G1, the deployment controller tracks retargeted motion capture clips from a dance to a kick to a 180 degree turn.
Chaining
A controller that tracks skills can also chain them: change the commanded skill mid-run and the robot moves from one behavior to the next, with no separate transition policy. In simulation we ran 42 chained sequences, each switch taking between 2 and 30 control steps, and measured the RMS joint acceleration around the switch. Ours is smoother than SONIC in 36 of 42, including all 13 switches that happen within two steps.
Steering
Steering a walk along one direction mostly changes one attribute while the walk continues. At amplitude 2 in Isaac PhysX, raising the left or right arm moves shoulder pitch by about 0.9 rad, the high-steps direction raises the swing apex from 15 to 36 cm, and the turn direction turns the robot by about 130 degrees in 4 seconds. Effects also add: two directions held together change the joints like the sum of their individual changes.
Not every direction is clean everywhere. Of 19 sampled directions, 13 can be applied cleanly on the walk base and 7 on the squat base, and five arm directions work on all five bases we tested. The paper has a fuller catalog, including negative results for BFM-Zero and SONIC steering.
The same thing works on the robot. The clips below start from a walk and add a right arm and then turns of increasing size.
A natural worry is that steering only retrieves motions that were already in the training set. We checked this in the skill space. The map below shows all 129,785 training clips, and a walk steered along the arm, turn, and high-step directions, one at a time. The steered skill leaves the regions where corpus walks that raise an arm or turn are common, and its distance to the nearest corpus motion goes up as each direction is added and comes back down when it is released. The unsteered walk stays at a typical distance.
Language planning
Skills can also be the action space of a language-conditioned planner. We compare three routes with the same planner architecture and training budget: predict skills and hand them straight to the controller, predict explicit robot states and re-encode them, or predict explicit states for a controller with no encoder. At a matched replanning interval of 10 steps, predicting skills lowers local error by 29% relative to re-encoding predicted states and raises closed-loop success from 77.1% to 91.1%. Explicit-to-explicit reaches about 52%.
The ordering reverses at replanning every step: explicit predictions win there, 94.5% against 84.3%. So the picture is that explicit states are fine when you can correct them constantly, and skills are the more compact and more robust target when the planner has to commit to longer horizons.
Takeaways
The main thing I take from this project is that the training signal for a command space matters as much as its dimension. Reconstruction asks a code to describe a segment; prediction asks it to describe what the segment leads to, and that seems to be what a planner and a controller both need. Making the skill enter the model affinely then costs little and gives an exact, model-derived way to edit skills afterward.
There are limits. The directions are found by inspecting what each eigenvector does, so naming them is manual, and only some are clean on a given base skill. The Jacobian statement is about one denoising step at a fixed context, not a claim that the whole sampler or the robot’s executed motion is linear. Our hardware experiments also use a deployment version of the controller with extra regularization. Code will be released with the final paper.