Factorized Flow Policy Optimization for Efficient Humanoid Whole-Body Control

Anonymous Authors

FFPO decomposes a whole-body flow policy into functional body agents, enabling reliable credit assignment and efficient policy improvement while preserving parallel control at execution.

Method

Body-wise flow policy optimization

Overview of Factorized Flow Policy Optimization
FFPO factorizes coupled whole-body updates into functional body actors, combines precise sequential credit with efficient policy gradients, and allocates optimization effort using body-wise gradient reliability.

Benchmark

16 whole-body control tasks

The benchmark spans reference-free locomotion, motion-conditioned locomotion and balance, and loco-manipulation with sustained whole-body coordination.

Task familyTasksCapability tested
Reference-free locomotion Stand; Walk; Run; Step-Up Basic gait learning, command following, balance, and terrain traversal.
Reference-based locomotion Turn-Jog; One-Foot Balance; Normal Walk; March; Goose Step; Frog Jump; Zombie Hop Motion tracking, balance transitions, jumping, and coordinated inter-limb control.
Loco-manipulation Bimanual Lift; Door Open; Pick-Carry-Place; Crate Carry; Box Push Balance-preserving interaction, object transport, and forceful manipulation.

Experiments

Main results

5.2 Overall performance

Fast learning without late-training collapse

Across all 16 tasks, FFPO achieves the strongest aggregate learning efficiency and remains stable on coordination-intensive motions. It improves normalized reward AUC by 14.4% and the task-balanced final normalized reward by 14.1% over the strongest baseline on each metric.

Success-rate learning curves for all 16 tasks
Task success-rate curves across the 16-task benchmark and five random seeds.
AlgorithmNormalized AUC MeanNormalized AUC IQMFinal IQMMean Drawdown
Gaussian PPO0.4370.4800.1580.517
FPO-0.683-0.430-0.7321.595
FPO++0.7190.8170.9790.314
GenPO0.7420.8000.8970.005
PolicyFlow0.6230.6730.8900.030
FFPO0.8490.8750.9760.014
Aggregate statistical comparison of FFPO and baselines
Aggregate final performance, learning progress, and reward-AUC profiles.

5.3 Mechanism analysis

Factorization improves update reliability

Body factorization substantially raises sample efficiency while suppressing low-confidence updates and late-training degradation. Reliability-aware allocation then directs additional optimization toward body agents with informative gradients.

Evidence for update reliability under body factorization
Five-seed comparison on Stand, Turn-Jog, and Crate Carry.

5.4 Ablation studies

Each optimization component contributes

Removing policy-step control slows learning, removing prefix correction reduces both learning speed and final performance, and SNR-based epoch allocation provides a consistent efficiency gain. FFPO remains effective over a broad range of target policy-change magnitudes and maximum step scales.

Component ablation learning curves
Leave-one-component-out ablation.
Hyperparameter sensitivity analysis
Sensitivity to policy-change magnitude and maximum step scale.

5.5 Sim2Real evaluation

Coordinated behavior transfers to Unitree G1

FFPO learns stable Normal Walk behavior after roughly 200 policy updates, while FPO++ requires more than 400 updates for comparable completion. The deployed FFPO policy completes the full 12.7-second motion using proprioceptive observations at 50 Hz, with body-wise tracking profiles that remain close to simulation. March and Goose Step provide two additional hardware motions.

Simulation training.
Task success-rate curves during simulation training on Normal Walk.

Demonstrations

Simulation videos

One FFPO rollout is shown for each task, selected from the strongest available task checkpoint.

Stand

Walk

Run

Step-Up

Turn-Jog

One-Foot Balance

Normal Walk

March

Goose Step

Frog Jump

Zombie Hop

Bimanual Lift

Door Open

Pick-Carry-Place

Crate Carry

Box Push

Unitree G1

Hardware demonstrations

Proprioceptive FFPO policies execute three whole-body motions on the 29-DoF Unitree G1.

Normal Walk

March

Goose Step