DriftOPD Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

Preprint
Youngjun Jun1, Kyumin Choi2, Youngmin Kim1, Seonghyun Jin1, Sunwoo Park1, Jangho Park1, Jong Chul Ye1
1KAIST  ·  2Sungkyunkwan University
youngjun.jun@kaist.ac.kr  ·  jong.ye@kaist.ac.kr
Every clip is one environment setting run by six policies at once — the multi-step teacher, its one-step prediction, sCD, MFD, Drift and DriftOPD — played at 5×.
Chunk-level VLA training versus the sequence-level formulation used by DriftOPD
Chunk-level VLA training versus our sequence-level formulation.

Abstract

Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation.

We introduce DriftOPD, a teacher-free, rollout-free framework for sequence-level on-policy distillation of continuous VLA action experts. We show that the sequence-level reverse Kullback–Leibler divergence decomposes into a chunk-level reverse-KL term and a future-potential term that captures the long-horizon effect of the current action. DriftOPD optimizes these two terms using a one-step drifting objective and a Q-function critic learned from offline demonstrations, respectively, enabling sequence-level optimization with only offline data and one-step action generation.

Across multiple VLA architectures in simulation and real-world manipulation, DriftOPD consistently outperforms existing one-step distillation baselines while matching the task success of multi-step teacher policies. These results demonstrate that long-horizon behavior can be effectively distilled into one-step VLA action experts without online interaction or a separate teacher.

Rollout-free

Sequence-level supervision comes from a Q-function critic trained on the same offline demonstrations — no closed-loop data collection.

Teacher-free

The drifting objective replaces the multi-step teacher, so the student is not bounded by a teacher it has to query during distillation.

One-step inference

A single denoising step at deployment, matching the task success of the multi-step teacher on a real robot.

Method

The sequence-level reverse KLD splits into a chunk-level reverse-KL term, optimised by a one-step drifting objective, and a future-potential term, estimated by a demonstration-trained critic.

DriftOPD method overview
Comparison of VLA training paradigms. (a) Standard VLA training uses demonstration-supervised forward KLD, (b) existing one-step distillation optimises chunk-level reverse KLD, and (c) RL relies on closed-loop rollouts, whereas (d) DriftOPD approximates sequence-level reverse KLD using a drifting objective and critic guidance without online rollouts.

Real-robot rollouts

Every panel is the same environment setting executed by six policies: the 4-step teacher, its 1-step prediction, and four one-step distillation methods. A panel freezes with its outcome when its episode ends, so the grid also shows how long each policy took. Videos below are settings where DriftOPD either beat both teachers on time or succeeded where the 4-step teacher failed.

or press Enter — a set advances on its own once every clip has finished

Simulation rollouts — RoboCasa365

The same six-panel layout on RoboCasa365 atomic_seen, for both π0.5 and GR00T N1.6: the multi-step teacher, its one-step prediction, and four one-step distillation methods on a shared environment seed.

or press Enter — a set advances on its own once every clip has finished

Simulation rollouts — RoboTwin 2.0

One episode per task (48 tasks) with ABot-M0, in the same six-panel layout: the 10-step teacher, its one-step prediction, and four one-step distillation methods sharing the environment seed.

or press Enter — a set advances on its own once every clip has finished

Real-world results

GR00T N1.6 on the I2RT YAM arm: four single-arm pick-and-place tasks and two bimanual hand-over tasks, 15 paired environment settings per cell.

Real-world evaluation in paired environments on unimanual pick-and-place and bimanual handover tasks. Success rate (%) and oscillation (°). Oscillation is the RMS of the commanded joint trajectory after removing its smooth trend (Savitzky–Golay, 0.5 s window, order 2), averaged over the arm joints. Bold marks the best student; the teachers are references. Parentheses give the partial-success rate (e.g. picking the object but failing to place or hand it over).
Task Success rate (%) Oscillation (°)
Teacher
4-step
Teacher
1-step
sCDMFDDriftDriftOPD Teacher
4-step
Teacher
1-step
sCDMFDDriftDriftOPD
single-arm (right arm)
Easy 93.340.00.0 (60.0)0.0 (0.0)86.793.3 1.701.774.647.551.972.01
Moderate 73.346.76.7 (20.0)0.0 (0.0)60.060.0 1.611.904.817.771.831.89
Hard 66.713.30.0 (0.0)0.0 (0.0)60.066.7 1.651.794.877.691.701.89
Challenging 46.70.00.0 (0.0)0.0 (0.0)13.313.3 1.642.035.017.661.711.68
Avg 70.025.01.70.055.058.3 1.651.874.837.671.801.87
bimanual
Handover (LR) 73.360.00.0 (0.0)0.0 (0.0)53.380.0 1.692.083.895.442.052.19
Handover (RL) 86.753.30.0 (26.7)0.0 (20.0)40.073.3 1.691.933.755.162.112.02
Avg 80.056.70.00.046.776.7 1.692.003.825.302.082.10

Simulation results

RoboCasa365, 30 episodes per task under paired environment settings.

Average task success rate (%). Each task uses 30 episodes, totalling 540, 480 and 480 per suite. Teacher steps: 10 for π0.5, 4 for GR00T. Bold marks the best one-step method.
ModelFine-tuningSuite Teacher1-step
TeachersCDMFDDriftDriftOPD
π0.5Full FTatomic_seen (18) 39.631.3028.3334.4432.2238.33
composite_seen (16) 7.10.831.042.505.005.42
composite_unseen (16) 1.20.830.830.831.041.67
π0.5LoRA (r=16)atomic_seen (18) 39.631.3029.2628.3334.0739.26
composite_seen (16) 7.10.831.672.713.125.00
composite_unseen (16) 1.20.830.831.251.461.46
GR00T N1.5Action expertatomic_seen (18) 50.742.9639.6343.5243.8945.37
composite_seen (16) 14.87.714.178.548.128.54
composite_unseen (16) 2.72.501.672.712.292.92
GR00T N1.6Head + top-4 VLM layersatomic_seen (18) 51.144.8151.6732.9647.2249.44
composite_seen (16) 9.42.713.540.833.965.00
composite_unseen (16) 1.70.830.621.040.621.04
RoboTwin 2.0, grouped by task horizon. 30 episodes per task under paired environment settings. short, medium and long horizons are average trajectory lengths of <150, 150–279 and ≥280 steps. Bold marks the best one-step method.
ModelAction objectiveFine-tuning SuiteTeacher1-step
TeachersCDMFDDriftDriftOPD
π0.5-LoRAFlow matchingLoRA (r=16)short (18) 60.7461.3064.8154.4466.3067.78
medium (21) 67.3056.3561.1148.2561.9064.13
long (11) 55.1540.6136.9716.3648.4846.67
ABot-M0Direct clean actionJoint FT (VLM + action expert)short (18) 70.1970.3767.412.4169.0764.63
medium (21) 73.9767.7866.834.1365.7168.10
long (11) 36.9736.9730.610.9135.7638.79

GPU time to success

One inference step does not automatically mean faster task completion: a noisier one-step policy can need more policy calls, or fail outright. We therefore measure policy calls × per-call latency on 10 randomly selected RoboCasa365 atomic_seen tasks with π0.5. DriftOPD is 2.84× faster than the 10-step teacher.

GPU time to success per method
GPU time to success: policy calls × per-call latency (idle B200, batch 1). Bars are total GPU time per success including failed episodes; diamonds (parenthesised) are the mean over successful episodes.

Critic analysis

Critic analysis: Q over episode progress, and success rate versus the critic weight
Analysis of critic guidance for sequence-level VLA policy optimisation on RoboCasa365 with GR00T N1.6. (a) Critic Q over episode progress for demonstrations, successful rollouts and failed student rollouts. (b) Task success rate versus λ on the atomic_seen suite.

BibTeX

@article{jun2026driftopd,
  title   = {DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies},
  author  = {Jun, Youngjun and Choi, Kyumin and Kim, Youngmin and Jin, Seonghyun and
             Park, Sunwoo and Park, Jangho and Ye, Jong Chul},
  journal = {arXiv preprint arXiv:2610.00317},
  year    = {2026}
}