Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation.
We introduce DriftOPD, a teacher-free, rollout-free framework for sequence-level on-policy distillation of continuous VLA action experts. We show that the sequence-level reverse Kullback–Leibler divergence decomposes into a chunk-level reverse-KL term and a future-potential term that captures the long-horizon effect of the current action. DriftOPD optimizes these two terms using a one-step drifting objective and a Q-function critic learned from offline demonstrations, respectively, enabling sequence-level optimization with only offline data and one-step action generation.
Across multiple VLA architectures in simulation and real-world manipulation, DriftOPD consistently outperforms existing one-step distillation baselines while matching the task success of multi-step teacher policies. These results demonstrate that long-horizon behavior can be effectively distilled into one-step VLA action experts without online interaction or a separate teacher.
Sequence-level supervision comes from a Q-function critic trained on the same offline demonstrations — no closed-loop data collection.
The drifting objective replaces the multi-step teacher, so the student is not bounded by a teacher it has to query during distillation.
A single denoising step at deployment, matching the task success of the multi-step teacher on a real robot.
The sequence-level reverse KLD splits into a chunk-level reverse-KL term, optimised by a one-step drifting objective, and a future-potential term, estimated by a demonstration-trained critic.
Every panel is the same environment setting executed by six policies: the 4-step teacher, its 1-step prediction, and four one-step distillation methods. A panel freezes with its outcome when its episode ends, so the grid also shows how long each policy took. Videos below are settings where DriftOPD either beat both teachers on time or succeeded where the 4-step teacher failed.
The same six-panel layout on RoboCasa365 atomic_seen, for both
π0.5 and GR00T N1.6: the multi-step teacher, its one-step prediction, and four one-step
distillation methods on a shared environment seed.
One episode per task (48 tasks) with ABot-M0, in the same six-panel layout: the 10-step teacher, its one-step prediction, and four one-step distillation methods sharing the environment seed.
GR00T N1.6 on the I2RT YAM arm: four single-arm pick-and-place tasks and two bimanual hand-over tasks, 15 paired environment settings per cell.
| Task | Success rate (%) | Oscillation (°) | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Teacher 4-step | Teacher 1-step | sCD | MFD | Drift | DriftOPD | Teacher 4-step | Teacher 1-step | sCD | MFD | Drift | DriftOPD | |
| single-arm (right arm) | ||||||||||||
| Easy | 93.3 | 40.0 | 0.0 (60.0) | 0.0 (0.0) | 86.7 | 93.3 | 1.70 | 1.77 | 4.64 | 7.55 | 1.97 | 2.01 |
| Moderate | 73.3 | 46.7 | 6.7 (20.0) | 0.0 (0.0) | 60.0 | 60.0 | 1.61 | 1.90 | 4.81 | 7.77 | 1.83 | 1.89 |
| Hard | 66.7 | 13.3 | 0.0 (0.0) | 0.0 (0.0) | 60.0 | 66.7 | 1.65 | 1.79 | 4.87 | 7.69 | 1.70 | 1.89 |
| Challenging | 46.7 | 0.0 | 0.0 (0.0) | 0.0 (0.0) | 13.3 | 13.3 | 1.64 | 2.03 | 5.01 | 7.66 | 1.71 | 1.68 |
| Avg | 70.0 | 25.0 | 1.7 | 0.0 | 55.0 | 58.3 | 1.65 | 1.87 | 4.83 | 7.67 | 1.80 | 1.87 |
| bimanual | ||||||||||||
| Handover (LR) | 73.3 | 60.0 | 0.0 (0.0) | 0.0 (0.0) | 53.3 | 80.0 | 1.69 | 2.08 | 3.89 | 5.44 | 2.05 | 2.19 |
| Handover (RL) | 86.7 | 53.3 | 0.0 (26.7) | 0.0 (20.0) | 40.0 | 73.3 | 1.69 | 1.93 | 3.75 | 5.16 | 2.11 | 2.02 |
| Avg | 80.0 | 56.7 | 0.0 | 0.0 | 46.7 | 76.7 | 1.69 | 2.00 | 3.82 | 5.30 | 2.08 | 2.10 |
RoboCasa365, 30 episodes per task under paired environment settings.
| Model | Fine-tuning | Suite | Teacher | 1-step | ||||
|---|---|---|---|---|---|---|---|---|
| Teacher | sCD | MFD | Drift | DriftOPD | ||||
| π0.5 | Full FT | atomic_seen (18) | 39.6 | 31.30 | 28.33 | 34.44 | 32.22 | 38.33 |
| composite_seen (16) | 7.1 | 0.83 | 1.04 | 2.50 | 5.00 | 5.42 | ||
| composite_unseen (16) | 1.2 | 0.83 | 0.83 | 0.83 | 1.04 | 1.67 | ||
| π0.5 | LoRA (r=16) | atomic_seen (18) | 39.6 | 31.30 | 29.26 | 28.33 | 34.07 | 39.26 |
| composite_seen (16) | 7.1 | 0.83 | 1.67 | 2.71 | 3.12 | 5.00 | ||
| composite_unseen (16) | 1.2 | 0.83 | 0.83 | 1.25 | 1.46 | 1.46 | ||
| GR00T N1.5 | Action expert | atomic_seen (18) | 50.7 | 42.96 | 39.63 | 43.52 | 43.89 | 45.37 |
| composite_seen (16) | 14.8 | 7.71 | 4.17 | 8.54 | 8.12 | 8.54 | ||
| composite_unseen (16) | 2.7 | 2.50 | 1.67 | 2.71 | 2.29 | 2.92 | ||
| GR00T N1.6 | Head + top-4 VLM layers | atomic_seen (18) | 51.1 | 44.81 | 51.67 | 32.96 | 47.22 | 49.44 |
| composite_seen (16) | 9.4 | 2.71 | 3.54 | 0.83 | 3.96 | 5.00 | ||
| composite_unseen (16) | 1.7 | 0.83 | 0.62 | 1.04 | 0.62 | 1.04 | ||
| Model | Action objective | Fine-tuning | Suite | Teacher | 1-step | ||||
|---|---|---|---|---|---|---|---|---|---|
| Teacher | sCD | MFD | Drift | DriftOPD | |||||
| π0.5-LoRA | Flow matching | LoRA (r=16) | short (18) | 60.74 | 61.30 | 64.81 | 54.44 | 66.30 | 67.78 |
| medium (21) | 67.30 | 56.35 | 61.11 | 48.25 | 61.90 | 64.13 | |||
| long (11) | 55.15 | 40.61 | 36.97 | 16.36 | 48.48 | 46.67 | |||
| ABot-M0 | Direct clean action | Joint FT (VLM + action expert) | short (18) | 70.19 | 70.37 | 67.41 | 2.41 | 69.07 | 64.63 |
| medium (21) | 73.97 | 67.78 | 66.83 | 4.13 | 65.71 | 68.10 | |||
| long (11) | 36.97 | 36.97 | 30.61 | 0.91 | 35.76 | 38.79 | |||
One inference step does not automatically mean faster task completion: a noisier one-step policy can
need more policy calls, or fail outright. We therefore measure policy calls × per-call latency on 10 randomly
selected RoboCasa365 atomic_seen tasks with π0.5. DriftOPD is
2.84× faster than the 10-step teacher.
atomic_seen suite.@article{jun2026driftopd,
title = {DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies},
author = {Jun, Youngjun and Choi, Kyumin and Kim, Youngmin and Jin, Seonghyun and
Park, Sunwoo and Park, Jangho and Ye, Jong Chul},
journal = {arXiv preprint arXiv:2610.00317},
year = {2026}
}