Route by semantic region.
Use Data-Forcing Distillation on the non-mouth person region to recover diverse pose and gesture patterns. Keep DMD on the mouth and background.
Streaming avatars · arXiv 2026
More expressive motion. More diverse gestures.
Distillation that knows where and when to guide.
1 Tsinghua University2 Joy Future Academy, JD3 Southeast University
* Equal contribution † Corresponding authors
Motion dynamics
SpeakerVid · RAFT-Motion vs. Self ForcingMotion diversity
Four metrics · SpeakerVid & AVSpeechCausal streaming student
Four sampling steps per blockThe idea
Distilling a video diffusion model into a fast causal generator can compress its motion diversity. The loss is uneven: pose and gesture regions suffer most, while the audio-driven mouth and background are less affected.
Routed Forcing directs real-video supervision to the non-mouth person region at high noise stages. Distribution matching handles the rest, preserving lip synchronization and refining visual details.
Read the full abstract01 / Method
Match the objective to the region
and the stage of denoising.
Use Data-Forcing Distillation on the non-mouth person region to recover diverse pose and gesture patterns. Keep DMD on the mouth and background.
Activate DFD at high noise to guide motion. At low noise, use DMD to refine details and avoid blur from spatial differences between real and generated videos.

02 / See the difference
Same first frame. Same audio.
Three training strategies.
03 / Results
Controlled comparisons with the same
initialization, training data, and random seeds.
| Method | Video Quality | Lip Sync | Dynamics (×100) | Diversity (×100) | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| FID ↓ | FVD ↓ | Sync-C ↑ | Sync-D ↓ | RAFT ↑ | DINO-T ↑ | CLIP ↑ | DINO ↑ | LPIPS ↑ | L1 ↑ | |
| Self Forcing | 27.17 | 320.91 | 4.03 | 10.97 | 5.62 | 1.07 | 3.07 | 1.37 | 4.96 | 2.14 |
| DFD | 28.78 | 325.33 | 3.92 | 11.10 | 5.47 | 1.11 | 3.08 | 1.42 | 4.67 | 1.84 |
| DP-DMD | 27.22 | 320.27 | 4.00 | 10.98 | 5.67 | 1.06 | 3.00 | 1.33 | 4.89 | 2.13 |
| Reward Forcing | 26.99 | 315.84 | 4.04 | 10.97 | 5.79 | 1.09 | 3.14 | 1.41 | 5.20 | 2.21 |
| DynaForcing | 27.75 | 326.38 | 4.04 | 10.95 | 5.45 | 1.07 | 2.96 | 1.31 | 4.69 | 2.06 |
| Routed Forcing Ours | 25.43 | 309.92 | 4.07 | 10.94 | 7.80 | 1.25 | 3.55 | 1.58 | 6.09 | 2.29 |
Full results from Table 1. Dynamics and diversity metrics are scaled by ×100, as in the paper. View Table 1 in the paper ↗
38.8% higher motion dynamics on SpeakerVid. RAFT-Motion rises from 5.62 to 7.80, while FID improves from 27.17 to 25.43 and Sync-C from 4.03 to 4.07.
04 / Resources
If you find Routed Forcing useful,
please consider citing our paper.
Training and inference code coming soon.
@misc{su2026routedforcing,
title={Where and When to Force: Routed Forcing for Streaming Avatars},
author={Zihan Su and Siwen Lu and Junhao Zhuang and Zeyue Xue and Haoyang Huang and Guanghao Li and Xiaofeng Tan and Chun Yuan and Nan Duan},
year={2026},
eprint={2609.30963},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.30963}
}