Routed Forcing Read paper

Streaming avatars · arXiv 2026

Where and When to Force:
Routed Forcing
for Streaming Avatars

More expressive motion. More diverse gestures.
Distillation that knows where and when to guide.

Zihan Su1,2,*Siwen Lu1,*Junhao Zhuang2,†Zeyue Xue2Haoyang Huang2
Guanghao Li1Xiaofeng Tan3Chun Yuan1,†Nan Duan2

1 Tsinghua University2 Joy Future Academy, JD3 Southeast University

* Equal contribution   † Corresponding authors

+38.8%

Motion dynamics

SpeakerVid · RAFT-Motion vs. Self Forcing
+7–25%

Motion diversity

Four metrics · SpeakerVid & AVSpeech
1.3B

Causal streaming student

Four sampling steps per block

The idea

Not every region needs
the same supervision.

Distilling a video diffusion model into a fast causal generator can compress its motion diversity. The loss is uneven: pose and gesture regions suffer most, while the audio-driven mouth and background are less affected.

Routed Forcing directs real-video supervision to the non-mouth person region at high noise stages. Distribution matching handles the rest, preserving lip synchronization and refining visual details.

Read the full abstract

01 / Method

Two decisions. One routing rule.

Match the objective to the region
and the stage of denoising.

WHERE

Route by semantic region.

Use Data-Forcing Distillation on the non-mouth person region to recover diverse pose and gesture patterns. Keep DMD on the mouth and background.

WHEN

Route by noise stage.

Activate DFD at high noise to guide motion. At low noise, use DMD to refine details and avoid blur from spatial differences between real and generated videos.

Routed Forcing pipeline: a causal student receives first-frame, audio, and text conditions. Region and noise-stage routing combines DMD on generated rollouts with DFD on real videos.
Figure 2 DFD supervises the non-mouth person region at high noise; DMD is used elsewhere. Click the figure to view it at full resolution.
Non-mouth person + high noiseDFDAll other regions / noise stagesDMD
Why route? The regional diversity observation
Motivation figure comparing teacher and student motion diversity. The non-mouth person region shows the largest diversity drop.
Figure 1 Regional diversity is measured as mean pairwise LPIPS across four seeds over 100 SpeakerVid conditions. The non-mouth person region shows the largest drop.

02 / See the difference

Give motion room to unfold.

Same first frame. Same audio.
Three training strategies.

Paper Figure 3: two cases across four timesteps, comparing Self Forcing, Data-Forcing Distillation, and Routed Forcing. Routed Forcing shows more varied hand gestures with clearer hand details.
Routed Forcing produces richer gestures than Self Forcing and clearer hands than Data-Forcing Distillation.
Paper Figure 8: two additional cases across four timesteps, comparing Self Forcing, Data-Forcing Distillation, and Routed Forcing. The examples show changes in hand position and palm orientation.
Additional examples show more varied hand positions and palm orientations, while preserving clearer details during gestures.

03 / Results

More motion. Details preserved.

Controlled comparisons with the same
initialization, training data, and random seeds.

↓ lower is better   ↑ higher is better
SpeakerVid · Controlled training strategy comparison
MethodVideo QualityLip SyncDynamics (×100)Diversity (×100)
FID ↓FVD ↓Sync-C ↑Sync-D ↓RAFT ↑DINO-T ↑CLIP ↑DINO ↑LPIPS ↑L1 ↑
Self Forcing27.17320.914.0310.975.621.073.071.374.962.14
DFD28.78325.333.9211.105.471.113.081.424.671.84
DP-DMD27.22320.274.0010.985.671.063.001.334.892.13
Reward Forcing26.99315.844.0410.975.791.093.141.415.202.21
DynaForcing27.75326.384.0410.955.451.072.961.314.692.06
Routed Forcing Ours25.43309.924.0710.947.801.253.551.586.092.29

Full results from Table 1. Dynamics and diversity metrics are scaled by ×100, as in the paper. View Table 1 in the paper ↗

38.8% higher motion dynamics on SpeakerVid. RAFT-Motion rises from 5.62 to 7.80, while FID improves from 27.17 to 25.43 and Sync-C from 4.03 to 4.07.

04 / Resources

Build on this work.

If you find Routed Forcing useful,
please consider citing our paper.

GitHub repository ↗

Training and inference code coming soon.

BibTeX
@misc{su2026routedforcing,
  title={Where and When to Force: Routed Forcing for Streaming Avatars},
  author={Zihan Su and Siwen Lu and Junhao Zhuang and Zeyue Xue and Haoyang Huang and Guanghao Li and Xiaofeng Tan and Chun Yuan and Nan Duan},
  year={2026},
  eprint={2609.30963},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2609.30963}
}