Egocentric Body Motion Reconstruction
ReViV establishes a new state of the art for monocular egocentric body motion reconstruction. Even without camera trajectories as input, it outperforms baselines that rely on ground-truth cameras across local pose accuracy, semantic correspondence, and motion realism, while being 100× faster than EgoAllo and over 10× faster than UniEgoMotion — demonstrating that global ego-motion can be effectively learned implicitly from pure monocular video.
Baselines given a VIPE camera trajectory
Columns from left to right: Input RGB · Ground Truth · Ours (RGB only) · EgoAllo (RGB + VIPE camera) · UniEgoMotion (RGB + VIPE camera).
Baselines given the ground-truth camera trajectory
Columns from left to right: Input RGB · Ground Truth · Ours (RGB only) · EgoAllo (RGB + GT camera) · UniEgoMotion (RGB + GT camera).