MoSE3: Learning World-Space SE(3) at Every Pixel
Paper • 2610.03716 • Published • 1
Paper | Project page | Code
MoSE3 is a feed-forward model that predicts world-space SE(3) motion for every pixel of a monocular RGB video, generalizing across rigid, articulated and deformable objects.
from mose3.models.mose3 import MoSE3
model = MoSE3.from_pretrained("mose3-tracker/MoSE3", strict=True).cuda().eval()
The code repository has the full inference and visualization scripts.
The weights are released under CC BY-NC 4.0 (non-commercial). They include the weights of π³, which are released under the same license.
@inproceedings{cheng2026mose3,
title = {{MoSE3}: Learning World-Space {SE(3)} at Every Pixel},
author = {Cheng, Jiahuan and Li, Zhiyi and Xia, Tian and
Cai, Ruojin and Du, Yilun and Wang, Qianqian},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2026}
}