MoSE3: Learning World-Space SE(3) at Every Pixel

Jiahuan Cheng1,3*, Zhiyi Li4*, Tian Xia1*, Ruojin Cai1,2, Yilun Du1,2, Qianqian Wang1,2

1Harvard University   2Kempner Institute   3Johns Hopkins University   4MIT

*Equal contribution

NeurIPS 2026 (Spotlight)

TL;DRMoSE3 predicts dense, world-space SE(3) motion at every pixel of a monocular RGB video — not just where each pixel goes, but how it rotates and which pixels move together as one rigid body.

More results →

Abstract

Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel SE(3) through two jointly learned intermediates, 3D point tracks and rigidity embeddings, and recovers SE(3) by differentiably fitting transforms within each soft rigid cluster, enabling end-to-end prediction and supervision. To close the data gap, we introduce Art-Kubric, a large-scale synthetic dataset with dense SE(3) and rigidity labels for articulated objects with rich physical interactions. MoSE3 achieves state-of-the-art SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks, and state-of-the-art average 3D point tracking accuracy across three datasets, while showing strong generalization to real-world videos despite being trained solely on synthetic motion data.

Explore in 3D

The predictions in the world space they were made in. Drag to turn the scene and scroll to zoom; use the bar at the bottom to move through time.

Layers



0.0 / 0.0 s

Results

The predictions drawn back onto the image, six clips at a time. The view you pick stays as you page through.

Rigidity embeddings

A per-pixel embedding of which parts move together, projected to colour by PCA. It is learned without any segmentation supervision: nothing in training tells the model where one body ends and the next begins. Drag the divider to compare with the input video.

InputRigidity embedding
↔
InputRigidity embedding
↔
InputRigidity embedding
↔
InputRigidity embedding
↔
InputRigidity embedding
↔
InputRigidity embedding
↔

Comparison with ProxyPose

ProxyPose on the left, ours on the right, with each method's wall-clock time on one GPU. ProxyPose predicts one pose per hand-picked query point; ours predicts SE(3) at every pixel in one pass, and is drawn at the same query points only so the two can be compared.

11.8 min 3 queries5.1 s all pixels
11.8 min 3 queries6.4 s all pixels

Benchmarks

SE(3) at pixel, part and object level, and world-space 3D point tracking, against the strongest baselines. Lower is better for RRE and ADD, higher for the rest.

V-DPM4RCTrack4World-Pi3xTrack4World-DA3ProxyPoseOurs

HO3D · object / part SE(3)

RRE↓V-DPM: 38.344RC: 36.08Track4World-Pi3x: 30.46Track4World-DA3: 30.98ProxyPose: 31.22Ours: 17.9617.96
ADD (cm)↓V-DPM: 3.564RC: 3.34Track4World-Pi3x: 3.17Track4World-DA3: 3.2–Ours: 2.062.06
ADD">AUCADD↑V-DPM: 0.664RC: 0.688Track4World-Pi3x: 0.698Track4World-DA3: 0.694ProxyPose: 0.49Ours: 0.8020.802
IoU↑V-DPM: 0.4214RC: 0.409Track4World-Pi3x: 0.414Track4World-DA3: 0.383–Ours: 0.7090.709

iTACO · object / part SE(3)

RRE↓V-DPM: 28.254RC: 18.7Track4World-Pi3x: 18.8Track4World-DA3: 23.78ProxyPose: 28.51Ours: 10.7610.76
ADD (cm)↓V-DPM: 19.924RC: 17.26Track4World-Pi3x: 20.14Track4World-DA3: 22.07–Ours: 10.4910.49
ADD">AUCADD↑V-DPM: 0.4484RC: 0.526Track4World-Pi3x: 0.397Track4World-DA3: 0.399ProxyPose: 0.3Ours: 0.6250.625
IoU↑V-DPM: 0.2584RC: 0.349Track4World-Pi3x: 0.353Track4World-DA3: 0.319–Ours: 0.7840.784

YCBInEOAT · object / part SE(3)

RRE↓V-DPM: 31.64RC: 24.23Track4World-Pi3x: 19.91Track4World-DA3: 28.89ProxyPose: 17.01Ours: 11.9311.93
ADD (cm)↓V-DPM: 8.184RC: 4.52Track4World-Pi3x: 7.53Track4World-DA3: 6.11–Ours: 44
ADD">AUCADD↑V-DPM: 0.5014RC: 0.625Track4World-Pi3x: 0.549Track4World-DA3: 0.542ProxyPose: 0.515Ours: 0.6380.638
IoU↑V-DPM: 0.2344RC: 0.322Track4World-Pi3x: 0.36Track4World-DA3: 0.278–Ours: 0.4980.498

Full result tables →

Method

Method overview. MoSE3 extends the frozen π³ backbone with a trainable tracking branch that predicts dense 3D point tracks and rigidity embeddings. Per-pixel SE(3) motion is then recovered in closed form via a weighted Horn fit on the world-frame tracks, with weights given by rigidity-embedding similarity. Because the Horn fit is differentiable, the dense SE(3) prediction is supervised end-to-end together with the tracks and rigidity embeddings.

Art-Kubric

To the best of our knowledge, no existing public corpus combines dense SE(3) supervision with articulated, interacting scenes, so we built one. Art-Kubric is a synthetic dataset and generation pipeline: articulated assets from 46 PartNet-Mobility, 226 Articraft and 7 Infinigen-Articulated categories; dense long-range 2D tracks and metric world-space 3D tracks; link-level SE(3) pose for every kinematic link at every frame; a per-pixel partition of the image by shared rigid motion; rigid and articulated bodies interacting under physics in every scene; and a rendering pipeline that randomizes HDRI lighting, camera intrinsics and object composition, so new clips can be re-rendered with the same supervision. The release has 5,000 scenes of 13–30 objects observed by a moving camera, exported with depth, normals, three-granularity segmentation and camera parameters: 163M tracks, each with its rigid-group label, per-frame occlusion flag and its link's SE(3). Panels below: rendered frame, actor segmentation, rigid-part segmentation.

9 of the 5,000 scenes in the Art-Kubric dataset.

Art-Kubric overview. Two example scenes, one per row. Columns 1–3: three RGB frames from the clip. Columns 4–7, rendered at the third frame: sampled long-range 3D tracks coloured by their parent rigid body, per-pixel SE(3) motion as rotated coordinate frames at sparse anchors, rigid-part segmentation, and depth. Articulated objects from PartNet-Mobility, Articraft and Infinigen-Articulated interact with rigid Google Scanned Objects under physics simulation while a moving camera observes the scene under varying HDRI lighting.

Failure modes

Cases where the prediction is wrong, kept here rather than left out.

BibTeX

@inproceedings{cheng2026mose3,
  title     = {MoSE3: Learning World-Space SE(3) at Every Pixel},
  author    = {Cheng, Joanna Jiahuan and Li, Zhiyi and Xia, Tian and Cai, Ruojin and Du, Yilun and Wang, Qianqian},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}