Videos

Fast and Robust 4D Reconstruction via Point Track Processing and Attention Chains

Presenter
July 14, 2026
Abstract
Recovering dynamic point clouds and 4D meshes from casual videos is a central goal in computer vision, but existing optimization-based methods remain slow and computationally expensive. This talk presents two recent approaches that accelerate 4D reconstruction by leveraging neural correspondences and the underlying mathematical structures of real motion. The first approach, TracksTo4D, is a learning-based method that infers dynamic point clouds and cameras via a single feed-forward pass. Trained without 3D supervision by minimizing 2D reprojection errors, the network operates directly over sparse point tracks. By modeling movement with a low-rank approximation that explicitly decomposes scenes into static and dynamic components, our method achieves robust temporal point cloud and camera reconstruction at a fraction of the computational cost of traditional optimization techniques. The second approach introduces a training-free 4D mesh generation framework driven by Spatio-Temporal Attention Chains. By treating attention matrices within a generative backbone as soft transition maps, our method establishes dense correspondences across space and time without explicit feature matching. To animate the mesh, sparse landmarks are tracked through these chains and their trajectories are robustly smoothed, with the resulting motion propagated via geodesic rigid skinning. Solving a weighted Procrustes alignment to compute local rigid transformations ensures volume preservation and generates topologically consistent 4D meshes in seconds. Ultimately, our approach scales reliably to longer sequences and enables camera estimation alongside 4D tracking.