WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model
arXiv:2609.29106v1 Announce Type: cross Abstract: 3D foundation models recover video cameras and geometry in one forward pass, but some of the strongest are up to scale. Joint people-scene reconstruction then requires two missing outputs: metric scale and persistent person identity. We ask whether…