§1 Abstract
Recent image-to-3D scene methods recover high-fidelity 3D objects with plausible arrangements, but often leave floatings and interpenetrations that limit physical validity and downstream use in interactive environments.
We present $\phi$-Scene, a physically grounded approach for open-vocabulary and compositional image-to-3D scene reconstruction that treats a scene not merely as a set of objects with predicted poses, but as a globally stable physical system. $\phi$-Scene formulates reconstruction as topology-driven physical assembly: given an initial scene of reconstructed objects, it infers how objects support one another and settles them one by one in topological order. For each object, SDF-based optimization first resolves penetrations against the already-settled support context, and rigid-body simulation then settles the object into a stable equilibrium under real-world physical constraints. The resulting scene stays aligned to the reference image, with every object resting at a physically valid, stable contact configuration.
On the 3D-Front benchmark, $\phi$-Scene achieves the strongest overall performance among out-of-domain methods and remains highly competitive with in-domain baselines on standard reconstruction metrics. Human and MLLM studies prefer $\phi$-Scene in visual quality, reference alignment, and physical plausibility. Dedicated physical metrics show that it substantially reduces penetration artifacts and yields much lower post-simulation drift.
To our knowledge, $\phi$-Scene is among the first image-to-3D scene reconstruction methods that explicitly reaches dynamic rigid-body equilibrium while preserving reference alignment.
§2 Interactive Comparison: Static
Drag to rotate, scroll to zoom, and right-drag to pan.
§3 Interactive Comparison: Dynamic
Drag to rotate, scroll to zoom, and right-drag to pan.
§4 BibTeX
@article{li2026phi,
title={$\phi$-Scene: Physically Grounded Image-to-3D Scene Reconstruction},
author={Li, Haodong and Shao, Lulu and Lu, Haolin and Fu, Yu and Chen, Yen-Ru and Jain, Seemandhar and Chandraker, Manmohan},
journal={arXiv preprint arXiv:2606.21596},
year={2026}
}