Consistent depth of moving objects in video
Name
3450626.3459871.pdf
Size
24.42 MB
Format
Adobe PDF
Checksum (MD5)
9aefd24bdbc2cd68ee616be38f3205c5
Author(s) • • • •
Zhang, Zhoutong
Cole, Forrester
Tucker, Richard
Freeman, William T.
Dekel, Tali
Date Issued
August 2021
Publisher
Association for Computing Machinery (ACM)
Citation
ACM Transactions on Graphics, Volume 40, Issue 4August 2021 Article No.: 148 pp 1–12
Version
Final published version
Abstract
We present a method to estimate depth of a dynamic scene, containing arbitrary moving objects, from an ordinary video captured with a moving camera. We seek a geometrically and temporally consistent solution to this under-constrained problem: the depth predictions of corresponding points across frames should induce plausible, smooth motion in 3D. We formulate this objective in a new test-time training framework where a depth-prediction CNN is trained in tandem with an auxiliary scene-flow prediction MLP over the entire input video. By recursively unrolling the scene-flow prediction MLP over varying time steps, we compute both short-range scene flow to impose local smooth motion priors directly in 3D, and long-range scene flow to impose multi-view consistency constraints with wide baselines. We demonstrate accurate and temporally coherent results on a variety of challenging videos containing diverse moving objects (pets, people, cars), as well as camera motion. Our depth maps give rise to a number of depth-and-motion aware video editing effects such as object and lighting insertion.
Subjects
Computer Graphics and Computer-Aided Design
MIT Department
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Terms of Use
Creative Commons Attribution 4.0 International license
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1145/3450626.3459871