EDBT 2026 Demo / reviewers in the wild / expert
Linyi Jin
dblp:245/7581
· DBLP profile ↗
11ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-0841-6970ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Stereo4D: Learning How Things Move in 3D from Internet Stereo VideosabstractLearning to understand dynamic 3D scenes from imagery is crucial for applications ranging from robotics to scene reconstruction. Yet, unlike other problems where large-scale supervised training has enabled rapid progress, directly supervising methods for recovering 3D motion remains challenging due to the fundamental difficulty of obtaining ground truth annotations. We present a system for mining high-quality 4D reconstructions from internet stereoscopic, wide-angle videos. Our system fuses and filters the outputs of camera pose estimation, stereo depth estimation, and temporal tracking methods into high-quality dynamic 3D reconstructions. We use this method to generate large-scale data in the form of world-consistent, pseudo-metric 3D point clouds with long-term motion trajectories. We demonstrate the utility of this data by training a variant of DUSt3R to predict structure and 3D motion from real-world image pairs, showing that training on our reconstructed data enables generalization to diverse real-world scenes. Project page and data at https://stereo4d.github.io Linyi Jin, Richard Tucker 0001, Zhengqi Li, David F. Fouhey, Noah Snavely, Aleksander Holynski |
CVPR | 1 |
| 2025 | MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic VideosabstractWe present a system that allows for accurate, fast, and robust estimation of camera parameters and depth maps from casual monocular videos of dynamic scenes. Most conventional structure from motion and monocular SLAM techniques assume input videos that feature predominantly static scenes with large amounts of parallax. Such methods tend to produce erroneous estimates in the absence of these conditions. Recent neural network-based approaches attempt to overcome these challenges; however, such methods are either computationally expensive or brittle when run on dynamic videos with uncontrolled camera motion or unknown field of view. We demonstrate the surprising effectiveness of a deep visual SLAM framework: with careful modifications to its training and inference schemes, this system can scale to real-world videos of complex dynamic scenes with unconstrained camera paths, including videos with little camera parallax. Extensive experiments on both synthetic and real videos demonstrate that our system is significantly more accurate and robust at camera pose and depth estimation when compared with prior and concurrent work, with faster or comparable running times. See interactive results on our project page: mega-sam.github.io. Zhengqi Li, Richard Tucker 0001, Forrester Cole, Qianqian Wang 0002, Linyi Jin, Vickie Ye, Angjoo Kanazawa, Aleksander Holynski, Noah Snavely |
CVPR | 5 |
| 2024 | FAR: Flexible, Accurate and Robust 6DoF Relative Camera Pose EstimationabstractEstimating relative camera poses between images has been a central problem in computer vision. Methods that find correspondences and solve for the fundamental matrix offer high precision in most cases. Conversely, methods predicting pose directly using neural networks are more robust to limited overlap and can infer absolute translation scale, but at the expense of reduced precision. We show how to combine the best of both methods; our approach yields results that are both precise and robust, while also accurately inferring translation scales. At the heart of our model lies a Transformer that (1) learns to balance between solved and learned pose estimations, and (2) provides a prior to guide a solver. A comprehensive analy-sis supports our design choices and demonstrates that our method adapts flexibly to various feature extractors and correspondence estimators, showing state-of-the-art performance in 6DoF pose estimation on Matterport3D, Inte-rio rNet, StreetLearn, and Map-free Relocalization. Project page: https://crockwell.github.io/farl Chris Rockwell 0001, Nilesh Kulkarni, Linyi Jin, Jeong Joon Park, Justin Johnson 0001, David F. Fouhey |
CVPR | 3 |
| 2024 | 3DFIRES: Few Image 3D REconstruction for Scenes with Hidden SurfacesabstractThis paper introduces 3DFIRES, a novel system for scene-level 3D reconstruction from posed images. Designed to work with as few as one view, 3DFIRES reconstructs the complete geometry of unseen scenes, including hidden surfaces. With multiple view inputs, our method pro-duces full reconstruction within all camera frustums. A key feature of our approach is the fusion of multi-view information at the feature level, enabling the production of coherent and comprehensive 3D reconstruction. We train our system on non-watertight scans from large-scale real scene dataset. We show it matches the efficacy of single-view reconstruction methods with only one input and surpasses existing techniques in both quantitative and qualitative measures for sparse-view 3D reconstruction. Project page: https://jinlinyi.github.io/3DFIRES/ Linyi Jin, Nilesh Kulkarni, David F. Fouhey |
CVPR | 1 |
| 2023 | Perspective Fields for Single Image Camera CalibrationabstractGeometric camera calibration is often required for applications that understand the perspective of the image. We propose Perspective Fields as a representation that models the local perspective properties of an image. Perspective Fields contain per-pixel information about the camera view, parameterized as an Up-vector and a Latitude value. This representation has a number of advantages; it makes minimal assumptions about the camera model and is invariant or equivariant to common image editing operations like cropping, warping, and rotation. It is also more interpretable and aligned with human perception. We train a neural network to predict Perspective Fields and the predicted Perspective Fields can be converted to calibration parameters easily. We demonstrate the robustness of our approach under various scenarios compared with camera calibration-based methods and show example applications in image compositing. Project page: https://jinlinyi.github.io/PerspectiveFields/. Linyi Jin, Jianming Zhang 0001, Yannick Hold-Geoffroy, Oliver Wang, Kevin Matzen, Matthew Sticha, David F. Fouhey |
CVPR | 1 |
| 2023 | Learning to Predict Scene-Level Implicit 3D from Posed RGBD DataabstractWe introduce a method that can learn to predict scenelevel implicit functions for 3D reconstruction from posed RGBD data. At test time, our system maps a previously unseen RGB image to a 3D reconstruction of a scene via implicit functions. While implicit functions for 3D reconstruction have often been tied to meshes, we show that we can train one using only a set of posed RGBD images. This setting may help 3D reconstruction unlock the sea of accelerometer+RGBD data that is coming with new phones. Our system, D2-DRDF, can match and sometimes outperform current methods that use mesh supervision and shows better robustness to sparse data. Nilesh Kulkarni, Linyi Jin, Justin Johnson 0001, David F. Fouhey |
CVPR | 2 |
| 2022 | Understanding 3D Object Articulation in Internet VideosabstractWe propose to investigate detecting and characterizing the 3D planar articulation of objects from ordinary RGB videos. While seemingly easy for humans, this problem poses many challenges for computers. Our approach is based on a top-down detection system that finds planes that can be articulated. This approach is followed by optimizing for a 3D plane that explains a sequence of detected articulations. We show that this system can be trained on a combination of videos and 3D scan datasets. When tested on a dataset of challenging Internet videos and the Charades dataset, our approach obtains strong performance. Shengyi Qian 0001, Linyi Jin, Chris Rockwell 0001, Siyi Chen 0003, David F. Fouhey |
CVPR | 2 |
| 2022 | PlaneFormers: From Sparse View Planes to 3D Reconstruction
Samir Agarwala, Linyi Jin, Chris Rockwell 0001, David F. Fouhey |
ECCV (3) | 2 |
| 2021 | Planar Surface Reconstruction from Sparse ViewsabstractThe paper studies planar surface reconstruction of indoor scenes from two views with unknown camera poses. While prior approaches have successfully created object-centric reconstructions of many scenes, they fail to exploit other structures, such as planes, which are typically the dominant components of indoor scenes. In this paper, we re-construct planar surfaces from multiple views, while jointly estimating camera pose. Our experiments demonstrate that our method is able to advance the state of the art of re-construction from sparse views, on challenging scenes from Matterport3D. Linyi Jin, Shengyi Qian 0001, Andrew Owens, David F. Fouhey |
ICCV | 1 |
| 2020 | Associative3D: Volumetric Reconstruction from Sparse Views
Shengyi Qian 0001, Linyi Jin, David F. Fouhey |
ECCV (15) | 2 |
| 2019 | Inferring Occluded Geometry Improves Performance When Retrieving an Object from Dense Clutter
Linyi Jin, Dmitry Berenson |
ISRR | 2 |