VLDB 2026 Research / reviewers in the wild / expert
Chris Rockwell 0001
dblp:61/4160-1
· DBLP profile ↗
9ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0003-3510-5382ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamic Camera Poses and Where to Find ThemabstractAnnotating camera poses on dynamic Internet videos at scale is critical for advancing fields like realistic video generation and simulation. However, collecting such a dataset is difficult, as most Internet videos are unsuitable for pose estimation. Furthermore, annotating dynamic Internet videos present significant challenges even for state-of-the-art methods. In this paper, we introduce DynPose-100K, a large-scale dataset of dynamic Internet videos annotated with camera poses. Our collection pipeline addresses filtering using a carefully combined set of task-specific and generalist models. For pose estimation, we combine the latest techniques of point tracking, dynamic masking, and structure-from-motion to achieve improvements over the state-of-the-art approaches. Our analysis and experiments demonstrate that DynPose-100K is both large-scale and diverse across several key attributes, opening up avenues for advancements in various downstream applications. Chris Rockwell 0001, Joseph Tung, Tsung-Yi Lin, Ming-Yu Liu 0001, David F. Fouhey, Chen-Hsuan Lin 0001 |
CVPR | 1 |
| 2024 | FAR: Flexible, Accurate and Robust 6DoF Relative Camera Pose EstimationabstractEstimating relative camera poses between images has been a central problem in computer vision. Methods that find correspondences and solve for the fundamental matrix offer high precision in most cases. Conversely, methods predicting pose directly using neural networks are more robust to limited overlap and can infer absolute translation scale, but at the expense of reduced precision. We show how to combine the best of both methods; our approach yields results that are both precise and robust, while also accurately inferring translation scales. At the heart of our model lies a Transformer that (1) learns to balance between solved and learned pose estimations, and (2) provides a prior to guide a solver. A comprehensive analy-sis supports our design choices and demonstrates that our method adapts flexibly to various feature extractors and correspondence estimators, showing state-of-the-art performance in 6DoF pose estimation on Matterport3D, Inte-rio rNet, StreetLearn, and Map-free Relocalization. Project page: https://crockwell.github.io/farl Chris Rockwell 0001, Nilesh Kulkarni, Linyi Jin, Jeong Joon Park, Justin Johnson 0001, David F. Fouhey |
CVPR | 1 |
| 2023 | Scalable 3D Captioning with Pretrained ModelsabstractWe introduce Cap3D, an automatic approach for generating descriptive text for 3D objects. This approach utilizes pretrained models from image captioning, image-text alignment, and LLM to consolidate captions from multiple views of a 3D asset, completely side-stepping the time-consuming and costly process of manual annotation. We apply Cap3D to the recently introduced large-scale 3D dataset, Objaverse, resulting in 660k 3D-text pairs. Our evaluation, conducted using 41k human annotations from the same dataset, demonstrates that Cap3D surpasses human-authored descriptions in terms of quality, cost, and speed. Through effective prompt engineering, Cap3D rivals human performance in generating geometric descriptions on 17k collected annotations from the ABO dataset. Finally, we finetune Text-to-3D models on Cap3D and human captions, and show Cap3D outperforms; and benchmark the SOTA including Point·E, Shape·E, and DreamFusion. Tiange Luo, Chris Rockwell 0001, Honglak Lee, Justin Johnson 0001 |
NeurIPS | 2 |
| 2022 | The 8-Point Algorithm as an Inductive Bias for Relative Pose Prediction by ViTsabstractWe present a simple baseline for directly estimating the relative pose (rotation and translation, including scale) between two images. Deep methods have recently shown strong progress but often require complex or multi-stage architectures. We show that a handful of modifications can be applied to a Vision Transformer (ViT) to bring its computations close to the Eight-Point Algorithm. This inductive bias enables a simple method to be competitive in multiple settings, often substantially improving over the state of the art with strong performance gains in limited data regimes. Chris Rockwell 0001, Justin Johnson 0001, David F. Fouhey |
3DV | 1 |
| 2022 | Understanding 3D Object Articulation in Internet VideosabstractWe propose to investigate detecting and characterizing the 3D planar articulation of objects from ordinary RGB videos. While seemingly easy for humans, this problem poses many challenges for computers. Our approach is based on a top-down detection system that finds planes that can be articulated. This approach is followed by optimizing for a 3D plane that explains a sequence of detected articulations. We show that this system can be trained on a combination of videos and 3D scan datasets. When tested on a dataset of challenging Internet videos and the Charades dataset, our approach obtains strong performance. Shengyi Qian 0001, Linyi Jin, Chris Rockwell 0001, Siyi Chen 0003, David F. Fouhey |
CVPR | 3 |
| 2022 | FWD: Real-time Novel View Synthesis with Forward Warping and DepthabstractNovel view synthesis (NVS) is a challenging task requiring systems to generate photorealistic images of scenes from new viewpoints, where both quality and speed are important for applications. Previous image-based rendering (IBR) methods are fast, but have poor quality when input views are sparse. Recent Neural Radiance Fields (NeRF) and generalizable variants give impressive results but are not real-time. In our paper, we propose a generalizable NVS method with sparse inputs, called FWD, which gives high-quality synthesis in real-time. With explicit depth and differentiable rendering, it achieves competitive results to the SOTA methods with 130-1000× speedup and better perceptual quality. If available, we can seamlessly integrate sensor depth during either training or inference to improve image quality while retaining real-time speed. With the growing prevalence of depths sensors, we hope that methods making use of depth will become increasingly useful. Ang Cao, Chris Rockwell 0001, Justin Johnson 0001 |
CVPR | 2 |
| 2022 | PlaneFormers: From Sparse View Planes to 3D Reconstruction
Samir Agarwala, Linyi Jin, Chris Rockwell 0001, David F. Fouhey |
ECCV (3) | 3 |
| 2021 | PixelSynth: Generating a 3D-Consistent Experience from a Single ImageabstractRecent advancements in differentiable rendering and 3D reasoning have driven exciting results in novel view synthesis from a single image. Despite realistic results, methods are limited to relatively small view change. In order to synthesize immersive scenes, models must also be able to extrapolate. We present an approach that fuses 3D reasoning with autoregressive modeling to outpaint large view changes in a 3D-consistent manner, enabling scene synthesis. We demonstrate considerable improvement in single-image large-angle view synthesis results compared to a variety of methods and possible variants across simulated and real datasets. In addition, we show increased 3D consistency compared to alternative accumulation methods. Chris Rockwell 0001, David F. Fouhey, Justin Johnson 0001 |
ICCV | 1 |
| 2020 | Full-Body Awareness from Partial Observations
Chris Rockwell 0001, David F. Fouhey |
ECCV (17) | 1 |