VLDB 2026 Research / reviewers in the wild / expert
Shubham Goel 0001
dblp:194/2742-1
· DBLP profile ↗
9ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-1700-939XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2Theory of computation · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Spatial Cognition from Egocentric Video: Out of Sight, Not Out of MindabstractAs humans move around, performing their daily tasks, they are able to recall where they have positioned objects in their environment, even if these objects are currently out of their sight. In this paper, we aim to mimic this spatial cognition ability. We thus formulate the task of Out of Sight, Not Out of Mind - 3D tracking active objects using observations captured through an egocentric camera. We introduce a simple but effective approach to address this challenging problem, called Lift, Match, and Keep (LMK). LMK lifts partial 2D observations to 3D world coordinates, matches them over time using visual appearance, 3D location and interactions to form object tracks, and keeps these object tracks even when they go out-of-view of the camera. We benchmark LMK on 100 long videos from EPICKITCHENS. Our results demonstrate that spatial cognition is critical for correctly locating objects over short and long time scales. E.g., for one long egocentric video, we estimate the 3D location of 50 active objects. After 120 seconds, 57 % of the objects are correctly localised by LMK, compared to just 33% by a recent 3D method for egocentric videos and 17 % by a general 2D tracking method. Chiara Plizzari, Shubham Goel 0001, Toby Perrett, Jacob Chalk, Angjoo Kanazawa, Dima Damen |
3DV | 2 |
| 2024 | The More You See in 2D, the More You Perceive in 3DabstractHumans can infer 3D structure from 2D images of an object based on past experience and improve their 3D understanding as they see more images. Inspired by this be-havior, we introduce SAP3D, a system for 3D reconstruction and novel view synthesis from an arbitrary number of un-posed images. Given a few unposed images of an object, we adapt a pre-trained view-conditioned diffusion model together with the camera poses of the images via test-time fine-tuning. The adapted diffusion model and the obtained camera poses are then utilized as instance-specific priors for 3D reconstruction and novel view synthesis. We show that as the number of input images increases, the performance of our approach improves, bridging the gap between optimization-based prior-less 3D reconstruction methods and single-image-to-3D diffusion-based methods. We demon-strate our system on real images as well as standard synthetic benchmarks. Our ablation studies confirm that this adaption behavior is key for more accurate 3D understanding.1 Zelin Gao, Angjoo Kanazawa, Shubham Goel 0001, Yossi Gandelsman |
CVPR | 4 |
| 2023 | Humans in 4D: Reconstructing and Tracking Humans with TransformersabstractWe present an approach to reconstruct humans and track them over time. At the core of our approach, we propose a fully "transformerized" version of a network for human mesh recovery. This network, HMR 2.0, advances the state of the art and shows the capability to analyze unusual poses that have in the past been difficult to reconstruct from single images. To analyze video, we use 3D reconstructions from HMR 2.0 as input to a tracking system that operates in 3D. This enables us to deal with multiple people and maintain identities through occlusion events. Our complete approach, 4DHumans, achieves state-of-the-art results for tracking people from monocular video. Furthermore, we demonstrate the effectiveness of HMR 2.0 on the downstream task of action recognition, achieving significant improvements over previous pose-based action recognition approaches. Our code and models are available on the project website: https://shubham-goel.github.io/4dhumans/. Shubham Goel 0001, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa, Jitendra Malik |
ICCV | 1 |
| 2022 | Differentiable Stereopsis: Meshes from multiple views using differentiable renderingabstractWe propose Differentiable Stereopsis, a multi-view stereo approach that reconstructs shape and texture from few input views and noisy cameras. We pair traditional stereopsis and modern differentiable rendering to build an end-to-end model which predicts textured 3D meshes of objects with varying topologies and shape. We frame stereopsis as an optimization problem and simultaneously update shape and cameras via simple gradient descent. We run an extensive quantitative analysis and compare to traditional multi-view stereo techniques and state-of-the-art learning based methods. We show compelling reconstructions on challenging real-world scenes and for an abundance of object types with complex shape, topology and texture.11Project webpage: https://shubham-goel.github.io/ds/ Shubham Goel 0001, Georgia Gkioxari, Jitendra Malik |
CVPR | 1 |
| 2022 | ABO: Dataset and Benchmarks for Real-World 3D Object UnderstandingabstractWe introduce Amazon Berkeley Objects (ABO), a new large-scale dataset designed to help bridge the gap between real and virtual 3D worlds. ABO contains product catalog images, metadata, and artist-created 3D models with com-plex geometries and physically-based materials that cor-respond to real, household objects. We derive challenging benchmarks that exploit the unique properties of ABO and measure the current limits of the state-of-the-art on three open problems for real-world 3D object understanding: single-view 3D reconstruction, material estimation, and cross-domain multi-view object retrieval. Jasmine Collins, Shubham Goel 0001, Kenan Deng 0001, Achleshwar Luthra, Leon Xu, Erhan Gundogdu, Tomás F. Yago Vicente, Thomas Dideriksen, Himanshu Arora, Matthieu Guillaumin, Jitendra Malik |
CVPR | 2 |
| 2021 | Boolean functional synthesis: hardness and practical algorithms
S. Akshay 0001, Supratik Chakraborty, Shubham Goel 0001, Sumith Kulal, Shetal Shah |
Formal Methods Syst. Des. | 3 |
| 2020 | Shape and Viewpoint Without Keypoints
Shubham Goel 0001, Angjoo Kanazawa, Jitendra Malik |
ECCV (15) | 1 |
| 2018 | What's Hard About Boolean Functional Synthesis?abstractGiven a relational specification between Boolean inputs and outputs, the goal of Boolean functional synthesis is to synthesize each output as a function of the inputs such that the specification is met. In this paper, we first show that unless some hard conjectures in complexity theory are falsified, Boolean functional synthesis must generate large Skolem functions in the worst-case. Given this inherent hardness, what does one do to solve the problem? We present a two-phase algorithm, where the first phase is efficient both in terms of time and size of synthesized functions, and solves a large fraction of benchmarks. To explain this surprisingly good performance, we provide a sufficient condition under which the first phase must produce correct answers. When this condition fails, the second phase builds upon the result of the first phase, possibly requiring exponential time and generating exponential-sized functions in the worst-case. Detailed experimental evaluation shows our algorithm to perform better than other techniques for a large number of benchmarks. S. Akshay 0001, Supratik Chakraborty, Shubham Goel 0001, Sumith Kulal, Shetal Shah |
CAV (1) | 3 |
| 2017 | Computing Scores of Forwarding Schemes in Switched Networks with Probabilistic Faults
Guy Avni, Shubham Goel 0001, Thomas A. Henzinger, Guillermo Rodríguez-Navas |
TACAS (2) | 2 |