VLDB 2026 Research / reviewers in the wild / expert
Abhishek Badki
dblp:166/4040
· DBLP profile ↗
9ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0002-5559-2336ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | L4P: Towards Unified Low-Level 4D Vision PerceptionabstractThe spatio-temporal relationship between the pixels of a video carries critical information for low-level 4D perception tasks. A single model that reasons about it should be able to solve several such tasks well. Yet, most state-of-theart methods rely on architectures specialized for the task at hand. We present L4P, a feedforward, general-purpose architecture that solves low-level 4D perception tasks in a unified framework. LAP leverages a pre-trained ViT-based video encoder and combines it with per-task heads that are lightweight and therefore do not require extensive training. Despite its general and feedforward formulation, our method is competitive with existing specialized methods on both dense tasks, such as depth or optical flow estimation, and sparse tasks, such as 2D/3D tracking. Moreover, it solves all tasks at once in a time comparable to that of single-task methods. Abhishek Badki, Hang Su 0005, Bowen Wen, Orazio Gallo |
3DV | 1 |
| 2025 | Zero-Shot Monocular Scene Flow Estimation in the WildabstractLarge models have shown generalization across datasets for many low-level vision tasks, like depth estimation, but no such general models exist for scene flow. Even though scene flow prediction has wide potential, its practical use is limited because of the lack of generalization of current predictive models. We identify three key challenges and propose solutions for each. First, we create a method that jointly estimates geometry and motion for accurate prediction. Second, we alleviate scene flow data scarcity with a data recipe that affords us 1M annotated training samples across diverse synthetic scenes. Third, we evaluate different parameterizations for scene flow prediction and adopt a natural and effective parameterization. Our model outperforms existing methods as well as baselines built on large-scale models in terms of 3D end-point error, and shows zero-shot generalization to the casually captured videos from DAVIS and the robotic manipulation scenes from RoboTAP. Overall, our approach makes scene flow prediction more practical in-the-wild.Website: https://research.nvidia.com/labs/lpr/zeromsf/ Yiqing Liang, Abhishek Badki, Hang Su 0005, James Tompkin 0001, Orazio Gallo |
CVPR | 2 |
| 2024 | FoVA-Depth: Field-of-View Agnostic Depth Estimation for Cross-Dataset GeneralizationabstractWide field-of-view (FoV) cameras efficiently capture large portions of the scene, which makes them attractive in multiple domains, such as automotive and robotics. For such applications, estimating depth from multiple images is a critical task, and therefore, a large amount of ground truth (GT) data is available. Unfortunately, most of the GT data is for pinhole cameras, making it impossible to properly train depth estimation models for large-FoV cameras. We propose the first method to train a stereo depth estimation model on the widely available pinhole data, and to generalize it to data captured with larger FoVs. Our intuition is simple: We warp the training data to a canonical, large-FoV representation and augment it to allow a single network to reason about diverse types of distortions that otherwise would prevent generalization. We show strong generalization ability of our approach on both indoor and outdoor datasets, which was not possible with previous methods. Daniel Lichy, Hang Su 0005, Abhishek Badki, Jan Kautz, Orazio Gallo |
3DV | 3 |
| 2024 | BlobGEN-3D: Compositional 3D-Consistent Freeview Image Generation with 3D Blobs
Chao Liu 0064, Weili Nie, Sifei Liu, Abhishek Badki, Hang Su 0005, Morteza Mardani, Benjamin Eckart, Arash Vahdat |
SIGGRAPH Asia | 4 |
| 2021 | Binary TTC: A Temporal Geofence for Autonomous NavigationabstractTime-to-contact (TTC), the time for an object to collide with the observer’s plane, is a powerful tool for path planning: it is potentially more informative than the depth, velocity, and acceleration of objects in the scene—even for humans. TTC presents several advantages, including requiring only a monocular, uncalibrated camera. However, regressing TTC for each pixel is not straightforward, and most existing methods make over-simplifying assumptions about the scene. We address this challenge by estimating TTC via a series of simpler, binary classifications. We predict with low latency whether the observer will collide with an obstacle within a certain time, which is often more critical than knowing exact, per-pixel TTC. For such scenarios, our method offers a temporal geofence in 6.4 ms—over 25× faster than existing methods. Our approach can also estimate per-pixel TTC with arbitrarily fine quantization (including continuous values), when the computational budget allows for it. To the best of our knowledge, our method is the first to offer TTC information (binary or coarsely quantized) at sufficiently high frame-rates for practical use. Abhishek Badki, Orazio Gallo, Jan Kautz, Pradeep Sen |
CVPR | 1 |
| 2020 | Meshlet Priors for 3D Mesh ReconstructionabstractEstimating a mesh from an unordered set of sparse, noisy 3D points is a challenging problem that requires to carefully select priors. Existing hand-crafted priors, such as smoothness regularizers, impose an undesirable trade-off between attenuating noise and preserving local detail. Recent deep-learning approaches produce impressive results by learning priors directly from the data. However, the priors are learned at the object level, which makes these algorithms class-specific, and even sensitive to the pose of the object. We introduce meshlets, small patches of mesh that we use to learn local shape priors. Meshlets act as a dictionary of local features and thus allow to use learned priors to reconstruct object meshes in any pose and from unseen classes, even when the noise is large and the samples sparse. Abhishek Badki, Orazio Gallo, Jan Kautz, Pradeep Sen |
CVPR | 1 |
| 2020 | Bi3D: Stereo Depth Estimation via Binary ClassificationsabstractStereo-based depth estimation is a cornerstone of computer vision, with state-of-the-art methods delivering accurate results in real time. For several applications such as autonomous navigation, however, it may be useful to trade accuracy for lower latency. We present Bi3D, a method that estimates depth via a series of binary classifications. Rather than testing if objects are at a particular depth D, as existing stereo methods do, it classifies them as being closer or farther than D. This property offers a powerful mechanism to balance accuracy and latency. Given a strict time budget, Bi3D can detect objects closer than a given distance in as little as a few milliseconds, or estimate depth with arbitrarily coarse quantization, with complexity linear with the number of quantization levels. Bi3D can also use the allotted quantization levels to get continuous depth, but in a specific depth range. For standard stereo (i.e., continuous depth on the whole range), our method is close to or on par with state-of-the-art, finely tuned stereo methods. Abhishek Badki, Alejandro J. Troccoli, Jan Kautz, Pradeep Sen, Orazio Gallo |
CVPR | 1 |
| 2017 | Computational zoom: a framework for post-capture image compositionabstractCapturing a picture that "tells a story" requires the ability to create the right composition. The two most important parameters controlling composition are the camera position and the focal length of the lens. The traditional paradigm is for a photographer to mentally visualize the desired picture, select the capture parameters to produce it, and finally take the photograph, thus committing to a particular composition. We propose to change this paradigm. To do this, we introduce computational zoom , a framework that allows a photographer to manipulate several aspects of composition in post-processing from a stack of pictures captured at different distances from the scene. We further define a multi-perspective camera model that can generate compositions that are not physically attainable, thus extending the photographer's control over factors such as the relative size of objects at different depths and the sense of depth of the picture. We show several applications and results of the proposed computational zoom framework. Abhishek Badki, Orazio Gallo, Jan Kautz, Pradeep Sen |
ACM Trans. Graph. | 1 |
| 2015 | Robust Radiometric Calibration for Dynamic Scenes in the WildabstractThe camera response function (CRF) that maps linear irradiance to pixel intensities must be known for computational imaging applications that match features in images with different exposures. This function is scene dependent and is difficult to estimate in scenes with significant motion. In this paper, we present a novel algorithm for radiometric calibration from multiple exposure images of a dynamic scene. Our approach is based on two key ideas from the literature: (1) intensity mapping functions which map pixel values in one image to the other without the need for pixel correspondences, and (2) a rank minimization algorithm for radiometric calibration. Although each method has its problems, we show how to combine them in a formulation that leverages their benefits. Our algorithm recovers the CRFs for dynamic scenes better than previous methods, and we show how it can be applied to existing algorithms such as those for high-dynamic range imaging to improve their results. Abhishek Badki, Nima Khademi Kalantari, Pradeep Sen |
ICCP | 1 |