EDBT 2026 Demo / reviewers in the wild / expert
Sparsh Garg
dblp:313/1181
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Image recognition and object detection · 18% Transfer learning and domain adaptation · 17% 3D vision · 16% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Autoscape: Geometry-Consistent Long-Horizon Scene Generation · ICCV 2025 |
Robotics › Motion planning and robot control › robot learning
manipulation skill learning |
0.9 | 1 | 2025 | SplatSim: Zero-Shot Sim2Real Transfer of RGB Manipulation Policies Using Gaussian Splatting · ICRA 2025 |
Computer vision › 3D vision › depth estimation › monocular depth estimation
metric depth estimation |
0.9 | 1 | 2025 | Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera · CVPR 2025 |
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer |
0.9 | 1 | 2025 | SplatSim: Zero-Shot Sim2Real Transfer of RGB Manipulation Policies Using Gaussian Splatting · ICRA 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model reasoning
zero-shot reasoning |
0.9 | 1 | 2025 | iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning · NeurIPS 2025 |
Visual content generation and editing
video generation |
0.9 | 1 | 2025 | Autoscape: Geometry-Consistent Long-Horizon Scene Generation · ICCV 2025 |
Computer vision › Image recognition and object detection › object detection › robust object detection
long-tailed object detection |
0.8 | 1 | 2024 | AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving · CVPR 2024 |
Computer vision › Image recognition and object detection
object detection |
0.8 | 1 | 2024 | AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving · CVPR 2024 |
Computer vision › Image recognition and object detection › object detection
open-world object detection |
0.8 | 1 | 2024 | AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving · CVPR 2024 |
Robotics › Autonomous driving
perception |
0.8 | 1 | 2024 | AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving · CVPR 2024 |
Computer vision › Segmentation and scene understanding
3d semantic segmentation |
0.6 | 1 | 2022 | MM-TTA: Multi-Modal Test-Time Adaptation for 3D Semantic Segmentation · CVPR 2022 |
Machine learning › Transfer learning and domain adaptation › test-time adaptation
multi-modal test-time adaptation |
0.6 | 1 | 2022 | MM-TTA: Multi-Modal Test-Time Adaptation for 3D Semantic Segmentation · CVPR 2022 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.6 | 1 | 2022 | Learning Semantic Segmentation from Multiple Datasets with Label Shifts · ECCV (28) 2022 |
Machine learning › Transfer learning and domain adaptation
test-time adaptation |
0.6 | 1 | 2022 | MM-TTA: Multi-Modal Test-Time Adaptation for 3D Semantic Segmentation · CVPR 2022 |
Robotics › Autonomous driving › scenario generation
driving scene generation |
0.3 | 1 | 2025 | Autoscape: Geometry-Consistent Long-Horizon Scene Generation · ICCV 2025 |
Computer vision › 3D vision
photorealistic rendering |
0.3 | 1 | 2025 | SplatSim: Zero-Shot Sim2Real Transfer of RGB Manipulation Policies Using Gaussian Splatting · ICRA 2025 |
Machine learning › Efficient and distributed learning
auto labeling |
0.2 | 1 | 2024 | AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving · CVPR 2024 |
Machine learning › Efficient and distributed learning
data-efficient learning |
0.2 | 1 | 2024 | AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving · CVPR 2024 |
Machine learning › Transfer learning and domain adaptation › label shift
label shift adaptation |
0.2 | 1 | 2022 | Learning Semantic Segmentation from Multiple Datasets with Label Shifts · ECCV (28) 2022 |
Methods — techniques the papers use, named apart from their topics
warp-consistent guidance · 1.7video interpolation · 1.7diffusion model · 1.7zero-shot transfer · 0.9prompting · 0.9pretrained vision models · 0.9multi-resolution data augmentation · 0.9gaussian splatting · 0.9equirectangular projection · 0.9RGB manipulation policy · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Depth Any Camera: Zero-Shot Metric Depth Estimation from Any CameraabstractWhile recent depth foundation models exhibit strong zero-shot generalization, achieving accurate metric depth across diverse camera types—particularly those with large fields of view (FoV) such as fisheye and 360-degree cameras—remains a significant challenge. This paper presents Depth Any Camera (DAC), a powerful zero-shot metric depth estimation framework that extends a perspective-trained model to effectively handle cameras with varying FoVs. The framework is designed to ensure that all existing 3D data can be leveraged, regardless of the specific camera types used in new applications. Remarkably, DAC is trained exclusively on perspective images but generalizes seamlessly to fisheye and 360-degree cameras without the need for specialized training data. DAC employs Equi-Rectangular Projection (ERP) as a unified image representation, enabling consistent processing of images with diverse FoVs. Its core components include pitch-aware Image-to-ERP conversion with efficient online augmentation to simulate distorted ERP patches from undistorted inputs, FoV alignment operations to enable effective training across a wide range of FoVs, and multi-resolution data augmentation to further address resolution disparities between training and testing. DAC achieves state-of-the-art zero-shot metric depth estimation, improving δ1accuracy by up to 50% on multiple fisheye and 360-degree datasets compared to prior metric depth foundation models, demonstrating robust generalization across camera types. Yuliang Guo, Sparsh Garg, S. Mahdi H. Miangoleh, Xinyu Huang 0001, Liu Ren 0001 |
CVPR | 2 |
| 2025 | Autoscape: Geometry-Consistent Long-Horizon Scene GenerationabstractThis paper proposes AutoScape, a long-horizon driving scene generation framework. At its core is a novel RGB-D diffusion model that iteratively generates sparse, geometrically consistent keyframes, serving as reliable anchors for the scene's appearance and geometry. To maintain long-range geometric consistency, the model 1) jointly handles image and depth in a shared latent space, 2) explicitly conditions on the existing scene geometry (i.e., rendered point clouds) from previously generated keyframes, and 3) steers the sampling process with a warp-consistent guidance. Given high-quality RGB-D keyframes, a video diffusion model then interpolates between them to produce dense and coherent video frames. AutoScape generates realistic and geometrically consistent driving videos of over 20 seconds, improving the long-horizon FID and FVD scores over the prior state-of-the-art by 48.6\% and 43.0\%, respectively. Ziyu Jiang, Mingfu Liang, Bingbing Zhuang, Jong-Chyi Su, Sparsh Garg, Ying Wu 0001, Manmohan Krishna Chandraker |
ICCV | 6 |
| 2025 | SplatSim: Zero-Shot Sim2Real Transfer of RGB Manipulation Policies Using Gaussian SplattingabstractSim2Real transfer, particularly for manipulation policies relying on RGB images, remains a critical challenge in robotics due to the significant domain shift between syn-thetic and real-world visual data. In this paper, we propose SplatSim, a novel framework that leverages Gaussian Splatting as the primary rendering primitive to reduce the Sim2Real gap for RGB-based manipulation policies. By replacing traditional mesh representations with Gaussian Splats in simulators, SplatSim produces highly photorealistic synthetic data while maintaining the scalability and cost-efficiency of simulation. We demonstrate the effectiveness of our framework by training manipulation policies within SplatSim and deploying them in the real world in a zero-shot manner, achieving an average success rate of 86.25%, compared to 97.5% for policies trained on real-world data. Videos can be found on our project page: https://splatsim.github.io Mohammad Nomaan Qureshi, Sparsh Garg, Francisco Yandún, David Held, George Kantor, Abhisesh Silwal |
ICRA | 2 |
| 2025 | iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video ReasoningabstractGrounding large language models (LLMs) in domain-specific tasks like post-hoc dash-cam driving video analysis is challenging due to their general-purpose training and lack of structured inductive biases. As vision is often the sole modality available for such analysis (i.e., no LiDAR, GPS, etc.), existing video-based vision-language models (V-VLMs) struggle with spatial reasoning, causal inference, and explainability of events in the input video. To this end, we introduce iFinder, a structured semantic grounding framework that decouples perception from reasoning by translating dash-cam videos into a hierarchical, interpretable data structure for LLMs. iFinder operates as a modular, training-free pipeline that employs pretrained vision models to extract critical cues—object pose, lane positions, and object trajectories—which are hierarchically organized into frame- and video-level structures. Combined with a three-block prompting strategy, it enables step-wise, grounded reasoning for the LLM to refine a peer V-VLM's outputs and provide accurate reasoning.
Evaluations on four public dash-cam video benchmarks show that iFinder's proposed grounding with domain-specific cues—especially object orientation and global context—significantly outperforms end-to-end V-VLMs on four zero-shot driving benchmarks, with up to 39% gains in accident reasoning accuracy. By grounding LLMs with driving domain-specific representations, iFinder offers a zero-shot, interpretable, and reliable alternative to end-to-end V-VLMs for post-hoc driving video understanding. Manyi Yao, Bingbing Zhuang, Sparsh Garg, Amit K. Roy-Chowdhury, Christian R. Shelton, Manmohan Krishna Chandraker, Abhishek Aich |
NeurIPS | 3 |
| 2024 | AIDE: An Automatic Data Engine for Object Detection in Autonomous DrivingabstractAutonomous vehicle (AV) systems rely on robust perception models as a cornerstone of safety assurance. However, objects encountered on the road exhibit a long-tailed distri-bution, with rare or unseen categories posing challenges to a deployed perception model. This necessitates an expen-sive process of continuously curating and annotating data with significant human effort. We propose to leverage recent advances in vision-language and large language models to design an Automatic Data Engine (AIDE) that automati-cally identifies issues, efficiently curates data, improves the model through auto-labeling, and verifies the model through generation of diverse scenarios. This process operates it-eratively, allowing for continuous self-improvement of the model. We further establish a benchmark for open-world detection on AV datasets to comprehensively evaluate vari-ous learning paradigms, demonstrating our method's supe-rior performance at a reduced cost. Mingfu Liang, Jong-Chyi Su, Samuel Schulter, Sparsh Garg, Shiyu Zhao 0001, Ying Wu 0001, Manmohan Krishna Chandraker |
CVPR | 4 |
| 2022 | MM-TTA: Multi-Modal Test-Time Adaptation for 3D Semantic SegmentationabstractTest-time adaptation approaches have recently emerged as a practical solution for handling domain shift without access to the source domain data. In this paper, we propose and explore a new multi-modal extension of test-time adaptation for 3D semantic segmentation. We find that, directly applying existing methods usually results in performance instability at test time, because multi-modal input is not considered jointly. To design a framework that can take full advantage of multi-modality, where each modality provides regularized self-supervisory signals to other modalities, we propose two complementary modules within and across the modalities. First, Intra-modal Pseudo-label Generation (Intra-PG) is introduced to obtain reliable pseudo labels within each modality by aggregating information from two models that are both pre-trained on source data but updated with target data at different paces. Second, Inter-modal Pseudo-label Refinement (Inter-PR) adaptively selects more reliable pseudo labels from different modalities based on a proposed consistency scheme. Experiments demonstrate that our regularized pseudo labels produce stable self-learning signals in numerous multi-modal test-time adaptation scenarios for 3D semantic segmentation. Visit our project website at https://www.nec-labs.com/~mas/MM-TTA Inkyu Shin, Yi-Hsuan Tsai, Bingbing Zhuang, Samuel Schulter, Buyu Liu, Sparsh Garg, In-So Kweon, Kuk-Jin Yoon |
CVPR | 6 |
| 2022 | Learning Semantic Segmentation from Multiple Datasets with Label Shifts
Dongwan Kim, Yi-Hsuan Tsai, Yumin Suh, Masoud Faraki, Sparsh Garg, Manmohan Krishna Chandraker, Bohyung Han |
ECCV (28) | 5 |