Sparsh Garg

dblp:313/1181 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Image recognition and object detection · 18% Transfer learning and domain adaptation · 17% 3D vision · 16%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
Autoscape: Geometry-Consistent Long-Horizon Scene Generation · ICCV 2025
Robotics › Motion planning and robot control › robot learning
manipulation skill learning
0.912025
SplatSim: Zero-Shot Sim2Real Transfer of RGB Manipulation Policies Using Gaussian Splatting · ICRA 2025
Computer vision › 3D vision › depth estimation › monocular depth estimation
metric depth estimation
0.912025
Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera · CVPR 2025
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.912025
SplatSim: Zero-Shot Sim2Real Transfer of RGB Manipulation Policies Using Gaussian Splatting · ICRA 2025
Computer vision › Vision and language
vision-language model
0.912025
iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model reasoning
zero-shot reasoning
0.912025
iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning · NeurIPS 2025
Visual content generation and editing
video generation
0.912025
Autoscape: Geometry-Consistent Long-Horizon Scene Generation · ICCV 2025
Computer vision › Image recognition and object detection › object detection › robust object detection
long-tailed object detection
0.812024
AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving · CVPR 2024
Computer vision › Image recognition and object detection
object detection
0.812024
AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving · CVPR 2024
Computer vision › Image recognition and object detection › object detection
open-world object detection
0.812024
AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving · CVPR 2024
Robotics › Autonomous driving
perception
0.812024
AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving · CVPR 2024
Computer vision › Segmentation and scene understanding
3d semantic segmentation
0.612022
MM-TTA: Multi-Modal Test-Time Adaptation for 3D Semantic Segmentation · CVPR 2022
Machine learning › Transfer learning and domain adaptation › test-time adaptation
multi-modal test-time adaptation
0.612022
MM-TTA: Multi-Modal Test-Time Adaptation for 3D Semantic Segmentation · CVPR 2022
Computer vision › Segmentation and scene understanding
semantic segmentation
0.612022
Learning Semantic Segmentation from Multiple Datasets with Label Shifts · ECCV (28) 2022
Machine learning › Transfer learning and domain adaptation
test-time adaptation
0.612022
MM-TTA: Multi-Modal Test-Time Adaptation for 3D Semantic Segmentation · CVPR 2022
Robotics › Autonomous driving › scenario generation
driving scene generation
0.312025
Autoscape: Geometry-Consistent Long-Horizon Scene Generation · ICCV 2025
Computer vision › 3D vision
photorealistic rendering
0.312025
SplatSim: Zero-Shot Sim2Real Transfer of RGB Manipulation Policies Using Gaussian Splatting · ICRA 2025
Machine learning › Efficient and distributed learning
auto labeling
0.212024
AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving · CVPR 2024
Machine learning › Efficient and distributed learning
data-efficient learning
0.212024
AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving · CVPR 2024
Machine learning › Transfer learning and domain adaptation › label shift
label shift adaptation
0.212022
Learning Semantic Segmentation from Multiple Datasets with Label Shifts · ECCV (28) 2022

Methods — techniques the papers use, named apart from their topics

warp-consistent guidance · 1.7video interpolation · 1.7diffusion model · 1.7zero-shot transfer · 0.9prompting · 0.9pretrained vision models · 0.9multi-resolution data augmentation · 0.9gaussian splatting · 0.9equirectangular projection · 0.9RGB manipulation policy · 0.9
YearPublicationVenuePosition
2025 Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera
abstract
While recent depth foundation models exhibit strong zero-shot generalization, achieving accurate metric depth across diverse camera types—particularly those with large fields of view (FoV) such as fisheye and 360-degree cameras—remains a significant challenge. This paper presents Depth Any Camera (DAC), a powerful zero-shot metric depth estimation framework that extends a perspective-trained model to effectively handle cameras with varying FoVs. The framework is designed to ensure that all existing 3D data can be leveraged, regardless of the specific camera types used in new applications. Remarkably, DAC is trained exclusively on perspective images but generalizes seamlessly to fisheye and 360-degree cameras without the need for specialized training data. DAC employs Equi-Rectangular Projection (ERP) as a unified image representation, enabling consistent processing of images with diverse FoVs. Its core components include pitch-aware Image-to-ERP conversion with efficient online augmentation to simulate distorted ERP patches from undistorted inputs, FoV alignment operations to enable effective training across a wide range of FoVs, and multi-resolution data augmentation to further address resolution disparities between training and testing. DAC achieves state-of-the-art zero-shot metric depth estimation, improving δ1accuracy by up to 50% on multiple fisheye and 360-degree datasets compared to prior metric depth foundation models, demonstrating robust generalization across camera types.
Yuliang Guo, Sparsh Garg, S. Mahdi H. Miangoleh, Xinyu Huang 0001, Liu Ren 0001
CVPR2
2025 Autoscape: Geometry-Consistent Long-Horizon Scene Generation
abstract
This paper proposes AutoScape, a long-horizon driving scene generation framework. At its core is a novel RGB-D diffusion model that iteratively generates sparse, geometrically consistent keyframes, serving as reliable anchors for the scene's appearance and geometry. To maintain long-range geometric consistency, the model 1) jointly handles image and depth in a shared latent space, 2) explicitly conditions on the existing scene geometry (i.e., rendered point clouds) from previously generated keyframes, and 3) steers the sampling process with a warp-consistent guidance. Given high-quality RGB-D keyframes, a video diffusion model then interpolates between them to produce dense and coherent video frames. AutoScape generates realistic and geometrically consistent driving videos of over 20 seconds, improving the long-horizon FID and FVD scores over the prior state-of-the-art by 48.6\% and 43.0\%, respectively.
Ziyu Jiang, Mingfu Liang, Bingbing Zhuang, Jong-Chyi Su, Sparsh Garg, Ying Wu 0001, Manmohan Krishna Chandraker
ICCV6
2025 SplatSim: Zero-Shot Sim2Real Transfer of RGB Manipulation Policies Using Gaussian Splatting
abstract
Sim2Real transfer, particularly for manipulation policies relying on RGB images, remains a critical challenge in robotics due to the significant domain shift between syn-thetic and real-world visual data. In this paper, we propose SplatSim, a novel framework that leverages Gaussian Splatting as the primary rendering primitive to reduce the Sim2Real gap for RGB-based manipulation policies. By replacing traditional mesh representations with Gaussian Splats in simulators, SplatSim produces highly photorealistic synthetic data while maintaining the scalability and cost-efficiency of simulation. We demonstrate the effectiveness of our framework by training manipulation policies within SplatSim and deploying them in the real world in a zero-shot manner, achieving an average success rate of 86.25%, compared to 97.5% for policies trained on real-world data. Videos can be found on our project page: https://splatsim.github.io
Mohammad Nomaan Qureshi, Sparsh Garg, Francisco Yandún, David Held, George Kantor, Abhisesh Silwal
ICRA2
2025 iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning
abstract
Grounding large language models (LLMs) in domain-specific tasks like post-hoc dash-cam driving video analysis is challenging due to their general-purpose training and lack of structured inductive biases. As vision is often the sole modality available for such analysis (i.e., no LiDAR, GPS, etc.), existing video-based vision-language models (V-VLMs) struggle with spatial reasoning, causal inference, and explainability of events in the input video. To this end, we introduce iFinder, a structured semantic grounding framework that decouples perception from reasoning by translating dash-cam videos into a hierarchical, interpretable data structure for LLMs. iFinder operates as a modular, training-free pipeline that employs pretrained vision models to extract critical cues—object pose, lane positions, and object trajectories—which are hierarchically organized into frame- and video-level structures. Combined with a three-block prompting strategy, it enables step-wise, grounded reasoning for the LLM to refine a peer V-VLM's outputs and provide accurate reasoning. Evaluations on four public dash-cam video benchmarks show that iFinder's proposed grounding with domain-specific cues—especially object orientation and global context—significantly outperforms end-to-end V-VLMs on four zero-shot driving benchmarks, with up to 39% gains in accident reasoning accuracy. By grounding LLMs with driving domain-specific representations, iFinder offers a zero-shot, interpretable, and reliable alternative to end-to-end V-VLMs for post-hoc driving video understanding.
Manyi Yao, Bingbing Zhuang, Sparsh Garg, Amit K. Roy-Chowdhury, Christian R. Shelton, Manmohan Krishna Chandraker, Abhishek Aich
NeurIPS3
2024 AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving
abstract
Autonomous vehicle (AV) systems rely on robust perception models as a cornerstone of safety assurance. However, objects encountered on the road exhibit a long-tailed distri-bution, with rare or unseen categories posing challenges to a deployed perception model. This necessitates an expen-sive process of continuously curating and annotating data with significant human effort. We propose to leverage recent advances in vision-language and large language models to design an Automatic Data Engine (AIDE) that automati-cally identifies issues, efficiently curates data, improves the model through auto-labeling, and verifies the model through generation of diverse scenarios. This process operates it-eratively, allowing for continuous self-improvement of the model. We further establish a benchmark for open-world detection on AV datasets to comprehensively evaluate vari-ous learning paradigms, demonstrating our method's supe-rior performance at a reduced cost.
Mingfu Liang, Jong-Chyi Su, Samuel Schulter, Sparsh Garg, Shiyu Zhao 0001, Ying Wu 0001, Manmohan Krishna Chandraker
CVPR4
2022 MM-TTA: Multi-Modal Test-Time Adaptation for 3D Semantic Segmentation
abstract
Test-time adaptation approaches have recently emerged as a practical solution for handling domain shift without access to the source domain data. In this paper, we propose and explore a new multi-modal extension of test-time adaptation for 3D semantic segmentation. We find that, directly applying existing methods usually results in performance instability at test time, because multi-modal input is not considered jointly. To design a framework that can take full advantage of multi-modality, where each modality provides regularized self-supervisory signals to other modalities, we propose two complementary modules within and across the modalities. First, Intra-modal Pseudo-label Generation (Intra-PG) is introduced to obtain reliable pseudo labels within each modality by aggregating information from two models that are both pre-trained on source data but updated with target data at different paces. Second, Inter-modal Pseudo-label Refinement (Inter-PR) adaptively selects more reliable pseudo labels from different modalities based on a proposed consistency scheme. Experiments demonstrate that our regularized pseudo labels produce stable self-learning signals in numerous multi-modal test-time adaptation scenarios for 3D semantic segmentation. Visit our project website at https://www.nec-labs.com/~mas/MM-TTA
Inkyu Shin, Yi-Hsuan Tsai, Bingbing Zhuang, Samuel Schulter, Buyu Liu, Sparsh Garg, In-So Kweon, Kuk-Jin Yoon
CVPR6
2022 Learning Semantic Segmentation from Multiple Datasets with Label Shifts
Dongwan Kim, Yi-Hsuan Tsai, Yumin Suh, Masoud Faraki, Sparsh Garg, Manmohan Krishna Chandraker, Bohyung Han
ECCV (28)5