Siddharth Choudhary

dblp:119/1502 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
1since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 1 since 2021Systems, architecture and hardware · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 30% Robot navigation and mapping · 25% Trustworthy machine learning · 15%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 100%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › decoding
decoding strategy
0.812024
Multi-Modal Hallucination Control by Visual Information Grounding · CVPR 2024
Natural language and speech › Language models and text generation
hallucination mitigation
0.812024
Multi-Modal Hallucination Control by Visual Information Grounding · CVPR 2024
Machine learning › Trustworthy machine learning › hallucination
vision-language model hallucination
0.812024
Multi-Modal Hallucination Control by Visual Information Grounding · CVPR 2024
Computer vision › Vision and language
visual grounding
0.812024
Multi-Modal Hallucination Control by Visual Information Grounding · CVPR 2024
Computer vision › Video understanding and tracking › multi-object tracking
data association
0.312018
Data-Efficient Decentralized Visual SLAM · ICRA 2018
Robotics › Robot navigation and mapping › SLAM › multi-robot SLAM
distributed SLAM
0.312018
Data-Efficient Decentralized Visual SLAM · ICRA 2018
Robotics › Robot navigation and mapping › SLAM
multi-robot SLAM
0.312018
Data-Efficient Decentralized Visual SLAM · ICRA 2018
Robotics › Robot navigation and mapping › SLAM
visual SLAM
0.312018
Data-Efficient Decentralized Visual SLAM · ICRA 2018
Knowledge, reasoning and agents › Multi-agent systems
distributed estimation
0.212016
Distributed trajectory estimation with privacy and communication constraints: A two-stage distributed Gauss-Seidel approach · ICRA 2016
Robotics › Robot navigation and mapping
SLAM
0.212015
Information-based reduced landmark SLAM · ICRA 2015
Geometric modeling and processing › 3d reconstruction
structure from motion
0.112012
Visibility Probability Structure from SfM Datasets and Applications · ECCV (5) 2012
Distributed systems › distributed optimization
decentralized optimization
0.112018
Data-Efficient Decentralized Visual SLAM · ICRA 2018
Knowledge, reasoning and agents › Multi-agent systems › multi-robot systems
cooperative robots
0.112016
Distributed trajectory estimation with privacy and communication constraints: A two-stage distributed Gauss-Seidel approach · ICRA 2016
Robotics › Robot navigation and mapping › SLAM
graph-based SLAM
0.112015
Information-based reduced landmark SLAM · ICRA 2015
Computer vision › 3D vision
3d scene reconstruction
0.012012
Visibility Probability Structure from SfM Datasets and Applications · ECCV (5) 2012

Methods — techniques the papers use, named apart from their topics

mutual-information decoding · 0.8direct preference optimization · 0.8visibility analysis · 0.3structure from motion · 0.3maximum likelihood estimation · 0.2distributed gauss-seidel · 0.2information-theoretic reduction · 0.2incremental reduction · 0.2
YearPublicationVenuePosition
2024 Multi-Modal Hallucination Control by Visual Information Grounding
abstract
Generative Vision-Language Models (VLMs) are prone to generate plausible-sounding textual answers that, however, are not always grounded in the input image. We investigate this phenomenon, usually referred to as “hallucination” and show that it stems from an excessive reliance on the language prior. In particular, we show that as more tokens are generated, the reliance on the visual prompt decreases, and this behavior strongly correlates with the emergence of hallucinations. To reduce hallucinations, we introduce Multi-Modal Mutual-Information Decoding (M3ID), a new sampling method for prompt amplification. M3ID amplifies the influence of the reference image over the language prior, hence favoring the generation of tokens with higher mutual information with the visual prompt. M3ID can be applied to any pre-trained autoregressive VLM at inference time without necessitating further training and with minimal computational overhead. If training is an option, we show that M3ID can be paired with Direct Preference Optimization (DPO) to improve the model's reliance on the prompt image without requiring any labels. Our empirical findings show that our algorithms maintain the fluency and linguistic capabilities of pre-trained VLMs while reducing hallucinations by mitigating visually ungrounded answers. Specifically, for the LLaVA 13B model, M3ID and M3ID+DPO reduce the percentage of hallucinated objects in captioning tasks by 25% and 28%, respectively, and improve the accuracy on VQA benchmarks such as POPE by 21% and 24%.
Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pramuditha Perera, Alessandro Achille, Ashwin Swaminathan, Stefano Soatto
CVPR4
2018 Data-Efficient Decentralized Visual SLAM
abstract
Decentralized visual simultaneous localization and mapping (SLAM) is a powerful tool for multi-robot applications in environments where absolute positioning is not available. Being visual, it relies on cheap, lightweight and versatile cameras, and, being decentralized, it does not rely on communication to a central entity. In this work, we integrate state-of-the-art decentralized SLAM components into a new, complete decentralized visual SLAM system. To allow for data association and optimization, existing decentralized visual SLAM systems exchange the full map data among all robots, incurring large data transfers at a complexity that scales quadratically with the robot count. In contrast, our method performs efficient data association in two stages: first, a compact full-image descriptor is deterministically sent to only one robot. Then, only if the first stage succeeded, the data required for relative pose estimation is sent, again to only one robot. Thus, data association scales linearly with the robot count and uses highly compact place representations. For optimization, a state-of-the-art decentralized pose-graph optimization method is used. It exchanges a minimum amount of data which is linear with trajectory overlap. We characterize the resulting system and identify bottlenecks in its components. The system is evaluated on publicly available datasets and we provide open access to the code. Supplementary Material Data and code are at: https://github.com/uzh-rpg/dslam_open.
Titus Cieslewski, Siddharth Choudhary, Davide Scaramuzza 0001
ICRA2
2016 Distributed trajectory estimation with privacy and communication constraints: A two-stage distributed Gauss-Seidel approach
abstract
We propose a distributed algorithm to estimate the 3D trajectories of multiple cooperative robots from relative pose measurements. Our approach leverages recent results [1] which show that the maximum likelihood trajectory is well approximated by a sequence of two quadratic subproblems. The main contribution of the present work is to show that these subproblems can be solved in a distributed manner, using the distributed Gauss-Seidel (DGS) algorithm. Our approach has several advantages. It requires minimal information exchange, which is beneficial in presence of communication and privacy constraints. It has an anytime flavor: after few iterations the trajectory estimates are already accurate, and they asymptotically convergence to the centralized estimate. The DGS approach scales well to large teams, and it has a straightforward implementation. We test the approach in simulations and field tests, demonstrating its advantages over related techniques.
Siddharth Choudhary, Luca Carlone, Carlos Nieto-Granda, John G. Rogers III, Henrik I. Christensen, Frank Dellaert
ICRA1
2016 Active planning based extrinsic calibration of exteroceptive sensors in unknown environments
abstract
Existing Simultaneous Localization and Mapping systems require an extensive manual pre-calibration process. Non-manual calibration procedures use manipulators to create known patterns in order to estimate the unknown calibration. Calibration is often time-consuming and involves humans performing repetitive tasks such as aligning a known calibration target at different poses with respect to the sensor. We propose an algorithm that plans a trajectory which actively reduces the uncertainty of the robot's calibration given a rough initial calibration estimate. Calibration is performed autonomously in a previously unknown environment by maintaining the belief over landmarks, poses, and the calibration parameters. We present experimental results to demonstrate the approach's ability to autonomously calibrate the exteroceptive sensor in simulated and real environments. We show that even a greedy approach can reduce the effort needed to perform calibration every time the robot is reconfigured for autonomous tasks and mitigates the possibility of human error added into the calibration.
Varun Murali, Carlos Nieto-Granda, Siddharth Choudhary, Henrik I. Christensen
IROS3
2015 Information-based reduced landmark SLAM
abstract
In this paper, we present an information-based approach to select a reduced number of landmarks and poses for a robot to localize itself and simultaneously build an accurate map. We develop an information theoretic algorithm to efficiently reduce the number of landmarks and poses in a SLAM estimate without compromising the accuracy of the estimated trajectory. We also propose an incremental version of the reduction algorithm which can be used in SLAM framework resulting in information based reduced landmark SLAM. The results of reduced landmark based SLAM algorithm are shown on Victoria park dataset and a Synthetic dataset and are compared with standard graph SLAM (SAM [6]) algorithm. We demonstrate a reduction of 40-50% in the number of landmarks and around 55% in the number of poses with minimal estimation error as compared to standard SLAM algorithm.
Siddharth Choudhary, Vadim Indelman, Henrik I. Christensen, Frank Dellaert
ICRA1
2015 Exactly sparse memory efficient SLAM using the multi-block alternating direction method of multipliers
abstract
Large-scale SLAM demands for scalable techniques in which the computational burden and the memory consumption is shared among many processing units. While recent literature offers competitive approaches for scalable mapping, these usually involve approximations to preserve sparsity of the resulting subproblems. We present an approach to scalable SLAM that is exactly sparse. The main insight is that rather than eliminating variables (which induces dense cliques), we split the separators connecting subgraphs. Then, we enforce consistency of the separators in different subgraphs using hard constraints. The resulting constrained optimization problem can be solved in a decentralized manner using the multi-block Alternating Direction Method of Multipliers (ADMM). Our framework is appealing since (i) it preserves the sparsity structure of the original problem, (ii) it has a straightforward implementation, (iii) it allows to easily trade-off between computation time and accuracy. While our approach is currently slower than competitors, it is more accurate than other memory efficient alternatives. Moreover, we believe that the proposed framework can be of interest on its own as it draws connections with recent literature on decentralized optimization.
Siddharth Choudhary, Luca Carlone, Henrik I. Christensen, Frank Dellaert
IROS1
2014 SLAM with object discovery, modeling and mapping
abstract
Object discovery and modeling have been widely studied in the computer vision and robotics communities. SLAM approaches that make use of objects and higher level features have also recently been proposed. Using higher level features provides several benefits: these can be more discriminative, which helps data association, and can serve to inform service robotic tasks that require higher level information, such as object models and poses. We propose an approach for online object discovery and object modeling, and extend a SLAM system to utilize these discovered and modeled objects as landmarks to help localize the robot in an online manner. Such landmarks are particularly useful for detecting loop closures in larger maps. In addition to the map, our system outputs a database of detected object models for use in future SLAM or service robotic tasks. Experimental results are presented to demonstrate the approach's ability to detect and model objects, as well as to improve SLAM results by detecting loop closures.
Siddharth Choudhary, Alexander J. B. Trevor, Henrik I. Christensen, Frank Dellaert
IROS1
2013 An infrastructure for automating large-scale performance studies and data processing
abstract
The Cloud has enabled the computing model to shift from traditional data centers to publicly shared computing infrastructure; yet, applications leveraging this new computing model can experience performance and scalability issues, which arise from the hidden complexities of the cloud. The most reliable path for better understanding these complexities is an empirically based approach that relies on collecting data from a large number of performance studies. Armed with this performance data, we can understand what has happened, why it happened, and more importantly, predict what will happen in the future. However, this approach presents challenges itself, namely in the form of data management. We attempt to mitigate these data challenges by fully automating the performance measurement process. Concretely, we have developed an automated infrastructure, which reduces the complexity of the large-scale performance measurement process by generating all the necessary resources to conduct experiments, to collect and process data and to store and analyze data. In this paper, we focus on the performance data management aspect of our infrastructure.
Deepal Jayasinghe, Josh Kimball, Siddharth Choudhary, Calton Pu
IEEE BigData4
2012 Visibility Probability Structure from SfM Datasets and Applications
Siddharth Choudhary, P. J. Narayanan
ECCV (5)1