EDBT 2026 Demo / reviewers in the wild / expert
Paloma Sodhi
dblp:124/0550
· DBLP profile ↗
13ranked-venue papers
7as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 7 first-author · 7 since 2021Systems, architecture and hardware · 10 · 6 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Better than Your Teacher: LLM Agents that learn from Privileged AI FeedbackabstractWhile large language models (LLMs) show impressive decision-making abilities, current methods lack a mechanism for automatic self-improvement from errors during task execution. We propose LEAP, an iterative fine-tuning framework that continually improves LLM agents using feedback from AI expert teachers. Our key insight is to equip the expert teachers with a privileged state -- information available during training but hidden at test time. This allows even weak experts to provide precise guidance, significantly improving the student agent's performance without access to privileged information at test time.
We evaluate LEAP on multiple decision-making benchmarks, including text-based games (ALFWorld), web navigation (WebShop), and interactive coding (Intercode Bash). Our experiments show that LEAP (1) outperforms behavior cloning and ReAct baselines (2) enables weak student models (e.g., Llama3-8B) to exceed performance of strong teacher models (GPT-4o), and (3) allows weak models to self-improve using privileged versions of themselves.
We provide a theoretical analysis showing that LEAP's success hinges on balancing privileged information with student’s realizability, which we empirically validate. Our code is available at \url{https://leap-llm.github.io}. Sanjiban Choudhury, Paloma Sodhi |
ICLR | 2 |
| 2023 | On the Effectiveness of Offline RL for Dialogue Response GenerationabstractA common training technique for language models is teacher forcing (TF). TF attempts to match human language exactly, even though identical meanings can be expressed in different ways. This motivates use of sequence-level objectives for dialogue response generation. In this paper, we study the efficacy of various offline reinforcement learning (RL) methods to maximize such objectives. We present a comprehensive evaluation across multiple datasets, models, and metrics. Offline RL shows a clear performance improvement over teacher forcing while not inducing training instability or sacrificing practical training budgets. Paloma Sodhi, Felix Wu, Ethan R. Elenberg, Kilian Q. Weinberger, Ryan McDonald |
ICML | 1 |
| 2022 | PatchGraph: In-hand tactile tracking with learned surface normalsabstractWe address the problem of tracking 3D object poses from touch during in-hand manipulations. Specifically, we look at tracking small objects using vision-based tactile sensors that provide high-dimensional tactile image measurements at the point of contact. While prior work has relied on a-priori information about the object being localized, we remove this requirement. Our key insight is that an object is composed of several local surface patches, each informative enough to achieve reliable object tracking. Moreover, we can recover the geometry of this local patch online by extracting local surface normal information embedded in each tactile image. We propose a novel two-stage approach. First, we learn a mapping from tactile images to surface normals using an image translation network. Second, we use these surface normals within a factor graph to both reconstruct a local patch map and use it to infer 3D object poses. We demonstrate reliable object tracking for over 100 contact sequences across unique shapes with four objects in simulation and two objects in the real-world. Paloma Sodhi, Michael Kaess, Mustafa Mukadam |
ICRA | 1 |
| 2022 | InCOpt: Incremental Constrained Optimization using the Bayes TreeabstractIn this work, we investigate the problem of incre-mentally solving constrained non-linear optimization problems formulated as factor graphs. Prior incremental solvers were either restricted to the unconstrained case or required periodic batch relinearizations of the objective and constraints which are expensive and detract from the online nature of the algorithm. We present InCOpt, an Augmented Lagrangian-based incremental constrained optimizer that views matrix operations as message passing over the Bayes tree. We first show how the linear system, resulting from linearizing the constrained objective, can be represented as a Bayes tree. We then propose an algorithm that views forward and back substitutions, which naturally arise from solving the Lagrangian, as upward and downward passes on the tree. Using this formulation, In-COpt can exploit properties such as fluid/online relinearization leading to increased accuracy without a sacrifice in runtime. We evaluate our solver on different applications (navigation and manipulation) and provide an extensive evaluation against existing constrained and unconstrained solvers. Mohamad Qadri, Paloma Sodhi, Josh Mangelson, Frank Dellaert, Michael Kaess |
IROS | 2 |
| 2022 | Theseus: A Library for Differentiable Nonlinear OptimizationabstractWe present Theseus, an efficient application-agnostic open source library for differentiable nonlinear least squares (DNLS) optimization built on PyTorch, providing a common framework for end-to-end structured learning in robotics and vision. Existing DNLS implementations are application specific and do not always incorporate many ingredients important for efficiency. Theseus is application-agnostic, as we illustrate with several example applications that are built using the same underlying differentiable components, such as second-order optimizers, standard costs functions, and Lie groups. For efficiency, Theseus incorporates support for sparse solvers, automatic vectorization, batching, GPU acceleration, and gradient computation with implicit differentiation and direct loss minimization. We do extensive performance evaluation in a set of applications, demonstrating significant efficiency gains and better scalability when these features are incorporated. Project page: https://sites.google.com/view/theseus-ai/ Luis Pineda, Taosha Fan, Maurizio Monge, Shobha Venkataraman, Paloma Sodhi, Ricky T. Q. Chen, Joseph Ortiz, Daniel DeTone, Austin S. Wang, Jing Dong 0002, Brandon Amos, Mustafa Mukadam |
NeurIPS | 5 |
| 2021 | Learning Tactile Models for Factor Graph-based EstimationabstractWe’re interested in the problem of estimating object states from touch during manipulation under occlusions. In this work, we address the problem of estimating object poses from touch during planar pushing. Vision-based tactile sensors provide rich, local image measurements at the point of contact. A single such measurement, however, contains limited information and multiple measurements are needed to infer latent object state. We solve this inference problem using a factor graph. In order to incorporate tactile measurements in the graph, we need local observation models that can map highdimensional tactile images onto a low-dimensional state space. Prior work has used low-dimensional force measurements or engineered functions to interpret tactile measurements. These methods, however, can be brittle and difficult to scale across objects and sensors. Our key insight is to directly learn tactile observation models that predict the relative pose of the sensor given a pair of tactile images. These relative poses can then be incorporated as factors within a factor graph. We propose a two-stage approach: first we learn local tactile observation models supervised with ground truth data, and then integrate these models along with physics and geometric factors within a factor graph optimizer. We demonstrate reliable object tracking using only tactile feedback for ~150 real-world planar pushing sequences with varying trajectories across three object shapes. Paloma Sodhi, Michael Kaess, Mustafa Mukadam |
ICRA | 1 |
| 2021 | Ground Encoding: Learned Factor Graph-based Models for Localizing Ground Penetrating RadarabstractWe address the problem of robot localization using ground penetrating radar (GPR) sensors. Current approaches for localization with GPR sensors require a priori maps of the system’s environment as well as access to approximate global positioning (GPS) during operation. In this paper, we propose a novel, real-time GPR-based localization system for unknown and GPS-denied environments. We model the localization problem as an inference over a factor graph. Our approach combines 1D single-channel GPR measurements to form 2D image submaps. To use these GPR images in the graph, we need sensor models that can map noisy, high-dimensional image measurements into the state space. These are challenging to obtain a priori since image generation has a complex dependency on subsurface composition and radar physics, which itself varies with sensors and variations in subsurface electromagnetic properties. Our key idea is to instead learn relative sensor models directly from GPR data that map non-sequential GPR image pairs to relative robot motion. These models are incorporated as factors within the factor graph with relative motion predictions correcting for accumulated drift in the position estimates. We demonstrate our approach over datasets collected across multiple locations using a custom designed experimental rig. We show reliable, real-time localization using only GPR and odometry measurements for varying trajectories in three distinct GPS-denied environments. Alexander Baikovitz, Paloma Sodhi, Michael Dille, Michael Kaess |
IROS | 2 |
| 2020 | ICS: Incremental Constrained Smoothing for State EstimationabstractA robot operating in the world constantly receives information about its environment in the form of new measurements at every time step. Smoothing-based estimation methods seek to optimize for the most likely robot state estimate using all measurements up till the current time step. Existing methods solve for this smoothing objective efficiently by framing the problem as that of incremental unconstrained optimization. However, in many cases observed measurements and knowledge of the environment is better modeled as hard constraints derived from real-world physics or dynamics. A key challenge is that the new optimality conditions introduced by the hard constraints break the matrix structure needed for incremental factorization in these incremental optimization methods. Our key insight is that if we leverage primal-dual methods, we can recover a matrix structure amenable to incremental factorization. We propose a framework ICS that combines a primal-dual method like the Augmented Lagrangian with an incremental Gauss Newton approach that reuses previously computed matrix factorizations. We evaluate ICS on a set of simulated and real-world problems involving equality constraints like object contact and inequality constraints like collision avoidance. Paloma Sodhi, Sanjiban Choudhury, Josh Mangelson, Michael Kaess |
ICRA | 1 |
| 2020 | Active SLAM using 3D Submap Saliency for Underwater Volumetric ExplorationabstractIn this paper, we present an active SLAM framework for volumetric exploration of 3D underwater environments with multibeam sonar. Recent work in integrated SLAM and planning performs localization while maintaining volumetric free-space information. However, an absence of informative loop closures can lead to imperfect maps, and therefore unsafe behavior. To solve this, we propose a navigation policy that reduces vehicle pose uncertainty by balancing between volumetric exploration and revisitation. To identify locations to revisit, we build a 3D visual dictionary from real-world sonar data and compute a metric of submap saliency. Revisit actions are chosen based on propagated pose uncertainty and sensor information gain. Loop closures are integrated as constraints in our pose-graph SLAM formulation and these deform the global occupancy grid map. We evaluate our performance in simulation and real-world experiments, and highlight the advantages over an uncertainty-agnostic framework. Sudharshan Suresh, Paloma Sodhi, Josh Mangelson, David Wettergreen, Michael Kaess |
ICRA | 2 |
| 2019 | Online and Consistent Occupancy Grid Mapping for Planning in Unknown EnvironmentsabstractActively exploring and mapping an unknown environment requires integration of both simultaneous localization and mapping (SLAM) and path planning methods. Path planning relies on a map that contains free and occupied space information and is efficient to query, while the role of SLAM is to keep the map consistent as new measurements are continuously added. A key challenge, however, lies in ensuring a map representation compatible with both these objectives: that is, a map that maintains free space information for planning but can also adapt efficiently to dynamically changing pose estimates from a graph-based SLAM system. In this paper, we propose an online global occupancy map that can be corrected for accumulated drift efficiently based on incremental solutions from a sparse graph-based SLAM optimization. Our map maintains free space information for real-time path planning while undergoing a bounded number of updates in each loop closure iteration. We evaluate performance for both simulated and real-world datasets for an application involving underwater exploration and mapping. Paloma Sodhi, Bing-Jui Ho, Michael Kaess |
IROS | 1 |
| 2018 | Virtual Occupancy Grid Map for Submap-based Pose Graph SLAM and Planning in 3D EnvironmentsabstractIn this paper, we propose a mapping approach that constructs a globally deformable virtual occupancy grid map (VOG-map) based on local submaps. Such a representation allows pose graph SLAM systems to correct globally accumulated drift via loop closures while maintaining free space information for the purpose of path planning. We demonstrate use of such a representation for implementing an underwater SLAM system in which the robot actively plans paths to generate accurate 3D scene reconstructions. We evaluate performance on simulated as well as real-world experiments. Our work furthers capabilities of mobile robots actively mapping and exploring unstructured, three dimensional environments. Bing-Jui Ho, Paloma Sodhi, Ming Hsiao, Tushar Kusnur, Michael Kaess |
IROS | 2 |
| 2018 | Robust Plant Phenotyping via Model-Based OptimizationabstractPlant phenotyping is the measurement of observable plant traits. Current methods for phenotyping in the field are labour intensive and error prone. High throughput plant phenotyping in an automated and noninvasive manner is crucial to accelerating plant breeding methods. Occlusions and non-ideal sensing conditions is a major problem for high throughput plant phenotyping with most state-of-the-art 3D phenotyping algorithms relying heavily on heuristics or hand-tuned parameters. To address this problem, we present a novel model-based optimization approach for estimating plant physical traits from plant units called phytomers. The proposed approach involves sampling parameterized 3D plant models from an underlying probability distribution. It then optimizes, making the mass of this probability distribution approach true parameters of the model. Reformulating the phenotyping objective as a search in the space of plant models lets us reason about the plant structure in a holistic manner without having to rely on hand-tuned parameters. This makes our approach robust to noise and occlusions as frequently encountered in real world environments. We evaluate our approach for plant units taken across simulated, greenhouse and field environments. This work furthers field-based robotic phenotyping capabilities paving the way for plant biologists to study the coupled effect of genetics and environment on improving crop yields. Paloma Sodhi, Hanqi Sun, Barnabás Póczos, David Wettergreen |
IROS | 1 |
| 2017 | In-field segmentation and identification of plant structures using 3D imagingabstractAutomatically correlating plant observable characteristics to their underlying genetics will streamline selection methods in plant breeding. Measurement of plant observable characteristics is called phenotyping, and knowing plant phenotypes accurately and throughout a plant's growth is central to making breeding decisions. In-field plant phenotyping in an automated and noninvasive manner is hence crucial to accelerating plant breeding methods. However, most of the existing methods on plant phenotyping using visual imaging are confined to controlled greenhouse environments. This paper presents an automated method of mapping 2D images collected in an outdoor sorghum field to segmented 3D plant units that are of interest for phenotyping. This method leverages multiple horizontal and vertical viewpoints while capturing 2D images from a robotic platform so as to generate in-field 3D reconstructions of the sorghum plant. We develop and quantitatively evaluate segmentation methods on these 3D reconstructions and also compare against reconstructions obtained from a controlled greenhouse environment. We present analysis that contrasts the role of purely local geometric features and the effect of addition of global context in both datasets. This work furthers capabilities of in-field phenotyping which paves the way forward for plant biologists to study the coupled effect of genetics and environment on improving crop yields. Paloma Sodhi, Srinivasan Vijayarangan, David Wettergreen |
IROS | 1 |