EDBT 2026 Demo / reviewers in the wild / expert
Siddarth Jain
dblp:50/8810
· DBLP profile ↗
22ranked-venue papers
5as first author
16since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 4 first-author · 15 since 2021Systems, architecture and hardware · 16 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLMabstractHuman-robot collaboration towards a shared goal requires robots to understand human action and interaction with the surrounding environment. This paper focuses on humanrobot interaction (HRI) based on human-robot dialogue that relies on the robot action confirmation and action step generation using multimodal scene understanding. The state-of-theart approach uses multimodal transformers to generate robot action steps aligned with robot action confirmation from a single clip showing a task composed of multiple micro steps. Although actions towards a long-horizon task depend on each other throughout an entire video, the current approaches mainly focus on clip-level processing and do not leverage long-context information. This paper proposes a long-context Q-former incorporating left and right context dependency in full videos. Furthermore, this paper proposes a text-conditioning approach to feed text embeddings directly into the LLM decoder to mitigate the high abstraction of the information in text by Q-former. Experiments with the YouCook2 corpus show that the accuracy of confirmation generation is a major factor in the performance of action planning. Furthermore, we demonstrate that the longcontext Q-former improves the confirmation and action planning by integrating VideoLLaMA3. Chiori Hori, Yoshiki Masuyama, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux |
ASRU | 3 |
| 2025 | Interactive Robot Action Replanning using Multimodal LLM Trained from Human Demonstration VideosabstractUnderstanding human actions could allow robots to perform a large spectrum of complex manipulation tasks and make collaboration with humans easier. Recently, multimodal scene understanding using audio-visual Transformers has been used to generate robot action sequences from videos of human demonstrations. However, automatic action sequence generation is not always perfect due to the distribution gap between the training and test environments. To bridge this gap, human intervention could be very effective, such as telling the robot agent what should be done. Motivated by this, we propose an error-correction-based action replanning approach that regenerates better action sequences using (1) automatically generated actions from a pretrained action generator and (2) human error-correction in natural language. We collected singlearm robot action sequences aligned to human action instruction for the cooking video dataset YouCook2. We trained the proposed errorcorrection-based action replanning model using a pre-trained multimodal LLM model (AVBLIP-2), generating a pair of (a) single-arm robot micro-step action sequences and (b) action descriptions in natural language simultaneously. To assess the performance of error correction, we collected human feedback on correcting errors in the automatically generated robot actions. Experiments show that our proposed interactive replanning model trained in a multitask manner using action sequence and description outperformed the baseline model in all types of scores. Chiori Hori, Motonari Kambara, Komei Sugiura, Kei Ota, Sameer Khurana, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux |
ICASSP | 6 |
| 2025 | PACE: Proactive Assistance in Human-Robot Collaboration Through Action-Completion EstimationabstractThis paper introduces the Proactive Assistance through action-Completion Estimation (PACE) framework, designed to enhance human-robot collaboration through real-time monitoring of human progress. PACE incorporates a novel method that combines Dynamic Time Warping (DTW) with correlation analysis to track human task progression from hand movements. PACE trains a reinforcement learning policy from limited demonstrations to generate a proactive assistance policy that synchronizes robotic actions with human activities, minimizing idle time and enhancing collaboration efficiency. We validate the framework through user studies involving 12 participants, showing significant improvements in interaction fluency, reduced waiting times, and positive user feedback compared to traditional methods. Davide De Lazzari, Matteo Terreran, Giulio Giacomuzzo, Siddarth Jain, Pietro Falco, Ruggero Carli, Diego Romeres |
ICRA | 4 |
| 2024 | Interactive Planning Using Large Language Models for Partially Observable Robotic TasksabstractDesigning robotic agents to perform open vocabulary tasks has been the long-standing goal in robotics and AI. Recently, Large Language Models (LLMs) have achieved impressive results in creating robotic agents for performing open vocabulary tasks. However, planning for these tasks in the presence of uncertainties is challenging as it requires "chain-of-thought" reasoning, aggregating information from the environment, updating state estimates, and generating actions based on the updated state estimates. In this paper, we present an interactive planning technique for partially observable tasks using LLMs. In the proposed method, an LLM is used to collect missing information from the environment using a robot, and infer the state of the underlying problem from collected observations while guiding the robot to perform the required actions. We also use a fine-tuned Llama 2 model via self-instruct and compare its performance against a pre-trained LLM like GPT-4. Results are demonstrated on several tasks in simulation as well as real-world environments. Lingfeng Sun, Devesh K. Jha, Chiori Hori, Siddarth Jain, Radu Corcodel, Xinghao Zhu, Masayoshi Tomizuka, Diego Romeres |
ICRA | 4 |
| 2024 | Insert-One: One-Shot Robust Visual-Force Servoing for Novel Object Insertion with 6-DoF TrackingabstractRecent advancements in autonomous robotic assembly have shown promising results, especially in addressing the precision insertion challenge. However, achieving adaptability across diverse object categories and tasks often necessitates a learning phase that requires costly real-world data collection. Moreover, previous research often assumes either the rigid attachment of the inserted object to the robot’s end-effector or relies on precise calibration within structured environments. We propose a one-shot method for high-precision contact-rich manipulation assembly tasks, enabling a robot to perform insertions of new objects from randomly presented orientations using just a single demonstration image. Our method incorporates a hybrid framework that blends 6-DoF visual tracking-based iterative control and impedance control, facilitating high-precision tasks with real-time visual feedback. Importantly, our approach requires no pre-training and demonstrates resilience against uncertainties arising from camera pose calibration errors and disturbances in the object in-hand pose. We validate the effectiveness of the proposed framework through extensive experiments in real-world scenarios, encompassing various high-precision assembly tasks. Haonan Chang, Abdeslam Boularias, Siddarth Jain |
IROS | 3 |
| 2024 | Few-shot Transparent Instance Segmentation for Bin PickingabstractIn this paper, we consider the problem of segmenting multiple instances of a transparent object from RGB or gray scale camera images in a robotic bin picking setting. Prior methods for solving this task are usually built on the Mask-RCNN framework, but they require large annotated datasets for fine-tuning. Instead, we consider the task in a few-shot setting and present TrInSeg, a data-efficient and robust instance segmentation method for transparent objects based on Mask-RCNN. Our key innovations in TrInSeg are twofold: i) a novel method, dubbed TransMixup, for producing new training images using synthetic transparent object instances created by spatially transforming annotated examples; and ii) a method for scoring the consistency between the predicted segments and rotations of an ideal object template. In our new scoring method, the spatial transformations are produced by an auxiliary neural network, and the scores are then used to filter inconsistent instance predictions. To demonstrate the effectiveness of our method, we present experiments on a new few-shot dataset consisting of seven categories of non-opaque (transparent and translucent) objects, each category varying in the size, shape, and degree of transparency of the objects. Our results show that TrInSeg achieves state-of-the-art performance, improving fine-tuned Mask-RCNN by more than 14% in mIoU, while requiring very few annotated training samples. Anoop Cherian, Siddarth Jain, Tim K. Marks |
IROS | 2 |
| 2024 | DECAF: a Discrete-Event based Collaborative Human-Robot Framework for Furniture AssemblyabstractThis paper proposes a task planning framework for collaborative Human-Robot scenarios, specifically focused on assembling complex systems such as furniture. The human is characterized as an uncontrollable agent, implying for example that the agent is not bound by a pre-established sequence of actions and instead acts according to its own preferences. Meanwhile, the task planner computes reactively the optimal actions for the collaborative robot to efficiently complete the entire assembly task in the least time possible.We formalize the problem as a Discrete Event Markov Decision Problem (DE-MDP), a comprehensive framework that incorporates a variety of asynchronous behaviors, human change of mind, and failure recovery as stochastic events. Although the problem could theoretically be addressed by constructing a graph of all possible actions, such an approach would be constrained by computational limitations. The proposed formulation offers an alternative solution utilizing Reinforcement Learning to derive an optimal policy for the robot. Experiments were conducted both in simulation and on a real system with human subjects assembling a chair in collaboration with a 7-DoF manipulator. Giulio Giacomuzzo, Matteo Terreran, Siddarth Jain, Diego Romeres |
IROS | 3 |
| 2024 | Autonomous Robotic Assembly: From Part Singulation to Precise AssemblyabstractImagine a robot that can assemble a functional product from the individual parts presented in any configuration to the robot. Designing such a robotic system is a complex problem which presents several open challenges. To bypass these challenges, the current generation of assembly systems is built with a lot of system integration effort to provide the structure and precision necessary for assembly. These systems are mostly responsible for part singulation, part kitting, and part detection, which is accomplished by intelligent system design. In this paper, we present autonomous assembly of a gear box with minimum requirements on structure. The assembly parts are randomly placed in a two-dimensional work environment for the robot. The proposed system makes use of several different manipulation skills such as sliding for grasping, in-hand manipulation, and insertion to assemble the gear box. All these tasks are run in a closed-loop fashion using vision, tactile, and Force-Torque (F/T) sensors. We perform extensive hardware experiments to show the robustness of the proposed methods as well as the overall system. See supplementary video at https://www.youtube.com/watch?v=cZ9M1DQ23OI. Kei Ota, Devesh K. Jha, Siddarth Jain, William Yerazunis, Radu Corcodel, Yash Shukla, Antonia Bronars, Diego Romeres |
IROS | 3 |
| 2024 | Open Human-Robot Collaboration using Decentralized Inverse Reinforcement LearningabstractThe growing interest in human-robot collaboration (HRC), where humans and robots cooperate towards shared goals, has seen significant advancements over the past decade. While previous research has addressed various challenges, several key issues remain unresolved. Many domains within HRC involve activities that do not necessarily require human presence throughout the entire task. Existing literature typically models HRC as a closed system, where all agents are present for the entire duration of the task. In contrast, an open model offers flexibility by allowing an agent to enter and exit the collaboration as needed, enabling them to concurrently manage other tasks. In this paper, we introduce a novel multiagent framework called oDec-MDP, designed specifically to model open HRC scenarios where agents can join or leave tasks flexibly during execution. We generalize a recent multiagent inverse reinforcement learning method - Dec-AIRL to learn from open systems modeled using the oDec-MDP. Our method is validated through experiments conducted in both a simplified toy firefighting domain and a realistic dyadic human-robot collaborative assembly. Results show that our framework and learning method improves upon its closed system counterpart. Prasanth Sengadu Suresh, Siddarth Jain, Prashant Doshi, Diego Romeres |
IROS | 2 |
| 2023 | Discriminative 3D Shape Modeling for Few-Shot Instance SegmentationabstractIn this paper, we present a simple and efficient scheme for segmenting approximately convex 3D object instances in depth images in a few-shot setting via discriminatively modeling the 3D shape of the object using a neural network. Our key idea is to select pairs of 3D points on the depth image between which we compute surface geodesics. As the number of such geodesics is quadratic in the number of image pixels, we can create a large training set of geodesics using only very limited ground truth instance annotations. These annotations are used to create a binary label for each geodesic, which indicates whether or not that geodesic belongs entirely to one instance segment. A neural network is then trained to classify the geodesics using these labels. During inference, we create geodesics from selected seed points in the test depth image, then produce a convex hull of the points that are classified by the neural network as belonging to the same instance, thereby achieving instance segmentation. We present experiments applying our method to segmenting instances of food items in real-world depth images. Our results demonstrate promising performances compared to prior methods in accuracy and computational efficiency. Anoop Cherian, Siddarth Jain, Tim K. Marks, Alan Sullivan |
ICRA | 2 |
| 2023 | Task-Directed Exploration in Continuous POMDPs for Robotic Manipulation of Articulated ObjectsabstractRepresenting and reasoning about uncertainty is crucial for autonomous agents acting in partially observable environments with noisy sensors. Partially observable Markov decision processes (POMDPs) serve as a general framework for representing problems in which uncertainty is an important factor. Online sample-based POMDP methods have emerged as efficient approaches to solving large POMDPs and have been shown to extend to continuous domains. However, these solutions struggle to find long-horizon plans in problems with significant uncertainty. Exploration heuristics can help guide planning, but many real-world settings contain significant task-irrelevant uncertainty that might distract from the task objective. In this paper, we propose STRUG, an online POMDP solver capable of handling domains that require long-horizon planning with significant task-relevant and task-irrelevant uncertainty. We demonstrate our solution on several temporally extended versions of toy POMDP problems as well as robotic manipulation of articulated objects using a neural perception frontend to construct a distribution of possible models. Our results show that STRUG outperforms the current sample-based online POMDP solvers on several tasks. Aidan Curtis, Leslie Pack Kaelbling, Siddarth Jain |
ICRA | 3 |
| 2023 | Learning Generalizable Pivoting SkillsabstractThe skill of pivoting an object with a robotic system is challenging for the external forces that act on the system, mainly given by contact interaction. The complexity increases when the same skills are required to generalize across different objects. This paper proposes a framework for learning robust and generalizable pivoting skills, which consists of three steps. First, we learn a pivoting policy on an “unitary” object using Reinforcement Learning (RL). Then, we obtain the object's feature space by supervised learning to encode the kinematic properties of arbitrary objects. Finally, to adapt the unitary policy to multiple objects, we learn data-driven projections based on the object features to adjust the state and action space of the new pivoting task. The proposed approach is entirely trained in simulation. It requires only one depth image of the object and can zero-shot transfer to real-world objects. We demonstrate robustness to sim-to-real transfer and generalization to multiple objects. Xiang Zhang 0020, Siddarth Jain, Baichuan Huang, Masayoshi Tomizuka, Diego Romeres |
ICRA | 2 |
| 2023 | Style-transfer based Speech and Audio-visual Scene understanding for Robot Action Sequence Acquisition from Videos
Chiori Hori, Puyuan Peng, David F. Harwath, Kei Ota, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux |
INTERSPEECH | 6 |
| 2023 | EARL: Eye-on-Hand Reinforcement Learner for Dynamic Grasping with Active Pose EstimationabstractIn this paper, we explore the dynamic grasping of moving objects through active pose tracking and reinforcement learning for hand-eye coordination systems. Most existing vision-based robotic grasping methods implicitly assume target objects are stationary or moving predictably. Performing grasping of unpredictably moving objects presents a unique set of challenges. For example, a pre-computed robust grasp can become unreachable or unstable as the target object moves, and motion planning must also be adaptive. In this work, we present a new approach, Eye-on-hAnd Reinforcement Learner (EARL), for enabling coupled Eye-on-Hand (EoH) robotic manipulation systems to perform real-time active pose tracking and dynamic grasping of novel objects without explicit motion prediction. EARL readily addresses many thorny issues in automated hand-eye coordination, including fast-tracking of 6D object pose from vision, learning control policy for a robotic arm to track a moving object while keeping the object in the camera's field of view, and performing dynamic grasping. We demonstrate the effectiveness of our approach in extensive experiments validated on multiple commercial robotic arms in both simulations and complex real-world tasks. Baichuan Huang, Jingjin Yu, Siddarth Jain |
IROS | 3 |
| 2022 | Learning to Synthesize Volumetric Meshes from Vision-based Tactile ImprintsabstractVision-based tactile sensors typically utilize a deformable elastomer and a camera mounted above to provide high-resolution image observations of contacts. Obtaining accurate volumetric meshes for the deformed elastomer can provide direct contact information and benefit robotic grasping and manipulation. This paper focuses on learning to synthesize the volumetric mesh of the elastomer based on the image imprints acquired from vision-based tactile sensors. Synthetic image-mesh pairs and real-world images are gathered from 3D finite element methods (FEM) and physical sensors, respectively. A graph neural network (GNN) is introduced to learn the image-to-mesh mappings with supervised learning. A self-supervised adaptation method and image augmentation techniques are proposed to transfer networks from simulation to reality, from primitive contacts to unseen contacts, and from one sensor to another. Using these learned and adapted networks, our proposed method can accurately reconstruct the deformation of the real-world tactile sensor elastomer in various domains, as indicated by the quantitative and qualitative results. Xinghao Zhu, Siddarth Jain, Masayoshi Tomizuka, Jeroen van Baar |
ICRA | 2 |
| 2021 | InSeGAN: A Generative Approach to Segmenting Identical Instances in Depth ImagesabstractIn this paper, we present InSeGAN, an unsupervised 3D generative adversarial network (GAN) for segmenting (nearly) identical instances of rigid objects in depth images. Using an analysis-by-synthesis approach, we design a novel GAN architecture to synthesize a multiple-instance depth image with independent control over each instance. InSeGAN takes in a set of code vectors (e.g., random noise vectors), each encoding the 3D pose of an object that is represented by a learned implicit object template. The generator has two distinct modules. The first module, the instance feature generator, uses each encoded pose to transform the implicit template into a feature map representation of each object instance. The second module, the depth image renderer, aggregates all of the single-instance feature maps output by the first module and generates a multiple-instance depth image. A discriminator distinguishes the generated multiple-instance depth images from the distribution of true depth images. To use our model for instance segmentation, we propose an instance pose encoder that learns to take in a generated depth image and reproduce the pose code vectors for all of the object instances. To evaluate our approach, we introduce a new synthetic dataset, "Insta-10," consisting of 100,000 depth images, each with 5 instances of an object from one of 10 classes. Our experiments on Insta-10, as well as on real-world noisy depth images, show that InSeGAN achieves state-of-the-art performance, often outperforming prior methods by large margins. Anoop Cherian, Gonçalo Dias Pais, Siddarth Jain, Tim K. Marks, Alan Sullivan |
ICCV | 3 |
| 2020 | Interactive Tactile Perception for Classification of Novel Object InstancesabstractIn this paper, we present a novel approach for classification of unseen object instances from interactive tactile feedback. Furthermore, we demonstrate the utility of a low resolution tactile sensor array for tactile perception that can potentially close the gap between vision and physical contact for manipulation. We contrast our sensor to high-resolution camera-based tactile sensors. Our proposed approach interactively learns a one-class classification model using 3D tactile descriptors, and thus demonstrates an advantage over the existing approaches, which require pre-training on objects. We describe how we derive 3D features from the tactile sensor inputs, and exploit them for learning one-class classifiers. In addition, since our proposed method uses unsupervised learning, we do not require ground truth labels. This makes our proposed method flexible and more practical for deployment on robotic systems. We validate our proposed method on a set of household objects and results indicate good classification performance in real-world experiments. Radu Corcodel, Siddarth Jain, Jeroen van Baar |
IROS | 2 |
| 2020 | Probabilistic Human Intent Recognition for Shared Autonomy in Assistive RoboticsabstractEffective human-robot collaboration in shared autonomy requires reasoning about the intentions of the human partner. To provide meaningful assistance, the autonomy has to first correctly predict, or infer, the intended goal of the human collaborator. In this work, we present a mathematical formulation for intent inference during assistive teleoperation under shared autonomy. Our recursive Bayesian filtering approach models and fuses multiple non-verbal observations to probabilistically reason about the intended goal of the user without explicit communication. In addition to contextual observations, we model and incorporate the human agent's behavior as goal-directed actions with adjustable rationality to inform intent recognition. Furthermore, we introduce a user-customized optimization of this adjustable rationality to achieve user personalization. We validate our approach with a human subjects study that evaluates intent inference performance under a variety of goal scenarios and tasks. Importantly, the studies are performed using multiple control interfaces that are typically available to users in the assistive domain, which differ in the continuity and dimensionality of the issued control signals. The implications of the control interface limitations on intent inference are analyzed. The study results show that our approach in many scenarios outperforms existing solutions for intent inference in assistive teleoperation, and otherwise performs comparably. Our findings demonstrate the benefit of probabilistic modeling and the incorporation of human agent behavior as goal-directed actions where the adjustable rationality model is user customized. Results further show that the underlying intent inference approach directly affects shared autonomy performance, as do control interface limitations. Siddarth Jain, Brenna D. Argall |
ACM Trans. Hum. Robot Interact. | 1 |
| 2018 | Recursive Bayesian Human Intent Recognition in Shared-Control RoboticsabstractEffective human-robot collaboration in shared control requires reasoning about the intentions of the human user. In this work, we present a mathematical formulation for human intent recognition during assistive teleoperation under shared autonomy. Our recursive Bayesian filtering approach models and fuses multiple non-verbal observations to probabilistically reason about the intended goal of the user. In addition to contextual observations, we model and incorporate the human agent's behavior as goal-directed actions with adjustable rationality to inform the underlying intent. We examine human inference on robot motion and furthermore validate our approach with a human subjects study that evaluates autonomy intent inference performance under a variety of goal scenarios and tasks, by novice subjects. Results show that our approach outperforms existing solutions and demonstrates that the probabilistic fusion of multiple observations improves intent inference and performance for shared-control operation. Siddarth Jain, Brenna D. Argall |
IROS | 1 |
| 2016 | Grasp detection for assistive robotic manipulationabstractIn this paper, we present a novel grasp detection algorithm targeted towards assistive robotic manipulation systems. We consider the problem of detecting robotic grasps using only the raw point cloud depth data of a scene containing unknown objects, and apply a geometric approach that categorizes objects into geometric shape primitives based on an analysis of local surface properties. Grasps are detected without a priori models, and the approach can generalize to any number of novel objects that fall within the shape primitive categories. Our approach generates multiple candidate object grasps, which moreover are semantically meaningful and similar to what a human would generate when teleoperating the robot-and thus should be suitable manipulation goals for assistive robotic systems. An evaluation of our algorithm on 30 household objects includes a pilot user study, confirms the robustness of the detected grasps and was conducted in real-world experiments using an assistive robotic arm. Siddarth Jain, Brenna D. Argall |
ICRA | 1 |
| 2014 | Automated perception of safe docking locations with alignment information for assistive wheelchairsabstractThere are basic manuvering tasks with a powered wheelchair, like docking under a table and passage through a doorway or narrow hallway, which can be difficult for users with severe motor impairments - not only because of limitations in their own motor control, but also because of limitations in the control interfaces available to them. Robot automation can help transfer some of this control burden from the user to the machine. This work presents an algorithm for the automated detection of safe docking locations at rectangular and circular docking structures (tables, desks) with proper alignment information using 3D point cloud data. The safe docking locations can then be provided as goals to an autonomous path planner, within the context of providing adaptive driving assistance for powered wheelchair users. We evaluate the performance of our algorithm with systematic testing on several docking structures, observed from varied viewpoints. Siddarth Jain, Brenna D. Argall |
IROS | 1 |
| 2010 | Ad-hoc swarm robotics optimization in grid based navigationabstractA technology to identify a correct remote object and to carry it back to a base camp has become an important utility form, in the military for identifying improvised explosive devices (IEDs) and for space scientists to collect samples from mars. Multi-robot systems can accomplish tasks that would be impossible for a single robot to achieve. Developing multi-agent systems with self interested agents with a large behavioural repertoire is a great challenge. In this paper, we combine collaborative agent behaviour to deal with these vital issues. An ad-hoc design approach for swarm robotics optimization in grid based navigation is presented. Implementation of low power CC2500 wireless communication chip is done to accomplish multi-robot communication. Primitive behaviors called real-time grid search, move-to-goal, avoid-static-obstacle, avoid-robot behaviors are introduced for the path planning of multi-robot. The system is based on Real-time Grid Searching and Collision free Path Tracing & Retracing algorithm. A randomly placed remote object in the grid is identified and is carried to the base station with the effective coordination of the multi-robot system. Siddarth Jain, Manish Sawlani, Vijay Kumar Chandwani |
ICARCV | 1 |