EDBT 2026 Demo / reviewers in the wild / expert
Utkarsh A. Mishra
dblp:274/2706 · also Utkarsh Aashu Mishra
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-4977-5187ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RAIL: Reachability-Aided Imitation Learning for Safe Policy ExecutionabstractImitation learning (IL) has shown great success in learning complex robot manipulation tasks. However, there remains a need for practical safety methods to justify widespread deployment. In particular, it is important to certify that a system obeys hard constraints on unsafe behavior in settings when it is unacceptable to design a tradeoff between performance and safety via tuning the policy (i.e. soft constraints). This leads to the question, how does enforcing hard constraints impact the performance (meaning safely completing tasks) of an IL policy? To answer this question, this paper builds a reach ability - based safety filter to enforce hard constraints on IL, which we call Reachability-Aided Imitation Learning (RAIL). Through evaluations with state-of-the-art IL policies in mobile robots and manipulation tasks, we make two key findings. First, the highest-performing policies are sometimes only so because they frequently violate constraints, and significantly lose performance under hard constraints. Second, surprisingly, hard constraints on the lower-performing policies can occasionally increase their ability to perform tasks safely. Finally, hardware evaluation confirms the method can operate in real time. More results can be found at our website: https://safe-robotics-lab-gt.github.io/rail/. Wonsuhk Jung, Dennis Anthony, Utkarsh A. Mishra, Nadun Ranawaka Arachchige, Matthew Bronars, Danfei Xu, Shreyas Kousik |
ICRA | 3 |
| 2025 | Generative Trajectory Stitching through Diffusion CompositionabstractEffective trajectory stitching for long-horizon planning is a significant challenge in robotic decision-making. While diffusion models have shown promise in planning, they are limited to solving tasks similar to those seen in their training data. We propose CompDiffuser, a novel generative approach that can solve new tasks by learning to compositionally stitch together shorter trajectory chunks from previously seen tasks. Our key insight is modeling the trajectory distribution by subdividing it into overlapping chunks and learning their conditional relationships through a single bidirectional diffusion model. This allows information to propagate between segments during generation, ensuring physically consistent connections. We conduct experiments on benchmark tasks of various difficulties, covering different environment sizes, agent state dimension, trajectory types, training data quality, and show that CompDiffuser significantly outperforms existing methods. Yunhao Luo 0001, Utkarsh A. Mishra, Yilun Du, Danfei Xu |
NeurIPS | 2 |
| 2024 | ReorientDiff: Diffusion Model based Reorientation for Object ManipulationabstractThe ability to manipulate objects in desired configurations is a fundamental requirement for robots to complete various practical applications. While certain goals can be achieved by picking and placing the objects of interest directly, object reorientation is needed for precise placement in most of the tasks. In such scenarios, the object must be reoriented and re-positioned into intermediate poses that facilitate accurate placement at the target pose. To this end, we propose a reorientation planning method, ReorientDiff, that utilizes a diffusion model-based approach. The proposed method employs both visual inputs from the scene, and goal-specific language prompts to plan intermediate reorientation poses. Specifically, the scene and language-task information are mapped into a joint scene-task representation feature space, which is subsequently leveraged to condition the diffusion model. The diffusion model samples intermediate poses based on the representation using classifier-free guidance and then uses gradients of learned feasibility-score models for implicit iterative pose-refinement. The proposed method is evaluated using a set of YCB-objects and a suction gripper, demonstrating a success rate of 95.2% in simulation. Overall, we present a promising approach to address the reorientation challenge in manipulation by learning a conditional distribution, which is an effective way to move towards generalizable object manipulation. More results can be found on our website: https://utkarshmishra04.github.io/ReorientDiff. Utkarsh A. Mishra, Yongxin Chen 0002 |
ICRA | 1 |
| 2023 | Neural network temporal quantized lagrange dynamics with cycloidal trajectory for a toe-foot bipedal robot to climb stairs
Gaurav Bhardwaj, Utkarsh A. Mishra, Nagarajan Sukavanam, R. Balasubramanian |
Appl. Intell. | 2 |
| 2022 | Dynamic Mirror Descent based Model Predictive Control for Accelerating Robot LearningabstractRecent works in Reinforcement Learning (RL) combine model-free (Mf)-RL algorithms with model-based (Mb)-RL approaches to get the best from both: asymptotic performance of Mf-RL and high sample-efficiency of Mb-RL. Inspired by these works, we propose a hierarchical framework that integrates online learning for the Mb-trajectory optimization with off-policy methods for the Mf-RL. In particular, two loops are proposed, where the Dynamic Mirror Descent based Model Predictive Control (DMD-MPC) is used as the inner loop Mb-RL to obtain an optimal sequence of actions. These actions are in turn used to significantly accelerate the outer loop Mf-RL. We show that our formulation is generic for a broad class of MPC based policies and objectives, and includes some of the well-known Mb-Mf approaches. We finally introduce a new algorithm: Mirror-Descent Model Predictive RL (M-DeMoRL), which uses Cross-Entropy Method (CEM) with elite fractions for the inner loop. Our experiments show faster convergence of the proposed hierarchical approach on benchmark MuJoCo tasks. We also demonstrate hardware training for trajectory tracking in a 2R leg, and hardware transfer for robust walking in a quadruped. We show that the inner-loop Mb-RL significantly decreases the number of training iterations required in the hardware setting, thereby validating the proposed approach. Utkarsh A. Mishra, Soumya R. Samineni, Prakhar Goel, Chandravaran Kunjeti, Himanshu Lodha, Aditya Sagi, Shalabh Bhatnagar, Shishir Kolathaya |
ICRA | 1 |
| 2021 | Kinematic Stability based AFG-RRT* Path Planning for Cable-Driven Parallel Robots †abstractMotion planning for Cable-Driven Parallel Robots (CDPRs) is a challenging task due to various restrictions on cable tensions, collisions and obstacle avoidance. The presented work aims at proposing an optimal path planning strategy in order to both maximize the wrench capability and the dexterity of the robot in a cluttered environment. First, an asymptoticallyoptimal path finding method based on a variant of rapidly exploring random trees (RRT) is implemented along with the GilbertJohnsonKeerthi (GJK) algorithm to account for the collision detections. Then, a goal biased Artificial Field Guide (AFG) is employed to reduce convergence time and ensure directional exploration. Finally, a post-processing algorithm is added to get a short and smooth resultant path by fitting appropriate splines. The proposed path planning strategy is analyzed and demonstrated on a simulated and experimental setup of a six-DOF spatial CDPR. Utkarsh A. Mishra, Marceau Métillon, Stéphane Caro |
ICRA | 1 |
| 2021 | Learning Linear Policies for Robust Bipedal Locomotion on Terrains with Varying SlopesabstractIn this paper, with a view toward deployment of light-weight control frameworks for bipedal walking robots, we realize end-foot trajectories that are shaped by a single linear feedback policy. We learn this policy via a model-free and a gradient free learning algorithm, Augmented Random Search (ARS), in the two robot platforms Rabbit and Digit. Our contributions are two-fold: a) By using torso and support plane orientation as inputs, we achieve robust walking on slopes of upto 20° in simulation. b) We demonstrate additional behaviors like walking backwards, stepping-in-place, and recovery from external pushes of upto 120 N. The end-result is a robust and a fast feedback control law for bipedal walking on terrains with varying slopes. Towards the end, we also provide preliminary results of hardware transfer to Digit. Lokesh Krishna, Utkarsh A. Mishra, Guillermo A. Castillo, Ayonga Hereid, Shishir Kolathaya |
IROS | 2 |