Nima Fazeli

dblp:150/9043 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-0834-4767ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 6 since 2021Systems, architecture and hardware · 12 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2025 RACER: Rich Language-Guided Failure Recovery Policies for Imitation Learning
abstract
Developing robust and correctable visuomotor policies for robotic manipulation is challenging due to the lack of self-recovery mechanisms from failures and the limitations of simple language instructions in guiding robot actions. To address these issues, we propose a scalable data generation pipeline that automatically augments expert demonstrations with failure recovery trajectories and fine-grained language annotations for training. We then introduce Rich languAge-guided failure reCovERy (RACER), a supervisor-actor frame-work, which combines failure recovery data with rich language descriptions to enhance robot control. RACER features a vision-language model (VLM) that acts as an online supervisor, providing detailed language guidance for error correction and task execution, and a language-conditioned visuomotor policy as an actor to predict the next actions. Our experimental results show that RACER outperforms the state-of-the-art Robotic View Transformer (RVT) on RLbench across various evaluation settings, including standard long-horizon tasks, dynamic goal-change tasks and zero-shot unseen tasks, achieving superior performance in both simulated and real world environments. Videos and code are available at: https://rich-language-failure-recovery.github.io.
Yinpei Dai, Jayjun Lee, Nima Fazeli, Joyce Y. Chai
ICRA3
2025 Tactile Functasets: Neural Implicit Representations of Tactile Datasets
abstract
Modern incarnations of tactile sensors produce high-dimensional raw sensory feedback such as images, making it challenging to efficiently store, process, and generalize across sensors. To address these concerns, we introduce a novel implicit function representation for tactile sensor feedback. Rather than directly using raw tactile images, we propose neural implicit functions trained to reconstruct the tactile dataset, producing compact representations that capture the underlying structure of the sensory inputs. These representations offer several advantages over their raw counterparts: they are compact, enable probabilistically interpretable inference, and facilitate generalization across different sensors. We demonstrate the efficacy of this representation on the downstream task of in-hand object pose estimation, achieving improved performance over image-based methods while simplifying downstream models. We release code, demos and datasets at https://www.mmintlab.com/tactile-functasets.
Sikai Li, Samanta Rodriguez, Yiming Dou, Andrew Owens, Nima Fazeli
ICRA5
2025 Contrastive Touch-to-Touch Pretraining
abstract
Today's tactile sensors have a variety of different designs, making it challenging to develop general-purpose methods for processing touch signals. In this paper, we learn a unified representation that captures the shared information between different tactile sensors. Unlike current approaches that focus on reconstruction or task-specific supervision, we leverage contrastive learning to integrate tactile signals from two different sensors into a shared embedding space, using a dataset in which the same objects are probed with multiple sensors. We apply this approach to paired touch signals from GelSlim and Soft Bubble sensors. We show that our learned features provide strong pretraining for downstream pose estimation and classification tasks. We also show that our embedding enables models trained using one touch sensor to be deployed using another without additional training. Project details can be found at https://www.mmintlab.com/research/cttp/.
Samanta Rodriguez, Yiming Dou, William van den Bogert, Miquel Oller, Kevin So, Andrew Owens, Nima Fazeli
ICRA7
2025 This&That: Language-Gesture Controlled Video Generation for Robot Planning
abstract
Clear, interpretable instructions are invaluable for complex tasks, helping to clarify goals and anticipate necessary steps. In this work, we propose a robot learning framework for communicating, planning, and executing a wide range of tasks, dubbed This&That. This&That solves general tasks by leveraging video generative models, which, through training on internet-scale data, contain rich physical and semantic context. Through this work, we tackle three fundamental challenges in video-based planning: 1) unambiguous task communication with simple human instructions, 2) controllable video gen-eration that respects user intent, and 3) translating visual plans into robot actions. This& That adds gesture conditioning alongside language to generate video predictions as a suc-cinct and unambiguous alternative to existing language-only methods, especially in complex and uncertain environments. These video predictions are then fed into a behavior cloning architecture dubbed Diffusion Video to Action (DiVA), which outperforms prior state-of-the-art behavior cloning and video-based planning methods by substantial margins. Project web-site: https://this-and-that-vid.github.io/this-and-thatl.
Nikhil Sridhar, Mark Van der Merwe, Adam Fishman, Nima Fazeli, Jeong Joon Park
ICRA6
2025 RUMI: Rummaging Using Mutual Information
abstract
This paper presents Rummaging Using Mutual Information (RUMI), a method for online generation of robot action sequences to gather information about the pose of a known movable object in visually-occluded environments. Focusing on contact-rich rummaging, our approach leverages mutual information between the object pose distribution and robot trajectory for action planning. From an observed partial point cloud, RUMI deduces the compatible object pose distribution and approximates the mutual information of it with workspace occupancy in real time. Based on this, we develop an information gain cost function and a reachability cost function to keep the object within the robot's reach. These are integrated into a model predictive control (MPC) framework with a stochastic dynamics model, updating the pose distribution in a closed loop. Key contributions include a new belief framework for object pose estimation, an efficient information gain computation strategy, and a robust MPC-based control scheme. RUMI demonstrates superior performance in both simulated and real tasks compared to baseline methods.
Sheng Zhong 0004, Nima Fazeli, Dmitry Berenson
IEEE Trans. Robotics2
2022 VIRDO: Visio-tactile Implicit Representations of Deformable Objects
abstract
Deformable object manipulation requires computationally efficient representations that are compatible with robotic sensing modalities. In this paper, we present VIRDO: an implicit, multi-modal, and continuous representation for deformable-elastic objects. VIRDO operates directly on visual (point cloud) and tactile (reaction forces) modalities and learns rich latent embeddings of contact locations and forces to predict object deformations subject to external contacts. Here, we demonstrate VIRDOs ability to: i) produce high-fidelity cross-modal reconstructions with dense unsupervised correspondences, ii) generalize to unseen contact formations, and iii) state-estimation with partial visio-tactile feedback. https://github.com/MMintLab/VIRDO
Youngsun Wi, Peter R. Florence, Andy Zeng 0001, Nima Fazeli
ICRA4
2022 Simultaneous Contact Location and Object Pose Estimation Using Proprioception and Tactile Feedback
abstract
Joint estimation of grasped object pose and extrinsic contacts is central to robust and dexterous manipulation. In this paper, we propose a novel state-estimation algorithm that jointly estimates contact location and object pose in 3D using exclusively proprioception and tactile feedback. Our approach leverages two complementary particle filters: one to estimate contact location (CPFGrasp) and another to estimate object poses (SCOPE). We implement and evaluate our approach on real-world single-arm and dual-arm robotic systems. We demonstrate that by bringing two objects into contact, the robots can infer contact location and object poses simultaneously. Our proposed method can be applied to a number of downstream tasks that require accurate pose estimates, such as tool use and assembly. Code and data can be found at https://github.com/MMintLab/scope.
Andrea Sipos, Nima Fazeli
IROS2
2020 Long-Horizon Prediction and Uncertainty Propagation with Residual Point Contact Learners
abstract
The ability to simulate and predict the outcome of contacts is paramount to the successful execution of many robotic tasks. Simulators are powerful tools for the design of robots and their behaviors, yet the discrepancy between their predictions and observed data limit their usability. In this paper, we propose a self-supervised approach to learning residual models for rigid-body simulators that exploits corrections of contact models to refine predictive performance and propagate uncertainty. We empirically evaluate the framework by predicting the outcomes of planar dice rolls and compare it's performance to state-of-the-art techniques.
Nima Fazeli, Anurag Ajay, Alberto Rodriguez 0003
ICRA1
2019 Combining Physical Simulators and Object-Based Networks for Control
abstract
Physics engines play an important role in robot planning and control; however, many real-world control problems involve complex contact dynamics that cannot be characterized analytically. Most physics engines therefore employ approximations that lead to a loss in precision. In this paper, we propose a hybrid dynamics model, simulator-augmented interaction networks (SAIN), combining a physics engine with an object-based neural network for dynamics modeling. Compared with existing models that are purely analytical or purely data-driven, our hybrid model captures the dynamics of interacting objects in a more accurate and data-efficient manner. Experiments both in simulation and on a real robot suggest that it also leads to better performance when used in complex control tasks. Finally, we show that our model generalizes to novel environments with varying object shapes and materials.
Anurag Ajay, Maria Bauzá 0001, Jiajun Wu 0001, Nima Fazeli, Josh Tenenbaum, Alberto Rodriguez 0003, Leslie Pack Kaelbling
ICRA4
2018 Robotic Pick-and-Place of Novel Objects in Clutter with Multi-Affordance Grasping and Cross-Domain Image Matching
abstract
This paper presents a robotic pick-and-place system that is capable of grasping and recognizing both known and novel objects in cluttered environments. The key new feature of the system is that it handles a wide range of object categories without needing any task-specific training data for novel objects. To achieve this, it first uses a category-agnostic affordance prediction algorithm to select and execute among four different grasping primitive behaviors. It then recognizes picked objects with a cross-domain image classification framework that matches observed images to product images. Since product images are readily available for a wide range of objects (e.g., from the web), the system works out-of-the-box for novel objects without requiring any additional training data. Exhaustive experimental results demonstrate that our multi-affordance grasping achieves high success rates for a wide variety of objects in clutter, and our recognition algorithm achieves high accuracy for both known and novel grasped objects. The approach was part of the MIT-Princeton Team system that took 1st place in the stowing task at the 2017 Amazon Robotics Challenge. All code, datasets, and pre-trained models are available online at http://arc.cs.princeton.edu.
Andy Zeng 0001, Shuran Song, Kuan-Ting Yu, Elliott Donlon, Francois Robert Hogan, Maria Bauzá 0001, Daolin Ma, Orion Taylor, Melody Liu, Eudald Romo Grau, Nima Fazeli, Ferran Alet, Nikhil Chavan Dafle, Rachel M. Holladay, Isabella Morona, Prem Qu Nair, Druck Green, Ian H. Taylor, Weber Liu, Thomas A. Funkhouser, Alberto Rodriguez 0003
ICRA11
2018 Augmenting Physical Simulators with Stochastic Neural Networks: Case Study of Planar Pushing and Bouncing
abstract
An efficient, generalizable physical simulator with universal uncertainty estimates has wide applications in robot state estimation, planning, and control. In this paper, we build such a simulator for two scenarios, planar pushing and ball bouncing, by augmenting an analytical rigid-body simulator with a neural network that learns to model uncertainty as residuals. Combining symbolic, deterministic simulators with learnable, stochastic neural nets provides us with expressiveness, efficiency, and generalizability simultaneously. Our model outperforms both purely analytical and purely learned simulators consistently on real, standard benchmarks. Compared with methods that model uncertainty using Gaussian processes, our model runs much faster, generalizes better to new object shapes, and is able to characterize the complex distribution of object trajectories.
Anurag Ajay, Jiajun Wu 0001, Nima Fazeli, Maria Bauzá 0001, Leslie Pack Kaelbling, Josh Tenenbaum, Alberto Rodriguez 0003
IROS3
2017 Empirical evaluation of common contact models for planar impact
abstract
In this paper we evaluate the predictive performance of six commonly used rigid body impact models on real planar impacts captured with a motion tracking system. We propose a metric to evaluate the performance of impact models on a task (based on predicting post impact momentum) and use this metric to tune the six parametric models. We evaluate model performance in predicting impact outcomes against the defined metric and discuss the implications of uncertainty in geometric models and initial conditions. We show that the models can fairly effectively predict the outcomes of single impacts on our chosen task. We motivate further study into consensus and hybrid impact models by showing that a hypothetical hybrid model would significantly outperform the isolated models by providing a post-hoc model that demonstrates an upper bound on the combined predictive power of the models. We use perturbation analysis to compute the predictive range of the models and show that bifurcations can cause the model predictions to cluster into regions of the state space.
Nima Fazeli, Elliott Donlon, Evan M. Drumwright, Alberto Rodriguez 0003
ICRA1
2017 Fundamental Limitations in Performance and Interpretability of Common Planar Rigid-Body Contact Models
Nima Fazeli, Samuel Zapolsky, Evan M. Drumwright, Alberto Rodriguez 0003
ISRR1
2016 More than a million ways to be pushed. A high-fidelity experimental dataset of planar pushing
abstract
Pushing is a motion primitive useful to handle objects that are too large, too heavy, or too cluttered to be grasped. It is at the core of much of robotic manipulation, in particular when physical interaction is involved. It seems reasonable then to wish for robots to understand how pushed objects move. In reality, however, robots often rely on approximations which yield models that are computable, but also restricted and inaccurate. Just how close are those models? How reasonable are the assumptions they are based on? To help answer these questions, and to get a better experimental understanding of pushing, we present a comprehensive and high-fidelity dataset of planar pushing experiments. The dataset contains time-stamped poses of a circular pusher and a pushed object, as well as forces at the interaction. We vary the push interaction in 6 dimensions: surface material, shape of the pushed object, contact position, pushing direction, pushing speed, and pushing acceleration. An industrial robot automates the data capturing along precisely controlled position-velocity-acceleration trajectories of the pusher, which give dense samples of positions and forces of uniform quality. We finish the paper by characterizing the variability of friction, and evaluating the most common assumptions and simplifications made by models of frictional pushing in robotics.
Kuan-Ting Yu, Maria Bauzá 0001, Nima Fazeli, Alberto Rodriguez 0003
IROS3
2015 Identifiability Analysis of Planar Rigid-Body Frictional Contact
Nima Fazeli, Russ Tedrake, Alberto Rodriguez 0003
ISRR (2)1
2015 Quantification of Wave Reflection Using Peripheral Blood Pressure Waveforms
abstract
This paper presents a novel minimally invasive method for quantifying blood pressure (BP) wave reflection in the arterial tree. In this method, two peripheral BP waveforms are analyzed to obtain an estimate of central aortic BP waveform, which is used together with a peripheral BP waveform to compute forward and backward pressure waves. These forward and backward waves are then used to quantify the strength of wave reflection in the arterial tree. Two unique strengths of the proposed method are that 1) it replaces highly invasive central aortic BP and flow waveforms required in many existing methods by less invasive peripheral BP waveforms, and 2) it does not require estimation of characteristic impedance. The feasibility of the proposed method was examined in an experimental swine subject under a wide range of physiologic states and in 13 cardiac surgery patients. In the swine subject, the method was comparable to the reference method based on central aortic BP and flow. In cardiac surgery patients, the method was able to estimate forward and backward pressure waves in the absence of any central aortic waveforms: on the average, the root-mean-squared error between actual versus computed forward and backward pressure waves was less than 5 mmHg, and the error between actual versus computed reflection index was less than 0.03.
Chang-Sei Kim, Nima Fazeli, M. Sean McMurtry, Barry A. Finegan, Jin-Oh Hahn
IEEE J. Biomed. Health Informatics2