Claudia Pérez-D'Arpino

dblp:69/3922 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0002-2949-4214ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 5 since 2021Systems, architecture and hardware · 10 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Inference-Time Policy Steering Through Human Interactions
abstract
Generative policies trained with human demonstrations can autonomously accomplish multimodal, longhorizon tasks. However, during inference, humans are often removed from the policy execution loop, limiting the ability to guide a pre-trained policy towards a specific sub-goal or trajectory shape among multiple predictions. Naive human intervention may inadvertently exacerbate distribution shift, leading to constraint violations or execution failures. To better align policy output with human intent without inducing out-of-distribution errors, we propose an Inference-Time Policy Steering (ITPS) framework that leverages human interactions to bias the generative sampling process, rather than finetuning the policy on interaction data. We evaluate ITPS across three simulated and real-world benchmarks, testing three forms of human interaction and associated alignment distance metrics. Among six sampling strategies, our proposed stochastic sampling with diffusion policy achieves the best trade-off between alignment and distribution shift. Videos are available at https://yanweiw.github.io/itps/.
Lirui Wang, Yilun Du, Balakumar Sundaralingam, Xuning Yang, Yu-Wei Chao, Claudia Pérez-D'Arpino, Dieter Fox, Julie A. Shah
ICRA7
2025 Principles and Guidelines for Evaluating Social Robot Navigation Algorithms
abstract
A major challenge to deploying robots widely is navigation in human-populated environments, commonly referred to as social robot navigation . While the field of social navigation has advanced tremendously in recent years, the fair evaluation of algorithms that tackle social navigation remains hard because it involves not just robotic agents moving in static environments but also dynamic human agents and their perceptions of the appropriateness of robot behavior. In contrast, clear, repeatable, and accessible benchmarks have accelerated progress in fields like computer vision, natural language processing and traditional robot navigation by enabling researchers to fairly compare algorithms, revealing limitations of existing solutions and illuminating promising new directions. We believe the same approach can benefit social navigation. In this article, we pave the road toward common, widely accessible, and repeatable benchmarking criteria to evaluate social robot navigation. Our contributions include (a) a definition of a socially navigating robot as one that respects the principles of safety, comfort, legibility, politeness, social competency, agent understanding, proactivity, and responsiveness to context, (b) guidelines for the use of metrics, development of scenarios, benchmarks, datasets, and simulators to evaluate social navigation, and (c) a design of a social navigation metrics framework to make it easier to compare results from different simulators, robots, and datasets.
Anthony G. Francis, Claudia Pérez-D'Arpino, Chengshu Li 0002, Fei Xia 0002, Alexandre Alahi, Rachid Alami 0001, Aniket Bera, Abhijat Biswas, Joydeep Biswas, Rohan Chandra, Hao-Tien Chiang, Michael Everett, Sehoon Ha, Justin W. Hart, Jonathan P. How, Haresh Karnan, Tsang-Wei Edward Lee, Luis Manso, Reuth Mirsky, Sören Pirk, Phani-Teja Singamaneni, Peter Stone 0001, Ada V. Taylor, Pete Trautman, Nathan Tsoi, Marynel Vázquez, Xuesu Xiao, Peng Xu 0010, Naoki Yokoyama, Alexander Toshev, Roberto Martin Martin
ACM Trans. Hum. Robot Interact.2
2024 Fast Explicit-Input Assistance for Teleoperation in Clutter
abstract
The performance of prediction-based assistance for robot teleoperation degrades in unseen or goal-rich environments due to incorrect or quickly-changing intent inferences. Poor predictions can confuse operators or cause them to change their control input to implicitly signal their goal. We present a new assistance interface for robotic manipulation where an operator can explicitly communicate a manipulation goal by pointing the end-effector. The pointing target specifies a region for local pose generation and optimization, providing interactive control over grasp and placement pose candidates. We evaluate this explicit pointing interface against an implicit inference-based assistance scheme and an unassisted control condition in a within-subjects user study (N=20), where participants teleoperate a simulated robot to complete a multi-step singulation and stacking task in cluttered environments. We find that operators prefer the explicit interface, experience fewer pick failures and report lower cognitive workload. Our code is available at: github.com/NVlabs/fast-explicit-teleop.
Nick Walker 0001, Xuning Yang, Animesh Garg, Maya Cakmak, Dieter Fox, Claudia Pérez-D'Arpino
IROS6
2024 Experimental Assessment of Human-Robot Teaming for Multi-step Remote Manipulation with Expert Operators
abstract
Remote robot manipulation with human control enables applications in which safety and environmental constraints are adverse to humans (e.g., underwater, space robotics and disaster response) or the complexity of the task demands human-level cognition and dexterity (e.g., robotic surgery and manufacturing). These systems typically use direct teleoperation at the motion level and are usually limited to low-DOF arms and two-dimensional (2D) perception. Improving dexterity and situational awareness demands new interaction and planning workflows. We explore the use of human–robot teaming through teleautonomy with assisted planning for remote control of a dual-arm dexterous robot for multi-step manipulation, and conduct a within-subjects experimental assessment (n = 12 expert users) to compare it with direct teleoperation with an imitation controller with 2D and three-dimensional (3D) perception, as well as teleoperation through a teleautonomy interface. The proposed assisted planning approach achieves task times comparable with direct teleoperation while improving other objective and subjective metrics, including re-grasps, collisions, and TLX workload. Assisted planning in the teleautonomy interface achieves faster task execution and removes a significant interaction with the operator’s expertise level, resulting in a performance equalizer across users. Our study protocol, metrics, and models for statistical analysis might also serve as a general benchmarking framework in teleoperation domains. Accompanying video and reference R code: https://people.csail.mit.edu/cdarpino/THRIteleop/
Claudia Pérez-D'Arpino, Rebecca P. Khurshid, Julie A. Shah
ACM Trans. Hum. Robot Interact.1
2023 Learning Human-to-Robot Handovers from Point Clouds
abstract
We propose the first framework to learn control policies for vision-based human-to-robot handovers, a critical task for human-robot interaction. While research in Embodied AI has made significant progress in training robot agents in simulated environments, interacting with humans remains challenging due to the difficulties of simulating humans. Fortunately, recent research has developed realistic simulated environments for human-to-robot handovers. Leveraging this result, we introduce a method that is trained with a human-in-the-loop via a two-stage teacher-student framework that uses motion and grasp planning, reinforcement learning, and self-supervision. We show significant performance gains over baselines on a simulation benchmark, sim-to-sim transfer and sim-to-real transfer. Video and code are available at https://handover-sim2real.github.io.
Sammy Joe Christen, Wei Yang 0019, Claudia Pérez-D'Arpino, Otmar Hilliges, Dieter Fox, Yu-Wei Chao
CVPR3
2021 Robot Navigation in Constrained Pedestrian Environments using Reinforcement Learning
abstract
Navigating fluently around pedestrians is a necessary capability for mobile robots deployed in human environments, such as buildings and homes. While research on social navigation has focused mainly on the scalability with the number of pedestrians in open spaces, typical indoor environments present the additional challenge of constrained spaces such as corridors and doorways that limit maneuverability and influence patterns of pedestrian interaction. We present an approach based on reinforcement learning (RL) to learn policies capable of dynamic adaptation to the presence of moving pedestrians while navigating between desired locations in constrained environments. The policy network receives guidance from a motion planner that provides waypoints to follow a globally planned trajectory, whereas RL handles the local interactions. We explore a compositional principle for multi-layout training and find that policies trained in a small set of geometrically simple layouts successfully generalize to more complex unseen layouts that exhibit composition of the structural elements available during training. Going beyond walls-world like domains, we show transfer of the learned policy to unseen 3D reconstructions of two real environments. These results support the applicability of the compositional principle to navigation in real-world buildings and indicate promising usage of multi-agent simulation within reconstructed environments for tasks that involve interaction. https://ai.stanford.edu/∼cdarpino/socialnavconstrained/
Claudia Pérez-D'Arpino, Patrick Goebel, Roberto Martin Martin, Silvio Savarese
ICRA1
2021 iGibson 1.0: A Simulation Environment for Interactive Tasks in Large Realistic Scenes
abstract
We present iGibson 1.0, a novel simulation environment to develop robotic solutions for interactive tasks in large-scale realistic scenes. Our environment contains 15 fully interactive home-sized scenes with 108 rooms populated with rigid and articulated objects. The scenes are replicas of real-world homes, with distribution and the layout of objects aligned to those of the real world. iGibson 1.0 integrates several key features to facilitate the study of interactive tasks: i) generation of high-quality virtual sensor signals (RGB, depth, segmentation, LiDAR, flow and so on), ii) domain randomization to change the materials of the objects (both visual and physical) and/or their shapes, iii) integrated sampling-based motion planners to generate collision-free trajectories for robot bases and arms, and iv) intuitive human-iGibson interface that enables efficient collection of human demonstrations. Through experiments, we show that the full interactivity of the scenes enables agents to learn useful visual representations that accelerate the training of downstream manipulation tasks. We also show that iGibson features enable the generalization of navigation agents, and that the human-iGibson interface and integrated motion planners facilitate efficient imitation learning of human demonstrated (mobile) manipulation behaviors. iGibson 1.0 is open-source, equipped with comprehensive examples and documentation. For more information, visit our project website: http://svl.stanford.edu/igibson/.
Bokui Shen, Fei Xia 0002, Chengshu Li 0002, Roberto Martin Martin, Linxi Fan, Guanzhi Wang, Claudia Pérez-D'Arpino, Shyamal Buch, Sanjana Srivastava, Lyne Tchapmi, Micael Tchapmi, Kent Vainio, Josiah Wong, Li Fei-Fei 0001, Silvio Savarese
IROS7
2017 C-LEARN: Learning geometric constraints from demonstrations for multi-step manipulation in shared autonomy
abstract
Learning from demonstrations has been shown to be a successful method for non-experts to teach manipulation tasks to robots. These methods typically build generative models from demonstrations and then use regression to reproduce skills. However, this approach has limitations to capture hard geometric constraints imposed by the task. On the other hand, while sampling and optimization-based motion planners exist that reason about geometric constraints, these are typically carefully hand-crafted by an expert. To address this technical gap, we contribute with C-LEARN, a method that learns multi-step manipulation tasks from demonstrations as a sequence of keyframes and a set of geometric constraints. The system builds a knowledge base for reaching and grasping objects, which is then leveraged to learn multi-step tasks from a single demonstration. C-LEARN supports multi-step tasks with multiple end effectors; reasons about SE(3) volumetric and CAD constraints, such as the need for two axes to be parallel; and offers a principled way to transfer skills between robots with different kinematics. We embed the execution of the learned tasks within a shared autonomy framework, and evaluate our approach by analyzing the success rate when performing physical tasks with a dual-arm Optimas robot, comparing the contribution of different constraints models, and demonstrating the ability of C-LEARN to transfer learned tasks by performing them with a legged dual-arm Atlas robot in simulation.
Claudia Pérez-D'Arpino, Julie A. Shah
ICRA1
2016 Fast Motion Prediction for Collaborative Robotics
Claudia Pérez-D'Arpino, Julie A. Shah
IJCAI1
2015 Fast target prediction of human reaching motion for cooperative human-robot manipulation tasks using time series classification
abstract
Interest in human-robot coexistence, in which humans and robots share a common work volume, is increasing in manufacturing environments. Efficient work coordination requires both awareness of the human pose and a plan of action for both human and robot agents in order to compute robot motion trajectories that synchronize naturally with human motion. In this paper, we present a data-driven approach that synthesizes anticipatory knowledge of both human motions and subsequent action steps in order to predict in real-time the intended target of a human performing a reaching motion. Motion-level anticipatory models are constructed using multiple demonstrations of human reaching motions. We produce a library of motions from human demonstrations, based on a statistical representation of the degrees of freedom of the human arm, using time series analysis, wherein each time step is encoded as a multivariate Gaussian distribution. We demonstrate the benefits of this approach through offline statistical analysis of human motion data. The results indicate a considerable improvement over prior techniques in early prediction, achieving 70% or higher correct classification on average for the first third of the trajectory (<; 500msec). We also indicate proof-of-concept through the demonstration of a human-robot cooperative manipulation task performed with a PR2 robot. Finally, we analyze the quality of task-level anticipatory knowledge required to improve prediction performance early in the human motion trajectory.
Claudia Pérez-D'Arpino, Julie A. Shah
ICRA1
2015 Human-robot co-navigation using anticipatory indicators of human walking motion
abstract
Mobile, interactive robots that operate in human-centric environments need the capability to safely and efficiently navigate around humans. This requires the ability to sense and predict human motion trajectories and to plan around them. In this paper, we present a study that supports the existence of statistically significant biomechanical turn indicators of human walking motions. Further, we demonstrate the effectiveness of these turn indicators as features in the prediction of human motion trajectories. Human motion capture data is collected with predefined goals to train and test a prediction algorithm. Use of anticipatory features results in improved performance of the prediction algorithm. Lastly, we demonstrate the closed-loop performance of the prediction algorithm using an existing algorithm for motion planning within dynamic environments. The anticipatory indicators of human walking motion can be used with different prediction and/or planning algorithms for robotics; the chosen planning and prediction algorithm demonstrates one such implementation for human-robot co-navigation.
Vaibhav V. Unhelkar, Claudia Pérez-D'Arpino, Leia A. Stirling 0001, Julie A. Shah
ICRA2
2014 A summary of team MIT's approach to the virtual robotics challenge
abstract
The paper describes the system developed by researchers from MIT for the Defense Advanced Research Projects Agency's (DARPA) Virtual Robotics Challenge (VRC), held in June 2013. The VRC was the first competition in the DARPA Robotics Challenge (DRC), a program that aims to “develop ground robotic capabilities to execute complex tasks in dangerous, degraded, human-engineered environments”. The VRC required teams to guide a model of Boston Dynamics' humanoid robot, Atlas, through driving, walking, and manipulation tasks in simulation. Team MIT's user interface, the Viewer, provided the operator with a unified representation of all available information. A 3D rendering of the robot depicted its most recently estimated body state with respect to the surrounding environment, represented by point clouds and texture-mapped meshes as sensed by on-board LIDAR and fused over time.
Russ Tedrake, Maurice Fallon, Sisir Karumanchi, Scott Kuindersma, Matthew E. Antone, Toby Schneider, Thomas M. Howard, Matthew R. Walter, Hongkai Dai, Robin Deits, Michael Fleder, Dehann Fourie, Riad I. Hammoud, Sachithra Hemachandra, P. Ilardi, Claudia Pérez-D'Arpino, Sudeep Pillai, Andres Valenzuela, Cecilia Cantu, C. Dolan, I. Evans, S. Jorgensen, J. Kristeller, Julie A. Shah, Karl Iagnemma, Seth J. Teller
ICRA16
2010 Generalized Bilateral MIMO Control by States Convergence with time delay and application for the teleoperation of a 2-DOF helicopter
abstract
Bilateral Control by States Convergence is a novel and little exploited control strategy that has been successfully applied to the teleoperation of robotic manipulators using SISO control. Based on the state space representation, the main philosophy of this control strategy consists in achieving convergence of states between the master and the slave, by setting the dynamical behavior of the master-slave error as a states-independent autonomous system. This paper presents a generalization of this strategy to MIMO systems, with time delay in the master-slave communication channels. It is demonstrated how the feedback gains required by the state convergence control schema can be found by solving a set of ((m × n) + (m × m) + n ) nonlinear equations for a system with m inputs, n states and m outputs. Unlike previous research, which has been applied only to manipulators considering 1-DOF for the states-convergence control loop, the extension to the general MIMO case has allowed to apply the technique to the teleoperation of a 2-DOF helicopter. A decoupling network and states-feedback are used for local control, while the states-convergence control manages the bilateral issues. Simulation results are presented, showing a satisfactory performance of the control strategy.
Claudia Pérez-D'Arpino, Wilfredis Medina Meléndez, Leonardo Fermín-Leon, Juan M. Bogado, Rafael R. Torrealba, Gerardo Fernández-López
ICRA1
2010 Through the development of a biomechatronic knee prosthesis for transfemoral amputees: Mechanical design and manufacture, human gait characterization, intelligent control strategies and tests
abstract
This paper presents the development of a biomechatronic knee prosthesis for transfemoral amputees. This kind of prostheses are considered `intelligent' because they are able to automatically adapt their response at the knee axis, as a natural knee does. This behavior is achieved by characterizing the amputee's gait through the signals captured with instrumentation of a prosthesis, which provides feedback about its current state along the gait cycle and therefore responds with the corresponding control action. In this case, unlike other commercially available intelligent knee prostheses, gait cycle characterization is based on accelerometers signals processed by an events detection algorithm. Two intelligent control strategies are presented: a bio-inspired approach, that consists of using a central pattern generator to generate a knee angle reference to be followed by the prosthesis during walking, and an adaptive scheme, that applies a control action proportional to the knee angle according to an auto-adaptive parameter dependant on gait speed. The mechanical design of the prosthesis is also presented, showing the knee joint mechanism and part of the manufacturing process. Results obtained from walking tests with both able body and amputees are shown, demonstrating the positive performance of the prosthesis in several aspects. Future works aimed at a finished product are also stated.
Rafael R. Torrealba, Claudia Pérez-D'Arpino, José Cappelletto, Leonardo Fermín-Leon, Gerardo Fernández-López, Juan C. Grieco
ICRA2