VLDB 2026 Research / reviewers in the wild / expert
Stefanos Nikolaidis
dblp:62/6555
· DBLP profile ↗
59ranked-venue papers
8as first author
41since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 7 first-author · 37 since 2021Human-computer interaction and ubiquitous computing · 21 · 8 first-author · 9 since 2021Systems, architecture and hardware · 13 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Objective Covariance Matrix Adaptation MAP-AnnealingabstractQuality-Diversity (QD) optimization is an emerging field that focuses on finding a set of behaviorally diverse and high-quality solutions. While the quality is typically defined w.r.t. a single objective function, recent work on Multi-Objective Quality-Diversity (MOQD) extends QD optimization to simultaneously optimize multiple objective functions. This opens up multi-objective applications for QD, such as generating a diverse set of game maps that maximize difficulty, realism, or other properties. Existing MOQD algorithms use non-adaptive methods such as mutation and crossover to search for non-dominated solutions and construct an archive of Pareto Sets (PS). However, recent work in QD has demonstrated enhanced performance through the use of covariance-based evolution strategies for adaptive solution search. We propose bringing this insight into the MOQD problem, and introduce MO-CMA-MAE, a new MOQD algorithm that leverages Covariance Matrix AdaptationEvolution Strategies (CMA-ES) to optimize the hypervolume associated with every PS within the archive. We test MO-CMA-MAE on three MOQD domains, and for generating maps of a co-operative video game, showing significant improvements in performance. Shihan Zhao, Stefanos Nikolaidis |
GECCO | 2 |
| 2025 | Contrastive Learning from Exploratory Actions: Leveraging Natural Interactions for Preference ElicitationabstractPeople have a variety of preferences for how robots behave. To understand and reason about these preferences, robots aim to learn a reward function that describes how aligned robot behaviors are with a user's preferences. Good representations of a robot's behavior can significantly reduce the time and effort required for a user to teach the robot their preferences. Specifying these representations-what “features“ of the robot's behavior matter to users-remains a difficult problem; Features learned from raw data lack semantic meaning and features learned from user data require users to engage in tedious labeling processes. Our key insight is that users tasked with customizing a robot are intrinsically motivated to produce labels through exploratory search; they explore behaviors that they find interesting and ignore behaviors that are irrelevant. To harness this novel data source of exploratory actions, we propose contrastive learning from exploratory actions (CLEA) to learn trajectory features that are aligned with features that users care about. We learned CLEA features from exploratory actions users performed in an open-ended signal design activity$(N=25)$with a Kuri robot, and evaluated CLEA features through a second user study with a different set of users$(N=42)$. CLEA features outperformed self-supervised features when eliciting user preferences over four metrics: completeness, simplicity, minimality, and explainability. Nathaniel Dennler, Stefanos Nikolaidis, Maja J. Mataric |
HRI | 2 |
| 2025 | Soft and Compliant Contact-Rich Hair Manipulation and CareabstractHair care robots can help address labor shortages in elderly care while enabling those with limited mobility to maintain their hair-related identity. We present MOE-Hair, a soft robot system that performs three hair-care tasks: head patting, finger combing, and hair grasping. The system features a tendon-driven soft robot end-effector (MOE) with a wrist-mounted RGBD camera, leveraging both mechanical compliance for safety and visual force sensing through deformation. In testing with a force-sensorized mannequin head, MOE achieved comparable hair-grasping effectiveness while applying significantly less force than rigid grippers. Our novel force estimation method combines visual deformation data and tendon tensions from actuators to infer applied forces, reducing sensing errors by up to 60.1% and 20.3% compared to actuator current load-only and depth image-only baselines, respectively. A user study with 12 participants demonstrated statistically significant preferences for MOE-Hair over a baseline system in terms of comfort, effectiveness, and appropriate force application. These results demonstrate the unique advantages of soft robots in contact-rich hair-care tasks, while highlighting the importance of precise force control despite the inherent compliance of the system. Videos, data, and code are available at moehair.github.io. Uksang Yoo, Nathaniel Dennler, Eliot Xing, Maja J. Mataric, Stefanos Nikolaidis, Jeffrey Ichnowski, Jean Oh |
HRI | 5 |
| 2025 | Integrating Field of View in Human-Aware Collaborative PlanningabstractIn human-robot collaboration (HRC), it is crucial for robot agents to consider humans' knowledge of their surroundings. In reality, humans possess a narrow field of view (FOV), limiting their perception. However, research on HRC often overlooks this aspect and presumes an omniscient human collaborator. Our study addresses the challenge of adapting to the evolving subtask intent of humans while accounting for their limited FOV. We integrate FOV within the humanaware probabilistic planning framework. To account for large state spaces due to considering FOV, we propose a hierarchical online planner that efficiently finds approximate solutions while enabling the robot to explore low-level action trajectories that enter the human FOV, influencing their intended subtask. Through user study with our adapted cooking domain, we demonstrate our FOV-aware planner reduces human's interruptions and redundant actions during collaboration by adapting to human perception limitations. We extend these findings to a virtual reality kitchen environment, where we observe similar collaborative behaviors. Ya-Chuan Hsu, Michael Defranco, Rutvik Patel, Stefanos Nikolaidis |
ICRA | 4 |
| 2025 | Adaptively Coordinating with Novel Partners via Learned Latent StrategiesabstractAdaptation is the cornerstone of effective collaboration among heterogeneous team members. In human-agent teams, artificial agents need to adapt to their human partners in real time, as individuals often have unique preferences and policies that may change dynamically throughout interactions. This becomes particularly challenging in tasks with time pressure and complex strategic spaces, where identifying partner behaviors and selecting suitable responses is difficult.
In this work, we introduce a strategy-conditioned cooperator framework that learns to represent, categorize, and adapt to a broad range of potential partner strategies in real-time.
Our approach encodes strategies with a variational autoencoder to learn a latent strategy space from agent trajectory data, identifies distinct strategy types through clustering, and trains a cooperator agent conditioned on these clusters by generating partners of each strategy type.
For online adaptation to novel partners, we leverage a fixed-share regret minimization algorithm that dynamically infers and adjusts the partner's strategy estimation during interaction.
We evaluate our method in a modified version of the Overcooked domain, a complex collaborative cooking environment that requires effective coordination among two players with a diverse potential strategy space.
Through these experiments and an online user study, we demonstrate that our proposed agent achieves state of the art performance compared to existing baselines when paired with novel human, and agent teammates. Benjamin Li, Shuyang Shi, Lucia Romero, Huao Li, Yaqi Xie 0001, Woojun Kim, Stefanos Nikolaidis, Charles Lewis, Katia P. Sycara, Simon Stepputtis |
NeurIPS | 7 |
| 2025 | Proactive Contingency-Aware Task Allocation and Scheduling in Multi-Robot Multi-Human Cells via Hindsight OptimizationabstractMulti-robot systems are becoming more common in various real-world applications, such as manufacturing and warehouse logistics. However, task allocation and scheduling for a multi-agent team face complex challenges due to the need to simultaneously consider time-extended tasks, task constraints, and uncertainties in execution. Potential task failures or contingencies can add additional tasks to recover from the failures, and reactively addressing contingencies can decrease teaming efficiency. To efficiently and proactively consider contingencies, this paper proposes treating the problem as a multi-robot task allocation under uncertainty problem. We suggest a hierarchical approach that divides the problem into two layers. We use mathematical program formulation for the lower layer to find the optimal solution for a deterministic multi-robot task allocation problem with known task outcomes. The higher-layer search intelligently generates more likely combinations of contingency scenarios and calls the inner-level search repeatedly to find the optimal task allocation sequence for the given scenario. We validate our results in simulation for manufacturing applications and demonstrate that our method can reduce the effect of potential delays from contingencies.Note to Practitioners—Automation engineers interested in deploying robotic cells in low-volume applications need to consider contingency handling. When the occurrence of contingencies can be characterized as probability distributions, it is often useful to consider using a proactive approach for task allocation and scheduling. To implement our algorithm, automation engineers will need to develop a hierarchical task network specified by domain experts that models task constraints and a task-agent duration model, which may be generated from simulation environments. Furthermore, they must identify tasks that can result in contingencies and describe them with a probabilistic model. This model can be generated from historical data and/or real-world experiments. Lastly, for addressing the contingency, the practitioner will need to specify a task procedure to recover from a specific contingency type. To run the algorithm, we found that repeatedly approximating the best proactive task allocation for a fixed computation budget and dispatching the best tasks worked well. The computation budget required to approximate the best task allocation is directly affected by the number of contingency scenarios that can be sampled. Therefore, the practitioner must determine a suitable computational budget empirically based on the number of contingencies that can occur. Neel Dhanaraj, Heramb Nemlekar, Stefanos Nikolaidis, Satyandra K. Gupta |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Covariance Matrix Adaptation MAP-Annealing: Theory and ExperimentsabstractSingle-objective optimization algorithms search for the single highest quality solution with respect to an objective. Quality diversity (QD) optimization algorithms, such as Covariance Matrix Adaptation MAP-Elites (CMA-ME), search for a collection of solutions that are both high quality with respect to an objective and diverse with respect to specified measure functions. However, CMA-ME suffers from three major limitations highlighted by the QD community: prematurely abandoning the objective in favor of exploration, struggling to explore flat objectives, and having poor performance for low-resolution archives. We propose a new QD algorithm, CMA MAP-Annealing (CMA-MAE), and its differentiable QD variant, CMA-MAE via a Gradient Arborescence (CMA-MAEGA), that address all three limitations. We provide theoretical justifications for the new algorithm with respect to each limitation. Our theory informs our experiments, which support the theory and show that CMA-MAE achieves state-of-the-art performance and robustness on standard QD benchmark and reinforcement learning domains. Shihan Zhao, Bryon Tjanaka, Matthew C. Fontaine, Stefanos Nikolaidis |
ACM Trans. Evol. Learn. Optim. | 4 |
| 2024 | Quality-Diversity Generative Sampling for Learning with Synthetic DataabstractGenerative models can serve as surrogates for some real data sources by creating synthetic training datasets, but in doing so they may transfer biases to downstream tasks. We focus on protecting quality and diversity when generating synthetic training datasets. We propose quality-diversity generative sampling (QDGS), a framework for sampling data uniformly across a user-defined measure space, despite the data coming from a biased generator. QDGS is a model-agnostic framework that uses prompt guidance to optimize a quality objective across measures of diversity for synthetically generated data, without fine-tuning the generative model. Using balanced synthetic datasets generated by QDGS, we first debias classifiers trained on color-biased shape datasets as a proof-of-concept. By applying QDGS to facial data synthesis, we prompt for desired semantic concepts, such as skin tone and age, to create an intersectional dataset with a combined blend of visual features. Leveraging this balanced data for training classifiers improves fairness while maintaining accuracy on facial recognition benchmarks. Code available at: https://github.com/Cylumn/qd-generative-sampling. Allen Chang, Matthew C. Fontaine, Serena Booth, Maja J. Mataric, Stefanos Nikolaidis |
AAAI | 5 |
| 2024 | Density Descent for Diversity OptimizationabstractDiversity optimization seeks to discover a set of solutions that elicit diverse features. Prior work has proposed Novelty Search (NS), which, given a current set of solutions, seeks to expand the set by finding points in areas of low density in the feature space. However, to estimate density, NS relies on a heuristic that considers the k-nearest neighbors of the search point in the feature space, which yields a weaker stability guarantee. We propose Density Descent Search (DDS), an algorithm that explores the feature space via CMA-ES on a continuous density estimate of the feature space that also provides a stronger stability guarantee. We experiment with DDS and two density estimation methods: kernel density estimation (KDE) and continuous normalizing flow (CNF). On several standard diversity optimization benchmarks, DDS outperforms NS, the recently proposed MAP-Annealing algorithm, and other state-of-the-art baselines. Additionally, we prove that DDS with KDE provides stronger stability guarantees than NS, making it more suitable for adaptive optimizers. Furthermore, we prove that NS is a special case of DDS that descends a KDE of the feature space. David H. Lee, Anishalakshmi V. Palaparthi, Matthew C. Fontaine, Bryon Tjanaka, Stefanos Nikolaidis |
GECCO | 5 |
| 2024 | Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement LearningabstractTraining generally capable agents that thoroughly explore their environment and
learn new and diverse skills is a long-term goal of robot learning. Quality Diversity
Reinforcement Learning (QD-RL) is an emerging research area that blends the
best aspects of both fields – Quality Diversity (QD) provides a principled form
of exploration and produces collections of behaviorally diverse agents, while
Reinforcement Learning (RL) provides a powerful performance improvement
operator enabling generalization across tasks and dynamic environments. Existing
QD-RL approaches have been constrained to sample efficient, deterministic off-
policy RL algorithms and/or evolution strategies and struggle with highly stochastic
environments. In this work, we, for the first time, adapt on-policy RL, specifically
Proximal Policy Optimization (PPO), to the Differentiable Quality Diversity (DQD)
framework and propose several changes that enable efficient optimization and
discovery of novel skills on high-dimensional, stochastic robotics tasks. Our new
algorithm, Proximal Policy Gradient Arborescence (PPGA), achieves state-of-
the-art results, including a 4x improvement in best reward over baselines on the
challenging humanoid domain. Sumeet Batra, Bryon Tjanaka, Matthew C. Fontaine, Aleksei Petrenko, Stefanos Nikolaidis, Gaurav S. Sukhatme |
ICLR | 5 |
| 2024 | Multi-Robot Task Allocation Under Uncertainty Via Hindsight OptimizationabstractMulti-robot systems are becoming increasingly prevalent in various real-world applications, such as manufacturing and warehouse logistics. These systems face complex challenges in 1) task allocation due to factors like time-extended tasks, and agent specialization, and 2) uncertainties in task execution. Potential task failures can add further contingency tasks to recover from the failure, thereby causing delays. This paper addresses the problem of Multi-Robot Task Allocation under Uncertainty by proposing a hierarchical approach that decouples the problem into two levels. We use a low-level optimization formulation to find the optimal solution for a deterministic multi-robot task allocation problem with known task outcomes. The higher-level search intelligently generates more likely combinations of failures and calls the inner-level search repeatedly to find the optimal task allocation sequence, given the known outcomes. We validate our results in simulation for a manufacturing domain and demonstrate that our method can reduce the effect of potential delays from contingencies. We show that our algorithm is computationally efficient while improving average makespan compared to other baselines. Neel Dhanaraj, Jeon Ho Kang, Heramb Nemlekar, Stefanos Nikolaidis, Satyandra K. Gupta |
ICRA | 5 |
| 2024 | Guidance Graph Optimization for Lifelong Multi-Agent Path Finding
Yulun Zhang 0002, Varun Bhatt, Stefanos Nikolaidis, Jiaoyang Li 0001 |
IJCAI | 4 |
| 2024 | BayRnTune: Adaptive Bayesian Domain Randomization via Strategic Fine-tuningabstractDomain randomization (DR), which entails training a policy with randomized dynamics, has proven to be a simple yet effective algorithm for reducing the gap between simulation and the real world. However, DR often requires careful tuning of randomization parameters. Methods like Bayesian Domain Randomization (Bayesian DR) and Active Domain Randomization (Adaptive DR) address this issue by automating parameter range selection using real-world experience. While effective, these algorithms often require long computation time, as a new policy is trained from scratch every iteration. In this work, we propose Adaptive Bayesian Domain Randomization via Strategic Fine-tuning (BayRnTune), which inherits the spirit of BayRn but aims to significantly accelerate the learning processes by fine-tuning from previously learned policy. This idea leads to a critical question: which previous policy should we use as a prior during fine-tuning? We investigated four different fine-tuning strategies and compared them against baseline algorithms in five simulated environments, ranging from simple benchmark tasks to more complex legged robot environments. Our analysis demonstrates that our method yields better rewards in the same amount of timesteps compared to vanilla domain randomization or Bayesian DR. Tianle Huang, Nitish Sontakke, K. Niranjan Kumar, Irfan A. Essa, Stefanos Nikolaidis, Dennis W. Hong, Sehoon Ha |
IROS | 5 |
| 2024 | Signal Temporal Logic-Guided Apprenticeship LearningabstractApprenticeship learning crucially depends on effectively learning rewards, and hence control policies from user demonstrations. Of particular difficulty is the setting where the desired task consists of a number of sub-goals with temporal dependencies. The quality of inferred rewards and hence policies are typically limited by the quality of demonstrations, and poor inference of these can lead to undesirable outcomes. In this paper, we show how temporal logic specifications that describe high level task objectives, are encoded in a graph to define a temporal-based metric that reasons about behaviors of demonstrators and the learner agent to improve the quality of inferred rewards and policies. Through experiments on a diverse set of robot manipulator simulations, we show how our framework overcomes the drawbacks of prior literature by drastically improving the number of demonstrations required to learn a control policy. Aniruddh Gopinath Puranic, Jyotirmoy V. Deshmukh, Stefanos Nikolaidis |
IROS | 3 |
| 2024 | Enabling Adaptive Agent Training in Open-Ended Simulators by Targeting DiversityabstractThe wider application of end-to-end learning methods to embodied decision-making domains remains bottlenecked by their reliance on a superabundance of training data representative of the target domain.
Meta-reinforcement learning (meta-RL) approaches abandon the aim of zero-shot *generalization*—the goal of standard reinforcement learning (RL)—in favor of few-shot *adaptation*, and thus hold promise for bridging larger generalization gaps.
While learning this meta-level adaptive behavior still requires substantial data, efficient environment simulators approaching real-world complexity are growing in prevalence.
Even so, hand-designing sufficiently diverse and numerous simulated training tasks for these complex domains is prohibitively labor-intensive.
Domain randomization (DR) and procedural generation (PG), offered as solutions to this problem, require simulators to possess carefully-defined parameters which directly translate to meaningful task diversity—a similarly prohibitive assumption.
In this work, we present **DIVA**, an evolutionary approach for generating diverse training tasks in such complex, open-ended simulators.
Like unsupervised environment design (UED) methods, DIVA can be applied to arbitrary parameterizations, but can additionally incorporate realistically-available domain knowledge—thus inheriting the *flexibility* and *generality* of UED, and the supervised *structure* embedded in well-designed simulators exploited by DR and PG.
Our empirical results showcase DIVA's unique ability to overcome complex parameterizations and successfully train adaptive agent behavior, far outperforming competitive baselines from prior literature.
These findings highlight the potential of such *semi-supervised environment design* (SSED) approaches, of which DIVA is the first humble constituent, to enable training in realistic simulated domains, and produce more robust and capable adaptive agents.
Our code is available at [https://github.com/robbycostales/diva](https://github.com/robbycostales/diva). Robby Costales, Stefanos Nikolaidis |
NeurIPS | 2 |
| 2024 | Multi-Robot Coordination and Layout Design for Automated Warehousing (Extended Abstract)abstractWith the rapid progress in Multi-Agent Path Finding (MAPF), researchers have studied how MAPF algorithms can be deployed to coordinate hundreds of robots in large automated warehouses. While most works try to improve the throughput of such warehouses by developing better MAPF algorithms, we focus on improving the throughput by optimizing the warehouse layout. We show that, even with state-of-the-art MAPF algorithms, commonly used human-designed layouts can lead to congestion for warehouses with large numbers of robots and thus have limited scalability. We extend existing automatic scenario generation methods to optimize warehouse layouts. Results show that our optimized warehouse layouts (1) reduce traffic congestion and thus improve throughput, (2) improve the scalability of the automated warehouses by doubling the number of robots in some cases, and (3) are capable of generating layouts with user-specified diversity measures. We include the source code at: https://github.com/lunjohnzhang/warehouse_env_gen_public Yulun Zhang 0002, Matthew C. Fontaine, Varun Bhatt, Stefanos Nikolaidis, Jiaoyang Li 0001 |
SOCS | 4 |
| 2024 | Arbitrarily Scalable Environment Generators via Neural Cellular Automata (Extended Abstract)abstractWe study the problem of generating arbitrarily large environments to improve the throughput of multi-robot systems. Prior work proposes Quality Diversity (QD) algorithms as an effective method for optimizing the environments of automated warehouses. However, these approaches optimize only relatively small environments, falling short when it comes to replicating real-world warehouse sizes. The challenge arises from the exponential increase in the search space as the environment size increases. Additionally, the previous methods have only been tested with up to 350 robots in simulations, while practical warehouses could host thousands of robots. In this paper, instead of optimizing environments, we propose to optimize Neural Cellular Automata (NCA) environment generators via QD algorithms. We train a collection of NCA generators with QD algorithms in small environments and then generate arbitrarily large environments from the generators at test time. We show that NCA environment generators maintain consistent, regularized patterns regardless of environment size, significantly enhancing the scalability of multi-robot systems in two different domains with up to 2,350 robots. Additionally, we demonstrate that our method scales a single-agent reinforcement learning policy to arbitrarily large environments with similar patterns. We include the source code at https://github.com/lunjohnzhang/warehouse_env_gen_nca_public. Yulun Zhang 0002, Matthew C. Fontaine, Varun Bhatt, Stefanos Nikolaidis, Jiaoyang Li 0001 |
SOCS | 4 |
| 2023 | Covariance Matrix Adaptation MAP-AnnealingabstractSingle-objective optimization algorithms search for the single highestquality solution with respect to an objective.Quality diversity (QD) optimization algorithms, such as Covariance Matrix Adaptation MAP-Elites (CMA-ME), search for a collection of solutions that are both high-quality with respect to an objective and diverse with respect to specified measure functions.However, CMA-ME suffers from three major limitations highlighted by the QD community: prematurely abandoning the objective in favor of exploration, struggling to explore flat objectives, and having poor performance for lowresolution archives.We propose a new quality diversity algorithm, Covariance Matrix Adaptation MAP-Annealing (CMA-MAE), that addresses all three limitations.We provide theoretical justifications for the new algorithm with respect to each limitation.Our theory informs our experiments, which support the theory and show that CMA-MAE achieves state-of-the-art performance and robustness. Matthew C. Fontaine, Stefanos Nikolaidis |
GECCO | 2 |
| 2023 | pyribs: A Bare-Bones Python Library for Quality Diversity OptimizationabstractRecent years have seen a rise in the popularity of quality diversity (QD) optimization, a branch of optimization that seeks to find a collection of diverse, high-performing solutions to a given problem. To grow further, we believe the QD community faces two challenges: developing a framework to represent the field's growing array of algorithms, and implementing that framework in software that supports a range of researchers and practitioners. To address these challenges, we have developed pyribs, a library built on a highly modular conceptual QD framework. By replacing components in the conceptual framework, and hence in pyribs, users can compose algorithms from across the QD literature; equally important, they can identify unexplored algorithm variations. Furthermore, pyribs makes this framework simple, flexible, and accessible, with a user-friendly API supported by extensive documentation and tutorials. This paper overviews the creation of pyribs, focusing on the conceptual framework that it implements and the design principles that have guided the library's development. Pyribs is available at https://pyribs.org Bryon Tjanaka, Matthew C. Fontaine, David H. Lee, Yulun Zhang 0002, Nivedit Reddy Balam, Nathaniel Dennler, Sujay S. Garlanka, Nikitas Dimitri Klapsis, Stefanos Nikolaidis |
GECCO | 9 |
| 2023 | Transfer Learning of Human Preferences for Proactive Robot Assistance in Assembly TasksabstractWe focus on enabling robots to proactively assist humans in assembly tasks by adapting to their preferred sequence of actions. Much work on robot adaptation requires human demonstrations of the task. However, human demonstrations of real-world assemblies can be tedious and time-consuming. Thus, we propose learning human preferences from demonstrations in a shorter, canonical task to predict user actions in the actual assembly task. The proposed system uses the preference model learned from the canonical task as a prior and updates the model through interaction when predictions are inaccurate. We evaluate the proposed system in simulated assembly tasks and in a real-world human-robot assembly study and we show that both transferring the preference model from the canonical task, as well as updating the model online, contribute to improved accuracy in human action prediction. This enables the robot to proactively assist users, significantly reduce their idle time, and improve their experience working with the robot, compared to a reactive robot. Heramb Nemlekar, Neel Dhanaraj, Angelos Guan, Satyandra K. Gupta, Stefanos Nikolaidis |
HRI | 5 |
| 2023 | Contingency-Aware Task Assignment and Scheduling for Human-Robot TeamsabstractWe consider the problem of task assignment and scheduling for human-robot teams to enable the efficient completion of complex problems, such as satellite assembly. In high-mix, low volume settings, we must enable the human-robot team to handle uncertainty due to changing task requirements, potential failures, and delays to maintain task completion efficiency. We make two contributions: (1) we account for the complex interaction of uncertainty that stems from the tasks and the agents using a multi-agent concurrent MDP framework, and (2) we use Mixed Integer Linear Programs and contingency sampling to approximate action values for task assignment. Our results show that our online algorithm is computationally efficient while making optimal task assignments compared to a value iteration baseline. We evaluate our method on a 24-task representative assembly and a real-world 60-task satellite assembly, and we show that we can find an assignment that results in a near-optimal makespan. Neel Dhanaraj, Santosh V. Narayan, Stefanos Nikolaidis, Satyandra K. Gupta |
ICRA | 3 |
| 2023 | Inverse Reinforcement Learning Framework for Transferring Task Sequencing Policies from Humans to Robots in Manufacturing ApplicationsabstractIn this work, we present an inverse reinforcement learning approach for solving the problem of task sequencing for robots in complex manufacturing processes. Our proposed framework is adaptable to variations in process and can perform sequencing for entirely new parts. We prescribe an approach to capture feature interactions in a demonstration dataset based on a metric that computes feature interaction coverage. We then actively learn the expert's policy by keeping the expert in the loop. Our training and testing results reveal that our model can successfully learn the expert's policy. We demonstrate the performance of our method on a real-world manufacturing application where we transfer the policy for task sequencing to a manipulator. Our experiments show that the robot can perform these tasks to produce human-competitive performance. Code and video can be found at: https://sites.google.com/usc.edu/irlfortasksequencing Omey M. Manyar, Zachary McNulty, Stefanos Nikolaidis, Satyandra K. Gupta |
ICRA | 3 |
| 2023 | Multi-Robot Coordination and Layout Design for Automated WarehousingabstractWith the rapid progress in Multi-Agent Path Finding (MAPF), researchers have studied how MAPF algorithms can be deployed to coordinate hundreds of robots in large automated warehouses. While most works try to improve the throughput of such warehouses by developing better MAPF algorithms, we focus on improving the throughput by optimizing the warehouse layout. We show that, even with state-of-the-art MAPF algorithms, commonly used human-designed layouts can lead to congestion for warehouses with large numbers of robots and thus have limited scalability. We extend existing automatic scenario generation methods to optimize warehouse layouts. Results show that our optimized warehouse layouts (1) reduce traffic congestion and thus improve throughput, (2) improve the scalability of the automated warehouses by doubling the number of robots in some cases, and (3) are capable of generating layouts with user-specified diversity measures. Yulun Zhang 0002, Matthew C. Fontaine, Varun Bhatt, Stefanos Nikolaidis, Jiaoyang Li 0001 |
IJCAI | 4 |
| 2023 | Arbitrarily Scalable Environment Generators via Neural Cellular AutomataabstractWe study the problem of generating arbitrarily large environments to improve the throughput of multi-robot systems. Prior work proposes Quality Diversity (QD) algorithms as an effective method for optimizing the environments of automated warehouses. However, these approaches optimize only relatively small environments, falling short when it comes to replicating real-world warehouse sizes. The challenge arises from the exponential increase in the search space as the environment size increases. Additionally, the previous methods have only been tested with up to 350 robots in simulations, while practical warehouses could host thousands of robots. In this paper, instead of optimizing environments, we propose to optimize Neural Cellular Automata (NCA) environment generators via QD algorithms. We train a collection of NCA generators with QD algorithms in small environments and then generate arbitrarily large environments from the generators at test time. We show that NCA environment generators maintain consistent, regularized patterns regardless of environment size, significantly enhancing the scalability of multi-robot systems in two different domains with up to 2,350 robots. Additionally, we demonstrate that our method scales a single-agent reinforcement learning policy to arbitrarily large environments with similar patterns. We include the source code at https://github.com/lunjohnzhang/warehouse_env_gen_nca_public. Yulun Zhang 0002, Matthew C. Fontaine, Varun Bhatt, Stefanos Nikolaidis, Jiaoyang Li 0001 |
NeurIPS | 4 |
| 2023 | Design Metaphors for Understanding User Expectations of Socially Interactive Robot EmbodimentsabstractThe physical design of a robot suggests expectations of that robot’s functionality for human users and collaborators. When those expectations align with the robot’s true capabilities, users are more likely to adopt the technologies for their intended use. However, the relationship between expectations and socially interactive robot design is not well understood. This article applies the concept of design metaphors to robot design and contributes the Metaphors for Understanding Functional and Social Anticipated Affordances dataset of 165 extant robots and the expectations users place on them. We used Mechanical Turk to crowd-source user expectation over three user studies. The first study ( N = 382) associated crowd-sourced design metaphors to different robot embodiments. The second study ( N = 803) assessed initial social expectations of robot embodiments. The final study ( N = 805) addressed the degree of abstraction of the design metaphors and the functional expectations projected on robot embodiments. We performed analyses to gain insights into how design metaphors can be used to understand social and functional expectations of robots and how these data can be visualized to be useful for study designers and robot designers. Together, these results can serve to guide robot designers toward aligning user expectations with true robot capabilities, facilitating positive human–robot interaction. Nathaniel Dennler, Changxiao Ruan, Jessica Hadiwijoyo, Brenna Chen, Stefanos Nikolaidis, Maja J. Mataric |
ACM Trans. Hum. Robot Interact. | 5 |
| 2022 | Illuminating diverse neural cellular automata for level generationabstractWe present a method of generating diverse collections of neural cellular automata (NCA) to design video game levels. While NCAs have so far only been trained via supervised learning, we present a quality diversity (QD) approach to generating a collection of NCA level generators. By framing the problem as a QD problem, our approach can train diverse level generators, whose output levels vary based on aesthetic or functional criteria. To efficiently generate NCAs, we train generators via Covariance Matrix Adaptation MAP-Elites (CMA-ME), a quality diversity algorithm which specializes in continuous search spaces. We apply our new method to generate level generators for several 2D tile-based games: a maze game, Sokoban, and Zelda. Our results show that CMA-ME can generate small NCAs that are diverse yet capable, often satisfying complex solvability criteria for deterministic agents. We compare against a Compositional Pattern-Producing Network (CPPN) baseline trained to produce diverse collections of generators and show that the NCA representation yields a better exploration of level-space. Sam Earle, Justin Snider, Matthew C. Fontaine, Stefanos Nikolaidis, Julian Togelius |
GECCO | 4 |
| 2022 | Approximating gradients for differentiable quality diversity in reinforcement learningabstractConsider the problem of training robustly capable agents. One approach is to generate a diverse collection of agent polices. Training can then be viewed as a quality diversity (QD) optimization problem, where we search for a collection of performant policies that are diverse with respect to quantified behavior. Recent work shows that differentiable quality diversity (DQD) algorithms greatly accelerate QD optimization when exact gradients are available. However, agent policies typically assume that the environment is not differentiable. To apply DQD algorithms to training agent policies, we must approximate gradients for performance and behavior. We propose two variants of the current state-of-the-art DQD algorithm that compute gradients via approximation methods common in reinforcement learning (RL). We evaluate our approach on four simulated locomotion tasks. One variant achieves results comparable to the current state-of-the-art in combining QD and RL, while the other performs comparably in two locomotion tasks. These results provide insight into the limitations of current DQD algorithms in domains where gradients must be approximated. Source code is available at https://github.com/icaros-usc/dqd-rl Bryon Tjanaka, Matthew C. Fontaine, Julian Togelius, Stefanos Nikolaidis |
GECCO | 4 |
| 2022 | Deep surrogate assisted MAP-elites for automated hearthstone deckbuildingabstractWe study the problem of efficiently generating high-quality and diverse content in games. Previous work on automated deckbuilding in Hearthstone shows that the quality diversity algorithm MAP-Elites can generate a collection of high-performing decks with diverse strategic gameplay. However, MAP-Elites requires a large number of expensive evaluations to discover a diverse collection of decks. We propose assisting MAP-Elites with a deep surrogate model trained online to predict game outcomes with respect to candidate decks. MAP-Elites discovers a diverse dataset to improve the surrogate model accuracy while the surrogate model helps guide MAP-Elites towards promising new content. In a Hearthstone deck-building case study, we show that our approach improves the sample efficiency of MAP-Elites and outperforms a model trained offline with random decks, as well as a linear surrogate model baseline, setting a new state-of-the-art for quality diversity approaches in automated Hearthstone deckbuilding. We include the source code for all the experiments at: https://github.com/icaros-usc/EvoStone2. Yulun Zhang 0002, Matthew C. Fontaine, Amy K. Hoover, Stefanos Nikolaidis |
GECCO | 4 |
| 2022 | Machine Learning in Human-Robot Collaboration: Bridging the GapabstractThis workshop aims to bring together researchers to explore and identify ways in which human-robot collaboration can reap the benefits of modern machine learning. The intended outcome is a roadmap that identifies key milestones that will lead us towards fluent effective human-robot teaming. In addition to focus groups and creative brainstorming exercises, this workshop will comprise invited talks, contributed paper talks, a poster session, and a debate. The papers, talks, posters, and roadmap will be made publicly available on our website: https://sites.google.com/view/mlhrc-hri-2022/home. Cynthia Matuszek, Harold Soh, Matthew C. Gombolay, Nakul Gopalan, Reid G. Simmons, Stefanos Nikolaidis |
HRI | 6 |
| 2022 | Poster Abstract: Learning from Demonstrations with Temporal LogicsabstractLearning-from-demonstrations (LfD) is a popular paradigm to obtain effective robot control policies for complex tasks via reinforcement learning without the need to explicitly design reward functions. However, it is susceptible to imperfections in demonstrations and also raises concerns of safety and interpretability in the learned control policies. To address these issues, we propose to use Signal Temporal Logic (STL) to express high-level robotic tasks and use its quantitative semantics to evaluate and rank the quality of demonstrations. Temporal logic-based specifications allow us to create non-Markovian rewards, and are also capable of defining interesting causal dependencies between tasks such as sequential task specifications. We present our completed work that proposed LfD-STL framework that learns from even suboptimal/imperfect demonstrations and STL specifications to infer rewards for reinforcement learning tasks. We have validated our approach through various experimental setups to show how our method outperforms prior LfD methods. We then discuss future directions for tackling the problem of explainability and interpretability in such learning-based systems. Aniruddh Gopinath Puranic, Jyotirmoy V. Deshmukh, Stefanos Nikolaidis |
HSCC | 3 |
| 2022 | Deep Surrogate Assisted Generation of EnvironmentsabstractRecent progress in reinforcement learning (RL) has started producing generally capable agents that can solve a distribution of complex environments. These agents are typically tested on fixed, human-authored environments. On the other hand, quality diversity (QD) optimization has been proven to be an effective component of environment generation algorithms, which can generate collections of high-quality environments that are diverse in the resulting agent behaviors. However, these algorithms require potentially expensive simulations of agents on newly generated environments. We propose Deep Surrogate Assisted Generation of Environments (DSAGE), a sample-efficient QD environment generation algorithm that maintains a deep surrogate model for predicting agent behaviors in new environments. Results in two benchmark domains show that DSAGE significantly outperforms existing QD environment generation algorithms in discovering collections of environments that elicit diverse behaviors of a state-of-the-art RL agent and a planning agent. Our source code and videos are available at https://dsagepaper.github.io/. Varun Bhatt, Bryon Tjanaka, Matthew C. Fontaine, Stefanos Nikolaidis |
NeurIPS | 4 |
| 2022 | Human-Guided Goal Assignment to Effectively Manage Workload for a Smart Robotic AssistantabstractManaging robot workloads in human robot teams is critical for efficient team operation. If robots are overloaded with work, then they will miss deadlines and force humans to take on extra work. This paper presents a framework for a robot to assess its own workload based on an initial goal assignment. The robot does this by generating task and motion plans and computing the probability of missing deadlines due to the possibility of delays in task execution. A branch and bound based search is used to generate task and motion plans by minimizing task execution effort. The robot presents a diverse set of task and motion plans to the humans to offer multiple different options. Humans can either approve a plan or provide guidance to reduce the workload by either relaxing deadlines or removing goal(s) assigned to the robots. Neel Dhanaraj, Rishi K. Malhan, Heramb Nemlekar, Stefanos Nikolaidis, Satyandra K. Gupta |
RO-MAN | 4 |
| 2022 | Towards Transferring Human Preferences from Canonical to Actual Assembly TasksabstractTo assist human users according to their individual preference in assembly tasks, robots typically require user demonstrations in the given task. However, providing demonstrations in actual assembly tasks can be tedious and time-consuming. Our thesis is that we can learn the preference of users in actual assembly tasks from their demonstrations in a representative canonical task. Inspired by prior work in economy of human movement, we propose to represent user preferences as a linear reward function over abstract task-agnostic features, such as movement and physical and mental effort required by the user. For each user, we learn the weights of the reward function from their demonstrations in a canonical task and use the learned weights to anticipate their actions in the actual assembly task; without any user demonstrations in the actual task. We evaluate our proposed method in a model-airplane assembly study and show that preferences can be effectively transferred from canonical to actual assembly tasks, enabling robots to anticipate user actions. Heramb Nemlekar, Runyu Guan 0001, Guanyang Luo, Satyandra K. Gupta, Stefanos Nikolaidis |
RO-MAN | 5 |
| 2022 | Evaluating Human-Robot Interaction Algorithms in Shared Autonomy via Quality Diversity Scenario GenerationabstractThe growth of scale and complexity of interactions between humans and robots highlights the need for new computational methods to automatically evaluate novel algorithms and applications. Exploring diverse scenarios of humans and robots interacting in simulation can improve understanding of the robotic system and avoid potentially costly failures in real-world settings. We formulate this problem as a quality diversity (QD) problem, of which the goal is to discover diverse failure scenarios by simultaneously exploring both environments and human actions. We focus on the shared autonomy domain, in which the robot attempts to infer the goal of a human operator, and adopt the QD algorithms CMA-ME and MAP-Elites to generate scenarios for two published algorithms in this domain: shared autonomy via hindsight optimization and linear policy blending. Some of the generated scenarios confirm previous theoretical findings, while others are surprising and bring about a new understanding of state-of-the-art implementations. Our experiments show that the QD algorithms CMA-ME and MAP-Elites outperform Monte-Carlo simulation and optimization-based methods in effectively searching the scenario space, highlighting their promise for automatic evaluation of algorithms in human–robot interaction. Matthew C. Fontaine, Stefanos Nikolaidis |
ACM Trans. Hum. Robot Interact. | 2 |
| 2021 | Illuminating Mario Scenes in the Latent Space of a Generative Adversarial NetworkabstractGenerative adversarial networks (GANs) are quickly becoming a ubiquitous approach to procedurally generating video game levels. While GAN generated levels are stylistically similar to human-authored examples, human designers often want to explore the generative design space of GANs to extract interesting levels. However, human designers find latent vectors opaque and would rather explore along dimensions the designer specifies, such as number of enemies or obstacles. We propose using state-of-the-art quality diversity algorithms designed to optimize continuous spaces, i.e. MAP-Elites with a directional variation operator and Covariance Matrix Adaptation MAP-Elites, to efficiently explore the latent space of a GAN to extract levels that vary across a set of specified gameplay measures. In the benchmark domain of Super Mario Bros, we demonstrate how designers may specify gameplay measures to our system and extract high-quality (playable) levels with a diverse range of level mechanics, while still maintaining stylistic similarity to human authored examples. An online user study shows how the different mechanics of the automatically generated levels affect subjective ratings of their perceived difficulty and appearance. Matthew C. Fontaine, Ahmed Khalifa 0001, Jignesh Modi, Julian Togelius, Amy K. Hoover, Stefanos Nikolaidis |
AAAI | 7 |
| 2021 | Two-Stage Clustering of Human Preferences for Action Prediction in Assembly TasksabstractTo effectively assist human workers in assembly tasks a robot must proactively offer support by inferring their preferences in sequencing the task actions. Previous work has focused on learning the dominant preferences of human workers for simple tasks largely based on their intended goal. However, people may have preferences at different resolutions: they may share the same high-level preference for the order of the sub-tasks but differ in the sequence of individual actions. We propose a two-stage approach for learning and inferring the preferences of human operators based on the sequence of sub-tasks and actions. We conduct an IKEA assembly study and demonstrate how our approach is able to learn the dominant preferences in a complex task. We show that our approach improves the prediction of human actions through cross-validation. Lastly we show that our two-stage approach improves the efficiency of task execution in an online experiment and demonstrate its applicability in a real-world robot-assisted IKEA assembly. Heramb Nemlekar, Jignesh Modi, Satyandra K. Gupta, Stefanos Nikolaidis |
ICRA | 4 |
| 2021 | Learning Collaborative Pushing and Grasping Policies in Dense ClutterabstractRobots must reason about pushing and grasping in order to engage in flexible manipulation in cluttered environments. Earlier works on learning pushing and grasping only consider each operation in isolation or are limited to top-down grasping and bin-picking. We train a robot to learn joint planar pushing and 6-degree-of-freedom (6-DoF) grasping policies by self-supervision. Two separate deep neural networks are trained to map from 3D visual observations to actions with a Q-learning framework. With collaborative pushes and expanded grasping action space, our system can deal with cluttered scenes with a wide variety of objects (e.g. grasping a plate from the side after pushing away surrounding obstacles). We compare our system to the state-of-the-art baseline model VPG [1] in simulation and outperform it with 10% higher action efficiency and 20% higher grasp success rate. We then demonstrate our system on a KUKA LBR iiwa arm with a Robotiq 3-finger gripper. Bingjie Tang, Matthew Corsaro, George Dimitri Konidaris, Stefanos Nikolaidis, Stefanie Tellex |
ICRA | 4 |
| 2021 | Design and Evaluation of a Hair Combing System Using a General-Purpose Robotic ArmabstractThis work introduces an approach for automatic hair combing by a lightweight robot. For people living with limited mobility, dexterity, or chronic fatigue, combing hair is often a difficult task that negatively impacts personal routines. We propose a modular system for enabling general robot manipulators to assist with a hair-combing task. The system consists of three main components. The first component is the segmentation module, which segments the location of hair in space. The second component is the path planning module that proposes automatically-generated paths through hair based on user input. The final component creates a trajectory for the robot to execute. We quantitatively evaluate the effectiveness of the paths planned by the system with 48 users and qualitatively evaluate the system with 30 users watching videos of the robot performing a hair-combing task in the physical world. The system is shown to effectively comb different hairstyles. Nathaniel Dennler, Eura Nofshin, Maja J. Mataric, Stefanos Nikolaidis |
IROS | 4 |
| 2021 | Robotic Lime Picking by Considering Leaves as Permeable ObstaclesabstractThe problem of robotic lime picking is challenging; lime plants have dense foliage which makes it difficult for a robotic arm to grasp a lime without coming in contact with leaves. Existing approaches either do not consider leaves, or treat them as obstacles and completely avoid them, often resulting in undesirable or infeasible plans. We focus on reaching a lime in the presence of dense foliage by considering the leaves of a plant as permeable obstacles with a collision cost. We then adapt the rapidly exploring random tree star (RRT*) algorithm for the problem of fruit harvesting by incorporating the cost of collision with leaves into the path cost. To reduce the time required for finding low-cost paths to goal, we bias the growth of the tree using an artificial potential field (APF). We compare our proposed method with prior work in a 2-D environment and a 6-DOF robot simulation. Our experiments and a real-world demonstration on a robotic lime picking task demonstrate the applicability of our approach. Heramb Nemlekar, Ziang Liu 0002, Suraj Kothawade, Sherdil Niyaz, Barath Raghavan, Stefanos Nikolaidis |
IROS | 6 |
| 2021 | Differentiable Quality DiversityabstractQuality diversity (QD) is a growing branch of stochastic optimization research that studies the problem of generating an archive of solutions that maximize a given objective function but are also diverse with respect to a set of specified measure functions. However, even when these functions are differentiable, QD algorithms treat them as "black boxes", ignoring gradient information. We present the differentiable quality diversity (DQD) problem, a special case of QD, where both the objective and measure functions are first order differentiable. We then present MAP-Elites via a Gradient Arborescence (MEGA), a DQD algorithm that leverages gradient information to efficiently explore the joint range of the objective and measure functions. Results in two QD benchmark domains and in searching the latent space of a StyleGAN show that MEGA significantly outperforms state-of-the-art QD algorithms, highlighting DQD's promise for efficient quality diversity optimization when gradient information is available. Source code is available at https://github.com/icaros-usc/dqd. Matthew C. Fontaine, Stefanos Nikolaidis |
NeurIPS | 2 |
| 2021 | Personalizing User Engagement Dynamics in a Non-Verbal Communication Game for Cerebral PalsyabstractChildren and adults with cerebral palsy (CP) can have involuntary upper limb movements as a consequence of the symptoms that characterize their motor disability, leading to difficulties in communicating with caretakers and peers. We describe how a socially assistive robot may help individuals with CP to practice non-verbal communicative gestures using an active orthosis in a one-on-one number-guessing game. We performed a user study and data collection with participants with CP; we found that participants preferred an embodied robot over a screen-based agent, and we used the participant data to train personalized models of participant engagement dynamics that can be used to select personalized robot actions. Our work highlights the benefit of personalized models in the engagement of users with CP with a socially assistive robot and offers design insights for future work in this area. Nathaniel Dennler, Catherine Yunis, Jonathan Realmuto, Terence D. Sanger, Stefanos Nikolaidis, Maja J. Mataric |
RO-MAN | 5 |
| 2020 | Covariance matrix adaptation for the rapid illumination of behavior spaceabstractWe focus on the challenge of finding a diverse collection of quality solutions on complex continuous domains. While quality diversity (QD) algorithms like Novelty Search with Local Competition (NSLC) and MAP-Elites are designed to generate a diverse range of solutions, these algorithms require a large number of evaluations for exploration of continuous spaces. Meanwhile, variants of the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) are among the best-performing derivative-free optimizers in single-objective continuous domains. This paper proposes a new QD algorithm called Covariance Matrix Adaptation MAP-Elites (CMA-ME). Our new algorithm combines the self-adaptation techniques of CMA-ES with archiving and mapping techniques for maintaining diversity in QD. Results from experiments based on standard continuous optimization benchmarks show that CMA-ME finds better-quality solutions than MAP-Elites; similarly, results on the strategic game Hearthstone show that CMA-ME finds both a higher overall quality and broader diversity of strategies than both CMA-ES and MAP-Elites. Overall, CMA-ME more than doubles the performance of MAP-Elites using standard QD performance metrics. These results suggest that QD algorithms augmented by operators from state-of-the-art optimization algorithms can yield high-performing methods for simultaneously exploring and optimizing continuous search spaces, with significant applications to design, testing, and reinforcement learning among other domains. Matthew C. Fontaine, Julian Togelius, Stefanos Nikolaidis, Amy K. Hoover |
GECCO | 3 |
| 2020 | Multi-Armed Bandits with Fairness Constraints for Distributing Resources to Human TeammatesabstractHow should a robot that collaborates with multiple people decide upon the distribution of resources (e.g. social attention, or parts needed for an assembly)? People are uniquely attuned to how resources are distributed. A decision to distribute more resources to one team member than another might be perceived as unfair with potentially detrimental effects for trust. We introduce a multi-armed bandit algorithm with fairness constraints, where a robot distributes resources to human teammates of different skill levels. In this problem, the robot does not know the skill level of each human teammate, but learns it by observing their performance over time. We define fairness as a constraint on the minimum rate that each human teammate is selected throughout the task. We provide theoretical guarantees on performance and perform a large-scale user study, where we adjust the level of fairness in our algorithm. Results show that fairness in resource distribution has a significant effect on users' trust in the system. Houston Claure, Yifang Chen 0004, Jignesh Modi, Malte F. Jung, Stefanos Nikolaidis |
HRI | 5 |
| 2020 | Robot Learning in Mixed Adversarial and Collaborative SettingsabstractPrevious work has shown that interacting with a human adversary can significantly improve the efficiency of the learning process in robot grasping. However, people are not consistent in applying adversarial forces; instead they may alternate between acting antagonistically with the robot or helping the robot achieve its tasks. We propose a physical framework for robot learning in a mixed adversarial/collaborative setting, where a second agent may act as a collaborator or as an antagonist, unbeknownst to the robot. The framework leverages prior estimates of the reward function to infer whether the actions of the second agent are collaborative or adversarial. Integrating the inference in an adversarial learning algorithm can significantly improve the robustness of learned grasps in a manipulation task. Seung Hee Yoon, Stefanos Nikolaidis |
IROS | 2 |
| 2020 | Fair Contextual Multi-Armed Bandits: Theory and ExperimentsabstractWhen an AI system interacts with multiple users, it frequently needs to make allocation decisions. For instance, a virtual agent decides whom to pay attention to in a group, or a factory robot selects a worker to deliver a part.Demonstrating fairness in decision making is essential for such systems to be broadly accepted. We introduce a Multi-Armed Bandit algorithm with fairness constraints, where fairness is defined as a minimum rate at which a task or a resource is assigned to a user. The proposed algorithm uses contextual information about the users and the task and makes no assumptions on how the losses capturing the performance of different users are generated. We provide theoretical guarantees of performance and empirical results from simulation and an online user study. The results highlight the benefit of accounting for contexts in fair decision making, especially when users perform better at some contexts and worse at others. Yifang Chen 0004, Alex Cuellar, Jignesh Modi, Heramb Nemlekar, Stefanos Nikolaidis |
UAI | 6 |
| 2020 | Trust-Aware Decision Making for Human-Robot Collaboration: Model Learning and PlanningabstractTrust in autonomy is essential for effective human-robot collaboration and user adoption of autonomous systems such as robot assistants. This article introduces a computational model that integrates trust into robot decision making. Specifically, we learn from data a partially observable Markov decision process (POMDP) with human trust as a latent variable. The trust-POMDP model provides a principled approach for the robot to (i) infer the trust of a human teammate through interaction, (ii) reason about the effect of its own actions on human trust, and (iii) choose actions that maximize team performance over the long term. We validated the model through human subject experiments on a table clearing task in simulation (201 participants) and with a real robot (20 participants). In our studies, the robot builds human trust by manipulating low-risk objects first. Interestingly, the robot sometimes fails intentionally to modulate human trust and achieve the best team performance. These results show that the trust-POMDP calibrates trust to improve human-robot team performance over the long term. Further, they highlight that maximizing trust alone does not always lead to the best performance. Min Chen 0018, Stefanos Nikolaidis, Harold Soh, David Hsu, Siddhartha S. Srinivasa |
ACM Trans. Hum. Robot Interact. | 2 |
| 2019 | Robot Object Referencing through Legible Situated ProjectionsabstractThe ability to reference objects in the environment is a key communication skill that robots need for complex, task-oriented human-robot collaborations. In this paper we explore the use of projections, which are a powerful communication channel for robot-to-human information transfer as they allow for situated, instantaneous, and parallelized visual referencing. We focus on the question of what makes a good projection for referencing a target object. To that end, we mathematically formulatelegibility of projections intended to reference an object, and propose alternative arrow-object match functions for optimally computing the placement of an arrow to indicate a target object in a cluttered scene. We implement our approach on a PR2 robot with a head-mounted projector. Through an online (48 participants) and an in-person (12 participants) user study we validate the effectiveness of our approach, identify the types of scenes where projections may fail, and characterize the differences between alternative match functions. Thomas Weng, Leah Perlmutter, Stefanos Nikolaidis, Siddhartha S. Srinivasa, Maya Cakmak |
ICRA | 3 |
| 2019 | Robot Learning via Human Adversarial GamesabstractMuch work in robotics has focused on “humanin-the-loop” learning techniques that improve the efficiency of the learning process. However, these algorithms have made the strong assumption of a cooperating human supervisor that assists the robot. In reality, human observers tend to also act in an adversarial manner towards deployed robotic systems. We show that this can in fact improve the robustness of the learned models by proposing a physical framework that leverages perturbations applied by a human adversary, guiding the robot towards more robust models. In a manipulation task, we show that grasping success improves significantly when the robot trains with a human adversary as compared to training in a self-supervised manner. Jiali Duan, Lerrel Pinto, C.-C. Jay Kuo, Stefanos Nikolaidis |
IROS | 5 |
| 2019 | Learning Collaborative Action Plans from YouTube Videos
Po-Jen Lai, Sayan Paul, Suraj Kothawade, Stefanos Nikolaidis |
ISRR | 5 |
| 2019 | Surprise! Predicting Infant Visual Attention in a Socially Assistive Robot Contingent Learning ParadigmabstractEarly intervention to address developmental disability in infants has the potential to promote improved outcomes in neurodevelopmental structure and function [1]. Researchers are starting to explore Socially Assistive Robotics (SAR) as a tool for delivering early interventions that are synergistic with and enhance human-administered therapy. For SAR to be effective, the robot must be able to consistently attract the attention of the infant in order to engage the infant in a desired activity. This work presents the analysis of eye gaze tracking data from five 6-8 month old infants interacting with a Nao robot that kicked its leg as a contingent reward for infant leg movement. We evaluate a Bayesian model of low-level surprise on video data from the infants' head-mounted camera and on the timing of robot behaviors as a predictor of infant visual attention. The results demonstrate that over 67% of infant gaze locations were in areas the model evaluated to be more surprising than average. We also present an initial exploration using surprise to predict the extent to which the robot attracts infant visual attention during specific intervals in the study. This work is the first to validate the surprise model on infants; our results indicate the potential for using surprise to inform robot behaviors that attract infant attention during SAR interactions. Lauren Klein, Laurent Itti, Beth A. Smith, Marcelo R. Rosales, Stefanos Nikolaidis, Maja J. Mataric |
RO-MAN | 5 |
| 2018 | Planning with Trust for Human-Robot CollaborationabstractTrust is essential for human-robot collaboration and user adoption of autonomous systems, such as robot assistants. This paper introduces a computational model which integrates trust into robot decision-making. Specifically, we learn from data a partially observable Markov decision process (POMDP) with human trust as a latent variable. The trust-POMDP model provides a principled approach for the robot to (i) infer the trust of a human teammate through interaction, (ii) reason about the effect of its own actions on human behaviors, and (iii) choose actions that maximize team performance over the long term. We validated the model through human subject experiments on a table-clearing task in simulation (201 participants) and with a real robot (20 participants). The results show that the trust-POMDP improves human-robot team performance in this task. They further suggest that maximizing trust in itself may not improve team performance. Min Chen 0018, Stefanos Nikolaidis, Harold Soh, David Hsu, Siddhartha S. Srinivasa |
HRI | 2 |
| 2018 | Planning with Verbal Communication for Human-Robot CollaborationabstractHuman collaborators coordinate effectively their actions through both verbal and non-verbal communication. We believe that the the same should hold for human-robot teams. We propose a formalism that enables a robot to decide optimally between taking a physical action toward task completion and issuing an utterance to the human teammate. We focus on two types of utterances: verbal commands, where the robot asks the human to take a physical action, and state-conveying actions, where the robot informs the human about its internal state, which captures the information that the robot uses in its decision making. Human subject experiments show that enabling the robot to issue verbal commands is the most effective form of communicating objectives, while retaining user trust in the robot. Communicating information about the robot’s state should be done judiciously, since many participants questioned the truthfulness of the robot statements when the robot did not provide sufficient explanation about its actions. Stefanos Nikolaidis, Minae Kwon, Jodi Forlizzi, Siddhartha S. Srinivasa |
ACM Trans. Hum. Robot Interact. | 1 |
| 2017 | Game-Theoretic Modeling of Human Adaptation in Human-Robot CollaborationabstractIn human-robot teams, humans often start with an inaccurate model of the robot capabilities. As they interact with the robot, they infer the robot's capabilities and partially adapt to the robot, i.e., they might change their actions based on the observed outcomes and the robot's actions, without replicating the robot's policy. We present a game-theoretic model of human partial adaptation to the robot, where the human responds to the robot's actions by maximizing a reward function that changes stochastically over time, capturing the evolution of their expectations of the robot's capabilities. The robot can then use this model to decide optimally between taking actions that reveal its capabilities to the human and taking the best action given the information that the human currently has. We prove that under certain observability assumptions, the optimal policy can be computed efficiently. We demonstrate through a human subject experiment that the proposed model significantly improves human-robot team performance, compared to policies that assume complete adaptation of the human to the robot. Stefanos Nikolaidis, Swaprava Nath, Ariel D. Procaccia, Siddhartha S. Srinivasa |
HRI | 1 |
| 2017 | Human-Robot Mutual Adaptation in Shared AutonomyabstractShared autonomy integrates user input with robot autonomy in order to control a robot and help the user to complete a task. Our work aims to improve the performance of such a human-robot team: the robot tries to guide the human towards an effective strategy, sometimes against the human's own preference, while still retaining his trust. We achieve this through a principled human-robot mutual adaptation formalism. We integrate a bounded-memory adaptation model of the human into a partially observable stochastic decision model, which enables the robot to adapt to an adaptable human. When the human is adaptable, the robot guides the human towards a good strategy, maybe unknown to the human in advance. When the human is stubborn and not adaptable, the robot complies with the human's preference in order to retain their trust. In the shared autonomy setting, unlike many other common human-robot collaboration settings, only the robot actions can change the physical state of the world, and the human and robot goals are not fully observable. We address these challenges and show in a human subject experiment that the proposed mutual adaptation formalism improves human-robot team performance, while retaining a high level of user trust in the robot, compared to the common approach of having the robot strictly following participants' preference. Stefanos Nikolaidis, Yu Xiang Zhu, David Hsu, Siddhartha S. Srinivasa |
HRI | 1 |
| 2016 | Viewpoint-Based Legibility OptimizationabstractMuch robotics research has focused on intent-expressive (legible) motion. However, algorithms that can autonomously generate legible motion have implicitly made the strong assumption of an omniscient observer, with access to the robot's configuration as it changes across time. In reality, human observers have a particular viewpoint, which biases the way they perceive the motion. In this work, we free robots from this assumption and introduce the notion of an observer with a specific point of view into legibility optimization. In doing so, we account for two factors: (1) depth uncertainty induced by a particular viewpoint, and (2) occlusions along the motion, during which (part of) the robot is hidden behind some object. We propose viewpoint and occlusion models that enable autonomous generation of viewpoint-based legible motions, and show through large-scale user studies that the produced motions are significantly more legible compared to those generated assuming an omniscient observer. Stefanos Nikolaidis, Anca D. Dragan, Siddhartha S. Srinivasa |
HRI | 1 |
| 2016 | Formalizing Human-Robot Mutual Adaptation: A Bounded Memory ModelabstractMutual adaptation is critical for effective team collaboration. This paper presents a formalism for human-robot mutual adaptation in collaborative tasks. We propose the bounded-memory adaptation model (BAM), which captures human adaptive behaviors based on a bounded memory assumption. We integrate BAM into a partially observable stochastic model, which enables robot adaptation to the human. When the human is adaptive, the robot will guide the human towards a new, optimal collaborative strategy unknown to the human in advance. When the human is not willing to change their strategy, the robot adapts to the human in order to retain human trust. Human subject experiments indicate that the proposed formalism can significantly improve the effectiveness of human-robot teams, while human subject ratings on the robot performance and trust are comparable to those achieved by cross training, a state-of-the-art human-robot team training practice. Stefanos Nikolaidis, Anton Kuznetsov, David Hsu, Siddhartha S. Srinivasa |
HRI | 1 |
| 2015 | Efficient Model Learning from Joint-Action Demonstrations for Human-Robot Collaborative TasksabstractWe present a framework for automatically learning human user models from joint-action demonstrations that enables a robot to compute a robust policy for a collaborative task with a human. First, the demonstrated action sequences are clustered into different human types using an unsupervised learning algorithm. A reward function is then learned for each type through the employment of an inverse reinforcement learning algorithm. The learned model is then incorporated into a mixed-observability Markov decision process (MOMDP) formulation, wherein the human type is a partially observable variable. With this framework, we can infer online the human type of a new user that was not included in the training set, and can compute a policy for the robot that will be aligned to the preference of this user. In a human subject experiment (n=30), participants agreed more strongly that the robot anticipated their actions when working with a robot incorporating the proposed framework (p<0.01), compared to manually annotating robot actions. In trials where participants faced difficulty annotating the robot actions to complete the task, the proposed framework significantly improved team efficiency (p<0.01). The robot incorporating the framework was also found to be more responsive to human actions compared to policies computed using a hand-coded reward function by a domain expert (p<0.01). These results indicate that learning human user models from joint-action demonstrations and encoding them in a MOMDP formalism can support effective teaming in human-robot collaborative tasks. Stefanos Nikolaidis, Ramya Ramakrishnan, Keren Gu, Julie A. Shah |
HRI | 1 |
| 2013 | Human-robot cross-training: computational formulation, modeling and evaluation of a human team training strategy
Stefanos Nikolaidis, Julie A. Shah |
HRI | 1 |
| 2009 | Optimal arrangement of ceiling cameras for home service robots using genetic algorithmsabstractIn the near future robots will be used in home environments to provide assistance for the elderly and challenged people. As home environments are complicated, external sensors like ceiling cameras need to be placed on the environment to provide the robot with information about its position. The pose of cameras influences the area covered by the cameras, as well as the error of the robot localization. We examine the problem of the finding the arrangement of ceiling cameras at home environments that maximizes the area covered and minimizes the localization error. Genetic algorithms are proposed for the single and multi-objective optimization problem. Simulation results indicate that we can obtain the optimal arrangement of cameras that satisfies the given objectives and the required constraints. Stefanos Nikolaidis, Tamio Arai |
RO-MAN | 1 |