Yiannis Demiris

dblp:47/1473 · DBLP profile ↗
← Back
138ranked-venue papers
4as first author
46since 2021 · last 2026
0000-0003-4917-3343ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 109 · 3 first-author · 30 since 2021Systems, architecture and hardware · 47 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 7 since 2021Human-computer interaction and ubiquitous computing · 24 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3
YearPublicationVenuePosition
2026 Continuous Real-time Adaptation Framework for Enhancing Trust and Technology Acceptance: An Assistive Feeding Study
abstract
Assistive feeding robots operate in close proximity to users often unfamiliar with robotic systems, making trust a crucial element of Human-Robot Interaction (HRI). While technical challenges in assistive feeding are actively researched, human factors such as trust remain underexplored. We propose the Continuous Real-time Adaptation Framework for Trust and Technology Acceptance (CRAFTT), which adapts the robot’s behaviour in real-time based on user reactions to enhance trust and technology acceptance, thereby giving the user greater control over the interaction. This adaptation is modelled as a multi-objective optimisation problem, minimising emotional and physical discomfort while maximising efficiency. In a within-participants study, 26 participants interacted with both adaptive (CRAFTT-equipped) and non-adaptive robot behaviours during a simulated meal. Dependent variables included user comfort cost and time efficiency, along with self-reported HRI metrics—trust, reliance intention, perceived safety, fluency, technology acceptance and cognitive workload—assessed using validated questionnaires. The robot’s impact was evaluated via Bayesian Data Analysis. Findings indicate that CRAFTT enhances trust, reliance, technology acceptance and fluency, with minimal efficiency loss, while maintaining perceived safety and cognitive workload. Future work will explore long-term interactions, expand the problem space (e.g., real food, multi-arm coordination) and apply CRAFTT to other assistive scenarios such as dressing and bathing.
Dimitra Tsakona, Yiannis Demiris
ACM Trans. Hum. Robot Interact.2
2026 Large (Vision) Language Models for Autonomous Vehicles: Current Trends and Future Directions
abstract
As autonomous vehicles (AVs) advance, the integration of Large (Vision) Language Models (LLMs and VLMs) has emerged as a promising approach to enhance AV capabilities in perception, planning, decision-making, and data generation. However, the practical challenges of incorporating LLMs and VLMs into AV systems, including computational efficiency, real-time processing, and ethical considerations, remain underexplored. This survey aims to provide a comprehensive review of the current research on LLM and VLM applications in AVs, focusing on the following key areas: modular integration, end-to-end integration, data generation, evaluation platforms, datasets, and benchmarks. We systematically analyse 77 recent papers published before Sep 2025, detailing their methodologies and models. Our findings highlight the potential of LLMs and VLMs to improve AV system performance while acknowledging limitations. This survey offers researchers and practitioners a panoptic view of the classification and progression of LLMs and VLMs in the AV sphere, while systematically distilling models to their core components. We envision this survey as a central reference for AV researchers navigating this rapidly evolving landscape to accelerate future research.
Hanlin Tian, Kethan Reddy, Mohammed A. Quddus 0001, Yiannis Demiris, Panagiotis Angeloudis
IEEE Trans. Intell. Transp. Syst.5
2026 Navigating Uncertainty: Diffusion-Based User Intention Estimation for Wheelchair Assistance
abstract
User intention estimation is essential in shared control systems for powered wheelchairs. It enables seamless navigation assistance that enhances safety, efficiency, and usability, while preserving user autonomy and reducing effort. This paper presents Diffusion-based Wheelchair User Intention Estimation (DIWIE), a novel multimodal learning framework that leverages a Denoising Diffusion Probabilistic Model (DDPM) to forecast multiple plausible future trajectories, addressing uncertainty in human behaviour. DIWIE conditions on diverse inputs, including obstacle information, user attention cues from eye gaze and head pose, semantic context, wheelchair kinematics, and joystick commands, operating without predefined maps or target destinations. Evaluated on a large new dataset of natural navigation by multiple drivers, DIWIE outperforms state-of-the-art methods, achieving lower displacement errors and collision rates, making it a valuable component for integration into shared control systems. This work also analyses the relevance of different data sources for intention estimation and aligns evaluation metrics with related fields to foster reproducibility.
Fernando E. Casado, Rodrigo Chacon, Yiannis Demiris
IEEE Trans. Robotics3
2025 Interface Matters: Comparing First and Third-Person Perspective Interfaces for Bi-Manual Robot Behavioural Cloning
abstract
Despite the growing interest in Behavioural Cloning for robots, few existing research has explicitly explored the impact of user interfaces on the effectiveness of expert demonstrations. We investigate the importance of user interface design in Behavioural Cloning, highlighting the critical role that interfaces play in conveying human demonstrations and robotics capabilities. This study compares the effectiveness of first and third-person perspective interfaces for robot shoe-lacing, a highly dexterous, bi-manual manipulation task that involves deformable objects and requires high precision. Our study highlights the importance of considering the impact of interface design on expert demonstration quality in Behavioural Cloning applications. By providing a first-person perspective, we observed significant differences in demonstration execution time and consistency compared to the third-person perspective. These findings suggest that the choice of interface can influence the quality of expert demonstrations, which in turn affects the performance of learning algorithms.
Haining Luo, Rodrigo Chacon, Fernando E. Casado, Nico Lingg, Yiannis Demiris
ICRA5
2025 Toward Shared Control for Mobile Bimanual Manipulation on a Robotic Wheelchair
abstract
Assistance through wheelchair-mounted manipulators has the potential to enhance the independence of individuals with disabilities. However, existing approaches primarily focus on single-arm systems or require extensive user input and demonstrations to infer intentions. In this study, we present a shared control framework for intuitive dual-arm operation through a standard 2D joystick, focusing on pick-and-place tasks. Our approach infers user intent in real-time, eliminating the need for specifying prior goals or beliefs. To address the challenges of controlling two 7-DoF (Kinova Gen3) manipulators, we propose two distinct control methods: one for pre-grasp positioning and one for grasp execution. The first method employs a shared control policy to optimize the pre-grasp positioning of the wheelchair base, ensuring ergonomic alignment for front-grasping tasks while incorporating mobile manipulation to reduce task completion time. The second method allows users to maintain high-level goal control and fine-tuning through task-specific arbitration. Experimental results demonstrate high grasp quality and task efficiency across pick-and-place scenarios, establishing the feasibility of shared control for bimanual manipulation and wheelchair navigation. We believe this is the first unified framework for mobile bimanual manipulation using standard wheelchair controls.
Rohan Gandhi, Fernando E. Casado, Yiannis Demiris
RO-MAN3
2025 Alignment Strategies for Language-Model-Driven Robots in Human-Robot Collaboration: Effectiveness and Impact on Trust
abstract
This study examines the alignment of Large Language Models (LLMs)-controlled assistive robots with human intentions. We conducted a comparative analysis of various alignment strategies for action selection using 13 different locally-run LLMs, demonstrating that: (1) alignment strategies significantly influence robot action choices, and (2) off-the-shelf LLMs exhibit varying default alignments. Therefore, selecting an LLM for robot control should consider not only its performance but also the alignment strategy employed. Additionally, we conducted a user study (N=24, 17h53 of data collection), indicating that alignment strategies significantly impact human trust. Participants reported higher trust levels when the robot prioritised following their instructions over ensuring their wellbeing. Furthermore, we observed a wide range of participant behaviours, from complete delegation to the robot after a few successful actions to annoyance with the robot as soon as it displayed initiative. Exploring the role of personality in this context, we found that reported personality traits using the BFI model did not significantly explain reported trust, while observed personality traits, such as the number of requests made to the robot, were more predictive. This suggests that the robot’s alignment strategy should be tailored to individual users, taking into account their specific needs and preferences.
Cedric Goubard, Yiannis Demiris
RO-MAN2
2025 Trust-Act: Integrating Trust in Imitation Learning
abstract
In human-robot interaction, trust traditionally serves as a performance metric rather than a control variable, limiting its potential to shape robot behavior. In this paper, we present Trust-Act, a framework that integrates trust labels into a conditional variational autoencoder (CVAE) policy for imitation learning. Our approach trains on demonstrations where operators explicitly executed movements with varying characteristics—fast, smooth, and confident for high-trust, moderate with occasional pauses for medium-trust, and deliberately hesitant, jerky, and error-prone for low-trust—enabling the robot to generate trajectories that reflect these distinct trust-specific behaviors. We address a significant challenge in CVAE architectures, posterior collapse, through a partial input masking technique that preserves a meaningful and diverse latent space. Furthermore, we develop a trust prediction model that acts as a reward function, enabling a best-of-n strategy to identify trajectories that maintain trust characteristics while maximizing reliability. Experiments show our trust-conditioned policies maintain distinct motion characteristics across trust levels while our best-of-n sampling approach consistently improves success rates in all trust conditions. Our results demonstrate the promise of trust conditioning as a pathway to more controllable and human-aligned policy generation.
Nico Lingg, Yiannis Demiris
RO-MAN2
2025 Keystate-Driven Long-Term Generation of Bimanual Object Manipulation Sequences
abstract
Learning to forecast or synthesize bimanual object manipulation sequences has broad applications in assistive robotics and extended reality. Previous methods have several limitations: (1) They can only forecast for short durations as the output deteriorates with longer predictions. (2) They minimize the MSE of fine motions, such as cutting or stirring, which are treated as noise and averaged out, leading to static outputs. (3) They model hand-object contact implicitly, resulting in unrealistic motion where the objects float in the air. We address long-term forecasting degradation by decomposing long sequences of bimanual actions into shorter subsequences defined by keystates, minimizing output quality deterioration. Segmenting sequences into meaningful keystates allows us to treat fine periodic motions as primitives without optimizing their MSE in raw trajectories. We construct a motion dictionary to store representative dynamics for each action category, queried at test time to generate fine motions. Lastly, we improve hand-object contact using a novel neural network that forecasts the pose for objects in motion, while encouraging hand-object contact through generative models for 3D hand grasps. We evaluate our approach on publicly available bimanual manipulation datasets, showing significant improvements over the state of the art.
Haziq Razali, Yiannis Demiris
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 An Integrated 3D Eye-Gaze Tracking Framework for Assessing Trust in Human-Robot Interaction
abstract
We introduce a comprehensive approach to examining the complexities of trust during Human–Robot Interactions (HRIs) through an innovative 3D eye-gaze tracking framework. Trust is a fundamental psychological factor in HRI studies, influencing how humans perceive and interact with robots. Although researchers have previously highlighted eye-tracking as a promising tool for capturing behavioural manifestations of trust continuously and non-intrusively, traditional approaches have been limited to 2D setups, leaving their applicability to real-world HRI largely unexplored. Thus, there still is limited evidence for the feasibility and validity of using eye-tracking to assess human–robot trust in more realistic settings. To this end, our framework employs Head-Mounted Displays with 3D eye-gaze and spatial tracking capabilities to gather continuous eye-gaze data alongside real-time user and robot positions. In addition to 3D eye-gaze tracking capabilities, we designed and incorporated a Bayesian model to evaluate experimental treatments’ effectiveness while identifying eye-gaze features correlating with participants’ subjective trust scores. The latter are measured using Likert-type instruments, widely used in HRI research. We applied our framework to a user study involving 25 participants performing an inspection task with a robot under two reliability conditions—high versus low. Our results revealed significant differences in subjective trust between conditions. Moreover, the results show that participants exposed to the low-reliability condition fixate for longer and have higher fixation and saccade amplitudes when compared to those in the high-reliability condition. Additionally, the group with low reliability had a greater rate of transitions between fixations. These findings are consistent with previous research on 2D settings. However, we observed differences in scan-path length and total fixation count compared to previous studies. Lastly, our results show that incorporating multiple eye-gaze feature categories simultaneously into our Bayesian model can lead to a more nuanced comprehension of the intricate connections between eye-gaze patterns and subjective trust in HRI. A supplementary video providing additional details is available online as supplementary material and can also be accessed at https://www.imperial.ac.uk/personal-robotics/videos/ .
Rodrigo Chacon, Fernando E. Casado, Yiannis Demiris
ACM Trans. Hum. Robot Interact.3
2025 Cognitive Modelling of Visual Attention Captures Trust Dynamics in Human-Robot Collaboration
abstract
Understanding how humans perceive and interact with robots is crucial for collaborative scenarios. Trust, a pivotal factor in such interactions, is inherently volatile and subjective, posing significant challenges for robots. However, trust has also been shown to influence specific human bio-signals and behaviours, suggesting that it could be inferred from those indicators. One such indicator is visual attention, the cognitive process of focusing on distinct environmental elements, often manifested through eye gaze. Despite recent research connecting eye gaze and trust in Human–Robot Collaboration scenarios, this relationship remains largely unexplored. This article presents a novel signal, the Attention Arbitration Ratio (AAR), which is shown to be a promising real-time predictor of subjective and objective trust measures. We obtain this signal using a visual attention modelling framework that explicitly emulates the Bottom-Up and Top-Down processes, two key cognitive components. We demonstrate the connection between the AAR and trust using Bayesian data analysis, and we analyse the sensitivity of that connection with different visual attention models. For evaluation purposes, we collected gaze data and trust questionnaires from 49 interactions where 29 participants engaged in a collaborative assistive cooking task with a robot, for a total duration of 24h53 of data collection. The video for this article is available at https://youtu.be/LQ9Oi88YmWk.
Cedric Goubard, Yiannis Demiris
ACM Trans. Hum. Robot Interact.2
2024 Forecasting Bimanual Object Manipulation Sequences from Unimanual Observations
abstract
Learning to forecast bimanual object manipulation sequences from unimanual observations has broad applications in assistive robots and augmented reality. This challenging task requires us to first infer motion from the missing arm and the object it would have been manipulating were the person bimanual, then forecast the human and object motion while maintaining hand-object contact during manipulation. Previous attempts model the hand-object interactions only implicitly, and thus tend to produce unrealistic motion where the objects float in air. We address this with a novel neural network that (i) identifies and forecasts the pose for only the objects undergoing motion through an object motion module and (ii) refines human pose predictions by encouraging hand-object contact during manipulation through an ensemble of human pose predictors. The components are also designed to be generic enough for use in both unimanual and bimanual contexts. Our approach outperforms the state-of-the-art pose forecasting methods on bimanual manipulation datasets.
Haziq Razali, Yiannis Demiris
AAAI2
2024 DanceMVP: Self-Supervised Learning for Multi-Task Primitive-Based Dance Performance Assessment via Transformer Text Prompting
abstract
Dance is generally considered to be complex for most people as it requires coordination of numerous body motions and accurate responses to the musical content and rhythm. Studies on automatic dance performance assessment could help people improve their sensorimotor skills and promote research in many fields, including human motion analysis and motion generation. Recent papers on dance performance assessment usually evaluate simple dance motions with a single task - estimating final performance scores. In this paper, we propose DanceMVP: multi-task dance performance assessment via text prompting that solves three related tasks - (i) dance vocabulary recognition, (ii) dance performance scoring and (iii) dance rhythm evaluation. In the pre-training phase, we contrastively learn the primitive-based features of complex dance motion and music using the InfoNCE loss. For the downstream task, we propose a transformer-based text prompter to perform multi-task evaluations for the three proposed assessment tasks. Also, we build a multimodal dance-music dataset named ImperialDance. The novelty of our ImperialDance is that it contains dance motions for diverse expertise levels and a significant amount of repeating dance sequences for the same choreography to keep track of the dance performance progression. Qualitative results show that our pre-trained feature representation could cluster dance pieces for different dance genres, choreographies, expertise levels and primitives, which generalizes well on both ours and other dance-music datasets. The downstream experiments demonstrate the robustness and improvement of our method over several ablations and baselines across all three tasks, as well as monitoring the users' dance level progression.
Yun Zhong, Yiannis Demiris
AAAI2
2024 Learning Self-Confidence from Semantic Action Embeddings for Improved Trust in Human-Robot Interaction
abstract
In Human-Robot Interaction (HRI) scenarios, human factors like trust can greatly impact task performance and interaction quality. Recent research has confirmed that perceived robot proficiency is a major antecedent of trust. By making robots aware of their capabilities, we can allow them to choose when to perform low-confidence actions, thus actively controlling the risk of trust reduction. In this paper, we propose Self-Confidence through Observed Novel Experiences (SCONE), a policy to learn self-confidence from experience using semantic action embeddings. Using an assistive cooking setting, we show that the semantic aspect allows SCONE to learn self-confidence faster than existing approaches, while also achieving promising performance in simple instructions following. Finally, we share results from a pilot study with 31 participants, showing that such a self-confidence-aware policy increases capability-based human trust.
Cedric Goubard, Yiannis Demiris
ICRA2
2024 Model Predictive Control with Graph Dynamics for Garment Opening Insertion during Robot-Assisted Dressing
abstract
Robots have a great potential to help people with movement limitations in activities of daily living, such as dressing. A common problem in almost all dressing tasks is the insertion of a garment’s opening around a part of the human body. The rich contact environment and the deformations of the garment make the task a challenging problem for robots. In this paper, we propose a bi-manual control method for garment opening insertion during robot-assisted dressing. Specifically, we propose a model predictive controller that uses an Attention-based Relational Graph Convolutional Network (ARGCN) for modeling the dynamics of the opening in the presence of the body. We train the model entirely in simulation and validate our method in four real-world dressing scenarios of a medical training manikin. We show that our method generalizes well in the real-world opening insertion tasks achieving an overall success rate of 97.5%, even though the dynamics and the shapes vastly differ from the simulation setup.
Stelios Kotsovolis, Yiannis Demiris
ICRA2
2024 Learning Bimanual Manipulation Policies for Bathing Bed-bound People
abstract
Assistive robots hold promise in enhancing the quality of life for older adults and people with mobility impairments in daily bed bathing routines. When providing bathing assistance to bed-bound people, human caregivers often support the joints when lifting the arms and legs to properly wash and dry occluded areas. This research introduces a novel approach to robotic bed bathing manipulation, where a bimanual robot learns to lift a target limb while controlling a cleaning tool to bath the surface within safe force bounds. To ensure safe, cooperative bath manipulation, our work combines Multi-Agent Reinforcement Learning (MARL) framework with a variable impedance action space enabling adaptive interaction with the environment and carefully-designed reward functions regulating contact force on the human body. Simulation results demonstrate improved bathing area coverage compared to unimanual models and exhibit great adaptability to contact-rich interaction within a safe force boundary. We validate our approach across various human body sizes, showcasing its generalizability. We also transfer our models to a physical Baxter robot bathing a medical-grade manikin. We further incorporate a force tracking controller with the trained models to enhance adaptation to noisy real-world bathing scenarios. To the best of our knowledge, this is the first robot-assisted bed bathing application that performs autonomous bathing around the human body using bimanual robot arms.
Yijun Gu, Yiannis Demiris
IROS2
2024 On the Effect of Augmented-Reality Multi-User Interfaces and Shared Mental Models on Human-Robot Trust
abstract
Augmented Reality multi-user interfaces facilitate communication, coordination and collaboration among teams. Moreover, these interfaces can help to align the team’s perceptions and expectations under a shared mental model. This model is a psychological construct that represents the common knowledge, beliefs, and understandings held by team members. In this paper, we study to what extent, if any, the combination of Augmented Reality multi-user interfaces and shared mental models affects human-robot trust. To this end, we developed an Augmented Reality multi-user interface to perform a user study (N = 37) comparing non-dyadic human-robot interactions with a quadruped robot exhibiting low reliability (Group 3), against dyadic interactions while the robot exhibited high-reliability (Group 1) or low-reliability (Group 2). We made this comparison using validated trust questionnaires relevant to HRI. Our results, obtained via Bayesian data analysis methods, show differences in the distribution of answers between groups 1 and 2. Notably, this difference is smaller between groups 1 and 3, which suggests that the combination of shared mental models and multi-user interfaces holds promise as an effective way to manage and calibrate human-robot trust.
Rodrigo Chacon, Fernando E. Casado, Yiannis Demiris
RO-MAN3
2024 A framework for trust-related knowledge transfer in human-robot interaction
abstract
Abstract Trustworthy human–robot interaction (HRI) during activities of daily living (ADL) presents an interesting and challenging domain for assistive robots, particularly since methods for estimating the trust level of a human participant towards the assistive robot are still in their infancy. Trust is a multifaced concept which is affected by the interactions between the robot and the human, and depends, among other factors, on the history of the robot’s functionality, the task and the environmental state. In this paper, we are concerned with the challenge of trust transfer, i.e. whether experiences from interactions on a previous collaborative task can be taken into consideration in the trust level inference for a new collaborative task. This has the potential of avoiding re-computing trust levels from scratch for every new situation. The key challenge here is to automatically evaluate the similarity between the original and the novel situation, then adapt the robot’s behaviour to the novel situation using previous experience with various objects and tasks. To achieve this, we measure the semantic similarity between concepts in knowledge graphs (KGs) and adapt the robot’s actions towards a specific user based on personalised interaction histories. These actions are grounded and then verified before execution using a geometric motion planner to generate feasible trajectories in novel situations. This framework has been experimentally tested in human–robot handover tasks in different kitchen scene contexts. We conclude that trust-related knowledge positively influences and improves collaboration in both performance and time aspects.
Mohammed Diab, Yiannis Demiris
Auton. Agents Multi Agent Syst.2
2024 Naturalistic Robot-to-Human Bimanual Handover in Complex Environments Through Multi-Sensor Fusion
abstract
Robot-human object handover has been extensively studied in recent years for a wide range of applications. However, it is still far from being as natural as human-human handovers, largely due to the robots’ limited sensing capabilities. Previous approaches in the literature typically simplify the handover scenarios, including one or more of (a) conducting handovers at fixed locations, (b) not adapting to human preferences, or (c) only focusing on single-arm handover with small objects due to the sensor occlusions caused by large objects. To advance the state of the art toward a human-human level of handover fluency, this paper investigates a bimanual handover scenario in a naturalistic, complex setup. Specifically, we target robot-to-human box transfer while the human partner is on a ladder, and ensure that the object is adaptively delivered based on human preferences. To address the occlusion problem that arises in a complex environment, we develop an onboard multi-sensor perception system for the bimanual robot, introduce a measurement confidence estimation technique, and propose an occlusion-resilient multi-sensor fusion technique by positioning visual perception sensors in distinct locations on the robot with different fields of view. In addition, we establish a Cartesian space controller with a quaternion approach and a leader-follower control structure for compliant motion. Four distinct experiments are conducted, covering different human preferences (such as the box delivered above or below the hands) and significant handover location changes once the process has begun. For validation, the proposed multi-sensor fusion technique was compared to a single-sensor approach for both top and bottom sensors separately, and to simple averaging of both sensors. 30 repetitions were performed for each experiment (four experiments, four methods), the equivalent of 480 handover repetitions in total. Multi-sensor fusion approach achieved a handover success rate above$\textbf{86.7\%}$for all experiments by successfully combining the strengths of both fields of view for human pose tracking under significant occlusions without sacrificing handover duration. In contrast, due to the occlusions, the single-sensor and simple averaging approaches completely failed during challenging experiments, illustrating the importance of multi-sensor fusion in complex handover scenarios.Note to Practitioners—This paper is motivated by enabling naturalistic robot-to-human bimanual object handovers in complex environments, which is a challenging problem due to occlusions. Existing approaches in the literature do not benefit from multi-sensor fusion to handle occlusions, which is essential in such physical human-robot interaction scenarios. To this aim, we have developed a multi-sensor fusion technique to improve the perception capabilities of robots with respect to human co-workers. The developed framework has been tested with Microsoft Azure Kinect sensors and a bimanual mobile Baxter robot, but it can be adapted to any depth perception sensor and bimanual robotic platform. Furthermore, the introduced multi-sensor fusion technique is comprehensive and generic, as it can be applied to any intermittent sensor data, such as human pose tracking via RGBD sensors. The presented approach shows that increasing the field of view of robots‘ perception used with enhanced data fusion could drastically improve the robot‘s sensing capability. For future work, data fusion can be improved by introducing Bayesian filters, and the system can be validated with different sensors and robotic platforms. Moreover, the handover detection method of physical interaction could further benefit from the incorporation of force sensors.
Salih Ertug Ovur, Yiannis Demiris
IEEE Trans Autom. Sci. Eng.2
2024 Multi-Dimensional Evaluation of an Augmented Reality Head-Mounted Display User Interface for Controlling Legged Manipulators
abstract
Controlling assistive robots can be challenging for some users, especially those lacking relevant experience. Augmented Reality (AR) User Interfaces (UIs) have the potential to facilitate this task. Although extensive research regarding legged manipulators exists, comparatively little is on their UIs. Most existing UIs leverage traditional control interfaces such as joysticks, Hand-Held (HH) controllers and 2D UIs. These interfaces not only risk being unintuitive, thus discouraging interaction with the robot partner, but also draw the operator’s focus away from the task and towards the UI. This shift in attention raises additional safety concerns, particularly in potentially hazardous environments where legged manipulators are frequently deployed. Moreover, traditional interfaces limit the operators’ availability to use their hands for other tasks. Towards overcoming these limitations, in this article, we provide a user study comparing an AR Head-Mounted Display (HMD) UI we developed for controlling a legged manipulator against off-the-shelf control methods for such robots. This user study involved 27 participants and 135 trials, from which we gathered over 405 completed questionnaires. These trials involved multiple navigation and manipulation tasks with varying difficulty levels using a Boston Dynamics’s Spot, a 7 df Kinova robot arm and a Robotiq 2F-85 gripper that we integrated into a legged manipulator. We made the comparison between UIs across multiple dimensions relevant to a successful human–robot interaction. These dimensions include cognitive workload, technology acceptance, fluency, system usability, immersion and trust. Our study employed a factorial experimental design with participants undergoing five different conditions, generating longitudinal data. Due to potential unknown distributions and outliers in such data, using parametric methods for its analysis is questionable, and while non-parametric alternatives exist, they may lead to reduced statistical power. Therefore, to analyse the data that resulted from our experiment, we chose Bayesian data analysis as an effective alternative to address these limitations. Our results show that AR UIs can outpace HH-based control methods and reduce the cognitive requirements when designers include hands-free interactions and cognitive offloading principles into the UI. Furthermore, the use of the AR UI together with our cognitive offloading feature resulted in higher usability scores and significantly higher fluency and Technology Acceptance Model scores. Regarding immersion, our results revealed that the response values for the AR Immersion questionnaire associated with the AR UI are significantly higher than those associated with the HH UI, regardless of the main interaction method with the former, i.e., hand gestures or cognitive offloading. Derived from the participants’ qualitative answers, we believe this is due to a combination of factors, of which the most important is the free use of the hands when using the HMD, as well as the ability to see the real environment without the need to divert their attention to the UI. Regarding trust, our findings did not display discernible differences in reported trust scores across UI options. However, during the manipulation phase of our user study, where participants were given the choice to select their preferred UI, they consistently reported higher levels of trust compared to the navigation category. Moreover, there was a drastic change in the percentage of participants that selected the AR UI for completing this manipulation stage after incorporating the cognitive offloading feature. Thus, trust seems to have mediated the use and non-use of the UIs in a dimension different from the ones considered in our study, i.e., delegation and reliance. Therefore, our AR HMD UI for the control of legged manipulators was found to improve human–robot interaction across several relevant dimensions, underscoring the critical role of UI design in the effective and trustworthy utilisation of robotic systems.
Rodrigo Chacon, Yiannis Demiris
ACM Trans. Hum. Robot Interact.2
2024 3PFS: Protecting Pedestrian Privacy Through Face Swapping
abstract
In the era of artificial intelligence, privacy has become a paramount concern, especially within intelligent transportation systems (ITS) where pedestrians are frequently captured by vehicle-mounted cameras for deep learning model training. To address this, we introduce 3PFS, a novel method designed to protect pedestrian privacy via face swapping while preserving the utility of processed images. Our method consists of a pedestrian detector, a face detector, a pre-processing module, a source face selection algorithm, and a face swapping algorithm. After detecting pedestrians and their corresponding faces, the pre-processing module enhances image quality. Our unique source face selection algorithm then chooses an appropriate face from our source face library, which is subsequently swapped with the target face using a face swapping algorithm. Notably, with the combination of a pedestrian tracking algorithm, our 3PFS is well-suited for video anonymization. Additionally, we propose a comprehensive evaluation strategy to evaluate the performance of pedestrian anonymization methods. We validate the effectiveness of 3PFS through extensive experiments on a dataset we created based on the publicly available JAAD dataset and on videos captured using our robotic wheelchair.
Zixian Zhao 0001, Yiannis Demiris
IEEE Trans. Intell. Transp. Syst.3
2023 Action-Conditioned Generation of Bimanual Object Manipulation Sequences
abstract
The generation of bimanual object manipulation sequences given a semantic action label has broad applications in collaborative robots or augmented reality. This relatively new problem differs from existing works that generate whole-body motions without any object interaction as it now requires the model to additionally learn the spatio-temporal relationship that exists between the human joints and object motion given said label. To tackle this task, we leverage the varying degree each muscle or joint is involved during object manipulation. For instance, the wrists act as the prime movers for the objects while the finger joints are angled to provide a firm grip. The remaining body joints are the least involved in that they are positioned as naturally and comfortably as possible. We thus design an architecture that comprises 3 main components: (i) a graph recurrent network that generates the wrist and object motion, (ii) an attention-based recurrent network that estimates the required finger joint angles given the graph configuration, and (iii) a recurrent network that reconstructs the body pose given the locations of the wrist. We evaluate our approach on the KIT Motion Capture and KIT RGBD Bimanual Manipulation datasets and show improvements over a simplified approach that treats the entire body as a single entity, and existing whole-body-only methods.
Haziq Razali, Yiannis Demiris
AAAI2
2023 Contrastive Self-Supervised Learning for Automated Multi-Modal Dance Performance Assessment
abstract
A fundamental challenge of analyzing human motion is to effectively represent human movements both spatially and temporally. We propose a contrastive self-supervised strategy to tackle this challenge. Particularly, we focus on dancing, which involves a high level of physical and intellectual abilities. Firstly, we deploy Graph and Residual Neural Networks with Siamese architecture to represent the dance motion and music features respectively. Secondly, we apply the InfoNCE loss to contrastively embed the high-dimensional multimedia signals onto the latent space without label supervision. Finally, our proposed framework is evaluated on a multi-modal Dance- Music-Level dataset composed of various dance motions, music, genres and choreographies with dancers of different expertise levels. Experimental results demonstrate the robustness and improvements of our proposed method over 3 baselines and 6 ablation studies across tasks of dance genres, choreographies classification and dancer expertise level assessment.
Yun Zhong, Fan Zhang 0030, Yiannis Demiris
ICASSP3
2023 Design and Evaluation of an Augmented Reality Head-Mounted Display User Interface for Controlling Legged Manipulators
abstract
Designing an intuitive User Interface (UI) for controlling assistive robots remains challenging. Most existing UIs leverage traditional control interfaces such as joysticks, hand-held controllers, and 2D UIs. Thus, users have limited availability to use their hands for other tasks. Furthermore, although there is extensive research regarding legged manipulators, comparatively little is on their UIs. Towards extending the state-of-art in this domain, we provide a user study comparing an Augmented Reality (AR) Head-Mounted Display (HMD) UI we developed for controlling a legged manipulator against off-the-shelf control methods for such robots. We made this comparison baseline across multiple factors relevant to a successful interaction. The results from our user study ($N=17$) show that although the AR UI increases immersion, off-the-shelf control methods outperformed the AR UI in terms of time performance and cognitive workload. Nonetheless, a follow-up pilot study incorporating the lessons learned shows that AR UIs can outpace hand-held-based control methods and reduce the cognitive requirements when designers include hands-free interactions and cognitive offloading principles into the UI.
Rodrigo Chacon, Yiannis Demiris
ICRA2
2023 Bi-Manual Manipulation of Multi-Component Garments towards Robot-Assisted Dressing
abstract
In this paper, we propose a strategy for robot-assisted dressing with multi-component garments, such as gloves. Most studies in robot-assisted dressing usually experiment with single-component garments, such as sleeves, while multi-component tasks are often approached as sequential single-component problems. In dressing scenarios with more complex garments, robots should estimate the alignment of the human body to the manipulated garments, and revise their dressing strategy. In this paper, we focus on a glove dressing scenario and propose a decision process for selecting dressing action primitives on the different components of the garment, based on a hierarchical representation of the task and a set of environmental conditions. To complement this process, we propose a set of bi-manual control strategies, based on hybrid position, visual, and force feedback, in order to execute the dressing action primitives with the deformable object. The experimental results validate our method, enabling the Baxter robot to dress a mannequin's hand with a gardening glove.
Stelios Kotsovolis, Yiannis Demiris
ICRA2
2023 Bi-Manual Robot Shoe Lacing
abstract
Shoe lacing (SL) is a challenging sensorimotor task in daily life and a complex engineering problem in the shoe-making industry. In this paper, we propose a system for autonomous SL. It contains a mathematical definition of the SL task and searches for the best lacing pattern corresponding to the shoe configuration and the user preferences. We propose a set of action primitives and generate plans of action sequences according to the designed pattern. Our system plans the trajectories based on the perceived position of the eyelets and aglets with an active perception strategy, and deploys the trajectories on a bi-manual robot. Experiments demonstrate that the proposed system can successfully lace 3 different shoes in different configurations, with a completion rate of 92.0%, 91.6% and 77.5% for 6, 8 and 10-eyelet patterns respectively. To the best of our knowledge, this is the first demonstration of autonomous SL using a bi-manual robot.
Haining Luo, Yiannis Demiris
IROS2
2023 Beyond Self-Report: A Continuous Trust Measurement Device for HRI
abstract
Trust is a crucial part of human-robot interactions, and its accurate measurement is a challenging task. We introduce Trusty, a handheld continuous trust level measurement device and investigate its validity by analysing the correlation between its measurements and self-reported trust scores. In a study with 29 participants, we evaluated the effectiveness of the device with an autonomous wheelchair in a mobile navigation task. The participants collaborated with an autonomous wheelchair to deliver packages to predefined checkpoints in an unstructured environment, and the performance of the wheelchair was manipulated to be either under a good-performing condition or a bad-performing condition. Our first finding reveals a notable influence of wheelchair performance on self-reported trust. Participants interacting with a good-performing wheelchair exhibited increased trust levels, as evidenced by higher scores on post-experiment trust questionnaires and verbal self-reported trust measures. Additionally, our study proposes Trusty as a continuous measurement tool for assessing trust during HRI, demonstrating its equivalence to self-report measures and traditional questionnaire scores.
Nico Lingg, Yiannis Demiris
RO-MAN2
2023 Risk-aware controller for autonomous vehicles using model-based collision prediction and reinforcement learning
abstract
Autonomous Vehicles (AVs) have the potential to save millions of lives and increase the efficiency of transportation services. However, the successful deployment of AVs requires tackling multiple challenges related to modeling and certifying safety. State-of-the-art decision-making methods usually rely on end-to-end learning or imitation learning approaches, which still pose significant safety risks. Hence the necessity of risk-aware AVs that can better predict and handle dangerous situations. Furthermore, current approaches tend to lack explainability due to their reliance on end-to-end Deep Learning, where significant causal relationships are not guaranteed to be learned from data. This paper introduces a novel risk-aware framework for training AV agents using a bespoke collision prediction model and Reinforcement Learning (RL). The collision prediction model is based on Gaussian Processes and vehicle dynamics, and is used to generate the RL state vector. Using an explicit risk model increases the post-hoc explainability of the AV agent, which is vital for reaching and certifying the high safety levels required for AVs and other safety-sensitive applications. Experimental results obtained with a simulator and state-of-the-art RL algorithms show that the risk-aware RL framework decreases average collision rates by 15%, makes AVs more robust to sudden harsh braking situations, and achieves better performance in both safety and speed when compared to a standard rule-based method (the Intelligent Driver Model). Moreover, the proposed collision prediction model outperforms other models in the literature.
Eduardo Candela, Olivier Doustaly, Leandro Parada, Felix Feng, Yiannis Demiris, Panagiotis Angeloudis
Artif. Intell.5
2023 Visible and Infrared Image Fusion Using Deep Learning
abstract
Visible and infrared image fusion (VIF) has attracted a lot of interest in recent years due to its application in many tasks, such as object detection, object tracking, scene segmentation, and crowd counting. In addition to conventional VIF methods, an increasing number of deep learning-based VIF methods have been proposed in the last five years. Different types of methods, such as CNN-based, autoencoder-based, GAN-based, and transformer-based methods, have been proposed. Deep learning-based methods have undoubtedly become dominant methods for the VIF task. However, while much progress has been made, the field will benefit from a systematic review of these deep learning-based methods. In this paper we present a comprehensive review of deep learning-based VIF methods. We discuss motivation, taxonomy, recent development characteristics, datasets, and performance evaluation methods in detail. We also discuss future prospects of the VIF field. This paper can serve as a reference for VIF researchers and those interested in entering this fast-developing field.
Yiannis Demiris
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Dual-branch spatio-temporal graph neural networks for pedestrian trajectory prediction
abstract
Pedestrian trajectory prediction is an important area in computer vision, with wide applications in autonomous driving, robot path planning, and surveillance systems. The core underlying technique of these applications is pattern recognition. A key challenge in this area is modeling social interactions between pedestrians, such as pedestrian view area and group behaviors. However, although many methods have been proposed to model social interactions, pedestrian view area and group behaviors have not been explored together to account for complex situations. Additionally, most existing studies require additional detectors and manual annotations to handle view area and group interactions, respectively. In this paper, we propose a dual-branch spatio-temporal graph neural network to automatically model view area and grouping together. Specifically, a spatio-temporal graph attention network (STGAT) branch is designed to handle pedestrian view area, and a spatio-temporal graph convolutional network (STGCN) branch is designed to model group interactions. The features of these branches are then fused to provide better feature representations, on which a temporal convolution operation (TCN) is performed for trajectory prediction. Experiments on public standard datasets demonstrate that the proposed method achieves very competitive performance and predicts socially acceptable trajectories in different challenging scenarios.
Panagiotis Angeloudis, Yiannis Demiris
Pattern Recognit.3
2022 Faster, Better Blink Detection through Curriculum Learning by Augmentation
abstract
Blinking is a useful biological signal that can gate gaze regression models to avoid the use of incorrect data in downstream tasks. Existing datasets are imbalanced both in frequency of class but also in intra-class difficulty which we demonstrate is a barrier for curriculum learning. We thus propose a novel curriculum augmentation scheme that aims to address frequency and difficulty imbalances implicitly which are are terming Curriculum Learning by Augmentation (CLbA).
Ahmed Al-Hindawi, Marcela P. Vizcaychipi, Yiannis Demiris
ETRA3
2022 What Is The Patient Looking At? Robust Gaze-Scene Intersection Under Free-Viewing Conditions
abstract
Locating the user’s gaze in the scene, also known as Point of Regard (PoR) estimation, following gaze regression is important for many downstream tasks. Current techniques either require the user to wear and calibrate instruments, require significant pre-processing of the scene information, or place restrictions on user’s head movements.We propose a geometrically inspired algorithm that, despite its simplicity, provides high accuracy and O(J) performance under a variety of challenging situations including sparse depth maps, high noise, and high dynamic parallax between the user and the scene camera. We demonstrate the utility of the proposed algorithm in regressing the PoR from scenes captured in the Intensive Care Unit (ICU) at Chelsea & Westminster Hospital NHS Foundation Trusta.
Ahmed Al-Hindawi, Marcela P. Vizcaychipi, Yiannis Demiris
ICASSP3
2022 Using a Single Input to Forecast Human Action Keystates in Everyday Pick and Place Actions
abstract
We define action keystates as the start or end of an action that contains information such as the human pose and time. Existing methods that forecast the human pose use recurrent networks that input and output a sequence of poses. In this paper, we present a method tailored for everyday pick and place actions where the object of interest is known. In contrast to existing methods, ours uses an input from a single timestep to directly forecast (i) the key pose the instant the pick or place action is performed and (ii) the time it takes to get to the predicted key pose. Experimental results show that our method outperforms the state-of-the-art for key pose forecasting and is comparable for time forecasting while running at least an order of magnitude faster. Further ablative studies reveal the significance of the object of interest in enabling the total number of parameters across all existing methods to be reduced by at least 90% without any degradation in performance.a
Haziq Razali, Yiannis Demiris
ICASSP2
2022 Message Passing Framework for Vision Prediction Stability in Human Robot Interaction
abstract
In Human Robot Interaction (HRI) scenarios, robot systems would benefit from an understanding of the user's state, actions and their effects on the environments to enable better interactions. While there are specialised vision algorithms for different perceptual channels, such as objects, scenes, human pose, and human actions, it is worth considering how their interaction can help improve each other's output. In computer vision, individual prediction modules for these perceptual channels frequently produce noisy outputs due to the limited datasets used for training and the compartmentalisation of the perceptual channels, often resulting in noisy or unstable prediction outcomes. To stabilise vision prediction results in HRI, this paper presents a novel message passing framework that uses the memory of individual modules to correct each other's outputs. The proposed framework is designed utilising common-sense rules of physics (such as the law of gravity) to reduce noise while introducing a pipeline that helps to effectively improve the output of each other's modules. The proposed framework aims to analyse primitive human activities such as grasping an object in a video captured from the perspective of a robot. Experimental results show that the proposed framework significantly reduces the output noise of individual modules compared to the case of running independently. This pipeline can be used to measure human reactions when interacting with a robot in various HRI scenarios.
Youngkyoon Jang, Yiannis Demiris
ICRA2
2022 Kinematic Structure Estimation of Arbitrary Articulated Rigid Objects for Event Cameras
abstract
We propose a novel method that estimates the Kinematic Structure (KS) of arbitrary articulated rigid objects from event-based data. Event cameras are emerging sensors that asynchronously report brightness changes with a time resolution of microseconds, making them suitable candidates for motion-related perception. By assuming that an articulated rigid object is composed of body parts whose shape can be approximately described by a Gaussian distribution, we jointly segment the different parts by combining an adapted Bayesian inference approach and incremental event-based motion estimation. The respective KS is then generated based on the segmented parts and their respective biharmonic distance, which is estimated by building an affinity matrix of points sampled from the estimated Gaussian distributions. The method outperforms frame-based methods in sequences obtained by simulating events from video sequences and achieves a solid performance on new high-speed motions sequences, which frame-based KS estimation methods can not handle.
Urbano Miguel Nunes, Yiannis Demiris
ICRA2
2022 Using Eye Gaze to Forecast Human Pose in Everyday Pick and Place Actions
abstract
Collaborative robots that operate alongside humans require the ability to understand their intent and forecast their pose. Among the various indicators of intent, the eye gaze is particularly important as it signals action towards the gazed object. By observing a person's gaze, one can effectively predict the object of interest and subsequently, forecast the person's pose. We leverage this and present a method that forecasts the human pose using gaze information for everyday pick and place actions in a home environment. Our method first attends to fixations to locate the coordinates of the object of interest before inputting said coordinates to a pose forecasting network. Experiments on the MoGaze dataset show that our gaze network lowers the errors of existing pose forecasting methods and that incorporating prior in the form of textual instructions further lowers the errors by a significant amount. Furthermore, the use of eye gaze now allows a simple multilayer perceptron network to directly forecast the keypose.aaCode available at www.imperial.ac.uk/personal-robotics/software
Haziq Razali, Yiannis Demiris
ICRA2
2022 Transferring Multi-Agent Reinforcement Learning Policies for Autonomous Driving using Sim-to-Real
abstract
Autonomous Driving requires high levels of coordination and collaboration between agents. Achieving effective coordination in multi-agent systems is a difficult task that remains largely unresolved. Multi-Agent Reinforcement Learning has arisen as a powerful method to accomplish this task because it considers the interaction between agents and also allows for decentralized training—which makes it highly scalable. However, transferring policies from simulation to the real world is a big challenge, even for single-agent applications. Multi-agent systems add additional complexities to the Sim-to-Real gap due to agent collaboration and environment synchronization. In this paper, we propose a method to transfer multi-agent autonomous driving policies to the real world. For this, we create a multi-agent environment that imitates the dynamics of the Duckietown multi-robot testbed, and train multi-agent policies using the MAPPO algorithm with different levels of domain randomization. We then transfer the trained policies to the Duckietown testbed and show that when using our method, domain randomization can reduce the reality gap by 90%. Moreover, we show that different levels of parameter randomization have a substantial impact on the Sim-to-Real gap. Finally, our approach achieves significantly better results than a rule-based benchmark.
Eduardo Candela, Leandro Parada, Luís Marques 0002, Tiberiu-Andrei Georgescu, Yiannis Demiris, Panagiotis Angeloudis
IROS5
2022 Federated Learning from Demonstration for Active Assistance to Smart Wheelchair Users
abstract
Learning from Demonstration (LfD) is a very appealing approach to empower robots with autonomy. Given some demonstrations provided by a human teacher, the robot can learn a policy to solve the task without explicit programming. A promising use case is to endow smart robotic wheelchairs with active assistance to navigation. By using LfD, it is possible to learn to infer short-term destinations anywhere, without the need of building a map of the environment beforehand. Nevertheless, it is difficult to generalize robot behaviors to environments other than those used for training. We believe that one possible solution is learning from crowds, involving a broad number of teachers (the end users themselves) who perform demonstrations in diverse and real environments. To this end, in this work we consider Federated Learning from Demonstration (FLfD), a distributed approach based on a Federated Learning architecture. Our proposal allows the training of a global deep neural network using sensitive local data (images and laser readings) with privacy guarantees. In our experiments we pose a scenario involving different clients working in heterogeneous domains. We show that the federated model is able to generalize and deal with non Independent and Identically Distributed (non-IID) data.
Fernando E. Casado, Yiannis Demiris
IROS2
2022 Holo-SpoK: Affordance-Aware Augmented Reality Control of Legged Manipulators
abstract
Although there is extensive research regarding legged manipulators, comparatively little focuses on their User Interfaces (UIs). Towards extending the state-of-art in this domain, in this work, we integrate a Boston Dynamics (BD) Spot® with a light-weight 7 DoF Kinova® robot arm and a Robotiq® 2F-85 gripper into a legged manipulator. Furthermore, we jointly control the robotic platform using an affordance-aware Augmented Reality (AR) Head-Mounted Display (HMD) UI developed for the Microsoft HoloLens 2. We named the combined platform Holo-SpoK. Moreover, we explain how this manipulator colocalises with the HoloLens 2 for its control through AR. In addition, we present the details of our algorithms for autonomously detecting grasp-ability affordances and for the refinement of the positions obtained via vision-based colocalisation. We validate the suitability of our proposed methods with multiple navigation and manipulation experiments. To the best of our knowledge, this is the first demonstration of an AR HMD UI for controlling legged manipulators.
Rodrigo Chacon, Yiannis Demiris
IROS2
2022 Disentangled Sequence Clustering for Human Intention Inference
abstract
Equipping robots with the ability to infer human intent is a vital precondition for effective collaboration. Most computational approaches towards this objective derive a probability distribution of “intent” conditioned on the robot's perceived state. However, these approaches typically assume task-specific labels of human intent are known a priori. To overcome this constraint, we propose the Disentangled Sequence Clustering Variational Autoencoder (DiSCVAE), a clustering framework capable of learning such a distribution of intent in an unsupervised manner. The proposed framework leverages recent advances in unsupervised learning to disentangle latent representations of sequence data, separating time-varying local features from time-invariant global attributes. As a novel extension, the DiSCVAE also infers a discrete variable to form a latent mixture model and thus enable clustering over these global sequence concepts, e.g. high-level intentions. We evaluate the DiSCVAE on a real-world human-robot interaction dataset collected using a robotic wheelchair. Our findings reveal that the inferred discrete variable coincides with human intent, holding promise for collaborative settings, such as shared control.
Mark Zolotas, Yiannis Demiris
IROS2
2022 Robust Event-Based Vision Model Estimation by Dispersion Minimisation
abstract
We propose a novel Dispersion Minimisation framework for event-based vision model estimation, with applications to optical flow and high-speed motion estimation. The framework extends previous event-based motion compensation algorithms by avoiding computing an optimisation score based on an explicit image-based representation, which provides three main benefits: i) The framework can be extended to perform incremental estimation, i.e., on an event-by-event basis. ii) Besides purely visual transformations in 2D, the framework can readily use additional information, e.g., by augmenting the events with depth, to estimate the parameters of motion models in higher dimensional spaces. iii) The optimisation complexity only depends on the number of events. We achieve this by modelling the event alignment according to candidate parameters and minimising the resultant dispersion, which is computed by a family of suitable entropy-based measures. Data whitening is also proposed as a simple and effective pre-processing step to make the framework's accuracy performance more robust, as well as other event-based motion-compensation methods. The framework is evaluated on several challenging motion estimation problems, including 6-DOF transformation, rotational motion, and optical flow estimation, achieving state-of-the-art performance.
Urbano Miguel Nunes, Yiannis Demiris
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 A Cloud-based Robot System for Long-term Interaction: Principles, Implementation, Lessons Learned
abstract
Making the transition to long-term interaction with social-robot systems has been identified as one of the main challenges in human-robot interaction. This article identifies four design principles to address this challenge and applies them in a real-world implementation: cloud-based robot control, a modular design, one common knowledge base for all applications, and hybrid artificial intelligence for decision making and reasoning. The control architecture for this robot includes a common Knowledge-base (ontologies), Data-base, “Hybrid Artificial Brain” (dialogue manager, action selection and explainable AI), Activities Centre (Timeline, Quiz, Break and Sort, Memory, Tip of the Day, \( \ldots \) ), Embodied Conversational Agent (ECA, i.e., robot and avatar), and Dashboards (for authoring and monitoring the interaction). Further, the ECA is integrated with an expandable set of (mobile) health applications. The resulting system is a Personal Assistant for a healthy Lifestyle (PAL), which supports diabetic children with self-management and educates them on health-related issues (48 children, aged 6–14, recruited via hospitals in the Netherlands and in Italy). It is capable of autonomous interaction “in the wild” for prolonged periods of time without the need for a “Wizard-of-Oz” (up until 6 months online). PAL is an exemplary system that provides personalised, stable and diverse, long-term human-robot interaction.
Frank Kaptein, Bernd Kiefer, Antoine Cully, Oya Çeliktutan, Bert P. B. Bierman, Rifca Rijgersberg-Peters, Joost Broekens, Willeke van Vught, Michael van Bekkum, Yiannis Demiris, Mark A. Neerincx
ACM Trans. Hum. Robot Interact.10
2022 HammerDrive: A Task-Aware Driving Visual Attention Model
abstract
We introduce HammerDrive, a novel architecture for task-aware visual attention prediction in driving. The proposed architecture is learnable from data and can reliably infer the current focus of attention of the driver in real-time, while only requiring limited and easy-to-access telemetry data from the vehicle. We build the proposed architecture on two core concepts: 1) driving can be modeled as a collection of sub-tasks (maneuvers), and 2) each sub-task affects the way a driver allocates visual attention resources, i.e., their eye gaze fixation. HammerDrive comprises two networks: a hierarchical monitoring network of forward-inverse model pairs for sub-task recognition and an ensemble network of task-dependent convolutional neural network modules for visual attention modeling. We assess the ability of HammerDrive to infer driver visual attention on data we collected from 20 experienced drivers in a virtual reality-based driving simulator experiment. We evaluate the accuracy of our monitoring network for sub-task recognition and show that it is an effective and light-weight network for reliable real-time tracking of driving maneuvers with above 90% accuracy. Our results show that HammerDrive outperforms a comparable state-of-the-art deep learning model for visual attention prediction on numerous metrics with ~13% improvement for both Kullback-Leibler divergence and similarity, and demonstrate that task-awareness is beneficial for driver visual attention prediction.
Pierluigi Vito Amadori, Tobias Fischer 0001, Yiannis Demiris
IEEE Trans. Intell. Transp. Syst.3
2022 ST CrossingPose: A Spatial-Temporal Graph Convolutional Network for Skeleton-Based Pedestrian Crossing Intention Prediction
abstract
Pedestrian crossing intention prediction is crucial for the safety of pedestrians in the context of both autonomous and conventional vehicles and has attracted widespread interest recently. Various methods have been proposed to perform pedestrian crossing intention prediction, among which the skeleton-based methods have been very popular in recent years. However, most existing studies utilize manually designed features to handle skeleton data, limiting the performance of these methods. To solve this issue, we propose to predict pedestrian crossing intention based on spatial-temporal graph convolutional networks using skeleton data (ST CrossingPose). The proposed method can learn both spatial and temporal patterns from skeleton data, thus having a good feature representation ability. Extensive experiments on a public dataset demonstrate that the proposed method achieves very competitive performance in predicting crossing intention while maintaining a fast inference speed. We also analyze the effect of several factors, e.g., size of pedestrians, time to event, and occlusion, on the proposed method.
Panagiotis Angeloudis, Yiannis Demiris
IEEE Trans. Intell. Transp. Syst.3
2022 Monocular Visual Traffic Surveillance: A Review
abstract
To facilitate the monitoring and management of modern transportation systems, monocular visual traffic surveillance systems have been widely adopted for speed measurement, accident detection, and accident prediction. Thanks to the recent innovations in computer vision and deep learning research, the performance of visual traffic surveillance systems has been significantly improved. However, despite this success, there is a lack of survey papers that systematically review these new methods. Therefore, we conduct a systematic review of relevant studies to fill this gap and provide guidance to future studies. This paper is structured along the visual information processing pipeline that includes object detection, object tracking, and camera calibration. Moreover, we also include important applications of visual traffic surveillance systems, such as speed measurement, behavior learning, accident detection and prediction. Finally, future research directions of visual traffic surveillance systems are outlined.
Panagiotis Angeloudis, Yiannis Demiris
IEEE Trans. Intell. Transp. Syst.4
2021 Embodied Reasoning for Discovering Object Properties via Manipulation
abstract
In this paper, we present an integrated system that includes reasoning from visual and natural language inputs, action and motion planning, executing tasks by a robotic arm, manipulating objects, and discovering their properties. A vision to action module recognises the scene with objects and their attributes and analyses enquiries formulated in natural language. It performs multi-modal reasoning and generates a sequence of simple actions that can be executed by a robot. The scene model and action sequence are sent to a planning and execution module that generates a motion plan with collision avoidance, simulates the actions, and executes them. We use synthetic data to train various components of the system and test on a real robot to show the generalization capabilities. We focus on a tabletop scenario with objects that can be grasped by our embodied agent i.e. a 7DoF manipulator with a two-finger gripper. We evaluate the agent on 60 representative queries repeated 3 times (e.g., ’Check what is on the other side of the soda can’) concerning different objects and tasks in the scene. We perform experiments in a simulated and real environment and report the success rate for various components of the system. Our system achieves up to 80.6% success rate on challenging scenes and queries. We also analyse and discuss the challenges that such an intelligent embodied system faces.
Jan Kristof Behrens, Michal Nazarczuk, Karla Stépánová, Matej Hoffmann, Yiannis Demiris, Krystian Mikolajczyk
ICRA5
2021 Multitask Variational Autoencoding of Human-to-Human Object Handover
abstract
Assistive robots that operate alongside humans require the ability to understand and replicate human behaviours during a handover. A handover is defined as a joint action between two participants in which a giver hands an object over to the receiver. In this paper, we present a method for learning human-to-human handovers observed from motion capture data. Given the giver and receiver pose from a single timestep, and the object label in the form of a word embedding, our Multitask Variational Autoencoder jointly forecasts their pose as well as the orientation of the object held by the giver at handover. Our method is in large contrast to existing works for human pose forecasting that employ deep autoregressive models requiring a sequence of inputs. Furthermore, our method is novel in that it learns both the human pose and object orientation in a joint manner. Experimental results on the publicly available Handover Orientation and Motion Capture Dataset show that our proposed method outperforms the autoregressive baselines for handover pose forecasting by approximately 20% while being on-par for object orientation prediction with a runtime that is 5x faster.a
Haziq Razali, Yiannis Demiris
IROS2
2020 D2D: Keypoint Extraction with Describe to Detect Approach
Yurun Tian, Vassileios Balntas, Tony Ng, Axel Barroso Laguna, Yiannis Demiris, Krystian Mikolajczyk
ACCV (3)5
2020 Entropy Minimisation Framework for Event-Based Vision Model Estimation
Urbano Miguel Nunes, Yiannis Demiris
ECCV (5)2
2020 Rotational Adjoint Methods for Learning-Free 3D Human Pose Estimation from IMU Data
abstract
We present a new framework for learning-free 3D human pose estimation from Inertial Measurement Unit (IMU) data. The proposed method does not rely on a full motion sequence to calculate a pose for any particular time point and thus can operate in real-time. A cost function based only on joint rotations is used, removing the need for frequent transformations between rotations and 3D Cartesian coordinates. A Jacobian that preserves skeleton structure is derived using adjoint methods from Variational Data Assimilation. To facilitate further research in IMU-based Motion Capture, we provide a dataset that combines RGB and depth images from an Intel RealSense camera, marker-based motion capture from an Optitrack system and Xsens IMU data. We have evaluated our method on both our dataset and the Total Capture dataset, showing an average error across 24 joints of 0.45 and 0.48 radians respectively.
Caterina Buizza, Yiannis Demiris
ICPR2
2020 Learning Grasping Points for Garment Manipulation in Robot-Assisted Dressing
abstract
Assistive robots have the potential to provide tremendous support for disabled and elderly people in their daily dressing activities. Recent studies on robot-assisted dressing usually simplify the setup of the initial robot configuration by manually attaching the garments on the robot end-effector and positioning them close to the user's arm. A fundamental challenge in automating such a process for robots is computing suitable grasping points on garments that facilitate robotic manipulation. In this paper, we address this problem by introducing a supervised deep neural network to locate a predefined grasping point on the garment, using depth images for their invariance to color and texture. To reduce the amount of real data required, which is costly to collect, we leverage the power of simulation to produce large amounts of labeled data. The network is jointly trained with synthetic datasets of depth images and a limited amount of real data. We introduce a robot-assisted dressing system that combines the grasping point prediction method, with a grasping and manipulation strategy which takes grasping orientation computation and robot-garment collision avoidance into account. The experimental results demonstrate that our method is capable of yielding accurate grasping point estimations. The proposed dressing system enables the Baxter robot to autonomously grasp a hospital gown hung on a rail, bring it close to the user and successfully dress the upper-body.
Fan Zhang 0030, Yiannis Demiris
ICRA2
2020 Improving Generalisation in Learning Assistance by Demonstration for Smart Wheelchairs
abstract
Learning Assistance by Demonstration (LAD) is concerned with using demonstrations of a human agent to teach a robot how to assist another human. The concept has previously been used with smart wheelchairs to provide customised assistance to individuals with driving difficulties. A basic premise of this technique is that the learned assistive policy should be able to generalise to environments different than the ones used for training; but this has not been tested before. In this work we evaluate the assistive power and the generalisation capability of LAD using our custom teleoperation and learning system for smart wheelchairs, while seeking to improve it by experimenting with different combinations of dimensionality reduction techniques and machine learning models. Using Autoencoders to reduce the dimension of laserscan data and a Gaussian Process as the learning model, we achieved a 23% improvement in prediction performance against the combination used by the latest work on the field. Using this model to assist a driver exposed to a simulated disability, we observed a 9.8% reduction in track completion times when compared to driving without assistance.
Vinícius Barbosa Schettino, Yiannis Demiris
ICRA2
2020 Transparent Intent for Explainable Shared Control in Assistive Robotics
abstract
Robots supplied with the ability to infer human intent have many applications in assistive robotics. In these applications, robots rely on accurate models of human intent to administer appropriate assistance. However, the effectiveness of this assistance also heavily depends on whether the human can form accurate mental models of robot behaviour. The research problem is to therefore establish a transparent interaction, such that both the robot and human understand each other’s underlying "intent". We situate this problem in our Explainable Shared Control paradigm and present ongoing efforts to achieve transparency in human-robot collaboration.
Mark Zolotas, Yiannis Demiris
IJCAI2
2020 Augmented Reality User Interfaces for Heterogeneous Multirobot Control
abstract
Recent advances in the design of head-mounted augmented reality (AR) interfaces for assistive human-robot interaction (HRI) have allowed untrained users to rapidly and fluently control single-robot platforms. In this paper, we investigate how such interfaces transfer onto multirobot architectures, as several assistive robotics applications need to be distributed among robots that are different both physically and in terms of software. As part of this investigation, we introduce a novel head-mounted AR interface for heterogeneous multirobot control. This interface generates and displays dynamic joint-affordance signifiers, i.e. signifiers that combine and show multiple actions from different robots that can be applied simultaneously to an object. We present a user study with 15 participants analysing the effects of our approach on their perceived fluency. Participants were given the task of filling-out a cup with water making use of a multirobot platform. Our results show a clear improvement in standard HRI fluency metrics when users applied dynamic joint-affordance signifiers, as opposed to a sequence of independent actions.
Rodrigo Chacon, Yiannis Demiris
IROS2
2020 Structured Prediction for Conditional Meta-Learning
abstract
The goal of optimization-based meta-learning is to find a single initialization shared across a distribution of tasks to speed up the process of learning new tasks. Conditional meta-learning seeks task-specific initialization to better capture complex task distributions and improve performance. However, many existing conditional methods are difficult to generalize and lack theoretical guarantees. In this work, we propose a new perspective on conditional meta-learning via structured prediction. We derive task-adaptive structured meta-learning (TASML), a principled framework that yields task-specific objective functions by weighing meta-training data on target tasks. Our non-parametric approach is model-agnostic and can be combined with existing meta-learning methods to achieve conditioning. Empirically, we show that TASML improves the performance of existing meta-learning models, and outperforms the state-of-the-art on benchmark datasets.
Yiannis Demiris, Carlo Ciliberto
NeurIPS2
2020 Real-Time Multi-Person Pose Tracking using Data Assimilation
abstract
We propose a framework for the integration of data assimilation and machine learning methods in human pose estimation, with the aim of enabling any pose estimation method to be run in real-time, whilst also increasing consistency and accuracy. Data assimilation and machine learning are complementary methods: the former allows us to make use of information about the underlying dynamics of a system but lacks the flexibility of a data-based model, which we can instead obtain with the latter. Our framework presents a real-time tracking module for any single or multi-person pose estimation system. Specifically, tracking is performed by a number of Kalman filters initiated for each new person appearing in a motion sequence. This permits tracking of multiple skeletons and reduces the frequency that computationally expensive pose estimation has to be run, enabling online pose tracking. The module tracks for N frames while the pose estimates are calculated for frame N + 1 . This also results in increased consistency of person identification and reduced inaccuracies due to missing joint locations and inversion of left-and right-side joints.
Caterina Buizza, Tobias Fischer 0001, Yiannis Demiris
WACV3
2020 Online Knowledge Level Tracking with Data-Driven Student Models and Collaborative Filtering
abstract
Intelligent Tutoring Systems are promising tools for delivering optimal and personalized learning experiences to students. A key component for their personalization is the student model, which infers the knowledge level of the students to balance the difficulty of the exercises. While important advances have been achieved, several challenges remain. In particular, the models should be able to track in real-time the evolution of the students' knowledge levels. These evolutions are likely to follow different profiles for each student, while measuring the exact knowledge level remains difficult given the limited and noisy information provided by the interactions. This paper introduces a novel model that addresses these challenges with three contributions:. 1) the model relies on Gaussian Processes to track online the evolution of the student's knowledge level over time, 2) it uses collaborative filtering to rapidly provide long-term predictions by leveraging the information from previous users, and 3) it automatically generates abstract representations of knowledge components via automatic relevance determination of covariance matrices. The model is evaluated on three datasets, including real users. The results demonstrate that the model converges to accurate predictions in average four times faster than the compared methods.
Antoine Cully, Yiannis Demiris
IEEE Trans. Knowl. Data Eng.2
2019 Online Unsupervised Learning of the 3D Kinematic Structure of Arbitrary Rigid Bodies
abstract
This work addresses the problem of 3D kinematic structure learning of arbitrary articulated rigid bodies from RGB-D data sequences. Typically, this problem is addressed by offline methods that process a batch of frames, assuming that complete point trajectories are available. However, this approach is not feasible when considering scenarios that require continuity and fluidity, for instance, human-robot interaction. In contrast, we propose to tackle this problem in an online unsupervised fashion, by recursively maintaining the metric distance of the scene's 3D structure, while achieving real-time performance. The influence of noise is mitigated by building a similarity measure based on a linear embedding representation and incorporating this representation into the original metric distance. The kinematic structure is then estimated based on a combination of implicit motion and spatial properties. The proposed approach achieves competitive performance both quantitatively and qualitatively in terms of estimation accuracy, even compared to offline methods.
Urbano Miguel Nunes, Yiannis Demiris
ICCV2
2019 Random Expert Distillation: Imitation Learning via Expert Policy Support Estimation
abstract
We consider the problem of imitation learning from a finite set of expert trajectories, without access to reinforcement signals. The classical approach of extracting the expert’s reward function via inverse reinforcement learning, followed by reinforcement learning is indirect and may be computationally expensive. Recent generative adversarial methods based on matching the policy distribution between the expert and the agent could be unstable during training. We propose a new framework for imitation learning by estimating the support of the expert policy to compute a fixed reward function, which allows us to re-frame imitation learning within the standard reinforcement learning setting. We demonstrate the efficacy of our reward function on both discrete and continuous domains, achieving comparable or better performance than the state of the art under different reinforcement learning algorithms.
Carlo Ciliberto, Pierluigi Vito Amadori, Yiannis Demiris
ICML4
2019 Augmented Reality Controlled Smart Wheelchair Using Dynamic Signifiers for Affordance Representation
abstract
The design of augmented reality interfaces for people with mobility impairments is a novel area with great potential, as well as multiple outstanding research challenges. In this paper we present an augmented reality user interface for controlling a smart wheelchair with a head-mounted display to provide assistance for mobility restricted people. Our motivation is to reduce the cognitive requirements needed to control a smart wheelchair. A key element of our platform is the ability to control the smart wheelchair using the concepts of affordances and signifiers. In addition to the technical details of our platform, we present a baseline study by evaluating our platform through user-trials of able-bodied individuals and two different affordances: 1) Door Go Through and 2) People Approach. To present these affordances to the user, we evaluated fixed symbol based signifiers versus our novel dynamic signifiers in terms of ease to understand the suggested actions and its relation with the objects. Our results show a clear preference for dynamic signifiers. In addition, we show that the task load reported by participants is lower when controlling the smart wheelchair with our augmented reality user interface compared to using the joystick, which is consistent with their qualitative answers.
Rodrigo Chacon, Yiannis Demiris
IROS2
2019 Inference of user-intention in remote robot wheelchair assistance using multimodal interfaces
abstract
Shared control methodologies have the potential of enabling wheelchair-bound users with limited motor abilities to perform tasks that would usually be beyond their capabilities. Deriving such methodologies in advance is challenging, since they are frequently heavily dependent on unique characteristics of users. Learning Assistance by Demonstration paradigms allow derivation of customized policies by recording how remote human assistants assist particular users. However, for accurate determination of the optimal policies for each user and context, the remote assistant needs to infer the intention of the driver, which is frequently obscured by noisy signals dependent on the user's motor impairment. In this paper, we propose a multimodal teleoperation interface, incorporating map information, haptic feedback and user eye-gaze data, and examine which of these factors are most important for allowing accurate determination of user intention in a simulated tremor experiment. Our study indicates that, for expert assistants, presence of additional haptic and gaze information increases their ability to accurately infer the user's intention, providing supporting evidence for the utility of multimodal interfaces in remote assistance scenarios for Learning Assistance by Demonstration. Our study also reveals strong individual preferences on the different modalities, with large variations of performance occurring depending on whether supplemental eye-gaze or haptic information was given.
Vinícius Barbosa Schettino, Yiannis Demiris
IROS2
2019 Towards Explainable Shared Control using Augmented Reality
abstract
Shared control plays a pivotal role in establishing effective human-robot interactions. Traditional control-sharing methods strive to complement a human's capabilities at safely completing a task, and thereby rely on users forming a mental model of the expected robot behaviour. However, these methods can often bewilder or frustrate users whenever their actions do not elicit the intended system response, forming a misalignment between the respective internal models of the robot and human. To resolve this model misalignment, we introduce Explainable Shared Control as a paradigm in which assistance and information feedback are jointly considered. Augmented reality is presented as an integral component of this paradigm, by visually unveiling the robot's inner workings to human operators. Explainable Shared Control is instantiated and tested for assistive navigation in a setup involving a robotic wheelchair and a Microsoft HoloLens with add-on eye tracking. Experimental results indicate that the introduced paradigm facilitates transparent assistance by improving recovery times from adverse events associated with model misalignment.
Mark Zolotas, Yiannis Demiris
IROS2
2019 Probabilistic Real-Time User Posture Tracking for Personalized Robot-Assisted Dressing
abstract
Robotic solutions to dressing assistance have the potential to provide tremendous support for elderly and disabled people. However, unexpected user movements may lead to dressing failures or even pose a risk to the user. Tracking such user movements with vision sensors is challenging due to severe visual occlusions created by the robot and clothes. In this paper, we propose a probabilistic tracking method using Bayesian networks in latent spaces, which fuses robot end-effector positions and force information to enable cameraless and real-time estimation of the user postures during dressing. The latent spaces are created before dressing by modeling the user movements with a Gaussian process latent variable model, taking the user's movement limitations into account. We introduce a robot-assisted dressing system that combines our tracking method with hierarchical multitask control to minimize the force between the user and the robot. The experimental results demonstrate the robustness and accuracy of our tracking method. The proposed method enables the Baxter robot to provide personalized dressing assistance in putting on a sleeveless jacket for users with (simulated) upper-body impairments.
Fan Zhang 0030, Antoine Cully, Yiannis Demiris
IEEE Trans. Robotics3
2018 3D Motion Segmentation of Articulated Rigid Bodies based on RGB-D Data
Urbano Miguel Nunes, Yiannis Demiris
BMVC2
2018 Context-Aware Deep Feature Compression for High-Speed Visual Tracking
abstract
We propose a new context-aware correlation filter based tracking framework to achieve both high computational speed and state-of-the-art performance among real-time trackers. The major contribution to the high computational speed lies in the proposed deep feature compression that is achieved by a context-aware scheme utilizing multiple expert auto-encoders; a context in our framework refers to the coarse category of the tracking target according to appearance patterns. In the pre-training phase, one expert auto-encoder is trained per category. In the tracking phase, the best expert auto-encoder is selected for a given target, and only this auto-encoder is used. To achieve high tracking performance with the compressed feature map, we introduce extrinsic denoising processes and a new orthogonality loss term for pre-training and fine-tuning of the expert autoencoders. We validate the proposed context-aware framework through a number of experiments, where our method achieves a comparable performance to state-of-the-art trackers which cannot run in real-time, while running at a significantly fast speed of over 100 fps.
Jongwon Choi 0002, Hyung Jin Chang, Tobias Fischer 0001, Sangdoo Yun, Kyuewang Lee, Jiyeoup Jeong, Yiannis Demiris, Jin Young Choi 0002
CVPR7
2018 RT-GENE: Real-Time Eye Gaze Estimation in Natural Environments
Tobias Fischer 0001, Hyung Jin Chang, Yiannis Demiris
ECCV (10)3
2018 Hierarchical behavioral repertoires with unsupervised descriptors
abstract
Enabling artificial agents to automatically learn complex, versatile and high-performing behaviors is a long-lasting challenge. This paper presents a step in this direction with hierarchical behavioral repertoires that stack several behavioral repertoires to generate sophisticated behaviors. Each repertoire of this architecture uses the lower repertoires to create complex behaviors as sequences of simpler ones, while only the lowest repertoire directly controls the agent's movements. This paper also introduces a novel approach to automatically define behavioral descriptors thanks to an unsupervised neural network that organizes the produced high-level behaviors. The experiments show that the proposed architecture enables a robot to learn how to draw digits in an unsupervised manner after having learned to draw lines and arcs. Compared to traditional behavioral repertoires, the proposed architecture reduces the dimensionality of the optimization problems by orders of magnitude and provides behaviors with a twice better fitness. More importantly, it enables the transfer of knowledge between robots: a hierarchical repertoire evolved for a robotic arm to draw digits can be transferred to a humanoid robot by simply changing the lowest layer of the hierarchy. This enables the humanoid to draw digits although it has never been trained for this task.
Antoine Cully, Yiannis Demiris
GECCO2
2018 Augmented Reality for Feedback in a Shared Control Spraying Task
abstract
Using industrial robots to spray structures has been investigated extensively, however interesting challenges emerge when using handheld spraying robots. In previous work we have demonstrated the use of shared control of a handheld spraying robot to assist a user in a 3D spraying task. In this paper we demonstrate the use of Augmented Reality Interfaces to increase the user's progress and task awareness. We describe our solutions to challenging calibration issues between the Microsoft Hololens system and a motion capture system without the need for well defined markers or careful alignment on the part of the user. Error relative to the motion capture system was shown to be 10mm after only a 4 second calibration routine. Secondly we outline a logical approach for visualising liquid density for an augmented reality spraying task, this system allows the user to see target regions to complete, areas that are complete and areas that have been overdosed clearly. Finally we produced a user study to investigate the level of assistance that a handheld robot utilising shared control methods should provide during a spraying task. Using a handheld spraying robot with a moving spray head did not aid the user much over simply actuating spray nozzle for them. Compared to manual control the automatic modes significantly reduced the task load experienced by the user and significantly increased the quality of the result of the spraying task, reducing the error by 33-45%.
Joshua Elsdon, Yiannis Demiris
ICRA2
2018 Transferring Visuomotor Learning from Simulation to the Real World for Robotics Manipulation Tasks
abstract
Hand-eye coordination is a requirement for many manipulation tasks including grasping and reaching. However, accurate hand-eye coordination has shown to be especially difficult to achieve in complex robots like the iCub humanoid. In this work, we solve the hand-eye coordination task using a visuomotor deep neural network predictor that estimates the arm's joint configuration given a stereo image pair of the arm and the underlying head configuration. As there are various unavoidable sources of sensing error on the physical robot, we train the predictor on images obtained from simulation. The images from simulation were modified to look realistic using an image-to-image translation approach. In various experiments, we first show that the visuomotor predictor provides accurate joint estimates of the iCub's hand in simulation. We then show that the predictor can be used to obtain the systematic error of the robot's joint measurements on the physical iCub robot. We demonstrate that a calibrator can be designed to automatically compensate this error. Finally, we validate that this enables accurate reaching of objects while circumventing manual fine-calibration of the robot.
Phuong D. H. Nguyen, Tobias Fischer 0001, Hyung Jin Chang, Ugo Pattacini, Giorgio Metta, Yiannis Demiris
IROS6
2018 Real-Time Workload Classification during Driving using HyperNetworks
abstract
Classifying human cognitive states from behavioral and physiological signals is a challenging problem with important applications in robotics. The problem is challenging due to the data variability among individual users, and sensor artefacts. In this work, we propose an end-to-end framework for real-time cognitive workload classification with mixture Hyper Long Short Term Memory Networks (m-HyperLSTM), a novel variant of HyperNetworks. Evaluating the proposed approach on an eye-gaze pattern dataset collected from simulated driving scenarios of different cognitive demands, we show that the proposed framework outperforms previous baseline methods and achieves 83.9% precision and 87.8% recall during test. We also demonstrate the merit of our proposed architecture by showing improved performance over other LSTM-based methods.
Pierluigi Vito Amadori, Yiannis Demiris
IROS3
2018 Head-Mounted Augmented Reality for Explainable Robotic Wheelchair Assistance
abstract
Robotic wheelchairs with built-in assistive features, such as shared control, are an emerging means of providing independent mobility to severely disabled individuals. However, patients often struggle to build a mental model of their wheelchair's behaviour under different environmental conditions. Motivated by the desire to help users bridge this gap in perception, we propose a novel augmented reality system using a Microsoft Hololens as a head-mounted aid for wheelchair navigation. The system displays visual feedback to the wearer as a way of explaining the underlying dynamics of the wheelchair's shared controller and its predicted future states. To investigate the influence of different interface design options, a pilot study was also conducted. We evaluated the acceptance rate and learning curve of an immersive wheelchair training regime, revealing preliminary insights into the potential beneficial and adverse nature of different augmented reality cues for assistive navigation. In particular, we demonstrate that care should be taken in the presentation of information, with effort-reducing cues for augmented information acquisition (for example, a rear-view display) being the most appreciated.
Mark Zolotas, Joshua Elsdon, Yiannis Demiris
IROS3
2018 Highly Articulated Kinematic Structure Estimation Combining Motion and Skeleton Information
abstract
In this paper, we present a novel framework for unsupervised kinematic structure learning of complex articulated objects from a single-view 2D image sequence. In contrast to prior motion-based methods, which estimate relatively simple articulations, our method can generate arbitrarily complex kinematic structures with skeletal topology via a successive iterative merging strategy. The iterative merge process is guided by a density weighted skeleton map which is generated from a novel object boundary generation method from sparse 2D feature points. Our main contributions can be summarised as follows: (i) An unsupervised complex articulated kinematic structure estimation method that combines motion segments with skeleton information. (ii) An iterative fine-to-coarse merging strategy for adaptive motion segmentation and structural topology embedding. (iii) A skeleton estimation method based on a novel silhouette boundary generation from sparse feature points using an adaptive model selection method. (iv) A new highly articulated object dataset with ground truth annotation. We have verified the effectiveness of our proposed method in terms of computational time and estimation accuracy through rigorous experiments with multiple datasets. Our experiments show that the proposed method outperforms state-of-the-art methods both quantitatively and qualitatively.
Hyung Jin Chang, Yiannis Demiris
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Learning Kinematic Structure Correspondences Using Multi-Order Similarities
abstract
In this paper, we present a novel framework for finding the kinematic structure correspondences between two articulated objects in videos via hypergraph matching. In contrast to appearance and graph alignment based matching methods, which have been applied among two similar static images, the proposed method finds correspondences between two dynamic kinematic structures of heterogeneous objects in videos. Thus our method allows matching the structure of objects which have similar topologies or motions, or a combination of the two. Our main contributions can be summarised as follows: (i) casting the kinematic structure correspondence problem into a hypergraph matching problem by incorporating multi-order similarities with normalising weights, (ii) introducing a structural topology similarity measure by aggregating topology constrained subgraph isomorphisms, (iii) measuring kinematic correlations between pairwise nodes, and (iv) proposing a combinatorial local motion similarity measure using geodesic distance on the Riemannian manifold. We demonstrate the robustness and accuracy of our method through a number of experiments on synthetic and real data, outperforming various other state of the art methods. Our method is not limited to a specific application nor sensor, and can be used as building block in applications such as action recognition, human motion retargeting to robots, and articulated object manipulation amongst others.
Hyung Jin Chang, Tobias Fischer 0001, Maxime Petit, Martina Zambelli, Yiannis Demiris
IEEE Trans. Pattern Anal. Mach. Intell.5
2018 Quality and Diversity Optimization: A Unifying Modular Framework
abstract
The optimization of functions to find the best solution according to one or several objectives has a central role in many engineering and research fields. Recently, a new family of optimization algorithms, named quality-diversity (QD) optimization, has been introduced, and contrasts with classic algorithms. Instead of searching for a single solution, QD algorithms are searching for a large collection of both diverse and high-performing solutions. The role of this collection is to cover the range of possible solution types as much as possible, and to contain the best solution for each type. The contribution of this paper is threefold. First, we present a unifying framework of QD optimization algorithms that covers the two main algorithms of this family (multidimensional archive of phenotypic elites and the novelty search with local competition), and that highlights the large variety of variants that can be investigated within this family. Second, we propose algorithms with a new selection mechanism for QD algorithms that outperforms all the algorithms tested in this paper. Lastly, we present a new collection management that overcomes the erosion issues observed when using unstructured collections. These three contributions are supported by extensive experimental comparisons of QD algorithms on three different experimental scenarios.
Antoine Cully, Yiannis Demiris
IEEE Trans. Evol. Comput.2
2017 Attentional Correlation Filter Network for Adaptive Visual Tracking
abstract
We propose a new tracking framework with an attentional mechanism that chooses a subset of the associated correlation filters for increased robustness and computational efficiency. The subset of filters is adaptively selected by a deep attentional network according to the dynamic properties of the tracking target. Our contributions are manifold, and are summarised as follows: (i) Introducing the Attentional Correlation Filter Network which allows adaptive tracking of dynamic targets. (ii) Utilising an attentional network which shifts the attention to the best candidate modules, as well as predicting the estimated accuracy of currently inactive modules. (iii) Enlarging the variety of correlation filters which cover target drift, blurriness, occlusion, scale changes, and flexible aspect ratio. (iv) Validating the robustness and efficiency of the attentional mechanism for visual tracking through a number of experiments. Our method achieves similar performance to non real-time trackers, and state-of-the-art performance amongst real-time trackers.
Jongwon Choi 0002, Hyung Jin Chang, Sangdoo Yun, Tobias Fischer 0001, Yiannis Demiris, Jin Young Choi 0002
CVPR5
2017 Variational Autoencoded Regression: High Dimensional Regression of Visual Data on Complex Manifold
abstract
This paper proposes a new high dimensional regression method by merging Gaussian process regression into a variational autoencoder framework. In contrast to other regression methods, the proposed method focuses on the case where output responses are on a complex high dimensional manifold, such as images. Our contributions are summarized as follows: (i) A new regression method estimating high dimensional image responses, which is not handled by existing regression algorithms, is proposed. (ii) The proposed regression method introduces a strategy to learn the latent space as well as the encoder and decoder so that the result of the regressed response in the latent space coincide with the corresponding response in the data space. (iii) The proposed regression is embedded into a generative model, and the whole procedure is developed by the variational autoencoder framework. We demonstrate the robustness and effectiveness of our method through a number of experiments on various visual data regression problems.
Young Joon Yoo, Sangdoo Yun, Hyung Jin Chang, Yiannis Demiris, Jin Young Choi 0002
CVPR4
2017 Assisted painting of 3D structures using shared control with a hand-held robot
abstract
We present a shared control method of painting 3D geometries, using a handheld robot which has a single autonomously controlled degree of freedom. The user scans the robot near to the desired painting location, the single movement axis moves the spray head to achieve the required paint distribution. A simultaneous simulation of the spraying procedure is performed, giving an open loop approximation of the current state of the painting. An online prediction of the best path for the spray nozzle actuation is calculated in a receding horizon fashion. This is calculated by producing a map of the paint required in the 2D space defined by nozzle position on the gantry and the time into the future. A directed graph then extracts its edge weights from this paint density map and Dijkstra's algorithm is then used to find the candidate for the most effective path. Due to the heavy parallelisation of this approach and the majority of the calculations taking place on a GPU we can run the prediction loop in 32.6ms for a prediction horizon of 1 second, this approach is computationally efficient, outperforming a greedy algorithm. The path chosen by the proposed method on average chooses a path in the top 15% of all paths as calculated by exhaustive testing. This approach enables development of real time path planning for assisted spray painting onto complicated 3D geometries. This method could be applied to applications such as assistive painting for people with disabilities, or accurate placement of liquid when large scale positioning of the head is too expensive.
Joshua Elsdon, Yiannis Demiris
ICRA2
2017 Personalized robot-assisted dressing using user modeling in latent spaces
abstract
Robots have the potential to provide tremendous support to disabled and elderly people in their everyday tasks, such as dressing. Many recent studies on robotic dressing assistance usually view dressing as a trajectory planning problem. However, the user movements during the dressing process are rarely taken into account, which often leads to the failures of the planned trajectory and may put the user at risk. The main difficulty of taking user movements into account is caused by severe occlusions created by the robot, the user, and the clothes during the dressing process, which prevent vision sensors from accurately detecting the postures of the user in real time. In this paper, we address this problem by introducing an approach that allows the robot to automatically adapt its motion according to the force applied on the robot's gripper caused by user movements. There are two main contributions introduced in this paper: 1) the use of a hierarchical multi-task control strategy to automatically adapt the robot motion and minimize the force applied between the user and the robot caused by user movements; 2) the online update of the dressing trajectory based on the user movement limitations modeled with the Gaussian Process Latent Variable Model in a latent space, and the density information extracted from such latent space. The combination of these two contributions leads to a personalized dressing assistance that can cope with unpredicted user movements during the dressing while constantly minimizing the force that the robot may apply on the user. The experimental results demonstrate that the proposed method allows the Baxter humanoid robot to provide personalized dressing assistance for human users with simulated upper-body impairments.
Fan Zhang 0030, Antoine Cully, Yiannis Demiris
IROS3
2017 Multi-task and multi-kernel Gaussian process dynamical systems
Dimitrios Korkinof, Yiannis Demiris
Pattern Recognit.2
2017 Adaptive user modelling in car racing games using behavioural and physiological data
abstract
Personalised content adaptation has great potential to increase user engagement in video games. Procedural generation of user-tailored content increases the self-motivation of players as they immerse themselves in the virtual world. An adaptive user model is needed to capture the skills of the player and enable automatic game content altering algorithms to fit the individual user. We propose an adaptive user modelling approach using a combination of unobtrusive physiological data to identify strengths and weaknesses in user performance in car racing games. Our system creates user-tailored tracks to improve driving habits and user experience, and to keep engagement at high levels. The user modelling approach adopts concepts from the Trace Theory framework; it uses machine learning to extract features from the user’s physiological data and game-related actions, and cluster them into low level primitives. These primitives are transformed and evaluated into higher level abstractions such as experience , exploration and attention . These abstractions are subsequently used to provide track alteration decisions for the player. Collection of data and feedback from 52 users allowed us to associate key model variables and outcomes to user responses, and to verify that the model provides statistically significant decisions personalised to the individual player. Tailored game content variations between users in our experiments, as well as the correlations with user satisfaction demonstrate that our algorithm is able to automatically incorporate user feedback in subsequent procedural content generation.
Theodosis Georgiou, Yiannis Demiris
User Model. User Adapt. Interact.2
2016 Kinematic Structure Correspondences via Hypergraph Matching
abstract
In this paper, we present a novel framework for finding the kinematic structure correspondence between two objects in videos via hypergraph matching. In contrast to prior appearance and graph alignment based matching methods which have been applied among two similar static images, the proposed method finds correspondences between two dynamic kinematic structures of heterogeneous objects in videos. Our main contributions can be summarised as follows: (i) casting the kinematic structure correspondence problem into a hypergraph matching problem, incorporating multi-order similarities with normalising weights, (ii) a structural topology similarity measure by a new topology constrained subgraph isomorphism aggregation, (iii) a kinematic correlation measure between pairwise nodes, and (iv) a combinatorial local motion similarity measure using geodesic distance on the Riemannian manifold. We demonstrate the robustness and accuracy of our method through a number of experiments on complex articulated synthetic and real data.
Hyung Jin Chang, Tobias Fischer 0001, Maxime Petit, Martina Zambelli, Yiannis Demiris
CVPR5
2016 Visual Tracking Using Attention-Modulated Disintegration and Integration
abstract
In this paper, we present a novel attention-modulated visual tracking algorithm that decomposes an object into multiple cognitive units, and trains multiple elementary trackers in order to modulate the distribution of attention according to various feature and kernel types. In the integration stage it recombines the units to memorize and recognize the target object effectively. With respect to the elementary trackers, we present a novel attentional feature-based correlation filter (AtCF) that focuses on distinctive attentional features. The effectiveness of the proposed algorithm is validated through experimental comparison with state-of-theart methods on widely-used tracking benchmark datasets.
Jongwon Choi 0002, Hyung Jin Chang, Jiyeoup Jeong, Yiannis Demiris, Jin Young Choi 0002
CVPR4
2016 Markerless perspective taking for humanoid robots in unconstrained environments
abstract
Perspective taking enables humans to imagine the world from another viewpoint. This allows reasoning about the state of other agents, which in turn is used to more accurately predict their behavior. In this paper, we equip an iCub humanoid robot with the ability to perform visuospatial perspective taking (PT) using a single depth camera mounted above the robot. Our approach has the distinct benefit that the robot can be used in unconstrained environments, as opposed to previous works which employ marker-based motion capture systems. Prior to and during the PT, the iCub learns the environment, recognizes objects within the environment, and estimates the gaze of surrounding humans. We propose a new head pose estimation algorithm which shows a performance boost by normalizing the depth data to be aligned with the human head. Inspired by psychological studies, we employ two separate mechanisms for the two different types of PT. We implement line of sight tracing to determine whether an object is visible to the humans (level 1 PT). For more complex PT tasks (level 2 PT), the acquired point cloud is mentally rotated, which allows algorithms to reason as if the input data was acquired from an egocentric perspective. We show that this can be used to better judge where object are in relation to the humans. The multifaceted improvements to the PT pipeline advance the state of the art, and move PT in robots to markerless, unconstrained environments.
Tobias Fischer 0001, Yiannis Demiris
ICRA2
2016 Hierarchical action learning by instruction through interactive grounding of body parts and proto-actions
abstract
Learning by instruction allows humans programming a robot to achieve a task using spoken language, without the requirement of being able to do the task themselves, which can be problematic for users with motor impairments. We provide a developmental framework to program the humanoid robot iCub without any hand-coded a-priori knowledge about any motor skills. Inspired by child development theories, the system involves hierarchical learning, starting with the human verbally labelling robot body parts. The robot can then focus its attention on a precise body part during robot motor babbling, and link the on-the-fly spoken descriptions of proto-actions to angle values of a specific joint. The direct grounding of proto-actions is possible through the use of a linear model which calculates the effects on the joint of the proto-action and the body part used, allowing a generalisation of the proto-action if the joint has never been used before. Eventually, transferring the grounding is allowed via learning by instructions where humans can combine the newly acquired proto-actions to build primitives and more complex actions by scaffolding them. The framework has been validated using a humanoid robot iCub, which is able to learn without any prior knowledge: 1) the name of its fingers and the corresponding joint number, 2) how to fold and unfold them and 3) how to close or open its hand and how to show numbers with its fingers.
Maxime Petit, Yiannis Demiris
ICRA2
2016 Iterative path optimisation for personalised dressing assistance using vision and force information
abstract
We propose an online iterative path optimisation method to enable a Baxter humanoid robot to assist human users to dress. The robot searches for the optimal personalised dressing path using vision and force sensor information: vision information is used to recognise the human pose and model the movement space of upper-body joints; force sensor information is used for the robot to detect external force resistance and to locally adjust its motion. We propose a new stochastic path optimisation method based on adaptive moment estimation. We first compare the proposed method with other path optimisation algorithms on synthetic data. Experimental results show that the performance of the method achieves the smallest error with fewer iterations and less computation time. We also evaluate real-world data by enabling the Baxter robot to assist real human users with their dressing.
Yixing Gao 0001, Hyung Jin Chang, Yiannis Demiris
IROS3
2016 Multimodal imitation using self-learned sensorimotor representations
abstract
Although many tasks intrinsically involve multiple modalities, often only data from a single modality are used to improve complex robots acquisition of new skills. We present a method to equip robots with multimodal learning skills to achieve multimodal imitation on-the-fly on multiple concurrent task spaces, including vision, touch and proprioception, only using self-learned multimodal sensorimotor relations, without the need of solving inverse kinematic problems or explicit analytical models formulation. We evaluate the proposed method on a humanoid iCub robot learning to interact with a piano keyboard and imitating a human demonstration. Since no assumptions are made on the kinematic structure of the robot, the method can be also applied to different robotic platforms.
Martina Zambelli, Yiannis Demiris
IROS2
2016 Towards long-term social child-robot interaction: using multi-activity switching to engage young users
abstract
Social robots have the potential to provide support in a number of practical domains, such as learning and behaviour change. This potential is particularly relevant for children, who have proven receptive to interactions with social robots. To reach learning and therapeutic goals, a number of issues need to be investigated, notably the design of an effective child-robot interaction (cHRI) to ensure the child remains engaged in the relationship and that educational goals are met. Typically, current cHRI research experiments focus on a single type of interaction activity (e.g. a game). However, these can suffer from a lack of adaptation to the child, or from an increasingly repetitive nature of the activity and interaction. In this paper, we motivate and propose a practicable solution to this issue: an adaptive robot able to switch between multiple activities within single interactions. We describe a system that embodies this idea, and present a case study in which diabetic children collaboratively learn with the robot about various aspects of managing their condition. We demonstrate the ability of our system to induce a varied interaction and show the potential of this approach both as an educational tool and as a research method for long-term cHRI.
Miranda Coninx, Paul Baxter 0001, Elettra Oleari, Sara Bellini, Bert P. B. Bierman, Olivier A. Blanson Henkemans, Lola Cañamero, Piero Cosi, Valentin Enescu, Raquel Ros, Antoine Hiolle, Rémi Humbert, Bernd Kiefer, Ivana Kruijff-Korbayová, Rosemarijn Looije, Marco Mosconi, Mark A. Neerincx, Giulio Paci, Yorgos Patsis, Clara Pozzi, Francesca Sacchitelli, Hichem Sahli, Alberto Sanna, Giacomo Sommavilla, Fabio Tesser, Yiannis Demiris, Tony Belpaeme
J. Hum. Robot Interact.26
2015 Unsupervised learning of complex articulated kinematic structures combining motion and skeleton information
abstract
In this paper we present a novel framework for unsupervised kinematic structure learning of complex articulated objects from a single-view image sequence. In contrast to prior motion information based methods, which estimate relatively simple articulations, our method can generate arbitrarily complex kinematic structures with skeletal topology by a successive iterative merge process. The iterative merge process is guided by a skeleton distance function which is generated from a novel object boundary generation method from sparse points. Our main contributions can be summarised as follows: (i) Unsupervised complex articulated kinematic structure learning by combining motion and skeleton information. (ii) Iterative fine-to-coarse merging strategy for adaptive motion segmentation and structure smoothing. (iii) Skeleton estimation from sparse feature points. (iv) A new highly articulated object dataset containing multi-stage complexity with ground truth. Our experiments show that the proposed method out-performs state-of-the-art methods both quantitatively and qualitatively.
Hyung Jin Chang, Yiannis Demiris
CVPR2
2015 Encoderless position control of a two-link robot manipulator
abstract
Encoders have been an inseparable part of robots since the very beginning of modern robotics in the 1950s. As a result, the foundations of robot control are built on the concepts of kinematics and dynamics of articulated rigid bodies, which rely on explicitly measuring the robot configuration in terms of joint angles - done by encoders. In this paper, we propose a radically new concept for controlling robots called Encoderless Robot Control (EnRoCo). The concept is based on our hypothesis that it is possible to control a robot without explicitly measuring its joint angles, by measuring instead the effects of the actuation on its end-effector. To prove the feasibility of this unconventional control approach, we propose a proof-of-concept control algorithm for encoderless position control of a robot's end-effector in task space. We demonstrate a prototype implementation of this controller in a dynamics simulation of a two-link robot manipulator. The prototype controller is able to successfully control the robot's end-effector to reach a reference position, as well as to track continuously a desired trajectory. Notably, we demonstrate how this novel controller can cope with something that traditional control approaches fail to do: adapt on-the-fly to changes in the kinematics of the robot, such as changing the lengths of the links.
Petar Kormushev, Yiannis Demiris, Darwin G. Caldwell
ICRA2
2015 User modelling for personalised dressing assistance by humanoid robots
abstract
Assistive robots can improve the well-being of disabled or frail human users by reducing the burden that activities of daily living impose on them. To enable personalised assistance, such robots benefit from building a user-specific model, so that the assistance is customised to the particular set of user abilities. In this paper, we present an end-to-end approach for home-environment assistive humanoid robots to provide personalised assistance through a dressing application for users who have upper-body movement limitations. We use randomised decision forests to estimate the upper-body pose of users captured by a top-view depth camera, and model the movement space of upper-body joints using Gaussian mixture models. The movement space of each upper-body joint consists of regions with different reaching capabilities. We propose a method which is based on real-time upper-body pose and user models to plan robot motions for assistive dressing. We validate each part of our approach and test the whole system, allowing a Baxter humanoid robot to assist human to wear a sleeveless jacket.
Yixing Gao 0001, Hyung Jin Chang, Yiannis Demiris
IROS3
2015 Kinematic-free position control of a 2-DOF planar robot arm
abstract
This paper challenges the well-established assumption in robotics that in order to control a robot it is necessary to know its kinematic information, that is, the arrangement of links and joints, the link dimensions and the joint positions. We propose a kinematic-free robot control concept that does not require any prior kinematic knowledge. The concept is based on our hypothesis that it is possible to control a robot without explicitly measuring its joint angles, by measuring instead the effects of the actuation on its end-effector. We implement a proof-of-concept encoderless robot controller and apply it for the position control of a physical 2-DOF planar robot arm. The prototype controller is able to successfully control the robot to reach a reference position, as well as to track a continuous reference trajectory. Notably, we demonstrate how this novel controller can cope with something that traditional control approaches fail to do: adapt to drastic kinematic changes such as 100% elongation of a link, 35-degree angular offset of a joint, and even a complete overhaul of the kinematics involving the addition of new joints and links.
Petar Kormushev, Yiannis Demiris, Darwin G. Caldwell
IROS2
2015 Predicting car states through learned models of vehicle dynamics and user behaviours
abstract
The ability to predict forthcoming car states is crucial for the development of smart assistance systems. Forthcoming car states do not only depend on vehicle dynamics but also on user behaviour. In this paper, we describe a novel prediction methodology by combining information from both sources - vehicle and user - using Gaussian Processes. We then apply this method in the context of high speed car racing. Results show that the forthcoming position and speed of the car can be predicted with low Root Mean Square Error through the trained model.
Theodosis Georgiou, Yiannis Demiris
Intelligent Vehicles Symposium2
2015 One-shot assistance estimation from expert demonstrations for a shared control wheelchair system
abstract
An emerging research problem in the field of assistive robotics is the design of methodologies that allow robots to provide human-like assistance to the users. Especially within the rehabilitation domain, a grand challenge is to program a robot to mimic the operation of an occupational therapist, intervening with the user when necessary so as to improve the therapeutic power of the assistive robotic system. We propose a method to estimate assistance policies from expert demonstrations to present human-like intervention during navigation in a powered wheelchair setup. For this purpose, we constructed a setting, where a human offers assistance to the user over a haptic shared control system. The robot learns from human assistance demonstrations while the user is actively driving the wheelchair in an unconstrained environment. We train a Gaussian process regression model to learn assistance commands given past and current actions of the user and the state of the environment. The results indicate that the model can estimate human assistance after only a single demonstration, i.e. in one-shot, so that the robot can help the user by selecting the appropriate assistance in a human-like fashion.
Ayse Küçükyilmaz, Yiannis Demiris
RO-MAN2
2015 Towards a synchronised Grammars framework for adaptive musical human-robot collaboration
abstract
We present an adaptive musical collaboration framework for interaction between a human and a robot. The aim of our work is to develop a system that receives feedback from the user in real time and learns the music progression style of the user over time. To tackle this problem, we represent a song as a hierarchically structured sequence of music primitives. By exploiting the sequential constraints of these primitives inferred from the structural information combined with user feedback, we show that a robot can play music in accordance with the user's anticipated actions. We use Stochastic Context-Free Grammars augmented with the knowledge of the learnt user's preferences. We provide synthetic experiments as well as a pilot study with a Baxter robot and a tangible music table. The synthetic results show the synchronisation and adaptivity features of our framework and the pilot study suggest these are applicable to create an effective musical collaboration experience.
Miguel Sarabia, Kyuhwa Lee, Yiannis Demiris
RO-MAN3
2015 Learning assistance by demonstration: smart mobility with shared control and paired haptic controllers
abstract
In this paper, we present a framework, probabilistic model, and algorithm for learning shared control policies by observing an assistant. This is a methodology we refer to as Learning Assistance by Demonstration (LAD). As a subset of robot Learning by Demonstration (LbD), LAD focuses on the assistive element by explicitly capturing how and when to help. The latter is especially important in assistive scenarios---such as rehabilitation and training---where there exists multiple and possibly conflicting goals. We formalize these notions in a probabilistic model and develop an efficient online mixture of experts (OME) algorithm, based on sparse Gaussian processes (GPs), for learning the assistive policy. Focusing on smart mobility, we couple the LAD methodology with a novel paired-haptic-controllers setup for helping smart wheelchair users navigate their environment. Experimental results with 15 able-bodied participants demonstrate that our learned shared control policy improved driving performance (as measured in lap seconds) by 43 s (a speedup of 191%). Furthermore, survey results indicate that the participants not only performed better quantitatively, but also qualitatively felt the model assistance helped them complete the task.
Harold Soh, Yiannis Demiris
J. Hum. Robot Interact.2
2015 STARE: Spatio-Temporal Attention Relocation for Multiple Structured Activities Detection
abstract
We present a spatio-temporal attention relocation (STARE) method, an information-theoretic approach for efficient detection of simultaneously occurring structured activities. Given multiple human activities in a scene, our method dynamically focuses on the currently most informative activity. Each activity can be detected without complete observation, as the structure of sequential actions plays an important role on making the system robust to unattended observations. For such systems, the ability to decide where and when to focus is crucial to achieving high detection performances under resource bounded condition. Our main contributions can be summarized as follows: 1) information-theoretic dynamic attention relocation framework that allows the detection of multiple activities efficiently by exploiting the activity structure information and 2) a new high-resolution data set of temporally-structured concurrent activities. Our experiments on applications show that the STARE method performs efficiently while maintaining a reasonable level of accuracy.
Kyuhwa Lee, Dimitri Ognibene, Hyung Jin Chang, Tae-Kyun Kim 0001, Yiannis Demiris
IEEE Trans. Image Process.5
2015 Spatio-Temporal Learning With the Online Finite and Infinite Echo-State Gaussian Processes
abstract
Successful biological systems adapt to change. In this paper, we are principally concerned with adaptive systems that operate in environments where data arrives sequentially and is multivariate in nature, for example, sensory streams in robotic systems. We contribute two reservoir inspired methods: 1) the online echostate Gaussian process (OESGP) and 2) its infinite variant, the online infinite echostate Gaussian process (OIESGP) Both algorithms are iterative fixed-budget methods that learn from noisy time series. In particular, the OESGP combines the echo-state network with Bayesian online learning for Gaussian processes. Extending this to infinite reservoirs yields the OIESGP, which uses a novel recursive kernel with automatic relevance determination that enables spatial and temporal feature weighting. When fused with stochastic natural gradient descent, the kernel hyperparameters are iteratively adapted to better model the target system. Furthermore, insights into the underlying system can be gleamed from inspection of the resulting hyperparameters. Experiments on noisy benchmark problems (one-step prediction and system identification) demonstrate that our methods yield high accuracies relative to state-of-the-art methods, and standard kernels with sliding windows, particularly on problems with irrelevant dimensions. In addition, we describe two case studies in robotic learning-by-demonstration involving the Nao humanoid robot and the Assistive Robot Transport for Youngsters (ARTY) smart wheelchair.
Harold Soh, Yiannis Demiris
IEEE Trans. Neural Networks Learn. Syst.2
2014 Behavioral accommodation towards a dance robot tutor
abstract
We report first results on children adaptive behavior towards a dance tutoring robot. We can observe that children behavior rapidly evolves through few sessions in order to accommodate with the robotic tutor rhythm and instructions.
Raquel Ros, Miranda Coninx, Yiannis Demiris, Yorgos Patsis, Valentin Enescu, Hichem Sahli
HRI3
2014 Increasing the accuracy and the repeatability of position control for micromanipulations using Heteroscedastic Gaussian Processes
abstract
Many recent studies describe micromanipulation systems by using complex Analytic Forward Models (AFM), but such models are difficult to build and incapable of describing unmodelable factors, such as manufacturing defects. In this work, we propose the Enhanced Analytic Forward Model (EAFM), an integrated model of the AFM and the Heteroscedastic Gaussian Processes (HGP). The EAFM can compensate the shortfalls of the AFM by training the HGP on the residual of the AFM. This also allows the HGP to learn the repeatability of the micromanipulation system. Based on the EAFM, we further contribute an optimal position controller for improving the accuracy and the repeatability. This optimal EAFM controller is implemented and tested on a three degree-of-freedom micromanipulator based micromanipulation system. Two sets of real-world experiments are carried out to verify our method. The results demonstrate that the controller using EAFM can statistically achieve higher accuracy and repeatability than solely using the AFM.
Yanyu Su, Wei Dong 0004, Yan Wu 0002, Zhijiang Du, Yiannis Demiris
ICRA5
2013 Towards Active Event Recognition
Dimitri Ognibene, Yiannis Demiris
IJCAI2
2013 Online quantum mixture regression for trajectory learning by demonstration
abstract
In this work, we present the online Quantum Mixture Model (oQMM), which combines the merits of quantum mechanics and stochastic optimization. More specifically it allows for quantum effects on the mixture states, which in turn become a superposition of conventional mixture states. We propose an efficient stochastic online learning algorithm based on the online Expectation Maximization (EM), as well as a generation and decay scheme for model components. Our method is suitable for complex robotic applications, where data is abundant or where we wish to iteratively refine our model and conduct predictions during the course of learning. With a synthetic example, we show that the algorithm can achieve higher numerical stability. We also empirically demonstrate the efficacy of our method in well-known regression benchmark datasets. Under a trajectory Learning by Demonstration setting we employ a multi-shot learning application in joint angle space, where we observe higher quality of learning and reproduction. We compare against popular and well-established methods, widely adopted across the robotics community.
Dimitrios Korkinof, Yiannis Demiris
IROS2
2013 When and how to help: An iterative probabilistic model for learning assistance by demonstration
abstract
Crafting a proper assistance policy is a difficult endeavour but essential for the development of robotic assistants. Indeed, assistance is a complex issue that depends not only on the task-at-hand, but also on the state of the user, environment and competing objectives. As a way forward, this paper proposes learning the task of assistance through observation; an approach we term Learning Assistance by Demonstration (LAD). Our methodology is a subclass of Learning-by-Demonstration (LbD), yet directly addresses difficult issues associated with proper assistance such as when and how to appropriately assist. To learn assistive policies, we develop a probabilistic model that explicitly captures these elements and provide efficient, online, training methods. Experimental results on smart mobility assistance - using both simulation and a real-world smart wheelchair platform - demonstrate the effectiveness of our approach; the LAD model quickly learns when to assist (achieving an AUC score of 0.95 after only one demonstration) and improves with additional examples. Results show that this translates into better task-performance; our LAD-enabled smart wheelchair improved participant driving performance (measured in lap seconds) by 20.6s (a speedup of 137%), after a single teacher demonstration.
Harold Soh, Yiannis Demiris
IROS2
2013 Enhanced kinematic model for dexterous manipulation with an underactuated hand
abstract
Recent studies on underactuated manipulation usually describe the system with a Kinematic Model (KM), which is built by adding external constraints to the standard manipulation analysis method. However, such external constraints are easily violated in a real-world dexterous manipulation task which results in significant control errors. In this work, the Enhanced Kinematic Model (E-KM), an integrated model of the KM and the Sparse Online Gaussian Process (SOGP) is proposed. The E-KM can compensate the shortfalls of the KM by on-the-fly training the SOGP on the residual between the prediction of the KM and the ground truth data. Based on the E-KM, we further contribute an optimal controller for underactuated manipulations. This optimal E-KM controller is implemented and tested on the iCub, a humanoid robot with two anthropomorphic underactuated hands. Two sets of real-world experiments are carried out to verify our method. The results demonstrate that the controller using E-KM statistically can achieve higher control accuracy than using solely using the KM for a wide range of objects.
Yanyu Su, Yan Wu 0002, Harold Soh, Zhijiang Du, Yiannis Demiris
IROS5
2013 The Infinite-Order Conditional Random Field Model for Sequential Data Modeling
abstract
Sequential data labeling is a fundamental task in machine learning applications, with speech and natural language processing, activity recognition in video sequences, and biomedical data analysis being characteristic examples, to name just a few. The conditional random field (CRF), a log-linear model representing the conditional distribution of the observation labels, is one of the most successful approaches for sequential data labeling and classification, and has lately received significant attention in machine learning as it achieves superb prediction performance in a variety of scenarios. Nevertheless, existing CRF formulations can capture only one- or few-timestep interactions and neglect higher order dependences, which are potentially useful in many real-life sequential data modeling applications. To resolve these issues, in this paper we introduce a novel CRF formulation, based on the postulation of an energy function which entails infinitely long time-dependences between the modeled data. Building blocks of our novel approach are: 1) the sequence memoizer (SM), a recently proposed nonparametric Bayesian approach for modeling label sequences with infinitely long time dependences, and 2) a mean-field-like approximation of the model marginal likelihood, which allows for the derivation of computationally efficient inference algorithms for our model. The efficacy of the so-obtained infinite-order CRF (CRF(∞)) model is experimentally demonstrated.
Sotirios Chatzis, Yiannis Demiris
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Multimodal child-robot interaction: building social bonds
Tony Belpaeme, Paul Baxter 0001, Robin Read, Rachel Wood, Heriberto Cuayáhuitl, Bernd Kiefer, Stefania Racioppa, Ivana Kruijff-Korbayová, Georgios Athanasopoulos, Valentin Enescu, Rosemarijn Looije, Mark A. Neerincx, Yiannis Demiris, Raquel Ros, Aryel Beck, Lola Cañamero, Antoine Hiolle, Matthew Lewis 0001, Ilaria Baroni, Marco Nalin, Piero Cosi, Giulio Paci, Fabio Tesser, Giacomo Sommavilla, Rémi Humbert
J. Hum. Robot Interact.13
2012 Learning action symbols for hierarchical grammar induction
Kyuhwa Lee, Tae-Kyun Kim 0001, Yiannis Demiris
ICPR3
2012 Learning reusable task components using hierarchical activity grammars with uncertainties
abstract
We present a novel learning method using activity grammars capable of learning reusable task components from a reasonably small number of samples under noisy conditions. Our linguistic approach aims to extract the hierarchical structure of activities which can be recursively applied to help recognize unforeseen, more complicated tasks that share the same underlying structures. To achieve this goal, our method 1) actively searches for frequently occurring action symbols that are subset of input samples to effectively discover the hierarchy, and 2) explicitly takes into account the uncertainty values associated with input symbols due to the noise inherent in low-level detectors. In addition to experimenting with a synthetic dataset to systematically analyze the algorithm's performance, we apply our method in human-led imitation learning environment where a robot learns reusable components of the task from short demonstrations to correctly imitate more complicated, longer demonstrations of the same task category. The results suggest that under reasonable amount of noise, our method is capable to capture the reusable structures of tasks and generalize to cope with recursions.
Kyuhwa Lee, Tae-Kyun Kim 0001, Yiannis Demiris
ICRA3
2012 Iterative temporal learning and prediction with the sparse online echo state gaussian process
abstract
In this work, we contribute the online echo state gaussian process (OESGP), a novel Bayesian-based online method that is capable of iteratively learning complex temporal dynamics and producing predictive distributions (instead of point predictions). Our method can be seen as a combination of the echo state network with a sparse approximation of Gaussian processes (GPs). Extensive experiments on the one-step prediction task on well-known benchmark problems show that OESGP produced statistically superior results to current online ESNs and state-of-the-art regression methods. In addition, we characterise the benefits (and drawbacks) associated with the considered online methods, specifically with regards to the trade-off between computational cost and accuracy. For a high-dimensional action recognition task, we demonstrate that OESGP produces high accuracies comparable to a recently published graphical model, while being fast enough for real-time interactive scenarios.
Harold Soh, Yiannis Demiris
IJCNN2
2012 Online spatio-temporal Gaussian process experts with application to tactile classification
abstract
In this work, we are primarily concerned with robotic systems that learn online and continuously from multi-variate data-streams. Our first contribution is a new recursive kernel, which we have integrated into a sparse Gaussian Process to yield the Spatio-Temporal Online Recursive Kernel Gaussian Process (STORK-GP). This algorithm iteratively learns from time-series, providing both predictions and uncertainty estimates. Experiments on benchmarks demonstrate that our method achieves high accuracies relative to state-of-the-art methods. Second, we contribute an online tactile classifier which uses an array of STORK-GP experts. In contrast to existing work, our classifier is capable of learning new objects as they are presented, improving itself over time. We show that our approach yields results comparable to highly-optimised offline classification methods. Moreover, we conducted experiments with human subjects in a similar online setting with true-label feedback and present the insights gained.
Harold Soh, Yanyu Su, Yiannis Demiris
IROS3
2012 A sparse nonparametric hierarchical Bayesian approach towards inductive transfer for preference modeling
Sotirios Chatzis, Yiannis Demiris
Expert Syst. Appl.2
2012 The echo state conditional random field model for sequential data modeling
Sotirios Chatzis, Yiannis Demiris
Expert Syst. Appl.2
2012 A spatially-constrained normalized Gamma process prior
Sotirios Chatzis, Dimitrios Korkinof, Yiannis Demiris
Expert Syst. Appl.3
2012 The copula echo state network
Sotirios Chatzis, Yiannis Demiris
Pattern Recognit.2
2012 A reservoir-driven non-stationary hidden Markov model
Sotirios Chatzis, Yiannis Demiris
Pattern Recognit.2
2012 Nonparametric Mixtures of Gaussian Processes With Power-Law Behavior
abstract
Gaussian processes (GPs) constitute one of the most important Bayesian machine learning approaches, based on a particularly effective method for placing a prior distribution over the space of regression functions. Several researchers have considered postulating mixtures of GPs as a means of dealing with nonstationary covariance functions, discontinuities, multimodality, and overlapping output signals. In existing works, mixtures of GPs are based on the introduction of a gating function defined over the space of model input variables. This way, each postulated mixture component GP is effectively restricted in a limited subset of the input space. In this paper, we follow a different approach. We consider a fully generative nonparametric Bayesian model with power-law behavior, generating GPs over the whole input space of the learned task. We provide an efficient algorithm for model inference, based on the variational Bayesian framework, and prove its efficacy using benchmark and real-world datasets.
Sotirios Chatzis, Yiannis Demiris
IEEE Trans. Neural Networks Learn. Syst.2
2012 A Quantum-Statistical Approach Toward Robot Learning by Demonstration
abstract
Statistical machine learning approaches have been at the epicenter of the ongoing research work in the field of robot learning by demonstration over the past few years. One of the most successful methodologies used for this purpose is a Gaussian mixture regression (GMR). In this paper, we propose an extension of GMR-based learning by demonstration models to incorporate concepts from the field of quantum mechanics. Indeed, conventional GMR models are formulated under the notion that all the observed data points can be assigned to a distinct number of model states (mixture components). In this paper, we reformulate GMR models, introducing some quantum states constructed by superposing conventional GMR states by means of linear combinations. The so-obtained quantum statistics-inspired mixture regression algorithm is subsequently applied to obtain a novel robot learning by demonstration methodology, offering a significantly increased quality of regenerated trajectories for computational costs comparable with currently state-of-the-art trajectory-based robot learning by demonstration approaches. We experimentally demonstrate the efficacy of the proposed approach.
Sotirios Chatzis, Dimitrios Korkinof, Yiannis Demiris
IEEE Trans. Robotics3
2012 Collaborative Control for a Robotic Wheelchair: Evaluation of Performance, Attention, and Workload
abstract
Powered wheelchair users often struggle to drive safely and effectively and, in more critical cases, can only get around when accompanied by an assistant. To address these issues, we propose a collaborative control mechanism that assists users as and when they require help. The system uses a multiple-hypothesis method to predict the driver's intentions and, if necessary, adjusts the control signals to achieve the desired goal safely. The main emphasis of this paper is on a comprehensive evaluation, where we not only look at the system performance but also, perhaps more importantly, characterize the user performance in an experiment that combines eye tracking with a secondary task. Without assistance, participants experienced multiple collisions while driving around the predefined route. Conversely, when they were assisted by the collaborative controller, not only did they drive more safely but also they were able to pay less attention to their driving, resulting in a reduced cognitive workload. We discuss the importance of these results and their implications for other applications of shared control, such as brain-machine interfaces, where it could be used to compensate for both the low frequency and the low resolution of the user input.
Tom Carlson, Yiannis Demiris
IEEE Trans. Syst. Man Cybern. Part B2
2011 Evolving policies for multi-reward partially observable markov decision processes (MR-POMDPs)
abstract
Plans and decisions in many real-world scenarios are made under uncertainty and to satisfy multiple, possibly conflicting, objectives. In this work, we contribute the multi-reward partially-observable Markov decision process (MR-POMDP) as a general modelling framework. To solve MR-POMDPs, we present two hybrid (memetic) multi-objective evolutionary algorithms that generate non-dominated sets of policies (in the form of stochastic finite state controllers). Performance comparisons between the methods on multi-objective problems in robotics (with 2, 3 and 5 objectives), web-advertising (with 3, 4 and 5 objectives) and infectious disease control (with 3 objectives), revealed that memetic variants outperformed their original counterparts. We anticipate that the MR-POMDP along with multi-objective evolutionary solvers will prove useful in a variety of theoretical and real-world applications.
Harold Soh, Yiannis Demiris
GECCO2
2011 Adapting robot behavior to user's capabilities: a dance instruction study
abstract
The ALIZ-E1 project's goal is to design a robot companion able to maintain affective interactions with young users over a period of time. One of these interactions consists in teaching a dance to hospitalized children according to their capabilities. We propose a methodology for adapting both, the movements used in the dance based on the user's cognitive and physical capabilities through a set of metrics, and the robot's interaction based on the user's personality traits.
Raquel Ros, Ilaria Baroni, Marco Nalin, Yiannis Demiris
HRI4
2011 Child-robot interaction in the wild: advice to the aspiring experimenter
abstract
We present insights gleaned from a series of child-robot interaction experiments carried out in a hospital paediatric department. Our aim here is to share good practice in experimental design and lessons learned about the implementation of systems for social HRI with child users towards application in "the wild", rather than in tightly controlled and constrained laboratory environments: a trade-off between the structures imposed by experimental design and the desire for removal of such constraints that inhibit interaction depth, and hence engagement, requires a careful balance.
Raquel Ros, Marco Nalin, Rachel Wood, Paul Baxter 0001, Rosemarijn Looije, Yiannis Demiris, Tony Belpaeme, Alessio Giusti, Clara Pozzi
ICMI6
2011 The One-Hidden Layer Non-parametric Bayesian Kernel Machine
abstract
In this paper, we present a nonparametric Bayesian approach towards one-hidden-layer feed forward neural networks. Our approach is based on a random selection of the weights of the synapses between the input and the hidden layer neurons, and a Bayesian marginalization over the weights of the connections between the hidden layer neurons and the output neurons, giving rise to a kernel-based nonparametric Bayesian inference procedure for feed forward neural networks. Compared to existing approaches, our method presents a number of advantages, with the most significant being: (i) it offers a significant improvement in terms of the obtained generalization capabilities, (ii) being a nonparametric Bayesian learning approach, it entails inference instead of fitting to data, thus resolving the over fitting issues of non-Bayesian approaches, and (iii) it yields a full predictive posterior distribution, thus naturally providing a measure of uncertainty on the generated predictions (expressed by means of the variance of the predictive distribution), without the need of applying computationally intensive methods, e.g., bootstrap. We exhibit the merits of our approach by investigating its application to two difficult multimedia content classification applications: semantic characterization of audio scenes based on content, and yearly song classification, as well as a set of benchmark classification and regression tasks.
Sotirios Chatzis, Dimitrios Korkinof, Yiannis Demiris
ICTAI3
2011 Generalising human demonstration data by identifying affordance symmetries in object interaction trajectories
abstract
This paper concerns modelling human hand or tool trajectories when interacting with everyday objects. In these interactions symmetries may be exhibited in portions of the trajectories which can be used to identify task space redundancy. This paper presents a formal description of a set of these symmetries, which we term affordance symmetries, and a method to identify them in multiple demonstration recordings. The approach is robust to arbitrary motion before and after the symmetry artifact and relies only on recorded trajectory data. To illustrate the method's performance two examples are discussed involving two different types of symmetries. An simple illustration of the application of the concept in reproduction planning is also provided.
Jonathan Claassens, Yiannis Demiris
IROS2
2011 Echo State Gaussian Process
abstract
Echo state networks (ESNs) constitute a novel approach to recurrent neural network (RNN) training, with an RNN (the reservoir) being generated randomly, and only a readout being trained using a simple computationally efficient algorithm. ESNs have greatly facilitated the practical application of RNNs, outperforming classical approaches on a number of benchmark tasks. In this paper, we introduce a novel Bayesian approach toward ESNs, the echo state Gaussian process (ESGP). The ESGP combines the merits of ESNs and Gaussian processes to provide a more robust alternative to conventional reservoir computing networks while also offering a measure of confidence on the generated predictions (in the form of a predictive distribution). We exhibit the merits of our approach in a number of applications, considering both benchmark datasets and real-world applications, where we show that our method offers a significant enhancement in the dynamical data modeling capabilities of ESNs. Additionally, we also show that our method is orders of magnitude more computationally efficient compared to existing Gaussian process-based methods for dynamical data modeling, without compromises in the obtained predictive performance.
Sotirios Chatzis, Yiannis Demiris
IEEE Trans. Neural Networks2
2010 Hierarchical learning approach for one-shot action imitation in humanoid robots
abstract
We consider the issue of segmenting an action in the learning phase into a logical set of smaller primitives in order to construct a generative model for imitation learning using a hierarchical approach. Our proposed framework, addressing the “how-to” question in imitation, is based on a one-shot imitation learning algorithm. It incorporates segmentation of a demonstrated template into a series of subactions and takes a hierarchical approach to generate the task action by using a finite state machine in a generative way. Two sets of experiments have been conducted to evaluate the performance of the framework, both statistically and in practice, through playing a tic-tac-toe game. The experiments demonstrate that the proposed framework can effectively improve the performance of the one-shot learning algorithm and reduce the size of primitive space, without compromising the learning quality.
Yan Wu 0002, Yiannis Demiris
ICARCV2
2010 Increasing robotic wheelchair safety with collaborative control: Evidence from secondary task experiments
abstract
Powered wheelchairs play a vital role in bringing independence to the severely mobility-impaired. Our robotic wheelchair aims to assist users in driving safely, without undermining their capabilities or curtailing the natural development of their skills. An important research question is to determine the conditions under which shared control is most beneficial. In this paper, we describe an experiment, where a distracting secondary task caused the majority of participants to crash the wheelchair when driving without assistance. However, when they were assisted by our collaborative controller, not only did they drive safely, but they also increased their performance in the secondary task. We demonstrate that a degree of shared control is beneficial even to proficient drivers under certain circumstances, for instance when they are under a heightened workload.
Tom Carlson, Yiannis Demiris
ICRA2
2010 Towards One Shot Learning by imitation for humanoid robots
abstract
Teaching a robot to learn new knowledge is a repetitive and tedious process. In order to accelerate the process, we propose a novel template-based approach for robot arm movement imitation. This algorithm selects a previously observed path demonstrated by a human and generates a path in a novel situation based on pairwise mapping of invariant feature locations present in both the demonstrated and the new scenes using a combination of minimum distortion and minimum energy strategies. This One-Shot Learning algorithm is capable of not only mapping simple point-to-point paths but also adapting to more complex tasks such as those involving forced waypoints. As compared to traditional methodologies, our work require neither extensive training for generalisation nor expensive run-time computation for accuracy. This algorithm has been statistically validated using cross-validation of grasping experiments as well as tested for practical implementation on the iCub humanoid robot for playing the tic-tac-toe game.
Yan Wu 0002, Yiannis Demiris
ICRA2
2010 Spectral clustering in multi-agent systems
Bálint Takács, Yiannis Demiris
Knowl. Inf. Syst.2
2009 Multi-robot plan adaptation by constrained minimal distortion feature mapping
abstract
We propose a novel method for multi-robot plan adaptation which can be used for adapting existing spatial plans of robotic teams to new environments or imitating collaborative spatial teamwork of robots in novel situations. The algorithm selects correspondences between previous and current spatial features by the application of pairwise constraints, and generates the transformation function with a fast regular grid approximation which minimizes distortion. The algorithm requires minimal domain knowledge, is capable of transforming the spatial aspects of collaborative team behavior and performs better in noisy problems with large displacements than the most generally used quadratic differences method. The algorithm can be utilized for rapid plan adaptation, plan generalization or team behavior imitation. Methods are demonstrated on a multi-robot control problem in a random environment.
Bálint Takács, Yiannis Demiris
ICRA2
2009 A Groovy Virtual Drumming Agent
Axel Tidemann, Pinar Öztürk, Yiannis Demiris
IVA3
2008 Groovy Neural Networks
abstract
The drum machine has been an important tool in music production for decades. However, its flawless way of playing drum patterns is often perceived as mechanical and rigid, far from the groove provided by a human drummer. This paper presents research towards enhancing the drum machine with learning capabilities. The drum machine learns user-specific variations (i.e. the groove) from human drummers, and stores the groove as attractors in Echo State Networks (ESNs). The ESNs are purely generative (i.e. not driven by an input signal) and the output is used by the drum machine to imitate the playing style of human drummers, making it a cost-effective way of achieving life-like drums.
Axel Tidemann, Yiannis Demiris
ECAI2
2008 Balancing Spectral Clustering for Segmenting Spatio-temporal Observations of Multi-agent Systems
abstract
We examine the application of spectral clustering for breaking up the behavior of a multi-agent system in space and time into smaller, independent elements. We cluster observations of individualentities in order to identify significant changes in the parameter space (like spatial position)and detect temporal alterations of behavior within the same framework. Data is also influenced byknowledge about important events. Clusters are pre-processed at each step of the iterative subdivision to make the algorithm invariant against spatial scaling, rotation, replay speed andvarying sampling frequency. A method is presented to balance spatial and temporal segmentation based on the expected group size. We demonstrate our results by analyzing the outcomes of acomputer game.
Bálint Takács, Yiannis Demiris
ICDM2
2008 Human-wheelchair collaboration through prediction of intention and adaptive assistance
abstract
Powered wheelchair users want to be active drivers, not just passengers. However, in some situations (varying from person to person), they may require assistance; hence, research is being carried out into the development of 'smart' wheelchairs. Predominantly, this research has been derived from the field of mobile robotics, focussing on creating autonomous systems, which unfortunately tend to treat the human as little more than a precious piece of cargo. Instead, the design should be based around each individual user's abilities and desires, maximising the amount of control they are given. In this paper, we look at how collaborative control techniques can be used to achieve this, offering the user help, as and when it is required. We then evaluate the effects of this collaboration, which is built by predicting user intentions and responding to these predictions with adaptable levels of assistance.
Tom Carlson, Yiannis Demiris
ICRA2
2007 Special Issue on Robot Learning by Observation, Demonstration, and Imitation
abstract
This special issue contains selected extended contributions from both the Adaptation in Artificial and Biological Systems symposium held in Hertforshire in 2006 and the wider academic community following a public call for papers in 2006. The papers presented serve as a good illustration of the challenges faced by robotics researchers today in the field of programming by observation, demonstration, and imitation.
Yiannis Demiris, Aude Billard
IEEE Trans. Syst. Man Cybern. Part B1
2006 Content-based control of goal-directed attention during human action perception
abstract
During the perception of human actions by robotic assistants, the robotic assistant needs to direct its computational and sensor resources to relevant parts of the human action. In previous work we have introduced HAMMER (Hierarchical Attentive Multiple Models for Execution and Recognition) in Demiris, Y. and Khadhouri, B., (2006), a computational architecture that forms multiple hypotheses with respect to what the demonstrated task is, and multiple predictions with respect to the forthcoming states of the human action. To confirm their predictions, the hypotheses request information from an attentional mechanism, which allocates the robot's resources as a function of the saliency of the hypotheses. In this paper we augment the attention mechanism with a component that considers the content of the hypotheses' requests, with respect to reliability, utility and cost. This content-based attention component further optimises the utilisation of the resources while remaining robust to noise. Such computational mechanisms are important for the development of robotic devices that will rapidly respond to human actions, either for imitation or collaboration purposes
Yiannis Demiris, Bassam Khadhouri
RO-MAN1
2006 Perceiving the unusual: Temporal properties of hierarchical motor representations for action perception
Yiannis Demiris, Gavin Simmons
Neural Networks1
2005 Learning Forward Models for Robots
Anthony M. Dearden, Yiannis Demiris
IJCAI2
2005 Compound Effects of Top-down and Bottom-up Influences on Visual Attention During Action Recognition
Bassam Khadhouri, Yiannis Demiris
IJCAI2
2004 Biologically inspired optimal robot arm control with signal-dependent noise
abstract
Progress in the field of humanoid robotics and the need to find simpler ways to program such robots has prompted research into computational models for robotic learning from human demonstration. To further investigate biologically inspired human-like robotic movement and imitation, we have constructed a framework based on three key features of human movement and planning: optimality, modularity and learning. In this paper we focus on the application of optimality principles to the production of human-like movement by a robot arm. Among computational theories of human movement, the signal-dependent noise, or minimum variance, model was chosen as a biologically realistic control scheme to produce human-like movement. A well known optimal control algorithm, the linear quadratic regulator, was adapted to implement this model. The scheme was applied both in simulation and on a real robot arm, which demonstrated human-like movement profiles in a point-to-point reaching experiment.
Gavin Simmons, Yiannis Demiris
IROS2
2003 Distributed, predictive perception of actions: a biologically inspired robotics architecture for imitation and learning
abstract
One of the most important abilities for an agent's cognitive development in a social environment is the ability to recognize and imitate actions of others. In this paper we describe a cognitive architecture for action recognition and imitation, and present experiments demonstrating its implementation in robots. Inspired by neuroscientific and psychological data, and adopting a ‘simulation theory of mind’ approach, the architecture uses the motor systems of the imitator in a dual role, both for generating actions, and for understanding actions when performed by others. It consists of a distributed system of inverse and forward models that uses prediction accuracy as a means to classify demonstrated actions. The architecture is also shown to be capable of learning new composite actions from demonstration.
Yiannis Demiris
Connect. Sci.1