VLDB 2026 Research / reviewers in the wild / expert
Marta Romeo
dblp:222/7940
· DBLP profile ↗
22ranked-venue papers
5as first author
17since 2021 · last 2025
0000-0003-4438-0255ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 15 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 3rd Workshop on Explainability in Human-Robot Collaboration: Real-World ConcernsabstractRobots powered by AI and machine learning are increasingly capable of collaboration and social interaction with humans, leading to a demand to develop new approaches to ensure their transparency and explainable behaviour. As explainable AI (XAI) seeks to clarify AI decisions, its integration into physical robots often creates an illusion of explainability—raising questions about whether current approaches truly enhance understanding. The 3rd Workshop on Explainability in Human-Robot Collaboration aims to address the real-world concerns associated with developing explainable and transparent robots through a focused, multi-faceted panel discussion and a series of paper presentations. In this workshop, we will focus on refining when and how explanations should be provided, integrating human communication principles to enhance trust and transparency in human-robot collaboration through both technical and user-centred solutions. Elmira Yadollahi, Fethiye Irmak Dogan, Marta Romeo, Dimosthenis Kontogiorgos, Peizhu Qian, Yan Zhang 0122 |
HRI | 3 |
| 2025 | Multimodal Emotion Recognition in Conversation via Possible Speaker's Audio and Visual Sequence SelectionabstractMultimodal Emotion Recognition in Conversation (MERC) is an important element in human-machine interaction. It allows machines to automatically identify and track the emotional status of speakers during a conversation in a multimodal setting. However, the conversations involving various audio and visual cues aligned with textual cues are very complex. Recent works have tried integrating the audio and visual modalities with textual to improve the performance of emotion recognition in conversation. Although many MERC models leverage textual, audio, and visual modalities, those models assume that the speaker’s textual utterance, audio speech, and facial sequences are present. However, a conversation may contain multiple parties, among which only one is the speaker. Previous MERC assumed the availability of all modalities, but in many instances, one or more modalities may be unavailable during multiparty conversations. To tackle these issues, we propose the Possible Speaker Informed Multimodal Emotion Recognition in Conversation framework (PSI). PSI is specifically tasked to extract audio (speech) and visual (face) sequences of a possible speaker in the presence of multiple parties. Further, PSI seamlessly extracts the rich unimodal features and fuses them while addressing the unavailability of specific modalities. PSI demonstrates competitive performance with existing state-of-the-art models through experiments with a benchmark dataset. Rahul Singh Maharjan, Niyati Rawal, Marta Romeo, Lorenzo Baraldi 0001, Rita Cucchiara, Angelo Cangelosi |
ICASSP | 3 |
| 2025 | Policy Learning for Social Robot-Led PhysiotherapyabstractSocial robots offer a promising solution for autonomously guiding patients through physiotherapy exercise sessions, but effective deployment requires advanced decision-making to adapt to patient needs. A key challenge is the scarcity of patient behavior data for developing robust policies. To address this, we engaged 33 expert healthcare practitioners as patient proxies, using their interactions with our robot to inform a patient behavior model capable of generating exercise performance metrics and subjective scores on perceived exertion. We trained a reinforcement learning-based policy in simulation, demonstrating that it can adapt exercise instructions to individual exertion tolerances and fluctuating performance, while also being applicable to patients at different recovery stages with varying exercise plans. Carl Bettosi, Lynne Baillie, Susan D. Shenkin, Marta Romeo |
IROS | 4 |
| 2025 | Brain-Robot Interface for Exercise MimicryabstractFor social robots to maintain long-term engagement as exercise instructors, rapport-building is essential. Motor mimicry—imitating one’s physical actions—during social interaction has long been recognized as a powerful tool for fostering rapport, and it is widely used in rehabilitation exercises where patients mirror a physiotherapist or video demonstration. We developed a novel Brain-Robot Interface (BRI) that allows a social robot instructor to mimic a patient’s exercise movements in real-time, using mental commands derived from the patient’s intention. The system was evaluated in an exploratory study with 14 participants (3 physiotherapists and 11 hemiparetic patients recovering from stroke or other injuries). We found our system successfully demonstrated exercise mimicry in 12 sessions, however, accuracy varied. Participants had positive perceptions of the robot instructor, with high trust and acceptance levels, which were not affected by the introduction of BRI technology. Carl Bettosi, Emilyann Nault, Lynne Baillie, Markus Garschall, Marta Romeo, Beatrix Zechmann, Nicole Binderlehner, Theodoros Georgiou 0002 |
RO-MAN | 5 |
| 2025 | Continual Facial Features Transfer for Facial Expression RecognitionabstractFacial Expression Recognition (FER) models based on deep learning mostly rely on a supervised train-once-test-all approach. These approaches assume that a model trained on an in-the-wild facial expression dataset with one type of domain distribution will perform well on a test dataset with a domain distribution shift. However, facial images in real-world can be from different domain distributions from which the model has been trained. However, re-training models on only new domain distributions will severely affect the performance of the previous domain. Re-training on all previous and new data can improve overall performance but is computationally expansive. In this study, we oppose the train-once-test-all approach and propose a buffer-based continual learning approach to enhance the performance of multiple in-the-wild datasets. We propose a model that continually leverages attention to important facial features from the pre-trained model to improve performance in multiple datasets. We validated our model using split-in-the-wild datasets where the dataset is provided to the model in an incremental setting instead of all at once. Furthermore, to evaluate the model performance, we continually used three in-the-wild datasets representing different domains (Domain-FER). Extensive experiments on these datasets reveal that the proposed model achieves better results than other Continual FER models. Rahul Singh Maharjan, Lorenzo Bonicelli, Marta Romeo, Simone Calderara, Angelo Cangelosi, Rita Cucchiara |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | The Effect of Voice and Repair Strategy on Trust Formation and Repair in Human-Robot InteractionabstractTrust is essential for social interactions, including those between humans and social artificial agents, such as robots. Several factors and combinations thereof can contribute to the formation of trust and, importantly in the case of machines that work with a certain margin of error, to its maintenance and repair after it has been breached. In this article, we present the results of a study aimed at investigating the role of robot voice and chosen repair strategy on trust formation and repair in a collaborative task. People helped a robot navigate through a maze, and the robot made mistakes at pre-defined points during the navigation. Via in-game behaviour and follow-up questionnaires, we could measure people’s trust towards the robot. We found that people trusted the robot speaking with a state-of-the-art synthetic voice more than with the default robot voice in the game, even though they indicated the opposite in the questionnaires. Additionally, we found that three repair strategies that people use in human-human interaction (justification of the mistake, promise to be better and denial of the mistake) work also in human-robot interaction. Marta Romeo, Ilaria Torre 0002, Sébastien Le Maguer, Alexander Sleat, Angelo Cangelosi, Iolanda Leite |
ACM Trans. Hum. Robot Interact. | 1 |
| 2024 | Data collection towards socially inspired interactive motion planningabstractIn public and social spaces shared by humans and robots, it is essential for both parties to be aware of each other’s goals and to communicate their intentions effectively. This work in progress aims to develop a motion planner that not only finds the shortest path to a goal while adhering to social norms and spatial constraints but also generates communicative gestures and movements, such as hesitations or prompting motions. Extensive datasets are required to teach robots socially compliant navigation and effective communication with their human counterparts. Here, we outline our data collection efforts to train such a planner and to inform the broader community about potential interactions between humans and robots in these scenarios. The primary goal of this research is to establish a foundation for user-centered design of interaction strategies and avoidance mechanisms during social navigation. Meriam Moujahid, Daniel Hernández García, Marta Romeo, Christian Dondrup |
HAI | 3 |
| 2024 | Broken Trust: Does the Agent Matter?abstractTrust is a key part of any social interaction, whether that be between humans, or humans interacting with different artificial agents. This paper investigates how an agent’s repeated incongruence failure might impact users’ trust. We augment a previously published human-robot interaction study (Nesset et al., 2023), by replacing the robot condition with a human actor. Here, we explore how users’ trust can be impacted by repeated failure depending on the agent involved and how to best repair trust once the failures take place. Our study found a significant decrease in users’ trust when a human makes an incongruence failure, but not when this failure was repeated, regardless of the repair strategy implemented. When comparing this to the previous robot condition, we found a significant difference in the trust measured in the human and the robot condition. Additionally, the repair strategy used had a significant effect on the users’ trust when the robot repeated its failure but not when the actor did. Our findings contribute to research on broken trust with repeated failures and highlight the importance of including a human comparison to better understand research findings in human-robot interactions. Birthe Nesset, Gnanathusharan Rajendran, Marta Romeo |
HAI | 3 |
| 2024 | A Holistic Evaluation Methodology for Multi-Party Spoken Conversational AgentsabstractWhile research in multi-party spoken conversation with intelligent embodied agents has made significant progress in sub-tasks like speaker identification and non-verbal cues, there’s a gap in fully autonomous applications users can directly interact with. This lack translates to the absence of a standard methodology for evaluating multi-party conversational speech agents that considers both task-based system performance and user experience. Nancie Gunson, Angus Addlesee, Daniel Hernández García, Marta Romeo, Christian Dondrup, Oliver Lemon |
IVA | 4 |
| 2024 | Sigh!!! There is more than just faces and verbal speech to recognize emotion in human-robot interactionabstractUnderstanding human emotions is paramount for effective human-human interactions. As technology advances, social robots are increasingly being developed with the capability to discern and respond to human emotions, with the ultimate aim of providing assistance and companionship. However, existing research on emotion recognition for human-robot interaction predominantly focuses on facial expressions or verbal speech, neglecting other potential mediums of emotional expression. In this study, we shed light on the significance of considering various forms of emotional expression, mainly nonverbal vocalization known as vocal bursts, which have been overlooked in emotion modeling for human-robot interaction. Vocal bursts, characterized by brief and intense vocal utterances, represent a rich source of emotional cues that can significantly enhance the capabilities of social robots in understanding and responding to human emotions. Driven by the increasing interest in vocal bursts within speech and affective computing research, we propose a baseline model for affective vocal burst recognition that can outperform large audio models. The proposed baseline model achieves weighted F1 scores of 0.606, 0.342, and 0.287 on 10, 24, and 30 emotion classes, respectively. Additionally, we identify challenges that must be addressed to enhance affective vocal burst recognition for human-robot interaction. Code available at /github.com/rahullabs/Sigh Rahul Singh Maharjan, Marta Romeo, Angelo Cangelosi |
RO-MAN | 2 |
| 2024 | Deep Learning-Based Adaptation of Robot Behaviour for Assistive RoboticsabstractRobot behaviour models in socially assistive robotics are typically trained using high-level features, such as a user’s engagement, such that inaccuracies in the feature extraction can have a significant effect on a robot’s subsequent performance. In this paper, we study whether a behaviour model can be meaningfully represented using an end-to-end approach, where multimodal input, concretely visual data and activity information, is directly processed by a neural network. This paper concretely analyses the different building blocks of such a model, such that the aim is to identify a suitable architecture that can meaningfully combine the different modalities for guiding a robot’s behaviour. We conduct the analysis in the context of a sequence learning game, such that we compare different vision-only models that are then combined with an activity processing network into a joint multimodal model. The results of our evaluation on a dedicated dataset from the sequence learning game demonstrate that a multimodal end-to-end behaviour model has potential for assistive robotics — we report an F1 score of around 0.88 across different dataset-based test scenarios — but the real-life transferability strongly depends on whether the data is diverse enough for capturing meaningful variations in real-world scenarios, such as users being at different distances from a robot. Michal Stolarz, Marta Romeo, Alex Mitrevski, Paul-Gerhard Plöger |
RO-MAN | 2 |
| 2023 | Why is my Agent so Slow? Deploying Human-Like Conversational Turn-TakingabstractThe emphasis on one-to-one speak/wait spoken conversational interaction with intelligent agents leads to long pauses between conversational turns, undermines the flow and naturalness of the interaction, and undermines the user experience. Despite ground breaking advances in the area of generating and understanding natural language with techniques such as LLMs, conversational interaction has remained relatively overlooked. In this workshop we will discuss and review the challenges, recent work and potential impact of improving conversational interaction with artificial systems. We hope to share experiences of poor human/system interaction, best practices with third party tools, and generate design guidance for the community. Matthew P. Aylett, Éva Székely, Donald McMillan, Gabriel Skantze, Marta Romeo, Joel E. Fischer, Gisela Reyes-Cruz |
HAI | 5 |
| 2023 | Faces are Domains: Domain Incremental Learning for Expression RecognitionabstractSince most existing facial expression recognition methods depend on deep learning models trained in isolation on a facial expression image corpora, once employed in scenarios that are different from those in the corpora, they usually demand ad-hoc retraining to be able to perform better in the expression recognition task for new scenarios. Furthermore, most of these facial expression recognition methods are inconsistent when recognising person-specific expressions or are incapable of adjusting to real-world scenarios where data is exclusively obtainable incrementally. In this paper, we present a face incremental expression recognition model, where we utilise domain incremental learning methods to learn individual facial features of facial expressions. We assume that each individual's facial expression (domain) is presented to the model one domain at a time. We assessed our model's ability to remember previously seen domains (individual's facial expression) and incrementally perform on new face domains. Our model improves performance compared to a non-incremental learning model and an incremental learning model in facial expression recognition for individual data with different expression classes. Rahul Singh Maharjan, Marta Romeo, Angelo Cangelosi |
IJCNN | 2 |
| 2023 | To Whom are You Talking? A Deep Learning Model to Endow Social Robots with Addressee Estimation SkillsabstractCommunicating shapes our social word. For a robot to be considered social and being consequently integrated in our social environment it is fundamental to understand some of the dynamics that rule human-human communication. In this work, we tackle the problem of Addressee Estimation, the ability to understand an utterance's addressee, by interpreting and exploiting non-verbal bodily cues from the speaker. We do so by implementing an hybrid deep learning model composed of convolutional layers and LSTM cells taking as input images portraying the face of the speaker and 2D vectors of the speaker's body posture. Our implementation choices were guided by the aim to develop a model that could be deployed on social robots and be efficient in ecological scenarios. We demonstrate that our model is able to solve the Addressee Estimation problem in terms of addressee localisation in space, from a robot ego-centric point of view. Carlo Mazzola, Marta Romeo, Francesco Rea, Alessandra Sciutti, Angelo Cangelosi |
IJCNN | 2 |
| 2023 | Robot Broken Promise? Repair strategies for mitigating loss of trust for repeated failuresabstractTrust repair strategies are an important part of human-robot interaction. In this study, we investigate how repeated failures impact users’ trust and how we might mitigate them. Specifically, we look at different repair strategies in the form of apologies, with additional features to them such as warnings and promises. Through an online study, we explore these repair strategies for repeated failures in the form of robot incongruence, where there is a mismatch of verbal and non-verbal information given by the robot. Our results show that such incongruent robot behaviour has a significant overall negative impact on participants’ trust. We found that the robot making a promise, and then breaking it, results in a significant decrease in participants’ trust, when compared to a general apology as a repair strategy. These findings contribute to the research on trust repair strategies and, additionally, shed light on how robot failures, in the form of incongruences, impact participants’ trust. Birthe Nesset, Marta Romeo, Gnanathusharan Rajendran, Helen Hastie |
RO-MAN | 2 |
| 2023 | Putting Robots in Context: Challenging the Influence of Voice and Empathic Behaviour on TrustabstractTrust is essential for social interactions, including those between humans and social artificial agents, such as robots. Several robot-related factors can contribute to the formation of trust. However, previous work has often treated trust as an absolute concept, whereas it is highly context-dependent, and it is possible that some robot-related features will influence trust in some contexts, but not in others. In this paper, we present the results of two video-based online studies aimed at investigating the role of robot voice and empathic behaviour on trust formation in a general context as well as in a task-specific context. We found that voice influences trust in the specific context, with no effect of voice or empathic behaviour in the general context. Thus, context mediated whether robot-related features play a role in people’s trust formation towards robots. Marta Romeo, Ilaria Torre 0002, Sébastien Le Maguer, Angelo Cangelosi, Iolanda Leite |
RO-MAN | 1 |
| 2022 | Exploring Theory of Mind for Human-Robot CollaborationabstractThe ability to impute mental states to oneself or others, or Theory of Mind (ToM), has been intrinsically linked to trust between humans. However, less is known about how a robot mimicking ToM affects users’ trust and behaviour. We explore this through an online study, where we compare three robot personas in a cooperative maze navigation task: one neutral, one that explains its reasoning in technical terms, and one that mimics ToM. We show that ToM influences human decision-making behaviour and trust in a way that makes it more appropriate with respect to the competencies of the robot. This is key for human-robot collaboration and adoption of robotics moving forward. Marta Romeo, Peter E. McKenna, David A. Robb 0001, Gnanathusharan Rajendran, Birthe Nesset, Angelo Cangelosi, Helen Hastie |
RO-MAN | 1 |
| 2019 | Deploying a Deep Learning Agent for HRI with PotentialabstractWith the global population aging at an alarming rate, the need to find alternative ways to deliver quality assistance is becoming a pressing concern for health and care systems. To promptly provide companion-like assistance, robots need to gain social intelligence in an autonomous way, without relying on human operators. The work described in this paper aims to develop a deep learning agent that, by means of convolutional neural network architecture in the decision making loop, could understand when and how, to interact with one, or more people, gathered in a room. This was done by training a robot to assess the level of user engagement at the initiation of the interaction, so that the robot could detect the person most willing to start interacting. The robot's performance as a deep learning agent was tested through an experiment with potential ''end-users'', following an iterative process, over four days. The deep learning agent was able to take the right decision 59% of the times by the end of the experiment, from an initial success rate of 44% on the first day, proving the potential of such technologies in this application field. Marta Romeo, Daniel Hernández García, Ray Jones, Angelo Cangelosi |
HAI | 1 |
| 2019 | Social Robots in Therapy and CareabstractThe Social Robots in Therapy workshop series aims at advancing research topics related to the use of robots in the contexts of Social Care and Robot-Assisted Therapy (RAT). Robots in social care and therapy have been a long time promise in HRI as they have the opportunity to improve patients life significantly. Multiple challenges have to be addressed for this, such as building platforms that work in proximity with patients, therapists and health-care professionals; understanding user needs; developing adaptive and autonomous robot interactions; and addressing ethical questions regarding the use of robots with a vulnerable population. The full-day workshop follows last year's edition which centered on how social robots can improve health-care interventions, how increasing the degree of autonomy of the robots might affect therapies, and how to overcome the ethical challenges inherent to the use of robot assisted technologies. This 2ndedition of the workshop will be focused on the importance of equipping social robots with socio-emotional intelligence and the ability to perform meaningful and personalized interactions. This workshop aims to bring together researchers and industry experts in the fields of Human-Robot Interaction, Machine Learning and Robots in Health and Social Care. It will be an opportunity for all to share and discuss ideas, strategies and findings to guide the design and development of robot-assisted systems for therapy and social care implementations that can provide personalize, natural, engaging and autonomous interactions with patients (and health-care providers). Daniel Hernández García, Pablo Gómez Esteban, Hee Rin Lee, Marta Romeo, Emmanuel Senft, Erik Billing |
HRI | 4 |
| 2019 | Evaluating the Acceptability of Assistive Robots for Early Detection of Mild Cognitive ImpairmentabstractThe employment of Social Assistive Robots (SARs) for monitoring elderly users represents a valuable gateway for at-home assistance. Their deployment in the house of the users can provide effective opportunities for early detection of Mild Cognitive Impairment (MCI), a condition of increasing impact in our aging society, by means of digitalized cognitive tests. In this work, we present a system where a specific set of cognitive tests is selected, digitalized, and integrated with a robotic assistant, whose task is the guidance and supervision of the users during the completion of such tests. The system is then evaluated by means of an experimental study involving potential future users, in order to assess its acceptability and identify key directions for technical improvements. Matteo Luperto, Marta Romeo, Francesca Lunardini, Nicola Basilico, Carlo Abbate, Ray Jones, Angelo Cangelosi, Simona Ferrante, N. Alberto Borghese |
IROS | 2 |
| 2018 | Digitalized Cognitive Assessment mediated by a Virtual CaregiverabstractThe ageing of the population deeply impacts on the social costs relative to health care. The use of modern technologies is one of the most promising approaches, under current study, to reduce such impact. In this demonstration, we propose a framework that can be employed for at-home assessment of Mild Cognitive Impairment (MCI). It is composed by a set of digitalized cognitive tests, developed from their paper-and-pencil counterparts, and by a Virtual Caregiver, which oversees the test execution and provides instructions. Matteo Luperto, Marta Romeo, Francesca Lunardini, Nicola Basilico, Ray Jones, Angelo Cangelosi, Simona Ferrante, N. Alberto Borghese |
IJCAI | 2 |
| 2018 | Developing a Deep Learning Agent for HRI: Dataset Collection and TrainingabstractThe world population is ageing at a dramatic rate, raising new challenges for social and health care systems. Sometimes, assistance can simply derive from a social interaction between a robotic platform and human users. In these cases, robots cannot rely on human operators. Therefore, they need to gain social intelligence in a fully autonomous way. The focus of this paper is on the initial steps needed to implement a completely autonomous robotic agent able to adapt itself to its users. For this reason, an interactive data collection was carried out to gather a dataset from which the robot could learn how to respond to its users in different situations. From these data, a first evaluation of the performances of the deep learning agent, embodied in the robot, has been completed. The agent was able to generalize to new sets of test data. The study explored how, using modern machine learning algorithms, a robot could learn to understand if, and how, to interact with one, or more people, gathered in a room. This was done by training a robot to read the level of the engagement of the users at the initiation of the interaction. Marta Romeo, Angelo Cangelosi, Ray Jones |
RO-MAN | 1 |