VLDB 2026 Research / reviewers in the wild / expert
Dimosthenis Kontogiorgos
dblp:166/4673
· DBLP profile ↗
19ranked-venue papers
10as first author
5since 2021 · last 2025
0000-0002-8874-6629ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 6 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 12 · 7 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Questioning the Robot: Using Human Non-verbal Cues to Estimate the Need for ExplanationsabstractAs black-box AI systems become increasingly complex, understanding when and how to provide explanations to users is crucial. Multimodal signals, such as facial expressions, offer novel insights into how frequently explanations should be given. This paper explores whether users’ facial features can help estimate the need for explanations in a collaborative robot task. We applied three state-of-the-art eXplainable AI (XAI) methods, addressing how, why, and what-if questions, explaining the robot's failure detection model. Each explanation type conveyed information differently: how-explanations described how the model functions, why-explanations prowided personalised insights into input-feature-related cues, and what-if-explanations explored alternative scenarios. In a mixed-design study ($\mathrm{N}=33$), participants performed a robot-assisted pick-and-place task, receiving different explanation types. Our results show that users responded differently to these explanations, with why-explanations being the most preferred and prompting closer alignment in facial expressions with the robot Contrary to expectations, what-if explanations led to the least alignment and required greater vocal effort. These findings demonstrate how non-verbal cues can guide the frequency and type of explanations (personalised or general) and further highlight the importance of model transparency in human-robot collaboration. Dimosthenis Kontogiorgos, Julie A. Shah |
HRI | 1 |
| 2025 | 3rd Workshop on Explainability in Human-Robot Collaboration: Real-World ConcernsabstractRobots powered by AI and machine learning are increasingly capable of collaboration and social interaction with humans, leading to a demand to develop new approaches to ensure their transparency and explainable behaviour. As explainable AI (XAI) seeks to clarify AI decisions, its integration into physical robots often creates an illusion of explainability—raising questions about whether current approaches truly enhance understanding. The 3rd Workshop on Explainability in Human-Robot Collaboration aims to address the real-world concerns associated with developing explainable and transparent robots through a focused, multi-faceted panel discussion and a series of paper presentations. In this workshop, we will focus on refining when and how explanations should be provided, integrating human communication principles to enhance trust and transparency in human-robot collaboration through both technical and user-centred solutions. Elmira Yadollahi, Fethiye Irmak Dogan, Marta Romeo, Dimosthenis Kontogiorgos, Peizhu Qian, Yan Zhang 0122 |
HRI | 4 |
| 2025 | Versatile Demonstration Interface: Toward More Flexible Robot Demonstration CollectionabstractPrevious methods for Learning from Demonstration leverage several approaches for a human to teach motions to a robot, including teleoperation, kinesthetic teaching, and natural demonstrations. However, little previous work has explored more general interfaces that allow for multiple demonstration types. Given the varied preferences of human demonstrators and task characteristics, a flexible tool that enables multiple demonstration types could be crucial for broader robot skill training. In this work, we propose Versatile Demonstration Interface (VDI), an attachment for collaborative robots that simplifies the collection of three common types of demonstrations. Designed for flexible deployment in industrial settings, our tool requires no additional instrumentation of the environment. Our prototype interface captures human demonstrations through a combination of vision, force sensing, and state tracking (e.g., through the robot proprioception or AprilTag tracking). Through a user study where we deployed our prototype VDI at a local manufacturing innovation center with manufacturing experts, we demonstrated VDI in representative industrial tasks. Interactions from our study highlight the practical value of VDI’s varied demonstration types, expose a range of industrial use cases for VDI, and provide insights for future tool design. Michael Hagenow, Dimosthenis Kontogiorgos, Julie A. Shah |
IROS | 2 |
| 2022 | Robo-Identity: Exploring Artificial Identity and Emotion via Speech InteractionsabstractFollowing the success of the first edition of Robo-Identity, the second edition will provide an opportunity to expand the discussion about artificial identity. This year, we are focusing on emotions that are expressed through speech and voice. Synthetic voices of robots can resemble and are becoming indistinguishable from expressive human voices. This can be an opportunity and a constraint in expressing emotional speech that can (falsely) convey a human-like identity that can mislead people, leading to ethical issues. How should we envision an agent's artificial identity? In what ways should we have robots that maintain a machine-like stance, e.g., through robotic speech, and should emotional expressions that are increasingly human-like be seen as design opportunities? These are not mutually exclusive concerns. As this discussion needs to be conducted in a multidisciplinary manner, we welcome perspectives on challenges and opportunities from variety of fields. For this year's edition, the special theme will be “speech, emotion and artificial identity”. Guy Laban, Sébastien Le Maguer, Minha Lee, Dimosthenis Kontogiorgos, Samantha Reig, Ilaria Torre 0002, Ravi Tejwani, Matthew J. Dennis, André Pereira 0001 |
HRI | 4 |
| 2021 | A Systematic Cross-Corpus Analysis of Human Reactions to Robot Conversational FailuresabstractIn this paper, we analyze multimodal behavioral responses to robot failures across different tasks. Two multimodal datasets are examined in which humans interact with guided-task robots in task-oriented dialogues. In both datasets, the robots simulated failures of conversational breakdown and miscommunication typically observed in human-robot interactions. We closely examine human reactions to these failures looking at facial and acoustic features. Our analyses identify the significant behavioral features for automatic detection of such failures in interaction. We also examine human responses to different types of robot failures and if failures occurred early or late in the interaction cause variation in the responses. Our findings indicate that several nonverbal behaviors are consistently present in responses to robots’ failures, e.g., gaze and speech prosody, whereas, linguistic features appear to be task-dependent. We discuss how these findings may generalize to other tasks, and how autonomous robots may identify opportunities to detect and recover from failures in interactions with humans. Dimosthenis Kontogiorgos, Minh Tran 0004, Joakim Gustafson, Mohammad Soleymani 0001 |
ICMI | 1 |
| 2020 | Embodiment Effects in Interactions with Failing RobotsabstractThe increasing use of robots in real-world applications will inevitably cause users to encounter more failures in interactions. While there is a longstanding effort in bringing human-likeness to robots, how robot embodiment affects users' perception of failures remains largely unexplored. In this paper, we extend prior work on robot failures by assessing the impact that embodiment and failure severity have on people's behaviours and their perception of robots. Our findings show that when using a smart-speaker embodiment, failures negatively affect users' intention to frequently interact with the device, however not when using a human-like robot embodiment. Additionally, users significantly rate the human-like robot higher in terms of perceived intelligence and social presence. Our results further suggest that in higher severity situations, human-likeness is distracting and detrimental to the interaction. Drawing on quantitative findings, we discuss benefits and drawbacks of embodiment in robot failures that occur in guided tasks. Dimosthenis Kontogiorgos, Sanne van Waveren, Olle Wallberg, André Pereira 0001, Iolanda Leite, Joakim Gustafson |
CHI | 1 |
| 2020 | Behavioural Responses to Robot Conversational FailuresabstractHumans and robots will increasingly collaborate in domestic environments which will cause users to encounter more failures in interactions. Robots should be able to infer conversational failures by detecting human users' behavioural and social signals. In this paper, we study and analyse these behavioural cues in response to robot conversational failures. Using a guided task corpus, where robot embodiment and time pressure are manipulated, we ask human annotators to estimate whether user affective states differ during various types of robot failures. We also train a random forest classifier to detect whether a robot failure has occurred and compare results to human annotator benchmarks. Our findings show that human-like robots augment users' reactions to failures, as shown in users' visual attention, in comparison to non-human-like smart-speaker embodiments. The results further suggest that speech behaviours are utilised more in responses to failures when non-human-like designs are present. This is particularly important to robot failure detection mechanisms that may need to consider the robot's physical design in its failure detection model. Dimosthenis Kontogiorgos, André Pereira 0001, Boran Sahindal, Sanne van Waveren, Joakim Gustafson |
HRI | 1 |
| 2020 | Chinese Whispers: A Multimodal Dataset for Embodied Language GroundingabstractIn this paper, we introduce a multimodal dataset in which subjects are instructing each other how to assemble IKEA furniture. Using the concept of ‘Chinese Whispers’, an old children’s game, we employ a novel method to avoid implicit experimenter biases. We let subjects instruct each other on the nature of the task: the process of the furniture assembly. Uncertainty, hesitations, repairs and self-corrections are naturally introduced in the incremental process of establishing common ground. The corpus consists of 34 interactions, where each subject first assembles and then instructs. We collected speech, eye-gaze, pointing gestures, and object movements, as well as subjective interpretations of mutual understanding, collaboration and task recall. The corpus is of particular interest to researchers who are interested in multimodal signals in situated dialogue, especially in referential communication and the process of language grounding. Dimosthenis Kontogiorgos, Elena Sibirtseva, Joakim Gustafson |
LREC | 1 |
| 2019 | The Effects of Embodiment and Social Eye-Gaze in Conversational Agents
Dimosthenis Kontogiorgos, Gabriel Skantze, André Pereira 0001, Joakim Gustafson |
CogSci | 1 |
| 2019 | Estimating Uncertainty in Task-Oriented DialogueabstractSituated multimodal systems that instruct humans need to handle user uncertainties, as expressed in behaviour, and plan their actions accordingly. Speakers’ decision to reformulate or repair previous utterances depends greatly on the listeners’ signals of uncertainty. In this paper, we estimate uncertainty in a situated guided task, as leveraged in non-verbal cues expressed by the listener, and predict that the speaker will reformulate their utterance. We use a corpus where people instruct how to assemble furniture, and extract their multimodal features. While uncertainty is in cases verbally expressed, most instances are expressed non-verbally, which indicates the importance of multimodal approaches. In this work, we present a model for uncertainty estimation. Our findings indicate that uncertainty estimation from non-verbal cues works well, and can exceed human annotator performance when verbal features cannot be perceived. Dimosthenis Kontogiorgos, André Pereira 0001, Joakim Gustafson |
ICMI | 1 |
| 2019 | The Effects of Anthropomorphism and Non-verbal Social Behaviour in Virtual AssistantsabstractThe adoption of virtual assistants is growing at a rapid pace. However, these assistants are not optimised to simulate key social aspects of human conversational environments. Humans are intellectually biased toward social activity when facing anthropomorphic agents or when presented with subtle social cues. In this paper, we test whether humans respond the same way to assistants in guided tasks, when in different forms of embodiment and social behaviour. In a within-subject study (N=30), we asked subjects to engage in dialogue with a smart speaker and a social robot. We observed shifting of interactive behaviour, as shown in behavioural and subjective measures. Our findings indicate that it is not always favourable for agents to be anthropomorphised or to communicate with nonverbal cues. We found a trade-off between task performance and perceived sociability when controlling for anthropomorphism and social behaviour. Dimosthenis Kontogiorgos, André Pereira 0001, Olle Andersson, Marco Koivisto, Elena Gonzalez Rabal, Ville Vartiainen, Joakim Gustafson |
IVA | 1 |
| 2019 | Modeling of Human Visual Attention in Multiparty Open-World DialoguesabstractThis study proposes, develops, and evaluates methods for modeling the eye-gaze direction and head orientation of a person in multiparty open-world dialogues, as a function of low-level communicative signals generated by his/hers interlocutors. These signals include speech activity, eye-gaze direction, and head orientation, all of which can be estimated in real time during the interaction. By utilizing these signals and novel data representations suitable for the task and context, the developed methods can generate plausible candidate gaze targets in real time. The methods are based on Feedforward Neural Networks and Long Short-Term Memory Networks. The proposed methods are developed using several hours of unrestricted interaction data and their performance is compared with a heuristic baseline method. The study offers an extensive evaluation of the proposed methods that investigates the contribution of different predictors to the accurate generation of candidate gaze targets. The results show that the methods can accurately generate candidate gaze targets when the person being modeled is in a listening state. However, when the person being modeled is in a speaking state, the proposed methods yield significantly lower performance. Kalin Stefanov, Giampiero Salvi, Dimosthenis Kontogiorgos, Hedvig Kjellström, Jonas Beskow |
ACM Trans. Hum. Robot Interact. | 3 |
| 2018 | FARMI: A FrAmework for Recording Multi-Modal Interactions
Patrik Jonell, Mattias Bystedt, Per Fallgren, Dimosthenis Kontogiorgos, José Lopes 0001, Zofia Malisz, Samuel Mascarenhas, Catharine Oertel, Eran Raveh, Todd Shore |
LREC | 4 |
| 2018 | Crowdsourced Multimodal Corpora Collection Tool
Patrik Jonell, Catharine Oertel, Dimosthenis Kontogiorgos, Jonas Beskow, Joakim Gustafson |
LREC | 3 |
| 2018 | A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction
Dimosthenis Kontogiorgos, Vanya Avramova, Simon Alexanderson, Patrik Jonell, Catharine Oertel, Jonas Beskow, Gabriel Skantze, Joakim Gustafson |
LREC | 1 |
| 2018 | A Comparison of Visualisation Methods for Disambiguating Verbal Requests in Human-Robot InteractionabstractPicking up objects requested by a human user is a common task in human-robot interaction. When multiple objects match the user's verbal description, the robot needs to clarify which object the user is referring to before executing the action. Previous research has focused on perceiving user's multimodal behaviour to complement verbal commands or minimising the number of follow up questions to reduce task time. In this paper, we propose a system for reference disambiguation based on visualisation and compare three methods to disambiguate natural language instructions. In a controlled experiment with a YuMi robot, we investigated realtime augmentations of the workspace in three conditions - head-mounted display, projector, and a monitor as the baseline - using objective measures such as time and accuracy, and subjective measures like engagement, immersion, and display interference. Significant differences were found in accuracy and engagement between the conditions, but no differences were found in task time. Despite the higher error rates in the head-mounted display condition, participants found that modality more engaging than the other two, but overall showed preference for the projector condition over the monitor and head-mounted display conditions. Elena Sibirtseva, Dimosthenis Kontogiorgos, Olov Nykvist, Hakan Karaoguz, Iolanda Leite, Joakim Gustafson, Danica Kragic |
RO-MAN | 2 |
| 2017 | Multimodal language grounding for improved human-robot collaboration: exploring spatial semantic representations in the shared space of attentionabstractThere is an increased interest in artificially intelligent technology that surrounds us and takes decisions on our behalf. This creates the need for such technology to be able to communicate with humans and understand natural language and non-verbal behaviour that may carry information about our complex physical world. Artificial agents today still have little knowledge about the physical space that surrounds us and about the objects or concepts within our attention. We are still lacking computational methods in understanding the context of human conversation that involves objects and locations around us. Can we use multimodal cues from human perception of the real world as an example of language learning for robots? Can artificial agents and robots learn about the physical world by observing how humans interact with it and how they refer to it and attend during their conversations? This PhD project’s focus is on combining spoken language and non-verbal behaviour extracted by multi-party dialogue in order to increase context awareness and spatial understanding for artificial agents. Dimosthenis Kontogiorgos |
ICMI | 1 |
| 2017 | Crowd-Sourced Design of Artificial Attentive ListenersabstractFeedback generation is an important component of humanhuman communication. Humans can choose to signal support, understanding, agreement or also sceptiscism by means of feedback tokens. Many studie ... Catharine Oertel, Patrik Jonell, Dimosthenis Kontogiorgos, Joseph Mendelson, Jonas Beskow, Joakim Gustafson |
INTERSPEECH | 3 |
| 2017 | Crowd-Powered Design of Virtual Attentive Listeners
Patrik Jonell, Catharine Oertel, Dimosthenis Kontogiorgos, Jonas Beskow, Joakim Gustafson |
IVA | 3 |