José Lopes 0001

dblp:76/5350-1 · also José David Águas Lopes · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
6since 2021 · last 2022
0000-0002-8773-9216ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 6 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2022 Demonstration of a Robo-Barista for In the Wild Interactions
abstract
We present a demonstration of a Robo-Barista: a social robot that takes hot beverage orders through verbal interaction and completes them via a Bluetooth enabled coffee machine. The demonstration is highly robust and it is the intention that this could be installed as a permanent feature, enabling “In the Wild” experimentation and long term studies. In the demonstration video, we show a user interacting with a Furhat robot to order a coffee. The robot has a novel architecture that allows it to exhibit both verbal and non-verbal cues, such as shared attention and chitchat. Furthermore, it is enabled with a unique tiredness detector based on visual facial features.
Mei Yii Lim, José Lopes 0001, David A. Robb 0001, Bruce W. Wilson, Meriam Moujahid, Helen Hastie
HRI2
2022 Sensitivity of Trust Scales in the Face of Errors
abstract
Trust between humans and robots is a complex, multifaceted phenomenon and measuring it subjectively and reliably is challenging. It is also context dependent and so choosing the right tool for a specific study can prove difficult. This paper aims to evaluate various trust measures and compare them in terms of sensitivity to changes in trust. This is done by comparing two validated trust questionnaires (TAS and MDMT) and one single item assessment in a COVID-19 triage scenario. We found that trust measures are equivalent in terms of sensitivity to changes in trust. Furthermore, the study showed that trust could be measured similarly through a single item assessment in comparison with other lengthier scales, in scenarios with distinct breaks in trust. This finding would be of use for experiments where lengthy questionnaires are not appropriate, such as those in the wild.
Birthe Nesset, Gnanathusharan Rajendran, José Lopes 0001, Helen Hastie
HRI3
2022 We are all Individuals: The Role of Robot Personality and Human Traits in Trustworthy Interaction
abstract
As robots take on roles in our society, it is important that their appearance, behaviour and personality are appropriate for the job they are given and are perceived favourably by the people with whom they interact. Here, we provide an extensive quantitative and qualitative study exploring robot personality but, importantly, with respect to individual human traits. Firstly, we show that we can accurately portray personality in a social robot, in terms of extroversion-introversion using vocal cues and linguistic features. Secondly, through garnering preferences and trust ratings for these different robot personalities, we establish that, for a Robo-Barista, an extrovert robot is preferred and trusted more than an introvert robot, regardless of the subject’s own personality. Thirdly, we find that individual attitudes and predispositions towards robots do impact trust in the Robo-Baristas, and are therefore important considerations in addition to robot personality, roles and interaction context when designing any human-robot interaction study.
Mei Yii Lim, José Lopes 0001, David A. Robb 0001, Bruce W. Wilson, Meriam Moujahid, Emanuele De Pellegrin, Helen Hastie
RO-MAN2
2022 Identification of Low-engaged Learners in Robot-led Second Language Conversations with Adults
abstract
The main aim of this study is to investigate if verbal, vocal, and facial information can be used to identify low-engaged second language learners in robot-led conversation practice. The experiments were performed on voice recordings and video data from 50 conversations, in which a robotic head talks with pairs of adult language learners using four different interaction strategies with varying robot-learner focus and initiative. It was found that these robot interaction strategies influenced learner activity and engagement. The verbal analysis indicated that learners with low activity rated the robot significantly lower on two out of four scales related to social competence. The acoustic vocal and video-based facial analysis, based on manual annotations or machine learning classification, both showed that learners with low engagement rated the robot’s social competencies consistently, and in several cases significantly, lower, and in addition rated the learning effectiveness lower. The agreement between manual and automatic identification of low-engaged learners based on voice recordings or face videos was further found to be adequate for future use. These experiments constitute a first step towards enabling adaption to learners’ activity and engagement through within- and between-strategy changes of the robot’s interaction with learners.
Olov Engwall, Ronald Cumbal, José Lopes 0001, Mikael Ljung, Linnea Månsson
ACM Trans. Hum. Robot Interact.3
2021 Robot Gaze Can Mediate Participation Imbalance in Groups with Different Skill Levels
abstract
Many small group activities, like working teams or study groups, have a high dependency on the skill of each group member. Differences in skill level among participants can affect not only the performance of a team but also influence the social interaction of its members. In these circumstances, an active member could balance individual participation without exerting direct pressure on specific members by using indirect means of communication, such as gaze behaviors. Similarly, in this study, we evaluate whether a social robot can balance the level of participation in a language skill-dependent game, played by a native speaker and a second language learner. In a between-subjects study (N = 72), we compared an adaptive robot gaze behavior, that was targeted to increase the level of contribution of the least active player, with a non-adaptive gaze behavior. Our results imply that, while overall levels of speech participation were influenced predominantly by personal traits of the participants, the robot's adaptive gaze behavior could shape the interaction among participants which lead to more even participation during the game.
Sarah Gillet, Ronald Cumbal, André Pereira 0001, José Lopes 0001, Olov Engwall, Iolanda Leite
HRI4
2021 "You don't understand me!": Comparing ASR Results for L1 and L2 Speakers of Swedish
abstract
The performance of Automatic Speech Recognition (ASR)systems has constantly increased in state-of-the-art develop-ment. However, performance tends to decrease considerably inmore challenging conditions (e.g., background noise, multiplespeaker social conversations) and with more atypical speakers(e.g., children, non-native speakers or people with speech dis-orders), which signifies that general improvements do not nec-essarily transfer to applications that rely on ASR, e.g., educa-tional software for younger students or language learners. Inthis study, we focus on the gap in performance between recog-nition results for native and non-native, read and spontaneous,Swedish utterances transcribed by different ASR services. Wecompare the recognition results using Word Error Rate and an-alyze the linguistic factors that may generate the observed tran-scription errors.
Ronald Cumbal, Birger Moëll, José Lopes 0001, Olov Engwall
Interspeech3
2020 Detection of Listener Uncertainty in Robot-Led Second Language Conversation Practice
abstract
Uncertainty is a frequently occurring affective state that learners experience during the acquisition of a second language. This state can constitute both a learning opportunity and a source of learner frustration. An appropriate detection could therefore benefit the learning process by reducing cognitive instability. In this study, we use a dyadic practice conversation between an adult second-language learner and a social robot to elicit events of uncertainty through the manipulation of the robot's spoken utterances (increased lexical complexity or prosody modifications). The characteristics of these events are then used to analyze multi-party practice conversations between a robot and two learners. Classification models are trained with multimodal features from annotated events of listener (un)certainty. We report the performance of our models on different settings, (sub)turn segments and multimodal inputs.
Ronald Cumbal, José Lopes 0001, Olov Engwall
ICMI2
2020 CRWIZ: A Framework for Crowdsourcing Real-Time Wizard-of-Oz Dialogues
abstract
Large corpora of task-based and open-domain conversational dialogues are hugely valuable in the field of data-driven dialogue systems. Crowdsourcing platforms, such as Amazon Mechanical Turk, have been an effective method for collecting such large amounts of data. However, difficulties arise when task-based dialogues require expert domain knowledge or rapid access to domain-relevant information, such as databases for tourism. This will become even more prevalent as dialogue systems become increasingly ambitious, expanding into tasks with high levels of complexity that require collaboration and forward planning, such as in our domain of emergency response. In this paper, we propose CRWIZ: a framework for collecting real-time Wizard of Oz dialogues through crowdsourcing for collaborative, complex tasks. This framework uses semi-guided dialogue to avoid interactions that breach procedures and processes only known to experts, while enabling the capture of a wide variety of interactions.
Francisco Javier Chiyah Garcia, José Lopes 0001, Xingkun Liu, Helen Hastie
LREC2
2019 Exploring Interaction with Remote Autonomous Systems using Conversational Agents
abstract
Autonomous vehicles and robots are increasingly being deployed to remote, dangerous environments in the energy sector, search and rescue and the military. As a result, there is a need for humans to interact with these robots to monitor their tasks, such as inspecting and repairing offshore wind-turbines. Conversational Agents can improve situation awareness and transparency, while being a hands-free medium to communicate key information quickly and succinctly. As part of our user-centered design of such systems, we conducted an in-depth immersive qualitative study of twelve marine research scientists and engineers, interacting with a prototype Conversational Agent. Our results expose insights into the appropriate content and style for the natural language interaction and, from this study, we derive nine design recommendations to inform future Conversational Agent design for remote autonomous systems.
David A. Robb 0001, José Lopes 0001, Stefano Padilla, Atanas Laskov, Francisco Javier Chiyah Garcia, Xingkun Liu, Jonatan Scharff Willners, Nicolas Valeyrie, Katrin S. Lohan, David Lane, Pedro Patrón, Yvan R. Petillot, Mike J. Chantler, Helen Hastie
Conference on Designing Interactive Systems2
2019 Towards a Conversational Agent for Remote Robot-Human Teaming
abstract
There are many challenges when it comes to deploying robots remotely including lack of operator situation awareness and decreased trust. Here, we present a conversational agent embodied in a Furhat robot that can help with the deployment of such remote robots by facilitating teaming with varying levels of operator control.
José Lopes 0001, David A. Robb 0001, Muneeb Imtiaz Ahmad, Xingkun Liu, Katrin S. Lohan, Helen Hastie
HRI1
2019 A Digital Twin for Human-Robot Interaction
abstract
To avoid putting humans at risk, there is an imminent need to pursue autonomous robotized facilities with maintenance capabilities in the energy industry. This paper presents a video of the ORCA Hub simulator, a framework that unifies three types of autonomous systems (Husky, ANYmal and UAVs) on an offshore platform digital twin for training and testing human-robot collaboration scenarios, such as inspection and emergency response.
Èric Pairet, Paola Ardón Ramirez, Xingkun Liu, José Lopes 0001, Helen Hastie, Katrin S. Lohan
HRI4
2018 FARMI: A FrAmework for Recording Multi-Modal Interactions
Patrik Jonell, Mattias Bystedt, Per Fallgren, Dimosthenis Kontogiorgos, José Lopes 0001, Zofia Malisz, Samuel Mascarenhas, Catharine Oertel, Eran Raveh, Todd Shore
LREC5
2018 The Spot the Difference corpus: a multi-modal corpus of spontaneous task oriented spoken interactions
José Lopes 0001, Nils Hemmingsson, Oliver Åstrand
LREC1
2016 Towards building an attentive artificial listener: on the perception of attentiveness in audio-visual feedback tokens
abstract
Current dialogue systems typically lack a variation of audio-visual feedback tokens. Either they do not encompass feedback tokens at all, or only support a limited set of stereotypical functions. However, this does not mirror the subtleties of spontaneous conversations. If we want to be able to build an artificial listener, as a first step towards building an empathetic artificial agent, we also need to be able to synthesize more subtle audio-visual feedback tokens. In this study, we devised an array of monomodal and multimodal binary comparison perception tests and experiments to understand how different realisations of verbal and visual feedback tokens influence third-party perception of the degree of attentiveness. This allowed us to investigate i) which features (amplitude, frequency, duration...) of the visual feedback influences attentiveness perception; ii) whether visual or verbal backchannels are perceived to be more attentive iii) whether the fusion of unimodal tokens with low perceived attentiveness increases the degree of perceived attentiveness compared to unimodal tokens with high perceived attentiveness taken alone; iv) the automatic ranking of audio-visual feedback token in terms of conveyed degree of attentiveness.
Catharine Oertel, José Lopes 0001, Yu Yu 0003, Kenneth Alberto Funes Mora, Joakim Gustafson, Alan W. Black, Jean-Marc Odobez
ICMI2
2016 Root Cause Analysis of Miscommunication Hotspots in Spoken Dialogue Systems
abstract
A major challenge in Spoken Dialogue Systems (SDS) is the detection of problematic communication (hotspots), as well as the classification of these hotspots into different types (root cause analysi ...
Spiros Georgiladakis, Georgia Athanasopoulou, Raveesh Meena, José Lopes 0001, Arodami Chorianopoulou, Elisavet Palogiannidi, Elias Iosif, Gabriel Skantze, Alexandros Potamianos
INTERSPEECH4
2016 The SpeDial datasets: datasets for Spoken Dialogue Systems analytics
José Lopes 0001, Arodami Chorianopoulou, Elisavet Palogiannidi, Helena Moniz, Alberto Abad, Katerina Louka, Elias Iosif, Alexandros Potamianos
LREC1
2015 Detecting repetitions in spoken dialogue systems using phonetic distances
abstract
This paper addresses the problem of automatic detection of re-peated turns in Spoken Dialogue Systems. Repetitions can be a symptom of problematic communication between users and systems. Such repetitions are often due to speech recognition errors, which in turn makes it hard to use speech recognition to detect repetitions. We present an approach to detect rep-etition using the phonetic distance to find the best alignment between turns in the same dialogue. The alignment score ob-tained is combined with different features to improve repeti-tion detection. To evaluate the method proposed we compare several alignment techniques from edit distance to DTW-based distance, previously used in Spoken-Term detection tasks. We also compare two different methods to compute the phonetic distance: the first one using the phoneme sequence, and the second one using the distance between the phone posterior vec-tors. Two different datasets were used in this evaluation: a bus-schedule information system (in English) and a call routing system (in Swedish). The results show that approaches using phoneme distances over-perform approaches using Levenshtein distances between ASR outputs for repetition detection. Index Terms: spoken dialogue systems, repetition detection, phonetic distance
José Lopes 0001, Giampiero Salvi, Gabriel Skantze, Alberto Abad, Joakim Gustafson, Fernando Batista, Raveesh Meena, Isabel Trancoso
INTERSPEECH1
2015 Automatic Detection of Miscommunication in Spoken Dialogue Systems
abstract
In this paper, we present a data-driven approach for detecting instances of miscommunication in dialogue system interactions.A range of generic features that are both automatically extractable and manually annotated were used to train two models for online detection and one for offline analysis.Online detection could be used to raise the error awareness of the system, whereas offline detection could be used by a system designer to identify potential flaws in the dialogue design.In experimental evaluations on system logs from three different dialogue systems that vary in their dialogue strategy, the proposed models performed substantially better than the majority class baseline models.
Raveesh Meena, José Lopes 0001, Gabriel Skantze, Joakim Gustafson
SIGDIAL Conference2
2015 From rule-based to data-driven lexical entrainment models in spoken dialog systems
José Lopes 0001, Maxine Eskénazi, Isabel Trancoso
Comput. Speech Lang.1
2014 Human-robot collaborative tutoring using multiparty multimodal spoken dialogue
abstract
In this paper, we describe a project that explores a novel experimental setup towards building a spoken, multi-modally rich, and human-like multiparty tutoring robot. A human-robot interaction setup is designed, and a human-human dialogue corpus is collected. The corpus targets the development of a dialogue system platform to study verbal and nonverbal tutoring strategies in multiparty spoken interactions with robots which are capable of spoken dialogue. The dialogue task is centered on two participants involved in a dialogue aiming to solve a card-ordering game. Along with the participants sits a tutor (robot) that helps the participants perform the task, and organizes and balances their interaction. Different multimodal signals captured and auto-synchronized by different audio-visual capture technologies, such as a microphone array, Kinects, and video cameras, were coupled with manual annotations. These are used build a situated model of the interaction based on the participants personalities, their state of attention, their conversational engagement and verbal dominance, and how that is correlated with the verbal and visual feed-back, turn-management, and conversation regulatory actions generated by the tutor. Driven by the analysis of the corpus, we will show also the detailed design methodologies for an affective, and multimodally rich dialogue system that allows the robot to measure incrementally the attention states, and the dominance for each participant, allowing the robot head Furhat to maintain a well-coordinated, balanced, and engaging conversation, that attempts to maximize the agreement and the contribution to solve the task.
Samer Al Moubayed, Jonas Beskow, Bajibabu Bollepalli, Joakim Gustafson, Ahmed Hussen Abdelaziz, Martin Johansson, Maria Koutsombogera, José Lopes 0001, Jekaterina Novikova, Catharine Oertel, Gabriel Skantze, Kalin Stefanov, Gül Varol
HRI8
2014 The Tutorbot Corpus ― A Corpus for Studying Tutoring Behaviour in Multiparty Face-to-Face Spoken Dialogue
Maria Koutsombogera, Samer Al Moubayed, Bajibabu Bollepalli, Ahmed Hussen Abdelaziz, Martin Johansson, José Lopes 0001, Jekaterina Novikova, Catharine Oertel, Kalin Stefanov, Gül Varol
LREC6
2013 Automated two-way entrainment to improve spoken dialog system performance
abstract
This paper proposes an approach to the use of lexical entrainment in Spoken Dialog Systems. This approach aims to increase the dialog success rate by adapting the lexical choices of the system to the user's lexical choices. If the system finds that the users lexical choice degrades the performance, it will try to establish a new conceptual pact, proposing other words that the user may adopt, in order to be more successful in task completion. The approach was implemented and tested in two different systems. Tests showed a relative dialog estimated error rate reduction of 10% and a relative reduction in the average number of turns per session of 6%.
José Lopes 0001, Maxine Eskénazi, Isabel Trancoso
ICASSP1
2011 Towards choosing better primes for spoken dialog systems
abstract
When humans and computers use the same terms (primes, when they entrain to one another), spoken dialogs proceed more smoothly. The goal of this paper is to describe initial steps we have found that will enable us to eventually automatically choose better primes in spoken dialog system prompts. Two different sets of prompts were used to understand what makes one prime more suitable than another. The impact of the primes chosen in speech recognition was evaluated. In addition, results reveal that users did adopt the new vocabulary introduced in the new system prompts. As a result of this, performance of the system improved, providing clues for the trade off needed when choosing between adequate primes in prompts and speech recognition performance.
José Lopes 0001, Maxine Eskénazi, Isabel Trancoso
ASRU1
2011 A nativeness classifier for TED Talks
abstract
This paper presents a nativeness classifier for English. The detector was developed and tested with TED Talks collected from the web, where the major non-native cues are in terms of segmental aspects and prosody. The first experiments were made using only acoustic features, with Gaussian supervectors for training a classifier based on support vector machines. These experiments resulted in an equal error rate of 13.11%. The following experiments based on prosodic features alone did not yield good results. However, a fused system, combining acoustic and prosodic cues, achieved an equal error rate of 10.58%. A small human benchmark was conducted, showing an inter-rater agreement of 0.88. This value is also very close to the agreement value between humans and the best fused system.
José Lopes 0001, Isabel Trancoso, Alberto Abad
ICASSP1
2010 Multimedia learning materials
abstract
This paper describes the integration of multimedia documents in the Portuguese version of REAP, a tutoring system for vocabulary learning. The documents result from the pipeline processing of Broadcast News videos that automatically segments the audio files, transcribes them, adds punctuation and capitalization, and breaks them into stories classified by topics. The integration of these materials in REAP was done in a way that tries to decrease the impact of potential errors of the automatic chain in the learning process.
José Lopes 0001, Isabel Trancoso, Rui Correia, Thomas Pellegrini, Hugo Meinedo, Nuno J. Mamede, Maxine Eskénazi
SLT1