VLDB 2026 Research / reviewers in the wild / expert
Divesh Lala
dblp:11/10643
· DBLP profile ↗
32ranked-venue papers
13as first author
13since 2021 · last 2026
0009-0005-3181-6410ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 6 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 10 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous LevelsabstractIn multi-party dialogues between a dialogue system and multiple users, identifying to whom an utterance is addressed is a key challenge. Prior work has typically treated addressee detection as a multi-class classification task, selecting a single label representing an individual participant or the group. This formulation assumes that address is inherently discrete and has primarily been used for predicting turn-taking. In this paper, we revisit this assumption by analyzing address as a continuous phenomenon. Using a multi-party human dialogue corpus annotated by multiple annotators, we construct both binary address labels derived from majority-vote addressee labels and continuous address levels inferred from annotator judgments using a latent-variable model. We then examine how these representations relate to turn-taking as well as listener behaviors, including gaze and backchannels. Our results show that, in addition to turn-taking, both gaze and backchannels are associated with address. Furthermore, models using continuous address levels achieve better predictive fit than those using discrete labels, suggesting that address may exhibit graded structure. Finally, we discuss the future directions of addressee detection research based on the findings of this study. Taiga Mori, Koji Inoue, Divesh Lala, Tatsuya Kawahara |
SIGDIAL | 3 |
| 2026 | Robot-Mediated Multi-Party Conversation Aimed at Affect Improvement for Psychiatric PatientsabstractThis paper describes a multi-party attentive listening system that interacts with two persons to familiarize each other through conversation. We are mainly targeting social implementation in hospitals to contribute to the rehabilitation of people with psychiatric disorders to promote affect improvement in terms of pleasure and arousal. We conducted an experiment in a psychiatric outpatient-daycare program. Twenty daycare attendees participated in a three-party conversation session between a pair of two humans and a humanoid robot. One of the paired participants talked about his/her favorite topic and was attentively listened to by the other and the robot. In a subsequent session, the human pairs switched each other's roles. The subjective evaluations showed that both the pleasure and arousal of the participants were significantly improved after the conversation. The participants rated the impression of the robot as easier to talk with than strangers. They also rated that they could understand and feel familiar with significantly more their human talk partner after the conversational session. The multiple linear regression analysis showed that participants became more pleasant and verbal when stimulated by backchannels and questions from both the human listener and the robot. It suggests that those who expressed themself using more words received positive impressions. Keiko Ochi, Divesh Lala, Koji Inoue, Tatsuya Kawahara, Hirokazu Kumazaki |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | CCMI 2025: Cross-Cultural Multimodal Interaction
Koji Inoue, Shogo Okada, Divesh Lala, Sahba Zojaji, Nancy F. Chen, Tatsuya Kawahara |
ICMI | 3 |
| 2025 | Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System
Kazushi Kato, Koji Inoue, Divesh Lala, Keiko Ochi, Tatsuya Kawahara |
ICMI | 3 |
| 2025 | Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue SystemsabstractTurn-taking is a fundamental component of spoken dialogue, however conventional studies mostly involve dyadic settings. This work focuses on applying voice activity projection (VAP) to predict upcoming turn-taking in triadic multi-party scenarios. The goal of VAP models is to predict the future voice activity for each speaker utilizing only acoustic data. This is the first study to extend VAP into triadic conversation. We trained multiple models on a Japanese triadic dataset where participants discussed a variety of topics. We found that the VAP trained on triadic conversation outperformed the baseline for all models but that the type of conversation affected the accuracy. This study establishes that VAP can be used for turn-taking in triadic dialogue scenarios. Future work will incorporate this triadic VAP turn-taking model into spoken dialogue systems. Mikey Elmers, Koji Inoue, Divesh Lala, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2025 | Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity ProjectionabstractKoji Inoue, Divesh Lala, Gabriel Skantze, Tatsuya Kawahara. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Koji Inoue, Divesh Lala, Gabriel Skantze, Tatsuya Kawahara |
NAACL (Long Papers) | 2 |
| 2025 | Prompt-Guided Turn-Taking PredictionabstractTurn-taking prediction models are essential components in spoken dialogue systems and conversational robots. Recent approaches leverage transformer-based architectures to predict speech activity continuously and in real-time. In this study, we propose a novel model that enables turn-taking prediction to be dynamically controlled via textual prompts. This approach allows intuitive and explicit control through instructions such as “faster” or “calmer,” adapting dynamically to conversational partners and contexts. The proposed model builds upon a transformer-based voice activity projection (VAP) model, incorporating textual prompt embeddings into both channel-wise transformers and a cross-channel transformer. We evaluated the feasibility of our approach using over 950 hours of human-human spoken dialogue data. Since textual prompt data for the proposed approach was not available in existing datasets, we utilized a large language model (LLM) to generate synthetic prompt sentences. Experimental results demonstrated that the proposed model improved prediction accuracy and effectively varied turn-taking timing behaviors according to the textual prompts. Koji Inoue, Mikey Elmers, Yahui Fu 0001, Zi Haur Pang, Divesh Lala, Keiko Ochi, Tatsuya Kawahara |
SIGDIAL | 5 |
| 2024 | Entrainment Analysis and Prosody Prediction of Subsequent Interlocutor's Backchannels in Dialogue
Keiko Ochi, Koji Inoue, Divesh Lala, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2022 | Backchannel Generation Model for a Third Party Listener AgentabstractIn this work we propose a listening agent which can be used in a conversation between two humans. We firstly conduct a corpus analysis to identify three different categories of backchannel which the agent can use - responsive interjections, expressive interjections and shared laughs. From this data we train and evaluate a continuous backchannel generation model consisting of separate timing and form prediction models. We then conduct a subjective experiment to compare our model to random, dyadic, and ground truth models. We find that our model outperforms a random baseline and is comparable to the dyadic model despite the low evaluation of expressive interjections. We suggest that the perception of expressive interjections contribute significantly to the perception of the agent’s empathy and understanding of the conversation. The results also show the need for a more robust model to generate expressive interjections, perhaps aided by the use of linguistic features. Divesh Lala, Koji Inoue, Tatsuya Kawahara, Kei Sawada |
HAI | 1 |
| 2022 | Alzheimer's Dementia Detection through Spontaneous Dialogue with Proactive Robotic ListenersabstractAs the aging of society continues to accelerate, Alzheimer's Disease (AD) has received more and more attention from not only medical but also other fields, such as computer science, over the past decade. Since speech is considered one of the effective ways to diagnose cognitive decline, AD detection from speech has emerged as a hot topic. Nevertheless, such approaches fail to tackle several key issues: 1) AD is a complex neurocognitive disorder which means it is inappropriate to conduct AD detection using utterance information alone while ignoring dialogue infor-mation; 2) Utterances of AD patients contain many disfluencies that affect speech recognition yet are helpful to diagnosis; 3) AD patients tend to speak less, causing dialogue breakdown as the disease progresses. This fact leads to a small number of utterances, which may cause detection bias. Therefore, in this paper, we propose a novel AD detection architecture consisting of two major modules: an ensemble AD detector and a proactive listener. This architecture can be embedded in the dialogue system of conversational robots for healthcare. Yuanchao Li, Catherine Lai, Divesh Lala, Koji Inoue, Tatsuya Kawahara |
HRI | 3 |
| 2022 | Simultaneous Job Interview System Using Multiple Semi-autonomous AgentsabstractIn recent years, spoken dialogue systems have been used in job interviews where an applicant talks to a system that asks pre-defined questions, called on-demand and self-paced job interviews.We propose a simultaneous job interview system, where one interviewer can conduct one-on-one interviews with multiple applicants simultaneously by cooperating with multiple autonomous interview dialogue systems.However, it is challenging for interviewers to monitor and understand all parallel interviews done by the autonomous system simultaneously.To address this issue, we implement two automatic dialogue understanding functions: (1) response evaluation of each applicant's responses and (2) keyword extraction for a summary of the responses.In this system, interviewers can intervene in a dialogue session when needed and smoothly ask a proper question that elaborates the interview.We have conducted a pilot experiment where an interviewer conducted simultaneous job interviews with three candidates. Haruki Kawai, Yusuke Muraki, Kenta Yamamoto, Divesh Lala, Koji Inoue, Tatsuya Kawahara |
SIGDIAL | 4 |
| 2021 | A multi-party attentive listening robot which stimulates involvement from side participantsabstractWe demonstrate the moderating abilities of a multi-party attentive listening robot system when multiple people are speaking in turns.Our conventional one-on-one attentive listening system generates listener responses such as backchannels, repeats, elaborating questions, and assessments.In this paper, additional robot responses that stimulate a listening user (side participant) to become more involved in the dialogue are proposed.The additional responses elicit assessments and questions from the side participant, making the dialogue more empathetic and lively. Koji Inoue, Hiromi Sakamoto, Kenta Yamamoto, Divesh Lala, Tatsuya Kawahara |
SIGDIAL | 4 |
| 2021 | ERICA: An Empathetic Android Companion for Covid-19 QuarantineabstractOver the past year, research in various domains, including Natural Language Processing (NLP), has been accelerated to fight against the COVID-19 pandemic, yet such research has just started on dialogue systems.In this paper, we introduce an end-to-end dialogue system which aims to ease the isolation of people under self-quarantine.We conduct a control simulation experiment to assess the effects of the user interface, a web-based virtual agent called Nora vs. the android ERICA via a video call.The experimental results show that the android offers a more valuable user experience by giving the impression of being more empathetic and engaging in the conversation due to its nonverbal information, such as facial expressions and body gestures.Demo video available at https://youtu.be/PLPEBXLeKJI. Etsuko Ishii, Genta Indra Winata, Samuel Cahyawijaya, Divesh Lala, Tatsuya Kawahara, Pascale Fung |
SIGDIAL | 4 |
| 2020 | Designing Precise and Robust Dialogue Response EvaluatorsabstractAutomatic dialogue response evaluator has been proposed as an alternative to automated metrics and human evaluation.However, existing automatic evaluators achieve only moderate correlation with human judgement and they are not robust.In this work, we propose to build a reference-free evaluator and exploit the power of semi-supervised training and pretrained (masked) language models.Experimental results demonstrate that the proposed evaluator achieves a strong correlation (> 0.6) with human judgement and generalizes robustly to diverse responses and corpora.We open-source the code and data in https://github.com/ ZHAOTING/dialog-processing. Tianyu Zhao 0001, Divesh Lala, Tatsuya Kawahara |
ACL | 2 |
| 2020 | Job Interviewer Android with Elaborate Follow-up Question GenerationabstractA job interview is a domain that takes advantage of an android robot's human-like appearance and behaviors. In this work, our goal is to implement a system in which an android plays the role of an interviewer so that users may practice for a real job interview. Our proposed system generates elaborate follow-up questions based on responses from the interviewee. We conducted an interactive experiment to compare the proposed system against a baseline system that asked only fixed-form questions. We found that this system was significantly better than the baseline system with respect to the impression of the interview and the quality of the questions, and that the presence of the android interviewer was enhanced by the follow-up questions. We also found a similar result when using a virtual agent interviewer, except that presence was not enhanced. Koji Inoue, Kohei Hara, Divesh Lala, Kenta Yamamoto, Shizuka Nakamura, Katsuya Takanashi, Tatsuya Kawahara |
ICMI | 3 |
| 2020 | An Attentive Listening System with Android ERICA: Comparison of Autonomous and WOZ InteractionsabstractWe describe an attentive listening system for the autonomous android robot ERICA.The proposed system generates several types of listener responses: backchannels, repeats, elaborating questions, assessments, generic sentimental responses, and generic responses.In this paper, we report a subjective experiment with 20 elderly people.First, we evaluated each system utterance excluding backchannels and generic responses, in an offline manner.It was found that most of the system utterances were linguistically appropriate, and they elicited positive reactions from the subjects.Furthermore, 58.2% of the responses were acknowledged as being appropriate listener responses.We also compared the proposed system with a WOZ system where a human operator was operating the robot.From the subjective evaluation, the proposed system achieved comparable scores in basic skills of attentive listening such as encouragement to talk, focused on the talk, and actively listening.It was also found that there is still a gap between the system and the WOZ for more sophisticated skills such as dialogue understanding, showing interest, and empathy towards the user. Koji Inoue, Divesh Lala, Kenta Yamamoto, Shizuka Nakamura, Katsuya Takanashi, Tatsuya Kawahara |
SIGdial | 2 |
| 2019 | Smooth Turn-taking by a Robot Using an Online Continuous Model to Generate Turn-taking CuesabstractTurn-taking in human-robot interaction is a crucial part of spoken dialogue systems, but current models do not allow for human-like turn-taking speed seen in natural conversation. In this work we propose combining two independent prediction models. A continuous model predicts the upcoming end of the turn in order to generate gaze aversion and fillers as turn-taking cues. This prediction is done while the user is speaking, so turn-taking can be done with little silence between turns, or even overlap. Once a speech recognition result has been received at a later time, a second model uses the lexical information to decide if or when the turn should actually be taken. We constructed the continuous model using the speaker’s prosodic features as inputs and evaluated its online performance. We then conducted a subjective experiment in which we implemented our model in an android robot and asked participants to compare it to one without turn-taking cues, which produces a response when a speech recognition result is received. We found that using both gaze aversion and a filler was preferred when the continuous model correctly predicted the upcoming end of turn, while using only gaze aversion was better if the prediction was wrong. Divesh Lala, Koji Inoue, Tatsuya Kawahara |
ICMI | 1 |
| 2019 | ERICA and WikiTalkabstractThe demo shows ERICA, a highly realistic female android robot, and WikiTalk, an application that helps robots to talk about thousands of topics using information from Wikipedia. The combination of ERICA and WikiTalk results in more natural and engaging human-robot conversations. Divesh Lala, Graham Wilcock, Kristiina Jokinen, Tatsuya Kawahara |
IJCAI | 1 |
| 2019 | Analysis of Effect and Timing of Fillers in Natural Turn-Taking
Divesh Lala, Shizuka Nakamura, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2018 | Audio-Visual Conversation Analysis by Smart Posterboard and Humanoid RobotabstractThis paper addresses audio-visual signal processing for conversation analysis, which involves multi-modal behavior detection and mental-state recognition. We have investigated prediction of turn-taking by the audience in a poster session from their multi-modal behaviors, and found out that the eye-gaze provides an important cue compared with head nodding and verbal backchannels. This finding has been applied to audio-visual speaker diarization by combining eye-gaze information. We are now investigating engagement recognition in human-robot interaction based on the same scheme. Robust and realtime detection of laughing, backchannels and nodding is realized based on LSTM-CTC. We introduce a latent “character” model to cope with the subjectivity and variations of engagement annotations. Experimental evaluations demonstrate that (1) the latent character model is effective, (2) automatic behavior detection is robust and does not degrade the engagement recognition accuracy, and (3) the eye-gaze is the most important feature among others. Tatsuya Kawahara, Koji Inoue, Divesh Lala, Katsuya Takanashi |
ICASSP | 3 |
| 2018 | Evaluation of Real-time Deep Learning Turn-taking Models for Multiple Dialogue ScenariosabstractThe task of identifying when to take a conversational turn is an important function of spoken dialogue systems. The turn-taking system should also ideally be able to handle many types of dialogue, from structured conversation to spontaneous and unstructured discourse. Our goal is to determine how much a generalized model trained on many types of dialogue scenarios would improve on a model trained only for a specific scenario. To achieve this goal we created a large corpus of Wizard-of-Oz conversation data which consisted of several different types of dialogue sessions, and then compared a generalized model with scenario-specific models. For our evaluation we go further than simply reporting conventional metrics, which we show are not informative enough to evaluate turn-taking in a real-time system. Instead, we process results using a performance curve of latency and false cut-in rate, and further improve our model's real-time performance using a finite-state turn-taking machine. Our results show that the generalized model greatly outperformed the individual model for attentive listening scenarios but was worse in job interview scenarios. This implies that a model based on a large corpus is better suited to conversation which is more user-initiated and unstructured. We also propose that our method of evaluation leads to more informative performance metrics in a real-time system. Divesh Lala, Koji Inoue, Tatsuya Kawahara |
ICMI | 1 |
| 2018 | Engagement Recognition in Spoken Dialogue via Neural Network by Aggregating Different Annotators' Models
Koji Inoue, Divesh Lala, Katsuya Takanashi, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2017 | Utterance Behavior of Users While Playing Basketball with a Virtual Teammate
Divesh Lala, Yuanchao Li, Tatsuya Kawahara |
ICAART (1) | 1 |
| 2017 | Attentive listening system with backchanneling, response generation and flexible turn-takingabstractAttentive listening systems are designed to let people, especially senior people, keep talking to maintain communication ability and mental health.This paper addresses key components of an attentive listening system which encourages users to talk smoothly.First, we introduce continuous prediction of end-of-utterances and generation of backchannels, rather than generating backchannels after end-point detection of utterances.This improves subjective evaluations of backchannels.Second, we propose an effective statement response mechanism which detects focus words and responds in the form of a question or partial repeat.This can be applied to any statement.Moreover, a flexible turn-taking mechanism is designed which uses backchannels or fillers when the turnswitch is ambiguous.These techniques are integrated into a humanoid robot to conduct attentive listening.We test the feasibility of the system in a pilot experiment and show that it can produce coherent dialogues during conversation. Divesh Lala, Pierrick Milhorat, Koji Inoue, Masanari Ishida, Katsuya Takanashi, Tatsuya Kawahara |
SIGDIAL Conference | 1 |
| 2017 | A data-driven passing interaction model for embodied basketball agents
Divesh Lala, Toyoaki Nishida |
J. Intell. Inf. Syst. | 1 |
| 2016 | Heat map visualization of multi-slice medical images through correspondence matching of video framesabstractVisual inspection of medical imagery such as MRI and CT scans is a major task for medical professionals who must diagnose and treat patients without error. Given this goal, visualizing search behavior patterns used to recognize abnormalities in these images is of interest. In this paper we describe the development of a system which automatically generates multiple image-dependent heat maps from eye gaze data of users viewing medical image slices. This system only requires the use of a non-wearable eye gaze tracker and video capturing system. The main automated features are the identification of a medical image slice located inside a video frame and calculation of the correspondence between display screen and raw image eye gaze locations. We propose that the system can be used for eye gaze analysis and diagnostic training in the medical field. Divesh Lala, Atsushi Nakazawa |
ETRA | 1 |
| 2016 | Multimodal interaction with the autonomous Android ERICAabstractWe demonstrate an interactive conversation with an android named ERICA. In this demonstration the user can converse with ERICA on a number of topics. We demonstrate both the dialog management system and the eye gaze behavior of ERICA used for indicating attention and turn taking. Divesh Lala, Pierrick Milhorat, Koji Inoue, Tianyu Zhao 0001, Tatsuya Kawahara |
ICMI | 1 |
| 2016 | Managing Dialog and Joint Actions for Virtual Basketball Teammates
Divesh Lala, Tatsuya Kawahara |
IVA | 1 |
| 2016 | Talking with ERICA, an autonomous androidabstractWe demonstrate dialogues with an autonomous android ERICA, who has an appearance like a human being.Currently, ERICA plays two social roles: a laboratory guide and a counselor.It is designed to follow the protocols of human dialogue to make the user comfortable: (1) having a chat before the main talk, (2) proactively asking questions, and (3) conveying proper feedbacks.The combination of the human-like appearance and the appropriate behaviors according to her social roles allows for symbiotic human-robot interaction. Koji Inoue, Pierrick Milhorat, Divesh Lala, Tianyu Zhao 0001, Tatsuya Kawahara |
SIGDIAL Conference | 3 |
| 2015 | Synthetic Evidential Study as Augmented Collective Thought Process - Preliminary Report
Toyoaki Nishida, Masakazu Abe, Takashi Ookaki, Divesh Lala, Sutasinee Thovutikul, Hengjie Song, Yasser Mohammad, Christian Nitschke, Yoshimasa Ohmoto, Atsushi Nakazawa, Takaaki Shochi, Jean-Luc Rouas, Aurélie Bugeau, Fabien Lotte, Zuheng Ming, Geoffrey Letournel, Marine Guerry, Dominique Fourer |
ACIIDS (1) | 4 |
| 2015 | User Perceptions of Communicative and Task-competent Agents in a Virtual Basketball Game
Divesh Lala, Christian Nitschke, Toyoaki Nishida |
ICAART (1) | 1 |
| 2015 | An imitation-based approach to represent common ground in a virtual basketball agent team mateabstractOne barrier to creating truly intelligent autonomous virtual characters is the lack of common ground knowledge contained between the agent and its human interaction partner. Predefining this knowledge for the agent is infeasible so such agents generally can only interact within a limited task domain. This is particularly true for conversational agents, who must handle a wide range of topics. Rather than attempting to tackle the difficult problem of conversation, we instead create an environment where the common ground knowledge is the actions to perform in a virtual basketball game. These actions are not internally predefined, but instead extracted and imitated during interaction. Imitation is crucial because common ground can only be said to exist if there is evidence that the parties have this knowledge [Clark 1996]. Divesh Lala, Toyoaki Nishida |
VRST | 1 |