VLDB 2026 Research / reviewers in the wild / expert
Catherine Lai
dblp:65/400
· DBLP profile ↗
46ranked-venue papers
9as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 8 first-author · 20 since 2021Artificial intelligence and machine learning · 30 · 8 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MaiChat: A Text-based Dialogue Corpus Rich in Conversational Features
Mai Hoang Dao, Catherine Lai, Peter Bell 0001 |
LREC | 2 |
| 2025 | Can We "Cherry-Pick"? Investigating Multiple Renditions from a Generative Speech Synthesis ModelabstractGenerative Speech Models (GSMs) have seen a surge in popularity due to their ability to generate diverse and high-quality speech. Evaluating models that generate many different renditions for a given input sentence presents a new challenge. Listening tests are still the gold standard for evaluating synthetic speech, but current paradigms only consider a single arbitrary rendition: this fails to give a complete picture of best/typical/worst-rendition performance. We propose a general framework for evaluating and deploying generative speech models. This involves selecting amongst renditions using a sequence of filtering or ranking steps, each using either an objective or subjective (listening) method. The framework is not tied to a particular generative model, and so could be applied to any such model. In this paper, we provide a demonstration of a simple version of this framework which would apply to use-cases where best-rendition performance matters. We explore the concept of "cherry-picking", and ask the question "Is there a rendition that is consistently preferred above all others by listeners?". In a subjective listening test, participants ranked several renditions of the same sentence, from which we measured the prevalence of exceptional renditions. We find that there is indeed a preferred rendition in many, but not all cases. Our framework is flexible. In particular, the use of listeners is optional. In future, they could be replaced with model-based objective measures, for example. Adaeze Adigwe, Sarenne Wallbridge, Zehai Tu, Catherine Lai |
ICASSP | 4 |
| 2025 | Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-LabelingabstractThe lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervised Learning (SSL) framework, introducing a novel multi-view pseudo-labeling method that leverages both acoustic and linguistic characteristics to select the most confident data for training the classification model. Acoustically, unlabeled data are compared to labeled data using the Fréchet audio distance, calculated from embeddings generated by multiple audio encoders. Linguistically, large language models are prompted to revise automatic speech recognition transcriptions and predict labels based on our proposed task-specific knowledge. High-confidence data are identified when pseudo-labels from both sources align, while mismatches are treated as low-confidence data. A bimodal classifier is then trained to iteratively label the low-confidence data until a predefined criterion is met. We evaluate our SSL framework on emotion recognition and dementia detection tasks. Experimental results demonstrate that our method achieves competitive performance compared to fully supervised learning using only 30% of the labeled data and significantly outperforms two selected baselines. Yuanchao Li, Zixing Zhang 0001, Jing Han 0010, Peter Bell 0001, Catherine Lai |
ICASSP | 5 |
| 2025 | Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error CorrectionabstractAnnotating and recognizing speech emotion using prompt engineering has recently emerged with the advancement of Large Language Models (LLMs), yet its efficacy and reliability remain questionable. In this paper, we conduct a systematic study on this topic, beginning with the proposal of novel prompts that incorporate emotion-specific knowledge from acoustics, linguistics, and psychology. Subsequently, we examine the effectiveness of LLM-based prompting on Automatic Speech Recognition (ASR) transcription, contrasting it with ground-truth transcription. Furthermore, we propose a Revise-Reason-Recognize prompting pipeline for robust LLM-based emotion recognition from spoken language with ASR errors. Additionally, experiments on context-aware learning, in-context learning, and instruction tuning are performed to examine the usefulness of LLM training schemes in this direction. Finally, we investigate the sensitivity of LLMs to minor prompt variations. Experimental results demonstrate the efficacy of the emotion-specific prompts, ASR error correction, and LLM training schemes for LLM-based emotion recognition. Our study aims to refine the use of LLMs in emotion recognition and related domains. Yuanchao Li, Yuan Gong 0001, Chao-Han Huck Yang, Peter Bell 0001, Catherine Lai |
ICASSP | 5 |
| 2025 | HRAI 2025: The 1st Workshop on Holistic and Responsible Affective IntelligenceabstractThe ICMI 2025 Workshop on Holistic and Responsible Affective Intelligence (HRAI 2025) aims to advance research in affective intelligence by fostering discussions on the holistic development of affective computing and the ethical challenges it entails. The workshop aims to strengthen interdisciplinary connections within the affective computing community, promoting better integration of methodologies and enhancing real-world applicability. By tackling both technical and ethical issues, HRAI 2025 aspires to shape the future of affective AI, ensuring it is not only powerful but also fair, safe, and socially responsible. Yuanchao Li, Dimitris Kollias, Guillaume Chanel, Marios A. Fanourakis, Michal Muszynski, Brandon M. Booth, Leimin Tian, Madhawa Perera, Catherine Lai, Huili Chen |
ICMI | 9 |
| 2025 | Conveying Gender Through Speech: Insights from Trans Men
Alice Ross, Cliodhna Hughes, Eddie L. Ungless, Catherine Lai |
INTERSPEECH | 4 |
| 2025 | Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
Sarenne Wallbridge, Christoph Minixhofer, Catherine Lai, Peter Bell 0001 |
INTERSPEECH | 3 |
| 2024 | Language Technologies as If People Mattered: Centering Communities in Language Technology DevelopmentabstractIn this position paper we argue that researchers interested in language and/or language technologies should attend to challenges of linguistic and algorithmic injustice together with language communities. We put forward that this can be done by drawing together diverse scholarly and experiential insights, building strong interdisciplinary teams, and paying close attention to the wider social, cultural and historical contexts of both language communities and the technologies we aim to develop. Nina Markl, Lauren Hall-Lew, Catherine Lai |
LREC/COLING | 3 |
| 2024 | Well, what can you do with messy data? Exploring the prosody and pragmatic function of the discourse marker "well" with found data and speech synthesis
Johannah O'Mahony, Catherine Lai, Éva Székely |
INTERSPEECH | 2 |
| 2024 | Speech Emotion Recognition With ASR Transcripts: a Comprehensive Study on Word Error Rate and Fusion TechniquesabstractText data is commonly utilized as a primary input to enhance Speech Emotion Recognition (SER) performance and reliability. However, the reliance on human-transcribed text in most studies impedes the development of practical SER systems, creating a gap between in-lab research and real-world scenarios where Automatic Speech Recognition (ASR) serves as the text source. Hence, this study benchmarks SER performance using ASR transcripts with varying Word Error Rates (WERs) from eleven models on three well-known corpora: IEMOCAP, CMU-MOSI, and MSP-Podcast. Our evaluation includes both text-only and bimodal SER with six fusion techniques, aiming for a comprehensive analysis that uncovers novel findings and challenges faced by current SER research. Additionally, we propose a unified ASR error-robust framework integrating ASR error correction and modality-gated fusion, achieving lower WER and higher SER results compared to the best-performing ASR transcript. These findings provide insights into SER with ASR assistance, especially for real-world applications. Yuanchao Li, Peter Bell 0001, Catherine Lai |
SLT | 3 |
| 2024 | Crossmodal ASR Error Correction With Discrete Speech UnitsabstractASR remains unsatisfactory in scenarios where the speaking style diverges from that used to train ASR systems, resulting in erroneous transcripts. To address this, ASR Error Correction (AEC), a post-ASR processing approach, is required. In this work, we tackle an understudied issue: the Low-Resource Out-of-Domain (LROOD) problem, by investigating crossmodal AEC on very limited downstream data with 1-best hypothesis transcription. We explore pretraining and fine-tuning strategies and uncover an ASR domain discrepancy phenomenon, shedding light on appropriate training schemes for LROOD data. Moreover, we propose the incorporation of discrete speech units to align with and enhance the word embeddings for improving AEC quality. Results from multiple corpora and several evaluation metrics demonstrate the feasibility and efficacy of our proposed AEC approach on LROOD data as well as its generalizability and superiority on large-scale data. Finally, a study on speech emotion recognition confirms that our model produces ASR error-robust transcripts suitable for downstream applications. Yuanchao Li, Pinzhen Chen, Peter Bell 0001, Catherine Lai |
SLT | 4 |
| 2024 | Large Language Model Based Generative Error Correction: A Challenge and Baselines For Speech Recognition, Speaker Tagging, and Emotion RecognitionabstractGiven recent advances in generative AI technology, a key question is how large language models (LLMs) can enhance acoustic modeling tasks using text decoding results from a frozen, pretrained automatic speech recognition (ASR) model. To explore new capabilities in language modeling for speech processing, we introduce the generative speech transcription error correction (GenSEC) challenge. This challenge comprises three post-ASR language modeling tasks: (i) post-ASR transcription correction, (ii) speaker tagging, and (iii) emotion recognition. These tasks aim to emulate future LLM-based agents handling voice-based interfaces while remaining accessible to a broad audience by utilizing open pretrained language models or agent-based APIs. We also discuss insights from baseline evaluations, as well as lessons learned for designing future evaluations. Chao-Han Huck Yang, Taejin Park, Yuan Gong 0001, Yuanchao Li, Zhehuai Chen, Chen Chen 0075, Kunal Dhawan, Piotr Zelasko, Chao Zhang 0031, Yun-Nung Chen, Yu Tsao 0001, Jagadeesh Balam, Boris Ginsburg, Sabato Marco Siniscalchi, Chng Eng Siong, Peter Bell 0001, Catherine Lai, Shinji Watanabe 0001, Andreas Stolcke |
SLT | 19 |
| 2023 | Do dialogue representations align with perception? An empirical studyabstractThere has been a surge of interest regarding the alignment of large-scale language models with human language comprehension behaviour.The majority of this research investigates comprehension behaviours from reading isolated, written sentences.We propose studying the perception of dialogue, focusing on an intrinsic form of language use: spoken conversations.Using the task of predicting upcoming dialogue turns, we ask whether turn plausibility scores produced by state-of-the-art language models correlate with human judgements.We find a strong correlation for some but not all models: masked language models produce stronger correlations than autoregressive models.In doing so, we quantify human performance on the response selection task for open-domain spoken conversation.To the best of our knowledge, this is the first such quantification.We find that response selection performance can be used as a coarse proxy for the strength of correlation with human judgements, however humans and models make different response selection mistakes.The model which produces the strongest correlation also outperforms human response selection performance.Through ablation studies, we show that pre-trained language models provide a useful basis for turn representations; however, finegrained contextualisation, inclusion of dialogue structure information, and fine-tuning towards response selection all boost response selection accuracy by over 30 absolute points. Sarenne Wallbridge, Peter Bell 0001, Catherine Lai |
EACL | 3 |
| 2023 | Multimodal Dyadic Impression Recognition via Listener Adaptive Cross-Domain FusionabstractAs a sub-branch of affective computing, impression recognition, e.g., perception of speaker characteristics such as warmth or competence, is potentially a critical part of both human-human conversations and spoken dialogue systems. Most research has studied impressions only from the behaviors expressed by the speaker or the response from the listener, yet ignored their latent connection. In this paper, we perform impression recognition using a proposed listener adaptive cross-domain architecture, which consists of a listener adaptation function to model the causality between speaker and listener behaviors and a cross-domain fusion function to strengthen their connection. The experimental evaluation on the dyadic IMPRESSION dataset verified the efficacy of our method, producing concordance correlation coefficients of 78.8% and 77.5% in the competence and warmth dimensions, outperforming previous studies. The proposed method is expected to be generalized to similar dyadic interaction scenarios in affective computing. Yuanchao Li, Peter Bell 0001, Catherine Lai |
ICASSP | 3 |
| 2023 | Transfer Learning for Personality Perception via Speech Emotion RecognitionabstractHolistic perception of affective attributes is an important human perceptual ability. However, this ability is far from being realized in current affective computing, as not all of the attributes are well studied and their interrelationships are poorly understood. In this work, we investigate the relationship between two affective attributes: personality and emotion, from a transfer learning perspective. Specifically, we transfer Transformer-based and wav2vec-based emotion recognition models to perceive personality from speech across corpora. Compared with previous studies, our results show that transferring emotion recognition is effective for personality perception. Moreoever, this allows for better use and exploration of small personality corpora. We also provide novel findings on the relationship between personality and emotion that will aid future research on holistic affect recognition. Yuanchao Li, Peter Bell 0001, Catherine Lai |
INTERSPEECH | 3 |
| 2023 | ASR and Emotional Speech: A Word-Level Investigation of the Mutual Impact of Speech and Emotion RecognitionabstractIn Speech Emotion Recognition (SER), textual data is often used alongside audio signals to address their inherent variability. However, the reliance on human annotated text in most research hinders the development of practical SER systems. To overcome this challenge, we investigate how Automatic Speech Recognition (ASR) performs on emotional speech by analyzing the ASR performance on emotion corpora and examining the distribution of word errors and confidence scores in ASR transcripts to gain insight into how emotion affects ASR. We utilize four ASR systems, namely Kaldi ASR, wav2vec, Conformer, and Whisper, and three corpora: IEMOCAP, MOSI, and MELD to ensure generalizability. Additionally, we conduct text-based SER on ASR transcripts with increasing word error rates to investigate how ASR affects SER. The objective of this study is to uncover the relationship and mutual impact of ASR and SER, in order to facilitate ASR adaptation to emotional speech and the use of SER in real world. Yuanchao Li, Zeyu Zhao 0004, Ondrej Klejch, Peter Bell 0001, Catherine Lai |
INTERSPEECH | 5 |
| 2023 | Everyone has an accentabstractIn this paper, we consider how the notion "accent" in particular in the context of "accented speech" has been discussed in Interspeech publications between 2004 and 2022. We contrast the way speech technology research published in the conference has conceptualised these terms with their usage in linguistics. The point of this comparison is to: highlight significant inter-disciplinary differences in the way apparently core terms are used, discuss disadvantages of using inexact language in research, and encourage researchers to be more mindful about the use of particular short-hands. Nina Markl, Catherine Lai |
INTERSPEECH | 2 |
| 2023 | Quantifying the perceptual value of lexical and non-lexical channels in speechabstractSpeech is a fundamental means of communication that can be seen to provide two channels for transmitting information: the lexical channel of \textit{which} words are said, and the non-lexical channel of \textit{how} they are spoken. Both channels shape listener expectations of upcoming communication; however, directly quantifying their relative effect on expectations is challenging. Previous attempts require spoken variations of lexically-equivalent dialogue turns or conspicuous acoustic manipulations. This paper introduces a generalised paradigm to study the value of non-lexical information in dialogue across unconstrained lexical content.By quantifying the perceptual value of the non-lexical channel with both accuracy and entropy reduction, we show that non-lexical information produces a consistent effect on expectations of upcoming dialogue: even when it leads to poorer discriminative turn judgements than lexical content alone, it yields higher consensus among participants. Sarenne Wallbridge, Peter Bell 0001, Catherine Lai |
INTERSPEECH | 3 |
| 2023 | Synthesising Personality with Neural Speech SynthesisabstractMatching the personality of conversational agents to the personality of the user can significantly improve the user experience, with many successful examples in text-based chatbots.It is also important for a voice-based system to be able to alter the personality of the speech as perceived by the users.In this pilot study, fifteen voices were rated using Big Five personality traits.Five content-neutral sentences were chosen for the listening tests.The audio data, together with two rated traits (Extroversion and Agreeableness), were used to train a neural speech synthesiser based on one male and one female voices.The effect of altering the personality trait features was evaluated by a second listening test.Both perceived extroversion and agreeableness in the synthetic voices were affected significantly.The controllable range was limited due to a lack of variance in the source audio data.The perceived personality traits correlated with each other and with the naturalness of the speech. Shilin Gao, Matthew P. Aylett, David A. Braude, Catherine Lai |
SIGDIAL | 4 |
| 2022 | Alzheimer's Dementia Detection through Spontaneous Dialogue with Proactive Robotic ListenersabstractAs the aging of society continues to accelerate, Alzheimer's Disease (AD) has received more and more attention from not only medical but also other fields, such as computer science, over the past decade. Since speech is considered one of the effective ways to diagnose cognitive decline, AD detection from speech has emerged as a hot topic. Nevertheless, such approaches fail to tackle several key issues: 1) AD is a complex neurocognitive disorder which means it is inappropriate to conduct AD detection using utterance information alone while ignoring dialogue infor-mation; 2) Utterances of AD patients contain many disfluencies that affect speech recognition yet are helpful to diagnosis; 3) AD patients tend to speak less, causing dialogue breakdown as the disease progresses. This fact leads to a small number of utterances, which may cause detection bias. Therefore, in this paper, we propose a novel AD detection architecture consisting of two major modules: an ensemble AD detector and a proactive listener. This architecture can be embedded in the dialogue system of conversational robots for healthcare. Yuanchao Li, Catherine Lai, Divesh Lala, Koji Inoue, Tatsuya Kawahara |
HRI | 2 |
| 2022 | Fusing ASR Outputs in Joint Training for Speech Emotion RecognitionabstractAlongside acoustic information, linguistic features based on speech transcripts have been proven useful in Speech Emotion Recognition (SER). However, due to the scarcity of emotion labelled data and the difficulty of recognizing emotional speech, it is hard to obtain reliable linguistic features and models in this research area. In this paper, we propose to fuse Automatic Speech Recognition (ASR) outputs into the pipeline for joint training SER. The relationship between ASR and SER is understudied, and it is unclear what and how ASR features benefit SER. By examining various ASR outputs and fusion methods, our experiments show that in joint ASR-SER training, incorporating both ASR hidden and text output using a hierarchical co-attention fusion approach improves the SER performance the most. On the IEMOCAP corpus, our approach achieves 63.4% weighted accuracy, which is close to the baseline results achieved by combining ground-truth transcripts. In addition, we also present novel word error rate analysis on IEMOCAP and layer-difference analysis of the Wav2vec 2.0 model to better understand the relationship between ASR and SER. Yuanchao Li, Peter Bell 0001, Catherine Lai |
ICASSP | 3 |
| 2022 | Combining conversational speech with read speech to improve prosody in Text-to-Speech synthesisabstractFor isolated utterances, speech synthesis quality has improved immensely thanks to the use of sequence-to-sequence models. However, these models are generally trained on read speech and fail to generalise to unseen speaking styles. Recently, more re-search is focused on the synthesis of expressive and conversa-tional speech. Conversational speech contains many prosodic phenomena that are not present in read speech. We would like to learn these prosodic patterns from data, but unfortunately, many large conversational corpora are unsuitable for speech synthesis due to low audio quality. We investigate whether a data mixing strategy can improve conversational prosody for a target voice based on monologue data from audiobooks by adding real con-versational data from podcasts. We filter the podcast data to create a set of 26k question and answer pairs. We evaluate two FastPitch models: one trained on 20 hours of monologue speech from a single speaker, and another trained on 5 hours of monologue speech from that speaker plus 15 hours of ques-tions and answers spoken by nearly 15k speakers. Results from three listening tests show that the second model generates more preferred question prosody. Johannah O'Mahony, Catherine Lai, Simon King 0001 |
INTERSPEECH | 2 |
| 2022 | Voice Puppetry with FastPitch
Emelie Van De Vreken, Korin Richmond, Catherine Lai |
INTERSPEECH | 3 |
| 2022 | Investigating perception of spoken dialogue acceptability through surprisalabstractSurprisal is used throughout computational psycholinguistics to model a range of language processing behaviour. There is growing evidence that language model (LM) estimates of surprisal correlate with human performance on a range of written language comprehension tasks. Although communicative interaction is arguably the primary form of language use, most studies of surprisal are based on monological, written data. Towards the goal of understanding perception in spontaneous, natural language, we present an exploratory investigation into whether the relationship between human comprehension behaviour and LM-estimated surprisal holds when applied to dialogue, considering both written dialogue, and the lexical component of spoken dialogue. We use a novel judgement task of dialogue utterance acceptability to ask two questions: “How well can people make predictions about written dialogue and transcripts of spoken dialogue?” and “Does surprisal correlate with these acceptability judgements?”. We demonstrate that people can make accurate predictions about upcoming dialogue and that their ability differs between spoken transcripts and written conversation. We investigate the relationship between global and local operationalisations of surprisal and human acceptability judgements, finding a combination of both to provide the most predictive power Sarenne Wallbridge, Catherine Lai, Peter Bell 0001 |
INTERSPEECH | 2 |
| 2022 | Exploration of a Self-Supervised Speech Model: A Study on Emotional CorporaabstractSelf-supervised speech models have grown fast during the past few years and have proven feasible for use in various downstream tasks. Some recent work has started to look at the characteristics of these models, yet many concerns have not been fully addressed. In this work, we conduct a study on emotional corpora to explore a popular self-supervised model - wav2vec 2.0. Via a set of quantitative analysis, we mainly demonstrate that: 1) wav2vec 2.0 appears to discard paralinguistic information that is less useful for word recognition purposes; 2) for emotion recognition, representations from the middle layer alone perform as well as those derived from layer averaging, while the final layer results in the worst performance in some cases; 3) current self-supervised models may not be the optimal solution for downstream tasks that make use of non-lexical features. Our work provides novel findings that will aid future research in this area and theoretical basis for the use of existing models. Yuanchao Li, Yumnah Mohamied, Peter Bell 0001, Catherine Lai |
SLT | 4 |
| 2021 | It's Not What You Said, it's How You Said it: Discriminative Perception of Speech as a Multichannel Communication SystemabstractPeople convey information extremely effectively through spoken interaction using multiple channels of information transmission: the lexical channel of what is said, and the non-lexical channel of how it is said. We propose studying human perception of spoken communication as a means to better understand how information is encoded across these channels, focusing on the question 'What characteristics of communicative context affect listener's expectations of speech?'. To investigate this, we present a novel behavioural task testing whether listeners can discriminate between the true utterance in a dialogue and utterances sampled from other contexts with the same lexical content. We characterize how perception - and subsequent discriminative capability - is affected by different degrees of additional contextual information across both the lexical and non-lexical channel of speech. Results demonstrate that people can effectively discriminate between different prosodic realisations, that non-lexical context is informative, and that this channel provides more salient information than the lexical channel, highlighting the importance of the non-lexical channel in spoken interaction. Sarenne Wallbridge, Peter Bell 0001, Catherine Lai |
Interspeech | 3 |
| 2021 | Recognizing Induced Emotions of Movie Audiences from Multimodal InformationabstractRecognizing emotional reactions of movie audiences to affective movie content is a challenging task in affective computing. Previous research on induced emotion recognition has mainly focused on using audio-visual movie content. Nevertheless, the relationship between the perceptions of the affective movie content (perceived emotions) and the emotions evoked in the audiences (induced emotions) is unexplored. In this work, we studied the relationship between perceived and induced emotions of movie audiences. Moreover, we investigated multimodal modelling approaches to predict movie induced emotions from movie content based features, as well as physiological and behavioral reactions of movie audiences. To carry out analysis of induced and perceived emotions, we first extended an existing database for movie affect analysis by annotating perceived emotions in a crowd-sourced manner. We find that perceived and induced emotions are not always consistent with each other. In addition, we show that perceived emotions, movie dialogues, and aesthetic highlights are discriminative for movie induced emotion recognition besides spectators' physiological and behavioral reactions. Also, our experiments revealed that induced emotion recognition could benefit from including temporal information and performing multimodal fusion. Moreover, our work deeply investigated the gap between affective content analysis and induced emotion recognition by gaining insight into the relationships between aesthetic highlights, induced emotions, and perceived emotions. Michal Muszynski, Leimin Tian, Catherine Lai, Johanna D. Moore, Theodoros Kostoulas, Patrizia Lombardo, Thierry Pun, Guillaume Chanel |
IEEE Trans. Affect. Comput. | 3 |
| 2020 | Integrating lexical and prosodic features for automatic paragraph segmentation
Catherine Lai, Mireia Farrús, Johanna D. Moore |
Speech Commun. | 1 |
| 2019 | Detecting Topic-Oriented Speaker Stance in Conversational SpeechabstractBeing able to detect topics and speaker stances in conversations is a key requirement for developing spoken language understanding systems that are personalized and adaptive. In this work, we explore how topic-oriented speaker stance is expressed in conversational speech. To do this, we present a new set of topic and stance annotations of the CallHome corpus of spontaneous dialogues. Specifically, we focus on six stances-positivity, certainty, surprise, amusement, interest, and comfort-which are useful for characterizing important aspects of a conversation, such as whether a conversation is going well or not. Based on this, we investigate the use of neural network models for automatically detecting speaker stance from speech in multi-turn, multi-speaker contexts. In particular, we examine how performance changes depending on how input feature representations are constructed and how this is related to dialogue structure. Our experiments show that incorporating both lexical and acoustic features is beneficial for stance detection. However, we observe variation in whether using hierarchical models for encoding lexical and acoustic information improves performance, suggesting that some aspects of speaker stance are expressed more locally than others. Overall, our findings highlight the importance of modelling interaction dynamics and non-lexical content for stance detection. Catherine Lai, Beatrice Alex, Johanna D. Moore, Leimin Tian, Tatsuro Hori, Gianpiero Francesca |
INTERSPEECH | 1 |
| 2018 | Group Interaction Frontiers in TechnologyabstractAnalysis of group interaction and team dynamics is an important topic in a wide variety of fields, owing to the amount of time that individuals typically spend in small groups for both professional and personal purposes, and given how crucial group cohesion and productivity are to the success of businesses and other organizations. This fact is attested by the rapid growth of fields such as People Analytics and Human Resource Analytics, which in turn have grown out of many decades of research in social psychology, organizational behaviour, computing, and network science, amongst other fields. The goal of this workshop is to bring together researchers from diverse fields related to group interaction, team dynamics, people analytics, multi-modal speech and language processing, social psychology, and organizational behaviour. Gabriel Murray, Hayley Hung, Joann Keyton, Catherine Lai, Nale Lehmann-Willenbrock, Catharine Oertel |
ICMI | 4 |
| 2017 | Recognizing induced emotions of movie audiences: Are induced and perceived emotions the same?abstractPredicting the emotional response of movie audiences to affective movie content is a challenging task in affective computing. Previous work has focused on using audiovisual movie content to predict movie induced emotions. However, the relationship between the audience's perceptions of the affective movie content (perceived emotions) and the emotions evoked in the audience (induced emotions) remains unexplored. In this work, we address the relationship between perceived and induced emotions in movies, and identify features and modelling approaches effective for predicting movie induced emotions. First, we extend the LIRIS-ACCEDE database by annotating perceived emotions in a crowd-sourced manner, and find that perceived and induced emotions are not always consistent. Second, we show that dialogue events and aesthetic highlights are effective predictors of movie induced emotions. In addition to movie based features, we also study physiological and behavioural measurements of audiences. Our experiments show that induced emotion recognition can benefit from including temporal context and from including multimodal information. Our study bridges the gap between affective content analysis and induced emotion prediction. Leimin Tian, Michal Muszynski, Catherine Lai, Johanna D. Moore, Theodoros Kostoulas, Patrizia Lombardo, Thierry Pun, Guillaume Chanel |
ACII | 3 |
| 2017 | A System for Real Time Collaborative Transcription Correction
Peter Bell 0001, Joachim Fainberg, Catherine Lai, Mark Sinclair |
INTERSPEECH | 3 |
| 2017 | Using Prosody to Classify Discourse RelationsabstractComunicació presentada a: The 18th Annual Conference of the International Speech Communication Association (INTERSPEECH 2017), celebrada a Estocolm, Suència, del 20 al 24 d'agost de 2017. Janine Kleinhans, Mireia Farrús, Agustín Gravano, Juan Manuel Pérez, Catherine Lai, Leo Wanner |
INTERSPEECH | 5 |
| 2016 | Automatic Paragraph Segmentation with Lexical and Prosodic FeaturesabstractAs long-form spoken documents become more ubiquitous in everyday life, so does the need for automatic discourse segmentation in spoken language processing tasks. Although previous work has focused on broad topic segmentation, detection of finer-grained discourse units, such as paragraphs, is highly desirable for presenting and analyzing spoken content. To better understand how different aspects of speech cue these subtle discourse transitions, we investigate automatic paragraph segmentation of TED talks. We build lexical and prosodic paragraph segmenters using Support Vector Machines, AdaBoost, and Long Short Term Memory (LSTM) recurrent neural networks. In general, we find that induced cue words and supra-sentential prosodic features outperform features based on topical coherence, syntactic form and complexity. However, our best performance is achieved by combining a wide range of individually weak lexical and prosodic features, with the sequence modelling LSTM generally outperforming the other classifiers by a large margin. Moreover, we find that models that allow lower level interactions between different feature types produce better results than treating lexical and prosodic contributions as separate, independent information sources. Catherine Lai, Mireia Farrús, Johanna D. Moore |
INTERSPEECH | 1 |
| 2016 | Recognizing emotions in spoken dialogue with hierarchically fused acoustic and lexical featuresabstractAutomatic emotion recognition is vital for building natural and engaging human-computer interaction systems. Combining information from multiple modalities typically improves emotion recognition performance. In previous work, features from different modalities have generally been fused at the same level with two types of fusion strategies: Feature-Level fusion, which concatenates feature sets before recognition; and Decision-Level fusion, which makes the final decision based on outputs of the unimodal models. However, different features may describe data at different time scales or have different levels of abstraction. Cognitive Science research also indicates that when perceiving emotions, humans use information from different modalities at different cognitive levels and time steps. Therefore, we propose a Hierarchical fusion strategy for multimodal emotion recognition, which incorporates global or more abstract features at higher levels of its knowledge-inspired structure. We build multimodal emotion recognition models combining state-of-the-art acoustic and lexical features to study the performance of the proposed Hierarchical fusion. Experiments on two emotion databases of spoken dialogue show that this fusion strategy consistently outperforms both Feature-Level and Decision-Level fusion. The multimodal emotion recognition models using the Hierarchical fusion strategy achieved state-of-the-art performance on recognizing emotions in both spontaneous and acted dialogue. Leimin Tian, Johanna D. Moore, Catherine Lai |
SLT | 3 |
| 2015 | Emotion recognition in spontaneous and acted dialoguesabstractIn this work, we compare emotion recognition on two types of speech: spontaneous and acted dialogues. Experiments were conducted on the AVEC2012 database of spontaneous dialogues and the IEMOCAP database of acted dialogues. We studied the performance of two types of acoustic features for emotion recognition: knowledge-inspired disfluency and nonverbal vocalisation (DIS-NV) features, and statistical Low-Level Descriptor (LLD) based features. Both Support Vector Machines (SVM) and Long Short-Term Memory Recurrent Neural Networks (LSTM-RNN) were built using each feature set on each emotional database. Our work aims to identify aspects of the data that constrain the effectiveness of models and features. Our results show that the performance of different types of features and models is influenced by the type of dialogue and the amount of training data. Because DIS-NVs are less frequent in acted dialogues than in spontaneous dialogues, the DIS-NV features perform better than the LLD features when recognizing emotions in spontaneous dialogues, but not in acted dialogues. The LSTM-RNN model gives better performance than the SVM model when there is enough training data, but the complex structure of a LSTM-RNN model may limit its performance when there is less training data available, and may also risk over-fitting. Additionally, we find that long distance contexts may be more useful when performing emotion recognition at the word level than at the utterance level. Leimin Tian, Johanna D. Moore, Catherine Lai |
ACII | 3 |
| 2015 | Recognizing emotions in dialogues with acoustic and lexical featuresabstractAutomatic emotion recognition has long been a focus of Affective Computing. We aim at improving the performance of state-of-the-art emotion recognition in dialogues using novel knowledge-inspired features and modality fusion strategies. We propose features based on disfluencies and nonverbal vocalisations (DIS-NVs), and show that they are highly predictive for recognizing emotions in spontaneous dialogues. We also propose the hierarchical fusion strategy as an alternative to current feature-level and decision-level fusion. This fusion strategy combines features from different modalities at different layers in a hierarchical structure. It is expected to overcome limitations of feature-level and decision-level fusion by including knowledge on modality differences, while preserving information of each modality. Leimin Tian, Johanna D. Moore, Catherine Lai |
ACII | 3 |
| 2015 | A system for automatic broadcast news summarisation, geolocation and translation
Peter Bell 0001, Catherine Lai, Clare Llewellyn, Alexandra Birch, Mark Sinclair |
INTERSPEECH | 2 |
| 2015 | Towards automatic detection of reported speech in dialogue using prosodic cuesabstractThe phenomenon of reported speech -- whereby we quote the words, thoughts and opinions of others, or recount past dialogue -- is widespread in conversational speech. Detecting such quotations automatically has numerous applications: for example, in enhancing automatic transcription or spoken language understanding applications. However, the task is challenging, not least because lexical cues of quotations are frequently ambiguous or not present in spoken language. The aim of this paper is to identify potential prosodic cues of reported speech which could be used, along with the lexical ones, to automatically detect quotations and ascribe them to their rightful source, that is reconstructing their attribution relations. In order to do so we analyze SARC, a small corpus of telephone conversations that we have annotated with attribution relations. The results of the statistical analysis performed on the data show how variations in pitch, intensity, and timing features can be exploited as cues of quotations. Furthermore, we build a SVM classifier which integrates lexical and prosodic cues to automatically detect quotations in speech that performs significantly better than chance. Alessandra Cervone, Catherine Lai, Silvia Pareti, Peter Bell 0001 |
INTERSPEECH | 2 |
| 2014 | Word-Level Emotion Recognition Using High-Level Features
Johanna D. Moore, Leimin Tian, Catherine Lai |
CICLing (2) | 3 |
| 2014 | Incorporating lexical and prosodic information at different levels for meeting summarizationabstractThis paper investigates how prosodic features can be used to augment lexical features for meeting summarization. Auto-matic detection of summary-worthy content using non-lexical features, like prosody, has generally focused on features cal-culated over dialogue acts. However, a salient role of prosody is to distinguish important words within utterances. To exam-ine whether including more fine grained prosodic information can help extractive summarization, we perform experiments incorporating lexical and prosodic features at different levels. For ICSI and AMI meeting corpora, we find that combining prosodic and lexical features at a lower level has better AUROC performance than adding in prosodic features derived over di-alogue acts. ROUGE F-scores also show the same pattern for the ICSI data. However, the differences are less clear for the AMI data where the range of scores is much more compressed. In order to understand the relationship between the generated summaries and differences in standard measures, we look at the distribution of extracted content over meeting as well as sum-mary redundancy. We find that summaries based on dialogue act level prosody better reflect the amount of human annotated summary content in meeting segments, while summaries de-rived from prosodically augmented lexical features exhibit less redundancy. Index Terms: meeting summarization, prosody, dialogue. 1. Catherine Lai, Steve Renals |
INTERSPEECH | 1 |
| 2013 | Detecting summarization hot spots in meetings using group level involvement and turn-taking featuresabstractIn this paper we investigate how participant involvement and turn-taking features relate to extractive summarization of meeting dialogues. In particular, we examine whether automatically derived measures of group level involvement, like participation equality and turn-taking freedom, can help detect where summarization relevant meeting segments will be. Results show that classification using turn-taking features performed better than the majority class baseline for data from both AMI and ICSI meeting corpora in identifying whether meeting segments contain extractive summary dialogue acts. The feature based approach also provided better recall than using manual ICSI involvement hot spot annotations. Turn-taking features were additionally found to be predictive of the amount of extractive summary content in a segment. In general, we find that summary content decreases with higher participation equality and overlap, while it increases with the number of very short utterances. Differences in results between the AMI and ICSI data sets suggest how group participatory structure can be used to understand what makes meetings easy or difficult to summarize. Index Terms: Turn-taking, involvement, hot spots, summarization, meetings, dialogue Catherine Lai, Jean Carletta, Steve Renals |
INTERSPEECH | 1 |
| 2010 | What do you mean, you're uncertain?: the interpretation of cue words and rising intonation in dialogueabstractThis paper investigates how rising intonation affects the interpretation of cue words in dialogue. Both cue words and rising intonation express a range of speaker attitudes like uncertainty and surprise. However, it is unclear how the perception of these attitudes relates to dialogue structure and belief co-ordination. Perception experiment results suggest that rises reflect difficulty integrating new information rather than signaling a lack of credibility. This leads to a general analysis of rising intonation as signaling that the current question under discussion is unresolved. However, the interaction with cue word semantics restricts how much their interpretation can vary with prosody. Catherine Lai |
INTERSPEECH | 1 |
| 2009 | Perceiving surprise on cue words: prosody and semantics interact on right and reallyabstractCue words in dialogue have different interpretations depending context and prosody. This paper presents a corpus study and perception experiment investigating when prosody causes right and really to be perceived as questioning or expressing surprise. Pitch range is found to be the best cue for surprise. This extends to the question rating for really but not for right. In fact, prosody appears to interact with semantics so ratings differ for these two types of cue word even when prosodic features are similar. So, different semantics appears to result in different surprise/question rating thresholds. Catherine Lai |
INTERSPEECH | 1 |
| 2007 | Perception of disfluency: language differences and listener biasabstractThis paper describes a crosslinguistic disfluency perception experiment. We tested the recognizability of pause fillers and partial words in English, German and Mandarin. Subjects were speakers of English with no knowledge of Mandarin or German. We found that subjects could identify disfluent from fluent utterances at a level above chance. Pause fillers were easier to identify than partial words. Accuracy rates were highest for English, followed by German and then Mandarin. Although German accuracy rates were higher than those for Mandarin, discriminability analysis suggests that this is due to conservative bias towards false negatives rather than non-recognition of the acoustic material. The fact that subjects could identify disfluent speech in languages they did not know shows that there are real phonetic crosslinguistic cues to disfluency. Index Terms: crosslinguistic perception, disfluency, pause filler, partial words. Catherine Lai, Kyle Gorman, Jiahong Yuan, Mark Y. Liberman |
INTERSPEECH | 1 |
| 2005 | LPath+: A First-Order Complete Language for Linguistic Tree Query
Catherine Lai, Steven Bird |
PACLIC | 1 |