EDBT 2026 Demo / reviewers in the wild / expert
Yuanchao Li
dblp:71/11316
· DBLP profile ↗
25ranked-venue papers
17as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 11 first-author · 16 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 first-author · 3 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond manual transcripts: Exploring the potential of automatic speech recognition errors in improving Alzheimer's disease detection
Yin-Long Liu, Yuanchao Li, Jiahong Yuan, Zhen-Hua Ling |
J. Biomed. Informatics | 2 |
| 2025 | Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised ModelsabstractUtilizing Self-Supervised Learning (SSL) models for Speech Emotion Recognition (SER) has proven effective, yet limited research has explored cross-lingual scenarios. This study presents a comparative analysis between human performance and SSL models, beginning with a layer-wise analysis and an exploration of parameter-efficient fine-tuning strategies in monolingual, cross-lingual, and transfer learning contexts. We further compare the SER ability of models and humans at both utterance- and segment-levels. Additionally, we investigate the impact of dialect on cross-lingual SER through human evaluation. Our findings reveal that models, with appropriate knowledge transfer, can adapt to the target language and achieve performance comparable to native speakers. We also demonstrate the significant effect of dialect on SER for individuals without prior linguistic and paralinguistic background. Moreover, both humans and models exhibit distinct behaviors across different emotions. These results offer new insights into the cross-lingual SER capabilities of SSL models, underscoring both their similarities to and differences from human emotion perception. Zhichen Han, Tianqi Geng, Jiahong Yuan, Korin Richmond, Yuanchao Li |
ICASSP | 6 |
| 2025 | Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-LabelingabstractThe lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervised Learning (SSL) framework, introducing a novel multi-view pseudo-labeling method that leverages both acoustic and linguistic characteristics to select the most confident data for training the classification model. Acoustically, unlabeled data are compared to labeled data using the Fréchet audio distance, calculated from embeddings generated by multiple audio encoders. Linguistically, large language models are prompted to revise automatic speech recognition transcriptions and predict labels based on our proposed task-specific knowledge. High-confidence data are identified when pseudo-labels from both sources align, while mismatches are treated as low-confidence data. A bimodal classifier is then trained to iteratively label the low-confidence data until a predefined criterion is met. We evaluate our SSL framework on emotion recognition and dementia detection tasks. Experimental results demonstrate that our method achieves competitive performance compared to fully supervised learning using only 30% of the labeled data and significantly outperforms two selected baselines. Yuanchao Li, Zixing Zhang 0001, Jing Han 0010, Peter Bell 0001, Catherine Lai |
ICASSP | 1 |
| 2025 | Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error CorrectionabstractAnnotating and recognizing speech emotion using prompt engineering has recently emerged with the advancement of Large Language Models (LLMs), yet its efficacy and reliability remain questionable. In this paper, we conduct a systematic study on this topic, beginning with the proposal of novel prompts that incorporate emotion-specific knowledge from acoustics, linguistics, and psychology. Subsequently, we examine the effectiveness of LLM-based prompting on Automatic Speech Recognition (ASR) transcription, contrasting it with ground-truth transcription. Furthermore, we propose a Revise-Reason-Recognize prompting pipeline for robust LLM-based emotion recognition from spoken language with ASR errors. Additionally, experiments on context-aware learning, in-context learning, and instruction tuning are performed to examine the usefulness of LLM training schemes in this direction. Finally, we investigate the sensitivity of LLMs to minor prompt variations. Experimental results demonstrate the efficacy of the emotion-specific prompts, ASR error correction, and LLM training schemes for LLM-based emotion recognition. Our study aims to refine the use of LLMs in emotion recognition and related domains. Yuanchao Li, Yuan Gong 0001, Chao-Han Huck Yang, Peter Bell 0001, Catherine Lai |
ICASSP | 1 |
| 2025 | Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised RepresentationsabstractEmotion recognition from speech and music shares similarities due to their acoustic overlap, which has led to interest in transferring knowledge between these domains. However, the shared acoustic cues between speech and music, particularly those encoded by Self-Supervised Learning (SSL) models, remain largely unexplored, given the fact that SSL models for speech and music have rarely been applied in cross-domain research. In this work, we revisit the acoustic similarity between emotion speech and music, starting with an analysis of the layerwise behavior of SSL models for Speech Emotion Recognition (SER) and Music Emotion Recognition (MER). Furthermore, we perform cross-domain adaptation by comparing several approaches in a two-stage fine-tuning process, examining effective ways to utilize music for SER and speech for MER. Lastly, we explore the acoustic similarities between emotional speech and music using Fréchet audio distance for individual emotions, uncovering the issue of emotion bias in both speech and music SSL models. Our findings reveal that while speech and music SSL models do capture shared acoustic features, their behaviors can vary depending on different emotions due to their training strategies and domain-specificities. Additionally, parameter-efficient fine-tuning can enhance SER and MER performance by leveraging knowledge from each other. This study provides new insights into the acoustic similarity between emotional speech and music, and highlights the potential for cross-domain generalization to improve SER and MER systems. Korin Richmond, Yuanchao Li |
ICASSP | 4 |
| 2025 | Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio DistanceabstractThe complex nature of musical emotion introduces inherent bias in both recognition and generation, particularly when relying on a single audio encoder, emotion classifier, or evaluation metric. In this work, we conduct a study on Music Emotion Recognition (MER) and Emotional Music Generation (EMG), employing diverse audio encoders alongside Frechet Audio Distance (FAD), a reference-free evaluation metric. Our study begins with a benchmark evaluation of MER, highlighting the limitations of using a single audio encoder and the disparities observed across different measurements. We then propose assessing MER performance using FAD derived from multiple encoders to provide a more objective measure of musical emotion. Furthermore, we introduce an enhanced EMG approach designed to improve both the variability and prominence of generated musical emotion, thereby enhancing its realism. Additionally, we investigate the differences in realism between the emotions conveyed in real and synthetic music, comparing our EMG model against two baseline models. Experimental results underscore the issue of emotion bias in both MER and EMG and demonstrate the potential of using FAD and diverse audio encoders to evaluate musical emotion more objectively and effectively. Yuanchao Li, Azalea Gui, Dimitra Emmanouilidou, Hannes Gamper |
ICME | 1 |
| 2025 | HRAI 2025: The 1st Workshop on Holistic and Responsible Affective IntelligenceabstractThe ICMI 2025 Workshop on Holistic and Responsible Affective Intelligence (HRAI 2025) aims to advance research in affective intelligence by fostering discussions on the holistic development of affective computing and the ethical challenges it entails. The workshop aims to strengthen interdisciplinary connections within the affective computing community, promoting better integration of methodologies and enhancing real-world applicability. By tackling both technical and ethical issues, HRAI 2025 aspires to shape the future of affective AI, ensuring it is not only powerful but also fair, safe, and socially responsible. Yuanchao Li, Dimitris Kollias, Guillaume Chanel, Marios A. Fanourakis, Michal Muszynski, Brandon M. Booth, Leimin Tian, Madhawa Perera, Catherine Lai, Huili Chen |
ICMI | 1 |
| 2025 | Leveraging Cascaded Binary Classification and Multimodal Fusion for Dementia Detection through Spontaneous Speech
Yin-Long Liu, Yuanchao Li, Yu-Ang Chen, Yan-Han Peng, Jia-Hong Yuan, Zhen-Hua Ling |
INTERSPEECH | 2 |
| 2025 | Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
Nicholas Sanders, Yuanchao Li, Korin Richmond, Simon King 0001 |
INTERSPEECH | 2 |
| 2024 | Can Textual Semantics Mitigate Sounding Object Segmentation Preference?
Yaoting Wang, Peiwen Sun, Yuanchao Li, Honggang Zhang 0002, Di Hu 0001 |
ECCV (74) | 3 |
| 2024 | Speech Emotion Recognition With ASR Transcripts: a Comprehensive Study on Word Error Rate and Fusion TechniquesabstractText data is commonly utilized as a primary input to enhance Speech Emotion Recognition (SER) performance and reliability. However, the reliance on human-transcribed text in most studies impedes the development of practical SER systems, creating a gap between in-lab research and real-world scenarios where Automatic Speech Recognition (ASR) serves as the text source. Hence, this study benchmarks SER performance using ASR transcripts with varying Word Error Rates (WERs) from eleven models on three well-known corpora: IEMOCAP, CMU-MOSI, and MSP-Podcast. Our evaluation includes both text-only and bimodal SER with six fusion techniques, aiming for a comprehensive analysis that uncovers novel findings and challenges faced by current SER research. Additionally, we propose a unified ASR error-robust framework integrating ASR error correction and modality-gated fusion, achieving lower WER and higher SER results compared to the best-performing ASR transcript. These findings provide insights into SER with ASR assistance, especially for real-world applications. Yuanchao Li, Peter Bell 0001, Catherine Lai |
SLT | 1 |
| 2024 | Crossmodal ASR Error Correction With Discrete Speech UnitsabstractASR remains unsatisfactory in scenarios where the speaking style diverges from that used to train ASR systems, resulting in erroneous transcripts. To address this, ASR Error Correction (AEC), a post-ASR processing approach, is required. In this work, we tackle an understudied issue: the Low-Resource Out-of-Domain (LROOD) problem, by investigating crossmodal AEC on very limited downstream data with 1-best hypothesis transcription. We explore pretraining and fine-tuning strategies and uncover an ASR domain discrepancy phenomenon, shedding light on appropriate training schemes for LROOD data. Moreover, we propose the incorporation of discrete speech units to align with and enhance the word embeddings for improving AEC quality. Results from multiple corpora and several evaluation metrics demonstrate the feasibility and efficacy of our proposed AEC approach on LROOD data as well as its generalizability and superiority on large-scale data. Finally, a study on speech emotion recognition confirms that our model produces ASR error-robust transcripts suitable for downstream applications. Yuanchao Li, Pinzhen Chen, Peter Bell 0001, Catherine Lai |
SLT | 1 |
| 2024 | Large Language Model Based Generative Error Correction: A Challenge and Baselines For Speech Recognition, Speaker Tagging, and Emotion RecognitionabstractGiven recent advances in generative AI technology, a key question is how large language models (LLMs) can enhance acoustic modeling tasks using text decoding results from a frozen, pretrained automatic speech recognition (ASR) model. To explore new capabilities in language modeling for speech processing, we introduce the generative speech transcription error correction (GenSEC) challenge. This challenge comprises three post-ASR language modeling tasks: (i) post-ASR transcription correction, (ii) speaker tagging, and (iii) emotion recognition. These tasks aim to emulate future LLM-based agents handling voice-based interfaces while remaining accessible to a broad audience by utilizing open pretrained language models or agent-based APIs. We also discuss insights from baseline evaluations, as well as lessons learned for designing future evaluations. Chao-Han Huck Yang, Taejin Park, Yuan Gong 0001, Yuanchao Li, Zhehuai Chen, Chen Chen 0075, Kunal Dhawan, Piotr Zelasko, Chao Zhang 0031, Yun-Nung Chen, Yu Tsao 0001, Jagadeesh Balam, Boris Ginsburg, Sabato Marco Siniscalchi, Chng Eng Siong, Peter Bell 0001, Catherine Lai, Shinji Watanabe 0001, Andreas Stolcke |
SLT | 4 |
| 2023 | Multimodal Dyadic Impression Recognition via Listener Adaptive Cross-Domain FusionabstractAs a sub-branch of affective computing, impression recognition, e.g., perception of speaker characteristics such as warmth or competence, is potentially a critical part of both human-human conversations and spoken dialogue systems. Most research has studied impressions only from the behaviors expressed by the speaker or the response from the listener, yet ignored their latent connection. In this paper, we perform impression recognition using a proposed listener adaptive cross-domain architecture, which consists of a listener adaptation function to model the causality between speaker and listener behaviors and a cross-domain fusion function to strengthen their connection. The experimental evaluation on the dyadic IMPRESSION dataset verified the efficacy of our method, producing concordance correlation coefficients of 78.8% and 77.5% in the competence and warmth dimensions, outperforming previous studies. The proposed method is expected to be generalized to similar dyadic interaction scenarios in affective computing. Yuanchao Li, Peter Bell 0001, Catherine Lai |
ICASSP | 1 |
| 2023 | Transfer Learning for Personality Perception via Speech Emotion RecognitionabstractHolistic perception of affective attributes is an important human perceptual ability. However, this ability is far from being realized in current affective computing, as not all of the attributes are well studied and their interrelationships are poorly understood. In this work, we investigate the relationship between two affective attributes: personality and emotion, from a transfer learning perspective. Specifically, we transfer Transformer-based and wav2vec-based emotion recognition models to perceive personality from speech across corpora. Compared with previous studies, our results show that transferring emotion recognition is effective for personality perception. Moreoever, this allows for better use and exploration of small personality corpora. We also provide novel findings on the relationship between personality and emotion that will aid future research on holistic affect recognition. Yuanchao Li, Peter Bell 0001, Catherine Lai |
INTERSPEECH | 1 |
| 2023 | ASR and Emotional Speech: A Word-Level Investigation of the Mutual Impact of Speech and Emotion RecognitionabstractIn Speech Emotion Recognition (SER), textual data is often used alongside audio signals to address their inherent variability. However, the reliance on human annotated text in most research hinders the development of practical SER systems. To overcome this challenge, we investigate how Automatic Speech Recognition (ASR) performs on emotional speech by analyzing the ASR performance on emotion corpora and examining the distribution of word errors and confidence scores in ASR transcripts to gain insight into how emotion affects ASR. We utilize four ASR systems, namely Kaldi ASR, wav2vec, Conformer, and Whisper, and three corpora: IEMOCAP, MOSI, and MELD to ensure generalizability. Additionally, we conduct text-based SER on ASR transcripts with increasing word error rates to investigate how ASR affects SER. The objective of this study is to uncover the relationship and mutual impact of ASR and SER, in order to facilitate ASR adaptation to emotional speech and the use of SER in real world. Yuanchao Li, Zeyu Zhao 0004, Ondrej Klejch, Peter Bell 0001, Catherine Lai |
INTERSPEECH | 1 |
| 2022 | Alzheimer's Dementia Detection through Spontaneous Dialogue with Proactive Robotic ListenersabstractAs the aging of society continues to accelerate, Alzheimer's Disease (AD) has received more and more attention from not only medical but also other fields, such as computer science, over the past decade. Since speech is considered one of the effective ways to diagnose cognitive decline, AD detection from speech has emerged as a hot topic. Nevertheless, such approaches fail to tackle several key issues: 1) AD is a complex neurocognitive disorder which means it is inappropriate to conduct AD detection using utterance information alone while ignoring dialogue infor-mation; 2) Utterances of AD patients contain many disfluencies that affect speech recognition yet are helpful to diagnosis; 3) AD patients tend to speak less, causing dialogue breakdown as the disease progresses. This fact leads to a small number of utterances, which may cause detection bias. Therefore, in this paper, we propose a novel AD detection architecture consisting of two major modules: an ensemble AD detector and a proactive listener. This architecture can be embedded in the dialogue system of conversational robots for healthcare. Yuanchao Li, Catherine Lai, Divesh Lala, Koji Inoue, Tatsuya Kawahara |
HRI | 1 |
| 2022 | Fusing ASR Outputs in Joint Training for Speech Emotion RecognitionabstractAlongside acoustic information, linguistic features based on speech transcripts have been proven useful in Speech Emotion Recognition (SER). However, due to the scarcity of emotion labelled data and the difficulty of recognizing emotional speech, it is hard to obtain reliable linguistic features and models in this research area. In this paper, we propose to fuse Automatic Speech Recognition (ASR) outputs into the pipeline for joint training SER. The relationship between ASR and SER is understudied, and it is unclear what and how ASR features benefit SER. By examining various ASR outputs and fusion methods, our experiments show that in joint ASR-SER training, incorporating both ASR hidden and text output using a hierarchical co-attention fusion approach improves the SER performance the most. On the IEMOCAP corpus, our approach achieves 63.4% weighted accuracy, which is close to the baseline results achieved by combining ground-truth transcripts. In addition, we also present novel word error rate analysis on IEMOCAP and layer-difference analysis of the Wav2vec 2.0 model to better understand the relationship between ASR and SER. Yuanchao Li, Peter Bell 0001, Catherine Lai |
ICASSP | 1 |
| 2022 | Exploration of a Self-Supervised Speech Model: A Study on Emotional CorporaabstractSelf-supervised speech models have grown fast during the past few years and have proven feasible for use in various downstream tasks. Some recent work has started to look at the characteristics of these models, yet many concerns have not been fully addressed. In this work, we conduct a study on emotional corpora to explore a popular self-supervised model - wav2vec 2.0. Via a set of quantitative analysis, we mainly demonstrate that: 1) wav2vec 2.0 appears to discard paralinguistic information that is less useful for word recognition purposes; 2) for emotion recognition, representations from the middle layer alone perform as well as those derived from layer averaging, while the final layer results in the worst performance in some cases; 3) current self-supervised models may not be the optimal solution for downstream tasks that make use of non-lexical features. Our work provides novel findings that will aid future research in this area and theoretical basis for the use of existing models. Yuanchao Li, Yumnah Mohamied, Peter Bell 0001, Catherine Lai |
SLT | 1 |
| 2021 | Semi-Supervised Learning for Multimodal Speech and Emotion RecognitionabstractSpeech Emotion Recognition (SER) is becoming necessary for interactive spoken dialogue systems as users are expecting empathy from computers. Recent work has shown the importance of approaching this problem from a multimodal perspective, with models that combine visual, acoustic, and lexical features performing better than models based on single modalities. However, current SER models are not robust to out of domain data, partly due to the fact that emotion labeled corpora are generally small. This paper outlines my PhD research plan that aims to improve the SER model by proposing to jointly train with an Automatic Speech Recognition (ASR) model using a novel cross-task semi-supervised learning approach on unlabeled data. The ASR model would be benefit from the training approach and serve as the lexical features provider. This joint ASR-SER model is expected to alleviate the lack of data problem and to be applied in real-life applications such as human-computer interaction and digital health. Yuanchao Li |
ICMI | 1 |
| 2019 | Improved End-to-End Speech Emotion Recognition Using Self Attention Mechanism and Multitask Learning
Yuanchao Li, Tianyu Zhao 0001, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2018 | Towards Improving Speech Emotion Recognition for In-Vehicle Agents: Preliminary Results of Incorporating Sentiment Analysis by Using Early and Late Fusion MethodsabstractThe automotive industry is integrating a number of Artificial Intelligence (AI) functions, such as speech recognition and face recognition into in-vehicle agents, aiming to make them able to conduct human-like interaction with the driver [1, 5, 8, 12]. Among the many cutting-edge functions, emotion recognition has been proved significant to improve driving experience and road safety [14]. However, due to the noisy in-vehicle environment, trustworthy Speech Emotion Recognition (SER) is still hard to achieve. The voice from other passengers and the outside noise make SER unreliable. Positive emotion is sometimes recognized as negative and vice versa. Yuanchao Li |
HAI | 1 |
| 2018 | Adaptive Transmission Design in Fog Radio Access Networks with Partition-Based CachingabstractThis paper investigates the transmission design in fog radio access networks where each edge node (EN) has a local cache and can pre-store contents based on partition-based caching. In the considered partition- based caching, each file is partitioned into multiple subfiles and cached at different ENs. Upon user requests, the ENs can jointly transmit all the subfiles to the target users simultaneously. Each user adopts a successive interference cancellation receiver to decode each desired subfile with certain order. To improve the network performance, we propose a novel cache-aware scheme to determine the decoding order. Furthermore, we formulate a joint beamforming and dynamic EN clustering optimization problem to minimize the weighted sum of transmission power cost and fronthaul cost under the quality-of-service constraint for each user. An efficient algorithm by adopting smooth function approximation and convex-concave procedure is proposed to solve this problem. To exploit the advantages of partition-based caching, we further design a user-aware caching strategy. Numerical results show that the proposed transmission scheme together with the proposed caching strategy can strike a better balance between power consumption and fronthaul consumption than existing schemes. Yuanchao Li, Erkai Chen, Meixia Tao |
ICC | 1 |
| 2017 | Utterance Behavior of Users While Playing Basketball with a Virtual Teammate
Divesh Lala, Yuanchao Li, Tatsuya Kawahara |
ICAART (1) | 2 |
| 2016 | Application of linear canonical transform correlation for detection of linear frequency modulated signalsabstractLinear canonical transform (LCT), which can be deemed to be a generalisation of the fractional Fourier transform, has been used in several areas, including signal processing and optics. Motivated by the operator theory, a new unitary operator associated with the LCT is introduced. This new operator generalises the unitary fractional operator which is proposed by Akay et al . recently. Via operator manipulations, the authors also derive a new definition, the LCT correlation operation, and present an alternative and efficient implementation of it. It is shown that the proposed LCT autocorrelation corresponds to radial slices of the ambiguity function in the ambiguity plane. On the basis of this relationship, an application of the fast LCT autocorrelation for detection and parameter estimation with respect to the chirp rates of linear frequency modulated signals corrupted by noise is proposed. Finally, the validity of the proposed method is verified by simulation results. Yuanchao Li, Feng Zhang 0011, Ran Tao 0003 |
IET Signal Process. | 1 |