EDBT 2026 Demo / reviewers in the wild / expert
Wen Wu 0007
dblp:92/382-7
· DBLP profile ↗
17ranked-venue papers
10as first author
17since 2021 · last 2025
0000-0001-8116-7715ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bayesian WeakS-to-Strong from Text Classification to GenerationabstractAdvances in large language models raise the question of how alignment techniques will adapt as models become increasingly complex and humans will only be able to supervise them weakly. Weak-to-Strong mimics such a scenario where weak model supervision attempts to harness the full capabilities of a much stronger model. This work extends Weak-to-Strong to WeakS-to-Strong by exploring an ensemble of weak models which simulate the variability in human opinions. Confidence scores are estimated using a Bayesian approach to guide the WeakS-to-Strong generalization. Furthermore, we extend the application of WeakS-to-Strong from text classification tasks to text generation tasks where more advanced strategies are investigated for supervision. Moreover, direct preference optimization is applied to advance the student model's preference learning, beyond the basic learning framework of teacher forcing. Results demonstrate the effectiveness of the proposed approach for the reliability of a strong student model, showing potential for superalignment. Ziyun Cui, Guangzhi Sun, Wen Wu 0007, Chao Zhang 0031 |
ICLR | 4 |
| 2025 | The 1st SpeechWellness Challenge: Detecting Suicide Risk Among AdolescentsabstractThe 1st SpeechWellness Challenge (SW1) aims to advance methods for detecting current suicide risk in adolescents using speech analysis techniques. Suicide among adolescents is a critical public health issue globally. Early detection of suicidal tendencies can lead to timely intervention and potentially save lives. Traditional methods of assessment often rely on self-reporting or clinical interviews, which may not always be accessible. The SW1 challenge addresses this gap by exploring speech as a non-invasive and readily available indicator of mental health. We release the SW1 dataset which contains speech recordings from 600 adolescents aged 10-18 years. By focusing on speech generated from natural tasks, the challenge seeks to uncover patterns and markers that correlate with current suicide risk. Wen Wu 0007, Ziyun Cui, Chang Lei, Yinan Duan, Diyang Qu, Ji Wu 0002, Bowen Zhou 0001, Runsen Chen, Chao Zhang 0031 |
INTERSPEECH | 1 |
| 2025 | BrainOmni: A Brain Foundation Model for Unified EEG and MEG SignalsabstractElectroencephalography (EEG) and magnetoencephalography (MEG) measure neural activity non-invasively by capturing electromagnetic fields generated by dendritic currents.
Although rooted in the same biophysics, EEG and MEG exhibit distinct signal patterns, further complicated by variations in sensor configurations across modalities and recording devices.
Existing approaches typically rely on separate, modality- and dataset-specific models, which limits the performance and cross-domain scalability.
This paper proposes BrainOmni, the first brain foundation model that generalises across heterogeneous EEG and MEG recordings.
To unify diverse data sources, we introduce BrainTokenizer, the first tokeniser that quantises spatiotemporal brain activity into discrete representations.
Central to BrainTokenizer is a novel Sensor Encoder that encodes sensor properties such as spatial layout, orientation, and type, enabling compatibility across devices and modalities.
Building upon the discrete representations, BrainOmni learns unified semantic embeddings of brain signals by self-supervised pretraining. To the best of our knowledge, it is the first foundation model to support both EEG and MEG signals, as well as the first to incorporate large-scale MEG pretraining.
A total of 1,997 hours of EEG and 656 hours of MEG data are curated and standardised from publicly available sources for pretraining.
Experiments show that BrainOmni outperforms both existing foundation models and state-of-the-art task-specific models on a range of downstream tasks. It also demonstrates strong generalisation to unseen EEG and MEG devices. Further analysis reveals that joint EEG-MEG (EMEG) training yields consistent improvements across both modalities. Code and checkpoints are publicly available at https://github.com/OpenTSLab/BrainOmni Qinfan Xiao, Ziyun Cui, Wen Wu 0007, Andrew Thwaites, Alexandra Woolgar, Bowen Zhou 0001, Chao Zhang 0031 |
NeurIPS | 5 |
| 2024 | Handling Ambiguity in Emotion: From Out-of-Domain Detection to Distribution EstimationabstractWen Wu, Bo Li, Chao Zhang, Chung-Cheng Chiu, Qiujia Li, Junwen Bai, Tara Sainath, Phil Woodland. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wen Wu 0007, Bo Li 0028, Chao Zhang 0031, Chung-Cheng Chiu, Qiujia Li, Junwen Bai, Tara N. Sainath, Philip C. Woodland |
ACL (1) | 1 |
| 2024 | Parameter Efficient Finetuning for Speech Emotion Recognition and Domain AdaptationabstractFoundation models have shown superior performance for speech emotion recognition (SER). However, given the limited data in emotion corpora, finetuning all parameters of large pre-trained models for SER can be both resource-intensive and susceptible to overfitting. This paper investigates parameter-efficient finetuning (PEFT) for SER. Various PEFT adaptors are systematically studied for both classification of discrete emotion categories and prediction of dimensional emotional attributes. The results demonstrate that the combination of PEFT methods surpasses full finetuning with a significant reduction in the number of trainable parameters. Furthermore, a two-stage adaptation strategy is proposed to adapt models trained on acted emotion data, which is more readily available, to make the model more adept at capturing natural emotional expressions. Both intra- and cross-corpus experiments validate the efficacy of the proposed approach in enhancing the performance on both the source and target domains. Nineli Lashkarashvili, Wen Wu 0007, Guangzhi Sun, Philip C. Woodland |
ICASSP | 2 |
| 2024 | Leveraging Speech PTM, Text LLM, And Emotional TTS For Speech Emotion RecognitionabstractIn this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS. First, we investigated the representation ability of different speech self-supervised pre-trained models, and we found that data2vec has a good representation ability on the SER task. Second, we employed a powerful large language model (LLM), GPT-4, and emotional text-to-speech (TTS) model, Azure TTS, to generate emotionally congruent text and speech. We carefully designed the text prompt and dataset construction, to obtain the synthetic emotional speech data with high quality. Third, we studied different ways of data augmentation to promote the SER task with synthetic speech, including random mixing, adversarial training, transfer learning, and curriculum learning. Experiments and ablation studies on the IEMOCAP dataset demonstrate the effectiveness of our method, compared with other data augmentation methods, and data augmentation with other synthetic data. Ziyang Ma 0001, Wen Wu 0007, Zhisheng Zheng, Qian Chen 0003, Shiliang Zhang, Xie Chen 0001 |
ICASSP | 2 |
| 2024 | Confidence Estimation for Automatic Detection of Depression and Alzheimer's Disease Based on Clinical Interviews
Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland |
INTERSPEECH | 1 |
| 2024 | Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
Ziyun Cui, Chang Lei, Wen Wu 0007, Yinan Duan, Diyang Qu, Ji Wu 0002, Runsen Chen, Chao Zhang 0031 |
INTERSPEECH | 3 |
| 2024 | Affect Recognition in Conversations Using Large Language ModelsabstractAffect recognition, encompassing emotions, moods, and feelings, plays a pivotal role in human communication.In the realm of conversational artificial intelligence, the ability to discern and respond to human affective cues is a critical factor for creating engaging and empathetic interactions.This study investigates the capacity of large language models (LLMs) to recognise human affect in conversations, with a focus on both open-domain chit-chat dialogues and task-oriented dialogues.Leveraging three diverse datasets, namely IEMOCAP (Busso et al., 2008), EmoWOZ (Feng et al., 2022), and DAIC-WOZ (Gratch et al., 2014), covering a spectrum of dialogues from casual conversations to clinical interviews, we evaluate and compare LLMs' performance in affect recognition.Our investigation explores the zero-shot and few-shot capabilities of LLMs through incontext learning as well as their model capacities through task-specific fine-tuning.Additionally, this study takes into account the potential impact of automatic speech recognition errors on LLM predictions.With this work, we aim to shed light on the extent to which LLMs can replicate human-like affect recognition capabilities in conversations. Shutong Feng, Guangzhi Sun, Nurul Lubis, Wen Wu 0007, Chao Zhang 0031, Milica Gasic |
SIGDIAL | 4 |
| 2023 | Estimating the Uncertainty in Emotion Attributes using Deep Evidential RegressionabstractIn automatic emotion recognition (AER), labels assigned by different human annotators to the same utterance are often inconsistent due to the inherent complexity of emotion and the subjectivity of perception.Though deterministic labels generated by averaging or voting are often used as the ground truth, it ignores the intrinsic uncertainty revealed by the inconsistent labels.This paper proposes a Bayesian approach, deep evidential emotion regression (DEER), to estimate the uncertainty in emotion attributes.Treating the emotion attribute labels of an utterance as samples drawn from an unknown Gaussian distribution, DEER places an utterance-specific normal-inverse gamma prior over the Gaussian likelihood and predicts its hyper-parameters using a deep neural network model.It enables a joint estimation of emotion attributes along with the aleatoric and epistemic uncertainties.AER experiments on the widely used MSP-Podcast and IEMOCAP datasets showed DEER produced state-of-theart results for both the mean values and the distribution of emotion attributes 1 . Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland |
ACL (1) | 1 |
| 2023 | Transferring Speech-Generic and Depression-Specific Knowledge for Alzheimer's Disease DetectionabstractThe detection of Alzheimer’s disease (AD) from spontaneous speech has attracted increasing attention while the sparsity of training data remains an important issue. This paper handles the issue by knowledge transfer, specifically from both speech-generic and depression-specific knowledge. The paper first studies sequential knowledge transfer from generic foundation models pretrained on large amounts of speech and text data. A block-wise analysis is performed for AD diagnosis based on the representations extracted from different intermediate blocks of different foundation models. Apart from the knowledge from speech-generic representations, this paper also proposes to simultaneously transfer the knowledge from a speech depression detection task based on the high comorbidity rates of depression and AD. A parallel knowledge transfer framework is studied that jointly learns the information shared between these two tasks. Experimental results show that the proposed method improves AD and depression detection, and produces a state-of-the-art F1 score of 0.928 for AD diagnosis on the commonly used ADReSSo dataset. Ziyun Cui, Wen Wu 0007, Weiqiang Zhang 0001, Ji Wu 0002, Chao Zhang 0031 |
ASRU | 2 |
| 2023 | Self-Supervised Representations in Speech-Based Depression DetectionabstractThis paper proposes handling training data sparsity in speech-based automatic depression detection (SDD) using foundation models pre-trained with self-supervised learning (SSL). An analysis of SSL representations derived from different layers of pre-trained foundation models is first presented for SDD, which provides insight to suitable indicator for depression detection. Knowledge transfer is then performed from automatic speech recognition (ASR) and emotion recognition to SDD by fine-tuning the foundation models. Results show that the uses of oracle and ASR transcriptions yield similar SDD performance when the hidden representations of the ASR model is incorporated along with the ASR textual information. By integrating representations from multiple foundation models, state-of-the-art SDD results based on real ASR were achieved on the DAIC-WOZ dataset. Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland |
ICASSP | 1 |
| 2023 | Integrating Emotion Recognition with Speech Recognition and Speaker Diarisation for ConversationsabstractAlthough automatic emotion recognition (AER) has recently drawn significant research interest, most current AER studies use manually segmented utterances, which are usually unavailable for dialogue systems. This paper proposes integrating AER with automatic speech recognition (ASR) and speaker diarisation (SD) in a jointly-trained system. Distinct output layers are built for four sub-tasks including AER, ASR, voice activity detection and speaker classification based on a shared encoder. Taking the audio of a conversation as input, the integrated system finds all speech segments and transcribes the corresponding emotion classes, word sequences, and speaker identities. Two metrics are proposed to evaluate AER performance with automatic segmentation based on time-weighted emotion and speaker classification errors. Results on the IEMOCAP dataset show that the proposed system consistently outperforms two baselines with separately trained single-task systems on AER, ASR and SD. Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland |
INTERSPEECH | 1 |
| 2023 | Estimating the Uncertainty in Emotion Class Labels With Utterance-Specific Dirichlet PriorsabstractEmotion recognition is a key attribute for artificial intelligence systems that need to naturally interact with humans. However, the task definition is still an open problem due to the inherent ambiguity of emotions. In this paper, a novel Bayesian training loss based on per-utterance Dirichlet prior distributions is proposed for verbal emotion recognition, which models the uncertainty in one-hot labels created when human annotators assign the same utterance to different emotion classes. An additional metric is used to evaluate the performance by detecting test utterances with high labelling uncertainty. This removes a major limitation that emotion classification systems only consider utterances with labels where the majority of annotators agree on the emotion class. Furthermore, a frequentist approach is studied to leverage the continuous-valued “soft” labels obtained by averaging the one-hot labels. We propose a two-branch model structure for emotion classification on a per-utterance basis, which achieves state-of-the-art classification results on the widely used IEMOCAP dataset. Based on this, uncertainty estimation experiments were performed. The best performance in terms of the area under the precision-recall curve when detecting utterances with high uncertainty was achieved by interpolating the Bayesian training loss with the Kullback-Leibler divergence training loss for the soft labels. The generality of the proposed approach was verified using the MSP-Podcast dataset which yielded the same pattern of results. Wen Wu 0007, Chao Zhang 0031, Xixin Wu, Philip C. Woodland |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Climate and Weather: Inspecting Depression Detection via Emotion RecognitionabstractAutomatic depression detection has attracted increasing amount of attention but remains a challenging task. Psychological research suggests that depressive mood is closely related with emotion expression and perception, which motivates the investigation of whether knowledge of emotion recognition can be transferred for depression detection. This paper uses pretrained features extracted from the emotion recognition model for depression detection, further fuses emotion modality with audio and text to form multimodal depression detection. The proposed emotion transfer improves depression detection performance on DAIC-WOZ as well as increases the training stability. The analysis of how the emotion expressed by de-pressed individuals is further perceived provides clues for further understanding of the relationship between depression and emotion. Wen Wu 0007, Mengyue Wu, Kai Yu 0004 |
ICASSP | 1 |
| 2022 | Distribution-Based Emotion Recognition in ConversationabstractAutomatic emotion recognition in conversation (ERC) is crucial for emotion-aware conversational artificial intelligence. This paper proposes a distribution-based framework that formulates ERC as a sequence-to-sequence problem for emotion distribution estimation. The inherent ambiguity of emotions and the subjectivity of human perception lead to disagreements in emotion labels, which is handled naturally in our framework from the perspective of uncertainty estimation in emotion distributions. A Bayesian training loss is introduced to improve the uncertainty estimation by conditioning each emotional state on an utterance-specific Dirichlet prior distribution. Experimental results on the IEMOCAP dataset show that ERC outperformed the single-utterance-based system, and the proposed distribution-based ERC methods have not only better classification accuracy, but also show improved uncertainty estimation. Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland |
SLT | 1 |
| 2021 | Emotion Recognition by Fusing Time Synchronous and Time Asynchronous RepresentationsabstractIn this paper, a novel two-branch neural network model structure is proposed for multimodal emotion recognition, which consists of a time synchronous branch (TSB) and a time asynchronous branch (TAB). To capture correlations between each word and its acoustic realisation, the TSB combines speech and text modalities at each input window frame and then uses pooling across time to form a single embedding vector. The TAB, by contrast, provides cross-utterance information by integrating sentence text embeddings from a number of context utterances into another embedding vector. The final emotion classification uses both the TSB and the TAB embeddings. Experimental results on the IEMOCAP dataset demonstrate that the two-branch structure achieves state-of-the-art results in 4-way classification with all common test setups. When using automatic speech recognition (ASR) output instead of manually transcribed reference text, it is shown that the cross-utterance information considerably improves robustness against ASR errors. Furthermore, by incorporating an extra class for all the other emotions, the final 5-way classification system with ASR hypotheses can be viewed as a prototype for more realistic emotion recognition systems. Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland |
ICASSP | 1 |