EDBT 2026 Demo / reviewers in the wild / expert
Ya Li 0001
dblp:51/4056-1
· DBLP profile ↗
55ranked-venue papers
7as first author
33since 2021 · last 2026
0000-0002-6284-5039ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 6 first-author · 23 since 2021Artificial intelligence and machine learning · 30 · 2 first-author · 19 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource ScenariosabstractZero-shot singing voice conversion (SVC) transforms a source singer's timbre to an unseen target speaker's voice while preserving melodic content without fine-tuning. Existing methods model speaker timbre and vocal content separately, losing essential acoustic information that degrades output quality while requiring significant computational resources. To overcome these limitations, we propose HQ-SVC, an efficient framework for high-quality zero-shot SVC. HQ-SVC first extracts jointly content and speaker features using a decoupled codec. It then enhances fidelity through pitch and volume modeling, preserving critical acoustic information typically lost in separate modeling approaches, and progressively refines outputs via differentiable signal processing and diffusion techniques. Evaluations confirm HQ-SVC significantly outperforms state-of-the-art zero-shot SVC methods in conversion quality and efficiency. Beyond voice conversion, HQ-SVC achieves superior voice naturalness compared to specialized audio super-resolution methods while natively supporting voice super-resolution tasks. Bingsong Bai, Yizhong Geng, Fengping Wang, Puyuan Guo, Yingming Gao, Ya Li 0001 |
AAAI | 7 |
| 2025 | Controllable 3D Dance Generation Using Diffusion-Based Transformer U-NetabstractRecently, dance generation has attracted increasing interest. In particular, the success of diffusion models in image generation has led to the emergence of dance generation systems based on the diffusion framework. However, these systems lack controllability, which limits their practical applications. In this paper, we propose a controllable dance generation method based on the diffusion model, which can generate 3D dance motions controlled by 2D keypoint sequences. Specifically, we design a transformer-based U-Net model to predict actual motions. Then, we fix the parameters of the U-Net model and train an additional control network, enabling the generated motions to be controlled by 2D keypoints. We conduct extensive experiments and compared our method with existing works on the widely used AIST++ dataset, demonstrating that our approach has certain advantages and controllability. Moreover, we also test our model on in-the-wild videos and find that it is capable of generating dance movements similar to the motions in the videos as well. Puyuan Guo, Tuo Hao, Wenxin Fu, Yingming Gao, Ya Li 0001 |
AAAI | 5 |
| 2025 | Beyond Surface Simplicity: Revealing Hidden Reasoning Attributes for Precise Commonsense DiagnosisabstractCommonsense question answering (QA) are widely used to evaluate the commonsense abilities of large language models. However, answering commonsense questions correctly requires not only knowledge but also reasoning—even for seemingly simple questions. We demonstrate that such hidden reasoning attributes in commonsense questions can lead evaluation accuracy differences of up to 24.8% across different difficulty levels in the same benchmark. Current benchmarks overlook these hidden reasoning attributes, making it difficult to assess a model’s specific levels of commonsense knowledge and reasoning ability. To address this issue, we introduce ReComSBench, a novel framework that reveals hidden reasoning attributes behind commonsense questions by leveraging the knowledge generated during the reasoning process. Additionally, ReComSBench proposes three new metrics for decoupled evaluation: Knowledge Balanced Accuracy, Marginal Sampling Gain, and Knowledge Coverage Ratio. Experiments show that ReComSBench provides insights into model performance that traditional benchmarks cannot offer. The difficulty stratification based on revealed hidden reasoning attributes performs as effectively as the model-probability-based approach but is more generalizable and better suited for improving a model’s commonsense reasoning abilities. By uncovering and analyzing the hidden reasoning attributes in commonsense data, ReComSBench offers a new approach to enhancing existing commonsense benchmarks. Huijun Lian, Zekai Sun, Yingming Gao, Ya Li 0001 |
ACL (1) | 5 |
| 2025 | Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio GenerationabstractMainstream Text-to-Audio (TTA) models that rely on Mel-spectrograms often struggle to generate audio with rich content, leading to blurred or incoherent outputs. This stems from an inability to model intricate spectral details and textures. We investigate the role of U-Net components in generation and find that high-frequency components in skip-connections and the backbone are crucial for texture, while low-frequency backbone components are vital for the denoising process. Based on this, we propose “Mel-Refine,” a plug-and-play approach that enhances Mel-spectrogram quality by adjusting component weights during inference. Our method requires no additional training or fine-tuning and is fully compatible with any diffusion-based TTA architecture. Experiments show that Mel-Refine boosts the performance of the latest TTA model, Tango2, by $25 \%$, demonstrating its effectiveness. Hongming Guo, Ruibo Fu, Yizhong Geng, Shuchen Shi, Tao Wang 0074, Chunyu Qiang, Ya Li 0001, Zhengqi Wen, Xuefei Liu, Chenxing Li |
ASRU | 7 |
| 2025 | DetailTTS: Learning Residual Detail Information for Zero-shot Text-to-speechabstractTraditional text-to-speech (TTS) systems often face challenges in aligning text and speech, leading to the omission of critical linguistic and acoustic details. This misalignment creates an information gap, which existing methods attempt to address by incorporating additional inputs, but these often introduce data inconsistencies and increase complexity. To address these issues, we propose DetailTTS, a zero-shot TTS system based on a conditional variational autoencoder. It incorporates two key components: the Prior Detail Module and the Duration Detail Module, which capture residual detail information missed during alignment. These modules effectively enhance the model’s ability to retain fine-grained details, significantly improving speech quality while simplifying the model by obviating the need for additional inputs. Experiments on the WenetSpeech4TTS dataset show that DetailTTS outperforms traditional TTS systems in both naturalness and speaker similarity, even in zero-shot scenarios. Our source code and demo page are available at https://detailtts.github.io/. Yichen Han, Yizhong Geng, Yingming Gao, Fengping Wang, Bingsong Bai, Jinlong Xue, Yayue Deng, Zhengqi Wen, Ya Li 0001 |
ICASSP | 11 |
| 2025 | OV-MER: Towards Open-Vocabulary Multimodal Emotion RecognitionabstractMultimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to capture the inherent complexity, subtlety, and multi-appraisal nature of human emotional experiences, as demonstrated by studies in psychology and cognitive science. To overcome this limitation, we advocate for introducing the concept of open vocabulary into MER. This paradigm shift aims to enable models to predict emotions beyond a fixed label space, accommodating a flexible set of categories to better reflect the nuanced spectrum of human emotions. To achieve this, we propose a novel paradigm: Open-Vocabulary MER (OV-MER), which enables emotion prediction without being confined to predefined spaces. However, constructing a dataset that encompasses the full range of emotions for OV-MER is practically infeasible; hence, we present a comprehensive solution including a newly curated database, novel evaluation metrics, and a preliminary benchmark. By advancing MER from basic emotions to more nuanced and diverse emotional states, we hope this work can inspire the next generation of MER, enhancing its generalizability and applicability in real-world scenarios. Code and dataset are available at: https://github.com/zeroQiaoba/AffectGPT. Zheng Lian 0004, Haiyang Sun 0004, Licai Sun, Haoyu Chen 0001, Lan Chen 0005, Zhuofan Wen 0001, Hailiang Yao, Bin Liu 0041, Rui Liu 0008, Shan Liang 0007, Ya Li 0001, Jiangyan Yi, Jianhua Tao 0001 |
ICML | 14 |
| 2025 | EEG-based Voice Conversion : Hearing the Voice of Your Brain
Yizhong Geng, Wenxin Fu, Qihang Lu, Bingsong Bai, Yingming Gao, Ya Li 0001 |
INTERSPEECH | 7 |
| 2025 | MER 2025: When Affective Computing Meets Large Language ModelsabstractMER2025 is the third year of our MER series of challenges. Previously, MER2023 (http://merchallenge.cn/mer2023) focused on multi-label learning, noise robustness, and semi-supervised learning, while MER2024 (https://zeroqiaoba.github.io/MER2024-website) introduced a new track dedicated to open-vocabulary emotion recognition. This year, MER2025 centers on the theme ''When Affective Computing Meets Large Language Models (LLMs)''. We aim to shift the paradigm from traditional categorical frameworks reliant on predefined emotion taxonomies to LLM-driven generative methods, offering innovative solutions for more accurate and reliable emotion understanding. The challenge contains four tracks: MER-SEMI focuses on fixed categorical emotion recognition enhanced by semi-supervised learning; MER-FG explores fine-grained emotions, expanding recognition from basic to nuanced emotional states; MER-DES incorporates multimodal cues (beyond emotion words) into predictions to enhance model interpretability; MER-PR reveals whether emotion prediction results can improve personality recognition performance. For the first three tracks, the baseline code is available at MERTools (https://github.com/zeroQiaoba/MERTools) and datasets can be accessed via Hugging Face (https://huggingface.co/datasets/MERChallenge/MER2025). For the last track, the dataset and baseline code are available on GitHub (https://github.com/cai-cong/MER25_personality). Zheng Lian 0004, Rui Liu 0008, Kele Xu, Bin Liu 0041, Xuefei Liu, Yazhou Zhang 0001, Xin Liu 0012, Yong Li 0032, Zebang Cheng, Haolin Zuo, Ziyang Ma 0001, Xiaojiang Peng, Xie Chen 0001, Ya Li 0001, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001 |
ACM Multimedia | 14 |
| 2025 | EMO-Avatar: An LLM-Agent-Orchestrated Framework for Multimodal Emotional Support in Human Animation
Wenxin Fu, Qihang Lu, Zekai Sun, Yizhong Geng, Puyuan Guo, Yingming Gao, Ya Li 0001 |
ACM Multimedia | 9 |
| 2025 | Video Demoireing Using Focused-Defocused Dual-Camera SystemabstractMoire patterns, unwanted color artifacts in images and videos, arise from the interference between spatially high-frequency scene contents and the spatial discrete sampling of digital cameras. Existing demoireing methods primarily rely on single-camera image/video processing, which faces two critical challenges: 1) distinguishing moire patterns from visually similar real textures, and 2) preserving tonal consistency and temporal coherence while removing moire artifacts. To address these issues, we propose a dual-camera framework that captures synchronized videos of the same scene: one in focus (retaining high-quality textures but may exhibit moire patterns) and one defocused (with significantly reduced moire patterns but blurred textures). We use the defocused video to help distinguish moire patterns from real texture, so as to guide the demoireing of the focused video. We propose a frame-wise demoireing pipeline, which begins with an optical flow based alignment step to address any discrepancies in displacement and occlusion between the focused and defocused frames. Then, we leverage the aligned defocused frame to guide the demoireing of the focused frame using a multi-scale CNN and a multi-dimensional training loss. To maintain tonal and temporal consistency, our final step involves a joint bilateral filter to leverage the demoireing result from the CNN as the guide to filter the input focused frame to obtain the final output. Experimental results demonstrate that our proposed framework largely outperforms state-of-the-art image and video demoireing methods. Xuan Dong 0001, Xiangyuan Sun, Ya Li 0001, Weixin Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | SeeNet: A Soft Emotion Expert and Data Augmentation Method to Enhance Speech Emotion RecognitionabstractSpeech emotion recognition (SER) systems are designed to enable machines to recognize emotional states in human speech during human-computer interactions, enhancing the interactive experience. While considerable progress has been achieved in this field recently, SER systems still encounter challenges related to performance and robustness, primarily stemming from the limited labeled data. To this end, we propose a novel multitask learning framework to learn a distinctive and robust emotional representation by our “Soft Emotion Expert Network (SeeNet)”. SeeNet consists of three components: a pretrained model, an auxiliary task soft emotion expert (SEE) module and an energy-based mixup (EBM) data augmentation module. The pretrained model and EBM module are employed to mitigate the challenges arising from limited labeled data, thereby enhancing the model performance and bolstering robustness. The SEE module as an auxiliary task is designed to assist the main task of SER by enhancing the distinction between samples exhibiting high similarity across categories. This aims to further improve the performance and robustness of the system. Comprehensive experiments on three different settings and multiple datasets are conducted to evaluate the performance and robustness of our proposed method. The experimental results demonstrate that SeeNet surpasses the state-of-the-art (SOTA) methods in both performance and robustness. Yingming Gao, Yuhua Wen, Ziping Zhao 0001, Ya Li 0001, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | DashFusion: Dual-Stream Alignment With Hierarchical Bottleneck Fusion for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) integrates various modalities, such as text, image, and audio, to provide a more comprehensive understanding of sentiment. However, effective MSA is challenged by alignment and fusion issues. Alignment requires synchronizing both temporal and semantic information across modalities, while fusion involves integrating these aligned features into a unified representation. Existing methods often address alignment or fusion in isolation, leading to limitations in performance and efficiency. To tackle these issues, we propose a novel framework called dual-stream alignment with hierarchical bottleneck fusion (DashFusion). First, the dual-stream alignment module synchronizes multimodal features through temporal and semantic alignment. Temporal alignment employs cross-modal attention (CA) to establish frame-level correspondences among multimodal sequences. Semantic alignment ensures consistency across the feature space through contrastive learning. Second, supervised contrastive learning (SCL) leverages label information to refine the modality features. Finally, hierarchical bottleneck fusion (HBF) progressively integrates multimodal information through compressed bottleneck tokens, which achieves a balance between performance and computational efficiency. We evaluate DashFusion on three datasets: CMU-MOSI, CMU-MOSEI, and CH-SIMS. Experimental results demonstrate that DashFusion achieves state-of-the-art (SOTA) performance across various metrics, and ablation studies confirm the effectiveness of our alignment and fusion techniques. The codes for our experiments are available at https://github.com/ultramarineX/DashFusion. Yuhua Wen, Yingying Zhou, Yingming Gao, Zhengqi Wen, Jianhua Tao 0001, Ya Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Concss: Contrastive-based Context Comprehension for Dialogue-Appropriate Prosody in Conversational Speech SynthesisabstractConversational speech synthesis (CSS) incorporates historical dialogue as supplementary information with the aim of generating speech that has dialogue-appropriate prosody. While previous methods have already delved into enhancing context comprehension, context representation still lacks effective representation capabilities and context-sensitive discriminability. In this paper, we introduce a contrastive learning-based CSS framework, CONCSS. Within this framework, we define an innovative pretext task specific to CSS that enables the model to perform self-supervised learning on unlabeled conversational datasets to boost the model’s context understanding. Additionally, we introduce a sampling strategy for negative sample augmentation to enhance context vectors’ discriminability. This is the first attempt to integrate contrastive learning into CSS. We conduct ablation studies on different contrastive learning strategies and comprehensive experiments in comparison with prior CSS systems. Results demonstrate that the synthesized speech from our proposed method exhibits more contextually appropriate and sensitive prosody. Yayue Deng, Jinlong Xue, Yukang Jia, Yichen Han, Fengping Wang, Yingming Gao, Dengfeng Ke, Ya Li 0001 |
ICASSP | 9 |
| 2024 | Frame-Level Emotional State Alignment Method for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) systems aim to recognize human emotional state during human-computer interaction. Most existing SER systems are trained based on utterance-level labels. However, not all frames in an audio have affective states consistent with utterance-level label, which makes it difficult for the model to distinguish the true emotion of the audio and perform poorly. To address this problem, we propose a frame-level emotional state alignment method for SER. First, we fine-tune HuBERT model to obtain an SER system with task-adaptive pretraining (TAPT) method, and extract embeddings from its transformer layers to form frame-level pseudo-emotion labels with clustering. Then, the pseudo labels are used to pretrain HuBERT. Hence, each frame from the output of HuBERT has corresponding emotional information. Finally, we fine-tune the above pretrained HuBERT for SER by adding an attention layer on the top of it, which can focus only on those frames that are emotionally more consistent with utterance-level label. The experimental results performed on IEMOCAP indicate that our proposed method performs better than state-of-the-art (SOTA) methods. The codes are available at github repository1. Yingming Gao, Yayue Deng, Jinlong Xue, Yichen Han, Ya Li 0001 |
ICASSP | 7 |
| 2024 | SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
Bingsong Bai, Fengping Wang, Yingming Gao, Ya Li 0001 |
INTERSPEECH | 4 |
| 2024 | Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
Yingming Gao, Yuhua Wen, Ya Li 0001 |
INTERSPEECH | 5 |
| 2024 | Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
Jinlong Xue, Yayue Deng, Yingming Gao, Ya Li 0001 |
INTERSPEECH | 4 |
| 2024 | Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
Jinlong Xue, Yayue Deng, Yicheng Han, Yingming Gao, Ya Li 0001 |
INTERSPEECH | 5 |
| 2024 | WavDepressionNet: Automatic Depression Level Prediction via Raw Speech SignalsabstractPhysiological reports have confirmed that there are differences in speech signals between depressed and healthy individuals. Therefore, as an application in the field of affective computing, automatic depression level prediction through speech signals has received the attention of researchers, which often estimate the depression severity of individuals by the Fourier or Mel spectrograms of speech signals. However, some studies on speech emotion recognition suggest that directly modeling the raw speech signal is more helpful for extracting emotion-related information. Inspired by this fact, we develop a WavDepressionNet to model raw speech signals for the improvement of prediction accuracy. In our method, a representation block is proposed to find a set of basis vectors to construct the optimal transformation space and generate the transformation result (named Depression Feature Map, DFM) of speech signal for facilitating the perception of depression cues. We further propose an assessment block, which cannot only use the designed spatiotemporal self-calibration mechanism to calibrate the DFM and highlight the useful elements, but also aggregates the calibrated DFM across various temporal ranges with the dilated convolution. Experimental results on the AVEC 2013 and AVEC 2014 depression databases demonstrate the effectiveness of our approach over previous works. Mingyue Niu, Jianhua Tao 0001, Ya Li 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | Articulatory Copy Synthesis Based on the Speech Synthesizer VocalTractLab and Convolutional Recurrent Neural NetworksabstractArticulatory copy synthesis (ACS) refers to the synthetic reproduction of natural utterances. The existing methods of ACS have the limitations of poor generalizability for unknown speakers, high computing costs, the lack of systematic evaluation, etc. Here we propose an ACS method based on the articulatory speech synthesizer VocalTractLab (VTL) and convolutional recurrent neural networks. We first created paired articulatory-acoustic samples using VTL, and then trained neural-network-based ACS models with acoustic features and articulatory trajectories as inputs and outputs, respectively. The basic approach for training relied on fully synthetic training data (and was later supplemented with natural speech and corresponding synthetic articulatory data). In addition, to represent as much of the articulatory and acoustic space as possible, the training samples were augmented by varying the phonation type, speaking effort, and the vocal tract length of the synthetic utterances. Furthermore, two regularization methods were proposed: one based on the smoothness loss of articulatory trajectories and another based on the acoustic loss between original and estimated acoustic features. For given new utterances of arbitrary length, the trained ACS models could estimate articulatory trajectories that were then fed into VTL to synthesize new speech. Experiments showed that our proposed ACS method achieved an average correlation coefficient of 0.983 between the reference and estimated VTL articulatory parameters for speaker-dependent German utterances. When applied to speaker-independent German, English, and Mandarin Chinese utterances, the copy-synthesized speech achieved recognition rates of 73.88%, 52.92%, and 52.41%, respectively, using the automatic speech recognizer Google Speech-to-Text. Yingming Gao, Peter Birkholz, Ya Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio GenerationabstractRecent advancements in diffusion models and large language models (LLMs) have significantly propelled the field of generation tasks. Text-to-Audio (TTA), a burgeoning generation application designed to generate audio from natural language prompts, is attracting increasing attention. However, existing TTA studies often struggle with generation quality and text-audio alignment, especially for complex textual inputs. Drawing inspiration from state-of-the-art Text-to-Image (T2I) diffusion models, we introduce Auffusion, a TTA system adapting T2I model frameworks to TTA task, by effectively leveraging their inherent generative strengths and precise cross-modal alignment. Our objective and subjective evaluations demonstrate that Auffusion surpasses previous TTA approaches using limited data and computational resources. Furthermore, the text encoder serves as a critical bridge between text and audio, since it acts as an instruction for the diffusion model to generate coherent content. Previous studies in T2I recognize the significant impact of encoder choice on cross-modal alignment, like fine-grained details and object bindings, while similar evaluation is lacking in prior TTA works. Through comprehensive ablation studies and innovative cross-attention map visualizations, we provide insightful assessments, being the first to reveal the internal mechanisms in the TTA field and intuitively explain how different text encoders influence the diffusion process. Our findings reveal Auffusion's superior capability in generating audios that accurately match textual descriptions, which is further demonstrated in several related tasks, such as audio style transfer, inpainting, and other manipulations. Jinlong Xue, Yayue Deng, Yingming Gao, Ya Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | DepressionMLP: A Multi-Layer Perceptron Architecture for Automatic Depression Level Prediction via Facial Keypoints and Action UnitsabstractPhysiological studies have confirmed that there are differences in facial activities between depressed and healthy individuals. Therefore, while protecting the privacy of subjects, substantial efforts are made to predict the depression severity of individuals by analyzing Facial Keypoints Representation Sequences (FKRS) and Action Units Representation Sequences (AURS). However, those works has struggled to examine the spatial distribution and temporal changes of Facial Keypoints (FKs) and Action Units (AUs) simultaneously, which is limited in extracting the facial dynamics characterizing depressive cues. Besides, those works don’t realize the complementarity of effective information extracted from FKRS and AURS, which reduces the prediction accuracy. To this end, we intend to use the recently proposed Multi-Layer Perceptrons with gating (gMLP) architecture to process FKRS and AURS for predicting depression levels. However, the channel projection in the gMLP disrupts the spatial distribution of FKs and AUs, leading to input and output sequences not having the same spatiotemporal attributes. This discrepancy hinders the additivity of residual connections in a physical sense. Therefore, we construct a novel MLP architecture named DepressionMLP. In this model, we propose the Dual Gating (DG) and Mutual Guidance (MG) modules. The DG module embeds cross-location and cross-frame gating results into the input sequence to maintain the physical properties of data to make up for the shortcomings of gMLP. The MG module takes the global information of FKRS (AURS) as a guidance mask to filter the AURS (FKRS) to achieve the interaction between FKRS and AURS. Experimental results on several benchmark datasets show the effectiveness of our method. Mingyue Niu, Ya Li 0001, Jianhua Tao 0001, Xiuzhuang Zhou, Björn W. Schuller |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech SynthesisabstractConversational text-to-speech (TTS) aims to synthesize speech with proper prosody of reply based on the historical conversation. However, it is still a challenge to comprehensively model the conversation, and a majority of conversational TTS systems only focus on extracting global information and omit local prosody features, which contain important fine-grained information like keywords and emphasis. Moreover, it is insufficient to only consider the textual features, and acoustic features also contain various prosody information. Hence, we propose M2-CTTS, an end-to-end multi-scale multi-modal conversational text-to-speech system, aiming to comprehensively utilize historical conversation and enhance prosodic expression. More specifically, we design a textual context module and an acoustic context module with both coarse-grained and fine-grained modeling. Experimental results demonstrate that our model mixed with fine-grained context information and additionally considering acoustic features achieves better prosody performance and naturalness in CMOS tests. Jinlong Xue, Yayue Deng, Fengping Wang, Ya Li 0001, Yingming Gao, Jianhua Tao 0001, Jianqing Sun, Jiaen Liang |
ICASSP | 4 |
| 2023 | FTA-net: A Frequency and Time Attention Network for Speech Depression Detection
Yiming Ren 0001, Yingming Gao, Ya Li 0001 |
INTERSPEECH | 5 |
| 2023 | CMCU-CSS: Enhancing Naturalness via Commonsense-based Multi-modal Context Understanding in Conversational Speech SynthesisabstractConversational Speech Synthesis (CSS) aims to produce speech appropriate for oral communication. However, the complexity of context dependency modeling poses significant challenges in the field of CSS, especially the mutual psychological influence between interlocutors. Previous studies have verified that prior commonsense knowledge helps machines understand subtle psychological information (e.g., feelings and intentions) in spontaneous oral dialogues. Therefore, to enhance context understanding and improve the naturalness of synthesized speech, we propose a novel conversational speech synthesis system (CMCU-CSS) that incorporates the Commonsense-based Multi-modal Context Understanding (CMCU) module to model the dynamic emotional interaction among interlocutors. Specifically, we first utilize three implicit states (intent state, internal state and external state) in CMCU to model the context dependency between inter/intra speakers with the help of commonsense knowledge. Furthermore, we infer emotion vectors from the fusion of these implicit states and multi-modal features to enhance the emotion discriminability of synthesized speech. This is the first attempt to combine commonsense knowledge with conversational speech synthesis, and its effect in terms of emotion discriminability of synthetic speech is evaluated by emotion recognition in conversation task. The results of subjective and objective evaluations demonstrate that the CMCU-CSS model achieves more natural speech with context-appropriate emotion and is equipped with the best emotion discriminability, surpassing that of other conversational speech synthesis models. Yayue Deng, Jinlong Xue, Fengping Wang, Yingming Gao, Ya Li 0001 |
ACM Multimedia | 5 |
| 2023 | Mining High-quality Samples from Raw Data and Majority Voting Method for Multimodal Emotion RecognitionabstractAutomatic emotion recognition has a wide range of applications in human-computer interaction. In this paper, we present our work in the Multimodel Emotion Recognition (MER) 2023, which contains three sub-challenges: MER-MULTI, MER-NOISE, and MER-SEMI. We first use a vanilla semi-supervised method to mine high quality samples from the MER-SEMI unlabeled dataset to expand the training set. Specifically, we ensemble three models trained with the official training set by a majority voting method, which is used to select samples with high prediction consistency. The selected samples together with the original training set are further augmented by adding noise. Then, the features of different modalities of expanded dataset are extracted from several pre-trained or fine-tuned models, and they are subsequently used to create different feature combinations to capture more effective emotion representations. Besides, we employ early fusion of different modal features and late fusion of different recognition models to obtain the final prediction. Experimental results show that our proposed method improves the performance over the official baselines by 30.4%, 55.3% and 1.57% for the three sub-challenges and ranks 4, 3, and 5, respectively. The present work sheds light on high-quality data mining and model ensemble by majority voting for multimodal emotion recognition. Yingming Gao, Ya Li 0001 |
ACM Multimedia | 3 |
| 2023 | Dual Attention and Element Recalibration Networks for Automatic Depression Level PredictionabstractPhysiological studies have identified that facial dynamics can be considered as biomarkers to analyze depression severity. This paper accordingly develops a Dual Attention and Element Recalibration (DAER) network to extract facial changes to predict the depression level. In this model, we propose two blocks: a Dual Attention (DA) block and Element Recalibration (ER) block. The DA block uses the self-attention to investigate the dynamic changes in the representation sequence of a facial video segment. It further examines the influence of feature components of the representation sequence on depression level prediction through bilinear-attention. Moreover, to improve the representation ability of network, the ER block is used to obtain the global information to recalibrate each element of the tensor. Adopting this approach, for the depression level prediction task, we first divide the long-term video into fixed-length segments and use the trained ResNet50 to encode each frame to generate the representation sequences of video segments. Second, the representation sequences are input into DAER network to obtain the depression level scores. Finally, the average of these scores yields the prediction result corresponding to the long-term video. Experiments on publicly available AVEC 2013 and AVEC 2014 depression databases illustrate the effectiveness of our method. Mingyue Niu, Ziping Zhao 0001, Jianhua Tao 0001, Ya Li 0001, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | Dual-Lens HDR using Guided 3D Exposure CNN and Guided Denoising TransformerabstractWe study the high dynamic range (HDR) imaging problem in dual-lens systems. Existing methods usually treat the HDR imaging problem as an image fusion problem and the HDR result is estimated by fusing the aligned short exposure image and long exposure image. However, the image fusion pipeline depends highly on the image alignment, which is difficult to be perfect. We propose to transfer the dual-lens HDR imaging problem into the disentangled enhancement of exposure correction and denoising for the short exposure image, guided by the long exposure image. In the guided exposure correction module, we make use of the guidance image and 3D color transformation to propose a guided 3D exposure CNN (GEC) to get the rough HDR result from the short exposure image. Then, in the guided denoising module, we make use of the cross-attention mechanism to propose a guided denoising transformer (GDT) to directly use the long exposure image as guidance to denoise the rough HDR result in a pyramid way. And in both modules, we bypass the difficult image alignment processing. Experimental results demonstrate the superiority of our method over the state-of-the-art ones. Weixin Li 0001, Chang Liu 0071, Xue Tian, Ya Li 0001, Xiaojie Wang 0006, Xuan Dong 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | Automatic Depression Level Assessment from Speech By Long-Term Global Information EmbeddingabstractDepression is a serious mood disorder which brings negative effects on people's social activities. Therefore, growing attention has been paid to automatic depression assessment, especially from speech. However, most of the previous work uses hand-crafted features or deep neural network-based feature extractors to obtain deep features and then feed them into a classifier or a regression, which ignores the temporal relation of these features. To address this issue, this paper proposes a global information embedding (GIE) to make use of the long-term global information of depression and re-weight the LSTM output sequence. The short-term features are then pooled into long-term features by LASSO optimization to further improve the accuracy of depression recognition. Experiments on AVEC 2013 and AVEC 2014 verified the proposed method, and the RMSEs are 9.63 and 9.40, respectively. Ya Li 0001, Mingyue Niu, Ziping Zhao 0001, Jianhua Tao 0001 |
ICASSP | 1 |
| 2022 | Automatic Respiratory Sound Classification Via Multi-Branch Temporal Convolutional NetworkabstractAutomated classification of respiratory sounds has become an active research area in recent years. While recent studies have utilised deep learning methods to aid with respiratory sound classification, the performance is heavily influenced by the datasets available for respiratory sound classification tasks, which tend to be smaller and imbalanced. In this paper, we propose to explore the effectiveness of a multi-branch Temporal Convolutional Network (TCN) architecture integrated with Squeeze-and-Excitation Network (SEnet), a system denoted herein as MBTCNSE, for respiratory sound classification. To the best of the authors’ knowledge, this is the first time that such a hybrid architecture has been employed for respiratory sounds classification. Experiments based on the ICBHI challenge respiratory sound dataset demonstrate the effectiveness of our method. Ziping Zhao 0001, Mingyue Niu, Haishuai Wang, Zixing Zhang 0001, Ya Li 0001 |
ICASSP | 7 |
| 2022 | A Keypoint Based Enhancement Method for Audio Driven Free View Talking Head SynthesisabstractAudio driven talking head synthesis is a challenging task that attracts increasing attention in recent years. Although existing methods based on 2D landmarks or 3D face models can synthesize accurate lip synchronization and rhythmic head pose for arbitrary identity, they still have limitations, such as the cut feeling in the mouth mapping and the lack of skin highlights. The morphed region is blurry compared to the surrounding face. A Keypoint Based Enhancement (KPBE) method is proposed for audio driven free view talking head synthesis to improve the naturalness of the generated video. Firstly, existing methods were used as the backend to synthesize intermediate results. Then we used keypoint decomposition to extract video synthesis controlling parameters from the backend output and the source image. After that, the controlling parameters were composited to the source keypoints and the driving keypoints. A motion field based method was used to generate the final image from the keypoint representation. With keypoint representation, we overcame the cut feeling in the mouth mapping and the lack of skin highlights. Experiments show that our proposed enhancement method improved the quality of talking-head videos in terms of mean opinion score. Yichen Han, Ya Li 0001, Yingming Gao, Jinlong Xue, Songpo Wang |
MMSP | 2 |
| 2022 | Depressioner: Facial dynamic representation for automatic depression level prediction
Mingyue Niu, Ya Li 0001, Bin Liu 0041 |
Expert Syst. Appl. | 3 |
| 2022 | Selective Element and Two Orders Vectorization Networks for Automatic Depression Severity Diagnosis via Facial ChangesabstractPhysiological studies have shown that healthy and depressed individuals present different facial changes. Thus, many researchers have attempted to use Convolutional Neural Networks (CNNs) to extract high-level facial dynamic representations for predicting depression severity. However, the max-pooling (or average-pooling) layers in the CNN lead to the loss of subtle depression cues. Without pooling layers, the CNN cannot extract multi-scale information and has difficulties for tensor vectorization. To this end, we propose a Selective Element and Two Orders Vectorization (SE-TOV) network. For the SE-TOV network, an SE block is constructed to adaptively select the effective elements from the tensors obtained by receptive fields of different sizes. Moreover, we propose a TOV block for vectorizing a high-dimensional tensor. On the one hand, TOV block inputs a tensor into the Global Average Pooling layer to obtain the first-order vectorization result. On the other hand, it takes principal components of the correlation matrix of channels in a tensor as the second-order vectorization result. Experimental results on AVEC 2013 (RMSE$=7.42$, MAE$=6.09$) and AVEC 2014 (RMSE$=7.39$, MAE$=5.87$) depression databases illustrate the superiority of our approach over previous works. Mingyue Niu, Ziping Zhao 0001, Jianhua Tao 0001, Ya Li 0001, Björn W. Schuller |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Discriminative Video Representation with Temporal Order for Micro-expression RecognitionabstractMicro-expression recognition is a challenging task due to its low intensity and short duration and how to extract the subtle facial changes is a key issue in this field. Although there are many methods attempt to cope with this problem, they are difficult to encode the temporal order of all frames in the video clips. For these reasons, this paper employs rank pooling and ℓ2,1-norm to obtain the discriminative video representation with temporal order. In particular, we extract Local Two-Order Gradient Pattern (LTOGP) feature of each frame to describe the subtle information. Then, the video representation is generated by using rank pooling, which captures the temporal order among all frames. Furthermore, considering the sparsity of ℓ2,1-norm, we can select those discriminant features. Finally, micro-expression classification is accomplished using SVM. Experiments are conducted on two publicly available micro-expression databases i.e. CASME and CASME2. The results demonstrate that our method achieves better performance than the state-of-the-art algorithms. Mingyue Niu, Jianhua Tao 0001, Ya Li 0001, Jian Huang 0014, Zheng Lian 0004 |
ICASSP | 3 |
| 2018 | End-to-End Continuous Emotion Recognition from Video Using 3D Convlstm NetworksabstractConventional continuous emotion recognition consists of feature extraction step followed by regression step. However, the objective of the two steps is not consistent as they are parted. Besides, there is still no consensus about appropriate emotional features. In this study, we propose an end-to-end continuous emotion recognition framework which merges feature extraction and regressor into a unified system. We employ 3D convolutional networks with Long Short-Term Memory Neutral Network (ConvLSTM) to handle spatiotemporal information for continuous emotion recognition. This model is applied on AVEC 2017 database. The experiment results reveal that ConvLSTM model makes a positive effect on the performance improvement, which outperforms the baseline results for arousal of 0.583 vs 0.525 (baseline) and for valence of 0.h54 vs 0.507. Jian Huang 0014, Ya Li 0001, Jianhua Tao 0001, Zheng Lian 0004, Jiangyan Yi |
ICASSP | 2 |
| 2018 | Speech Emotion Recognition from Variable-Length Inputs with Triplet Loss Function
Jian Huang 0014, Ya Li 0001, Jianhua Tao 0001, Zhen Lian |
INTERSPEECH | 2 |
| 2018 | BLSTM-CRF Based End-to-End Prosodic Boundary Prediction with Context Sensitive Embeddings in a Text-to-Speech Front-End
Yibin Zheng, Jianhua Tao 0001, Zhengqi Wen, Ya Li 0001 |
INTERSPEECH | 4 |
| 2017 | Distilling Knowledge from an Ensemble of Models for Punctuation Prediction
Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen, Ya Li 0001 |
INTERSPEECH | 4 |
| 2017 | Investigating Efficient Feature Representation Methods and Training Objective for BLSTM-Based Phone Duration Prediction
Yibin Zheng, Jianhua Tao 0001, Zhengqi Wen, Ya Li 0001, Bin Liu 0041 |
INTERSPEECH | 4 |
| 2017 | Quantitative intonation modeling of interrogative sentences for Mandarin speech synthesis
Ya Li 0001, Jianhua Tao 0001, Xiaoying Xu |
Speech Commun. | 1 |
| 2016 | Long short term memory recurrent neural network based encoding method for emotion recognition in videoabstractHuman emotion is a temporally dynamic event which can be inferred from both audio and video feature sequences. In this paper we investigate the long short term memory recurrent neural network (LSTM-RNN) based encoding method for category emotion recognition in the video. LSTM-RNN is able to incorporate knowledge about how emotion evolves over long range successive frames and emotion clues from isolated frame. After encoding, each video clip can be represented by a vector for each input feature sequence. The vectors contain both frame level and sequence level emotion information. These vectors are then concatenated and fed into support vector machine (SVM) to get the final prediction result. Extensive evaluations on Emotion Challenge in the Wild (EmotiW2015) dataset show the efficiency of the proposed encoding method and competitive results are obtained. The final recognition accuracy achieves 46.38% for audio-video emotion recognition sub-challenge, where the challenge baseline is 39.33%. Linlin Chao, Jianhua Tao 0001, Ya Li 0001, Zhengqi Wen |
ICASSP | 4 |
| 2016 | The Parameterized Phoneme Identity Feature as a Continuous Real-Valued Vector for Neural Network Based Speech Synthesis
Zhengqi Wen, Ya Li 0001, Jianhua Tao 0001 |
INTERSPEECH | 2 |
| 2016 | Improving Prosodic Boundaries Prediction for Mandarin Speech Synthesis by Using Enhanced Embedding Feature and Model Fusion Approach
Yibin Zheng, Ya Li 0001, Zhengqi Wen, Xingguang Ding, Jianhua Tao 0001 |
INTERSPEECH | 2 |
| 2015 | Multi task sequence learning for depression scale prediction from videoabstractDepression is a typical mood disorder, which affects people in mental and even physical problems. People who suffer depression always behave abnormal in visual behavior and the voice. In this paper, an audio visual based multimodal depression scale prediction system is proposed. Firstly, features are extracted from video and audio are fused in feature level to represent the audio visual behavior. Secondly, long short memory recurrent neural network (LSTM-RNN) is utilized to encode the dynamic temporal information of the abnormal audio visual behavior. Thirdly, emotion information is utilized by multi-task learning to boost the performance further. The proposed approach is evaluated on the Audio-Visual Emotion Challenge (AVEC2014) dataset. Experiments results show the dimensional emotion recognition helps to depression scale prediction. Linlin Chao, Jianhua Tao 0001, Ya Li 0001 |
ACII | 4 |
| 2015 | From simulated speech to natural speech, what are the robust features for emotion recognition?abstractThe earliest research on emotion recognition starts with simulated/acted stereotypical emotional corpus, and then extends to elicited corpus. Recently, the demanding for real application forces the research shift to natural and spontaneous corpus. Previous research shows that accuracies of emotion recognition are gradual decline from simulated speech, to elicited and totally natural speech. This paper aims to investigate the effects of the common utilized spectral, prosody and voice quality features in emotion recognition with the three types of corpus, and finds out the robust feature for emotion recognition with natural speech. Emotion recognition by several common machine learning methods are carried out and thoroughly compared. Three feature selection methods are performed to find the robust features. The results on six common used corpora confirm that recognition accuracies decrease when the corpus changing from simulated to natural corpus. In addition, prosody and voice quality features are robust for emotion recognition on simulated corpus, while spectral feature is robust in elicited and natural corpus. Ya Li 0001, Linlin Chao, Yazhu Liu, Jianhua Tao 0001 |
ACII | 1 |
| 2015 | Voice quality: Not only about "you" but also about "your interlocutor"abstractThis paper investigates the effect of voice quality in commutative speech. Voice quality is often considered as the characteristic auditory colouring of an individual speaker's voice, but in our study, we find that voice quality can also reveal information about the interlocutor in everyday social interactions. In the correlation analysis between acoustic measures and interlocutors, the effect caused by the linguistic content was reduced by focusing on the corpus of the commonly-used Japanese word “yes”. The distributions of voice quality features, e.g., Normalized Amplitude Quotient (NAQ), jitter and shimmer, showed a clear difference among different interlocutors, e.g., friend or business partner. Automatic classification of interlocutors was conducted using random forest method with voice quality, prosodic and spectral features. The best classification accuracy was 76.9% over four interlocutors' data. Ya Li 0001, Nick Campbell 0001, Jianhua Tao 0001 |
ICASSP | 1 |
| 2015 | A novel method of artificial bandwidth extension using deep architecture
Bin Liu 0041, Jianhua Tao 0001, Zhengqi Wen, Ya Li 0001, Danish Bukhari |
INTERSPEECH | 4 |
| 2015 | Hierarchical stress modeling and generation in mandarin for expressive Text-to-Speech
Ya Li 0001, Jianhua Tao 0001, Keikichi Hirose, Xiaoying Xu |
Speech Commun. | 1 |
| 2014 | A novel hybrid mandarin speech synthesis system using different base units for model training and concatenationabstractThe hybrid speech synthesis system, which uses the acoustic model trained according to the criterion of Maximum Likelihood to select the proper candidates from the corpus, has become a hot topic in recent days. For this hybrid system, the performance is affected by the size of the base training unit and the base candidate unit. Most of existed hybrid systems use the same kind of base unit such as syllable or phone for both model training and concatenation. In Mandarin, initials and finals form the fundamental elements of pronunciation, and are always chosen as the base training unit for statistical parametric TTS system. In this paper a new hybrid Mandarin TTS system is proposed, which uses initial/final for model training and syllable for concatenation. Objective and subjective evaluations are conducted and the comparison results show that the hybrid system we proposed outperforms the traditional systems which use the same base unit for both processes with 4000 and 6000 sentences' corpus. Jianhua Tao 0001, Ya Li 0001, Zhengqi Wen |
ICASSP | 3 |
| 2014 | Improving Mandarin prosodic boundary prediction with rich syntactic featuresabstractPrevious researches indicated that the performance of automatic prosodic boundary labeling benefited from syntactic phrase information for Mandarin. However, the influence of other syntactic features such as dependency has not been studied in-depth yet, especially on large scale corpus. This paper demonstrates the usefulness of rich syntactic features for Mandarin phrase boundary prediction. Both syntactic phrase and dependency features are considered in our methods. The experimental results show that rich syntactic features improve the performance of prosodic boundary prediction effectively. Index Terms: prosodic boundary, syntactic feature, syntactic phrase structure, dependency. Hao Che, Jianhua Tao 0001, Ya Li 0001 |
INTERSPEECH | 3 |
| 2014 | A hierarchical viterbi algorithm for Mandarin hybrid speech synthesis systemabstractThe hybrid speech synthesis system, which combines the
hidden Markov model and unit selection method, has become
an additional main stream in state-of-the-art TTS systems.
However, traditional Viterbi algorithm is based on global
minimization of a cost function and the procedure can end up
selecting some poor-quality units with larger local errors,
which can hardly be tolerated by the listeners. In Mandarin
and many other languages, the naturalness of the region of
consecutive voiced speech segments (CVS) is more essential
to the overall quality of the synthetic speech. Consequently, in
this paper, we proposed to use a hierarchical Viterbi algorithm
which involves two rounds of Viterbi search: one is for the
sub-paths in the CVS regions; the other is for the utterance
path connecting all the sub-paths. In the proposed technique,
we defined CVS Region as a region which is formed by two or
more voiced phones, and whose observation of pitch has a
continuous value. Subjective evaluations suggest that the use
of hierarchical Viterbi algorithm in the Mandarin hybrid
speech synthesis system outperforms the use of traditional
algorithm in both the naturalness and speech quality of
synthetic speech. Zhengqi Wen, Jianhua Tao 0001, Ya Li 0001, Xiaoyan Lou |
INTERSPEECH | 4 |
| 2013 | Bayesian Inference Based Temporal Modeling for Naturalistic Affective Expression ClassificationabstractIn real life, the affective state of human beings changes gradually and smoothly. There is a high probability that the affective state of a certain moment depends on the states of a previous period. In this study, we propose to explicitly model the temporal relationship using a Bayesian inference based two-stage classification approach. This approach could involve knowledge about the dynamics of affective states during a period, so that the inferred affective states are predicted by considering a certain amount of context. Evaluations on the Audio Sub-Challenge of the 2011 Audio/Visual Emotion Challenge show our approach obtains competitive results to those of Audio Sub-Challenge winners. The temporal context modeling method proposed in this paper is also helpful for other sequential pattern recognition problems. Linlin Chao, Jianhua Tao 0001, Ya Li 0001 |
ACII | 4 |
| 2011 | The CASIA Audio Emotion Recognition Method for Audio/Visual Emotion Challenge 2011
Shifeng Pan, Jianhua Tao 0001, Ya Li 0001 |
ACII (2) | 3 |
| 2011 | Hierarchical Stress Modeling in Mandarin Text-to-SpeechabstractAutomatic stress prediction is helpful for both speech synthesis and natural speech understanding. This paper proposes a novel hierarchical Mandarin stress modeling method. The top level emphasizes stressed syllables, while the bottom level focuses on unstressed syllables for the first time due to its importance in both naturalness and expressiveness of synthetic speech. Maximum Entropy model is adopted to predict stress structure from textual features. Experiments show that the modeling method could capture the macro- and micro-characteristics of stress successfully. The F-score of two-level stress predictions are 73.3% and 78.7%, respectively, which are satisfactory compared to other prosody predictions. Index Terms: Text-to-Speech, prosody, stress, Mandarin Ya Li 0001, Jianhua Tao 0001, Xiaoying Xu |
INTERSPEECH | 1 |
| 2010 | Text-based unstressed syllable prediction in MandarinabstractRecently, an increasing attention has been paid to Mandarin word stress which is important for improving the naturalness of speech synthesis. Most of the research on Mandarin speech synthesis focuses on three stress levels: stressed, regular and unstressed. This paper emphasizes the unstressed syllable prediction because the unstressed syllable is also important to the intelligibility of the synthetic speech. Similar as the prosodic structure, it is not easy to detect stress from text analysis due to the complicated context information. A method based on Classification and Regression Tree (CART) model has been proposed to predict the unstressed syllables with the high accuracy of 85%. The method has been finally applied into the TTS system. The experiment shows that the MOS score of synthetic speech has been improved by 0.35; the pitch contour of the new synthesized speech is also closer to natural speech. Index Terms: Text-to-Speech, stress, unstressed syllable, prosody Ya Li 0001, Jianhua Tao 0001, Shifeng Pan, Xiaoying Xu |
INTERSPEECH | 1 |