VLDB 2026 Research / reviewers in the wild / expert
Yingming Gao
dblp:173/6544
· DBLP profile ↗
29ranked-venue papers
4as first author
25since 2021 · last 2026
0000-0001-5881-3723ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 21 · 4 first-author · 17 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource ScenariosabstractZero-shot singing voice conversion (SVC) transforms a source singer's timbre to an unseen target speaker's voice while preserving melodic content without fine-tuning. Existing methods model speaker timbre and vocal content separately, losing essential acoustic information that degrades output quality while requiring significant computational resources. To overcome these limitations, we propose HQ-SVC, an efficient framework for high-quality zero-shot SVC. HQ-SVC first extracts jointly content and speaker features using a decoupled codec. It then enhances fidelity through pitch and volume modeling, preserving critical acoustic information typically lost in separate modeling approaches, and progressively refines outputs via differentiable signal processing and diffusion techniques. Evaluations confirm HQ-SVC significantly outperforms state-of-the-art zero-shot SVC methods in conversion quality and efficiency. Beyond voice conversion, HQ-SVC achieves superior voice naturalness compared to specialized audio super-resolution methods while natively supporting voice super-resolution tasks. Bingsong Bai, Yizhong Geng, Fengping Wang, Puyuan Guo, Yingming Gao, Ya Li 0001 |
AAAI | 6 |
| 2025 | Controllable 3D Dance Generation Using Diffusion-Based Transformer U-NetabstractRecently, dance generation has attracted increasing interest. In particular, the success of diffusion models in image generation has led to the emergence of dance generation systems based on the diffusion framework. However, these systems lack controllability, which limits their practical applications. In this paper, we propose a controllable dance generation method based on the diffusion model, which can generate 3D dance motions controlled by 2D keypoint sequences. Specifically, we design a transformer-based U-Net model to predict actual motions. Then, we fix the parameters of the U-Net model and train an additional control network, enabling the generated motions to be controlled by 2D keypoints. We conduct extensive experiments and compared our method with existing works on the widely used AIST++ dataset, demonstrating that our approach has certain advantages and controllability. Moreover, we also test our model on in-the-wild videos and find that it is capable of generating dance movements similar to the motions in the videos as well. Puyuan Guo, Tuo Hao, Wenxin Fu, Yingming Gao, Ya Li 0001 |
AAAI | 4 |
| 2025 | Beyond Surface Simplicity: Revealing Hidden Reasoning Attributes for Precise Commonsense DiagnosisabstractCommonsense question answering (QA) are widely used to evaluate the commonsense abilities of large language models. However, answering commonsense questions correctly requires not only knowledge but also reasoning—even for seemingly simple questions. We demonstrate that such hidden reasoning attributes in commonsense questions can lead evaluation accuracy differences of up to 24.8% across different difficulty levels in the same benchmark. Current benchmarks overlook these hidden reasoning attributes, making it difficult to assess a model’s specific levels of commonsense knowledge and reasoning ability. To address this issue, we introduce ReComSBench, a novel framework that reveals hidden reasoning attributes behind commonsense questions by leveraging the knowledge generated during the reasoning process. Additionally, ReComSBench proposes three new metrics for decoupled evaluation: Knowledge Balanced Accuracy, Marginal Sampling Gain, and Knowledge Coverage Ratio. Experiments show that ReComSBench provides insights into model performance that traditional benchmarks cannot offer. The difficulty stratification based on revealed hidden reasoning attributes performs as effectively as the model-probability-based approach but is more generalizable and better suited for improving a model’s commonsense reasoning abilities. By uncovering and analyzing the hidden reasoning attributes in commonsense data, ReComSBench offers a new approach to enhancing existing commonsense benchmarks. Huijun Lian, Zekai Sun, Yingming Gao, Ya Li 0001 |
ACL (1) | 4 |
| 2025 | DetailTTS: Learning Residual Detail Information for Zero-shot Text-to-speechabstractTraditional text-to-speech (TTS) systems often face challenges in aligning text and speech, leading to the omission of critical linguistic and acoustic details. This misalignment creates an information gap, which existing methods attempt to address by incorporating additional inputs, but these often introduce data inconsistencies and increase complexity. To address these issues, we propose DetailTTS, a zero-shot TTS system based on a conditional variational autoencoder. It incorporates two key components: the Prior Detail Module and the Duration Detail Module, which capture residual detail information missed during alignment. These modules effectively enhance the model’s ability to retain fine-grained details, significantly improving speech quality while simplifying the model by obviating the need for additional inputs. Experiments on the WenetSpeech4TTS dataset show that DetailTTS outperforms traditional TTS systems in both naturalness and speaker similarity, even in zero-shot scenarios. Our source code and demo page are available at https://detailtts.github.io/. Yichen Han, Yizhong Geng, Yingming Gao, Fengping Wang, Bingsong Bai, Jinlong Xue, Yayue Deng, Zhengqi Wen, Ya Li 0001 |
ICASSP | 4 |
| 2025 | EEG-based Voice Conversion : Hearing the Voice of Your Brain
Yizhong Geng, Wenxin Fu, Qihang Lu, Bingsong Bai, Yingming Gao, Ya Li 0001 |
INTERSPEECH | 6 |
| 2025 | EMO-Avatar: An LLM-Agent-Orchestrated Framework for Multimodal Emotional Support in Human Animation
Wenxin Fu, Qihang Lu, Zekai Sun, Yizhong Geng, Puyuan Guo, Yingming Gao, Ya Li 0001 |
ACM Multimedia | 8 |
| 2025 | SeeNet: A Soft Emotion Expert and Data Augmentation Method to Enhance Speech Emotion RecognitionabstractSpeech emotion recognition (SER) systems are designed to enable machines to recognize emotional states in human speech during human-computer interactions, enhancing the interactive experience. While considerable progress has been achieved in this field recently, SER systems still encounter challenges related to performance and robustness, primarily stemming from the limited labeled data. To this end, we propose a novel multitask learning framework to learn a distinctive and robust emotional representation by our “Soft Emotion Expert Network (SeeNet)”. SeeNet consists of three components: a pretrained model, an auxiliary task soft emotion expert (SEE) module and an energy-based mixup (EBM) data augmentation module. The pretrained model and EBM module are employed to mitigate the challenges arising from limited labeled data, thereby enhancing the model performance and bolstering robustness. The SEE module as an auxiliary task is designed to assist the main task of SER by enhancing the distinction between samples exhibiting high similarity across categories. This aims to further improve the performance and robustness of the system. Comprehensive experiments on three different settings and multiple datasets are conducted to evaluate the performance and robustness of our proposed method. The experimental results demonstrate that SeeNet surpasses the state-of-the-art (SOTA) methods in both performance and robustness. Yingming Gao, Yuhua Wen, Ziping Zhao 0001, Ya Li 0001, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Disentanglement of Prosody Representations via Diffusion Models and Scheduled Gradient ReversalabstractProsody plays a fundamental role in human speech and communication, facilitating intelligibility and conveying emotional and cognitive states. Extracting accurate prosodic information from speech is vital for building assistive technology, such as controllable speech synthesis, speaking style transfer, and speech emotion recognition (SER). However, it is challenging to disentangle speaker-independent prosody representations since prosodic attributes, such as intonation, excessively entangle with speaker-specific attributes, e.g., pitch. In this article, we propose a novel model, called Diffsody, to disentangle and refine prosody representations: 1) to disentangle prosody representations, we leverage the expressive generative ability of a diffusion model by conditioning it on quantified semantic information and pretrained speaker embeddings. Additionally, a prosody encoder automatically learns prosody representations used for spectrogram reconstruction in an unsupervised fashion; and 2) to refine and learn speaker-invariant prosody representations, a scheduled gradient reversal layer (sGRL) is proposed and integrated into the prosody encoder of Diffsody. We extensively evaluate Diffsody through qualitative and quantitative means. t-SNE visualization and speaker verification experiments demonstrate the efficacy of the sGRL method in preventing speaker-specific information leakage. Experimental results on speaker-independent SER and automatic depression detection (ADD) tasks demonstrate that Diffsody can efficiently factorize speaker-independent prosody representations, resulting in a significant boost in SER and ADD. In addition, Diffsody synergistically integrates with the semantic representation model WavLM, which leads to a discernibly elevated performance, outperforming contemporary methods in both SER and ADD tasks. Furthermore, the Diffsody model exhibits promising potential for various practical applications, such as voice or style conversion. Some audio samples can be found on our https://leyuanqu.github.io/Diffsody/demo website. Leyuan Qu, Cornelius Weber, Wei Wang 0310, Jia Jin, Yingming Gao, Taihao Li, Stefan Wermter |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | DashFusion: Dual-Stream Alignment With Hierarchical Bottleneck Fusion for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) integrates various modalities, such as text, image, and audio, to provide a more comprehensive understanding of sentiment. However, effective MSA is challenged by alignment and fusion issues. Alignment requires synchronizing both temporal and semantic information across modalities, while fusion involves integrating these aligned features into a unified representation. Existing methods often address alignment or fusion in isolation, leading to limitations in performance and efficiency. To tackle these issues, we propose a novel framework called dual-stream alignment with hierarchical bottleneck fusion (DashFusion). First, the dual-stream alignment module synchronizes multimodal features through temporal and semantic alignment. Temporal alignment employs cross-modal attention (CA) to establish frame-level correspondences among multimodal sequences. Semantic alignment ensures consistency across the feature space through contrastive learning. Second, supervised contrastive learning (SCL) leverages label information to refine the modality features. Finally, hierarchical bottleneck fusion (HBF) progressively integrates multimodal information through compressed bottleneck tokens, which achieves a balance between performance and computational efficiency. We evaluate DashFusion on three datasets: CMU-MOSI, CMU-MOSEI, and CH-SIMS. Experimental results demonstrate that DashFusion achieves state-of-the-art (SOTA) performance across various metrics, and ablation studies confirm the effectiveness of our alignment and fusion techniques. The codes for our experiments are available at https://github.com/ultramarineX/DashFusion. Yuhua Wen, Yingying Zhou, Yingming Gao, Zhengqi Wen, Jianhua Tao 0001, Ya Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Concss: Contrastive-based Context Comprehension for Dialogue-Appropriate Prosody in Conversational Speech SynthesisabstractConversational speech synthesis (CSS) incorporates historical dialogue as supplementary information with the aim of generating speech that has dialogue-appropriate prosody. While previous methods have already delved into enhancing context comprehension, context representation still lacks effective representation capabilities and context-sensitive discriminability. In this paper, we introduce a contrastive learning-based CSS framework, CONCSS. Within this framework, we define an innovative pretext task specific to CSS that enables the model to perform self-supervised learning on unlabeled conversational datasets to boost the model’s context understanding. Additionally, we introduce a sampling strategy for negative sample augmentation to enhance context vectors’ discriminability. This is the first attempt to integrate contrastive learning into CSS. We conduct ablation studies on different contrastive learning strategies and comprehensive experiments in comparison with prior CSS systems. Results demonstrate that the synthesized speech from our proposed method exhibits more contextually appropriate and sensitive prosody. Yayue Deng, Jinlong Xue, Yukang Jia, Yichen Han, Fengping Wang, Yingming Gao, Dengfeng Ke, Ya Li 0001 |
ICASSP | 7 |
| 2024 | Frame-Level Emotional State Alignment Method for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) systems aim to recognize human emotional state during human-computer interaction. Most existing SER systems are trained based on utterance-level labels. However, not all frames in an audio have affective states consistent with utterance-level label, which makes it difficult for the model to distinguish the true emotion of the audio and perform poorly. To address this problem, we propose a frame-level emotional state alignment method for SER. First, we fine-tune HuBERT model to obtain an SER system with task-adaptive pretraining (TAPT) method, and extract embeddings from its transformer layers to form frame-level pseudo-emotion labels with clustering. Then, the pseudo labels are used to pretrain HuBERT. Hence, each frame from the output of HuBERT has corresponding emotional information. Finally, we fine-tune the above pretrained HuBERT for SER by adding an attention layer on the top of it, which can focus only on those frames that are emotionally more consistent with utterance-level label. The experimental results performed on IEMOCAP indicate that our proposed method performs better than state-of-the-art (SOTA) methods. The codes are available at github repository1. Yingming Gao, Yayue Deng, Jinlong Xue, Yichen Han, Ya Li 0001 |
ICASSP | 2 |
| 2024 | SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
Bingsong Bai, Fengping Wang, Yingming Gao, Ya Li 0001 |
INTERSPEECH | 3 |
| 2024 | Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
Yingming Gao, Yuhua Wen, Ya Li 0001 |
INTERSPEECH | 2 |
| 2024 | Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
Jinlong Xue, Yayue Deng, Yingming Gao, Ya Li 0001 |
INTERSPEECH | 3 |
| 2024 | Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
Jinlong Xue, Yayue Deng, Yicheng Han, Yingming Gao, Ya Li 0001 |
INTERSPEECH | 4 |
| 2024 | Articulatory Copy Synthesis Based on the Speech Synthesizer VocalTractLab and Convolutional Recurrent Neural NetworksabstractArticulatory copy synthesis (ACS) refers to the synthetic reproduction of natural utterances. The existing methods of ACS have the limitations of poor generalizability for unknown speakers, high computing costs, the lack of systematic evaluation, etc. Here we propose an ACS method based on the articulatory speech synthesizer VocalTractLab (VTL) and convolutional recurrent neural networks. We first created paired articulatory-acoustic samples using VTL, and then trained neural-network-based ACS models with acoustic features and articulatory trajectories as inputs and outputs, respectively. The basic approach for training relied on fully synthetic training data (and was later supplemented with natural speech and corresponding synthetic articulatory data). In addition, to represent as much of the articulatory and acoustic space as possible, the training samples were augmented by varying the phonation type, speaking effort, and the vocal tract length of the synthetic utterances. Furthermore, two regularization methods were proposed: one based on the smoothness loss of articulatory trajectories and another based on the acoustic loss between original and estimated acoustic features. For given new utterances of arbitrary length, the trained ACS models could estimate articulatory trajectories that were then fed into VTL to synthesize new speech. Experiments showed that our proposed ACS method achieved an average correlation coefficient of 0.983 between the reference and estimated VTL articulatory parameters for speaker-dependent German utterances. When applied to speaker-independent German, English, and Mandarin Chinese utterances, the copy-synthesized speech achieved recognition rates of 73.88%, 52.92%, and 52.41%, respectively, using the automatic speech recognizer Google Speech-to-Text. Yingming Gao, Peter Birkholz, Ya Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2024 | Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio GenerationabstractRecent advancements in diffusion models and large language models (LLMs) have significantly propelled the field of generation tasks. Text-to-Audio (TTA), a burgeoning generation application designed to generate audio from natural language prompts, is attracting increasing attention. However, existing TTA studies often struggle with generation quality and text-audio alignment, especially for complex textual inputs. Drawing inspiration from state-of-the-art Text-to-Image (T2I) diffusion models, we introduce Auffusion, a TTA system adapting T2I model frameworks to TTA task, by effectively leveraging their inherent generative strengths and precise cross-modal alignment. Our objective and subjective evaluations demonstrate that Auffusion surpasses previous TTA approaches using limited data and computational resources. Furthermore, the text encoder serves as a critical bridge between text and audio, since it acts as an instruction for the diffusion model to generate coherent content. Previous studies in T2I recognize the significant impact of encoder choice on cross-modal alignment, like fine-grained details and object bindings, while similar evaluation is lacking in prior TTA works. Through comprehensive ablation studies and innovative cross-attention map visualizations, we provide insightful assessments, being the first to reveal the internal mechanisms in the TTA field and intuitively explain how different text encoders influence the diffusion process. Our findings reveal Auffusion's superior capability in generating audios that accurately match textual descriptions, which is further demonstrated in several related tasks, such as audio style transfer, inpainting, and other manipulations. Jinlong Xue, Yayue Deng, Yingming Gao, Ya Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech SynthesisabstractConversational text-to-speech (TTS) aims to synthesize speech with proper prosody of reply based on the historical conversation. However, it is still a challenge to comprehensively model the conversation, and a majority of conversational TTS systems only focus on extracting global information and omit local prosody features, which contain important fine-grained information like keywords and emphasis. Moreover, it is insufficient to only consider the textual features, and acoustic features also contain various prosody information. Hence, we propose M2-CTTS, an end-to-end multi-scale multi-modal conversational text-to-speech system, aiming to comprehensively utilize historical conversation and enhance prosodic expression. More specifically, we design a textual context module and an acoustic context module with both coarse-grained and fine-grained modeling. Experimental results demonstrate that our model mixed with fine-grained context information and additionally considering acoustic features achieves better prosody performance and naturalness in CMOS tests. Jinlong Xue, Yayue Deng, Fengping Wang, Ya Li 0001, Yingming Gao, Jianhua Tao 0001, Jianqing Sun, Jiaen Liang |
ICASSP | 5 |
| 2023 | Dual Audio Encoders Based Mandarin Prosodic Boundary Prediction by Using Multi-Granularity Prosodic Representations
Ruishan Li, Yingming Gao, Yanlu Xie, Dengfeng Ke, Jinsong Zhang 0001 |
INTERSPEECH | 2 |
| 2023 | FTA-net: A Frequency and Time Attention Network for Speech Depression Detection
Yiming Ren 0001, Yingming Gao, Ya Li 0001 |
INTERSPEECH | 4 |
| 2023 | CMCU-CSS: Enhancing Naturalness via Commonsense-based Multi-modal Context Understanding in Conversational Speech SynthesisabstractConversational Speech Synthesis (CSS) aims to produce speech appropriate for oral communication. However, the complexity of context dependency modeling poses significant challenges in the field of CSS, especially the mutual psychological influence between interlocutors. Previous studies have verified that prior commonsense knowledge helps machines understand subtle psychological information (e.g., feelings and intentions) in spontaneous oral dialogues. Therefore, to enhance context understanding and improve the naturalness of synthesized speech, we propose a novel conversational speech synthesis system (CMCU-CSS) that incorporates the Commonsense-based Multi-modal Context Understanding (CMCU) module to model the dynamic emotional interaction among interlocutors. Specifically, we first utilize three implicit states (intent state, internal state and external state) in CMCU to model the context dependency between inter/intra speakers with the help of commonsense knowledge. Furthermore, we infer emotion vectors from the fusion of these implicit states and multi-modal features to enhance the emotion discriminability of synthesized speech. This is the first attempt to combine commonsense knowledge with conversational speech synthesis, and its effect in terms of emotion discriminability of synthetic speech is evaluated by emotion recognition in conversation task. The results of subjective and objective evaluations demonstrate that the CMCU-CSS model achieves more natural speech with context-appropriate emotion and is equipped with the best emotion discriminability, surpassing that of other conversational speech synthesis models. Yayue Deng, Jinlong Xue, Fengping Wang, Yingming Gao, Ya Li 0001 |
ACM Multimedia | 4 |
| 2023 | Mining High-quality Samples from Raw Data and Majority Voting Method for Multimodal Emotion RecognitionabstractAutomatic emotion recognition has a wide range of applications in human-computer interaction. In this paper, we present our work in the Multimodel Emotion Recognition (MER) 2023, which contains three sub-challenges: MER-MULTI, MER-NOISE, and MER-SEMI. We first use a vanilla semi-supervised method to mine high quality samples from the MER-SEMI unlabeled dataset to expand the training set. Specifically, we ensemble three models trained with the official training set by a majority voting method, which is used to select samples with high prediction consistency. The selected samples together with the original training set are further augmented by adding noise. Then, the features of different modalities of expanded dataset are extracted from several pre-trained or fine-tuned models, and they are subsequently used to create different feature combinations to capture more effective emotion representations. Besides, we employ early fusion of different modal features and late fusion of different recognition models to obtain the final prediction. Experimental results show that our proposed method improves the performance over the official baselines by 30.4%, 55.3% and 1.57% for the three sub-challenges and ranks 4, 3, and 5, respectively. The present work sheds light on high-quality data mining and model ensemble by majority voting for multimodal emotion recognition. Yingming Gao, Ya Li 0001 |
ACM Multimedia | 2 |
| 2022 | A study of production error analysis for Mandarin-speaking Children with Hearing Impairment
Jingwen Cheng, Yingming Gao, Xiaoli Feng, Yannan Wang, Jinsong Zhang 0001 |
INTERSPEECH | 3 |
| 2022 | A Keypoint Based Enhancement Method for Audio Driven Free View Talking Head SynthesisabstractAudio driven talking head synthesis is a challenging task that attracts increasing attention in recent years. Although existing methods based on 2D landmarks or 3D face models can synthesize accurate lip synchronization and rhythmic head pose for arbitrary identity, they still have limitations, such as the cut feeling in the mouth mapping and the lack of skin highlights. The morphed region is blurry compared to the surrounding face. A Keypoint Based Enhancement (KPBE) method is proposed for audio driven free view talking head synthesis to improve the naturalness of the generated video. Firstly, existing methods were used as the backend to synthesize intermediate results. Then we used keypoint decomposition to extract video synthesis controlling parameters from the backend output and the source image. After that, the controlling parameters were composited to the source keypoints and the driving keypoints. A motion field based method was used to generate the final image from the keypoint representation. With keypoint representation, we overcame the cut feeling in the mouth mapping and the lack of skin highlights. Experiments show that our proposed enhancement method improved the quality of talking-head videos in terms of mean opinion score. Yichen Han, Ya Li 0001, Yingming Gao, Jinlong Xue, Songpo Wang |
MMSP | 3 |
| 2022 | Articulatory Synthesis of Vocalized /r/ Allophones in GermanabstractArticulatory synthesis relies on precise, parametric vocal tract shapes to generate natural-sounding speech. In German, a particular challenge is the accurate synthesis of the vocalic /r/ allophones following vowels or in syllable coda position. Using established phonetic conventions, no satisfying results could be achieved so far, implying a possible shortcoming of these existing conventions. This study therefore analyzed a large number of natural recordings of the sounds in question from a single speaker to find the optimal number of target vocal tract shapes. Applying clustering techniques, the manifold of vocalic /r/ allophones could be reduced to two prototypical [ɐ] variants previously undescribed in the literature. As shown by a listening experiment, which of these two sounds was preferred in which context depended not only on the respective context vowel’s tenseness, but also on its openness and acoustic distance to other context vowels ending in the same [ɐ] variant. This indicates that the two different allophones might serve as a contrastive cue to help differentiate between otherwise similar context vowels. Simon Stone, Yingming Gao, Peter Birkholz |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Formant Tracking Using Dilated Convolutional Networks Through Dense Connection with Gating MechanismabstractFormant tracking is one of the most fundamental problems in speech processing.Traditionally, formants are estimated using signal processing methods.Recent studies showed that generic convolutional architectures can outperform recurrent networks on temporal tasks such as speech synthesis and machine translation.In this paper, we explored the use of Temporal Convolutional Network (TCN) for formant tracking.In addition to the conventional implementation, we modified the architecture from three aspects.First, we turned off the "causal" mode of dilated convolution, making the dilated convolution see the future speech frames.Second, each hidden layer reused the output information from all the previous layers through dense connection.Third, we also adopted a gating mechanism to alleviate the problem of gradient disappearance by selectively forgetting unimportant information.The model was validated on the open access formant database VTR.The experiment showed that our proposed model was easy to converge and achieved an overall mean absolute percent error (MAPE) of 8.2% on speech-labeled frames, compared to three competitive baselines of 9.4% (LSTM), 9.1% (Bi-LSTM) and 8.9% (TCN). Wang Dai, Jinsong Zhang 0001, Yingming Gao, Dengfeng Ke, Binghuai Lin, Yanlu Xie |
INTERSPEECH | 3 |
| 2020 | An Investigation of the Target Approximation Model for Tone Modeling and Recognition in Continuous Mandarin SpeechabstractThe complex f0 variations in continuous speech make it rather difficult to perform automatic recognition of tones in a language like Mandarin Chinese.In this study, we tested the use of target approximation model (TAM) for continuous tone recognition on two datasets.TAM simulates f0 production from the articulatory point of view and so allow to discover the underlying pitch targets from the surface f0 contour.The f0 contour of each tone represented by 30 equidistant points in the first dataset was simulated by the TAM model.Using a support vector machine (SVM) to classify tones showed that, compared to the representation by 30 f0 values, the estimated three-dimensional TAM parameters had a comparable performance in characterizing tone patterns.The TAM model was further tested on the second dataset containing more complex tonal variations.With equal or a fewer number of features, the TAM parameters provided better performance than the coefficients of the cosine transform and a slightly worse performance than the statistical f0 parameters for tone recognition.Furthermore, we investigated bidirectional LSTM neural network for modelling the sequential tonal variations, which proved to be more powerful than the SVM classifier.The BLSTM system incorporating TAM and statistical f0 parameters achieved the best accuracy of 87.56%. Yingming Gao, Jinsong Zhang 0001, Peter Birkholz |
INTERSPEECH | 1 |
| 2019 | Articulatory Copy Synthesis Based on a Genetic Algorithm
Yingming Gao, Simon Stone, Peter Birkholz |
INTERSPEECH | 1 |
| 2015 | A study on robust detection of pronunciation erroneous tendency based on deep neural network
Yingming Gao, Yanlu Xie, Jinsong Zhang 0001 |
INTERSPEECH | 1 |