EDBT 2026 Demo / reviewers in the wild / expert
Yanlu Xie
dblp:92/6633
· DBLP profile ↗
24ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0001-6765-4808ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mispronunciation detection and diagnosis based on large language modelsabstractThis study explores the potential application of Large Language Models (LLMs) in the Mispronunciation Detection and Diagnosis (MDD) system, which includes pronunciation error detection, feedback, and diagnosis. Accurate detection of incorrect pronunciation, along with comprehensive and effective diagnosis, is key to guiding learners in corrective exercises. Traditional MDD research requires collecting data for specific tasks and training models similar to those used in speech recognition. Moreover, most previous research focuses on identifying the types of errors rather than providing specific pronunciation guidance, resulting in textual feedback on pronunciation errors being relatively limited and the content not sufficiently rich. The recent breakthroughs in LLMs have created new opportunities for pronunciation learning through their ability to generate fluent, educationally valuable feedback, such as explaining error types, demonstrating correct pronunciation, and providing personalized practice guidance. This study explores the potential of multimodal speech models in end-to-end pronunciation error detection and feedback generation. Our experiment results show that comprehensive fine-tuning of the Whisper model using second language (L2) speech data can improve its ability to model L2 speech, thereby increasing the accuracy of mispronunciation detection. The feedback text generated by this model is comparable in quality to the current state-of-the-art (SOTA) level based on LLMs (G-Score of 0.52, compared to SOTA’s 0.54). In addition, this study proposes a pronunciation error feedback method based on pronunciation attribute features using LLMs. The LLMs effectively improve the accuracy and effectiveness of feedback text by analyzing the pronunciation attribute features of incorrect phoneme positions. The evaluation of feedback from LLMs by L2 learners indicates a significant improvement in comprehensibility and helpfulness when using these pronunciation representations. These results confirm the potential of articulatory feature engineering and strategic model optimization in CAPT systems by enhancing learner engagement while reducing instructor workload. • Leveraging Large Language Models to Refine Automatic Feedback Generation. • Fine-tuning multimodal LLM and adding Linguistic knowledge to enhance the mispronunciation detection and diagnosis system. • Substituting edit distance with the Viterbi algorithm for precise phoneme-level error localization between predicted and reference sequences. • Erroneous phonemes are mapped to articulatory attribute features, enabling detailed feedback mechanisms. Yanlu Xie, Huihang Zhong, Xuhui Lan, Wenwei Dong |
Comput. Speech Lang. | 1 |
| 2025 | Speech stimulus continuum synthesis using deep learning methods
Yuqing Zhang 0003, Yanlu Xie |
Speech Commun. | 3 |
| 2024 | A Study of Mispronunciation Detection and Diagnosis Based on Meta-LearningabstractThe majority of the current mispronunciation detection and diagnosis (MD&D) methods rely on manually annotated data for model training. However, annotating mispronunciations produced by second language (L2) learners is costly. Consequently, data scarcity emerges as a significant challenge in MD&D tasks. In this paper, we employ model-agnostic meta-learning (MAML) to train a phoneme recognition model for MD&D. We conduct experiments using varied meta-learning task partitioning and training strategies to endow the model’s ability to rapidly adapt to unfamiliar speakers. Our best-performing method achieves an F-measure of 61.45%, surpassing both the method using fine-tuned pre-trained model wav2vec2.0 and the approach of incorporating reference text during training. These related works also aim to address the challenge of data scarcity in MD&D. Notably, with few-shot fine-tuning, our model still yielded some remarkable results on F-measure, which suggest that in MD&D tasks, meta-learning is indeed effective. Yukai Wan, Yuqi Shi, Binghuai Lin, Yanlu Xie |
ICASSP | 4 |
| 2024 | Leveraging Large Language Models to Refine Automatic Feedback Generation at Articulatory Level in Computer Aided Pronunciation Training
Huihang Zhong, Yanlu Xie, ZiJin Yao |
INTERSPEECH | 2 |
| 2023 | LIMI-VC: A Light Weight Voice Conversion Model with Mutual Information DisentanglementabstractVoice conversion(VC) model aims to convert the source timbre to the target one. Recently, many VC models utilize pre-trained models to enhance the performance and achieve good results. However, pre-trained models could not somehow disentangle the timbre and linguistic information, thus resulting in a redundancy, which may hurt the conversion performance. In this paper we proposed LIMI-VC, reducing the redundancy between the linguistic content and the timbre information with mutual information disentanglement. We design the model in a light weight form, for the sake of parameter and computation efficiency when pre-trained models are commonly used nowadays. Experiments show that the proposed model can still improve the performance, with 15 times smaller size, compared to baseline. An out-of-domain cross-lingual inference also shows that our model greatly outperforms the baseline. Our source code and audio examples will be available at: https://github.com/WongLaw/LIMI-VC. Liangjie Huang, Yunming Liang, Can Wen, Yanlu Xie, Jinsong Zhang 0001, Dengfeng Ke |
ICASSP | 6 |
| 2023 | Dual Audio Encoders Based Mandarin Prosodic Boundary Prediction by Using Multi-Granularity Prosodic Representations
Ruishan Li, Yingming Gao, Yanlu Xie, Dengfeng Ke, Jinsong Zhang 0001 |
INTERSPEECH | 3 |
| 2022 | A VR Interactive 3D Mandarin Pronunciation Teaching Model
Yujia Jin, Yanlu Xie, Jinsong Zhang 0001 |
INTERSPEECH | 2 |
| 2021 | Relationships Between Perceptual Distinctiveness, Articulatory Complexity and Functional Load in Speech Communication
Yuqing Zhang 0003, Yanlu Xie, Binghuai Lin, Jinsong Zhang 0001 |
Interspeech | 4 |
| 2020 | A multi-view approach for Mandarin non-native mispronunciation verificationabstractTraditionally, the performance of non-native mispronunciation verification systems relied on effective phone-level labelling of non-native corpora. In this study, a multi-view approach is proposed to incorporate discriminative feature representations which requires less annotation for non-native mispronunciation verification of Mandarin. Here, models are jointly learned to embed acoustic sequence and multi-source information for speech attributes and bottleneck features. Bidirectional LSTM embedding models with contrastive losses are used to map acoustic sequences and multi-source information into fixed-dimensional embeddings. The distance between acoustic embeddings is taken as the similarity between phones. Accordingly, examples of mispronounced phones are expected to have a small similarity score with their canonical pronunciations. The approach shows improvement over GOP-based approach by +11.23% and single-view approach by +1.47% in diagnostic accuracy for a mispronunciation verification task. Zhenyu Wang 0011, John H. L. Hansen, Yanlu Xie |
ICASSP | 3 |
| 2020 | Formant Tracking Using Dilated Convolutional Networks Through Dense Connection with Gating MechanismabstractFormant tracking is one of the most fundamental problems in speech processing.Traditionally, formants are estimated using signal processing methods.Recent studies showed that generic convolutional architectures can outperform recurrent networks on temporal tasks such as speech synthesis and machine translation.In this paper, we explored the use of Temporal Convolutional Network (TCN) for formant tracking.In addition to the conventional implementation, we modified the architecture from three aspects.First, we turned off the "causal" mode of dilated convolution, making the dilated convolution see the future speech frames.Second, each hidden layer reused the output information from all the previous layers through dense connection.Third, we also adopted a gating mechanism to alleviate the problem of gradient disappearance by selectively forgetting unimportant information.The model was validated on the open access formant database VTR.The experiment showed that our proposed model was easy to converge and achieved an overall mean absolute percent error (MAPE) of 8.2% on speech-labeled frames, compared to three competitive baselines of 9.4% (LSTM), 9.1% (Bi-LSTM) and 8.9% (TCN). Wang Dai, Jinsong Zhang 0001, Yingming Gao, Dengfeng Ke, Binghuai Lin, Yanlu Xie |
INTERSPEECH | 7 |
| 2020 | A Dynamic 3D Pronunciation Teaching Model Based on Pronunciation Attributes and Anatomy
Xiaoli Feng, Yanlu Xie, Yayue Deng, Boxue Li |
INTERSPEECH | 2 |
| 2020 | A Mandarin L2 Learning APP with Mispronunciation Detection and Feedback
Yanlu Xie, Xiaoli Feng, Boxue Li, Jinsong Zhang 0001, Yujia Jin |
INTERSPEECH | 1 |
| 2018 | Interactions between Vowels and Nasal Codas in Mandarin Speakers' Perception of Nasal Finals
Yanlu Xie, Jinsong Zhang 0001 |
INTERSPEECH | 4 |
| 2018 | A Preliminary Study on Tonal Coarticulation in Continuous Speech
Lixia Hao, Wei Zhang 0190, Yanlu Xie, Jinsong Zhang 0001 |
INTERSPEECH | 3 |
| 2018 | Improving Mandarin Tone Recognition Using Convolutional Bidirectional Long Short-Term Memory with Attention
Longfei Yang, Yanlu Xie, Jinsong Zhang 0001 |
INTERSPEECH | 2 |
| 2017 | The Influence on Realization and Perception of Lexical Tones from Affricate's Aspiration
Yanlu Xie, Jinsong Zhang 0001 |
INTERSPEECH | 2 |
| 2017 | Reanalyze Fundamental Frequency Peak Delay in Mandarin
Lixia Hao, Wei Zhang 0190, Yanlu Xie, Jinsong Zhang 0001 |
INTERSPEECH | 3 |
| 2016 | Landmark of Mandarin nasal codas and its application in pronunciation error detectionabstractL2 learners of Mandarin have difficulty learning native-like pronunciation of nasal codas. In order to help them learn native-like pronunciation, we propose to develop targeted classifiers for automatic pronunciation error detection. In this paper, perceptual experiments with modified speech are designed to analyze the exact position of the landmark of a nasal coda. Based on perceptual results from isolated words, we propose that information about nasal coda place of articulation is most dense near a landmark at the center of the nasalized vowel. Landmarks detected in a database of Japanese learners of Mandarin, and classified as correct vs. incorrect using an SVM. The result shows that the detection performance of the SVM+Landmark system is similar to that of a DNN-HMM+MFCC system. When the two systems are combined, an FRR of 4.6% is achieved at DA of 83.9%. This performance is comparable to that of previously developed classifiers for 16 common Mandarin pronunciation errors. Yanlu Xie, Mark Hasegawa-Johnson, Leyuan Qu, Jinsong Zhang 0001 |
ICASSP | 1 |
| 2016 | Automatic Pronunciation Evaluation of Non-Native Mandarin Tone by Using Multi-Level Confidence Measures
Ju Lin, Yanlu Xie, Jinsong Zhang 0001 |
INTERSPEECH | 2 |
| 2015 | A study on robust detection of pronunciation erroneous tendency based on deep neural network
Yingming Gao, Yanlu Xie, Jinsong Zhang 0001 |
INTERSPEECH | 2 |
| 2014 | A preliminary study on ASR-based detection of Chinese mispronunciation by Japanese learners
Richeng Duan, Jinsong Zhang 0001, Yanlu Xie |
INTERSPEECH | 4 |
| 2007 | A Segmentation Posterior Based Endpointing AlgorithmabstractA segmentation posterior probability based endpointing algorithm for robust ASR is proposed. First, each speech signal is partitioned into homogeneous segments via auto-segmentation. Then posterior probabilities of all possible endpoints are computed, based on the segmentation likelihoods of all levels in a selected range. Endpoints with the highest posterior probabilities are finally selected. The new method differs from the previous auto-segmentation and clustering based algorithm on that the former considers hypotheses from several levels, while the latter depends only on one appropriate level. Another potential benefit of the proposed method is that any endpointing or VAD results can be integrated, as hypotheses, into the posterior probability framework. Experiments based on the AURORA2 digit database show the robustness of the proposed method. Yanlu Xie, Yu Shi 0001, Frank K. Soong, Beiqian Dai |
ICASSP (4) | 1 |
| 2006 | Improved GMM-UBM/SVM For Speaker VerificationabstractThis paper combines Gaussian Mixture Model-Universal Background Model (GMM-UBM) and Support Vector Machine (SVM) through post processing the GMM-UBM scores of different dimension feature parameter with SVM in speaker verification. Because different dimension feature makes different contribution to recognition performance and SVM has good discriminability, this combining approach yields significant performance improvements on decision-making. Experiments on text-independent speaker verification in NIST05 8conv4w-1 conv4w data showed that the actual detection cost function (DCF) of the test system was reduced to 0.0290 from 0.0343. Beiqian Dai, Yanlu Xie |
ICASSP (1) | 3 |
| 2006 | Kurtosis Normalization in Feature Space for Robust Speaker VerificationabstractThe acoustic mismatch between the training and test environments will lead to the difference of the statistical characteristics of speech parameters. Since the statistical characteristics of the kurtosis can measure the non-Gaussianity of a random variable, kurtosis normalization will make the training and test speech parameters match the standard normal distribution in some sense. In this paper, a kurtosis normalization method using sigmoid functions (logit functions) in feature space is presented for GMM-UBM based text-independent speaker verification system. Experimental results on the 2004 NIST SRE database show that with the new method significant improvement can be achieved in not only equal error rate but also minimum detection cost compared with baseline system (more than 33% relative reduction for long speech). Yanlu Xie, Beiqian Dai |
ICASSP (1) | 1 |