EDBT 2026 Demo / reviewers in the wild / expert
Guanlong Zhao
dblp:201/7257
· DBLP profile ↗
21ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0002-6059-4053ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Personalizing Keyword Spotting with Speaker InformationabstractKeyword spotting systems often struggle to generalize to a diverse population with various accents and age groups. To address this challenge, we propose a novel approach that integrates speaker information into keyword spotting using Feature-wise Linear Modulation (FiLM), a recent method that allows models to learn from different data inputs and features. We explore both Text-Dependent and Text-Independent speaker recognition systems to extract speaker information, and we experiment on extracting this information from both the input audio and pre-enrolled user audio. Evaluating our systems on a diverse dataset, our primary approach yields a notable 2.6% relative improvement on Equal Error Rate overall, particularly improving performance by 5.9% for children under 12 years old and up to 24% for underrepresented speaker groups. Moreover, our proposed approach only requires a small 1% increase in the number of parameters, with a minimum impact on latency and computational cost, which makes it a practical solution for real-world applications. Beltran Labrador, Pai Zhu, Guanlong Zhao, Angelo Scorza Scarpati, Alicia Lozano-Diez, Ignacio López-Moreno |
ICASSP | 3 |
| 2024 | USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation ModelsabstractWe introduce a multilingual speaker change detection model (USM-SCD) that can simultaneously detect speaker turns and perform ASR for 96 languages. This model is adapted from a speech foundation model trained on a large quantity of supervised and unsupervised data, demonstrating the utility of fine-tuning from a large generic foundation model for a downstream task. We analyze the performance of this multilingual speaker change detection model through a series of ablation studies. We show that the USM-SCD model can achieve more than 75% average speaker change detection F1 score across a test set that consists of data from 96 languages. On American English, the USM-SCD model can achieve an 85.8% speaker change detection F1 score across various public and internal test sets, beating the previous monolingual baseline model by 21% relative. We also show that we only need to fine-tune one-quarter of the trainable model parameters to achieve the best model performance. The USM-SCD model exhibits state-of-the-art ASR quality compared with a strong public ASR baseline, making it suitable to handle both tasks with negligible additional computational cost. Guanlong Zhao, Jason Pelecanos, Hank Liao, Han Lu 0003 |
ICASSP | 1 |
| 2024 | On the Success and Limitations of Auxiliary Network Based Word-Level End-to-End Neural Speaker Diarization
Guanlong Zhao, Hank Liao |
INTERSPEECH | 3 |
| 2024 | DiarizationLM: Speaker Diarization Post-Processing with Large Language ModelsabstractIn this paper, we introduce DiarizationLM, a framework to leverage large language models (LLM) to post-process the outputs from a speaker diarization system.In this framework, the outputs of the automatic speech recognition (ASR) and speaker diarization systems are represented as a compact textual format, which is included in the prompt to an optionally finetuned LLM.The outputs of the LLM can be used as the refined diarization results with the desired enhancement.As a post-processing step, this framework can be easily applied to any off-the-shelf ASR and speaker diarization systems without retraining existing components.Our experiments show that a finetuned PaLM 2-S model can reduce the WDER by rel.55.5% on the Fisher telephone conversation dataset, and rel.44.9% on the Callhome English dataset. Guanlong Zhao, Evan Clark, Hank Liao |
INTERSPEECH | 3 |
| 2023 | Exploring Sequence-to-Sequence Transformer-Transducer Models for Keyword SpottingabstractIn this paper, we present a novel approach to adapt a sequence-to-sequence Transformer-Transducer ASR system to the keyword spotting (KWS) task. We achieve this by replacing the keyword in the text transcription with a special tokenand training the system to detect thetoken in an audio stream. At inference time, we create a decision function inspired by conventional KWS approaches, to make our approach more suitable for the KWS task. Furthermore, we introduce a specific keyword spotting loss by adapting the sequence-discriminative Minimum Bayes-Risk training technique. We find that our approach significantly outperforms ASR based KWS systems. When compared with a conventional keyword spotting system, our proposal has similar performance while bringing the advantages and flexibility of sequence-to-sequence training. Additionally, when combined with the conventional KWS system, our approach can improve the performance at any operation point. Beltran Labrador, Guanlong Zhao, Ignacio López-Moreno, Angelo Scorza Scarpati, Liam Fowl |
ICASSP | 2 |
| 2023 | Augmenting Transformer-Transducer Based Speaker Change Detection with Token-Level Training LossabstractIn this work we propose a novel token-based training strategy that improves Transformer-Transducer (T-T) based speaker change detection (SCD) performance. The conventional T-T based SCD model loss optimizes all output tokens equally. Due to the sparsity of the speaker changes in the training data, the conventional T-T based SCD model loss leads to sub-optimal detection accuracy. To mitigate this issue, we use a customized edit-distance algorithm to estimate the token-level SCD false accept (FA) and false reject (FR) rates during training and optimize model parameters to minimize a weighted combination of the FA and FR, focusing the model on accurately predicting speaker changes. We also propose a set of evaluation metrics that align better with commercial use cases. Experiments on a group of challenging real-world datasets show that the proposed training method can significantly improve the overall performance of the SCD model with the same number of parameters. Guanlong Zhao, Han Lu 0003, Ignacio López-Moreno |
ICASSP | 1 |
| 2022 | Accentron: Foreign accent conversion to arbitrary non-native speakers using zero-shot learning
Shaojin Ding, Guanlong Zhao, Ricardo Gutierrez-Osuna |
Comput. Speech Lang. | 2 |
| 2021 | Assessing Posterior-Based Mispronunciation Detection on Field-Collected Recordings from Child Speech Therapy Sessions
Adam Hair, Guanlong Zhao, Beena Ahmed, Kirrie J. Ballard, Ricardo Gutierrez-Osuna |
Interspeech | 2 |
| 2021 | Effects of Voice Type and Task on L2 Learners' Awareness of Pronunciation ErrorsabstractResearch suggests learners may improve their second language (L2) pronunciation by imitating voices with similar acoustic profiles. However, previously reported improvements have been in suprasegmentals (prosodic features such as intonation). It remains unclear if voice similarity applies to L2 segmentals (consonants and vowels). To address this issue, this study investigates how voice similarity facilitates awareness of pronunciation errors, a necessary step in pronunciation improvement. In two experiments, advanced L2 learners identified their pronunciation errors by comparing their production to the production of a resynthesized model voice using learners’ voices as the base (Golden Speaker voice), or to an unfamiliar resynthesized voice with the same gender as the learner (Silver Speaker voice). In Experiment 1, L2 learners identified all syllables with vowel and consonant errors when comparing their production to the model voice. Their choices were compared to identifications by expert judges. In Experiment 2, learners were told how many errors the expert judges had identified before identifying the same number of errors. Results did not support facilitative effects of Golden Speaker voices in either experiment, but Experiment 2 resulted in higher identification percentages. Discussion of the challenges in self-identification of errors in relation to voice similarity are offered. Alif Silpachai, Ivana Rehman, Taylor Anne Barriuso, John Levis, Evgeny Chukharev-Hudilainen, Guanlong Zhao, Ricardo Gutierrez-Osuna |
Interspeech | 6 |
| 2021 | Converting Foreign Accent Speech Without a ReferenceabstractForeign accent conversion (FAC) is the problem of generating a synthetic voice that has the voice identity of a second-language (L2) learner and the pronunciation patterns of a native (L1) speaker. This synthetic voice has been referred to as a “golden-speaker” in the pronunciation-training literature. FAC is generally achieved by building a voice-conversion model that maps utterances from a source (L1) speaker onto the target (L2) speaker. As such, FAC requires that a reference utterance from the L1 speaker be available at synthesis time. This greatly restricts the application scope of the FAC system. In this work, we propose a “reference-free” FAC system that eliminates the need for reference L1 utterances at synthesis time, and transforms L2 utterances directly. The system is trained in two steps. First, a conventional FAC procedure is used to create a golden-speaker using utterances from a reference L1 speaker (which are then discarded) and the L2 speaker. Second, a pronunciation-correction model is trained to convert L2 utterances to match the golden-speaker utterances obtained in the first step. At synthesis time, the pronunciation-correction model directly transforms a novel L2 utterance into its golden-speaker counterpart. Our results show that the system reduces foreign accents in novel L2 utterances, achieving a 20.5% relative reduction in word-error-rate of an American English automatic speech recognizer and a 19% reduction in perceptual ratings of foreign accentedness obtained through listening tests. Over 73% of the listeners also rated golden-speaker utterances as having the same voice identity as the original L2 utterances. Guanlong Zhao, Shaojin Ding, Ricardo Gutierrez-Osuna |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Understanding the Effect of Voice Quality and Accent on Talker SimilarityabstractThis paper presents a methodology to study the role of nonnative accents on talker recognition by humans. The methodology combines a state-of-the-art accent-conversion system to resynthesize the voice of a speaker with a different accent of her/his own, and a protocol for perceptual listening tests to measure the relative contribution of accent and voice quality on speaker similarity. Using a corpus of non-native and native speakers, we generated accent conversions in two different directions: non-native speakers with native accents, and native speakers with non-native accents. Then, we asked listeners to rate the similarity between 50 pairs of real or synthesized speakers. Using a linear mixed effects model, we find that (for our corpus) the effect of voice quality is five times as large as that of non-native accent, and that the effect goes away when speakers share the same (native) accent. We discuss the potential significance of this work in earwitness identification and sociophonetics. Anurag Das, Guanlong Zhao, John Levis, Evgeny Chukharev-Hudilainen, Ricardo Gutierrez-Osuna |
INTERSPEECH | 2 |
| 2020 | Improving the Speaker Identity of Non-Parallel Many-to-Many Voice Conversion with Adversarial Speaker Recognition
Shaojin Ding, Guanlong Zhao, Ricardo Gutierrez-Osuna |
INTERSPEECH | 2 |
| 2020 | Learning Structured Sparse Representations for Voice ConversionabstractSparse-coding techniques for voice conversion assume that an utterance can be decomposed into a sparse code that only carries linguistic contents, and a dictionary of atoms that captures the speakers' characteristics. However, conventional dictionary-construction and sparse-coding algorithms rarely meet this assumption. The result is that the sparse code is no longer speaker-independent, which leads to lower voice-conversion performance. In this paper, we propose a Cluster-Structured Sparse Representation (CSSR) that improves the speaker independence of the representations. CSSR consists of two complementary components: a Cluster-Structured Dictionary Learning module that groups atoms in the dictionary into clusters, and a Cluster-Selective Objective Function that encourages each speech frame to be represented by atoms from a small number of clusters. We conducted four experiments on the CMU ARCTIC corpus to evaluate the proposed method. In a first ablation study, results show that each of the two CSSR components enhances speaker independence, and that combining both components leads to further improvements. In a second experiment, we find that CSSR uses increasingly larger dictionaries more efficiently than phoneme-based representations by allowing finer-grained decompositions of speech sounds. In a third experiment, results from objective and subjective measurements show that CSSR outperforms prior voice-conversion methods, improving the acoustic quality of the synthesized speech while retaining the target speaker's voice identity. Finally, we show that the CSSR captures latent (i.e., phonetic) information in the speech signal. Shaojin Ding, Guanlong Zhao, Christopher Liberatore, Ricardo Gutierrez-Osuna |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Foreign Accent Conversion by Synthesizing Speech from Phonetic Posteriorgrams
Guanlong Zhao, Shaojin Ding, Ricardo Gutierrez-Osuna |
INTERSPEECH | 1 |
| 2019 | Golden speaker builder - An interactive tool for pronunciation training
Shaojin Ding, Christopher Liberatore, Sinem Sonsaat, Ivana Lucic, Alif Silpachai, Guanlong Zhao, Evgeny Chukharev-Hudilainen, John Levis, Ricardo Gutierrez-Osuna |
Speech Commun. | 6 |
| 2019 | Using Phonetic Posteriorgram Based Frame Pairing for Segmental Accent ConversionabstractAccent conversion (AC) aims to transform non-native utterances to sound as if the speaker had a native accent. This can be achieved by mapping source speech spectra from a native speaker into the acoustic space of the target non-native speaker. In prior work, we proposed an AC approach that matches frames between the two speakers based on their acoustic similarity after compensating for differences in vocal tract length. In this paper, we propose a new approach that matches frames between the two speakers based on their phonetic (rather than acoustic) similarity. Namely, we map frames from the two speakers into a phonetic posteriorgram using speaker-independent acoustic models trained on native speech. We thoroughly evaluate the approach on a speech corpus containing multiple native and non-native speakers. The proposed algorithm outperforms the prior approach, improving ratings of acoustic quality (22% increase in mean opinion score) and native accent (69% preference) while retaining the voice quality of the non-native speaker. Furthermore, we show that the approach can be used in the reverse conversion direction, i.e., generating speech with a native speaker's voice quality and a non-native accent. Finally, we show that this approach can be applied to non-parallel training data, achieving the same accent conversion performance. Guanlong Zhao, Ricardo Gutierrez-Osuna |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Voice Conversion Through Residual Warping in a Sparse, Anchor-Based Representation of SpeechabstractIn previous work we presented a Sparse, Anchor-Based Representation of speech (SABR) that uses phonemic “anchors” to represent an utterance with a set of sparse non-negative weights. SABR is speaker-independent: combining weights from a source speaker with anchors from a target speaker can be used for voice conversion. Here, we present an extension of the original SABR that significantly improves voice conversion synthesis. Namely, we take the residual signal from the SABR decomposition of the source speaker's utterance, and warp it to the target speaker's space using a weighted warping function learned from pairs of source-target anchors. Using subjective and objective evaluations, we examine the performance of adding the warped residual (SABR+Res) to the original synthesis (SABR). Specifically, listeners rated SABR+Res with an average mean opinion score (MOS) of 3.6, a significant improvement compared to 2.2 MOS for SABR alone (p <; 0.01) and 2.5 MOS for a baseline GMM method (p <; 0.01). In an XAB speaker identity test, listeners correctly identified the identity of SABR+Res (81 %) and SABR (84%) as frequently as a GMM method (82%) (p = 0.70, P = 0.35). These results indicate that adding the warped residual can dramatically improve synthesis while retaining the desirable independent qualities of SABR models. Christopher Liberatore, Guanlong Zhao, Ricardo Gutierrez-Osuna |
ICASSP | 2 |
| 2018 | Accent Conversion Using Phonetic PosteriorgramsabstractAccent conversion (AC) aims to transform non-native speech to sound as if the speaker had a native accent. This can be achieved by mapping source spectra from a native speaker into the acoustic space of the non-native speaker. In prior work, we proposed an AC approach that matches frames between the two speakers based on their acoustic similarity after compensating for differences in vocal tract length. In this paper, we propose an approach that matches frames between the two speakers based on their phonetic (rather than acoustic) similarity. Namely, we map frames from the two speakers into a phonetic posteriorgram using speaker-independent acoustic models trained on native speech. We evaluate the proposed algorithm on a corpus containing multiple native and non-native speakers. Compared to the previous AC algorithm, the proposed algorithm improves the ratings of acoustic quality (20% increase in mean opinion score) and native accent (69% preference) while retaining the voice identity of the non-native speaker. Guanlong Zhao, Sinem Sonsaat, John Levis, Evgeny Chukharev-Hudilainen, Ricardo Gutierrez-Osuna |
ICASSP | 1 |
| 2018 | Improving Sparse Representations in Exemplar-Based Voice Conversion with a Phoneme-Selective Objective Function
Shaojin Ding, Guanlong Zhao, Christopher Liberatore, Ricardo Gutierrez-Osuna |
INTERSPEECH | 2 |
| 2018 | L2-ARCTIC: A Non-native English Speech CorpusabstractIn this paper, we introduce L2-ARCTIC, a speech corpus of non-native English that is intended for research in voice conversion, accent conversion, and mispronunciation detection. This initial release includes recordings from ten non-native speakers of English whose first languages (L1s) are Hindi, Korean, Mandarin, Spanish, and Arabic, each L1 containing recordings from one male and one female speaker. Each speaker recorded approximately one hour of read speech from the Carnegie Mellon University ARCTIC prompts, from which we generated orthographic and forced-aligned phonetic transcriptions. In addition, we manually annotated 150 utterances per speaker to identify three types of mispronunciation errors: substitutions, deletions, and additions, making it a valuable resource not only for research in voice conversion and accent conversion but also in computer-assisted pronunciation training. The corpus is publicly accessible at https://psi.engr.tamu.edu/l2-arctic-corpus/. Guanlong Zhao, Sinem Sonsaat, Alif Silpachai, Ivana Lucic, Evgeny Chukharev-Hudilainen, John Levis, Ricardo Gutierrez-Osuna |
INTERSPEECH | 1 |
| 2017 | Exemplar selection methods in voice conversionabstractExemplar-based methods for voice conversion often use a large number of randomly-selected exemplars to ensure good coverage. As a result, the factorization step can be costly. This paper presents two algorithms that can be used to construct compact sets of exemplars. The first algorithm uses a forward selection procedure to build the exemplar set sequentially, selecting exemplar pairs that minimize the joint reconstruction error on source and target frames. The second algorithm uses a backward elimination procedure to remove exemplars that contribute the least to the factorization. We evaluate both selection strategies on voice conversion tasks using the ARCTIC corpus. Our results using objective measures and subjective listening tests show that both strategies can significantly reduce the size of the exemplar set (five-fold, in our experiments) while achieving the same performance on voice conversion. Guanlong Zhao, Ricardo Gutierrez-Osuna |
ICASSP | 1 |