Ngoc-Quan Pham

dblp:151/8723 · also Quan Ngoc Pham · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
14since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 10 since 2021
YearPublicationVenuePosition
2025 Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS
abstract
Previous approaches on accent conversion (AC) mainly aimed at making non-native speech sound more native while maintaining the original content and speaker identity. However, non-native speakers sometimes have pronunciation issues, which can make it difficult for listeners to understand them. Hence, we developed a new AC approach that not only focuses on accent conversion but also improves pronunciation of non-native accented speaker. By providing the non-native audio and the corresponding transcript, we generate the ideal ground-truth audio with native-like pronunciation with original duration and prosody. This ground-truth data aids the model in learning a direct mapping between accented and native speech. We utilize the end-to-end VITS framework to achieve high-quality waveform reconstruction for the AC task. As a result, our system not only produces audio that closely resembles native accents and while retaining the original speaker’s identity but also improve pronunciation, as demonstrated by evaluation results.
Tuan Nam Nguyen, Seymanur Akti, Ngoc-Quan Pham, Alex Waibel
ICASSP3
2025 PIER: A Novel Metric for Evaluating What Matters in Code-Switching
abstract
Code-switching, the alternation of languages within a single discourse, presents a significant challenge for Automatic Speech Recognition. Despite the unique nature of the task, performance is commonly measured with established metrics such as Word-Error-Rate (WER). However, in this paper, we question whether these general metrics accurately assess performance on code-switching. Specifically, using both Connectionist-Temporal-Classification and Encoder-Decoder models, we show fine-tuning on non-code-switched data from both matrix and embedded language improves classical metrics on code-switching test sets, although actual code-switched words worsen (as expected). Therefore, we propose Point-of-Interest Error Rate (PIER), a variant of WER that focuses only on specific words of interest. We instantiate PIER on code-switched utterances and show that this more accurately describes the code-switching performance, showing huge room for improvement in future work. This focused evaluation allows for a more precise assessment of model performance, particularly in challenging aspects such as inter-word and intra-word code-switching.
Enes Yavuz Ugan, Ngoc-Quan Pham, Leonard Bärmann, Alex Waibel
ICASSP2
2025 Promoting Ensemble Diversity with Interactive Bayesian Distributional Robustness for Fine-tuning Foundation Models
abstract
We introduce Interactive Bayesian Distributional Robustness (IBDR), a novel Bayesian inference framework that allows modeling the interactions between particles, thereby enhancing ensemble quality through increased particle diversity. IBDR is grounded in a generalized theoretical framework that connects the distributional population loss with the approximate posterior, motivating a practical dual optimization procedure that enforces distributional robustness while fostering particle diversity. We evaluate IBDR's performance against various baseline methods using the VTAB-1K benchmark and the common reasoning language task. The results consistently show that IBDR outperforms these baselines, underscoring its effectiveness in real-world applications.
Ngoc-Quan Pham, Tuan Truong, Quyen Tran, Tan M. Nguyen, Dinh Q. Phung, Trung Le 0001
ICML1
2025 Improving Generalization with Flat Hilbert Bayesian Inference
abstract
We introduce Flat Hilbert Bayesian Inference (FHBI), an algorithm designed to enhance generalization in Bayesian inference. Our approach involves an iterative two-step procedure with an adversarial functional perturbation step and a functional descent step within the reproducing kernel Hilbert spaces. This methodology is supported by a theoretical analysis that extends previous findings on generalization ability from finite-dimensional Euclidean spaces to infinite-dimensional functional spaces. To evaluate the effectiveness of FHBI, we conduct comprehensive comparisons against nine baseline methods on the VTAB-1K benchmark, which encompasses 19 diverse datasets across various domains with diverse semantics. Empirical results demonstrate that FHBI consistently outperforms the baselines by notable margins, highlighting its practical efficacy.
Tuan Truong, Quyen Tran, Ngoc-Quan Pham, Nhat Ho, Dinh Q. Phung, Trung Le 0001
ICML3
2025 Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
Tuan-Nam Nguyen, Ngoc-Quan Pham, Seymanur Akti, Alex Waibel
INTERSPEECH2
2025 Cocktail-Party Audio-Visual Speech Recognition
Thai-Binh Nguyen, Ngoc-Quan Pham, Alex Waibel
INTERSPEECH2
2025 Weight Factorization and Centralization for Continual Learning in Speech Recognition
Enes Yavuz Ugan, Ngoc-Quan Pham, Alex Waibel
INTERSPEECH2
2024 Decoupled Vocabulary Learning Enables Zero-Shot Translation from Unseen Languages
abstract
Multilingual neural machine translation systems learn to map sentences of different languages into a common representation space.Intuitively, with a growing number of seen languages the encoder sentence representation grows more flexible and easily adaptable to new languages.In this work, we test this hypothesis by zero-shot translating from unseen languages.To deal with unknown vocabularies from unknown languages we propose a setup where we decouple learning of vocabulary and syntax, i.e. for each language we learn word representations in a separate step (using cross-lingual word embeddings), and then train to translate while keeping those word representations frozen.We demonstrate that this setup enables zero-shot translation from entirely unseen languages.Zero-shot translating with a model trained on Germanic and Romance languages we achieve scores of 42.6 BLEU for Portuguese-English and 20.7 BLEU for Russian-English on TED domain.We explore how this zero-shot translation capability develops with varying number of languages seen by the encoder.Lastly, we explore the effectiveness of our decoupled learning strategy for unsupervised machine translation.By exploiting our model's zero-shot translation capability for iterative back-translation we attain near parity with a supervised setting.
Carlos Mullov, Ngoc-Quan Pham, Alex Waibel
ACL (1)2
2024 DECM: Evaluating Bilingual ASR Performance on a Code-switching/mixing Benchmark
abstract
Automatic Speech Recognition has made significant progress, but challenges persist. Code-switched (CSW) Speech presents one such challenge, involving the mixing of multiple languages by a speaker. Even when multilingual ASR models are trained, each utterance on its own usually remains monolingual. We introduce an evaluation dataset for German-English CSW, with German as the matrix language and English as the embedded language. The dataset comprises spontaneous speech from diverse domains, enabling realistic CSW evaluation in German-English. It includes splits with varying degrees of CSW to facilitate specialized model analysis. As it is difficult to collect CSW data for all language pairs, the provision of such evaluation data, is crucial for developing and analyzing ASR models capable of generalizing across unseen pairs. Detailed data statistics are presented, and state-of-the-art (SOTA) multilingual models are evaluated showing challanges of CSW speech.
Enes Yavuz Ugan, Ngoc-Quan Pham, Alex Waibel
LREC/COLING2
2023 SYNTACC : Synthesizing Multi-Accent Speech By Weight Factorization
abstract
Conventional multi-speaker text-to-speech synthesis (TTS) is known to be capable of synthesizing speech for multiple voices, yet it cannot generate speech in different accents. This limitation has motivated us to develop SYNTACC (Synthesizing speech with accents) which adapts conventional multi-speaker TTS to produce multi-accent speech. Our method uses the YourTTS model and involves a novel multi-accent training mechanism. The method works by decomposing each weight matrix into a shared component and an accent-dependent component, with the former being initialized by the pretrained multi-speaker TTS model and the latter being factorized into vectors using rank-1 matrices to reduce the number of training parameters per accent. This weight factorization method proves to be effective in fine-tuning the SYNTACC on multi-accent data sets in a low-resource condition. Our SYNTACC model eventually allows speech synthesis in not only different voices but also in different accents.
Tuan-Nam Nguyen, Ngoc-Quan Pham, Alex Waibel
ICASSP2
2023 Towards continually learning new languages
Ngoc-Quan Pham, Jan Niehues, Alex Waibel
INTERSPEECH1
2022 Accent Conversion using Pre-trained Model and Synthesized Data from Voice Conversion
abstract
Accent conversion (AC) aims to generate synthetic audios by changing the pronunciation pattern and prosody of source speakers (in source audios) while preserving voice quality and linguistic content.There has not been a parallel corpus that contains pairs of audios having the same contents yet coming from the same speakers in different accents, the authors hence work on a solution to synthesize one as training input.The training pipeline is conducted via two steps.First, a voice conversion (VC) model is constructed to synthesize a training data set, containing pairs of audios in the same voice but two different accents.Second, an AC model is trained with the synthesized data to convert a source accented speech to a target accented speech.Given the recognized success of self-supervised learning speech representation (wav2vec 2.0) on certain speech problems such as VC, speech recognition, speech translation, and speech-tospeech translation, we adopt this architecture with some customization to train the AC model in the second step.With just 9-hour synthesized training data, the encoder initialized by the weight of the pre-trained wav2vec 2.0 model outperforms the LSTM-based encoder.
Tuan-Nam Nguyen, Ngoc-Quan Pham, Alex Waibel
INTERSPEECH2
2022 Adaptive multilingual speech recognition with pretrained models
abstract
Multilingual speech recognition with supervised learning has achieved great results as reflected in recent research.With the development of pretraining methods on audio and text data, it is imperative to transfer the knowledge from unsupervised multilingual models to facilitate recognition, especially in many languages with limited data.Our work investigated the effectiveness of using two pretrained models for two modalities: wav2vec 2.0 for audio and MBART50 for text, together with the adaptive weight techniques to massively improve the recognition quality on the public datasets containing CommonVoice and Europarl.Overall, we noticed an 44% improvement over purely supervised learning, and more importantly, each technique provides a different reinforcement in different languages.We also explore other possibilities to potentially obtain the best model by slightly adding either depth or relative attention to the architecture.
Ngoc-Quan Pham, Alex Waibel, Jan Niehues
INTERSPEECH1
2021 Efficient Weight Factorization for Multilingual Speech Recognition
abstract
End-to-end multilingual speech recognition involves using a single model training on a compositional speech corpus including many languages, resulting in a single neural network to handle transcribing different languages. Due to the fact that each language in the training data has different characteristics, the shared network may struggle to optimize for all various languages simultaneously. In this paper we propose a novel multilingual architecture that targets the core operation in neural networks: linear transformation functions. The key idea of the method is to assign fast weight matrices for each language by decomposing each weight matrix into a shared component and a language dependent component. The latter is then factorized into vectors using rank-1 assumptions to reduce the number of parameters per language. This efficient factorization scheme is proved to be effective in two multilingual settings with $7$ and $27$ languages, reducing the word error rates by $26\%$ and $27\%$ rel. for two popular architectures LSTM and Transformer, respectively.
Ngoc-Quan Pham, Tuan-Nam Nguyen, Sebastian Stüker, Alex Waibel
Interspeech1
2020 High Performance Sequence-to-Sequence Model for Streaming Speech Recognition
abstract
Recently sequence-to-sequence models have started to achieve state-of-the-art performance on standard speech recognition tasks when processing audio data in batch mode, i.e., the complete audio data is available when starting processing. However, when it comes to performing run-on recognition on an input stream of audio data while producing recognition results in real-time and with low word-based latency, these models face several challenges. For many techniques, the whole audio sequence to be decoded needs to be available at the start of the processing, e.g., for the attention mechanism or the bidirectional LSTM (BLSTM). In this paper, we propose several techniques to mitigate these problems. We introduce an additional loss function controlling the uncertainty of the attention mechanism, a modified beam search identifying partial, stable hypotheses, ways of working with BLSTM in the encoder, and the use of chunked BLSTM. Our experiments show that with the right combination of these techniques, it is possible to perform run-on speech recognition with low word-based latency without sacrificing in word error rate performance.
Thai Son Nguyen, Ngoc-Quan Pham, Sebastian Stüker, Alex Waibel
INTERSPEECH2
2020 Relative Positional Encoding for Speech Recognition and Direct Translation
abstract
Transformer models are powerful sequence-to-sequence architectures that are capable of directly mapping speech inputs to transcriptions or translations. However, the mechanism for modeling positions in this model was tailored for text modeling, and thus is less ideal for acoustic inputs. In this work, we adapt the relative position encoding scheme to the Speech Transformer, where the key addition is relative distance between input states in the self-attention network. As a result, the network can better adapt to the variable distributions present in speech data. Our experiments show that our resulting model achieves the best recognition result on the Switchboard benchmark in the non-augmentation condition, and the best published result in the MuST-C speech translation benchmark. We also show that this model is able to better utilize synthetic data than the Transformer, and adapts better to variable sentence segmentation quality for speech translation.
Ngoc-Quan Pham, Thanh-Le Ha, Tuan-Nam Nguyen, Thai Son Nguyen, Elizabeth Salesky, Sebastian Stüker, Jan Niehues, Alex Waibel
INTERSPEECH1
2019 Self-Attentional Models for Lattice Inputs
abstract
Lattices are an efficient and effective method to encode ambiguity of upstream systems in natural language processing tasks, for example to compactly capture multiple speech recognition hypotheses, or to represent multiple linguistic analyses.Previous work has extended recurrent neural networks to model lattice inputs and achieved improvements in various tasks, but these models suffer from very slow computation speeds.This paper extends the recently proposed paradigm of self-attention to handle lattice inputs.Self-attention is a sequence modeling technique that relates inputs to one another by computing pairwise similarities and has gained popularity for both its strong results and its computational efficiency.To extend such models to handle lattices, we introduce probabilistic reachability masks that incorporate lattice structure into the model and support lattice scores if available.We also propose a method for adapting positional embeddings to lattice structures.We apply the proposed model to a speech translation task and find that it outperforms all examined baselines while being much faster to compute than previous neural lattice models during both training and inference.
Matthias Sperber, Graham Neubig, Ngoc-Quan Pham, Alex Waibel
ACL (1)3
2019 Modeling Confidence in Sequence-to-Sequence Models
abstract
Recently, significant improvements have been achieved in various natural language processing tasks using neural sequence-to-sequence models.While aiming for the best generation quality is important, ultimately it is also necessary to develop models that can assess the quality of their output.In this work, we propose to use the similarity between training and test conditions as a measure for models' confidence.We investigate methods solely using the similarity as well as methods combining it with the posterior probability.While traditionally only target tokens are annotated with confidence measures, we also investigate methods to annotate source tokens with confidence.By learning an internal alignment model, we can significantly improve confidence projection over using stateof-the-art external alignment tools.We evaluate the proposed methods on downstream confidence estimation for machine translation (MT).We show improvements on segmentlevel confidence estimation as well as on confidence estimation for source tokens.In addition, we show that the same methods can also be applied to other tasks using sequence-tosequence models.On the automatic speech recognition (ASR) task, we are able to find 60% of the errors by looking at 20% of the data.
Jan Niehues, Ngoc-Quan Pham
INLG2
2019 Very Deep Self-Attention Networks for End-to-End Speech Recognition
abstract
Recently, end-to-end sequence-to-sequence models for speech recognition have gained significant interest in the research community. While previous architecture choices revolve around time-delay neural networks (TDNN) and long short-term memory (LSTM) recurrent neural networks, we propose to use self-attention via the Transformer architecture as an alternative. Our analysis shows that deep Transformer networks with high learning capacity are able to exceed performance from previous end-to-end approaches and even match the conventional hybrid systems. Moreover, we trained very deep models with up to 48 Transformer layers for both encoder and decoders combined with stochastic residual connections, which greatly improve generalizability and training efficiency. The resulting models outperform all previous end-to-end ASR approaches on the Switchboard benchmark. An ensemble of these models achieve 9.9% and 17.7% WER on Switchboard and CallHome test sets respectively. This finding brings our end-to-end models to competitive levels with previous hybrid systems. Further, with model ensembling the Transformers can outperform certain hybrid systems, which are more complicated in terms of both structure and training procedure.
Ngoc-Quan Pham, Thai Son Nguyen, Jan Niehues, Markus Müller 0001, Alex Waibel
INTERSPEECH1
2018 Low-Latency Neural Speech Translation
abstract
Through the development of neural machine translation, the quality of machine translation systems has been improved significantly.By exploiting advancements in deep learning, systems are now able to better approximate the complex mapping from source sentences to target sentences.But with this ability, new challenges also arise.An example is the translation of partial sentences in low-latency speech translation.Since the model has only seen complete sentences in training, it will always try to generate a complete sentence, though the input may only be a partial sentence.We show that NMT systems can be adapted to scenarios where no task-specific training data is available.Furthermore, this is possible without losing performance on the original training data.We achieve this by creating artificial data and by using multi-task learning.After adaptation, we are able to reduce the number of corrections displayed during incremental output construction by 45%, without a decrease in translation quality.
Jan Niehues, Ngoc-Quan Pham, Thanh-Le Ha, Matthias Sperber, Alex Waibel
INTERSPEECH2
2018 KIT-Multi: A Translation-Oriented Multilingual Embedding Corpus
Thanh-Le Ha, Jan Niehues, Matthias Sperber, Ngoc-Quan Pham, Alex Waibel
LREC4
2016 The LAMBADA dataset: Word prediction requiring a broad discourse context
abstract
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, Raquel Fernández. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016.
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, Raquel Fernández
ACL (1)4
2016 Convolutional Neural Network Language Models
abstract
Convolutional Neural Networks (CNNs) have shown to yield very strong results in several Computer Vision tasks. Their application to language has received much less attention, and it has mainly focused on static classification tasks, such as sentence classification for Sentiment Analysis or relation extraction. In this work, we study the application of CNNs to language modeling, a dynamic, sequential prediction task that needs models to capture local as well as long-range dependency information. Our contribution is twofold. First, we show that CNNs achieve 11-26% better absolute performance than feed-forward neural\nlanguage models, demonstrating their potential for language representation even in sequential tasks. As for recurrent models, our model outperforms RNNs but is below state of the art LSTM models. Second, we gain some understanding of the behavior of the model, showing that CNNs in language act as feature detectors at a high level of abstraction, like in Computer Vision, and that the model can profitably use information from as far as 16 words before the target.
Ngoc-Quan Pham, Germán Kruszewski, Gemma Boleda
EMNLP1