VLDB 2026 Research / reviewers in the wild / expert
Trung Hieu Nguyen 0001
dblp:44/3549-1
· DBLP profile ↗
26ranked-venue papers
2as first author
12since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Are Soft Prompts Good Zero-Shot Learners for Speech Recognition?abstractLarge self-supervised pre-trained speech models require computationally expensive fine-tuning for downstream tasks. Soft prompt tuning offers a simple parameter-efficient alternative by utilizing minimal soft prompt guidance, enhancing portability while also maintaining competitive performance. However, not many people understand how and why this is so. In this study, we aim to deepen our understanding of this emerging method by investigating the role of soft prompts in automatic speech recognition (ASR). Our findings highlight their role as zero-shot learners in improving ASR performance while also exposing them to the risk of malicious modifications. Soft prompts aid generalization but are not obligatory for inference. We also identify two primary roles of soft prompts: content refinement and noise information enhancement, which enhances robustness against background noise. Additionally, we propose an effective modification on noise prompts to show that they are capable of zero-shot learning on adapting to out-of-distribution noise environments. Dianwen Ng, Chong Zhang 0003, Ruixi Zhang, Fabian Ritter Gutierrez, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Chng Eng Siong, Bin Ma 0001 |
ICASSP | 6 |
| 2024 | SPGM: Prioritizing Local Features for Enhanced Speech Separation PerformanceabstractDual-path is a popular architecture for speech separation models (e.g. Sepformer) which splits long sequences into overlapping chunks for its intra- and inter-blocks that separately model intra-chunk local features and inter-chunk global relationships. However, it has been found that inter-blocks, which comprise half a dual-path model’s parameters, contribute minimally to performance. Thus, we propose the Single-Path Global Modulation (SPGM) block to replace inter-blocks. SPGM is named after its structure consisting of a parameter-free global pooling module followed by a modulation module comprising only 2% of the model’s total parameters. The SPGM block allows all transformer layers in the model to be dedicated to local feature modelling, making the overall model single-path. SPGM achieves 22.1 dB SI-SDRi on WSJ0-2Mix and 20.4 dB SI-SDRi on Libri2Mix, exceeding the performance of Sepformer by 0.5 dB and 0.3 dB respectively and matches the performance of recent SOTA models with up to 8 times fewer parameters. Model and weights are available at huggingface.co/yipjiaqi/spgm Jia Qi Yip, Shengkui Zhao, Chongjia Ni, Chong Zhang 0003, Hao Wang 0199, Trung Hieu Nguyen 0001, Kun Zhou 0003, Dianwen Ng, Chng Eng Siong, Bin Ma 0001 |
ICASSP | 7 |
| 2024 | MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech SeparationabstractOur previously proposed MossFormer has achieved promising performance in monaural speech separation. However, it predominantly adopts a self-attention-based MossFormer module, which tends to emphasize longer-range, coarser-scale dependencies, with a deficiency in effectively modelling finer-scale recurrent patterns. In this paper, we introduce a novel hybrid model that provides the capabilities to model both long-range, coarse-scale dependencies and fine-scale recurrent patterns by integrating a recurrent module into the MossFormer framework. Instead of applying the recurrent neural networks (RNNs) that use traditional recurrent connections, we present a recurrent module based on a feedforward sequential memory network (FSMN), which is considered "RNN-free" recurrent network due to the ability to capture recurrent patterns without using recurrent connections. Our recurrent module mainly comprises an enhanced dilated FSMN block by using gated convolutional units (GCU) and dense connections. In addition, a bottleneck layer and an output layer are also added for controlling information flow. The recurrent module relies on linear projections and convolutions for seamless, parallel processing of the entire sequence. The integrated MossFormer2 hybrid model demonstrates remarkable enhancements over MossFormer and surpasses other state-of-the-art methods in WSJ0-2/3mix, Libri2Mix, and WHAM!/WHAMR! benchmarks. Shengkui Zhao, Chongjia Ni, Chong Zhang 0003, Hao Wang 0199, Trung Hieu Nguyen 0001, Kun Zhou 0003, Jia Qi Yip, Dianwen Ng, Bin Ma 0001 |
ICASSP | 6 |
| 2024 | Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
Kun Zhou 0003, Shengkui Zhao, Chong Zhang 0003, Hao Wang 0199, Dianwen Ng, Chongjia Ni, Trung Hieu Nguyen 0001, Jia Qi Yip, Bin Ma 0001 |
INTERSPEECH | 8 |
| 2023 | Auxiliary Pooling Layer For Spoken Language UnderstandingabstractEnd-to-end spoken language understanding requires speech data annotated with semantic information and may suffer from the shortage of annotated data. Recent progresses leverage unlabelled speech data to pre-train a speech encoder. However, it remains a challenge for the pre-trained speech encoder to encode semantic information. Existing works explore transferring knowledge from a pre-trained text model with different alignment losses at a fixed granularity. In this paper, we address the variable granularity in transferring knowledge from texts to speech representation via APLY, an auxiliary pooling layer, that fuses the global information with the adaptively encoded local context. We demonstrate the effectiveness of APLY on three benchmarks of spoken language understanding. Trung Hieu Nguyen 0001, Jinjie Ni, Wen Wang 0001, Qian Chen 0003, Chong Zhang 0003, Bin Ma 0001 |
ICASSP | 2 |
| 2023 | Contrastive Speech Mixup for Low-Resource Keyword SpottingabstractMost of the existing neural-based models for keyword spotting (KWS) in smart devices require thousands of training samples to learn a decent audio representation. However, with the rising demand for smart devices to become more person-alized, KWS models need to adapt quickly to smaller user samples. To tackle this challenge, we propose a contrastive speech mixup (CosMix) learning algorithm for low-resource KWS. CosMix introduces an auxiliary contrastive loss to the existing mixup augmentation technique to maximize the relative similarity between the original pre-mixed samples and the augmented samples. The goal is to inject enhancing constraints to guide the model towards simpler but richer content-based speech representations from two augmented views (i.e. noisy mixed and clean pre-mixed utterances). We conduct our experiments on the Google Speech Command dataset, where we trim the size of the training set to as small as 2.5 mins per keyword to simulate a low-resource condition. Our experimental results show a consistent improvement in the performance of multiple models, which exhibits the effectiveness of our method. Dianwen Ng, Ruixi Zhang, Jia Qi Yip, Chong Zhang 0003, Trung Hieu Nguyen 0001, Chongjia Ni, Chng Eng Siong, Bin Ma 0001 |
ICASSP | 6 |
| 2023 | Adaptive Knowledge Distillation Between Text and Speech Pre-Trained ModelsabstractLearning on a massive amount of speech corpus leads to the recent success of many self-supervised speech models. With knowledge distillation, these models may also benefit from the knowledge encoded by language models that are pre-trained on rich sources of texts. The distillation process, however, is challenging due to the modal disparity between textual and speech embedding spaces. This paper studies metric-based distillation to align the embedding space of text and speech with only a small amount of data without modifying the model structure. Since the semantic and granularity gap between text and speech has been omitted in literature, which impairs the distillation, we propose the Prior-informed Adaptive knowledge Distillation (PAD) that adaptively leverages text/speech units of variable granularity and prior distributions to achieve better global and local alignments between text and speech pre-trained models. We evaluate on three spoken language understanding benchmarks to show that PAD is more effective in transferring linguistic knowledge than other metric-based distillation approaches. Jinjie Ni, Wen Wang 0001, Qian Chen 0033, Dianwen Ng, Han Lei, Trung Hieu Nguyen 0001, Chong Zhang 0003, Bin Ma 0001, Erik Cambria |
ICASSP | 7 |
| 2023 | Adapter-tuning with Effective Token-dependent Representation Shift for Automatic Speech Recognition
Dianwen Ng, Chong Zhang 0003, Ruixi Zhang, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Qian Chen 0003, Wen Wang 0001, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 5 |
| 2023 | ACA-Net: Towards Lightweight Speaker Verification using Asymmetric Cross Attention
Jia Qi Yip, Duc-Tuan Truong, Dianwen Ng, Chong Zhang 0003, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 6 |
| 2022 | CPT: Cross-Modal Prefix-Tuning for Speech-To-Text TranslationabstractSpeech translation models benefit from adapting multilingual pretrained language models. However, such adaptation modifies the parameters in the pretrained model to favor a specific task. Prefix-tuning, as a lightweight adaptation technique, has recently emerged as an efficient adaptation method that significantly reduces the number of trainable parameters and has demonstrated great potential in low-resource settings. It inserts prefixes into the output of each layer of a pretrained model, without modifying its parameters. During training, only the parameters of prefixes are updated while the rest of the model are being frozen. In this paper, we improve the performance of speech translation in medium-/low-resource settings by a cross-modal prefix that bridges the gap between speech input and translation modules to reduce the information loss in the cascaded model. We show that the proposed cross-modal prefix-tuning is effective, robust and parameter-efficient for adapting a speech recognition and translation pipeline. Trung Hieu Nguyen 0001, Bin Ma 0001 |
ICASSP | 2 |
| 2021 | Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency LossesabstractDeep complex U-Net structure and convolutional recurrent network (CRN) structure achieve state-of-the-art performance for monaural speech enhancement. Both deep complex U-Net and CRN are encoder and decoder structures with skip connections, which heavily rely on the representation power of the complex-valued convolutional layers. In this paper, we propose a complex convolutional block attention module (CCBAM) to boost the representation power of the complex-valued convolutional layers by constructing more informative features. The CCBAM is a lightweight and general module which can be easily integrated into any complex-valued convolutional layers. We integrate CCBAM with the deep complex U-Net and CRN to enhance their performance for speech enhancement. We further propose a mixed loss function to jointly optimize the complex models in both time-frequency (TF) domain and time domain. By integrating CCBAM and the mixed loss, we form a new end-to-end (E2E) complex speech enhancement framework. Ablation experiments and objective evaluations show the superior performance of the proposed approaches. Shengkui Zhao, Trung Hieu Nguyen 0001, Bin Ma 0001 |
ICASSP | 2 |
| 2021 | Towards Natural and Controllable Cross-Lingual Voice Conversion Based on Neural TTS Model and Phonetic PosteriorgramabstractCross-lingual voice conversion (VC) is an important and challenging problem due to significant mismatches of the phonetic set and the speech prosody of different languages. In this paper, we build upon the neural text-to-speech (TTS) model, i.e., FastSpeech, and LPCNet neural vocoder to design a new cross-lingual VC framework named FastSpeech-VC. We address the mismatches of the phonetic set and the speech prosody by applying Phonetic PosteriorGrams (PPGs), which have been proved to bridge across speaker and language boundaries. Moreover, we add normalized logarithm-scale fundamental frequency (Log-F0) to further compensate for the prosodic mismatches and significantly improve naturalness. Our experiments on English and Mandarin languages demonstrate that with only mono-lingual corpus, the proposed FastSpeech-VC can achieve high quality converted speech with mean opinion score (MOS) close to the professional records while maintaining good speaker similarity. Compared to the baselines using Tacotron2 and Transformer TTS models, the FastSpeech-VC can achieve controllable converted speech rate and much faster inference speed. More importantly, the FastSpeech-VC can easily be adapted to a speaker with limited training utterances. Shengkui Zhao, Hao Wang 0199, Trung Hieu Nguyen 0001, Bin Ma 0001 |
ICASSP | 3 |
| 2020 | Towards Natural Bilingual and Code-Switched Speech Synthesis Based on Mix of Monolingual Recordings and Cross-Lingual Voice ConversionabstractRecent state-of-the-art neural text-to-speech (TTS) synthesis models have dramatically improved intelligibility and naturalness of generated speech from text. However, building a good bilingual or code-switched TTS for a particular voice is still a challenge. The main reason is that it is not easy to obtain a bilingual corpus from a speaker who achieves native-level fluency in both languages. In this paper, we explore the use of Mandarin speech recordings from a Mandarin speaker, and English speech recordings from another English speaker to build high-quality bilingual and code-switched TTS for both speakers. A Tacotron2-based cross-lingual voice conversion system is employed to generate the Mandarin speaker's English speech and the English speaker's Mandarin speech, which show good naturalness and speaker similarity. The obtained bilingual data are then augmented with code-switched utterances synthesized using a Transformer model. With these data, three neural TTS models -- Tacotron2, Transformer and FastSpeech are applied for building bilingual and code-switched TTS. Subjective evaluation results show that all the three systems can produce (near-)native-level speech in both languages for each of the speaker. Shengkui Zhao, Trung Hieu Nguyen 0001, Hao Wang 0199, Bin Ma 0001 |
INTERSPEECH | 2 |
| 2019 | Fast Learning for Non-Parallel Many-to-Many Voice Conversion with Residual Star Generative Adversarial Networks
Shengkui Zhao, Trung Hieu Nguyen 0001, Hao Wang 0199, Bin Ma 0001 |
INTERSPEECH | 2 |
| 2017 | The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016abstract18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017 Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah |
INTERSPEECH | 20 |
| 2016 | The 2015 NIST Language Recognition Evaluation: The Shared View of I2R, Fantastic4 and SingaMSabstractTechnical report for NIST LRE 2015 Workshop Kong-Aik Lee, Haizhou Li 0001, Li Deng 0001, Ville Hautamäki, Wei Rao 0002, Anthony Larcher, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Aleksandr Sizov, Jianshu Chen, Ivan Kukanov, Amir Hossein Poorjam, Trung Ngo Trong, Chenglin Xu, Haihua Xu 0001, Bin Ma 0001, Chng Eng Siong, Sylvain Meignier |
INTERSPEECH | 9 |
| 2016 | Joint Speaker and Lexical Modeling for Short-Term Characterization of Speaker
Guangsen Wang, Kong-Aik Lee, Trung Hieu Nguyen 0001, Hanwu Sun, Bin Ma 0001 |
INTERSPEECH | 3 |
| 2015 | The reddots platform for mobile crowd-sourcing of speech data
Kong-Aik Lee, Guangsen Wang, Kam Pheng Ng, Hanwu Sun, Trung Hieu Nguyen 0001, Ngoc Thuy Huong Thai, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2014 | Extended RSR2015 for text-dependent speaker verification over VHF channelabstractInternational audience Anthony Larcher, Kong-Aik Lee, Pablo Luis Sordo Martinez, Trung Hieu Nguyen 0001, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2014 | On the use of Bhattacharyya based GMM distance and neural net features for identification of cognitive load levelsabstractThis paper presents a method for detecting cognitive load levels from speech. When speech is modulated by different lev-els of cognitive load, acoustic characteristics of speech change. In this paper, we measure acoustic distance of a stressed ut-terance from the baseline stress free speech using GMM-SVM kernel with Bhattacharyya based GMM distance. In addition, it is believed that airflow structure of speech production is non-linear. This motivates us to investigate better techniques to cap-ture nonlinear characteristic of stress information in acoustic features. Inspired by the recent success of neural networks for representation learning, we employ a single hidden layer feed forward network with non-linear activation to extract the fea-ture vectors. Furthermore, people have different reactions to a particular task load. This inter-speaker difference in stress re-sponses presents a major challenge for stress level detection. We use a bootstrapped training process to learn the stress re-sponse of a particular speaker. We perform experiments using data sets from Cognitive Load with Speech and EGG (CLSE) provided for the Cognitive Load Sub-Challenge of the INTER-SPEECH 2014 Computational Paralinguistics Challenge. The results show that the system with our proposed strategies per-forms well on validation and test sets. Index Terms: cognitive load, GMM-supervector, neural net features Tin Lay Nwe, Trung Hieu Nguyen 0001, Bin Ma 0001 |
INTERSPEECH | 2 |
| 2013 | Bhattacharyya distance based emotional dissimilarity measure for emotion classificationabstractSpeech is one of the most important signals that can be used to detect human emotions. When speech is modulated by different emotions, spectral distribution of speech is changed accordingly. A Gaussian Mixture Model(GMM) can model the changes in spectral distributions effectively. A GMM-supervector characterizes the spectral distribution of an emotion utterance by the GMM parameters such as the mean vectors and covariance matrices. In this paper, we propose to use the GMM-supervectors that characterize the emotional spectral dissimilarity measure for emotion classification. We employ the GMM-SVM kernel with Bhattacharyya based GMM distance to obtain dissimilarity measure. Beside the first-order statistics of mean, we consider dissimilarity measure using second-order statistics of covariance which describe the shape of the distribution. Experiments are conducted using SVM classifier to classify emotions of anger, happiness, neutral and sadness. We achieve average accuracy of 78.14% for speaker independent emotion classification. Tin Lay Nwe, Trung Hieu Nguyen 0001, Dilip Kumar Limbu |
ICASSP | 2 |
| 2013 | Bhattacharyya distance based emotional dissimilarity measure in multi-dimensional space for emotion classification
Tin Lay Nwe, Trung Hieu Nguyen 0001, Dilip Kumar Limbu |
INTERSPEECH | 2 |
| 2010 | The IIR NIST SRE 2008 and 2010 summed channel speaker recognition systemsabstractThis paper reports the IIR speaker recognition system for the summed channel evaluation tasks in the NIST SRE 2008 and 2010. The system includes three main modules: voice activity detection, speaker diarization and speaker recognition. The front-end process employs a voice activity detection algorithm for effective speech frame selection. The speaker diarization system that was developed for 2007 and 2009 NIST RT Evaluations is adopted for summed channel speech segmentation. A hybrid purifying and clustering algorithm is developed to segregate the summed channel speech by speakers. The GMM-SVM speaker recognition system is adopted to evaluate the performance with both MFCC and LPCC features. The system achieves an overall EER of 3.46% in the 1conv-summed task and 1.87% in the 8conv-summed task, respectively, where only all English trials are involved. Hanwu Sun, Bin Ma 0001, Chien-Lin Huang, Trung Hieu Nguyen 0001, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2009 | Cluster criterion functions in spectral subspace and their application in speaker clusteringabstractIn this paper, we propose two cluster criterion functions which aim to maximize the separation between intra-cluster distances and inter-cluster distances. These criteria can automatically deduce the desired number of clusters based on their extremized values. We then propose an algorithm to apply our criterion functions in conjunction with spectral clustering. By exploiting the characteristic of spectral subspace, we show that the speakers are more separable in this subspace which will further enhance the effectiveness of our proposed criteria. The algorithm is used in our agglomerative hierarchical speaker diarization system to test on Rich Transcription 2007 conference data set and obtains very good results. Trung Hieu Nguyen 0001, Haizhou Li 0001, Chng Eng Siong |
ICASSP | 1 |
| 2008 | T-test distance and clustering criterion for speaker diarizationabstractIn this paper, we present an application of student’s t-test to measure the similarity between two speaker models. The mea-sure is evaluated by comparing with other distance metrics: the Generalized Likelihood Ratio, the Cross Likelihood Ratio and the Normalized Cross Likelihood Ratio in speaker detec-tion task. We also propose an objective criterion for speaker clustering. The criterion deduces the number of speakers auto-matically by maximizing the separation between intra-speaker distances and inter-speaker distances. It requires no develop-ment data and works well with various distance metrics. We then report the performance of our proposed similarity distance measure and objective criterion in speaker diarization task. The system produces competitive results: low speaker diarization error rate and high accuracy in detecting number of speakers. Index Terms: speaker diarization, speaker detection, intra-speaker, inter-speaker. Trung Hieu Nguyen 0001, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2007 | Using direction of arrival estimate and acoustic feature information in speaker diarization
Chin-Wei Eugene Koh, Hanwu Sun, Tin Lay Nwe, Trung Hieu Nguyen 0001, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001, Susanto Rahardja |
INTERSPEECH | 4 |