EDBT 2026 Demo / reviewers in the wild / expert
Haihua Xu 0001
dblp:26/1181
· DBLP profile ↗
46ranked-venue papers
12as first author
10since 2021 · last 2026
0000-0002-2220-8465ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 10 first-author · 9 since 2021Artificial intelligence and machine learning · 24 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LCMA-SRT: Language-Conditional Mixture-of-Experts Adapters for Joint Multilingual Speech Recognition and TranslationabstractNeural transducers offer an alignment-free framework for speech-to-text modeling, and hierarchical transducer architectures further improve multilingual joint automatic speech recognition (ASR) and speech translation (ST) by stacking a translation-focused encoder on top of an ASR encoder.However, extending hierarchical transducers to multilingual manyto-many settings remains challenging: fully shared models often suffer from negative transfer and unstable target-language generation, while training separate models for each direction is computationally prohibitive.We propose LCMA-SRT (Language-Conditional Mixtureof-Experts Adapters for Speech Recognition and Translation), which augments a hierarchical transducer with language-conditional Mixture-of-Experts (MoE) adapters.A sourceconditioned MoE adapter (SRC-MoE) uses source-language embeddings to reduce crosslanguage interference and improve multilingual ASR.A target-conditioned MoE adapter (TGT-MoE) uses the desired target language to reduce cross-target interference and stabilize targetlanguage generation in many-to-many ST.Experiments on Europarl-ST (9 languages, 72 directions) show that LCMA-SRT improves both ASR and ST within a single joint model, reducing average WER and improving BLEU and COMET over strong hierarchical transducer baselines.We release our code and models at https://github.com/linanjie0820/ LCMA-SRT. Nanjie Li, Xiaoyong Guo, Hao Huang 0009, Haihua Xu 0001 |
ACL (1) | 4 |
| 2023 | Reducing Language Confusion for Code-Switching Speech Recognition with Token-Level Language DiarizationabstractCode-switching (CS) occurs when languages switch within a speech signal and leads to language confusion for automatic speech recognition (ASR). We address the problem of language confusion for improving CS-ASR from two perspectives: incorporating and disentangling language information. We incorporate language information within the CS-ASR model by dynamically biasing the model with token-level language posteriors corresponding to outputs of a sequence-to-sequence auxiliary language diarization (LD) module. In contrast, the disentangling process reduces the difference between languages via adversarial training so as to normalize two languages. We conduct experiments on the SEAME dataset. Compared to the baseline model, both the joint optimization with LD and the language posterior bias achieve performance improvement. Comparison of the proposed methods indicates that incorporating language information is more effective than disentangling for reducing language confusion in CS speech. Hexin Liu, Haihua Xu 0001, L. Paola García-Perera, Andy W. H. Khong, Sanjeev Khudanpur |
ICASSP | 2 |
| 2023 | Internal Language Model Estimation Based Adaptive Language Model Fusion for Domain AdaptationabstractASR model deployment environment is ever-changing, and the incoming speech can be switched across different domains during a session. This brings a challenge for effective domain adaptation when only target domain text data is available, and our objective is to obtain obviously improved performance on the target domain while the performance on the general domain is less undermined. In this paper, we propose an adaptive LM fusion approach called internal language model estimation based adaptive domain adaptation (ILME-ADA). To realize such an ILME-ADA, an interpolated log-likelihood score is calculated based on the maximum of the scores from the internal LM and the external LM (ELM) respectively. We demonstrate the efficacy of the proposed ILME-ADA method with both RNN-T and LAS modeling frameworks employing neural network and n-gram LMs as ELMs respectively on two domain specific (target) test sets. The proposed method can achieve significantly better performance on the target test sets while it gets minimal performance degradation on the general test set, compared with both shallow and ILME-based LM fusion methods. Rao Ma, Jin Qiu, Yanan Qin, Haihua Xu 0001, Peihao Wu, Zejun Ma 0001 |
ICASSP | 5 |
| 2023 | Speech-Text Based Multi-Modal Training with Bidirectional Attention for Improved Speech RecognitionabstractTo let the state-of-the-art end-to-end ASR model enjoy data efficiency, as well as much more unpaired text data by multi-modal training, one needs to address two problems: 1) the synchronicity of feature sampling rates between speech and language (aka text data); 2) the homogeneity of the learned representations from two encoders. In this paper we propose to employ a novel bidirectional attention mechanism (BiAM) to jointly learn both ASR encoder (bottom layers) and text encoder with a multi-modal learning method. The BiAM is to facilitate feature sampling rate exchange, realizing the quality of the transformed features for the one kind to be measured in another space, with diversified objective functions. As a result, the speech representations are enriched with more linguistic information, while the representations generated by the text encoder are more similar to corresponding speech ones, and therefore the shared ASR models are more amenable for unpaired text data pretraining. To validate the efficacy of the proposed method, we perform two categories of experiments with or without extra unpaired text data. Experimental results on Librispeech corpus show it can achieve up to 6.15% word error rate reduction (WERR) with only paired data learning, while 9.23% WERR when more unpaired text data is employed1. Haihua Xu 0001, Hao Huang 0009, Chng Eng Siong, Sheng Li 0010 |
ICASSP | 2 |
| 2023 | Knowledge Distillation Approach for Efficient Internal Language Model Estimation
Haihua Xu 0001, Yerbolat Khassanov, Lu Lu 0015, Zejun Ma 0001, Ji Wu 0002 |
INTERSPEECH | 2 |
| 2023 | Self-supervised Learning Representation based Accent Recognition with Persistent Accent Memory
Zhiwei Xie 0007, Haihua Xu 0001, Yizhou Peng, Hexin Liu, Hao Huang 0009, Chng Eng Siong |
INTERSPEECH | 3 |
| 2023 | Random Utterance Concatenation Based Data Augmentation for Improving Short-video Speech RecognitionabstractOne of limitations in end-to-end automatic speech recognition (ASR) framework is its performance would be compromised if train-test utterance lengths are mismatched.In this paper, we propose an on-the-fly random utterance concatenation (RUC) based data augmentation method to alleviate train-test utterance length mismatch issue for short-video ASR task.Specifically, we are motivated by observations that our human-transcribed training utterances tend to be much shorter for short-video spontaneous speech (∼3 seconds on average), while our test utterance generated from voice activity detection front-end is much longer (∼10 seconds on average).Such a mismatch can lead to suboptimal performance.Empirically, it's observed the proposed RUC method significantly improves long utterance recognition without performance drop on short one.Overall, it achieves 5.72% word error rate reduction on average for 15 languages and improved robustness to various utterance length. Yist Y. Lin, Haihua Xu 0001, Van Tung Pham, Yerbolat Khassanov, Tze Yuang Chong, Lu Lu 0015, Zejun Ma 0001 |
INTERSPEECH | 3 |
| 2022 | Minimum Word Error Training For Non-Autoregressive Transformer-Based Code-Switching ASRabstractNon-autoregressive end-to-end ASR framework might be potentially appropriate for code-switching recognition task thanks to its inherent property that present output token being independent of historical ones. However, it still under-performs the state-of-the-art autoregressive ASR frameworks. In this paper, we propose various approaches to boosting the performance of a CTC-mask-based non-autoregressive Transformer under code-switching ASR scenario. To begin with, we attempt diversified masking method that are closely related with code-switching point, yielding an improved baseline model. More importantly, we employ Minimum Word Error (MWE) criterion to train the model. One of the challenges is how to generate a diversified hypothetical space, so as to obtain the average loss for a given ground truth. To address such a challenge, we explore different approaches to yielding desired N-best-based hypothetical space. We demonstrate the efficacy of the proposed methods on SEAME corpus, a challenging English-Mandarin code-switching corpus for Southeast Asia community. Compared with the cross-entropy-trained strong baseline, the proposed MWE training method achieves consistent performance improvement on the test sets. Yizhou Peng, Haihua Xu 0001, Hao Huang 0009, Chng Eng Siong |
ICASSP | 3 |
| 2022 | Internal Language Model Estimation Through Explicit Context Vector Learning for Attention-based Encoder-decoder ASR
Rao Ma, Haihua Xu 0001, Zejun Ma 0001 |
INTERSPEECH | 3 |
| 2021 | E2E-Based Multi-Task Learning Approach to Joint Speech and Accent RecognitionabstractIn this paper, we propose a single multi-task learning framework to perform End-to-End (E2E) speech recognition (ASR) and accent recognition (AR) simultaneously.The proposed framework is not only more compact but can also yield comparable or even better results than standalone systems.Specifically, we found that the overall performance is predominantly determined by the ASR task, and the E2E-based ASR pretraining is essential to achieve improved performance, particularly for the AR task.Additionally, we conduct several analyses of the proposed method.First, though the objective loss for the AR task is much smaller compared with its counterpart of ASR task, a smaller weighting factor with the AR task in the joint objective function is necessary to yield better results for each task.Second, we found that sharing only a few layers of the encoder yields better AR results than sharing the overall encoder.Experimentally, the proposed method produces WER results close to the best standalone E2E ASR ones, while it achieves 7.7% and 4.2% relative improvement over standalone and single-task-based joint recognition methods on test set for accent recognition respectively. Yizhou Peng, Van Tung Pham, Haihua Xu 0001, Hao Huang 0009, Chng Eng Siong |
Interspeech | 4 |
| 2020 | Independent Language Modeling Architecture for End-To-End ASRabstractThe attention-based end-to-end (E2E) automatic speech recognition (ASR) architecture allows for joint optimization of acoustic and language models within a single network. However, in a vanilla E2E ASR architecture, the decoder sub-network (subnet), which incorporates the role of the language model (LM), is conditioned on the encoder output. This means that the acoustic encoder and the language model are entangled that doesn’t allow language model to be trained separately from external text data. To address this problem, in this work, we propose a new architecture that separates the decoder subnet from the encoder output. In this way, the decoupled subnet becomes an independently trainable LM subnet, which can easily be updated using the external text data. We study two strategies for updating the new architecture. Experimental results show that, 1) the independent LM architecture benefits from external text data, achieving 9.3% and 22.8% relative character and word error rate reduction on Mandarin HKUST and English NSC datasets respectively; 2) the proposed architecture works well with external LM and can be generalized to different amount of labelled data. Van Tung Pham, Haihua Xu 0001, Yerbolat Khassanov, Zhiping Zeng, Chng Eng Siong, Chongjia Ni, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 2 |
| 2020 | Monolingual Data Selection Analysis for English-Mandarin Hybrid Code-Switching Speech RecognitionabstractIn this paper, we conduct data selection analysis in building an English-Mandarin code-switching (CS) speech recognition (CSSR) system, which is aimed for a real CSSR contest in China.The overall training sets have three subsets, i.e., a codeswitching data set, an English (LibriSpeech) and a Mandarin data set respectively.The code-switching data are Mandarin dominated.First of all, it is found using the overall data yields worse results, and hence data selection study is necessary.Then to exploit monolingual data, we find data matching is crucial.Mandarin data is closely matched with the Mandarin part in the code-switching data, while English data is not.However, Mandarin data only helps on those utterances that are significantly Mandarin-dominated.Besides, there is a balance point, over which more monolingual data will divert the CSSR system, degrading results.Finally, we analyze the effectiveness of combining monolingual data to train a CSSR system with the HMM-DNN hybrid framework.The CSSR system can perform within-utterance code-switch recognition, but it still has a margin with the one trained on code-switching data. Haihua Xu 0001, Van Tung Pham, Hao Huang 0009, Chng Eng Siong |
INTERSPEECH | 2 |
| 2019 | Cross-lingual Voice Conversion with Bilingual Phonetic Posteriorgram and Average ModelingabstractThis paper presents a cross-lingual voice conversion approach using bilingual Phonetic PosteriorGram (PPG) and average modeling. The proposed approach makes use of bilingual PPGs to represent speaker-independent features of speech signals from different languages in the same feature space. In particular, a bilingual PPG is formed by stacking two monolingual PPG vectors, which are extracted from two monolingual speech recognition systems. The conversion model is trained to learn the relationship between bilingual PPGs and the corresponding acoustic features. To leverage the linguistic and acoustic information from other speakers in different languages, an average model is trained with multiple speakers in both source and target languages. I-vector is utilized as an additional input feature of the average model for network adaptation. Experiments are performed for intralingual and cross-lingual voice conversion between English and Mandarin speakers. Both objective and subjective evaluations demonstrate the effectiveness of our proposed approach. Yi Zhou 0020, Xiaohai Tian, Haihua Xu 0001, Rohan Kumar Das, Haizhou Li 0001 |
ICASSP | 3 |
| 2019 | Constrained Output Embeddings for End-to-End Code-Switching Speech Recognition with Only Monolingual DataabstractThe lack of code-switch training data is one of the major concerns in the development of end-to-end code-switching automatic speech recognition (ASR) models. In this work, we propose a method to train an improved end-to-end code-switching ASR using only monolingual data. Our method encourages the distributions of output token embeddings of monolingual languages to be similar, and hence, promotes the ASR model to easily code-switch between languages. Specifically, we propose to use Jensen-Shannon divergence and cosine distance based constraints. The former will enforce output embeddings of monolingual languages to possess similar distributions, while the later simply brings the centroids of two distributions to be close to each other. Experimental results demonstrate high effectiveness of the proposed method, yielding up to 4.5% absolute mixed error rate improvement on Mandarin-English code-switching ASR task. Yerbolat Khassanov, Haihua Xu 0001, Van Tung Pham, Zhiping Zeng, Chng Eng Siong, Chongjia Ni, Bin Ma 0001 |
INTERSPEECH | 2 |
| 2019 | Enriching Rare Word Representations in Neural Language Models by Embedding Matrix AugmentationabstractThe neural language models (NLM) achieve strong generalization capability by learning the dense representation of words and using them to estimate probability distribution function. However, learning the representation of rare words is a challenging problem causing the NLM to produce unreliable probability estimates. To address this problem, we propose a method to enrich representations of rare words in pre-trained NLM and consequently improve its probability estimation performance. The proposed method augments the word embedding matrices of pre-trained NLM while keeping other parameters unchanged. Specifically, our method updates the embedding vectors of rare words using embedding vectors of other semantically and syntactically similar words. To evaluate the proposed method, we enrich the rare street names in the pre-trained NLM and use it to rescore 100-best hypotheses output from the Singapore English speech recognition system. The enriched NLM reduces the word error rate by 6% relative and improves the recognition accuracy of the rare words by 16% absolute as compared to the baseline NLM. Yerbolat Khassanov, Zhiping Zeng, Van Tung Pham, Haihua Xu 0001, Chng Eng Siong |
INTERSPEECH | 4 |
| 2019 | On the End-to-End Solution to Mandarin-English Code-Switching Speech RecognitionabstractCode-switching (CS) refers to a linguistic phenomenon where a speaker uses different languages in an utterance or between alternating utterances.In this work, we study end-to-end (E2E) approaches to the Mandarin-English code-switching speech recognition task.We first examine the effectiveness of using data augmentation and byte-pair encoding (BPE) subword units.More importantly, we propose a multitask learning recipe, where a language identification task is explicitly learned in addition to the E2E speech recognition task.Furthermore, we introduce an efficient word vocabulary expansion method for language modeling to alleviate data sparsity issues under the code-switching scenario.Experimental results on the SEAME data, a Mandarin-English code-switching corpus, demonstrate the effectiveness of the proposed methods. Zhiping Zeng, Yerbolat Khassanov, Van Tung Pham, Haihua Xu 0001, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2018 | Study of Semi-supervised Approaches to Improving English-Mandarin Code-Switching Speech RecognitionabstractIn this paper, we present our overall efforts to improve the performance of a code-switching speech recognition system using semi-supervised training methods from lexicon learning to acoustic modeling, on the South East Asian Mandarin-English (SEAME) data.We first investigate semi-supervised lexicon learning approach to adapt the canonical lexicon, which is meant to alleviate the heavily accented pronunciation issue within the code-switching conversation of the local area.As a result, the learned lexicon yields improved performance.Furthermore, we attempt to use semi-supervised training to deal with those transcriptions that are highly mismatched between human transcribers and ASR system.Specifically, we conduct semi-supervised training assuming those poorly transcribed data as unsupervised data.We found the semi-supervised acoustic modeling can lead to improved results.Finally, to make up for the limitation of the conventional n-gram language models due to data sparsity issue, we perform lattice rescoring using neural network language models, and significant WER reduction is obtained. Haihua Xu 0001, Lei Xie 0001, Chng Eng Siong |
INTERSPEECH | 2 |
| 2018 | Mandarin-English Code-switching Speech Recognition
Haihua Xu 0001, Van Tung Pham, Kyaw Zin Tun, Zhi Hao Lim, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2018 | Re-ranking spoken term detection with acoustic exemplars of keywords
Van Tung Pham, Haihua Xu 0001, Nancy F. Chen, Chng Eng Siong, Haizhou Li 0001 |
Speech Commun. | 2 |
| 2017 | The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016abstract18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017 Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah |
INTERSPEECH | 48 |
| 2016 | Exemplar-inspired strategies for low-resource spoken keyword search in SwahiliabstractWe present exemplar-inspired low-resource spoken keyword search strategies for acoustic modeling, keyword verification, and system combination. This state-of-the-art system was developed by the SINGA team in the context of the 2015 NIST Open Keyword Search Evaluation (OpenKWS15) using conversational Swahili provided by the IARPA Babel program. In this work, we elaborate on the following: (1) exploiting exemplar training samples to construct a non-parametric acoustic model using kernel density estimation at test time; (2) rescoring hypothesized keyword detections through quantifying their acoustic similarity with exemplar training samples; (3 ) extending our previously proposed system combination approach to incorporate prosody features of exemplar keyword samples. Nancy F. Chen, Van Tung Pham, Haihua Xu 0001, Van Hai Do, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Chin-Hui Lee 0001, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 3 |
| 2016 | Keyword search using query expansion for graph-based rescoring of hypothesized detectionsabstractIn this work, we propose a novel framework for rescoring keyword search (KWS) detections using acoustic samples extracted from the training data. We view the keyword rescoring task as an information retrieval task and adopt the idea of query expansion. We expand a textual keyword with multiple speech keyword samples extracted from the training data. In this way, the hypothesized detections are compared with the multiple keywords using non-parametric approaches such as dynamic time warping (DTW). The obtained similarity scores are used in a graph based method to re-rank the original confidence scores estimated by the automatic speech recognition (ASR) systems. Experimental results on the NIST OpenKWS15 Evaluation show that our rescoring method is effective, especially for the subword system. For subword experiments, the graph-based rescoring with training samples obtains 5.1% and 1.5% absolute improvement over two baseline systems. One is a standard parametric ASR system, while the other is the graph-based rescoring without training samples. Van Tung Pham, Haihua Xu 0001, Nancy F. Chen, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 2 |
| 2016 | Approximate search of audio queries by using DTW with phone time boundary and data augmentationabstractDynamic Time Warping (DTW) is widely used in language independent query-by-example (QbE) spoken term detection (STD) tasks due to its high performance. However, there are two limitations of DTW based template matching, 1) it is not straightforward to perform approximate match of audio queries; 2) DTW is sensitive to the mismatch of signal conditions between the query and the speech search data. To allow approximate search, we propose a partial template matching strategy using phone time boundary information generated by a phone recognizer. To have more invariant representation of audio signals, we use bottleneck features (BNF) as the input of DTW. The BNF network is trained from augmented data, which is generated by adding reverberation and additive noises to the clean training data. Experimental results on QUESST 2015 task shows the effectiveness of the proposed methods for QbE-STD when the queries and search data are both distorted by reverberation and noises. Haihua Xu 0001, Jingyong Hou, Van Tung Pham, Cheung-Chi Leung, Lei Wang 0020, Van Hai Do, Hang Lv 0001, Lei Xie 0001, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 1 |
| 2016 | The 2015 NIST Language Recognition Evaluation: The Shared View of I2R, Fantastic4 and SingaMSabstractTechnical report for NIST LRE 2015 Workshop Kong-Aik Lee, Haizhou Li 0001, Li Deng 0001, Ville Hautamäki, Wei Rao 0002, Anthony Larcher, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Aleksandr Sizov, Jianshu Chen, Ivan Kukanov, Amir Hossein Poorjam, Trung Ngo Trong, Chenglin Xu, Haihua Xu 0001, Bin Ma 0001, Chng Eng Siong, Sylvain Meignier |
INTERSPEECH | 17 |
| 2016 | Toward High-Performance Language-Independent Query-by-Example Spoken Term Detection for MediaEval 2015: Post-Evaluation Analysis
Cheung-Chi Leung, Lei Wang 0020, Haihua Xu 0001, Jingyong Hou, Van Tung Pham, Hang Lv 0001, Lei Xie 0001, Chongjia Ni, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2016 | Rescoring Hypothesized Detections of Out-of-Vocabulary Keywords Using Subword Samples
Van Tung Pham, Haihua Xu 0001, Nancy F. Chen, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2016 | Semi-Supervised and Cross-Lingual Knowledge Transfer Learnings for DNN Hybrid Acoustic Models Under Low-Resource Conditions
Haihua Xu 0001, Chongjia Ni, Hao Huang 0009, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2015 | Low-resource keyword search strategies for tamilabstractWe propose strategies for a state-of-the-art keyword search (KWS) system developed by the SINGA team in the context of the 2014 NIST Open Keyword Search Evaluation (OpenKWS14) using conversational Tamil provided by the IARPA Babel program. To tackle low-resource challenges and the rich morphological nature of Tamil, we present highlights of our current KWS system, including: (1) Submodular optimization data selection to maximize acoustic diversity through Gaussian component indexed N-grams; (2) Keywordaware language modeling; (3) Subword modeling of morphemes and homophones. Nancy F. Chen, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Van Tung Pham, Haihua Xu 0001, Tze Siong Lau, Su Jun Leow, Boon Pang Lim, Cheung-Chi Leung, Lei Wang 0020, Chin-Hui Lee 0001, Alvina Goh, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 6 |
| 2015 | Language independent query-by-example spoken term detection using N-best phone sequences and partial matchingabstractIn this paper, we propose a partial sequence matching based symbolic search (SS) method for the task of language independent query-by-example spoken term detection. One main drawback of conventional SS approach is the high miss rate for long queries. This is due to high variations in symbol representation of query and search audios, especially in language independent scenario. The successful matching of a query with its instances in search audio becomes exponentially more difficult as the query grows longer. To reduce miss rate, we propose a partial matching strategy, in which all partial phone sequences of a query are used to search for query instances. The partial matching is also suitable for real life applications where exact match is usually not necessary and word prefix, suffix, and order should not affect the search result. When applied to the QUESST 2014 task, results show the partial matching of phone sequences is able to reduce miss rate of long queries significantly compared with conventional full matching method. In addition, for the most challenging inexact matching queries (type 3), it also shows clear advantage over DTW-based methods. Haihua Xu 0001, Lei Xie 0001, Cheung-Chi Leung, Hongjie Chen 0001, Jia Yu 0002, Hang Lv 0001, Lei Wang 0020, Su Jun Leow, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 1 |
| 2015 | Multi-softmax deep neural network for semi-supervised trainingabstractIn this paper we propose a Shared Hidden Layer Multi-softmax Deep Neural Network (SHL-MDNN) approach for semi-supervised training (SST). This approach aims to boost low-resource speech recognition where limited training data is available. Supervised data and unsupervised data share the same hidden layers but are fed into different softmax layers so that erroneous automatic speech recognition (ASR) transcrip-tions of the unsupervised data have less effect on shared hid-den layers. Experimental results on Babel data indicate that this approach always outperform naive SST on DNN, and it can yield 1.3 % word error rate (WER) reduction compared with su-pervised DNN hybrid system. In addition, if softmax layer is retrained with supervised data, it can lead up to another 0.8% WER reduction. Confidence based data selection is also studied in this setup. Experiments show that this method is not sensitive to ASR transcription errors. Haihua Xu 0001 |
INTERSPEECH | 2 |
| 2015 | Spoofing speech detection using high dimensional magnitude and phase features: the NTU approach for ASVspoof 2015 challenge
Xiaohai Tian, Steven Du, Haihua Xu 0001, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2015 | A comparative study of BNF and DNN multilingual training on cross-lingual low-resource speech recognition
Haihua Xu 0001, Van Hai Do, Chng Eng Siong |
INTERSPEECH | 1 |
| 2015 | Maximum F1-Score Discriminative Training Criterion for Automatic Mispronunciation DetectionabstractWe carry out an in-depth investigation on a newly proposed Maximum F1-score Criterion (MFC) discriminative training objective function for Goodness of Pronunciation (GOP) based automatic mispronunciation detection that makes use of Gaussian Mixture Model-hidden Markov model (GMM-HMM) as acoustic models. The formulation of MFC seeks to directly optimize F1-score by converting the non-differentiable F1-score function into a continuous objective function to facilitate optimization. We present model-space training algorithm according to MFC using extended Baum–Welch form like update equations based on the weak-sense auxiliary function method. We then present MFC based feature-space discriminative training. We train a matrix projecting from posteriors of Gaussians to a normal size feature space, and add the projected features to traditional spectral features. Mispronunciation detection experiments show MFC based model-space training and feature-space training are effective in improving F1-score and other commonly used evaluation metrics. It is also shown MFC training in both the feature-space and model-space outperforms either model-space training or feature-space training alone, and is about 11.6% better than the maximum likelihood (ML) trained baseline in terms of F1-score. Further, we review and compare mispronunciation detection results with the use of MFC and some traditional training criteria that minimize word error rate in speech recognition. The experimental analysis and comparison provide useful insight into the correlations between F1-score maximization and optimization of these training criteria. Hao Huang 0009, Haihua Xu 0001, Wushour Slamu |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Strategies for Vietnamese keyword searchabstractWe propose strategies for a state-of-the-art Vietnamese keyword search (KWS) system developed at the Institute for Infocomm Research (I2R). The KWS system exploits acoustic features characterizing creaky voice quality peculiar to lexical tones in Vietnamese, a minimal-resource transliteration framework to alleviate out-of-vocabulary issues from foreign loan words, and a proposed system combination scheme FusionX. We show that the proposed creaky voice quality features complement pitch-related features, reaching fusion gains of 17.7% relative (6.9% absolute). To the best of our knowledge, the proposed transliteration framework is the first reported rule-based system for Vietnamese; it outperforms statistical-approach baselines up to 14.93–36.73% relative on foreign loan word search tasks. Using FusionX to combine 3 sub-systems, the actual term-weighted value (ATWV) reaches 0.4742, exceeding the ATWV=0.3 benchmark for IARPA Babel participants in the NIST OpenKWSB Evaluation. Nancy F. Chen, Sunil Sivadas, Boon Pang Lim, Hoang Gia Ngo, Haihua Xu 0001, Van Tung Pham, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 5 |
| 2014 | Discriminative score normalization for keyword search decisionabstractMany keyword search (KWS) systems make “hit/false alarm (FA)” decisions based on the lattice-based posterior probability, which is incomparable across keywords. Therefore, score normalization is essential for a KWS system. In this paper, we investigate the integration of two novel features, ranking-score and relative-to-max, into a discriminative score normalization method. These features are extracted by considering all competing hypotheses of a putative detection. A metric-based normalization method is also applied as a post-processing step to further optimize the term-weighted value (TWV) evaluation metric. We report empirical improvements over standard baselines using the Vietnamese data from IARPA's Babel program in the NIST OpenKWS13 Evaluation setup. Van Tung Pham, Haihua Xu 0001, Nancy F. Chen, Sunil Sivadas, Boon Pang Lim, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 2 |
| 2014 | Semi-supervised training for bottle-neck feature based DNN-HMM hybrid systemsabstractIn this paper, we investigate semi-supervised training (SST) method in various state-of-the-art acoustic modeling tech-niques, using bottle-neck and corresponding tandem features. These techniques include subspace GMM, tanh-neuron deep neural network (DNN), and a generalized soft-maxout (p-norm) DNN. We demonstrate that SST may lead up to 2 % Word Error Rate (WER) reduction using all these techniques in each case, and the best one comes from tandem feature based p-norm DNN system. In addition to recognition performance, effectiveness of the SST on keyword search performance is also investigated. Results on Actual Term Weighted Value (ATWV) are reported, with an analysis on lattice density. It is shown that SST may not necessarily increase ATWV due to the shrink of lattices size. Haihua Xu 0001, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2014 | System and keyword dependent fusion for spoken term detectionabstractSystem combination (or data fusion1) is known to provide significant improvement for spoken term detection (STD). The key issue of the system combination is how to effectively fuse the various scores of participant systems. Currently, most system combination methods are system and keyword independent, i.e. they use the same arithmetic functions to combine scores for all keywords. Although such strategy improve keyword search performance, the improvement is limited. In this paper we first propose an arithmetic-based system combination method to incorporate the system and keyword characteristics into the fusion procedure to enhance the effectiveness of system combination. The method incorporates a system-keyword dependent property, which is the number of acceptances in this paper, into the combination procedure. We then introduce a discriminative model to combine various useful system and keyword characteristics into a general framework. Improvements over standard baselines are observed on the Vietnamese data from IARPA Babel program with the NIST OpenKWS13 Evaluation setup. Van Tung Pham, Nancy F. Chen, Sunil Sivadas, Haihua Xu 0001, I-Fan Chen, Chongjia Ni, Chng Eng Siong, Haizhou Li 0001 |
SLT | 4 |
| 2011 | Minimum Bayes Risk decoding and system combination based on a recursion for edit distance
Haihua Xu 0001, Daniel Povey, Lidia Mangu, Jie Zhu 0006 |
Comput. Speech Lang. | 1 |
| 2011 | Aniterative approach to Bayes risk decoding and system combinationabstractWe describe a novel approach to Bayes risk (BR) decoding for speech recognition, in which we attempt to find the hypothesis that minimizes an estimate of the BR with regard to the minimum word error (MWE) metric. To achieve this, we propose improved forward and backward algorithms on the lattices and the whole procedure is optimized recursively. The remarkable characteristics of the proposed approach are that the optimization procedure is expectation-maximization (EM) like and the formation of the updated result is similar to that obtained with the confusion network (CN) decoding method. Experimental results indicated that the proposed method leads to an error reduction for both lattice rescoring and lattice-based system combinations, compared with CN decoding, confusion network combination (CNC), and ROVER methods. Haihua Xu 0001, Jie Zhu 0006 |
J. Zhejiang Univ. Sci. C | 1 |
| 2010 | An improved consensus-like method for Minimum Bayes Risk decoding and lattice combinationabstractIn this paper we describe a method for Minimum Bayes Risk decoding for speech recognition. This is a technique similar to Consensus a.k.a. Confusion Network Decoding, in which we attempt to find the hypothesis that minimizes the Bayes' Risk with respect to the word error rate, based on a lattice of alternative outputs. Our method is an E-M like technique which makes approximations which we believe are less severe than the approximations made in Consensus, and our experimental results show an improvement in WER both for lattice rescoring and lattice-based system combination, versus baselines such as Consensus, Confusion Network Combination and ROVER. Haihua Xu 0001, Daniel Povey, Lidia Mangu, Jie Zhu 0006 |
ICASSP | 1 |
| 2009 | A hybrid visual feature extraction method for audio-visual speech recognitionabstractIn this paper, a hybrid visual feature extraction method that combines the extended locally linear embedding (LLE) with visemic linear discriminant analysis (LDA) was presented for the audio-visual speech recognition (AVSR). Firstly the extended LLE is presented to reduce the dimension of the mouth images, which constrains the scope of finding mouth data neighborhood to the corresponding individual's dataset instead of the whole dataset, and then maps the high dimensional mouth image matrices into a low-dimensional Euclidean space. Secondly we project the feature vectors on the visemic linear discriminant space to find the optimal classification. Finally, in the audio-visual fusion period, the minimum classification error (MCE) training based on the segmental generalized probabilistic descent (GPD) is applied to audio and visual stream weights optimization. Experimental results conducted the CUAVE database show that the proposed method achieves a significant performance than that of the classical PCA and LDA based method in visual-only speech recognition. Further experimental results show the robustness of the MCE based discriminative training method in noisy environment. Guanyong Wu, Jie Zhu 0006, Haihua Xu 0001 |
ICIP | 3 |
| 2009 | Minimum phone error based stream weight training for mandarin audio-visual Speech recognitionabstractStream weight training is one of the key issues in the bimodal integration for the audio-visual speech recognition. In this paper, the audio- and video-only HMM classifiers are combined to recognize audio-visual speech recognition. More specifically, a discriminative training method is provided, in which the state-dependent stream weights are trained based on lattice rescoring by the minimum phone error using the extended Baum Welch algorithm. The proposed method is evaluated on our Mandarin large vocabulary audio-visual database. Experimental results show the proposed method has achieved significant error reduction than traditional global stream weight based approach and outperforms the minimum classification error based discriminative stream weight training method. Guanyong Wu, Jie Zhu 0006, Haihua Xu 0001 |
ICME | 3 |
| 2009 | An efficient multistage Rover method for Automatic Speech recognitionabstractIn this paper, we implemented a multistage recognizer output voting error reduction (ROVER) method for better automatic speech recognition (ASR). The first stage ROVER is conducted by combining three recognizers, which are respectively trained with maximum likelihood estimation (MLE), minimum phone error (MPE) and recently proposed boosted maximum mutual information (BMMI) criteria. After that the second stage ROVER is performed on two groups of recognizers, which are separately adapted with maximum a posteriori (MAP) and maximum likelihood linear regression (MLLR) methods based on results from the first stage ROVER. It is found MAP adapted recognizers based ROVER does not lead to word error rate (WER) reduction while MLLR adapted recognizers based ROVER is still effective on recognition accuracy improvement. Finally, the third stage ROVER is constructed by combining adapted and unadapted recognizers, and the best WER is obtained by combining MLLR adapted recognizers with BMMI trained recognizers that are not adapted, which achieved 6.0% and 3.0% relative WER reduction, as opposed to the best result from the one-best decoding method and the single stage ROVER method accordingly. Haihua Xu 0001, Jie Zhu 0006, Guanyong Wu |
ICME | 1 |
| 2009 | Minimum hypothesis phone error as a decoding method for speech recognitionabstractIn this paper we show how methods for approximating phone error as normally used for Minimum Phone Error (MPE) dis-criminative training, can be used instead as a decoding criterion for lattice rescoring. This is an alternative to Confusion Net-works (CN) which are commonly used in speech recognition. The standard (Maximum A Posteriori) decoding approach is a Minimum Bayes Risk estimate with respect to the Sentence Er-ror Rate (SER); however, we are typically more interested in the Word Error Rate (WER). Methods such as CN and our pro-posed Minimum Hypothesis Phone Error (MHPE) aim to get closer to minimizing the expected WER. Based on preliminary experiments we find that our approach gives more improvement than CN, and is conceptually simpler. Haihua Xu 0001, Daniel Povey, Jie Zhu 0006, Guanyong Wu |
INTERSPEECH | 1 |
| 2009 | Minimum tag error for discriminative training of conditional random fields
Jie Zhu 0006, Hao Huang 0009, Haihua Xu 0001 |
Inf. Sci. | 4 |
| 2008 | Towards more efficient and accurate methods for Mandarin LVCSR discriminative trainingabstractDiscriminative training of Mandarin large vocabulary continuous speech recognition (LVCSR) has been remarkably improved in speech community recent years. However, much work still needs further investigating. In this work, we focus on improvements to two aspects of discriminative training method, in particular related to minimum phone error (MPE) training method in Mandarin speech recognition. One is to use syllable, not multi-character word, as speech recognition unit (SRU) to generate phone lattice to train models. The other is to investigate better objective functions related to MPE with comparisons on recent proposed methods. Experimental results showed that the proposed methods improved both efficiency and accuracy for discriminative training in Mandarin speech recognition. Haihua Xu 0001, Jie Zhu 0006 |
ICME | 1 |