EDBT 2026 Demo / reviewers in the wild / expert
Chongjia Ni
dblp:48/9230 · also Chong-Jia Ni
· DBLP profile ↗
41ranked-venue papers
11as first author
15since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 24 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A-SMiLE: Affective Sparse Mixture-of-Experts Adapter with Multi-Task Learning for Spoken Dialogue Models
Yi-Wen Chao, Yizhou Peng, Dianwen Ng, Chongjia Ni, Bin Ma 0001, Chng Eng Siong |
INTERSPEECH | 5 |
| 2025 | FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
Yizhou Peng, Yi-Wen Chao, Dianwen Ng, Chongjia Ni, Bin Ma 0001, Chng Eng Siong |
INTERSPEECH | 5 |
| 2024 | Are Soft Prompts Good Zero-Shot Learners for Speech Recognition?abstractLarge self-supervised pre-trained speech models require computationally expensive fine-tuning for downstream tasks. Soft prompt tuning offers a simple parameter-efficient alternative by utilizing minimal soft prompt guidance, enhancing portability while also maintaining competitive performance. However, not many people understand how and why this is so. In this study, we aim to deepen our understanding of this emerging method by investigating the role of soft prompts in automatic speech recognition (ASR). Our findings highlight their role as zero-shot learners in improving ASR performance while also exposing them to the risk of malicious modifications. Soft prompts aid generalization but are not obligatory for inference. We also identify two primary roles of soft prompts: content refinement and noise information enhancement, which enhances robustness against background noise. Additionally, we propose an effective modification on noise prompts to show that they are capable of zero-shot learning on adapting to out-of-distribution noise environments. Dianwen Ng, Chong Zhang 0003, Ruixi Zhang, Fabian Ritter Gutierrez, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Chng Eng Siong, Bin Ma 0001 |
ICASSP | 7 |
| 2024 | SPGM: Prioritizing Local Features for Enhanced Speech Separation PerformanceabstractDual-path is a popular architecture for speech separation models (e.g. Sepformer) which splits long sequences into overlapping chunks for its intra- and inter-blocks that separately model intra-chunk local features and inter-chunk global relationships. However, it has been found that inter-blocks, which comprise half a dual-path model’s parameters, contribute minimally to performance. Thus, we propose the Single-Path Global Modulation (SPGM) block to replace inter-blocks. SPGM is named after its structure consisting of a parameter-free global pooling module followed by a modulation module comprising only 2% of the model’s total parameters. The SPGM block allows all transformer layers in the model to be dedicated to local feature modelling, making the overall model single-path. SPGM achieves 22.1 dB SI-SDRi on WSJ0-2Mix and 20.4 dB SI-SDRi on Libri2Mix, exceeding the performance of Sepformer by 0.5 dB and 0.3 dB respectively and matches the performance of recent SOTA models with up to 8 times fewer parameters. Model and weights are available at huggingface.co/yipjiaqi/spgm Jia Qi Yip, Shengkui Zhao, Chongjia Ni, Chong Zhang 0003, Hao Wang 0199, Trung Hieu Nguyen 0001, Kun Zhou 0003, Dianwen Ng, Chng Eng Siong, Bin Ma 0001 |
ICASSP | 4 |
| 2024 | MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech SeparationabstractOur previously proposed MossFormer has achieved promising performance in monaural speech separation. However, it predominantly adopts a self-attention-based MossFormer module, which tends to emphasize longer-range, coarser-scale dependencies, with a deficiency in effectively modelling finer-scale recurrent patterns. In this paper, we introduce a novel hybrid model that provides the capabilities to model both long-range, coarse-scale dependencies and fine-scale recurrent patterns by integrating a recurrent module into the MossFormer framework. Instead of applying the recurrent neural networks (RNNs) that use traditional recurrent connections, we present a recurrent module based on a feedforward sequential memory network (FSMN), which is considered "RNN-free" recurrent network due to the ability to capture recurrent patterns without using recurrent connections. Our recurrent module mainly comprises an enhanced dilated FSMN block by using gated convolutional units (GCU) and dense connections. In addition, a bottleneck layer and an output layer are also added for controlling information flow. The recurrent module relies on linear projections and convolutions for seamless, parallel processing of the entire sequence. The integrated MossFormer2 hybrid model demonstrates remarkable enhancements over MossFormer and surpasses other state-of-the-art methods in WSJ0-2/3mix, Libri2Mix, and WHAM!/WHAMR! benchmarks. Shengkui Zhao, Chongjia Ni, Chong Zhang 0003, Hao Wang 0199, Trung Hieu Nguyen 0001, Kun Zhou 0003, Jia Qi Yip, Dianwen Ng, Bin Ma 0001 |
ICASSP | 3 |
| 2024 | Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
Kun Zhou 0003, Shengkui Zhao, Chong Zhang 0003, Hao Wang 0199, Dianwen Ng, Chongjia Ni, Trung Hieu Nguyen 0001, Jia Qi Yip, Bin Ma 0001 |
INTERSPEECH | 7 |
| 2023 | De'hubert: Disentangling Noise in a Self-Supervised Model for Robust Speech RecognitionabstractExisting self-supervised pre-trained speech models have offered an effective way to leverage massive unannotated corpora to build good automatic speech recognition (ASR). However, many current models are trained on a clean corpus from a single source, which tends to do poorly when noise is present during testing. Nonetheless, it is crucial to overcome the adverse influence of noise for real-world applications. In this work, we propose a novel training framework, called deHuBERT, for noise reduction encoding inspired by H. Barlow’s redundancy-reduction principle. The new framework improves the HuBERT training algorithm by introducing auxiliary losses that drive the self- and cross-correlation matrix between pairwise noise-distorted embeddings towards identity matrix. This encourages the model to produce noise- agnostic speech representations. With this method, we report improved robustness in noisy environments, including unseen noises, without impairing the performance on the clean set. Dianwen Ng, Ruixi Zhang, Jia Qi Yip, Jinjie Ni, Chong Zhang 0003, Chongjia Ni, Chng Eng Siong, Bin Ma 0001 |
ICASSP | 8 |
| 2023 | Contrastive Speech Mixup for Low-Resource Keyword SpottingabstractMost of the existing neural-based models for keyword spotting (KWS) in smart devices require thousands of training samples to learn a decent audio representation. However, with the rising demand for smart devices to become more person-alized, KWS models need to adapt quickly to smaller user samples. To tackle this challenge, we propose a contrastive speech mixup (CosMix) learning algorithm for low-resource KWS. CosMix introduces an auxiliary contrastive loss to the existing mixup augmentation technique to maximize the relative similarity between the original pre-mixed samples and the augmented samples. The goal is to inject enhancing constraints to guide the model towards simpler but richer content-based speech representations from two augmented views (i.e. noisy mixed and clean pre-mixed utterances). We conduct our experiments on the Google Speech Command dataset, where we trim the size of the training set to as small as 2.5 mins per keyword to simulate a low-resource condition. Our experimental results show a consistent improvement in the performance of multiple models, which exhibits the effectiveness of our method. Dianwen Ng, Ruixi Zhang, Jia Qi Yip, Chong Zhang 0003, Trung Hieu Nguyen 0001, Chongjia Ni, Chng Eng Siong, Bin Ma 0001 |
ICASSP | 7 |
| 2023 | Adapter-tuning with Effective Token-dependent Representation Shift for Automatic Speech Recognition
Dianwen Ng, Chong Zhang 0003, Ruixi Zhang, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Qian Chen 0003, Wen Wang 0001, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 6 |
| 2023 | Dual Acoustic Linguistic Self-supervised Representation Learning for Cross-Domain Speech Recognition
Dianwen Ng, Chong Zhang 0003, Xiao Fu 0001, Wei Xi 0003, Chongjia Ni, Chng Eng Siong, Bin Ma 0001, Jizhong Zhao |
INTERSPEECH | 8 |
| 2023 | A Unified Recognition and Correction Model under Noisy and Accent Speech Conditions
Dianwen Ng, Chong Zhang 0003, Wei Xi 0003, Chongjia Ni, Jizhong Zhao, Bin Ma 0001, Chng Eng Siong |
INTERSPEECH | 7 |
| 2023 | Dual-Memory Multi-Modal Learning for Continual Spoken Keyword Spotting with Confidence Selection and Diversity Enhancement
Dianwen Ng, Xizhe Li, Chong Zhang 0003, Wei Xi 0003, Chongjia Ni, Jizhong Zhao, Bin Ma 0001, Chng Eng Siong |
INTERSPEECH | 8 |
| 2023 | ACA-Net: Towards Lightweight Speaker Verification using Asymmetric Cross Attention
Jia Qi Yip, Duc-Tuan Truong, Dianwen Ng, Chong Zhang 0003, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 7 |
| 2021 | A Unified Speaker Adaptation Approach for ASRabstractTransformer models have been used in automatic speech recognition (ASR) successfully and yields state-of-the-art results. However, its performance is still affected by speaker mismatch between training and test data. Further finetuning a trained model with target speaker data is the most natural approach for adaptation, but it takes a lot of compute and may cause catastrophic forgetting to the existing speakers. In this work, we propose a unified speaker adaptation approach consisting of feature adaptation and model adaptation. For feature adaptation, we employ a speaker-aware persistent memory model which generalizes better to unseen test speakers by making use of speaker i-vectors to form a persistent memory. For model adaptation, we use a novel gradual pruning method to adapt to target speakers without changing the model architecture, which to the best of our knowledge, has never been explored in ASR. Specifically, we gradually prune less contributing parameters on model encoder to a certain sparsity level, and use the pruned parameters for adaptation, while freezing the unpruned parameters to keep the original model performance. We conduct experiments on the Librispeech dataset. Our proposed approach brings relative 2.74-6.52% word error rate (WER) reduction on general speaker adaptation. On target speaker adaptation, our method outperforms the baseline with up to 20.58% relative WER reduction, and surpasses the finetuning method by up to relative 2.54%. Besides, with extremely low-resource adaptation data (e.g., 1 utterance), our method could improve the WER by relative 6.53% with only a few epochs of training. Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001 |
EMNLP (1) | 2 |
| 2021 | Preventing Early Endpointing for Online Automatic Speech RecognitionabstractWith the recent development of end-to-end models in speech recognition, there have been more interests in adapting these models for online speech recognition. However, using end-to-end models for online speech recognition is known to suffer from an early endpointing problem, which brings in many deletion errors. In this paper, we propose to address the early endpointing problem from the gradient perspective. Specifically, we leverage on the recently proposed ScaleGrad technique, which was proposed to mitigate the text degeneration issue. Different from ScaleGrad, we adapt it to discourage the early generation of the end-of-sentence () token. A scaling term is added to directly maneuver the gradient of the training loss to encourage the model to learn to keep generating non-tokens. Compared with previous approaches such as voice-activity-detection and end-of-query detection, the proposed method does not rely on various types of silence, and it also saves the trouble from obtaining the ground truth endpoint with forced alignment. Nevertheless, it can be jointly applied with other techniques. Experiments on AISHELL-1 dataset show that our model brings relative 5.4%-10.1% CER reductions over the baseline, and surpasses the unlikelihood training method which directly reduces the generation probability oftoken. Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001 |
ICASSP | 2 |
| 2020 | Independent Language Modeling Architecture for End-To-End ASRabstractThe attention-based end-to-end (E2E) automatic speech recognition (ASR) architecture allows for joint optimization of acoustic and language models within a single network. However, in a vanilla E2E ASR architecture, the decoder sub-network (subnet), which incorporates the role of the language model (LM), is conditioned on the encoder output. This means that the acoustic encoder and the language model are entangled that doesn’t allow language model to be trained separately from external text data. To address this problem, in this work, we propose a new architecture that separates the decoder subnet from the encoder output. In this way, the decoupled subnet becomes an independently trainable LM subnet, which can easily be updated using the external text data. We study two strategies for updating the new architecture. Experimental results show that, 1) the independent LM architecture benefits from external text data, achieving 9.3% and 22.8% relative character and word error rate reduction on Mandarin HKUST and English NSC datasets respectively; 2) the proposed architecture works well with external LM and can be generalized to different amount of labelled data. Van Tung Pham, Haihua Xu 0001, Yerbolat Khassanov, Zhiping Zeng, Chng Eng Siong, Chongjia Ni, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 6 |
| 2020 | Speech Transformer with Speaker Aware Persistent Memory
Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 2 |
| 2020 | Universal Speech Transformer
Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 2 |
| 2020 | Cross Attention with Monotonic Alignment for Speech Transformer
Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 2 |
| 2019 | Constrained Output Embeddings for End-to-End Code-Switching Speech Recognition with Only Monolingual DataabstractThe lack of code-switch training data is one of the major concerns in the development of end-to-end code-switching automatic speech recognition (ASR) models. In this work, we propose a method to train an improved end-to-end code-switching ASR using only monolingual data. Our method encourages the distributions of output token embeddings of monolingual languages to be similar, and hence, promotes the ASR model to easily code-switch between languages. Specifically, we propose to use Jensen-Shannon divergence and cosine distance based constraints. The former will enforce output embeddings of monolingual languages to possess similar distributions, while the later simply brings the centroids of two distributions to be close to each other. Experimental results demonstrate high effectiveness of the proposed method, yielding up to 4.5% absolute mixed error rate improvement on Mandarin-English code-switching ASR task. Yerbolat Khassanov, Haihua Xu 0001, Van Tung Pham, Zhiping Zeng, Chng Eng Siong, Chongjia Ni, Bin Ma 0001 |
INTERSPEECH | 6 |
| 2019 | Multi-Task Multi-Network Joint-Learning of Deep Residual Networks and Cycle-Consistency Generative Adversarial Networks for Robust Speech Recognition
Shengkui Zhao, Chongjia Ni, Rong Tong, Bin Ma 0001 |
INTERSPEECH | 2 |
| 2017 | Efficient methods to train multilingual bottleneck feature extractors for low resource keyword searchabstractTraining a bottleneck feature (BNF) extractor with multilingual data has been common in low resource keyword search. In a low resource application, the amount of transcribed target language data is limited while there are usually plenty of multilingual data. In this paper, we investigated two methods to train efficient multilingual BNF extractors for low resource keyword search. One method is to use the target language data to update an existing BNF extractor, and another method is to combine the target language data to train a new multilingual BNF extractor from the start. In these two methods, we proposed to use long short-term memory recurrent neural network based language identification to select utterances in the multilingual training data that are acoustically close to the target language. Experiments on Swahili in the OpenKWS15 data demonstrated the efficiency of our proposed methods. The first method facilitates rapid system development, while both methods outperform using baseline BNF extractors in terms of accuracy. Chongjia Ni, Cheung-Chi Leung, Lei Wang 0020, Nancy F. Chen, Bin Ma 0001 |
ICASSP | 1 |
| 2017 | Modification on LSA speech enhancement for speech recognitionabstractSpeech recognition performance deteriorates in face of unknown noise. Speech enhancement offers a solution by reducing the noise in speech at runtime. However, it also introduces artificial distortions to the speech signals. In this paper, we aim at reducing the artifacts that has adverse effects on speech recognition. With this motivation, we propose a modification scheme including smoothing adaptation to frame SNR and reestimation of a priori SNR for spectral-domain log-spectral-amplitude (LSA) speech enhancement. The experiments show that the proposed scheme of enhancement significantly improves the performance of the state-of-the-art speech recognition over the baseline speech enhancement. Chang Huai You, Bin Ma 0001, Chongjia Ni |
ICASSP | 3 |
| 2016 | Exemplar-inspired strategies for low-resource spoken keyword search in SwahiliabstractWe present exemplar-inspired low-resource spoken keyword search strategies for acoustic modeling, keyword verification, and system combination. This state-of-the-art system was developed by the SINGA team in the context of the 2015 NIST Open Keyword Search Evaluation (OpenKWS15) using conversational Swahili provided by the IARPA Babel program. In this work, we elaborate on the following: (1) exploiting exemplar training samples to construct a non-parametric acoustic model using kernel density estimation at test time; (2) rescoring hypothesized keyword detections through quantifying their acoustic similarity with exemplar training samples; (3 ) extending our previously proposed system combination approach to incorporate prosody features of exemplar keyword samples. Nancy F. Chen, Van Tung Pham, Haihua Xu 0001, Van Hai Do, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Chin-Hui Lee 0001, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 6 |
| 2016 | Cross-lingual deep neural network based submodular unbiased data selection for low-resource keyword searchabstractIn this paper, we propose a cross-lingual deep neural network (DNN) based submodular unbiased data selection approach for low-resource keyword search (KWS). A small amount (e.g. one hour) of transcribed data is used to conduct cross-lingual transfer. The frame-level senone sequence activated by the cross-lingual DNN is used to represent each untranscribed speech utterance. The proposed submodular function considers utterance length normalization and the feature distribution matched to a development set. Experiments are conducted by selecting 9 hours of Tamil speech for the 2014 NIST Open Keyword Search Evaluation (OpenKWS14). The proposed data selection approach provides 35.8% relative actual term weighted value (ATWV) improvement over random selection on the OpenKWS14 Evalpartl data set. Further analysis of the experimental results shows that both utterance length normalization and the feature distribution estimated from a development set deployed in the submodular function can suppress the preference to select long utterances. The selected utterances can cover a more diverse range of tri-phones, words, and acoustic variations from a wider set of utterances. Moreover, the wider coverage of words also benefits the acquired linguistic knowledge, which also contributes to improving KWS performance. Chongjia Ni, Cheung-Chi Leung, Lei Wang 0020, Feng Rao, Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 1 |
| 2016 | Toward High-Performance Language-Independent Query-by-Example Spoken Term Detection for MediaEval 2015: Post-Evaluation Analysis
Cheung-Chi Leung, Lei Wang 0020, Haihua Xu 0001, Jingyong Hou, Van Tung Pham, Hang Lv 0001, Lei Xie 0001, Chongjia Ni, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 9 |
| 2016 | Rapid Update of Multilingual Deep Neural Network for Low-Resource Keyword Search
Chongjia Ni, Lei Wang 0020, Cheung-Chi Leung, Feng Rao, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2016 | Semi-Supervised and Cross-Lingual Knowledge Transfer Learnings for DNN Hybrid Acoustic Models Under Low-Resource Conditions
Haihua Xu 0001, Chongjia Ni, Hao Huang 0009, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2015 | Low-resource keyword search strategies for tamilabstractWe propose strategies for a state-of-the-art keyword search (KWS) system developed by the SINGA team in the context of the 2014 NIST Open Keyword Search Evaluation (OpenKWS14) using conversational Tamil provided by the IARPA Babel program. To tackle low-resource challenges and the rich morphological nature of Tamil, we present highlights of our current KWS system, including: (1) Submodular optimization data selection to maximize acoustic diversity through Gaussian component indexed N-grams; (2) Keywordaware language modeling; (3) Subword modeling of morphemes and homophones. Nancy F. Chen, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Van Tung Pham, Haihua Xu 0001, Tze Siong Lau, Su Jun Leow, Boon Pang Lim, Cheung-Chi Leung, Lei Wang 0020, Chin-Hui Lee 0001, Alvina Goh, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 2 |
| 2015 | A keyword-aware grammar framework for LVCSR-based spoken keyword searchabstractIn this paper, we proposed a method to realize the recently developed keyword-aware grammar for LVCSR-based keyword search using weight finite-state automata (WFSA). The approach creates a compact and deterministic grammar WFSA by inserting keyword paths to an existing n-gram WFSA. Tested on the evalpart1 data of the IARPA Babel OpenKWS13 Vietnamese and OpenKWS14 Tamil limited language pack tasks, the experimental results indicate the proposed keyword-aware framework achieves significant improvement, with about 50% relative actual term weighted value (ATWV) enhancement for both languages. Comparisons between the keyword-aware grammar and our previously proposed n-gram LM based approximation approach for the grammar also show that the KWS performances of these two realizations are complementary. I-Fan Chen, Chongjia Ni, Boon Pang Lim, Nancy F. Chen, Chin-Hui Lee 0001 |
ICASSP | 2 |
| 2015 | Unsupervised data selection and word-morph mixed language model for tamil low-resource keyword searchabstractThis paper considers an unsupervised data selection problem for the training data of an acoustic model and the vocabulary coverage of a keyword search system in low-resource settings. We propose to use Gaussian component index based n-grams as acoustic features in a submodular function for unsupervised data selection. The submodular function provides a near-optimal solution in terms of the objective being optimized. Moreover, to further resolve the high out-of-vocabulary (OOV) rate for morphologically-rich languages like Tamil, word-morph mixed language modeling is also considered. Our experiments are conducted on the Tamil speech provided by the IAPRA Babel program for the 2014 NIST Open Keyword Search Evaluation (OpenKWS14). We show that the selection of data plays an important role to the word error rate of the speech recognition system and the actual term weighted value (ATWV) of the keyword search system. The 10 hours of speech selected from the full language pack (FLP) using the proposed algorithm provides a relative 23.2% and 20.7% ATWV improvement over two other data subsets, the 10-hour data from the limited language pack (LLP) defined by IARPA and the 10 hours of speech randomly selected from the FLP, respectively. The proposed algorithm also increases the vocabulary coverage, implicitly alleviating the OOV problem: The number of OOV search terms drops from 1,686 and 1,171 in the two baseline conditions to 972. Chongjia Ni, Cheung-Chi Leung, Lei Wang 0020, Nancy F. Chen, Bin Ma 0001 |
ICASSP | 1 |
| 2015 | Submodular data selection with acoustic and phonetic features for automatic speech recognitionabstractIn this paper, we propose to use acoustic feature based submodular function optimization to select a subset of untranscribed data for manual transcription, and retrain the initial acoustic model with the additional transcribed data. The acoustic features are obtained from an unsupervised Gaussian mixture model. We also integrate the acoustic features with the phonetic features, which are obtained from an initial ASR system, in the submodular function. Submodular function optimization has been theoretically shown its near-optimal guarantee. We performed the experiments on 1000 hours of Mandarin mobile phone speech, in which 300 hours of initial data was for the training of an initial acoustic model. The experimental results show that the acoustic feature based approach, which does not rely on an initial ASR system, performs as well as the phonetic feature based approach. Moreover, there is complementary effect between the acoustic feature based and the phonetic feature based data selection. The submodular function with the combined features provides a relative 4.8% character error rate (CER) reduction over the corresponding ASR system using random selection. We also include the desired feature distribution obtained from a development set in a generalized function, but the improvement is insignificant. Chongjia Ni, Lei Wang 0020, Cheung-Chi Leung, Bin Ma 0001 |
ICASSP | 1 |
| 2015 | "multilingual" deep neural network for music genre classification
Jia Dai, Chongjia Ni, Like Dong |
INTERSPEECH | 3 |
| 2015 | Smarter driving with IDA, the intelligent driving assistant for singapore
Andreea I. Niculescu, Ngoc Thuy Huong Thai, Chongjia Ni, Boon Pang Lim, Kheng Hui Yeo, Rafael E. Banchs |
INTERSPEECH | 3 |
| 2014 | System and keyword dependent fusion for spoken term detectionabstractSystem combination (or data fusion1) is known to provide significant improvement for spoken term detection (STD). The key issue of the system combination is how to effectively fuse the various scores of participant systems. Currently, most system combination methods are system and keyword independent, i.e. they use the same arithmetic functions to combine scores for all keywords. Although such strategy improve keyword search performance, the improvement is limited. In this paper we first propose an arithmetic-based system combination method to incorporate the system and keyword characteristics into the fusion procedure to enhance the effectiveness of system combination. The method incorporates a system-keyword dependent property, which is the number of acceptances in this paper, into the combination procedure. We then introduce a discriminative model to combine various useful system and keyword characteristics into a general framework. Improvements over standard baselines are observed on the Vietnamese data from IARPA Babel program with the NIST OpenKWS13 Evaluation setup. Van Tung Pham, Nancy F. Chen, Sunil Sivadas, Haihua Xu 0001, I-Fan Chen, Chongjia Ni, Chng Eng Siong, Haizhou Li 0001 |
SLT | 6 |
| 2012 | From English pitch accent detection to Mandarin stress detection, where is the difference?
Chongjia Ni, Bo Xu 0002 |
Comput. Speech Lang. | 1 |
| 2012 | Automatic Prosodic Break Detection and Feature Analysis
Chongjia Ni, Aiying Zhang, Bo Xu 0002 |
J. Comput. Sci. Technol. | 1 |
| 2011 | Prosody dependent Mandarin speech recognitionabstractIn this paper, we discuss how to model and train Mandarin prosody dependent acoustic model based on automatic prosody annotation corpus. Based on prosody annotation corpus, we first utilize our proposed methods to train prosody dependent and prosody independent tonal syllable model, and then use these models to get the mixed acoustic models. In this paper, we also utilize tone model to improve the correct rate of tonal syllable through revising the tone of the tonal syllable at certain significant level. When compared with the baseline system, the performance of our proposed mixed speech recognition system improves the correct rate of tonal syllable significantly. Chongjia Ni, Bo Xu 0002 |
IJCNN | 1 |
| 2011 | Automatic Prosodic Events Detection by Using Syllable-Based Acoustic, Lexical and Syntactic FeaturesabstractAutomatic prosodic events detection and annotation are important for both speech understanding and natural speech synthesis. In this paper, the complementary model method is proposed to detect prosodic events. This method discards the independent assumption between the acoustic features and the lexical and syntactic features, models not only the features of the current syllable but also the contextual features of the current syllable at the model level, and realizes the complementarities by taking the advantages of each model. The experiments on Boston University Radio News Corpus show that the complementary model can yield 91.40% pitch accent detection accuracy rate, 95.19% intonational phrase boundaries (IPB) detection accuracy rate and 93.96% break index detection accuracy rate. When compared with the previous work, the results for pitch accent, IPB and break index detection are significantly better. Index Terms: complementary model, boosting classification and regression tree (CART), conditional random fields (CRFs) Chongjia Ni, Bo Xu 0002 |
INTERSPEECH | 1 |
| 2010 | Mandarin stress detection using hierarchical model based boosting classification and regression treeabstractAutomatic stress detection is important for both speech understanding and natural speech synthesis. In this paper, we develop hierarchical model based boosting classification and regression tree (CART) to detect Mandarin stress by using acoustic evidence and text information. When comparing with previous proposed method at the same training and test sets, there are 2.52% and 1.09% absolute accuracy rate improvements respectively. We also analyze the differences between Mandarin stress detection and English pitch accent prediction, and prove some linguistic conclusions based on the large corpus in a different way. Chongjia Ni, Bo Xu 0002 |
IJCNN | 1 |
| 2010 | Using prosody to improve Mandarin automatic speech recognition
Chongjia Ni, Bo Xu 0002 |
INTERSPEECH | 1 |