I-Fan Chen

dblp:53/463 · DBLP profile ↗
← Back
30ranked-venue papers
9as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 20 · 7 first-author · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2024 Hot-Fixing Wake Word Recognition for End-to-End ASR Via Neural Model Reprogramming
abstract
This paper proposes two novel variants of neural reprogramming to enhance wake word recognition in streaming end-to-end ASR models without updating model weights. The first, "trigger-frame reprogramming", prepends the input speech feature sequence with the learned trigger-frames of the target wake word to adjust ASR model’s hidden states for improved wake word recognition. The second, "predictor-state initialization", trains only the initial state vectors (cell and hidden states) of the LSTMs in the prediction network. When applying to a baseline LibriSpeech Emformer RNN-T model with a 98% wake word verification false rejection rate (FRR) on unseen wake words, the proposed approaches achieve 76% and 97% relative FRR reductions with no increase on false acceptance rate. In-depth characteristic analyses of the proposed approaches are also conducted to provide deeper insights. These approaches offer an effective hot-fixing methods to improve wake word recognition performance in deployed production ASR models without the need for model updates.
Pin-Jui Ku, I-Fan Chen, Chao-Han Huck Yang, Anirudh Raju, Pranav Dheram, Pegah Ghahremani, Brian King, Roger Ren, Phani S. Nidadavolu
ICASSP2
2023 Low-Rank Adaptation of Large Language Model Rescoring for Parameter-Efficient Speech Recognition
abstract
We propose a neural language modeling system based on low-rank adaptation (LoRA) for speech recognition output rescoring. Although pretrained language models (LMs) like BERT have shown superior performance in second-pass rescoring, the high computational cost of scaling up the pretraining stage and adapting the pretrained models to specific domains limit their practical use in rescoring. Here we present a method based on low-rank decomposition to train a rescoring BERT model and adapt it to new domains using only a fraction (0.08%) of the pretrained parameters. These inserted matrices are optimized through a discriminative training objective along with a correlation-based regularization loss. The proposed low-rank adaptation RescoreBERT (LoRB) architecture is evaluated on LibriSpeech and internal datasets with decreased training times by factors between 5.4 and 3.6.
Chao-Han Huck Yang, Jari Kolehmainen, Prashanth Gurunath Shivakumar, Yile Gu, Sungho Ryu, Roger Ren, Aditya Gourav, I-Fan Chen, Yi-Chieh Liu, Tuan Dinh, Ankur Gandhe, Denis Filimonov, Shalini Ghosh, Andreas Stolcke, Ariya Rastrow, Ivan Bulyko
ASRU10
2023 Prune Then Distill: Dataset Distillation with Importance Sampling
abstract
The development of large datasets for various tasks has driven the success of deep learning models but at the cost of increased label noise, duplication, collection challenges, storage capabilities, and training requirements. In this work, we investigate whether all samples in large datasets contribute equally to better model accuracy. We study statistical and mathematical techniques to reduce redundancies in datasets by directly optimizing data samples for the generalization accuracy of deep learning models. Existing dataset optimization approaches include analytic methods that remove unimportant samples and synthetic methods that generate new datasets to maximize the generalization accuracy. We develop Prune then distill, a combination of analytic and synthetic dataset optimization algorithms, and demonstrate up to 15% relative improvement in generalization accuracy over either approach used independently on standard image and audio classification tasks. Additionally, we demonstrate up to 38% improvement in generalization accuracy of dataset pruning algorithms by maintaining class balance while pruning.
Anirudh S. Sundar, Gökçe Keskin, Chander Chandak, I-Fan Chen, Pegah Ghahremani, Shalini Ghosh
ICASSP4
2022 Toward Fairness in Speech Recognition: Discovery and mitigation of performance disparities
abstract
As for other forms of AI, speech recognition has recently been examined with respect to performance disparities across different user cohorts.One approach to achieve fairness in speech recognition is to (1) identify speaker cohorts that suffer from subpar performance and (2) apply fairness mitigation measures targeting the cohorts discovered.In this paper, we report on initial findings with both discovery and mitigation of performance disparities using data from a product-scale AI assistant speech recognition system.We compare cohort discovery based on geographic and demographic information to a more scalable method that groups speakers without human labels, using speaker embedding technology.For fairness mitigation, we find that oversampling of underrepresented cohorts, as well as modeling speaker cohort membership by additional input variables, reduces the gap between top-and bottom-performing cohorts, without deteriorating overall recognition accuracy.
Pranav Dheram, Murugesan Ramakrishnan, Anirudh Raju, I-Fan Chen, Brian King, Katherine Powell, Melissa Saboowala, Karan Shetty, Andreas Stolcke
INTERSPEECH4
2022 Learning to rank with BERT-based confidence models in ASR rescoring
Ting-Wei Wu, I-Fan Chen, Ankur Gandhe
INTERSPEECH2
2022 An Experimental Study on Private Aggregation of Teacher Ensemble Learning for End-to-End Speech Recognition
abstract
Differential privacy (DP) is one data protection avenue to safeguard user information used for training deep models by imposing noisy distortion on privacy data. Such a noise perturbation often results in a severe performance degradation in automatic speech recognition (ASR) in order to meet a privacy budget ε. Private aggregation of teacher ensemble (PATE) utilizes ensemble probabilities to improve ASR accuracy when dealing with the noise effects controlled by small values of ε. We extend PATE learning to work with dynamic patterns, namely speech utterances, and perform a first experimental demonstration that it prevents acoustic data leakage in ASR training. We evaluate three end-to-end deep models, including LAS, hybrid CTC/attention, and RNN transducer, on the open-source LibriSpeech and TIMIT corpora. PATE learning-enhanced ASR models outperform the benchmark DP-SGD mechanisms, especially under strict DP budgets, giving relative word error rate reductions between 26.2% and 27.5% for an RNN transducer model evaluated with LibriSpeech. We also introduce a DP-preserving ASR solution for pretraining on public speech corpora.
Chao-Han Huck Yang, I-Fan Chen, Andreas Stolcke, Sabato Marco Siniscalchi, Chin-Hui Lee 0001
SLT2
2019 End-to-end Anchored Speech Recognition
abstract
Voice-controlled house-hold devices, like Amazon Echo or Google Home, face the problem of performing speech recognition of device-directed speech in the presence of interfering background speech, i.e., background noise and interfering speech from another person or media device in proximity need to be ignored. We propose two end-to-end models to tackle this problem with information extracted from the anchored segment. The anchored segment refers to the wake-up word part of an audio stream, which contains valuable speaker information that can be used to suppress interfering speech and background noise. The first method is called Multi-source Attention where the attention mechanism takes both the speaker information and decoder state into consideration. The second method directly learns a frame-level mask on top of the encoder output. We also explore a multi-task learning setup where we use the ground truth of the mask to guide the learner. Given that audio data with interfering speech is rare in our training data, we also propose a way to synthesize "noisy" speech from "clean" speech to mitigate the mismatch between training and test data. Our proposed methods show up to 15% relative reduction in WER for Amazon Alexa live data with interfering background speech without significantly degrading on clean speech.
Yiming Wang 0006, I-Fan Chen, Yuzong Liu, Tongfei Chen, Björn Hoffmeister
ICASSP3
2019 Two Tiered Distributed Training Algorithm for Acoustic Modeling
Pranav Ladkat, Oleg Rybakov, Radhika Arava, Sree Hari Krishnan Parthasarathi, I-Fan Chen, Nikko Strom
INTERSPEECH5
2017 Robust Speech Recognition via Anchor Word Representations
Brian John King, I-Fan Chen, Yonatan Vaizman, Yuzong Liu, Roland Maas, Sree Hari Krishnan Parthasarathi, Björn Hoffmeister
INTERSPEECH2
2016 Exemplar-inspired strategies for low-resource spoken keyword search in Swahili
abstract
We present exemplar-inspired low-resource spoken keyword search strategies for acoustic modeling, keyword verification, and system combination. This state-of-the-art system was developed by the SINGA team in the context of the 2015 NIST Open Keyword Search Evaluation (OpenKWS15) using conversational Swahili provided by the IARPA Babel program. In this work, we elaborate on the following: (1) exploiting exemplar training samples to construct a non-parametric acoustic model using kernel density estimation at test time; (2) rescoring hypothesized keyword detections through quantifying their acoustic similarity with exemplar training samples; (3 ) extending our previously proposed system combination approach to incorporate prosody features of exemplar keyword samples.
Nancy F. Chen, Van Tung Pham, Haihua Xu 0001, Van Hai Do, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Chin-Hui Lee 0001, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001
ICASSP7
2015 Low-resource keyword search strategies for tamil
abstract
We propose strategies for a state-of-the-art keyword search (KWS) system developed by the SINGA team in the context of the 2014 NIST Open Keyword Search Evaluation (OpenKWS14) using conversational Tamil provided by the IARPA Babel program. To tackle low-resource challenges and the rich morphological nature of Tamil, we present highlights of our current KWS system, including: (1) Submodular optimization data selection to maximize acoustic diversity through Gaussian component indexed N-grams; (2) Keywordaware language modeling; (3) Subword modeling of morphemes and homophones.
Nancy F. Chen, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Van Tung Pham, Haihua Xu 0001, Tze Siong Lau, Su Jun Leow, Boon Pang Lim, Cheung-Chi Leung, Lei Wang 0020, Chin-Hui Lee 0001, Alvina Goh, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001
ICASSP3
2015 A keyword-aware grammar framework for LVCSR-based spoken keyword search
abstract
In this paper, we proposed a method to realize the recently developed keyword-aware grammar for LVCSR-based keyword search using weight finite-state automata (WFSA). The approach creates a compact and deterministic grammar WFSA by inserting keyword paths to an existing n-gram WFSA. Tested on the evalpart1 data of the IARPA Babel OpenKWS13 Vietnamese and OpenKWS14 Tamil limited language pack tasks, the experimental results indicate the proposed keyword-aware framework achieves significant improvement, with about 50% relative actual term weighted value (ATWV) enhancement for both languages. Comparisons between the keyword-aware grammar and our previously proposed n-gram LM based approximation approach for the grammar also show that the KWS performances of these two realizations are complementary.
I-Fan Chen, Chongjia Ni, Boon Pang Lim, Nancy F. Chen, Chin-Hui Lee 0001
ICASSP1
2015 Rapid adaptation for deep neural networks through multi-task learning
abstract
We propose a novel approach to addressing the adaptation effectiveness issue in parameter adaptation for deep neural network (DNN) based acoustic models for automatic speech recognition by adding one or more small auxiliary output layers modeling broad acoustic units, such as mono-phones or tied-state (often called senone) clusters. In scenarios with a limited amount of available adaptation data, most senones are usually rarely seen or not observed, and consequently the ability to model them in a new condition is often not fully exploited. With the original senone classification task as the primary task, and adding auxiliary mono-phone/senone-cluster classification as the secondary tasks, multi-task learning (MTL) is employed to adapt the DNN parameters. With the proposed MTL adaptation framework, we improve the learning ability of the original DNN structure, then enlarge the coverage of the acoustic space to deal with the unseen senone problem, and thus enhance the discrimination power of the adapted DNN models. Experimental results on the 20,000-word open vocabulary WSJ task demonstrate that the proposed framework consistently outperforms the conventional linear hidden layer adaptation schemes without MTL by providing 5.4% relative reduction in word error rate (WERR) with only 1 single adaptation utterance, and 10.7% WERR with 40 adaptation utterances against the un-adapted DNN models.
Zhen Huang 0001, Jinyu Li 0001, Sabato Marco Siniscalchi, I-Fan Chen, Ji Wu 0002, Chin-Hui Lee 0001
INTERSPEECH4
2015 Maximum a posteriori adaptation of network parameters in deep models
abstract
We present a Bayesian approach to adapting parameters of a well-trained context-dependent, deep-neural-network, hidden Markov model (CD-DNN-HMM) to improve automatic speech recognition performance. Given an abundance of DNN parameters but with only a limited amount of data, the effectiveness of the adapted DNN model can often be compromised. We formulate maximum a posteriori (MAP) adaptation of parameters of a specially designed CD-DNN-HMM with an augmented linear hidden networks connected to the output tied states, or senones, and compare it to feature space MAP linear regression previously proposed. Experimental evidences on the 20,000-word open vocabulary Wall Street Journal task demonstrate the feasibility of the proposed framework. In supervised adaptation, the proposed MAP adaptation approach provides more than 10% relative error reduction and consistently outperforms the conventional transformation based methods. Furthermore, we present an initial attempt to generate hierarchical priors to improve adaptation efficiency and effectiveness with limited adaptation data by exploiting similarities among senones.
Zhen Huang 0001, Sabato Marco Siniscalchi, I-Fan Chen, Jinyu Li 0001, Jiadong Wu, Chin-Hui Lee 0001
INTERSPEECH3
2015 Tunable keyword-aware language modeling and context dependent fillers for LVCSR-based spoken keyword search
abstract
We explore the potential of using keyword-aware language modeling to extend the ability of trading higher false alarm rates in exchange for lower miss detection rates in LVCSRbased keyword search (KWS). A context-dependent keyword language modeling method is also proposed to further enhance the keyword-aware language modeling framework by reducing the number of false alarms often sacrificed in order to achieve the desirable low miss detection rates. We demonstrate that by using keyword-aware language modeling, a KWS system is able to achieve different operating points (misses vs. false alarms) by tuning a parameter in language modeling. We observe a relative gain of 20% in actual term weighted value (ATWV) performance with the keyword-aware KWS systems over the conventional LVCSR-based KWS systems when testing on the English Switchboard data. Moreover the proposed context-dependent keyword language modeling could further achieve a 9% relative ATWV improvement over the original keyword-aware KWS systems for single-word keywords which cause the most false alarms.
Tze Siong Lau, I-Fan Chen
INTERSPEECH2
2014 Attribute based lattice rescoring in spontaneous speech recognition
abstract
In this paper we extend attribute-based lattice rescoring to spontaneous speech recognition. This technique is based on two key features: (i) an attribute-based frontend, which consists of a bank of speech attribute detectors followed up by an evidence merger that generates confidence scores (e.g., sub-word posterior probabilities), and (ii) a rescoring module that integrates information generated by the frontend into an existing ASR engine through lattice rescoring. The speech attributes used in this work are phonetic features, such as frication and palatalization. Experimental results on the Switchboard part of the NIST 2000 Hub5 data set demonstrate that the proposed approach outperforms LVCSR systems based on Gaussian mixture model/ hidden Markov model (GMM/HMM) that does not use attribute related information. Furthermore, a small yet promising improvement is also observed when rescoring word-lattices generated by a state-of-the-art ASR system using deep neural networks. Different frontend configuration are investigated and tested.
I-Fan Chen, Sabato Marco Siniscalchi, Chin-Hui Lee 0001
ICASSP1
2014 A keyword-boosted sMBR criterion to enhance keyword search performance in deep neural network based acoustic modeling
I-Fan Chen, Nancy F. Chen, Chin-Hui Lee 0001
INTERSPEECH1
2014 Feature space maximum a posteriori linear regression for adaptation of deep neural networks
abstract
We propose a feature space maximum a posteriori (MAP) linear regression framework to adapt parameters for context dependent deep neural network hidden Markov models (CD-DNN-HMMs). Due to the huge amount of parameters used in DNN acoustic models in large vocabulary continuous speech recognition, the problem of over-fitting can be severe in DNN adaptation, thus often impair the robustness of the adapted DNN model. Linear input network (LIN) as a straight-forward feature space adaptation method for DNN, similar to feature space maximum likelihood linear regression (fMLLR), can potentially suffer from the same robustness situation. The proposed adaptation framework is built based on MAP estimation of the LIN parameters by incorporating prior knowledge into the adaptation process. Experimental results on the Switchboard task show that against the speaker independent CD-DNN-HMM systems, LIN provides 4.28% relative word error rate reduction (WERR) and the proposed fMAPLIN method is able to provide further 1.15% (totally 5.43%) WERR on top of LIN.
Zhen Huang 0001, Jinyu Li 0001, Sabato Marco Siniscalchi, I-Fan Chen, Chao Weng, Chin-Hui Lee 0001
INTERSPEECH4
2014 System and keyword dependent fusion for spoken term detection
abstract
System combination (or data fusion1) is known to provide significant improvement for spoken term detection (STD). The key issue of the system combination is how to effectively fuse the various scores of participant systems. Currently, most system combination methods are system and keyword independent, i.e. they use the same arithmetic functions to combine scores for all keywords. Although such strategy improve keyword search performance, the improvement is limited. In this paper we first propose an arithmetic-based system combination method to incorporate the system and keyword characteristics into the fusion procedure to enhance the effectiveness of system combination. The method incorporates a system-keyword dependent property, which is the number of acceptances in this paper, into the combination procedure. We then introduce a discriminative model to combine various useful system and keyword characteristics into a general framework. Improvements over standard baselines are observed on the Vietnamese data from IARPA Babel program with the NIST OpenKWS13 Evaluation setup.
Van Tung Pham, Nancy F. Chen, Sunil Sivadas, Haihua Xu 0001, I-Fan Chen, Chongjia Ni, Chng Eng Siong, Haizhou Li 0001
SLT5
2013 A hybrid HMM/DNN approach to keyword spotting of short words
abstract
An HMM/DNN framework is proposed to address the issues of short-word detection. The first-stage keyword hypothesizer is redesigned with a context-aware keyword model and a 9state filler model to reduce the miss rate from 80% to 6% and increase the figure-of-merit (FOM) from 6.08% to 21.88% for short words. The hypothesizer is followed by a MLP-based second-stage keyword verifier to further reduce its putative hits. To enhance short word detection, three new techniques, including an HMM-based feature transformation for the MLPs, knowledge-based features, and deep neural networks, are incorporated into redesigning the verifier. With a set of nine short keywords from the TIMIT set the best FOM we had achieved for the proposed KWS system was 42.79%, which is comparable with that of 42.6% for long content words and much better than the FOM of 18.4% for short keywords reported in previous research [10].
I-Fan Chen
INTERSPEECH1
2013 A resource-dependent approach to word modeling for keyword spotting
abstract
A hierarchical framework is proposed to address the issues of modeling different type of words in keyword spotting (KWS). Keyword models are built at various levels according to the availability of training set resources for each individual word. The proposed approach improves the performance of KWS even when no training speech is available for the keywords. It also suggests an easier way to collect training data for these resource-limited words. Experimental results show that the proposed framework improves performance in KWS in a figure-of-merit (FOM) metric regardless of the number of training instances for each keyword. For words with abundant speech data, the proposed method exploits the training data better than the conventional modeling technique and boosts the system FOM from 9.79% to 42.78%. For words with a small amount of training data, the new method increases the system FOM from 29.05% to 49.06%. Even for keywords without any training examples, the new modeling scheme improves the system FOM from 60.96% to 66.51%.
I-Fan Chen
INTERSPEECH1
2012 A Study on Using Word-Level HMMs to Improve ASR Performance over State-of-the-Art Phone-Level Acoustic Modeling for LVCSR
abstract
In this paper, we propose word-level hidden Markov models (HMMs) to supplement state-of-the-art phone-based acoustic modeling in order to enhance the performance of automatic speech recognition (ASR) system. Each word in a vocabulary is initially modeled by well-trained triphone models. Maximum a posteriori adaptation is then applied to generate models for words with a large number of occurrences in the training set so that the acoustic distribution of the words can be modeled more precisely. Experimental results show that the proposed wordbased systems outperform phone-based systems on the TIMIT task with a small training corpus. While in tasks with plenty of training data, word-based systems still show improvements over phone-based systems, such as the WSJ task. Furthermore the word-based systems have a better discriminating ability on short words and homophones. They are also more robust to language model weight variation than conventional phone-based systems.
I-Fan Chen
INTERSPEECH1
2010 Phonetic subspace mixture model for speaker diarization
abstract
This paper presents an improved distance measure for speaker clustering in speaker diarization systems. The proposed phonetic subspace mixture (PSM) model introduces phonetic information to the ΔBIC distance measure. Therefore, the new PSM model-based ΔBIC distance measure can remove the effect of phonetic content on the diarization results. The typical ΔBIC distance measure can be seen as a special case of the new ΔBIC distance measure. Our experiment results show that the new distance measurement consistently improves the speaker diarization performance on three datasets. Index Terms: BIC, phonetic information, speaker diarization 1.
I-Fan Chen, Shih-Sian Cheng, Hsin-Min Wang
INTERSPEECH1
2010 Bayesian speaker recognition using Gaussian mixture model and laplace approximation
abstract
This paper presents a Bayesian approach for Gaussian mix-ture model (GMM)-based speaker identification. Some ap-proaches evaluate the speaker score of a test speech utterance using a single data likelihood over the GMM learned by point estimation methods according to the maximum likelihood or maximum a posteriori criteria. In contrast, the Bayesian ap-proach evaluates the score by using the expectation of the data likelihood over the posterior distribution of the model parame-ters, which is depicted by Bayesian integration. However, as the integration can not be derived analytically, we apply Laplace approximation to the derivations. Theoretically, we show that the proposed Bayesian approach is equivalent to the GMM-UBM approach when infinite training data is available for each speaker. The results of speaker identification experiments on the TIMIT corpus show that the proposed Bayesian approach consistently outperforms the GMM-UBM approach under very limited training data conditions, although the improvement is not very significant. Index Terms: speaker identification, speaker recognition, Bayesian inference, GMM-UBM
Shih-Sian Cheng, I-Fan Chen, Hsin-Min Wang
INTERSPEECH2
2010 An intelligent resource management scheme for heterogeneous WiFi and WiMAX multi-hop relay networks
Chenn-Jung Huang, Kai-Wen Hu, I-Fan Chen, You-Jia Chen, Hong-Xin Chen
Expert Syst. Appl.3
2009 Articulatory feature asynchrony analysis and compensation in detection-based ASR
abstract
This paper investigates the effects of two types of imperfection, namely detection errors and articulatory feature asynchrony, of the front-end articulatory feature detector on the performance of a detection-based ASR system. Based on a set of variable-controlled experiments, we find that articulatory feature asynchrony is the major issue that should be addressed in detection-based ASR. To this end, we propose several methods to reduce the asynchrony or the effects of asynchrony. The results are quite promising; for example, currently, we can achieve 67.67 % phone accuracy in the TIMIT free phone recognition task with only 11 binary-valued articulatory features. Index Terms: articulatory feature asynchrony, detection-based ASR, speech recognition
I-Fan Chen, Hsin-Min Wang
INTERSPEECH1
2009 QoS-aware roadside base station assisted routing in vehicular networks
Chenn-Jung Huang, Yi-Ta Chuang, You-Jia Chen, Dian-Xiu Yang, I-Fan Chen
Eng. Appl. Artif. Intell.5
2009 An intelligent infotainment dissemination scheme for heterogeneous vehicular networks
Chenn-Jung Huang, You-Jia Chen, I-Fan Chen, Tsung-Hsien Wu
Expert Syst. Appl.3
2006 Using File Grouping to Improve the Disk Performance (Extended Abstract)
abstract
As the speed gap between CPU and the secondary storage device will not be narrowing in the foreseeable future, file grouping can be a promising way to reduce the disk I/O latency. The order of sequential access among files observed during the execution of individual programs is very predictable. Based on this idea, we propose a new file grouping model called program-based grouping (PBG). Through the Reiser file system, we implemented PBG into Linux kernel. The experiments demonstrate that PBG out performs both ReiserFS and Ext3. Compared with ReiserFS, PBG can improve the ReiserFS performance by up to 66%
Tsozen Yeh, Joseph Arul, Jia-Shian Wu, I-Fan Chen, Kuo-Hsin Tan
HPDC4
2006 A new framework for system combination based on integrated hypothesis space
abstract
In this paper, a new concept of integrated hypothesis space for large vocabulary continuous speech recognition (LVCSR) system combination is proposed. Unlike the conventional systems combination approaches such as ROVER, the hypothesis spaces are directly integrated here without string alignment. In this way the timing information for all word hypotheses is well preserved and the new framework is more flexible on rescoring approaches used. Four rescoring criteria on the integrated hypothesis space were further explored and experiments on Chinese broadcast news corpus indicated improved performance.
I-Fan Chen, Lin-Shan Lee
INTERSPEECH1