Ching-Feng Yeh

dblp:145/6612 · also Ching-feng Yeh · DBLP profile ↗
← Back
33ranked-venue papers
9as first author
12since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 16 · 5 first-author · 5 since 2021Systems, architecture and hardware · 4
YearPublicationVenuePosition
2025 Meta CLIP 2: A Worldwide Scaling Recipe
abstract
Contrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs). Although CLIP is successfully trained on billion-scale image-text pairs from the English world, scaling CLIP's training further to learning from the worldwide web data is still challenging: (1) no curation method is available to handle data points from non-English world; (2) the English performance from existing multilingual CLIP is worse than its English-only counterpart, i.e., "curse of multilinguality" that is common in LLMs. Here, we present Meta CLIP 2, the first recipe training CLIP from scratch on worldwide web-scale image-text pairs. To generalize our findings, we conduct rigorous ablations with minimal changes that are necessary to address the above challenges and present a recipe enabling mutual benefits from English and non-English world data. In zero-shot ImageNet classification, Meta CLIP 2 ViT-H/14 surpasses its English-only counterpart by 0.8% and mSigLIP by 0.7%, and surprisingly sets new state-of-the-art without system-level confounding factors (e.g., translation, bespoke architecture changes) on multilingual benchmarks, such as CVQA with 57.4%, Babel-ImageNet with 50.2% and XM3600 with 64.3% on image-to-text retrieval. Code and model are available at https://github.com/facebookresearch/MetaCLIP.
Yung-Sung Chuang, Ching-Feng Yeh, Kehan Lyu, Ramya Raghavendra, James R. Glass, Lifei Huang, Jason Weston, Luke Zettlemoyer, Xinlei Chen, Zhuang Liu 0003, Saining Xie, Scott Yih, Shang-Wen Li 0001, Hu Xu 0001
NeurIPS4
2024 Altogether: Image Captioning via Re-aligning Alt-text
abstract
Hu Xu, Po-Yao Huang, Xiaoqing Tan, Ching-Feng Yeh, Jacob Kahn, Christine Jou, Gargi Ghosh, Omer Levy, Luke Zettlemoyer, Wen-tau Yih, Shang-Wen Li, Saining Xie, Christoph Feichtenhofer. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Hu Xu 0001, Po-Yao Huang 0001, Xiaoqing Ellen Tan, Ching-Feng Yeh, Jacob Kahn, Christine Jou, Gargi Ghosh, Omer Levy, Luke Zettlemoyer, Scott Yih, Shang-Wen Li 0001, Saining Xie, Christoph Feichtenhofer
EMNLP4
2023 Flap: Fast Language-Audio Pre-Training
abstract
We propose Fast Language-Audio Pre-training (FLAP), a self-supervised approach that efficiently and effectively learns aligned audio and language representations through masking, contrastive learning and reconstruction. For efficiency, FLAP randomly drops audio spectrogram tokens, focusing solely on the remaining ones for self-supervision. Through inter-modal contrastive learning, FLAP learns to align paired audio and text representations in a shared latent space. Notably, FLAP leverages multiple augmented views via masking for intermodal contrast and learns to reconstruct the masked portion of audio tokens. Moreover, FLAP leverages large language models (LLMs) to augment the text inputs, contributing to improved performance. These approaches lead to more robust and informative audio-text representations, enabling FLAP to achieve state-of-the-art (SoTA) performance on audio-text retrieval tasks on AudioCaps (achieving 53.0% R@1) and Clotho (achieving 25.5% R@1).
Ching-Feng Yeh, Po-Yao Huang 0001, Vasu Sharma, Shang-Wen Li 0001, Gargi Ghosh
ASRU1
2023 Continual Learning for On-Device Speech Recognition Using Disentangled Conformers
abstract
Automatic speech recognition research focuses on training and evaluating on static datasets. Yet, as speech models are increasingly deployed on personal devices, such models encounter user-specific distributional shifts. To simulate this real-world scenario, we introduce LibriContinual, a continual learning benchmark for speaker-specific domain adaptation derived from LibriVox audiobooks, with data corresponding to 118 individual speakers and 6 train splits per speaker of different sizes. Additionally, current speech recognition models and continual learning algorithms are not optimized to be compute-efficient. We adapt a general-purpose training algorithm NetAug for ASR and create a novel Conformer variant called the DisConformer (Disentangled Conformer). This algorithm produces ASR models consisting of a frozen ‘core’ network for general-purpose use and several tunable ‘augment’ networks for speaker-specific tuning. Using such models, we propose a novel compute-efficient continual learning algorithm called DisentangledCL. Our experiments show that the DisConformer models significantly outperform base-lines on general ASR i.e. LibriSpeech (15.58% rel. WER on test-other). On speaker-specific LibriContinual they significantly outper-form trainable-parameter-matched baselines (by 20.65% rel. WER on test) and even match fully finetuned baselines in some settings.
Anuj Diwan, Ching-Feng Yeh, Wei-Ning Hsu, Paden Tomasello, Eunsol Choi, David F. Harwath, Abdel-rahman Mohamed
ICASSP2
2022 Superb @ SLT 2022: Challenge on Generalization and Efficiency of Self-Supervised Speech Representation Learning
abstract
We present the SUPERB challenge at SLT 2022, which aims at learning self-supervised speech representation for better performance, generalization, and efficiency. The challenge builds upon the SUPERB benchmark and implements metrics to measure the computation requirements of self-supervised learning (SSL) representation and to evaluate its generalizability and performance across the diverse SUPERB tasks. The SUPERB benchmark provides comprehensive coverage of popular speech processing tasks, from speech and speaker recognition to audio generation and semantic understanding. As SSL has gained interest in the speech community and showed promising outcomes, we envision the challenge to uplevel the impact of SSL techniques by motivating more practical designs of techniques beyond task performance. We summarize the results of 14 submitted models in this paper. We also discuss the main findings from those submissions and the future directions of SSL research.
Tzu-hsun Feng, Shuyan Dong, Ching-Feng Yeh, Shu-Wen Yang, Tzu-Quan Lin, Jiatong Shi, Kai-Wei Chang 0001, Zili Huang, Xuankai Chang, Shinji Watanabe 0001, Abdel-rahman Mohamed, Shang-Wen Li 0001, Hung-yi Lee
SLT3
2021 Emformer: Efficient Memory Transformer Based Acoustic Model for Low Latency Streaming Speech Recognition
abstract
This paper proposes an efficient memory transformer Emformer for low latency streaming speech recognition. In Emformer, the long-range history context is distilled into an augmented memory bank to reduce self-attention’s computation complexity. A cache mechanism saves the computation for the key and value in self-attention for the left context. Emformer applies a parallelized block processing in training to support low latency models. We carry out experiments on benchmark LibriSpeech data. Under average latency of 960 ms, Emformer gets WER 2.50% on test-clean and 5.62% on test-other. Comparing with a strong baseline augmented memory transformer (AM-TRF), Emformer gets 4.6 folds training speedup and 18% relative real-time factor (RTF) reduction in decoding with relative WER reduction 17% on test-clean and 9% on test-other. For a low latency scenario with an average latency of 80 ms, Emformer achieves WER 3.01% on test-clean and 7.09% on test-other. Comparing with the LSTM baseline with the same latency and model size, Emformer gets relative WER reduction 9% and 16% on test-clean and test-other, respectively.
Yangyang Shi, Yongqiang Wang 0005, Chunyang Wu, Ching-Feng Yeh, Julian Chan, Frank Zhang 0001, Mike Seltzer
ICASSP4
2021 Transformer in Action: A Comparative Study of Transformer-Based Acoustic Models for Large Scale Speech Recognition Applications
abstract
Transformer-based acoustic models have shown promising results very recently. In this paper, we summarize the application of transformer and its streamable variant, Emformer based acoustic model [1] for large scale speech recognition applications. We compare the transformer based acoustic models with their LSTM counterparts on industrial scale tasks. Specifically, we compare Emformer with latency-controlled BLSTM (LCBLSTM) on medium latency tasks and LSTM on low latency tasks. On a low latency voice assistant task, Emformer gets 24% to 26% relative word error rate reductions (WERRs). For medium latency scenarios, comparing with LCBLSTM with similar model size and latency, Emformer gets significant WERR across four languages in video captioning datasets with 2-3 times inference real-time factors reduction.
Yongqiang Wang 0005, Yangyang Shi, Frank Zhang 0001, Chunyang Wu, Julian Chan, Ching-Feng Yeh, Alex Xiao
ICASSP6
2021 Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding
abstract
Word Error Rate (WER) has been the predominant metric used to evaluate the performance of automatic speech recognition (ASR) systems. However, WER is sometimes not a good indicator for downstream Natural Language Understanding (NLU) tasks, such as intent recognition, slot filling, and semantic parsing in task-oriented dialog systems. This is because WER takes into consideration only literal correctness instead of semantic correctness, the latter of which is typically more important for these downstream tasks. In this study, we propose a novel Semantic Distance (SemDist) measure as an alternative evaluation metric for ASR systems to address this issue. We define SemDist as the distance between a reference and hypothesis pair in a sentence-level embedding space. To represent the reference and hypothesis as a sentence embedding, we exploit RoBERTa, a state-of-the-art pre-trained deep contextualized language model based on the transformer architecture. We demonstrate the effectiveness of our proposed metric on various downstream tasks, including intent recognition, semantic parsing, and named entity recognition.
Suyoun Kim, Abhinav Arora, Ching-Feng Yeh, Christian Fügen, Ozlem Kalinli, Michael L. Seltzer
Interspeech4
2021 Dynamic Encoder Transducer: A Flexible Solution for Trading Off Accuracy for Latency
abstract
We propose a dynamic encoder transducer (DET) for on-device speech recognition. One DET model scales to multiple devices with different computation capacities without retraining or finetuning. To trading off accuracy and latency, DET assigns different encoders to decode different parts of an utterance. We apply and compare the layer dropout and the collaborative learning for DET training. The layer dropout method that randomly drops out encoder layers in the training phase, can do on-demand layer dropout in decoding. Collaborative learning jointly trains multiple encoders with different depths in one single model. Experiment results on Librispeech and in-house data show that DET provides a flexible accuracy and latency trade-off. Results on Librispeech show that the full-size encoder in DET relatively reduces the word error rate of the same size baseline by over 8%. The lightweight encoder in DET trained with collaborative learning reduces the model size by 25% but still gets similar WER as the full-size baseline. DET gets similar accuracy as a baseline model with better latency on a large in-house data set by assigning a lightweight encoder for the beginning part of one utterance and a full-size encoder for the rest.
Yangyang Shi, Varun Nagaraja, Chunyang Wu, Jay Mahadeokar, Rohit Prabhavalkar, Alex Xiao, Ching-Feng Yeh, Julian Chan, Christian Fügen, Ozlem Kalinli, Michael L. Seltzer
Interspeech8
2021 Alignment Restricted Streaming Recurrent Neural Network Transducer
abstract
There is a growing interest in the speech community in developing Recurrent Neural Network Transducer (RNN-T) models for automatic speech recognition (ASR) applications. RNN-T is trained with a loss function that does not enforce temporal alignment of the training transcripts and audio. As a result, RNN-T models built with uni-directional long short term memory (LSTM) encoders tend to wait for longer spans of input audio, before streaming already decoded ASR tokens. In this work, we propose a modification to the RNN-T loss function and develop Alignment Restricted RNN-T (Ar-RNN-T) models, which utilize audio-text alignment in-formation to guide the loss computation. We compare the proposed method with existing works, such as monotonic RNN-T, on LibriSpeech and in-house datasets. We show that the Ar-RNN-T loss provides a refined control to navigate the trade-offs between the token emission delays and the Word Error Rate (WER). The Ar-RNN-T models also improve downstream applications such as the ASR End-pointing by guaranteeing token emissions within any given range of latency. Moreover, the Ar-RNN-T loss allows for bigger batch sizes and 4 times higher throughput for our LSTM model architecture, enabling faster training and convergence on GPUs.
Jay Mahadeokar, Yuan Shangguan, Gil Keren, Thong Le, Ching-Feng Yeh, Christian Fügen, Michael L. Seltzer
SLT7
2021 Streaming Attention-Based Models with Augmented Memory for End-To-End Speech Recognition
abstract
Attention-based models have been gaining popularity recently for their strong performance demonstrated in fields such as machine translation [1] and automatic speech recognition [2]. One major challenge of attention-based models is the need of access to the full sequence and the quadratically growing computational cost concerning the sequence length. These characteristics pose challenges, especially for low-latency scenarios, where the system is often required to be streaming. In this paper, we build a compact and streaming speech recognition system on top of the end-to-end neural transducer architecture [3] with attention-based modules augmented with convolution [2]. The proposed system equips the end-to-end models with the streaming capability and reduces the large footprint from the streaming attention-based model using augmented memory [4], [5]. On the LibriSpeech [6] dataset, our proposed system achieves word error rates 2.7% on test-clean and 5.8% on test-other, to our best knowledge the lowest among streaming approaches reported so far.
Ching-Feng Yeh, Yongqiang Wang 0005, Yangyang Shi, Chunyang Wu, Frank Zhang 0001, Julian Chan, Michael L. Seltzer
SLT1
2021 Benchmarking LF-MMI, CTC And RNN-T Criteria For Streaming ASR
abstract
In this work, to measure the accuracy and efficiency for a latency-controlled streaming automatic speech recognition (ASR) application, we perform comprehensive evaluations on three popular training criteria: LF-MMI, CTC and RNN-T. In transcribing social media videos of 7 languages with training data 3K - 14K hours, we conduct large-scale controlled experimentation across each criterion using identical datasets and encoder model architecture. We find that RNN-T has consistent wins in ASR accuracy, while CTC models excel at inference efficiency. Moreover, we selectively examine various modeling strategies for different training criteria, including modeling units, encoder architectures, pre-training, etc. Given such large-scale real-world streaming ASR application, to our best knowledge, we present the first comprehensive benchmark on these three widely used training criteria across a great many languages.
Xiaohui Zhang 0007, Frank Zhang 0001, Chunxi Liu, Kjell Schubert, Julian Chan, Pradyot Prakash, Ching-Feng Yeh, Fuchun Peng, Yatharth Saraf, Geoffrey Zweig
SLT8
2020 Aipnet: Generative Adversarial Pre-Training of Accent-Invariant Networks for End-To-End Speech Recognition
abstract
As one of the major sources in speech variability, accents have posed a grand challenge to the robustness of speech recognition systems. In this paper, our goal is to build a unified end-to-end speech recognition system that generalizes well across accents. For this purpose, we propose a novel pre-training framework AIPNetbased on generative adversarial nets (GAN) for accent-invariant representation learning: Accent Invariant Pre-training Networks. We pre-train AIPNetto disentangle accent-invariant and accent-specific characteristics from acoustic features through adversarial training on accented data for which transcriptions are not necessarily available. We further fine-tune AIPNetby connecting the accent-invariant module with an attention-based encoder-decoder model for multi-accent speech recognition. In the experiments, our approach is compared against four baselines including both accent-dependent and accent-independent models. Experimental results on 9 English accents show that the proposed approach outperforms all the baselines by 2.3 ~ 4.5% relative reduction on average WER when transcriptions are available in all accents and by 1.6 ~ 6.1% relative reduction when transcriptions are only available in US accent.
Ching-Feng Yeh, Mahaveer Jain, Michael L. Seltzer
ICASSP3
2020 Weak-Attention Suppression for Transformer Based Speech Recognition
abstract
Transformers, originally proposed for natural language processing (NLP) tasks, have recently achieved great success in automatic speech recognition (ASR). However, adjacent acoustic units (i.e., frames) are highly correlated, and long-distance dependencies between them are weak, unlike text units. It suggests that ASR will likely benefit from sparse and localized attention. In this paper, we propose Weak-Attention Suppression (WAS), a method that dynamically induces sparsity in attention probabilities. We demonstrate that WAS leads to consistent Word Error Rate (WER) improvement over strong transformer baselines. On the widely used LibriSpeech benchmark, our proposed method reduced WER by 10%$ on test-clean and 5% on test-other for streamable transformers, resulting in a new state-of-the-art among streaming models. Further analysis shows that WAS learns to suppress attention of non-critical and redundant continuous acoustic frames, and is more likely to suppress past frames rather than future ones. It indicates the importance of lookahead in attention-based ASR models.
Yangyang Shi, Yongqiang Wang 0005, Chunyang Wu, Christian Fügen, Frank Zhang 0001, Ching-Feng Yeh, Michael L. Seltzer
INTERSPEECH7
2020 Streaming Transformer-Based Acoustic Models Using Self-Attention with Augmented Memory
abstract
Transformer-based acoustic modeling has achieved great success for both hybrid and sequence-to-sequence speech recognition.However, it requires access to the full sequence, and the computational cost grows quadratically with respect to the input sequence length.These factors limit its adoption for streaming applications.In this work, we proposed a novel augmented memory self-attention, which attends on a short segment of the input sequence and a bank of memories.The memory bank stores the embedding information for all the processed segments.On the librispeech benchmark, our proposed method outperforms all the existing streamable transformer methods by a large margin and achieved over 15% relative error reduction, compared with the widely used LC-BLSTM baseline.Our findings are also confirmed on some large internal datasets.
Chunyang Wu, Yongqiang Wang 0005, Yangyang Shi, Ching-Feng Yeh, Frank Zhang 0001
INTERSPEECH4
2018 Domain Adversarial Training for Accented Speech Recognition
abstract
In this paper, we propose a domain adversarial training (DAT) algorithm to alleviate the accented speech recognition problem. In order to reduce the mismatch between labeled source domain data (“standard” accent) and unlabeled target domain data (with heavy accents), we augment the learning objective for a Kaldi TDNN network with a domain adversarial training (DAT) objective to encourage the model to learn accent-invariant features. In experiments with three Mandarin accents, we show that DAT yields up to 7.45% relative character error rate reduction when we do not have transcriptions of the accented speech, compared with the baseline trained on standard accent data only. We also find a benefit from DAT when used in combination with training from automatic transcriptions on the accented data. Furthermore, we find that DAT is superior to multi-task learning for accented speech recognition.
Sining Sun, Ching-Feng Yeh, Mei-Yuh Hwang, Mari Ostendorf, Lei Xie 0001
ICASSP2
2018 Training Augmentation with Adversarial Examples for Robust Speech Recognition
abstract
This paper explores the use of adversarial examples in training speech recognition systems to increase robustness of deep neural network acoustic models.During training, the fast gradient sign method is used to generate adversarial examples augmenting the original training data.Different from conventional data augmentation based on data transformations, the examples are dynamically generated based on current acoustic model parameters.We assess the impact of adversarial data augmentation in experiments on the Aurora-4 and CHiME-4 single-channel tasks, showing improved robustness against noise and channel variation.Further improvement is obtained when combining adversarial examples with teacher/student training, leading to a 23% relative word error rate reduction on Aurora-4.
Sining Sun, Ching-Feng Yeh, Mari Ostendorf, Mei-Yuh Hwang, Lei Xie 0001
INTERSPEECH2
2017 An Efficient Two-Phase ILP-Based Algorithm for Precise CMOS RFIC Layout Generation
abstract
With advancing process technologies and booming Internet of Things markets, millimeter-wave CMOS RFICs have evolved rapidly and been widely applied in recent years. The performance of CMOS RFICs is very sensitive to the chip layout, and a tiny variation of the microstrip length can cause a large impact to the circuit performance. This results in a time-consuming tuning process including much simulation effort for chip design, which becomes the major bottleneck for time to market. This paper introduces a progressive integer-linear-programming-based method consisting of two phases: 1) global layout generation and 2) iterative validation. In the global layout generation phase, we focus on the most critical constraints such as layout planarity and device connection relations to determine the topology of the final design. This provides a basis for constructing the accurate model in the iterative validation phase. The layouts generated by applying our method can satisfy very stringent routing requirements of microstrip lines, including spacing/noncrossing rules, precise length, and bend number minimization, within a given layout area. The resulting RFIC layouts excel in both performance and area with much fewer bends compared with the simulation-tuning based manual layout, while the layout generation time is significantly reduced from weeks to a few minutes.
Tsun-Ming Tseng, Bing Li 0005, Ching-Feng Yeh, Hsiang-Chieh Jhan, Zuo-Min Tsai, Mark Po-Hung Lin, Ulf Schlichtmann
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2016 Novel CMOS RFIC layout generation with concurrent device placement and fixed-length microstrip routing
abstract
With advancing process technologies and booming IoT markets, millimeter-wave CMOS RFICs have been widely developed in recent years. Since the performance of CMOS RFICs is very sensitive to the precision of the layout, precise placement of devices and precisely matched microstrip lengths to given values have been a labor-intensive and time-consuming task, and thus become a major bottleneck for time to market. This paper introduces a progressive integer-linear-programming-based method to generate high-quality RFIC layouts satisfying very stringent routing requirements of microstrip lines, including spacing/non-crossing rules, precise length, and bend number minimization, within a given layout area. The resulting RFIC layouts excel in both performance and area with much fewer bends compared with the simulation-tuning based manual layout, while the layout generation time is significantly reduced from weeks to half an hour.
Tsun-Ming Tseng, Bing Li 0005, Ching-Feng Yeh, Hsiang-Chieh Jhan, Zuo-Min Tsai, Mark Po-Hung Lin, Ulf Schlichtmann
DAC3
2015 Personalized speech recognizer with keyword-based personalized lexicon and language model using word vector representations
Ching-Feng Yeh, Yuan-ming Liou, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH1
2015 An Improved Framework for Recognizing Highly Imbalanced Bilingual Code-Switched Lectures with Cross-Language Acoustic Modeling and Frame-Level Language Identification
abstract
This paper considers the recognition of a widely observed type of bilingual code-switched speech: the speaker speaks primarily the host language (usually his native language), but with a few words or phrases in the guest language (usually his second language) inserted in many utterances of the host language. In this case, not only the languages are switched back and forth within an utterance so the language identification is difficult, but much less data are available for the guest language, which results in poor recognition accuracy for the guest language part. Unit merging approaches on three levels of acoustic modeling (triphone models, HMM states and Gaussians) have been proposed for cross-lingual data sharing for such highly imbalanced bilingual code-switched speech. In this paper, we present an improved overall framework on top of the previously proposed unit merging approaches for recognizing such code-switched speech. This includes unit recovery for reconstructing the identity for units of the two languages after being merged, unit occupancy ranking to offer much more flexible data sharing between units both across languages and within the language based on the accumulated occupancy of the HMM states, and estimation of frame-level language posteriors using blurred posteriorgram features (BPFs) to be used in decoding. We also present a complete set of experimental results comparing all approaches involved for a real-world application scenario under unified conditions, and show very good improvement achieved with the proposed approaches.
Ching-Feng Yeh, Lin-Shan Lee
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 A Novel Analog Physical Synthesis Methodology Integrating Existent Design Expertise
abstract
Analog layout design has been a manual, time-consuming, and error-prone task for decades. To speed up layout design time for a new design, analog layout designers prefer referring to legacy designs and layouts rather than starting from scratch, or thoroughly applying placement and routing tools because legacy layouts contain pretty much design expertise. Motivated by such layout design process, this paper presents the first knowledge-based physical synthesis methodology to generate new layouts by integrating existent design expertise. The proposed approach can automatically analyze legacy design data including circuits, layouts, and constraints, extract matched sub-circuits between new and legacy designs, and generate multiple layouts for the new design by utilizing the quality-approved legacy layouts as much as possible. Experimental results show that the proposed methodology can achieve high layout reusage rate, and hence the designers' layout preference can be successfully reserved.
Po-Hsun Wu, Mark Po-Hung Lin, Tung-Chieh Chen, Ching-Feng Yeh, Xin Li 0001, Tsung-Yi Ho
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2014 Transcribing code-switched bilingual lectures using deep neural networks with unit merging in acoustic modeling
abstract
This paper considers the transcription of the widely observed yet less investigated bilingual code-switched speech: the words or phrases of the guest language are inserted within the utterances of the host language, so the languages are switched back and forth within an utterance, and much less data are available for the guest language. Two approaches utilizing the deep neural network (DNN) were tested and analyzed, including using DNN bottleneck features in HMM/GMM (BF-HMM/GMM) and modeling context-dependent HMM senones by DNN (CD-DNN-HMM). In both cases the unit merging (and recovery) techniques in acoustic modeling were used to handle the data imbalance problem. Improved recognition accuracies were observed with unit merging (and recovery) for the two approaches under different conditions.
Ching-Feng Yeh, Lin-Shan Lee
ICASSP1
2014 Spoken Knowledge Organization by Semantic Structuring and a Prototype Course Lecture System for Personalized Learning
abstract
It takes very long time to go through a complete online course. Without proper background, it is also difficult to understand retrieved spoken paragraphs. This paper therefore presents a new approach of spoken knowledge organization for course lectures for efficient personalized learning. Automatically extracted key terms are taken as the fundamental elements of the semantics of the course. Key term graph constructed by connecting related key terms forms the backbone of the global semantic structure. Audio/video signals are divided into multi-layer temporal structure including paragraphs, sections and chapters, each of which includes a summary as the local semantic structure. The interconnection between semantic structure and temporal structure together with spoken term detection jointly offer to the learners efficient ways to navigate across the course knowledge with personalized learning paths considering their personal interests, available time and background knowledge. A preliminary prototype system has also been successfully developed.
Hung-yi Lee, Sz-Rung Shiang, Ching-Feng Yeh, Yun-Nung Chen, Sheng-yi Kong, Lin-Shan Lee
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Exploring Feasibilities of Symmetry Islands and Monotonic Current Paths in Slicing Trees for Analog Placement
abstract
Although modern analog placement algorithms aimed to minimize area and wirelength while satisfying symmetry, proximity, and other placement constraints, the generated layout does not reflect the circuit performance very well because of the routing-induced parasitics on the critical current/signal paths. To simultaneously consider symmetry, wirelength, area utilization, and current/signal paths during analog placement, this paper explores the feasibilities of symmetry islands and monotonic current paths in slicing trees for analog placement optimization. Experimental results show that the proposed formulation and algorithms can generate much more compact layouts resulting in similar or even better circuit performance compared with the previous work.
Po-Hsun Wu, Mark Po-Hung Lin, Tung-Chieh Chen, Ching-Feng Yeh, Tsung-Yi Ho, Bin-Da Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2013 Speaking rate normalization with lattice-based context-dependent phoneme duration modeling for personalized speech recognizers on mobile devices
Ching-Feng Yeh, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH1
2012 Recognition of highly imbalanced code-mixed bilingual speech with frame-level language detection based on blurred posteriorgram
abstract
In this work, we proposed a new framework for recognition of highly imbalanced code-mixed bilingual speech using an additional frame-level language detector in the conventional recognition system. Blurred posteriorgram features (BPFs) are also proposed to be used in the language detector. The approach was evaluated with real spontaneous lectures offered at National Taiwan University. The highly imbalanced language distribution in code-mixed speech makes the task difficult. Preliminary experimental results showed not only very good performance improvement, but the improvement is complementary to that brought by better acoustic models, whether due to better adaptation approach or increased training data. The code-mixed bilingual speech is frequently used in the daily lives of many people in the globalized world today.
Ching-Feng Yeh, Aaron Heidel, Hung-yi Lee, Lin-Shan Lee
ICASSP1
2011 Bilingual acoustic modeling with state mapping and three-stage adaptation for transcribing unbalanced code-mixed lectures
abstract
This paper presents a bilingual acoustic modeling approach for transcribing Mandarin-English code-mixed lectures with highly unbalanced language distribution. Special terminologies for the content were produced in the guest language of English (about 15%) and embedded in the utterances produced in the host language of Mandarin (about 85%). The code-mixing nature of the target corpus and the very small percentage of the English data made the task difficult. State mapping and merging approaches plus three stages of model adaptation handles the above problem. Significant improvements in recognition accuracy were obtained in the experiment with a real bilingual code-mixed lecture corpus recorded at National Taiwan University. The code-mixing situation considered is actually very natural in the spoken language of the daily lives of many people in the globalized world today.
Ching-Feng Yeh, Liang-Che Sun, Chao-Yu Huang, Lin-Shan Lee
ICASSP1
2011 Spoken Lecture Summarization by Random Walk over a Graph Constructed with Automatically Extracted Key Terms
abstract
This paper proposes an improved approach for spoken lecture summarization, in which random walk is performed on a graph constructed with automatically extracted key terms and proba-bilistic latent semantic analysis (PLSA). Each sentence of the document is represented as a node of the graph and the edge be-tween two nodes is weighted by the topical similarity between the two sentences. The basic idea is that sentences topically similar to more important sentences should be more important. In this way all sentences in the document can be jointly consid-ered more globally rather than individually. Experimental re-sults showed significant improvement in terms of ROUGE eval-uation. Index Terms: summarization, course lecture, probabilistic la-tent semantic analysis (PLSA), random walk, key term
Yun-Nung Chen, Ching-Feng Yeh, Lin-Shan Lee
INTERSPEECH3
2011 Bilingual Acoustic Model Adaptation by Unit Merging on Different Levels and Cross-Level Integration
Ching-Feng Yeh, Chao-Yu Huang, Lin-Shan Lee
INTERSPEECH1
2010 Improved spoken term detection by feature space pseudo-relevance feedback
abstract
Abstract In this paper, we propose an improved approach for spokenterm detection using pseudo-relevance feedback. To remedy theproblem of unmatched acoustic models with respect to spokenutterances produced under different acoustic conditions, whichmay give relatively poor recognition output, we integrate therelevance scores derived from the lattices with the DTW dis-tances derived from the feature space of MFCC parametersor phonetic posteriorgrams. These DTW distances are evalu-ated for a carefully selected set of pseudo-relevant utterances,which obtained from the first-pass returned list given by thesearch engine. The utterances on the first-pass returned list arethen reranked accordingly and finally shown to the user. Veryencouraging, performance improvements were obtained in thepreliminary experiments, especially when the acoustic modelsare poorly matched to the spoken utterances.Index Terms: spoken term detection, pseudo-relevance feed-back 1. Introduction Spoken term detection is to return a list of spoken utterancescontaining the term requested by the user. In many approachesof spoken term detection, the spoken utterances are first recog-nized and transformed into transcriptions or lattices by speechrecognition technologies, and then the search engine looksthrough all the transcriptions or lattices very similar to the text-based information retrieval. In this process much of the in-formation in the acoustic signals may be lost in the stage ofspeech recognition, especially when the acoustic models usedare not well matched to the characteristics of the acoustic sig-nals, which naturally results in degraded recognition accuracyand poor detection performance. This is very common in thescenario of spoken term detection, because the huge quantitiesof spoken utterances available over the Internet are naturallyproduced by many different people under many different acous-tic conditions, it is thus very difficult to train a set of acousticmodels well matched to so many different acoustic conditions.As a result, when the relevance scores such as the posteriorprobabilities of the query term derived from transcriptions orlattices are used to rank the retrieved utterances, it is hard tojudge whether a word hypothesis of the query in the transcrip-tions or lattices is a positive target or a false alarm when therecognition output is unreliable. Although many efficient ap-proaches [1, 2, 3] have been proposed to enhance the detectionperformance due to the relatively poor recognition output, thecompensative information straightly from the feature space isnecessary.In text-based information retrieval, even if the texts to beretrieved include all precise words, it is still difficult to retrieveall documents relevant to the query term because many of themdo not include the very short query term entered by the user.However, because many related terms may co-occur in manyrelated documents, a document containing some words appear-ing in some documents identified to be relevant by the searchengine may have high probability to be relevant, even if it doesnot include the query term. For example, a document includingthe words ”George Bush”, ”US”, ”Middle East” may be relevantto a query term of ”White House”, even if it does not includethe query term of ”White House”. In other words, it is possi-ble to retrieve the relevant documents without the query termsince they are ”similar” to some retrieved relevant documentsin some way. Pseudo-relevance feedback, also known as blindrelevance feedback, is one way to realize the above idea. In thisapproach, it is assumed that the set of documents appearing onthe top of the retrieved document list are relevant (or ”pseudo-relevant”), so documents somehow similar to those ”pseudo-relevant” documents can be retrieved, for example, by expand-ing the query with keywords from those ”pseudo-relevant” doc-uments [4]. Similar idea of pseudo-relevance feedback has beenapplied on spoken term detection [5].In this paper, we try to perform similar pseudo-relevancefeedback for spoken term detection as shown in Figure 1. Theupper half of Figure 1 is the conventional spoken term detec-tion. MFCC features were obtained from all spoken utterancesin the archive, speech recognition produces lattices for the ut-terances, and the retrieved engine selects the utterances basedon the relevance scores evaluated from the lattices with respectto the query Qentered by the user. The approach proposedhere in this paper is shown in the lower half of Figure 1. Thefirst-pass returned list is not shown to the user, but instead a”pseudo-relevant utterance set X
Chia-Ping Chen, Hung-yi Lee, Ching-Feng Yeh, Lin-Shan Lee
INTERSPEECH3
2010 Improved spoken term detection by discriminative training of acoustic models based on user relevance feedback
Hung-yi Lee, Chia-Ping Chen, Ching-Feng Yeh, Lin-Shan Lee
INTERSPEECH3
2010 A framework integrating different relevance feedback scenarios and approaches for spoken term detection
abstract
This paper presents a new framework integrating different relevance feedback scenarios (pseudo relevance feedback and user relevance feedback in short- and long-term context) and different approaches (model- and example-based) in a spoken term detection system, and shows the retrieval performance can be improved step by step. It is found that short-term context user relevance feedback can further improve the retrieval performance after pseudo relevance feedback, regardless of whether the acoustic models have been adapted by matched data or long-term context user relevance feedback or not. Moreover, model-based and example-based methods are shown to be additive when integrated in short-term context user relevance feedback scenario.
Hung-yi Lee, Chia-Ping Chen, Ching-Feng Yeh, Lin-Shan Lee
SLT3