VLDB 2026 Research / reviewers in the wild / expert
Berlin Chen
dblp:c/BerlinChen
· DBLP profile ↗
162ranked-venue papers
27as first author
43since 2021 · last 2026
0000-0003-0693-8932ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 120 · 15 first-author · 34 since 2021Artificial intelligence and machine learning · 105 · 17 first-author · 26 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Explainable Detection of Human-Machine Texts: Methods and Performance Comparison
Cen-Chieh Chen, De-Cian Liu, Berlin Chen, Tao-Hsing Chang, Yao-Ting Sung |
IEA/AIE (1) | 3 |
| 2026 | Efficient Dialect-Aware Modeling and Conditioning for Low-Resource Taiwanese Hakka Speech Processing
An-Ci Peng, Kuan-Tang Huang, Tien-Hong Lo, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
LREC | 6 |
| 2026 | MASA: A Novel Multimodal Foundation Model for L2 Speaking Assessment in Picture-description Scenarios
Bi-Cheng Yan, Fu-An Chao, Hong-Yun Lin, Berlin Chen |
LREC | 4 |
| 2026 | TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
Cheng-Yeh Yang, Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
LREC | 6 |
| 2026 | DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality AlignmentabstractEvaluating text-to-music (TTM) systems remains expensive because music impression (MI) and text alignment (TA) scores rely on human mean opinion scores (MOS). Most automatic MOS estimators are trained with point-wise regression or distributional classification. These objectives do not directly optimize rank-based metrics and provide weak geometric constraints for cross-modal coherence. To address these gaps, we propose DeRA-MOS, a decoupled optimization framework for TTM evaluation. For MI, we introduce a batch-aware listwise ranking loss that models relative order within each mini-batch and better aligns with evaluation based on Spearman's rank correlation coefficient (SRCC). For TA, we introduce a score-anchored modality alignment loss that maps human scores to target audio-text similarity and regularizes the latent space before fusion. By effectively mitigating the point-wise training mismatch and modality drift, experiments on MusicEval demonstrate that our decoupled framework yields substantial improvements in both MI and TA ranking metrics, establishing a robust paradigm for large-scale TTM evaluation. Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
IEEE Signal Process. Lett. | 4 |
| 2025 | Beyond Modality Limitations: A Unified MLLM Approach to Automated Speaking Assessment with Effective Curriculum LearningabstractTraditional Automated Speaking Assessment (ASA) systems exhibit inherent modality limitations: text-based approaches lack acoustic information while audio-based methods miss semantic context. Multimodal Large Language Models (MLLM) offer unprecedented opportunities for comprehensive ASA by simultaneously processing audio and text within unified frameworks. This paper presents a very first systematic study of MLLM for comprehensive ASA, demonstrating the superior performance of MLLM across the aspects of content and language use. However, assessment on the delivery aspect reveals unique challenges, which is deemed to require specialized training strategies. We thus propose Speech-First Multimodal Training (SFMT), leveraging a curriculum learning principle to establish more robust modeling foundations of speech before cross-modal synergetic fusion. A series of experiments on a benchmark dataset show MLLM-based systems can elevate the holistic assessment performance from a PCC value of 0.783 to 0.846. In particular, SFMT excels in the evaluation of the delivery aspect, achieving an absolute accuracy improvement of 4% over conventional training approaches, which also paves a new avenue for ASA. Yu-Hsuan Fang, Tien-Hong Lo, Yao-Ting Sung, Berlin Chen |
ASRU | 4 |
| 2025 | Revealing the Role of Audio Channels in ASR Performance DegradationabstractPre-trained automatic speech recognition (ASR) models have demonstrated strong performance on a variety of tasks. However, their performance can degrade substantially when the input audio comes from different recording channels. While previous studies have demonstrated this phenomenon, it is often attributed to the mismatch between training and testing corpora. This study argues that variations in speech characteristics caused by different recording channels can fundamentally harm ASR performance. To address this limitation, we propose a normalization technique designed to mitigate the impact of channel variation by aligning internal feature representations in the ASR model with those derived from a clean reference channel. This approach significantly improves ASR performance on previously unseen channels and languages, highlighting its ability to generalize across channel and language differences. Kuan-Tang Huang, Hung-Shin Lee, Berlin Chen, Hsin-Min Wang |
ASRU | 4 |
| 2025 | PRIME: Novel Prompting Strategies for Effective Biasing Word Recognition in Contextualized ASRabstractAccurately recognizing domain-specific words remains an arduous challenge facing current automatic speech recognition (ASR) systems. While there are a number of prior arts managing to incorporate domain context through crossattention mechanisms, these methods often add additional components that increase model complexity and focus solely on wordlevel cues, overlooking broader domain-level topic information. To address these limitations, we put forward PRIME, a simple yet effective prompt-tuning method that enriches ASR with domain context using LLM-generated topic descriptions. To mitigate the limited context window of the ASR decoder, we introduce a biasing word retriever that selects the most relevant domain-specific words to construct informative prompts. Notably, PRIME requires no architectural modifications and offers a lightweight, scalable solution for contextualized ASR. A series of experiments counducted on the AISHELL and SlideSpeech benchmark datasets show that PRIME considerably promotes biasing word recognition, outperforming some strong baselines. Yu-Chun Liu, Li-Ting Pai, Yi-Cheng Wang, Bi-Cheng Yan, Hsin-Wei Wang, Chi-Han Lin, Juan-Wei Xu, Berlin Chen |
ASRU | 8 |
| 2025 | An Effective Strategy for Modeling Score Ordinality and Non-uniform Intervals in Automated Speaking AssessmentabstractA recent line of research on automated speaking assessment (ASA) has benefited from self-supervised learning (SSL) representations, which capture rich acoustic and linguistic patterns in non-native speech without underlying assumptions of feature curation. However, speech-based SSL models capture acoustic-related traits but overlook linguistic content, while text-based SSL models rely on ASR output and fail to encode prosodic nuances. Moreover, most prior arts treat proficiency levels as nominal classes, ignoring their ordinal structure and non-uniform intervals between proficiency labels. To address these limitations, we propose an effective ASA approach combining SSL with handcrafted indicator features via a novel modeling paradigm. We further introduce a multi-margin ordinal loss that jointly models both the score ordinality and non-uniform intervals of proficiency labels. Extensive experiments on the TEEMI corpus show that our method consistently outperforms strong baselines and generalizes well to unseen prompts. Tien-Hong Lo, Szu-Yu Chen, Yao-Ting Sung, Berlin Chen |
ASRU | 4 |
| 2025 | QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation SystemsabstractEvaluating audio generation systems, including text-to-music (TTM), text-to-speech (TTS), and text-to-audio (TTA), remains challenging due to the subjective and multi-dimensional nature of human perception. Existing methods treat mean opinion score (MOS) prediction as a regression problem, but standard regression losses overlook the relativity of perceptual judgments. To address this limitation, we introduce QAMRO, a novel Quality-aware Adaptive Margin Ranking Optimization framework that seamlessly integrates regression objectives from different perspectives, aiming to highlight perceptual differences and prioritize accurate ratings. Our framework leverages pre-trained audio-text models such as CLAP and Audiobox-Aesthetics, and is trained exclusively on the official AudioMOS Challenge 2025 dataset. It demonstrates superior alignment with human evaluations across all dimensions, significantly outperforming robust baseline models. Chien-Chun Wang, Kuan-Tang Huang, Cheng-Yeh Yang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
ASRU | 6 |
| 2025 | SincQDR-VAD: A Noise-Robust Voice Activity Detection Framework Leveraging Learnable Filters and Ranking-Aware OptimizationabstractVoice activity detection (VAD) is essential for speech-driven applications, but remains far from perfect in noisy and resource-limited environments. Existing methods often lack robustness to noise, and their frame-wise classification losses are only loosely coupled with the evaluation metric of VAD. To address these challenges, we propose SincQDR-VAD, a compact and robust framework that combines a Sinc-extractor front-end with a novel quadratic disparity ranking loss. The Sinc-extractor uses learnable bandpass filters to capture noise-resistant spectral features, while the ranking loss optimizes the pairwise score order between speech and non-speech frames to improve the area under the receiver operating characteristic curve (AUROC). A series of experiments conducted on representative benchmark datasets show that our framework considerably improves both AUROC and $F_{2}$-Score, while using only $69 \%$ of the parameters compared to prior arts, confirming its efficiency and practical viability. Chien-Chun Wang, En-Lun Yu, Jeih-Weih Hung, Shih-Chieh Huang, Berlin Chen |
ASRU | 5 |
| 2025 | Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech RecognitionabstractWhile pre-trained automatic speech recognition (ASR) systems demonstrate impressive performance on matched domains, their performance often degrades when confronted with channel mismatch stemming from unseen recording environments and conditions. To mitigate this issue, we propose a novel channel-aware data simulation method for robust ASR training. Our method harnesses the synergistic power of channel-extractive techniques and generative adversarial networks (GANs). We first train a channel encoder capable of extracting embeddings from arbitrary audio. On top of this, channel embeddings are extracted using a minimal amount of target-domain data and used to guide a GAN-based speech synthesizer. This synthesizer generates speech that faithfully preserves the phonetic content of the input while mimicking the channel characteristics of the target domain. We evaluate our method on the challenging Hakka Across Taiwan (HAT) and Taiwanese Across Taiwan (TAT) corpora, achieving relative character error rate (CER) reductions of 20.02% and 9.64%, respectively, compared to the baselines. These results highlight the efficacy of our channel-aware data simulation method for bridging the gap between source- and target-domain acoustics. Chien-Chun Wang, Cheng-Kang Chou, Hung-Shin Lee, Berlin Chen, Hsin-Min Wang |
ICASSP | 5 |
| 2025 | ConPCO: Preserving Phoneme Characteristics For Automatic Pronunciation Assessment Leveraging Contrastive Ordinal RegularizationabstractAutomatic pronunciation assessment (APA) manages to evaluate the pronunciation proficiency of a second language (L2) learner in a target language. Existing efforts typically draw on regression models for proficiency score prediction, wherein the models are trained to estimate target values without explicitly accounting for phoneme-awareness in the feature space. In this paper, we propose a contrastive phonemic ordinal regularizer (ConPCO) tailored for regression-based APA models to generate more phoneme-discriminative features while factoring in the ordinal relationships among the regression targets. The proposed ConPCO first aligns the phoneme representations of an APA model and textual embeddings of phonetic transcriptions via contrastive learning. Afterward, the phoneme characteristics are retained by regulating the distances between inter- and intra-phoneme categories in the feature space while allowing for the ordinal relationships among the output targets. We further design and develop a hierarchical APA model to evaluate the effectiveness of our regularizer. A series of experiments conducted on the speechocean762 benchmark dataset suggests the feasibility and effectiveness of our approach in relation to several competitive baselines. Bi-Cheng Yan, Yi-Cheng Wang, Jiun-Ting Li, Meng-Shin Lin, Hsin-Wei Wang, Wei-Cheng Chao, Berlin Chen |
ICASSP | 7 |
| 2025 | Long-Context State-Space Video World ModelsabstractVideo diffusion models have recently shown promise for world modeling through autoregressive frame prediction conditioned on actions. However, they struggle to maintain long-term memory due to the high computational cost associated with processing extended sequences in attention layers. To overcome this limitation, we propose a novel architecture leveraging state-space models (SSMs) to extend temporal memory without compromising computational efficiency. Unlike previous approaches that retrofit SSMs for non-causal vision tasks, our method fully exploits the inherent advantages of SSMs in causal sequence modeling. Central to our design is a block-wise SSM scanning scheme, which strategically trades off spatial consistency for extended temporal memory, combined with dense local attention to ensure coherence between consecutive frames. We evaluate the long-term memory capabilities of our model through spatial retrieval and reasoning tasks over extended horizons. Experiments on Memory Maze and Minecraft datasets demonstrate that our approach surpasses baselines in preserving long-range memory, while maintaining practical inference speeds suitable for interactive applications. Ryan Po, Yotam Nitzan, Richard Zhang 0001, Berlin Chen, Tri Dao, Eli Shechtman, Gordon Wetzstein |
ICCV | 4 |
| 2025 | Flexible VAD-PVAD Transition: A Detachable PVAD Module for Dynamic Encoder RNN VAD
En-Lun Yu, Chien-Chun Wang, Jeih-Weih Hung, Shih-Chieh Huang, Berlin Chen |
INTERSPEECH | 5 |
| 2025 | Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy LossabstractFu-An Chao, Berlin Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Fu-An Chao, Berlin Chen |
NAACL (Long Papers) | 2 |
| 2024 | An Effective Pronunciation Assessment Approach Leveraging Hierarchical Transformers and Pre-training StrategiesabstractBi-Cheng Yan, Jiun-Ting Li, Yi-Cheng Wang, Hsin Wei Wang, Tien-Hong Lo, Yung-Chang Hsu, Wei-Cheng Chao, Berlin Chen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Bi-Cheng Yan, Jiun-Ting Li, Yi-Cheng Wang, Hsin-Wei Wang, Tien-Hong Lo, Yung-Chang Hsu, Wei-Cheng Chao, Berlin Chen |
ACL (1) | 8 |
| 2024 | DANCER: Entity Description Augmented Named Entity Corrector for Automatic Speech RecognitionabstractEnd-to-end automatic speech recognition (E2E ASR) systems often suffer from mistranscription of domain-specific phrases, such as named entities, sometimes leading to catastrophic failures in downstream tasks. A family of fast and lightweight named entity correction (NEC) models for ASR have recently been proposed, which normally build on pho-netic-level edit distance algorithms and have shown impressive NEC performance. However, as the named entity (NE) list grows, the problems of phonetic confusion in the NE list are exacerbated; for example, homophone ambiguities increase substantially. In view of this, we proposed a novel Description Augmented Named entity CorrEctoR (dubbed DANCER), which leverages entity descriptions to provide additional information to facilitate mitigation of phonetic con-fusion for NEC on ASR transcription. To this end, an efficient entity description augmented masked language model (EDA-MLM) comprised of a dense retrieval model is introduced, enabling MLM to adapt swiftly to domain-specific entities for the NEC task. A series of experiments conducted on the AISHELL-1 and Homophone datasets confirm the effectiveness of our modeling approach. DANCER outperforms a strong baseline, the phonetic edit-distance-based NEC model (PED-NEC), by a character error rate (CER) reduction of about 7% relatively on AISHELL-1 for named entities. More notably, when tested on Homophone that contain named entities of high phonetic confusion, DANCER offers a more pronounced CER reduction of 46% relatively over PED-NEC for named entities. The code is available at https://github.com/Amiannn/Dancer. Yi-Cheng Wang, Hsin-Wei Wang, Bi-Cheng Yan, Chi-Han Lin, Berlin Chen |
LREC/COLING | 5 |
| 2024 | What Do Neural Networks Listen to? Exploring the Crucial Bands in Speech Enhancement Using SINC-ConvolutionabstractThis study introduces a reformed Sinc-convolution (Sincconv) framework tailored for the encoder component of deep networks for speech enhancement (SE). The reformed Sinc-conv, based on parametrized sinc functions as band-pass filters, offers notable advantages in terms of training efficiency, filter diversity, and interpretability. The reformed Sinc-conv is evaluated in conjunction with various SE models, showcasing its ability to boost SE performance. Furthermore, the reformed Sincconv provides valuable insights into the specific frequency components that are prioritized in an SE scenario. This opens up a new direction of SE research and improving our knowledge of their operating dynamics. Kuan-Hsun Ho, Jeih-Weih Hung, Berlin Chen |
ICASSP | 3 |
| 2024 | An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder DisentanglementabstractWith the massive developments of end-to-end (E2E) neural networks, recent years have witnessed unprecedented breakthroughs in automatic speech recognition (ASR). However, the code-switching phenomenon remains a major obstacle that hinders ASR from perfection, as the lack of labeled data and the variations between languages often lead to degradation of ASR performance. In this paper, we focus exclusively on improving the acoustic encoder of E2E ASR to tackle the challenge caused by the code-switching phenomenon. Our main contributions are threefold: First, we introduce a novel disentanglement loss to enable the lower-layer of the encoder to capture inter-lingual acoustic information while mitigating linguistic confusion at the higher-layer of the encoder. Second, through comprehensive experiments, we verify that our proposed method outperforms the prior-art methods using pre-trained dual-encoders, meanwhile having access only to the code-switching corpus and consuming half of the parameterization. Third, the apparent differentiation of the encoders’ output features also corroborates the complementarity between the disentanglement loss and the mixture-of-experts (MoE) architecture. Tzu-Ting Yang, Hsin-Wei Wang, Yi-Cheng Wang, Chi-Han Lin, Berlin Chen |
ICASSP | 5 |
| 2024 | TEEMI: a speaking practice tool for L2 English learners
Szu-Yu Chen, Tien-Hong Lo, Yao-Ting Sung, Ching-Yu Tseng, Berlin Chen |
INTERSPEECH | 5 |
| 2024 | EZTalking: English assessment platform for teachers and students
Yu-Sheng Tsao, Yung-Chang Hsu, Jiun-Ting Li, Siang-Hong Weng, Tien-Hong Lo, Berlin Chen |
INTERSPEECH | 6 |
| 2024 | Optimizing Automatic Speech Assessment: W-RankSim Regularization and Hybrid Feature Fusion Strategies
Chung-Wen Wu, Berlin Chen |
INTERSPEECH | 2 |
| 2024 | Speaker Conditional Sinc-Extractor for Personal VAD
En-Lun Yu, Kuan-Hsun Ho, Jeih-Weih Hung, Shih-Chieh Huang, Berlin Chen |
INTERSPEECH | 5 |
| 2024 | Automated Speaking Assessment of Conversation Tests with Novel Graph-Based Modeling on Spoken Response CoherenceabstractAutomated speaking assessment in conversation tests (ASAC) aims to evaluate the overall speaking proficiency of an L2 (second-language) speaker in a setting where an interlocutor interacts with one or more candidates. Although prior ASAC approaches have shown promising performance on their respective datasets, there is still a dearth of research specifically focused on incorporating the coherence of the logical flow within a conversation into the grading model. To address this critical challenge, we propose a hierarchical graph model that aptly incorporates both broad inter-response interactions (e.g., discourse relations) and nuanced semantic information (e.g., semantic words and speaker intents), which is subsequently fused with contextual information for the final prediction. Extensive experimental results on the NICT-JLE benchmark dataset suggest that our proposed modeling approach can yield considerable improvements in prediction accuracy with respect to various assessment metrics, as compared to some strong baselines. This also sheds light on the importance of investigating coherence-related facets of spoken responses in ASAC. Jiun-Ting Li, Bi-Cheng Yan, Tien-Hong Lo, Yi-Cheng Wang, Yung-Chang Hsu, Berlin Chen |
SLT | 6 |
| 2024 | Enhancing Automatic Speech Assessment Leveraging Heterogeneous Features and Soft Labels For Ordinal ClassificationabstractThe general goal of automated speech assessment (ASA) is to provide a consistent and objective evaluation on the spoken language proficiency of an L2 learner or test-taker. In contrast to most previous work that treats ASA as a nominal multi-classification task and thus neglects the sequential nature of proficiency grades, this paper explores the notion of soft labels for use in ASA. In particular, we strive to enhance ASA performance by examining two critical issues: (1) the impact of applying soft labels instead of hard labels in the optimization of ordinal classification for ASA, and (2) the effects of combining self-supervised learning (SSL) with handcrafted indicator features via a novel modeling paradigm. Our results demonstrate that the proposed model can considerably enhance performance compared to existing strong baselines. The improvement is evident not only in the test dataset of seen prompts but also in those of unseen prompts, suggesting the robust generalization and adaptability of our method. Wen-Hsuan Peng, Sally Chen, Berlin Chen |
SLT | 3 |
| 2024 | Effective Noise-Aware Data Simulation For Domain-Adaptive Speech Enhancement Leveraging Dynamic Stochastic PerturbationabstractCross-domain speech enhancement (SE) is often faced with severe challenges due to the scarcity of noise and background information in an unseen target domain, leading to a mismatch between training and test conditions. This study puts forward a novel data simulation method to address this issue, leveraging noise-extractive techniques and generative adversarial networks (GANs) with only limited target noisy speech data. Notably, our method employs a noise encoder to extract noise embeddings from target-domain data. These embeddings aptly guide the generator to synthesize utterances acoustically fitted to the target domain while authentically preserving the phonetic content of the input clean speech. Furthermore, we introduce the notion of dynamic stochastic perturbation, which can inject controlled perturbations into the noise embeddings during inference, thereby enabling the model to generalize well to unseen noise conditions. Experiments on the VoiceBank-DEMAND benchmark dataset demonstrate that our domain-adaptive SE method outperforms an existing strong baseline based on data simulation. Chien-Chun Wang, Hung-Shin Lee, Berlin Chen, Hsin-Min Wang |
SLT | 4 |
| 2024 | An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech RecognitionabstractEnd-to-end (E2E) automatic speech recognition (ASR) models have become standard practice for various commercial applications. However, in real-world scenarios, the long-tailed nature of word distribution often leads E2E ASR models to perform well on common words but fall short in recognizing uncommon ones. Recently, the notion of a contextual adapter (CA) was proposed to infuse external knowledge represented by a context word list into E2E ASR models. Although CA can improve recognition performance on rare words, two crucial data imbalance problems remain. First, when using low-frequency words as context words during training, since these words rarely occur in the utterance, CA becomes prone to overfit on attending to thetoken due to higher-frequency words not being present in the context list. Second, the long-tailed distribution within the context list itself still causes the model to perform poorly on low-frequency context words. In light of this, we explore in-depth the impact of altering the context list to have words with different frequency distributions on model performance, and meanwhile extend CA with a simple yet effective context-balanced learning objective1. A series of experiments conducted on the AISHELL-1 benchmark dataset suggests that using all vocabulary words from the training corpus as the context list and pairing them with our balanced objective yields the best performance, demonstrating a significant reduction in character error rate (CER) by up to 1.21% and a more pronounced 9.44% reduction in the error rate of zero-shot words.1The code is available at: https://github.com/Amiannn/espnet/tree/context_balanced_adapter Yi-Cheng Wang, Li-Ting Pai, Bi-Cheng Yan, Hsin-Wei Wang, Chi-Han Lin, Berlin Chen |
SLT | 6 |
| 2024 | Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior InjectionabstractCode-switching—where multilingual speakers alternately switch between languages during conversations—still poses significant challenges to end-to-end (E2E) automatic speech recognition (ASR) systems due to phenomena of both acoustic and semantic confusion. This issue arises because ASR systems struggle to handle the rapid alternation of languages effectively, which often leads to significant performance degradation. Our main contributions are at least threefold: First, we incorporate language identification (LID) information into several intermediate layers of the encoder, aiming to enrich output embeddings with more detailed language information. Secondly, through the novel application of language boundary alignment loss, the subsequent ASR modules are enabled to more effectively utilize the knowledge of internal language posteriors. Third, we explore the feasibility of using language posteriors to facilitate deep interaction between shared encoder and language-specific encoders. Through comprehensive experiments on the SEAME corpus, we have verified that our proposed method outperforms the prior-art method, disentangle based mixture-of-experts (D-MoE), further enhancing the acuity of the encoder to languages. Tzu-Ting Yang, Hsin-Wei Wang, Yi-Cheng Wang, Berlin Chen |
SLT | 4 |
| 2024 | An Effective Hierarchical Graph Attention Network Modeling Approach for Pronunciation AssessmentabstractAutomatic pronunciation assessment (APA) manages to quantify second language (L2) learners’ pronunciation proficiency in a target language by providing fine-grained feedback with multiple aspect scores (e.g., accuracy, fluency, and completeness) at various linguistic levels (i.e., phone, word, and utterance). Most of the existing efforts commonly follow a parallel modeling framework, which takes a sequence of phone-level pronunciation feature embeddings of a learner's utterance as input and then predicts multiple aspect scores across various linguistic levels. However, these approaches neither take the hierarchy of linguistic units into account nor consider the relatedness among the pronunciation aspects in an explicit manner. In light of this, we put forward an effective modeling approach for APA, termed HierGAT, which is grounded on a hierarchical graph attention network. Our approach facilitates hierarchical modeling of the input utterance as a heterogeneous graph that contains linguistic nodes at various levels of granularity. On top of the tactfully designed hierarchical graph message passing mechanism, intricate interdependencies within and across different linguistic levels are encapsulated and the language hierarchy of an utterance is factored in as well. Furthermore, we also design a novel aspect attention module to encode relatedness among aspects. To our knowledge, we are the first to introduce multiple types of linguistic nodes into graph-based neural networks for APA and perform a comprehensive qualitative analysis to investigate their merits. A series of experiments conducted on the speechocean762 benchmark dataset suggests the feasibility and effectiveness of our approach in relation to several competitive baselines. Bi-Cheng Yan, Berlin Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Preserving Phonemic Distinctions For Ordinal Regression: A Novel Loss Function For Automatic Pronunciation AssessmentabstractAutomatic pronunciation assessment (APA) manages to quantify the pronunciation proficiency of a second language (L2) learner in a language. Prevailing approaches to APA normally leverage neural models trained with a regression loss function, such as the mean-squared error (MSE) loss, for proficiency level prediction. Despite most regression models can effectively capture the ordinality of proficiency levels in the feature space, they are confronted with a primary obstacle that different phoneme categories with the same proficiency level are inevitably forced to be close to each other, retaining less phoneme-discriminative information. On account of this, we devise a phonemic contrast ordinal (PCO) loss for training regression-based APA models, which aims to preserve better phonemic distinctions between phoneme categories meanwhile considering ordinal relationships of the regression target output. Specifically, we introduce a phoneme-distinct regularizer into the MSE loss, which encourages feature representations of different phoneme categories to be far apart while simultaneously pulling closer the representations belonging to the same phoneme category by means of weighted distances. An extensive set of experiments carried out on the speechocean 762 benchmark dataset demonstrate the feasibility and effectiveness of our model in relation to some existing state-of-the-art models. Bi-Cheng Yan, Hsin-Wei Wang, Yi-Cheng Wang, Jiun-Ting Li, Chi-Han Lin, Berlin Chen |
ASRU | 6 |
| 2023 | Effective Graph-Based Modeling of Articulation Traits for Mispronunciation Detection and DiagnosisabstractMispronunciation detection and diagnosis (MDD) manages to pinpoint phone-level erroneous pronunciation segmentations and provide instant and informative diagnostic feedback to L2 (second-language) learners. Among the various modeling paradigms for MDD, dictation-based neural methods have recently become a de facto standard, which identifies pronunciation errors and returns diagnostic feedback at the same time by aligning the recognized phone sequence uttered by an L2 learner to the corresponding canonical phone sequence of a given text prompt. Despite their decent efficacy, dictation-based methods have at least two downsides. First, the dictation process and alignment process are made independent of each other, often resulting in a poor diagnostic feedback. Second, prior knowledge about the articulation traits of the canonical phones in the text prompt is not fully utilized in MDD. On account of this, we propose a novel end-to-end MDD method that can streamline the dictation process and the alignment process in a non-autoregressive manner. In addition, knowledge about phone-level articulation traits are extracted with a graph convolutional network (GCN) to obtain more discriminative phonetic embeddings so as to promote the MDD performance. An extensive set of experiments conducted on the L2-ARCTIC benchmark dataset suggest the feasibility and effectiveness of our approach in relation to competitive baselines. Bi-Cheng Yan, Hsin-Wei Wang, Yi-Cheng Wang, Berlin Chen |
ICASSP | 4 |
| 2023 | A Hierarchical Context-aware Modeling Approach for Multi-aspect and Multi-granular Pronunciation Assessment
Fu-An Chao, Tien-Hong Lo, Tzu-I Wu, Yao-Ting Sung, Berlin Chen |
INTERSPEECH | 5 |
| 2022 | Exploring Non-Autoregressive End-to-End Neural Modeling for English Mispronunciation Detection and DiagnosisabstractEnd-to-end (E2E) neural modeling has emerged as one predominant school of thought to develop computer-assisted pronunciation training (CAPT) systems, showing competitive performance to conventional pronunciation-scoring based methods. However, current E2E neural methods for CAPT are faced with at least two pivotal challenges. On one hand, most of the E2E methods operate in an autoregressive manner with left-to-right beam search to dictate the pronunciations of an L2 learners. This however leads to very slow inference speed, which inevitably hinders their practical use. On the other hand, E2E neural methods are normally data-hungry and meanwhile an insufficient amount of nonnative training data would often reduce their efficacy on mispronunciation detection and diagnosis (MD&D). In response, we put forward a novel MD&D method that leverages non-autoregressive (NAR) E2E neural modeling to dramatically speed up the inference time while maintaining performance in line with the conventional E2E neural methods. In addition, we design and develop a pronunciation modeling network stacked on top of the NAR E2E models of our method to further boost the effectiveness of MD&D. Empirical experiments conducted on the L2-ARCTIC English dataset seems to validate the feasibility of our method, in comparison to some top-of-the-line E2E models and an iconic pronunciation-scoring based method built on a DNN-HMM acoustic model. Hsin-Wei Wang, Bi-Cheng Yan, Hsuan-Sheng Chiu, Yung-Chang Hsu, Berlin Chen |
ICASSP | 5 |
| 2022 | Maximum F1-Score Training for End-to-End Mispronunciation Detection and Diagnosis of L2 English SpeechabstractEnd-to-end (E2E) neural models are increasingly attracting attention as a promising modeling approach for mispronunciation detection and diagnosis (MDD). Typically, these models are trained by optimizing a cross-entropy criterion, which corresponds to improving the log-likelihood of the training data. However, there is a discrepancy between the objectives of model training and the MDD evaluation, since the performance of an MDD model is commonly evaluated in terms of F1-score instead of phone or word error rate (PER/WER). In view of this, we in this paper explore the use of a discriminative objective function for training E2E MDD models, which aims to maximize the expected F1-score directly. A series of experiments conducted on the L2-ARCTIC dataset show that our proposed method can yield considerable performance improvements in relation to some state-of-the-art E2E MDD approaches and the celebrated GOP method. Bi-Cheng Yan, Hsin-Wei Wang, Shao-Wei Fan-Jiang, Fu-An Chao, Berlin Chen |
ICME | 5 |
| 2022 | Effective Cross-Utterance Language Modeling for Conversational Speech RecognitionabstractConversational speech normally is embodied with loose syntactic structures at the utterance level but simultaneously exhibits topical coherence relations across consecutive utterances. Prior work has shown that capturing longer context information with a recurrent neural network or long short-term memory language model (LM) may suffer from the recent bias while excluding the long-range context. In order to capture the long-term semantic interactions among words and across utterances, we put forward disparate conversation history fusion methods for language modeling in automatic speech recognition (ASR) of conversational speech. Furthermore, a novel audio-fusion mechanism is introduced, which manages to fuse and utilize the acoustic embeddings of a current utterance and the semantic content of its corresponding conversation history in a cooperative way. To flesh out our ideas, we frame the ASR N-best hypothesis rescoring task as a prediction problem, leveraging BERT, an iconic pre-trained LM, as the ingredient vehicle to facilitate selection of the oracle hypothesis from a given N-best hypothesis list. Empirical experiments conducted on the AMI benchmark dataset seem to demonstrate the feasibility and efficacy of our methods in relation to some current top-of-line methods. The proposed methods not only achieve significant inference time reduction but also improve the ASR performance for conversational speech. Bi-Cheng Yan, Hsin-Wei Wang, Shih-Hsuan Chiu, Hsuan-Sheng Chiu, Berlin Chen |
IJCNN | 5 |
| 2022 | Adaptive-FSN: Integrating Full-Band Extraction and Adaptive Sub-Band Encoding for Monaural Speech EnhancementabstractAn important more recent thread of speech enhancement work is to utilize fine-grinded local spectral patterns with sub-band processing that complement full-band features nicely. To extend the efficacy of sub-band spectral information, we propose Adaptive-FSN, a fully convolutional real-time speech enhancement framework, to dynamically acquire a sub-band embedding within a wide range of sub-band frequencies. We exploit an adaptive subband encoder to portray sub-band processing that encapsulates a wide range of sub-band units. Then we build this effective sub-band embedding with a Conformer-based structure and multi-view attention. As for the full-band features, we make use of the FullSubNet+ architecture with its full-band extractor to get global spectral information. Finally, a Conformer-based fusion model combines the above information sources to predict the complex ideal ratio mask (cIRM). Experimental results on the VoiceBank-DEMAND benchmark task reveal that this novel framework outperforms FullSubNet+ by promoting the quality of processed utterances and reducing the implementation complexity for faster real-time computation. Yu-sheng Tsao, Kuan-Hsun Ho, Jeih-Weih Hung, Berlin Chen |
SLT | 4 |
| 2022 | Peppanet: Effective Mispronunciation Detection and Diagnosis Leveraging Phonetic, Phonological, and Acoustic CuesabstractMispronunciation detection and diagnosis (MDD) aims to detect erroneous pronunciation segments in an L2 learner's articulation and subsequently provide informative diagnostic feedback. Most existing neural methods follow a dictation-based modeling paradigm that finds out pronunciation errors and returns diagnostic feedback at the same time by aligning the recognized phone sequence uttered by an L2 learner to the corresponding canonical phone sequence of a given text prompt. However, the main downside of these methods is that the dictation process and alignment process are mostly made independent of each other. In view of this, we present a novel end-to-end neural method, dubbed PeppaNet, building on a unified structure that can jointly model the dictation process and the alignment process. The model of our method learns to directly predict the pronunciation correctness of each canonical phone of the text prompt and in turn provides its corresponding diagnostic feedback. In contrast to the conventional dictation-based methods that rely mainly on a free-phone recognition process, PeppaNet makes good use of an effective selective gating mechanism to simultaneously incorporate phonetic, phonological and acoustic cues to generate corrections that are more proper and phonetically related to the canonical pronunciations. Extensive sets of experiments conducted on the L2-ARCTIC benchmark dataset seem to show the merits of our proposed method in comparison to some recent top-of-the-line methods. Bi-Cheng Yan, Hsin-Wei Wang, Berlin Chen |
SLT | 3 |
| 2021 | TENET: A Time-Reversal Enhancement Network for Noise-Robust ASRabstractDue to the unprecedented breakthroughs brought about by deep learning, speech enhancement (SE) techniques have been developed rapidly and play an important role prior to acoustic modeling so as to mitigate noise effects on speech. To increase the perceptual quality of speech, the current state-of-the-art in the realm of SE adopts adversarial training by connecting an objective metric to the discriminator. However, there is no guarantee that optimizing the perceptual quality of speech will necessarily lead to improved automatic speech recognition (ASR) performance. In this study, we present TENET††Inspired by the movie - TENET, Christopher Nolan, 2020., **Some of the enhanced audio samples can be found from https://fuann.github.io/TENET., a novel Time-reversal Enhancement NETwork, which leverages the transformation of an input noisy signal itself, i.e., the time-reversed version, in conjunction with a Siamese network and a complex dual-path Transformer to promote SE performance for noise-robust ASR. Extensive experiments conducted on the Voicebank-DEMAND dataset show that TENET can achieve stellar results compared to a few top-of-the-line methods in terms of both SE and ASR evaluation metrics. To demonstrate the model generalization ability, we further evaluate TENET on the test set of scenarios contaminated with unseen noise, and the results also confirm the superiority of this promising method. Fu-An Chao, Shao-Wei Fan-Jiang, Bi-Cheng Yan, Jeih-Weih Hung, Berlin Chen |
ASRU | 5 |
| 2021 | Towards Robust Mispronunciation Detection and Diagnosis for L2 English Learners with Accent-Modulating MethodsabstractWith the acceleration of globalization, more and more people are willing or required to learn second languages (L2). One of the major remaining challenges facing current mispronunciation and diagnosis (MDD) models for use in computer-assisted pronunciation training (CAPT) is to handle speech from L2 learners with a diverse set of accents. In this paper, we set out to mitigate the adverse effects of accent variety in building an L2 English MDD system with end-to-end (E2E) neural models. To this end, we first propose an effective modeling framework that infuses accent features into an E2E MDD model, thereby making the model more accent-aware. Going a step further, we design and present disparate accent-aware modules to perform accent-aware modulation of acoustic features in a finer-grained manner, so as to enhance the discriminating capability of the resulting MDD model. Extensive sets of experiments conducted on the L2-ARCTIC benchmark dataset show the merits of our MDD model, in comparison to some existing E2E-based strong baselines and the celebrated pronunciation scoring based method. Shao-Wei Fan-Jiang, Bi-Cheng Yan, Tien-Hong Lo, Fu-An Chao, Berlin Chen |
ASRU | 5 |
| 2021 | Cross-Domain Single-Channel Speech Enhancement Model with BI-Projection Fusion Module for Noise-Robust ASRabstractIn recent decades, many studies have suggested that phase information is crucial for speech enhancement (SE), and time-domain single-channel speech enhancement techniques have shown promise in noise suppression and robust automatic speech recognition (ASR). This paper presents a continuation of the above lines of research and explores two effective SE methods that consider phase information in time domain and frequency domain of speech signals, respectively. Going one step further, we put forward a novel cross-domain speech enhancement model and a bi-projection fusion (BPF) mechanism for noise-robust ASR. To evaluate the effectiveness of our proposed method, we conduct an extensive set of experiments on the publicly-available Aishell-1 Mandarin benchmark speech corpus. The evaluation results confirm the superiority of our proposed method in relation to a few current top-of-the-line time-domain and frequency-domain SE methods in both enhancement and ASR evaluation metrics for the test set of scenarios contaminated with seen and unseen noise, respectively. Fu-An Chao, Jeih-Weih Hung, Berlin Chen |
ICME | 3 |
| 2021 | Cross-sentence Neural Language Models for Conversational Speech RecognitionabstractAn important research direction in automatic speech recognition (ASR) has centered around the development of effective methods to rerank the output hypotheses of an ASR system with more sophisticated language models (LMs) for further gains. A current mainstream school of thoughts for ASR N-best hypothesis reranking is to employ a recurrent neural network (RNN)-based LM or its variants, with performance superiority over the conventional n-gram LMs across a range of ASR tasks. In real scenarios such as a long conversation, a sequence of consecutive sentences may jointly contain ample cues of conversation-level information such as topical coherence, lexical entrainment and adjacency pairs, which however remains to be underexplored. In view of this, we first formulate ASR N-best reranking as a prediction problem, putting forward an effective cross-sentence neural LM approach that reranks the ASR N-best hypotheses of an upcoming sentence by taking into consideration the word usage in its precedent sentences. Furthermore, we also explore to extract task-specific global topical information of the cross-sentence history in an unsupervised manner for better ASR performance. Extensive experiments conducted on the AMI conversational benchmark corpus indicate the effectiveness and feasibility of our methods in comparison to several state-of-the-art reranking methods. Shih-Hsuan Chiu, Tien-Hong Lo, Berlin Chen |
IJCNN | 3 |
| 2021 | Innovative Bert-Based Reranking Language Models for Speech RecognitionabstractMore recently, Bidirectional Encoder Representations from Transformers (BERT) was proposed and has achieved impressive success on many natural language processing (NLP) tasks such as question answering and language understanding, due mainly to its effective pre-training then fine-tuning paradigm as well as strong local contextual modeling ability. In view of the above, this paper presents a novel instantiation of the BERT-based contextualized language models (LMs) for use in reranking of N-best hypotheses produced by automatic speech recognition (ASR). To this end, we frame N-best hypothesis reranking with BERT as a prediction problem, which aims to predict the oracle hypothesis that has the lowest word error rate (WER) given the N-best hypotheses (denoted by PBERT). In particular, we also explore to capitalize on task-specific global topic information in an unsupervised manner to assist PBERT in N-best hypothesis reranking (denoted by TPBERT). Extensive experiments conducted on the AMI benchmark corpus demonstrate the effectiveness and feasibility of our methods in comparison to the conventional autoregressive models like the recurrent neural network (RNN) and a recently proposed method that employed BERT to compute pseudo-log-likelihood (PLL) scores for N-best hypothesis reranking. Shih-Hsuan Chiu, Berlin Chen |
SLT | 2 |
| 2020 | Spoken Document Retrieval Leveraging Bert-Based Modeling and Query ReformulationabstractSpoken document retrieval (SDR) has long been deemed a fundamental and important step towards efficient organization of, and access to multimedia associated with spoken content. In this paper, we present a novel study of SDR leveraging the Bidirectional Encoder Representations from Transformers (BERT) model for query and document representations (embeddings), as well as for relevance scoring. BERT has produced extremely promising results for various tasks in natural language understanding, but relatively little research on it is devoted to text information retrieval (IR), let alone SDR. We further tackle one of the critical problems facing SDR, viz. a query is often too short to convey a user's information need, via the process of pseudo-relevance feedback (PRF), showing how information cues induced from PRF can be aptly incorporated into BERT for query expansion. In addition, such query reformulation through PRF also works in conjunction with additional augmentation of lexical features and confidence scores into the document embeddings learned from BERT. The merits of our approach are attested through extensive sets of experiments, which compare it with several classic and cutting-edge (deep learning-based) retrieval approaches. Shao-Wei Fan-Jiang, Tien-Hong Lo, Berlin Chen |
ICASSP | 3 |
| 2020 | The NTNU System at the Interspeech 2020 Non-Native Children's Speech ASR ChallengeabstractThis paper describes the NTNU ASR system participating in the Interspeech 2020 Non-Native Children's Speech ASR Challenge supported by the SIG-CHILD group of ISCA. This ASR shared task is made much more challenging due to the coexisting diversity of non-native and children speaking characteristics. In the setting of closed-track evaluation, all participants were restricted to develop their systems merely based on the speech and text corpora provided by the organizer. To work around this under-resourced issue, we built our ASR system on top of CNN-TDNNF-based acoustic models, meanwhile harnessing the synergistic power of various data augmentation strategies, including both utterance- and word-level speed perturbation and spectrogram augmentation, alongside a simple yet effective data-cleansing approach. All variants of our ASR system employed an RNN-based language model to rescore the first-pass recognition hypotheses, which was trained solely on the text dataset released by the organizer. Our system with the best configuration came out in second place, resulting in a word error rate (WER) of 17.59 %, while those of the top-performing, second runner-up and official baseline systems are 15.67%, 18.71%, 35.09%, respectively. Tien-Hong Lo, Fu-An Chao, Shi-Yan Weng, Berlin Chen |
INTERSPEECH | 4 |
| 2020 | An Effective End-to-End Modeling Approach for Mispronunciation DetectionabstractRecently, end-to-end (E2E) automatic speech recognition (ASR) systems have garnered tremendous attention because of their great success and unified modeling paradigms in comparison to conventional hybrid DNN-HMM ASR systems. Despite the widespread adoption of E2E modeling frameworks on ASR, there still is a dearth of work on investigating the E2E frameworks for use in computer-assisted pronunciation learning (CAPT), particularly for Mispronunciation detection (MD). In response, we first present a novel use of hybrid CTCAttention approach to the MD task, taking advantage of the strengths of both CTC and the attention-based model meanwhile getting around the need for phone-level forced alignment. Second, we perform input augmentation with text prompt information to make the resulting E2E model more tailored for the MD task. On the other hand, we adopt two MD decision methods so as to better cooperate with the proposed framework: 1) decision-making based on a recognition confidence measure or 2) simply based on speech recognition results. A series of Mandarin MD experiments demonstrate that our approach not only simplifies the processing pipeline of existing hybrid DNN-HMM systems but also brings about systematic and substantial performance improvements. Furthermore, input augmentation with text prompts seems to hold excellent promise for the E2E-based MD approach. Tien-Hong Lo, Shi-Yan Weng, Hsiu-Jui Chang, Berlin Chen |
INTERSPEECH | 4 |
| 2020 | An End-to-End Mispronunciation Detection System for L2 English Speech Leveraging Novel Anti-Phone ModelingabstractMispronunciation detection and diagnosis (MDD) is a core component of computer-assisted pronunciation training (CAPT).Most of the existing MDD approaches focus on dealing with categorical errors (viz.one canonical phone is substituted by another one, aside from those mispronunciations caused by deletions or insertions).However, accurate detection and diagnosis of non-categorial or distortion errors (viz.approximating L2 phones with L1 (first-language) phones, or erroneous pronunciations in between) still seems out of reach.In view of this, we propose to conduct MDD with a novel endto-end automatic speech recognition (E2E-based ASR) approach.In particular, we expand the original L2 phone set with their corresponding anti-phone set, making the E2E-based MDD approach have a better capability to take in both categorical and non-categorial mispronunciations, aiming to provide better mispronunciation detection and diagnosis feedback.Furthermore, a novel transfer-learning paradigm is devised to obtain the initial model estimate of the E2E-based MDD system without resource to any phonological rules.Extensive sets of experimental results on the L2-ARCTIC dataset show that our best system can outperform the existing E2E baseline system and pronunciation scoring based method (GOP) in terms of the F1-score, by 11.05% and 27.71%, respectively.Index Bi-Cheng Yan, Meng-Che Wu, Hsiao-Tsung Hung, Berlin Chen |
INTERSPEECH | 4 |
| 2020 | Enhanced Language Modeling with Proximity and Sentence Relatedness Information for Extractive Broadcast News SummarizationabstractThe primary task of extractive summarization is to automatically select a set of representative sentences from a text or spoken document that can concisely express the most important theme of the original document. Recently, language modeling (LM) has been proven to be a promising modeling framework for performing this task in an unsupervised manner. However, there still remain three fundamental challenges facing the existing LM-based methods, which we set out to tackle in this article. The first one is how to construct a more accurate sentence model in this framework without resorting to external sources of information. The second is how to take into account sentence-level structural relationships, in addition to word-level information within a document, for important sentence selection. The last one is how to exploit the proximity cues inherent in sentences to obtain a more accurate estimation of respective sentence models. Specifically, for the first and second challenges, we explore a novel, principled approach that generates overlapped clusters to extract sentence relatedness information from the document to be summarized, which can be used not only to enhance the estimation of various sentence models but also to render sentence-level structural relationships within the document, leading to better summarization effectiveness. For the third challenge, we investigate several formulations of proximity cues for use in sentence modeling involved in the LM-based summarization framework, free of the strict bag-of-words assumption. Furthermore, we also present various ensemble methods that seamlessly integrate proximity and sentence relatedness information into sentence modeling. Extensive experiments conducted on a Mandarin broadcast news summarization task show that such integration of proximity and sentence relatedness information is indeed beneficial for speech summarization. Our proposed summarization methods can significantly boost the performance of an LM-based strong baseline (e.g., with a maximum ROUGE-2 improvement of 26.7% relative) and also outperform several state-of-the-art unsupervised methods compared in the article. Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2020 | Multi-Instrument Automatic Music Transcription With Self-Attention-Based Instance SegmentationabstractMulti-instrument automatic music transcription (AMT) is a critical but less investigated problem in the field of music information retrieval (MIR). With all the difficulties faced by traditional AMT research, multi-instrument AMT needs further investigation on high-level music semantic modeling, efficient training methods for multiple attributes, and a clear problem scenario for system performance evaluation. In this article, we propose a multi-instrument AMT method, with signal processing techniques specifying pitch saliency, novel deep learning techniques, and concepts partly inspired by multi-object recognition, instance segmentation, and image-to-image translation in computer vision. The proposed method is flexible for all the sub-tasks in multi-instrument AMT, including multi-instrument note tracking, a task that has rarely been investigated before. State-of-the-art performance is also reported in the sub-task of multi-pitch streaming. Yu-Te Wu, Berlin Chen, Li Su 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language ModelingabstractAlex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Jan Hula, Patrick Xia 0002, Raghavendra Pappagari, Tom McCoy 0001, Roma Patel, Najoung Kim, Ian Tenney, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman |
ACL (1) | 12 |
| 2019 | Enhanced Bert-Based Ranking Models for Spoken Document RetrievalabstractThe Bidirectional Encoder Representations from Transformers (BERT) model has recently achieved record-breaking success on many natural language processing (NLP) tasks such as question answering and language understanding. However, relatively little work has been done on ad-hoc information retrieval (IR), especially for spoken document retrieval (SDR). This paper adopts and extends BERT for SDR, while its contributions are at least three-fold. First, we augment BERT with extra language features such as unigram and inverse document frequency (IDF) statistics to make it more applicable to SDR. Second, we also explore the incorporation of confidence scores into document representations to see if they could help alleviate the negative effects resulting from imperfect automatic speech recognition (ASR). Third, we conduct a comprehensive set of experiments to compare our BERT-based ranking methods with other state-of-the-art ones and investigate the synergy effect of them as well. Hsiao-Yun Lin, Tien-Hong Lo, Berlin Chen |
ASRU | 3 |
| 2019 | A Hierarchical Neural Summarization Framework for Spoken DocumentsabstractExtractive text or speech summarization seeks to select indicative sentences from a source document and assemble them together to form a succinct summary, so as to help people to browse and understand the main theme of the document efficiently. A more recent trend is towards developing supervised deep learning based methods for extractive summarization. This paper extends and contextualizes this line of research for spoken document summarization, while its contributions are at least three-fold. First, we propose a neural summarization framework with the flexibility to incorporate extra acoustic/prosodic and lexical features, for which the ROUGE evaluation metric is embedded into the training objective function and can be optimized with reinforcement learning. Second, disparate ways to integrate acoustic features into this framework are investigated. Third, the utility of our proposed summarization methods and several widely-used state-of-the-art ones are extensively compared and evaluated. A series of empirical experiments seem to demonstrate the effectiveness of our summarization methods. Tzu-En Liu, Shih-Hung Liu, Berlin Chen |
ICASSP | 3 |
| 2019 | Polyphonic Music Transcription with Semantic SegmentationabstractThe multi-instrument transcription task refers to joint recognition of instrument and pitch of every event in polyphonic music signals generated by one or more classes of music instruments. In this paper, we leverage multi-object semantic segmentation techniques to solve this problem. We design a time-frequency representation, which has multiple channels to jointly represent the harmonic structure and pitch saliency of a pitch activation. The transcription task therefore becomes a pixel-wise multi-task classification problem including pitch activity detection and instrument recognition. Experiments on both single- and multi-instrument data verify the competitiveness of the proposed method. Yu-Te Wu, Berlin Chen, Li Su 0004 |
ICASSP | 2 |
| 2019 | What do you learn from context? Probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia 0002, Berlin Chen, Adam Poliak, Tom McCoy 0001, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das 0001, Ellie Pavlick |
ICLR (Poster) | 3 |
| 2019 | Integrating LSA-based hierarchical conceptual space and machine learning methods for leveling the readability of domain-specific textsabstractAbstract Text readability assessment is a challenging interdisciplinary endeavor with rich practical implications. It has long drawn the attention of researchers internationally, and the readability models since developed have been widely applied to various fields. Previous readability models have only made use of linguistic features employed for general text analysis and have not been sufficiently accurate when used to gauge domain-specific texts. In view of this, this study proposes a latent-semantic-analysis (LSA)-constructed hierarchical conceptual space that can be used to train a readability model to accurately assess domain-specific texts. Compared with a baseline reference using a traditional model, the new model improves by 13.88% to achieve 68.98% of accuracy when leveling social science texts, and by 24.61% to achieve 73.96% of accuracy when assessing natural science texts. We then combine the readability features developed for the current study with general linguistic features, and the accuracy of leveling social science texts improves by an even higher degree of 31.58% to achieve 86.68%, and that of natural science texts by 26.56% to achieve 75.91%. These results indicate that the readability features developed in this study can be used both to train a readability model for leveling domain-specific texts and also in combination with the more common linguistic features to enhance the efficacy of the model. Future research can expand the generalizability of the model by assessing texts from different fields and grade levels using the proposed method, thus enhancing the practical applications of this new method. Hou-Chiang Tseng, Berlin Chen, Tao-Hsing Chang, Yao-Ting Sung |
Nat. Lang. Eng. | 2 |
| 2018 | Essence Vector-Based Query Modeling for Spoken Document RetrievalabstractSpoken document retrieval (SDR) has become a prominently required application since unprecedented volumes of multimedia data along with speech have become available in our daily life. As far as we are aware, there has been relatively less work in launching unsupervised paragraph embedding methods and investigating the effectiveness of these methods on the SDR task. This paper first presents a novel paragraph embedding method, named the essence vector (EV) model, which aims at inferring a representation for a given paragraph by encapsulating the most representative information from the paragraph and excluding the general background information at the same time. On top of the EV model, we develop three query language modeling mechanisms to improve the retrieval performance. A series of empirical SDR experiments conducted on two benchmark collections demonstrate the good efficacy of the proposed framework, compared to several existing strong baseline systems. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang |
ICASSP | 3 |
| 2018 | Automatic Music Transcription Leveraging Generalized Cepstral Features and Deep LearningabstractSpectral features are limited in modeling musical signals with multiple concurrent pitches due to the challenge to suppress the interference of the harmonic peaks from one pitch to another. In this paper, we show that using multiple features represented in both the frequency and time domains with deep learning modeling can reduce such interference. These features are derived systematically from conventional pitch detection functions that relate to one another through the discrete Fourier transform and a nonlinear scaling function. Neural networks modeled with these features outperform state-of-the-art methods while using less training data. Yu-Te Wu, Berlin Chen, Li Su 0004 |
ICASSP | 2 |
| 2018 | An Information Distillation Framework for Extractive SummarizationabstractIn the context of natural language processing, representation learning has emerged as a newly active research subject because of its excellent performance in many applications. Learning representations of words is a pioneering study in this school of research. However, paragraph (or sentence and document) embedding learning is more suitable/reasonable for some realistic tasks such as document summarization. Nevertheless, classic paragraph embedding methods infer the representation of a given paragraph by considering all of the words occurring in the paragraph. Consequently, those stop or function words that occur frequently may mislead the embedding learning process to produce a misty paragraph representation. Motivated by these observations, our major contributions in this paper are threefold. First, we propose a novel unsupervised paragraph embedding method, named the essence vector (EV) model, which aims at not only distilling the most representative information from a paragraph but also excluding the general background information to produce a more informative low-dimensional vector representation for the paragraph of interest. Second, in view of the increasing importance of spoken content processing, an extension of the EV model, named the denoising essence vector (D-EV) model, is proposed. The D-EV model not only inherits the advantages of the EV model but also can infer a more robust representation for a given spoken paragraph against imperfect speech recognition. Third, a new summarization framework, which can take both relevance and redundancy information into account simultaneously, is also introduced. We evaluate the proposed embedding methods (i.e., EV and D-EV) and the summarization framework on two benchmark summarization corpora. The experimental results demonstrate the effectiveness and applicability of the proposed framework in relation to several well-practiced and state-of-the-art summarization methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Neural relevance-aware query modeling for spoken document retrievalabstractSpoken document retrieval (SDR) is becoming a much-needed application due to that unprecedented volumes of audio-visual media have been made available in our daily life. As far as we are aware, most of the wide variety of SDR methods mainly focus on exploring robust indexing and effective retrieval methods to quantify the relevance degree between a pair of query and document. However, similar to information retrieval (IR), a fundamental challenge facing SDR is that a query is usually too short to convey a user's information need, such that a retrieval system cannot always achieve prospective efficacy when with the existing retrieval methods. In order to further boost retrieval performance, several studies turn their attention to reformulating the original query by leveraging an online pseudo-relevance feedback (PRF) process, which often comes at the price of taking significant time. Motivated by these observations, this paper presents a novel extension of the general line of SDR research and its contribution is at least two-fold. First, building on neural network-based techniques, we put forward a neural relevance-aware query modeling (NRM) framework, which is designed to not only infer a discriminative query language model automatically for a given query, but also get around the time-consuming PRF process. Second, the utility of the methods instantiated from our proposed framework and several widely-used retrieval methods are extensively analyzed and compared on a standard SDR task, which suggests the superiority of our methods. Tien-Hong Lo, Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen |
ASRU | 5 |
| 2017 | A locality-preserving essence vector modeling framework for spoken document retrievalabstractBecause unprecedented volumes of multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research area in the past decades. Recently, representation learning has emerged as an active research topic in many machine learning applications owing largely to its excellent performance. In the context of natural language processing, the pioneering work can date back to the word embedding methods. However, learning of paragraph (or sentence and document) representations is more reasonable and suitable for some tasks, such as information retrieval and document summarization. Nevertheless, as far as we are aware, there is relatively less work focusing on launching paragraph embedding methods into SDR. Motivated by these observations, this paper proposes a novel paragraph embedding method, named the locality-preserving essence vector (LPEV) model. LPEV is designed with consideration to two aspects. First, the model aims at not only distilling the most representative information from a paragraph but also getting rid of the general background information. Second, inspired by the local invariance perspective, which is a celebrated principle used in manifold learning techniques, LPEV also manages to preserve semantic locality in the learned low-dimensional embedding space for producing more informative and discriminative vector representations of paragraphs. On top of the proposed framework, a series of empirical SDR experiments conducted on the TDT-2 (Topic Detection and Tracking) collection demonstrate the good efficacy of our SDR methods as compared to existing strong baselines. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang |
ICASSP | 3 |
| 2017 | Leveraging manifold learning for extractive broadcast news summarizationabstractExtractive speech summarization is intended to produce a condensed version of the original spoken document by selecting a few salient sentences from the document and concatenate them together to form a summary. In this paper, we study a novel use of manifold learning techniques for extractive speech summarization. Manifold learning has experienced a surge of research interest in various domains concerned with dimensionality reduction and data representation recently, but has so far been largely under-explored in extractive text or speech summarization. Our contributions in this paper are at least twofold. First, we explore the use of several manifold learning algorithms to capture the latent semantic information of sentences for enhanced extractive speech summarization, including isometric feature mapping (ISOMAP), locally linear embedding (LLE) and Laplacian eigenmap. Second, the merits of our proposed summarization methods and several widely-used methods are extensively analyzed and compared. The empirical results demonstrate the effectiveness of our unsupervised summarization methods, in relation to several state-of-the-art methods. In particular, a synergy of the manifold learning based methods and state-of-the-art methods, such as the integer linear programming (ILP) method, contributes to further gains in summarization performance. Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen, Hsin-Min Wang, Wen-Lian Hsu |
ICASSP | 3 |
| 2017 | Enhancing feature modulation spectra with dictionary learning approaches for robust speech recognitionabstractNoise robustness has long garnered much interest from researchers and practitioners of the automatic speech recognition (ASR) community due to its paramount importance to the success of ASR systems. This paper presents a novel approach to improving the noise robustness of speech features, building on top of the dictionary learning paradigm. To this end, we employ the K-SVD method and its variants to create sparse representations with respect to a common set of basis spectral vectors that captures the intrinsic temporal structure inherent in the modulation spectra of clean training speech features. The enhanced modulation spectra of speech features, constructed by mapping the original modulation spectra into the space spanned by these representative basis vectors, can better carry noise-resistant acoustic characteristics. In addition, considering the nonnegative property of the modulation spectrum amplitudes, we utilize the nonnegative K-SVD method, in combination with the nonnegative sparse coding method, to generate more noise-robust speech features. All experiments were conducted and verified using the standard Aurora-2 database and task. The empirical results show that the proposed dictionary learning based approach can provide significant average word error reductions when being integrated with either a GMM-HMM or a DNN-HMM based ASR system. Bi-Cheng Yan, Chin-Hong Shih, Shih-Hung Liu, Berlin Chen |
ICME | 4 |
| 2017 | Exploring the Use of Significant Words Language Modeling for Spoken Document Retrieval
Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen |
INTERSPEECH | 4 |
| 2017 | Exploring Low-Dimensional Structures of Modulation Spectra for Robust Speech Recognition
Bi-Cheng Yan, Chin-Hong Shih, Shih-Hung Liu, Berlin Chen |
INTERSPEECH | 4 |
| 2017 | Discriminative Autoencoders for Acoustic Modeling
Ming-Han Yang, Hung-Shin Lee, Yu-Ding Lu, Kuan-Yu Chen 0002, Yu Tsao 0001, Berlin Chen, Hsin-Min Wang |
INTERSPEECH | 6 |
| 2017 | A Position-Aware Language Modeling Framework for Extractive Broadcast News Speech SummarizationabstractExtractive summarization, a process that automatically picks exemplary sentences from a text (or spoken) document with the goal of concisely conveying key information therein, has seen a surge of attention from scholars and practitioners recently. Using a language modeling (LM) approach for sentence selection has been proven effective for performing unsupervised extractive summarization. However, one of the major difficulties facing the LM approach is to model sentences and estimate their parameters more accurately for each text (or spoken) document. We extend this line of research and make the following contributions in this work. First, we propose a position-aware language modeling framework using various granularities of position-specific information to better estimate the sentence models involved in the summarization process. Second, we explore disparate ways to integrate the positional cues into relevance models through a pseudo-relevance feedback procedure. Third, we extensively evaluate various models originated from our proposed framework and several well-established unsupervised methods. Empirical evaluation conducted on a broadcast news summarization task further demonstrates performance merits of the proposed summarization methods. Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2016 | Learning to Distill: The Essence Vector Modeling FrameworkabstractIn the context of natural language processing, representation learning has emerged as a newly active research subject because of its excellent performance in many applications. Learning representations of words is a pioneering study in this school of research. However, paragraph (or sentence and document) embedding learning is more suitable/reasonable for some tasks, such as sentiment classification and document summarization. Nevertheless, as far as we are aware, there is only a dearth of research focusing on launching unsupervised paragraph embedding methods. Classic paragraph embedding methods infer the representation of a given paragraph by considering all of the words occurring in the paragraph. Consequently, those stop or function words that occur frequently may mislead the embedding learning process to produce a misty paragraph representation. Motivated by these observations, our major contributions are twofold. First, we propose a novel unsupervised paragraph embedding method, named the essence vector (EV) model, which aims at not only distilling the most representative information from a paragraph but also excluding the general background information to produce a more informative low-dimensional vector representation for the paragraph. We evaluate the proposed EV model on benchmark sentiment classification and multi-document summarization tasks. The experimental results demonstrate the effectiveness and applicability of the proposed embedding method. Second, in view of the increasing importance of spoken content processing, an extension of the EV model, named the denoising essence vector (D-EV) model, is proposed. The D-EV model not only inherits the advantages of the EV model but also can infer a more robust representation for a given spoken paragraph against imperfect speech recognition. The utility of the D-EV model is evaluated on a spoken document summarization task, confirming the effectiveness of the proposed embedding method in relation to several well-practiced and state-of-the-art summarization methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang |
COLING | 3 |
| 2016 | Improved spoken document summarization with coverage modeling techniquesabstractExtractive summarization aims at selecting a set of indicative sentences from a source document as a summary that can express the major theme of the document. A general consensus on extractive summarization is that both relevance and coverage are critical issues to address. The existing methods designed to model coverage can be characterized by either reducing redundancy or increasing diversity in the summary. Maximal margin relevance (MMR) is a widely-cited method since it takes both relevance and redundancy into account when generating a summary for a given document. In addition to MMR, there is only a dearth of research concentrating on reducing redundancy or increasing diversity for the spoken document summarization task, as far as we are aware. Motivated by these observations, two major contributions are presented in this paper. First, in contrast to MMR, which considers coverage by reducing redundancy, we propose two novel coverage-based methods, which directly increase diversity. With the proposed methods, a set of representative sentences, which not only are relevant to the given document but also cover most of the important sub-themes of the document, can be selected automatically. Second, we make a step forward to plug in several document/sentence representation methods into the proposed framework to further enhance the summarization performance. A series of empirical evaluations demonstrate the effectiveness of our proposed methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang |
ICASSP | 3 |
| 2016 | Mispronunciation Detection Leveraging Maximum Performance Criterion Training of Acoustic Models and Decision Functions
Yao-Chi Hsu, Ming-Han Yang, Hsiao-Tsung Hung, Berlin Chen |
INTERSPEECH | 4 |
| 2016 | Exploring Word Mover's Distance and Semantic-Aware Embedding Techniques for Extractive Broadcast News Summarization
Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
INTERSPEECH | 4 |
| 2016 | Novel Word Embedding and Translation-based Language Modeling for Extractive Speech SummarizationabstractWord embedding methods revolve around learning continuous distributed vector representations of words with neural networks, which can capture semantic and/or syntactic cues, and in turn be used to induce similarity measures among words, sentences and documents in context. Celebrated methods can be categorized as prediction-based and count-based methods according to the training objectives and model architectures. Their pros and cons have been extensively analyzed and evaluated in recent studies, but there is relatively less work continuing the line of research to develop an enhanced learning method that brings together the advantages of the two model families. In addition, the interpretation of the learned word representations still remains somewhat opaque. Motivated by the observations and considering the pressing need, this paper presents a novel method for learning the word representations, which not only inherits the advantages of classic word embedding methods but also offers a clearer and more rigorous interpretation of the learned word representations. Built upon the proposed word embedding method, we further formulate a translation-based language modeling framework for the extractive speech summarization task. A series of empirical evaluations demonstrate the effectiveness of the proposed word representation learning and language modeling techniques in extractive speech summarization. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Hsin-Hsi Chen |
ACM Multimedia | 3 |
| 2016 | Extractive speech summarization leveraging convolutional neural network techniquesabstractExtractive text or speech summarization endeavors to select representative sentences from a source document and assemble them into a concise summary, so as to help people to browse and assimilate the main theme of the document efficiently. The recent past has seen a surge of interest in developing deep learning- or deep neural network-based supervised methods for extractive text summarization. This paper presents a continuation of this line of research for speech summarization and its contributions are three-fold. First, we exploit an effective framework that integrates two convolutional neural networks (CNNs) and a multilayer perceptron (MLP) for summary sentence selection. Specifically, CNNs encode a given document-sentence pair into two discriminative vector embeddings separately, while MLP in turn takes the two embeddings of a document-sentence pair and their similarity measure as the input to induce a ranking score for each sentence. Second, the input of MLP is augmented by a rich set of prosodic and lexical features apart from those derived from CNNs. Third, the utility of our proposed summarization methods and several widely-used methods are extensively analyzed and compared. The empirical results seem to demonstrate the effectiveness of our summarization method in relation to several state-of-the-art methods. Chun-I Tsai, Hsiao-Tsung Hung, Kuan-Yu Chen 0002, Berlin Chen |
SLT | 4 |
| 2016 | Exploring the use of unsupervised query modeling techniques for speech recognition and summarization
Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Hsin-Hsi Chen |
Speech Commun. | 3 |
| 2016 | Robust Speech Recognition via Enhancing the Complex-Valued Acoustic Spectrum in Modulation DomainabstractThe purpose of this paper is to develop a novel speech feature extraction framework for independently compensating the real and imaginary acoustic spectra of speech signals in the modulation domain with the techniques of histogram equalization (HEQ) and non-negative matrix factorization (NMF). By doing so, we can enhance not only the magnitude but also the phase components of the acoustic spectra, thereby creating noise-robust speech features. More specifically, the proposed framework makes the following three major contributions: First, via either of the HEQ and NMF operations, the long-term cross-frame correlation among the acoustic spectra at the same frequency can be captured to compensate for the spectral distortion caused by noise. Second, the noise effect can be handled in a high acoustic frequency resolution. Finally, the distortion dwelt in the acoustic spectra can be more extensively mitigated due to the independent processes for the respective real and imaginary parts. The evaluation experiments were carried out on the Aurora-2 and Aurora-4 benchmark tasks, and the corresponding results suggest that our proposed methods can achieve performance competitive to or better than many widely used noise robustness methods, including the well-known advanced front-end (AFE) extraction scheme, in speech recognition. Jeih-Weih Hung, Hsin-Ju Hsieh, Berlin Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Incorporating paragraph embeddings and density peaks clustering for spoken document summarizationabstractRepresentation learning has emerged as a newly active research subject in many machine learning applications because of its excellent performance. As an instantiation, word embedding has been widely used in the natural language processing area. However, as far as we are aware, there are relatively few studies investigating paragraph embedding methods in extractive text or speech summarization. Extractive summarization aims at selecting a set of indicative sentences from a source document to express the most important theme of the document. There is a general consensus that relevance and redundancy are both critical issues for users in a realistic summarization scenario. However, most of the existing methods focus on determining only the relevance degree between sentences and a given document, while the redundancy degree is calculated by a post-processing step. Based on these observations, three contributions are proposed in this paper. First, we comprehensively compare the word and paragraph embedding methods for spoken document summarization. Next, we propose a novel summarization framework which can take both relevance and redundancy information into account simultaneously. Consequently, a set of representative sentences can be automatically selected through a one-pass process. Third, we further plug in paragraph embedding methods into the proposed framework to enhance the summarization performance. Experimental results demonstrate the effectiveness of our proposed methods, compared to existing state-of-the-art methods. Kuan-Yu Chen 0002, Kai-Wun Shih, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang |
ASRU | 4 |
| 2015 | I-vector based language modeling for query representationabstractSince more and more multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research subject in the past two decades. Following the research tendency, many efforts have been devoted towards developing indexing and modeling techniques for representing spoken documents, but only few have been made on improving query formulation for better representing users' information needs. The i-vector based language modeling (IVLM) framework, stemming from the state-of-the-art i-vector framework for language identification and speaker recognition, has been proposed and formulated to represent documents in SDR with good promise recently. However, a major challenge of using IVLM for query modeling is that a query usually consists of only a few words; thus, it is hard to learn a reliable representation accordingly. In this paper, we focus our attention on query reformulation and propose three novel methods on top of IVLM to more accurately represent users' information needs. In addition, we also explore the use of multi-levels of index features, including word- and subword-level units, to work in concert with the proposed methods. A series of empirical SDR experiments conducted on the TDT-2 (Topic Detection and Tracking) collection demonstrate the good effectiveness of our proposed methods as compared to existing state-of-the-art methods. Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen, Hsin-Hsi Chen |
ICASSP | 3 |
| 2015 | Leveraging word embeddings for spoken document summarizationabstractOwing to the rapidly growing multimedia content available on the Internet, extractive spoken document summarization, with the purpose of automatically selecting a set of representative sentences from a spoken document to concisely express the most important theme of the document, has been an active area of research and experimentation. On the other hand, word embedding has emerged as a newly favorite research subject because of its excellent performance in many natural language processing (NLP)-related tasks. However, as far as we are aware, there are relatively few studies investigating its use in extractive text or speech summarization. A common thread of leveraging word embeddings in the summarization process is to represent the document (or sentence) by averaging the word embeddings of the words occurring in the document (or sentence). Then, intuitively, the cosine similarity measure can be employed to determine the relevance degree between a pair of representations. Beyond the continued efforts made to improve the representation of words, this paper focuses on building novel and efficient ranking models based on the general word embedding methods for extractive speech summarization. Experimental results demonstrate the effectiveness of our proposed methods, compared to existing state-of-the-art methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Hsin-Min Wang, Berlin Chen, Hsin-Hsi Chen |
INTERSPEECH | 4 |
| 2015 | Positional language modeling for extractive broadcast news speech summarizationabstractExtractive summarization, with the intention of automatically selecting a set of representative sentences from a text (or spoken) document so as to concisely express the most important theme of the document, has been an active area of experimentation and development.A recent trend of research is to employ the language modeling (LM) approach for important sentence selection, which has proven to be effective for performing extractive summarization in an unsupervised fashion.However, one of the major challenges facing the LM approach is how to formulate the sentence models and estimate their parameters more accurately for each text (or spoken) document to be summarized.This paper extends this line of research and its contributions are three-fold.First, we propose a positional language modeling framework using different granularities of position-specific information to better estimate the sentence models involved in summarization.Second, we also explore to integrate the positional cues into relevance modeling through a pseudo-relevance feedback procedure.Third, the utilities of the various methods originated from our proposed framework and several well-established unsupervised methods are analyzed and compared extensively.Empirical evaluations conducted on a broadcast news summarization task seem to demonstrate the performance merits of our summarization methods. Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
INTERSPEECH | 3 |
| 2015 | Histogram equalization of contextual statistics of speech features for robust speech recognition
Hsin-Ju Hsieh, Berlin Chen, Jeih-Weih Hung |
Multim. Tools Appl. | 2 |
| 2015 | Extractive Broadcast News Summarization Leveraging Recurrent Neural Network Language Modeling TechniquesabstractExtractive text or speech summarization manages to select a set of salient sentences from an original document and concatenate them to form a summary, enabling users to better browse through and understand the content of the document. A recent stream of research on extractive summarization is to employ the language modeling (LM) approach for important sentence selection, which has proven to be effective for performing speech summarization in an unsupervised fashion. However, one of the major challenges facing the LM approach is how to formulate the sentence models and accurately estimate their parameters for each sentence in the document to be summarized. In view of this, our work in this paper explores a novel use of recurrent neural network language modeling (RNNLM) framework for extractive broadcast news summarization. On top of such a framework, the deduced sentence models are able to render not only word usage cues but also long-span structural information of word co-occurrence relationships within broadcast news documents, getting around the need for the strict bag-of-words assumption. Furthermore, different model complexities and combinations are extensively analyzed and compared. Experimental results demonstrate the performance merits of our summarization methods when compared to several well-studied state-of-the-art unsupervised methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Ea-Ee Jan, Wen-Lian Hsu, Hsin-Hsi Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Combining Relevance Language Modeling and Clarity Measure for Extractive Speech SummarizationabstractExtractive speech summarization, which purports to select an indicative set of sentences from a spoken document so as to succinctly represent the most important aspects of the document, has garnered much research over the years. In this paper, we cast extractive speech summarization as an ad-hoc information retrieval (IR) problem and investigate various language modeling (LM) methods for important sentence selection. The main contributions of this paper are four-fold. First, we explore a novel sentence modeling paradigm built on top of the notion of relevance, where the relationship between a candidate summary sentence and a spoken document to be summarized is discovered through different granularities of context for relevance modeling. Second, not only lexical but also topical cues inherent in the spoken document are exploited for sentence modeling. Third, we propose a novel clarity measure for use in important sentence selection, which can help quantify the thematic specificity of each individual sentence that is deemed to be a crucial indicator orthogonal to the relevance measure provided by the LM-based methods. Fourth, in an attempt to lessen summarization performance degradation caused by imperfect speech recognition, we investigate making use of different levels of index features for LM-based sentence modeling, including words, subword-level units, and their combination. Experiments on broadcast news summarization seem to demonstrate the performance merits of our methods when compared to several existing well-developed and/or state-of-the-art methods. Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Leveraging Effective Query Modeling Techniques for Speech Recognition and SummarizationabstractStatistical language modeling (LM) that purports to quantify the acceptability of a given piece of text has long been an interesting yet challenging research area.In particular, language modeling for information retrieval (IR) has enjoyed remarkable empirical success; one emerging stream of the LM approach for IR is to employ the pseudo-relevance feedback process to enhance the representation of an input query so as to improve retrieval effectiveness.This paper presents a continuation of such a general line of research and the main contribution is threefold.First, we propose a principled framework which can unify the relationships among several widely-used query modeling formulations.Second, on top of the successfully developed framework, we propose an extended query modeling formulation by incorporating critical query-specific information cues to guide the model estimation.Third, we further adopt and formalize such a framework to the speech recognition and summarization tasks.A series of empirical experiments reveal the feasibility of such an LM framework and the performance merits of the deduced models on these two tasks. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Ea-Ee Jan, Hsin-Min Wang, Wen-Lian Hsu, Hsin-Hsi Chen |
EMNLP | 3 |
| 2014 | I-vector based language modeling for spoken document retrievalabstractSince more and more multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research subject in the past two decades. The i-vector based framework has been proposed and introduced to language identification (LID) and speaker recognition (SR) tasks recently. The major contribution of the i-vector framework is to reduce a series of acoustic feature vectors of a speech utterance to a low-dimensional vector representation, and then numbers of well-developed postprocessing techniques (such as probabilistic linear discriminative analysis, PLDA) can be readily and effectively used. However, to our best knowledge, there is no research up to date on applying the i-vector framework for SDR or information retrieval (IR). In this paper, we make a step forward to formulate an i-vector based language modeling (IVLM) framework for SDR. Furthermore, we evaluate the proposed IVLM framework with both inductive and transductive learning strategies. We also exploit multi-levels of index features, including word- and subword-level units, in concert with the proposed framework. The results of SDR experiments conducted on the TDT-2 (Topic Detection and Tracking) collection demonstrate the performance merits of our proposed framework when compared to several existing approaches. Kuan-Yu Chen 0002, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen, Hsin-Hsi Chen |
ICASSP | 4 |
| 2014 | Effective pseudo-relevance feedback for language modeling in extractive speech summarizationabstractExtractive speech summarization, aiming to automatically select an indicative set of sentences from a spoken document so as to concisely represent the most important aspects of the document, has become an active area for research and experimentation. An emerging stream of work is to employ the language modeling (LM) framework along with the Kullback-Leibler divergence measure for extractive speech summarization, which can perform important sentence selection in an unsupervised manner and has shown preliminary success. This paper presents a continuation of such a general line of research and its main contribution is two-fold. First, by virtue of pseudo-relevance feedback, we explore several effective sentence modeling formulations to enhance the sentence models involved in the LM-based summarization framework. Second, the utilities of our summarization methods and several widely-used methods are analyzed and compared extensively, which demonstrates the effectiveness of our methods. Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
ICASSP | 4 |
| 2014 | A recurrent neural network language modeling framework for extractive speech summarizationabstractExtractive speech summarization, with the purpose of automatically selecting a set of representative sentences from a spoken document so as to concisely express the most important theme of the document, has been an active area of research and development. A recent school of thought is to employ the language modeling (LM) approach for important sentence selection, which has proven to be effective for performing speech summarization in an unsupervised fashion. However, one of the major challenges facing the LM approach is how to formulate the sentence models and accurately estimate their parameters for each spoken document to be summarized. This paper presents a continuation of this general line of research and its contribution is two-fold. First, we propose a novel and effective recurrent neural network language modeling (RNNLM) framework for speech summarization, on top of which the deduced sentence models are able to render not only word usage cues but also long-span structural information of word co-occurrence relationships within spoken documents, getting around the need for the strict bag-of-words assumption. Second, the utilities of the method originated from our proposed framework and several widely-used unsupervised methods are analyzed and compared extensively. A series of experiments conducted on a broadcast news summarization task seem to demonstrate the performance merits of our summarization method when compared to several state-of-the-art existing unsupervised methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Wen-Lian Hsu, Hsin-Hsi Chen |
ICME | 3 |
| 2014 | Effective modulation spectrum factorization for robust speech recognition
Yu-Chen Kao, Berlin Chen |
INTERSPEECH | 3 |
| 2014 | Enhanced language modeling for extractive speech summarization with sentence relatedness informationabstractExtractive summarization is intended to automatically select a set of representative sentences from a text or spoken document that can concisely express the most important topics of the document. Language modeling (LM) has been proven to be a promising framework for performing extractive summarization in an unsupervised manner. However, there remain two fundamental challenges facing existing LM-based methods. One is how to construct sentence models involved in the LM framework more accurately without resorting to external information sources. The other is how to additionally take into account the sentence-level structural relationships embedded in a document for important sentence selection. To address these two challenges, in this paper we explore a novel approach that generates overlapped clusters to extract sentence relatedness information from the document to be summarized, which can be used not only to enhance the estimation of various sentence models but also to allow for the sentence-level structural relationships for better summarization performance. Further, the utilities of our proposed methods and several state-of-the-art unsupervised methods are analyzed and compared extensively. A series of experiments conducted on a Mandarin broadcast news summarization task demonstrate the effectiveness and viability of our method. Index Terms: speech summarization, language modeling, clustering, relevance, sentence relatedness Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
INTERSPEECH | 4 |
| 2014 | Leveraging topical and positional cues for language modeling in speech recognition
Hsuan-Sheng Chiu, Kuan-Yu Chen 0002, Berlin Chen |
Multim. Tools Appl. | 3 |
| 2013 | Effective pseudo-relevance feedback for language modeling in speech recognitionabstractA part and parcel of any automatic speech recognition (ASR) system is language modeling (LM), which helps to constrain the acoustic analysis, guide the search through multiple candidate word strings, and quantify the acceptability of the final output hypothesis given an input utterance. Despite the fact that the n-gram model remains the predominant one, a number of novel and ingenious LM methods have been developed to complement or be used in place of the n-gram model. A more recent line of research is to leverage information cues gleaned from pseudo-relevance feedback (PRF) to derive an utterance-regularized language model for complementing the n-gram model. This paper presents a continuation of this general line of research and its main contribution is two-fold. First, we explore an alternative and more efficient formulation to construct such an utterance-regularized language model for ASR. Second, the utilities of various utterance-regularized language models are analyzed and compared extensively. Empirical experiments on a large vocabulary continuous speech recognition (LVCSR) task demonstrate that our proposed language models can offer substantial improvements over the baseline n-gram system, and achieve performance competitive to, or better than, some state-of-the-art language models. Berlin Chen, Yi-Wen Chen, Kuan-Yu Chen 0002, Ea-Ee Jan |
ASRU | 1 |
| 2013 | Effective pseudo-relevance feedback for spoken document retrievalabstractWith the exponential proliferation of multimedia associated with spoken documents, research on spoken document retrieval (SDR) has emerged and attracted much attention in the past two decades. Apart from much effort devoted to developing robust indexing and modeling techniques for representing spoken documents, a recent line of thought targets at the improvement of query modeling for better reflecting the user's information need. Pseudo-relevance feedback is by far the most commonly-used paradigm for query reformulation, which assumes that a small amount of top-ranked feedback documents obtained from the initial round of retrieval are relevant and can be utilized for this purpose. Nevertheless, simply taking all of the top-ranked feedback documents obtained from the initial retrieval for query modeling (reformulation) does not always work well, especially when the top-ranked documents contain much redundant or non-relevant information. In the view of this, we explore in this paper an interesting problem of how to effectively glean useful cues from the top-ranked documents so as to achieve more accurate query modeling. To do this, different kinds of information cues are considered and integrated into the process of feedback document selection so as to improve query effectiveness. Experiments conducted on the TDT (Topic Detection and Tracking) task show the advantages of our retrieval methods for SDR. Yi-Wen Chen, Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen |
ICASSP | 4 |
| 2013 | Weighted matrix factorization for spoken document retrievalabstractSince more and more multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research subject in the past two decades. Recently, topic models have been successfully used in SDR as well as general information retrieval (IR). These models fall into two categories: probabilistic topic models (PTM) and non-probabilistic topic models (NPTM). One major difference between PTM and NPTM is that the former only takes the words occurring in a document into account, whereas the latter, such as latent semantic analysis (LSA), explicitly models all the words in the vocabulary (including both occurring and non-occurring words). We believe that the non-occurring words can provide additional information that is also useful for SDR. However, to our best knowledge, there is a dearth of work investigating the effectiveness of the non-occurring words for SDR and IR. In order to make effective use of those non-occurring words of documents for semantic analysis, we propose a weighted matrix factorization (WMF) framework, in which the impact of the non-occurring words on the semantic analysis can be modulated properly. The results of SDR experiments conducted on the TDT-2 (Topic Detection and Tracking) collection highlight the performance merits of our proposed framework when compared to several existing topic models. Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen, Hsin-Hsi Chen |
ICASSP | 3 |
| 2013 | Sentence modeling for extractive speech summarizationabstractExtractive speech summarization, aiming to select an indicative set of sentences from a spoken document so as to concisely represent the most important aspects of the document, has emerged as an attractive area of research and experimentation. A recent school of thought is to employ the language modeling (LM) framework along with the Kullback-Leibler (KL) divergence measure for important sentence selection, which has shown preliminary promise for extractive speech summarization. Our work in this paper continues this general line of research in two significant aspects. First, we explore a novel sentence modeling approach built on top of the notion of relevance, where the relationship between a candidate summary sentence and the spoken document to be summarized is discovered through various granularities of context for relevance modeling. Second, not only lexical but also topical cues inherent in the spoken document are exploited for sentence modeling. Experiments on broadcast news summarization seem to demonstrate the performance merits of our methods when compared to several existing methods. Berlin Chen, Hao-Chin Chang, Kuan-Yu Chen 0002 |
ICME | 1 |
| 2013 | Incorporating proximity information for relevance language modeling in speech recognition
Yi-Wen Chen, Bo-Han Hao, Kuan-Yu Chen 0002, Berlin Chen |
INTERSPEECH | 4 |
| 2013 | Histogram equalization of real and imaginary modulation spectra for noise-robust speech recognition
Hsin-Ju Hsieh, Berlin Chen, Jeih-Weih Hung |
INTERSPEECH | 2 |
| 2013 | Distribution-based feature normalization for robust speech recognition leveraging context and dynamics cues
Yu-Chen Kao, Berlin Chen |
INTERSPEECH | 2 |
| 2013 | Leveraging relevance cues for language modeling in speech recognition
Berlin Chen, Kuan-Yu Chen 0002 |
Inf. Process. Manag. | 1 |
| 2013 | Extractive speech summarization using evaluation metric-related training criteria
Berlin Chen, Shih-Hsiang Lin, Yu-Mei Chang, Jia-Wen Liu |
Inf. Process. Manag. | 1 |
| 2012 | Constructing effective ranking models for speech summarizationabstractSpeech summarization, facilitating users to better browse through and understand speech information (especially, spoken documents), has become an active area of intensive research recently. Many of the existing machine-learning approaches to speech summarization cast important sentence selection as a two-class classification problem and have shown empirical success for a wide array of summarization tasks. One common deficiency of these approaches is that the corresponding learning criteria are loosely related to the final evaluation metric. To cater for this problem, we present a novel probabilistic framework to learn the summarization models, building on top of the Bayes decision theory. Two effective training criteria, viz. maximum relevance estimation (MRE) and minimum ranking loss estimation (MRLE), deduced from such a framework are introduced to characterize the pair-wise preference relationships between spoken sentences. Experiments on a broadcast news speech summarization task exhibit the performance merits of our summarization methods when compared to existing methods. Yueng-Tien Lo, Shih-Hsiang Lin, Berlin Chen |
ICASSP | 3 |
| 2012 | Word Relevance Modeling for Speech RecognitionabstractLanguage models for speech recognition tend to be brittle across domains, since their performance is vulnerable to changes in the genre or topic of the text on which they are trained. A number of adaptation methods, discovering either lexical co-occurrence or topic cues, have been developed to mitigate this problem with varying degrees of success. Among them, a more recent thread of work is the relevance modeling approach, which has shown promise to capture the lexical co-occurrence relationship between the entire search history and an upcoming word. However, a potential downside to such an approach is the need of resorting to a retrieval procedure to obtain relevance information; this is usually complex and time-consuming for practical applications. In this paper, we propose a word relevance modeling framework, which introduces a novel use of relevance information for dynamic language model adaptation in speech recognition. It not only inherits the merits of several existing techniques but also provides a flexible yet systematic way to render the lexical, topical, and proximity relationships between the search history and the upcoming word. Experiments on large vocabulary continuous speech recognition demonstrate the performance merits of the methods instantiated from this framework when compared to several existing methods. Kuan-Yu Chen 0002, Hao-Chin Chang, Berlin Chen, Hsin-Min Wang |
INTERSPEECH | 3 |
| 2012 | Exploring Joint Equalization of Spatial-Temporal Contextual Statistics of Speech Features for Robust Speech Recognition
Hsin-Ju Hsieh, Jeih-Weih Hung, Berlin Chen |
INTERSPEECH | 3 |
| 2012 | Spoken Document Retrieval With Unsupervised Query Modeling TechniquesabstractEver-increasing amounts of publicly available multimedia associated with speech information have motivated spoken document retrieval (SDR) to be an active area of intensive research in the speech processing community. Much work has been dedicated to developing elaborate indexing and modeling techniques for representing spoken documents, but only little to improving query formulations for better representing the information needs of users. The latter is critical to the success of a SDR system. In view of this, we present in this paper a novel use of a relevance language modeling framework for SDR. It not only inherits the merits of several existing techniques but also provides a principled way to render the lexical and topical relationships between a query and a spoken document. We further explore various ways to glean both relevance and non-relevance cues from the spoken document collection so as to enhance query modeling in an unsupervised fashion. In addition, we also investigate representing the query and documents with different granularities of index features to work in conjunction with the various relevance and/or non-relevance cues. Empirical evaluations performed on the TDT (Topic Detection and Tracking) collections reveal that the methods derived from our modeling framework hold good promise for SDR and are very competitive with existing retrieval methods. Berlin Chen, Kuan-Yu Chen 0002, Pei-Ning Chen, Yi-Wen Chen |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | A Risk-Aware Modeling Framework for Speech SummarizationabstractExtractive speech summarization attempts to select a representative set of sentences from a spoken document so as to succinctly describe the main theme of the original document. In this paper, we adapt the notion of risk minimization for extractive speech summarization by formulating the selection of summary sentences as a decision-making problem. To this end, we develop several selection strategies and modeling paradigms that can leverage supervised and unsupervised summarization models to inherit their individual merits as well as to overcome their inherent limitations. On top of that, various component models are introduced, providing a principled way to render the redundancy and coherence relationships among sentences and between sentences and the whole document, respectively. A series of experiments on speech summarization seem to demonstrate that the methods deduced from our summarization framework are very competitive with existing summarization methods. Berlin Chen, Shih-Hsiang Lin |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Query modeling for spoken document retrievalabstractSpoken document retrieval (SDR) has recently become a more interesting research avenue due to increasing volumes of publicly available multimedia associated with speech information. Many efforts have been devoted to developing elaborate indexing and modeling techniques for representing spoken documents, but only few to improving query formulations for better representing the users' information needs. In view of this, we recently presented a language modeling framework exploring a novel use of relevance information cues for improving query effectiveness. Our work in this paper continues this general line of research in two main aspects. We further explore various ways to glean both relevance and non-relevance cues from the spoken document collection so as to enhance query modeling in an unsupervised fashion. Furthermore, we also investigate representing the query and documents with different granularities of index features to work in conjunction with the various relevance and/or non-relevance cues. Experiments conducted on the TDT (Topic Detection and Tracking) SDR task demonstrate the performance merits of the methods instantiated from our retrieval framework when compared to other existing retrieval methods. Berlin Chen, Pei-Ning Chen, Kuan-Yu Chen 0002 |
ASRU | 1 |
| 2011 | Relevance language modeling for speech recognitionabstractLanguage models for speech recognition tend to be brittle across domains, since their performance is vulnerable to changes in the genre or topic of the text on which they are trained. A number of adaptation methods, exploring either lexical co-occurrence or topic cues, have been developed to mitigate this problem with varying degrees of success. In this paper, we study a novel use of relevance information for dynamic language model adaptation in speech recognition. It not only inherits the merits of several existing techniques but also provides a flexible but systematic way to render the lexical and topical relationships between a search history and an upcoming word. Empirical results on large vocabulary continuous speech recognition show that the methods deduced from our framework represent promising alternatives to the other existing language model adaptation methods compared in this paper. Kuan-Yu Chen 0002, Berlin Chen |
ICASSP | 2 |
| 2011 | Handling verbose queries for spoken document retrievalabstractQuery-by-example information retrieval provides users a flexible but efficient way to accurately describe their information needs. The query exemplars are usually long and in the form of either a partial or even a full document. However, they may contain extraneous terms that would have potential negative impacts on the retrieval performance. In order to alleviate those negative impacts, we propose a novel term-based query reduction mechanism so as to improve the informativeness of verbose query exemplars. We also explore the notion of term discrimination power to select a salient subset of query terms automatically. Experiments on the TDT Chinese collection show that the proposed approach is indeed effective and promising. Shih-Hsiang Lin, Ea-Ee Jan, Berlin Chen |
ICASSP | 3 |
| 2011 | Discriminative language modeling for speech recognition with relevance informationabstractDiscriminative language modeling (DLM) attempts to improve speech recognition performance by reranking the recognition hypotheses output from a baseline system. Most of the existing DLM methods assume that the reranking task can be treated as a linear discrimination problem and all testing utterances share the same parameter vector for reranking of hypotheses. However, the latter assumption sometimes results in a trained DLM model with weak generalizability and unsatisfactory performance. In view of this problem, we hence propose a relevance-based DLM (RDLM) framework that can efficiently infer the DLM model parameters of each testing utterance on-the-fly for better recognition performance. The structures and characteristics of the RDLM framework are extensively investigated, while the performance is thoroughly analyzed and verified by comparison with the existing DLM methods. Berlin Chen, Jia-Wen Liu |
ICME | 1 |
| 2011 | An Effective and Robust Framework for Transliteration Exploration
Ea-Ee Jan, Niyu Ge, Shih-Hsiang Lin, Berlin Chen |
IJCNLP | 4 |
| 2011 | Leveraging Relevance Cues for Improved Spoken Document Retrieval
Pei-Ning Chen, Kuan-Yu Chen 0002, Berlin Chen |
INTERSPEECH | 3 |
| 2011 | Robust speech recognition using spatial-temporal feature distribution characteristics
Berlin Chen, Wei-Hau Chen, Shih-Hsiang Lin, Wen-Yi Chu |
Pattern Recognit. Lett. | 1 |
| 2011 | Leveraging Kullback-Leibler Divergence Measures and Information-Rich Cues for Speech SummarizationabstractImperfect speech recognition often leads to degraded performance when exploiting conventional text-based methods for speech summarization. To alleviate this problem, this paper investigates various ways to robustly represent the recognition hypotheses of spoken documents beyond the top scoring ones. Moreover, a summarization framework, building on the Kullback-Leibler (KL) divergence measure and exploring both the relevance and topical information cues of spoken documents and sentences, is presented to work with such robust representations. Experiments on broadcast news speech summarization tasks appear to demonstrate the utility of the presented approaches. Shih-Hsiang Lin, Yao-Ming Yeh, Berlin Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2010 | A Risk Minimization Framework for Extractive Speech Summarization
Shih-Hsiang Lin, Berlin Chen |
ACL | 2 |
| 2010 | Latent topic modeling of word vicinity information for speech recognitionabstractTopic language models, mostly revolving around the discovery of “word-document” co-occurrence dependence, have attracted significant attention and shown good performance in a wide variety of speech recognition tasks over the years. In this paper, a new topic language model, named word vicinity model (WVM), is proposed to explore the co-occurrence relationship between words, as well as the long-span latent topical information for language model adaptation. A search history is modeled as a composite WVM model for predicting a decoded word. The underlying characteristics and different kinds of model structures are extensively investigated, while the performance of WVM is thoroughly analyzed and verified by comparison with a few existing topic language models. Moreover, we also present a new modeling approach to our recently proposed word topic model (WTM), and design an efficient way to simultaneously extract “word-document” and “word-word” co-occurrence characteristics through the sharing of the same set of latent topics. Experiments on broadcast news transcription seem to demonstrate the utility of the presented models. Kuan-Yu Chen 0002, Hsuan-Sheng Chiu, Berlin Chen |
ICASSP | 3 |
| 2010 | Leveraging evaluation metric-related training criteria for speech summarizationabstractMany of the existing machine-learning approaches to speech summarization cast important sentence selection as a two-class classification problem and have shown empirical success for a wide variety of summarization tasks. However, the imbalanced-data problem sometimes results in a trained speech summarizer with unsatisfactory performance. On the other hand, training the summarizer by improving the associated classification accuracy does not always lead to better summarization evaluation performance. In view of such phenomena, we hence investigate two different training criteria to alleviate the negative effects caused by them, as well as to boost the summarizer's performance. One is to learn the classification capability of a summarizer on the basis of the pair-wise ordering information of sentences in a training document according to a degree of importance. The other is to train the summarizer by directly maximizing the associated evaluation score. Experimental results on the broadcast news summarization task show that these two training criteria can give substantial improvements over the baseline SVM summarization system. Shih-Hsiang Lin, Yu-Mei Chang, Jia-Wen Liu, Berlin Chen |
ICASSP | 4 |
| 2010 | A Discriminative and Heteroscedastic Linear Feature Transformation for Multiclass ClassificationabstractThis paper presents a novel discriminative feature transformation, named full-rank generalized likelihood ratio discriminant analysis (fGLRDA), on the grounds of the likelihood ratio test (LRT). fGLRDA attempts to seek a feature space, which is linearly isomorphic to the original n-dimensional feature space and is characterized by a full-rank (n×n) transformation matrix, under the assumption that all the class-discrimination information resides in a d-dimensional subspace (d <; n), through making the most confusing situation, described by the null hypothesis, as unlikely as possible to happen without the homoscedastic assumption on class distributions. Our experimental results demonstrate that fGLRDA can yield moderate performance improvements over other existing methods, such as linear discriminant analysis (LDA) for the speaker identification task. Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
ICPR | 3 |
| 2010 | Extractive speech summarization - from the view of decision theory
Shih-Hsiang Lin, Yao-Ming Yeh, Berlin Chen |
INTERSPEECH | 3 |
| 2009 | Generalized likelihood ratio discriminant analysisabstractIn the past several decades, classifier-independent front-end feature extraction, where the derivation of acoustic features is lightly associated with the back-end model training or classification, has been prominently used in various pattern recognition tasks, including automatic speech recognition (ASR). In this paper, we present a novel discriminative feature transformation, named generalized likelihood ratio discriminant analysis (GLRDA), on the basis of the likelihood ratio test (LRT). It attempts to seek a lower dimensional feature subspace by making the most confusing situation, described by the null hypothesis, as unlikely to happen as possible without the homoscedastic assumption on class distributions. We also show that the classical linear discriminant analysis (LDA) and its well-known extension - heteroscedastic linear discriminant analysis (HLDA) can be regarded as two special cases of our proposed method. The empirical class confusion information can be further incorporated into GLRDA for better recognition performance. Experimental results demonstrate that GLRDA and its variant can yield moderate performance improvements over HLDA and LDA for the large vocabulary continuous speech recognition (LVCSR) task. Hung-Shin Lee, Berlin Chen |
ASRU | 2 |
| 2009 | Latent topic modelling of word co-occurence information for spoken document retrievalabstractIn this paper, we present a word topic model (WTM) approach, discovering the co-occurrence relationship between words as well as the long-span latent topic information, for spoken document retrieval (SDR). A given document as a whole is modeled as a composite WTM model for generating an observed query. The underlying characteristics and different kinds of model structures are extensively investigated, while the performance of WTM is thoroughly analyzed and verified by comparison with a few existing retrieval models on the TDT-2 SDR task. We also attempt to incorporate part-of-speech (POS) weighting into the representations of the query observations and the WTM models for obtaining better retrieval performance. Berlin Chen |
ICASSP | 1 |
| 2009 | Empirical error rate minimization based linear discriminant analysisabstractLinear discriminant analysis (LDA) is designed to seek a linear transformation that projects a data set into a lower-dimensional feature space while retaining geometrical class separability. However, LDA cannot always guarantee better classification accuracy. One of the possible reasons lies in that its formulation is not directly associated with the classification error rate, so that it is not necessarily suited for the allocation rule governed by a given classifier, such as that employed in automatic speech recognition (ASR). In this paper, we extend the classical LDA by leveraging the relationship between the empirical classification error rate and the Mahalanobis distance for each respective class pair, and modify the original between-class scatter from a measure of the squared Euclidean distance to the pairwise empirical classification accuracy for each class pair, while preserving the lightweight solvability and taking no distributional assumption, just as what LDA does. Experimental results seem to demonstrate that our approach yields moderate improvements over LDA on the large vocabulary continuous speech recognition (LVCSR) task. Hung-Shin Lee, Berlin Chen |
ICASSP | 2 |
| 2009 | Improved speech summarization with multiple-hypothesis representations and kullback-leibler divergence measuresabstractImperfect speech recognition often leads to degraded performance when leveraging existing text-based methods for speech summarization. To alleviate this problem, this paper investigates various ways to robustly represent the recognition hypotheses of spoken documents beyond the top scoring ones. Moreover, a new summarization method stemming from the Kullback-Leibler (KL) divergence measure and exploring both the sentence and document relevance information is proposed to work with such robust representations. Experiments on broadcast news speech summarization seem to demonstrate the utility of the presented approaches. Index Terms: speech summarization, multiple recognition hypotheses, KL divergence, relevance information Shih-Hsiang Lin, Berlin Chen |
INTERSPEECH | 2 |
| 2009 | Hybrids of supervised and unsupervised models for extractive speech summarizationabstractSpeech summarization, distilling important information and removing redundant and incorrect information from spoken documents, has become an active area of intensive research in the recent past. In this paper, we consider hybrids of supervised and unsupervised models for extractive speech summarization. Moreover, we investigate the use of the unsupervised summarizer to improve the performance of the supervised summarizer when manual labels are not available for training the latter. A novel training data selection and relabeling approach designed to leverage the inter-document or/and the inter-sentence similarity information is explored as well. Encouraging results were initially demonstrated. Index Terms — Speech summarization, hybrid summarizer, unsupervised training Shih-Hsiang Lin, Yueng-Tien Lo, Yao-Ming Yeh, Berlin Chen |
INTERSPEECH | 4 |
| 2009 | Training data selection for improving discriminative training of acoustic models
Berlin Chen, Shih-Hung Liu, Fang-Hui Chu |
Pattern Recognit. Lett. | 1 |
| 2009 | Word Topic Models for Spoken Document Retrieval and TranscriptionabstractStatistical language modeling (LM), which aims to capture the regularities in human natural language and quantify the acceptability of a given word sequence, has long been an interesting yet challenging research topic in the speech and language processing community. It also has been introduced to information retrieval (IR) problems, and provided an effective and theoretically attractive probabilistic framework for building IR systems. In this article, we propose a word topic model (WTM) to explore the co-occurrence relationship between words, as well as the long-span latent topical information, for language modeling in spoken document retrieval and transcription. The document or the search history as a whole is modeled as a composite WTM model for generating a newly observed word. The underlying characteristics and different kinds of model structures are extensively investigated, while the performance of WTM is thoroughly analyzed and verified by comparison with the well-known probabilistic latent semantic analysis (PLSA) model as well as the other models. The IR experiments are performed on the TDT Chinese collections (TDT-2 and TDT-3), while the large vocabulary continuous speech recognition (LVCSR) experiments are conducted on the Mandarin broadcast news collected in Taiwan. Experimental results seem to indicate that WTM is a promising alternative to the existing models. Berlin Chen |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2009 | A Comparative Study of Probabilistic Ranking Models for Chinese Spoken Document SummarizationabstractExtractive document summarization automatically selects a number of indicative sentences, passages, or paragraphs from an original document according to a target summarization ratio, and sequences them to form a concise summary. In this article, we present a comparative study of various probabilistic ranking models for spoken document summarization, including supervised classification-based summarizers and unsupervised probabilistic generative summarizers. We also investigate the use of unsupervised summarizers to improve the performance of supervised summarizers when manual labels are not available for training the latter. A novel training data selection approach that leverages the relevance information of spoken sentences to select reliable document-summary pairs derived by the probabilistic generative summarizers is explored for training the classification-based summarizers. Encouraging initial results on Mandarin Chinese broadcast news data are demonstrated. Shih-Hsiang Lin, Berlin Chen, Hsin-Min Wang |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2009 | A Probabilistic Generative Framework for Extractive Broadcast News Speech SummarizationabstractIn this paper, we consider extractive summarization of broadcast news speech and propose a unified probabilistic generative framework that combines the sentence generative probability and the sentence prior probability for sentence ranking. Each sentence of a spoken document to be summarized is treated as a probabilistic generative model for predicting the document. Two matching strategies, namely literal term matching and concept matching, are thoroughly investigated. We explore the use of the language model (LM) and the relevance model (RM) for literal term matching, while the sentence topical mixture model (STMM) and the word topical mixture model (WTMM) are used for concept matching. In addition, the lexical and prosodic features, as well as the relevance information of spoken sentences, are properly incorporated for the estimation of the sentence prior probability. An elegant feature of our proposed framework is that both the sentence generative probability and the sentence prior probability can be estimated in an unsupervised manner, without the need for handcrafted document-summary pairs. The experiments were performed on Chinese broadcast news collected in Taiwan, and very encouraging results were obtained. Berlin Chen, Hsin-Min Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Exploring the Use of Speech Features and Their Corresponding Distribution Characteristics for Robust Speech RecognitionabstractThe performance of current automatic speech recognition (ASR) systems often deteriorates radically when the input speech is corrupted by various kinds of noise sources. Several methods have been proposed to improve ASR robustness over the last few decades. The related literature can be generally classified into two categories according to whether the methods are directly based on the feature domain or consider some specific statistical feature characteristics. In this paper, we present a polynomial regression approach that has the merit of directly characterizing the relationship between speech features and their corresponding distribution characteristics to compensate for noise interference. The proposed approach and a variant were thoroughly investigated and compared with a few existing noise robustness approaches. All experiments were conducted using the Aurora-2 database and task. The results show that our approaches achieve considerable word error rate reductions over the baseline system and are comparable to most of the conventional robustness approaches discussed in this paper. Shih-Hsiang Lin, Berlin Chen, Yao-Ming Yeh |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Exploiting spatial-temporal feature distribution characteristics for robust speech recognition
Wei-Hau Chen, Shih-Hsiang Lin, Berlin Chen |
INTERSPEECH | 3 |
| 2008 | Linear discriminant feature extraction using weighted classification confusion informationabstractLinear discriminant analysis (LDA) can be viewed as a twostage procedure geometrically. The first stage conducts an orthogonal and whitening transformation of the variables. The second stage involves a principal component analysis (PCA) on the transformed class means, which is intended to maximize the class separability along the principal axes. In this paper, we demonstrate that the second stage does not necessarily guarantee better classification accuracy. Furthermore, we propose a generalization of LDA, weighted LDA (WLDA), by integrating the empirical classification confusion information between each class pair, such that the separability and the classification error rate can be taken into consideration simultaneously. WLDA can be efficiently solved by a lightweight eigen-decomposition and easily combined with other modifications to the LDA criterion. The experiment results show that WLDA can yield an relative character error reduction of 4.6% over LDA on the Mandarin LVCSR task. Hung-Shin Lee, Berlin Chen |
INTERSPEECH | 2 |
| 2008 | Extractive spoken document summarization for information retrieval
Berlin Chen |
Pattern Recognit. Lett. | 1 |
| 2007 | Spoken document summarization using relevant informationabstractExtractive summarization usually automatically selects indicative sentences from a document according to a certain target summarization ratio, and then sequences them to form a summary. In this paper, we investigate the use of information from relevant documents retrieved from a contemporary text collection for each sentence of a spoken document to be summarized in a probabilistic generative framework for extractive spoken document summarization. In the proposed methods, the probability of a document being generated by a sentence is modeled by a hidden Markov model (HMM), while the retrieved relevant text documents are used to estimate the HMM’s parameters and the sentence’s prior probability. The results of experiments on Chinese broadcast news compiled in Taiwan show that the new methods outperform the previous HMM approach. Shih-Hsiang Lin, Hsin-Min Wang, Berlin Chen |
ASRU | 4 |
| 2007 | Investigating the use of speech features and their corresponding distribution characteristics for robust speech recognitionabstractThe performance of current automatic speech recognition (ASR) systems often deteriorates radically when the input speech is corrupted by various kinds of noise sources. Quite a few of techniques have been proposed to improve ASR robustness over the last few decades. Related work reported in the literature can be generally divided into two aspects according to whether the orientation of the methods is either from the feature domain or from the corresponding probability distributions. In this paper, we present a polynomial regression approach which has the merit of directly characterizing the relationship between the speech features and their corresponding probability distributions to compensate the noise effects. Two variants of the proposed approach are also extensively investigated as well. All experiments are conducted on the Aurora-2 database and task. Experimental results show that for clean-condition training, our approaches achieve considerable word error rate reductions over the baseline system, and also significantly outperform other conventional methods. Shih-Hsiang Lin, Yao-Ming Yeh, Berlin Chen |
ASRU | 3 |
| 2007 | Training data selection for improving discriminative training of acoustic modelsabstractThis paper considers training data selection for discriminative training of acoustic models for broadcast news speech recognition. Three novel data selection approaches were proposed. First, the average phone accuracy over all hypothesized word sequences in the word lattice of a training utterance was utilized for utterancelevel data selection. Second, phone-level data selection based on the difference between the expected accuracy of a phone arc and the average phone accuracy of the word lattice was investigated. Finally, frame-level data selection based on the normalized frame-level entropy of Gaussian posterior probabilities obtained from the word lattice was explored. The underlying characteristics of the presented approaches were extensively investigated and their performance was verified by comparison with the standard discriminative training approaches. Experiments conducted on the Mandarin broadcast news collected in Taiwan shown that both phone- and frame-level data selection could achieve slight but consistent improvements over the baseline systems at lower training iterations. Shih-Hung Liu, Fang-Hui Chu, Shih-Hsiang Lin, Hung-Shin Lee, Berlin Chen |
ASRU | 5 |
| 2007 | Word Topical Mixture Models for Dynamic Language Model AdaptationabstractThis paper considers dynamic language model adaptation for Mandarin broadcast news recognition. A word topical mixture model (TMM) is proposed to explore the co-occurrence relationship between words, as well as the long-span latent topical information, for language model adaptation. The search history is modeled as a composite word TMM model for predicting the decoded word. The underlying characteristics and different kinds of model structures were extensively investigated, while the performance of word TMM was analyzed and verified by comparison with the conventional probabilistic latent semantic analysis-based language model (PLSALM) and trigger-based language model (TBLM) adaptation approaches. The large vocabulary continuous speech recognition (LVCSR) experiments were conducted on the Mandarin broadcast news collected in Taiwan. Very promising results in perplexity as well as character error rate reductions were initially obtained. Hsuan-Sheng Chin, Berlin Chen |
ICASSP (4) | 2 |
| 2007 | Word Topical Mixture Models for Extractive Spoken Document SummarizationabstractThis paper considers extractive summarization of Chinese spoken documents. In contrast to conventional approaches, we attempt to deal with the extractive summarization problem under a probabilistic generative framework. A word topical mixture model (w-TMM) was proposed to explore the cooccurrence relationship between words of the language. Each sentence of the spoken document to be summarized was treated as a composite word TMM model for generating the document, and sentences were ranked and selected according to their likelihoods. Various kinds of modeling structures and learning approaches were extensively investigated. In addition, the summarization capabilities were verified by comparison with the other conventional summarization approaches. The experiments were performed on the Chinese broadcast news collected in Taiwan. Noticeable performance gains were obtained. The proposed summarization technique has also been properly integrated into our prototype system for voice retrieval of broadcast news via mobile devices. Berlin Chen |
ICME | 1 |
| 2007 | Improved Histogram Equalzaiton (HEQ) for Robust Speech RecogntionabstractWith the rapid development of intelligent transportation systems (ITS), how to provide users with a natural and efficient human-machine interface is now becoming a crucial issue for driver safety. It is no doubt that speech will be one of the best mediators of human-machine interaction; however, the performance of automatic speech recognition (ASR) always radically degrades when the input speech is corrupted by varying noises. In this paper, we consider the use of histogram equalization (HEQ) for robust ASR. A novel data fitting scheme was presented to efficiently approximate the inverse of the cumulative density function of training speech for HEQ, which has the merits of lower storage and time consumption compared to the conventional table-lookup or quantile based HEQ approaches. Moreover, a more elaborate attempt of using multiple inverse functions for different noise conditions was investigated as well. All experiments were carried out on the Aurora-2 standard database and task. Very encouraging results were obtained. The proposed robustness technique has also been properly integrated into our prototype system for in-vehicle traffic information retrieval using spoken queries. Shih-Hsiang Lin, Hung-Bin Chen, Yao-Ming Yeh, Berlin Chen |
ICME | 4 |
| 2007 | Investigating Data Selection for Minimum Phone Error Training of Acoustic ModelsabstractThis paper considers minimum phone error (MPE) based discriminative training of acoustic models for Mandarin broadcast news recognition. A novel data selection approach based on the normalized frame-level entropy of Gaussian posterior probabilities obtained from the word lattice of the training utterance was explored. It has the merit of making the training algorithm focus much more on the training statistics of those frame samples that center nearly around the decision boundary for better discrimination. Moreover, we presented a new phone accuracy function based on the frame-level accuracy of hypothesized phone arcs instead of using the raw phone accuracy function of MPE training. The underlying characteristics of the presented approaches were extensively investigated and their performance was verified by comparison with the original MPE training approach. Experiments conducted on the broadcast news collected in Taiwan showed that the integration of the frame-level data selection and accuracy calculation could achieve slight but consistent improvements over the baseline system. Shih-Hung Liu, Fang-Hui Chu, Shih-Hsiang Lin, Berlin Chen |
ICME | 4 |
| 2007 | A unified probabilistic generative framework for extractive spoken document summarizationabstractIn this paper, we consider extractive summarization of Chinese broadcast news speech. A unified probabilistic generative framework that combined the sentence generative probability and the sentence prior probability for sentence ranking was proposed. Each sentence of a spoken document to be summarized was treated as a probabilistic generative model for predicting the document. Two different matching strategies, i.e., literal term matching and concept matching, were extensively investigated. We explored the use of the hidden Markov model (HMM) and relevance model (RM) for literal term matching, while the word topical mixture model (WTMM) for concept matching. On the other hand, the confidence scores, structural features, and a set of prosodic features were properly incorporated together using the whole sentence maximum entropy model (WSME) for the estimation of the sentence prior probability. The experiments were performed on the Chinese broadcast news collected in Taiwan. Very promising and encouraging results were initially obtained. Hsuan-Sheng Chiu, Hsin-Min Wang, Berlin Chen |
INTERSPEECH | 4 |
| 2007 | Cluster-based polynomial-fit histogram equalization (CPHEQ) for robust speech recognitionabstractNoise robustness is one of the primary challenges facing most automatic speech recognition (ASR) systems. A vast amount of research efforts on preventing the degradation of ASR performance under various noisy environments have been made during the past several years. In this paper, we consider the use of histogram equalization (HEQ) for robust ASR. In contrast to conventional methods, a novel data fitting method based on polynomial regression was presented to efficiently approximate the inverse of the cumulative density functions of speech feature vectors for HEQ. Moreover, a more elaborate attempt of using such polynomial regression models to directly characterizing the relationship between the speech feature vectors and their corresponding probability distributions, under various noise conditions, was proposed as well. All experiments were carried out on the Aurora-2 database and task. The performance of the presented methods were extensively tested and verified by comparison with the other methods. Experimental results shown that for cleancondition training, our method achieved a considerable word error rate reduction over the baseline system, and also significantly outperformed the other methods. Index Terms: noise robustness, speech recognition, histogram equalization, polynomial regression model Shih-Hsiang Lin, Yao-Ming Yeh, Berlin Chen |
INTERSPEECH | 3 |
| 2007 | Subword-based position specific posterior lattices (s-PSPL) for indexing speech information
Yi-Cheng Pan, Hung-lin Chang, Berlin Chen, Lin-Shan Lee |
INTERSPEECH | 3 |
| 2006 | Chinese Spoken Document Summarization Using Probabilistic Latent Topical InformationabstractThe purpose of extractive summarization is to automatically select a number of indicative sentences, passages, or paragraphs from the original document according to a target summarization ratio and then sequence them to form a concise summary. In the paper, we proposed the use of probabilistic latent topical information for extractive summarization of spoken documents. Various kinds of modeling structures and learning approaches were extensively investigated. In addition, the summarization capabilities were verified by comparison with the conventional vector space model and latent semantic indexing model, as well as the HMM model. The experiments were performed on the Chinese broadcast news collected in Taiwan. Noticeable performance gains were obtained. Berlin Chen, Yao-Ming Yeh, Yao-Min Huang |
ICASSP (1) | 1 |
| 2006 | Exploiting polynomial-fit histogram equalization and temporal average for robust speech recognitionabstractThe performance of current automatic speech recognition (ASR) systems radically deteriorates when the input speech is corrupted by various kinds of noise sources. Quite a few of techniques have been proposed to improve ASR robustness in the past several years. Histogram equalization (HEQ) is one of the most efficient techniques that have been used to compensate the nonlinear distortion. In this paper, we explored the use of the data fitting scheme to efficiently approximate the inverse of the cumulative density function of training speech for HEQ, in contrast to the conventional tablelookup or quantile based approaches. Moreover, the temporal average operation was also performed on the feature vector components to alleviate the influence of sharp peaks and valleys that were caused by non-stationary noises. Finally, we also investigated the possibility of combining our approaches with other feature discrimination and decorrelation methods. All experiments were carried out on the Aurora-2 database and task. Encouraging results were initially demonstrated. Index Terms: histogram equalization, data fitting, temporal average, robustness Shih-Hsiang Lin, Yao-Ming Yeh, Berlin Chen |
INTERSPEECH | 3 |
| 2006 | Voice retrieval of Mandarin broadcast news speechabstractThis paper presents an improved framework for voice retrieval of Mandarin broadcast news speech. First, several unsupervised and data-driven approaches for broadcast news transcription were proposed to improve the speech recognition accuracy and efficiency. Then, a multiscale indexing paradigm for broadcast news retrieval was exploited to alleviate the problems caused by the speech recognition errors and the flexible wording structure of the Chinese language. Finally, we used the PDA as the platform and broadcast radio programs collected in Taiwan as the document collection to establish a speech-based multimedia information retrieval prototype system. Very encouraging results were obtained. Berlin Chen |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2006 | Exploring the use of latent topical information for statistical Chinese spoken document retrieval
Berlin Chen |
Pattern Recognit. Lett. | 1 |
| 2005 | Dynamic language model adaptation using latent topical information and automatic transcriptsabstractThis paper considers dynamic language model adaptation for Mandarin broadcast news recognition. Both contemporary newswire texts and in-domain automatic transcripts were exploited in language model adaptation. A topical mixture model was presented to dynamically explore the long-span latent topical information for language model adaptation. The underlying characteristics and different kinds of model structures were extensively investigated, while their performance was analyzed and verified by comparison with the conventional MAP-based adaptation approaches, which are devoted to extracting the short-span n-gram information. The fusion of global topical and local contextual information was investigated as well. The speech recognition experiments were conducted on the broadcast news collected in Taiwan. Very promising results in perplexity as well as character error rate reductions were initially obtained. Berlin Chen |
ICME | 1 |
| 2005 | Speech retrieval of Mandarin broadcast news via mobile devicesabstractThis paper presents a system for speech retrieval of Mandarin broadcast news. First, several data-driven and unsupervised approaches are integrated into the broadcast news transcription system to improve the speech recognition accuracy and efficiency. Then, a multi-scale indexing paradigm for broadcast news retrieval is proposed to make use of the special structural properties of the Chinese language as well as to alleviate the problems caused by the speech recognition errors. Finally, we use the PDA as the platform and Mandarin broadcast news stories collected in Taiwan as the document collection to establish a speech-based multimedia information retrieval prototype system. Very encouraging results are obtained. 1. Berlin Chen, Chih-Hao Chang, Hung-Bin Chen |
INTERSPEECH | 1 |
| 2005 | Minimum word error based discriminative training of language modelsabstractThis paper considers discriminative training of language models for large vocabulary continuous speech recognition. The minimum word error (MWE) criterion was explored to make use of the word confusion information as well as the local lexical constraints inherent in the acoustic training corpus, in conjunction with those constraints obtained from the background text corpus, for properly guiding the speech recognizer to separate the correct hypothesis from the competing ones. The underlying characteristics of the MWEbased approach were extensively investigated, and its performance was verified by comparison with the conventional maximum likelihood (ML) approaches as well. The speech recognition experiments were performed on the broadcast news collected in Taiwan. Jen-Wei Kuo, Berlin Chen |
INTERSPEECH | 2 |
| 2005 | Hierarchical topic organization and visual presentation of spoken documents using probabilistic latent semantic analysis (PLSA) for efficient retrieval/browsing applicationsabstractThe most attractive form of future network content will be multi-media including speech information, and such speech information usually carries the core concepts for the content. As a result, the spoken documents associated with the multi-media content very possibly can serve as the key for retrieval and browsing. This paper presents a new approach of hierarchical topic organization and visual presentation of spoken documents for such a purpose based on the Probabilistic Latent Semantic Analysis (PLSA). With this approach the spoken documents can be organized into a two-dimensional tree (or multi-layered map) of topic clusters, and the user can very efficiently retrieve or browse the network content or associated spoken documents. Different from the conventional document clustering approaches, with PLSA the relationships among the topic clusters and the appropriate terms as the topic labels can be very well derived. An initial prototype system with Chinese broadcast news as the example spoken documents including automatic generation of titles and summaries and retrieval/browsing functionalities is also presented. Choice of different units other than words to be used as the terms in the processing is also considered in the system based on the special structure of the Chinese language. 1. Te-Hsuan Li, Ming-Han Lee, Berlin Chen, Lin-Shan Lee |
INTERSPEECH | 3 |
| 2004 | Lightly supervised and data-driven approaches to Mandarin broadcast news transcriptionabstractThis paper investigates the use of several lightly supervised and data-driven approaches to Mandarin broadcast news transcription. First, with a consideration of the special structural properties of the Chinese language, a fast acoustic took-ahead technique for estimating the unexplored part of speech utterance was integrated into the lexical tree search to improve the search efficiency, in conjunction with the conventional language model look-ahead technique. Then, a verification-based method for automatic acoustic training data acquisition was developed to make use of the large amount of untranscribed speech data. Finally, two alternative strategies for language model adaptation were further studied for accurate language model estimation. With the above approaches, the system yielded an 11.94% character error rate on the Mandarin broadcast news collected in Taiwan. Berlin Chen, Jen-Wei Kuo, Wen-Hung Tsai |
ICASSP (1) | 1 |
| 2004 | Statistical Chinese spoken document retrieval using latent topical informationabstractInformation retrieval which aims to provide people with easy access to all kinds of information is now becoming more and more emphasized. However, most approaches to information retrieval are primarily based on literal term matching and operate in a deterministic manner. Thus their performance is often limited due to the problems of vocabulary mismatch and not able to be steadily improved through use. In order to overcome these drawbacks as well as to enhance the retrieval performance, in this paper we explore the use of topical mixture model for statistical Chinese spoken document retrieval. Various kinds of model structures and learning approaches were extensively investigated. In addition, the retrieval capabilities were verified by comparison with the conventional vector space model and latent semantic indexing model, as well as our previously presented HMM/N-gram retrieval model. The experiments were performed on the TDT2 Chinese collection. Noticeable improvements in retrieval performance were obtained. Jen-Wei Kuo, Yao-Min Huang, Berlin Chen, Hsin-Min Wang |
INTERSPEECH | 3 |
| 2004 | Mandarin-English Information (MEI): investigating translingual speech retrieval
Helen M. Meng, Berlin Chen, Sanjeev Khudanpur, Gina-Anne Levow, Wai Kit Lo, Douglas W. Oard, Patrick Schone, Karen Tang, Hsin-Min Wang, Jianqiang Wang 0002 |
Comput. Speech Lang. | 2 |
| 2004 | A discriminative HMM/N-gram-based retrieval approach for mandarin spoken documentsabstractIn recent years, statistical modeling approaches have steadily gained in popularity in the field of information retrieval. This article presents an HMM/N-gram-based retrieval approach for Mandarin spoken documents. The underlying characteristics and the various structures of this approach were extensively investigated and analyzed. The retrieval capabilities were verified by tests with word- and syllable-level indexing features and comparisons to the conventional vector-space model approach. To further improve the discrimination capabilities of the HMMs, both the expectation-maximization (EM) and minimum classification error (MCE) training algorithms were introduced in training. Fusion of information via indexing word- and syllable-level features was also investigated. The spoken document retrieval experiments were performed on the Topic Detection and Tracking Corpora (TDT-2 and TDT-3). Very encouraging retrieval performance was obtained. Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2002 | Improved Chinese spoken document retrieval with hybrid modeling and data-driven indexing featuresabstractDifferent models retrieve the documents based on different approaches of extracting the underlying content. Different levels of indexing features also offer different functionalities and discriminabilities when retrieving the documents. In this paper, we present results for Chinese spoken document retrieval with hybrid models to integrate the knowledge obtainable from three basic retrieval models, namely, the standard vector space model (VSM), the hidden Markov model (HMM), and the latent semantic indexing (LSI) model. The characteristics of retrieval performance using both word-level and syllable-level indexing features were extensively explored. In addition, a data-driven approach to derive variable-length indexing features is also presented. Very satisfactory performance can be achieved with these data-driven features while retaining very compact feature set size. Experiments showed that this approach has the potential to identify domain-specific terminologies or newly-generated phrases. It is therefore very useful not only in Chinese document retrieval, but also in detecting out of vocabulary (OOV) words in Chinese. Very encouraging results were obtained when the hybrid models were used with the data-driven indexing features as well. 1. Chun-Jen Wang, Berlin Chen, Lin-Shan Lee |
INTERSPEECH | 2 |
| 2002 | A hierarchical tag-graph search scheme with layered grammar rules for spontaneous speech understanding
Bor-Shen Lin, Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
Pattern Recognit. Lett. | 2 |
| 2002 | Discriminating capabilities of syllable-based features and approaches of utilizing them for voice retrieval of speech information in Mandarin ChineseabstractWith the rapidly growing use of the audio and multimedia information over the Internet, the technology for retrieving speech information using voice queries is becoming more and more important. In this paper, considering the monosyllabic structure of the Chinese language, a whole class of syllable-based indexing features, including overlapping segments of syllables and syllable pairs separated by a few syllables, is extensively investigated based on a Mandarin broadcast news database. The strong discriminating capabilities of such syllable-based features were verified by comparing with the word- or character-based features. Good approaches for better utilizing such capabilities, including fusion with the word- and character-level information and improved approaches to obtain better syllable-based features and query expressions, were extensively investigated. Very encouraging experimental results were obtained. Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
IEEE Trans. Speech Audio Process. | 1 |
| 2001 | Multi-scale-audio indexing for translingual spoken document retrievalabstractMEI (Mandarin-English Information) is an English-Chinese crosslingual spoken document retrieval (CL-SDR) system developed during the Johns Hopkins University Summer Workshop 2000. We integrate speech recognition, machine translation, and information retrieval technologies to perform CL-SDR. MEI advocates a multi-scale paradigm, where both Chinese words and subwords (characters and syllables) are used in retrieval. The use of subword units can complement the word unit in handling the problems of Chinese word tokenization ambiguity, Chinese homophone ambiguity, and out-of-vocabulary words in audio indexing. This paper focuses on multi-scale audio indexing in MEI. Experiments are based on the Topic Detection and Tracking Corpora (TDT-2 and TDT-3), where we indexed Voice of America Mandarin news broadcasts by speech recognition on both the word and subword scales. We discuss the development of the MEI syllable recognizer, the representations of spoken documents using overlapping subword n-grams and lattice structures. Results show that augmenting words with subwords is beneficial to CL-SDR performance. Hsin-Min Wang, Helen M. Meng, Patrick Schone, Berlin Chen, Wai Kit Lo |
ICASSP | 4 |
| 2001 | Improved spoken document retrieval by exploring extra acoustic and linguistic cuesabstractIn this paper, we explored the use of various extra information to improve the performance of spoken document retrieval (SDR). From the speech recognition perspective, we incorporated the acoustic stress and word confusion information into the audio indexing. From the linguistic perspective, we applied the part-of-speech information in both the audio indexing and the query representation. From the information retrieval perspective, we integrated techniques such as the query expansion by word associations and the blind relevance feedback into the retrieval process. The SDR experiments were based on the Topic Detection and Tracking Corpora (TDT-2 and TDT-3). We used the Chinese newswire text stories as query exemplars and the Mandarin Chinese audio news stories as the spoken documents. With all the above acoustic and linguistic cues applied, the average precision was improved from 0.5122 to 0.6312 for the TDT-2 collection and from 0.6216 to 0.7172 for the TDT-3 collection. 1. Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
INTERSPEECH | 1 |
| 2001 | An HMM/n-gram-based linguistic processing approach for Mandarin spoken document retrievalabstractIn this paper an HMM/N-gram-based linguistic processing approach for Mandarin spoken document retrieval is presented. The underlying characteristics and different structures of this approach were extensively investigated. The retrieval capabilities were verified by tests with indexing features of word- and syllable(subword)-levels and comparison with the conventional vector space model approach. To further improve the discrimination capabilities of the HMMs, both the expectation-maximization (EM) and minimum classification error (MCE) training algorithms were introduced in training. The information fusion of indexing features of word- and syllable-levels was also investigated. The spoken document retrieval experiments were performed on the Topic Detection and Tracking Corpora (TDT-2 and TDT-3). Very encouraging retrieval performance was obtained. Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
INTERSPEECH | 1 |
| 2000 | Retrieval of broadcast news speech in Mandarin Chinese collected in Taiwan using syllable-level statistical characteristicsabstractSpoken document retrieval has been extensively studied over the years because of its high potential in various applications in the near future. Considering the monosyllabic structure of the Chinese language, a whole class of indexing features for retrieval of spoken documents in Mandarin Chinese using syllable-level statistical characteristics has been studied, and very encouraging experimental results on retrieval of broadcast news speech collected in Taiwan were obtained. This paper reports some interesting initial results and findings obtained in this research. Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
ICASSP | 1 |
| 2000 | Retrieval of mandarin broadcast news using spoken queriesabstractConsidering the monosyllabic structure of the Chinese language, a whole class of indexing features for retrieval of Mandarin broadcast news using syllable-level statistical characteristics has been previously investigated. This paper presents the improvements achieved over the previous results. The major differences are: (1) Multi-scale character- and word-level indexing terms have been integrated with the syllable-level information. (2) Information cues from the contemporary newswire text corpus have been used to create more accurate syllable indexing terms. (3) Automatic document expansion, blind relevance feedback, and query expansion via the term association matrix have been applied in retrieval. With all these schemes, the average precision can be improved from 55.46 % to 71.29%. 1. Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
INTERSPEECH | 1 |
| 2000 | Syllable-Based Chinese Text/Spoken Document Retrieval Using Text/Speech QueriesabstractIn light of the rapid growth of Chinese information resources on the Internet, this study investigates a novel approach that deals with the problem of Chinese text and spoken document retrieval using both text and speech queries. By properly utilizing the monosyllabic structure of the Chinese language, the proposed approach estimates the statistical similarity between the text/speech queries and the text/spoken documents at the phonetic level using the syllable-based statistical information. The investigation successfully implemented a prototype system with an interface supporting some user-friendly functions and the initial test results demonstrate the feasibility of the proposed approach. Bo-Ren Bai, Berlin Chen, Hsin-Min Wang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 1998 | A*-admissible key-phrase spotting with sub-syllable level utterance verificationabstractIn this paper, we propose an A*-admissible key-phrase spotting framework, which needs little domain knowledge and is capable of extracting salient key-phrase fragments from an input utterance in real-time. There are two key features in our approach. Firstly, the acoustic models and the search framework are specially designed such that very high degree vocabulary flexibility can be achieved for any desired application tasks. Secondly, the search framework uses an efficient two-pass A* search to generate N-best key-phrase candidates and then several sub-syllable level verification functions are properly weighted and used to further improve the recognition accuracy. Experimental results show that the A*-admissible key-phrase spotting with sub-word level utterance method outperforms the baseline methods used in common approaches. 1. INTRODUCTION In recent years, various spoken dialog systems have been widely investigated for the fast growing demand for real-world applications. It is diff... Berlin Chen, Hsin-Min Wang, Lee-Feng Chien, Lin-Shan Lee |
ICSLP | 1 |
| 1998 | Hierarchical tag-graph search for spontaneous speech understanding in spoken dialog systemsabstractIt has been relatively difficult to develop natural language parsers for spoken dialog systems, not only because of the possible recognition errors, pauses, hesitations, out-ofvocabulary words, and the grammatically incorrect sentence structures, but because of the great efforts required to develop a general enough grammar with satisfactory coverage and flexibility to handle different applications. In this paper, a new hierarchical graph-based search scheme with layered structure is presented, which is shown to provide more robust and flexible spontaneous speech understanding for spoken dialog systems. 1.INTRODUCTION Traditionally, natural language understanding is integrated with the speech recognizer with a N-best interface in spoken dialog systems [1][2], that is, the recognizer sequentially generates its best N sentence hypotheses until any one is accepted by the natural language understanding part. However, for spontaneous speech with fragments, disfluencies, OOV words, and ill-... Bor-Shen Lin, Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
ICSLP | 2 |
| 1998 | Towards a Mandarin voice memo systemabstractUsing voice memos in stead of text memos is believed to be more natural, convenient, and attractive. This paper presents a working Mandarin voice memo system that provides functions of automatic notification and voice retrieval. The main techniques include the content-based spoken document retrieval approach and the date-time expression detection and understanding approach. Extensive preliminary experiments were performed and encouraging results were demonstrated. 1. INTRODUCTION Using voice memos in stead of text memos is believed to be more natural, convenient, and attractive because it is definitely much easier for people to speak memos than to write down memos or to type memos into computers using a keyboard. Furthermore, users are more likely to record detailed information if all they need to do is just speak. However, it's far more difficult to retrieve these voice memos than to retrieve the text ones. With advances in the speech recognition technology, voice retrieval of spoke... Hsin-Min Wang, Bor-Shen Lin, Berlin Chen, Bo-Ren Bai |
ICSLP | 3 |