Kuan-Yu Chen 0002

dblp:35/6313-2 · DBLP profile ↗
← Back
69ranked-venue papers
28as first author
10since 2021 · last 2026
0000-0001-9656-7551ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 15 first-author · 2 since 2021Artificial intelligence and machine learning · 30 · 8 first-author · 4 since 2021Theory of computation · 12 · 8 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Toward Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model Representations
abstract
Perceptual voice quality assessment plays a vital role in diagnosing and monitoring voice disorders.Traditional methods, such as the Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V) and the Grade, Roughness, Breathiness, Asthenia, and Strain (GRBAS) scales, rely on expert raters and are prone to inter-rater variability, emphasizing the need for objective solutions. This study introduces the Voice Quality Assessment Network (VOQANet), a deep learning framework that employs an attention mechanism and Speech Foundation Model (SFM) embeddings to extract high-level features. To further enhance performance, we propose VOQANet+, which integrates self-supervised SFM embeddings with low-level acoustic descriptors-namely jitter, shimmer, and harmonics-to-noise ratio (HNR). Unlike previous approaches that focus solely on vowel-based phonation (PVQD-A), our models are evaluated on both vowel-level and sentence-level speech (PVQD-S) to assess generalizability. Experimental results demonstrate that sentence-based inputs yield higher accuracy, particularly at the patient level. Overall, VOQANet consistently outperforms baseline models in terms of root mean squared error (RMSE) and Pearson correlation coefficient across CAPE-V and GRBAS dimensions, with VOQANet+ achieving even greater performance gains. Additionally, VOQANet+ maintains consistent performance under noisy conditions, suggesting enhanced robustness for real-world and telehealth applications. This work highlights the value of combining SFM embeddings with low-level features for accurate and robust pathological voice assessment.
Whenty Ariyanti, Kuan-Yu Chen 0002, Sabato Marco Siniscalchi, Hsin-Min Wang, Yu Tsao 0001
IEEE J. Biomed. Health Informatics2
2025 RATE: A Retrieval-Augmented Transformer for Regional Earthquake Early Warning
abstract
Accurate and timely seismic intensity prediction is essential for effective regional earthquake early warning (EEW). This study presents a retrieval-augmented Transformer (RATE) model that leverages historical seismic events to enhance regional ground motion predictions. Upon receiving initial P-phase signals, RATE retrieves similar past events based on waveform similarity and integrates them into a Transformer-based prediction pipeline. This design allows the model to adapt the diverse seismic contexts and generalize across regions. The experiments on datasets from Japan and Taiwan demonstrate that the RATE consistently outperforms baseline models in terms of intensity estimation accuracy and alert precision. These results highlight the potential of a retrieval-augmented (RA) framework to enhance real-time EEW capabilities in diverse seismic regions.
Wen-Wei Lin, Kuan-Yu Chen 0002, Da-Yi Chen
IEEE Geosci. Remote. Sens. Lett.2
2025 An Attention-Based Framework With Multistation Information for Earthquake Early Warnings
abstract
Earthquake early warning systems play crucial roles in reducing the risk of seismic disasters. Previously, the dominant modeling system was the single-station models. Such models digest signal data received at a given station and predict earthquake parameters, such as the p-phase arrival time, intensity, and magnitude at that location. Various methods have demonstrated adequate performance. However, most of these methods present the challenges of the difficulty of speeding up the alarm time, providing early warning for distant areas, and considering global information to enhance performance. Recently, deep learning has significantly impacted many fields, including seismology. Thus, this paper proposes a deep learning-based framework, called SENSE1, for the intensity prediction task of earthquake early warning systems. To explicitly consider global information from a regional or national perspective, the input to the SENSE model includes features derived from multiple stations within a specific region or country. The SENSE model is designed to learn the relationships among these input stations and to derive locality-specific embeddings for each station individually, further enhancing prediction performance. Thus, SENSE is not only expected to provide more reliable forecasts by considering multistation data but also has the ability to provide early warnings to distant areas that have not yet received signals. This study conducted extensive experiments on datasets from Japan and Taiwan. The results revealed that SENSE can deliver competitive or even better performances compared with other advanced deep learning-based methods.
Kuan-Yu Chen 0002, Wen-Wei Lin, Da-Yi Chen
IEEE Trans. Geosci. Remote. Sens.2
2024 A Small-Footprint Keyword Spotting Modeling with Novel Data Augmentation Methods
abstract
Keyword spotting (KWS) system is an important human-computer interaction medium in smart devices. However, requiring the KWS model to maintain robust performance with a small number of parameters is extremely challenging. In addition to model architecture research, efficiently leveraging speech data to develop more powerful KWS systems is also an attractive research topic. In the speech-related community, various data augmentation methods have been studied. Mainstream methods mainly focus on manipulating speech waveforms or acoustic features. Regardless of the approach taken, a simple conclusion is that data augmentation strategies can further improve task performance. Motivated by practical needs and following interesting research lines, on the one hand, the paper presents an efficient KWS model that is designed to leverage the spatial and channel information of the input speech simultaneously while keeping the lightweight property of the model. On the other hand, we develop two novel data augmentation methods for training the KWS model to boost performance further. A series of experiments are conducted using the Google Speech Commands V2 dataset to evaluate the proposed KWS framework. Based on the experiments, the proposed framework can achieve 98.81% accuracy with fewer parameters than other state-of-the-art models.
Wei-Kai Huang, Kuan-Yu Chen 0002
IJCNN2
2022 A Context-Aware Knowledge Transferring Strategy for CTC-Based ASR
abstract
Non-autoregressive automatic speech recognition (ASR) modeling has received increasing attention recently because of its fast decoding speed and superior performance. Among representatives, methods based on the connectionist temporal classification (CTC) are still a dominating stream. However, the theoretically inherent flaw, the assumption of independence between tokens, creates a performance barrier for the school of works. To mitigate the challenge, we propose a context-aware knowledge transferring strategy, consisting of a knowledge transferring module and a context-aware training strategy, for CTC-based ASR. The former is designed to distill linguistic information from a pre-trained language model, and the latter is framed to modulate the limitations caused by the conditional independence assumption. As a result, a knowledge-injected context-aware CTC-based ASR built upon the wav2vec2.0 is presented in this paper. A series of experiments on the AISHELL-1 and AISHELL-2 datasets demonstrate the effectiveness of the proposed method.
Ke-Han Lu, Kuan-Yu Chen 0002
SLT2
2022 Non-Autoregressive ASR Modeling Using Pre-Trained Language Models for Chinese Speech Recognition
abstract
Transformer-based models have led to significant innovation in various classic and practical subjects, including speech processing, natural language processing, and computer vision. On top of the Transformer, attention-based end-to-end automatic speech recognition (ASR) models have become a popular fashion in recent years. Specifically, an emergent research topic is non-autoregressive modeling, which can achieve fast inference speed and obtain competitive performance when compared with conventional autoregressive methods. In addition, in the context of natural language processing, the bidirectional encoder representations from Transformers (BERT) model and its variants have received widespread attention, partially due to their ability to infer contextualized word representations and obtain superior performances of downstream tasks through simple fine-tuning. However, to our knowledge, leveraging the synergistic power of non-autoregressive modeling and pre-trained language model for ASR remains relatively underexplored. In this regard, this study presents a novel pre-trained language model-based non-autoregressive ASR framework. A series of experiments were conducted on two publicly available Chinese datasets, AISHELL-1 and AISHELL-2, to demonstrate competitive or superior results of the proposed ASR models when compared with well-practiced baseline systems. In addition, a set of comparative experiments is likewise carried out with different settings to analyze the performance of the proposed framework.
Fu-Hao Yu, Kuan-Yu Chen 0002, Ke-Han Lu
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 An Attention-Based Hypocenter Estimator for Earthquake Localization
abstract
The accuracy of earthquake localization is of great importance for earthquake monitoring systems. Traditionally, numerical optimization methods are used to estimate the hypocenter location and the origin time of an earthquake in an iterative manner. The traditional methods usually depend on certain theoretical models, but the geological conditions in practice can be quite different from the presumed models. In this study, an attention-based hypocenter estimation (AHE) model was proposed to locate the hypocenter and origin time of the earthquake. Rather than using the raw waveforms, the phase picking times and the positions of the triggered stations are used as the input. The attention mechanism is adapted to reveal the correlations among the input sequence. An experimental model was trained using data collected from earthquakes in Taiwan in 2016 and 2017 and tested on data in 2018. From the results, AHE is capable of locating the hypocenter with a high degree of accuracy in terms of the distance, depth, and origin time of earthquakes.
Tai-Lin Chin, Kuan-Yu Chen 0002, Da-Yi Chen, Te-Hsiu Wang
IEEE Trans. Geosci. Remote. Sens.2
2021 A Re-thinking ASR Modeling Framework using Attention Mechanisms
abstract
Several reasons have led to the widespread adoption of neural-based algorithms for end-to-end automatic speech recognition (ASR), including their high performance, elegant model designs, and parallel computing capabilities. Numerous ASR models have been proposed to improve the recognition results, but the gains are still insufficient. This paper proposes a re-thinking ASR model, which aims to bridge the gap by rethinking the regularities of a given hypothesis and the relationship between text-level and acoustic-level characteristics of the input speech utterance. For the re-thinking ASR model, a mixed attention mechanism, a self-and-mixed attention mechanism, and a deep acoustic feature extractor are meticulously designed to enable the notion to be realized. A publicly available benchmark corpus is used to evaluate the proposed model. As the experimental results demonstrate, the proposed re-thinking ASR model can provide significant and consistent improvements over popular baseline systems.
Chih-Ying Yang, Kuan-Yu Chen 0002
IEEE BigData2
2021 Speech Recognition by Simply Fine-Tuning Bert
abstract
We propose a simple method for automatic speech recognition (ASR) by fine-tuning BERT, which is a language model (LM) trained on large-scale unlabeled text data and can generate rich contextual representations. Our assumption is that given a history context sequence, a powerful LM can narrow the range of possible choices and the speech signal can be used as a simple clue. Hence, comparing to conventional ASR systems that train a powerful acoustic model (AM) from scratch, we believe that speech recognition is possible by simply fine-tuning a BERT model. As an initial study, we demonstrate the effectiveness of the proposed idea on the AISHELL dataset and show that stacking a very simple AM on top of BERT can yield reasonable performance.
Wen-Chin Huang, Chia-Hua Wu, Shang-Bao Luo, Kuan-Yu Chen 0002, Hsin-Min Wang, Tomoki Toda
ICASSP4
2021 Audio-Aware Spoken Multiple-Choice Question Answering With Pre-Trained Language Models
abstract
Spoken multiple-choice question answering (SMCQA) requires machines to select the correct choice to answer the question by referring to the passage, where the passage, the question, and multiple choices are all in the form of speech. While the audio could contain useful cues for SMCQA, usually only the auto-transcribed text is utilized in model development. Thanks to the large-scaled pre-trained language representation models, such as the bidirectional encoder representations from Transformers (BERT), systems with only auto-transcribed text can still achieve a certain level of performance. However, previous studies have evidenced that acoustic-level statistics can offset text inaccuracies caused by the automatic speech recognition systems or representation inadequacy lurking in word embedding generators, thereby making the SMCQA system robust. Along the line of research, in this study, an audio-aware SMCQA framework is proposed. Two different mechanisms are introduced to distill the useful cues from speech, and then a BERT-based SMCQA framework is presented. In other words, the proposed SMCQA framework not only inherits the advantages of contextualized language representations learned by BERT but integrates the complementary acoustic-level information distilled from audio with the text-level information. A series of experiments demonstrates remarkable improvements in accuracy over selected baselines and SOTA systems on a published Chinese SMCQA dataset.
Chia-Chih Kuo, Kuan-Yu Chen 0002, Shang-Bao Luo
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 An Audio-Enriched BERT-Based Framework for Spoken Multiple-Choice Question Answering
abstract
In a spoken multiple-choice question answering (SMCQA) task, given a passage, a question, and multiple choices all in the form of speech, the machine needs to pick the correct choice to answer the question. While the audio could contain useful cues for SMCQA, usually only the auto-transcribed text is utilized in system development. Thanks to the large-scaled pre-trained language representation models, such as the bidirectional encoder representations from transformers (BERT), systems with only auto-transcribed text can still achieve a certain level of performance. However, previous studies have evidenced that acoustic-level statistics can offset text inaccuracies caused by the automatic speech recognition systems or representation inadequacy lurking in word embedding generators, thereby making the SMCQA system robust. Along the line of research, this study concentrates on designing a BERT-based SMCQA framework, which not only inherits the advantages of contextualized language representations learned by BERT, but integrates the complementary acoustic-level information distilled from audio with the text-level information. Consequently, an audio-enriched BERT-based SMCQA framework is proposed. A series of experiments demonstrates remarkable improvements in accuracy over selected baselines and SOTA systems on a published Chinese SMCQA dataset.
Chia-Chih Kuo, Shang-Bao Luo, Kuan-Yu Chen 0002
INTERSPEECH3
2020 Enhanced Language Modeling with Proximity and Sentence Relatedness Information for Extractive Broadcast News Summarization
abstract
The primary task of extractive summarization is to automatically select a set of representative sentences from a text or spoken document that can concisely express the most important theme of the original document. Recently, language modeling (LM) has been proven to be a promising modeling framework for performing this task in an unsupervised manner. However, there still remain three fundamental challenges facing the existing LM-based methods, which we set out to tackle in this article. The first one is how to construct a more accurate sentence model in this framework without resorting to external sources of information. The second is how to take into account sentence-level structural relationships, in addition to word-level information within a document, for important sentence selection. The last one is how to exploit the proximity cues inherent in sentences to obtain a more accurate estimation of respective sentence models. Specifically, for the first and second challenges, we explore a novel, principled approach that generates overlapped clusters to extract sentence relatedness information from the document to be summarized, which can be used not only to enhance the estimation of various sentence models but also to render sentence-level structural relationships within the document, leading to better summarization effectiveness. For the third challenge, we investigate several formulations of proximity cues for use in sentence modeling involved in the LM-based summarization framework, free of the strict bag-of-words assumption. Furthermore, we also present various ensemble methods that seamlessly integrate proximity and sentence relatedness information into sentence modeling. Extensive experiments conducted on a Mandarin broadcast news summarization task show that such integration of proximity and sentence relatedness information is indeed beneficial for speech summarization. Our proposed summarization methods can significantly boost the performance of an LM-based strong baseline (e.g., with a maximum ROUGE-2 improvement of 26.7% relative) and also outperform several state-of-the-art unsupervised methods compared in the article.
Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2020 Intelligent Real-Time Earthquake Detection by Recurrent Neural Networks
abstract
Taiwan that is located at the junction of the Eurasian Plate and the Philippine Sea Plate is one of the most active seismic zones in the world. Devastating earthquakes have occurred around the island and have caused severe damages from time to time. To avoid the severe loss, earthquake early warning (EEW) is of great importance, and one of the most critical issues of EEW is fast and reliable detection for the presence of earthquakes. Traditional methods for earthquake detection usually use criterion-based algorithms to detect the onset of the earthquake waves. Currently, the thresholds for those criteria are usually decided empirically and may result in excessive false alarms. Obviously, false alarms can cause undue panics and diminish the credibility of the system. In this article, the recurrent neural network (RNN) models are adopted to develop a real-time EEW system. The developed system is designed to identify the occurrence of an earthquake event, and the duration of the P-wave and the S-wave. It was trained and tested using the seismograms recorded in Taiwan from 2016 to 2017. From the simulation results, the proposed scheme outperforms the traditional criterion-based schemes in terms of detection accuracy and processing time.
Tai-Lin Chin, Kuan-Yu Chen 0002, Da-Yi Chen, De-En Lin
IEEE Trans. Geosci. Remote. Sens.2
2019 Spoken Multiple-Choice Question Answering Using Multimodal Convolutional Neural Networks
abstract
In a spoken multiple-choice question answering (MCQA) task, where passages, questions, and choices are given in the form of speech, usually only the auto-transcribed text is considered in system development. The acoustic-level information may contain useful cues for answer prediction. However, to the best of our knowledge, only a few studies focus on using the acoustic-level information or fusing the acoustic-level information with the text-level information for a spoken MCQA task. Therefore, this paper presents a hierarchical multistage multimodal (HMM) framework based on convolutional neural networks (CNNs) to integrate text- and acoustic-level statistics into neural modeling for spoken MCQA. Specifically, the acoustic-level statistics are expected to offset text inaccuracies caused by automatic speech recognition (ASR) systems or representation inadequacy lurking in word embedding generators, thereby making the spoken MCQA system robust. In the proposed HMM framework, two modalities are first manipulated to separately derive the acoustic- and text-level representations for the passage, question, and choices. Next, these clever features are jointly involved in inferring the relationships among the passage, question, and choices. Then, a final representation is derived for each choice, which encodes the relationship of the choice to the passage and question. Finally, the most likely answer is determined based on the individual final representations of all choices. Evaluated on the data of “Formosa Grand Challenge - Talk to AI”, a Mandarin Chinese spoken MCQA contest held in 2018, the proposed HMM framework achieves remarkable improvements in accuracy over the text-only baseline.
Shang-Bao Luo, Hung-Shin Lee, Kuan-Yu Chen 0002, Hsin-Min Wang
ASRU3
2019 Exploring the Encoder Layers of Discriminative Autoencoders for LVCSR
Pin-Tuan Huang, Hung-Shin Lee, Syu-Siang Wang, Kuan-Yu Chen 0002, Yu Tsao 0001, Hsin-Min Wang
INTERSPEECH4
2018 Essence Vector-Based Query Modeling for Spoken Document Retrieval
abstract
Spoken document retrieval (SDR) has become a prominently required application since unprecedented volumes of multimedia data along with speech have become available in our daily life. As far as we are aware, there has been relatively less work in launching unsupervised paragraph embedding methods and investigating the effectiveness of these methods on the SDR task. This paper first presents a novel paragraph embedding method, named the essence vector (EV) model, which aims at inferring a representation for a given paragraph by encapsulating the most representative information from the paragraph and excluding the general background information at the same time. On top of the EV model, we develop three query language modeling mechanisms to improve the retrieval performance. A series of empirical SDR experiments conducted on two benchmark collections demonstrate the good efficacy of the proposed framework, compared to several existing strong baseline systems.
Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang
ICASSP1
2018 An Information Distillation Framework for Extractive Summarization
abstract
In the context of natural language processing, representation learning has emerged as a newly active research subject because of its excellent performance in many applications. Learning representations of words is a pioneering study in this school of research. However, paragraph (or sentence and document) embedding learning is more suitable/reasonable for some realistic tasks such as document summarization. Nevertheless, classic paragraph embedding methods infer the representation of a given paragraph by considering all of the words occurring in the paragraph. Consequently, those stop or function words that occur frequently may mislead the embedding learning process to produce a misty paragraph representation. Motivated by these observations, our major contributions in this paper are threefold. First, we propose a novel unsupervised paragraph embedding method, named the essence vector (EV) model, which aims at not only distilling the most representative information from a paragraph but also excluding the general background information to produce a more informative low-dimensional vector representation for the paragraph of interest. Second, in view of the increasing importance of spoken content processing, an extension of the EV model, named the denoising essence vector (D-EV) model, is proposed. The D-EV model not only inherits the advantages of the EV model but also can infer a more robust representation for a given spoken paragraph against imperfect speech recognition. Third, a new summarization framework, which can take both relevance and redundancy information into account simultaneously, is also introduced. We evaluate the proposed embedding methods (i.e., EV and D-EV) and the summarization framework on two benchmark summarization corpora. The experimental results demonstrate the effectiveness and applicability of the proposed framework in relation to several well-practiced and state-of-the-art summarization methods.
Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Neural relevance-aware query modeling for spoken document retrieval
abstract
Spoken document retrieval (SDR) is becoming a much-needed application due to that unprecedented volumes of audio-visual media have been made available in our daily life. As far as we are aware, most of the wide variety of SDR methods mainly focus on exploring robust indexing and effective retrieval methods to quantify the relevance degree between a pair of query and document. However, similar to information retrieval (IR), a fundamental challenge facing SDR is that a query is usually too short to convey a user's information need, such that a retrieval system cannot always achieve prospective efficacy when with the existing retrieval methods. In order to further boost retrieval performance, several studies turn their attention to reformulating the original query by leveraging an online pseudo-relevance feedback (PRF) process, which often comes at the price of taking significant time. Motivated by these observations, this paper presents a novel extension of the general line of SDR research and its contribution is at least two-fold. First, building on neural network-based techniques, we put forward a neural relevance-aware query modeling (NRM) framework, which is designed to not only infer a discriminative query language model automatically for a given query, but also get around the time-consuming PRF process. Second, the utility of the methods instantiated from our proposed framework and several widely-used retrieval methods are extensively analyzed and compared on a standard SDR task, which suggests the superiority of our methods.
Tien-Hong Lo, Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen
ASRU3
2017 A locality-preserving essence vector modeling framework for spoken document retrieval
abstract
Because unprecedented volumes of multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research area in the past decades. Recently, representation learning has emerged as an active research topic in many machine learning applications owing largely to its excellent performance. In the context of natural language processing, the pioneering work can date back to the word embedding methods. However, learning of paragraph (or sentence and document) representations is more reasonable and suitable for some tasks, such as information retrieval and document summarization. Nevertheless, as far as we are aware, there is relatively less work focusing on launching paragraph embedding methods into SDR. Motivated by these observations, this paper proposes a novel paragraph embedding method, named the locality-preserving essence vector (LPEV) model. LPEV is designed with consideration to two aspects. First, the model aims at not only distilling the most representative information from a paragraph but also getting rid of the general background information. Second, inspired by the local invariance perspective, which is a celebrated principle used in manifold learning techniques, LPEV also manages to preserve semantic locality in the learned low-dimensional embedding space for producing more informative and discriminative vector representations of paragraphs. On top of the proposed framework, a series of empirical SDR experiments conducted on the TDT-2 (Topic Detection and Tracking) collection demonstrate the good efficacy of our SDR methods as compared to existing strong baselines.
Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang
ICASSP1
2017 Leveraging manifold learning for extractive broadcast news summarization
abstract
Extractive speech summarization is intended to produce a condensed version of the original spoken document by selecting a few salient sentences from the document and concatenate them together to form a summary. In this paper, we study a novel use of manifold learning techniques for extractive speech summarization. Manifold learning has experienced a surge of research interest in various domains concerned with dimensionality reduction and data representation recently, but has so far been largely under-explored in extractive text or speech summarization. Our contributions in this paper are at least twofold. First, we explore the use of several manifold learning algorithms to capture the latent semantic information of sentences for enhanced extractive speech summarization, including isometric feature mapping (ISOMAP), locally linear embedding (LLE) and Laplacian eigenmap. Second, the merits of our proposed summarization methods and several widely-used methods are extensively analyzed and compared. The empirical results demonstrate the effectiveness of our unsupervised summarization methods, in relation to several state-of-the-art methods. In particular, a synergy of the manifold learning based methods and state-of-the-art methods, such as the integer linear programming (ILP) method, contributes to further gains in summarization performance.
Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen, Hsin-Min Wang, Wen-Lian Hsu
ICASSP2
2017 Exploring the Use of Significant Words Language Modeling for Spoken Document Retrieval
Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen
INTERSPEECH2
2017 Discriminative Autoencoders for Acoustic Modeling
Ming-Han Yang, Hung-Shin Lee, Yu-Ding Lu, Kuan-Yu Chen 0002, Yu Tsao 0001, Berlin Chen, Hsin-Min Wang
INTERSPEECH4
2017 A Position-Aware Language Modeling Framework for Extractive Broadcast News Speech Summarization
abstract
Extractive summarization, a process that automatically picks exemplary sentences from a text (or spoken) document with the goal of concisely conveying key information therein, has seen a surge of attention from scholars and practitioners recently. Using a language modeling (LM) approach for sentence selection has been proven effective for performing unsupervised extractive summarization. However, one of the major difficulties facing the LM approach is to model sentences and estimate their parameters more accurately for each text (or spoken) document. We extend this line of research and make the following contributions in this work. First, we propose a position-aware language modeling framework using various granularities of position-specific information to better estimate the sentence models involved in the summarization process. Second, we explore disparate ways to integrate the positional cues into relevance models through a pseudo-relevance feedback procedure. Third, we extensively evaluate various models originated from our proposed framework and several well-established unsupervised methods. Empirical evaluation conducted on a broadcast news summarization task further demonstrates performance merits of the proposed summarization methods.
Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2016 Learning to Distill: The Essence Vector Modeling Framework
abstract
In the context of natural language processing, representation learning has emerged as a newly active research subject because of its excellent performance in many applications. Learning representations of words is a pioneering study in this school of research. However, paragraph (or sentence and document) embedding learning is more suitable/reasonable for some tasks, such as sentiment classification and document summarization. Nevertheless, as far as we are aware, there is only a dearth of research focusing on launching unsupervised paragraph embedding methods. Classic paragraph embedding methods infer the representation of a given paragraph by considering all of the words occurring in the paragraph. Consequently, those stop or function words that occur frequently may mislead the embedding learning process to produce a misty paragraph representation. Motivated by these observations, our major contributions are twofold. First, we propose a novel unsupervised paragraph embedding method, named the essence vector (EV) model, which aims at not only distilling the most representative information from a paragraph but also excluding the general background information to produce a more informative low-dimensional vector representation for the paragraph. We evaluate the proposed EV model on benchmark sentiment classification and multi-document summarization tasks. The experimental results demonstrate the effectiveness and applicability of the proposed embedding method. Second, in view of the increasing importance of spoken content processing, an extension of the EV model, named the denoising essence vector (D-EV) model, is proposed. The D-EV model not only inherits the advantages of the EV model but also can infer a more robust representation for a given spoken paragraph against imperfect speech recognition. The utility of the D-EV model is evaluated on a spoken document summarization task, confirming the effectiveness of the proposed embedding method in relation to several well-practiced and state-of-the-art summarization methods.
Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang
COLING1
2016 Improved spoken document summarization with coverage modeling techniques
abstract
Extractive summarization aims at selecting a set of indicative sentences from a source document as a summary that can express the major theme of the document. A general consensus on extractive summarization is that both relevance and coverage are critical issues to address. The existing methods designed to model coverage can be characterized by either reducing redundancy or increasing diversity in the summary. Maximal margin relevance (MMR) is a widely-cited method since it takes both relevance and redundancy into account when generating a summary for a given document. In addition to MMR, there is only a dearth of research concentrating on reducing redundancy or increasing diversity for the spoken document summarization task, as far as we are aware. Motivated by these observations, two major contributions are presented in this paper. First, in contrast to MMR, which considers coverage by reducing redundancy, we propose two novel coverage-based methods, which directly increase diversity. With the proposed methods, a set of representative sentences, which not only are relevant to the given document but also cover most of the important sub-themes of the document, can be selected automatically. Second, we make a step forward to plug in several document/sentence representation methods into the proposed framework to further enhance the summarization performance. A series of empirical evaluations demonstrate the effectiveness of our proposed methods.
Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang
ICASSP1
2016 Exploring Word Mover's Distance and Semantic-Aware Embedding Techniques for Extractive Broadcast News Summarization
Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu
INTERSPEECH2
2016 Novel Word Embedding and Translation-based Language Modeling for Extractive Speech Summarization
abstract
Word embedding methods revolve around learning continuous distributed vector representations of words with neural networks, which can capture semantic and/or syntactic cues, and in turn be used to induce similarity measures among words, sentences and documents in context. Celebrated methods can be categorized as prediction-based and count-based methods according to the training objectives and model architectures. Their pros and cons have been extensively analyzed and evaluated in recent studies, but there is relatively less work continuing the line of research to develop an enhanced learning method that brings together the advantages of the two model families. In addition, the interpretation of the learned word representations still remains somewhat opaque. Motivated by the observations and considering the pressing need, this paper presents a novel method for learning the word representations, which not only inherits the advantages of classic word embedding methods but also offers a clearer and more rigorous interpretation of the learned word representations. Built upon the proposed word embedding method, we further formulate a translation-based language modeling framework for the extractive speech summarization task. A series of empirical evaluations demonstrate the effectiveness of the proposed word representation learning and language modeling techniques in extractive speech summarization.
Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Hsin-Hsi Chen
ACM Multimedia1
2016 Extractive speech summarization leveraging convolutional neural network techniques
abstract
Extractive text or speech summarization endeavors to select representative sentences from a source document and assemble them into a concise summary, so as to help people to browse and assimilate the main theme of the document efficiently. The recent past has seen a surge of interest in developing deep learning- or deep neural network-based supervised methods for extractive text summarization. This paper presents a continuation of this line of research for speech summarization and its contributions are three-fold. First, we exploit an effective framework that integrates two convolutional neural networks (CNNs) and a multilayer perceptron (MLP) for summary sentence selection. Specifically, CNNs encode a given document-sentence pair into two discriminative vector embeddings separately, while MLP in turn takes the two embeddings of a document-sentence pair and their similarity measure as the input to induce a ranking score for each sentence. Second, the input of MLP is augmented by a rich set of prosodic and lexical features apart from those derived from CNNs. Third, the utility of our proposed summarization methods and several widely-used methods are extensively analyzed and compared. The empirical results seem to demonstrate the effectiveness of our summarization method in relation to several state-of-the-art methods.
Chun-I Tsai, Hsiao-Tsung Hung, Kuan-Yu Chen 0002, Berlin Chen
SLT3
2016 Exploring the use of unsupervised query modeling techniques for speech recognition and summarization
Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Hsin-Hsi Chen
Speech Commun.1
2015 Incorporating paragraph embeddings and density peaks clustering for spoken document summarization
abstract
Representation learning has emerged as a newly active research subject in many machine learning applications because of its excellent performance. As an instantiation, word embedding has been widely used in the natural language processing area. However, as far as we are aware, there are relatively few studies investigating paragraph embedding methods in extractive text or speech summarization. Extractive summarization aims at selecting a set of indicative sentences from a source document to express the most important theme of the document. There is a general consensus that relevance and redundancy are both critical issues for users in a realistic summarization scenario. However, most of the existing methods focus on determining only the relevance degree between sentences and a given document, while the redundancy degree is calculated by a post-processing step. Based on these observations, three contributions are proposed in this paper. First, we comprehensively compare the word and paragraph embedding methods for spoken document summarization. Next, we propose a novel summarization framework which can take both relevance and redundancy information into account simultaneously. Consequently, a set of representative sentences can be automatically selected through a one-pass process. Third, we further plug in paragraph embedding methods into the proposed framework to enhance the summarization performance. Experimental results demonstrate the effectiveness of our proposed methods, compared to existing state-of-the-art methods.
Kuan-Yu Chen 0002, Kai-Wun Shih, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang
ASRU1
2015 I-vector based language modeling for query representation
abstract
Since more and more multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research subject in the past two decades. Following the research tendency, many efforts have been devoted towards developing indexing and modeling techniques for representing spoken documents, but only few have been made on improving query formulation for better representing users' information needs. The i-vector based language modeling (IVLM) framework, stemming from the state-of-the-art i-vector framework for language identification and speaker recognition, has been proposed and formulated to represent documents in SDR with good promise recently. However, a major challenge of using IVLM for query modeling is that a query usually consists of only a few words; thus, it is hard to learn a reliable representation accordingly. In this paper, we focus our attention on query reformulation and propose three novel methods on top of IVLM to more accurately represent users' information needs. In addition, we also explore the use of multi-levels of index features, including word- and subword-level units, to work in concert with the proposed methods. A series of empirical SDR experiments conducted on the TDT-2 (Topic Detection and Tracking) collection demonstrate the good effectiveness of our proposed methods as compared to existing state-of-the-art methods.
Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen, Hsin-Hsi Chen
ICASSP1
2015 Leveraging word embeddings for spoken document summarization
abstract
Owing to the rapidly growing multimedia content available on the Internet, extractive spoken document summarization, with the purpose of automatically selecting a set of representative sentences from a spoken document to concisely express the most important theme of the document, has been an active area of research and experimentation. On the other hand, word embedding has emerged as a newly favorite research subject because of its excellent performance in many natural language processing (NLP)-related tasks. However, as far as we are aware, there are relatively few studies investigating its use in extractive text or speech summarization. A common thread of leveraging word embeddings in the summarization process is to represent the document (or sentence) by averaging the word embeddings of the words occurring in the document (or sentence). Then, intuitively, the cosine similarity measure can be employed to determine the relevance degree between a pair of representations. Beyond the continued efforts made to improve the representation of words, this paper focuses on building novel and efficient ranking models based on the general word embedding methods for extractive speech summarization. Experimental results demonstrate the effectiveness of our proposed methods, compared to existing state-of-the-art methods.
Kuan-Yu Chen 0002, Shih-Hung Liu, Hsin-Min Wang, Berlin Chen, Hsin-Hsi Chen
INTERSPEECH1
2015 Positional language modeling for extractive broadcast news speech summarization
abstract
Extractive summarization, with the intention of automatically selecting a set of representative sentences from a text (or spoken) document so as to concisely express the most important theme of the document, has been an active area of experimentation and development.A recent trend of research is to employ the language modeling (LM) approach for important sentence selection, which has proven to be effective for performing extractive summarization in an unsupervised fashion.However, one of the major challenges facing the LM approach is how to formulate the sentence models and estimate their parameters more accurately for each text (or spoken) document to be summarized.This paper extends this line of research and its contributions are three-fold.First, we propose a positional language modeling framework using different granularities of position-specific information to better estimate the sentence models involved in summarization.Second, we also explore to integrate the positional cues into relevance modeling through a pseudo-relevance feedback procedure.Third, the utilities of the various methods originated from our proposed framework and several well-established unsupervised methods are analyzed and compared extensively.Empirical evaluations conducted on a broadcast news summarization task seem to demonstrate the performance merits of our summarization methods.
Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu
INTERSPEECH2
2015 A Probabilistic Framework for Chinese Spelling Check
abstract
Chinese spelling check (CSC) is still an unsolved problem today since there are many homonymous or homomorphous characters. Recently, more and more CSC systems have been proposed. To the best of our knowledge, language modeling is one of the major components among these systems because of its simplicity and moderately good predictive power. After deeply analyzing the school of research, we are aware that most of the systems only employ the conventional n -gram language models. The contributions of this article are threefold. First, we propose a novel probabilistic framework for CSC, which naturally combines several important components, such as the substitution model and the language model, to inherit their individual merits as well as to overcome their limitations. Second, we incorporate the topic language models into the CSC system in an unsupervised fashion. The topic language models can capture the long-span semantic information from a word (character) string while the conventional n -gram language models can only preserve the local regularity information. Third, we further integrate Web resources with the proposed framework to enhance the overall performance. Our rigorously empirical experiments demonstrate the consistent and utility performance of the proposed framework in the CSC task.
Kuan-Yu Chen 0002, Hsin-Min Wang, Hsin-Hsi Chen
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2015 Extractive Broadcast News Summarization Leveraging Recurrent Neural Network Language Modeling Techniques
abstract
Extractive text or speech summarization manages to select a set of salient sentences from an original document and concatenate them to form a summary, enabling users to better browse through and understand the content of the document. A recent stream of research on extractive summarization is to employ the language modeling (LM) approach for important sentence selection, which has proven to be effective for performing speech summarization in an unsupervised fashion. However, one of the major challenges facing the LM approach is how to formulate the sentence models and accurately estimate their parameters for each sentence in the document to be summarized. In view of this, our work in this paper explores a novel use of recurrent neural network language modeling (RNNLM) framework for extractive broadcast news summarization. On top of such a framework, the deduced sentence models are able to render not only word usage cues but also long-span structural information of word co-occurrence relationships within broadcast news documents, getting around the need for the strict bag-of-words assumption. Furthermore, different model complexities and combinations are extensively analyzed and compared. Experimental results demonstrate the performance merits of our summarization methods when compared to several well-studied state-of-the-art unsupervised methods.
Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Ea-Ee Jan, Wen-Lian Hsu, Hsin-Hsi Chen
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Combining Relevance Language Modeling and Clarity Measure for Extractive Speech Summarization
abstract
Extractive speech summarization, which purports to select an indicative set of sentences from a spoken document so as to succinctly represent the most important aspects of the document, has garnered much research over the years. In this paper, we cast extractive speech summarization as an ad-hoc information retrieval (IR) problem and investigate various language modeling (LM) methods for important sentence selection. The main contributions of this paper are four-fold. First, we explore a novel sentence modeling paradigm built on top of the notion of relevance, where the relationship between a candidate summary sentence and a spoken document to be summarized is discovered through different granularities of context for relevance modeling. Second, not only lexical but also topical cues inherent in the spoken document are exploited for sentence modeling. Third, we propose a novel clarity measure for use in important sentence selection, which can help quantify the thematic specificity of each individual sentence that is deemed to be a crucial indicator orthogonal to the relevance measure provided by the LM-based methods. Fourth, in an attempt to lessen summarization performance degradation caused by imperfect speech recognition, we investigate making use of different levels of index features for LM-based sentence modeling, including words, subword-level units, and their combination. Experiments on broadcast news summarization seem to demonstrate the performance merits of our methods when compared to several existing well-developed and/or state-of-the-art methods.
Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Leveraging Effective Query Modeling Techniques for Speech Recognition and Summarization
abstract
Statistical language modeling (LM) that purports to quantify the acceptability of a given piece of text has long been an interesting yet challenging research area.In particular, language modeling for information retrieval (IR) has enjoyed remarkable empirical success; one emerging stream of the LM approach for IR is to employ the pseudo-relevance feedback process to enhance the representation of an input query so as to improve retrieval effectiveness.This paper presents a continuation of such a general line of research and the main contribution is threefold.First, we propose a principled framework which can unify the relationships among several widely-used query modeling formulations.Second, on top of the successfully developed framework, we propose an extended query modeling formulation by incorporating critical query-specific information cues to guide the model estimation.Third, we further adopt and formalize such a framework to the speech recognition and summarization tasks.A series of empirical experiments reveal the feasibility of such an LM framework and the performance merits of the deduced models on these two tasks.
Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Ea-Ee Jan, Hsin-Min Wang, Wen-Lian Hsu, Hsin-Hsi Chen
EMNLP1
2014 I-vector based language modeling for spoken document retrieval
abstract
Since more and more multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research subject in the past two decades. The i-vector based framework has been proposed and introduced to language identification (LID) and speaker recognition (SR) tasks recently. The major contribution of the i-vector framework is to reduce a series of acoustic feature vectors of a speech utterance to a low-dimensional vector representation, and then numbers of well-developed postprocessing techniques (such as probabilistic linear discriminative analysis, PLDA) can be readily and effectively used. However, to our best knowledge, there is no research up to date on applying the i-vector framework for SDR or information retrieval (IR). In this paper, we make a step forward to formulate an i-vector based language modeling (IVLM) framework for SDR. Furthermore, we evaluate the proposed IVLM framework with both inductive and transductive learning strategies. We also exploit multi-levels of index features, including word- and subword-level units, in concert with the proposed framework. The results of SDR experiments conducted on the TDT-2 (Topic Detection and Tracking) collection demonstrate the performance merits of our proposed framework when compared to several existing approaches.
Kuan-Yu Chen 0002, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen, Hsin-Hsi Chen
ICASSP1
2014 Effective pseudo-relevance feedback for language modeling in extractive speech summarization
abstract
Extractive speech summarization, aiming to automatically select an indicative set of sentences from a spoken document so as to concisely represent the most important aspects of the document, has become an active area for research and experimentation. An emerging stream of work is to employ the language modeling (LM) framework along with the Kullback-Leibler divergence measure for extractive speech summarization, which can perform important sentence selection in an unsupervised manner and has shown preliminary success. This paper presents a continuation of such a general line of research and its main contribution is two-fold. First, by virtue of pseudo-relevance feedback, we explore several effective sentence modeling formulations to enhance the sentence models involved in the LM-based summarization framework. Second, the utilities of our summarization methods and several widely-used methods are analyzed and compared extensively, which demonstrates the effectiveness of our methods.
Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu
ICASSP2
2014 A recurrent neural network language modeling framework for extractive speech summarization
abstract
Extractive speech summarization, with the purpose of automatically selecting a set of representative sentences from a spoken document so as to concisely express the most important theme of the document, has been an active area of research and development. A recent school of thought is to employ the language modeling (LM) approach for important sentence selection, which has proven to be effective for performing speech summarization in an unsupervised fashion. However, one of the major challenges facing the LM approach is how to formulate the sentence models and accurately estimate their parameters for each spoken document to be summarized. This paper presents a continuation of this general line of research and its contribution is two-fold. First, we propose a novel and effective recurrent neural network language modeling (RNNLM) framework for speech summarization, on top of which the deduced sentence models are able to render not only word usage cues but also long-span structural information of word co-occurrence relationships within spoken documents, getting around the need for the strict bag-of-words assumption. Second, the utilities of the method originated from our proposed framework and several widely-used unsupervised methods are analyzed and compared extensively. A series of experiments conducted on a broadcast news summarization task seem to demonstrate the performance merits of our summarization method when compared to several state-of-the-art existing unsupervised methods.
Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Wen-Lian Hsu, Hsin-Hsi Chen
ICME1
2014 Enhanced language modeling for extractive speech summarization with sentence relatedness information
abstract
Extractive summarization is intended to automatically select a set of representative sentences from a text or spoken document that can concisely express the most important topics of the document. Language modeling (LM) has been proven to be a promising framework for performing extractive summarization in an unsupervised manner. However, there remain two fundamental challenges facing existing LM-based methods. One is how to construct sentence models involved in the LM framework more accurately without resorting to external information sources. The other is how to additionally take into account the sentence-level structural relationships embedded in a document for important sentence selection. To address these two challenges, in this paper we explore a novel approach that generates overlapped clusters to extract sentence relatedness information from the document to be summarized, which can be used not only to enhance the estimation of various sentence models but also to allow for the sentence-level structural relationships for better summarization performance. Further, the utilities of our proposed methods and several state-of-the-art unsupervised methods are analyzed and compared extensively. A series of experiments conducted on a Mandarin broadcast news summarization task demonstrate the effectiveness and viability of our method. Index Terms: speech summarization, language modeling, clustering, relevance, sentence relatedness
Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu
INTERSPEECH2
2014 Leveraging topical and positional cues for language modeling in speech recognition
Hsuan-Sheng Chiu, Kuan-Yu Chen 0002, Berlin Chen
Multim. Tools Appl.2
2014 One-dimensional approximate point set pattern matching with Lp-norm
Hung-Lung Wang, Kuan-Yu Chen 0002
Theor. Comput. Sci.2
2013 Effective pseudo-relevance feedback for language modeling in speech recognition
abstract
A part and parcel of any automatic speech recognition (ASR) system is language modeling (LM), which helps to constrain the acoustic analysis, guide the search through multiple candidate word strings, and quantify the acceptability of the final output hypothesis given an input utterance. Despite the fact that the n-gram model remains the predominant one, a number of novel and ingenious LM methods have been developed to complement or be used in place of the n-gram model. A more recent line of research is to leverage information cues gleaned from pseudo-relevance feedback (PRF) to derive an utterance-regularized language model for complementing the n-gram model. This paper presents a continuation of this general line of research and its main contribution is two-fold. First, we explore an alternative and more efficient formulation to construct such an utterance-regularized language model for ASR. Second, the utilities of various utterance-regularized language models are analyzed and compared extensively. Empirical experiments on a large vocabulary continuous speech recognition (LVCSR) task demonstrate that our proposed language models can offer substantial improvements over the baseline n-gram system, and achieve performance competitive to, or better than, some state-of-the-art language models.
Berlin Chen, Yi-Wen Chen, Kuan-Yu Chen 0002, Ea-Ee Jan
ASRU3
2013 Effective pseudo-relevance feedback for spoken document retrieval
abstract
With the exponential proliferation of multimedia associated with spoken documents, research on spoken document retrieval (SDR) has emerged and attracted much attention in the past two decades. Apart from much effort devoted to developing robust indexing and modeling techniques for representing spoken documents, a recent line of thought targets at the improvement of query modeling for better reflecting the user's information need. Pseudo-relevance feedback is by far the most commonly-used paradigm for query reformulation, which assumes that a small amount of top-ranked feedback documents obtained from the initial round of retrieval are relevant and can be utilized for this purpose. Nevertheless, simply taking all of the top-ranked feedback documents obtained from the initial retrieval for query modeling (reformulation) does not always work well, especially when the top-ranked documents contain much redundant or non-relevant information. In the view of this, we explore in this paper an interesting problem of how to effectively glean useful cues from the top-ranked documents so as to achieve more accurate query modeling. To do this, different kinds of information cues are considered and integrated into the process of feedback document selection so as to improve query effectiveness. Experiments conducted on the TDT (Topic Detection and Tracking) task show the advantages of our retrieval methods for SDR.
Yi-Wen Chen, Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen
ICASSP2
2013 Weighted matrix factorization for spoken document retrieval
abstract
Since more and more multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research subject in the past two decades. Recently, topic models have been successfully used in SDR as well as general information retrieval (IR). These models fall into two categories: probabilistic topic models (PTM) and non-probabilistic topic models (NPTM). One major difference between PTM and NPTM is that the former only takes the words occurring in a document into account, whereas the latter, such as latent semantic analysis (LSA), explicitly models all the words in the vocabulary (including both occurring and non-occurring words). We believe that the non-occurring words can provide additional information that is also useful for SDR. However, to our best knowledge, there is a dearth of work investigating the effectiveness of the non-occurring words for SDR and IR. In order to make effective use of those non-occurring words of documents for semantic analysis, we propose a weighted matrix factorization (WMF) framework, in which the impact of the non-occurring words on the semantic analysis can be modulated properly. The results of SDR experiments conducted on the TDT-2 (Topic Detection and Tracking) collection highlight the performance merits of our proposed framework when compared to several existing topic models.
Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen, Hsin-Hsi Chen
ICASSP1
2013 Sentence modeling for extractive speech summarization
abstract
Extractive speech summarization, aiming to select an indicative set of sentences from a spoken document so as to concisely represent the most important aspects of the document, has emerged as an attractive area of research and experimentation. A recent school of thought is to employ the language modeling (LM) framework along with the Kullback-Leibler (KL) divergence measure for important sentence selection, which has shown preliminary promise for extractive speech summarization. Our work in this paper continues this general line of research in two significant aspects. First, we explore a novel sentence modeling approach built on top of the notion of relevance, where the relationship between a candidate summary sentence and the spoken document to be summarized is discovered through various granularities of context for relevance modeling. Second, not only lexical but also topical cues inherent in the spoken document are exploited for sentence modeling. Experiments on broadcast news summarization seem to demonstrate the performance merits of our methods when compared to several existing methods.
Berlin Chen, Hao-Chin Chang, Kuan-Yu Chen 0002
ICME3
2013 Semantic Naïve Bayes Classifier for Document Classification
How Jing, Yu Tsao 0001, Kuan-Yu Chen 0002, Hsin-Min Wang
IJCNLP3
2013 Incorporating proximity information for relevance language modeling in speech recognition
Yi-Wen Chen, Bo-Han Hao, Kuan-Yu Chen 0002, Berlin Chen
INTERSPEECH3
2013 A Fully Compressed Algorithm for Computing the Edit Distance of Run-Length Encoded Strings
Kuan-Yu Chen 0002, Kun-Mao Chao
Algorithmica1
2013 Leveraging relevance cues for language modeling in speech recognition
Berlin Chen, Kuan-Yu Chen 0002
Inf. Process. Manag.2
2012 Word Relevance Modeling for Speech Recognition
abstract
Language models for speech recognition tend to be brittle across domains, since their performance is vulnerable to changes in the genre or topic of the text on which they are trained. A number of adaptation methods, discovering either lexical co-occurrence or topic cues, have been developed to mitigate this problem with varying degrees of success. Among them, a more recent thread of work is the relevance modeling approach, which has shown promise to capture the lexical co-occurrence relationship between the entire search history and an upcoming word. However, a potential downside to such an approach is the need of resorting to a retrieval procedure to obtain relevance information; this is usually complex and time-consuming for practical applications. In this paper, we propose a word relevance modeling framework, which introduces a novel use of relevance information for dynamic language model adaptation in speech recognition. It not only inherits the merits of several existing techniques but also provides a flexible yet systematic way to render the lexical, topical, and proximity relationships between the search history and the upcoming word. Experiments on large vocabulary continuous speech recognition demonstrate the performance merits of the methods instantiated from this framework when compared to several existing methods.
Kuan-Yu Chen 0002, Hao-Chin Chang, Berlin Chen, Hsin-Min Wang
INTERSPEECH1
2012 Spoken Document Retrieval With Unsupervised Query Modeling Techniques
abstract
Ever-increasing amounts of publicly available multimedia associated with speech information have motivated spoken document retrieval (SDR) to be an active area of intensive research in the speech processing community. Much work has been dedicated to developing elaborate indexing and modeling techniques for representing spoken documents, but only little to improving query formulations for better representing the information needs of users. The latter is critical to the success of a SDR system. In view of this, we present in this paper a novel use of a relevance language modeling framework for SDR. It not only inherits the merits of several existing techniques but also provides a principled way to render the lexical and topical relationships between a query and a spoken document. We further explore various ways to glean both relevance and non-relevance cues from the spoken document collection so as to enhance query modeling in an unsupervised fashion. In addition, we also investigate representing the query and documents with different granularities of index features to work in conjunction with the various relevance and/or non-relevance cues. Empirical evaluations performed on the TDT (Topic Detection and Tracking) collections reveal that the methods derived from our modeling framework hold good promise for SDR and are very competitive with existing retrieval methods.
Berlin Chen, Kuan-Yu Chen 0002, Pei-Ning Chen, Yi-Wen Chen
IEEE Trans. Speech Audio Process.2
2012 Efficient retrieval of approximate palindromes in a run-length encoded string
Kuan-Yu Chen 0002, Ping-Hui Hsu, Kun-Mao Chao
Theor. Comput. Sci.1
2011 Query modeling for spoken document retrieval
abstract
Spoken document retrieval (SDR) has recently become a more interesting research avenue due to increasing volumes of publicly available multimedia associated with speech information. Many efforts have been devoted to developing elaborate indexing and modeling techniques for representing spoken documents, but only few to improving query formulations for better representing the users' information needs. In view of this, we recently presented a language modeling framework exploring a novel use of relevance information cues for improving query effectiveness. Our work in this paper continues this general line of research in two main aspects. We further explore various ways to glean both relevance and non-relevance cues from the spoken document collection so as to enhance query modeling in an unsupervised fashion. Furthermore, we also investigate representing the query and documents with different granularities of index features to work in conjunction with the various relevance and/or non-relevance cues. Experiments conducted on the TDT (Topic Detection and Tracking) SDR task demonstrate the performance merits of the methods instantiated from our retrieval framework when compared to other existing retrieval methods.
Berlin Chen, Pei-Ning Chen, Kuan-Yu Chen 0002
ASRU3
2011 Relevance language modeling for speech recognition
abstract
Language models for speech recognition tend to be brittle across domains, since their performance is vulnerable to changes in the genre or topic of the text on which they are trained. A number of adaptation methods, exploring either lexical co-occurrence or topic cues, have been developed to mitigate this problem with varying degrees of success. In this paper, we study a novel use of relevance information for dynamic language model adaptation in speech recognition. It not only inherits the merits of several existing techniques but also provides a flexible but systematic way to render the lexical and topical relationships between a search history and an upcoming word. Empirical results on large vocabulary continuous speech recognition show that the methods deduced from our framework represent promising alternatives to the other existing language model adaptation methods compared in this paper.
Kuan-Yu Chen 0002, Berlin Chen
ICASSP1
2011 Leveraging Relevance Cues for Improved Spoken Document Retrieval
Pei-Ning Chen, Kuan-Yu Chen 0002, Berlin Chen
INTERSPEECH2
2011 Approximate Point Set Pattern Matching with L p -Norm
Hung-Lung Wang, Kuan-Yu Chen 0002
SPIRE2
2010 A Fully Compressed Algorithm for Computing the Edit Distance of Run-Length Encoded Strings
Kuan-Yu Chen 0002, Kun-Mao Chao
ESA (1)1
2010 Latent topic modeling of word vicinity information for speech recognition
abstract
Topic language models, mostly revolving around the discovery of “word-document” co-occurrence dependence, have attracted significant attention and shown good performance in a wide variety of speech recognition tasks over the years. In this paper, a new topic language model, named word vicinity model (WVM), is proposed to explore the co-occurrence relationship between words, as well as the long-span latent topical information for language model adaptation. A search history is modeled as a composite WVM model for predicting a decoded word. The underlying characteristics and different kinds of model structures are extensively investigated, while the performance of WVM is thoroughly analyzed and verified by comparison with a few existing topic language models. Moreover, we also present a new modeling approach to our recently proposed word topic model (WTM), and design an efficient way to simultaneously extract “word-document” and “word-word” co-occurrence characteristics through the sharing of the same set of latent topics. Experiments on broadcast news transcription seem to demonstrate the utility of the presented models.
Kuan-Yu Chen 0002, Hsuan-Sheng Chiu, Berlin Chen
ICASSP1
2010 Identifying Approximate Palindromes in Run-Length Encoded Strings
Kuan-Yu Chen 0002, Ping-Hui Hsu, Kun-Mao Chao
ISAAC (2)1
2010 Hardness of comparing two run-length encoded strings
Kuan-Yu Chen 0002, Ping-Hui Hsu, Kun-Mao Chao
J. Complex.1
2009 Approximate Matching for Run-Length Encoded Strings Is 3sum-Hard
Kuan-Yu Chen 0002, Ping-Hui Hsu, Kun-Mao Chao
CPM1
2009 Finding All Approximate Gapped Palindromes
Ping-Hui Hsu, Kuan-Yu Chen 0002, Kun-Mao Chao
ISAAC2
2007 On the range maximum-sum segment query problem
Kuan-Yu Chen 0002, Kun-Mao Chao
Discret. Appl. Math.1
2006 Improved algorithms for the k maximum-sums problems
Chih-Huai Cheng, Kuan-Yu Chen 0002, Wen-Chin Tien, Kun-Mao Chao
Theor. Comput. Sci.2
2005 Improved Algorithms for the k Maximum-Sums Problems
Chih-Huai Cheng, Kuan-Yu Chen 0002, Wen-Chin Tien, Kun-Mao Chao
ISAAC2
2005 Optimal algorithms for locating the longest and shortest segments satisfying a sum or an average constraint
Kuan-Yu Chen 0002, Kun-Mao Chao
Inf. Process. Lett.1
2004 On the Range Maximum-Sum Segment Query Problem
Kuan-Yu Chen 0002, Kun-Mao Chao
ISAAC1