VLDB 2026 Research / reviewers in the wild / expert
Hung-Shin Lee
dblp:13/8052
· DBLP profile ↗
33ranked-venue papers
11as first author
15since 2021 · last 2026
0000-0001-7044-9434ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 10 first-author · 13 since 2021Artificial intelligence and machine learning · 22 · 7 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Dialect-Aware Modeling and Conditioning for Low-Resource Taiwanese Hakka Speech Processing
An-Ci Peng, Kuan-Tang Huang, Tien-Hong Lo, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
LREC | 4 |
| 2026 | TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
Cheng-Yeh Yang, Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
LREC | 4 |
| 2026 | DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality AlignmentabstractEvaluating text-to-music (TTM) systems remains expensive because music impression (MI) and text alignment (TA) scores rely on human mean opinion scores (MOS). Most automatic MOS estimators are trained with point-wise regression or distributional classification. These objectives do not directly optimize rank-based metrics and provide weak geometric constraints for cross-modal coherence. To address these gaps, we propose DeRA-MOS, a decoupled optimization framework for TTM evaluation. For MI, we introduce a batch-aware listwise ranking loss that models relative order within each mini-batch and better aligns with evaluation based on Spearman's rank correlation coefficient (SRCC). For TA, we introduce a score-anchored modality alignment loss that maps human scores to target audio-text similarity and regularizes the latent space before fusion. By effectively mitigating the point-wise training mismatch and modality drift, experiments on MusicEval demonstrate that our decoupled framework yields substantial improvements in both MI and TA ranking metrics, establishing a robust paradigm for large-scale TTM evaluation. Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
IEEE Signal Process. Lett. | 2 |
| 2025 | Revealing the Role of Audio Channels in ASR Performance DegradationabstractPre-trained automatic speech recognition (ASR) models have demonstrated strong performance on a variety of tasks. However, their performance can degrade substantially when the input audio comes from different recording channels. While previous studies have demonstrated this phenomenon, it is often attributed to the mismatch between training and testing corpora. This study argues that variations in speech characteristics caused by different recording channels can fundamentally harm ASR performance. To address this limitation, we propose a normalization technique designed to mitigate the impact of channel variation by aligning internal feature representations in the ASR model with those derived from a clean reference channel. This approach significantly improves ASR performance on previously unseen channels and languages, highlighting its ability to generalize across channel and language differences. Kuan-Tang Huang, Hung-Shin Lee, Berlin Chen, Hsin-Min Wang |
ASRU | 3 |
| 2025 | QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation SystemsabstractEvaluating audio generation systems, including text-to-music (TTM), text-to-speech (TTS), and text-to-audio (TTA), remains challenging due to the subjective and multi-dimensional nature of human perception. Existing methods treat mean opinion score (MOS) prediction as a regression problem, but standard regression losses overlook the relativity of perceptual judgments. To address this limitation, we introduce QAMRO, a novel Quality-aware Adaptive Margin Ranking Optimization framework that seamlessly integrates regression objectives from different perspectives, aiming to highlight perceptual differences and prioritize accurate ratings. Our framework leverages pre-trained audio-text models such as CLAP and Audiobox-Aesthetics, and is trained exclusively on the official AudioMOS Challenge 2025 dataset. It demonstrates superior alignment with human evaluations across all dimensions, significantly outperforming robust baseline models. Chien-Chun Wang, Kuan-Tang Huang, Cheng-Yeh Yang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
ASRU | 4 |
| 2025 | Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech RecognitionabstractWhile pre-trained automatic speech recognition (ASR) systems demonstrate impressive performance on matched domains, their performance often degrades when confronted with channel mismatch stemming from unseen recording environments and conditions. To mitigate this issue, we propose a novel channel-aware data simulation method for robust ASR training. Our method harnesses the synergistic power of channel-extractive techniques and generative adversarial networks (GANs). We first train a channel encoder capable of extracting embeddings from arbitrary audio. On top of this, channel embeddings are extracted using a minimal amount of target-domain data and used to guide a GAN-based speech synthesizer. This synthesizer generates speech that faithfully preserves the phonetic content of the input while mimicking the channel characteristics of the target domain. We evaluate our method on the challenging Hakka Across Taiwan (HAT) and Taiwanese Across Taiwan (TAT) corpora, achieving relative character error rate (CER) reductions of 20.02% and 9.64%, respectively, compared to the baselines. These results highlight the efficacy of our channel-aware data simulation method for bridging the gap between source- and target-domain acoustics. Chien-Chun Wang, Cheng-Kang Chou, Hung-Shin Lee, Berlin Chen, Hsin-Min Wang |
ICASSP | 4 |
| 2024 | Effective Noise-Aware Data Simulation For Domain-Adaptive Speech Enhancement Leveraging Dynamic Stochastic PerturbationabstractCross-domain speech enhancement (SE) is often faced with severe challenges due to the scarcity of noise and background information in an unseen target domain, leading to a mismatch between training and test conditions. This study puts forward a novel data simulation method to address this issue, leveraging noise-extractive techniques and generative adversarial networks (GANs) with only limited target noisy speech data. Notably, our method employs a noise encoder to extract noise embeddings from target-domain data. These embeddings aptly guide the generator to synthesize utterances acoustically fitted to the target domain while authentically preserving the phonetic content of the input clean speech. Furthermore, we introduce the notion of dynamic stochastic perturbation, which can inject controlled perturbations into the noise embeddings during inference, thereby enabling the model to generalize well to unseen noise conditions. Experiments on the VoiceBank-DEMAND benchmark dataset demonstrate that our domain-adaptive SE method outperforms an existing strong baseline based on data simulation. Chien-Chun Wang, Hung-Shin Lee, Berlin Chen, Hsin-Min Wang |
SLT | 3 |
| 2023 | A Training and Inference Strategy Using Noisy and Enhanced Speech as Target for Speech Enhancement without Clean Speech
Yao-Fei Cheng, Hung-Shin Lee, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 3 |
| 2023 | Multi-Target Extractor and Detector for Unknown-Number Speaker DiarizationabstractStrong representations of target speakers can help extract important information about speakers and detect corresponding temporal regions in multi-speaker conversations. In this study, we propose a neural architecture that simultaneously extracts speaker representations consistent with the speaker diarization objective and detects the presence of each speaker on a frame-by-frame basis regardless of the number of speakers in a conversation. A speaker representation (called z-vector) extractor and a time-speaker contextualizer, implemented by a residual network and processing data in both temporal and speaker dimensions, are integrated into a unified framework. Tests on the CALLHOME corpus show that our model outperforms most of the methods proposed so far. Evaluations in a more challenging case with simultaneous speakers ranging from 2 to 7 show that our model achieves 6.4% to 30.9% relative diarization error rate reductions over several typical baselines. Chin-Yi Cheng, Hung-Shin Lee, Yu Tsao 0001, Hsin-Min Wang |
IEEE Signal Process. Lett. | 2 |
| 2022 | Chain-based Discriminative Autoencoders for Speech RecognitionabstractIn our previous work, we proposed a discriminative autoencoder (DcAE) for speech recognition.DcAE combines two training schemes into one.First, since DcAE aims to learn encoderdecoder mappings, the squared error between the reconstructed speech and the input speech is minimized.Second, in the code layer, frame-based phonetic embeddings are obtained by minimizing the categorical cross-entropy between ground truth labels and predicted triphone-state scores.DcAE is developed based on the Kaldi toolkit by treating various TDNN models as encoders.In this paper, we further propose three new versions of DcAE.First, a new objective function that considers both categorical cross-entropy and mutual information between ground truth and predicted triphone-state sequences is used.The resulting DcAE is called a chain-based DcAE (c-DcAE).For application to robust speech recognition, we further extend c-DcAE to hierarchical and parallel structures, resulting in hc-DcAE and pc-DcAE.In these two models, both the error between the reconstructed noisy speech and the input noisy speech and the error between the enhanced speech and the reference clean speech are taken into the objective function.Experimental results on the WSJ and Aurora-4 corpora show that our DcAE models outperform baseline systems. Hung-Shin Lee, Pin-Tuan Huang, Yao-Fei Cheng, Hsin-Min Wang |
INTERSPEECH | 1 |
| 2022 | Disentangling the Impacts of Language and Channel Variability on Speech Separation NetworksabstractBecause the performance of speech separation is excellent for speech in which two speakers completely overlap, research attention has been shifted to dealing with more realistic scenarios.However, domain mismatch between training/test situations due to factors, such as speaker, content, channel, and environment, remains a severe problem for speech separation.Speaker and environment mismatches have been studied in the existing literature.Nevertheless, there are few studies on speech content and channel mismatches.Moreover, the impacts of language and channel in these studies are mostly tangled.In this study, we create several datasets for various experiments.The results show that the impacts of different languages are small enough to be ignored compared to the impacts of different channels.In our experiments, training on data recorded by Android phones leads to the best generalizability.Moreover, we provide a new solution for channel mismatch by evaluating projection, where the channel similarity can be measured and used to effectively select additional training data to improve the performance of in-the-wild test data. Fan-Lin Wang, Hung-Shin Lee, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 2 |
| 2021 | Melody Harmonization Using Orderless Nade, Chord Balancing, and Blocked Gibbs SamplingabstractCoherence and interestingness are two criteria for evaluating the performance of melody harmonization, which aims to generate a chord progression from a symbolic melody. In this study, we apply the concept of orderless NADE, which takes the melody and its partially masked chord sequence as the input of the BiLSTM-based networks to learn the masked ground truth, to the training process. In addition, the class weights are used to compensate for some reasonable chord labels that are rarely seen in the training set. Consistent with the stochasticity in training, blocked Gibbs sampling with proper numbers of masking/generating loops is used in the inference phase to progressively trade the coherence of the generated chord sequence off against its interestingness. The experiments were conducted on a dataset of 18,005 melody/chord pairs. Our proposed model outperforms the state-of-the-art system MTHarmonizer in five of six different objective metrics based on chord/melody harmonicity and chord progression. The subjective test results with more than 100 participants also show the superiority of our model. Chung-En Sun, Yi-Wei Chen, Hung-Shin Lee, Yen-Hsing Chen, Hsin-Min Wang |
ICASSP | 3 |
| 2021 | AlloST: Low-Resource Speech Translation Without Source TranscriptionabstractThe end-to-end architecture has made promising progress in speech translation (ST). However, the ST task is still challenging under low-resource conditions. Most ST models have shown unsatisfactory results, especially in the absence of word information from the source speech utterance. In this study, we survey methods to improve ST performance without using source transcription, and propose a learning framework that utilizes a language-independent universal phone recognizer. The framework is based on an attention-based sequence-to-sequence model, where the encoder generates the phonetic embeddings and phone-aware acoustic representations, and the decoder controls the fusion of the two embedding streams to produce the target token sequence. In addition to investigating different fusion strategies, we explore the specific usage of byte pair encoding (BPE), which compresses a phone sequence into a syllable-like segmented sequence. Due to the conversion of symbols, a segmented sequence represents not only pronunciation but also language-dependent information lacking in phones. Experiments conducted on the Fisher Spanish-English and Taigi-Mandarin drama corpora show that our method outperforms the conformer-based baseline, and the performance is close to that of the existing best method using source transcription. Yao-Fei Cheng, Hung-Shin Lee, Hsin-Min Wang |
Interspeech | 2 |
| 2021 | Dual-Path Filter Network: Speaker-Aware Modeling for Speech SeparationabstractSpeech separation has been extensively studied to deal with the cocktail party problem in recent years.All related approaches can be divided into two categories: time-frequency domain methods and time domain methods.In addition, some methods try to generate speaker vectors to support source separation.In this study, we propose a new model called dualpath filter network (DPFN).Our model focuses on the postprocessing of speech separation to improve speech separation performance.DPFN is composed of two parts: the speaker module and the separation module.First, the speaker module infers the identities of the speakers.Then, the separation module uses the speakers' information to extract the voices of individual speakers from the mixture.DPFN constructed based on DPRNN-TasNet is not only superior to DPRNN-TasNet, but also avoids the problem of permutation-invariant training (PIT). Fan-Lin Wang, Yu-Huai Peng, Hung-Shin Lee, Hsin-Min Wang |
Interspeech | 3 |
| 2021 | Relational Data Selection for Data Augmentation of Speaker-Dependent Multi-Band MelGAN VocoderabstractNowadays, neural vocoders can generate very high-fidelity speech when a bunch of training data is available.Although a speaker-dependent (SD) vocoder usually outperforms a speaker-independent (SI) vocoder, it is impractical to collect a large amount of data of a specific target speaker for most realworld applications.To tackle the problem of limited target data, a data augmentation method based on speaker representation and similarity measurement of speaker verification is proposed in this paper.The proposed method selects utterances that have similar speaker identity to the target speaker from an external corpus, and then combines the selected utterances with the limited target data for SD vocoder adaptation.The evaluation results show that, compared with the vocoder adapted using only limited target data, the vocoder adapted using augmented data improves both the quality and similarity of synthesized speech. Yi-Chiao Wu, Cheng-Hung Hu, Hung-Shin Lee, Yu-Huai Peng, Wen-Chin Huang, Yu Tsao 0001, Hsin-Min Wang, Tomoki Toda |
Interspeech | 3 |
| 2020 | Subspace-Based Representation and Learning for Phonotactic Spoken Language RecognitionabstractPhonotactic constraints can be employed to distinguish languages by representing a speech utterance as a multinomial distribution or phone events. In the present study, we propose a new learning mechanism based on subspace-based representation, which can extract concealed phonotactic structures from utterances, for language verification and dialect/accent identification. The framework mainly involves two successive parts. The first part involves subspace construction. Specifically, it decodes each utterance into a sequence of vectors filled with phone-posteriors and transforms the vector sequence into a linear orthogonal subspace based on low-rank matrix factorization or dynamic linear modeling. The second part involves subspace learning based on kernel machines, such as support vector machines and the newly developed subspace-based neural networks (SNNs). The input layer of SNNs is specifically designed for the sample represented by subspaces. The topology ensures that the same output can be derived from identical subspaces by modifying the conventional feed-forward pass to fit the mathematical definition of subspace similarity. Evaluated on the “General LR” test of NIST LRE 2007, the proposed method achieved up to 52%, 46%, 56%, and 27% relative reductions in equal error rates over the sequence-based PPR-LM, PPR-VSM, and PPR-IVEC methods and the lattice-based PPR-LM method, respectively. Furthermore, on the dialect/accent identification task of NIST LRE 2009, the SNN-based system performed better than the aforementioned four baseline methods. Hung-Shin Lee, Yu Tsao 0001, Shyh-Kang Jeng, Hsin-Min Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Spoken Multiple-Choice Question Answering Using Multimodal Convolutional Neural NetworksabstractIn a spoken multiple-choice question answering (MCQA) task, where passages, questions, and choices are given in the form of speech, usually only the auto-transcribed text is considered in system development. The acoustic-level information may contain useful cues for answer prediction. However, to the best of our knowledge, only a few studies focus on using the acoustic-level information or fusing the acoustic-level information with the text-level information for a spoken MCQA task. Therefore, this paper presents a hierarchical multistage multimodal (HMM) framework based on convolutional neural networks (CNNs) to integrate text- and acoustic-level statistics into neural modeling for spoken MCQA. Specifically, the acoustic-level statistics are expected to offset text inaccuracies caused by automatic speech recognition (ASR) systems or representation inadequacy lurking in word embedding generators, thereby making the spoken MCQA system robust. In the proposed HMM framework, two modalities are first manipulated to separately derive the acoustic- and text-level representations for the passage, question, and choices. Next, these clever features are jointly involved in inferring the relationships among the passage, question, and choices. Then, a final representation is derived for each choice, which encodes the relationship of the choice to the passage and question. Finally, the most likely answer is determined based on the individual final representations of all choices. Evaluated on the data of “Formosa Grand Challenge - Talk to AI”, a Mandarin Chinese spoken MCQA contest held in 2018, the proposed HMM framework achieves remarkable improvements in accuracy over the text-only baseline. Shang-Bao Luo, Hung-Shin Lee, Kuan-Yu Chen 0002, Hsin-Min Wang |
ASRU | 2 |
| 2019 | Exploring the Encoder Layers of Discriminative Autoencoders for LVCSR
Pin-Tuan Huang, Hung-Shin Lee, Syu-Siang Wang, Kuan-Yu Chen 0002, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 2 |
| 2017 | Discriminative autoencoders for speaker verificationabstractThis paper presents a learning and scoring framework based on neural networks for speaker verification. The framework employs an autoencoder as its primary structure while three factors are jointly considered in the objective function for speaker discrimination. The first one, relating to the sample reconstruction error, makes the structure essentially a generative model, which benefits to learn most salient and useful properties of the data. Functioning in the middlemost hidden layer, the other two attempt to ensure that utterances spoken by the same speaker are mapped into similar identity codes in the speaker discriminative subspace, where the dispersion of all identity codes are maximized to some extent so as to avoid the effect of over-concentration. Finally, the decision score of each utterance pair is simply computed by cosine similarity of their identity codes. Dealing with utterances represented by i-vectors, the results of experiments conducted on the male portion of the core task in the NIST 2010 Speaker Recognition Evaluation (SRE) significantly demonstrate the merits of our approach over the conventional PLDA method. Hung-Shin Lee, Yu-Ding Lu, Chin-Cheng Hsu, Yu Tsao 0001, Hsin-Min Wang, Shyh-Kang Jeng |
ICASSP | 1 |
| 2017 | Discriminative Autoencoders for Acoustic Modeling
Ming-Han Yang, Hung-Shin Lee, Yu-Ding Lu, Kuan-Yu Chen 0002, Yu Tsao 0001, Berlin Chen, Hsin-Min Wang |
INTERSPEECH | 2 |
| 2016 | Minimization of Regression and Ranking Losses with Shallow Neural Networks on Automatic Sincerity Evaluation
Hung-Shin Lee, Yu Tsao 0001, Chi-Chun Lee, Hsin-Min Wang, Wei-Chen Chen, Shan-Wen Hsiao, Shyh-Kang Jeng |
INTERSPEECH | 1 |
| 2014 | I-vector based language modeling for spoken document retrievalabstractSince more and more multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research subject in the past two decades. The i-vector based framework has been proposed and introduced to language identification (LID) and speaker recognition (SR) tasks recently. The major contribution of the i-vector framework is to reduce a series of acoustic feature vectors of a speech utterance to a low-dimensional vector representation, and then numbers of well-developed postprocessing techniques (such as probabilistic linear discriminative analysis, PLDA) can be readily and effectively used. However, to our best knowledge, there is no research up to date on applying the i-vector framework for SDR or information retrieval (IR). In this paper, we make a step forward to formulate an i-vector based language modeling (IVLM) framework for SDR. Furthermore, we evaluate the proposed IVLM framework with both inductive and transductive learning strategies. We also exploit multi-levels of index features, including word- and subword-level units, in concert with the proposed framework. The results of SDR experiments conducted on the TDT-2 (Topic Detection and Tracking) collection demonstrate the performance merits of our proposed framework when compared to several existing approaches. Kuan-Yu Chen 0002, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen, Hsin-Hsi Chen |
ICASSP | 2 |
| 2014 | Speaker verification using kernel-based binary classifiers with binary operation derived featuresabstractIn this paper, we study the use of two kinds of kernel-based discriminative models, namely support vector machine (SVM) and deep neural network (DNN), for speaker verification. We treat the verification task as a binary classification problem, in which a pair of two utterances, each represented by an i-vector, is assumed to belong to either the “within-speaker” group or the “between-speaker” group. To solve the problem, we employ various binary operations to retain the basic relationship between any pair of i-vectors to form a single vector for training the discriminative models. This study also investigates the correlation of achievable performances with the number of training pairs and the various combinations of basic binary operations, using the SVM and DNN binary classifiers. The experiments are conducted on the male portion of the core task in the NIST 2005 Speaker Recognition Evaluation (SRE), and the results are competitive or even better, in terms of normalized decision cost function (minDCF) and equal error rate (EER), while compared to other non-probabilistic based models, such as the conventional speaker SVMs and the LDA-based cosine distance scoring. Hung-Shin Lee, Yu Tso, Yun-Fan Chang, Hsin-Min Wang, Shyh-Kang Jeng |
ICASSP | 1 |
| 2014 | Ensemble of machine learning algorithms for cognitive and physical speaker load detectionabstractWe present our methods and results on participating in the Interspeech 2014 Computational Paralinguistics ChallengE (ComParE) of which the goal is to detect certain type of load of a speaker using acoustic features. There are in total seven classification models contributing to our final prediction, namely, neural network with rectified linear unit and dropout (ReLUNet), conditional restricted Boltzmann machine (CRBM), logistic regression (LR), support vector machine (SVM), Gaussian discriminant analysis (GDA), k-nearest neighbors (KNN), and random forest (RF). When linearly blending the predictions of these models, we are able to get significant improvements over the challenge baseline. Index Terms: Physical Load Detection, Cognitive Load Detection, Neural Network, Classification Models How Jing, Ting-Yao Hu, Hung-Shin Lee, Wei-Chen Chen, Chi-Chun Lee, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 3 |
| 2014 | Clustering-based i-vector formulation for speaker recognition
Hung-Shin Lee, Yu Tsao 0001, Hsin-Min Wang, Shyh-Kang Jeng |
INTERSPEECH | 1 |
| 2013 | Subspace-based phonotactic language recognition using multivariate dynamic linear modelsabstractPhonotactics, dealing with permissible phone patterns and their frequencies of occurrence in a specific language, is acknowledged to be related to spoken language recognition (SLR). With the assistance of phone recognizers, each speech utterance can be decoded into an ordered sequence of phone vectors filled with likelihood scores contributed by all possible phone models. In this paper, we propose a novel approach to dig the concealed phonotactic structure out of the phone-likelihood vectors through a kind of multivariate time series analysis: dynamic linear models (DLM). In these models, treating the generation of phone patterns in each utterance as a dynamic system, the relationship between adjacent vectors is linearly and time-invariantly modeled, and unobserved states are introduced to capture a temporal coherence intrinsic in the system. Each utterance expressed by the DLM is further transformed into a fixed-dimensional linear subspace so that well-developed distance measures between two subspaces can be applied to linear discriminant analysis (LDA) in a dissimilarity-based fashion. The results of SLR experiments on the OGI-TS corpus demonstrate that the proposed framework outperforms the well-known vector space modeling (VSM)-based methods and achieves comparable performance to our previous subspace-based method. Hung-Shin Lee, Yu-Chin Shih, Hsin-Min Wang, Shyh-Kang Jeng |
ICASSP | 1 |
| 2012 | Subspace-Based Feature Representation and Learning for Language RecognitionabstractThis paper presents a novel subspace-based approach for phonotactic language recognition. The whole framework is divided into two parts: the speech feature representation and the subspacebased learning algorithm. First, the phonetic information as well as the contextual relationship, possessed by spoken utterances, are more abundantly retrieved by likelihood computation and feature concatenation through the decoding processed by an automatic speech recognizer. It is assumed that the extracted phone frames reside in a lower dimensional eigen-subspace, in which the structure of data can be approximately captured. Each utterance is further represented by a fixed-dimensional linear subspace. Second, to measure the similarity between two utterances, suitable non-Euclidean metrics are explored and applied to non-linear discriminant analysis in a kernel fashion, followed by a back-end classifier, such as the k-nearest neighbor (K-NN) classifier. The results of experiments on the OGI-TS database demonstrate that the proposed framework outperforms the well-known vector space modeling based method with relative reductions of 38.90% and 27.13% on the 1-to-50-second and 3-second data sets respectively in equal error rate (EER). Yu-Chin Shih, Hung-Shin Lee, Hsin-Min Wang, Shyh-Kang Jeng |
INTERSPEECH | 2 |
| 2010 | A Discriminative and Heteroscedastic Linear Feature Transformation for Multiclass ClassificationabstractThis paper presents a novel discriminative feature transformation, named full-rank generalized likelihood ratio discriminant analysis (fGLRDA), on the grounds of the likelihood ratio test (LRT). fGLRDA attempts to seek a feature space, which is linearly isomorphic to the original n-dimensional feature space and is characterized by a full-rank (n×n) transformation matrix, under the assumption that all the class-discrimination information resides in a d-dimensional subspace (d <; n), through making the most confusing situation, described by the null hypothesis, as unlikely as possible to happen without the homoscedastic assumption on class distributions. Our experimental results demonstrate that fGLRDA can yield moderate performance improvements over other existing methods, such as linear discriminant analysis (LDA) for the speaker identification task. Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
ICPR | 1 |
| 2010 | Exploiting semantic associative information in topic modelingabstractTopic modeling has been widely applied in a variety of text modeling tasks as well as in speech recognition systems for effectively capturing the semantic and statistic information in documents or speech utterances. Most topic models rely on the bag-of-words assumption that results in learned latent topics composed of lists of individual words. Unfortunately, these words may convey topical information but lack accurate semantic knowledge of the text. In this paper, we present the semantic associative topic model, where the concept of the semantic association terms is extended to topic modeling, which provides guidance on modeling the semantic associations that occur among single words by expressing a document as an association of multiple words. Further, the pointwise KL-divergence metric is used to measure the significance of the association. We also integrate original PLSA and SATM models, which have mixed feature representations. Experimental results on WSJ and AP datasets show that the proposed approaches achieved higher performance compared to other methods. Meng-Sung Wu, Hung-Shin Lee, Hsin-Min Wang |
SLT | 2 |
| 2009 | Generalized likelihood ratio discriminant analysisabstractIn the past several decades, classifier-independent front-end feature extraction, where the derivation of acoustic features is lightly associated with the back-end model training or classification, has been prominently used in various pattern recognition tasks, including automatic speech recognition (ASR). In this paper, we present a novel discriminative feature transformation, named generalized likelihood ratio discriminant analysis (GLRDA), on the basis of the likelihood ratio test (LRT). It attempts to seek a lower dimensional feature subspace by making the most confusing situation, described by the null hypothesis, as unlikely to happen as possible without the homoscedastic assumption on class distributions. We also show that the classical linear discriminant analysis (LDA) and its well-known extension - heteroscedastic linear discriminant analysis (HLDA) can be regarded as two special cases of our proposed method. The empirical class confusion information can be further incorporated into GLRDA for better recognition performance. Experimental results demonstrate that GLRDA and its variant can yield moderate performance improvements over HLDA and LDA for the large vocabulary continuous speech recognition (LVCSR) task. Hung-Shin Lee, Berlin Chen |
ASRU | 1 |
| 2009 | Empirical error rate minimization based linear discriminant analysisabstractLinear discriminant analysis (LDA) is designed to seek a linear transformation that projects a data set into a lower-dimensional feature space while retaining geometrical class separability. However, LDA cannot always guarantee better classification accuracy. One of the possible reasons lies in that its formulation is not directly associated with the classification error rate, so that it is not necessarily suited for the allocation rule governed by a given classifier, such as that employed in automatic speech recognition (ASR). In this paper, we extend the classical LDA by leveraging the relationship between the empirical classification error rate and the Mahalanobis distance for each respective class pair, and modify the original between-class scatter from a measure of the squared Euclidean distance to the pairwise empirical classification accuracy for each class pair, while preserving the lightweight solvability and taking no distributional assumption, just as what LDA does. Experimental results seem to demonstrate that our approach yields moderate improvements over LDA on the large vocabulary continuous speech recognition (LVCSR) task. Hung-Shin Lee, Berlin Chen |
ICASSP | 1 |
| 2008 | Linear discriminant feature extraction using weighted classification confusion informationabstractLinear discriminant analysis (LDA) can be viewed as a twostage procedure geometrically. The first stage conducts an orthogonal and whitening transformation of the variables. The second stage involves a principal component analysis (PCA) on the transformed class means, which is intended to maximize the class separability along the principal axes. In this paper, we demonstrate that the second stage does not necessarily guarantee better classification accuracy. Furthermore, we propose a generalization of LDA, weighted LDA (WLDA), by integrating the empirical classification confusion information between each class pair, such that the separability and the classification error rate can be taken into consideration simultaneously. WLDA can be efficiently solved by a lightweight eigen-decomposition and easily combined with other modifications to the LDA criterion. The experiment results show that WLDA can yield an relative character error reduction of 4.6% over LDA on the Mandarin LVCSR task. Hung-Shin Lee, Berlin Chen |
INTERSPEECH | 1 |
| 2007 | Training data selection for improving discriminative training of acoustic modelsabstractThis paper considers training data selection for discriminative training of acoustic models for broadcast news speech recognition. Three novel data selection approaches were proposed. First, the average phone accuracy over all hypothesized word sequences in the word lattice of a training utterance was utilized for utterancelevel data selection. Second, phone-level data selection based on the difference between the expected accuracy of a phone arc and the average phone accuracy of the word lattice was investigated. Finally, frame-level data selection based on the normalized frame-level entropy of Gaussian posterior probabilities obtained from the word lattice was explored. The underlying characteristics of the presented approaches were extensively investigated and their performance was verified by comparison with the standard discriminative training approaches. Experiments conducted on the Mandarin broadcast news collected in Taiwan shown that both phone- and frame-level data selection could achieve slight but consistent improvements over the baseline systems at lower training iterations. Shih-Hung Liu, Fang-Hui Chu, Shih-Hsiang Lin, Hung-Shin Lee, Berlin Chen |
ASRU | 4 |