EDBT 2026 Demo / reviewers in the wild / expert
Jeih-Weih Hung
dblp:49/272 · also Jeih-weih Hung
· DBLP profile ↗
47ranked-venue papers
17as first author
8since 2021 · last 2025
0000-0001-9366-3070ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 13 first-author · 7 since 2021Artificial intelligence and machine learning · 27 · 9 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SincQDR-VAD: A Noise-Robust Voice Activity Detection Framework Leveraging Learnable Filters and Ranking-Aware OptimizationabstractVoice activity detection (VAD) is essential for speech-driven applications, but remains far from perfect in noisy and resource-limited environments. Existing methods often lack robustness to noise, and their frame-wise classification losses are only loosely coupled with the evaluation metric of VAD. To address these challenges, we propose SincQDR-VAD, a compact and robust framework that combines a Sinc-extractor front-end with a novel quadratic disparity ranking loss. The Sinc-extractor uses learnable bandpass filters to capture noise-resistant spectral features, while the ranking loss optimizes the pairwise score order between speech and non-speech frames to improve the area under the receiver operating characteristic curve (AUROC). A series of experiments conducted on representative benchmark datasets show that our framework considerably improves both AUROC and $F_{2}$-Score, while using only $69 \%$ of the parameters compared to prior arts, confirming its efficiency and practical viability. Chien-Chun Wang, En-Lun Yu, Jeih-Weih Hung, Shih-Chieh Huang, Berlin Chen |
ASRU | 3 |
| 2025 | Flexible VAD-PVAD Transition: A Detachable PVAD Module for Dynamic Encoder RNN VAD
En-Lun Yu, Chien-Chun Wang, Jeih-Weih Hung, Shih-Chieh Huang, Berlin Chen |
INTERSPEECH | 3 |
| 2024 | What Do Neural Networks Listen to? Exploring the Crucial Bands in Speech Enhancement Using SINC-ConvolutionabstractThis study introduces a reformed Sinc-convolution (Sincconv) framework tailored for the encoder component of deep networks for speech enhancement (SE). The reformed Sinc-conv, based on parametrized sinc functions as band-pass filters, offers notable advantages in terms of training efficiency, filter diversity, and interpretability. The reformed Sinc-conv is evaluated in conjunction with various SE models, showcasing its ability to boost SE performance. Furthermore, the reformed Sincconv provides valuable insights into the specific frequency components that are prioritized in an SE scenario. This opens up a new direction of SE research and improving our knowledge of their operating dynamics. Kuan-Hsun Ho, Jeih-Weih Hung, Berlin Chen |
ICASSP | 2 |
| 2024 | Speaker Conditional Sinc-Extractor for Personal VAD
En-Lun Yu, Kuan-Hsun Ho, Jeih-Weih Hung, Shih-Chieh Huang, Berlin Chen |
INTERSPEECH | 3 |
| 2023 | Improving Speech Enhancement Performance by Leveraging Contextual Broad Phonetic Class InformationabstractPrevious studies have confirmed that by augmenting acoustic features with the place/manner of articulatory features, the speech enhancement (SE) process can be guided to consider the broad phonetic properties of the input speech when performing enhancement to attain performance improvements. In this paper, we explore the contextual information of articulatory attributes as additional information to further benefit SE. More specifically, we propose to improve the SE performance by leveraging losses from an end-to-end automatic speech recognition (E2E-ASR) model that predicts the sequence of broad phonetic classes (BPCs). We also developed multi-objective training with ASR and perceptual losses to train the SE system based on a BPC-based E2E-ASR. Experimental results from speech denoising, speech dereverberation, and impaired speech enhancement tasks confirmed that contextual BPC information improves SE performance. Moreover, the SE model trained with the BPC-based E2E-ASR outperforms that with the phoneme-based E2E-ASR. The results suggest that objectives with misclassification of phonemes by the ASR system may lead to imperfect feedback, and BPC could be a potentially better choice. Finally, it is noted that combining the most-confusable phonetic targets into the same BPC when calculating the additional objective can effectively improve the SE performance. Yen-Ju Lu, Chia-Yu Chang, Ching-Feng Liu, Jeih-Weih Hung, Shinji Watanabe 0001, Yu Tsao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Adaptive-FSN: Integrating Full-Band Extraction and Adaptive Sub-Band Encoding for Monaural Speech EnhancementabstractAn important more recent thread of speech enhancement work is to utilize fine-grinded local spectral patterns with sub-band processing that complement full-band features nicely. To extend the efficacy of sub-band spectral information, we propose Adaptive-FSN, a fully convolutional real-time speech enhancement framework, to dynamically acquire a sub-band embedding within a wide range of sub-band frequencies. We exploit an adaptive subband encoder to portray sub-band processing that encapsulates a wide range of sub-band units. Then we build this effective sub-band embedding with a Conformer-based structure and multi-view attention. As for the full-band features, we make use of the FullSubNet+ architecture with its full-band extractor to get global spectral information. Finally, a Conformer-based fusion model combines the above information sources to predict the complex ideal ratio mask (cIRM). Experimental results on the VoiceBank-DEMAND benchmark task reveal that this novel framework outperforms FullSubNet+ by promoting the quality of processed utterances and reducing the implementation complexity for faster real-time computation. Yu-sheng Tsao, Kuan-Hsun Ho, Jeih-Weih Hung, Berlin Chen |
SLT | 3 |
| 2021 | TENET: A Time-Reversal Enhancement Network for Noise-Robust ASRabstractDue to the unprecedented breakthroughs brought about by deep learning, speech enhancement (SE) techniques have been developed rapidly and play an important role prior to acoustic modeling so as to mitigate noise effects on speech. To increase the perceptual quality of speech, the current state-of-the-art in the realm of SE adopts adversarial training by connecting an objective metric to the discriminator. However, there is no guarantee that optimizing the perceptual quality of speech will necessarily lead to improved automatic speech recognition (ASR) performance. In this study, we present TENET††Inspired by the movie - TENET, Christopher Nolan, 2020., **Some of the enhanced audio samples can be found from https://fuann.github.io/TENET., a novel Time-reversal Enhancement NETwork, which leverages the transformation of an input noisy signal itself, i.e., the time-reversed version, in conjunction with a Siamese network and a complex dual-path Transformer to promote SE performance for noise-robust ASR. Extensive experiments conducted on the Voicebank-DEMAND dataset show that TENET can achieve stellar results compared to a few top-of-the-line methods in terms of both SE and ASR evaluation metrics. To demonstrate the model generalization ability, we further evaluate TENET on the test set of scenarios contaminated with unseen noise, and the results also confirm the superiority of this promising method. Fu-An Chao, Shao-Wei Fan-Jiang, Bi-Cheng Yan, Jeih-Weih Hung, Berlin Chen |
ASRU | 4 |
| 2021 | Cross-Domain Single-Channel Speech Enhancement Model with BI-Projection Fusion Module for Noise-Robust ASRabstractIn recent decades, many studies have suggested that phase information is crucial for speech enhancement (SE), and time-domain single-channel speech enhancement techniques have shown promise in noise suppression and robust automatic speech recognition (ASR). This paper presents a continuation of the above lines of research and explores two effective SE methods that consider phase information in time domain and frequency domain of speech signals, respectively. Going one step further, we put forward a novel cross-domain speech enhancement model and a bi-projection fusion (BPF) mechanism for noise-robust ASR. To evaluate the effectiveness of our proposed method, we conduct an extensive set of experiments on the publicly-available Aishell-1 Mandarin benchmark speech corpus. The evaluation results confirm the superiority of our proposed method in relation to a few current top-of-the-line time-domain and frequency-domain SE methods in both enhancement and ASR evaluation metrics for the test set of scenarios contaminated with seen and unseen noise, respectively. Fu-An Chao, Jeih-Weih Hung, Berlin Chen |
ICME | 2 |
| 2020 | Incorporating Broad Phonetic Information for Speech EnhancementabstractIn noisy conditions, knowing speech contents facilitates listeners to more effectively suppress background noise components and to retrieve pure speech signals.Previous studies have also confirmed the benefits of incorporating phonetic information in a speech enhancement (SE) system to achieve better denoising performance.To obtain the phonetic information, we usually prepare a phoneme-based acoustic model, which is trained using speech waveforms and phoneme labels.Despite performing well in normal noisy conditions, when operating in very noisy conditions, however, the recognized phonemes may be erroneous and thus misguide the SE process.To overcome the limitation, this study proposes to incorporate the broad phonetic class (BPC) information into the SE process.We have investigated three criteria to build the BPC, including two knowledgebased criteria: place and manner of articulatory and one datadriven criterion.Moreover, the recognition accuracies of BPCs are much higher than that of phonemes, thus providing more accurate phonetic information to guide the SE process under very noisy conditions.Experimental results demonstrate that the proposed SE with the BPC information framework can achieve notable performance improvements over the baseline system and an SE system using monophonic information in terms of both speech quality intelligibility on the TIMIT dataset. Yen-Ju Lu, Chien-Feng Liao, Xugang Lu, Jeih-Weih Hung, Yu Tsao 0001 |
INTERSPEECH | 4 |
| 2020 | Time-Domain Multi-Modal Bone/Air Conducted Speech EnhancementabstractPrevious studies have proven that integrating video signals, as a complementary modality, can facilitate improved performance for speech enhancement (SE). However, video clips usually contain large amounts of data and pose a high cost in terms of computational resources and thus may complicate the SE system. As an alternative source, a bone-conducted speech signal has a moderate data size while manifesting speech-phoneme structures, and thus complements its air-conducted counterpart. In this study, we propose a novel multi-modal SE structure in the time domain that leverages bone- and air-conducted signals. In addition, we examine two ensemble-learning-based strategies, early fusion (EF) and late fusion (LF), to integrate the two types of speech signals, and adopt a deep learning-based fully convolutional network to conduct the enhancement. The experiment results on the Mandarin corpus indicate that this newly presented multi-modal (integrating bone- and air-conducted signals) SE structure significantly outperforms the single-source SE counterparts (with a bone- or air-conducted signal only) in various speech evaluation metrics. In addition, the adoption of an LF strategy other than an EF in this novel SE multi-modal structure achieves better results. Kuo-Hsuan Hung, Syu-Siang Wang, Yu Tsao 0001, Jeih-Weih Hung |
IEEE Signal Process. Lett. | 5 |
| 2019 | Speaker-Aware Deep Denoising Autoencoder with Embedded Speaker Identity for Speech Enhancement
Fu-Kai Chuang, Syu-Siang Wang, Jeih-Weih Hung, Yu Tsao 0001, Shih-Hau Fang |
INTERSPEECH | 3 |
| 2018 | Suppression by Selecting Wavelets for Feature Compression in Distributed Speech RecognitionabstractDistributed speech recognition (DSR) splits the processing of data between a mobile device and a network server. In the front-end, features are extracted and compressed to transmit over a wireless channel to a back-end server, where the incoming stream is received and reconstructed for recognition tasks. In this paper, we propose a feature compression algorithm termed suppression by selecting wavelets (SSW) to achieve the two main goals of DSR: Minimizing memory and device requirements while also maintaining or even improving the recognition performance. The SSW approach first applies the discrete wavelet transform (DWT) to filter the incoming speech feature sequence into two temporal subsequences at the client terminal. Feature compression is achieved by keeping the low (modulation) frequency subsequence while discarding the high frequency counterpart. The low-frequency subsequence is then transmitted across the remote network for specific feature statistics normalization. Wavelets are favorable for resolving the temporal properties of the feature sequence, and the down-sampling process in DWT achieves data compression by reducing the amount of data at the terminal prior to transmission across the network. Once the compressed features have arrived at the server, the feature sequence can be enhanced by statistics normalization, reconstructed with inverse DWT, and compensated with a simple post filter to alleviate any over-smoothing effects from the compression stage. Results on a standard robustness task (Aurora-4) and on a Mandarin Chinese news corpus showed SSW outperforms conventional noise-robustness techniques while also providing nearly a 50% compression rate during the transmission stage of DSR systems. Syu-Siang Wang, Payton Lin, Yu Tsao 0001, Jeih-Weih Hung, Borching Su |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Wavelet Speech Enhancement Based on Robust Principal Component Analysis
Chia-Lung Wu, Hsiang-Ping Hsu, Syu-Siang Wang, Jeih-Weih Hung, Ying-Hui Lai, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 4 |
| 2016 | Wavelet Speech Enhancement Based on Nonnegative Matrix FactorizationabstractFor the state-of-the-art speech enhancement (SE) techniques, a spectrogram is usually preferred than the respective time-domain raw data, since it reveals more compact presentation together with conspicuous temporal information over a long time span. However, two problems can cause distortions in the conventional nonnegative matrix factorization (NMF)-based SE algorithms. One is related to the overlap-and-add operation used in the short-time Fourier transform (STFT)-based signal reconstruction, and the other is concerned with directly using the phase of the noisy speech as that of the enhanced speech in signal reconstruction. These two problems can cause information loss or discontinuity when comparing the clean signal with the reconstructed signal. To solve these two problems, we propose a novel SE method that adopts discrete wavelet packet transform (DWPT) and NMF. In brief, the DWPT is first applied to split a time-domain speech signal into a series of subband signals. Then, we exploit NMF to highlight the speech component for each subband. These enhanced subband signals are joined together via the inverse DWPT to reconstruct a noise-reduced signal in time domain. We evaluate the proposed DWPT-NMF-based SE method on the Mandarin hearing in noise test (MHINT) task. Experimental results show that this new method effectively enhances speech quality and intelligibility and outperforms the conventional STFT-NMF-based SE system. Syu-Siang Wang, Alan Chern, Yu Tsao 0001, Jeih-Weih Hung, Xugang Lu, Ying-Hui Lai, Borching Su |
IEEE Signal Process. Lett. | 4 |
| 2016 | Robust Speech Recognition via Enhancing the Complex-Valued Acoustic Spectrum in Modulation DomainabstractThe purpose of this paper is to develop a novel speech feature extraction framework for independently compensating the real and imaginary acoustic spectra of speech signals in the modulation domain with the techniques of histogram equalization (HEQ) and non-negative matrix factorization (NMF). By doing so, we can enhance not only the magnitude but also the phase components of the acoustic spectra, thereby creating noise-robust speech features. More specifically, the proposed framework makes the following three major contributions: First, via either of the HEQ and NMF operations, the long-term cross-frame correlation among the acoustic spectra at the same frequency can be captured to compensate for the spectral distortion caused by noise. Second, the noise effect can be handled in a high acoustic frequency resolution. Finally, the distortion dwelt in the acoustic spectra can be more extensively mitigated due to the independent processes for the respective real and imaginary parts. The evaluation experiments were carried out on the Aurora-2 and Aurora-4 benchmark tasks, and the corresponding results suggest that our proposed methods can achieve performance competitive to or better than many widely used noise robustness methods, including the well-known advanced front-end (AFE) extraction scheme, in speech recognition. Jeih-Weih Hung, Hsin-Ju Hsieh, Berlin Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | Histogram equalization of contextual statistics of speech features for robust speech recognition
Hsin-Ju Hsieh, Berlin Chen, Jeih-Weih Hung |
Multim. Tools Appl. | 3 |
| 2014 | Speech enhancement using segmental nonnegative matrix factorizationabstractThe conventional NMF-based speech enhancement algorithm analyzes the magnitude spectrograms of both clean speech and noise in the training data via NMF and estimates a set of spectral basis vectors. These basis vectors are used to span a space to approximate the magnitude spectrogram of the noise-corrupted testing utterances. Finally, the components associated with the clean-speech spectral basis vectors are used to construct the updated magnitude spectrogram, producing an enhanced speech utterance. Considering that the rich spectral-temporal structure may be explored in local frequency and time-varying spectral patches, this study proposes a segmental NMF (SNMF) speech enhancement scheme to improve the conventional frame-wise NMF-based method. Two algorithms are derived to decompose the original nonnegative matrix associated with the magnitude spectrogram; the first algorithm is used in the spectral domain and the second algorithm is used in the temporal domain. When using the decomposition processes, noisy speech signals can be modeled more precisely, and spectrograms regarding the speech part can be constituted more favorably compared with using the conventional NMF-based method. Objective evaluations using perceptual evaluation of speech quality (PESQ) indicate that the proposed SNMF strategy increases the sound quality in noise conditions and outperforms the well-known MMSE log-spectral amplitude (LSA) estimation. Hao-Teng Fan, Jeih-Weih Hung, Xugang Lu, Syu-Siang Wang, Yu Tsao 0001 |
ICASSP | 2 |
| 2013 | Filtering on the temporal probability sequence in histogram equalization for robust speech recognitionabstractIn this paper, we propose a filter-based histogram equalization (FHEQ) approach for robust speech recognition. The FHEQ approach first represents the original acoustic feature sequence with statistic probability. Then, a temporal average (TA) filter is applied to smooth the statistic probability sequence. Finally, the filtered statistic probability sequence is transformed to form a new acoustic feature stream. Filtering on statistic probability of a feature sequence is a novel concept that can incorporate the advantages of the conventional histogram equalization (HEQ) and temporal filtering techniques for better noise robustness. Our experimental results on the Aurora-2 and Aurora-4 tasks show that FHEQ outperforms the conventional cepstral mean subtraction (CMS), cepstral mean and variance normalization (CMVN), and HEQ. Furthermore, we conducted a comparison test on TA-HEQ and HEQ-TA, which apply a TA filter to smooth acoustic features before and after the HEQ processing, respectively. The test results show that FHEQ outperforms both TA-HEQ and HEQ-TA, suggesting that filtering in probability is more effective than filtering in acoustic feature. Syu-Siang Wang, Yu Tsao 0001, Jeih-Weih Hung |
ICASSP | 3 |
| 2013 | Histogram equalization of real and imaginary modulation spectra for noise-robust speech recognition
Hsin-Ju Hsieh, Berlin Chen, Jeih-Weih Hung |
INTERSPEECH | 3 |
| 2012 | Exploring Joint Equalization of Spatial-Temporal Contextual Statistics of Speech Features for Robust Speech Recognition
Hsin-Ju Hsieh, Jeih-Weih Hung, Berlin Chen |
INTERSPEECH | 2 |
| 2012 | Improved modulation spectrum enhancement methods for robust speech recognition
Jeih-Weih Hung, Wen-Hsiang Tu, Chien-chou Lai |
Signal Process. | 1 |
| 2010 | Magnitude spectrum enhancement for robust speech recognitionabstractIn this paper, an effective compensation scheme for the spectra of speech signals is proposed in order to improve their noise robustness. In this compensation scheme, named magnitude spectrum enhancement (MSE), a voice activity detection (VAD) process is first processed for the frame sequence of the utterance, and then the magnitude spectra of non-speech frames are set to be small while those of speech frames are amplified. In experiments conducted on the Aurora-2 noisy digits database, MSE achieves a relative error reduction rate of nearly 50% from the baseline processing, which outperforms the well-known spectral-domain speech enhancement techniques, spectral subtraction (SS) and Wiener filtering (WF). In addition, the proposed MSE can be integrated with cepstral-domain robustness methods, like mean and variance normalization (MVN) and histogram normalization (HEQ), to achieve further improved recognition accuracy under noise-corrupted environments. Wen-Hsiang Tu, Jeih-Weih Hung |
ICASSP | 2 |
| 2009 | Sub-band modulation spectrum compensation for robust speech recognitionabstractThis paper proposes a novel scheme in performing feature statistics normalization techniques for robust speech recognition. In the proposed approach, the processed temporal-domain feature sequence is first converted into the modulation spectral domain. The magnitude part of the modulation spectrum is decomposed into non-uniform sub-band segments, and then each sub-band segment is individually processed by the well-known normalization methods, like mean normalization (MN), mean and variance normalization (MVN) and histogram equalization (HEQ). Finally, we reconstruct the feature stream with all the modified sub-band magnitude spectral segments and the original phase spectrum using the inverse DFT. With this process, the components that correspond to more important modulation spectral bands in the feature sequence can be processed separately. For the Aurora-2 clean-condition training task, the new proposed sub-band spectral MVN and HEQ provide relative error rate reductions of 18.66% and 23.58% over the conventional temporal MVN and HEQ, respectively. Wen-Hsiang Tu, Sheng-Yuan Huang, Jeih-Weih Hung |
ASRU | 3 |
| 2009 | Sub-band feature statistics compensation techniques based on discrete wavelet transform for robust speech recognitionabstractThis paper proposes a novel scheme in performing feature statistics normalization techniques for robust speech recognition. In the proposed approach, the processed temporal-domain feature sequence is first decomposed into non-uniform sub-bands using discrete wavelet transform (DWT), and then each sub-band stream is individually processed by the well-known normalization methods, like mean and variance normalization (MVN) and histogram equalization (HEQ). Finally, we reconstruct the feature stream with all the modified sub-band streams using inverse DWT. With this process, the components that correspond to more important modulation spectral bands in the feature sequence can be processed separately. For the Aurora-2 clean-condition training task, the new proposed sub-band MVN and HEQ provide relative error rate reductions of 20.18% and 19.65% over the conventional MVN and HEQ. Hao-Teng Fan, Jeih-Weih Hung |
ICME | 2 |
| 2009 | Integrating codebook and utterance information in cepstral statistics normalization techniques for robust speech recognition
Guan-min He, Jeih-Weih Hung |
INTERSPEECH | 2 |
| 2009 | Subband Feature Statistics Normalization Techniques Based on a Discrete Wavelet Transform for Robust Speech RecognitionabstractThis letter proposes a novel scheme that applies feature statistics normalization techniques for robust speech recognition. In the proposed approach, the processed temporal-domain feature sequence is first decomposed into nonuniform subbands using the discrete wavelet transform (DWT), and then each subband stream is individually processed by well-known normalization methods, such as mean and variance normalization (MVN) and histogram equalization (HEQ). Finally, we reconstruct the feature stream with all of the modified subband streams using the inverse DWT. With this process, the components that correspond to more important modulation spectral bands in the feature sequence can be processed separately. Jeih-Weih Hung, Hao-Teng Fan |
IEEE Signal Process. Lett. | 1 |
| 2009 | Incorporating Codebook and Utterance Information in Cepstral Statistics Normalization Techniques for Robust Speech Recognition in Additive Noise EnvironmentsabstractCepstral statistics normalization techniques have been shown to be very successful at improving the noise robustness of speech features. This letter proposes a hybrid-based scheme to achieve a more accurate estimate of the statistical information of features in these techniques. By properly integrating codebook and utterance knowledge, the resulting hybrid-based approach significantly outperforms conventional utterance-based, segment-based and codebook-based approaches in additive noise environments. Furthermore, the high-performance CS-HEQ can be implemented with a short delay and can thus be applied in real-time online systems. Jeih-Weih Hung, Wen-Hsiang Tu |
IEEE Signal Process. Lett. | 1 |
| 2008 | Improved modulation spectrum normalization techniques for robust speech recognitionabstractThe modulation spectra of speech features are often distorted due to environmental interferences. In order to reduce this distortion, in this paper we propose several approaches to normalize the power spectral density (PSD) of the feature stream to a reference function. These approaches include least-squares temporal filtering (LSTF), least-squares spectrum fitting (LSSF) and magnitude spectrum interpolation (MSI). It is shown that all the proposed approaches can effectively improve the speech recognition accuracy in various noise corrupted environments. In experiments conducted on the Aurora-2 noisy digits database with a complex back-end, these new approaches provide an average relative error reduction rate of over 40% when compared with the baseline MFCC processing. Chi-an Pan, Chih-Cheng Wang, Jeih-Weih Hung |
ICASSP | 3 |
| 2008 | Silence feature normalization for robust speech recognition in additive noise environments
Chih-Cheng Wang, Chi-an Pan, Jeih-Weih Hung |
INTERSPEECH | 3 |
| 2008 | Constructing Modulation Frequency Domain-Based Features for Robust Speech RecognitionabstractData-driven temporal filtering approaches based on a specific optimization technique have been shown to be capable of enhancing the discrimination and robustness of speech features in speech recognition. The filters in these approaches are often obtained with the statistics of the features in the temporal domain. In this paper, we derive new data-driven temporal filters that employ the statistics of the modulation spectra of the speech features. Three new temporal filtering approaches are proposed and based on constrained versions of linear discriminant analysis (LDA), principal component analysis (PCA), and minimum class distance (MCD), respectively. It is shown that these proposed temporal filters can effectively improve the speech recognition accuracy in various noise-corrupted environments. In experiments conducted on Test Set A of the Aurora-2 noisy digits database, these new temporal filters, together with cepstral mean and variance normalization (CMVN), provide average relative error reduction rates of over 40% and 27% when compared with baseline Mel frequency cepstral coefficient (MFCC) processing and CMVN alone, respectively. Jeih-Weih Hung, Wei-Yi Tsai |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Optimization of Temporal Filters in the Modulation Frequency Domain via Constrained Linear Discriminant Analysis (C-LDA) for Constructing Robust Features in Speech RecognitionabstractData-driven temporal filtering approaches based on a specific optimization criterion have been shown to be capable of enhancing the discrimination and robustness of speech features in speech recognition. The filters in these approaches are often obtained with the statistics of the features in the temporal domain. In this paper, we derive new data-driven temporal filters that employ the statistics of the modulation spectra of the speech features. The new temporal filtering approach is based on the constrained version of linear discriminant analysis (C-LDA). It is shown that the proposed C-LDA temporal filters can effectively improve the speech recognition accuracy in various noise corrupted environments. In experiments conducted on Test Set A of the Aurora-2 noisy digits database, these new temporal filters, together with cepstral mean and variance normalization (CMVN), provides average relative error reduction rates of over 47% and 30%, when compared with the baseline MFCC processing and CMVN alone, respectively. Jeih-Weih Hung |
ICASSP (4) | 1 |
| 2007 | Speech feature compensation based on pseudo stereo codebooks for robust speech recognition in additive noise environments
Tsung-hsueh Hsieh, Jeih-Weih Hung |
INTERSPEECH | 2 |
| 2007 | Optimization of temporal filters in the modulation frequency domain for constructing robust features in speech recognition
Jeih-Weih Hung |
INTERSPEECH | 1 |
| 2006 | Cepstral Statistics Compensation Using Online Pseudo Stereo Codebooks for Robust Speech Recognition in Additive Noise EnvironmentsabstractIn this paper, we propose the cepstral statistics compensation (CSC) algorithm, which alleviates the effect of additive noise on the cepstral features for speech recognition. It is a simple but quite efficient noise reduction technique that makes use of online constructed pseudo stereo codebooks. The statistics, such as mean and variance, for the cepstral features in both clean and noisy environments are evaluated using pseudo stereo codebooks. Then a transform is obtained for the noise-corrupted cepstra so that the statistics of the transformed ones are close to those of clean cepstra. Experimental results show that CSC provided a 13% reduction in word error rate when compared to the results obtained using cepstral mean and variance normalization (CMVN), and a 34% reduction in error rate compared to baseline processing in the noise range of 0-20dB in experiments conducted on Aurora-2 Test Set A noisy digits database. In addition, we also provide some other noise robustness approaches based on pseudo stereo codebooks and show their effectiveness in noisy speech recognition. Jeih-Weih Hung |
ICASSP (1) | 1 |
| 2006 | Silence energy normalization for robust speech recognition in additive noise environment
Chung-fu Tai, Jeih-Weih Hung |
INTERSPEECH | 2 |
| 2006 | Optimization of temporal filters for constructing robust features in speech recognitionabstractLinear discriminant analysis (LDA) has long been used to derive data-driven temporal filters in order to improve the robustness of speech features used in speech recognition. In this paper, we proposed the use of new optimization criteria of principal component analysis (PCA) and the minimum classification error (MCE) for constructing the temporal filters. Detailed comparative performance analysis for the features obtained using the three optimization criteria, LDA, PCA, and MCE, with various types of noise and a wide range of SNR values is presented. It was found that the new criteria lead to superior performance over the original MFCC features, just as LDA-derived filters can. In addition, the newly proposed MCE-derived filters can often do better than the LDA-derived filters. Also, it is shown that further performance improvements are achievable if any of these LDA/PCA/MCE-derived filters are integrated with the conventional approach of cepstral mean and variance normalization (CMVN). The performance improvements obtained in recognition experiments are further supported by analyses conducted using two different distance measures. Jeih-Weih Hung, Lin-Shan Lee |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | Data-driven temporal filters based on multi-eigenvectors for robust features in speech recognitionabstractIt was previously proposed to use the principal component analysis (PCA) to derive the data-driven temporal filters for obtaining robust features in speech recognition, in which the first principal components are taken as the filter coefficients. In this paper, a multi-eigenvector approach is proposed instead, in which the first M eigenvectors obtained in PCA are weighted by their corresponding eigenvalues and summed to be used as the filter coefficients. Experimental results showed that the multi-eigenvector filters offer significant recognition performance as compared to the previously proposed PCA-derived filters under all different conditions tested with the AURORA2 database, especially when the training and testing environments are highly mismatched. Ni-Chun Wang, Jeih-Weih Hung, Lin-Shan Lee |
ICASSP (1) | 2 |
| 2002 | Data-driven temporal filters for robust features in speech recognition obtained via Minimum Classification Error (MCE)abstractIn deriving the data-driven temporal filters for speech features, the Linear Discriminant Analysis (LDA) and the Principal Component Analysis (PCA) have been shown to be successful in improving the feature robustness. In this paper, it's proposed that the criterion of Minimum Classification Error (MCE) can also be used to obtain the data-driven temporal filters. Two versions of MCE-derived temporal filters, Feature-based and Model-based, are proposed and it is shown that both of them can significantly improve the recognition performance of the original MFCC features as the LDA/PCA-derived filters do. Detailed comparative analysis among the different temporal filtering approaches is presented. It is also shown that the proposed MCE filters can be integrated with the conventional temporal filters, RASTA or CMS, to obtain improved recognition performance regardless of whether the training and testing environments are matched or mismatched, compressed or noise corrupted. Jeih-Weih Hung, Lin-Shan Lee |
ICASSP | 1 |
| 2002 | Data-driven temporal filters obtained via different optimization criteria evaluated on Aurora2 database
Jeih-Weih Hung, Lin-Shan Lee |
INTERSPEECH | 1 |
| 2001 | Comparative analysis for data-driven temporal filters obtained via principal component analysis (PCA) and linear discriminant analysis (LDA) in speech recognitionabstractThe Linear Discriminant Analysis (LDA) has been widely used to derive the data-driven temporal filtering of speech feature vectors. In this paper, we proposed that the Principal Component Analysis (PCA) can also be used in the optimization process just as LDA to obtain the temporal filters, and detailed comparative analysis between these two approaches are presented and discussed. It's found that the PCA-derived temporal filters significantly improve the recognition performance of the original MFCC features as LDA-derived filters do. Also, while PCA/LDA filters are combined with the conventional temporal filters, RASTA or CMS, the recognition performance will be further improved regardless the training and testing environments are matched or mismatched, compressed or noise corrupted. 1. Jeih-Weih Hung, Hsin-Min Wang, Lin-Shan Lee |
INTERSPEECH | 1 |
| 2001 | New approaches for domain transformation and parameter combination for improved accuracy in parallel model combination (PMC) techniquesabstractParallel model combination (PMC) techniques have been very successful and popularly used in many applications to improve the performance of speech recognition systems under noisy environments. However, it is believed that some assumptions and approximations made in this approach, primarily in the domain transformation and parameter combination processes, are not necessarily accurate enough in certain practical situations, which may degrade the achievable performance of PMC. In this paper, the possible sources that cause the performance degradation in these processes are carefully analyzed and discussed. Three new approaches, including the truncated Gaussian approach and the split mixture approach for the domain transformation process and the estimated cross-term approach for parameter combination process, are proposed in this paper in order to handle these problems, minimize such degradation, and improve the accuracy of the PMC techniques. These proposed approaches were analyzed and discussed with two recognition tasks, one relatively simple, and the other more complicated and realistic. Both sets of experiments showed that these proposed approaches are able to provide significant improvements over the original PMC method, especially when the SNR condition is worse. Jeih-Weih Hung, Jia-Lin Shen, Lin-Shan Lee |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | Automatic metric-based speech segmentation for broadcast news via principal component analysisabstractIn this paper, we proposed an algorithm used to improve the performance of the metric-based segmentation techniques, by which the segmentation points are found at maxima of a distance measured between two contiguous windows shifted along the stream of speech features. In our proposed method, the PCA processes are first performed on the speech features to obtain more robust features, and then the above metric-based segmentation was applied on the PCA-derived features to decide the segmentation points. Experiment results show that our proposed method can efficiently improve the detection rates of the segmentation points up to 7% while the false alarm rates remain unchanged. Jeih-Weih Hung, Hsin-Min Wang, Lin-Shan Lee |
INTERSPEECH | 1 |
| 1999 | Improved parallel model combination techniques with split Gaussian mixtures for speech recognition under noisy conditionsabstractThe parallel model combination (PMC) technique has been very successful and frequently used to improve the performance of a speech recognition system under noisy environments. In this approach it is assumed that the log spectrum of speech signals is Gaussian-distributed, which is not always valid especially when the number of mixtures in the HMMs is few. In this paper, a simple approach is proposed to improve the PMC method by splitting the mixtures before the domain transformation process in the PMC is performed, and merging the mixtures back to the original number after the PMC processes are completed. Preliminary experimental results show that the increased number of mixtures during the PMC processes can in fact provide significant improvements over the original PMC method in terms of the recognition accuracies, especially when the SNR is low. Jeih-Weih Hung, Jia-Lin Shen, Lin-Shan Lee |
ICASSP | 1 |
| 1998 | Improved robustness for speech recognition under noisy conditions using correlated parallel model combinationabstractThe parallel model combination (PMC) technique has been shown to achieve very good performance for speech recognition under noisy conditions. In this approach, the speech signal and the noise are assumed uncorrelated during modeling. A new correlated PMC is proposed by properly estimating and modeling the nonzero correlation between the speech signal and the noise. Preliminary experimental results show that this correlated PMC can provide significant improvements over the original PMC in terms of both the model differences and the recognition accuracies. Error rate reduction on the order of 14% can be achieved. Jeih-Weih Hung, Jia-Lin Shen, Lin-Shan Lee |
ICASSP | 1 |
| 1998 | Improved parallel model combination based on better domain transformation for speech recognition under noisy environments
Jeih-Weih Hung, Jia-Lin Shen, Lin-Shan Lee |
ICSLP | 1 |
| 1998 | Improved robust speech recognition considering signal correlation approximated by taylor series
Jia-Lin Shen, Jeih-Weih Hung, Lin-Shan Lee |
ICSLP | 2 |
| 1998 | Robust entropy-based endpoint detection for speech recognition in noisy environmentsabstractThis paper presents an entropy-based algorithm for accurate and robust endpoint detection for speech recognition under noisy environments. Instead of using the conventional energy-based features, the spectral entropy is developed to identify the speech segments accurately. Experimental results show that this algorithm outperforms the energy-based algorithms in both detection accuracy and recognition performance under noisy environments, with an average error rate reduction of more than 16%. Jia-Lin Shen, Jeih-Weih Hung, Lin-Shan Lee |
ICSLP | 2 |