Fei Chen 0011

dblp:81/4345-11 · DBLP profile ↗
← Back
45ranked-venue papers
14as first author
16since 2021 · last 2026
0000-0002-6988-492XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 13 first-author · 11 since 2021Artificial intelligence and machine learning · 28 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021
YearPublicationVenuePosition
2026 Unsupervised joint domain adaptive framework for patient-independent seizure classification
Sunday Timothy Aboyeji, Xin Wang 0088, Oluwarotimi Williams Samuel, Juanjuan Li, Fei Chen 0011, Shengyun Liang, Michael C. F. Tong, Lina Men, Xianhai Zeng, Shixiong Chen
Expert Syst. Appl.7
2025 Attention-Based Beamformer For Multi-Channel Speech Enhancement
abstract
Minimum Variance Distortionless Response (MVDR) is a classical adaptive beamformer that theoretically ensures the distortionless transmission of signals in the target direction, which makes it popular in real applications. Its noise reduction performance actually depends on the accuracy of the noise and speech spatial covariance matrices (SCMs) estimation. Time-frequency masks are often used to compute these SCMs. However, most mask-based beamforming methods typically assume that the sources are stationary, ignoring the case of moving sources, which leads to performance degradation. In this paper, we propose an attention-based mechanism to calculate the speech and noise SCMs and then apply MVDR to obtain the enhanced speech. To fully incorporate spatial information, the inplace convolution operator and frequency-independent LSTM are applied to facilitate SCMs estimation. The model is optimized in an end-to-end manner. Experiments demonstrate that the proposed method outperforms baselines with reduced computation and fewer parameters under various conditions.
Jinglin Bai, Hao Li 0046, Xueliang Zhang 0001, Fei Chen 0011
ICASSP4
2025 Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
abstract
Given the critical role of non-intrusive speech intelligibility assessment in hearing aids (HA), this paper enhances its performance by introducing Feature Importance across Domains (FiDo). We estimate feature importance on spectral and time-domain acoustic features as well as latent representations of Whisper. Importance weights are calculated per frame, and based on these weights, features are projected into new spaces, allowing the model to focus on important areas early. Next, feature concatenation is performed to combine the features before the assessment module processes them. Experimental results show that when FiDo is incorporated into the improved multi-branched speech intelligibility model MBI-Net+, RMSE can be reduced by 7.62% (from 26.10 to 24.11). MBI-Net+ with FiDo also achieves a relative RMSE reduction of 3.98% compared to the best system in the 2023 Clarity Prediction Challenge. These results validate FiDo's effectiveness in enhancing neural speech assessment in HA.
Ryandhimas E. Zezario, Sabato Marco Siniscalchi, Fei Chen 0011, Hsin-Min Wang, Yu Tsao 0001
INTERSPEECH3
2024 3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications
abstract
Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters which are computational intensive and impractical for real-time applications, espetially on resource-constrained platforms. In this paper, we address the TSE task using microphone array and introduce a novel three-stage solution that systematically decouples the process: First, a neural network is trained to estimate the direction of the target speaker. Second, with the direction determined, the Generalized Sidelobe Canceller (GSC) is used to extract the target speech. Third, an Inplace Convolutional Recurrent Neural Network (ICRN) acts as a denoising post-processor, refining the GSC output to yield the final separated speech. Our approach delivers superior performance while drastically reducing computational load, setting a new standard for efficient real-time target speaker extraction.
Shulin He, Hao Li 0046, Yang Yang 0121, Fei Chen 0011, Xueliang Zhang 0001
ICASSP5
2024 Cross-Attention-Guided WaveNet for EEG-to-MEL Spectrogram Reconstruction
Hao Li 0046, Xueliang Zhang 0001, Fei Chen 0011, Guanglai Gao
INTERSPEECH4
2024 Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
Ryandhimas E. Zezario, Fei Chen 0011, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001
INTERSPEECH2
2024 Age-Related Changes in Blood Volume Pulse Wave at Fingers and Ears
abstract
OBJECTIVE: The decline in vascular elasticity with aging can be manifested in the shape of pulse wave. The study investigated the pulse wave features that are sensitive to age and the pattern of these features change with increasing age were examined. METHODS: Five features were proposed and extracted from the photoplethysmography (PPG)-based pulse wave or its first derivative wave. The correlation between these PPG features and ages was studied in 100 healthy subjects with a wide range of ages (20-71 years). Piecewise regression coefficients were calculated to examine the rates of change of the PPG features with age at different age stages. RESULTS: The proposed PPG features obtained from the finger showed a strong and significant correlation with age (with r = 0.76 - 0.77, p < 0.01), indicating higher sensitivity to age changes compared to the PPG features reported in previous studies (with r = 0.66 - 0.75). The correlation remained significant even after correcting for other clinical variables. The rate of change of the PPG feature values was found to be significantly faster in subjects aged ≥40 years compared to those aged < 40 years in the healthy population. This rate of change was similar to the age-related progression of arterial stiffness evaluated by pulse wave velocity (PWV), which is considered a gold standard for evaluating vascular stiffness. CONCLUSIONS: The proposed PPG features showed a high correlation with chronological age in healthy subjects and exhibited a similar age-related change trend as PWV. SIGNIFICANCE: With the convenience of PPG measures, the proposed age-related features have the potential to be used as biomarkers for vascular aging and estimating the risk of cardiovascular disease.
Wan-Hua Lin, Dingchang Zheng, Guanglin Li 0001, Fei Chen 0011
IEEE J. Biomed. Health Informatics4
2023 Deep Learning-Based Non-Intrusive Multi-Objective Speech Assessment Model With Cross-Domain Features
abstract
This study proposes a cross-domain multi-objective speech assessment model, called MOSA-Net, which can simultaneously estimate the speech quality, intelligibility, and distortion assessment scores of an input speech signal. MOSA-Net comprises a convolutional neural network and bidirectional long short-term memory architecture for representation extraction, and a multiplicative attention layer and a fully connected layer for each assessment metric prediction. Additionally, cross-domain features (spectral and time-domain features) and latent representations from self-supervised learned (SSL) models are used as inputs to combine rich acoustic information to obtain more accurate assessments. Experimental results show that in both seen and unseen noise environments, MOSA-Net can improve the linear correlation coefficient (LCC) scores in perceptual evaluation of speech quality (PESQ) prediction, compared to Quality-Net, an existing single-task model for PESQ prediction, and improve LCC scores in short-time objective intelligibility (STOI) prediction, compared to STOI-Net, an existing single-task model for STOI prediction. Moreover, MOSA-Net can be used as a pre-trained model to be effectively adapted to an assessment model for predicting subjective quality and intelligibility scores with a limited amount of training data. Experimental results show that MOSA-Net can improve LCC scores in mean opinion score (MOS) predictions, compared to MOS-SSL, a strong single-task model for MOS prediction. We further adopt the latent representations of MOSA-Net to guide the speech enhancement (SE) process and derive a quality-intelligibility (QI)-aware SE (QIA-SE) approach. Experimental results show that QIA-SE outperforms the baseline SE system with improved PESQ scores in both seen and unseen noise environments over a baseline SE model.
Ryandhimas E. Zezario, Szu-Wei Fu, Fei Chen 0011, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Fast and Calibrationless Low-Rank Parallel Imaging Reconstruction Through Unrolled Deep Learning Estimation of Multi-Channel Spatial Support Maps
abstract
Low-rank technique has emerged as a powerful calibrationless alternative for parallel magnetic resonance (MR) imaging. Calibrationless low-rank reconstruction, such as low-rank modeling of local k-space neighborhoods (LORAKS), implicitly exploits both coil sensitivity modulations and the finite spatial support constraint of MR images through an iterative low-rank matrix recovery process. Although powerful, this slow iteration process is computationally demanding and reconstruction requires empirical rank optimization, hampering its robust applications for high-resolution volume imaging. This paper proposes a fast and calibrationless low-rank reconstruction of undersampled multi-slice MR brain data, based on the finite spatial support constraint reformulation with a direct deep learning estimation of spatial support maps. The iteration process of low-rank reconstruction is unrolled into a complex-valued network by training on fully-sampled multi-slice axial brain datasets acquired from the same MR coil system. To utilize coil-subject geometric parameters available for datasets, the model minimizes a hybrid loss on two sets of spatial support maps, corresponding to brain data at the original slice locations as actually acquired and nearby locations within the standard reference coordinate. This deep learning framework was integrated with LORAKS reconstruction and was evaluated with publically available gradient-echo T1-weighted brain datasets. It directly produced high-quality multi-channel spatial support maps from undersampled data, enabling rapid reconstruction without iteration. Moreover, it led to effective reductions of artifacts and noise amplification at high acceleration. In summary, our proposed deep learning framework offers a new strategy to advance the existing calibrationless low-rank reconstruction, rendering it computationally efficient, simple, and robust in practice.
Zheyuan Yi, Jiahao Hu 0003, Yujiao Zhao 0003, Linfang Xiao, Alex T. L. Leong, Fei Chen 0011, Ed X. Wu
IEEE Trans. Medical Imaging7
2022 ConferencingSpeech 2022 Challenge: Non-intrusive Objective Speech Quality Assessment (NISQA) Challenge for Online Conferencing Applications
abstract
With the advances in speech communication systems such as online conferencing applications, we can seamlessly work with people regardless of where they are. However, during online meetings, speech quality can be significantly affected by background noise, reverberation, packet loss, network jitter, etc. Because of its nature, speech quality is traditionally assessed in subjective tests in laboratories and lately also in crowdsourcing following the international standards from ITU-T Rec. P.800 series. However, those approaches are costly and cannot be applied to customer data. Therefore, an effective objective assessment approach is needed to evaluate or monitor the speech quality of the ongoing conversation. The ConferencingSpeech 2022 challenge targets the non-intrusive deep neural network models for the speech quality assessment task. We open-sourced a training corpus with more than 86K speech clips in different languages, with a wide range of synthesized and live degradations and their corresponding subjective quality scores through crowdsourcing. 18 teams submitted their models for evaluation in this challenge. The blind test sets included about 4300 clips from wide ranges of degradations. This paper describes the challenge, the datasets, and the evaluation methods and reports the final results.
Gaoxiong Yi, Babak Naderi, Sebastian Möller 0001, Wafaa Wardah, Gabriel Mittag, Ross Cutler, Zhuohuang Zhang, Donald S. Williamson, Fei Chen 0011, Shidong Shang
INTERSPEECH11
2022 MBI-Net: A Non-Intrusive Multi-Branched Speech Intelligibility Prediction Model for Hearing Aids
Ryandhimas E. Zezario, Fei Chen 0011, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001
INTERSPEECH2
2022 MTI-Net: A Multi-Target Speech Intelligibility Prediction Model
abstract
Recently, deep learning (DL)-based non-intrusive speech assessment models have attracted great attention.Many studies report that these DL-based models yield satisfactory assessment performance and good flexibility, but their performance in unseen environments remains a challenge.Furthermore, compared to quality scores, fewer studies elaborate deep learning models to estimate intelligibility scores.This study proposes a multi-task speech intelligibility prediction model, called MTI-Net, for simultaneously predicting human and machine intelligibility measures.Specifically, given a speech utterance, MTI-Net is designed to predict human subjective listening test results and word error rate (WER) scores.We also investigate several methods that can improve the prediction performance of MTI-Net.First, we compare different features (including low-level features and embeddings from self-supervised learning (SSL) models) and prediction targets of MTI-Net.Second, we explore the effect of transfer learning and multi-tasking learning on training MTI-Net.Finally, we examine the potential advantages of fine-tuning SSL embeddings.Experimental results demonstrate the effectiveness of using cross-domain features, multi-task learning, and fine-tuning SSL embeddings.Furthermore, it is confirmed that the intelligibility and WER scores predicted by MTI-Net are highly correlated with the ground-truth scores.
Ryandhimas E. Zezario, Szu-Wei Fu, Fei Chen 0011, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001
INTERSPEECH3
2022 Single-Channel Selection for EEG-Based Emotion Recognition Using Brain Rhythm Sequencing
abstract
Recently, electroencephalography (EEG) signals have shown great potential for emotion recognition. Nevertheless, multichannel EEG recordings lead to redundant data, computational burden, and hardware complexity. Hence, efficient channel selection, especially single-channel selection, is vital. For this purpose, a technique termed brain rhythm sequencing (BRS) that interprets EEG based on a dominant brain rhythm having the maximum instantaneous power at each 0.2 s timestamp has been proposed. Then, dynamic time warping (DTW) is used for rhythm sequence classification through the similarity measure. After evaluating the rhythm sequences for the emotion recognition task, the representative channel that produces impressive accuracy can be found, which realizes single-channel selection accordingly. In addition, the appropriate time segment for emotion recognition is estimated during the assessments. The results from the music emotion recognition (MER) experiment and three emotional datasets (SEED, DEAP, and MAHNOB) indicate that the classification accuracies achieve 70-82% by single-channel data with a 10 s time length. Such performances are remarkable when considering minimum data sources as the primary concerns. Furthermore, the individual characteristics in emotion recognition are investigated based on the channels and times found. Therefore, this study provides a novel method to solve single-channel selection for emotion recognition.
Jia Wen Li 0001, Shovan Barma, Peng Un Mak, Fei Chen 0011, Ming Tao Li, Mang I Vai, Sio-Hang Pun
IEEE J. Biomed. Health Informatics4
2021 Perceptual Contributions of Vowels and Consonant-Vowel Transitions in Understanding Time-Compressed Mandarin Sentences
Changjie Pan, Fei Chen 0011
Interspeech3
2021 Effect of Carrier Bandwidth on Understanding Mandarin Sentences in Simulated Electric-Acoustic Hearing
Jing Chen 0019, Fei Chen 0011
Interspeech3
2021 The effect of speech and noise levels on the quality perceived by cochlear implant and normal hearing listeners
abstract
Electrical hearing by cochlear implants (CIs) may be fundamentally different from the acoustic hearing by normal-hearing (NH) listeners, presumably showing unequal speech quality perception in various noise environments. Noise reduction (NR) algorithms used in CIs reduce the noise in favor of signal-to-noise ratio (SNR), regardless of plausible accompanying distortions that may degrade the speech quality perception. To gain a better understanding of CI speech quality perception, the present work aimed at investigating speech quality perception in diverse noise conditions, including factors of speech/noise levels, type of noise, and distortions caused by NR models. Fifteen NH and seven CI subjects participated in this study. Speech sentences were set to two different levels (65 and 75 dB SPL (Sound Pressure Level)). Two types of noise (Cafeteria and Babble) at three levels (55, 65, and 75 dB SPL) were used. Sentences were processed using two NR algorithms to investigate the perceptual sensitivity of CI and NH listeners to the distortion. All sentences processed with the combinations of these sets were presented to CI and NH listeners, and they were asked to rate the sound quality of speech as they perceived. The effect of each factor on the perceived speech quality was investigated based on the group-averaged quality rated by CI and NH listeners. Consistent with previous studies, CI listeners were not as sensitive as NH listeners to the distortion made by NR algorithms. Statistical analysis showed that the speech level has a significant effect on quality perception. At the same SNR, the quality of 65 dB speech was rated higher than that of 75 dB for CI users, but vice versa for NH listeners. Therefore, the present study showed that the perceived speech quality patterns were different between CI and NH listeners in terms of their sensitivity to distortion and speech level in a complex listening environment.
Sara Akbarzadeh, Sungmin Lee 0003, Fei Chen 0011, Chin-Tuan Tan
Speech Commun.3
2020 Contribution of RMS-Level-Based Speech Segments to Target Speech Decoding Under Noisy Conditions
Lei Wang 0074, Ed X. Wu, Fei Chen 0011
INTERSPEECH3
2020 Guest Editorial Flexible Sensing and Medical Imaging for Cerebro-Cardiovascular Health
abstract
The articles in this special section focus on flexible sensing and medical imaging for cerebro-cardiovascular health care services. Healthcare and disease management are receiving increasing attention. Cerebro-cardiovascular diseases (CCVDs) are the leading cause of death globally. Cerebrocardiovascular diseases include a variety of medical conditions that affect the blood vessels of the brain, the cerebral circulation, and the heart. The common presentations of CCVDs include an ischemic stroke or mini-stroke and sometimes a hemorrhagic stroke, heart failure, hypertensive heart disease, etc. The important contributing risk factors include high blood pressure, smoking, diabetes, lack of exercise, obesity, high blood cholesterol, and excessive alcohol consumption, among others. A rapidly growing field, biomedical and health engineering research for CCVDs is unique in that it involves a variety of specialties such as neurology, surgery, cardiology, psychology and rehabilitation, and must meet the growing need for sophisticated, up-to-date biomedical and health informatics on clinical data, diagnostic testing, and therapeutic issues.
Paolo Bonato, Yifan Chen 0001, Fei Chen 0011, Yuan-Ting Zhang
IEEE J. Biomed. Health Informatics3
2019 Segmental contributions to cochlear implant speech perception
Fei Chen 0011
Speech Commun.1
2018 Speech Dereverberation Based on Integrated Deep and Ensemble Learning Algorithm
abstract
Reverberation, which is generally caused by sound reflections from walls, ceilings, and floors, can result in severe performance degradation of acoustic applications. Due to a complicated combination of attenuation and time-delay effects, the reverberation property is difficult to characterize, and it remains a challenging task to effectively retrieve the anechoic speech signals from reverberation ones. In the present study, we proposed a novel integrated deep and ensemble learning algorithm (IDEA) for speech dereverberation. The IDEA consists of offline and online phases. In the offline phase, we train multiple dereverberation models, each aiming to precisely dereverb speech signals in a particular acoustic environment; then a unified fusion function is estimated that aims to integrate the information of multiple dereverberation models. In the online phase, an input utterance is first processed by each of the dereverberation models. The outputs of all models are integrated accordingly to generate the final anechoic signal. We evaluated the IDEA on designed acoustic environments, including both matched and mismatched conditions of the training and testing data. Experimental results confirm that the proposed IDEA outperforms single deep-neural-network-based dereverberation model with the same model architecture and training data.
Wei-Jen Lee, Syu-Siang Wang, Fei Chen 0011, Xugang Lu, Shao-Yi Chien, Yu Tsao 0001
ICASSP3
2018 A Time-Weighted Method for Predicting the Intelligibility of Speech in the Presence of Interfering Sounds
abstract
The speech intelligibility index (SII) has been widely used as an objective method of predicting speech intelligibility, but its traditional form is most effective predicting speech intelligibility scores under stationary noise but not more challenging conditions (e.g., competing noise interference). To address this limitation, the present work extended the SII model to predict the intelligibility of speech in both steady speech-spectral noise (SSN) and dual-talker speech (DTS), by using a time-weighted function that accounted for the relative perceptual importance of vowels and consonants in speech intelligibility. The performance of the new time-weighted SII (TW-SII) was compared to the other two well-known methods, i.e., the time-averaged SII (TA-SII) and coherence SII (CSII). Experimental results showed the intelligibility prediction accuracy of the three methods was similar for speech in SSN, but the prediction by TW-SII was more accurate than those by TA-SII and CSII for speech in DTS. The possible applications and limitations of the present intelligibility model were analyzed and discussed.
Mingjie Song, Fei Chen 0011, Xihong Wu, Jing Chen 0019
ICASSP2
2017 Factors Affecting the Intelligibility of Low-Pass Filtered Speech
Lei Wang 0074, Fei Chen 0011
INTERSPEECH2
2017 Phonetic Restoration of Temporally Reversed Speech
Shiyu Wang 0003, Fei Chen 0011
INTERSPEECH2
2017 Multi-style learning with denoising autoencoders for acoustic modeling in the internet of things (IoT)
Payton Lin, Dau-Cheng Lyu, Fei Chen 0011, Syu-Siang Wang, Yu Tsao 0001
Comput. Speech Lang.3
2016 Modeling Noise Influence to Speech Intelligibility Non-Intrusively by Reduced Speech Dynamic Range
Fei Chen 0011
INTERSPEECH1
2016 Comparing the Contributions of Amplitude and Phase to Speech Intelligibility in a Vocoder-Based Speech Synthesis Model
Fei Chen 0011, Benson C. L. Chiao
INTERSPEECH1
2016 Factors Affecting the Intelligibility of Sine-Wave Speech
Fei Chen 0011, Daniel Fogerty
INTERSPEECH1
2016 Vowel Fundamental and Formant Frequency Contributions to English and Mandarin Sentence Intelligibility
Daniel Fogerty, Fei Chen 0011
INTERSPEECH2
2016 Assessing Level-Dependent Segmental Contribution to the Intelligibility of Speech Processed by Single-Channel Noise-Suppression Algorithms
Tian Guan, Guangxing Chu, Fei Chen 0011
INTERSPEECH3
2016 Understanding Periodically Interrupted Mandarin Speech
Rosanna H. N. Tong, Fei Chen 0011
INTERSPEECH3
2016 Relative Contributions of Amplitude and Phase to the Intelligibility Advantage of Ideal Binary Masked Sentences
Lei Wang 0074, Shufeng Zhu, Diliang Chen, Fei Chen 0011
INTERSPEECH5
2016 Modeling speech intelligibility with recovered envelope from temporal fine structure stimulus
Fei Chen 0011, Yu Tsao 0001, Ying-Hui Lai
Speech Commun.1
2014 Objective quality evaluation of noise-suppressed speech: effects of temporal envelope and fine-structure cues
Fei Chen 0011
INTERSPEECH1
2014 Effect of spectral degradation to the intelligibility of vowel sentences
Fei Chen 0011, Sharon W. K. Wong, Lena L. N. Wong
INTERSPEECH1
2014 An adaptive envelope compression strategy for speech processing in cochlear implants
abstract
Hearing-impaired patients have limited hearing dynamic range for speech perception, which partially accounts for their poor speech understanding abilities, particularly in noise. Wide dynamic range compression aims to compress speech signal into the usable hearing dynamic range of hearing-impaired listeners; however, it normally uses a static compression based strategy. This work proposed a strategy to continuously adjust the envelope compression ratio for speech processing in cochlear implants. This adaptive envelope compression (AEC) strategy aims to keep the compression processing as close to linear as possible, while still confine the compressed amplitude envelope within the pre-set dynamic range. Vocoder simulation experiments showed that, when narrowed down to a small dynamic range, the intelligibility of AEC-processed sentences was significantly better than those processed by static envelope compression. This makes the proposed AEC strategy a promising way to improve speech recognition performance for implanted patients in the future.
Ying-Hui Lai, Fei Chen 0011, Yu Tsao 0001
INTERSPEECH2
2014 Automatic speech recognition with primarily temporal envelope information
abstract
The aim of this study is to devise a computational method to predict cochlear implant (CI) speech recognition. Here, we describe a high-throughput screening system for optimizing CI speech processing strategies using hidden Markov model (HMM)-based automatic speech recognition (ASR). Word accuracy was computed on vocoded CI speech synthesized from primarily multi-channel temporal envelope information. The ASR performance increased with the number of channels in a similar manner displayed in human recognition scores. Results showed the computational method of HMM-based ASR offers better process control for comparing signal carrier type. Training-test mismatch reduction provided a novel platform for reevaluating the relative contributions of spectral and temporal cues to human speech recognition.
Payton Lin, Fei Chen 0011, Syu-Siang Wang, Ying-Hui Lai, Yu Tsao 0001
INTERSPEECH2
2013 Contributions of the high-RMS-level segments to the intelligibility of mandarin sentences
abstract
Recent evidence suggests that segments carrying more spectral changes [e.g., consonant-vowel boundaries in the middle root-mean-square (RMS) level segments] are important to predict the intelligibility of English sentences. Nevertheless, considering the difference between Mandarin and English languages, it is hypothesized that the high-RMS-level segments might provide more perceptual information to the intelligibility of Mandarin speech. Two studies were conducted in this paper to assess the relative contributions of the high-RMS-level segments to the intelligibility of Mandarin sentences, i.e., speech perception and intelligibility prediction. Results show that 1) Mandarin sentences containing the high-RMS-level (i.e., above the overall RMS level of the whole utterance) segments are more intelligible (i.e., recognition rate up to 91%) than those with the middle-RMS-level segments; and 2) the high-RMS-level segments, which carry more vowel and tonal information, contribute more in predicting the intelligibility of Mandarin sentences in noise.
Fei Chen 0011, Lena L. N. Wong
ICASSP1
2013 Effect of linguistic masker on the intelligibility of Mandarin sentences
Fei Chen 0011, Lena L. N. Wong, Yonghong Yan 0002
INTERSPEECH1
2013 Comparative investigation of objective speech intelligibility prediction measures for noise-reduced signals in Mandarin and Japanese
abstract
In this paper, eight state-of-the-art objective speech intelligibility prediction measures are comparatively investigated for noisy signals before and after noise-reduction processing between Mandarin and Japanese. Clean speech signals (Chinese words and Japanese words) were first corrupted by three types of noise at two signal-to-noise ratios and then processed by normal-hearing listeners for recognition, whose intelligibility was subsequently predicted by objective measures. Further investigations were conducted for objective measures in predicting speech intelligibility of noise-reduced signals between subjective evaluation scores and objective prediction results, and of noisy signals before and after noise-reduction processing, in terms of correlation analysis and prediction errors. Results showed that the majority of objective measures behave differently for Mandarin and Japanese in predicting the subjective ratings, and the STOI measure consistently provided the best ability in predicting the effect on speech intelligibility of the noise-reduction processing for both Mandarin and Japanese.
Fei Chen 0011, Masato Akagi, Yonghong Yan 0002
INTERSPEECH2
2013 A Hilbert-fine-structure-derived physical metric for predicting the intelligibility of noise-distorted and noise-suppressed speech
Fei Chen 0011, Lena L. N. Wong
Speech Commun.1
2012 Impact of SNR and gain-function over- and under-estimation on speech intelligibility
Fei Chen 0011, Philipos C. Loizou
Speech Commun.1
2010 Speech enhancement using a frequency-specific composite Wiener function
abstract
This paper introduces a new speech enhancement approach based on the design of a frequency-specific composite gain function for Wiener filtering. Motivated by the recently established finding that the acoustic cues at low frequencies can improve speech recognition in noise by the combined electric and acoustic stimulation technique, a less aggressive gain function is applied to replace the Wiener filter's conventional rigid gain function at low frequencies. With this modification, the proposed approach is able to recover more low frequency (LF) components and enhance the speech quality in noisy environments. An adaptive procedure is utilized to determine the low frequency boundary. Objective evaluation, based on the PESQ measure, revealed that the proposed approach with adaptive LF boundary improved significantly the PESQ scores compared to the scores obtained with the conventional Wiener filter in steady-state noise, white noise and babble interference conditions. Speech quality was also found to be significantly enhanced in the car noise conditions when the less aggressive gain function was selectively applied only to the consonants.
Fei Chen 0011, Philipos C. Loizou
ICASSP1
2008 A novel temporal fine structure-based speech synthesis model for cochlear implant
Fei Chen 0011, Yuan-Ting Zhang
Signal Process.1
2007 An integrate-and-fire-based auditory nerve model and its response to high-rate pulse train
Fei Chen 0011, Yuan-Ting Zhang
Neurocomputing1
2002 Signal processing techniques in genomic engineering
abstract
Now that the human genome has been sequenced, the measurement, processing, and analysis of specific genomic information in real time are gaining considerable interest because of their importance to better the understanding of the inherent genomic function, the early diagnosis of disease, and the discovery of new drugs. Traditional methods to process and analyze deoxyribonucleic acid (DNA) or ribonucleic acid data, based on the statistical or Fourier theories, are not robust enough and are time-consuming, and thus not well suited for future routine and rapid medical applications, particularly for emergency cases. In this paper, we present an overview of some recent applications of signal processing techniques for DNA structure prediction, detection, feature extraction, and classification of differentially expressed genes. Our emphasis is placed on the application of wavelet transform in DNA sequence analysis and on cellular neural networks in microarray image analysis, which can have a potentially large effect on the real-time realization of DNA analysis. Finally, some interesting areas for possible future research are summarized, which include a biomodel-based signal processing technique for genomic feature extraction and hybrid multidimensional approaches to process the dynamic genomic information in real time.
Fei Chen 0011, Yuan-Ting Zhang, Shannon Agner, Metin Akay, Zu-Hong Lu, Mary Miu Yee Waye, Stephen Kwok-Wing Tsui
Proc. IEEE2