Nengheng Zheng

dblp:97/5158 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-4493-1160ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Effects of harmonicity on Mandarin speech perception in cochlear implant users
Mingyue Shi, Huali Zhou, Yefei Mo, Nengheng Zheng
Speech Commun.6
2024 Two-Stage Acoustic Echo Cancellation Network with Dual-Path Alignment
abstract
Deep learning has become a popular approach for improving acoustic echo cancellation (AEC) in communication systems. However, existing systems mostly rely on traditional delay estimation methods, which often result in performance degradation due to inaccurate delay estimation. Furthermore, most deep learning-based methods use single-stage networks, the limited learning capability of which hinders the performance under harsh echo conditions. To address these challenges, this paper proposes a two-stage system with dual-path alignment. The two-stage strategy performs echo suppression on the magnitude spectrum in the first stage, followed by phase correction in the second one. In addition, the first stage processes two parallel features: the magnitude spectrum and its exponential compressed version, with which dual-path alignment is conducted to improve delay estimation. Experiment results demonstrate the effectiveness of the proposed system for echo suppression in challenging scenarios involving long delay, double-talk and nonlinear distortion.
Zhijian Jiang, Haoming Li 0020, Nengheng Zheng
ICASSP3
2024 DBD-CI: Doubling the Band Density for Bilateral Cochlear Implants
Mingyue Shi, Huali Zhou, Nengheng Zheng
INTERSPEECH4
2023 Cochlear-implant Listeners Listening to Cochlear-implant Simulated Speech
Fanhui Kong, Nengheng Zheng, Xianren Wang, Jan W. H. Schnupp
INTERSPEECH2
2023 Effects of hearing loss and amplification on Mandarin consonant perception
Huali Zhou, Xianming Bei, Nengheng Zheng
INTERSPEECH3
2023 F0inTFS: A lightweight periodicity enhancement strategy for cochlear implants
Huali Zhou, Fanhui Kong, Nengheng Zheng
INTERSPEECH3
2022 Internet Streaming Audio Based Speech Reception Threshold Measurement in Cochlear Implant Users
abstract
Traditional face-to-face subjective listening test has become a challenge due to the COVID-19 pandemic. We developed a remote assessment system with Tencent Meeting, a video conferencing application, to address this issue. This paper presents our work on evaluating the reliability of the remote assessment system. Two speech reception threshold (SRT) experiments were conducted to study the effects of noise suppression and maxima selection number on cochlear implant (CI) hearing. Both experiments were conducted locally and remotely, the correlations between the respective results were analyzed. Results showed that remote tests replicated the differences among testing conditions observed in local tests, but the absolute SRT values for individual conditions varied significantly between the two modes. The variations could be attributed to multiple reasons, such as online data transmission issues, audio playback devices, environmental conditions, and the training of participants. In conclusion, the relative variation of SRTs for CIs can be measured reliably, but the absolute SRT values should be carefully compared and explained according to objective and subjective experimental conditions.
Yefei Mo, Kang Ouyang, Mingyue Shi, Huali Zhou, Yupeng Shi, Shidong Shang, Nengheng Zheng
ICASSP10
2022 An EEG Source Imaging-based Feature Extraction Method for Motor Imagery Classification
abstract
This paper presents a new feature extraction method for Electroencephalogram (EEG)-based motor imagery (MI) classification. Current researches mostly classify different MIs by detecting the event-related desynchronization (ERD) phenomenon from the EEG signals. Due to the poor spatial resolution of the MI-EEG signals, the cortical area (source) activating the MI cannot be located accurately with the EEG (sensor) signal, which might degrade the classification accuracy. This study adopts the EEG source imaging (ESI) technique to estimate the cortical area where source ERD happens from the EEG signal. An improved ESI method based on the linearly constrained minimum variance (LCMV) algorithm, in which an average LCMV filter and an average baseline covariance are constructed for the ESI, is proposed to locate the activated cortical area from the noisy EEG signals. The source ERD features are then extracted. Analytical results show that, for subjects with obvious average source ERD phenomenon, their activated cortical area in a single-trial MI can be well located. MI classification results also support the feasibility of the proposed method for MI-EEG signal processing.
Junhan Li, Nengheng Zheng
SMC2
2021 A Noise-Robust Signal Processing Strategy for Cochlear Implants Using Neural Networks
abstract
Signal processing strategies in most clinical cochlear implants (CIs) extract and transmit speech envelopes to stimulate the auditory neurons. The incomplete representation of the rich fine structures in speech has significantly degraded the CI recipients’ ability in high- level perception, including their speech understanding in noise. This paper presents a noise-robust signal processing strategy to deal with this problem. Neural networks (NN) are built and trained to simulate the advanced combination encoder (ACE, a strategy for CI products of Cochlear Corporation). The NN-based ACE (namely, NNACE) is trained with a sophisticatedly designed loss function to output envelope-like signals that 1) is compatible with ACE-based CI system and can serve as the modulator to generate the electric stimuli, 2) is more noise-robust, and 3) might bear a certain degree of the temporal fine structures of speech. Subjective and objective evaluations with vocoder simulated speech show that NNACE outperforms the other methods and further actual CI experiments are warranted.
Nengheng Zheng, Yupeng Shi, Yuyong Kang
ICASSP1
2021 Feature Fusion by Attention Networks for Robust DOA Estimation
Rongliang Liu, Nengheng Zheng
Interspeech2
2020 Enhancing the Interaural Time Difference of Bilateral Cochlear Implants with the Temporal Limits Encoder
Yangyang Wan, Huali Zhou, Nengheng Zheng
INTERSPEECH4
2018 Weighting Pitch Contour and Loudness Contour in Mandarin Tone Perception in Cochlear Implant Listeners
Nengheng Zheng, Ambika Prasad Mishra, Jacinta Dan Luo, Jan W. H. Schnupp
INTERSPEECH2
2015 A temporal limits encoder for cochlear implants
abstract
Cochlear implant (CI) strategies extract multi-channel temporal envelopes to stimulate an array of electrodes. However, temporal pitch perception ability is known to be limited to a low frequency range at single channels. Therefore, this study proposes a temporal limits encoder (TLE) to make full use of the temporal processing abilities at each individual channels. First the total bandwidth is allocated to uniform narrow-bands, which are then down-shifted to the CI user's low frequency range (about 50–250 Hz) of temporal pitch. After the slowly varying signals are processed by half-wave rectification, the rectified signals are compressed and used to amplitude modulate interleaved high constant rate pulse sequences. By first deriving slowly varying signals, more information or even novel information about the original signal can be preserved based on the perceptual abilities of the implantees. Preliminary psychoacoustic frequency discrimination experiments suggest that the proposed TLE strategy can offer finer frequency resolution abilities.
Nengheng Zheng
ICASSP2
2014 Opportunistic user scheduling in MIMO cognitive radio networks
abstract
This paper studies multiuser diversity of uplink MIMO cognitive radio network and proposes a two-stage opportunistic user scheduling scheme. In the first stage, a cognitive beamforming design is proposed to ensure the interference caused by secondary signals is canceled or minimized on the spatial dimensions occupied by primary MIMO system. Then, some secondary users that cause minimal interference leakage at primary system are pre-selected as candidate users. In the second stage, some candidate users that produce maximum sum secondary rate are further selected for uplink scheduling. The proposed scheme enables the secondary link to take advantage of multiuser diversity while ensuring that the interference on primary link is within a certain threshold. Analytical results show that the sum rate of secondary uplink scales as Nslog logK for K secondary users and Nsantennas on secondary receiver for very large K.
Lu Yang 0001, Wei Zhang 0001, Nengheng Zheng, Pak-Chung Ching
ICASSP3
2011 Robust Speaker Recognition Using Denoised Vocal Source and Vocal Tract Features
abstract
To alleviate the problem of severe degradation of speaker recognition performance under noisy environments because of inadequate and inaccurate speaker-discriminative information, a method of robust feature estimation that can capture both vocal source- and vocal tract-related characteristics from noisy speech utterances is proposed. Spectral subtraction, a simple yet useful speech enhancement technique, is employed to remove the noise-specific components prior to the feature extraction process. It has been shown through analytical derivation, as well as by simulation results, that the proposed feature estimation method leads to robust recognition performance, especially at low signal-to-noise ratios. In the context of Gaussian mixture model-based speaker recognition with the presence of additive white Gaussian noise, the new approach produces consistent reduction of both identification error rate and equal error rate at signal-to-noise ratios ranging from 0 to 15 dB.
Ning Wang 0052, Pak-Chung Ching, Nengheng Zheng, Tan Lee
IEEE Trans. Speech Audio Process.3
2008 Language modeling for speech recognition of spoken Cantonese
Yu Ting Yeung, Houwei Cao, Nengheng Zheng, Tan Lee, Pak-Chung Ching
INTERSPEECH3
2007 Enhancement of Chinese speech based on nonlinear dynamics
Nengheng Zheng
Signal Process.2
2007 Integration of Complementary Acoustic Features for Speaker Recognition
abstract
This letter describes a speaker verification system that uses complementary acoustic features derived from the vocal source excitation and the vocal tract system. A new feature set, named the wavelet octave coefficients of residues (WOCOR), is proposed to capture the spectro-temporal source excitation characteristics embedded in the linear predictive residual signal. WOCOR is used to supplement the conventional vocal tract-related features, in this case, the Mel-frequency cepstral coefficients (MFCC), for speaker verification. A novel confidence measure-based score fusion technique is applied to integrate WOCOR and MFCC. Speaker verification experiments are carried out on the NIST 2001 database. The equal error rate (EER) attained with the proposed method is 7.67%, in comparison to 9.30% of the conventional MFCC-based system
Nengheng Zheng, Tan Lee, Pak-Chung Ching
IEEE Signal Process. Lett.1
2007 Discrimination Power of Vocal Source and Vocal Tract Related Features for Speaker Segmentation
abstract
This paper presents an analysis of the speaker discrimination power of vocal source related features, in comparison to the conventional vocal tract related features. The vocal source features, named wavelet octave coefficients of residues (WOCOR), are extracted by pitch-synchronous wavelet transform of the linear predictive (LP) residual signals. Using a series of controlled experiments, it is shown that WOCOR is less sensitive to spoken content than the conventional MFCC features and thus more discriminative when the amount of training data is limited. These advantages of WOCOR are exploited in the task of speaker segmentation for telephone conversation, in which statistical speaker models need to be built upon short speech segments. Experimental results show that the proposed use of WOCOR leads to noticeable reduction of segmentation errors.
Wai Nang Chan, Nengheng Zheng, Tan Lee
IEEE Trans. Speech Audio Process.2
2006 Use of Vocal Source Features in Speaker Segmentation
abstract
This paper addresses the problem of speaker segmentation in telephone conversation. The segmentation is done in three steps: 1) preliminary segmentation to hypothesize speaker turning points; 2) clustering of segments; and 3) re-segmentation to determine speaker identity of each segment. It is found that vocal source related features are more speaker-discriminative than the conventional vocal tract related features for small amount of data. This motivates us to thoughtfully incorporate vocal source features into early stages of the speaker segmentation process, where decisions have to be made with limited data. Speaker segmentation experiments are carried out on 36 summed channel conversations in the NIST 2004 Speaker Recognition Evaluation. The proposed use of vocal source features leads to noticeable performance improvement
Wai Nang Chan, Tan Lee, Nengheng Zheng, Hua Ouyang
ICASSP (1)3
2004 Using Haar transformed vocal source information for automatic speaker recognition
abstract
This paper attempts to investigate the effectiveness of incorporating vocal source information for enhancing automatic speaker recognition accuracy. We propose a new method to extract discriminative features from the linear prediction (LP) residual signal, which are closely related to the glottal excitation of individual speaker. A complementary parameter set in addition to the commonly used linear predictive cepstral coefficients (LPCC), called Haar octave coefficients of residue (HOCOR), is obtained by applying a Haar transform to the LP residue. This additional feature vector retains the spectro-temporal characteristics of the source excitation sequences that are related to the fundamental frequency, harmonics, as well as their phases. Experimental evaluation over the YOHO corpus demonstrates the high speaker discriminative power and high inter-speaker variability of HOCOR. Speaker recognition tests with both vocal tract feature (LPCC) and vocal source information (HOCOR) outperform the conventional methods of using LPCC only.
Nengheng Zheng, Pak-Chung Ching
ICASSP (1)1
2004 Time -frequency analysis of vocal source signal for speaker recognition
abstract
This paper investigates the importance of spectro-temporal characteristics of the source excitation signal for speaker recognition. We propose an effective feature extraction technique for obtaining essential time-frequency information from the linear prediction (LP) residual signal, which are closely related to the glottal excitation of individual speaker. With pitch synchro-nous analysis, wavelet transform is applied to every two pitch cycles of the LP residual signal to generate a new feature vector, called Wavelet Octave Coefficients of Residues (WOCOR), which provides additional speaker discriminative power to the commonly used linear predictive Cepstral coefficients (LPCC). Experimental evaluation over a Cantonese speaker recognition corpus demonstrates the effectiveness of WOCOR for speaker recognition. Recognition tests with WOCOR and LPCC outperforms the conventional methods of using Mel Frequency Cepstral Coefficients (MFCC). 1.
Nengheng Zheng, Pak-Chung Ching, Tan Lee
INTERSPEECH1