Xiongwei Zhang

dblp:86/2678 · DBLP profile ↗
← Back
40ranked-venue papers
2as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 5 since 2021Security and privacy · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Enhancing network monitoring in IoT with an energy-efficient collaborative framework using edge computing and federated learning
Zhaoyang Cui, Xiongwei Zhang, Shang Zhou
Expert Syst. Appl.2
2026 AMID: Audio-visual deepfake detection via adaptive multi-dimensional interaction modeling
Chenlong Xue, Kunyuan Li, Meng Sun 0001, Qiang Zhang 0054, Xiongwei Zhang, Kui Yao
Knowl. Based Syst.5
2026 Non-Iterative Reversible Information Hiding in the Sharing Domain With Adjustable Capacity
Kaili Qi, Jibin Yang, Tieyong Cao, Xiangli Xiao, Xiongwei Zhang, Yushu Zhang 0001, Zehang Wang
IEEE Trans. Dependable Secur. Comput.5
2025 Bayesian Nonparametric Clustering for Source Counting with a Small Aperture Microphone Array
abstract
Source counting (SC) in an indoor environment is an important problem in computational auditory scene analysis. However, the problem is challenging, especially when reverberation and ambient noise are present in the environment. To address this problem, we propose an augmented Bayesian non-parametric (ABNP) clustering algorithm for source counting based on sound intensity (SI) captured by a small aperture microphone array. The core idea is to incorporate an infinite Gaussian mixture model (IGMM) and a time-frequency (TF) augmented weight selection and update scheme for sound intensity estimation. The use of IGMM enables the exemption of the maximum number of sources assumed in previous methods. Experiments on both simulated and real-world data show the improved performance by the proposed method as compared with the state of the art baseline methods.
Kunkun SongGong, Pufen Zhang, Xiongwei Zhang, Wenwu Wang 0001, Meng Sun 0001, Chong Jia
ICASSP3
2025 Cross-domain redundancy exploration by a deep encoder-decoder network for speech steganography
abstract
The technique of speech steganography involves embedding messages within openly transmitted speech channels without arousing suspicion. Nevertheless, current methods for embedding speech in speech suffer from weak imperceptibility and low message speech intelligibility. In this paper, we introduce a novel approach that explores cross-domain redundancy by leveraging a deep encoder–decoder neural network architecture to embed Mel-spectrograms into magnitude spectrograms. Specifically, the message is transformed into its Mel-spectrogram, while the cover is transformed into its magnitude spectrogram. Subsequently, the Mel-spectrogram is embedded as residuals in the magnitude spectrogram through an encoder known as the spectrogram super-resolution network (SSRN). Upon receiving the stego, a decoder network recoveres the Mel-spectrograms of the messages, and a high-fidelity HiFi-GAN vocoder then recovers the message waveform. The encoder–decoder network’s parameters are optimized to ensure imperceptibility and high quality. To validate the superiority of our proposed method, we compare it with recently proposed baselines using common databases such as the LJ Speech and VCTK datasets. Experimental results demonstrate that our method achieves SNRs of 33.83 dB and 30.28 dB for the cover signals on these two datasets, respectively. Furthermore, both the content and speaker identity of the recovered messages are well preserved, and the experiments also confirm the robustness against noises and the security of our approach.
Xiaoyi Ge, Xiongwei Zhang, Meng Sun 0001, Kunkun SongGong
J. Inf. Secur. Appl.2
2025 A dynamically interactable framework with dual-channel security: GAN-based speech steganography for concealed dialogues
Xiaoyi Ge, Xiongwei Zhang, Meng Sun 0001
Knowl. Based Syst.2
2025 Introducing Euclidean distance optimization into Softmax loss under neural collapse
Qiang Zhang 0054, Xiongwei Zhang, Jibin Yang, Meng Sun 0001, Tieyong Cao
Pattern Recognit.2
2025 A Soft-Contrastive Pseudo Learning Approach Toward Open-World Forged Speech Attribution
abstract
Anti-spoofing of deepfake or forged speech is an important technique for the security usage of generative artificial intelligence. Beyond binary classification of real and forged speech, method attribution of forged speech is becoming a practical solution of interpretable anti-spoofing strategies. However, existing related methods have poor performance on analyzing speech forgery methods unseen in their training data, which is inefficient in open-world scenarios with emerging new forgery methods. In this paper, Open-World Forged Speech Attribution (OW-FSA) is firstly defined towards the attribution of forged speech on the methods generating it, where the recognized methods are not limited to the seen ones in training data and the properties of the unseen methods should also be depicted adequately. A novel algorithm, Soft-contrastive Pseudo Learning (SPL), is proposed to address the challenges outlined in OW-FSA, which introduces two key innovations: 1) Based on similarities between features at different scales, the proposed similarity-based soft filtering module filters and matches utterances from the same forgery class to enhance the intra-class compactness of features through contrastive learning. 2) The proposed similarity-based soft pseudo-labeling module integrates label-smoothing-like and similarity weighting techniques to mitigate possible errors in pseudo-labeling. Besides, an iterative algorithm based on SPL is proposed to predict the number of unseen classes. Extensive experiments have validated the superiority of the proposed algorithm over other recently proposed methods on the task of OW-FSA with or without the knowledge of the number of unseen classes. Intuitive visualization and ablation studies have also been conducted to illustrate the advantages of the proposed algorithm. The newly defined task OW-FSA and the proposed algorithm SPL in this paper will help advance the research in speech anti-spoofing.
Qiang Zhang 0054, Xiongwei Zhang, Meng Sun 0001, Jibin Yang
IEEE Trans. Inf. Forensics Secur.2
2024 Multi-Speaker Localization in the Circular Harmonic Domain on Small Aperture Microphone Arrays Using Deep Convolutional Networks
abstract
Acoustic signal processing in the circular harmonic domain (CHD) is an appealing method for speaker localization, since it inherently supports wideband acoustic sources and provides frequency invariant beampatterns. However, the performance of existing circular harmonic direction-of-arrival (DOA) estimation approaches can be degraded by a variety of factors, including background noise and reverberation in the acoustic environments, small aperture size of the circular array and the presence of multiple active sources. This paper addresses these issues by proposing a novel multi-speaker CHD localization method with small-sized microphone arrays using deep convolutional neural networks (CNN). The core idea is to construct circular harmonic features through joining the selected time-frequency (TF) bins of higher power and the operation of a randomization process by mimicking the sparsity property of speech signals. After that, we implement multi-speaker estimation as a multi-label classification task, and propose to use CNN with binary cross-entropy as the loss function. Experimental results show that our method performs significantly better than the baseline methods, on both simulated and real data, in terms of the accuracy of DOA estimation.
Kunkun SongGong, Pufen Zhang, Xiongwei Zhang, Meng Sun 0001, Wenwu Wang 0001
ICASSP3
2024 Scale-aware dual-branch complex convolutional recurrent network for monaural speech enhancement
Meng Sun 0001, Xiongwei Zhang, Hugo Van hamme
Comput. Speech Lang.3
2024 Noise-robust voice conversion using adversarial training with multi-feature decoupling
Xiongwei Zhang, Meng Sun 0001
Eng. Appl. Artif. Intell.2
2024 Target speaker filtration by mask estimation for source speaker traceability in voice conversion
Xiongwei Zhang, Meng Sun 0001, Xia Zou, Chong Jia
Eng. Appl. Artif. Intell.2
2024 An attack-agnostic defense method against adversarial attacks on speaker verification by fusing downsampling and upsampling of speech signals
Xiongwei Zhang, Meng Sun 0001, Yinan Li 0006
Inf. Sci.2
2024 An efficient low-perceptual environmental sound classification adversarial method based on GAN
Jibin Yang, Xiongwei Zhang, Tieyong Cao
Multim. Tools Appl.3
2023 Self-distillation object segmentation via pyramid knowledge representation and transfer
Meng Sun 0001, Tieyong Cao, Xiongwei Zhang, Lixing Xing, Zheng Fang 0012
Multim. Syst.5
2023 DRRNets: Dynamic Recurrent Routing via Low-Rank Regularization in Recurrent Neural Networks
abstract
Recurrent neural networks (RNNs) continue to show outstanding performance in sequence learning tasks such as language modeling, but it remains difficult to train RNNs for long sequences. The main challenges lie in the complex dependencies, gradient vanishing or exploding, and low resource requirement in model deployment. In order to address these challenges, we propose dynamic recurrent routing neural networks (DRRNets), which can: 1) shorten the recurrent lengths by allocating recurrent routes dynamically for different dependencies and 2) reduce the number of parameters significantly by imposing low-rank constraints on the fully connected layers. A novel optimization algorithm via low-rank constraint and sparsity projection is developed to train the network. We verify the effectiveness of the proposed method by comparing it with multiple competitive approaches in several popular sequential learning tasks, such as language modeling and speaker recognition. The results in terms of different criteria demonstrate the superiority of our proposed method.
Dongjing Shan, Yong Luo 0002, Xiongwei Zhang, Chao Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Waveform level adversarial example generation for joint attacks against both automatic speaker verification and spoofing countermeasures
Xiongwei Zhang, Wei Liu 0005, Xia Zou, Meng Sun 0001, Jian Zhao 0006
Eng. Appl. Artif. Intell.2
2022 Spatial-temporal slowfast graph convolutional network for skeleton-based action recognition
abstract
Abstract In skeleton‐based action recognition, the graph convolutional network (GCN) has achieved great success. Modelling skeleton data in a suitable spatial‐temporal way and designing the adjacency matrix are crucial aspects for GCN‐based methods to capture joint relationships. In this study, we propose the spatial‐temporal slowfast graph convolutional network (STSF‐GCN) and design the adjacency matrices for the skeleton data graphs in STSF‐GCN. STSF‐GCN contains two pathways: (1) the fast pathway is in a high frame rate, and joints of adjacent frames are unified to build ‘small’ spatial‐temporal graphs. A new spatial‐temporal adjacency matrix is proposed for these ‘small’ spatial‐temporal graphs. Ablation studies verify the effectiveness of the proposed adjacency matrix. (2) The slow pathway is in a low frame rate, and joints from all frames are unified to build one ‘big’ spatial‐temporal graph. The adjacency matrix for the ‘big’ spatial‐temporal graph is obtained by computing self‐attention coefficients of each joint. Finally, outputs from two pathways are fused to predict the action category. STSF‐GCN can efficiently capture both long‐range and short‐range spatial‐temporal joint relationships. On three datasets for skeleton‐based action recognition, STSF‐GCN can achieve state‐of‐the‐art performance with much less computational cost.
Zheng Fang 0012, Xiongwei Zhang, Tieyong Cao, Meng Sun 0001
IET Comput. Vis.2
2022 SO-softmax loss for discriminable embedding learning in CNNs
Qiang Zhang 0054, Jibin Yang, Xiongwei Zhang, Tieyong Cao
Pattern Recognit.3
2022 Underdetermined Blind Direction-of-Arrival Estimation Using a Moving Platform
abstract
Underdetermined blind DOA estimation for linear arrays without the prior knowledge of array configuration is a challenging problem. Two methods are proposed to solve this problem based on a moving platform. The first one exploits the structure of the equivalent array manifold matrix constructed by using several different time intervals for the array motions and utilizes iterative optimization to obtain the estimation of DOAs. By considering secondary time intervals for the array motions further, the other combines the DOA matrix and multistage Wiener filter (MSWF) to estimate DOAs in a more cost-effective way. Simulation results and complexity analysis testify the effectiveness and superiority of the proposed algorithms.
Qiao Su, Xiongwei Zhang, Nan Sha, Kui Xu 0001
IEEE Signal Process. Lett.2
2021 The spectra and reticulation of EQ-algebras
Xiongwei Zhang
Soft Comput.1
2021 When Automatic Voice Disguise Meets Automatic Speaker Verification
abstract
The technique of transforming voices in order to hide the real identity of a speaker is called voice disguise, among which automatic voice disguise (AVD) by modifying the spectral and temporal characteristics of voices with miscellaneous algorithms are easily conducted with softwares accessible to the public. AVD has posed great threat to both human listening and automatic speaker verification (ASV). In this paper, we have found that ASV is not only a victim of AVD but could be a tool to beat some simple types of AVD. Firstly, three types of AVD, pitch scaling, vocal tract length normalization (VTLN) and voice conversion (VC), are introduced as representative methods. State-of-the-art ASV methods are subsequently utilized to objectively evaluate the impact of AVD on ASV by equal error rates (EER). Moreover, an approach to restore disguised voice to its original version is proposed by minimizing a function of ASV scores w.r.t. restoration parameters. Experiments are then conducted on disguised voices from Voxceleb, a dataset recorded in real-world noisy scenario. The results have shown that, for the voice disguise by pitch scaling, the proposed approach obtains an EER around 7% comparing to the 30% EER of a recently proposed baseline using the ratio of fundamental frequencies. The proposed approach generalizes well to restore the disguise with nonlinear frequency warping in VTLN by reducing its EER from 34.3% to 18.5%. However, it is difficult to restore the source speakers in VC by our approach, where more complex forms of restoration functions or other paralinguistic cues might be necessary to restore the nonlinear transform in VC. Finally, contrastive visualization on ASV features with and without restoration illustrate the role of the proposed approach in an intuitive way.
Linlin Zheng, Jiakang Li, Meng Sun 0001, Xiongwei Zhang, Thomas Fang Zheng
IEEE Trans. Inf. Forensics Secur.4
2019 Some weaker versions of topological residuated lattices
Xiongwei Zhang
Fuzzy Sets Syst.2
2019 Finite direct products of EQ-algebras
Xiongwei Zhang
Soft Comput.2
2019 Detection of People With Camouflage Pattern Via Dense Deconvolution Network
abstract
In this letter, we explore the detection of people with camouflage pattern in cluttered natural scenes. First, considering the lack of open evaluation and training data on camouflaged people detection, a specific dataset of camouflaged people in natural scenes is constructed by us for the first time to the best of our knowledge. Secon, due to the serious corruption of the discrimination of low-level features by the camouflage patterns and cluttered background, we extract the high-level semantic features in deep convolution network and introduce short connections in deconvolution phase, to construct the dense deconvolution network. In training procedure, we augment and shift the images of camouflaged people to generate the proper training data. Attributing to the usage and fusion of semantic information, the proposed network effectively labels the camouflaged people regions as a whole. Finally, we use the superpixel segmentation and spatial smoothness constraint for further improvement of the detection result. Experimental results demonstrate that the proposed method outperforms the classical camouflaged object detection method and typical CNN-based detection methods.
Xiongwei Zhang, Feng Wang 0032, Tieyong Cao, Meng Sun 0001
IEEE Signal Process. Lett.2
2018 Visual saliency based on extended manifold ranking and third-order optimization refinement
Dongjing Shan, Xiongwei Zhang, Chao Zhang 0001
Pattern Recognit. Lett.2
2017 Auditory mask estimation by RPCA for monaural speech enhancement
abstract
Mask estimation has shown a lot of promise in speech enhancement for its simplicity and large speech intelligibility improvement. In this paper, the gammachirp filter banks are applied on the contaminated speech signal to get the auditory time-frequency representation. Robust principal component analysis with non-negative constraint is employed to decompose the auditory time-frequency representation into sparse and low-rank components using alternating direction method of multipliers optimization algorithm. Auditory Mask is estimated by these two parts which are correspond to the speech and noise. Consider that binary mask produces separated sources with more distortion than soft mask estimation. Auditory mask estimation is based on the ideal ratio mask estimation. Experimental results show that the proposed method could achieve better performance in terms of PESQ and LSD compared with multiband spectral subtraction and Robust principal component analysis methods.
Wenhua Shi, Xiongwei Zhang, Xia Zou, Gang Min
ICIS2
2016 Adaptive extraction of repeating non-negative temporal patterns for single-channel speech enhancement
abstract
Estimating unknown background noise from single-channel noisy speech is a key yet challenging problem for speech enhancement. Given the fact that the background noises typically have the repeating property and the foreground speech is sparse and time-variant, many literatures decompose the noisy spectrogram directly in an unsupervised fashion when there is no isolated training example of the target speaker or particular noise types beforehand. However, recently proposed methods suffer from un-interpretable decomposed patterns, neglecting the temporal structure of the background noise or being constrained by the pre-fixed parameters. To settle these issues, we propose a novel method based on autocorrelation technique and convolutive non-negative matrix factorization. The proposed method can adaptively estimate the underlying non-negative repeating temporal patterns from noisy speech and identify the clean speech spectrogram simultaneously. Experiments on NOIZEUS dataset mixed with various real-world background noises showed that the proposed method performs better than some state-of-the-art methods.
Yinan Li 0006, Xiongwei Zhang, Meng Sun 0001, Gang Min, Jibin Yang
ICASSP2
2016 Joint optimization of audible noise suppression and deep neural networks for single-channel speech enhancement
abstract
Improving the perceptual quality of speech signals is a key yet challenging problem for many real world applications. Taking into account the good performance of deep learning in signal representation, a novel single-channel speech enhancement technique is presented based on joint Deep Neural Networks and audible noise suppression as a whole network architecture. This new deep neural network jointly trains an audible noise suppression function which is used to estimate the magnitude spectrum of the clean speech and shape the spectrum of the audible noise at the same time. Experimental results on TIMIT with 20 noise types at various noise levels demonstrate the superiority of the proposed method over the baselines, no matter whether the noise conditions are included in the training set or not.
Xiongwei Zhang, Gang Min, Meng Sun 0001, Jibin Yang
ICME2
2016 A perceptually motivated approach via sparse and low-rank model for speech enhancement
abstract
A perceptually motivated speech enhancement approach is proposed in this paper. Different from the conventional sparse and low-rank model based approaches, this new approach takes into account the perceptual differences in different frequency bands of the human auditory system, and separates speech from background noises in the Mel spectral domain. After two propositions for the Mel frequency weighted spectrogram are proved, speech enhancement can be modeled as a sparse and low-rank constrained optimization problem, which is solved efficiently by the alternating direction method of multipliers (ADMM). The proposed approach is totally unsupervised, neither the speech nor the noise dictionary needs to be trained beforehand. The experimental results have shown its promising performance under strong background noises. The performance can be further improved by information fusion technique at high input SNRs.
Gang Min, Xiongwei Zhang, Jibin Yang, Xia Zou
ICME2
2016 Video saliency detection using 3D shearlet transform
Xiongwei Zhang
Multim. Tools Appl.2
2016 Perceptually Weighted Analysis-by-Synthesis Vector Quantization for Low Bit Rate MFCC Codec
abstract
This letter presents a perceptually weighted analysis-by-synthesis vector quantization (VQ) algorithm for low bit rate MFCC codec. Different from conventional VQ of mel-frequency cepstral coefficients (MFCCs) vector, this algorithm uses an analysis-by-synthesis technique and aims to minimize the perceptually weighted spectral reconstruction distortion rather than the distortion of MFCCs vector itself. Also, to reduce the computational complexity, we propose a practical suboptimal codebook searching technique and embed it into the split and multistage VQ framework. Objective and subjective experimental results on Mandarin speech show that the proposed algorithm yields intelligible and natural sounding speech for speech coding at 600-2400 bit/s. Compared to current VQ in MFCC codec, the output speech quality is substantially improved in terms of frequency-weighted segmental SNR, short-time objective intelligibility score, perceptual evaluation of speech quality score, and mean opinion score.
Gang Min, Xiongwei Zhang, Xia Zou, Jibin Yang
IEEE Signal Process. Lett.2
2016 Unseen Noise Estimation Using Separable Deep Auto Encoder for Speech Enhancement
abstract
Unseen noise estimation is a key yet challenging step to make a speech enhancement algorithm work in adverse environments. At worst, the only prior knowledge we know about the encountered noise is that it is different from the involved speech. Therefore, by subtracting the components which cannot be adequately represented by a well defined speech model, the noises can be estimated and removed. Given the good performance of deep learning in signal representation, a deep auto encoder (DAE) is employed in this work for accurately modeling the clean speech spectrum. In the subsequent stage of speech enhancement, an extra DAE is introduced to represent the residual part obtained by subtracting the estimated clean speech spectrum (by using the pre-trained DAE) from the noisy speech spectrum. By adjusting the estimated clean speech spectrum and the unknown parameters of the noise DAE, one can reach a stationary point to minimize the total reconstruction error of the noisy speech spectrum. The enhanced speech signal is thus obtained by transforming the estimated clean speech spectrum back into time domain. The above proposed technique is called separable deep auto encoder (SDAE). Given the under-determined nature of the above optimization problem, the clean speech reconstruction is confined in the convex hull spanned by a pre-trained speech dictionary. New learning algorithms are investigated to respect the non-negativity of the parameters in the SDAE. Experimental results on TIMIT with 20 noise types at various noise levels demonstrate the superiority of the proposed method over the conventional baselines.
Meng Sun 0001, Xiongwei Zhang, Hugo Van hamme, Thomas Fang Zheng
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 SegBOMP: An efficient algorithm for block non-sparse signal recovery
abstract
Block sparse signal recovery methods have attracted great interests which take the block structure of the nonzero coefficients into account when clustering. Compared with traditional compressive sensing methods, it can obtain better recovery performance with fewer measurements by utilizing the block-sparsity explicitly. In this paper we propose a segmented-version of the block orthogonal matching pursuit algorithm in which it divides any vector into several sparse sub-vectors. By doing this, the original method can be significantly accelerated due to the dimension reduction of measurements for each segmented vector. Experimental results showed that with low complexity the proposed method yielded identical or even better reconstruction performance than the conventional methods which treated the signal in the standard block-sparsity fashion. Furthermore, in the specific case, where not all segments contain nonzero blocks, the performance improvement can be interpreted as a gain in “effective SNR” in noisy environment.
Xushan Chen, Xiongwei Zhang, Jibin Yang, Meng Sun 0001
ICME2
2015 Speech enhancement based on robust NMF solved by alternating direction method of multipliers
abstract
A robust version of non-negative matrix factorization (RNMF) with generalized Kullback-Leibler divergence designed for the task of unsupervised monaural speech enhancement is proposed. RNMF tackles unsupervised speech enhancement problem through factorizing the magnitude spectrum of mixture into the sum of a non-negative sparse matrix and a non-negative low-rank matrix. The parameters of nonnegative components are estimated through minimizing the reconstruction error defined by the divergence. The closed-from updating formulae of RNMF are derived using alternating direction method of multipliers. Experimental results demonstrated that the proposed algorithm yields superior results compared with the multiplicative updates at the expense of more computational complexity.
Yinan Li 0006, Xiongwei Zhang, Meng Sun 0001, Jingfeng Pan
MMSP2
2015 Speech reconstruction from mel-frequency cepstral coefficients via ℓ1-norm minimization
abstract
This paper presents a high quality speech reconstruction method from Mel-frequency cepstral coefficients (MFCC). Due to the sparse characteristic of the power spectrum of speech, the ℓ1-norm minimization method is used to tackle the under-determined nature of the speech reconstruction problem. The phase spectrum is recovered by the well-known LSE-ISTFTM algorithm. Experimental results demonstrate that the quality of the reconstructed speech is dramatically improved than the common ℓ2-norm minimization method, it sounds very close to the original speech when using the high-resolution MFCC, the PESQ score reaches 4.0.
Gang Min, Xiongwei Zhang, Jibin Yang, Xia Zou
MMSP2
2015 A stable approach for model order selection in nonnegative matrix factorization
Meng Sun 0001, Xiongwei Zhang, Hugo Van hamme
Pattern Recognit. Lett.2
2015 Speech Enhancement Under Low SNR Conditions Via Noise Estimation Using Sparse and Low-Rank NMF with Kullback-Leibler Divergence
abstract
A key stage in speech enhancement is noise estimation which usually requires prior models for speech or noise or both. However, prior models can sometimes be difficult to obtain. In this paper, without any prior knowledge of speech and noise, sparse and low-rank nonnegative matrix factorization (NMF) with Kullback-Leibler divergence is proposed to noise and speech estimation by decomposing the input noisy magnitude spectrogram into a low-rank noise part and a sparse speech-like part. This initial unsupervised speech-noise estimation allows us to set a subsequent regularized version of NMF or convolutional NMF to reconstruct the noise and speech spectrogram, either by estimating a speech dictionary on the fly (categorized as unsupervised approaches) or by using a pre-trained speech dictionary on utterances with disjoint speakers (categorized as semi-supervised approaches). Information fusion was investigated by taking the geometric mean of the outputs from multiple enhancement algorithms. The performance of the algorithms were evaluated on five metrics (PESQ, SDR, SNR, STOI, and OVERALL) by making experiments on TIMIT with 15 noise types. The geometric means of the proposed unsupervised approaches outperformed spectral subtraction (SS), minimum mean square estimation (MMSE) under low input SNR conditions. All the proposed semi-supervised approaches showed superiority over SS and MMSE and also obtained better performance than the state-of-the-art algorithms which utilized a prior noise or speech dictionary under low SNR conditions.
Meng Sun 0001, Yinan Li 0006, Jort F. Gemmeke, Xiongwei Zhang
IEEE ACM Trans. Audio Speech Lang. Process.4
2000 Research on Speech Recognition Based on Phase Space Reconstruction Theory
Yanxin Chen, Xiongwei Zhang
ICMI3
1992 A new excitation model for LPC vocoder at 2.4 kb/s
abstract
A novel excitation model called the multicategory vector excitation (MCVE) model for a linear predictive coding (LPC) vocoder at 2.4 kb/s is proposed. In this model, speech signal is classified into four categories: unvoiced, voiced, onset, and offset. For every category of speech, an excitation codebook is available. Different excitation codebooks hold different characteristics. The analysis-by-synthesis procedure is used to select the excitation vectors. The computer simulation has been carried out, and the results show that the vocoder with the new excitation model is capable of synthesizing more intelligible and more natural speech at 2.4 kb/s.>
Xiongwei Zhang, Chen Xianzhi
ICASSP1