VLDB 2026 Research / reviewers in the wild / expert
Xinmeng Xu
dblp:221/3983
· DBLP profile ↗
26ranked-venue papers
16as first author
26since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 12 first-author · 22 since 2021Artificial intelligence and machine learning · 13 · 10 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Speech Enhancement by Cross- and Sub-band Processing with State Space ModelabstractRecently, the state space model (SSM) represented by Mamba has shown remarkable performance in long-term sequence modeling tasks, including speech enhancement. However, due to substantial differences in sub-band features, applying the same SSM to all sub-bands limits its inference capability. Additionally, when processing each time frame of the time-frequency representation, the SSM may forget certain high-frequency information of low energy, making the restoration of structure in the high-frequency bands challenging. For this reason, we propose Cross- and Sub-band Mamba (CSMamba). To assist the SSM in handling different sub-band features flexibly, we propose a band split block that splits the full-band into four sub-bands with different widths based on their information similarity. We then allocate independent weights to each sub-band, thereby reducing the inference burden on the SSM. Furthermore, to mitigate the forgetting of low-energy information in the high-frequency bands by the SSM, we introduce a spectrum restoration block that enhances the representation of the cross-band features from multiple perspectives. Experimental results on the DNS Challenge 2021 dataset demonstrate that CSMamba outperforms several state-of-the-art (SOTA) speech enhancement methods in three objective evaluation metrics with fewer parameters. Jizhen Li, Weiping Tu, Yuhong Yang 0001, Xinmeng Xu, Yanzhen Ren |
ICASSP | 4 |
| 2025 | Attention Weighting and Conditional Entropy-driven Quantization Loss for Neural Audio CodecsabstractExisting end-to-end neural codecs have made great progress in preserving audio quality. Despite their success, they still face challenges in achieving accurate and efficient quantization. Specifically, these codecs often overlook which features have a greater impact on perceptual audio quality during quantization, leading to a quantization error distribution that fails to reflect the actual importance of latent features. They are also sensitive to unusual data points (outliers) because they use Mean Squared Error (MSE) to measure quantization errors, which can increase quantization noise or spectral artifacts. To address these limitations, we propose AW-CEQCodec which integrates an Attention Weighting (AW) module and a Conditional Entropy-driven Quantization (CEQ) loss. The AW enhances key regions of latent features before quantization, enabling more accurate quantizing critical features and reducing their quantization errors. After quantization, it restores global details from dequantized features, improving overall reconstruction. Moreover, the CEQ minimizes the uncertainty between latent and quantized features, effectively reflecting the distortion introduced by the quantization module. Experimental results on the CodecSuperb-STL dataset demonstrate that our method consistently outperforms baseline approaches, achieving superior audio quality at bitrates as low as 0.5 kbps, confirming its effectiveness in minimizing distortion and preserving perceptual quality. The reconstruction audio samples can be find at https://huazhi1024.github.io/first-page. Weiping Tu, Yuhong Yang 0001, Xinmeng Xu, Yanzhen Ren |
ICASSP | 4 |
| 2025 | FIRING-Net: A filtered feature recycling network for speech enhancementabstractCurrent deep neural networks for speech enhancement (SE) aim to minimize the distance between the output signal and the clean target by filtering out noise features from input features. However, when noise and speech components are highly similar, SE models struggle to learn effective discrimination patterns. To address this challenge, we propose a Filter-Recycle-Interguide framework termed Filter-Recycle-INterGuide NETwork (FIRING-Net) for SE, which filters the input features to extract target features and recycles the filtered-out features as non-target features. These two feature sets then guide each other to refine the features, leading to the aggregation of speech information within the target features and noise information within the non-target features. The proposed FIRING-Net mainly consists of a Local Module (LM) and a Global Module (GM). The LM uses outputs of the speech extraction network as target features and the residual between input and output as non-target features. The GM leverages the energy distribution of self-attention map to extract target and non-target features guided by highest and lowest energy regions. Both LM and GM include interaction modules to leverage the two feature sets in an inter-guided manner for collecting speech from non-target features and filtering out noise from target features. Experiments confirm the effectiveness of the Filter-Recycle-Interguide framework, with FIRING-Net achieving a strong balance between SE performance and computational efficiency, surpassing comparable models across various SNR levels and noise environments. Xinmeng Xu, Jizhen Li, Yuhong Yang 0001, Yong Luo 0002, Weiping Tu |
ICLR | 1 |
| 2025 | PGD-N2L: A Parameter-Guided Disentanglement Approach for Normal-To-Lombard Speech ConversionabstractThe Normal-To-Lombard (N2L) speech conversion can effectively improve speech intelligibility in noisy communication scenarios and serve as a data augmentation tool for various speech-related algorithms. However, existing N2L methods did not aim to disentangle the Lombard effect from other speech attributes, leading to incomplete conversions. In this paper, we propose a Parameter-Guided Disentanglement approach for N2L speech conversion (PGD-N2L) which decomposes speech into linguistic content, speaker identity, and Lombard effect. To extract disentangled linguistic content, we propose a DeLomb-Based content encoder. To extract disentangled speaker identity and Lombard effect, we propose a style encoder that combines a fine-tuned speaker encoder and a learnable Lombard encoder to form a personalized style embedding. Furthermore, an En-Lomb-Based injection module is designed to accurately integrate the target Lombard effect and speaker identity into the linguistic content based on personalized style embedding, ensuring complete Lombard conversion. Experimental results demonstrate that our proposed method outperforms existing N2L models in speech intelligibility, acoustic similarity, and speech quality. Ablation studies confirm that the fine-tuned speaker encoder and the De-Lomb block effectively improve speech intelligibility and acoustic similarity, while the En-Lomb block enables the converted speech to more closely match the target Lombard speech. Hongyang Chen 0004, Yuhong Yang 0001, Xinmeng Xu, Weiping Tu, Zhongyuan Wang 0001, Cedar Lin |
ICME | 3 |
| 2025 | Spatial information aided speech and noise feature discrimination for Monaural speech enhancement
Xinmeng Xu, Jizhen Li, Weiping Tu, Yuhong Yang 0001 |
Expert Syst. Appl. | 1 |
| 2024 | Curricular Contrastive Regularization for Speech Enhancement with Self-Supervised RepresentationsabstractExisting deep learning-based speech enhancement methods only adopt clean speech as positive samples to guide the training of speech enhancement networks while negative samples, i.e., noisy speech, are unexploited. In this paper, we adopt contrastive regularization (CR) built upon contrastive learning to exploit both the information of noisy and clean speech as negative and positive samples, respectively. Particularly, CR minimizes the distance between clean and enhanced speech and maximizes the distance between noisy and enhanced speech in the representation space of the self-supervised learning model. However, the contrastive samples are non-consensual, as the negatives are usually represented distantly from the clean speech, leaving the solution space still under-constricted. To tackle this issue, we provide the negative samples assembled from (1) the noisy speech, and (2) the corresponding enhanced speech without using CR, and we customize a curriculum learning strategy to define the importance of these negative samples to balance the learning difficulty caused by different similarities between the embeddings of the positive and negative samples. Experiments show that our proposal improves SE performance effectively without introducing additional computation/parameters. Xinmeng Xu, Chang Han, Weiping Tu, Yuhong Yang 0001 |
ICASSP | 1 |
| 2024 | An Efficient and Interpre Table Speech Enhancement Network Via Deep Dictionary LearningabstractSpeech enhancement is a vital and highly ill-posed problem for many speech downstream tasks. While currently existing deep learning based speech enhancement methods have held state-of-the-art results, they still possess apparent shortcomings in that most of the deep learning based models lack interpretability. This deficiency results in unsatisfied speech enhancement performance in many sophisticated scenarios. To tackle this problem, we integrate dictionary learning and sparse coding into deep learning networks for speech enhancement and present a deep dictionary learning based speech enhancement network (DicLSENet). Specifically, the proposed DicLSENet strictly follows the principle of dictionary learning, learns the priors for both representation coefficients and dictionaries, and adaptively adjusts the dictionary for each input. Experimental results show that the proposed model outperforms state-of-the-art fully deep learning based methods with attractive computational costs. Xinmeng Xu, Weiping Tu, Yuhong Yang 0001 |
ICASSP | 1 |
| 2024 | Improving Acoustic Echo Cancellation by Exploring Speech and Echo Affinity with Multi-Head AttentionabstractDeep learning-based approaches formulate acoustic echo cancellation (AEC) as a supervised speech separation task, where the mixture signal and the far-end signal are combined directly before or after the encoding stage. However, the mixture signal and the far-end signal are not integrated sufficiently due to the lack of interpretability for the affinity between speech and echo in a noisy mixture. In this paper, we propose DCA-Net, a dual-branch cross-attention neural network, to improve AEC performance by exploring the affinities between speech and echo in the representation space. In particular, the two branches predict speech and echo, respectively, and an interaction module is designed at several intermediate feature domains between the two branches to learn the correlations between these features of the two branches. Such an interaction can leverage features learned from one branch to restore missing information or counteract undesired information of the other by calculating the similarity between these features of two branches using multi-head cross attention. Evaluation results show that the proposed DCA-Net effectively suppresses acoustic echo and noise while preserving good speech quality. Xinmeng Xu, Weiping Tu |
ICASSP | 2 |
| 2024 | Srcodec: Split-Residual Vector Quantization for Neural Speech CodecabstractEnd-to-end neural speech coding achieves state-of-the-art performance by using residual vector quantization. However, it is a challenge to quantize the latent variables with as few bits as possible. In this paper, we propose SRCodec, a neural speech codec that relies on a fully convolutional encoder/decoder network with specifically proposed split-residual vector quantization. In particular, it divides the latent representation into two parts with the same dimensions. We utilize two different quantizers to quantize the low-dimensional features and the residual between the low- and high-dimensional features. Meanwhile, we propose a dual attention module in split-residual vector quantization to improve information sharing along both dimensions. Both subjective and objective evaluations demonstrate that the effectiveness of our proposed method can achieve a higher quality of reconstructed speech at 0.95 kbps than Lyra-v1 at 3 kbps and Encodec at 3 kbps. Youqiang Zheng, Weiping Tu, Li Xiao 0007, Xinmeng Xu |
ICASSP | 4 |
| 2024 | SuperCodec: A Neural Speech Codec with Selective Back-Projection NetworkabstractNeural speech coding is a rapidly developing topic, where state-of-the-art approaches now exhibit superior compression performance than conventional methods. Despite significant progress, existing methods still have limitations in preserving and reconstructing fine details for optimal reconstruction, especially at low bitrates. In this study, we introduce SuperCodec, a neural speech codec that achieves state-of-the-art performance at low bitrates. It employs a novel back projection method with selective feature fusion for augmented representation. Specifically, we propose to use Selective Up-sampling Back Projection (SUBP) and Selective Down-sampling Back Projection (SDBP) modules to replace the standard up- and down-sampling layers at the encoder and decoder, respectively. Experimental results show that our method outperforms the existing neural speech codecs operating at various bitrates. Specifically, our proposed method can achieve higher quality reconstructed speech at 1 kbps than Lyra V2 at 3.2 kbps and Encodec at 6 kbps. Youqiang Zheng, Weiping Tu, Li Xiao 0007, Xinmeng Xu |
ICASSP | 4 |
| 2024 | Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
Jizhen Li, Xinmeng Xu, Weiping Tu, Yuhong Yang 0001 |
INTERSPEECH | 2 |
| 2024 | Adaptive selection of local and non-local attention mechanisms for speech enhancement
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001 |
Neural Networks | 1 |
| 2024 | Improving Monaural Speech Enhancement by Mapping to Fixed Simulation Space With Knowledge DistillationabstractMonaural speech enhancement (SE) is a versatile and cost-effective approach that leverages recordings from a single microphone. However, it falls short of multi-channel SE due to the absence of spatial cues. These cues, present in multi-channel recordings, aid in distinguishing speech from noise more effectively. To bridge this gap, we introduce a method for mapping monaural speech into a fixed simulation space. Here, single-channel recordings are transformed into a predefined binaural format, enhancing the differentiation between target speech and noise components. This is achieved through knowledge distillation, enabling the monaural SE model to learn simulated binaural speech features from a pre-trained binaural SE model. It is important to note that we use a single type of binaural room impulse response and the monaural input of the student to simulate binaural speech. This way, our approach bypasses the paradox of generating virtual spatial information from monaural speech, while still benefiting from the spatial cues of binaural speech. Rigorous experiments demonstrate the effectiveness of our proposed method, showcasing its superior performance compared to recent monaural SE techniques in terms of PESQ and STOI scores. Xinmeng Xu |
IEEE Signal Process. Lett. | 1 |
| 2023 | Selector-Enhancer: Learning Dynamic Selection of Local and Non-local Attention Operation for Speech EnhancementabstractAttention mechanisms, such as local and non-local attention, play a fundamental role in recent deep learning based speech enhancement (SE) systems. However, a natural speech contains many fast-changing and relatively briefly acoustic events, therefore, capturing the most informative speech features by indiscriminately using local and non-local attention is challenged. We observe that the noise type and speech feature vary within a sequence of speech and the local and non-local can respectively process different types of corrupted speech regions. To leverage this, we propose Selector-Enhancer, a dual-attention based convolution neural network (CNN) with a feature-filter that can dynamically select regions from low-resolution speech features and feed them to local or non-local attention operations. In particular, the proposed feature-filter is trained by using reinforcement learning (RL) with a developed difficulty-regulated reward that related to network performance, model complexity and “the difficulty of the SE task”. The results show that our method achieves comparable or superior performance to existing approaches. In particular, Selector-Enhancer is effective for real-world denoising, where the number and types of noise are varies on a single noisy mixture. Xinmeng Xu, Weiping Tu, Yuhong Yang 0001 |
AAAI | 1 |
| 2023 | Improving Acoustic Echo Cancellation by Mixing Speech Local and Global Features with TransformerabstractWe propose MiT-Net, a novel mix-transformer neural network with a pyramid encoder operating in the time domain, for the task of acoustic echo cancellation. The MiT-Net formulates acoustic echo cancellation as a supervised speech separation problem, in which near-end speech is separated from a single microphone recording and sent to the far end, and consists of two key components. First, we apply a pyramid encoder, which adopts the coarse-to-fine structure, to extract the latent correlations between double-end signals and to fuse them in a multiscale manner. Second, we propose a mix-transformer, a combination of local and global attention in a parallel way, to leverage local and global speech information for separation. Experimental results show that the proposed method outperforms recent AEC methods in terms of objective evaluation metrics. In addition, exploring the correlation between speech local and global features by using the mix-transformer significantly improves the system performance and shows more robustness than the conventional transformer. Xinmeng Xu, Weiping Tu, Yuhong Yang 0001, Li Xiao 0007 |
ICASSP | 2 |
| 2023 | Leveraging Sound Local and Global Features for Language-Queried Target Sound Extraction
Xinmeng Xu, Yuhong Yang 0001, Weiping Tu |
ICONIP (4) | 1 |
| 2023 | Exploring the Interactions Between Target Positive and Negative Information for Acoustic Echo Cancellation
Chang Han, Xinmeng Xu, Weiping Tu, Yuhong Yang 0001 |
INTERSPEECH | 2 |
| 2023 | PCNN: A Lightweight Parallel Conformer Neural Network for Efficient Monaural Speech Enhancement
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001 |
INTERSPEECH | 1 |
| 2023 | CQNV: A Combination of Coarsely Quantized Bitstream and Neural Vocoder for Low Rate Speech Coding
Youqiang Zheng, Li Xiao 0007, Weiping Tu, Yuhong Yang 0001, Xinmeng Xu |
INTERSPEECH | 5 |
| 2023 | CASE-Net: Integrating local and non-local attention operations for speech enhancement
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001 |
Speech Commun. | 1 |
| 2022 | Improving Dual-Microphone Speech Enhancement by Learning Cross-Channel Features with Multi-Head AttentionabstractHand-crafted spatial features, such as inter-channel intensity difference (IID) and inter-channel phase difference (IPD), play a fundamental role in recent deep learning based dual-microphone speech enhancement (DMSE) systems. However, learning the mutual relationship between artificially designed spatial and spectral features is hard in the end-to-end DMSE. In this work, a novel architecture for DMSE using a multi-head cross-attention based convolutional recurrent network (MHCA-CRN) is presented. The proposed MHCA-CRN model includes a channel-wise encoding structure for preserving intra-channel features and a multi-head cross-attention mechanism for fully exploiting cross-channel features. In addition, the proposed approach specifically formulates the decoder with an extra SNR estimator to estimate frame-level SNR under a multi-task learning framework, which is expected to avoid speech distortion led by end-to-end DMSE module. Finally, a spectral gain function is adopted to further suppress the unnatural residual noise. Experiment results demonstrated superior performance of the proposed model against several state-of-the-art models. Xinmeng Xu, Rongzhi Gu, Yuexian Zou |
ICASSP | 1 |
| 2022 | VSEGAN: Visual Speech Enhancement Generative Adversarial NetworkabstractSpeech enhancement is an essential task of improving speech quality in noise scenario. Several state-of-the-art approaches have introduced visual information for speech enhancement, since the visual aspect of speech is essentially unaffected by acoustic environment. This paper proposes a novel framework that involves visual information for speech enhancement, by incorporating a Generative Adversarial Network (GAN). In particular, the proposed visual speech enhancement GAN consists of two networks trained in adversarial manner, i) a generator that adopts multi-layer feature fusion convolution network to enhance input noisy speech, and ii) a discriminator that attempts to minimize the discrepancy between the distributions of the clean speech signal and enhanced speech signal. Experiment results demonstrated superior performance of the proposed model against several state-of-the-art models. Xinmeng Xu, Dongxiang Xu, Yiyuan Peng, Jie Jia 0003, Binbin Chen 0006 |
ICASSP | 1 |
| 2022 | U-Former: Improving Monaural Speech Enhancement with Multi-head Self and Cross AttentionabstractFor supervised speech enhancement, contextual information is important for accurate spectral mapping. However, commonly used deep neural networks (DNNs) are limited in capturing temporal contexts. To leverage long-term contexts for tracking a target speaker, this paper treats the speech enhancement as sequence-to-sequence mapping, and propose a novel monaural speech enhancement U-net structure based on Transformer, dubbed U-Former. The key idea is to model long-term correlations and dependencies, which are crucial for accurate noisy speech modeling, through the multi-head attention mechanisms. For this purpose, U-Former incorporates multi-head attention mechanisms at two levels: 1) a multi-head self-attention module which calculate the attention map along both time-and frequency-axis to generate time and frequency sub-attention maps for leveraging global interactions between encoder features, while 2) multi-head cross-attention module which are inserted in the skip connections allows a fine recovery in the decoder by filtering out uncorrelated features. Experimental results illustrate that the U-Former obtains consistently better performance than recent models of PESQ, STOI, and SSNR scores. Xinmeng Xu, Jianjun Hao |
ICPR | 1 |
| 2022 | GLD-Net: Improving Monaural Speech Enhancement by Learning Global and Local Dependency Features with GLD BlockabstractFor monaural speech enhancement, contextual information is important for accurate speech estimation.However, commonly used convolution neural networks (CNNs) are weak in capturing temporal contexts since they only build blocks that process one local neighborhood at a time.To address this problem, we learn from human auditory perception to introduce a twostage trainable reasoning mechanism, referred as global-local dependency (GLD) block.GLD blocks capture long-term dependency of time-frequency bins both in global level and local level from the noisy spectrogram to help detecting correlations among speech part, noise part, and whole noisy input.What is more, we conduct a monaural speech enhancement network called GLD-Net, which adopts encoder-decoder architecture and consists of speech object branch, interference branch, and global noisy branch.The extracted speech feature at globallevel and local-level are efficiently reasoned and aggregated in each of the branches.We compare the proposed GLD-Net with existing state-of-art methods on WSJ0 and DEMAND dataset.The results show that GLD-Net outperforms the state-of-the-art methods in terms of PESQ and STOI. Xinmeng Xu, Jie Jia 0003, Binbin Chen 0006, Jianjun Hao |
INTERSPEECH | 1 |
| 2022 | Improving Visual Speech Enhancement Network by Learning Audio-visual Affinity with Multi-head AttentionabstractAudio-visual speech enhancement system is regarded as one of promising solutions for isolating and enhancing speech of desired speaker.Typical methods focus on predicting clean speech spectrum via a naive convolution neural network based encoderdecoder architecture, and these methods a) are not adequate to use data fully, b) are unable to effectively balance audio-visual features.The proposed model alleviates these drawbacks by a) applying a model that fuses audio and visual features layer by layer in encoding phase, and that feeds fused audio-visual features to each corresponding decoder layer, and more importantly, b) introducing a 2-stage multi-head cross attention (MHCA) mechanism to infer audio-visual speech enhancement for balancing the fused audio-visual features and eliminating irrelevant features.This paper proposes attentional audio-visual multi-layer feature fusion model, in which MHCA units are applied to feature mapping at every layer of decoder.The proposed model demonstrates the superior performance of the network against the state-of-the-art models.Speech samples are available at: https://XinmengXu.github.io/AVSE/AVCRN.html Xinmeng Xu, Jie Jia 0003, Binbin Chen 0006 |
INTERSPEECH | 1 |
| 2021 | Multi-Stage Progressive Speech Enhancement Network
Xinmeng Xu, Dongxiang Xu, Yiyuan Peng, Jie Jia 0003, Binbin Chen 0006 |
Interspeech | 1 |