EDBT 2026 Demo / reviewers in the wild / expert
Shiyun Xu
dblp:03/8131
· DBLP profile ↗
17ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MS-TTBA: Mid-side tokenized text-to-binaural audio with large language models
Changjun He, Lianyu Zhou, Shiyun Xu, Weiping Chen, Mingjiang Wang |
Neurocomputing | 4 |
| 2026 | MS-VBRVQ: Multi-scale variable bitrate speech residual vector quantization
Yukun Qian, Shiyun Xu, Xuyi Zhuang, Mingjiang Wang |
Speech Commun. | 2 |
| 2026 | ADS-BiMamba: Attentive Dynamic-Split Bidirectional Mamba for Multi-Channel Speech Enhancement
Shiyun Xu, Yinghan Cao, Changjun He, Mingjiang Wang |
IEEE Signal Process. Lett. | 1 |
| 2026 | CARVE: Content-Adaptive Rate-Variable Encoding for Neural Speech CodecsabstractNeural speech codecs that convert continuous wave forms into discrete tokens are an essential component for building compact speech representations in downstream applications. However, most existing codecs require both fixed and high frame rates to maintain reconstruction quality, which reduces the efficiency of downstream applications. To address this issue, we propose CARVE, a Content-Adaptive Rate-Variable Encoding strategy that performs compression at an externally specified frame rate within the latent space of a neural speech codec. CARVE comprises three modules: (i) Unsupervised Density Aware Scoring Module that combines inter-frame similarity with redundancy-aware cues to quantify the relative importance of each frame; (ii) Sparse Frame Segmentation Module that allocates finer segmentation to high-density information regions and coarser segmentation to low-density information regions based on these scores; and (iii) Context-Aware Frame Reconstruction Module that maps segment-level context back to frame-level latent representations. Integrating CARVE into the neural speech codec, extensive experiments show that it can perform compression at an externally specified frame rate while maintaining high reconstruction quality, thereby validating its effectiveness and practicality in low-frame-rate neural speech coding scenarios. Audio samples are available at: https://ethuil.github.io/CARVEdemo/ Yukun Qian, Yinghan Cao, Changjun He, Shiyun Xu, Mingjiang Wang |
IEEE Signal Process. Lett. | 5 |
| 2025 | Hybrid Feature Global Attention Network for Noisy-reverberant Speech EnhancementabstractDeep neural network-based speech enhancement methods have become widespread, with one of its fundamental aspects being the effective extraction and application of features in the time-frequency domain. This paper proposes a hybrid feature global attention network (HFGANet) designed to efficiently extract time-frequency domain features. HFGANet incorporates a hybrid gated multilayer perceptron (HgMLP) that effectively captures local, global, and inter-window features in the time-frequency domain to create hybrid representations. In contrast to traditional convolutional recurrent neural network architectures, this paper innovatively proposes a global attention structure to leverage these hybrid features. The proposed global attention block enhances the integration of local and global features. Additionally, we introduce Temporal Mamba and Frequency Mamba to further improve the model's ability to capture contextual information in both time and frequency dimensions. On the 1st Deep Noise Suppression Challenge blind test set with reverberation, HFGANet achieves 3.51 WB-PESQ, 95.03% STOI, and 17.72 SI-SDR, while maintaining a lower parameter count compared to state-of-the-art models. In the task of noisy-reverberant speech enhancement, our model achieved an improvement of 1.22 in PESQ, 17.4% in STOI, and 1.63 in DNSMOS. Shiyun Xu, Yinghan Cao, Changjun He |
ICASSP | 2 |
| 2025 | Joint Training Framework for Accent and Speech Recognition Based on Conformer Low-Rank AdaptationabstractIn real-world scenarios, accent variations often reduce Automatic Speech Recognition (ASR) accuracy. Addressing this typically involves a multi-task ASR and Accent Recognition (ASR-AR) framework, but there is limited research on optimizing task-specific feature extraction and enhancing ASR with AR information. This study introduces the Conformer Low-rank Adaptation for Joint Accent and Speech Recognition (CLAnSR), employing LoRA to augment both ASR and AR capabilities using a shared pre-trained base encoder. This approach significantly reduces the model’s parameter and training resource demands while facilitating the extraction of task-specific features. Additionally, we have incorporated accent-aware multi-channel embedding layers, which through spatially independent embeddings, enhance the model’s capacity to accurately represent tokens across diverse dialectical contexts. Tested on the KeSpeech dataset, CLAnSR reaches state-of-the-art AR accuracy 80.41% and competitive ASR CER 8.39%, outperforming non-LLM systems and matching those with LLMs. It reduces parameters by 34.02% and enhances both ASR and AR performance, effectively handling speech dialect variations and advancing the field. Xuyi Zhuang, Yukun Qian, Shiyun Xu, Mingjiang Wang |
ICASSP | 3 |
| 2025 | Gradient descent with generalized Newton's methodabstractWe propose the generalized Newton's method (GeN) --- a Hessian-informed approach that applies to any optimizer such as SGD and Adam, and covers the Newton-Raphson method as a sub-case. Our method automatically and dynamically selects the learning rate that accelerates the convergence, without the intensive tuning of the learning rate scheduler. In practice, our method is easily implementable, since it only requires additional forward passes with almost zero computational overhead (in terms of training time and memory cost), if the overhead is amortized over many iterations. We present extensive experiments on language and vision tasks (e.g. GPT and ResNet) to showcase that GeN optimizers match the state-of-the-art performance, which was achieved with carefully tuned learning rate schedulers. Zhiqi Bu, Shiyun Xu |
ICLR | 2 |
| 2025 | SF-AN: A lightweight shuffle Fourier attention network for multi-channel speech enhancement
Shiyun Xu, Yinghan Cao, Yukun Qian, Changjun He, Mingjiang Wang |
Speech Commun. | 1 |
| 2025 | Two-stage UNet with channel and temporal-frequency attention for multi-channel speech enhancement
Shiyun Xu, Yinghan Cao, Mingjiang Wang |
Speech Commun. | 1 |
| 2025 | FSTF-AN: Fused Sparse Temporal-Frequency Attentive Network for Multi-Channel Speech EnhancementabstractThe Transformer has achieved impressive performance in the multi-channel speech enhancement field; however, it struggles to capture local features, which leads to the loss of speech details. To enhance the extraction of local features in the network, we propose a fused sparse temporal-frequency attentive network (FSTF-AN), which aims to fully capture features across the temporal-frequency, frequency, and temporal dimensions. We propose top-$k$fused sparse self-attention, which employs a fusion strategy to adaptively retain the most crucial attention scores when computing self-attention maps, thereby eliminating irrelevant information interference and better aggregating features. Furthermore, we propose a multi-scale fused feed-forward network, which effectively captures multi-scale features, further enhancing the network's ability to capture local features. The experimental results demonstrate that FSTF-AN exhibits significant advantages over other SOTA models, effectively enhancing speech quality and intelligibility. Shiyun Xu, Yinghan Cao, Mingjiang Wang |
IEEE Signal Process. Lett. | 1 |
| 2023 | Half-Temporal and Half-Frequency Attention U2Net for Speech Signal ImprovementabstractDuring communication, volume changes, noise, and reverberation can disturb speech signals, significantly affecting the quality and intelligibility of speech. In the context of the ICASSP 2023 Signal Processing Grand Challenge, the first Speech Signal Improvement Grand Challenge (SIG) is organized to improve the quality of speech signals during communication. This paper proposes half-temporal and half-frequency attention U2Net for improving full-band speech signal. Channel-spectrum attention is proposed for the skip connection between the encoder and decoder. The proposed model achieves 0.353, 1.289, 0.604, 0.625, and 0.924 improvements in signal, noise, overall, reverberation, and loudness, respectively, in the SIG subjective test. The proposed model achieved fourth place in the SIG real-time track, showing excellent denoising and de-reverberation performance. Shiyun Xu, Xuyi Zhuang, Yukun Qian, Lianyu Zhou, Mingjiang Wang |
ICASSP | 2 |
| 2023 | Two-Stage UNet with Multi-Axis Gated Multilayer Perceptron for Monaural Noisy-Reverberant Speech EnhancementabstractIn denoising and de-reverberation tasks, the dominant methods are complex spectral masking and complex spectral mapping. To combine advantages and improve speech enhancement performance, we propose a two-stage UNet (TSUNet) to estimate complex spectral masking and complex spectral mapping. We use a multi-axis gated multilayer perceptron to build global and local attention modules of linear complexity for extracting speech features. Furthermore, we use the residual channel attention block to further filter out important speech features. On the blind test dataset of the Deep Noise Suppression Challenge, our proposed TSUNet has a massive advantage over other state-of-the-art models. TSUNet performs significantly better than the most recent models at noisy-reverberant speech enhancement. Shiyun Xu, Xuyi Zhuang, Lianyu Zhou, Heng Li 0013, Mingjiang Wang |
ICASSP | 2 |
| 2023 | Sparse Neural Additive Model: Interpretable Deep Learning with Feature Selection via Group Sparsity
Shiyun Xu, Zhiqi Bu, Pratik Chaudhari, Ian J. Barnett |
ECML/PKDD (3) | 1 |
| 2022 | Scalable and Efficient Training of Large Convolutional Neural Networks with Differential PrivacyabstractLarge convolutional neural networks (CNN) can be difficult to train in the differentially private (DP) regime, since the optimization algorithms require a computationally expensive operation, known as the per-sample gradient clipping. We propose an efficient and scalable implementation of this clipping on convolutional layers, termed as the mixed ghost clipping, that significantly eases the private training in terms of both time and space complexities, without affecting the accuracy. The improvement in efficiency is rigorously studied through the first complexity analysis for the mixed ghost clipping and existing DP training algorithms.Extensive experiments on vision classification tasks, with large ResNet, VGG, and Vision Transformers (ViT), demonstrate that DP training with mixed ghost clipping adds $1\sim 10\%$ memory overhead and $<2\times$ slowdown to the standard non-private training. Specifically, when training VGG19 on CIFAR10, the mixed ghost clipping is $3\times$ faster than state-of-the-art Opacus library with $18\times$ larger maximum batch size. To emphasize the significance of efficient DP training on convolutional layers, we achieve 96.7\% accuracy on CIFAR10 and 83.0\% on CIFAR100 at $\epsilon=1$ using BEiT, while the previous best results are 94.8\% and 67.4\%, respectively. We open-source a privacy engine (\url{https://github.com/woodyx218/private_vision}) that implements DP training of CNN (including convolutional ViT) with a few lines of code. Zhiqi Bu, Jialin Mao, Shiyun Xu |
NeurIPS | 3 |
| 2021 | A Dynamical View on Optimization Algorithms of Overparameterized Neural NetworksabstractWhen equipped with efficient optimization algorithms, the over-parameterized neural networks have demonstrated high level of performance even though the loss function is non-convex and non-smooth. While many works have been focusing on understanding the loss dynamics by training neural networks with the gradient descent (GD), in this work, we consider a broad class of optimization algorithms that are commonly used in practice. For example, we show from a dynamical system perspective that the Heavy Ball (HB) method can converge to global minimum on mean squared error (MSE) at a linear rate (similar to GD); however, the Nesterov accelerated gradient descent (NAG) may only converge to global minimum sublinearly. Our results rely on the connection between neural tangent kernel (NTK) and finitely-wide over-parameterized neural networks with ReLU activation, which leads to analyzing the limiting ordinary differential equations (ODE) for optimization algorithms. We show that, optimizing the non-convex loss over the weights corresponds to optimizing some strongly convex loss over the prediction error. As a consequence, we can leverage the classical convex optimization theory to understand the convergence behavior of neural networks. We believe our approach can also be extended to other optimization algorithms and network architectures. Zhiqi Bu, Shiyun Xu |
AISTATS | 2 |
| 2021 | DebiNet: Debiasing Linear Models with Nonlinear Overparameterized Neural NetworksabstractRecent years have witnessed strong empirical performance of over-parameterized neural networks on various tasks and many advances in the theory, e.g. the universal approximation and provable convergence to global minimum. In this paper, we incorporate over-parameterized neural networks into semi-parametric models to bridge the gap between inference and prediction, especially in the high dimensional linear problem. By doing so, we can exploit a wide class of networks to approximate the nuisance functions and to estimate the parameters of interest consistently. Therefore, we may offer the best of two worlds: the universal approximation ability from neural networks and the interpretability from classic ordinary linear model, leading to valid inference and accurate prediction. We show the theoretical foundations that make this possible and demonstrate with numerical experiments. Furthermore, we propose a framework, DebiNet, in which we plug-in arbitrary feature selection methods to our semi-parametric neural network and illustrate that our framework debiases the regularized estimators and performs well, in terms of the post-selection inference and the generalization error. Shiyun Xu, Zhiqi Bu |
AISTATS | 1 |
| 2021 | Asymptotic Statistical Analysis of Sparse Group LASSO via Approximate Message Passing
Zhiqi Bu, Shiyun Xu |
ECML/PKDD (3) | 3 |