EDBT 2026 Demo / reviewers in the wild / expert
Xianjun Xia
dblp:205/4003
· DBLP profile ↗
27ranked-venue papers
9as first author
17since 2021 · last 2025
0000-0001-5277-6634ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 17 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FAF-Filt: Frequency-aware Fourier Filter for Sound Event DetectionabstractCapturing time-frequency patterns along the frequency axis, is crucial for the precision of sound event detection systems. Frequency dynamic convolution (FDY) and a series of its variants which incorporate frequency-adaptive kernels in standard 2D convolutions, have demonstrated remarkable performance, yet also suffered from high computational costs. To address the issue, we propose an efficient and light-weighted frequency-aware Fourier filter (FAF-Filt), which performs a 2D Fourier transform on features to the frequency domain and employs a learnable frequency-aware filter to process the transformed features, thereby integrating global information more effectively to extract decisive frequency components. In addition, frequency-adaptive convolution (FA-Conv) is adopted to further strengthen the representative ability of convolution, which incorporates the frequency-aware attention mechanism into the inputs and outputs of the convolutions. Experimental results exhibit superiority of the proposed method, achieving comparable performance with FDY-CRNN in terms of polyphonic sound event scores (PSDS) with a significantly 56% reduction in parameters. Xiaohuai Le, Zhuangqi Chen, Xianjun Xia, Chuanzeng Huang |
ICASSP | 4 |
| 2025 | AF-Vocoder: Artifact-Free Neural Vocoder with Global Artifact Filter
Zhuangqi Chen, Xianjun Xia, Xiaohuai Le, Chuanzeng Huang |
INTERSPEECH | 2 |
| 2025 | Multistage Universal Speech Enhancement System for URGENT Challenge
Xiaohuai Le, Zhuangqi Chen, Xianjun Xia, Chuanzeng Huang |
INTERSPEECH | 4 |
| 2025 | CBA-Whisper: Curriculum Learning-Based AdaLoRA Fine-Tuning on Whisper for Low-Resource Dysarthric Speech Recognition
Tianyi Tan, Xiaohuai Le, Wenzhi Fan, Xianjun Xia, Chuanzeng Huang |
INTERSPEECH | 5 |
| 2025 | U-SAM: An Audio Language Model for Unified Speech, Audio, and Music Understanding
Xianjun Xia, Xinfa Zhu, Lei Xie 0001 |
INTERSPEECH | 2 |
| 2024 | RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
Mingshuai Liu, Zhuangqi Chen, Xiaopeng Yan, Yuanjun Lv, Xianjun Xia, Chuanzeng Huang, Yijian Xiao, Lei Xie 0001 |
INTERSPEECH | 5 |
| 2024 | BS-PLCNet 2: Two-stage Band-split Packet Loss Concealment Network with Intra-model Knowledge Distillation
Xianjun Xia, Chuanzeng Huang, Yijian Xiao, Lei Xie 0001 |
INTERSPEECH | 2 |
| 2023 | An Exploration of Task-Decoupling on Two-Stage Neural Post Filter for Real-Time Personalized Acoustic Echo CancellationabstractDeep learning based techniques have been popularly adopted in acoustic echo cancellation (AEC). Utilization of speaker representation has extended the frontier of AEC, thus attracting many researchers’ interest in personalized acoustic echo cancellation (PAEC). Meanwhile, task-decoupling strategies are widely adopted in speech enhancement. To further explore the task-decoupling approach, we propose to use a two-stage task-decoupling post-filter (TDPF) in PAEC. Furthermore, a multi-scale local-global speaker representation is applied to improve speaker extraction in PAEC. Experimental results indicate that the task-decoupling model can yield better performance than a single joint network. The optimal approach is to decouple the echo cancellation from noise and interference speech suppression. Based on the task-decoupling sequence, optimal training strategies for the two-stage model are explored afterwards. Jiayao Sun, Xianjun Xia, Xiaopeng Yan, Yijian Xiao, Lei Xie 0001 |
ASRU | 3 |
| 2023 | A Progressive Neural Network for Acoustic Echo CancellationabstractAcoustic echo cancellation is a key issue in hand-free communication systems. In this paper, we proposed a hybrid signal processing and deep echo cancellation method, where a two-stage neural network is designed to remove residual echo progressively. For the personalized acoustic echo cancellation, we proposed to decouple the tasks of echo cancellation and target speech extraction, and introduced a speaker attentive module for personalized separation, where the ECAPA-TDNN is used for speaker embedding generation. The proposed method (ByteAudio-18) ranked first on both Track 1 and Track 2 in ICASSP 2023 AEC Challenge. Zhuangqi Chen, Xianjun Xia, Guoliang Xie, Pingjian Zhang, Yijian Xiao |
ICASSP | 2 |
| 2023 | The Ajmide Topic Segmentation System for the ICASSP 2023 General Meeting Understanding and Generation ChallengeabstractThis paper describes our topic segmentation (TS) system submitted to the ICASSP2023 Signal Processing Grand Challenge - General Meeting Understanding and Generation challenge (MUG). We make three improvements to the official baseline system of the TS track. Firstly, considering that meeting transcriptions are usually long-form documents, we propose a PoNet-Svec-Transformers network to learn both sentence representations and document-level context. Secondly, we introduce a training data synthesis method that significantly increase the size of the training dataset. Finally, we leverage focal loss and adversarial training methods to improve system performance. Our best submission achieves a first-place score of 48.56/48.84 on the Eval/Test set in the TS task. Beibei Hu, Xianjun Xia |
ICASSP | 3 |
| 2023 | Personalized Speech Enhancement Combining Band-Split RNN and Speaker Attentive ModuleabstractTarget speaker information can be utilized in speech enhancement (SE) models to more effectively extract the desired speech. Previous works introduce the speaker embedding into speech enhancement models by means of concatenation or affine transformation. In this paper, we propose a speaker attentive module to calculate the attention scores between the speaker embedding and the intermediate features, which are used to rescale the features. By merging this module in the state-of-the-art SE model, we construct the personalized SE model for ICASSP Signal Processing Grand Challenge: DNS Challenge 5 (2023). Our system achieves a final score of 0.529 on the blind test set of track1 and 0.549 on track2. Xiaohuai Le, Yiqing Guo, Xianjun Xia |
ICASSP | 6 |
| 2023 | Two-Stage Neural Network for ICASSP 2023 Speech Signal Improvement ChallengeabstractIn ICASSP 2023 speech signal improvement challenge, we developed a dual-stage neural model which improves speech signal quality induced by different distortions in a stage-wise divide-and-conquer fashion. Specifically, in the first stage, the speech improvement network focuses on recovering the missing components of the spectrum, while in the second stage, our model aims to further suppress noise, reverberation, and artifacts introduced by the first-stage model. Achieving 0.446 in the final score and 0.517 in the P.835 score, our system ranks 4th in the non-real-time track. Mingshuai Liu, Shubo Lv, Runduo Han, Xianjun Xia, Yijian Xiao, Lei Xie 0001 |
ICASSP | 6 |
| 2023 | A Lightweight Fourier Convolutional Attention Encoder for Multi-Channel Speech EnhancementabstractBeamforming weights prediction via deep neural networks has been one of the main methods in multi-channel speech enhancement tasks. The spectral-spatial cues are crucial in beamforming weights estimation, however, many existing works fail to optimally predict the beamforming weights with an absence of adequate spectral-spatial information learning. To tackle this challenge, we propose a Fourier convolutional attention encoder (FCAE) to provide a global receptive field over the frequency axis and boost the learning of spectral contexts and cross-channel features. Besides, a new convolutional recurrent encoder-decoder (CRED) structure is proposed in this work, within which FCAEs, attention blocks with skip connections and a deep feedback sequential memory network (DFSMN) serving as recurrent module are involved. The proposed CRED structure is exploited to capture the spectral-spatial joint information to obtain accurate estimation of beamforming weights. Experimental results demonstrate the superiority of the proposed approach with only 0.74M parameters and a PESQ improvement from 2.225 to 2.359 on the ConferencingSpeech2021 challenge development test set. Xianjun Xia, Yijian Xiao, Piao Ding, Shenyi Song, Roberto Togneri |
ICASSP | 4 |
| 2023 | A Two-stage Progressive Neural Network for Acoustic Echo CancellationabstractRecent studies in deep learning based acoustic echo cancellation proves the benefits of introducing a linear echo cancellation module. However, the convergence problem and potential target speech distortion impose an additional learning burden for the neural network. In this paper, we propose a two-stage progressive neural network consisting of a coarse-stage and a fine-stage module. For the coarse-stage, a light-weighted network module is designed to suppress partial echo and potential noise, where a voice activity detection path is used to enhance the learned features. For the fine-stage, a larger network is employed to deal with the more complex echo path and restore the near-end speech. We have conducted extensive experiments to verify the proposed method, and the results show that the proposed two-stage method provides a superior performance to other state-of-the-art methods. Zhuangqi Chen, Xianjun Xia, Xianke Wang, Yanhong Leng, Roberto Togneri, Yijian Xiao, Piao Ding, Shenyi Song, Pingjian Zhang |
INTERSPEECH | 2 |
| 2023 | Harmonic enhancement using learnable comb filter for light-weight full-band speech enhancement model
Xiaohuai Le, Yiqing Guo, Xianjun Xia, Hua Gao, Yijian Xiao, Piao Ding, Shenyi Song |
INTERSPEECH | 7 |
| 2023 | An Intra-BRNN and GB-RVQ Based END-TO-END Neural Audio Codec
Linping Xu, Dejun Zhang, Xianjun Xia, Yijian Xiao, Piao Ding, Shenyi Song, Sixing Yin, Ferdous Sohel |
INTERSPEECH | 4 |
| 2021 | A Two-Stage Approach to Device-Robust Acoustic Scene ClassificationabstractTo improve device robustness, a highly desirable key feature of a competitive data-driven acoustic scene classification (ASC) system, a novel two-stage system based on fully convolutional neural networks (CNNs) is proposed. Our two-stage system leverages on an ad-hoc score combination based on two CNN classifiers: (i) the first CNN classifies acoustic inputs into one of three broad classes, and (ii) the second CNN classifies the same inputs into one of ten finergrained classes. Three different CNN architectures are explored to implement the two-stage classifiers, and a frequency sub-sampling scheme is investigated. Moreover, novel data augmentation schemes for ASC are also investigated. Evaluated on DCASE 2020 Task 1a, our results show that the proposed ASC system attains a state-of-the-art accuracy on the development set, where our best system, a two-stage fusion of CNN ensembles, delivers a 81.9% average accuracy among multi-device test data, and it obtains a significant improvement on unseen devices. Finally, neural saliency analysis with class activation mapping (CAM) gives new insights on the patterns learnt by our models. Hu Hu, Chao-Han Huck Yang, Xianjun Xia, Yajian Wang, Shutong Niu, Li Chai 0002, Juanjuan Li, Hongning Zhu, Sabato Marco Siniscalchi, Yannan Wang, Jun Du 0002, Chin-Hui Lee 0001 |
ICASSP | 3 |
| 2020 | Audio Sound Determination Using Feature Space Attention Based Convolution Recurrent Neural NetworkabstractThe classification framework has been popularly adopted to perform sound event detection. However, the existing neural network based classification based approaches treat each feature dimension equally and the varying influence of feature dimensions has not been taken into consideration. To deal with this, we propose a feature space attention based convolution recurrent neural network approach utilizing the varying importance of each feature dimension to perform acoustic event detection. The convolution layers are used to extract the high level information from the audio signals. Then the feature space attention scheme is applied to the extracted features to automatically determine the importance of each feature dimension. Experimental results on the latest TUT Sound Event 2017 dataset demonstrate the improved performance of the proposed approach compared to the existing acoustic event detection systems. Xianjun Xia, Jingjing Pan, Yannan Wang |
ICASSP | 1 |
| 2020 | Sound Event Detection Using Multiple Optimized KernelsabstractSound event detection (SED) has been widely applied in real world applications. Convolutional recurrent neural network based SED approaches have achieved state-of-the-art performance. However, the convolution process is typically performed by using a fixed sized kernel, which adversely affects the detection accuracy especially when the acoustic features of different event classes are characterized by high variations. To deal with this, this article proposes a sound event detection technique using a convolutional recurrent neural network framework with multiple convolutional kernels of different sizes. The top performing kernels are selected from a kernel pool based on the unsupervised clustering errors and the accuracies of the temporarily trained models. Afterwards, the selected kernels are fed to multiple convolution layers to deal with the acoustic feature variations. Experimental results on different subsets of AudioSet, namely the DCASE Challenge 2017 Task 4 and DCASE Challenge 2018 Task 4, demonstrate the performance of the proposed approach compared to state-of-the-art systems. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Multi-Task Learning for Acoustic Event Detection Using Event and Frame Position InformationabstractAcoustic event detection deals with the acoustic signals to determine the sound type and to estimate the audio event boundaries. Multi-label classification based approaches are commonly used to detect the frame wise event types with a median filter applied to determine the happening acoustic events. However, the multi-label classifiers are trained only on the acoustic event types ignoring the frame position within the audio events. To deal with this, this paper proposes to construct a joint learning based multi-task system. The first task performs the acoustic event type detection and the second task is to predict the frame position information. By sharing representations between the two tasks, we can enable the acoustic models to generalize better than the original classifier by averaging respective noise patterns to be implicitly regularized. Experimental results on the monophonic UPC-TALP and the polyphonic TUT Sound Event datasets demonstrate the superior performance of the joint learning method by achieving lower error rate and higher F-score compared to the baseline AED system. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
IEEE Trans. Multim. | 1 |
| 2019 | Auxiliary Classifier Generative Adversarial Network With Soft Labels in Imbalanced Acoustic Event DetectionabstractIn acoustic event detection, the training data size of some acoustic events is often small and imbalanced. To deal with this, this paper proposes generating the virtual training data categorically using the auxiliary classifier generative adversarial networks. Soft labels of acoustic events are first calculated to represent the acoustic event localization information. The closer the current frame is to the middle of the manually labeled acoustic event, the higher the soft label will be, which makes the soft labels positively correlated with the acoustic event localization. Then, the acoustic event class and the quantized soft labels are used as the input condition to the auxiliary classifier generative adversarial networks to generate an arbitrary number of training samples. Experimental results on the TUT Sound Event 2016 under the home environment and TUT Sound Event 2017 under the street environment demonstrate the improved performance of the proposed technique compared to existing acoustic event detection systems. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
IEEE Trans. Multim. | 1 |
| 2018 | Confidence Based Acoustic Event DetectionabstractAcoustic event detection, the determination of the acoustic event type and the localisation of the event, has been widely applied in many real-world applications. Many works adopt the multi-label classification technique to perform the polyphonic acoustic event detection with a global threshold to detect the active acoustic events. However, the manually labeled boundaries are error-prone and cannot always be accurate, especially when the frame length is too short to be accurately labeled by human annotators. To deal with this, a confidence is assigned to each frame and acoustic event detection is performed using a multi-variable regression approach in this paper. Experimental results on the latest TUT sound event 2017 database of polyphonic events demonstrate the superior performance of the proposed approach compared to the multi-label classification based AED method. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
ICASSP | 1 |
| 2018 | Local Binary Pattern with Random Forest for Acoustic Scene ClassificationabstractThis paper presents an approach for acoustic scene classification using the local binary pattern (LBP) and random forest (RF). The audio signal is converted to a Constant-Q transform (CQT) representation and LBP is used to extract the features from this time-frequency representation. The CQT representations are divided into a number of sub-bands to obtain more localized features relevant to the spectral information. We then use random forest to select the most important features for each band of extracted LBP features. For further performance enhancement, we use feature level fusion of LBP and HOG features. The proposed system has achieved an accuracy of 85% on the DCASE 2016 dataset. Shamsiah Abidin, Xianjun Xia, Roberto Togneri, Ferdous Sohel |
ICME | 2 |
| 2018 | Random forest classification based acoustic event detection utilizing contextual-information and bottleneck features
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
Pattern Recognit. | 1 |
| 2017 | Random forest regression based acoustic event detection with bottleneck featuresabstractThis paper deals with random forest regression based acoustic event detection (AED) by combining acoustic features with bottleneck features (BN). The bottleneck features have a good reputation of being inherently discriminative in acoustic signal processing. To deal with the unstructured and complex real-world acoustic events, an acoustic event detection system is constructed using bottleneck features combined with acoustic features. Evaluations were carried out on the UPC-TALP and ITC-Irst databases which consist of highly variable acoustic events. Experimental results demonstrate the usefulness of the low-dimensional and discriminative bottleneck features with relative 5.33% and 5.51% decreases in error rates respectively. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
ICME | 1 |
| 2017 | Random forest classification based acoustic event detectionabstractThis paper deals with the acoustic event detection (AED) to improve the detection accuracy of acoustic events. Acoustic event detection task is performed by a regression via classification (RvC) based approach along with the random forest technique. A discretization process is used to convert the continuous frame positions within acoustic events into event duration class labels. Outputs of the category-specific random forest classifiers are then reversed back to the event boundary information. Evaluations on the UPC-TALP database which consists of highly variable acoustic events demonstrate the efficiency of the proposed approaches with improvements in detection error rate compared to the best baseline system. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
ICME | 1 |
| 2017 | Frame-Wise Dynamic Threshold Based Polyphonic Acoustic Event DetectionabstractAcoustic event detection, the determination of the acoustic event type and the localisation of the event, has been widely applied in many real-world applications. Many works adopt multi-label classification techniques to perform the polyphonic acoustic event detection with a global threshold to detect the active acoustic events. However, the global threshold has to be set manually and is highly dependent on the database being tested. To deal with this, we replaced the fixed threshold method with a frame-wise dynamic threshold approach in this paper. Two novel approaches, namely contour and regressor based dynamic threshold approaches are proposed in this work. Experimental results on the popular TUT Acoustic Scenes 2016 database of polyphonic events demonstrated the superior performance of the proposed approaches. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
INTERSPEECH | 1 |