Xianjun Xia

dblp:205/4003 · DBLP profile ↗
← Back
27ranked-venue papers
9as first author
17since 2021 · last 2025
0000-0001-5277-6634ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 17 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 10 since 2021
YearPublicationVenuePosition
2025 FAF-Filt: Frequency-aware Fourier Filter for Sound Event Detection
abstract
Capturing time-frequency patterns along the frequency axis, is crucial for the precision of sound event detection systems. Frequency dynamic convolution (FDY) and a series of its variants which incorporate frequency-adaptive kernels in standard 2D convolutions, have demonstrated remarkable performance, yet also suffered from high computational costs. To address the issue, we propose an efficient and light-weighted frequency-aware Fourier filter (FAF-Filt), which performs a 2D Fourier transform on features to the frequency domain and employs a learnable frequency-aware filter to process the transformed features, thereby integrating global information more effectively to extract decisive frequency components. In addition, frequency-adaptive convolution (FA-Conv) is adopted to further strengthen the representative ability of convolution, which incorporates the frequency-aware attention mechanism into the inputs and outputs of the convolutions. Experimental results exhibit superiority of the proposed method, achieving comparable performance with FDY-CRNN in terms of polyphonic sound event scores (PSDS) with a significantly 56% reduction in parameters.
Xiaohuai Le, Zhuangqi Chen, Xianjun Xia, Chuanzeng Huang
ICASSP4
2025 AF-Vocoder: Artifact-Free Neural Vocoder with Global Artifact Filter
Zhuangqi Chen, Xianjun Xia, Xiaohuai Le, Chuanzeng Huang
INTERSPEECH2
2025 Multistage Universal Speech Enhancement System for URGENT Challenge
Xiaohuai Le, Zhuangqi Chen, Xianjun Xia, Chuanzeng Huang
INTERSPEECH4
2025 CBA-Whisper: Curriculum Learning-Based AdaLoRA Fine-Tuning on Whisper for Low-Resource Dysarthric Speech Recognition
Tianyi Tan, Xiaohuai Le, Wenzhi Fan, Xianjun Xia, Chuanzeng Huang
INTERSPEECH5
2025 U-SAM: An Audio Language Model for Unified Speech, Audio, and Music Understanding
Xianjun Xia, Xinfa Zhu, Lei Xie 0001
INTERSPEECH2
2024 RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
Mingshuai Liu, Zhuangqi Chen, Xiaopeng Yan, Yuanjun Lv, Xianjun Xia, Chuanzeng Huang, Yijian Xiao, Lei Xie 0001
INTERSPEECH5
2024 BS-PLCNet 2: Two-stage Band-split Packet Loss Concealment Network with Intra-model Knowledge Distillation
Xianjun Xia, Chuanzeng Huang, Yijian Xiao, Lei Xie 0001
INTERSPEECH2
2023 An Exploration of Task-Decoupling on Two-Stage Neural Post Filter for Real-Time Personalized Acoustic Echo Cancellation
abstract
Deep learning based techniques have been popularly adopted in acoustic echo cancellation (AEC). Utilization of speaker representation has extended the frontier of AEC, thus attracting many researchers’ interest in personalized acoustic echo cancellation (PAEC). Meanwhile, task-decoupling strategies are widely adopted in speech enhancement. To further explore the task-decoupling approach, we propose to use a two-stage task-decoupling post-filter (TDPF) in PAEC. Furthermore, a multi-scale local-global speaker representation is applied to improve speaker extraction in PAEC. Experimental results indicate that the task-decoupling model can yield better performance than a single joint network. The optimal approach is to decouple the echo cancellation from noise and interference speech suppression. Based on the task-decoupling sequence, optimal training strategies for the two-stage model are explored afterwards.
Jiayao Sun, Xianjun Xia, Xiaopeng Yan, Yijian Xiao, Lei Xie 0001
ASRU3
2023 A Progressive Neural Network for Acoustic Echo Cancellation
abstract
Acoustic echo cancellation is a key issue in hand-free communication systems. In this paper, we proposed a hybrid signal processing and deep echo cancellation method, where a two-stage neural network is designed to remove residual echo progressively. For the personalized acoustic echo cancellation, we proposed to decouple the tasks of echo cancellation and target speech extraction, and introduced a speaker attentive module for personalized separation, where the ECAPA-TDNN is used for speaker embedding generation. The proposed method (ByteAudio-18) ranked first on both Track 1 and Track 2 in ICASSP 2023 AEC Challenge.
Zhuangqi Chen, Xianjun Xia, Guoliang Xie, Pingjian Zhang, Yijian Xiao
ICASSP2
2023 The Ajmide Topic Segmentation System for the ICASSP 2023 General Meeting Understanding and Generation Challenge
abstract
This paper describes our topic segmentation (TS) system submitted to the ICASSP2023 Signal Processing Grand Challenge - General Meeting Understanding and Generation challenge (MUG). We make three improvements to the official baseline system of the TS track. Firstly, considering that meeting transcriptions are usually long-form documents, we propose a PoNet-Svec-Transformers network to learn both sentence representations and document-level context. Secondly, we introduce a training data synthesis method that significantly increase the size of the training dataset. Finally, we leverage focal loss and adversarial training methods to improve system performance. Our best submission achieves a first-place score of 48.56/48.84 on the Eval/Test set in the TS task.
Beibei Hu, Xianjun Xia
ICASSP3
2023 Personalized Speech Enhancement Combining Band-Split RNN and Speaker Attentive Module
abstract
Target speaker information can be utilized in speech enhancement (SE) models to more effectively extract the desired speech. Previous works introduce the speaker embedding into speech enhancement models by means of concatenation or affine transformation. In this paper, we propose a speaker attentive module to calculate the attention scores between the speaker embedding and the intermediate features, which are used to rescale the features. By merging this module in the state-of-the-art SE model, we construct the personalized SE model for ICASSP Signal Processing Grand Challenge: DNS Challenge 5 (2023). Our system achieves a final score of 0.529 on the blind test set of track1 and 0.549 on track2.
Xiaohuai Le, Yiqing Guo, Xianjun Xia
ICASSP6
2023 Two-Stage Neural Network for ICASSP 2023 Speech Signal Improvement Challenge
abstract
In ICASSP 2023 speech signal improvement challenge, we developed a dual-stage neural model which improves speech signal quality induced by different distortions in a stage-wise divide-and-conquer fashion. Specifically, in the first stage, the speech improvement network focuses on recovering the missing components of the spectrum, while in the second stage, our model aims to further suppress noise, reverberation, and artifacts introduced by the first-stage model. Achieving 0.446 in the final score and 0.517 in the P.835 score, our system ranks 4th in the non-real-time track.
Mingshuai Liu, Shubo Lv, Runduo Han, Xianjun Xia, Yijian Xiao, Lei Xie 0001
ICASSP6
2023 A Lightweight Fourier Convolutional Attention Encoder for Multi-Channel Speech Enhancement
abstract
Beamforming weights prediction via deep neural networks has been one of the main methods in multi-channel speech enhancement tasks. The spectral-spatial cues are crucial in beamforming weights estimation, however, many existing works fail to optimally predict the beamforming weights with an absence of adequate spectral-spatial information learning. To tackle this challenge, we propose a Fourier convolutional attention encoder (FCAE) to provide a global receptive field over the frequency axis and boost the learning of spectral contexts and cross-channel features. Besides, a new convolutional recurrent encoder-decoder (CRED) structure is proposed in this work, within which FCAEs, attention blocks with skip connections and a deep feedback sequential memory network (DFSMN) serving as recurrent module are involved. The proposed CRED structure is exploited to capture the spectral-spatial joint information to obtain accurate estimation of beamforming weights. Experimental results demonstrate the superiority of the proposed approach with only 0.74M parameters and a PESQ improvement from 2.225 to 2.359 on the ConferencingSpeech2021 challenge development test set.
Xianjun Xia, Yijian Xiao, Piao Ding, Shenyi Song, Roberto Togneri
ICASSP4
2023 A Two-stage Progressive Neural Network for Acoustic Echo Cancellation
abstract
Recent studies in deep learning based acoustic echo cancellation proves the benefits of introducing a linear echo cancellation module. However, the convergence problem and potential target speech distortion impose an additional learning burden for the neural network. In this paper, we propose a two-stage progressive neural network consisting of a coarse-stage and a fine-stage module. For the coarse-stage, a light-weighted network module is designed to suppress partial echo and potential noise, where a voice activity detection path is used to enhance the learned features. For the fine-stage, a larger network is employed to deal with the more complex echo path and restore the near-end speech. We have conducted extensive experiments to verify the proposed method, and the results show that the proposed two-stage method provides a superior performance to other state-of-the-art methods.
Zhuangqi Chen, Xianjun Xia, Xianke Wang, Yanhong Leng, Roberto Togneri, Yijian Xiao, Piao Ding, Shenyi Song, Pingjian Zhang
INTERSPEECH2
2023 Harmonic enhancement using learnable comb filter for light-weight full-band speech enhancement model
Xiaohuai Le, Yiqing Guo, Xianjun Xia, Hua Gao, Yijian Xiao, Piao Ding, Shenyi Song
INTERSPEECH7
2023 An Intra-BRNN and GB-RVQ Based END-TO-END Neural Audio Codec
Linping Xu, Dejun Zhang, Xianjun Xia, Yijian Xiao, Piao Ding, Shenyi Song, Sixing Yin, Ferdous Sohel
INTERSPEECH4
2021 A Two-Stage Approach to Device-Robust Acoustic Scene Classification
abstract
To improve device robustness, a highly desirable key feature of a competitive data-driven acoustic scene classification (ASC) system, a novel two-stage system based on fully convolutional neural networks (CNNs) is proposed. Our two-stage system leverages on an ad-hoc score combination based on two CNN classifiers: (i) the first CNN classifies acoustic inputs into one of three broad classes, and (ii) the second CNN classifies the same inputs into one of ten finergrained classes. Three different CNN architectures are explored to implement the two-stage classifiers, and a frequency sub-sampling scheme is investigated. Moreover, novel data augmentation schemes for ASC are also investigated. Evaluated on DCASE 2020 Task 1a, our results show that the proposed ASC system attains a state-of-the-art accuracy on the development set, where our best system, a two-stage fusion of CNN ensembles, delivers a 81.9% average accuracy among multi-device test data, and it obtains a significant improvement on unseen devices. Finally, neural saliency analysis with class activation mapping (CAM) gives new insights on the patterns learnt by our models.
Hu Hu, Chao-Han Huck Yang, Xianjun Xia, Yajian Wang, Shutong Niu, Li Chai 0002, Juanjuan Li, Hongning Zhu, Sabato Marco Siniscalchi, Yannan Wang, Jun Du 0002, Chin-Hui Lee 0001
ICASSP3
2020 Audio Sound Determination Using Feature Space Attention Based Convolution Recurrent Neural Network
abstract
The classification framework has been popularly adopted to perform sound event detection. However, the existing neural network based classification based approaches treat each feature dimension equally and the varying influence of feature dimensions has not been taken into consideration. To deal with this, we propose a feature space attention based convolution recurrent neural network approach utilizing the varying importance of each feature dimension to perform acoustic event detection. The convolution layers are used to extract the high level information from the audio signals. Then the feature space attention scheme is applied to the extracted features to automatically determine the importance of each feature dimension. Experimental results on the latest TUT Sound Event 2017 dataset demonstrate the improved performance of the proposed approach compared to the existing acoustic event detection systems.
Xianjun Xia, Jingjing Pan, Yannan Wang
ICASSP1
2020 Sound Event Detection Using Multiple Optimized Kernels
abstract
Sound event detection (SED) has been widely applied in real world applications. Convolutional recurrent neural network based SED approaches have achieved state-of-the-art performance. However, the convolution process is typically performed by using a fixed sized kernel, which adversely affects the detection accuracy especially when the acoustic features of different event classes are characterized by high variations. To deal with this, this article proposes a sound event detection technique using a convolutional recurrent neural network framework with multiple convolutional kernels of different sizes. The top performing kernels are selected from a kernel pool based on the unsupervised clustering errors and the accuracies of the temporarily trained models. Afterwards, the selected kernels are fed to multiple convolution layers to deal with the acoustic feature variations. Experimental results on different subsets of AudioSet, namely the DCASE Challenge 2017 Task 4 and DCASE Challenge 2018 Task 4, demonstrate the performance of the proposed approach compared to state-of-the-art systems.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Multi-Task Learning for Acoustic Event Detection Using Event and Frame Position Information
abstract
Acoustic event detection deals with the acoustic signals to determine the sound type and to estimate the audio event boundaries. Multi-label classification based approaches are commonly used to detect the frame wise event types with a median filter applied to determine the happening acoustic events. However, the multi-label classifiers are trained only on the acoustic event types ignoring the frame position within the audio events. To deal with this, this paper proposes to construct a joint learning based multi-task system. The first task performs the acoustic event type detection and the second task is to predict the frame position information. By sharing representations between the two tasks, we can enable the acoustic models to generalize better than the original classifier by averaging respective noise patterns to be implicitly regularized. Experimental results on the monophonic UPC-TALP and the polyphonic TUT Sound Event datasets demonstrate the superior performance of the joint learning method by achieving lower error rate and higher F-score compared to the baseline AED system.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
IEEE Trans. Multim.1
2019 Auxiliary Classifier Generative Adversarial Network With Soft Labels in Imbalanced Acoustic Event Detection
abstract
In acoustic event detection, the training data size of some acoustic events is often small and imbalanced. To deal with this, this paper proposes generating the virtual training data categorically using the auxiliary classifier generative adversarial networks. Soft labels of acoustic events are first calculated to represent the acoustic event localization information. The closer the current frame is to the middle of the manually labeled acoustic event, the higher the soft label will be, which makes the soft labels positively correlated with the acoustic event localization. Then, the acoustic event class and the quantized soft labels are used as the input condition to the auxiliary classifier generative adversarial networks to generate an arbitrary number of training samples. Experimental results on the TUT Sound Event 2016 under the home environment and TUT Sound Event 2017 under the street environment demonstrate the improved performance of the proposed technique compared to existing acoustic event detection systems.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
IEEE Trans. Multim.1
2018 Confidence Based Acoustic Event Detection
abstract
Acoustic event detection, the determination of the acoustic event type and the localisation of the event, has been widely applied in many real-world applications. Many works adopt the multi-label classification technique to perform the polyphonic acoustic event detection with a global threshold to detect the active acoustic events. However, the manually labeled boundaries are error-prone and cannot always be accurate, especially when the frame length is too short to be accurately labeled by human annotators. To deal with this, a confidence is assigned to each frame and acoustic event detection is performed using a multi-variable regression approach in this paper. Experimental results on the latest TUT sound event 2017 database of polyphonic events demonstrate the superior performance of the proposed approach compared to the multi-label classification based AED method.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
ICASSP1
2018 Local Binary Pattern with Random Forest for Acoustic Scene Classification
abstract
This paper presents an approach for acoustic scene classification using the local binary pattern (LBP) and random forest (RF). The audio signal is converted to a Constant-Q transform (CQT) representation and LBP is used to extract the features from this time-frequency representation. The CQT representations are divided into a number of sub-bands to obtain more localized features relevant to the spectral information. We then use random forest to select the most important features for each band of extracted LBP features. For further performance enhancement, we use feature level fusion of LBP and HOG features. The proposed system has achieved an accuracy of 85% on the DCASE 2016 dataset.
Shamsiah Abidin, Xianjun Xia, Roberto Togneri, Ferdous Sohel
ICME2
2018 Random forest classification based acoustic event detection utilizing contextual-information and bottleneck features
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
Pattern Recognit.1
2017 Random forest regression based acoustic event detection with bottleneck features
abstract
This paper deals with random forest regression based acoustic event detection (AED) by combining acoustic features with bottleneck features (BN). The bottleneck features have a good reputation of being inherently discriminative in acoustic signal processing. To deal with the unstructured and complex real-world acoustic events, an acoustic event detection system is constructed using bottleneck features combined with acoustic features. Evaluations were carried out on the UPC-TALP and ITC-Irst databases which consist of highly variable acoustic events. Experimental results demonstrate the usefulness of the low-dimensional and discriminative bottleneck features with relative 5.33% and 5.51% decreases in error rates respectively.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
ICME1
2017 Random forest classification based acoustic event detection
abstract
This paper deals with the acoustic event detection (AED) to improve the detection accuracy of acoustic events. Acoustic event detection task is performed by a regression via classification (RvC) based approach along with the random forest technique. A discretization process is used to convert the continuous frame positions within acoustic events into event duration class labels. Outputs of the category-specific random forest classifiers are then reversed back to the event boundary information. Evaluations on the UPC-TALP database which consists of highly variable acoustic events demonstrate the efficiency of the proposed approaches with improvements in detection error rate compared to the best baseline system.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
ICME1
2017 Frame-Wise Dynamic Threshold Based Polyphonic Acoustic Event Detection
abstract
Acoustic event detection, the determination of the acoustic event type and the localisation of the event, has been widely applied in many real-world applications. Many works adopt multi-label classification techniques to perform the polyphonic acoustic event detection with a global threshold to detect the active acoustic events. However, the global threshold has to be set manually and is highly dependent on the database being tested. To deal with this, we replaced the fixed threshold method with a frame-wise dynamic threshold approach in this paper. Two novel approaches, namely contour and regressor based dynamic threshold approaches are proposed in this work. Experimental results on the popular TUT Acoustic Scenes 2016 database of polyphonic events demonstrated the superior performance of the proposed approaches.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
INTERSPEECH1