Yijian Xiao

dblp:342/8000 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021
YearPublicationVenuePosition
2024 RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
Mingshuai Liu, Zhuangqi Chen, Xiaopeng Yan, Yuanjun Lv, Xianjun Xia, Chuanzeng Huang, Yijian Xiao, Lei Xie 0001
INTERSPEECH7
2024 BS-PLCNet 2: Two-stage Band-split Packet Loss Concealment Network with Intra-model Knowledge Distillation
Xianjun Xia, Chuanzeng Huang, Yijian Xiao, Lei Xie 0001
INTERSPEECH4
2023 An Exploration of Task-Decoupling on Two-Stage Neural Post Filter for Real-Time Personalized Acoustic Echo Cancellation
abstract
Deep learning based techniques have been popularly adopted in acoustic echo cancellation (AEC). Utilization of speaker representation has extended the frontier of AEC, thus attracting many researchers’ interest in personalized acoustic echo cancellation (PAEC). Meanwhile, task-decoupling strategies are widely adopted in speech enhancement. To further explore the task-decoupling approach, we propose to use a two-stage task-decoupling post-filter (TDPF) in PAEC. Furthermore, a multi-scale local-global speaker representation is applied to improve speaker extraction in PAEC. Experimental results indicate that the task-decoupling model can yield better performance than a single joint network. The optimal approach is to decouple the echo cancellation from noise and interference speech suppression. Based on the task-decoupling sequence, optimal training strategies for the two-stage model are explored afterwards.
Jiayao Sun, Xianjun Xia, Xiaopeng Yan, Yijian Xiao, Lei Xie 0001
ASRU6
2023 A Progressive Neural Network for Acoustic Echo Cancellation
abstract
Acoustic echo cancellation is a key issue in hand-free communication systems. In this paper, we proposed a hybrid signal processing and deep echo cancellation method, where a two-stage neural network is designed to remove residual echo progressively. For the personalized acoustic echo cancellation, we proposed to decouple the tasks of echo cancellation and target speech extraction, and introduced a speaker attentive module for personalized separation, where the ECAPA-TDNN is used for speaker embedding generation. The proposed method (ByteAudio-18) ranked first on both Track 1 and Track 2 in ICASSP 2023 AEC Challenge.
Zhuangqi Chen, Xianjun Xia, Guoliang Xie, Pingjian Zhang, Yijian Xiao
ICASSP8
2023 Two-Stage Neural Network for ICASSP 2023 Speech Signal Improvement Challenge
abstract
In ICASSP 2023 speech signal improvement challenge, we developed a dual-stage neural model which improves speech signal quality induced by different distortions in a stage-wise divide-and-conquer fashion. Specifically, in the first stage, the speech improvement network focuses on recovering the missing components of the spectrum, while in the second stage, our model aims to further suppress noise, reverberation, and artifacts introduced by the first-stage model. Achieving 0.446 in the final score and 0.517 in the P.835 score, our system ranks 4th in the non-real-time track.
Mingshuai Liu, Shubo Lv, Runduo Han, Xianjun Xia, Yijian Xiao, Lei Xie 0001
ICASSP8
2023 A Lightweight Fourier Convolutional Attention Encoder for Multi-Channel Speech Enhancement
abstract
Beamforming weights prediction via deep neural networks has been one of the main methods in multi-channel speech enhancement tasks. The spectral-spatial cues are crucial in beamforming weights estimation, however, many existing works fail to optimally predict the beamforming weights with an absence of adequate spectral-spatial information learning. To tackle this challenge, we propose a Fourier convolutional attention encoder (FCAE) to provide a global receptive field over the frequency axis and boost the learning of spectral contexts and cross-channel features. Besides, a new convolutional recurrent encoder-decoder (CRED) structure is proposed in this work, within which FCAEs, attention blocks with skip connections and a deep feedback sequential memory network (DFSMN) serving as recurrent module are involved. The proposed CRED structure is exploited to capture the spectral-spatial joint information to obtain accurate estimation of beamforming weights. Experimental results demonstrate the superiority of the proposed approach with only 0.74M parameters and a PESQ improvement from 2.225 to 2.359 on the ConferencingSpeech2021 challenge development test set.
Xianjun Xia, Yijian Xiao, Piao Ding, Shenyi Song, Roberto Togneri
ICASSP6
2023 A Two-stage Progressive Neural Network for Acoustic Echo Cancellation
abstract
Recent studies in deep learning based acoustic echo cancellation proves the benefits of introducing a linear echo cancellation module. However, the convergence problem and potential target speech distortion impose an additional learning burden for the neural network. In this paper, we propose a two-stage progressive neural network consisting of a coarse-stage and a fine-stage module. For the coarse-stage, a light-weighted network module is designed to suppress partial echo and potential noise, where a voice activity detection path is used to enhance the learned features. For the fine-stage, a larger network is employed to deal with the more complex echo path and restore the near-end speech. We have conducted extensive experiments to verify the proposed method, and the results show that the proposed two-stage method provides a superior performance to other state-of-the-art methods.
Zhuangqi Chen, Xianjun Xia, Xianke Wang, Yanhong Leng, Roberto Togneri, Yijian Xiao, Piao Ding, Shenyi Song, Pingjian Zhang
INTERSPEECH8
2023 Harmonic enhancement using learnable comb filter for light-weight full-band speech enhancement model
Xiaohuai Le, Yiqing Guo, Xianjun Xia, Hua Gao, Yijian Xiao, Piao Ding, Shenyi Song
INTERSPEECH9
2023 An Intra-BRNN and GB-RVQ Based END-TO-END Neural Audio Codec
Linping Xu, Dejun Zhang, Xianjun Xia, Yijian Xiao, Piao Ding, Shenyi Song, Sixing Yin, Ferdous Sohel
INTERSPEECH6