Jiayao Sun

dblp:155/6679 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
6since 2021 · last 2025
0009-0007-1002-6879ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Explainable Riemannian Manifold Learning for Application Scene Generalization With Distributed Acoustic Sensing System
abstract
Intrusion event detection in distributed fiber optic sensing systems plays a vital role in urban emergency management and public safety. However, current AI-based systems face high false positive and false negative rates, as well as limited generalizability, due to variations in dataset quality, processing methods, and model architectures. To address these issues, we propose the Riemannian-Transformer-OTDR (RTO) model, an explainable deep-learning approach based on Riemannian manifolds. By exploiting the manifold space, our method identifies stable latent representations of dynamic data distributions from similar events, embedding intrusion detection within this latent manifold. The RTO model accurately identifies identical event types across diverse scenarios, reducing misclassification rates and enhancing model stability. Field tests using a distributed fiber optic sensing system validated the effectiveness of RTO: it achieved 98.43% accuracy in a subway construction scenario and demonstrated its cross-scenario generalization ability by maintaining 94.21% accuracy in gas pipeline monitoring, with 80.79% accuracy in a 100% noisy environment. Additionally, signal structure analysis of misclassified samples revealed insights into the model’s learning process. Our analysis of performance degradation across scenarios and noise levels further clarifies the RTO mechanisms and explainability. Ultimately, the RTO enables precise, retraining-free intrusion event classification across multiple scenarios, offering a practical solution for rapid event detection deployment in various domains.
Puyu Han, Guanfang Shao, Zixing Huang, Jiayao Sun, Xiaobing Shi, Xiongji Yang, Li-Yang Shao
IEEE Internet Things J.6
2024 Dualsep: A Light-Weight Dual-Encoder Convolutional Recurrent Network For Real-Time In-Car Speech Separation
abstract
Advancements in deep learning and voice-activated technologies have driven the development of human-vehicle interaction. Distributed microphone arrays are widely used in incar scenarios because they can accurately capture the voices of passengers from different speech zones. However, the increase in the number of audio channels, coupled with the limited computational resources and low latency requirements of in-car systems, presents challenges for in-car multi-channel speech separation. To migrate the problems, we propose a lightweight framework that cascades digital signal processing (DSP) and neural networks (NN). We utilize fixed beamforming (BF) to reduce computational costs and independent vector analysis (IVA) to provide spatial prior. We employ dual encoders for dual-branch modeling, with spatial encoder capturing spatial cues and spectral encoder preserving spectral information, facilitating spatial-spectral fusion. Our proposed system supports both streaming and non-streaming modes. Experimental results demonstrate the superiority of the proposed system across various metrics. With only 0.83 M parameters and 0.39 real-time factor (RTF) on an Intel Core i7 $(2.6 \mathrm{GHz}) \mathrm{CPU}$, it effectively separates speech into distinct speech zones. Our demos are available at https://honeew.github.io/DualSep/.
Jiayao Sun, Jie Liu 0097, Lei Xie 0001
SLT2
2023 An Exploration of Task-Decoupling on Two-Stage Neural Post Filter for Real-Time Personalized Acoustic Echo Cancellation
abstract
Deep learning based techniques have been popularly adopted in acoustic echo cancellation (AEC). Utilization of speaker representation has extended the frontier of AEC, thus attracting many researchers’ interest in personalized acoustic echo cancellation (PAEC). Meanwhile, task-decoupling strategies are widely adopted in speech enhancement. To further explore the task-decoupling approach, we propose to use a two-stage task-decoupling post-filter (TDPF) in PAEC. Furthermore, a multi-scale local-global speaker representation is applied to improve speaker extraction in PAEC. Experimental results indicate that the task-decoupling model can yield better performance than a single joint network. The optimal approach is to decouple the echo cancellation from noise and interference speech suppression. Based on the task-decoupling sequence, optimal training strategies for the two-stage model are explored afterwards.
Jiayao Sun, Xianjun Xia, Xiaopeng Yan, Yijian Xiao, Lei Xie 0001
ASRU2
2023 Multi-Task Sub-Band Network For Deep Residual Echo Suppression
abstract
This paper introduces the SWANT team’s entry to the ICASSP 2023 AEC Challenge. We submit a system that cascades a linear filter with a neural post-filter. Particularly, we adopt sub-band processing to handle full-band signals and shape the network with multi-task learning, where dual signal voice activity detection (DSVAD) and echo estimation are adopted as auxiliary tasks. Moreover, we particularly improve the time frequency convolution module (TFCM) to increase the receptive field using small convolution kernels. Finally, our system has ranked 4th in ICASSP 2023 AEC Challenge Non-personalized track.
Jiayao Sun, Dawei Luo, Zhaoxia Li, Yukai Jv
ICASSP1
2022 S-DCCRN: Super Wide Band DCCRN with Learnable Complex Feature for Speech Enhancement
abstract
In speech enhancement, complex neural network has shown promising performance due to their effectiveness in processing complex-valued spectrum. Most of the recent speech enhancement approaches mainly focus on wide-band signal with a sampling rate of 16K Hz. However, research on super wide band (e.g., 32K Hz) or even full-band (48K) denoising using deep learning is still in its infancy due to the difficulty of modeling more frequency bands and particularly high frequency components. In this paper, we extend our previous deep complex convolution recurrent neural network (DCCRN) substantially to a super wide band version–S-DCCRN, to perform speech denoising on speech of 32K Hz sampling rate. We first employ a cascaded sub-band and full-band processing module, which consists of two small-footprint DCCRNs–one operates on sub-band signal and one operates on full-band signal, aiming at benefiting from both local and global frequency information. Moreover, instead of simply adopting the STFT feature as input, we use a complex feature encoder trained in an end-to-end manner to refine the information of different frequency bands. We also use a complex feature decoder to revert the feature to time-frequency domain. Finally, a learnable spectrum compression method is adopted to adjust the energy of different frequency bands, which is beneficial for neural network learning. The proposed model, S-DCCRN, has surpassed PercepNet as well as several competitive models and achieves state-of-the-art performance in terms of speech quality and intelligibility. Ablation studies also demonstrate the effectiveness of different contributions.
Shubo Lv, Yihui Fu, Mengtao Xing, Jiayao Sun, Lei Xie 0001, Yannan Wang
ICASSP4
2022 Multi-Task Deep Residual Echo Suppression with Echo-Aware Loss
abstract
This paper introduces the NWPU Team’s entry to the ICASSP 2022 AEC Challenge. We take a hybrid approach that cascades a linear AEC with a neural post-filter. The former is used to deal with the linear echo components while the latter suppresses the residual non-linear echo components. We use gated convolutional F-T-LSTM neural network (GFTNN) as the backbone and shape the post-filter by a multi-task learning (MTL) framework, where a voice activity detection (VAD) module is adopted as an auxiliary task along with echo suppression, with the aim to avoid over suppression that may cause speech distortion. Moreover, we adopt an echo-aware loss function, where the mean square error (MSE) loss can be optimized particularly for every time-frequency bin (TF-bin) according to the signal-to-echo ratio (SER), leading to further suppression on the echo. Extensive ablation study shows that the time delay estimation (TDE) module in neural post-filter leads to better perceptual quality, and an adaptive filter with better convergence will bring consistent performance gain for the post-filter. Besides, we find that using the linear echo as the input of our neural post-filter is a better choice than using the reference signal directly. In the ICASSP 2022 AEC-Challenge, our approach has ranked the 1st place on word accuracy (WAcc) (0.817) and the 3rd place on both mean opinion score (MOS) (4.502) and the final score (0.864).
Jiayao Sun, Yihui Fu, Lei Xie 0001
ICASSP3
2015 A semi-supervised incremental learning method based on adaptive probabilistic hypergraph for video semantic detection
Yongzhao Zhan 0001, Jiayao Sun, DeJiao Niu, Qirong Mao, Jianping Fan 0001
Multim. Tools Appl.2