Qiaoxi Zhu

dblp:236/4342 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0003-3942-0945ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Graph-Enhanced Dual-Stream Feature Fusion with Pre-Trained Model for Acoustic Traffic Monitoring
abstract
Microphone array techniques are widely used in sound source localization and smart city acoustic-based traffic monitoring, but these applications face significant challenges due to the scarcity of labeled real-world traffic audio data and the complexity and diversity of application scenarios. The DCASE Challenge’s Task 10 focuses on using multi-channel audio signals to count vehicles (cars or commercial vehicles) and identify their directions (left-to-right or vice versa). In this paper, we propose a graph-enhanced dual-stream feature fusion network (GEDF-Net) for acoustic traffic monitoring, which simultaneously considers vehicle type and direction to improve detection. We propose a graph-enhanced dual-stream feature fusion strategy which consists of a vehicle type feature extraction (VTFE) branch, a vehicle direction feature extraction (VDFE) branch, and a frame-level feature fusion module to combine the type and direction feature for enhanced performance. A pre-trained model (PANNs) is used in the VTFE branch to mitigate data scarcity and enhance the type features, followed by a graph attention mechanism to exploit temporal relationships and highlight important audio events within these features. The frame-level fusion of direction and type features enables fine-grained feature representation, resulting in better detection performance. Experiments demonstrate the effectiveness of our proposed method. GEDF-Net is our submission that achieved 1st place in the DCASE 2024 Challenge Task 10.
Shitong Fan, Feiyang Xiao, Shuhan Qi, Qiaoxi Zhu, Wenwu Wang 0001, Jian Guan 0001
ICASSP5
2025 Disentangling Hierarchical Features for Anomalous Sound Detection Under Domain Shift
abstract
Anomalous sound detection (ASD) encounters difficulties with domain shift, where the sounds of machines in target domains differ significantly from those in source domains due to varying operating conditions. Existing methods typically employ domain classifiers to enhance detection performance, but they often overlook the influence of domain-unrelated information. This oversight can hinder the model’s ability to clearly distinguish between domains, thereby weakening its capacity to differentiate normal from abnormal sounds. In this paper, we propose a Gradient Reversal-based Hierarchical feature Disentanglement (GRHD) method to address the above challenge. GRHD uses gradient reversal to separate domain-related features from domain-unrelated ones, resulting in more robust feature representations. Additionally, the method employs a hierarchical structure to guide the learning of fine-grained, domain-specific features by leveraging available metadata, such as section IDs and machine sound attributes. Experimental results on the DCASE 2022 Challenge Task 2 dataset demonstrate that the proposed method significantly improves ASD performance under domain shift.
Jian Guan 0001, Jiantong Tian, Qiaoxi Zhu, Feiyang Xiao, Hejing Zhang, Xubo Liu 0001
ICASSP3
2025 Spectral-Temporal Fusion Representation for Person-in-Bed Detection
abstract
This study is based on the ICASSP 2025 Signal Processing Grand Challenge’s Accelerometer-Based Person-in-Bed Detection Challenge, which aims to determine bed occupancy using accelerometer signals. The task is divided into two tracks: "in bed" and "not in bed" segmented detection and streaming detection, facing challenges such as individual differences, posture variations, and external disturbances. We propose a spectral-temporal fusion-based feature representation method with mixup data augmentation, and adopt Intersection over Union (IoU) loss to optimize detection accuracy. In the two tracks, our method achieved outstanding results of 100.00% and 95.55% in detection scores, securing first place and third place, respectively.
Xuefeng Yang, Shiheng Zhang, Feiyang Xiao, Qiaoxi Zhu
ICASSP6
2025 Attacking Voice Anonymization Systems with Augmented Feature and Speaker Identity Difference
abstract
This study focuses on the First VoicePrivacy Attacker Challenge within the ICASSP 2025 Signal Processing Grand Challenge, which aims to develop speaker verification systems capable of determining whether two anonymized speech signals are from the same speaker. However, differences between feature distributions of original and anonymized speech complicate this task. To address this challenge, we propose an attacker system that combines Data Augmentation enhanced feature representation and Speaker Identity Difference enhanced classifier to improve verification performance, termed DA-SID. Specifically, data augmentation strategies (i.e., data fusion and SpecAugment) are utilized to mitigate feature distribution gaps, while probabilistic linear discriminant analysis (PLDA) is employed to further enhance speaker identity difference. Our system significantly outperforms the baseline, demonstrating exceptional effectiveness and robustness against various voice anonymization systems, ultimately securing a top-5 ranking in the challenge.
Yanzhe Zhang 0001, Zhonghao Bi, Feiyang Xiao, Xuefeng Yang, Qiaoxi Zhu, Jian Guan 0001
ICASSP5
2024 Hierarchical Metadata Information Constrained Self-Supervised Learning for Anomalous Sound Detection under Domain Shift
abstract
Self-supervised learning methods have achieved promising performance for anomalous sound detection (ASD) under domain shift by incorporating the metadata of domain shift types and machine sound attributes in feature learning. However, the relation between domain shifts and machine sound attributes has yet to be fully utilised despite their potential benefits for characterising domain shifts. This paper presents a hierarchical metadata information constrained self-supervised ASD method, where the hierarchical relation between domain shift types (section IDs) and attributes is constructed and used as constraints to improve feature representation. In addition, we propose an attribute-group-centre based method for calculating the anomaly score under the domain shift condition. Experiments show improved audio feature learning over the state-of-the-art methods in DCASE 2022 challenge Task 2.
Haiyan Lan, Qiaoxi Zhu, Jian Guan 0001, Yuming Wei, Wenwu Wang 0001
ICASSP2
2024 First-Shot Unsupervised Anomalous Sound Detection with Unknown Anomalies Estimated by Metadata-Assisted Audio Generation
abstract
First-shot (FS) unsupervised anomalous sound detection (ASD) is a brand-new task introduced in DCASE 2023 Challenge Task 2, where the anomalous sounds for the target machine types are unseen in training. Existing methods often rely on the availability of normal and abnormal sound data from the target machines. However, due to the lack of anomalous sound data for the target machine types, it becomes challenging when adapting the existing ASD methods to the first-shot task. In this paper, we propose a new framework for the first-shot unsupervised ASD, where metadata-assisted audio generation is used to estimate unknown anomalies, by utilising the available machine information (i.e., metadata and sound data) to fine-tune a text-to-audio generation model for generating the anomalous sounds that contain unique acoustic characteristics accounting for each different machine type. We then use the method of Time-Weighted Frequency domain audio Representation with Gaussian Mixture Model (TWFRGMM) as the backbone to achieve the first-shot unsupervised ASD. Our proposed FS-TWFR-GMM method achieves competitive performance amongst top systems in DCASE 2023 Challenge Task 2, while requiring only 1% model parameters for detection, as validated in our experiments.
Hejing Zhang, Qiaoxi Zhu, Jian Guan 0001, Haohe Liu, Feiyang Xiao, Jiantong Tian, Xinhao Mei, Xubo Liu 0001, Wenwu Wang 0001
ICASSP2
2023 Time-Weighted Frequency Domain Audio Representation with GMM Estimator for Anomalous Sound Detection
abstract
Although deep learning is the mainstream method in unsupervised anomalous sound detection, Gaussian Mixture Model (GMM) with statistical audio frequency representation as input can achieve comparable results with much lower model complexity and fewer parameters. Existing statistical frequency representations, e.g. the log-Mel spectrogram’s average or maximum over time, do not always work well for different machines. This paper presents Time-Weighted Frequency Domain Representation (TWFR) with the GMM method (TWFR-GMM) for anomalous sound detection. The TWFR is a generalized statistical frequency domain representation that can adapt to different machine types, using the global weighted ranking pooling over time-domain. This allows GMM estimator to recognize anomalies, even under domain-shift conditions, as visualized with a Mahalanobis distance-based metric. Experiments on DCASE 2022 Challenge Task2 dataset show that our method has better detection performance than recent deep learning methods. TWFR-GMM is the core of our submission that achieved the 3rd place in DCASE 2022 Challenge Task2.
Jian Guan 0001, Youde Liu, Qiaoxi Zhu, Tieran Zheng, Jiqing Han 0001, Wenwu Wang 0001
ICASSP3
2023 Anomalous Sound Detection Using Audio Representation with Machine ID Based Contrastive Learning Pretraining
abstract
Existing contrastive learning methods for anomalous sound detection refine the audio representation of each audio sample by using the contrast between the samples’ augmentations (e.g., with time or frequency masking). However, they might be biased by the augmented data, due to the lack of physical properties of machine sound, thereby limiting the detection performance. This paper uses contrastive learning to refine audio representations for each machine ID, rather than for each audio sample. The proposed two-stage method uses contrastive learning to pretrain the audio representation model by incorporating machine ID and a self-supervised ID classifier to fine-tune the learnt model, while enhancing the relation between audio features from the same ID. Experiments show that our method outperforms the state-of-the-art methods using contrastive learning or self-supervised classification in overall anomaly detection performance and stability on DCASE 2020 Challenge Task2 dataset.
Jian Guan 0001, Feiyang Xiao, Youde Liu, Qiaoxi Zhu, Wenwu Wang 0001
ICASSP4
2023 Anomalous Sound Detection Using Self-Attention-Based Frequency Pattern Analysis of Machine Sounds
Hejing Zhang, Jian Guan 0001, Qiaoxi Zhu, Feiyang Xiao, Youde Liu
INTERSPEECH3
2023 Graph Attention for Automated Audio Captioning
abstract
State-of-the-art audio captioning methods typically use the encoder-decoder structure with pretrained audio neural networks (PANNs) as encoders for feature extraction. However, the convolution operation used in PANNs is limited in capturing the long-time dependencies within an audio signal, thereby leading to potential performance degradation in audio captioning. This letter presents a novel method using graph attention (GraphAC) for encoder-decoder based audio captioning. In the encoder, a graph attention module is introduced after the PANNs to learn contextual association (i.e. the dependency among the audio features over different time frames) through an adjacency graph, and a top-kmask is used to mitigate the interference from noisy nodes. The learnt contextual association leads to a more effective feature representation with feature node aggregation. As a result, the decoder can predict important semantic information about the acoustic scene and events based on the contextual associations learned from the audio signal. Experimental results show that GraphAC outperforms the state-of-the-art methods with PANNs as the encoders, thanks to the incorporation of the graph attention module into the encoder for capturing the long-time dependencies within the audio signal. The source code is available at https://github.com/LittleFlyingSheep/GraphAC.
Feiyang Xiao, Jian Guan 0001, Qiaoxi Zhu, Wenwu Wang 0001
IEEE Signal Process. Lett.3
2022 Anomalous Sound Detection Using Spectral-Temporal Information Fusion
abstract
Unsupervised anomalous sound detection aims to detect unknown abnormal sounds of machines from normal sounds. However, the state-of-the-art approaches are not always stable and perform dramatically differently even for machines of the same type, making it impractical for general applications. This paper proposes a spectral-temporal fusion based self-supervised method to model the feature of the normal sound, which improves the stability and performance consistency in detection of anomalous sounds from individual machines, even of the same type. Experiments on the DCASE 2020 Challenge Task 2 dataset show that the proposed method achieved 81.39%, 83.48%, 98.22% and 98.83% in terms of the minimum AUC (worst-case detection performance amongst individuals) in four types of real machines (fan, pump, slider and valve), respectively, giving 31.79%, 17.78%, 10.42% and 21.13% improvement compared to the state-of-the-art method, i.e., Glow_Aff. Moreover, the proposed method has improved AUC (average performance of individuals) for all the types of machines in the dataset. The source codes are available at https://github.com/liuyoude/STgram_MFN
Youde Liu, Jian Guan 0001, Qiaoxi Zhu, Wenwu Wang 0001
ICASSP3
2022 Local Information Assisted Attention-Free Decoder for Audio Captioning
abstract
Automated audio captioning aims to describe audio data with captions using natural language. Existing methods often employ an encoder-decoder structure, where the attention-based decoder (e.g., Transformer decoder) is widely used and achieves state-of-the-art performance. Although this method effectively captures global information within audio data via the self-attention mechanism, it may ignore the event with short time duration, due to its limitation in capturing local information in an audio signal, leading to inaccurate prediction of captions. To address this issue, we propose a method using the pretrained audio neural networks (PANNs) as the encoder and local information assisted attention-free Transformer (LocalAFT) as the decoder. The novelty of our method is in the proposal of the LocalAFT decoder, which allows local information within an audio signal to be captured while retaining the global information. This enables the events of different duration, including short duration, to be captured for more precise caption generation. Experiments show that our method outperforms the state-of-the-art methods in Task 6 of the DCASE 2021 Challenge with the standard attention-based decoder for caption generation.
Feiyang Xiao, Jian Guan 0001, Haiyan Lan, Qiaoxi Zhu, Wenwu Wang 0001
IEEE Signal Process. Lett.4
2020 An Acoustic Modelling Based Remote Error Sensing Approach for Quiet Zone Generation in a Noisy Environment
abstract
Remote error sensing is required in active noise control systems when they are used to create a quiet zone in a noisy environment with the constraint that the error microphones cannot be inside the zone. The challenge in remote error sensing is to estimate the sound pressure in the target zone with a small number of physical microphones outside it. The spatial harmonic decomposition method uses wave domain sound field parameterisation to reduce the required number of the error microphones but can only provide accurate estimation below a certain frequency. This paper presents an improved approach to increase the effective frequency range based on the acoustic modelling. The simulation results demonstrate the proposed method can provide more than 20 dB noise reduction up to 1650 Hz for a quiet zone with a radius of 0.1 m by using only three microphones under the studied situations.
Qiaoxi Zhu, Xiaojun Qiu, Ian S. Burnett
ICASSP1
2019 Robust Personal Audio Geometry Optimization in the SVD-Based Modal Domain
abstract
Personal audio generates sound zones in a shared space to provide private and personalized listening experiences with minimized interference between consumers. Regularization has been commonly used to increase the robustness of such systems against potential perturbations in the sound reproduction. However, the performance is limited by the system geometry such as the number and location of the loudspeakers and controlled zones. This paper proposes a geometry optimization method to find the most geometrically robust approach for personal audio amongst all available candidate system placements. The proposed method aims to approach the most “natural” sound reproduction so that the solo control of the listening zone coincidently accompanies the preferred quiet zone. Being formulated in the SVD-based modal domain, the method is demonstrated by applications in three typical personal audio optimizations, i.e., the acoustic contrast control, the pressure matching, and the planarity control. Simulation results show that the proposed method can obtain the system geometry with better avoidance of “occlusion,” improved robustness to regularization, and improved broadband equalization.
Qiaoxi Zhu, Philip Coleman, Xiaojun Qiu, Ming Wu 0005, Jun Yang 0004, Ian S. Burnett
IEEE ACM Trans. Audio Speech Lang. Process.1