Anbai Jiang

dblp:334/2161 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0009-0395-994XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning
abstract
The goal of the acoustic scene classification (ASC) task is to classify recordings into one of the predefined acoustic scene classes. However, in real-world scenarios, ASC systems often encounter challenges such as recording device mismatch, low-complexity constraints, and the limited availability of labeled data. To alleviate these issues, in this paper, a data-efficient and low-complexity ASC system is built with a new model architecture and better training strategies. Specifically, we firstly design a new low-complexity architecture named Rep-Mobile by integrating multi-convolution branches which can be reparameterized at inference. Compared to other models, it achieves better performance and less computational complexity. Then we apply the knowledge distillation strategy and provide a comparison of the data efficiency of the teacher model with different architectures. Finally, we propose a progressive pruning strategy, which involves pruning the model multiple times in small amounts, resulting in better performance compared to a single step pruning. Experiments are conducted on the TAU dataset. With Rep-Mobile and these training strategies, our proposed ASC system achieves the state-of-the-art (SOTA) results so far, while also winning the first place with a significant advantage over others in the DCASE2024 Challenge.
Bing Han 0008, Wen Huang 0004, Zhengyang Chen, Anbai Jiang, Pingyi Fan, Cheng Lu 0007, Zhiqiang Lv, Jia Liu 0001, Weiqiang Zhang 0001, Yanmin Qian
ICASSP4
2025 Adaptive Prototype Learning for Anomalous Sound Detection with Partially Known Attributes
abstract
Adapting pre-trained models has become the dominant approach for anomalous sound detection (ASD), where classifying the attributes of machine working status is commonly chosen as the deputy task for fine-tuning. However, attributes might be intractable to collect for some machines, causing the label to bear mixed granularity and thus deprecating the ASD performance. Therefore, we propose an adaptive proto-type learning scheme for fine-tuning pre-trained models, which adaptively scales coarse-grained labels to sub-centers so as to keep consistency with fine-grained labels. To deal with domain shift, we employ SMOTE to over-sample the prototypes of the target domain. The experiment on the DCASE 2024 ASD dataset demonstrates the efficacy of the proposed scheme, setting up a new milestone of 65.01% on both sets and outperforming the best system of the challenge. A detailed ablation study is also conducted to validate the effectiveness.
Anbai Jiang, Xinhu Zheng, Yihong Qiu, Pingyi Fan, Cheng Lu 0007, Jia Liu 0001
ICASSP1
2024 CoopASD: Cooperative Machine Anomalous Sound Detection with Privacy Concerns
abstract
Machine anomalous sound detection (ASD) has emerged as one of the most promising applications in the Industrial Internet of Things (IIoT) due to its unprecedented efficacy in mitigating risks of malfunctions and promoting production efficiency. Previous works mainly investigated the machine ASD task under centralized settings. However, developing the ASD system under decentralized settings is crucial in practice, since the machine data are dispersed in various factories and the data should not be explicitly shared due to privacy concerns. To enable these factories to cooperatively develop a scalable ASD model while preserving their privacy, we propose a novel framework named CoopASD, where each factory trains an ASD model on its local dataset, and a central server aggregates these local models periodically. We employ a pre-trained model as the backbone of the ASD model to improve its robustness and develop specialized techniques to stabilize the model under a completely non-iid and domain shift setting. Compared with previous state-of-the-art (SOTA) models trained in centralized settings, CoopASD showcases competitive results with negligible degradation of 0.08%. We also conduct extensive ablation studies to demonstrate the effectiveness of CoopASD.
Anbai Jiang, Pingyi Fan
GLOBECOM1
2024 Exploring Large Scale Pre-Trained Models for Robust Machine Anomalous Sound Detection
abstract
Machine anomalous sound detection is a useful technique for various applications, but it often suffers from poor generalization due to the challenges of data collection and complex acoustic environment. To address this issue, we propose a robust machine anomalous sound detection model that leverages self-supervised pre-trained models on large-scale speech data. Specifically, we assign different weights to the features from different layers of the pre-trained model and then use the working condition as the label for self-supervised classification fine-tuning. Moreover, we introduce a data augmentation method that simulates different operating states of the machine to enrich the dataset. Furthermore, we devise a transformer pooling method that fuses the features of different segments. Experiments on the DCASE2023 dataset show that our proposed method outperforms the commonly used reconstruction-based autoencoder and classification-based convolutional network by a large margin, demonstrating the effectiveness of large-scale pre-training for enhancing the generalization and robustness of machine anomalous sound detection. In Task2 of DCASE2023, we achieve 2nd place with these methods.
Bing Han 0008, Zhiqiang Lv, Anbai Jiang, Wen Huang 0004, Zhengyang Chen, Yufeng Deng, Cheng Lu 0007, Weiqiang Zhang 0001, Pingyi Fan, Jia Liu 0001, Yanmin Qian
ICASSP3
2024 AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection
Anbai Jiang, Bing Han 0008, Zhiqiang Lv, Yufeng Deng, Weiqiang Zhang 0001, Xie Chen 0001, Yanmin Qian, Jia Liu 0001, Pingyi Fan
INTERSPEECH1
2024 Improving Anomalous Sound Detection Via Low-Rank Adaptation Fine-Tuning of Pre-Trained Audio Models
abstract
Anomalous Sound Detection (ASD) has gained significant interest through the application of various Artificial Intelligence (AI) technologies in industrial settings. Though possessing great potential, ASD systems can hardly be readily deployed in real production sites due to the generalization problem, which is primarily caused by the difficulty of data collection and the complexity of environmental factors. This paper introduces a robust ASD model that leverages audio pre-trained models. Specifically, we fine-tune these models using machine operation data, employing SpecAug as a data augmentation strategy. Additionally, we investigate the impact of utilizing Low-Rank Adaptation (LoRA) tuning instead of full fine-tuning to address the problem of limited data for fine-tuning. Our experiments on the DCASE2023 Task 2 dataset establish a new benchmark of 77.75% on the evaluation set, with a significant improvement of 6.48% compared with previous state-of-the-art (SOTA) models, including top-tier traditional convolutional networks and speech pre-trained models, which demonstrates the effectiveness of audio pre-trained models with LoRA tuning. Ablation studies are also conducted to showcase the efficacy of the proposed scheme.
Xinhu Zheng, Anbai Jiang, Bing Han 0008, Yanmin Qian, Pingyi Fan, Jia Liu 0001, Weiqiang Zhang 0001
SLT2
2023 Decoupling Detectors for Scalable Anomaly Detection in AIoT Systems with Multiple Machines
abstract
The fast-developing Artificial Internet of Things (AIoT) technologies enable the consistent monitoring of multiple machines, by which machine failures can be detected in the early phases, and production efficiency and system management can be greatly promoted, bringing huge significance for anomaly detection. However, in most cases, anomalies are not provided for training, and the lack of direct supervision deprecates the anomaly detection performance. For the application viewpoint, the detector is required to generalize well on multiple machines, except for being computationally efficient. The computational cost is strictly limited, which is a great challenge for mobile and embedded devices. In face of these issues, we propose MobileAnoNet, which decouples an end-to-end detector into a front-end feature extractor and a back-end anomaly detector. The front-end extractor, consuming most computation, is unified for all machine types, while the back-end detector is specialized for each machine type, improving the detection capacity. The model is trained by handy labels of machine types and working conditions, in which multiple classification heads are attached behind the feature extractor during training. The performance of the model is evaluated on two DCASE datasets focusing on machine audio anomaly detection. It's shown that MobileAnoNet achieves a general improvement of 6.9% and 8.8% on two datasets, respectively. The ablation study demonstrates that multi-task learning promotes the general representation capacity. The source code is available at: www.github.com/hqj-les30/MobileAnoNet.
Qijun Hou, Anbai Jiang, Weiqiang Zhang 0001, Pingyi Fan, Jia Liu 0001
GLOBECOM2
2023 Unsupervised Anomaly Detection and Localization of Machine Audio: A Gan-Based Approach
abstract
Automatic detection of machine anomaly remains challenging for machine learning. We believe the capability of generative adversarial network (GAN) suits the need of machine audio anomaly detection, yet rarely has this been investigated by previous work. In this paper, we propose AEGAN-AD, a totally unsupervised approach in which the generator (also an autoencoder) is trained to reconstruct input spectrograms. It is pointed out that the denoising nature of reconstruction deprecates its capacity. Thus, the discriminator is redesigned to aid the generator during both training stage and detection stage. The performance of AEGAN-AD on the dataset of DCASE 2022 Challenge TASK 2 demonstrates the state-of-the-art result on five machine types. A novel anomaly localization method is also investigated. Source code available at: www.github.com/jianganbai/AEGAN-AD
Anbai Jiang, Weiqiang Zhang 0001, Yufeng Deng, Pingyi Fan, Jia Liu 0001
ICASSP1