EDBT 2026 Demo / reviewers in the wild / expert
Yang Xiao 0019
dblp:181/1848-19
· DBLP profile ↗
12ranked-venue papers
7as first author
12since 2021 · last 2025
0009-0005-9329-7425ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dark Experience for Incremental Keyword SpottingabstractSpoken keyword spotting (KWS) is crucial for identifying keywords within audio inputs and is widely used in applications like Apple Siri and Google Home, particularly on edge devices. Current deep learning-based KWS systems, which are typically trained on a limited set of keywords, can suffer from performance degradation when encountering new domains, a challenge often addressed through few-shot fine-tuning. However, this adaptation frequently leads to catastrophic forgetting, where the model’s performance on original data deteriorates. Progressive continual learning (CL) strategies have been proposed to overcome this, but they face limitations such as the need for task-ID information and increased storage, making them less practical for lightweight devices. To address these challenges, we introduce Dark Experience for Keyword Spotting (DE-KWS), a novel CL approach that leverages dark knowledge to distill past experiences throughout the training process. DE-KWS combines rehearsal and distillation, using both ground truth labels and logits stored in a memory buffer to maintain model performance across tasks. Evaluations on the Google Speech Command dataset show that DE-KWS outperforms existing CL baselines in average accuracy without increasing model size, offering an effective solution for resource-constrained edge devices. The scripts are available on GitHub1for future research. Tianyi Peng, Yang Xiao 0019 |
ICASSP | 2 |
| 2025 | UCIL: An Unsupervised Class Incremental Learning Approach for Sound Event DetectionabstractThis work explores class-incremental learning (CIL) for sound event detection (SED), advancing adaptability towards real-world scenarios. CIL’s success in domains like computer vision inspired our SED-tailored method, addressing the unique challenges of diverse and complex audio environments. Our approach employs an independent unsupervised learning frame-work with a distillation loss function to integrate new sound classes while preserving the SED model consistency across incremental tasks. We further enhance this framework with a sample selection strategy for unlabeled data and a balanced exemplar update mechanism, ensuring varied and illustrative sound representations. Evaluating various continual learning methods on the DCASE 2023 Task 4 dataset, our research offers insights into each method’s applicability for real-world SED systems that can have newly added sound classes. The findings also delineate future directions of CIL in dynamic audio settings. Yang Xiao 0019, Rohan Kumar Das |
ICASSP | 1 |
| 2025 | Exploring Text-Queried Sound Event Detection with Audio Source SeparationabstractIn sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we first pre-train a language-queried audio source separation (LASS) model to separate the audio tracks corresponding to different events from the input audio. Then, multiple target SED branches are employed to detect individual events. AudioSep is a state-of-the-art LASS model, but has limitations in extracting dynamic audio information because of its pure convolutional structure for separation. To address this, we integrate a dual-path recurrent neural network block into the model. We refer to this structure as AudioSep-DP, which achieves the first place in DCASE 2024 Task 9 on language-queried audio source separation (objective single model track). Experimental results show that TQ-SED can significantly improve the SED performance, with an improvement of 7.22% on F1 score over the conventional framework. Additionally, we setup comprehensive experiments to explore the impact of model complexity. The source code and pre-trained model are released at https://github.com/apple-yinhan/TQ-SED. Han Yin, Jisheng Bai, Yang Xiao 0019, Hui Wang 0030, Yafeng Chen, Rohan Kumar Das, Chong Deng |
ICASSP | 3 |
| 2025 | Where's That Voice Coming? Continual Learning for Sound Source LocalizationabstractSound source localization (SSL) is essential for many speech-processing applications. Deep learning models have achieved high performance, but often fail when the training and inference environments differ. Adapting SSL models to dynamic acoustic conditions faces a major challenge: catastrophic forgetting. In this work, we propose an exemplar-free continual learning strategy for SSL (CL-SSL) to address such a forgetting phenomenon. CL-SSL applies task-specific sub-networks to adapt across diverse acoustic environments while retaining previously learned knowledge. It also uses a scaling mechanism to limit parameter growth, ensuring consistent performance across incremental tasks. We evaluated CL-SSL on simulated data with varying microphone distances and real-world data with different noise levels. The results demonstrate CL-SSL’s ability to maintain high accuracy with minimal parameter increase, offering an efficient solution for SSL applications. Yang Xiao 0019, Rohan Kumar Das |
ICME | 1 |
| 2025 | TF-Mamba: A Time-Frequency Network for Sound Source Localization
Yang Xiao 0019, Rohan Kumar Das |
INTERSPEECH | 1 |
| 2025 | Listen, Analyze, and Adapt to Learn New Attacks: An Exemplar-Free Class Incremental Learning Method for Audio Deepfake Source Tracing
Yang Xiao 0019, Rohan Kumar Das |
INTERSPEECH | 1 |
| 2025 | AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation
Yang Xiao 0019, Tianyi Peng, Yanghao Zhou, Rohan Kumar Das |
INTERSPEECH | 1 |
| 2025 | EnvSDD: Benchmarking Environmental Sound Deepfake Detection
Han Yin, Yang Xiao 0019, Rohan Kumar Das, Jisheng Bai, Haohe Liu, Wenwu Wang 0001, Mark D. Plumbley |
INTERSPEECH | 2 |
| 2025 | Noise-Robust Sound Event Detection and Counting via Language-Queried Sound Separation
Yuanjian Chen, Yang Xiao 0019, Han Yin, Yadong Guan, Xubo Liu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2025 | XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack DetectionabstractTransformers and their variants have achieved great success in speech processing. However, their multi-head self-attention mechanism is computationally expensive. Therefore, one novel selective state space model, Mamba, has been proposed as an alternative. Building on its success in automatic speech recognition, we apply Mamba for spoofing attack detection. Mamba is well-suited for this task as it can capture the artifacts in spoofed speech signals by handling long-length sequences. However, Mamba's performance may suffer when it is trained with limited labeled data. To mitigate this, we propose combining a new structure of Mamba based on a dual-column architecture with self-supervised learning, using the pre-trained wav2vec 2.0 model. The experiments show that our proposed approach achieves competitive results and faster inference on the ASVspoof 2021 LA and DF datasets, and on the more challenging In-the-Wild dataset, it emerges as the strongest candidate for spoofing attack detection. Yang Xiao 0019, Rohan Kumar Das |
IEEE Signal Process. Lett. | 1 |
| 2023 | Small Footprint Multi-channel Network for Keyword Spotting with Centroid Based Awareness
Dianwen Ng, Yang Xiao 0019, Jia Qi Yip, Biao Tian 0002, Qiang Fu 0001, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 2 |
| 2022 | Rainbow Keywords: Efficient Incremental Learning for Online Spoken Keyword SpottingabstractCatastrophic forgetting is a thorny challenge when updating keyword spotting (KWS) models after deployment.This problem will be more challenging if KWS models are further required for edge devices due to their limited memory.To alleviate such an issue, we propose a novel diversity-aware incremental learning method named Rainbow Keywords (RK).Specifically, the proposed RK approach introduces a diversity-aware sampler to select a diverse set from historical and incoming keywords by calculating classification uncertainty.As a result, the RK approach can incrementally learn new tasks without forgetting prior knowledge.Besides, the RK approach also proposes data augmentation and knowledge distillation loss function for efficient memory management on the edge device.Experimental results show that the proposed RK approach achieves 4.2% absolute improvement in terms of average accuracy over the best baseline on Google Speech Command dataset with less required memory.The scripts are available on GitHub 1 . Yang Xiao 0019, Nana Hou, Chng Eng Siong |
INTERSPEECH | 1 |