EDBT 2026 Demo / reviewers in the wild / expert
Xueshuai Zhang
dblp:273/0525
· DBLP profile ↗
13ranked-venue papers
2as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Data-Efficient Semi-Supervised Few-Shot Speaker Verification via Prototype Space OptimizationabstractSpeaker verification technology has widespread applications across many domains, benefiting from deep learning advancements. However, due to the high cost of acquiring labeled data, semi-supervised learning has emerged as a prominent research focus. Current semi-supervised learning frameworks commonly suffer from two limitations: (1) the labeled data distribution is often restricted, and (2) they still rely on a considerable amount of labeled data. To address these issues, we propose three different distribution scenarios of labeled data and construct a general semi-supervised framework. Furthermore, to enhance the guidance efficacy of limited labeled data, we innovatively employ prototype space optimization to strengthen the model's discriminative capability under low-resource scenarios. Experimental results demonstrate that on the Vox1-o test set, our approach achieves a 41.7% relative reduction in equal error rate compared to self-supervised baselines, and a 28.9% improvement over conventional semi-supervised framework baselines. Zhenduo Zhao, Shisong Wu, Pengyuan Zhang, Xueshuai Zhang, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 6 |
| 2026 | SIM-IDE-FFM: Similarity-Based Residual-Identity Feature Fusion for Speaker VerificationabstractRecent advances in speaker verification have focused on two primary areas: deepening or widening backbone networks and introducing different attention mechanisms in the main residual branch to enhance discriminative power. However, the shortcut branches in residual architectures are still treated as simple identity mappings, leaving their representational potential underexplored. In this paper, we propose SIM-IDE-FFM, a similarity-based identity feature enhancement framework which enables the trivial shortcut branch to explore correlations between identity feature maps through a similarity-based attention mechanism. SIM-IDE introduces an attention mechanism at the shortcut branch that models inter-channel similarity of identity inputs, enhancing channel-wise representations. A non-linear Feature Fusion Module (FFM) is further designed to efficiently combine original and enhanced features through nonlinear depthwise interactions. Without altering the backbones, SIM-IDE-FFM can be seamlessly integrated into ResNet, EcapaTDNN, and other state-of-the-art (SOTA) architectures. Experiments on VoxCeleb1 and cross-age subsets demonstrate consistent relative improvements across different frameworks, validating its robustness and generalization. Baizhu Li, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Debiased Training For Semi-supervised Sound Event DetectionabstractRecently, semi-supervised sound event detection has attracted increasing research interest due to the scarcity of labeled data. However, traditional semi-supervised learning methods can lead to training instability and confirmation bias because of potentially incorrect pseudo labels. To address this issue, we propose the debiased training, a novel approach to reduce the inherent bias of pseudo labels. Debiased training can effectively decouple the generation and utilization of pseudo labels to mitigate the error accumulation and promote model’s robustness against biased pseudo labels. In addition, we introduce the channel restruction module (CRM) to decrease redundant computing and facilitate representation ability. Experimental results on DCASE 2023 task4 dataset show that the proposed methods significantly enhance the performance of semi-supervised methods while maintaining relatively low computational complexity. Shengchang Xiao, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 2 |
| 2025 | A comprehensive validation study on the influencing factors of cough-based COVID-19 detection through multi-center data with abundant metadata
Jiakun Shen, Xueshuai Zhang, Yanfen Tang, Pengyuan Zhang, Yonghong Yan 0002, Pengfei Ye, Shaoxing Zhang |
J. Biomed. Informatics | 2 |
| 2025 | Multi-Branch Coordinate Attention With Channel Dynamic Difference for Speaker VerificationabstractIn prior studies, researchers have proved the excellent performance of various deep neural networks on speaker verification (SV). However, most of the improvements of SV systems are aimed at modifying the specific network structure to enhance its robustness but with limited flexibility. In this paper, MCA-CDD, which is a novel universal residual block module and can easily replace the original residual block without increasing the block number, is proposed with adaptive multi-branch coordinate attention (MCA) and channel-level dynamic difference (CDD). The design of multiple branches enables CA to model time and frequency at multiple scales. In addition, CDD fusion is applied into the feature fusion process of the residual block. The CDD fusion at shallow positions of the model enables the model to learn the detailed speaker-related dynamic texture information of speech. Experiments are conducted on several ResNet backbones and the results on different seen and unseen test sets show significant improvements, outperforming the baseline by about relatively 15-20%. Zhenduo Zhao, Shisong Wu, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Fetal Heart Sounds Classification Using Time-Cyclic Frequency Spectrogram and Hybrid Attention NetworkabstractFetal heart monitoring is a crucial method for assessing fetal health status. However, commonly used techniques may pose potential medical risks and are not suitable for long-term monitoring. In this paper, we propose utilizing fetal heart sounds (FHS) for classifying fetal health status, taking advantage of its non-invasive, safe, straightforward, and cost-effective properties. Firstly, we introduce a novel acoustic feature for fetal heart sounds, termed the time-cyclic frequency spectrogram. This feature emphasizes the periodicity of heartbeats and effectively captures the changes in fetal heart rate. Additionally, we implement a frequency band energy-weighted algorithm to mitigate interference from periodic noises. Secondly, we propose a hybrid attention network that integrates both global-local attention and time-cyclic frequency attention. This network leverages medical prior knowledge to focus on the most critical aspects of the time-cyclic frequency spectrogram. Experimental results demonstrate that the proposed feature effectively characterizes variations in fetal heart rate, and the hybrid attention network can accurately capture spectral line variations, leading to improved classification of fetal health conditions. Shengchang Xiao, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002 |
BIBM | 2 |
| 2024 | Snore Sound Features Based on Percussive Enhancing and Positional Encoding Combined with Multi-Task Learning for Osahs DetectionabstractObstructive sleep apnea hypopnea syndrome (OSAHS) is a serious sleep disorder. As the typical symptom of OSAHS, snoring has been proved effective in OSAHS diagnosis and potential to replace the current laborious and expensive polysomnography. However, the lack of analysis on the characteristics of pathological snoring sounds limits the diagnosing performance. In this paper, we propose novel sound features for the classification of OSA, hypopnea and normal snores. The proposed features are based on percussive enhancing and positional encoding as the snores exhibit different percussive properties and temporal traits due to the disease generation mechanisms. To enhance the classification performance, we propose a multi-task learning framework to aid the main classification task by simultaneous learning of two related simple tasks. Experiments on real-recorded snoring sounds show that the proposed methods can greatly improve the classification AUC and ACC and the proposed system performs better than those in other literatures. Aolin Hu, Xueshuai Zhang, Shaoxing Zhang, Pengyuan Zhang, Pengfei Ye, Qingwei Zhao, Yonghong Yan 0002 |
ICASSP | 2 |
| 2024 | One-Epoch Training with Single Test Sample in Test Time for Better Generalization of Cough-Based Covid-19 Detection ModelabstractThe outbreak of COVID-19 has raised researchers’ attention to audio-based rapid disease detection. Most of the previous studies have obtained competitive detection performance. However, these results are usually obtained by testing data from the same source offline. When making cross-dataset testing, the performance may deteriorate dramatically due to inconsistent data distribution between different datasets. In addition, in practical application, the model has to make a prediction for the current test audio without any prior information, which requires good model generalization under limited training data. To address the above issues, we adopt a test-time training framework to achieve a cough-based COVID-19 detection model with better generalizability. In the model development stage, resnet18 serves as the backbone network and a self-supervised learning branch is added as an auxiliary task. In testing stage, the model parameters are first fine-tuned by the self-supervised branch with the single test audio as input, and then the classification head outputs predictions. The proposed method is validated on three open-source datasets using a variety of hyperparameters. In cross-dataset testing, AUC and UAR increase by 3.65% and 3% on average absolutely, respectively. The results show that the proposed framework is applicable to improve the model performance in practical application. Jiakun Shen, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002, Qingwei Zhao, Ta Li, Yanfen Tang, Shaoxing Zhang |
ICASSP | 2 |
| 2024 | Prototype Division for Self-Supervised Speaker VerificationabstractSelf-supervised learning has shown promising performance on speaker verification tasks, among which Self DIstillation with NO labels (DINO) is currently a widely adopted framework. As one of the unsupervised deep clustering methods, the number of valid prototypes in DINO is far less than the speakers in practical applications and remains unchanged throughout the training period, leading to severe speaker confusion and performance degradation. Therefore, a strategy named prototype division (PD) is proposed to iteratively generate fine-grained prototypes in the projection space based on the converged model to separate confused categories, where new prototypes are derived from the neighborhood of the existing valid prototypes by clustering or sampling. The results on Vox1O achieve significant improvements, relatively outperforming the baseline by 31.1% without any auxiliary loss. Further experiments on CN-Celeb also show stable improvement, proving the consistency of the proposed method. Zhenduo Zhao, Zhuo Li 0020, Xueshuai Zhang, Pengyuan Zhang |
IEEE Signal Process. Lett. | 3 |
| 2023 | Piecewise Position Encoding in Convolutional Neural Network for Cough-Based Covid-19 DetectionabstractA fast and efficient COVID-19 detection method is of vital importance to control the spread of the epidemic. Many studies have achieved good performance on cough-based COVID19 detection in the past two years. However, the effect of position information in time-frequency features of cough audio has been less considered in previous studies. Even the convolutional neural networks that are capable to learn position information may be affected by small transformations of input features. Therefore, we propose piecewise position encoding added to time-frequency features to provide supplementary position information explicitly. Considering the differences in recording devices among different people, we use modified instance normalization to achieve better generalization. The proposed methods are validated on three open-sourced datasets and achieve significant improvements in AUC and UAR. The proposed model also shows competitive results in detecting asymptomatic patients. Jiakun Shen, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002, Shaoxing Zhang, Yanfen Tang, Fujie Zhang, Aijun Sun |
ICASSP | 2 |
| 2023 | Multi-Dimensional Frequency Dynamic Convolution with Confident Mean Teacher for Sound Event DetectionabstractRecently, convolutional neural networks (CNNs) have been widely used in sound event detection (SED). However, traditional convolution is deficient in learning time-frequency domain representation of different sound events. To address this issue, we propose multi-dimensional frequency dynamic convolution (MFDConv), a new design that endows convolutional kernels with frequency-adaptive dynamic properties along multiple dimensions. MFDConv utilizes a novel multi-dimensional attention mechanism with a parallel strategy to learn complementary frequency-adaptive attentions, which substantially strengthen the feature extraction ability of convolutional kernels. Moreover, in order to promote the performance of mean teacher, we propose the confident mean teacher to increase the accuracy of pseudo-labels from the teacher and train the student with high confidence labels. Experimental results show that the proposed methods achieve 0.470 and 0.692 of PSDS1 and PSDS2 on the DESED real validation dataset. Shengchang Xiao, Xueshuai Zhang, Pengyuan Zhang |
ICASSP | 2 |
| 2022 | Robust Cough Feature Extraction and Classification Method for COVID-19 Cough Detection Based on Vocalization CharacteristicsabstractA fast, efficient and accurate detection method of COVID-19 remains a critical challenge.Many cough-based COVID-19 detection researches have shown competitive results through artificial intelligence.However, the lack of analysis on vocalization characteristics of cough sounds limits the further improvement of detection performance.In this paper, we propose two novel acoustic features of cough sounds and a convolutional neural network structure for COVID-19 detection.First, a time-frequency differential feature is proposed to characterize dynamic information of cough sounds in time and frequency domain.Then, an energy ratio feature is proposed to calculate the energy difference caused by the phonation characteristics in different cough phases.Finally, a convolutional neural network with two parallel branches which is pre-trained on a large amount of unlabeled cough data is proposed for classification.Experiment results show that our proposed method achieves state-of-the-art performance on Coswara dataset for COVID-19 detection.The results on an external clinical dataset Virufy also show the better generalization ability of our proposed method. Xueshuai Zhang, Jiakun Shen, Jun Zhou 0024, Pengyuan Zhang, Yonghong Yan 0002, Yanfen Tang, Fujie Zhang, Shaoxing Zhang, Aijun Sun |
INTERSPEECH | 1 |
| 2020 | Speaker Diarization System Based on DPCA Algorithm for Fearless Steps Challenge Phase-2
Xueshuai Zhang, Pengyuan Zhang |
INTERSPEECH | 1 |