Zhihua Fang

dblp:332/6588 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-3018-7414ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MED-Net: Leveraging multi-visual encoding and Manhattan-distance pooling for robust medical image segmentation
Zhaozhao Su, Shiji Song, Zhiqiang Zhu, Zhihua Fang, Bowen Wang 0029
Expert Syst. Appl.6
2026 DS3: Dual-space sample selection for speaker representation learning with noisy labels
Zhihua Fang, Liang He 0003
Pattern Recognit.1
2025 Noise Supervised Contrastive Learning and Feature-Perturbed for Anomalous Sound Detection
abstract
Unsupervised anomalous sound detection aims to detect unknown anomalous sounds by training a model using only normal audio data. Despite advancements in self-supervised methods, the issue of frequent false alarms when handling samples of the same type from different machines remains unresolved. This paper introduces a novel training technique called one-stage supervised contrastive learning (OS-SCL), which significantly addresses this problem by perturbing features in the embedding space and employing a one-stage noisy supervised contrastive learning approach. On the DCASE 2020 Challenge Task 2, it achieved 94.64% AUC, 88.42% pAUC, and 89.24% mAUC using only Log-Mel features. Additionally, a time-frequency feature named TFgram is proposed, which is extracted from raw audio. This feature effectively captures critical information for anomalous sound detection, ultimately achieving 95.71% AUC, 90.23% pAUC, and 91.23% mAUC. The source code is available at: www.github.com/huangswt/OS-SCL.
Shun Huang, Zhihua Fang, Liang He 0003
ICASSP2
2025 Stable Extended U-Net for Noise-Robust Speaker Verification
abstract
With advancements in deep learning, speaker verification systems have significantly improved their performance in noisy environments. Researchers typically demonstrate the effectiveness of their improved models by comparing performance on specific datasets, such as the VoxCeleb benchmark. However, in diverse real-world noise conditions, the out-of-domain generalization ability is also a crucial factor in evaluating a model’s performance improvement. Research on stable learning indicates that eliminating the spurious correlation between training and testing data can enhance the generalization of the model. Building on this idea, we propose an improved speaker verification system with high generalization based on the extended U-Net (ExU-Net). It uses the sample reweighting method from stable learning to eliminate sample correlations and retains more effective speaker information through subpixel convolutions and coordinate attention mechanisms. We validate the effectiveness of this approach through extensive evaluations on VoxCeleb1, VOiCES, and other out-of-domain noise test sets, highlighting its generalization capability and model robustness.
Zonghui Wang, Zhihua Fang, Liang He 0003
ICASSP2
2025 Self-supervised Speaker Verification with Batch-scale Pseudo-labels Correction
abstract
Self-supervised learning has shown outstanding performance on speaker verification, and the 2-stage frameworks have more comprehensive training schemes, which typically exhibit better performance. They utilize clustering to obtain pseudo-labels, which are then used as the supervision signal in stage 2. However, these pseudo-labels often contain a significant amount of noisy labels, severely impacting speaker verification performance. In this paper, we propose a dynamic self-supervised pseudo-label correction method based on batch-scale training. By filtering and correcting samples based on the loss and prediction distribution, our method better aligns with the dynamic training process and achieves EER(%) of 1.33, 1.56 and 2.78 on the test sets of Voxceleb-O, E, H.
Junxu Wang, Zhihua Fang, Liang He 0003
ICASSP2
2025 Alignment Losses for End-to-End Speaker Diarization
Simeng Shi, Zhida Song, Zhihua Fang, Liang He 0003
ICIC (17)3
2025 AGDAformer: Agent-Guidance Dual Attention Transformer for Climate-Aware Crop Yield Prediction
abstract
As climate change becomes widespread, rapid, and intensifying, accurate crop yield predictions are crucial for formulating effective agricultural development strategies. However, crop yield prediction is fairly challenging, as it is increasingly influenced by the intensifying climate change. In this paper, we propose the Agent-Guidance Dual Attention Transformer (AGDAformer), which aims to enhance the accuracy of yield predictions for climate-sensitive crops by integrating spatial attention mechanisms with multi-scale meteorological data. On the one hand, the Cross-Spatial Attention (CSA) mechanism optimizes the acquisition of broad-scale meteorological information by adaptively identifying and focusing on deep features closely related to crop yield prediction, thus improving the effectiveness of feature selection. On the other hand, the Agent-Enhanced Transformer (AET) introduces an agent matrix to facilitate the fusion of multiscale meteorological data, embedding meteorological knowledge into Agent Attention to effectively integrate information from different scales and better capture complex climatic factors. Extensive experiments on the CropNet dataset demonstrate that AGDAformer achieves superior performance across four crop types, significantly improving prediction accuracy and providing robust support for future agricultural decision-making.
Zhongyan Yi, Zhihua Fang, Jintai Wang, Liang He 0003
IJCNN2
2025 Federated Learning with Feature Space Separation for Speaker Recognition
Zhihua Fang
INTERSPEECH2
2024 Multi-View Speaker Embedding Learning for Enhanced Stability and Discriminability
abstract
Deep neural network models based on x-vector have become the most popular framework for speaker recognition, and the quality of speaker features (embeddings) is important for open-set tasks such as speaker verification and speaker diarization. Currently, the most popular loss function is based on margin penalty, however, it only considers enlarging the inter-class distance while neglecting to reduce the intra-class feature differences. Therefore, we propose a multi-view learning approach that divides the training process into two views from the speaker embedding level. The classification view focuses on distinguishing the discriminability of different speakers, while the clustering view focuses on shrinking the feature boundaries of the same speaker, making intra-class differences smaller. The combined effect of the two perspectives achieves large inter-class distance and small intra-class distances, resulting in the extraction of more discriminative and stable speaker embeddings. We test the performance of the method on both speaker verification and speaker diarization tasks, and the results demonstrate the effectiveness of our approach.
Liang He 0003, Zhihua Fang, Zuoer Chen, Minqiang Xu
ICASSP2
2024 Self-Supervised Speaker Verification with Mini-Batch Prediction Correction
Junxu Wang, Zhihua Fang, Liang He 0003
INTERSPEECH2
2024 Improving Speaker Verification With Noise-Aware Label Ensembling and Sample Selection: Learning and Correcting Noisy Speaker Labels
abstract
Supervised deep learning has achieved tremendous success in speaker verification. However, deep speaker models tend to overfit noisy labels when they are present in the speaker datasets. To mitigate the detrimental effects of noisy labels, in this paper, we propose a novelLabel Ensembling and Sample Selectionframework. Firstly, we select labels with high confidence rankings as clean samples. Additionally, we use predictions from different epochs during training to smoothly correct the noisy labels. Our method does not require staged training and achieves integration of learning from noisy labels, selecting clean labels, and correcting noisy labels. A significant number of experimental results demonstrate the robustness of our method under noisy labels. Even when the training data contains 50% noisy labels, our method can mitigate an average of 86.54% of the performance degradation compared to the standard training method. Furthermore, further ablation experiments and analysis validate the effectiveness of High Confidence Ranking for sample selection and the correctness of Label Ensembling for noisy label correction.
Zhihua Fang, Liang He 0003, Lin Li 0032, Ying Hu 0005
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Robust Training for Speaker Verification against Noisy Labels
Zhihua Fang, Liang He 0003, Hanhan Ma, Lin Li 0032
INTERSPEECH1