VLDB 2026 Research / reviewers in the wild / expert
Yunfei Zi
dblp:232/9194
· DBLP profile ↗
13ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0002-4778-7109ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Location-Oriented Sound Event Localization and Detection with Spatial Mapping and Regression LocalizationabstractSound Event Localization and Detection (SELD) combines the Sound Event Detection (SED) with the corresponding Direction Of Arrival (DOA). Recently, adopted event-oriented multi-track methods affect the generality in polyphonic environments due to the limitation of the number of tracks. To enhance the generality in polyphonic environments, we propose Spatial Mapping and Regression Localization for SELD (SMRL-SELD). SMRL-SELD segments the 3D spatial space, mapping it to a 2D plane, and a new regression localization loss is proposed to help the results converge toward the location of the corresponding event. SMRL-SELD is location-oriented, allowing the model to learn event features based on orientation. Thus, the method enables the model to process polyphonic sounds regardless of the number of overlapping events. We conducted experiments on STARSS23 and STARSS22 datasets and our proposed SMRL-SELD outperforms the existing SELD methods in overall evaluation and polyphony environments. Xueping Zhang, Yaxiong Chen, Ruilin Yao, Yunfei Zi, Shengwu Xiong 0001 |
ICME | 4 |
| 2024 | Adaptive Learning via a Negative Selection Strategy for Few-Shot Bioacoustic Event DetectionabstractAlthough the Prototypical Network (ProtoNet) has demonstrated effectiveness in few-shot biological event detection, two persistent issues remain. Firstly, there is difficulty in constructing a representative negative prototype due to the absence of explicitly annotated negative samples. Secondly, the durations of the target biological vocalisations vary across tasks, making it challenging for the model to consistently yield optimal results across all tasks. To address these issues, we propose a novel adaptive learning framework with an adaptive learning loss to guide classifier updates. Additionally, we propose a negative selection strategy to construct a more representative negative prototype for ProtoNet. All experiments ware performed on the DCASE 2023 TASK5 few-shot bioacoustic event detection dataset. The results show that our proposed method achieves an F-measure of 0.703, an improvement of 12.84%. Yaxiong Chen, Xueping Zhang, Yunfei Zi, Shengwu Xiong 0001 |
ICME | 3 |
| 2024 | Fisher ratio-based multi-domain frame-level feature aggregation for short utterance speaker verificationabstractAs the durations of the short utterances are small, it is difficult to learn sufficient information to distinguish the person, thus, short utterance speaker recognition is highly challenging. In this paper, we propose a multi-domain frame-level feature joint learning method to aggregate the discriminative information from multiple dimensions and domain, which is different domains of the speech, time-domain, frequency-domain, and spectral-domain, represent distinct physical characteristics and provide different dimension information, the time domain captures information about the temporal aspect of the physical signal, the frequency domain represents the signal strength in different frequency ranges, and the spectral domain reflects the overall information of the speech, then, based on the extracted multi-domain frame-level features, using the Multi-Fisher criterion aggregates feature parameters categorically and match the corresponding Multi-Fisher ratio weights to the feature parameters as a way to achieve effective feature aggregation and to preserve more effective information, termed Firm-Domain. Extensive experiments are carried out on short-duration text-independent speaker verification datasets derived from the VoxCeleb, SITW, and NIST SRE corpora, which contain speech samples of varying lengths and scenarios. The results demonstrate that the proposed method outperforms the state-of-the-art deep learning architectures by at least 13%, respectively, in the test set. The results of the ablation experiments demonstrate that our proposed methods can significantly outperform previous approaches. Yunfei Zi, Shengwu Xiong 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Multi-Fisher and Triple-Domain Feature Enhancement-Based Short Utterance Speaker Verification for IoT Smart ServiceabstractSpeech authentication in IoT smart services typically involves short utterances. However, due to the short duration of these utterances (e.g., less than 3 s) and the limited enrolment and/or test data available, it is difficult to learn enough information to accurately distinguish the person. As a result, speaker recognition from short utterances is very challenging. In this article, in the acoustic end, we propose a Multi-Fisher feature enhancement method to enrich short utterance effective information. The method aggregates the effective feature parameters by assigning the corresponding Multi-Fisher ratio weights to these parameters. In the architecture end, we propose a triple-domain feature joint learning method to enhance discriminative information from multiple dimensions. This approach provides different dimensions of information through the different physical meanings of speech in the time domain, the frequency domain, and the spectral domain. Extensive experiments were conducted on VoxCeleb, SITW, and NIST SRE short-duration text-independent speaker verification tasks, containing speech samples of varying lengths and scenarios. The results show that our proposed method outperforms existing acoustic feature extraction approaches and state-of-the-art deep learning architectures by at least 6% and 12%, respectively. The ablation experiments further illustrate that our proposed approaches can achieve substantial improvement over previous methods. Yunfei Zi, Shengwu Xiong 0001 |
IEEE Internet Things J. | 1 |
| 2024 | Scale-Aware Adaptive Refinement and Cross-Interaction for Remote Sensing Audio-Visual Cross-Modal Retrieval
Yaxiong Chen, Chuang Du, Yunfei Zi, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Improving Speaker Recognition by Time-Frequency Domain Feature Enhanced Method
Yunfei Zi, Shengwu Xiong 0001 |
PRICAI (2) | 2 |
| 2023 | Joint filter combination-based central difference feature extraction and attention-enhanced Dense-Res2Block network for short-utterance speaker recognition
Yunfei Zi, Shengwu Xiong 0001 |
Expert Syst. Appl. | 1 |
| 2023 | Exploration of multi-source discriminative acoustic feature for speaker recognition with short-duration audio signal
Yunfei Zi, Shengwu Xiong 0001 |
Multim. Tools Appl. | 1 |
| 2023 | BSML: Bidirectional Sampling Aggregation-based Metric Learning for Low-resource Uyghur Few-shot Speaker VerificationabstractIn recent years, text-independent speaker verification has remained a hot research topic, especially for the limited enrollment and/or test data. At the same time, due to the lack of sufficient training data, the study of low-resource few-shot speaker verification makes the models prone to overfitting and low accuracy of recognition. Therefore, a bidirectional sampling aggregation-based meta-metric learning method is proposed to solve the low-accuracy problem of speaker recognition in a low-resource environment with limited data, termed bidirectional sampling multi-scale Fisher feature fusion (BSML). First, the BSML method was used for effective feature enhancement in the feature extraction stage; second, a large number of similar and disjoint tasks were used to train the models to learn how to compare sample similarity; finally, new tasks were used to identify unknown samples by calculating the similarity of the samples. Extensive experiments are conducted on a short-duration text-independent speaker verification dataset generated from the THUYG-20 low-resource Uyghur with limited data, which comprised speech samples of diverse lengths. The experimental result has shown that the metric learning approach is effective in avoiding model overfitting and improving model generalization, with significant results in the identification of short-duration speaker verification in low-resource Uyghur with few-shot. It also demonstrates that BSML outperforms the state-of-the-art deep-embedding speaker recognition architectures and recent metric learning approach by at least 18%–67% in the few-shot test set. The ablation experiments further illustrate that our proposed approaches can achieve substantial improvement over prior methods and achieves better performance and generalization ability. Yunfei Zi, Shengwu Xiong 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2022 | A Frame Loss of Multiple Instance Learning for Weakly Supervised Sound Event DetectionabstractSound event detection(SED) consists of two subtasks: predicting the classes of sound events within an audio clip (audio tagging) and indicating the onset and offset times for each event (localization). One of the common approaches for SED with weak label is multiple instance learning (MIL) method. However, the general MIL method only optimizes the global loss calculated from the aggregated clip-wise predictions and weak clip labels, lacking a direct constraint on the frame-wise predictions, which leads to a large number of unreasonable prediction values. To address this issue, we explore the deterministic information that can be used to constrain the framewise predictions and based on which we design a frame loss with two terms. Experimental results on the DCASE2017 Task4 dataset demonstrate that the proposed loss can improve the performance of general MIL method. While this article focuses on SED applications, the proposed methods could be applied widely to MIL problems. Code will be available at WSSED. Xiangjinzi Zhang, Yunfei Zi, Shengwu Xiong 0001 |
ICASSP | 3 |
| 2022 | Fusing Acoustic and Text Emotional Features for Expressive Speech SynthesisabstractProminent methods based on Tacotron2 and advanced models have improved the quality of synthesized speech. However, most data-driven Text- To-Speech (TTS) synthesis methods only aim to achieve reasonable neutral prosody, so the synthesized speech is less expressive. In this paper, a method was proposed which fuses acoustic and text emotional features to produce more vivid and realistic speech. Specifically, to obtain acoustic features, two acoustic encoders are leveraged to extract utterance-level and phoneme-level vectors from the target speech, respectively. To obtain the objective sentiment features of the text, the sentiment analysis model is exploited to extract the sentiment vector from the text and expand it. The expanded vector is feature- fused with the output vector of the acoustic model. The experimental results on the LJSpeech dataset show that the naturalness and expressiveness of the MOS score are 3.63 and 3.45, respectively, and the similarity of the SMOS score is 4.14. Pengfei Duan 0005, Yunfei Zi, Yaxiong Chen, Shengwu Xiong 0001 |
ICME | 3 |
| 2021 | Variational Information Bottleneck Based Regularization for Speaker Recognition
Yuanjie Dong, Yaxing Li, Yunfei Zi, Xiaoqi Li 0011, Shengwu Xiong 0001 |
Interspeech | 4 |
| 2018 | Research of Personalized Recommendation System Based on Multi-view Deep Neural Networks
Yunfei Zi, Yeli Li, Huayan Sun |
ADMA | 1 |