VLDB 2026 Research / reviewers in the wild / expert
Yafeng Chen
dblp:220/8657
· DBLP profile ↗
19ranked-venue papers
8as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 13 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language ModelsabstractThe Speaker Diarization and Recognition (SDR) task aims to predict ``who spoke when and what'' within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and dialogue systems. Existing SDR systems typically adopt a cascaded framework, combining multiple modules such as speaker diarization (SD) and automatic speech recognition (ASR). The cascaded systems suffer from several limitations, such as error propagation, difficulty in handling overlapping speech, and lack of joint optimization for exploring the synergy between SD and ASR tasks. To address these limitations, we introduce SpeakerLM, a unified multimodal large language model for SDR that jointly performs SD and ASR in an end-to-end manner. Moreover, to facilitate diverse real-world scenarios, we incorporate a flexible speaker registration mechanism into SpeakerLM, enabling SDR under different speaker registration settings. SpeakerLM is progressively developed with a multi-stage training strategy on large-scale real data. Extensive experiments show that SpeakerLM demonstrates strong data scaling capability and generalizability, outperforming state-of-the-art cascaded baselines on both in-domain and out-of-domain public SDR benchmarks. Furthermore, experimental results show that the proposed speaker registration mechanism effectively ensures robust SDR performance of SpeakerLM across diverse speaker registration conditions and varying numbers of registered speakers. Han Yin, Yafeng Chen, Chong Deng, Luyao Cheng, Hui Wang 0030, Chao-Hong Tan, Qian Chen 0003, Wen Wang 0001, Xiangang Li |
AAAI | 2 |
| 2025 | Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization on Multi-party ConversationabstractLuyao Cheng, Hui Wang, Chong Deng, Siqi Zheng, Yafeng Chen, Rongjie Huang, Qinglin Zhang, Qian Chen, Xihao Li, Wen Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Luyao Cheng, Hui Wang 0030, Chong Deng, Yafeng Chen, Rongjie Huang 0001, Qian Chen 0003, Xihao Li, Wen Wang 0001 |
ACL (1) | 5 |
| 2025 | Self-Distillation Prototypes Network: Learning Robust Speaker Representations without SupervisionabstractTraining speaker-discriminative and robust speaker verification systems without explicit speaker labels remains a persistent challenge. In this paper, we propose a novel self-supervised speaker verification approach, Self-Distillation Prototypes Network (SDPN), which effectively facilitates self-supervised speaker representation learning. SDPN assigns the representation of the augmented views of an utterance to the same prototypes as the representation of the original view, thereby enabling effective knowledge transfer between the augmented and original views. Due to lack of negative pairs in the SDPN training process, the network tends to align positive pairs quite closely in the embedding space, a phenomenon known as model collapse. To mitigate this problem, we introduce a diversity regularization term to embeddings in SDPN. Comprehensive experiments on the VoxCeleb datasets demonstrate the superiority of SDPN among self-supervised speaker verification approaches. SDPN sets a new state-of-the-art on the VoxCeleb1 speaker verification evaluation benchmark, achieving Equal Error Rate 1.80%, 1.99%, and 3.62% for trial VoxCeleb1-O, VoxCeleb1-E and VoxCeleb1H respectively1, without using any speaker labels in training. Ablation studies show that both proposed learnable prototypes in self-distillation network and diversity regularization contribute to the verification performance. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003, Chong Deng, Shiliang Zhang, Wen Wang 0001 |
ICASSP | 1 |
| 2025 | 3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and DiarizationabstractWe introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the combined strengths of acoustic, semantic, and visual data, seamlessly fusing these modalities to offer robust speaker recognition capabilities. The acoustic module extracts speaker embeddings from acoustic features, employing both fully-supervised and self-supervised learning approaches. The semantic module leverages advanced language models to comprehend the substance and context of spoken language, thereby augmenting the system’s proficiency in distinguishing speakers through linguistic patterns. The visual module applies image processing technologies to scrutinize facial features, which bolsters the precision of speaker diarization in multi-speaker environments. Collectively, these modules empower the 3D-Speaker-Toolkit to achieve substantially improved accuracy and reliability in speaker-related tasks. With 3D-Speaker-Toolkit, we establish a new benchmark for multimodal speaker analysis. The toolkit also includes a handful of open-source state-of-the-art models and a large-scale dataset containing over 10,000 speakers. The toolkit is publicly available at https://github.com/modelscope/3D-Speaker. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Tinglong Zhu, Rongjie Huang 0001, Chong Deng, Qian Chen 0003, Shiliang Zhang, Wen Wang 0001, Xihao Li |
ICASSP | 1 |
| 2025 | Exploring Text-Queried Sound Event Detection with Audio Source SeparationabstractIn sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we first pre-train a language-queried audio source separation (LASS) model to separate the audio tracks corresponding to different events from the input audio. Then, multiple target SED branches are employed to detect individual events. AudioSep is a state-of-the-art LASS model, but has limitations in extracting dynamic audio information because of its pure convolutional structure for separation. To address this, we integrate a dual-path recurrent neural network block into the model. We refer to this structure as AudioSep-DP, which achieves the first place in DCASE 2024 Task 9 on language-queried audio source separation (objective single model track). Experimental results show that TQ-SED can significantly improve the SED performance, with an improvement of 7.22% on F1 score over the conventional framework. Additionally, we setup comprehensive experiments to explore the impact of model complexity. The source code and pre-trained model are released at https://github.com/apple-yinhan/TQ-SED. Han Yin, Jisheng Bai, Yang Xiao 0019, Hui Wang 0030, Yafeng Chen, Rohan Kumar Das, Chong Deng |
ICASSP | 6 |
| 2025 | Pushing the Frontiers of Self-Distillation Prototypes Network with Dimension Regularization and Score Normalization
Yafeng Chen, Chong Deng, Hui Wang 0030, Yiheng Jiang, Han Yin, Qian Chen 0003, Wen Wang 0001 |
INTERSPEECH | 1 |
| 2025 | Exploring Efficient Directional and Distance Cues for Regional Speech Separation
Yiheng Jiang, Haoxu Wang, Yafeng Chen, Gang Qiao |
INTERSPEECH | 3 |
| 2025 | Speech Token Prediction via Compressed-to-fine Language Modeling for Speech GenerationabstractNeural audio codecs, used as speech tokenizers, have demonstrated remarkable potential in the field of speech generation. However, to ensure high-fidelity audio reconstruction, neural audio codecs typically encode audio into long sequences of speech tokens, posing a significant challenge for downstream language models in long-context modeling. We observe that speech token sequences exhibit short-range dependency: due to the monotonic alignment between text and speech in text-to-speech (TTS) tasks, the prediction of the current token primarily relies on its local context, while long-range tokens contribute less to the current token prediction and often contain redundant information. Inspired by this observation, we propose a compressed-to-fine language modeling approach to address the challenge of long sequence speech tokens within neural codec language models: (1) Fine-grained Initial and Short-range Information: Our approach retains the prompt and local tokens during prediction to ensure text alignment and the integrity of paralinguistic information; (2) Compressed Long-range Context: Our approach compresses long-range token spans into compact representations to reduce redundant information while preserving essential semantics. Extensive experiments on various neural audio codecs and downstream language models validate the effectiveness and generalizability of the proposed approach, highlighting the importance of token compression in improving speech generation within neural codec language models. The demo of audio samples will be available at https://anonymous.4open.science/r/SpeechTokenPredictionViaCompressedToFinedLM. Wenrui Liu 0003, Qian Chen 0003, Wen Wang 0019, Guanrou Yang, Minghui Fang 0002, Jialong Zuo, Xiaoda Yang, Tao Jin 0004, Jin Xu 0010, Yafeng Chen, Jionghao Bai, Zhifang Guo |
ACM Multimedia | 12 |
| 2024 | ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003, Shiliang Zhang |
INTERSPEECH | 1 |
| 2024 | Efficient Verifiable Cloud-Assisted PSI Cardinality for Privacy-Preserving Contact TracingabstractPrivate set intersection cardinality (PSI-CA) allows two parties to learn the size of the intersection between two private sets without revealing other additional information, which is a promising technique to solve privacy concerns in contact tracing. Efficient PSI protocols typically use oblivious transfer, involving multiple rounds of interaction and leading to heavy local computation overheads and protocol delays, especially when interacting with many receivers. Cloud-assisted PSI-CA is a better solution as it relieves participants' burdens of computation and communication. However, cloud servers may return incorrect or incomplete results for some reason, leading to an incorrectness issue. At present, to our knowledge, existing cloud-assisted PSI-CA protocols cannot address such a concern. To address this, we propose two specific verifiable cloud-assisted PSI-CA protocols: one based on a two-server protocol and the other on a single-server protocol. Further, we employ Cuckoo hashing to optimize these two protocols, enabling the receiver's computational costs independent of the size of the sender's set. We also prove the security of the protocols and implement them. Finally, we analyze and discuss their performance demonstrating that the single-server verifiable PSI-CA protocol does not introduce significant computation or communication costs while adding functionalities. Yafeng Chen, Axin Wu, Yuer Yang, Xiangjun Xin 0002 |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | Pushing the Limits of Self-Supervised Speaker Verification using Regularized Distillation FrameworkabstractTraining robust speaker verification systems without speaker labels has long been a challenging task. Previous studies observed a large performance gap between self-supervised and fully supervised methods. In this paper, we apply a non-contrastive self-supervised learning framework called DIstillation with NO labels (DINO) and propose two regularization terms applied to embeddings in DINO. One regularization term guarantees the diversity of the embeddings, while the other regularization term decorrelates the variables of each embedding. The effectiveness of various data augmentation techniques are explored, on both time and frequency domain. A range of experiments conducted on the VoxCeleb datasets demonstrate the superiority of the regularized DINO framework in speaker verification. Our method achieves the stateof-the-art speaker verification performance under a singlestage self-supervised setting on VoxCeleb. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003 |
ICASSP | 1 |
| 2023 | An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification
Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003, Jiajun Qi |
INTERSPEECH | 1 |
| 2023 | CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking
Hui Wang 0030, Yafeng Chen, Luyao Cheng, Qian Chen 0003 |
INTERSPEECH | 3 |
| 2023 | Detecting and Removing Phase Jitters for the Phase Synchronization of LT-1 Bistatic SARabstractPhase synchronization plays a crucial role in the LuTan-1 (LT-1) bistatic synthetic aperture radar (BiSAR) system, as it aims to eliminate additional azimuthal phase modulation caused by oscillator differences. However, for the pulse alternating transmission system operating in the L-band, the presence of radio frequency interference (RFI) poses an inevitable challenge. Serious RFIs introduce phase jitters, compromising the accuracy of synchronization. In this letter, an effective method for detecting and removing phase jitters is proposed to enhance synchronization accuracy. In the proposed method, the jitter features are separated by the iteratively reweighted least squares (IRLS) in the instantaneous frequency domain based on the established synchronization phase model. Then the jitter positions are detected by correlated peaks between the designed convolutional kernels and jitters. Finally, a polynomial model is utilized to remove jitters and assist in phase unwrapping. The synchronization phases acquired by the LT-1 mission are used to verify the feasibility of the proposed algorithm. The improved imaging quality demonstrates the effectiveness of the proposed method and confirms its ability to ensure the high-precision generation of the LT-1 BiSAR images. Yonghua Cai, Yachao Wang, Qingyue Yang, Yanyan Zhang 0002, Yafeng Chen, Pingping Lu, Robert Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Graph Convolutional Network Based Semi-Supervised Learning on Multi-Speaker Meeting DataabstractUnsupervised clustering on speakers is becoming increasingly important for its potential uses in semi-supervised learning. In reality, we are often presented with enormous amounts of unlabeled data from multi-party meetings and discussions. An effective unsupervised clustering approach would allow us to significantly increase the amount of training data without additional costs for annotations. Recently, methods based on graph convolutional networks (GCN) have received growing attention for unsupervised clustering, as these methods exploit the connectivity patterns between nodes to improve learning performance. In this work, we present a GCN-based approach for semi-supervised learning. Given a pre-trained embedding extractor, a graph convolutional network is trained on the labeled data and clusters unlabeled data with "pseudo-labels". We present a self-correcting training mechanism that iteratively runs the cluster-train-correct process on pseudo-labels. We show that this proposed approach effectively uses unlabeled data and improves speaker recognition accuracy. Fuchuan Tong, Yafeng Chen, Hongbin Suo, Qingyang Hong, Lin Li 0032 |
ICASSP | 4 |
| 2021 | Improved Meta-Learning Training for Speaker VerificationabstractMeta-learning (ML) has recently become a research hotspot in speaker verification (SV).We introduce two methods to improve the meta-learning training for SV in this paper.For the first method, a backbone embedding network is first jointly trained with the conventional cross entropy loss and prototypical networks (PN) loss.Then, inspired by speaker adaptive training in speech recognition, additional transformation coefficients are trained with only the PN loss.The transformation coefficients are used to modify the original backbone embedding network in the x-vector extraction process.Furthermore, the random erasing (RE) data augmentation technique is applied to all support samples in each episode to construct positive pairs, and a contrastive loss between the augmented and the original support samples is added to the objective in model training.Experiments are carried out on the Speaker in the Wild (SITW) and VOiCES databases.Both of the methods can obtain consistent improvements over existing meta-learning training frameworks.By combining these two methods, we can observe further improvements on these two databases. Yafeng Chen, Wu Guo, Bin Gu 0004 |
Interspeech | 1 |
| 2021 | The Processing Framework and Experimental Verification for the Noninterrupted Synchronization Scheme of LuTan-1abstractThe bistatic synthetic aperture radar (BiSAR) plays an important role in remote sensing. However, the deviation between the two oscillators in BiSAR systems will cause a residual modulation of the echo signal. Therefore, the phase synchronization is an important issue that must be addressed in the BiSAR system. An advanced noninterrupted phase synchronization scheme is used for LuTan-1. The synchronization pulses are exchanged immediately after the ending time of the radar echo receiving window and before the starting time of the next pulse repetition interval, which will not interrupt the normal SAR operation. In order to evaluate the accuracy of the phase synchronization scheme, the model of phase synchronization is introduced at first. The hardware design and processing flow of LuTan-1 are introduced in detail. An innovative internal calibration strategy is also described. Then, the test data acquired by the ground validation system are analyzed to verify the effectiveness of the phase synchronization scheme. The signal-to-noise ratio (SNR) and the synchronization rate are the two most important factors to influence the accuracy in phase synchronization. The conclusions have guiding significance for the synchronization module design of LuTan-1 and the future BiSAR system. Da Liang, Kaiyu Liu, Heng Zhang 0007, Yafeng Chen, Haixia Yue, Dacheng Liu, Yunkai Deng, Haoyu Lin, Tingzhu Fang, Chuang Li 0001, Robert Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | A High-Accuracy Synchronization Phase-Compensation Method Based on Kalman Filter for Bistatic Synthetic Aperture RadarabstractPhase synchronization is one of the key issues that must be addressed for the bistatic synthetic aperture radar (BiSAR) system. LuTan-1 (LT-1) is an innovative spaceborne BiSAR mission based on the use of two radar satellites operating in the L-band to generate the global digital terrain models in the bistatic interferometry mode. An advanced synchronization scheme is used for the LT-1 system. The synchronization pulses are exchanged immediately after the ending time of the radar echo-receiving window and before the starting time of the next pulse-repetition interval. Therefore, it cannot interrupt the normal SAR data acquisition, further improving the synchronization accuracy and avoiding the data missing effect. In this letter, a robust phase-error estimation and compensation method is proposed to improve the accuracy of the synchronization by using the Kalman filter. The test data acquired from the ground validation system of the LT-1 synchronization module are used to demonstrate the feasibility of the proposed scheme. The results validate the effectiveness of the proposed scheme and prove the promise for its future application in LT-1. Da Liang, Kaiyu Liu, Heng Zhang 0007, Yunkai Deng, Dacheng Liu, Yafeng Chen, Chuang Li 0001, Haixia Yue, Robert Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2019 | An Advanced Non-Interrupted Synchronization Scheme for Bistatic Synthetic Aperture RadarabstractThe phase synchronization is one of the key issues to be addressed for the bistatic synthetic aperture radar system. In this paper, an advanced non-interrupted phase synchronization scheme is proposed. Both satellites are equipped with four synchronization antennas for a mutual exchange of synchronization pulses, which are transmitted rightly after the radar signal transmitting and before the echo receiving. Therefore, it can not interrupt the normal SAR data acquisition, which can further improve synchronization accuracy and avoid the data missing effect. The ground validation system for TwinSAR-L synchronization module is described in detail. The results are also evaluated to demonstrate feasibility of the proposed scheme. Da Liang, Kaiyu Liu, Haixia Yue, Yafeng Chen, Yunkai Deng, Heng Zhang 0007, Chuang Li 0001, Guodong Jin, Robert Wang 0001 |
IGARSS | 4 |