VLDB 2026 Research / reviewers in the wild / expert
Bing Han 0008
dblp:74/2721-8
· DBLP profile ↗
26ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0002-6319-6755ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 8 first-author · 21 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-trainingabstractYifan Yang, Bing Han, Hui Wang, Wei Wang, Ziyang Ma, Long Zhou, Zengrui Jin, Guanrou Yang, Tianrui Wang, Xu Tan, Xie Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yifan Yang 0005, Bing Han 0008, Hui Wang 0075, Wei Wang 0010, Ziyang Ma 0001, Zengrui Jin, Guanrou Yang, Tianrui Wang, Xu Tan 0003, Xie Chen 0001 |
ACL (1) | 2 |
| 2025 | Autoregressive Speech Synthesis without Vector QuantizationabstractLingwei Meng, Long Zhou, Shujie Liu, Sanyuan Chen, Bing Han, Shujie Hu, Yanqing Liu, Jinyu Li, Sheng Zhao, Xixin Wu, Helen M. Meng, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lingwei Meng, Shujie Liu 0001, Sanyuan Chen, Bing Han 0008, Shujie Hu, Jinyu Li 0001, Sheng Zhao 0002, Xixin Wu, Helen M. Meng, Furu Wei |
ACL (1) | 5 |
| 2025 | StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis
Kaicheng Luo, Xuefei Gong, Yutao Sun, Jinling He, Yujie Hou, Xiaoyang Xing, Huiyan Li, Bing Han 0008, Yanmin Qian |
ASRU | 8 |
| 2025 | Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching for Speaker DiarizationabstractSpeaker diarization is typically considered as a discriminative task, using discriminative approaches to produce fixed diarization results. In this paper, we explore for the first time the use of neural network-based generative methods for speaker diarization. We implement a Flow-Matching (FM) based generative algorithm within the sequenceto-sequence target speaker voice activity detection (Seq2Seq-TSVAD) diarization system. Our experiments reveal that applying the generative method directly to the original binary label sequence space of the TS-VAD output is ineffective. To address this issue, we propose mapping the binary label sequence into a dense latent space before applying the generative algorithm, and our proposed Flow-TSVAD method can significantly outperform the traditional Seq2Seq-TSVAD system. Additionally, we observe that the FM algorithm converges rapidly during the inference stage, only requiring two inference steps to achieve promising results. Moreover, as a generative model, Flow-TSVAD allows for sampling different diarization results by running the model multiple times, so the ensemble system combining the results from various sampling instances can further boost the diarization performance. Zhengyang Chen, Bing Han 0008, Shuai Wang 0016, Yidi Jiang, Yanmin Qian |
ICASSP | 2 |
| 2025 | Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive PruningabstractThe goal of the acoustic scene classification (ASC) task is to classify recordings into one of the predefined acoustic scene classes. However, in real-world scenarios, ASC systems often encounter challenges such as recording device mismatch, low-complexity constraints, and the limited availability of labeled data. To alleviate these issues, in this paper, a data-efficient and low-complexity ASC system is built with a new model architecture and better training strategies. Specifically, we firstly design a new low-complexity architecture named Rep-Mobile by integrating multi-convolution branches which can be reparameterized at inference. Compared to other models, it achieves better performance and less computational complexity. Then we apply the knowledge distillation strategy and provide a comparison of the data efficiency of the teacher model with different architectures. Finally, we propose a progressive pruning strategy, which involves pruning the model multiple times in small amounts, resulting in better performance compared to a single step pruning. Experiments are conducted on the TAU dataset. With Rep-Mobile and these training strategies, our proposed ASC system achieves the state-of-the-art (SOTA) results so far, while also winning the first place with a significant advantage over others in the DCASE2024 Challenge. Bing Han 0008, Wen Huang 0004, Zhengyang Chen, Anbai Jiang, Pingyi Fan, Cheng Lu 0007, Zhiqiang Lv, Jia Liu 0001, Weiqiang Zhang 0001, Yanmin Qian |
ICASSP | 1 |
| 2025 | SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech TranslationabstractSimultaneous Speech Translation (SimulST) enables real-time cross-lingual communication by jointly optimizing speech recognition and machine translation under strict latency constraints. Existing systems struggle to balance translation quality, latency, and semantic coherence, particularly in multilingual many-to-many scenarios where divergent read/write policies hinder unified strategy learning. In this paper, we present SimulMEGA(Simultaneous Generation by Mixture-of-Experts GAting), an unsupervised policy learning framework that combines prefix-based training with a Mixture-of-Experts refiner to learn effective read/write decisions in an implicit manner, without adding inference-time overhead. Our design requires only minimal modifications to standard transformer architectures and generalizes across both speech-to-text and text-to-speech streaming tasks. Through comprehensive evaluation on six language pairs, our 500 M-parameter speech-to-text model outperforms the Seamless baseline, achieving under 7% BLEU degradation at 1.5 s average lag and under 3% at 3 s. We further demonstrate SimulMEGA’s versatility by extending it to streaming TTS via a unidirectional backbone, yielding superior latency–quality trade-offs. Chenyang Le, Bing Han 0008, Jinshun Li, Songyong Chen, Yanmin Qian |
NeurIPS | 2 |
| 2024 | Exploring Large Scale Pre-Trained Models for Robust Machine Anomalous Sound DetectionabstractMachine anomalous sound detection is a useful technique for various applications, but it often suffers from poor generalization due to the challenges of data collection and complex acoustic environment. To address this issue, we propose a robust machine anomalous sound detection model that leverages self-supervised pre-trained models on large-scale speech data. Specifically, we assign different weights to the features from different layers of the pre-trained model and then use the working condition as the label for self-supervised classification fine-tuning. Moreover, we introduce a data augmentation method that simulates different operating states of the machine to enrich the dataset. Furthermore, we devise a transformer pooling method that fuses the features of different segments. Experiments on the DCASE2023 dataset show that our proposed method outperforms the commonly used reconstruction-based autoencoder and classification-based convolutional network by a large margin, demonstrating the effectiveness of large-scale pre-training for enhancing the generalization and robustness of machine anomalous sound detection. In Task2 of DCASE2023, we achieve 2nd place with these methods. Bing Han 0008, Zhiqiang Lv, Anbai Jiang, Wen Huang 0004, Zhengyang Chen, Yufeng Deng, Cheng Lu 0007, Weiqiang Zhang 0001, Pingyi Fan, Jia Liu 0001, Yanmin Qian |
ICASSP | 1 |
| 2024 | Robust Cross-Domain Speaker Verification with Multi-Level Domain AdaptersabstractSpeaker verification encounters significant challenges when confronted with diverse domain data, often resulting in performance degradation due to domain mismatch. To enhance performance in cross-domain scenarios, we introduce the Domain Adapter, an adaptable module designed for specific domains. This module learns and integrates domain-specific information with speaker-related data, mitigating domain-related variations and promoting convergence of utterance embeddings from the same speaker across diverse domains. It offers configurability across multiple levels and is adaptable to various backbone architectures. Our proposed module substantially enhances cross-domain performance with minimal parameter increments while effectively generalizing to previously unseen domains. In our experiments, we present results on the 3D-Speaker dataset, which provides acoustically-relevant attributes crucial for domain categorization and the subsequent learning of domain information. The top-performing system integrated with domain adapters achieved 10.8%, 14.8%, and 21.1% EER improvements over the baseline across three 3D-Speaker dataset trials. Wen Huang 0004, Bing Han 0008, Shuai Wang 0016, Zhengyang Chen, Yanmin Qian |
ICASSP | 2 |
| 2024 | Leveraging in-the-wild Data for Effective Self-supervised Pretraining in Speaker RecognitionabstractCurrent speaker recognition systems primarily rely on supervised approaches, constrained by the scale of labeled datasets. To boost the system performance, researchers leverage large pretrained models such as WavLM to transfer learned high-level features to the downstream speaker recognition task. However, this approach introduces extra parameters as the pretrained model remains in the inference stage. Another group of researchers directly apply self-supervised methods such as DINO to speaker embedding learning, yet they have not explored its potential on large-scale in-the-wild datasets. In this paper, we present the effectiveness of DINO training on the large-scale WenetSpeech dataset and its transferability in enhancing the supervised system performance on the CNCeleb dataset. Additionally, we introduce a confidence-based data filtering algorithm to remove unreliable data from the pretraining dataset, leading to better performance with less training data. The associated pretrained models, confidence files, pretraining and finetuning scripts will be made available in the Wespeaker toolkit. Shuai Wang 0016, Qibing Bai, Qi Liu 0018, Jianwei Yu 0001, Zhengyang Chen, Bing Han 0008, Yanmin Qian, Haizhou Li 0001 |
ICASSP | 6 |
| 2024 | InstructME: An Instruction Guided Music Edit Framework with Latent Diffusion Models
Bing Han 0008, Junyu Dai, Weituo Hao, Xinyan He, Jitong Chen, Yuxuan Wang 0002, Yanmin Qian, Xuchen Song |
IJCAI | 1 |
| 2024 | AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection
Anbai Jiang, Bing Han 0008, Zhiqiang Lv, Yufeng Deng, Weiqiang Zhang 0001, Xie Chen 0001, Yanmin Qian, Jia Liu 0001, Pingyi Fan |
INTERSPEECH | 2 |
| 2024 | Improving Anomalous Sound Detection Via Low-Rank Adaptation Fine-Tuning of Pre-Trained Audio ModelsabstractAnomalous Sound Detection (ASD) has gained significant interest through the application of various Artificial Intelligence (AI) technologies in industrial settings. Though possessing great potential, ASD systems can hardly be readily deployed in real production sites due to the generalization problem, which is primarily caused by the difficulty of data collection and the complexity of environmental factors. This paper introduces a robust ASD model that leverages audio pre-trained models. Specifically, we fine-tune these models using machine operation data, employing SpecAug as a data augmentation strategy. Additionally, we investigate the impact of utilizing Low-Rank Adaptation (LoRA) tuning instead of full fine-tuning to address the problem of limited data for fine-tuning. Our experiments on the DCASE2023 Task 2 dataset establish a new benchmark of 77.75% on the evaluation set, with a significant improvement of 6.48% compared with previous state-of-the-art (SOTA) models, including top-tier traditional convolutional networks and speech pre-trained models, which demonstrates the effectiveness of audio pre-trained models with LoRA tuning. Ablation studies are also conducted to showcase the efficacy of the proposed scheme. Xinhu Zheng, Anbai Jiang, Bing Han 0008, Yanmin Qian, Pingyi Fan, Jia Liu 0001, Weiqiang Zhang 0001 |
SLT | 3 |
| 2024 | Advancing speaker embedding learning: Wespeaker toolkit for research and production
Shuai Wang 0016, Zhengyang Chen, Bing Han 0008, Chengdong Liang, Xu Xiang, Wen Ding 0005, Johan Rohdin, Anna Silnova, Yanmin Qian, Haizhou Li 0001 |
Speech Commun. | 3 |
| 2024 | Attention-Based Encoder-Decoder End-to-End Neural Diarization With Embedding EnhancerabstractDeep neural network-based systems have significantly improved the performance of speaker diarization tasks. However, end-to-end neural diarization (EEND) systems often struggle to generalize to scenarios with an unseen number of speakers, while target speaker voice activity detection (TS-VAD) systems tend to be overly complex. In this paper, we propose a simple attention-based encoder-decoder network for end-to-end neural diarization (AED-EEND). In our training process, we introduce a teacher-forcing strategy to address the speaker permutation problem, leading to faster model convergence. For evaluation, we propose an iterative decoding method that outputs diarization results for each speaker sequentially. Additionally, we propose an Enhancer module to enhance the frame-level speaker embeddings, enabling the model to handle scenarios with an unseen number of speakers. We also explore replacing the transformer encoder with a Conformer architecture, which better models local information. Furthermore, we discovered that commonly used simulation datasets for speaker diarization have a much higher overlap ratio compared to real data. We found that using simulated training data that is more consistent with real data can achieve an improvement in consistency. Extensive experimental validation demonstrates the effectiveness of our proposed methodologies. Our best system achieved a new state-of-the-art diarization error rate (DER) performance on all the CALLHOME (10.08%), DIHARD II (24.64%), and AMI (13.00%) evaluation benchmarks when overlap is considered and no oracle voice activity detection (VAD) is used. Beyond speaker diarization, our AED-EEND system also shows remarkable competitiveness as a speech type detection model. Zhengyang Chen, Bing Han 0008, Shuai Wang 0016, Yanmin Qian |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Self-Supervised Learning With Cluster-Aware-DINO for High-Performance Robust Speaker VerificationabstractThe automatic speaker verification task has achieved great success using deep learning approaches with a large-scale, manually annotated dataset. However, collecting a significant amount of well-labeled data for system building is very difficult and expensive. Recently, self-supervised speaker verification has attracted a lot of interest due to its no dependency on labeled data. In this article, we propose a novel and advanced self-supervised learning framework based on our prior work, which can construct a powerful speaker verification system with high performance without using any labeled data. To avoid the impact of false negative pairs, we adopt the self-distillation with no labels (DINO) framework as the initial model, which can be trained without exploiting negative pairs. Then, we further introduce a cluster-aware training strategy for DINO to improve the diversity of data. In the iterative learning stage, due to a mass of unreliable labels from unsupervised clustering, the quality of pseudo labels is important for the system performance. This motivates us to propose dynamic loss-gate and label correction (DLG-LC) methods to alleviate the performance degradation caused by unreliable labels. Furthermore, we extend the DLG-LC from single-modality to multi-modality on the audio-visual dataset to further improve the performance. The experiments were conducted using the widely-used Voxceleb dataset. Compared to the best-known self-supervised speaker verification system, our proposed method achieve relative EER improvement of 22.17%, 27.94% and 25.56% on Vox-O, Vox-E and Vox-H test sets, even with fewer iterations, smaller models, and simpler clustering methods. Importantly, the newly proposed self-supervised learning system even achieves comparable results with the fully supervised system, but without using any human-labeled data. Bing Han 0008, Zhengyang Chen, Yanmin Qian |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Exploring Binary Classification Loss for Speaker VerificationabstractThe mismatch between close-set training and open-set testing usually leads to significant performance degradation for speaker verification task. For existing loss functions, metric learning-based objectives depend strongly on searching effective pairs which might hinder further improvements. And popular multi-classification methods are usually observed with degradation when evaluated on unseen speakers. In this work, we introduce SphereFace2 framework which uses several binary classifiers to train the speaker model in a pair-wise manner instead of performing multi-classification. Benefiting from this learning paradigm, it can efficiently alleviate the gap between training and evaluation. Experiments conducted on Voxceleb show that the SphereFace2 outperforms other existing loss functions, especially on hard trials. Besides, large margin fine-tuning strategy is proven to be compatible with it for further improvements. Finally, SphereFace2 also shows its strong robustness to class-wise noisy labels which has the potential to be applied in the semi-supervised training scenario with inaccurate estimated pseudo labels. Bing Han 0008, Zhengyang Chen, Yanmin Qian |
ICASSP | 1 |
| 2023 | Attention-based Encoder-Decoder Network for End-to-End Neural Speaker Diarization with Target Speaker Attractor
Zhengyang Chen, Bing Han 0008, Shuai Wang 0016, Yanmin Qian |
INTERSPEECH | 2 |
| 2023 | Build a SRE Challenge System: Lessons from VoxSRC 2022 and CNSRC 2022
Zhengyang Chen, Bing Han 0008, Xu Xiang, Houjun Huang, Bei Liu 0003, Yanmin Qian |
INTERSPEECH | 2 |
| 2022 | MLP-SVNET: A Multi-Layer Perceptrons Based Network for Speaker VerificationabstractConvolution and self-attention based neural networks have both obtained excellent performance in automatic speaker verification. However, the convolution model often lacks the ability of long-term dependency modeling due to the limitation of receptive field, while the self-attention model is insufficient to model local information. To tackle this limitation, we propose a new multi-layer perceptrons based speaker verification network (MLP-SVNet) which can apply MLPs across temporal and frequency dimensions to capture the local and global information at the same time. The experimental results conducted on Voxceleb show that the proposed model is very competitive when compared to other systems based on convolution or self-attention. In addition, we demonstrate that MLP-SVNet based on multi-layer per-ceptrons can produce complementary embeddings, which can be fused with the state-of-the-art system to further improve the performance. Bing Han 0008, Zhengyang Chen, Bei Liu 0003, Yanmin Qian |
ICASSP | 1 |
| 2022 | Local Information Modeling with Self-Attention for Speaker VerificationabstractTransformer based on self attention mechanism has demonstrated its state-of-the-art performance in most natural language processing (NLP) tasks, but it’s not very competitive when applied for speaker verification in previous works. Generally, speaker identity is mostly reflected by the relationship between adjacent tokens, whose extraction mainly depends on local modeling ability. However, the self-attention module, as the key component of transformer, can help the model make full use of global information but insufficient to capture the local information. To tackle this limitation, in this paper, we strengthen the local information modeling from two different aspects: restricting the attention context to be local and introducing convolution operation into transformer. Experiments conducted on Voxceleb illustrate that our proposed methods can notably improve system performance, verifying the significance of local information for speaker verification task. Bing Han 0008, Zhengyang Chen, Yanmin Qian |
ICASSP | 1 |
| 2022 | The Sjtu System For Multimodal Information Based Speech Processing Challenge 2021abstractThis paper describes the SJTU system for ICASSP Multi-modal Information based Speech Processing Challenge (MISP) 2021. To solve the speech recognition problem in real complex environments where time-synchronized near- and far-field signals are available for training an enhancement frontend. We build a joint system with speech enhancement frontend and speech recognition backend. These two modules are optimized jointly by both ASR and enhancement criteria. Audio-visual fusion is explored to further boost the ASR performance. ROVER and test time augmentation techniques are used to combine recognition results from multiple systems. The final system achieves Chinese character error rates (CCER) of 34.9% on dev set and 34.0% on test set, which achieved third place in the MISP challenge. The absolute CCER reduction compared with the official baseline system is 26.9% on dev set and 28.7% on test set. Wei Wang 0010, Xun Gong 0005, Zhikai Zhou, Chenda Li, Wangyou Zhang, Bing Han 0008, Yanmin Qian |
ICASSP | 7 |
| 2022 | Self-Supervised Speaker Verification Using Dynamic Loss-Gate and Label CorrectionabstractFor self-supervised speaker verification, the quality of pseudo labels decides the upper bound of the system due to the massive unreliable labels.In this work, we propose dynamic loss-gate and label correction (DLG-LC) to alleviate the performance degradation caused by unreliable estimated labels.In DLG, we adopt Gaussian Mixture Model (GMM) to dynamically model the loss distribution and use the estimated GMM to distinguish the reliable and unreliable labels automatically.Besides, to better utilize the unreliable data instead of dropping them directly, we correct the unreliable label with model predictions.Moreover, we apply the negative-pairs-free DINO framework in our experiments for further improvement.Compared to the best-known speaker verification system with self-supervised learning, our proposed DLG-LC converges faster and achieves 11.45%, 18.35% and 15.16% relative improvement on Vox-O, Vox-E and Vox-H trials of Voxceleb1 evaluation dataset. Bing Han 0008, Zhengyang Chen, Yanmin Qian |
INTERSPEECH | 1 |
| 2022 | DF-ResNet: Boosting Speaker Verification Performance with Depth-First Design
Bei Liu 0003, Zhengyang Chen, Shuai Wang 0016, Haoyu Wang 0007, Bing Han 0008, Yanmin Qian |
INTERSPEECH | 5 |
| 2022 | A Comprehensive Study on Self-Supervised Distillation for Speaker Representation LearningabstractIn real application scenarios, it is often challenging to obtain a large amount of labeled data for speaker representation learning due to speaker privacy concerns. Self-supervised learning with no labels has become a more and more promising way to solve it. Compared with contrastive learning, self-distilled approaches use only positive samples in the loss function and thus are more attractive. In this paper, we present a comprehensive study on self-distilled self-supervised speaker representation learning, especially on critical data augmentation. Our proposed strategy of audio perturbation augmentation has pushed the performance of the speaker representation to a new limit. The experimental results show that our model can achieve a new SoTA on Voxceleb 1 speaker verification evaluation benchmark (i.e., equal error rate (EER) 2.505%, 2.473%, and 4.791 % for trial Vox1-O, Vox1-E and Vox1-H, respectively), discarding any speaker labels in the training phase. Zhengyang Chen, Yao Qian, Bing Han 0008, Yanmin Qian, Michael Zeng 0001 |
SLT | 3 |
| 2021 | SynAug: Synthesis-Based Data Augmentation for Text-Dependent Speaker VerificationabstractText-dependent speaker verification systems trained on large amount of labelled data exhibit remarkable performance. However, collecting the speech from a lot of speakers with target transcript is a lengthy and expensive process. In this work, we propose a synthesis based data augmentation method (SynAug) to expand the training set with more speakers and text-controlled synthesized speech. The performance of SynAug is evaluated on the RSR2015 dataset. Experimental results show that for i-vector framework, the proposed methods can boost the system performance significantly, especially for the low-resource condition where the amount of genuine speech is extremely limited. Moreover, combined with traditional data augmentation methods such as adding noises and reverberation, the systems could be further strengthened in extremely limited resource situation. Chenpeng Du, Bing Han 0008, Shuai Wang 0016, Yanmin Qian, Kai Yu 0004 |
ICASSP | 2 |
| 2021 | The SJTU System for Short-Duration Speaker Verification Challenge 2021abstractThis paper presents the SJTU system for both text-dependent and text-independent tasks in short-duration speaker verification (SdSV) challenge 2021.In this challenge, we explored different strong embedding extractors to extract robust speaker embedding.For text-independent task, language-dependent adaptive snorm is explored to improve the system performance under the cross-lingual verification condition.For text-dependent task, we mainly focus on the in-domain fine-tuning strategies based on the model pre-trained on large-scale out-of-domain data.In order to improve the distinction between different speakers uttering the same phrase, we proposed several novel phrase-aware fine-tuning strategies and phrase-aware neural PLDA.With such strategies, the system performance is further improved.Finally, we fused the scores of different systems, and our fusion systems achieved 0.0473 in Task1 (rank 3) and 0.0581 in Task2 (rank 8) on the primary evaluation metric. Bing Han 0008, Zhengyang Chen, Zhikai Zhou, Yanmin Qian |
Interspeech | 1 |