EDBT 2026 Demo / reviewers in the wild / expert
Bin Gu 0004
dblp:29/1758-4
· DBLP profile ↗
16ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0003-1621-7311ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Noise-Conditioned Mixture-of-Experts Framework for Robust Speaker VerificationabstractRobust speaker verification under noisy conditions remains an open challenge. Conventional deep learning methods learn a robust unified speaker representation space against diverse background noise and achieve significant improvement. In contrast, this paper presents a noise-conditioned mixture-of-experts framework that decomposes the feature space into specialized noise-aware subspaces for speaker verification. Specifically, we propose a noise-conditioned expert routing mechanism, a universal model based expert specialization strategy, and an SNR-decaying curriculum learning protocol, collectively improving model robustness and generalization under diverse noise conditions. The proposed method can automatically route inputs to expert networks based on noise information derived from the inputs, where each expert targets distinct noise characteristics while preserving speaker identity information. Comprehensive experiments demonstrate consistent superiority over baselines. Bin Gu 0004, Haitao Zhao 0001, Jibo Wei |
IEEE Signal Process. Lett. | 1 |
| 2025 | Aligning Noisy-Clean Speech Pairs at Feature and Embedding Levels for Learning Noise-Invariant Speaker RepresentationsabstractIn this paper, we propose a noise-invariant speaker representation learning (SRL) approach by aligning noisy-clean speech pairs at both the feature and embedding levels for model training. Specifically, we first construct noisy-clean pairs using data augmentation during training. The noisy features are then processed by a Conformer-based enhancement module. The feature-level alignment is achieved by minimizing the mean squared error between the enhanced and original clean data. At the embedding level, we introduce a supervised contrastive learning loss with noise-adaptive margin to simultaneously enhance the intra-speaker compactness and the inter-speaker separability and better adapt different noise levels, in combination with the Barlow Twins self-supervised loss to align the noisy-clean data pairs and reduce noise redundancy in the embedding space. Finally, these loss components are integrated with conventional classification loss to train the SRL network. Experimental results on various VoxCeleb1 test sets synthesized with noise sources demonstrate the effectiveness of the proposed method. Zuoliang Li, Yang Ai, Jie Zhang 0042, Shengyu Peng, Bin Gu 0004, Wu Guo |
ICASSP | 6 |
| 2025 | A Study of Multi-Scale Feature Learning From Pre-Trained Models on Speaker VerificationabstractIn this paper, a multi-scale feature fusion paradigm is proposed to fully exploit the power of the pre-trained models for text-independent speaker verification. It contains a front-end feature extractor and an enhanced ECAPA-TDNN backend in a cascade manner. The feature extractor incorporates local representations of the CNN layers as well as the global clues of the Transformer layers of the pre-trained models, which are combined to construct the multi-scale discriminative features. The outputs of the feature extractor are then fed into the back-end model (tailored from ECAPA-TDNN) to obtain the final speaker embedding. Results on VoxCeleb datasets validate the superiority of the proposed method with equal error rates of 0.633% and 0.457% on the official trials of Vox1-O using the base and large pre-trained models, respectively. Shengyu Peng, Wu Guo, Jie Zhang 0042, Zuoliang Li, Bin Gu 0004, Yang Ai |
ICASSP | 6 |
| 2023 | Introducing Self-Supervised Phonetic Information for Text-Independent Speaker Verification
Wu Guo, Bin Gu 0004 |
INTERSPEECH | 3 |
| 2023 | A Detection-based Attention Alignment Method for Document-level Neural Machine TranslationabstractPrevious works have shown that inter-sentential contextual information can lead to substantial improvements in document-level neural machine translation (DocNMT).Most existing DocNMT models focus on methods of introducing inter-sentential contextual information through attention mechanisms.Compared to intra-sentential attention, however, the long-range dependency in documentlevel attention calculation inevitably introduces meaningless contextual noise, resulting in significant performance deterioration.To address this problem, this paper proposes a detection-based attention alignment method, to help each translating word focus on relevant informative contextual words.We first introduce a context detector that automatically evaluates each source-side word's effect on the model's prediction.Based on the detection results, we align the original attention weights by integrating the cosine similarity between the aligned and original attention weights into the loss function, under a multi-task framework, which allows DocNMT to more effectively capture the document-level context.The results for three English-German (En-De) public translation datasets show that the proposed method can obtain consistent improvements over a strong G-Transformer baseline. Kang Zhong, Wu Guo, Bin Gu 0004, Weitai Zhang |
SEKE | 3 |
| 2023 | Memory Storable Network Based Feature Aggregation for Speaker Representation LearningabstractLearning fixed-dimensional speaker representation using deep neural networks is a key step in speaker verification. In this work, we propose an auxiliary memory storable network (MSN) to assist a backbone network for learning discriminative features, which are sequentially aggregated from lower to deeper layers of the backbone. The proposed MSN has a similar architecture to the ResNet and contains a set of cascaded feature aggregation (FA) blocks. Each FA block first aggregates the multi-level features from the previous block and the features from the corresponding backbone layer. The output features of each intermediate layer within the backbone are then refined by the multi-level features of the corresponding FA block through masking and biasing operations. Finally, the features from the last layers of both MSN and the backbone are concatenated to form more discriminative speaker representations. Experimental results on five public datasets show significant and consistent improvements over conventional approaches. The effectiveness of the proposed method is also validated using ablation studies, showing a robust generalization capacity in combination with different backbone networks. Bin Gu 0004, Wu Guo, Jie Zhang 0042 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | A Dynamic Convolution Framework for Session-Independent Speaker Embedding LearningabstractSpeaker verification (SV) has suffered from session variability in complex acoustic scenarios, and learning session independent speaker representations remains a challenging problem. To tackle this, we propose a dynamic convolution framework for SV in this article, which dynamically adapts the model parameters to each input feature during inference, such that the model can flexibly extract robust speaker characteristics under different acoustic conditions. Specifically, we combine an adaptive context vector extraction (ACVE) module and a sub-kernel scaling (SKS) module in a cascaded manner. The ACVE uses a moving weighted sum and a sub-band self-attention in parallel to extract global-local information, which is then passed through the subsequent SKS module for efficient dynamic kernel generation. The proposed method can be easily implemented on the backbone network by replacing the conventional static counterparts. Experimental results on five public SV datasets show significant and consistent improvements over comparison approaches, and ablation studies and visual analysis further demonstrate the effectiveness of the proposed method. Bin Gu 0004, Jie Zhang 0042, Wu Guo |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Multi-Level Contrastive Learning for Cross-Lingual AlignmentabstractCross-language pre-trained models such as multilingual BERT (mBERT) have achieved significant performance in various cross-lingual downstream NLP tasks. This paper proposes a multi-level contrastive learning (ML-CTL) framework to further improve the cross-lingual ability of pre-trained models. The proposed method uses translated parallel data to encourage the model to generate similar semantic embeddings for different languages. However, unlike the sentence-level alignment used in most previous studies, in this paper, we explicitly integrate the word-level information of each pair of parallel sentences into contrastive learning. Moreover, cross-zero noise contrastive estimation (CZ-NCE) loss is proposed to alleviate the impact of the floating-point error in the training process with a small batch size. The proposed method significantly improves the cross-lingual transfer ability of our basic model (mBERT) and outperforms on multiple zero-shot cross-lingual downstream tasks compared to the same-size models in the Xtreme benchmark. Beiduo Chen, Wu Guo, Bin Gu 0004, Quan Liu 0003 |
ICASSP | 3 |
| 2022 | Deep speaker embedding with frame-constrained training strategy for speaker verification
Bin Gu 0004 |
INTERSPEECH | 1 |
| 2022 | Dynamic Convolution With Global-Local Information for Session-Invariant Speaker Representation LearningabstractVarious mismatchedconditions result in performance degradation of the speaker verification (SV) systems. To address this issue, we extract robust speaker representations by devising a global-local information-based dynamic convolution neural network. In the proposed method, both global and local information of the input features are exploited to dynamically modify the convolution kernel values. This increases the model capability of capturing speaker characteristics by compensating both the inter- and intra-session variabilities. Extensive experiments on four publicly available SV datasets show significant and consistent improvements over the conventional approaches. The effectiveness of the proposed method is further investigated using ablation studies and visualizations. Bin Gu 0004, Wu Guo |
IEEE Signal Process. Lett. | 1 |
| 2021 | Improved Meta-Learning Training for Speaker VerificationabstractMeta-learning (ML) has recently become a research hotspot in speaker verification (SV).We introduce two methods to improve the meta-learning training for SV in this paper.For the first method, a backbone embedding network is first jointly trained with the conventional cross entropy loss and prototypical networks (PN) loss.Then, inspired by speaker adaptive training in speech recognition, additional transformation coefficients are trained with only the PN loss.The transformation coefficients are used to modify the original backbone embedding network in the x-vector extraction process.Furthermore, the random erasing (RE) data augmentation technique is applied to all support samples in each episode to construct positive pairs, and a contrastive loss between the augmented and the original support samples is added to the objective in model training.Experiments are carried out on the Speaker in the Wild (SITW) and VOiCES databases.Both of the methods can obtain consistent improvements over existing meta-learning training frameworks.By combining these two methods, we can observe further improvements on these two databases. Yafeng Chen, Wu Guo, Bin Gu 0004 |
Interspeech | 3 |
| 2021 | Bidirectional Multiscale Feature Aggregation for Speaker VerificationabstractIn this paper, we propose a novel bidirectional multiscale feature aggregation (BMFA) network with attentional fusion modules for text-independent speaker verification.The feature maps from different stages of the backbone network are iteratively combined and refined in both a bottom-up and top-down manner.Furthermore, instead of simple concatenation or elementwise addition of feature maps from different stages, an attentional fusion module is designed to compute the fusion weights.Experiments are conducted on the NIST SRE16 and VoxCeleb1 datasets.The experimental results demonstrate the effectiveness of the bidirectional aggregation strategy and show that the proposed attentional fusion module can further improve the performance. Jiajun Qi, Wu Guo, Bin Gu 0004 |
Interspeech | 3 |
| 2020 | An Improved Deep Neural Network for Modeling Speaker Characteristics at Different Temporal ScalesabstractThis paper presents an improved deep embedding learning method based on a convolutional neural network (CNN) for text-independent speaker verification. Two improvements are proposed for x-vector embedding learning: (1) a multiscale convolution (MSCNN) is adopted in the frame-level layers to capture the complementary speaker information in different receptive fields; (2) a Baum-Welch statistics attention (BWSA) mechanism is applied in the pooling layer, which can integrate more useful long-term speaker characteristics in the temporal pooling layer. Experiments are carried out on the NIST SRE16 evaluation set. The results demonstrate the effectiveness of the MSCNN and show that the proposed BWSA can further improve the performance of the DNN embedding system. Bin Gu 0004, Wu Guo, Li-Rong Dai 0001, Jun Du 0002 |
ICASSP | 1 |
| 2020 | Unsupervised Regularization-Based Adaptive Training for Speech Recognition
Fenglin Ding, Wu Guo, Bin Gu 0004, Zhen-Hua Ling, Jun Du 0002 |
INTERSPEECH | 3 |
| 2020 | Adaptive Speaker Normalization for CTC-Based Speech Recognition
Fenglin Ding, Wu Guo, Bin Gu 0004, Zhen-Hua Ling, Jun Du 0002 |
INTERSPEECH | 3 |
| 2020 | An Adaptive X-Vector Model for Text-Independent Speaker VerificationabstractIn this paper, adaptive mechanisms are applied in deep neural network (DNN) training for x-vector-based text-independent speaker verification.First, adaptive convolutional neural networks (ACNNs) are employed in frame-level embedding layers, where the parameters of the convolution filters are adjusted based on the input features.Compared with conventional CNNs, ACNNs have more flexibility in capturing speaker information.Moreover, we replace conventional batch normalization (BN) with adaptive batch normalization (ABN).By dynamically generating the scaling and shifting parameters in BN, ABN adapts models to the acoustic variability arising from various factors such as channel and environmental noises.Finally, we incorporate these two methods to further improve performance.Experiments are carried out on the speaker in the wild (SITW) and VOiCES databases.The results demonstrate that the proposed methods significantly outperform the original xvector approach. Bin Gu 0004, Wu Guo, Fenglin Ding, Zhen-Hua Ling, Jun Du 0002 |
INTERSPEECH | 1 |