Wu Guo

dblp:20/4480 · DBLP profile ↗
← Back
68ranked-venue papers
2as first author
31since 2021 · last 2026
0000-0002-3779-7944ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 59 · 2 first-author · 23 since 2021Artificial intelligence and machine learning · 39 · 19 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Study of the Removability of Speaker-Adversarial Perturbations
abstract
Recent advancements in adversarial attacks have demonstrated their effectiveness in misleading speaker recognition models, making wrong predictions about speaker identities. On the other hand, defense techniques against speaker-adversarial attacks focus on reducing the effects of speaker-adversarial perturbations on speaker attribute extraction. These techniques do not seek to fully remove the perturbations and restore the original speech. To this end, this paper studies the removability of speaker-adversarial perturbations. Specifically, the investigation is conducted assuming various degrees of awareness of the perturbation generator across three scenarios: ignorant, semi-informed, and well-informed. Besides, we consider both the optimization-based and feedforward perturbation generation methods. Experiments conducted on the LibriSpeech dataset demonstrated that: 1) in the ignorant scenario, speaker-adversarial perturbations cannot be eliminated, although their impact on speaker attribute extraction is reduced, 2) in the semi-informed scenario, the speaker-adversarial perturbations cannot be fully removed, while those generated by the feedforward model can be considerably reduced, and 3) in the well-informed scenario, speaker-adversarial perturbations are nearly eliminated, allowing for the restoration of the original speech.
Chenyang Guo, Kong-Aik Lee, Zhen-Hua Ling, Wu Guo
IEEE Trans. Dependable Secur. Comput.5
2025 Recursive Feature Learning from Pre-Trained Models for Spoofing Speech Detection
abstract
It was recently revealed that using features extracted from pre-trained models can achieve much better performance than using conventional hand-crafted acoustic features for spoofing speech detection. In this paper, we therefore enhance the features from pre-trained model based on recursive learning. Specifically, we modify the pre-trained model by feeding the features from the topmost transformer layer to bottom layers recursively, and the obtained recursive features from the bottom layers are fused with that from topmost layer. The fused features are then fed into the backend classifiers. Experiments are carried out on two benchmark datasets (i.e., ASVspoof 2019 LA and ASVspoof 2021 LA), which show the superiority of the proposed method over state-of-the-art systems.
Yang Ai, Zuoliang Li, Shengyu Peng, Wu Guo
ICASSP5
2025 Aligning Noisy-Clean Speech Pairs at Feature and Embedding Levels for Learning Noise-Invariant Speaker Representations
abstract
In this paper, we propose a noise-invariant speaker representation learning (SRL) approach by aligning noisy-clean speech pairs at both the feature and embedding levels for model training. Specifically, we first construct noisy-clean pairs using data augmentation during training. The noisy features are then processed by a Conformer-based enhancement module. The feature-level alignment is achieved by minimizing the mean squared error between the enhanced and original clean data. At the embedding level, we introduce a supervised contrastive learning loss with noise-adaptive margin to simultaneously enhance the intra-speaker compactness and the inter-speaker separability and better adapt different noise levels, in combination with the Barlow Twins self-supervised loss to align the noisy-clean data pairs and reduce noise redundancy in the embedding space. Finally, these loss components are integrated with conventional classification loss to train the SRL network. Experimental results on various VoxCeleb1 test sets synthesized with noise sources demonstrate the effectiveness of the proposed method.
Zuoliang Li, Yang Ai, Jie Zhang 0042, Shengyu Peng, Bin Gu 0004, Wu Guo
ICASSP7
2025 A Study of Multi-Scale Feature Learning From Pre-Trained Models on Speaker Verification
abstract
In this paper, a multi-scale feature fusion paradigm is proposed to fully exploit the power of the pre-trained models for text-independent speaker verification. It contains a front-end feature extractor and an enhanced ECAPA-TDNN backend in a cascade manner. The feature extractor incorporates local representations of the CNN layers as well as the global clues of the Transformer layers of the pre-trained models, which are combined to construct the multi-scale discriminative features. The outputs of the feature extractor are then fed into the back-end model (tailored from ECAPA-TDNN) to obtain the final speaker embedding. Results on VoxCeleb datasets validate the superiority of the proposed method with equal error rates of 0.633% and 0.457% on the official trials of Vox1-O using the base and large pre-trained models, respectively.
Shengyu Peng, Wu Guo, Jie Zhang 0042, Zuoliang Li, Bin Gu 0004, Yang Ai
ICASSP2
2025 Leveraging Multi-Level Features of ATST with Conformer-Based Dual-Branch Network for Sound Event Detection
Lipeng Dai, Shengyu Peng, Wu Guo
INTERSPEECH6
2025 Parameter-Efficient Fine-tuning with Instance-Aware Prompt and Parallel Adapters for Speaker Verification
Shengyu Peng, Wu Guo, Jie Zhang 0042, Lipeng Dai, Zuoliang Li
INTERSPEECH2
2024 Modeling Pseudo-Speaker Uncertainty in Voice Anonymization
abstract
Voice anonymization refers to the goal of suppressing personally identifiable voice attributes in speech. State-of-the-art models based on the voice conversion framework accomplish this goal by replacing the voice attributes of the speaker with those of a pseudo-speaker. This paper proposes to exploit the uncertainty estimate of pseudo-speaker in voice anonymization. For each target speaker, a pseudo-speaker distribution, characterized by a point estimate and its uncertainty, is estimated from a selected set of cohort speakers. Based on this distribution, a pseudo-speaker vector is sampled and used to replace the voice attributes in an anonymized speech. The efficacy of the proposed method was validated in the framework as provided by VoicePrivacy Challenge 2022. Audio samples can be found in https://voiceprivacy.github.io/pseudo-speaker-vector/.
Kong-Aik Lee, Wu Guo, Zhen-Hua Ling
ICASSP3
2024 Generating High-Quality Adversarial Examples with Universal Perturbation-Based Adaptive Network and Improved Perceptual Loss
abstract
Deep neural network-based speaker identification systems are vulnerable to adversarial attacks. However, the distortions of the adversarial examples are still obvious in most cases. In this work, we therefore propose a universal perturbation-based adaptive network (UPAN) to generate high-quality adversarial examples. Specifically, the UPAN first uses a universal perturbation generative module to generate an input-independent perturbation, which is then fed into an adaptation module to generate the input-specific perturbation. Finally, the sum of the generated perturbation and input audio is used as the adversarial example to attack the speaker identification system. To further enhance the speech quality, we introduce an improved perceptual loss that combines the mean square error and frame-wise cosine similarity of the MFCC features between the input audio and adversarial examples. Experimental results on the VoxCeleb1 dataset demonstrate that the proposed approach is effective for both non-targeted and targeted attacks.
Zhuhai Li, Wu Guo
ICASSP2
2024 Robust Spoof Speech Detection Based on Multi-Scale Feature Aggregation and Dynamic Convolution
abstract
Spoof speech detection (SSD) can help to protect an automatic speaker recognition system against malicious attacks. However, there exists a great diversity in the spoof utterances generated by different text-to-speech and voice conversion algorithms, resulting in a poor generality of an SSD system to unseen spoofing attacks. To address this problem, we integrate multi-scale feature aggregation (MFA) and dynamic convolution operations into the anti-spoofing framework to detect different local and global artifacts of unseen spoofing attacks. The proposed framework mainly contains eight stacked MFA blocks, where in each block the light-Res2Net module is used to capture multi-scale features and the convolutional kernel is dynamically generated by the local and global statistical information of the inputs. Results on two benchmark datasets (i.e., ADD 2023 Fake Audio Detection and ASVspoof 2021 Logical Access) show the superiority of the proposed method over existing state-of-the-art systems.
Wu Guo
ICASSP6
2024 Meta Representation Learning Method for Robust Speaker Verification in Unseen Domains
abstract
This paper presents a meta representation learning method for robust speaker verification (SV) in unseen domains. It is known that the existing embedding learning based SV systems may suffer from domain mismatch issues. To address this, we propose an episodic training procedure to compensate domain mismatch conditions at runtime. Specifically, episodes are constructed with domain balanced episodic sampling from two different domains, and a new domain alignment (DA) module is added besides the feature extractor (FE) and classifier to existing network structures. In each episodic training iteration, FE and DA modules are optimized separately with different objectives to improve the robustness of learning. Besides, a cross-domain inter-class alignment (CDICA) loss is proposed for improving the domain generalization ability. Experimental results on CNCeleb and VoxCeleb benchmarks demonstrate significant performance gains for unseen domains in SV.
Jian-Tao Zhang, Yan Song 0001, Wu Guo, Hao-Yu Song, Ian McLoughlin 0001
ICASSP4
2024 Multi-layer Feature Augmentation Based Transferable Adversarial Examples Generation for Speaker Recognition
Zhuhai Li, Wu Guo
ICIC (4)3
2024 Document-Level Machine Translation with Effective Batch-Level Context Representation
abstract
It is critical to provide inter-sentential context for document-level neural machine translation (DocNMT) to achieve higher-quality translations. As the document-level information is naturally preserved in mini-batches in case sentences are not shuffled, in this work we propose an effective batch-level context representation (EBCR) for DocNMT by leveraging structural contextual clues in the mini-batches. The EBCR is a plug-in module that is added to each encoder layer of the conventional Transformer model and can condense the inter-sentential contextual information within the mini-batch and reinforce the inter-sentential local context through gating operation. The proposed method is evaluated on three English-German document translation datasets, and results show that our model can present the wide-range context more effectively than existing methods.
Kang Zhong, Wu Guo
IJCNN3
2024 Contrastive Learning and Inter-Speaker Distribution Alignment Based Unsupervised Domain Adaptation for Robust Speaker Verification
Zuoliang Li, Wu Guo, Shengyu Peng
INTERSPEECH2
2024 Boosting the Transferability of Adversarial Examples with Gradient-Aligned Ensemble Attack for Speaker Recognition
Zhuhai Li, Wu Guo
INTERSPEECH3
2024 Fine-tune Pre-Trained Models with Multi-Level Feature Fusion for Speaker Verification
Shengyu Peng, Wu Guo, Zuoliang Li
INTERSPEECH2
2024 Adapter Learning from Pre-trained Model for Robust Spoof Speech Detection
Wu Guo, Shengyu Peng, Zhuhai Li
INTERSPEECH2
2024 Spoofing Speech Detection by Modeling Local Spectro-Temporal and Long-term Dependency
Wu Guo, Shengyu Peng
INTERSPEECH2
2024 On The Generation and Removal of Speaker Adversarial Perturbation For Voice-Privacy Protection
abstract
Neural networks are commonly known to be vulnerable to adversarial attacks mounted through subtle perturbation on the input data. Recent development in voice-privacy protection has shown the positive use cases of the same technique to conceal speaker’s voice attribute with additive perturbation signal generated by an adversarial network. This paper examines the reversibility property where an entity generating the adversarial perturbations is authorized to remove them and restore original speech (e.g., the speaker him/herself). A similar technique could also be used by an investigator to deanonymize a voice-protected speech to restore criminals’ identities in security and forensic analysis. In this setting, the perturbation generative module is assumed to be known in the removal process. To this end, a joint training of perturbation generation and removal modules is proposed. Experimental results on the LibriSpeech dataset demonstrated that the subtle perturbations added to the original speech can be predicted from the anonymized speech while achieving the goal of privacy protection. By removing these perturbations from the anonymized sample, the original speech can be restored. Audio samples can be found in https://voiceprivacy.github.io/Perturbation-Generation-Removal/.
Chenyang Guo, Zhuhai Li, Kong-Aik Lee, Zhen-Hua Ling, Wu Guo
SLT6
2023 FaMeSumm: Investigating and Improving Faithfulness of Medical Summarization
abstract
Summaries of medical text shall be faithful by being consistent and factual with source inputs, which is an important but understudied topic for safety and efficiency in healthcare.In this paper, we investigate and improve faithfulness in summarization on a broad range of medical summarization tasks.Our investigation reveals that current summarization models often produce unfaithful outputs for medical input text.We then introduce FAMESUMM, a framework to improve faithfulness by fine-tuning pre-trained language models based on medical knowledge.FAMESUMM performs contrastive learning on designed sets of faithful and unfaithful summaries, and it incorporates medical terms and their contexts to encourage faithful generation of medical terms.We conduct comprehensive experiments on three datasets in two languages: health question and radiology report summarization datasets in English, and a patient-doctor dialogue dataset in Chinese.Results demonstrate that FAMESUMM is flexible and effective by delivering consistent improvements over mainstream language models such as BART, T5, mT5, and PEGASUS, yielding state-of-the-art performances on metrics for faithfulness and general quality.Human evaluation by doctors also shows that FAMESUMM generates more faithful outputs.
Yusen Zhang 0001, Wu Guo, Prasenjit Mitra 0001, Rui Zhang 0037
EMNLP3
2023 Introducing Self-Supervised Phonetic Information for Text-Independent Speaker Verification
Wu Guo, Bin Gu 0004
INTERSPEECH2
2023 A Detection-based Attention Alignment Method for Document-level Neural Machine Translation
abstract
Previous works have shown that inter-sentential contextual information can lead to substantial improvements in document-level neural machine translation (DocNMT).Most existing DocNMT models focus on methods of introducing inter-sentential contextual information through attention mechanisms.Compared to intra-sentential attention, however, the long-range dependency in documentlevel attention calculation inevitably introduces meaningless contextual noise, resulting in significant performance deterioration.To address this problem, this paper proposes a detection-based attention alignment method, to help each translating word focus on relevant informative contextual words.We first introduce a context detector that automatically evaluates each source-side word's effect on the model's prediction.Based on the detection results, we align the original attention weights by integrating the cosine similarity between the aligned and original attention weights into the loss function, under a multi-task framework, which allows DocNMT to more effectively capture the document-level context.The results for three English-German (En-De) public translation datasets show that the proposed method can obtain consistent improvements over a strong G-Transformer baseline.
Kang Zhong, Wu Guo, Bin Gu 0004, Weitai Zhang
SEKE2
2023 Memory Storable Network Based Feature Aggregation for Speaker Representation Learning
abstract
Learning fixed-dimensional speaker representation using deep neural networks is a key step in speaker verification. In this work, we propose an auxiliary memory storable network (MSN) to assist a backbone network for learning discriminative features, which are sequentially aggregated from lower to deeper layers of the backbone. The proposed MSN has a similar architecture to the ResNet and contains a set of cascaded feature aggregation (FA) blocks. Each FA block first aggregates the multi-level features from the previous block and the features from the corresponding backbone layer. The output features of each intermediate layer within the backbone are then refined by the multi-level features of the corresponding FA block through masking and biasing operations. Finally, the features from the last layers of both MSN and the backbone are concatenated to form more discriminative speaker representations. Experimental results on five public datasets show significant and consistent improvements over conventional approaches. The effectiveness of the proposed method is also validated using ablation studies, showing a robust generalization capacity in combination with different backbone networks.
Bin Gu 0004, Wu Guo, Jie Zhang 0042
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 A Dynamic Convolution Framework for Session-Independent Speaker Embedding Learning
abstract
Speaker verification (SV) has suffered from session variability in complex acoustic scenarios, and learning session independent speaker representations remains a challenging problem. To tackle this, we propose a dynamic convolution framework for SV in this article, which dynamically adapts the model parameters to each input feature during inference, such that the model can flexibly extract robust speaker characteristics under different acoustic conditions. Specifically, we combine an adaptive context vector extraction (ACVE) module and a sub-kernel scaling (SKS) module in a cascaded manner. The ACVE uses a moving weighted sum and a sub-band self-attention in parallel to extract global-local information, which is then passed through the subsequent SKS module for efficient dynamic kernel generation. The proposed method can be easily implemented on the backbone network by replacing the conventional static counterparts. Experimental results on five public SV datasets show significant and consistent improvements over comparison approaches, and ablation studies and visual analysis further demonstrate the effectiveness of the proposed method.
Bin Gu 0004, Jie Zhang 0042, Wu Guo
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Wider & Closer: Mixture of Short-channel Distillers for Zero-shot Cross-lingual Named Entity Recognition
abstract
Zero-shot cross-lingual named entity recognition (NER) aims at transferring knowledge from annotated and rich-resource data in source languages to unlabeled and lean-resource data in target languages.Existing mainstream methods based on the teacher-student distillation framework ignore the rich and complementary information lying in the intermediate layers of pre-trained language models, and domaininvariant information is easily lost during transfer.In this study, a mixture of short-channel distillers (MSD) method is proposed to fully interact the rich hierarchical information in the teacher model and to transfer knowledge to the student model sufficiently and efficiently.Concretely, a multi-channel distillation framework is designed for sufficient information transfer by aggregating multiple distillers as a mixture.Besides, an unsupervised method adopting parallel domain adaptation is proposed to shorten the channels between the teacher and student models to preserve domaininvariant features.Experiments on four datasets across nine languages demonstrate that the proposed method achieves new state-of-the-art performance on zero-shot cross-lingual NER and shows great generalization and compatibility across languages and fields.
Jun-Yu Ma, Beiduo Chen, Jia-Chen Gu, Zhen-Hua Ling, Wu Guo, Quan Liu 0003, Zhigang Chen 0003, Cong Liu 0006
EMNLP5
2022 Multi-Level Contrastive Learning for Cross-Lingual Alignment
abstract
Cross-language pre-trained models such as multilingual BERT (mBERT) have achieved significant performance in various cross-lingual downstream NLP tasks. This paper proposes a multi-level contrastive learning (ML-CTL) framework to further improve the cross-lingual ability of pre-trained models. The proposed method uses translated parallel data to encourage the model to generate similar semantic embeddings for different languages. However, unlike the sentence-level alignment used in most previous studies, in this paper, we explicitly integrate the word-level information of each pair of parallel sentences into contrastive learning. Moreover, cross-zero noise contrastive estimation (CZ-NCE) loss is proposed to alleviate the impact of the floating-point error in the training process with a small batch size. The proposed method significantly improves the cross-lingual transfer ability of our basic model (mBERT) and outperforms on multiple zero-shot cross-lingual downstream tasks compared to the same-size models in the Xtreme benchmark.
Beiduo Chen, Wu Guo, Bin Gu 0004, Quan Liu 0003
ICASSP2
2022 Feature Aggregation in Zero-Shot Cross-Lingual Transfer Using Multilingual BERT
abstract
Multilingual BERT (mBERT), a language model pre-trained on large multilingual corpora, has impressive zeroshot cross-lingual transfer capabilities and performs surprisingly well on zero-shot POS tagging and Named Entity Recognition (NER), as well as on cross-lingual model transfer. At present, the mainstream methods to solve the cross-lingual downstream tasks are always using the last transformer layer’s output of mBERT as the representation of linguistic information. In this work, we explore the complementary property of lower layers to the last transformer layer of mBERT. A feature aggregation module based on an attention mechanism is proposed to fuse the information contained in different layers of mBERT. The experiments are conducted on four zero-shot cross-lingual transfer datasets, and the proposed method obtains performance improvements on key multilingual benchmark tasks XNLI (+1.5 %), PAWS-X (+2.4 %), NER (+1.2 F1), and POS (+1.5 F1). Through the analysis of the experimental results, we prove that the layers before the last layer of mBERT can provide extra useful information for cross-lingual downstream tasks and explore the interpretability of mBERT empirically.
Beiduo Chen, Wu Guo, Quan Liu 0003, Kun Tao
ICPR2
2022 An Improved Deliberation Network with Text Pre-training for Code-Switching Automatic Speech Recognition
Zhijie Shen, Wu Guo
INTERSPEECH2
2022 Dynamic Convolution With Global-Local Information for Session-Invariant Speaker Representation Learning
abstract
Various mismatchedconditions result in performance degradation of the speaker verification (SV) systems. To address this issue, we extract robust speaker representations by devising a global-local information-based dynamic convolution neural network. In the proposed method, both global and local information of the input features are exploited to dynamically modify the convolution kernel values. This increases the model capability of capturing speaker characteristics by compensating both the inter- and intra-session variabilities. Extensive experiments on four publicly available SV datasets show significant and consistent improvements over the conventional approaches. The effectiveness of the proposed method is further investigated using ablation studies and visualizations.
Bin Gu 0004, Wu Guo
IEEE Signal Process. Lett.2
2021 Topic Classification on Spoken Documents Using Deep Acoustic and Linguistic Features
abstract
Topic classification systems on spoken documents usually consist of two modules: an automatic speech recognition (ASR) module to convert speech into text and a text topic classification (TTC) module to predict the topic class from the decoded text. In this paper, instead of using the ASR transcripts, the fusion of deep acoustic and linguistic features is used for topic classification on spoken documents. More specifically, a conventional CTC-based acoustic model (AM) using phonemes as output units is first trained, and the outputs of the layer before the linear phoneme classifier in the trained AM are used as the deep acoustic features of spoken documents. Furthermore, these deep acoustic features are fed to a phoneme-to-word (P2W) module to obtain deep linguistic features. Finally, a local multi-head attention module is proposed to fuse these two types of deep features for topic classification. Experiments conducted on a subset selected from Switchboard corpus show that our proposed framework outperforms the conventional$\text{ASR}+\text{TTC}$systems and achieves a 3.13% improvement in ACC.
Tan Liu, Wu Guo
ASRU2
2021 Improved Meta-Learning Training for Speaker Verification
abstract
Meta-learning (ML) has recently become a research hotspot in speaker verification (SV).We introduce two methods to improve the meta-learning training for SV in this paper.For the first method, a backbone embedding network is first jointly trained with the conventional cross entropy loss and prototypical networks (PN) loss.Then, inspired by speaker adaptive training in speech recognition, additional transformation coefficients are trained with only the PN loss.The transformation coefficients are used to modify the original backbone embedding network in the x-vector extraction process.Furthermore, the random erasing (RE) data augmentation technique is applied to all support samples in each episode to construct positive pairs, and a contrastive loss between the augmented and the original support samples is added to the objective in model training.Experiments are carried out on the Speaker in the Wild (SITW) and VOiCES databases.Both of the methods can obtain consistent improvements over existing meta-learning training frameworks.By combining these two methods, we can observe further improvements on these two databases.
Yafeng Chen, Wu Guo, Bin Gu 0004
Interspeech2
2021 Bidirectional Multiscale Feature Aggregation for Speaker Verification
abstract
In this paper, we propose a novel bidirectional multiscale feature aggregation (BMFA) network with attentional fusion modules for text-independent speaker verification.The feature maps from different stages of the backbone network are iteratively combined and refined in both a bottom-up and top-down manner.Furthermore, instead of simple concatenation or elementwise addition of feature maps from different stages, an attentional fusion module is designed to compute the fusion weights.Experiments are conducted on the NIST SRE16 and VoxCeleb1 datasets.The experimental results demonstrate the effectiveness of the bidirectional aggregation strategy and show that the proposed attentional fusion module can further improve the performance.
Jiajun Qi, Wu Guo, Bin Gu 0004
Interspeech2
2020 Attention-Based Gated Scaling Adaptive Acoustic Model for CTC-Based Speech Recognition
abstract
In this paper, we propose a novel adaptive technique that uses an attention-based gated scaling (AGS) scheme to improve deep feature learning for connectionist temporal classification (CTC) acoustic modeling. In AGS, the outputs of each hidden layer of the main network are scaled by an auxiliary gate matrix extracted from the lower layer by using an attention mechanism. Furthermore, the auxiliary AGS layer and the main network are jointly trained without requiring second-pass model training or additional speaker information, such as i-vector. On the Mandarin AISHELL-1 dataset, the proposed AGS yields a 7.94% character error rate (CER). To the best of our knowledge, the results obtained when training on the full AISHELL-1 training set, are the best published currently for the end-to-end systems.
Fenglin Ding, Wu Guo, Li-Rong Dai 0001, Jun Du 0002
ICASSP2
2020 An Improved Deep Neural Network for Modeling Speaker Characteristics at Different Temporal Scales
abstract
This paper presents an improved deep embedding learning method based on a convolutional neural network (CNN) for text-independent speaker verification. Two improvements are proposed for x-vector embedding learning: (1) a multiscale convolution (MSCNN) is adopted in the frame-level layers to capture the complementary speaker information in different receptive fields; (2) a Baum-Welch statistics attention (BWSA) mechanism is applied in the pooling layer, which can integrate more useful long-term speaker characteristics in the temporal pooling layer. Experiments are carried out on the NIST SRE16 evaluation set. The results demonstrate the effectiveness of the MSCNN and show that the proposed BWSA can further improve the performance of the DNN embedding system.
Bin Gu 0004, Wu Guo, Li-Rong Dai 0001, Jun Du 0002
ICASSP2
2020 Unsupervised Regularization-Based Adaptive Training for Speech Recognition
Fenglin Ding, Wu Guo, Bin Gu 0004, Zhen-Hua Ling, Jun Du 0002
INTERSPEECH2
2020 Adaptive Speaker Normalization for CTC-Based Speech Recognition
Fenglin Ding, Wu Guo, Bin Gu 0004, Zhen-Hua Ling, Jun Du 0002
INTERSPEECH2
2020 An Adaptive X-Vector Model for Text-Independent Speaker Verification
abstract
In this paper, adaptive mechanisms are applied in deep neural network (DNN) training for x-vector-based text-independent speaker verification.First, adaptive convolutional neural networks (ACNNs) are employed in frame-level embedding layers, where the parameters of the convolution filters are adjusted based on the input features.Compared with conventional CNNs, ACNNs have more flexibility in capturing speaker information.Moreover, we replace conventional batch normalization (BN) with adaptive batch normalization (ABN).By dynamically generating the scaling and shifting parameters in BN, ABN adapts models to the acoustic variability arising from various factors such as channel and environmental noises.Finally, we incorporate these two methods to further improve performance.Experiments are carried out on the speaker in the wild (SITW) and VOiCES databases.The results demonstrate that the proposed methods significantly outperform the original xvector approach.
Bin Gu 0004, Wu Guo, Fenglin Ding, Zhen-Hua Ling, Jun Du 0002
INTERSPEECH2
2019 Topic Detection in Conversational Telephone Speech Using CNN with Multi-stream Inputs
abstract
Topic detection for conversational telephone speech (CTS) is addressed in this paper. The low accuracy of automatic speech recognition (ASR) will cause severe performance deterioration for topic detection. To make up for this, we adopt two ASR systems, HMM-BiLSTM and CTC systems, to provide complementary information for topic detection. After obtaining two sets of different recognized transcriptions, a CNN with multi-stream inputs is trained, and the pooling layer serves as document representations. Finally, element-wise summation of document representations from two streams is used as distributed representations of the documents, which are fed into agglomerative hierarchical clustering (AHC) algorithms to obtain clustering results. The experiments on a Japanese speech corpus demonstrate that the proposed approach can significantly improve the performance of topic detection.
Wu Guo, Yan Song 0001
ICASSP2
2019 A Region Based Attention Method for Weakly Supervised Sound Event Detection and Classification
abstract
Recently, an attention based convolutional recurrent neural network (CRNN) with learnable gated linear units (GLUs) has achieved state-of-the-art performance for audio tagging (AT) and sound event detection (SED) tasks in the Detection and Classification of Acoustic Scenes and Events (DCASE) challenges. The introduction of GLU and temporal attention-based localization mechanisms plays an important role for both AT and SED tasks. In this paper, we propose a novel region based attention method to further boost the representation power of the existing GLU based CRNN. Specifically, we insert a feature selection (FS) structure after each GLU to create what we term a GLU-F. block, to exploit channel relationships. Furthermore, we extract region features (or the prototypes of certain sound events) from multi-scale sliding windows over higher convolutional layers, which are fed into an attention-based recurrent neural network to model their context information for AT and SED tasks. To evaluate the proposed region based attention method, we conduct extensive experiments on SED and AT tasks in DCASE2017. We achieve 59.5% and 60.1% AT F1-score, 51.3% and 55.1% SED F1-score for development and evaluation sets respectively, significantly outperforming state-of-the-art results.
Yan Song 0001, Wu Guo, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP3
2019 Neural Text Clustering with Document-Level Attention Based on Dynamic Soft Labels
Wu Guo, Li-Rong Dai 0001, Zhen-Hua Ling, Jun Du 0002
INTERSPEECH2
2019 Multi-Task Learning with High-Order Statistics for x-Vector Based Text-Independent Speaker Verification
abstract
The x-vector based deep neural network (DNN) embedding systems have demonstrated effectiveness for text-independent speaker verification.This paper presents a multi-task learning architecture for training the speaker embedding DNN with the primary task of classifying the target speakers, and the auxiliary task of reconstructing the first-and higher-order statistics of the original input utterance.The proposed training strategy aggregates both the supervised and unsupervised learning into one framework to make the speaker embeddings more discriminative and robust.Experiments are carried out using the NIST SRE16 evaluation dataset and the VOiCES dataset.The results demonstrate that our proposed method outperforms the original x-vector approach with very low additional complexity added.
Lanhua You, Wu Guo, Li-Rong Dai 0001, Jun Du 0002
INTERSPEECH2
2019 Deep Neural Network Embeddings with Gating Mechanisms for Text-Independent Speaker Verification
abstract
In this paper, gating mechanisms are applied in deep neural network (DNN) training for x-vector-based text-independent speaker verification. First, a gated convolution neural network (GCNN) is employed for modeling the frame-level embedding layers. Compared with the time-delay DNN (TDNN), the GCNN can obtain more expressive frame-level representations through carefully designed memory cell and gating mechanisms. Moreover, we propose a novel gated-attention statistics pooling strategy in which the attention scores are shared with the output gate. The gated-attention statistics pooling combines both gating and attention mechanisms into one framework; therefore, we can capture more useful information in the temporal pooling layer. Experiments are carried out using the NIST SRE16 and SRE18 evaluation datasets. The results demonstrate the effectiveness of the GCNN and show that the proposed gated-attention statistics pooling can further improve the performance.
Lanhua You, Wu Guo, Li-Rong Dai 0001, Jun Du 0002
INTERSPEECH2
2018 Pseudo-Supervised Approach for Text Clustering Based on Consensus Analysis
abstract
In recent years, neural networks (NN) have achieved remarkable performance improvement in text classification due to their powerful ability to encode discriminative features by incorporating label information into model training. Inspired by the success of NN in text classification, we propose a pseudo-supervised neural network approach for text clustering. The neural network is trained in a supervised fashion with pseudo-labels, which are provided by the cluster labels of pre-clustering on unsupervised document representations. To enhance the quality of pseudo-labels, a consensus analysis is employed to select training samples for the neural network. The experimental results demonstrate that the proposed approach can improve the clustering performance significantly.
Peixin Chen, Wu Guo, Li-Rong Dai 0001, Zhen-Hua Ling
ICASSP2
2018 Gated Convolutional Neural Network for Sentence Matching
Peixin Chen, Wu Guo, Lanhua You
INTERSPEECH2
2018 An Improved Deep Embedding Learning Method for Short Duration Speaker Verification
abstract
This paper presents an improved deep embedding learning method based on convolutional neural networks (CNN) for short-duration speaker verification (SV). Existing deep learning-based SV methods generally extract frontend embeddings from a feed-forward deep neural network, in which the long-term speaker characteristics are captured via a pooling operation over the input speech. The extracted embeddings are then scored via a backend model, such as Probabilistic Linear Discriminative Analysis (PLDA). Two improvements are proposed for frontend embedding learning based on the CNN structure: (1) Motivated by the WaveNet for speech synthesis, dilated filters are designed to achieve a tradeoff between computational efficiency and receptive-filter size; and (2) A novel cross-convolutional-layer pooling method is exploited to capture $1^{st}$-order statistics for modelling long-term speaker characteristics. Specifically, the activations of one convolutional layer are aggregated with the guidance of the feature maps from the successive layer. To evaluate the effectiveness of our proposed methods, extensive experiments are conducted on the modified female portion of NIST SRE 2010 evaluations, with conditions ranging from 10s-10s to 5s-4s. Excellent performance has been achieved on each evaluation condition, significantly outperforming existing SV systems using i-vector and d-vector embeddings.
Zhifu Gao, Yan Song 0001, Ian McLoughlin 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH4
2018 An Attention Pooling Based Representation Learning Method for Speech Emotion Recognition
abstract
This paper proposes an attention pooling based representation learning method for speech emotion recognition (SER). The emotional representation is learned in an end-to-end fashion by applying a deep convolutional neural network (CNN) directly to spectrograms extracted from speech utterances. Motivated by the success of GoogleNet, two groups of filters with different shapes are designed to capture both temporal and frequency domain context information from the input spectrogram. The learned features are concatenated and fed into the subsequent convolutional layers. To learn the final emotional representation, a novel attention pooling method is further proposed. Compared with the existing pooling methods, such as max-pooling and average-pooling, the proposed attention pooling can effectively incorporate class-agnostic bottom-up, and class-specific top-down, attention maps. We conduct extensive evaluations on benchmark IEMOCAP data to assess the effectiveness of the proposed representation. Results demonstrate a recognition performance of 71.8% weighted accuracy (WA) and 68% unweighted accuracy (UA) over four emotions, which outperforms the state-of-the-art method by about 3% absolute for WA and 4% for UA.
Yan Song 0001, Ian McLoughlin 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH4
2018 Improved Supervised Locality Preserving Projection for I-vector Based Speaker Verification
Lanhua You, Wu Guo, Yan Song 0001
INTERSPEECH2
2017 Exploring universal speech attributes for speaker verification
abstract
The universal speech attributes for speaker verification (SV) are addressed in this paper. The aim of this work is to exploit fundamental characteristics across different speakers within the deep neural network (DNN)/i-vector framework. The manner and place of articulation form the fundamental speech attribute unit inventory, and new attribute units for acoustic modelling are generated by a two-step automatic clustering method in this paper. The DNN based on universal attribute units is used to generate posterior probability in total variability modelling and i-vector extracting for the speaker recognition procedure. Furthermore, Gaussian mixture models (GMMs) are used to fit the distribution of the features associated with a given context-dependent attribute unit to improve performance. The experiments are carried out on the core test from the NIST SRE 2008 corpus; the proposed system can obtain better performance than all other state-of-the-art systems.
Wu Guo
ICASSP2
2017 Feature mapping for speaker diarization in noisy conditions
abstract
Speaker diarization in noisy conditions is addressed in this paper. The regression-based DNN is first adopted to map the noisy acoustic features to the clean features, and then consensus clustering of the original and mapped features is used to fuse the diarization results. The experiments are conducted on the IFLY-DIAR-II database, which is a Chinese talk show database with various noise types, such as music, applause and laughter. Compared to the baseline system using PLP features, a 21.26% relative DER improvement can be achieved using the proposed algorithm.
Weixin Zhu, Wu Guo
ICASSP2
2017 End-to-End Language Identification Using High-Order Utterance Representation with Bilinear Pooling
abstract
A key problem in spoken language identification (LID) is how to design effective representations which are specific to language information. Recent advances in deep neural networks have led to significant improvements in results, with deep end-to-end methods proving effective. This paper proposes a novel network which aims to model an effective representation for high (first and second)-order statistics of LID-senones, defined as being LID analogues of senones in speech recognition. The high-order information extracted through bilinear pooling is robust to speakers, channels and background noise. Evaluation with NIST LRE 2009 shows improved performance compared to current state-of-the-art DBF/i-vector systems, achieving over 33% and 20% relative equal error rate (EER) improvement for 3s and 10s utterances and over 40% relative Cavg improvement for all durations.
Ma Jin, Yan Song 0001, Ian McLoughlin 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH4
2016 Web Data Selection Based on Word Embedding for Low-Resource Speech Recognition
Chuandong Xie, Wu Guo
INTERSPEECH2
2016 Intra-Topic Variability Normalization based on Linear Projection for Topic Classification
abstract
This paper proposes a variability normalization algorithm to reduce the variability between intra-topic documents for topic classification.Firstly, an optimization problem is constructed based on linear variability removable assumption.Secondly, a new feature space for document representation is found by solving the optimization problem with kernel principle component analysis (KPCA).Finally, effective feature transformation is taken through linear projection.As for experiments, state-of-the-art SVM and KNN algorithm are adopted for topic classification respectively.Experimental results on a free-style conversational corpus show that the proposed variability normalization algorithm for topic classification achieves 3.8% absolute improvement for micro-F 1 measure.
Quan Liu 0003, Wu Guo, Zhen-Hua Ling, Hui Jiang 0001, Yu Hu 0003
HLT-NAACL2
2015 Channel adaptation of plda for text-independent speaker verification
abstract
Probabilistic linear discriminant analysis (PLDA) has shown to be effective for modeling channel variability in the i-vector space for text-independent speaker verification. Speaker verification is a binary hypothesis testing. Given a test segment, the verification score could be computed as the log-likelihood ratio between a speaker-adapted PLDA and the universal PLDA model. This work proposes to infer the channel factor specific to each test segment and to include the channel estimate in the PLDA models, which essentially shifts the scoring function to better match that of the test channel. We also explore the influence of covariance adaptation in both speaker and channel adaptations. Experimental results on NIST SRE'08 and SRE'10 dataset confirm that the proposed channel adaptation can be effective when the covariance is kept un-adapted, while the covariance adaptation is necessary in the speaker adaptation.
Kong-Aik Lee, Bin Ma 0001, Wu Guo, Haizhou Li 0001, Li-Rong Dai 0001
ICASSP4
2015 Phone-centric local variability vector for text-constrained speaker verification
Kong-Aik Lee, Bin Ma 0001, Wu Guo, Haizhou Li 0001, Li-Rong Dai 0001
INTERSPEECH4
2014 Minimum divergence estimation of speaker prior in multi-session PLDA scoring
abstract
Probabilistic linear discriminant analysis (PLDA) has shown to be effective for modeling speaker and channel variability in the i-vector space for text-independent speaker verification. This paper shows that the PLDA scoring function could be formulated as model comparison between an adapted PLDA model and the universal PLDA. Based on this formulation, we show that a more robust adaptation could be attained by adapting the PLDA model through the use of minimum divergence estimate of speaker prior in the latent subspace. Experimental results on NIST SRE'10 and SRE'12 dataset confirm that the proposed method is effective in handling multi-session task. Notably, it is free from the covariance shrinkage problem typically found in the standard multi-session PLDA scoring.
Kong-Aik Lee, Bin Ma 0001, Wu Guo, Haizhou Li 0001, Li-Rong Dai 0001
ICASSP4
2014 Lattice based optimization of bottleneck feature extractor with linear transformation
abstract
This paper proposes a lattice-based sequential discriminative training method to extract more discriminative bottleneck features. In our method, the bottleneck neural network is first trained with cross entropy criteria, and then only the weights of bottleneck layer are retrained with sequential criteria. If the outputs of the layer before bottleneck are treated as the raw features, the new method is an equivalent to a linear feature transformation algorithm. This linearity makes the optimization much easier than updating the whole neural network. Just like the fMPE and RDLT, the neural network is retrained with batch mode gradient descent, making the training to be easily implemented in parallel. Meanwhile, batch mode optimization can naturally deal with the indirect gradient to make the optimization more precise. Experimental results on a Mandarin transcription task and the Switchboard task have shown the effectiveness of the proposed method with the CER decreases from 12.2% to 11.3% and the WER from 16.1% to 15.0%, respectively.
Diyuan Liu, Si Wei, Wu Guo, Yebo Bao, Shifu Xiong, Li-Rong Dai 0001
ICASSP3
2013 Phoneme variation based synthesized speech discrimination for speaker verification
abstract
How to discriminate the synthesized speech from the natural speech for speaker verification is addressed in this paper. With the development of HMM-based speech synthesis, it is easy to obtain high quality synthesized speech which sounds like target speaker, the robustness of synthesized speech become important for speaker verification. In this paper, a method based on the phoneme variation is proposed to discriminate synthesized speech from natural speech, which could be used as front-end module to detect the synthesized speech for speaker verification system. The experimental results show the effectiveness of the proposed method.
LianWu Chen, Wu Guo, Yan Song 0001, Li-Rong Dai 0001
ICASSP2
2012 Exemplar-Based Sparse Representation for Language Recognition on I-Vectors
abstract
In this paper, a new automatic language identification method using sparse representation on i-vectors in low-dimensional total variability space is proposed. It is mainly based on the recently proposed i-vector based language recognition systems. In our proposed method, an over-complete dictionary is first constructed by randomly sampling of the low-dimensional total variability space after Within-Class Covariance Normalization (WCCN) and Linear Discriminate Analysis (LDA). And then for each test sample, the classification score is derived from sparse linear representation with respect to the over-complete dictionary. Furthermore, a random subspace method, which combines different sparse representation classifiers, is introduced to address the possible over-fitting issue. Evaluations on NIST LRE 2007 dataset show that the proposed method outperforms the state-of-the-art i-vector based language recognition system. Especially for 30s test condition, our proposed method achieves relative reduction of 29.6 % on Equal Error Rate (EER) compared with the baseline system.
Bing Jiang, Yan Song 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH3
2011 Speaker characterization using spectral subband energy ratio based on Harmonic plus Noise Model
abstract
This paper proposes a feature extraction for speaker characterization by exploring the relationship between the two distinct components of the speech signal, one is harmonics accounting for the periodicity of the signal and the other is modulated noise accounting for the turbulences of the glottal airflow. The harmonic and noise parts of the speech signal are decomposed based on the Harmonic plus Noise Model approach. We estimate the spectral subband energy ratios (SSERs) as the speaker characteristic features, which are expected to reflect the interaction property of the vocal tract and glottal airflow of individual speakers for speaker verification. The speaker verification experiments based on a GMM-UBM system have shown the efficiency of the SSER features, reducing the error equal rate by 27.2% by combining with the conventional MFCC features.
Yanhua Long, Zhijie Yan, Frank K. Soong, Li-Rong Dai 0001, Wu Guo
ICASSP5
2011 Factored covariance modeling for text-independent speaker verification
abstract
Gaussian mixture models (GMMs) are commonly used to model the spectral distribution of speech signals for text-independent speaker verification. Mean vectors of the GMM, used in conjunction with support vector machine (SVM), have shown to be effective in characterizing speaker information. In addition to the mean vectors, covariance matrices capture the correlation between spectral features, which also represent some salient information about speaker identity. This paper investigates the use of local correlation between different dimensions of acoustic vector by using factor analysis and linear Gaussian model. Log-Euclidean inner product kernel is used to measure the similarity between two speech utterances in the form of covariance matrices. Experiments carried on NIST 2006 speaker verification tasks shows promising results.
Eryu Wang, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001, Wu Guo, Li-Rong Dai 0001
ICASSP5
2011 Improvements in Speaker Characterization Using Spectral Subband Energy Based on Harmonic plus Noise Model
Yanhua Long, Zhijie Yan, Frank K. Soong, Li-Rong Dai 0001, Wu Guo
INTERSPEECH5
2010 N-gram nearest neighbor algorithm for voice password system
abstract
A specific issue in the voice password system is addressed in this paper: When the text content of target speaker's enrollment password has been already known by imposters, they can do a well-behaved impersonation using the same text content as the target speaker. This results in a much higher false acceptance than the traditional voice password system. N-gram based nearest neighbor algorithm is proposed here to improve the speaker detection accuracy. Furthermore, correlation coefficient is adopted as the distance measurement between two acoustic features instead of the traditional Euclidean distance. Experimental results show that the proposed method outperforms the DTW and GMM-UBM algorithms.
Wu Guo, Yanhua Long, Li-Rong Dai 0001
ICASSP1
2010 Effects of the phonological relevance in speaker verification
Yanhua Long, Li-Rong Dai 0001, Bin Ma 0001, Wu Guo
INTERSPEECH4
2010 The estimation and kernel metric of spectral correlation for text-independent speaker verification
Eryu Wang, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH5
2009 iFLY system for the NIST 2008 speaker recognition evaluation
abstract
The description of iFLY system submitted for NIST 2008 speaker recognition evaluation (SRE), which has achieved excellent performance in the 2008 SRE evaluation, is presented in this paper. Our primary system is a fusion of two subsystems GMM-UBM and GMM-SVM. For each sub-system, two kinds of short-time acoustic features PLP and LPCC are adopted. We focus on three key issues in this evaluation: channel compensation, multi-lingual or bi-lingual cues and the voice activity detection. We also point out that data selection and factor analysis play key roles in the system improvement.
Wu Guo, Yanhua Long, Yijie Li 0001, Eryu Wang, Li-Rong Dai 0001
ICASSP1
2009 The I4U system in NIST 2008 speaker recognition evaluation
abstract
This paper describes the performance of the I4U speaker recognition system in the NIST 2008 Speaker Recognition Evaluation. The system consists of seven subsystems, each with different cepstral features and classifiers. We describe the I4U Primary system and report on its core test results as they were submitted, which were among the best-performing submissions. The I4U effort was led by the Institute for Infocomm Research, Singapore (IIR), with contributions from the University of Science and Technology of China (USTC), the University of New South Wales, Australia (UNSW), Nanyang Technological University, Singapore (NTU) and Carnegie Mellon University, USA (CMU).
Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee, Hanwu Sun, Donglai Zhu, Khe Chai Sim, Chang Huai You, Rong Tong, Ismo Kärkkäinen, Chien-Lin Huang, Vladimir Pervouchine, Wu Guo, Yijie Li 0001, Li-Rong Dai 0001, Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Chng Eng Siong, Tanja Schultz, Qin Jin
ICASSP12
2009 Exploiting prosodic information for Speaker Recognition
abstract
In this paper, we study speaker characterization using prosodic supervectors with negative within-class covariance normalization (NWCCN) projection and speaker modeling with support vector regression (SVR). We also propose a segmental weight fusion (SWF) technique that combines acoustic and prosodic subsystems effectively, despite the big performance gap between the subsystems. We validate the effectiveness of our proposed techniques on the NIST 2006 Speaker Recognition Evaluation (SRE) in comparison with other prominent solutions. The experiments have reported competitive results of 17.72% Equal Error Rate for the prosodic subsystem alone and 4.50% for the fusion system on NIST 2006 SRE core test condition.
Yanhua Long, Bin Ma 0001, Haizhou Li 0001, Wu Guo, Chng Eng Siong, Li-Rong Dai 0001
ICASSP4
2007 Hilbert-Huang Transform-based Local Regions Descriptors
abstract
This paper presents a new interest local regions descriptors method based on Hilbert-Huang Transform. The neighborhood of the interest local region is decomposed adaptively into oscillatory components called intrinsic mode functions (IMFs). Then the Hilbert transform is applied to each component and get the phase and amplitude information. The proposed descriptors samples the phase angles information and amalgamates them into 10 overlap squares with 8-bin orientation histograms. The experiments show that the proposed descriptors are better than SIFT and other standard descriptors. Essentially, the Hilbert-Huang Transform based descriptors can belong to the class of phase-based descriptors. So it can provides a better way to overcome the illumination changes. Additionally, the Hilbert-Huang transform is a new tool for analyzing signals and the proposed descriptors is a new attempt to the Hilbert-Huang transform. 1
Dongfeng Han, Wenhui Li 0002, Wu Guo
BMVC3
2006 Minimum generation error criterion for tree-based clustering of context dependent HMMs
Yi-Jian Wu, Wu Guo, Renhua Wang
INTERSPEECH2