Meng Liu 0017

dblp:41/7841-17 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
13since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2023 Self-Supervised Audio-Visual Speaker Representation with Co-Meta Learning
abstract
In self-supervised speaker verification, the quality of pseudo labels determines the upper bound of its performance and it is not uncommon to end up with massive amount of unreliable pseudo labels. We observe that the complementary information in different modalities ensures a robust supervisory signal for audio and visual representation learning. This motivates us to propose an audio-visual self-supervised learning framework named Co-Meta Learning. Inspired by the Coteaching+, we design a strategy that allows the information of two modalities to be coordinated through the Update by Disagreement. Moreover, we use the idea of modelagnostic meta learning (MAML) to update the network parameters, which makes the hard samples of two modalities to be better resolved by the other modality through gradient regularization. Compared to the baseline, our proposed method achieves a 29.8%, 11.7% and 12.9% relative improvement on Vox-O, Vox-E and Vox-H trials of Voxceleb1 evaluation dataset respectively.
Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001
ICASSP5
2023 Leveraging Positional-Related Local-Global Dependency for Synthetic Speech Detection
abstract
Automatic speaker verification (ASV) systems are vulnerable to spoofing attacks. As synthetic speech exhibits local and global artifacts compared to natural speech, incorporating local-global dependency would lead to better anti-spoofing performance. To this end, we propose the Rawformer that leverages positional-related local-global dependency for synthetic speech detection. The two-dimensional convolution and Transformer are used in our method to capture local and global dependency, respectively. Specifically, we design a novel positional aggregator that integrates local-global dependency by adding positional information and flattening strategy with less information loss. Furthermore, we propose the squeeze-and-excitation Rawformer (SE-Rawformer), which introduces squeeze-and-excitation operation to acquire local dependency better. The results demonstrate that our proposed SE-Rawformer leads to 37% relative improvement compared to the single state-of-the-art system on ASVspoof 2019 LA and generalizes well on ASVspoof 2021 LA. Especially, using the positional aggregator in the SE-Rawformer brings a 43% improvement on average.
Meng Liu 0017, Longbiao Wang, Kong-Aik Lee, Hanyi Zhang, Jianwu Dang 0001
ICASSP2
2023 Cross-Modal Audio-Visual Co-Learning for Text-Independent Speaker Verification
abstract
Visual speech (i.e., lip motion) is highly related to auditory speech due to the co-occurrence and synchronization in speech production. This paper investigates this correlation and proposes a cross-modal speech co-learning paradigm. The primary motivation of our cross-modal co-learning method is modeling one modality aided by exploiting knowledge from another modality. Specifically, two cross-modal boosters are introduced based on an audio-visual pseudo-siamese structure to learn the modality-transformed correlation. Inside each booster, a max-feature-map embedded Transformer variant is proposed for modality alignment and enhanced feature generation. The network is co-learned both from scratch and with pretrained models. Experimental results on the test scenarios demonstrate that our proposed method achieves around 60% and 20% average relative performance improvement over baseline unimodal and fusion systems, respectively.
Meng Liu 0017, Kong-Aik Lee, Longbiao Wang, Hanyi Zhang, Chang Zeng, Jianwu Dang 0001
ICASSP1
2023 Noise-Disentanglement Metric Learning for Robust Speaker Verification
abstract
Automatic speaker verification (ASV) suffers from performance degradation in noisy environments. To solve this problem, we propose the noise-disentanglement metric learning to reduce the speaker-irrelevant noisy components and build a noise-invariant embedding space. Specifically, the disentanglement module, including the speaker encoder and re-construction module, is dedicated to decoupling speech signals. The speaker encoder is used to disentangle speaker-related components, and the reconstruction module increases the model’s ability to constrain the noise information by re-constructing the signal. In addition, distribution optimization is introduced to supervise the spatial structure of speaker embeddings under noisy environments. Experiments on Vox-Celeb1 indicate that the proposed method improves the performance of the speaker verification system in both clean and noisy conditions.
Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001
ICASSP5
2023 Meta-Generalization for Domain-Invariant Speaker Verification
abstract
Automatic speaker verification (ASV) exhibits unsatisfactory performance under domain mismatch conditions owing to intrinsic and extrinsic factors, such as variations in speaking styles and recording devices encountered in real-world applications. To ensure robust performance under unseen conditions, domain generalization has been explored. However, an inherent contradiction exists between model discrimination and domain generalization, in which the discrimination ability may be reduced while learning to generalize. In this paper, to extract discriminative yet domain-invariant representations, we propose the meta-generalized speaker verification (MGSV) via meta-learning. Specifically, we propose a metric-based distribution optimization and a gradient-based meta-optimization to simultaneously supervise the spatial relationship between embeddings and improve the generalization ability of the model on unseen domains. In addition, we design multiple-single (MS) and simulated speaker verification (SSV) sampling strategies based on single-domain (SD) and single-single (SS) strategies to simulate the train/test domain mismatch more relevantly, thereby mining transferable speaker-related knowledge. SSV is chosen as the most effective method, as it substantially improves the domain generalization by ensuring that the model has learned to discriminate efficiently. Additionally, to intuitively reflect the model performance on the unseen domains, the proposed method is validated on cross-genre, cross-device, and cross-dataset tasks. The experimental results demonstrate that our proposed method achieves remarkable performance in handling domain mismatch issues in speaker verification.
Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001, Helen M. Meng
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Learning Domain-Invariant Transformation for Speaker Verification
abstract
Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors such as recording device and speaking style in real-world applications, which leads to unsatisfactory performance. To this end, we propose the meta generalized transformation via meta-learning to build a domain-invariant embedding space. Specifically, the transformation module is motivated to learn the domain generalization knowledge by executing meta-optimization on the meta-train and meta-test sets which are designed to simulate domain shift. Furthermore, distribution optimization is incorporated to supervise the metric structure of embeddings. In terms of the transformation module, we investigate various instantiations and observe the multilayer perceptron with gating (gMLP) is the most effective given its extrapolation capability. The experimental results on cross-genre and cross-dataset settings demonstrate that the meta generalized transformation dramatically improves the robustness of ASV systems to domain shift, while outperforms the state-of-the-art methods.
Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001
ICASSP4
2022 Data Augmentation Using McAdams-Coefficient-Based Speaker Anonymization for Fake Audio Detection
abstract
Fake audio detection (FAD) is a technique to distinguish synthetic speech from natural speech.In most FAD systems, removing irrelevant features from acoustic speech while keeping only robust discriminative features is essential.Intuitively, speaker information entangled in acoustic speech should be suppressed for the FAD task.Particularly in a deep neural network (DNN)-based FAD system, the learning system may learn speaker information from a training dataset and cannot generalize well on a testing dataset.In this paper, we propose to use the speaker anonymization (SA) technique to suppress speaker information from acoustic speech before inputting it into a DNN-based FAD system.We adopted the McAdamscoefficient-based SA (MC-SA) algorithm, and this is expected that the entangled speaker information will not be involved in the DNN-based FAD learning.Based on this idea, we implemented a light convolutional neural network bidirectional long short-term memory (LCNN-BLSTM)-based FAD system and conducted experiments on the Audio Deep Synthesis Detection Challenge (ADD2022) datasets.The results showed that removing the speaker information from acoustic speech improved the relative performance in the first track of ADD2022 by 17.66%.
Kai Li 0018, Sheng Li 0010, Xugang Lu, Masato Akagi, Meng Liu 0017, Lin Zhang 0054, Chang Zeng, Longbiao Wang, Jianwu Dang 0001, Masashi Unoki
INTERSPEECH5
2022 Spoofing-Aware Attention based ASV Back-end with Multiple Enrollment Utterances and a Sampling Strategy for the SASV Challenge 2022
abstract
Current state-of-the-art automatic speaker verification (ASV) systems are vulnerable to presentation attacks, and several countermeasures (CMs), which distinguish bona fide trials from spoofing ones, have been explored to protect ASV. However, ASV systems and CMs are generally developed and optimized independently without considering their inter-relationship. In this paper, we propose a new spoofing-aware ASV back-end module that efficiently computes a combined ASV score based on speaker similarity and CM score. In addition to the learnable fusion function of the two scores, the proposed back-end module has two types of attention components, scaled-dot and feed-forward self-attention, so that intra-relationship information of multiple enrollment utterances can also be learned at the same time. Moreover, a new effective trials-sampling strategy is designed for simulating new spoofing-aware verification scenarios introduced in the Spoof-Aware Speaker Verification (SASV) challenge 2022.
Chang Zeng, Lin Zhang 0054, Meng Liu 0017, Junichi Yamagishi
INTERSPEECH3
2021 DeepLip: A Benchmark for Deep Learning-Based Audio-Visual Lip Biometrics
abstract
Audio-visual lip biometrics (AV-LB) has been an emerging biometrics technology that straddles auditory and visual speech processing. Previous works mainly focused on the front-end lip-based feature engineering combined with a shallow statistical back-end model. Over the past decade, convolutional neural network (CNN, or ConvNet) has been widely used and achieved good performance in computer vision and speech processing tasks. However, the lack of a sizeable public AV-LB database led to a stagnation in deep-learning exploration on AV-LB tasks. In addition to the dual audio-visual streams, one essential requirement on the video stream is the region of interest (ROI) around the lips has to be of sufficient resolution. To this end, we compile a moderate-size database using existing public databases. Using this database, we present a deep learning-based AV-LB benchmark, dubbed DeepLip11https://github.com/DanielMengLiu/DeepLip, realized with convolutional video and audio unimodal modules, and a multimodal fusion module. Our experiments show that DeepLip outperforms the traditional lip-biometrics system in context modeling and achieves over 50% relative improvements compared with its unimodal system, with an equal error rate of 0.75% and 1.11% on the test datasets, respectively.
Meng Liu 0017, Longbiao Wang, Kong-Aik Lee, Hanyi Zhang, Chang Zeng, Jianwu Dang 0001
ASRU1
2021 Replay-Attack Detection Using Features With Adaptive Spectro-Temporal Resolution
abstract
Variable-resolution processing aims to improve the feature representation ability by enlarging the local discriminative details. In previous anti-spoofing studies, different phones and frequency regions were both proven to have various levels of sensitivity to replay distortion. In this paper, an adaptive spectro-temporal resolution is proposed to obtain the optimal scale in the feature space: the frequency resolution is adaptive to frequency discrimination, while the temporal resolution is adaptive to continuous phones. In the process, phone-frequency F-ratio analysis is applied to investigate the sensitivity divergences to replay distortion among phones and frequencies. Then, attentive filters are designed to automatically adapt to the phone-frequency discrimination. Validation experiments for the proposed method are conducted on two well-acknowledged magnitude and phase features. A comparative analysis on the ASVspoof 2017 V2.0 database demonstrates that our proposed adaptive spectro-temporal resolution method attains considerably higher error reduction rates than the approaches involving the corresponding original resolution features.
Meng Liu 0017, Longbiao Wang, Kong-Aik Lee, Xuanda Chen, Jianwu Dang 0001
ICASSP1
2021 Meta-Learning for Cross-Channel Speaker Verification
abstract
Automatic speaker verification (ASV) has been successfully deployed for identity recognition. With increasing use of ASV technology in real-world applications, channel mismatch caused by the recording devices and environments severely degrade its performance, especially in the case of unseen channels. To this end, we propose a meta speaker embedding network (MSEN) via meta-learning to generate channel-invariant utterance embeddings. Specifically, we optimize the differences between the embeddings of a support set and a query set in order to learn a channel-invariant embedding space for utterances. Furthermore, we incorporate distribution optimization (DO) to stabilize the performance of MSEN. To quantitatively measure the effect of MSEN on unseen channels, we specially design the generalized cross-channel (GCC) evaluation. The experimental results on the HI-MIA corpus demonstrate that the proposed MSEN reduce considerably the impact of channel mismatch, while significantly outperforms other state-of-the-art methods.
Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001
ICASSP4
2021 Joint Feature Enhancement and Speaker Recognition with Multi-Objective Task-Oriented Network
Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001
Interspeech4
2021 Replay attack detection using variable-frequency resolution phase and magnitude features
Meng Liu 0017, Longbiao Wang, Jianwu Dang 0001, Kong-Aik Lee, Seiichi Nakagawa
Comput. Speech Lang.1
2020 Deep Discriminative Embedding with Ranked Weight for Speaker Verification
Dao Zhou, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001
ICONIP (5)4
2020 ARET: Aggregated Residual Extended Time-Delay Neural Networks for Speaker Verification
Ruiteng Zhang, Jianguo Wei, Wenhuan Lu, Longbiao Wang, Meng Liu 0017, Lin Zhang 0054, Jiayu Jin, Junhai Xu
INTERSPEECH5
2020 Adversarial Separation Network for Speaker Recognition
Hanyi Zhang, Longbiao Wang, Yunchun Zhang, Meng Liu 0017, Kong-Aik Lee, Jianguo Wei
INTERSPEECH4
2020 Dynamic Margin Softmax Loss for Speaker Verification
Dao Zhou, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001, Jianguo Wei
INTERSPEECH5
2019 Replay Attack Detection Using Magnitude and Phase Information with Attention-based Adaptive Filters
abstract
Automatic Speech Verification (ASV) systems are highly vulnerable to spoofing attacks, and replay attack poses the greatest threat among various spoofing attacks. In this paper, we propose a novel multi-channel feature extraction method with attention-based adaptive filters (AAF). Original phase information, discarded by conventional feature extraction techniques after Fast Fourier Transform (FFT), is promising in distinguishing genuine from replay spoofed speech. Accordingly, phase and magnitude information are respectively extracted as phase channel and magnitude channel complementary features in our system. First, we make discriminative ability analysis on full frequency bands with F-ratio methods. Then attention-based adaptive filters are implemented to maximize capturing of high discriminative information on frequency bands, and the results on ASVspoof 2017 challenge indicate that our proposed approach achieved relative error reduction rates of 78.7% and 59.8% on development and evaluation dataset than the baseline method.
Meng Liu 0017, Longbiao Wang, Jianwu Dang 0001, Seiichi Nakagawa, Haotian Guan, Xiangang Li
ICASSP1
2018 Multiple Phase Information Combination for Replay Attacks Detection
Dongbo Li, Longbiao Wang, Jianwu Dang 0001, Meng Liu 0017, Zeyan Oo, Seiichi Nakagawa, Haotian Guan, Xiangang Li
INTERSPEECH4