Lantian Li

dblp:153/0735 · DBLP profile ↗
← Back
46ranked-venue papers
15as first author
28since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 10 first-author · 20 since 2021Artificial intelligence and machine learning · 29 · 8 first-author · 18 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge
Zehua Liu, Xiaolou Li, Chen Chen 0075, Lantian Li, Dong Wang 0013
INTERSPEECH4
2025 E-code: Mastering efficient code generation through pretrained models and expert encoder group
Yue Pan 0011, Chen Lyu 0001, Lantian Li, Xiuting Shao
Inf. Softw. Technol.4
2025 Neural Scoring: A Refreshed End-to-End Approach for Speaker Verification in Complex Conditions
abstract
Modern speaker verification systems primarily rely on speaker embeddings, followed by verification based on cosine similarity between the embedding vectors of the enrollment and test utterances. While effective, these methods struggle with multi-talker speech due to the unidentifiability of embedding vectors. In this paper, we propose Neural Scoring, a refreshed end-to-end framework that directly estimates verification posterior probabilities without relying on test-side embeddings, making it more robust to complex conditions, e.g., with multiple talkers. To make the training of such an end-to-end model more efficient, we introduce a large-scale trial e2e training strategy, where each test utterance pairs with a set of enrolled speakers, thus enabling processing of large-scale verification trials per batch. Experiments on VoxCeleb dataset demonstrate that Neural Scoring consistently outperforms both the baseline and competitive methods across various conditions, achieving an overall 70.36% reduction in Equal Error Rate compared to the baseline.
Wan Lin, Junhui Chen, Lantian Li, Dong Wang 0013
IEEE Signal Process. Lett.5
2024 ViTA: A Highly Efficient Dataflow and Architecture for Vision Transformers
abstract
Transformer-based DNNs have dominated several AI fields with remarkable performance. However, the scaling up of Transformer models up to trillions of parameters and computation operations has made them both computationally and data-intensive. This poses a significant challenge to utilizing Transformer models, e.g., in the area- and power-constrained systems. This paper introduces ViTA, an efficient architecture for accelerating the entire workload of Vision Transformer (ViT), targeting enhanced area and power efficiency. ViTA adopts a novel memory-centric dataflow that reduces memory usage and data movement, exploiting computational parallelism and locality. This design results in a 76.71% reduction in memory requirements for Multi-Head Self Attention (MHA) compared to original dataflows with VGA resolution images. A fused configurable module is designed for supporting non-linear functions in ViT workloads, such as GELU, Softmax, and LayerNorm, optimizing hardware resource usage. Our results show that ViTA achieves 16.384 TOPS with area and power efficiencies of 2.13 TOPS/mm2and 1.57 TOPS/W at 1 GHz, surpassing current Transformer accelerators by 27.85× and 1.40×, respectively.
Chunyun Chen, Lantian Li, Mohamed M. Sabry
DATE2
2024 An Investigation of Distribution Alignment in Multi-Genre Speaker Recognition
abstract
Multi-genre speaker recognition is becoming increasingly popular due to its ability to better represent the complexities of real-world applications. However, a major challenge is the significant shift in the distribution of speaker vectors across different genres. While distribution alignment is a common approach to address this challenge, previous studies have mainly focused on aligning a source domain with a target domain, and the performance of multi-genre data is unknown.This paper presents a comprehensive study of mainstream distribution alignment methods on multi-genre data, where multiple distributions need to be aligned. We analyze various methods both qualitatively and quantitatively. Our experiments on the CN-Celeb dataset show that within-between distribution alignment (WBDA) performs relatively better. However, we also found that none of the investigated methods consistently improved performance in all test cases. This suggests that solely aligning the distributions of speaker vectors may not fully address the challenges posed by multi-genre speaker recognition. Further investigation is necessary to develop a more comprehensive solution.
Junhui Chen, Namin Wang, Lantian Li, Dong Wang 0013
ICASSP4
2024 Serialized Output Training by Learned Dominance
Ying Shi 0001, Lantian Li, Dong Wang 0013, Jiqing Han 0001
INTERSPEECH2
2024 CNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge
Chen Chen 0075, Zehua Liu, Xiaolou Li, Lantian Li, Dong Wang 0013
INTERSPEECH4
2024 Zero-Shot Fake Video Detection by Audio-Visual Consistency
Xiaolou Li, Zehua Liu, Chen Chen 0075, Lantian Li, Li Guo 0004, Dong Wang 0013
INTERSPEECH4
2024 SE/BN Adapter: Parametric Efficient Domain Adaptation for Speaker Recognition
Lantian Li, Dong Wang 0013
INTERSPEECH2
2024 Few-Shot Keyword Spotting from Mixed Speech
abstract
Few-shot keyword spotting (KWS) aims to detect unknown keywords with limited training samples.A commonly used approach is the pre-training and fine-tuning framework.While effective in clean conditions, this approach struggles with mixed keyword spotting -simultaneously detecting multiple keywords blended in an utterance, which is crucial in real-world applications.Previous research has proposed a Mix-Training (MT) approach to solve the problem, however, it has never been tested in the few-shot scenario.In this paper, we investigate the possibility of using MT and other relevant methods to solve the two practical challenges together: few-shot and mixed speech.Experiments conducted on the LibriSpeech and Google Speech Command corpora demonstrate that MT is highly effective on this task when employed in either the pre-training phase or the fine-tuning phase.Moreover, combining SSL-based large-scale pre-training (HuBert) and MT fine-tuning yields very strong results in all the test conditions.
Junming Yuan, Ying Shi 0001, Lantian Li, Dong Wang 0013, Askar Hamdulla
INTERSPEECH3
2024 A Comprehensive Investigation on Speaker Augmentation for Speaker Recognition
abstract
Data augmentation (DA) has played a pivotal role in the success of deep speaker recognition.Current DA techniques primarily focus on speaker-preserving augmentation, which does not change the speaker trait of the speech and does not create new speakers.Recent research has shed light on the potential of speaker augmentation, which generates new speakers to enrich the training dataset.In this study, we delve into two speaker augmentation approaches: speed perturbation (SP) and vocal tract length perturbation (VTLP).Despite the empirical utilization of both methods, a comprehensive investigation into their efficacy is lacking.Our study, conducted using two public datasets, VoxCeleb and CN-Celeb, revealed that both SP and VTLP are proficient at generating new speakers, leading to significant performance improvements in speaker recognition.Furthermore, they exhibit distinct properties in sensitivity to perturbation factors and data complexity, hinting at the potential benefits of their fusion.Our research underscores the substantial potential of speaker augmentation, highlighting the importance of in-depth exploration and analysis.
Shibiao Xu, Lantian Li, Dong Wang 0013
INTERSPEECH4
2024 On evaluation trials in speaker verification
Lantian Li, Di Wang 0039, Andrew Abel, Dong Wang 0013
Appl. Intell.1
2024 Maximum Gaussianality training for deep speaker vector normalization
Yunqi Cai, Lantian Li, Andrew Abel, Xiaoyan Zhu 0001, Dong Wang 0013
Pattern Recognit.2
2024 Keyword Guided Target Speech Recognition
abstract
This letter presents a new target speech recognition problem, where the target speech is defined by a keyword. For instance, when a person speaks “Hey Google” or “Help Me”, we hope the model can recognize the entire contextual speech of that person, even with strong interference speech from other people. The new problem is denoted by target content ASR (TC-ASR). The core challenge of TC-ASR is that the model needs to simultaneously detect the existence of the keyword from heavily mixed speech and recognize the target speech component using the information of the detected keyword segment. Surprisingly, our experiments show that an attention encoder-decoder (AED) model augmented with a keyword encoder can solve this problem pretty well. We also defined a key content spotting (KCS) task and tested the proposed model on it. Our experiments on the LibriMix dataset demonstrated that our approach could address the KCS task with a promising accuracy, outperforming two baseline models by a large margin. Further analysis shows that the proposed model identifies the target speech by a timbre cue, i.e., ensuring that the identified speech is coherent in speaker trait.
Ying Shi 0001, Lantian Li, Dong Wang 0013, Jiqing Han 0001
IEEE Signal Process. Lett.2
2023 Spot Keywords From Very Noisy and Mixed Speech
Ying Shi 0001, Dong Wang 0013, Lantian Li, Jiqing Han 0001
INTERSPEECH3
2023 Visualizing Data Augmentation in Deep Speaker Recognition
Pengqi Li, Lantian Li, Askar Hamdulla, Dong Wang 0013
INTERSPEECH2
2023 CN-Celeb-AV: A Multi-Genre Audio-Visual Dataset for Person Recognition
Lantian Li, Xiaolou Li, Chen Chen 0075, Ruihai Hou, Dong Wang 0013
INTERSPEECH1
2023 Ordered and Binary Speaker Embedding
Namin Wang, Lantian Li, Dong Wang 0013
INTERSPEECH4
2023 A Multi-Scale Attentive Transformer for Multi-Instrument Symbolic Music Generation
Xipin Wei, Junhui Chen, Zirui Zheng, Li Guo 0004, Lantian Li, Dong Wang 0013
INTERSPEECH5
2023 Understanding Solidity Event Logging Practices in the Wild
abstract
Writing logging messages is a well-established conventional programming practice, and it is of vital importance for a wide variety of software development activities. The logging mechanism in Solidity programming is enabled by the high-level event feature, but up to now there lacks study for understanding Solidity event logging practices in the wild. To fill this gap, we in this paper provide the first quantitative characteristic study of the current Solidity event logging practices using 2,915 popular Solidity projects hosted on GitHub. The study methodically explores the pervasiveness of event logging, the goodness of current event logging practices, and in particular the reasons for event logging code evolution, and delivers 8 original and important findings. The findings notably include the existence of a large percentage of independent event logging code modifications, and the underlying reasons for different categories of independent event logging code modifications are diverse (for instance, bug fixing and gas saving). We additionally give the implications of our findings, and these implications can enlighten developers, researchers, tool builders, and language designers to improve the event logging practices. To illustrate the potential benefits of our study, we develop a proof-of-concept checker on top of one of our findings and the checker effectively detects problematic event logging code that consumes extra gas in 35 popular GitHub projects and 9 project owners have already confirmed the detected issues.
Lantian Li, Yejian Liang, Zhongxing Yu
ESEC/SIGSOFT FSE1
2023 Random Cycle Loss and Its Application to Voice Conversion
abstract
Speech disentanglement aims to decompose independent causal factors of speech signals into separate codes. Perfect disentanglement benefits to a broad range of speech processing tasks. This paper presents a simple but effective disentanglement approach based on cycle consistency loss and random factor substitution. This leads to a novel random cycle (RC) loss that enforces analysis-and-resynthesis consistency, a main principle of reductionism. We theoretically demonstrate that the proposed RC loss can achieve independent codes if well optimized, which in turn leads to superior disentanglement when combined with information bottleneck (IB). Extensive simulation experiments were conducted to understand the properties of the RC loss, and experimental results on voice conversion further demonstrate the practical merit of the proposal. Source code and audio samples can be found on the webpage http://rc.cslt.org.
Dong Wang 0013, Lantian Li, Chen Chen 0075, Thomas Fang Zheng
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Real Additive Margin Softmax for Speaker Verification
abstract
The additive margin softmax (AM-Softmax) loss has delivered remarkable performance in speaker verification. A supposed behavior of AM-Softmax is that it can shrink within-class variation by putting emphasis on target logits, which in turn improves margin between target and non-target classes. In this paper, we conduct a careful analysis on the behavior of AM-Softmax loss, and show that this loss does not implement real max-margin training. Based on this observation, we present a Real AM-Softmax loss which involves a true margin function in the softmax training. Experiments conducted on VoxCeleb1, SITW and CNCeleb demonstrated that the corrected AM-Softmax loss consistently outperforms the original one. The code has been released at https://gitlab.com/csltstu/sunine.
Lantian Li, Ruiqian Nai, Dong Wang 0013
ICASSP1
2022 Reliable Visualization for Deep Speaker Recognition
abstract
In spite of the impressive success of convolutional neural networks (CNNs) in speaker recognition, our understanding to CNNs' internal functions is still limited. A major obstacle is that some popular visualization tools are difficult to apply, for example those producing saliency maps. The reason is that speaker information does not show clear spatial patterns in the temporal-frequency space, which makes it hard to interpret the visualization results, and hence hard to confirm the reliability of a visualization tool. In this paper, we conduct an extensive analysis on three popular visualization methods based on CAM: Grad-CAM, Score-CAM and Layer-CAM, to investigate their reliability for speaker recognition tasks. Experiments conducted on a state-of-the-art ResNet34SE model show that the Layer-CAM algorithm can produce reliable visualization, and thus can be used as a promising tool to explain CNN-based speaker models. The source code and examples are available in our project page: http://project.cslt.org/.
Pengqi Li, Lantian Li, Askar Hamdulla, Dong Wang 0013
INTERSPEECH2
2022 CN-Celeb: Multi-genre speaker recognition
Lantian Li, Jiawen Kang 0002, Yunqi Cai, Ravichander Vipperla, Thomas Fang Zheng, Dong Wang 0013
Speech Commun.1
2022 A Principle Solution for Enroll-Test Mismatch in Speaker Recognition
abstract
Mismatch between enrollment and test conditions causes serious performance degradation on speaker recognition systems. This paper presents a statistics decomposition (SD) approach to solve this problem. This approach decomposes the PLDA score into three components that corresponding to enrollment, prediction and normalization respectively. Given that correct statistics are used in each component, the resultant score is theoretically optimal. A comprehensive experimental study was conducted on three datasets with different types of mismatch: (1) physical channel mismatch, (2) long-term speaker characteristics mismatch, (3) near-far recording mismatch. The results demonstrated that the proposed SD approach is highly effective, and outperforms the ad-hoc multi-condition training approach that is commonly adopted but not optimal in theory.
Lantian Li, Dong Wang 0013, Jiawen Kang 0002, Renyu Wang, Zhendong Gao, Xiao Chen 0012
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Squeezing Value of Cross-Domain Labels: A Decoupled Scoring Approach for Speaker Verification
abstract
Domain mismatch often occurs in real applications and causes serious performance reduction on speaker verification systems. The common wisdom is to collect cross-domain data and train a multi-domain PLDA model, with the hope to learn a domain-independent speaker subspace. In this paper, we firstly present an empirical study to show that simply adding cross-domain data does not help performance in conditions with enrollment-test mismatch. Careful analysis shows that this striking result is caused by the incoherent statistics between the enrollment and test conditions. Based on this analysis, we present a decoupled scoring approach that can maximally squeeze the value of cross-domain labels and obtain optimal verification scores in the enrollment-test mismatch condition. When the statistics are coherent, the new formulation falls back to the conventional PLDA. Experimental results on cross-channel test show that the proposed approach is highly effective and is a principal solution to domain mismatch.
Lantian Li, Yang Zhang 0052, Jiawen Kang 0002, Thomas Fang Zheng, Dong Wang 0013
ICASSP1
2021 Can We Trust Deep Speech Prior?
abstract
Recently, speech enhancement (SE) based on deep speech prior has attracted much attention, such as the variational auto-encoder with non-negative matrix factorization (VAE-NMF) architecture. Compared to conventional approaches that represent clean speech by shallow models such as Gaussians with a low-rank covariance, the new approach employs deep generative models to represent the clean speech, which often provides a better prior. Despite the clear advantage in theory, we argue that deep priors must be used with much caution, since the likelihood produced by a deep generative model does not always coincide with the speech quality. We designed a comprehensive study on this issue and demonstrated that based on deep speech priors, a reasonable SE performance can be achieved, but the results might be suboptimal. A careful analysis showed that this problem is deeply rooted in the disharmony between the flexibility of deep generative models and the nature of the maximum-likelihood (ML) training.
Ying Shi 0001, Zhiyuan Tang, Lantian Li, Dong Wang 0013, Jiqing Han 0001
SLT4
2021 Deep Normalization for Speaker Vectors
abstract
Deep speaker embedding has demonstrated state-of-the-art performance in speaker recognition tasks. However, one potential issue with this approach is that the speaker vectors derived from deep embedding models tend to be non-Gaussian for each individual speaker, and non-homogeneous for distributions of different speakers. These irregular distributions can seriously impact speaker recognition performance, especially with the popular PLDA scoring method, which assumes homogeneous Gaussian distribution. In this article, we argue that deep speaker vectors require deep normalization, and propose a deep normalization approach based on a novel discriminative normalization flow (DNF) model. We demonstrate the effectiveness of the proposed approach with experiments using the widely used SITW and CNCeleb corpora. In these experiments, the DNF-based normalization delivered substantial performance gains and also showed strong generalization capability in out-of-domain tests.
Yunqi Cai, Lantian Li, Andrew Abel, Xiaoyan Zhu 0001, Dong Wang 0013
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 CN-Celeb: A Challenging Chinese Speaker Recognition Dataset
abstract
Recently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary. However, most publicly available datasets are collected under constrained environments, i.e., with little noise and limited channel variation. These datasets tend to deliver over-optimistic performance and do not meet the request of research on speaker recognition in unconstrained conditions.In this paper, we present CN-Celeb, a large-scale speaker recognition dataset collected ‘in the wild’. This dataset contains more than 130,000 utterances from 1,000 Chinese celebrities, and covers 11 different genres in real world. Experiments conducted with two state-of-the-art speaker recognition approaches (i-vector and x-vector) show that the performance on CN-Celeb is far inferior to the one obtained on Vox-Celeb, a widely used speaker recognition dataset. This result demonstrates that in real-life conditions, the performance of existing techniques might be much worse than it was thought. Our database is free for researchers and can be downloaded from http://project.cslt.org.
Jiawen Kang 0002, Lantian Li, Kaicheng Li, Sitong Cheng, Pengyuan Zhang, Ziya Zhou, Yunqi Cai, Dong Wang 0013
ICASSP3
2020 ASR-Free Pronunciation Assessment
abstract
Most of the pronunciation assessment methods are based on local features derived from automatic speech recognition (ASR), e.g., the Goodness of Pronunciation (GOP) score. In this paper, we investigate an ASR-free scoring approach that is derived from the marginal distribution of raw speech signals. The hypothesis is that even if we have no knowledge of the language (so cannot recognize the phones/words), we can still tell how good a pronunciation is, by comparatively listening to some speech data from the target language. Our analysis shows that this new scoring approach provides an interesting correction for the phone-competition problem of GOP. Experimental results on the ERJ dataset demonstrated that combining the ASR-free score and GOP can achieve better performance than the GOP baseline.
Sitong Cheng, Lantian Li, Zhiyuan Tang, Dong Wang 0013, Thomas Fang Zheng
INTERSPEECH3
2020 Domain-Invariant Speaker Vector Projection by Model-Agnostic Meta-Learning
abstract
Domain generalization remains a critical problem for speaker recognition, even with the state-of-the-art architectures based on deep neural nets. For example, a model trained on reading speech may largely fail when applied to scenarios of singing or movie. In this paper, we propose a domain-invariant projection to improve the generalizability of speaker vectors. This projection is a simple neural net and is trained following the Model-Agnostic Meta-Learning (MAML) principle, for which the objective is to classify speakers in one domain if it had been updated with speech data in another domain. We tested the proposed method on CNCeleb, a new dataset consisting of single-speaker multi-condition (SSMC) data. The results demonstrated that the MAML-based domain-invariant projection can produce more generalizable speaker vectors, and effectively improve the performance in unseen domains.
Jiawen Kang 0002, Lantian Li, Yunqi Cai, Dong Wang 0013, Thomas Fang Zheng
INTERSPEECH3
2020 Neural Discriminant Analysis for Deep Speaker Embedding
abstract
Probabilistic Linear Discriminant Analysis (PLDA) is a popular tool in open-set classification/verification tasks.However, the Gaussian assumption underlying PLDA prevents it from being applied to situations where the data is clearly non-Gaussian.In this paper, we present a novel nonlinear version of PLDA named as Neural Discriminant Analysis (NDA).This model employs an invertible deep neural network to transform a complex distribution to a simple Gaussian, so that the linear Gaussian model can be readily established in the transformed space.We tested this NDA model on a speaker recognition task where the deep speaker vectors (x-vectors) are presumably non-Gaussian.Experimental results on two datasets demonstrate that NDA consistently outperforms PLDA, by handling the non-Gaussian distributions of the x-vectors.
Lantian Li, Dong Wang 0013, Thomas Fang Zheng
INTERSPEECH1
2020 Character-level neural network model based on Nadam optimization and its application in clinical concept extraction
Lantian Li, Weizhi Xu 0001, Hui Yu 0010
Neurocomputing1
2019 Gaussian-constrained Training for Speaker Verification
abstract
Neural models, in particular the d-vector and x-vector architectures, have produced state-of-the-art performance on many speaker verification tasks. However, two potential problems of these neural models deserve more investigation. Firstly, both models suffer from `information leak', which means that some parameters participating in model training will be discarded during inference, i.e, the layers that are used as the classifier. Secondly, these models do not regulate the distribution of the derived speaker vectors. This `unconstrained distribution' may degrade the performance of the subsequent scoring component, e.g., PLDA. This paper proposes a Gaussian-constrained training approach that (1) discards the parametric classifier, and (2) enforces the distribution of the derived speaker vectors to be Gaussian. Our experiments on the VoxCeleb and SITW databases demonstrated that this new training approach produced more representative and regular speaker embeddings, leading to consistent performance improvement.
Lantian Li, Zhiyuan Tang, Ying Shi 0001, Dong Wang 0013
ICASSP1
2019 VAE-Based Regularization for Deep Speaker Embedding
abstract
Deep speaker embedding has achieved state-of-the-art performance in speaker recognition. A potential problem of these embedded vectors (called `x-vectors') are not Gaussian, causing performance degradation with the famous PLDA back-end scoring. In this paper, we propose a regularization approach based on Variational Auto-Encoder (VAE). This model transforms x-vectors to a latent space where mapped latent codes are more Gaussian, hence more suitable for PLDA scoring.
Yang Zhang 0052, Lantian Li, Dong Wang 0013
INTERSPEECH2
2018 Full-Info Training for Deep Speaker Feature Learning
abstract
In recent studies, it has shown that speaker patterns can be learned from very short speech segments (e.g., 0.3 seconds) by a carefully designed convolutional & time-delay deep neural network (CT-DNN) model. By enforcing the model to discriminate the speakers in the training data, frame-level speaker features can be derived from the last hidden layer. In spite of its good performance, a potential problem of the present model is that it involves a parametric classifier, i.e., the last affine layer, which may consume some discriminative knowledge, thus leading to `information leak' for the feature learning. This paper presents a full-info training approach that discards the parametric classifier and enforces all the discriminative knowledge learned by the feature net. Our experiments on the Fisher database demonstrate that this new training scheme can produce more coherent features, leading to consistent and notable performance improvement on the speaker verification task.
Lantian Li, Zhiyuan Tang, Dong Wang 0013, Thomas Fang Zheng
ICASSP1
2018 Deep Factorization for Speech Signal
abstract
Various informative factors mixed in speech signals, leading to great difficulty when decoding any of the factors. An intuitive idea is to factorize each speech frame into individual informative factors, though it turns out to be highly difficult. Recently, we found that speaker traits, which were assumed to be long-term distributional properties, are actually short-time patterns, and can be learned by a carefully designed deep neural network (DNN). This discovery motivated a cascade deep factorization (CDF) framework that will be presented in this paper. The proposed framework infers speech factors in a sequential way, where factors previously inferred are used as conditional variables when inferring other factors. We will show that this approach can effectively factorize speech signals, and using these factors, the original speech spectrum can be recovered with a high accuracy. This factorization and reconstruction approach provides potential values for many speech processing tasks, e.g., speaker recognition and emotion recognition, as will be demonstrated in the paper.
Lantian Li, Dong Wang 0013, Yixiang Chen 0003, Ying Shi 0001, Zhiyuan Tang, Thomas Fang Zheng
ICASSP1
2018 Human and Machine Speaker Recognition Based on Short Trivial Events
abstract
Human speech often has events that we will call trivial events, e.g., cough, laugh and sniff. Compared to regular speech, these trivial events are usually short and variable, thus generally regarded as not speaker discriminative and so are largely ignored by present speaker recognition research. However, these trivial events are highly valuable in some particular circumstances such as forensic examination, as they are less subjected to intentional change, so can be used to discover the genuine speaker from disguised speech. In this paper, we collect a trivial event speech database that involves 75 speakers and 6 types of events, and report preliminary speaker recognition results on this database, by both human listeners and machines. Particularly, the deep feature learning technique recently proposed by our group is utilized to analyze and recognize the trivial events, leading to acceptable equal error rates (EERs) ranging from 5% to 15% despite the extremely short durations (0.2-0.5 seconds) of these events. Comparing different types of events, `hmm' seems more speaker discriminative.
Xiaofei Kang, Lantian Li, Zhiyuan Tang, Haisheng Dai, Dong Wang 0013
ICASSP4
2018 Phonetic Temporal Neural Model for Language Identification
abstract
Deep neural models, particularly the long short-term memory recurrent neural network (LSTM-RNN) model, have shown great potential for language identification (LID). However, the use of phonetic information has been largely overlooked by most existing neural LID methods, although this information has been used very successfully in conventional phonetic LID systems. We present a phonetic temporal neural model for LID, which is an LSTM-RNN LID system that accepts phonetic features produced by a phone-discriminative DNN as the input, rather than raw acoustic features. This new model is similar to traditional phonetic LID methods, but the phonetic knowledge here is much richer: It is at the frame level and involves compacted information of all phones. Our experiments conducted on the Babel database and the AP16-OLR database demonstrate that the temporal phonetic neural approach is very effective, and significantly outperforms existing acoustic neural models. It also outperforms the conventional i-vector approach on short utterances and in noisy conditions.
Zhiyuan Tang, Dong Wang 0013, Yixiang Chen 0003, Lantian Li, Andrew Abel
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Speaker segmentation using deep speaker vectors for fast speaker change scenarios
abstract
A novel speaker segmentation approach based on deep neural network is proposed and investigated. This approach uses deep speaker vectors (d-vectors) to represent speaker characteristics and to find speaker change points. The d-vector is a kind of frame-level speaker discriminative feature, whose discriminative training process corresponds to the goal of discriminating a speaker change point from a single speaker speech segment in a short time window. Following the traditional metric-based segmentation, each analysis window contains two sub-windows and is shifting along the audio stream to detect speaker change points, where the speaker characteristics are represented by the means of deep speaker vectors for all frames in each window. Experimental investigations conducted in fast speaker change scenarios show that the proposed method can detect speaker change points more quickly and more effectively than the commonly used segmentation methods.
Renyu Wang, Mingliang Gu, Lantian Li, Mingxing Xu, Thomas Fang Zheng
ICASSP3
2017 Deep Speaker Feature Learning for Text-Independent Speaker Verification
abstract
Recently deep neural networks (DNNs) have been used to learn speaker features.However, the quality of the learned features is not sufficiently good, so a complex back-end model, either neural or probabilistic, has to be used to address the residual uncertainty when applied to speaker verification, just as with raw features.This paper presents a convolutional timedelay deep neural network structure (CT-DNN) for speaker feature learning.Our experimental results on the Fisher database demonstrated that this CT-DNN can produce highquality speaker features: even with a single feature (0.3 seconds including the context), the EER can be as low as 7.68%.This effectively confirmed that the speaker trait is largely a deterministic short-time property rather than a long-time distributional pattern, and therefore can be extracted from just dozens of frames.
Lantian Li, Yixiang Chen 0003, Ying Shi 0001, Zhiyuan Tang, Dong Wang 0013
INTERSPEECH1
2017 A Study on Replay Attack and Anti-Spoofing for Automatic Speaker Verification
abstract
For practical automatic speaker verification (ASV) systems, replay attack poses a true risk.By replaying a pre-recorded speech signal of the genuine speaker, ASV systems tend to be easily fooled.An effective replay detection method is therefore highly desirable.In this study, we investigate a major difficulty in replay detection: the over-fitting problem caused by variability factors in speech signal.An F-ratio probing tool is proposed and three variability factors are investigated using this tool: speaker identity, speech content and playback & recording device.The analysis shows that device is the most influential factor that contributes the highest over-fitting risk.A frequency warping approach is studied to alleviate the over-fitting problem, as verified on the ASV-spoof 2017 database.
Lantian Li, Yixiang Chen 0003, Dong Wang 0013, Thomas Fang Zheng
INTERSPEECH1
2017 Collaborative Joint Training With Multitask Recurrent Model for Speech and Speaker Recognition
abstract
Automatic speech and speaker recognition are traditionally treated as two independent tasks and are studied separately. The human brain in contrast deciphers the linguistic content, and the speaker traits from the speech in a collaborative manner. This key observation motivates the work presented in this paper. A collaborative joint training approach based on multitask recurrent neural network models is proposed, where the output of one task is backpropagated to the other tasks. This is a general framework for learning collaborative tasks and fits well with the goal of joint learning of automatic speech and speaker recognition. Through a comprehensive study, it is shown that the multitask recurrent neural net models deliver improved performance on both automatic speech and speaker recognition tasks as compared to single-task systems. The strength of such multitask collaborative learning is analyzed, and the impact of various training configurations is investigated.
Zhiyuan Tang, Lantian Li, Dong Wang 0013, Ravichander Vipperla
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Improving speaker verification performance against long-term speaker variability
Jun Wang 0073, Lantian Li, Thomas Fang Zheng, Frank K. Soong
Speech Commun.3
2016 Improving Short Utterance Speaker Recognition by Modeling Speech Unit Classes
abstract
Short utterance speaker recognition (SUSR) is highly challenging due to the limited enrollment and/or test data. We argue that the difficulty can be largely attributed to the mismatched prior distributions of the speech data used to train the universal background model (UBM) and those for enrollment and test. This paper presents a novel solution that distributes speech signals into a multitude of acoustic subregions that are defined by speech units, and models speakers within the subregions. To avoid data sparsity, a data-driven approach is proposed to cluster speech units into speech unit classes, based on which robust subregion models can be constructed. Further more, we propose a model synthesis approach based on maximum likelihood linear regression (MLLR) to deal with no-data speech unit classes. The experiments were conducted on a publicly available database SUD12. The results demonstrated that on a text-independent speaker recognition task where the test utterances are no longer than 2 seconds and mostly shorter than 0.5 seconds, the proposed subregion modeling offered a 21.51% relative reduction in equal error rate (EER), compared with the standard GMM-UBM baseline. In addition, with the model synthesis approach, the performance can be greatly improved in scenarios where no enrollment data are available for some speech unit classes.
Lantian Li, Dong Wang 0013, Thomas Fang Zheng
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Community detection with manifold learning on speaker i-vector space for Chinese
Hongcui Wang, Di Jin 0001, Lantian Li, Jianwu Dang 0001
INTERSPEECH3