EDBT 2026 Demo / reviewers in the wild / expert
Ming Li 0026
dblp:l/MingLi26
· DBLP profile ↗
155ranked-venue papers
20as first author
84since 2021 · last 2026
0000-0002-6406-1983ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 113 · 15 first-author · 56 since 2021Artificial intelligence and machine learning · 98 · 15 first-author · 44 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal laryngoscopic video analysis for assisted diagnosis of vocal fold paralysis
Yucong Zhang, Jinshan Yang, Juan Liu 0007, Faya Liang, Ming Li 0026 |
Comput. Speech Lang. | 7 |
| 2026 | Quantization compensation and decomposition GAN for high-fidelity underwater image compression
Xufei Hu, Jian Zhang 0082, Heng Zhang 0001, Ming Li 0026, Meng Huang 0003, Hengmin Zhang |
Neurocomputing | 4 |
| 2026 | Detecting children with autism spectrum disorder based on script-centric behavior understanding with emotional enhancement
Yueran Pan, Dong Zhang 0002, Hongzhu Deng, Xiaobing Zou, Ming Li 0026 |
Neurocomputing | 6 |
| 2025 | SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means QuantizationabstractVoice anonymization protects speaker privacy by concealing identity while preserving linguistic and paralinguistic content. Self-supervised learning (SSL) representations encode linguistic features but preserve speaker traits. We propose a novel speaker-embedding-free framework called SEF-MK. Instead of using a single k-means model trained on the entire dataset, SEF-MK anonymizes SSL representations for each utterance by randomly selecting one of multiple k-means models, each trained on a different subset of speakers. We explore this approach from both attacker and user perspectives. Extensive experiments show that, compared to a single k -means model, SEF-MK with multiple $\mathbf{k}$-means models better preserves linguistic and emotional content from the user’s viewpoint. However, from the attacker’s perspective, utilizing multiple $\mathbf{k}$-means models boosts the effectiveness of privacy attacks. These insights can aid users in designing voice anonymization systems to mitigate attacker threats.11Code and audio samples can be found at https://github.com/Beilong-Tang/sef-mk Beilong Tang, Xiaoxiao Miao, Xin Wang 0037, Ming Li 0026 |
ASRU | 4 |
| 2025 | LauraTSE: Target Speaker Extraction using Auto-Regressive Decoder-Only Language ModelsabstractWe propose LauraTSE, an Auto-Regressive Decoder-Only Language Model for Target Speaker Extraction built upon the LauraGPT backbone. LauraTSE employs a small-scale auto-regressive decoder-only language model that generates the initial layers of the target speech’s discrete codec representations from the continuous embeddings of both the mixture and reference speech. These outputs serve as coarsegrained predictions. To refine them, a one-step encoder-only language model reconstructs the full codec representation by integrating information from both the mixture and the reference speech, adding fine-grained details. Experimental results show that our approach can achieve promising performance. Additionally, we conduct ablation studies to investigate the data scalability and the contribution of the encoder-only model. Beilong Tang, Bang Zeng, Ming Li 0026 |
ASRU | 3 |
| 2025 | DiCoGRN: Inference of Colorectal Cancer Subtypespecific Gene Regulatory Networks Using Dual-View Contrastive TransformerabstractColorectal cancer (CRC) is a highly heterogeneous disease with distinct molecular subtypes, exhibiting unique transcriptional programs and clinical behaviors. The development of subtype-specific gene regulatory network (GRN) prediction is critical for uncovering the dysregulated transcriptional circuits driving CRC progression, metastasis and therapy resistance. However, the existing GRN prediction methods have limitations in addressing data sparsity and generalization across CRC subtypes, failing to effectively capture subtype-specific regulatory relationships and struggling to handle uncharacterized regulatory factors. In this study, we propose a novel computational framework, Dual-view Contrastive Gene Regulatory Network (DiCoGRN), which integrate structural and semantic information through dual-view contrastive Transformer to infer CRC subtype-specific GRNs. DiCoGRN identifies different cell subpopulations (CRC subtypes) and applies Graph Contrastive Learning (GraphCL) to learn structural embeddings for each subtype. Then, DiCoGRN retrieves the highly variable genes and driver genes of CRC subtypes from the NCBI database, and uses a large language model (LLM) to extract the semantic embeddings of genes. These two-view information are fused through cross-attention and gating interaction mechanisms, ultimately inferring subtype-aware regulatory connections. Experimental results show that the proposed DiCoGRN outperforms existing methods in recovering regulatory interactions and generalization across 12 subtypes. The performed vitro wet experiments illustrate that GATA3 from our predicted regulatory relationships may be a potential CRC driver gene. Our method not only advances CRC subtype-specific GRN inference but also helps to advance toward personalized CRC treatment. Codes and data are available at https://github.com/Fraid-H/DiCoGRN. Meng Huang 0003, Huijin Hu, Ming Li 0026, Jian Zhang 0082, Heng Zhang 0001, Xiucai Ye |
BIBM | 3 |
| 2025 | Codec-ASV: Exploring Neural Audio Codec For Speaker Representation LearningabstractDiscrete speech representations have gained significant success in a variety of speech-related tasks. Among these, Neural Audio Codec (NAC), which serves as a compressed form of audio signals, have proven effective in speech AIGC applications. Moreover, we believe that the speaker information can be largely preserved in the compression process since the reconstructed voice is almost the same in human listening. In this paper, we explore various training strategies and codec types for NAC-based speaker representation learning. Using ECAPA-TDNN as the model backbone, our approach achieves state-of-the-art performance with a 2.08% EER in NAC-based speaker verification scenarios. To better retain speaker information in early, more compressed layers, we introduce mask-layer augmentation and embedding fusion techniques during the training process. Experimental results show the effectiveness of our methods, particularly when inferring with limited codec layers. Yuke Lin, Fulin Zhang, Yingying Gao, Shilei Zhang, Ming Li 0026 |
ICASSP | 5 |
| 2025 | Adversarial Attacks and Robust Defenses in Speaker Embedding based Zero-Shot Text-to-Speech SystemabstractSpeaker embedding based zero-shot Text-to-Speech (TTS) systems enable high-quality speech synthesis for unseen speakers using minimal data. However, these systems are vulnerable to adversarial attacks, where an attacker introduces imperceptible perturbations to the original speaker’s audio waveform, leading to synthesized speech sounds like another person. This vulnerability poses significant security risks, including speaker identity spoofing and unauthorized voice manipulation. This paper investigates two primary defense strategies to address these threats: adversarial training and adversarial purification. Adversarial training enhances the model’s robustness by integrating adversarial examples during the training process, thereby improving resistance to such attacks. Adversarial purification, on the other hand, employs diffusion probabilistic models to revert adversarially perturbed audio to its clean form. Experimental results demonstrate that these defense mechanisms can significantly reduce the impact of adversarial perturbations, enhancing the security and reliability of speaker embedding based zero-shot TTS systems in adversarial environments. Ze Li 0003, Ming Li 0026 |
ICME | 4 |
| 2025 | Multi-scale Scanning Network for Machine Anomalous Sound Detection
Yucong Zhang, Juan Liu 0007, Ming Li 0026 |
ICONIP (3) | 3 |
| 2025 | Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge
Ming Cheng 0005, Cancan Li, Juan Liu 0007, Ming Li 0026 |
INTERSPEECH | 5 |
| 2025 | Exploring Pre-trained models on Ultrasound Modeling for Mice Autism Detection with Uniform Filter Bank and Attentive Scoring
Yucong Zhang, Ming Li 0026 |
INTERSPEECH | 3 |
| 2025 | VCapAV: A Video-Caption Based Audio-Visual Deepfake Detection Dataset
Yikang Wang, Qishan Zhang, Hiromitsu Nishizaki, Ming Li 0026 |
INTERSPEECH | 5 |
| 2025 | Selective Channel Attention based Target Speaker Voice Activity Detection for Speaker Diarization under AD-HOC Microphone Array Settings
Ming Cheng 0005, Ming Li 0026 |
INTERSPEECH | 4 |
| 2025 | SMIIP-NV: A Multi-Annotation Non-Verbal Expressive Speech Corpus in Mandarin for LLM-Based Speech SynthesisabstractIn natural language communication, emotions are often conveyed through non-verbal sounds (NVs), such as laughter, crying, cough and so on. However, most existing text-to-speech (TTS) corpora lack annotations for these non-verbal sounds, leading to a scarcity of systems capable of generating them. To address this gap, we introduce SMIIP-NV, a non-verbal speech synthesis corpus annotated with both emotions and non-verbal sounds, including laughter, crying, and cough. To the best of our knowledge, SMIIP-NV is the largest publicly available open-source expressive speech corpus that includes non-verbal speech and rich annotations. It comprises 33 hours of speech data, covering five distinct emotions and three types of non-verbal sounds, with detailed transcriptions and precise timestamps for each occurrence of non-verbal sounds. Additionally, the corpus provides annotations for speech segments that contain laughter or crying. To demonstrate the utility of this dataset, we establish a baseline for non-verbal speech synthesis by employing a lightweight large language model (LLM). The SMIIP-NV dataset and static audio demonstrations are publicly available at https://axunyii.github.io/SMIIP-NV. The interactive real-time demonstrations can be accessed at https://huggingface.co/spaces/xunyi/SMIIP-NV_Finetuned_CosyVoice2. Zhuojun Wu, Dong Liu 0028, Juan Liu 0007, Yechen Wang, Hui Bu, Pengyuan Zhang, Ming Li 0026 |
ACM Multimedia | 9 |
| 2025 | Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
Bang Zeng, Ming Li 0026 |
Comput. Speech Lang. | 2 |
| 2025 | Enhanced Secure Communication via Dual-Mode AAV Equipped With Reconfigurable Intelligent SurfacesabstractThe vulnerability of wireless communication links to eavesdropping poses significant challenges in securing AAV-assisted networks. To enhance security, reconfigurable intelligent surfaces (RIS) and artificial noise (AN) have emerged as promising technologies for mitigating eavesdropping by controlling wireless propagation environments and introducing interference against eavesdroppers. However, existing works have rarely combined transmitter beamforming, RIS, and AN integratedly considered, and leveraging their complementary characteristics for efficient security enhancement remains challenging. Additionally, optimizing such system security performance is complicated by the nonconvexity of secrecy rate maximization and the highly time-varying communication links caused by the mobility of AAVs and users. To address these challenges, we propose a secure communication framework that integrates RIS and AN transmission devices on AAVs. To solve the resulting nonconvex optimization problem, we develop a dual-mode framework based on twin delayed deep deterministic policy gradient (TD3), employing two subenvironments that interact independently before updating a global environment. Extensive simulations demonstrate that the proposed approach significantly enhances secrecy rate performance compared to other methods. Heng Zhang 0001, Zhemin Sun, Chaoqun Yang 0001, Xianghui Cao, Jian Zhang 0082, Ming Li 0026 |
IEEE Internet Things J. | 6 |
| 2025 | Assessing the Expressive Language Levels of Autistic Children in Home InterventionabstractThe World Health Organization (WHO) has established the caregiver skill training (CST) program, designed to empower families with children diagnosed with autism spectrum disorder the essential caregiving skills. The joint engagement rating inventory (JERI) protocol evaluates participants’ engagement levels within the CST initiative. Traditionally, rating the expressive language level and use (EXLA) item in JERI relies on retrospective video analysis conducted by qualified professionals, thus incurring substantial labor costs. This study introduces a multimodal behavioral signal-processing framework designed to analyze both child and caregiver behaviors automatically, thereby rating EXLA. Initially, raw audio and video signals are segmented into concise intervals via voice activity detection, speaker diarization and speaker age classification, serving the dual purpose of eliminating nonspeech content and tagging each segment with its respective speaker. Subsequently, we extract an array of audio-visual features, encompassing our proposed interpretable, hand-crafted textual features, end-to-end audio embeddings and end-to-end video embeddings. Finally, these features are fused at the feature level to train a linear regression model aimed at predicting the EXLA scores. Our framework has been evaluated on the largest in-the-wild database currently available under the CST program. Experimental results indicate that the proposed system achieves a Pearson correlation coefficient of 0.768 against the expert ratings, evidencing promising performance comparable to that of human experts. Yueran Pan, Biyuan Chen, Ming Cheng 0005, Dong Zhang 0002, Hongzhu Deng, Xiaobing Zou, Ming Li 0026 |
IEEE Trans. Comput. Soc. Syst. | 8 |
| 2025 | S2DBFT: Spectral-Spatial Dual-Branch Fusion Transformer for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) and Transformer-based models have achieved remarkable success in hyperspectral image (HSI) classification tasks due to their outstanding ability to extract spatial and spectral features. However, most existing methods process spatial and spectral features separately, making it difficult to effectively learn their interactive features. To address this issue, we propose a spectral-spatial dual-branch fusion Transformer (S2DBFT) for HSI classification. Initially, we construct a spectral feature extraction module (SPEEM) and a spatial feature extraction module (SPAEM) to extract low-level features. These two modules consist of a one-dimensional convolution layer and a two-dimensional convolution layer, respectively, performing shallow extraction of spectral and spatial features. Next, the two feature sets obtained are fused through a weighted fusion process. Additionally, we design a multi-head spectral-spatial self-attention (MHS3A) mechanism to enhance the interactive fusion of spectral and spatial features. Upon completion of feature fusion, a linear layer is used to obtain the sample labels. Extensive experiments on four HSI datasets demonstrate the effectiveness of the proposed S2DBFT, compared to existing state-of-the-art methods. In terms of performance evaluation, the overall accuracy and average accuracy indicate the superiority and generalizability of S2DBFT. Meng Huang 0003, Ming Li 0026, Jian Zhang 0082, Shandong Wang, Jinglin Zhang 0001, Heng Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | StarRescue: the Design and Evaluation of A Turn-Taking Collaborative Game for Facilitating Autistic Children's Social SkillsabstractAutism Spectrum Disorder (ASD) presents challenges in social interaction skill development, particularly in turn-taking. Digital interventions offer potential solutions for improving autistic children’s social skills but often lack addressing specific collaboration techniques. Therefore, we designed a prototype of a turn-taking collaborative tablet game, StarRescue, which encourages children’s distinct collaborative roles and interdependence while progressively enhancing sharing and mutual planning skills. We further conducted a controlled study with 32 autistic children to evaluate StarRescue’s usability and potential effectiveness in improving their social skills. Findings indicated that StarRescue has great potential to foster turn-taking skills and social communication skills (e.g., prompting, negotiation, task allocation) within the game and also extend beyond the game. Additionally, we discussed implications for future work, such as including parents as game spectators and understanding autistic children’s territory awareness in collaboration. Our study contributes a promising digital intervention for autistic children’s turn-taking social skill development via a scaffolding approach and valuable design implications for future research. Rongqi Bei, Ming Li 0026, Yuhang Zhao 0001, Xin Tong 0004 |
CHI | 5 |
| 2024 | Invertible Voice Conversion with Parallel DataabstractThis paper introduces an innovative deep learning framework for parallel voice conversion to mitigate inherent risks associated with such systems. Our approach focuses on developing an invertible model capable of countering potential spoofing threats. Specifically, we present a conversion model that allows for the retrieval of source voices, thereby facilitating the identification of the source speaker. This framework is constructed using a series of invertible modules composed of affine coupling layers to ensure the reversibility of the conversion process. We conduct comprehensive training and evaluation of the proposed framework using parallel training data. Our experimental results reveal that this approach achieves comparable performance to non-invertible systems in voice conversion tasks. Notably, the converted outputs can be seamlessly reverted to the original source inputs using the same parameters employed during the forwarding process. This advancement holds considerable promise for elevating the security and reliability of voice conversion. Zexin Cai, Ming Li 0026 |
ICASSP | 2 |
| 2024 | Multi-Objective Progressive Clustering for Semi-Supervised Domain Adaptation in Speaker VerificationabstractUtilizing the pseudo-labeling algorithm with large-scale unlabeled data becomes crucial for semi-supervised domain adaptation in speaker verification tasks. In this paper, we propose a novel pseudo-labeling method named Multi-objective Progressive Clustering (MoPC), specifically designed for semi-supervised domain adaptation. Firstly, we utilize limited labeled data from the target domain to derive domain-specific descriptors based on multiple distinct objectives, namely within-graph denoising, intra-class denoising and inter-class denoising. Then, the Infomap algorithm is adopted for embedding clustering, and the descriptors are leveraged to further refine the target domain’s pseudo-labels. Moreover, to further improve the quality of pseudo labels, we introduce the subcenter-purification and progressive-merging strategy for label denoising. Our proposed MoPC method achieves 4.95% EER and ranked the 1stplace on the evaluation set of VoxSRC 2023 track 3. We also conduct additional experiments on the FFSVC dataset and yield promising results. Ze Li 0003, Yuke Lin, Xiaoyi Qin, Haiying Wu, Ming Li 0026 |
ICASSP | 7 |
| 2024 | Voxblink: A Large Scale Speaker Verification Dataset on CameraabstractIn this paper, we introduce a large-scale and high-quality audiovisual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains 1.45M utterances from 38K speakers. Due to the inherent nature of automated data collection, introducing noisy data is inevitable. Therefore, we also utilize a multi-modal purification step to generate a cleaner version of the VoxBlink, named VoxBlink-clean, comprising 18K identities and 1.02M utterances. In contrast to the VoxCeleb, the VoxBlink sources from short videos of ordinary users, and the covered scenarios can better align with real-life situations. To our best knowledge, the VoxBlink dataset is one of the largest publicly available speaker verification datasets. Leveraging the VoxCeleb and VoxBlink-clean datasets together, we employ diverse speaker verification models with multiple architectural backbones to conduct comprehensive evaluations on the VoxCeleb test sets. Experimental results indicate a substantial enhancement in performance—ranging from 12% to 30% relatively—across various backbone architectures upon incorporating the VoxBlink-clean into the training process. The details of the dataset can be found on $\color{Fuchsia} {{\text{Site}}}$. Yuke Lin, Xiaoyi Qin, Ming Cheng 0005, Haiying Wu, Ming Li 0026 |
ICASSP | 7 |
| 2024 | Joint Inference of Speaker Diarization and ASR with Multi-Stage Information SharingabstractIn this paper, we introduce a novel approach that unifies Automatic Speech Recognition (ASR) and speaker diarization in a cohesive framework. Utilizing the synergies between the two tasks, our method effectively extracts speaker-specific information from the lower layers of a pretrained Conformer-based ASR model while leveraging the higher layers for enhanced diarization performance. In particular, the integration of ASR contextual details into the diarization process has been demonstrated to be effective. Results on the DIHARD III dataset indicate that our approach achieves a Diarization Error Rate (DER) of 10.52%, which can be further reduced to 10.39% when integrating ASR features into the diarization model. These findings highlight the potential of our approach, suggesting competitive performance against other state-of-the-art systems. Additionally, our framework’s ability to simultaneously generate text transcripts for each speaker marks a distinct advantage, which can further enhance ASR capabilities and transition towards an end-to-end multitask framework encompassing both ASR and speaker diarization. Weiqing Wang 0004, Danwei Cai, Ming Cheng 0005, Ming Li 0026 |
ICASSP | 4 |
| 2024 | Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual ConformerabstractIn recent years, neural network-based Wake Word Spotting achieves good performance on clean audio samples but struggles in noisy environments. Audio-Visual Wake Word Spotting (AVWWS) receives lots of attention because visual lip movement information is not affected by complex acoustic scenes. Previous works usually use simple addition or concatenation for multi-modal fusion. The inter-modal correlation remains relatively under-explored. In this paper, we propose a novel module called Frame-Level Cross-Modal Attention (FLCMA) to improve the performance of AVWWS systems. This module can help model multi-modal information at the frame-level through synchronous lip movements and speech signals. We train the end-to-end FLCMA based Audio-Visual Conformer and further improve the performance by fine-tuning pre-trained uni-modal models for the AVWWS task. The proposed system achieves a new state-of-the-art result (4.57% WWS score) on the far-field MISP dataset. Haoxu Wang, Ming Cheng 0005, Qiang Fu 0001, Ming Li 0026 |
ICASSP | 4 |
| 2024 | SlideSpeech: A Large Scale Slide-Enriched Audio-Visual CorpusabstractMulti-Modal automatic speech recognition (ASR) techniques aim to leverage additional modalities to improve the performance of speech recognition systems. While existing approaches primarily focus on video or contextual information, the utilization of extra supplementary textual information has been overlooked. Recognizing the abundance of online conference videos with slides, which provide rich domain-specific information in the form of text and images, we release SlideSpeech, a large-scale audio-visual corpus enriched with slides. The corpus contains 1,705 videos, 1,000+ hours, with 473 hours of high-quality transcribed speech. Moreover, the corpus contains a significant amount of real-time synchronized slides. In this work, we present the pipeline for constructing the corpus and propose baseline methods for utilizing text information in the visual slide context. Through the application of keyword extraction and contextual ASR methods in the benchmark system, we demonstrate the potential of improving speech recognition performance by incorporating textual information from supplementary video slides. Haoxu Wang, Fan Yu 0002, Xian Shi, Yuezhang Wang, Shiliang Zhang, Ming Li 0026 |
ICASSP | 6 |
| 2024 | Efficient Personal Voice Activity Detection with Wake Word Reference SpeechabstractPersonal voice activity detection (PVAD) is gradually used in speech assistants. Traditional PVAD schemes extract the target speaker’s embedding from existing query reference speech through a pre-trained speaker verification model. Consequently, the performance of the PVAD model may suffer if the quality of the extracted speaker embedding is poor, such as when only utilizing wake word speech as the reference. In this work, we introduce a novel and efficient PVAD model. In contrast to conventional approaches that rely on speaker embeddings extracted from a pre-trained speaker verification model, our proposed method directly uses the raw frame-level features of the reference speech as the target speaker’s attributes. In this way, our proposed model achieves an ultra-high recall rate, which is vital for speech assistant applications. The experimental results show the effectiveness of our proposed method in both cases of using existing query speech or wake word speech as reference. Bang Zeng, Ming Cheng 0005, Ming Li 0026 |
ICASSP | 5 |
| 2024 | A Dual-Path Framework with Frequency-and-Time Excited Network for Anomalous Sound DetectionabstractIn contrast to human speech, machine-generated sounds of the same type often exhibit consistent frequency characteristics and discernible temporal periodicity. However, leveraging these dual attributes in anomaly detection remains relatively under-explored. In this paper, we propose an automated dual-path framework that learns prominent frequency and temporal patterns for diverse machine types. One pathway uses a novel Frequency-and-Time Excited Network (FTE-Net) to learn the salient features across frequency and time axes of the spectrogram. It incorporates a Frequency-and-Time Chunkwise Encoder (FTC-Encoder) and an excitation network. The other pathway uses a 1D convolutional network for utterance-level spectrum. Experimental results on the DCASE 2023 task 2 dataset show the state-of-the-art performance of our proposed method. Moreover, visualizations of the intermediate feature maps in the excitation network are provided to illustrate the effectiveness of our method. Yucong Zhang, Juan Liu 0007, Ming Li 0026 |
ICASSP | 5 |
| 2024 | TMCSpeech: A Chinese TV and Movie Speech Dataset with Character Descriptions and a Character-Based Voice Generation Model
Dong Liu 0028, Yueqian Lin, Ming Li 0026 |
ICPR (6) | 4 |
| 2024 | KunquDB: An Attempt for Speaker Verification in the Chinese Opera Scenario
Huali Zhou, Yuke Lin, Dong Liu 0028, Ming Li 0026 |
ICPR (23) | 4 |
| 2024 | Enhancing Voice Wake-Up for Dysarthria: Mandarin Dysarthria Speech Corpus Release and Customized System Design
Hang Chen 0001, Jun Du 0002, Hongxiao Guo, Hui Bu, Jianxing Yang, Ming Li 0026, Chin-Hui Lee 0001 |
INTERSPEECH | 8 |
| 2024 | AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
Rong Gong, Hongfei Xue, Lezhi Wang, Qisheng Li, Lei Xie 0001, Hui Bu, Shaomei Wu, Jiaming Zhou 0001, Jun Du 0002, Jia Bin, Ming Li 0026 |
INTERSPEECH | 14 |
| 2024 | VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
Yuke Lin, Ming Cheng 0005, Fulin Zhang, Yingying Gao, Shilei Zhang, Ming Li 0026 |
INTERSPEECH | 6 |
| 2024 | Poster Abstract: Intrusion Detection for In-vehicle Networks Based on Parc-net ArchitectureabstractThe Controller Area Network (CAN) serves as a pivotal communication protocol for Electronic Control Units (ECUs) in modern automotive systems. However, the increasing interconnectivity and sophistication of these ECUs introduce significant vulnerabilities, rendering in-vehicle networks susceptible to a variety of cyber threats. This work presents a multi-class classification model for intrusion detection within vehicle CAN networks, utilizing Bidirectional Long Short-Term Memory (BILSTM) and Parc-Net architectures. We also introduce a novel feature extraction module that computes the rate of ID changes, significantly enhancing the model's detection capabilities. The proposed method was evaluated on the CAR-HACKING dataset, demonstrating remarkable performance. The proposed Intrusion Detection System (IDS) greatly contributes to the security of vehicle CAN networks by facilitating real-time detection and localization of potential intrusions. Mingming Tan, Heng Zhang 0001, Xin Wang 0037, Ming Li 0026, Meng Huang 0003, Jian Zhang 0082 |
MSN | 4 |
| 2024 | Summary of Low-Resource Dysarthria Wake-Up Word Spotting ChallengeabstractIn recent years, the rapid advancement and widespread adoption of speech technology have made smart home systems a common feature in many households. However, individuals with dysarthria face difficulties using these technologies due to inconsistent speech patterns. This paper summarizes the Low-Resource Dysarthria Wake-Up Word Spotting (LRDWWS) Challenge at SLT 2024, which aimed to develop effective voice wake-up systems for individuals with dysarthria. The challenge attracted 25 teams from 4 countries, with 7 teams submitting results and 5 providing detailed system descriptions. This paper presents an overview of the dataset, evaluation metrics, and key innovations from participating teams. Our findings highlight the potential of these systems to enhance the accessibility and usability of smart home technologies for individuals with dysarthria. The challenge results underscore the importance of developing specialized solutions to meet the unique needs of this user group. Hang Chen 0001, Jun Du 0002, Hongxiao Guo, Hui Bu, Ming Li 0026, Chin-Hui Lee 0001 |
SLT | 7 |
| 2024 | The Database and Benchmark For the Source Speaker Tracing Challenge 2024abstractVoice conversion (VC) systems can transform audio to mimic another speaker’s voice, thereby attacking speaker verification (SV) systems. However, ongoing studies on source speaker verification (SSV) are hindered by limited data availability and methodological constraints. This paper presents the Source Speaker Tracking Challenge (SSTC) on STL 2024, which aims to fill the gap in the database and benchmark for the SSV task. In this study, we generate a large-scale converted speech database with 16 common VC methods and train a batch of baseline systems based on the MFA-Conformer architecture. In addition, we introduced a related task called conversion method recognition, with the aim of assisting the SSV task. We expect SSTC to be a platform for advancing the development of the SSV task and provide further insights into the performance and limitations of current SV systems against VC attacks. Further details about SSTC can be found here1.1https://sstc-challenge.github.io/ Ze Li 0003, Yuke Lin, Hongbin Suo, Pengyuan Zhang, Yanzhen Ren, Zexin Cai, Hiromitsu Nishizaki, Ming Li 0026 |
SLT | 9 |
| 2024 | Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition ChallengeabstractThe StutteringSpeech Challenge focuses on advancing speech technologies for people who stutter, specifically targeting Stuttering Event Detection (SED) and Automatic Speech Recognition (ASR) in Mandarin. The challenge comprises three tracks: (1) SED, which aims to develop systems for detection of stuttering events; (2) ASR, which focuses on creating robust systems for recognizing stuttered speech; and (3) Research track for innovative approaches utilizing the provided dataset. We utilizes an open-source Mandarin stuttering dataset AS-70, which has been split into new training and test sets for the challenge. This paper presents the dataset, details the challenge tracks, and analyzes the performance of the top systems, highlighting improvements in detection accuracy and reductions in recognition error rates. Our findings underscore the potential of specialized models and augmentation strategies in developing stuttered speech technologies. Hongfei Xue, Rong Gong, Mingchen Shao, Lezhi Wang, Lei Xie 0001, Hui Bu, Jiaming Zhou 0001, Jun Du 0002, Ming Li 0026 |
SLT | 11 |
| 2024 | Integrating frame-level boundary detection and deepfake detection for locating manipulated regions in partially spoofed audio forgery attacks
Zexin Cai, Ming Li 0026 |
Comput. Speech Lang. | 2 |
| 2024 | Joint Training on Multiple Datasets With Inconsistent Labeling Criteria for Facial Expression RecognitionabstractOne potential way to enhance the performance of facial expression recognition (FER) is to augment the training set by increasing the number of samples. By incorporating multiple FER datasets, deep learning models can extract more discriminative features. However, the inconsistent labeling criteria and subjective biases found in annotated FER datasets can significantly hinder the recognition accuracy of deep learning models when handling mixed datasets. Effectively perform joint training on multiple datasets remains a challenging task. In this study, we propose a joint training method for training an FER model using multiple FER datasets. Our method consists of four steps: (1) selecting a subset from the additional dataset, (2) generating pseudo-continuous labels for the target dataset, (3) refining the labels of different datasets using continuous label mapping and discrete label relabeling according to the labeling criteria of the target dataset, and (4) jointly training the model using multi-task learning. We conduct joint training experiments on two popular in-the-wild FER benchmark databases, RAF-DB and CAER-S, while utilizing the AffectNet dataset as an additional dataset. The experimental results demonstrate that our proposed method outperforms the direct merging of different FER datasets into a single training set and achieves state-of-the-art performance on RAF-DB and CAER-S with accuracies of 92.24% and 94.57%, respectively. Chengyan Yu, Dong Zhang 0002, Ming Li 0026 |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Leveraging ASR Pretrained Conformers for Speaker Verification Through Transfer Learning and Knowledge DistillationabstractThis paper focuses on the application of Conformers in speaker verification. Conformers, initially designed for Automatic Speech Recognition (ASR), excel at modeling both local and global contexts within speech signals effectively. Building on this synergistic relationship, this study introduces three strategies for leveraging ASR-pretrained Conformers in speaker verification: (1) Transfer learning: We use a pretrained ASR Conformer encoder to initialize the speaker embedding network, thereby enhancing model generalization and mitigating the risk of overfitting. (2) Knowledge distillation: We distill the complex capabilities of an ASR Conformer into a speaker verification model. This not only allows for flexibility in the student mode's network architecture but also incorporates frame-level ASR distillation loss as an auxiliary task to reinforce speaker verification. (3) Parameter-efficient transfer learning with speaker adaptation: A lightweight speaker adaptation module is proposed to convert ASR-derived features into speaker-specific embeddings, without altering the core architecture of the original ASR Conformer. This strategy facilitates the concurrent execution of ASR and speaker verification tasks within a singular model. Experiments were conducted on VoxCeleb datasets. The best model using the ASR pretraining method achieved a 0.43% equal error rate (EER) on the VoxCeleb1-O test trial, while the knowledge distillation approach yielded a 0.38% EER. Furthermore, by adding a mere 4.92 million parameters to a 130.94 million-parameter ASR Conformer encoder, the speaker adaptation approach achieved a 0.45% EER, enabling parallel speech recognition and speaker verification within a single ASR Conformer encoder. Overall, our techniques successfully transfer rich ASR knowledge to advanced speaker modeling. Danwei Cai, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Investigating Long-Term and Short-Term Time-Varying Speaker VerificationabstractThe performance of speaker verification systems can be adversely affected by time domain variations. However, limited research has been conducted on time-varying speaker verification due to the absence of appropriate datasets. This paper aims to investigate the impact of long-term and short-term time-varying in speaker verification and proposes solutions to mitigate these effects. For long-term speaker verification (i.e., cross-age speaker verification), we introduce an age-decoupling adversarial learning method to learn age-invariant speaker representation by mining age information from the VoxCeleb dataset. For short-term speaker verification, we collect the SMIIP-TimeVarying (SMIIP-TV) Dataset, which includes recordings at multiple time slots every day from 373 speakers for 90 consecutive days and other relevant meta information. Using this dataset, we analyze the time-varying of speaker embeddings and propose a novel but realistic time-varying speaker verification task, termed incremental sequence-pair speaker verification. This task involves continuous interaction between enrollment audios and a sequence of testing audios with the aim of improving performance over time. We introduce the template updating method to counter the negative effects over time, and then formulate the template updating processing as a Markov Decision Process and propose a template updating method based on deep reinforcement learning (DRL). The policy network of DRL is treated as an agent to determine if and how much should the template be updated. In summary, this paper releases our collected database, investigates both the long-term and short-term time-varying scenarios and provides insights and solutions into time-varying speaker verification. Xiaoyi Qin, Na Li 0012, Shufei Duan, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Online Neural Speaker Diarization With Target Speaker TrackingabstractThis paper proposes an online target speaker voice activity detection (TS-VAD) system for speaker diarization tasks that does not rely on prior knowledge from clustering-based diarization systems to obtain target speaker embeddings. By adapting conventional TS-VAD for real-time operation, our framework identifies speaker activities using self-generated embeddings, ensuring consistent performance and avoiding permutation inconsistencies during inference. In the inference phase, we employ a front-end model to extract frame-level speaker embeddings for each incoming signal block. Subsequently, we predict each speaker's detection state based on these frame-level embeddings and the previously estimated target speaker embeddings. The target speaker embeddings are then updated by aggregating the frame-level embeddings according to the current block's predictions. Our model predicts results block-by-block and iteratively updates target speaker embeddings until reaching the end of the signal. Experimental results demonstrate that the proposed method outperforms offline clustering-based diarization systems on the DIHARD III and AliMeeting datasets. Additionally, this approach is extended to multi-channel data, achieving comparable performance to state-of-the-art offline diarization systems. Weiqing Wang 0004, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Transferable Adversarial Attack Against Deep Reinforcement Learning-Based Smart Grid Dynamic Pricing SystemabstractSevere damage caused by transferable adversarial attacks has emerged as a prominent concern in recent years, especially in the smart grid. The security issue of the deep reinforcement learning (DRL)-based dynamic pricing system is directly related to the grid's reliability. Previous works have primarily focused on the attacks' transferability from the perspective of model architecture, whereas the concept of distribution bias offers a novel and relatively underexplored viewpoint. In this work, we propose transferable adversarial attacks with distribution (TAD) targeting the DRL model. The adversary emphasizes destroying the target model with the masqueraded malicious dataset while ensuring stealthiness. Concretely, the masqueraded dataset generated by the attacker is required to have a similar distribution to the original dataset, while perturb some of the critical samples to help the target model misdirect to the nonoptimal policy. To this end, we propose an innovative model named Masquerader, which leverages a variational auto-encoder and incorporates three elaborate loss functions to constrain the distribution and deviation of malicious samples. Extensive experiments in a DRL-based dynamic pricing system indicate that our attack strategy TAD could successfully perturb the target model's output. The aberrant flatness of retail prices and the grid system's reduction in daily profits further validate the attack's transferability and harmfulness. Heng Zhang 0001, Wen Yang 0002, Ming Li 0026, Jian Zhang 0082, Hongran Li |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Haha-POD: An Attempt for Laughter-Based Non-Verbal Speaker VerificationabstractIt is widely acknowledged that discriminative representation for speaker verification can be extracted from verbal speech. However, how much speaker information that non-verbal vocalization carries is still a puzzle. This paper explores speaker verification based on the most ubiquitous form of non-verbal voice, laughter. First, we use a semi-automatic pipeline to collect a new Haha-Pod dataset from open-source podcast media. The dataset contains over 240 speakers’ laughter clips with corresponding high-quality verbal speech. Second, we propose a Two-Stage Teacher-Student (2S-TS) framework to minimize the within-speaker embedding distance between verbal and non-verbal (laughter) signals. Considering Haha-Pod as a test set, two trial sets (S2L-Eval) are designed to verify the speaker’s identity through laugh sounds. Experimental results demonstrate that our method can significantly improve the performance of the S2L-Eval test set with only a minor degradation on the VoxCeleb1 test set. The resources for the Haha-Pod dataset can be found at https://github.com/nevermoreLin/HahaPod. Yuke Lin, Xiaoyi Qin, Ming Li 0026 |
ASRU | 5 |
| 2023 | Bisinger: Bilingual Singing Voice SynthesisabstractAlthough Singing Voice Synthesis (SVS) has made great strides with Text-to-Speech (TTS) techniques, multilingual singing voice modeling remains relatively unexplored. This paper presents BiSinger, a bilingual pop SVS system for English and Chinese Mandarin. Current systems require separate models per language and cannot accurately represent both Chinese and English, hindering code-switch SVS. To address this gap, we design a shared representation between Chinese and English singing voices, achieved by using the CMU dictionary with mapping rules. We fuse monolingual singing datasets with open-source singing voice conversion techniques to generate bilingual singing voices while also exploring the potential use of bilingual speech data. Experiments affirm that our language-independent representation and incorporation of related datasets enable a single model with enhanced performance in English and code-switch SVS while maintaining Chinese song performance. Audio samples are available at https://bisinger-svs.github.io. Huali Zhou, Yueqian Lin, Ming Li 0026 |
ASRU | 5 |
| 2023 | Identifying Source Speakers for Voice Conversion Based Spoofing Attacks on Speaker Verification SystemsabstractAn automatic speaker verification system aims to verify the speaker identity of a speech signal. However, a voice conversion system could manipulate a person’s speech signal to make it sound like another speaker’s voice and deceive the speaker verification system. Most countermeasures for voice conversion-based spoofing attacks are designed to discriminate bona fide speech from spoofed speech for speaker verification systems. In this paper, we investigate the problem of source speaker identification – inferring the identity of the source speaker given the voice converted speech. To perform source speaker identification, we simply add voice-converted speech data with the label of source speaker identity to the genuine speech dataset during speaker embedding network training. Experimental results show the feasibility of source speaker identification when training and testing with converted speeches from the same voice conversion model(s). In addition, our results demonstrate that having more converted utterances from various voice conversion model for training helps improve the source speaker identification performance on converted utterances from unseen voice conversion models. Danwei Cai, Zexin Cai, Ming Li 0026 |
ICASSP | 3 |
| 2023 | Waveform Boundary Detection for Partially Spoofed AudioabstractThe present paper proposes a waveform boundary detection system for audio spoofing attacks containing partially manipulated segments. Partially spoofed/fake audio, where part of the utterance is replaced, either with synthetic or natural audio clips, has recently been reported as one scenario of audio deepfakes. As deepfakes can be a threat to social security, the detection of such spoofing audio is essential. Accordingly, we propose to address the problem with a deep learning-based frame-level detection system that can detect partially spoofed audio and locate the manipulated pieces. Our proposed method is trained and evaluated on data provided by the ADD2022 Challenge. We evaluate our detection model concerning various acoustic features and network configurations. As a result, our detection system achieves an equal error rate (EER) of 6.58% on the ADD2022 challenge test set, which is the best performance in partially spoofed audio detection systems that can locate manipulated clips. Zexin Cai, Weiqing Wang 0004, Ming Li 0026 |
ICASSP | 3 |
| 2023 | Pretraining Conformer with ASR for Speaker VerificationabstractThis paper proposes to pretrain Conformer with automatic speech recognition (ASR) task for speaker verification. Conformer combines convolution neural network (CNN) and Transformer model for modeling local and global features, respectively. Recently, multi-scale feature aggregation Conformer (MFA-Conformer) has been proposed for automatic speaker verification. MFA-Conformer concatenates frame-level outputs from all Conformer blocks for further pooling. However, our experiments show that Conformer can be easily overfitted with limited speaker recognition training data. To avoid overfitting, we propose to transfer the knowledge learned from ASR to speaker verification. Specifically, an ASR pretrained Conformer is used to initialize the training of MFA-Conformer for speaker verification. Our experiments show that pretraining Conformer with ASR leads to significant performance gains across model sizes. The best model achieves 0.48%, 0.71% and 1.54% EER on Voxceleb1-O, Voxceleb1-E, and Voxceleb1-H, respectively. Danwei Cai, Weiqing Wang 0004, Ming Li 0026, Chuanzeng Huang |
ICASSP | 3 |
| 2023 | The WHU-Alibaba Audio-Visual Speaker Diarization System for the MISP 2022 ChallengeabstractThis paper describes the system developed by the WHU-Alibaba team for the Multimodal Information Based Speech Processing (MISP) 2022 Challenge. We extend the Sequence-to-Sequence Target-Speaker Voice Activity Detection framework to simultaneously detect multiple speakers’ voice activities from audio-visual signals. The final system achieves a diarization error rate (DER) of 8.82% on the evaluation set of the competition database, which ranks 1st in the speaker diarization track of the MISP 2022, ICASSP Signal Processing Grand Challenge. Ming Cheng 0005, Haoxu Wang, Qiang Fu 0001, Ming Li 0026 |
ICASSP | 5 |
| 2023 | Target-Speaker Voice Activity Detection Via Sequence-to-Sequence PredictionabstractTarget-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence Target-Speaker Voice Activity Detection (Seq2Seq-TSVAD) method that can efficiently address the joint modeling of large-scale speakers and predict high-resolution voice activities. Experimental results show that larger speaker capacity and higher output resolution can significantly reduce the diarization error rate (DER), which achieves the new state-of-the-art performance of 4.55% on the VoxConverse test set and 10.77% on Track 1 of the DIHARD-III evaluation set under the widely-used evaluation metrics. Ming Cheng 0005, Weiqing Wang 0004, Yucong Zhang, Xiaoyi Qin, Ming Li 0026 |
ICASSP | 5 |
| 2023 | The DKU Post-Challenge Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge: Deep AnalysisabstractThis paper further explores our previous wake word spotting system ranked 2-nd in Track 1 of the MISP Challenge 2021. First, we investigate a robust unimodal approach based on 3D and 2D convolution and adopt the simple attention module (SimAM) for our system to improve performance. Second, we explore different combinations of data augmentation methods for better performance. Finally, we study the fusion strategies, including score-level, cascaded and neural fusion. Our proposed multimodal system leverages multimodal features and uses the complementary visual information to mitigate the performance degradation of audio-only systems in complex acoustic scenarios. Our system obtains a false reject rate of 2.15% and a false alarm rate of 3.44% in the evaluation set of the competition database, which achieves the new state-of-the-art performance by 21% relative improvement compared to previous systems. Related resource can be found at: https://github.com/Mashiro009/DKU_WWS_MISP. Haoxu Wang, Ming Cheng 0005, Qiang Fu 0001, Ming Li 0026 |
ICASSP | 4 |
| 2023 | Exploring Universal Singing Speech Language Identification Using Self-Supervised Learning Based Front-End FeaturesabstractDespite the great performance of language identification (LID), there is a lack of large-scale singing LID databases to support the research of singing language identification (SLID). This paper proposed a over 3200 hours dataset used for singing language identification, called Slingua. As the baseline, we explore two self-supervised learning (SSL) models, WavLM and Wav2vec2, as the feature extractors for both SLID and universal singing speech language identification (ULID), compared with the traditional handcraft feature. Moreover, by training with speech language corpus, we compare the performance difference of the universal singing speech language identification. The final results show that the SSL-based features exhibit more robust generalization, especially for low-resource and open-set scenarios. The database can be downloaded following this repository: https://github.com/Doctor-Do/Slingua. Xingming Wang, Chuanzeng Huang, Ming Li 0026 |
ICASSP | 5 |
| 2023 | Robust Audio Anti-spoofing Countermeasure with Joint Training of Front-end and Back-end Models
Xingming Wang, Bang Zeng, Hongbin Suo, Yulong Wan, Ming Li 0026 |
INTERSPEECH | 5 |
| 2023 | SEF-Net: Speaker Embedding Free Target Speaker Extraction Network
Bang Zeng, Hongbin Suo, Yulong Wan, Ming Li 0026 |
INTERSPEECH | 4 |
| 2023 | Outlier-aware Inlier Modeling and Multi-scale Scoring for Anomalous Sound Detection via Multitask LearningabstractThis paper proposes an approach for anomalous sound detection that incorporates outlier exposure and inlier modeling within a unified framework by multitask learning. While outlier exposure-based methods can extract features efficiently, it is not robust. Inlier modeling is good at generating robust features, but the features are not very effective. Recently, serial approaches are proposed to combine these two methods, but it still requires a separate training step for normal data modeling. To overcome these limitations, we use multitask learning to train a conformer-based encoder for outlier-aware inlier modeling. Moreover, our approach provides multi-scale scores for detecting anomalies. Experimental results on the MIMII and DCASE 2020 task 2 datasets show that our approach outperforms state-of-the-art single-model systems and achieves comparable results with top-ranked multi-system ensembles. Yucong Zhang, Hongbin Suo, Yulong Wan, Ming Li 0026 |
INTERSPEECH | 4 |
| 2023 | Maximizing Throughput in Unmanned Surface Vehicle Relay System under Jamming AttacksabstractIn this paper, we address the issue of jamming attacks in the field of maritime communication and propose the application of reconfigurable intelligent surfaces (RISs) in anti-jamming communication at sea. The RIS is installed on an Unmanned surface vehicle (USV) to construct a RIS-assisted USV relay communication system, which mitigates jamming attacks while enhancing legitimate transmissions. Compared to traditional static RIS, the use of mobile USV with RIS enables better performance and greater flexibility. We jointly optimize the trajectory of the USV, the passive beamforming of the RIS, and the source power allocation for each time slot based on proximal policy optimization (PPO), aiming to maximize the average downlink throughput. Simulation results demonstrate that the deployment of RIS on USV effectively suppresses jamming attacks and protects legitimate transmissions. Heng Zhang 0001, Zhemin Sun, Ming Li 0026, Hongran Li, Jian Zhang 0082 |
MSN | 5 |
| 2023 | Assessing the Social Skills of Children with Autism Spectrum Disorder via Language-Image Pre-training Models
Ming Cheng 0005, Yueran Pan, Lynn Yuan, Suxiu Hu, Ming Li 0026, Songtian Zeng |
PRCV (13) | 6 |
| 2023 | Cross-lingual multi-speaker speech synthesis with limited bilingual training data
Zexin Cai, Yaogen Yang, Ming Li 0026 |
Comput. Speech Lang. | 3 |
| 2023 | STCAM: Spatial-Temporal and Channel Attention Module for Dynamic Facial Expression RecognitionabstractCapturing the dynamics of facial expression progression in video is an essential and challenging task for facial expression recognition (FER). In this article, we propose an effective framework to address this challenge. We develop a C3D-based network architecture, 3D-Inception-ResNet, to extract spatial-temporal features from the dynamic facial expression image sequence. A Spatial-Temporal and Channel Attention Module (STCAM) is proposed to explicitly exploit the holistic spatial-temporal and channel-wise correlations among the extracted features. Specifically, the proposed STCAM calculates a channel-wise and a spatial-temporal-wise attention map to enhance the features along the corresponding feature dimensions for more representative features. We evaluate our method on three popular dynamic facial expression recognition datasets, CK+, Oulu-CASIA, and MMI. Experimental results show that our method achieves better or comparable performance compared to the state-of-the-art approaches. Weicong Chen 0003, Dong Zhang 0002, Ming Li 0026, Dah-Jye Lee |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Computer-Aided Autism Spectrum Disorder Diagnosis With Behavior Signal ProcessingabstractBehavioral observation plays an essential role in the diagnosis of Autism Spectrum Disorder (ASD) by analyzing children's atypical patterns in social activities (e.g., impaired social interaction, restricted interests, and repetitive behavior). To date, this process still heavily relies on the questionnaire survey, clinical observation, or retrospective video analysis, leading to high demand for professionals with massive labor costs. This article proposes a standardized platform for stimulating, gathering, analyzing, modeling, and interpreting human behavioral data in the application of computer-aided ASD diagnosis. By a structured assessment process, the proposed system can automatically evaluate children's multiple social interaction skills using the captured audio-visual data and provide the final diagnostic suggestions. We collect a multimodal behavioral database of 95 participants (71 children with ASD and 24 age-matched typical controls) in a real clinic environment, the Third Affiliated Hospital of Sun Yat-sen University, China. On the clinical database, our proposed computer-aided ASD diagnosis system obtains an accuracy of 88.42% for identifying ASD children with an average age of 24 months, representing a performance comparable to top-level human experts. As a unified and replicable solution, it has good potential to be promoted to less developed areas with limited high-quality medical resources. Ming Cheng 0005, Yixiang Xie, Yueran Pan, Xiao Li 0048, Chengyan Yu, Dong Zhang 0002, Xiaoqian Huang, Cong You, Yuanyuan Zou 0003, Yuchong Liu, Fengjing Liang, Huilin Zhu, Chun Tang, Hongzhu Deng, Xiaobing Zou, Ming Li 0026 |
IEEE Trans. Affect. Comput. | 20 |
| 2023 | Typical Facial Expression Network Using a Facial Feature Decoupler and Spatial-Temporal LearningabstractFacial expression recognition (FER) accuracy is often affected by an individual’s unique facial characteristics. Recognition performance can be improved if the influence from these physical characteristics is minimized. Using video instead of single image for FER provides better results but requires extracting temporal features and the spatial structure of facial expressions in an integrated manner. We propose a new network called Typical Facial Expression Network (TFEN) to address both challenges. TFEN uses two deep two-dimensional (2D) convolutional neural networks (CNNs) to extract facial and expression features from input video. A facial feature decoupler decouples facial features from expression features to minimize the influence from inter-subject face variations. These networks combine with a 3D CNN and form a spatial-temporal learning network to jointly explore the spatial-temporal features in a video. A facial recognition network works as an adversarial network to refine the facial feature decoupler and the network performance by minimizing the residual influence of facial features after decoupling. The whole network is trained with an adversarial algorithm to improve FER performance. TFEN was evaluated on four popular dynamic FER datasets. Experimental results show TFEN achieves or outperforms the recognition accuracy of state-of-the-art approaches. Jianing Teng, Dong Zhang 0002, Ming Li 0026, Dah-Jye Lee |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | Robust Multi-Channel Far-Field Speaker Verification Under Different In-Domain Data Availability ScenariosabstractThe popularity and application of smart home devices have made far-field speaker verification an urgent need. However, speaker verification performance is unsatisfactory under far-field environments despite its significant improvements enabled by deep neural networks (DNN). In this paper, we summarize our previous work and propose multiple training strategies and models for multi-channel far-field speaker verification with different in-domain data availability scenarios. The experiments are conducted on the FFSVC20 dataset, and we proposed the cross-device and cross-domain trials. We focus on single-channel and multi-channel speaker verification training based on the dataset. For single-channel speaker verification, considering the size of training data and availability of labels, we introduce three training scenarios and given our proposed training methods, including 1) given zero out-of-domain data and few in-domain labeled data; 2) given large-scale out-of-domain labeled data and few in-domain labeled data; 3) given large-scale out-of-domain labeled data and few in-domain unlabeled data. To this end, we propose a meta-learning approach, refined transfer learning methods, and semi-supervised learning for three scenarios, respectively. For multi-channel speaker verification, we first introduce two types of 3 dimension convolution (3D Conv) residual network (ResNet) models proposed in our previous works, including fully 3D ResNet and incorporating 3D Conv with 2D Conv ResNet (3D2D-ResNet). In this paper, we propose channel-wise 3D squeeze-and-excitation ResNet (C3DSE-ResNet) and spatial-wise 3D SE ResNet (S3DSE-ResNet) to further explore the channel dependencies and improve the 3D ConvNet performance. The results show that the proposed strategies and models can significantly boost performance under the far-field scenario. Xiaoyi Qin, Danwei Cai, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Accurate Head Pose Estimation Using Image Rectification and a Lightweight Convolutional Neural NetworkabstractHead pose estimation is an important step for many human-computer interaction applications such as face detection, facial recognition, and facial expression classification. Accurate head pose estimation benefits these applications that require face images as the input. Most head pose estimation methods suffer from perspective distortion because the users do not always align their face perfectly with the camera. This paper presents a new approach that uses image rectification to reduce the negative effect of perspective distortion and a lightweight convolutional neural network to obtain highly accurate head pose estimations. The proposed method calculates the angle between the optical axis of the camera and the projection vector of the center of the face. The face image is rectified using this estimated angle through perspective transformation. A lightweight network that is only 0.88 MB in size is designed to take the rectified face image as the input to perform head pose estimation. The output of the network, the head pose estimation of the rectified face image, is transformed back to the camera coordinate system as the final head pose estimation. Experiments on public benchmark datasets show that the proposed image rectification method and the newly designed lightweight network improve the accuracy of head pose estimation remarkably. Compared with state-of-the-art methods, our approach achieves both higher accuracy and faster processing speed. Xiao Li 0048, Dong Zhang 0002, Ming Li 0026, Dah-Jye Lee |
IEEE Trans. Multim. | 3 |
| 2022 | The DKU Audio-Visual Wake Word Spotting System for the 2021 MISP ChallengeabstractThis paper describes the system developed by the DKU team for the MISP Challenge 2021. We present a two-stage approach consisting of end-to-end neural networks for the audio-visual wake word spotting task. We first process audio and video data to give them a similar structure and then train two unimodal models with unified network architecture separately. Second, we propose a Hierarchical Modality Aggregation (HMA) module that fuses multi-scale audio-visual information from pre-trained unimodal models. Our system has a clear and concise framework consisting of end-to-end neural networks. With this framework and extensive data augmentation methods, our presented system achieves a false reject rate of 3.85% and a false alarm rate of 3.42% on far-field audio in the development set of the competition database, which ranks 2nd in the wake word spotting track of the MISP challenge. Ming Cheng 0005, Haoxu Wang, Yechen Wang, Ming Li 0026 |
ICASSP | 4 |
| 2022 | Towards Lightweight Applications: Asymmetric Enroll-Verify Structure for Speaker VerificationabstractWith the development of deep learning, automatic speaker verification has made considerable progress over the past few years. However, to design a lightweight and robust system with limited computational resources is still a challenging problem. Traditionally, a speaker verification system is symmetrical, indicating that the same embedding extraction model is applied for both enrollment and verification in inference. In this paper, we come up with an innovative asymmetric structure, which takes the large-scale ECAPA-TDNN model for enrollment and the small-scale ECAPA-TDNNLite model for verification. As a symmetrical system, our proposed ECAPA-TDNNLite model achieves an EER of 3.07% on the Voxceleb1 original test set with only 11.6M FLOPS. Moreover, the asymmetric structure further reduces the EER to 2.31%, without increasing any computational costs during verification. Qingjian Li, Lin Yang 0014, Xiaoyi Qin, Junjie Wang 0010, Ming Li 0026 |
ICASSP | 6 |
| 2022 | Simple Attention Module Based Speaker Verification with Iterative Noisy Label DetectionabstractRecently, the attention mechanism such as squeeze-and-excitation module (SE) and convolutional block attention module (CBAM) has achieved great success in deep learning-based speaker verification system. This paper introduces an alternative effective yet simple one, i.e., simple attention module (SimAM), for speaker verification. The SimAM module is a plug-and-play module without extra modal parameters. In addition, we propose a noisy label detection method to iteratively filter out the data samples with a noisy label from the training data, considering that a large-scale dataset labeled with human annotation or other automated processes may contain noisy labels. Data with the noisy label may over parameterize a deep neural network (DNN) and result in a performance drop due to the memorization effect of the DNN. Experiments are conducted on VoxCeleb dataset. The speaker verification model with SimAM achieves the 0.675% equal error rate (EER) on VoxCeleb1 original test trials. Our proposed iterative noisy label detection method further reduces the EER to 0.643%. Xiaoyi Qin, Na Li 0012, Chao Weng, Dan Su 0002, Ming Li 0026 |
ICASSP | 5 |
| 2022 | Incorporating End-to-End Framework Into Target-Speaker Voice Activity DetectionabstractIn this paper, we propose an end-to-end target-speaker voice activity detection (E2E-TS-VAD) method for speaker diarization. First, a ResNet-based network extracts the frame-level speaker embeddings from the acoustic features. Then, the L2-normalized frame-level speaker embeddings are fed to the transformer encoder which produces the initialization of the speaker diarization results. Later, the frame-level speaker embeddings are aggregated to several target-speaker embeddings based on the output from the transformer encoder. Finally, a BiLSTM-based TS-VAD model predicts the refined diarization results. Several aggregation methods are explored, including soft/hard decisions with/without normalization. Results show that E2E-TS-VAD achieves better performance than the original TS-VAD method with the clustering-based initialization. Weiqing Wang 0004, Ming Li 0026 |
ICASSP | 2 |
| 2022 | Cross-Channel Attention-Based Target Speaker Voice Activity Detection: Experimental Results for the M2met ChallengeabstractDukeECE. As the highly overlapped speech exists in the dataset, we employ an x-vector-based target-speaker voice activity detection (TS-VAD) to find the overlap between speakers. Firstly, we separately train a single-channel model for each of the 8 channels and fuse the results. In addition, we also employ the cross-channel self-attention to further improve the performance, where the non-linear spatial correlations between different channels are learned and fused. Experimental results on the evaluation set show that the single-channel TS-VAD reduces the DER by over 75% from 12.68% to 3.14%. The multi-channel TS-VAD further reduces the DER by 28% and achieves a DER of 2.26%. Our final submitted system achieves a DER of 2.98% on the AliMeeting test set, which ranks 1st in the M2MET challenge. In this challenge, our team is denoted as A41. Weiqing Wang 0004, Xiaoyi Qin, Ming Li 0026 |
ICASSP | 3 |
| 2022 | SIG-VC: A Speaker Information Guided Zero-Shot Voice Conversion System for Both Human Beings and MachinesabstractNowadays, as more and more systems achieve good performance in traditional voice conversion (VC) tasks, people’s attention gradually turns to VC tasks under extreme conditions. In this paper, we propose a novel method for zero-shot voice conversion. We aim to obtain intermediate representations for speaker-content disentanglement of speech to better remove speaker information and get pure content information. Accordingly, our proposed framework contains a module that removes the speaker information from the acoustic feature of the source speaker. Moreover, speaker information control is added to our system to maintain the voice cloning performance. The proposed system is evaluated by subjective and objective metrics. Results show that our proposed system significantly reduces the trade-off problem in zero-shot voice conversion, while it also manages to have high spoofing power to the speaker verification system. Zexin Cai, Xiaoyi Qin, Ming Li 0026 |
ICASSP | 4 |
| 2022 | A Multimodal Framework for Automated Teaching Quality Assessment of One-to-many Online Instruction VideosabstractIn the post-pandemic era, online courses have been adopted universally. Manually assessing online course teaching quality requires significant time and professional pedagogy experience. To address this problem, we design an evaluation protocol and propose a multimodal machine learning framework1for automated teaching quality assessment of one-to-many online instruction videos. Our framework evaluates online teaching quality from five aspects, namely Clarity, Classroom interaction, Technical management of online teaching, Empathy, and Time management. Our method includes mid-level behavior feature extraction, high-level interpretable feature extraction, and supervised learning prediction. Our automated multimodal teaching quality assessment system achieves comparable performance to human annotators on our one-to-many online instruction videos. For binary classification, the best average accuracy of five aspects is 0.898. For regression, the best average means square error is 0.527 on a 0-10 scale. Yueran Pan, Ran Ju, Ziang Zhou, Jiayue Gu, Songtian Zeng, Lynn Yuan, Ming Li 0026 |
ICPR | 8 |
| 2022 | Cross-Age Speaker Verification: Learning Age-Invariant Speaker EmbeddingsabstractAutomatic speaker verification has achieved remarkable progress in recent years.However, there is little research on cross-age speaker verification (CASV) due to insufficient relevant data.In this paper, we mine cross-age test sets based on the VoxCeleb dataset and propose our age-invariant speaker representation(AISR) learning method.Since the VoxCeleb is collected from the YouTube platform, the dataset consists of crossage data inherently.However, the meta-data does not contain the speaker age label.Therefore, we adopt the face age estimation method to predict the speaker age value from the associated visual data, then label the audio recording with the estimated age.We construct multiple Cross-Age test sets on VoxCeleb (Vox-CA), which deliberately select the positive trials with large age-gap.Also, the effect of nationality and gender is considered in selecting negative pairs to align with Vox-H cases.The baseline system performance drops from 1.939% EER on the Vox-H test set to 10.419% on the Vox-CA20 test set, which indicates how difficult the cross-age scenario is.Consequently, we propose an age-decoupling adversarial learning (ADAL) method to alleviate the negative effect of the age gap and reduce intra-class variance.Our method outperforms the baseline system by over 10% related EER reduction on the Vox-CA20 test set.The source code and trial resources are available on https://github.com/qinxiaoyi/Cross-AgeSpeaker Verification. Xiaoyi Qin, Na Li 0012, Chao Weng, Dan Su 0002, Ming Li 0026 |
INTERSPEECH | 5 |
| 2022 | Online Target Speaker Voice Activity Detection for Speaker Diarization
Weiqing Wang 0004, Ming Li 0026, Qingjian Lin |
INTERSPEECH | 2 |
| 2022 | The DKU-OPPO System for the 2022 Spoofing-Aware Speaker Verification ChallengeabstractThis paper describes our DKU-OPPO system for the 2022 Spoofing-Aware Speaker Verification (SASV) Challenge.First, we split the joint task into speaker verification (SV) and spoofing countermeasure (CM), these two tasks which are optimized separately.For ASV systems, four state-of-the-art methods are employed.For CM systems, we propose two methods on top of the challenge baseline to further improve the performance, namely Embedding Random Sampling Augmentation (ERSA) and One-Class Confusion Loss(OCCL).Second, we also explore whether SV embedding could help improve CM system performance.We observe a dramatic performance degradation of existing CM systems on the domain-mismatched Voxceleb2 dataset.Third, we compare different fusion strategies, including parallel score fusion and sequential cascaded systems.Compared to the 1.71% SASV-EER baseline, our submitted cascaded system obtains a 0.21% SASV-EER on the challenge official evaluation set. Xingming Wang, Xiaoyi Qin, Yikang Wang, Ming Li 0026 |
INTERSPEECH | 5 |
| 2022 | Incorporating Visual Information in Audio Based Self-Supervised Speaker RecognitionabstractThe currentsuccess of deep learning largely benefits from the availability of large amount of labeled data. However, collecting a large-scale dataset with human annotation can be expensive and sometimes difficult. Self-supervised learning thus attracts many research interests to train models without labels. In this paper, we propose a self-supervised learning framework for speaker recognition. Combining clustering with deep representation learning, the proposed framework generates pseudo labels for the unlabeled dataset and learns speaker representation without human annotation. Our method starts with training a speaker representation encoder with contrastive self-supervised learning. Clustering on the learned representation generates pseudo labels, which are used as the supervisory signal for the subsequent training of the representation encoder. The clustering and representation learning process is performed iteratively to bootstrap the discriminative power of the deep neural network. We apply this self-supervised learning framework to both single modal audio data and multi-modal audio-visual data. For audio-visual data, audio and visual representation encoders are employed to learn representations of the corresponding modality. A cluster ensemble algorithm is then used to fuse the clustering results of the two modalities. The complementary information in multi-modalities ensures a robust and fault-tolerant supervisory signal for audio and visual representation learning. Experimental results show that our proposed iterative self-supervised learning framework outperforms previous works with self-supervision by large margins. Training with single modal audio data on the development set of VoxCeleb 2, our proposed framework achieves an equal error rate (EER) of 2.8% on the original test trials of VoxCeleb 1. When training with additional visual modality, the EER further reduces to 1.8%, which is only 20% higher than the fully supervised audio-based system with an EER of 1.5%. Also, experimental analysis shows that the proposed framework generates pseudolabels that are highly correlated to ground truth labels. Danwei Cai, Weiqing Wang 0004, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Similarity Measurement of Segment-Level Speaker Embeddings in Speaker DiarizationabstractIn this paper, we propose a neural-network-based similarity measurement method to learn the similarity between any two speaker embeddings, where both previous and future contexts are considered. Moreover, we propose the segmental pooling strategy and jointly train the speaker embedding network along with the similarity measurement model. Later, this joint training framework is further extended to the target-speaker voice activity detection (TS-VAD), with only slight modification in the network architecture. Experimental results of the DIHARD II, DIHARD III and VoxConverse datasets show that our clustering-based system with the neural similarity measurement achieves superior performance to recent approaches on all three datasets. In addition, the segment-level TS-VAD method further improves the clustering-based results and achieves DER of 16.48%, 11.62% and 4.39% on the DIHARD II, DIHARD III and VoxConverse datasets, respectively. Weiqing Wang 0004, Qingjian Lin, Danwei Cai, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | An Iterative Framework for Self-Supervised Deep Speaker Representation LearningabstractIn this paper, we propose an iterative framework for self-supervised speaker representation learning based on a deep neural network (DNN). The framework starts with training a self-supervision speaker embedding network by maximizing agreement between different segments within an utterance via a contrastive loss. Taking advantage of DNN’s ability to learn from data with label noise, we propose to cluster the speaker embedding obtained from the previous speaker network and use the subsequent class assignments as pseudo labels to train a new DNN. Moreover, we iteratively train the speaker network with pseudo labels generated from the previous step to bootstrap the discriminative power of a DNN. Speaker verification experiments are conducted on the VoxCeleb dataset. The results show that our proposed iterative self-supervised learning framework outperformed previous works using self-supervision. The speaker network after 5 iterations obtains a 61% performance gain over the speaker embedding model trained with contrastive loss. Danwei Cai, Weiqing Wang 0004, Ming Li 0026 |
ICASSP | 3 |
| 2021 | Cross-modal Assisted Training for Abnormal Event Recognition in ElevatorsabstractGiven that very few action recognition datasets collected in elevators contain multimodal data, we collect and propose our multimodal dataset investigating passenger safety and inappropriate elevator usage. Moreover, we present a novel framework (RGBP) to utilize multimodal data to enhance unimodal test performance for the task of abnormal event recognition in elevators. Experimental results show that the best network architecture with the RGBP framework effectively improves the unimodal inference performance on the Elevator RGBD dataset by 4.71% (accuracy) and 4.95% (F1 score) with respect to the pure RGB model. In addition, our RGBP framework outperforms two other methods for ”multimodal training and unimodal inference”: MTUT [1] and the two-stage method based on depth estimation. Xinmeng Chen, Xuchen Gong, Ming Cheng 0005, Ming Li 0026 |
ICMI | 5 |
| 2021 | The 2020 Personalized Voice Trigger Challenge: Open Datasets, Evaluation Metrics, Baseline System and Results
Xingming Wang, Xiaoyi Qin, Yinping Zhang, Junjie Wang 0010, Dong Zhang 0002, Ming Li 0026 |
Interspeech | 8 |
| 2021 | Our Learned Lessons from Cross-Lingual Speaker Verification: The CRMI-DKU System Description for the Short-Duration Speaker Verification Challenge 2021
Xiaoyi Qin, Chao Wang 0111, Shilei Zhang, Ming Li 0026 |
Interspeech | 6 |
| 2021 | AISHELL-3: A Multi-Speaker Mandarin TTS Corpus
Hui Bu, Shaoji Zhang, Ming Li 0026 |
Interspeech | 5 |
| 2021 | The DKU-Duke-Lenovo System Description for the Fearless Steps Challenge Phase III
Weiqing Wang 0004, Danwei Cai, Qingjian Lin, Mi Hong, Ming Li 0026 |
Interspeech | 7 |
| 2021 | Binary Neural Network for Speaker VerificationabstractAlthough deep neural networks are successful for many tasks in the speech domain, the high computational and memory costs of deep neural networks make it difficult to directly deploy highperformance Neural Network systems on low-resource embedded devices. There are several mechanisms to reduce the size of the neural networks i.e. parameter pruning, parameter quantization, etc. This paper focuses on how to apply binary neural networks to the task of speaker verification. The proposed binarization of training parameters can largely maintain the performance while significantly reducing storage space requirements and computational costs. Experiment results show that, after binarizing the Convolutional Neural Network, the ResNet34-based network achieves an EER of around 5% on the Voxceleb1 testing dataset and even outperforms the traditional real number network on the text-dependent dataset: Xiaole while having a 32x memory saving. Tinglong Zhu, Xiaoyi Qin, Ming Li 0026 |
Interspeech | 3 |
| 2021 | Embedding Aggregation for Far-Field Speaker Verification with Distributed Microphone ArraysabstractWith the successful application of deep speaker embedding networks, the performance of speaker verification systems has significantly improved under clean and close-talking settings; however, unsatisfactory performance persists under noisy and far-field environments. This study aims at improving the performance of far-field speaker verification systems with distributed microphone arrays in the smart home scenario. The proposed learning framework consists of two modules: a deep speaker embedding module and an aggregation module. The former extracts a speaker embedding for each recording. The latter, based on either averaged pooling or attentive pooling, aggregates speaker embeddings and learns a unified representation for all recordings captured by distributed microphone arrays. The two modules are trained in an end-to-end manner. To evaluate this framework, we conduct experiments on the real text-dependent far-field datasets Hi Mia. Results show that our framework outperforms the naive averaged aggregation methods by 20% in terms of equal error rate (EER) with six distributed microphone arrays. Also, we find that the attention-based aggregation advocates high-quality recordings and repels low-quality ones. Danwei Cai, Ming Li 0026 |
SLT | 2 |
| 2021 | Facial Expression Recognition with Identity and Emotion Joint LearningabstractDifferent subjects may express a specific expression in different ways due to inter-subject variabilities. In this work, besides training deep-learned facial expression feature (emotional feature), we also consider the influence of latent face identity feature such as the shape or appearance of face. We propose an identity and emotion joint learning approach with deep convolutional neural networks (CNNs) to enhance the performance of facial expression recognition (FER) tasks. First, we learn the emotion and identity features separately using two different CNNs with their corresponding training data. Second, we concatenate these two features together as a deep-learned Tandem Facial Expression (TFE) Feature and feed it to the subsequent fully connected layers to form a new model. Finally, we perform joint learning on the newly merged network using only the facial expression training data. Experimental results show that our proposed approach achieves 99.31 and 84.29 percent accuracy on the CK+ and the FER+ database, respectively, which outperforms the residual network baseline as well as many other state-of-the-art methods. Ming Li 0026, Xingchang Huang, Zhanmei Song, Xin Li 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | Audio-Based Piano Performance Evaluation for Beginners With Convolutional Neural Network and Attention MechanismabstractIn this paper, we propose two different audio-based piano performance evaluation systems for beginners. The first is a sequential and modularized system, including three steps: Convolutional Neural Network (CNN)-based acoustic feature extraction, matching via dynamic time warping (DTW), and performance score regression. The second system is an end-to-end system with CNNs and the attention mechanism. It takes two acoustic feature sequences as input and directly predicts a performance score. We evaluate two proposed methods with our new open-access Yingcai Piano Performance Evaluation Phase III Dataset (YCU-PPE-III) that contains more than 2000 piano audio pieces recorded in multiple real test sessions. Experimental results show that the modularized system achieves a mean absolute error (MAE) of 3.79 in a 0-100-point range. Another end-to-end system also achieves an MAE of 4.40, which shows that it is possible to train a robust end-to-end piano performance evaluation system with only two thousand audio pieces. Weiqing Wang 0004, Hua Yi, Zhanmei Song, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | Within-Sample Variability-Invariant Loss for Robust Speaker Recognition Under Noisy EnvironmentsabstractDespite the significant improvements in speaker recognition enabled by deep neural networks, unsatisfactory performance persists under noisy environments. In this paper, we train the speaker embedding network to learn the "clean" embedding of the noisy utterance. Specifically, the network is trained with the original speaker identification loss with an auxiliary within-sample variability-invariant loss. This auxiliary variability-invariant loss is used to learn the same embedding among the clean utterance and its noisy copies and prevents the network from encoding the undesired noises or variabilities into the speaker representation. Furthermore, we investigate the data preparation strategy for generating clean and noisy utterance pairs on-the-fly. The strategy generates different noisy copies for the same clean utterance at each training step, helping the speaker embedding network generalize better under noisy environments. Experiments on VoxCeleb1 indicate that the proposed training framework improves the performance of the speaker verification system in both clean and noisy conditions. Danwei Cai, Weicheng Cai, Ming Li 0026 |
ICASSP | 3 |
| 2020 | HI-MIA: A Far-Field Text-Dependent Speaker Verification Database and the BaselinesabstractThis paper presents a far-field text-dependent speaker verification database named HI-MIA. We aim to meet the data requirement for far-field microphone array based speaker verification since most of the publicly available databases are single channel close-talking and text-independent. The database contains recordings of 340 people in rooms designed for the far-field scenario. Recordings are captured by multiple microphone arrays located in different directions and distance to the speaker and a high-fidelity close-talking microphone. Besides, we propose a set of end-to-end neural network based baseline systems that adopt single-channel data for training. Moreover, we propose a testing background aware enrollment augmentation strategy to further enhance the performance. Results show that the fusion systems could achieve 3.29% EER in the far-field enrollment far field testing task and 4.02% EER in the close-talking enrollment and far-field testing task. Xiaoyi Qin, Hui Bu, Ming Li 0026 |
ICASSP | 3 |
| 2020 | RWF-2000: An Open Large Scale Video Database for Violence DetectionabstractIn recent years, surveillance cameras are widely deployed in public places, and the general crime rate has been reduced significantly due to these ubiquitous devices. Usually, these cameras provide cues and evidence after crimes are conducted, while they are rarely used to prevent or stop criminal activities in time. It is both time and labor consuming to manually monitor a large amount of video data from surveillance cameras. Therefore, automatically recognizing violent behaviors from video signals becomes essential. This paper summarizes several existing video datasets for violence detection and proposes the RWF-2000 database with 2,000 videos captured by surveillance cameras in real-world scenes. Also, we present a new method that utilizes both the merits of 3D-CNNs and optical flow, namely Flow Gated Network. The proposed approach obtains an accuracy of 87.25% on the test set of our proposed database. The database and source codes are currently open to access. Ming Cheng 0005, Kunjing Cai, Ming Li 0026 |
ICPR | 3 |
| 2020 | Responsive Social Smile: A Machine Learning based Multimodal Behavior Assessment Framework towards Early Stage Autism ScreeningabstractAutism spectrum disorder (ASD) is a neuro-developmental disorder, which causes deficits in social lives. Early screening of ASD for young children is important to reduce the impact of ASD on people's lives. Traditional screening methods mainly rely on protocol-based interviews and subjective evaluations from clinicians and domain experts, which requires advanced expertise and intensive labor. To standardize the process of ASD screening, we design a “Responsive Social Smile” protocol and the associated experimental setup. Moreover, we propose a machine learning based assessment framework for early ASD screening. By integrating speech recognition and computer vision technologies, the proposed framework can quantitatively analyze children's behaviors under well-designed protocols. We collect 196 stimulus samples from 41 children with an average age of 23.34 months, and the proposed method obtains 85.20% accuracy for predicting stimulus scores and 80.49% accuracy for the final ASD prediction. This result indicates that our model approaches the average level of domain experts in this “Responsive Social Smile” protocol. Yueran Pan, Kunjing Cai, Ming Cheng 0005, Xiaobing Zou, Ming Li 0026 |
ICPR | 5 |
| 2020 | From Speaker Verification to Multispeaker Speech Synthesis, Deep Transfer with Feedback ConstraintabstractHigh-fidelity speech can be synthesized by end-to-end text-tospeech models in recent years.However, accessing and controlling speech attributes such as speaker identity, prosody, and emotion in a text-to-speech system remains a challenge.This paper presents a system involving feedback constraints for multispeaker speech synthesis.We manage to enhance the knowledge transfer from the speaker verification to the speech synthesis by engaging the speaker verification network.The constraint is taken by an added loss related to the speaker identity, which is centralized to improve the speaker similarity between the synthesized speech and its natural reference audio.The model is trained and evaluated on publicly available datasets.Experimental results, including visualization on speaker embedding space, show significant improvement in terms of speaker identity cloning in the spectrogram level.In addition, synthesized samples are available online for listening. 1 Zexin Cai, Chuxiong Zhang, Ming Li 0026 |
INTERSPEECH | 3 |
| 2020 | Atss-Net: Target Speaker Separation via Attention-Based Neural NetworkabstractRecently, Convolutional Neural Network (CNN) and Long short-term memory (LSTM) based models have been introduced to deep learning-based target speaker separation. In this paper, we propose an Attention-based neural network (Atss-Net) in the spectrogram domain for the task. It allows the network to compute the correlation between each feature parallelly, and using shallower layers to extract more features, compared with the CNN-LSTM architecture. Experimental results show that our Atss-Net yields better performance than the VoiceFilter, although it only contains half of the parameters. Furthermore, our proposed model also demonstrates promising performance in speech enhancement. Tingle Li, Qingjian Lin, Yuanyuan Bao, Ming Li 0026 |
INTERSPEECH | 4 |
| 2020 | Self-Attentive Similarity Measurement Strategies in Speaker Diarization
Qingjian Lin, Ming Li 0026 |
INTERSPEECH | 3 |
| 2020 | The DKU Speech Activity Detection and Speaker Identification Systems for Fearless Steps Challenge Phase-02
Qingjian Lin, Tingle Li, Ming Li 0026 |
INTERSPEECH | 3 |
| 2020 | The INTERSPEECH 2020 Far-Field Speaker Verification ChallengeabstractThe INTERSPEECH 2020 Far-Field Speaker Verification Challenge (FFSVC 2020) addresses three different research problems under well-defined conditions: far-field text-dependent speaker verification from single microphone array, far-field textindependent speaker verification from single microphone array, and far-field text-dependent speaker verification from distributed microphone arrays.All three tasks pose a cross-channel challenge to the participants.To simulate the real-life scenario, the enrollment utterances are recorded from close-talk cellphone, while the test utterances are recorded from the far-field microphone arrays.In this paper, we describe the database, the challenge, and the baseline system, which is based on a ResNetbased deep speaker network with cosine similarity scoring.For a given utterance, the speaker embeddings of different channels are equally averaged as the final embedding.The baseline system achieves minDCFs of 0.62, 0.66, and 0.64 and EERs of 6.27%, 6.55%, and 7.18% for task 1, task 2, and task 3, respectively. Xiaoyi Qin, Ming Li 0026, Hui Bu, Wei Rao 0002, Rohan Kumar Das, Shri Narayanan, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2020 | Domain Aware Training for Far-Field Small-Footprint Keyword SpottingabstractIn this paper, we focus on the task of small-footprint keyword spotting under the far-field scenario.Far-field environments are commonly encountered in real-life speech applications, causing severe degradation of performance due to room reverberation and various kinds of noises.Our baseline system is built on the convolutional neural network trained with pooled data of both far-field and close-talking speech.To cope with the distortions, we develop three domain aware training systems, including the domain embedding system, the deep CORAL system, and the multi-task learning system.These methods incorporate domain knowledge into network training and improve the performance of the keyword classifier on far-field conditions.Experimental results show that our proposed methods manage to maintain the performance on the close-talking speech and achieve significant improvement on the far-field test set. Haiwei Wu, Yuanfei Nie, Ming Li 0026 |
INTERSPEECH | 4 |
| 2020 | On-the-Fly Data Loader and Utterance-Level Aggregation for Speaker and Language RecognitionabstractIn this article, our recent efforts on directly modeling utterance-level aggregation for speaker and language recognition is summarized. First, an on-the-fly data loader for efficient network training is proposed. The data loader acts as a bridge between the full-length utterances and the network. It generates mini-batch samples on the fly, which allows batch-wise variable-length training and online data augmentation. Second, the traditional dictionary learning and Baum-Welch statistical accumulation mechanisms are applied to the network structure, and a learnable dictionary encoding (LDE) layer is introduced. The former accumulates discriminative statistics from the variable-length input sequence and outputs a single fixed-dimensional utterance-level representation. Experiments were conducted on four different datasets, namely NIST LRE 2007, AP17-OLR, SITW, and NIST SRE 2016. Experimental results show the effectiveness of the proposed batch-wise variable-length training with online data augmentation and the LDE layer, which significantly outperforms the baseline methods. Weicheng Cai, Jinkun Chen, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Utterance-level End-to-end Language Identification Using Attention-based CNN-BLSTMabstractIn this paper, we present an end-to-end language identification framework, the attention-based Convolutional Neural Network-Bidirectional Long-short Term Memory (CNN-BLSTM). The model is performed on the utterance level, which means the utterance-level decision can be directly obtained from the output of the neural network. To handle speech utterances with entire arbitrary and potentially long duration, we combine CNN-BLSTM model with a self-attentive pooling layer together. The front-end CNN-BLSTM module plays a role as local pattern extractor for the variable-length inputs, and the following self-attentive pooling layer is built on top to get the fixed-dimensional utterance-level representation. We conducted experiments on NIST LRE07 closed-set task, and the results reveal that the proposed attention-based CNN-BLSTM model achieves comparable error reduction with other state-of-the-art utterance-level neural network approaches for all 3 seconds, 10 seconds, 30 seconds duration tasks. Weicheng Cai, Danwei Cai, Shen Huang, Ming Li 0026 |
ICASSP | 4 |
| 2019 | F0 Contour Estimation Using Phonetic Feature in Electrolaryngeal Speech EnhancementabstractPitch plays a significant role in understanding a tone based language like Mandarin. In this paper, we present a new method that estimates F0 contour for electrolaryngeal (EL) speech enhancement in Mandarin. Our system explores the usage of phonetic feature to improve the quality of EL speech. First, we train an acoustic model for EL speech and generate the phoneme posterior probabilities feature sequence for each input EL speech utterance. Then we employ the phonetic feature for F0 contour generation rather than the acoustic feature. The experimental results indicate that the EL speech is significantly enhanced under the adoption of the phonetic feature. Experimental results demonstrate that the proposed method achieves notable improvement regarding the intelligibility and the similarity with normal speech. Zexin Cai, Ming Li 0026 |
ICASSP | 3 |
| 2019 | The DKU-SMIIP System for NIST 2018 Speaker Recognition EvaluationabstractIn this paper, we present the system submission for the NIST 2018 Speaker Recognition Evaluation by DKU Speech and Multi-Modal Intelligent Information Processing (SMIIP) Lab.We explore various kinds of state-of-the-art front-end extractors as well as back-end modeling for text-independent speaker verifications.Our submitted primary systems employ multiple state-of-the-art front-end extractors, including the MFCC i-vector, the DNN tandem i-vector, the TDNN x-vector, and the deep ResNet.After speaker embedding is extracted, we exploit several kinds of back-end modeling to perform variability compensation and domain adaptation for mismatch training and testing conditions.The final submitted system on the fixed condition obtains actual detection cost of 0.392 and 0.494 on CMN2 and VAST evaluation data respectively.After the official evaluation, we further extend our experiments by investigating multiple encoding layer designs and loss functions for the deep ResNet system. Danwei Cai, Weicheng Cai, Ming Li 0026 |
INTERSPEECH | 3 |
| 2019 | The DKU System for the Speaker Recognition Task of the 2019 VOiCES from a Distance ChallengeabstractIn this paper, we present the DKU system for the speaker recognition task of the VOiCES from a distance challenge 2019. We investigate the whole system pipeline for the far-field speaker verification, including data pre-processing, short-term spectral feature representation, utterance-level speaker modeling, back-end scoring, and score normalization. Our best single system employs a residual neural network trained with angular softmax loss. Also, the weighted prediction error algorithms can further improve performance. It achieves 0.3668 minDCF and 5.58% EER on the evaluation set by using a simple cosine similarity scoring. Finally, the submitted primary system obtains 0.3532 minDCF and 4.96% EER on the evaluation set. Danwei Cai, Xiaoyi Qin, Weicheng Cai, Ming Li 0026 |
INTERSPEECH | 4 |
| 2019 | Multi-Channel Training for End-to-End Speaker Recognition Under Reverberant and Noisy Environment
Danwei Cai, Xiaoyi Qin, Ming Li 0026 |
INTERSPEECH | 3 |
| 2019 | The DKU Replay Detection System for the ASVspoof 2019 Challenge: On Data Augmentation, Feature Representation, Classification, and FusionabstractThis paper describes our DKU replay detection system for the ASVspoof 2019 challenge.The goal is to develop spoofing countermeasure for automatic speaker recognition in physical access scenario.We leverage the countermeasure system pipeline from four aspects, including the data augmentation, feature representation, classification, and fusion.First, we introduce an utterance-level deep learning framework for antispoofing.It receives the variable-length feature sequence and outputs the utterance-level scores directly.Based on the framework, we try out various kinds of input feature representations extracted from either the magnitude spectrum or phase spectrum.Besides, we also perform the data augmentation strategy by applying the speed perturbation on the raw waveform.Our best single system employs a residual neural network trained by the speed-perturbed group delay gram.It achieves EER of 1.04% on the development set, as well as EER of 1.08% on the evaluation set.Finally, using the simple average score from several single systems can further improve the performance.EER of 0.24% on the development set and 0.66% on the evaluation set is obtained for our primary system. Weicheng Cai, Haiwei Wu, Danwei Cai, Ming Li 0026 |
INTERSPEECH | 4 |
| 2019 | Polyphone Disambiguation for Mandarin Chinese Using Conditional Neural Network with Multi-Level Embedding FeaturesabstractThis paper describes a conditional neural network architecture for Mandarin Chinese polyphone disambiguation.The system is composed of a bidirectional recurrent neural network component acting as a sentence encoder to accumulate the context correlations, followed by a prediction network that maps the polyphonic character embeddings along with the conditions to corresponding pronunciations.We obtain the word-level condition from a pre-trained word-to-vector lookup table.One goal of polyphone disambiguation is to address the homograph problem existing in the front-end processing of Mandarin Chinese textto-speech system.Our system achieves an accuracy of 94.69% on a publicly available polyphonic character dataset.To further validate our choices on the conditional feature, we investigate polyphone disambiguation systems with multi-level conditions respectively.The experimental results show that both the sentence-level and the word-level conditional embedding features are able to attain good performance for Mandarin Chinese polyphone disambiguation. Zexin Cai, Yaogen Yang, Chuxiong Zhang, Xiaoyi Qin, Ming Li 0026 |
INTERSPEECH | 5 |
| 2019 | Survey Talk: End-to-End Deep Neural Network Based Speaker and Language Recognition
Ming Li 0026, Weicheng Cai, Danwei Cai |
INTERSPEECH | 1 |
| 2019 | LSTM Based Similarity Measurement with Spectral Clustering for Speaker DiarizationabstractInternational audience Qingjian Lin, Ruiqing Yin, Ming Li 0026, Hervé Bredin, Claude Barras |
INTERSPEECH | 3 |
| 2019 | Far-Field End-to-End Text-Dependent Speaker Verification Based on Mixed Training Data with Transfer Learning and Enrollment Data Augmentation
Xiaoyi Qin, Danwei Cai, Ming Li 0026 |
INTERSPEECH | 3 |
| 2019 | The DKU-LENOVO Systems for the INTERSPEECH 2019 Computational Paralinguistic Challenge
Haiwei Wu, Weiqing Wang 0004, Ming Li 0026 |
INTERSPEECH | 3 |
| 2019 | An automated assessment framework for atypical prosody and stereotyped idiosyncratic phrases related to autism spectrum disorder
Ming Li 0026, Dengke Tang, Junlin Zeng, Tianyan Zhou, Huilin Zhu, Biyuan Chen, Xiaobing Zou |
Comput. Speech Lang. | 1 |
| 2018 | Insights in-to-End Learning Scheme for Language IdentificationabstractA novel interpretable end-to-end learning scheme for language identification is proposed. It is in line with the classical GMM i-vector methods both theoretically and practically. In the end-to-end pipeline, a general encoding layer is employed on top of the frontend CNN, so that it can encode the variable-length input sequence into an utterance level vector automatically. After comparing with the state-of-the-art GMM i-vector methods, we give insights into CNN, and reveal its role and effect in the whole pipeline. We further introduce a general encoding layer, illustrating the reason why they might be appropriate for language identification. We elaborate on several typical encoding layers, including a temporal average pooling layer, a recurrent encoding layer and a novel learnable dictionary encoding layer. We conducted experiment on NIST LRE07 closed-set task, and the results show that our proposed end-to-end systems achieve state-of-the-art performance. Weicheng Cai, Zexin Cai, Wenbo Liu 0002, Ming Li 0026 |
ICASSP | 5 |
| 2018 | A Novel Learnable Dictionary Encoding Layer for End-to-End Language IdentificationabstractA novel learnable dictionary encoding layer is proposed in this paper for end-to-end language identification. It is inline with the conventional GMM i-vector approach both theoretically and practically. We imitate the mechanism of traditional GMM training and Supervector encoding procedure on the top of CNN. The proposed layer can accumulate high-order statistics from variable-length input sequence and generate an utterance level fixed-dimensional vector representation. Unlike the conventional methods, our new approach provides an end-to-end learning framework, where the inherent dictionary are learned directly from the loss function. The dictionaries and the encoding representation for the classifier are learned jointly. The representation is orderless and therefore appropriate for language identification. We conducted a preliminary experiment on NIST LRE07 closed-set task, and the results reveal that our proposed dictionary encoding layer achieves significant error reduction comparing with the simple average pooling. Weicheng Cai, Zexin Cai, Xiang Zhang 0014, Ming Li 0026 |
ICASSP | 5 |
| 2018 | Analysis of Length Normalization in End-to-End Speaker Verification SystemabstractThe classical i-vectors and the latest end-to-end deep speaker embeddings are the two representative categories of utterancelevel representations in automatic speaker verification systems.Traditionally, once i-vectors or deep speaker embeddings are extracted, we rely on an extra length normalization step to normalize the representations into unit-length hyperspace before back-end modeling.In this paper, we explore how the neural network learns length-normalized deep speaker embeddings in an end-to-end manner.To this end, we add a length normalization layer followed by a scale layer before the output layer of the common classification network.We conducted experiments on the verification task of the Voxceleb1 dataset.The results show that integrating this simple step in the end-to-end training pipeline significantly boosts the performance of speaker verification.In the testing stage of our L2-normalized end-to-end system, a simple inner-product can achieve the state-of-the-art. Weicheng Cai, Jinkun Chen, Ming Li 0026 |
INTERSPEECH | 3 |
| 2018 | An End-to-End Deep Learning Framework for Speech Emotion Recognition of Atypical Individuals
Dengke Tang, Junlin Zeng, Ming Li 0026 |
INTERSPEECH | 3 |
| 2018 | Cancellable speech template via random binary orthogonal matrices projection hashing
Kong-Yik Chee, Zhe Jin 0001, Danwei Cai, Ming Li 0026, Wun-She Yap, Yen-Lung Lai, Bok-Min Goi |
Pattern Recognit. | 4 |
| 2017 | Automatic emotional spoken language text corpus construction from written dialogs in fictionsabstractIn this paper, we propose a novel method to automatically construct emotional spoken language text corpus from written dialogs, and release a large scale Chinese emotional text dataset with short conversations extracted from thousands of fictions using the proposed method. The emotional spoken language transcript resources in Chinese are relatively limited. However, constructing a large scale supervised corpus manually is neither efficient nor low-cost. This motivates us to try alternative efficient and effective approaches. First, we build a small scale emotion dictionary manually instead of a large scale corpus. Each word in dictionary has an emotion tag. Then, we use the emotional words to search emotional dialogs heuristically in fictions and classify them automatically. Second, we share our work to boost the performance of emotion recognition on spoken languages using the proposed new database. The labeled dialogs can be used for supervised learning while the unlabeled ones provide better word embeddings for the semantic level emotion recognition. We use the dialogs corpus as an auxiliary dataset in speech emotion recognition. We carry out experiments on automatic speech recognition (ASR) generated texts from the speech signals in Chinese Natural Emotional Audio-Visual Database (CHEAVD). It is an eight emotion states recognition task. We obtain a baseline average macro precision (MAP) of 37.08% and accuracy of 31.13% in terms of text-based method. With the labeled dialogs to pre-train neural networks and over-sampling the minority classes, we achieve an optimized MAP of 47.50% and the accuracy of 43.91%, which outperforms the baseline by 10.42% and 12.78% respectively. Jinkun Chen, Ming Li 0026 |
ACII | 3 |
| 2017 | Response to name: A dataset and a multimodal machine learning framework towards autism studyabstractIn this paper, we propose a “Response to Name Dataset” for autism spectrum disorder (ASD) study as well as a multimodal ASD auxiliary screening system based on machine learning. ASD children are characterized by their impaired interpersonal communication abilities and lack of response. In the proposed dataset, the reactions of children are recorded by cameras upon calling their names. The responsiveness of each child is then evaluated by a clinician with a score among 0, 1 and 2 following the Autism Diagnostic Observation Schedule (ADOS). We then develop a rule-based multimodal framework to quantitatively evaluate each child. Our system involves speech recognition based automatic name calling detection, face detection/alignment, head pose estimation, and considers the response speed, eye contact duration and head orientation to output the final prediction. Compared to existing work, our dataset characterizes a more precise and detailed scoring system with clinical trial standards, as well as a more spontaneous setting by incorporating less lab-controlled sessions with dynamic/cluttered environments, multi-pose mobile captured videos, and flexible number of accompanying adults. Experiments show that our machine predicted scores align closely with human professional diagnosis, showing promising potential in early screening of ASD, and shedding light on future clinical applications. Wenbo Liu 0002, Tianyan Zhou, Xiaobing Zou, Ming Li 0026 |
ACII | 5 |
| 2017 | SphereFace: Deep Hypersphere Embedding for Face RecognitionabstractThis paper addresses deep face recognition (FR) problem under open-set protocol, where ideal face features are expected to have smaller maximal intra-class distance than minimal inter-class distance under a suitably chosen metric space. However, few existing algorithms can effectively achieve this criterion. To this end, we propose the angular softmax (A-Softmax) loss that enables convolutional neural networks (CNNs) to learn angularly discriminative features. Geometrically, A-Softmax loss can be viewed as imposing discriminative constraints on a hypersphere manifold, which intrinsically matches the prior that faces also lie on a manifold. Moreover, the size of angular margin can be quantitatively adjusted by a parameter m. We further derive specific m to approximate the ideal feature criterion. Extensive analysis and experiments on Labeled Face in the Wild (LFW), Youtube Faces (YTF) and MegaFace Challenge 1 show the superiority of A-Softmax loss in FR tasks. Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li 0026, Bhiksha Raj |
CVPR | 4 |
| 2017 | Countermeasures for Automatic Speaker Verification Replay Spoofing Attack : On Data Augmentation, Feature Representation, Classification and Fusion
Weicheng Cai, Danwei Cai, Wenbo Liu 0002, Ming Li 0026 |
INTERSPEECH | 5 |
| 2017 | End-to-End Deep Learning Framework for Speech Paralinguistics Detection Based on Perception Aware Spectrum
Danwei Cai, Zhidong Ni, Wenbo Liu 0002, Weicheng Cai, Ming Li 0026 |
INTERSPEECH | 6 |
| 2017 | An Audio Based Piano Performance Evaluation Method Using Deep Neural Network Based Acoustic Modeling
Ming Li 0026, Zhanmei Song, Xin Li 0001, Hua Yi, Manman Zhu |
INTERSPEECH | 2 |
| 2016 | On Order-Constrained Transitive Distance ClusteringabstractWe consider the problem of approximating order-constrained transitive distance (OCTD) and its clustering applications. Given any pairwise data, transitive distance (TD) is defined as the smallest possible "gap" on the set of paths connecting them. While such metric definition renders significant capability of addressing elongated clusters, it is sometimes also an over-simplified representation which loses necessary regularization on cluster structure and overfits to short links easily. As a result, conventional TD often suffers from degraded performance given clusters with "thick" structures. Our key intuition is that the maximum (path) order, which is the maximum number of nodes on a path, controls the level of flexibility. Reducing this order benefits the clustering performance by finding a trade-off between flexibility and regularization on cluster structure. Unlike TD, finding OCTD becomes an intractable problem even though the number of connecting paths is reduced. We therefore propose a fast approximation framework, using random samplings to generate multiple diversified TD matrices and a pooling to output the final approximated OCTD matrix. Comprehensive experiments on toy, image and speech datasets show the excellent performance of OCTD, surpassing TD with significant gains and giving state-of-the-art performance on several datasets. Zhiding Yu, Weiyang Liu, Wenbo Liu 0002, Yingzhen Yang, Ming Li 0026, B. V. K. Vijaya Kumar |
AAAI | 5 |
| 2016 | Entity Disambiguation by Knowledge and Text Jointly EmbeddingabstractFor most entity disambiguation systems, the secret recipes are feature representations for mentions and entities, most of which are based on Bag-of-Words (BoW) representations.Commonly, BoW has several drawbacks: (1) It ignores the intrinsic meaning of words/entities; (2) It often results in high-dimension vector spaces and expensive computation; (3) For different applications, methods of designing handcrafted representations may be quite different, lacking of a general guideline.In this paper, we propose a different approach named EDKate.We first learn low-dimensional continuous vector representations for entities and words by jointly embedding knowledge base and text in the same vector space.Then we utilize these embeddings to design simple but effective features and build a two-layer disambiguation model.Extensive experiments on real-world data sets show that (1) The embedding-based features are very effective.Even a single one embedding-based feature can beat the combination of several BoW-based features.(2) The superiority is even more promising in a difficult set where the mention-entity prior cannot work well.(3) The proposed embedding method is much better than trivial implementations of some off-the-shelf embedding algorithms.(4) We compared our EDKate with existing methods/systems and the results are also positive. Dilin Wang, Zheng Chen 0001, Ming Li 0026 |
CoNLL | 5 |
| 2016 | Text-independent voice conversion using deep neural network based phonetic level featuresabstractThis paper presents a phonetically-aware joint density Gaussian mixture model (JD-GMM) framework for voice conversion that no longer requires parallel data from source speaker at the training stage. Considering that the phonetic level features contain text information which should be preserved in the conversion task, we propose a method that only concatenates phonetic discriminant features and spectral features extracted from the same target speakers speech to train a JD-GMM. After the mapping relationship of these two features is trained, we can use phonetic discriminant features from source speaker to estimate target speaker's spectral features at conversion stage. The phonetic discriminant features are extracted using PCA from the output layer of a deep neural network (DNN) in an automatic speaker recognition (ASR) system. It can be seen as a low dimensional representation of the senone posteriors. We compare the proposed phonetically-aware method with conventional JD-GMM method on the Voice Conversion Challenge 2016 training database. The experimental results show that our proposed phonetically-aware feature method can obtain similar performance compared to the conventional JD-GMM in the case of using only target speech as training data. Huadi Zheng, Weicheng Cai, Tianyan Zhou, Shilei Zhang, Ming Li 0026 |
ICPR | 5 |
| 2016 | Speaker verification based on the fusion of speech acoustics and inverted articulatory signals
Ming Li 0026, Jangwon Kim, Adam C. Lammert, Prasanta Kumar Ghosh, Vikram Ramanarayanan, Shri Narayanan |
Comput. Speech Lang. | 1 |
| 2015 | Efficient autism spectrum disorder prediction with eye movement: A machine learning frameworkabstractWe propose an autism spectrum disorder (ASD) prediction system based on machine learning techniques. Our work features the novel development and application of machine learning methods over traditional ASD evaluation protocols. Specifically, we are interested in discovering the latent patterns that possibly indicate the symptom of ASD underneath the observations of eye movement. A group of subjects (either ASD or non-ASD) are shown with a set of aligned human face images, with eye gaze locations on each image recorded sequentially. An image-level feature is then extracted from the recorded eye gaze locations on each face image. Such feature extraction process is expected to capture discriminative eye movement patterns related to ASD. In this work, we propose a variety of feature extraction methods, seeking to evaluate their prediction performance comprehensively. We further propose an ASD prediction framework in which the prediction model is learned on the labeled features. At testing stage, a test subject is also asked to view the face images with eye gaze locations recorded. The learned model predicts the image-level labels and a threshold is set to determine whether the test subject potentially has ASD or not. Despite the inherent difficulty of ASD prediction, experimental results indicates statistical significance of the predicted results, showing promising perspective of this framework. Wenbo Liu 0002, Zhiding Yu, Xiaobing Zou, Bhiksha Raj, Ming Li 0026 |
ACII | 6 |
| 2015 | Speaker verification with the mixture of Gaussian factor analysis based representationabstractThis paper presents a generalized i-vector representation framework using the mixture of Gaussian (MoG) factor analysis for speaker verification. Conventionally, a single standard factor analysis is adopted to generate a low rank total variability subspace where the mean supervector is assumed to be Gaussian distributed. The energy that can't be represented by the low rank space is modeled by a single multivariate Gaussian. However, due to the sparsity of the frame level posterior probability and the short duration characteristics, some dimensions of the first-order statistics may not be Gaussian distributed. Therefore, we replace the single Gaussian with a mixture of Gaussians to better represent the residual energy. Experimental results on the NIST SRE 2010 condition 5 female task and the RSR 2015 part 1 female task show that the MoG i-vector outperforms the i-vector baseline by more than 10% relatively for both text independent and text dependent speaker verification tasks, respectively. Ming Li 0026 |
ICASSP | 1 |
| 2015 | Duration dependent covariance regularization in PLDA modeling for speaker verification
Weicheng Cai, Ming Li 0026, Lin Li 0032, Qingyang Hong |
INTERSPEECH | 2 |
| 2015 | Modified-prior PLDA and score calibration for duration mismatch compensation in speaker recognition systemabstractTo deal with the performance degradation of speaker recognition due to duration mismatch between enrollment and test utterances, a novel strategy to modify the standard normal prior distribution of the i-vector during probabilistic linear discriminant analysis (PLDA) modeling is employed. This new modified-prior PLDA model incorporates the covariance matrix scaled with duration of each utterance for each speaker, which achieves more discriminative characteristics by learning the duration variability as well as session variation in the i-vector space. Furthermore, an efficient Quality Measure Function (QMF) method which adopts duration variation as a compensation technique is employed to eliminate the linear shift in the score domain. To evaluate the robustness of the proposed approach, experiments were conducted on the NIST SRE10 core-core task in condition-5 with varying test utterance duration, in which the i-vectors of test utterances were extracted from full segment and randomly truncated segments of duration 10s and 20s. The results demonstrated the efficiency of modified-prior PLDA in different duration conditions, and the combined score calibration further improved the performance of speaker recognition. Qingyang Hong, Lin Li 0032, Ming Li 0026, Lihong Wan |
INTERSPEECH | 3 |
| 2015 | Locality constrained transitive distance clustering on speech data
Wenbo Liu 0002, Zhiding Yu, Bhiksha Raj, Ming Li 0026 |
INTERSPEECH | 4 |
| 2015 | Speech bandwidth expansion based on deep neural networksabstractThis paper proposes a new speech bandwidth expansion method, which uses Deep Neural Networks (DNNs) to build high-order eigenspaces between the low frequency components and the high frequency components of the speech signal. A four-layer DNN is trained layer-by-layer from a cascade of Neural Networks (NNs) and two Gaussian-Bernoulli Restricted Boltzmann Machines (GBRBMs). The GBRBMs are adopted to model the distribution of spectral envelopes of the low frequency and the high frequency respectively. The NNs are used to model the joint distribution of hidden variables extracted from the two GBRBMs. The proposed method takes advantage of the strong modeling ability of GBRBMs in modeling the distribution of the spectral envelopes. And both the objective and subjective test results show that the proposed method outperforms the conventional GMM based method. Wenbo Liu 0002, Ming Li 0026, Jingming Kuang 0001 |
INTERSPEECH | 4 |
| 2015 | Automatic intelligibility classification of sentence-level pathological speech
Jangwon Kim, Naveen Kumar 0004, Andreas Tsiartas, Ming Li 0026, Shri Narayanan |
Comput. Speech Lang. | 4 |
| 2014 | Verification based ECG biometrics with cardiac irregular conditions using heartbeat level and segment level information fusionabstractWe propose an ECG based robust human verification system for both healthy and cardiac irregular conditions using the heartbeat level and segment level information fusion. At the heartbeat level, we first propose a novel beat normalization and outlier removal algorithm after peak detection to extract normalized representative beats. Then after principal component analysis (PCA), we apply linear discriminant analysis (LDA) and within-class covariance normalization (WCCN) for beat variability compensation followed by cosine similarity and Snorm as scoring. At the segment level, we adopt the hierarchical Dirichlet process auto-regressive hidden Markov model (HDP-AR-HMM) in the Bayesian non-parametric framework for unsupervised joint segmentation and clustering without any peak detection. It automatically decodes each raw signal into a string vector. We then apply n-gram language model and hypothesis testing for scoring. Combining the aforementioned two subsystems together further improved the performance and outperformed the PCA baseline by 25% relatively on the PTB database. Ming Li 0026, Xin Li 0001 |
ICASSP | 1 |
| 2014 | Simplified and supervised i-vector modeling for speaker age regressionabstractWe propose a simplified and supervised i-vector modeling scheme for the speaker age regression task. The supervised i-vector is obtained by concatenating the label vector and the linear regression matrix at the end of the mean super-vector and the i-vector factor loading matrix, respectively. Different label vector designs are proposed to increase the robustness of the supervised i-vector models. Finally, Support Vector Regression (SVR) is deployed to estimate the age of the speakers. The proposed method outperforms the conventional i-vector baseline for speaker age estimation. A relative 2.4% decrease in Mean Absolute Error and 3.33% increase in correlation coefficient is achieved using supervised i-vector modeling using different label designs on the NIST SRE 2008 dataset male part. Prashanth Gurunath Shivakumar, Ming Li 0026, Vedant Dhandhania, Shri Narayanan |
ICASSP | 2 |
| 2014 | Automatic recognition of speaker physical load using posterior probability based features from acoustic and phonetic tokensabstractThis paper presents an automatic speaker physical load recog-nition approach using posterior probability based features from acoustic and phonetic tokens. In this method, the tokens for calculating the posterior probability or zero-order statistics are extended from the conventional MFCC trained Gaussian Mix-ture Models (GMM) components to parallel phonetic phonemes and tandem feature trained GMM components. Phoneme rec-ognizers from five different languages are employed to extract the phoneme posterior probabilities. We show that these his-togram style features at both the acoustic and phonetic levels are effective and complementary for capturing the speaker phys-ical load information from short utterances. Support vector ma-chine is adopted as the supervised classifier. By combining the proposed methods with the OpenSMILE baseline which covers the acoustic and prosodic information further improves the fi-nal performance. The proposed fusion system achieves 70.18% and 72.81 % unweighted accuracy on the validation and test set of the Munich Bio-voice Corpus for the binary physical load level recognition task in the INTERSPEECH 2014 Computa-tional Paralinguistics Challenge. Index Terms: physical load sub-challenge, speaker physical load recognition, posterior probability features Ming Li 0026 |
INTERSPEECH | 1 |
| 2014 | Speaker verification and spoken language identification using a generalized i-vector framework with phonetic tokenizations and tandem featuresabstractThis paper presents a generalized i-vector framework with pho-netic tokenizations and tandem features for speaker verification as well as language identification. First, the tokens for cal-culating the zero-order statistics is extended from the MFCC trained Gaussian Mixture Models (GMM) components to pho-netic phonemes, 3-grams and tandem feature trained GMM components using phoneme posterior probabilities. Second, given the calculated zero-order statistics (posterior probabilities on tokens), the feature used to calculate the first-order statis-tics is also extended from MFCC to tandem features and is not necessarily the same feature employed by the tokenizer. Third, the zero-order and first-order statistics vectors are then concate-nated and represented by the simplified supervised i-vector ap-proach followed by the standard back end modeling methods. We study different system setups with different tokens and fea-tures. Finally, selected effective systems are fused at the score level to further improve the performance. Experimental results are reported on the NIST SRE 2010 common condition 5 fe-male part task and the NIST LRE 2007 closed set 30 seconds task for speaker verification and language identification, respec-tively. The proposed generalized i-vector framework outper-forms the i-vector baseline by relatively 45 % in terms of equal error rate (EER) and norm minDCF values. Index Terms: speaker verification, language identification, generalized i-vector, phonetic tokenization, tandem feature Ming Li 0026, Wenbo Liu 0002 |
INTERSPEECH | 1 |
| 2014 | Intoxicated speech detection: A fusion framework with speaker-normalized hierarchical functionals and GMM supervectors
Daniel Bone, Ming Li 0026, Matthew Black, Shri Narayanan |
Comput. Speech Lang. | 2 |
| 2014 | Simplified supervised i-vector modeling with application to robust and efficient language identification and speaker verification
Ming Li 0026, Shri Narayanan |
Comput. Speech Lang. | 1 |
| 2013 | Speaker verification using simplified and supervised i-vector modelingabstractThis paper presents a simplified and supervised i-vector modeling framework that is applied in the task of robust and efficient speaker verification (SRE). First, by concatenating the mean supervector and the i-vector factor loading matrix with respectively the label vector and the linear classifier matrix, the traditional i-vectors are then extended to label-regularized supervised i-vectors. These supervised i-vectors are optimized to not only reconstruct the mean supervectors well but also minimize the mean squared error between the original and the reconstructed label vectors, such that they become more discriminative. Second, factor analysis (FA) can be performed on the pre-normalized centered GMM first order statistics supervector to ensure that the Gaussian statistics sub-vector of each Gaussian component is treated equally in the FA, which reduces the computational cost significantly. Experimental results are reported on the female part of the NIST SRE 2010 task with common condition 5. The proposed supervised i-vector approach outperforms the i-vector baseline by relatively 12% and 7% in terms of equal error rate (EER) and norm old minDCF values, respectively. Ming Li 0026, Andreas Tsiartas, Maarten Van Segbroeck, Shri Narayanan |
ICASSP | 1 |
| 2013 | Classifying language-related developmental disorders from speech cues: the promise and the potential confoundsabstractSpeech and spoken language cues offer a valuable means to measure and model human behavior. Computational models of speech behavior have the potential to support health care through assistive technologies, informed intervention, and effi-cient long-term monitoring. The Interspeech 2013 Autism Sub-Challenge addresses two developmental disorders that manifest in speech: autism spectrum disorders and specific language im-pairment. We present classification results with an analysis on the development set including a discussion of potential con-founds in the data such as recording condition differences. We hence propose study of features within these domains that may inform realistic separability between groups as well as have the potential to be used for behavioral intervention and monitoring. We investigate template-based prosodic and formant modeling as well as goodness of pronunciation modeling, reporting above chance classification accuracies. Index Terms: autism spectrum disorders, intonation, specific language impairment, goodness of pronunciation Daniel Bone, Theodora Chaspari, Kartik Audhkhasi, James Gibson, Andreas Tsiartas, Maarten Van Segbroeck, Ming Li 0026, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 7 |
| 2013 | TRAP language identification system for RATS phase II evaluationabstractAutomatic language identification or detection of audio data has become an important preprocessing step for speech/speaker recognition and audio data mining. In many surveillance applications, language detection has to be performed on highly degraded audio inputs. In this paper, we present our work on language detection in highly degraded radio channel scenarios. We provide a brief description of the Targeted Robust Audio Processing (TRAP) language detection system builtfor the Phase II Evaluationof the RobustAutomatic Transcription of Speech (RATS) program. This system is a combination of 15 systems with different frontends and speech activity decisions. We also analyze the usefulness of multi-layer perceptron (MLP) based non-linear projection of i-vectors before SVM classification. The proposed backend reduces the Equal Error Rate (EER) by 11%–25% relative compared to the baseline PCA-based feature representation for SVM classification, on the RATS test data consisting of data from eight highfrequency radio communication channels. Index Terms: Language identification (detection), highly degraded radio channel, RATS, i-vector, multi-layer perceptron. Kyu Jeong Han, Sriram Ganapathy, Ming Li 0026, Mohamed Kamal Omar, Shri Narayanan |
INTERSPEECH | 3 |
| 2013 | Speaker verification based on fusion of acoustic and articulatory informationabstractWe propose a practical, feature-level fusion approach for com-bining acoustic and articulatory information in speaker ver-ification task. We find that concatenating articulation fea-tures obtained from the measured speech production data with conventional Mel-frequency cepstral coefficients (MFCCs) im-proves the overall speaker verification performance. However, since access to the measured articulatory data is impractical for real world speaker verification applications, we also ex-periment with estimated articulatory features obtained using acoustic-to-articulatory inversion technique. Specifically, we show that augmenting MFCCs with articulatory features ob-tained from subject-independent acoustic-to-articulatory inver-sion technique also significantly enhances the speaker verifi-cation performance. This performance boost could be due to the information about inter-speaker variation present in the es-timated articulatory features, especially at the mean and vari-ance level. Experimental results on the Wisconsin X-Ray Mi-crobeam database show that the proposed acoustic-estimated-articulatory fusion approach significantly outperforms the tra-ditional acoustic-only baseline, providing up to 10 % relative re-duction in Equal Error Rate (EER). We further show that we can achieve an additional 5 % relative reduction in EER after score-level fusion. Index Terms: speech production, speaker verification, articula-tion features, acoustic-to-articulatory inversion, biometrics Ming Li 0026, Jangwon Kim, Prasanta Kumar Ghosh, Vikram Ramanarayanan, Shri Narayanan |
INTERSPEECH | 1 |
| 2013 | Multi-band long-term signal variability features for robust voice activity detectionabstractIn this paper, we propose robust features for the problem of voice activity detection (VAD). In particular, we extend the long term signal variability (LTSV) feature to accommodate multiple spectral bands. The motivation of the multi-band approach stems from the non-uniform frequency scale of speech phonemes and noise characteristics. Our analysis shows that the multi-band approach offers advantages over the single band LTSV for voice activity detection. In terms of classification accuracy, we show 0.3%-61.2% relative improvement over the best accuracy of the baselines considered for 7 out 8 different noisy channels. Experimental results, and error analysis, are reported on the DARPA RATS corpora of noisy speech. Index Terms: noisy speech data, voice activity detection, robust feature extraction Andreas Tsiartas, Theodora Chaspari, Athanasios Katsamanis, Prasanta Kumar Ghosh, Ming Li 0026, Maarten Van Segbroeck, Alexandros Potamianos, Shri Narayanan |
INTERSPEECH | 5 |
| 2013 | Automatic speaker age and gender recognition using acoustic and prosodic level information fusion
Ming Li 0026, Kyu Jeong Han, Shri Narayanan |
Comput. Speech Lang. | 1 |
| 2012 | Speaker states recognition using latent factor analysis based Eigenchannel factor vector modelingabstractThis paper presents an automatic speaker state recognition approach which models the factor vectors in the latent factor analysis framework improving upon the Gaussian Mixture Model (GMM) baseline performance. We investigate both intoxicated and affective speaker states. We consider the affective speech signal as the original normal average speech signal being corrupted by the affective channel effects. Rather than reducing the channel variability to enhance the robustness as in the speaker verification task, we directly model the speaker state on the channel factors under the factor analysis framework. In this work, the speaker state factor vectors are extracted and modeled by the latent factor analysis approach in the GMM modeling framework and support vector machine classification method. Experimental results show that the proposed speaker state factor vector modeling system achieved 5.34% and 1.49% unweighted accuracy improvement over the GMM baseline on the intoxicated speech detection task (Alcohol Language Corpus) and the emotion recognition task (IEMOCAP database), respectively. Ming Li 0026, Angeliki Metallinou, Daniel Bone, Shri Narayanan |
ICASSP | 1 |
| 2012 | Speaker Personality Classification Using Systems Based on Acoustic-Lexical Cues and an Optimal Tree-Structured Bayesian NetworkabstractAutomatic classification of human personality along the Big Five dimensions is an interesting problem with several prac-tical applications. This paper makes some contributions in this regard. First, we propose a few automatically-derived personality-discriminating lexical features which provide infor-mation complementary to the conventional acoustic-prosodic cues. We also design a frame-level Gaussian mixture model based system which adds complimentary information to the sys-tems trained on global statistical functionals. Next, we note that the Big Five dimensions are correlated and thus model the de-pendency between these dimensions in the form of an optimal tree-structured Bayesian network. Our final sub-system con-sists of within class covariance normalization followed by L1-regularized logistic regression. Fusion of all these sub-systems achieves better classification performance than independently trained classifiers using just acoustic features. Kartik Audhkhasi, Angeliki Metallinou, Ming Li 0026, Shri Narayanan |
INTERSPEECH | 3 |
| 2012 | Intelligibility classification of pathological speech using fusion of multiple high level descriptors
Jangwon Kim, Naveen Kumar 0004, Andreas Tsiartas, Ming Li 0026, Shri Narayanan |
INTERSPEECH | 4 |
| 2012 | KNOWME: An Energy-Efficient Multimodal Body Area Network for Physical Activity MonitoringabstractThe use of biometric sensors for monitoring an individual’s health and related behaviors, continuously and in real time, promises to revolutionize healthcare in the near future. In an effort to better understand the complex interplay between one’s medical condition and social, environmental, and metabolic parameters, this article presents the KNOWME platform, a complete, end-to-end, body area sensing system that integrates off-the-shelf biometric sensors with a Nokia N95 mobile phone to continuously monitor the metabolic signals of a subject. With a current focus on pediatric obesity, KNOWME employs metabolic signals to monitor and evaluate physical activity. KNOWME development and in-lab deployment studies have revealed three major challenges: (1) the need for robustness to highly varying operating environments due to subject-induced variability, such as mobility or sensor placement; (2) balancing the tension between achieving high fidelity data collection and minimizing network energy consumption; and (3) accurate physical activity detection using a modest number of sensors. The KNOWME platform described herein directly addresses these three challenges. Design robustness is achieved by creating a three-tiered sensor data collection architecture. The system architecture is designed to provide robust, continuous, multichannel data collection and scales without compromising normal mobile device operation. Novel physical activity detection methods which exploit new representations of sensor signals provide accurate and efficient physical activity detection. The physical activity detection method employs personalized training phases and accounts for intersession variability. Finally, exploiting the features of the hardware implementation, a low-complexity sensor sampling algorithm is developed, resulting in significant energy savings without loss of performance. Gautam Thatte, Ming Li 0026, B. Adar Emken, Shri Narayanan, Urbashi Mitra, Donna Spruijt-Metz, Murali Annavaram |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2011 | Robust talking face video verification using joint factor analysis and sparse representation on GMM mean shifted supervectorsabstractIt has been previously demonstrated that systems based on block wise local features and Gaussian mixture models (GMM) are suitable for video based talking face verification due to the best trade-off in terms of complexity, robustness and performance. In this paper, we propose two methods to enhance the robustness and performance of the GMM-ZTnorm baseline system. First, joint factor analysis is performed to compensate the session variabilities due to different recording devices, lighting conditions, facial expressions, etc. Second, the difference between the universal background model (UBM) and the maximum a posteriori (MAP) adapted model is mapped into the GMM mean shifted supervector whose over-complete dictionary becomes more incoherent. Then, for verification purpose, the sparse representation computed by l1-minimization with quadratic constraints is employed to model these GMM mean shifted supervectors. Experimental results show that the proposed system achieved 8.4% (group 1) and 10.5% (group 2) equal error rate on the Banca talking face video database following the P protocol and outperformed the GMM-ZTnorm baseline by yielding more than 20% relative error reduction. Ming Li 0026, Shri Narayanan |
ICASSP | 1 |
| 2011 | Intoxicated Speech Detection by Fusion of Speaker Normalized Hierarchical Features and GMM SupervectorsabstractSpeaker state recognition is a challenging problem due to speaker and context variability. Intoxication detection is an important area of paralinguistic speech research with potential real-world applications. In this work, we build upon a base set of various static acoustic features by proposing the combination of several different methods for this learning task. The methods include extracting hierarchical acoustic features, performing iterative speaker normalization, and using a set of GMM supervectors. We obtain an optimal unweighted recall for intoxication recognition using score-level fusion of these subsystems. Unweighted average recall performance is 70.54 % on the test set, an improvement of 4.64 % absolute (7.04 % relative) over the baseline model accuracy of 65.9%. Index Terms: intoxication detection, speaker state, hierarchical features, speaker normalization, GMM supervectors 1. Daniel Bone, Matthew Black, Ming Li 0026, Angeliki Metallinou, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 3 |
| 2011 | Speaker Verification Using Sparse Representations on Total Variability i-vectorsabstractIn this paper, the sparse representation computed by l1-minimization with quadratic constraints is employed to model the i-vectors in the low dimensional total variability space af-ter performing the Within-Class Covariance Normalization and Linear Discriminate Analysis channel compensation. First, we propose the background normalized l2 residual as a scoring cri-terion. Second, we demonstrate that the Tnorm can be effi-ciently achieved by using the Tnorm data as the non-target sam-ples in the over-complete dictionary. Finally, by fusing with the conventional i-vector based support vector machine (SVM) and cosine distance scoring system, we demonstrate overall system performance improvement. Experimental results show that the proposed fusion system achieved 4.05 % (male) and 5.25 % (fe-male) equal error rate (EER) after Tnorm on the single-single multi-language handheld telephone task of NIST SRE 2008 and outperformed the SVM baseline by yielding 7.1 % and 4.9 % rel-ative EER reduction for the male and female tasks, respectively. Index Terms: speaker verification, sparse representation i-vector modeling Ming Li 0026, Xiang Zhang 0014, Yonghong Yan 0002, Shri Narayanan |
INTERSPEECH | 1 |
| 2010 | Robust ECG Biometrics by Fusing Temporal and Cepstral InformationabstractThe use of vital signs as a biometric is a potentially viable approach in a variety of application scenarios such as security and personalized health care. In this paper, a novel robust Electrocardiogram (ECG) biometric algorithm based on both temporal and cepstral information is proposed. First, in the time domain, after pre-processing and normalization, each heartbeat of the ECG signal is modeled by Hermite polynomial expansion (HPE) and support vector machine (SVM). Second, in the homomorphic domain, cepstral features are extracted from the ECG signals and modeled by Gaussian mixture modeling (GMM). In the GMM framework, heteroscedastic linear discriminant analysis and GMM super vector kernel is used to perform feature dimension reduction and discriminative modeling, respectively. Finally, fusion of both temporal and cepstral system outcomes at the score level is used to improve the overall performance. Experiment results show that the proposed hybrid approach achieves 98.3% accuracy and 0.5% equal error rate on the MIT-BIH Normal Sinus Rhythm Database. Ming Li 0026, Shri Narayanan |
ICPR | 1 |
| 2010 | Combining five acoustic level modeling methods for automatic speaker age and gender recognitionabstractAutomatic recognition of paralinguistic information from speech is important. Speaker identity, gender, age range, emotional state, etc. Guide human computer interaction systems to automatically adapt to different user needs. Ming Li 0026, Chi-Sang Jung, Kyu Jeong Han |
INTERSPEECH | 1 |
| 2009 | Optimal Allocation of Time-Resources for Multihypothesis Activity-Level Detection
Gautam Thatte, Viktor Rozgic, Ming Li 0026, Sabyasachi Ghosh, Urbashi Mitra, Shri Narayanan, Murali Annavaram, Donna Spruijt-Metz |
DCOSS | 3 |
| 2008 | An objective singing evaluation approach by relating acoustic measurements to perceptual ratings
Chuan Cao, Ming Li 0026, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 2008 | Cochannel speech separation using multi-pitch estimation and model based voiced sequential groupingabstractIn this paper, a new cochannel speech separation algorithm us-ing multi-pitch extraction and speaker model based sequential grouping is proposed. After auditory segmentation based on on-set and offset analysis, robust multi-pitch estimation algorithm is performed on each segment and the corresponding voiced portions are segregated. Then speaker pair model based on support vector machine (SVM) is employed to determine the optimal sequential grouping alignments and group the speaker homogeneous segments into pure speaker streams. Systematic evaluation on the speech separation challenge database shows significant improvement over the baseline performance. Index Terms: Auditory scene analysis, cochannel speech, multi-pitch estimation, sequential grouping Ming Li 0026, Chuan Cao, Ping Lu 0009, Qiang Fu 0001, Yonghong Yan 0002 |
INTERSPEECH | 1 |
| 2007 | Spoken language identification using score vector modeling and support vector machineabstractThe support vector machine (SVM) framework based on generalized linear discriminate sequence (GLDS) kernel has been shown effective and widely used in language identifica-tion tasks. In this paper, in order to compensate the distortions due to inter-speaker variability within the same language and solve the practical limitation of computer memory requested by large database training, multiple speaker group based discrim-inative classifiers are employed to map the cepstral features of speech utterances into discriminative language characterization score vectors (DLCSV). Furthermore, backend SVM classifiers are used to model the probability distribution of each target language in the DLCSV space and the output scores of back-end classifiers are calibrated as the final language recognition scores by a pair-wise posterior probability estimation algorithm. The proposed SVM framework is evaluated on 2003 NIST Lan-guage Recognition Evaluation databases, achieving an equal er-ror rate of 4.0 % in 30-second tasks, which outperformed the state-of-art SVM system by more than 30 % relative error re-duction. Index Terms: spoken language identification, support vector machine, score vector modeling Ming Li 0026, Hongbin Suo, Ping Lu 0009, Yonghong Yan 0002 |
INTERSPEECH | 1 |
| 2000 | Multi-group mixture weight HMM
Ming Li 0026, Tiecheng Yu |
INTERSPEECH | 1 |