EDBT 2026 Demo / reviewers in the wild / expert
Liang He 0003
dblp:42/963-3
· DBLP profile ↗
72ranked-venue papers
7as first author
50since 2021 · last 2026
0000-0003-4076-7479ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 7 first-author · 35 since 2021Artificial intelligence and machine learning · 40 · 1 first-author · 25 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DS3: Dual-space sample selection for speaker representation learning with noisy labels
Zhihua Fang, Liang He 0003 |
Pattern Recognit. | 2 |
| 2025 | Noise Supervised Contrastive Learning and Feature-Perturbed for Anomalous Sound DetectionabstractUnsupervised anomalous sound detection aims to detect unknown anomalous sounds by training a model using only normal audio data. Despite advancements in self-supervised methods, the issue of frequent false alarms when handling samples of the same type from different machines remains unresolved. This paper introduces a novel training technique called one-stage supervised contrastive learning (OS-SCL), which significantly addresses this problem by perturbing features in the embedding space and employing a one-stage noisy supervised contrastive learning approach. On the DCASE 2020 Challenge Task 2, it achieved 94.64% AUC, 88.42% pAUC, and 89.24% mAUC using only Log-Mel features. Additionally, a time-frequency feature named TFgram is proposed, which is extracted from raw audio. This feature effectively captures critical information for anomalous sound detection, ultimately achieving 95.71% AUC, 90.23% pAUC, and 91.23% mAUC. The source code is available at: www.github.com/huangswt/OS-SCL. Shun Huang, Zhihua Fang, Liang He 0003 |
ICASSP | 3 |
| 2025 | Stable Extended U-Net for Noise-Robust Speaker VerificationabstractWith advancements in deep learning, speaker verification systems have significantly improved their performance in noisy environments. Researchers typically demonstrate the effectiveness of their improved models by comparing performance on specific datasets, such as the VoxCeleb benchmark. However, in diverse real-world noise conditions, the out-of-domain generalization ability is also a crucial factor in evaluating a model’s performance improvement. Research on stable learning indicates that eliminating the spurious correlation between training and testing data can enhance the generalization of the model. Building on this idea, we propose an improved speaker verification system with high generalization based on the extended U-Net (ExU-Net). It uses the sample reweighting method from stable learning to eliminate sample correlations and retains more effective speaker information through subpixel convolutions and coordinate attention mechanisms. We validate the effectiveness of this approach through extensive evaluations on VoxCeleb1, VOiCES, and other out-of-domain noise test sets, highlighting its generalization capability and model robustness. Zonghui Wang, Zhihua Fang, Liang He 0003 |
ICASSP | 3 |
| 2025 | Self-supervised Speaker Verification with Batch-scale Pseudo-labels CorrectionabstractSelf-supervised learning has shown outstanding performance on speaker verification, and the 2-stage frameworks have more comprehensive training schemes, which typically exhibit better performance. They utilize clustering to obtain pseudo-labels, which are then used as the supervision signal in stage 2. However, these pseudo-labels often contain a significant amount of noisy labels, severely impacting speaker verification performance. In this paper, we propose a dynamic self-supervised pseudo-label correction method based on batch-scale training. By filtering and correcting samples based on the loss and prediction distribution, our method better aligns with the dynamic training process and achieves EER(%) of 1.33, 1.56 and 2.78 on the test sets of Voxceleb-O, E, H. Junxu Wang, Zhihua Fang, Liang He 0003 |
ICASSP | 3 |
| 2025 | Alignment Losses for End-to-End Speaker Diarization
Simeng Shi, Zhida Song, Zhihua Fang, Liang He 0003 |
ICIC (17) | 5 |
| 2025 | Enhanced Biomedical Relation Extraction with Semantic and Descriptive KnowledgeabstractBiomedical relation extraction is challenging due to its highly specialized nature and often requires extensive domain-specific knowledge. Existing methods often fail to effectively utilize the rich domain knowledge in external knowledge bases and the latent knowledge embedded in the data. They also suffer from inefficiency and fail to address the issue of class imbalance in relation distribution. In this paper, we propose a knowledge-enhanced biomedical relation extraction method (SDK-BRE), which aims to improve model performance by integrating both internal relational semantic knowledge and external entity descriptive knowledge, addressing the limitations of existing methods. Specifically, we transform relation types into natural language texts containing rich semantic information and use contrastive learning to predict relations, helping the model more accurately capture and distinguish different types of relations. Meanwhile, we extract entity descriptive knowledge from external knowledge bases and incorporate it as additional input to enhance the model’s contextual understanding. These types of knowledge are integrated into the model through embeddings, thus avoiding complex reasoning steps. Additionally, we introduce a rank-based loss function to handle the class imbalance issue, further improving the model’s ability to predict relation existence. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple representative biomedical relation extraction benchmark datasets. Further ablation studies and analyses confirm the effectiveness of each component in our method and verify that the integrated knowledge significantly enhances model performance. Liang He 0003 |
IJCNN | 2 |
| 2025 | TabMeta: A Novel Framework for Chart Generation Using Metadata and Enhanced Column AttentionabstractData analysis using tables and graphs is fundamental to research across disciplines. While pre-trained language models have shown potential in table understanding, their application to chart generation remains underexplored due to limited table-to-chart datasets and challenges in encoding two-dimensional tables into one-dimensional representations. To address these, we introduce SmallT2F, a curated dataset of tabular data from leading academic journals annotated with corresponding chart types. We also propose TabMeta, a hybrid metadata extraction framework that encodes table structures through row and column metadata and incorporates an enhanced column-aware attention mechanism to capture spatial relationships. By introducing a bias term in the attention-scoring function, TabMeta strengthens interactions between metric and data columns. Experimental results demonstrate that TabMeta significantly improves chart type selection, with a minimum improvement of 20% in Precision, 18.6% in Recall, and 19.4% in F1-Score, underscoring the importance of metadata and inter-column relationships in leveraging table structure. Yuxing Tan, Liang He 0003, Xiaozhe Qi, Zhida Song, Tengfei Feng |
IJCNN | 2 |
| 2025 | Data Valuation Method Based on Federated Learning and Shapley ValueabstractIn deep learning, the quality of data directly impacts model performance. Therefore, quantifying the value of data for model predictions or decisions is a critical issue, especially when enterprises or individuals need to identify the most beneficial datasets within limited budgets before actual training. This study proposes a valuation framework based on federated learning and Shapley value, addressing the problem of multiuser data valuation while ensuring privacy and data security. To rigorously assess the framework’s sensitivity to varying data quality characteristics, we systematically controlled the quality of datasets by manipulating key factors: label accuracy, noise intensity, and data volume. To validate the hypothesis that data with higher valuations contribute more significantly to model performance, we propose two federated learning algorithms, FedNFSV and FedNFSV+, which aggregate models based on their corresponding valuations. We compare these algorithms with existing federated learning baselines. Through a series of experiments, we concluded that: (1) The valuation results are accurately reflect the quality of the dataset. (2) In simulated scenarios, the valuation results can accurately identify users with poor-quality data. (3) The data from users with high valuation scores contributes more significantly to improving model performance. Jintai Wang, Liang He 0003 |
IJCNN | 3 |
| 2025 | AGDAformer: Agent-Guidance Dual Attention Transformer for Climate-Aware Crop Yield PredictionabstractAs climate change becomes widespread, rapid, and intensifying, accurate crop yield predictions are crucial for formulating effective agricultural development strategies. However, crop yield prediction is fairly challenging, as it is increasingly influenced by the intensifying climate change. In this paper, we propose the Agent-Guidance Dual Attention Transformer (AGDAformer), which aims to enhance the accuracy of yield predictions for climate-sensitive crops by integrating spatial attention mechanisms with multi-scale meteorological data. On the one hand, the Cross-Spatial Attention (CSA) mechanism optimizes the acquisition of broad-scale meteorological information by adaptively identifying and focusing on deep features closely related to crop yield prediction, thus improving the effectiveness of feature selection. On the other hand, the Agent-Enhanced Transformer (AET) introduces an agent matrix to facilitate the fusion of multiscale meteorological data, embedding meteorological knowledge into Agent Attention to effectively integrate information from different scales and better capture complex climatic factors. Extensive experiments on the CropNet dataset demonstrate that AGDAformer achieves superior performance across four crop types, significantly improving prediction accuracy and providing robust support for future agricultural decision-making. Zhongyan Yi, Zhihua Fang, Jintai Wang, Liang He 0003 |
IJCNN | 5 |
| 2025 | Deep Quantization Retrieval of Audio Fingerprints Based on Double Contrast LearningabstractAudio fingerprinting technology converts audio content into unique identifiers, enabling fast recognition and retrieval. Existing methods typically consist of two stages: first extracting audio features, and then using hashing or quantization techniques for retrieval. We propose a deep quantization retrieval framework for audio fingerprints based on dual contrastive learning. This framework employs a contrastive learning mechanism, initially calculating loss in the real-valued feature space, followed by a trainable quantization module that further computes the asymmetric similarity loss between quantized features and real-valued features. This dual-loss mechanism effectively preserves critical information in the real-valued features while efficiently compressing them into compact quantization codes, enabling end-to-end quantization retrieval. Compared to traditional two-stage methods, our framework not only simplifies the retrieval process but also improves retrieval accuracy. Additionally, to better capture the global structural information of audio fingerprint time-frequency features, we developed a spatial pyramid pooling module based on Graph Convolutional Network (GCN). This module leverages the information propagation mechanism between GCN nodes and integrates graph structure information with multi-scale information, effectively capturing time-frequency dependencies across different regions. Experimental results demonstrate that our method achieves end-to-end retrieval and enhances retrieval accuracy on music datasets. Weize Yue, Liang He 0003 |
IJCNN | 2 |
| 2025 | VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge
Zijing Zhao 0008, Hao Huang 0009, Ying Hu 0005, Liang He 0003 |
INTERSPEECH | 5 |
| 2025 | A Joint Network for Singing Melody Extraction from Polyphonic Music with Attention Aggregation and Self-Consistency Training
Jiabo Jing, Ying Hu 0005, Hao Huang 0009, Liang He 0003, Zhijian Ou |
INTERSPEECH | 4 |
| 2024 | C-LLM: Learn to Check Chinese Spelling Errors Character by CharacterabstractChinese Spell Checking (CSC) aims to detect and correct spelling errors in sentences.Despite Large Language Models (LLMs) exhibit robust capabilities and are widely applied in various tasks, their performance on CSC is often unsatisfactory.We find that LLMs fail to meet the Chinese character-level constraints of the CSC task, namely equal length and phonetic similarity, leading to a performance bottleneck.Further analysis reveals that this issue stems from the granularity of tokenization, as current mixed character-word tokenization struggles to satisfy these characterlevel constraints.To address this issue, we propose C-LLM, a Large Language Modelbased Chinese Spell Checking method that learns to check errors Character by Character.Character-level tokenization enables the model to learn character-level alignment, effectively mitigating issues related to character-level constraints.Furthermore, CSC is simplified to replication-dominated and substitutionsupplemented tasks.Experiments on two CSC benchmarks demonstrate that C-LLM achieves an average improvement of 10% over existing methods.Specifically, it shows a 2.1% improvement in general scenarios and a significant 12% improvement in vertical domain scenarios, establishing state-of-the-art performance.The source code can be accessed at https://github.com/ktlKTL/C-LLM. Kunting Li, Liang He 0003, Fandong Meng, Jie Zhou 0016 |
EMNLP | 3 |
| 2024 | Multi-View Speaker Embedding Learning for Enhanced Stability and DiscriminabilityabstractDeep neural network models based on x-vector have become the most popular framework for speaker recognition, and the quality of speaker features (embeddings) is important for open-set tasks such as speaker verification and speaker diarization. Currently, the most popular loss function is based on margin penalty, however, it only considers enlarging the inter-class distance while neglecting to reduce the intra-class feature differences. Therefore, we propose a multi-view learning approach that divides the training process into two views from the speaker embedding level. The classification view focuses on distinguishing the discriminability of different speakers, while the clustering view focuses on shrinking the feature boundaries of the same speaker, making intra-class differences smaller. The combined effect of the two perspectives achieves large inter-class distance and small intra-class distances, resulting in the extraction of more discriminative and stable speaker embeddings. We test the performance of the method on both speaker verification and speaker diarization tasks, and the results demonstrate the effectiveness of our approach. Liang He 0003, Zhihua Fang, Zuoer Chen, Minqiang Xu |
ICASSP | 1 |
| 2024 | A Study on Graph Embedding for Speaker RecognitionabstractCurrently, most speaker recognition systems make a decision by calculating the similarity between enrollment and test embeddings extracted with convolutional neural networks. However, for each embedding, the local structure between itself and its neighbors in the low-dimensional space is different, which is beneficial but is often ignored. We regard embeddings as nodes on a graph, compute edges among them by a distance function and nearest neighbor algorithm, and extract graph embeddings with graph neural networks (GNN) to further mine relational information for speaker recognition. Variants of GNN and different graph configurations are comparatively studied on NIST SRE14 i-vector challenging and VoxCeleb1 datasets. Experimental results demonstrate the excellent performance of the proposed graph embeddings . Liang He 0003, Ruida Li, Mengqi Niu |
ICASSP | 1 |
| 2024 | SMMA-Net: An Audio Clue-Based Target Speaker Extraction Network with Spectrogram Matching and Mutual AttentionabstractWe propose a deep neural network with spectrogram matching and mutual attention (SMMA-Net) for audio clue-based target speaker extraction (TSE). To effectively use the auxiliary speech, we proposed spectrogram matching (SM) strategy and mutual attention (MA) block. We conducted all experiments on the WSJ0-2mix-extr dataset. The ablation and comparison studies verified the effectiveness of SM strategy and MA block. The experimental results show that our proposed method outperforms the state-of-the-art methods by a sizable margin of 1.3 dB on the metric of scale-invariant signal-to-distortion ratio improvement. Additionally, SMMA-Net achieved that the performance of model for TSE task exceeds that for speaker separation task under the similar architecture. The main code will be available at https://github.com/Ht-Xu/SMMA-Net. Ying Hu 0005, Zhongcun Guo, Hao Huang 0009, Liang He 0003 |
ICASSP | 5 |
| 2024 | Introducing Multilingual Phonetic Information to Speaker Embedding for Speaker VerificationabstractIncorporating frame-level phonetic information during the extraction of speaker embeddings has been shown to enhance the performance of speaker verification systems. However, previous studies have primarily relied on phonetic information obtained from pre-trained models of monolingual automatic speech recognition (ASR). Considering that speaker verification datasets typically consist of multiple languages, there are instances where speakers are proficient in multiple languages, resulting in discrepancies between the languages used in the enrolled and test utterances. To address these challenges, we employ a pre-trained multilingual ASR Conformer encoder to initialize the MFA-Conformer network for speaker verification. Experimental results on the VoxCeleb dataset demonstrate a significant improvement in the performance of the system that incorporates multilingual phonetic information across different evaluation sets, including VoxCeleb1-O, E, and H, as well as the VoxSRC21 validation set, which focuses on multilingual verification. The source code is released at https://github.com/zds-potato/multilingual-phonetic-sv. Zhida Song, Liang He 0003, Ying Hu 0005, Hao Huang 0009 |
ICASSP | 2 |
| 2024 | Phase Continuity-Aware Self-Attentive Recurrent Network with Adaptive Feature Selection for Robust VADabstractDeep neural network (DNN) applications have significantly progressed in voice activity detection (VAD). Most current DNN-based VAD methods ignore the rich audio information in the phase domain. Therefore, applying this auxiliary information rationally and coping with low signal-to-noise ratio (SNR) background noise environments remains one of the challenges for VAD. To address this problem, we propose a VAD model robust to noise called phase continuity-aware self-attentive recurrent network (PC-ARN). For the input of PC-ARN, we draw inspiration from recent speech enhancement research by introducing phase-related features and further employing an adaptive feature selection module (AFSM) to combine magnitude features with it efficiently. The backbone network is an ARN module combining the attention mechanism and recurrent neural network (RNN), which can consider the relationship between local and global information to improve VAD performance competently. Experimental results show that our method has remarkable generalization ability and robustness compared to the traditional VAD techniques. Minjie Tang, Hao Huang 0009, Liang He 0003 |
ICASSP | 4 |
| 2024 | A Speaker Recognition Method Based on Stable LearningabstractWith the development of deep learning, speaker recognition systems have shown increasingly better performance. The generalization ability of the models is also an important aspect of performance evaluation. Typically, a baseline system is used to compare against the improved models to demonstrate performance enhancements. However, we cannot determine the differences in learned voiceprint features between the improved models and the baseline system. This paper introduces an improved speaker recognition system based on the ECAPA-TDNN model. It utilizes stable learning to eliminate sample correlation and employs attribution analysis to compare the differences in voiceprint feature learning between the improved and baseline systems. Experimental results demonstrate that stable learning improves the model’s generalization performance and helps it learn better voiceprint features. The effectiveness and generalization capability of the proposed method are verified through experiments on the VoxCeleb, CNCeleb, and LibriSpeech datasets. This work is important for enhancing speaker recognition performance, analyzing differences in voiceprint feature learning, and promoting advancements in the field. Lin Li 0032, Liang He 0003 |
ICASSP | 5 |
| 2024 | Speaker Recognition Based on Pre-Trained Model and Deep ClusteringabstractIn this paper, we propose a novel loss by integrating a deep clustering (DC) loss at the frame-level and a speaker recognition loss at the segment-level into a single network without additional data requirements and exhaustive computation. The DC loss implicitly generates soft pseudo-phoneme labels for each frame-level feature, which facilitates extracting more discriminant speaker representation by suppressing phonetic content information. We study the DC loss not only on the acoustic feature, but also on the features extracted by the pre-trained models, such as wav2vec 2.0, HuBERT and WavLM. Experimental results on the VoxCeleb dataset shows that the overall system performance based on the pre-trained model features are better than the one on the acoustic feature. The proposed loss is significantly effective for systems on the acoustic feature and has a marginal improvement for systems on the pre-trained model feature. Liang He 0003, Zhida Song, Shuanghong Liu, Mengqi Niu, Ying Hu 0005, Hao Huang 0009 |
ICME | 1 |
| 2024 | CSMA-CNER: Multi-modal Chinese NER task with Cross- and Self-Modality AttentionabstractMany scholars have employed dictionary-based word enhancement methods and multimodal information supplementation techniques to improve Chinese Named Entity Recognition models. However, these approaches primarily rely on static weights between different modalities, leading to a failure in capturing fine-grained correlations within the text modality and across modalities. As a result, they do not fully exploit multimodal information, leading to a loss of valuable data. To overcome these limitations, this paper proposes the Cross- and Self-Modality Attention network. This network dynamically captures correlations within the text modality and across modalities at multiple levels, effectively enhancing multimodal mutual information and reducing information loss. Additionally, we introduce two CNN structures to extract glyph visual and phonetic information. We conducted extensive experiments on Weibo, Resume, Ontonotes 4.0, and MSRA, and the results demonstrate that our approach outperforms state-of-the-art (SOTA) baseline methods. Bo Kong 0002, Shengquan Liu, Liang He 0003, Liruizhi Jia |
ICME | 3 |
| 2024 | Enhancing Abstractive Dialogue Summarization with Internal KnowledgeabstractThe task of dialogue summarization involves distilling a given dialogue into a concise and coherent summary. However, discrepancies in language styles between dialogues and summaries, scattered crucial information, incomplete utterances with ellipsis, and coreferences bring unique challenges to dialogue summarization. To tackle these challenges, we present multiple strategies in this study to extract crucial information with varying levels of granularity in dialogue from word-level and utterance-level semantics. This crucial information is used as valuable annotations on the dialogue text to train the model to recognize and utilize key information during the training phase, and enhance the model’s ability to identify crucial information during inference to generate better dialogue summaries. Experimental results on the SAMSum and DialogSum datasets shows that our method outperforms strong baseline models in terms of both ROUGE and BERTScore metrics. We corroborate these improvements through human evaluation. Shaolei Wang, Hanhan Ma, Liang He 0003 |
IJCNN | 5 |
| 2024 | Cross-modal Features Interaction-and-Aggregation Network with Self-consistency Training for Speech Emotion Recognition
Ying Hu 0005, Hao Huang 0009, Liang He 0003 |
INTERSPEECH | 4 |
| 2024 | YOLOPitch: A Time-Frequency Dual-Branch YOLO Model for Pitch Estimation
Hao Huang 0009, Ying Hu 0005, Liang He 0003, Yuyi Wang 0007 |
INTERSPEECH | 4 |
| 2024 | Speech Topic Classification Based on Multi-Scale and Graph Attention Networks
Fangjing Niu, Xiaozhe Qi, Xinya Chen, Liang He 0003 |
INTERSPEECH | 4 |
| 2024 | Self-Supervised Speaker Verification with Mini-Batch Prediction Correction
Junxu Wang, Zhihua Fang, Liang He 0003 |
INTERSPEECH | 3 |
| 2024 | Scene Text Recognition Via k-NN Attention-Based Decoder and Margin-Based Softmax Loss
Minqiang Xu, Liang He 0003 |
PRCV (7) | 3 |
| 2024 | Prompt for extraction: Multiple templates choice model for event extractionabstractEvent Extraction (EE) is an essential task in natural language processing that aims to mine events occurring in event mentions represent events using event records, which usually consist of event types, trigger words, argument elements corresponding to roles in the event types. Recently, prompt-based generative models have been developed to extract events. However, these prompt-based generative studies have ignored the fact that the strong language comprehension capability of the pre-trained language model (PLM) can analyze extract the potential role relationships in multiple templates for more information that can help extract argument elements. To determine the extended templates that can help the model for event extraction, we propose the multiple template choice model (MTCM), which designs an extended event type mining module to automatically mine the extended event types in the event mention uses the templates corresponding to the extended event types to interact with the template, corresponding to the currently to-be-extracted event type of event mention, in a multi-template information interaction, which gives the model more information guides the PLM for event extraction. To validate our model, we used two widely used datasets in the event extraction domain, ACE2005-EN ERE-EN. The experimental results show that our model achieves state-of-the-art performance on the ACE2005-EN dataset significantly improves the ERE dataset. In addition, according to the results, our model can be effectively adapted to low-resource environments. Jiaren Peng, Wenzhong Yang, Fuyuan Wei, Liang He 0003 |
Knowl. Based Syst. | 4 |
| 2024 | LMKG: A large-scale and multi-source medical knowledge graph for intelligent medicine applicationsabstractMedical Knowledge Graph (KG) has shown great potential in various healthcare scenarios, such as drug recommendation and clinical decision support system. The factors that determine the role of a medical KG in practical applications are the scale, coverage, and quality of the medical knowledge it can provide. Most existing medical KGs are extracted from a single or a few information sources. However, medical knowledge extracted from insufficient information sources is usually highly incomplete or even biased, which results in a lack of data completeness and may lessen their effectiveness in real-world scenarios. Besides, the coverage of entity and relation types is inadequate in most previous works, which also might restrict their potential usage in future applications. In this paper, we build a unified system that can extract and manage medical knowledge from heterogeneous information sources. We first employ named entity recognition and relation extraction methods to extract knowledge triplets from medical texts. Then we propose a hierarchical entity alignment framework for further knowledge refinement. Based on our system, we construct a large-scale, high-quality, multi-source, and multi-lingual medical KG named LMKG, which includes 13 entity types and 17 relation types, and contains 403,784 entity and 1,225,097 relation instances. We conduct extensive experiments to evaluate the quality of LMKG. Experimental results show that LMKG can effectively enhance the performance of both upstream and downstream intelligent medicine applications. We have publicly released the KG resources and corresponding management service interface to facilitate research and applications in the medical field. Peiru Yang, Yingzhuo Huang, Yuesong Zhang, Shizhong Yang, Liang He 0003, Yongfeng Huang 0001 |
Knowl. Based Syst. | 10 |
| 2024 | IIFC-Net: A Monaural Speech Enhancement Network With High-Order Information Interaction and Feature CalibrationabstractRecently, many Transformer-style dual-path models have achieved impressive performance for speech enhancement. However, their high parameters and computational complexity hinder their practical application. In this letter, we propose a monaural speech enhancement network with lower parameter count and complexity based on high-order information interaction and feature calibration (IIFC-Net). The network includes high-order information interaction Transformer (HOIIFormer) with high-order information interaction (HOII) block instead of a multi-head self-attention (MHSA) in Transformer. IIFC-Net leverages dual-path HOIIFormer (DPH) to model the distant dependency relation along time and frequency dimensions, respectively, and effectively captures deep-level information through the HOII block. We also design a feature calibration (FC) block to enhance the frequency components of target speech, which can be verified by a visualization analysis. The outcomes of experiments conducted on the VoiceBank+DEMAND and WHAMR! Datasets demonstrate that IIFC-Net achieves comparable performance in terms of denoising, dereverberation, and simultaneous denoising & dereverberation. Wenbing Wei, Ying Hu 0005, Hao Huang 0009, Liang He 0003 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Improving Speaker Verification With Noise-Aware Label Ensembling and Sample Selection: Learning and Correcting Noisy Speaker LabelsabstractSupervised deep learning has achieved tremendous success in speaker verification. However, deep speaker models tend to overfit noisy labels when they are present in the speaker datasets. To mitigate the detrimental effects of noisy labels, in this paper, we propose a novelLabel Ensembling and Sample Selectionframework. Firstly, we select labels with high confidence rankings as clean samples. Additionally, we use predictions from different epochs during training to smoothly correct the noisy labels. Our method does not require staged training and achieves integration of learning from noisy labels, selecting clean labels, and correcting noisy labels. A significant number of experimental results demonstrate the robustness of our method under noisy labels. Even when the training data contains 50% noisy labels, our method can mitigate an average of 86.54% of the performance degradation compared to the standard training method. Furthermore, further ablation experiments and analysis validate the effectiveness of High Confidence Ranking for sample selection and the correctness of Label Ensembling for noisy label correction. Zhihua Fang, Liang He 0003, Lin Li 0032, Ying Hu 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | SKD-NER: Continual Named Entity Recognition via Span-based Knowledge Distillation with Reinforcement LearningabstractContinual learning for named entity recognition (CL-NER) aims to enable models to continuously learn new entity types while retaining the ability to recognize previously learned ones.However, the current strategies fall short of effectively addressing the catastrophic forgetting of previously learned entity types.To tackle this issue, we propose the SKD-NER model, an efficient continual learning NER model based on the span-based approach, which innovatively incorporates reinforcement learning strategies to enhance the model's ability against catastrophic forgetting.Specifically, we leverage knowledge distillation (KD) to retain memory and employ reinforcement learning strategies during the KD process to optimize the soft labeling and distillation losses generated by the teacher model to effectively prevent catastrophic forgetting during continual learning.This approach effectively prevents or mitigates catastrophic forgetting during continuous learning, allowing the model to retain previously learned knowledge while acquiring new knowledge.Our experiments on two benchmark datasets demonstrate that our model significantly improves the performance of the CL-NER task, outperforming state-of-the-art methods. 1 Liang He 0003 |
EMNLP | 2 |
| 2023 | A Joint Network Based on Interactive Attention for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) has played a vital role in human-machine interaction. In this paper, we propose a separate spectrum-based SER model and a joint network combining pre-trained and spectrum-based models. In the joint network, we design an interactive attention module to effectively fuse the intermediate features from two models. Our proposed separate spectrum-based model is superior to four compared spectrum-based methods under the speaker-dependent setting. For the application in real scenarios, we compared our proposed joint network with six methods utilizing the pre-trained model under the speaker-independent setting. Experimental results show that our proposed joint network achieves the best performance among four unimodal models on the unweighted accuracy (UA) of 73.32 % and weighted accuracy (WA) of 72.48 %, respectively. Ying Hu 0005, Shijing Hou, Hao Huang 0009, Liang He 0003 |
ICME | 5 |
| 2023 | Speech Topic Classification Based on Pre-trained and Graph NetworksabstractSpeech Topic Classification (STC) automatically classifies audio clips into predefined categories, which is widely used in short video, personalized recommendation and other fields. At present, the common system is composed of two parts: first, the speech is converted into text by automatic speech recognition (ASR), and then the text topic is classified by natural language processing (NLP). Most of them have problems such as error propagation and lack of global structure. So in this paper, we propose a new end-to-end framework based on a pre-trained model and graph network. The pre-trained model is used to extract the semantic features with sequential structure instead of acoustic features, and the combination with the global features of conversational context constructed by graph network has achieved good results on the Fisher dataset. Fangjing Niu, Ying Hu 0005, Hao Huang 0009, Liang He 0003 |
ICME | 5 |
| 2023 | CRA-DIFFUSE: Improved Cross-Domain Speech Enhancement Based on Diffusion Model with T-F Domain Pre-DenoisingabstractSpeech enhancement (SE) methods in both the Time-Frequency (T-F) domain and time-domain domains have their own advantages. Leveraging both T-F domain and time-domain (cross domain) inputs has shown to be successful in the speech enhancement task. Recent SE methods based on diffusion models have shown promising results. However, little research effort has been made in the cross-domain speech enhancement using a diffusion model. We propose CRA-DiffuSE, a cross-domain SE model that uses a diffusion-based enhancement model as a refinement module after initial enhancement to achieve better results. For pre-enhance stage, we design CRANet, a T-F domain enhancement model combining channel attention and spatial attention. For the post-enhance stage, we design DiffuNet, a conditional generation model based on Denoising Diffusion Implicit Model (DDIM) for speech enhancement. Experiments demonstrate that the proposed CRA-DiffuSE is significantly superior to the baselines. Zhibin Qiu, Yachao Guo, Mengfan Fu, Hao Huang 0009, Ying Hu 0005, Liang He 0003, Fuchun Sun 0001 |
ICME | 6 |
| 2023 | Robust Training for Speaker Verification against Noisy Labels
Zhihua Fang, Liang He 0003, Hanhan Ma, Lin Li 0032 |
INTERSPEECH | 2 |
| 2023 | MTANet: Multi-band Time-frequency Attention Network for Singing Melody Extraction from Polyphonic Music
Ying Hu 0005, Liusong Wang, Hao Huang 0009, Liang He 0003 |
INTERSPEECH | 5 |
| 2023 | Dynamic Fully-Connected Layer for Large-Scale Speaker Verification
Zhida Song, Liang He 0003, Baowei Zhao, Minqiang Xu, Yu Zheng 0020 |
INTERSPEECH | 2 |
| 2023 | A Study on Visualization of Voiceprint Feature
Liang He 0003 |
INTERSPEECH | 2 |
| 2023 | Reprogramming Self-supervised Learning-based Speech Representations for Speaker AnonymizationabstractCurrent speaker anonymization methods, especially with self-supervised learning (SSL) models, require massive computational resources when hiding speaker identity. This paper proposes an effective and parameter-efficient speaker anonymization method based on recent End-to-End model reprogramming technology. To improve the anonymization performance, we first extract speaker representation from large SSL models as the speaker identifies. To hide the speaker’s identity, we reprogram the speaker representation by adapting the speaker to a pseudo domain. Extensive experiments are carried out on the VoicePrivacy Challenge (VPC) 2022 datasets to demonstrate the effectiveness of our proposed parameter-efficient learning anonymization methods. Additionally, while achieving comparable performance with the VPC 2022 strong baseline 1.b, our approach also consumes less computational resources during anonymization. Sheng Li 0010, Jiyi Li, Hao Huang 0009, Yang Cao 0011, Liang He 0003 |
MMAsia | 6 |
| 2023 | GhostVec: A New Threat to Speaker Privacy of End-to-End Speech Recognition SystemabstractSpeaker adaptation systems face privacy concerns, for such systems are trained on private datasets and often overfitting. This paper demonstrates that an attacker can extract speaker information by querying speaker-adapted speech recognition (ASR) systems. We focus on the speaker information of a transformer-based ASR and propose GhostVec, a simple and efficient attack method to extract the speaker information from an encoder-decoder-based ASR system without any external speaker verification system or natural human voice as a reference. To make our results quantitative, we pre-process GhostVec using singular value decomposition (SVD) and synthesize it into waveform. Experiment results show that the synthesized audio of GhostVec reaches 10.83% EER and 0.47 minDCF with target speakers, which suggests the effectiveness of the proposed method. We hope the preliminary discovery in this study to catalyze future speech recognition research on privacy-preserving topics. Sheng Li 0010, Jiyi Li, Yang Cao 0011, Hao Huang 0009, Liang He 0003 |
MMAsia | 6 |
| 2023 | MAKBQA: Multi-hop Knowledge Base Question Answering System Based on Sensors and Internet Agricultural DataabstractIn order to deeply integrate smart agriculture with the new generation of information technology, we have designed a system that can collect crop growth data and environmental data represented by soil moisture and temperature for digital modeling. At the same time, to make full use of artificial intelligence to extract effective information from big data to help agricultural planting and decision-making, we combine the data collected by sensors with the existing knowledge about agriculture on the internet to build an agricultural knowledge base with entities and relationships. A multi-hop question-answering model based on a knowledge base is designed to guide agricultural production. We are taking 100-hectare experimental farmland as an example, introducing the deployment method of agricultural sensors, and focusing on designing a multi-hop question-answering model for the knowledge base. Our question-answering model achieves SOTA on three publicly available datasets. Moreover, 100% accuracy was obtained on the 3-hop test set of MetaQA. Liang He 0003, Shaolei Wang, Hanhan Ma, Kan Feng |
SECON | 2 |
| 2022 | A Graph Isomorphism Network with Weighted Multiple Aggregators for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) is an essential part of human-computer interaction. In this paper, we propose an SER network based on a Graph Isomorphism Network with Weighted Multiple Aggregators (WMA-GIN), which can effectively handle the problem of information confusion when neighbour nodes' features are aggregated together in GIN structure. Moreover, a Full-Adjacent (FA) layer is adopted for alleviating the over-squashing problem, which is existed in all Graph Neural Network (GNN) structures, including GIN. Furthermore, a multi-phase attention mechanism and multi-loss training strategy are employed to avoid missing the useful emotional information in the stacked WMA-GIN layers. We evaluated the performance of our proposed WMA-GIN on the popular IEMOCAP dataset. The experimental results show that WMA-GIN outperforms other GNN-based methods and is comparable to some advanced non-graph-based methods by achieving 72.48% of weighted accuracy (WA) and 67.72% of unweighted accuracy (UA). Ying Hu 0005, Yuwu Tang, Hao Huang 0009, Liang He 0003 |
INTERSPEECH | 4 |
| 2022 | A Multi-grained based Attention Network for Semi-supervised Sound Event DetectionabstractSound event detection (SED) is an interesting but challenging task due to the scarcity of data and diverse sound events in real life. This paper presents a multi-grained based attention network (MGA-Net) for semi-supervised sound event detection. To obtain the feature representations related to sound events, a residual hybrid convolution (RH-Conv) block is designed to boost the vanilla convolution's ability to extract the time-frequency features. Moreover, a multi-grained attention (MGA) module is designed to learn temporal resolution features from coarse-level to fine-level. With the MGA module,the network could capture the characteristics of target events with short- or long-duration, resulting in more accurately determining the onset and offset of sound events. Furthermore, to effectively boost the performance of the Mean Teacher (MT) method, a spatial shift (SS) module as a data perturbation mechanism is introduced to increase the diversity of data. Experimental results show that the MGA-Net outperforms the published state-of-the-art competitors, achieving 53.27% and 56.96% event-based macro F1 (EB-F1) score, 0.709 and 0.739 polyphonic sound detection score (PSDS) on the validation and public set respectively. Ying Hu 0005, Xiujuan Zhu, Hao Huang 0009, Liang He 0003 |
INTERSPEECH | 5 |
| 2022 | How to Boost Anti-Spoofing with X-VectorsabstractWith the development of speech synthesis or voice conversion, speech spoofing countermeasures are increasingly required for protecting automatic speaker verification system. In our daily life, if we are familiar with the speaker, we tend to seek her/his traits in our memory to distinguish between bona fide and spoofed speech of her/him. Speaker label can not be directly used to guide the training of anti-spoofing network because it is difficult to obtain in real scenes. Motivated by this, we use x-vectors to represent speaker information and propose two novel methods by introducing x-vectors on the acoustic feature and embedding level into the two mainstream anti-spoofing methods (LightCNN and SeNet). An attention module is also added on the embedding level for further improvement. Experimental results on the ASVspoof 2019 logical access (LA) database show that the best EER and mintDCF in our methods are 0.98% and 0.0294, outperforming state-of-the-art single systems as far as we know. Shen Huang, Ji Gao, Ying Hu 0005, Liang He 0003 |
SLT | 6 |
| 2022 | Multi-stage music separation network with dual-branch attention and hybrid convolution
Yadong Chen 0003, Ying Hu 0005, Liang He 0003, Hao Huang 0009 |
J. Intell. Inf. Syst. | 3 |
| 2022 | A bimodal network based on Audio-Text-Interactional-Attention with ArcFace loss for speech emotion recognition
Yuwu Tang, Ying Hu 0005, Liang He 0003, Hao Huang 0009 |
Speech Commun. | 3 |
| 2022 | Hierarchic Temporal Convolutional Network With Cross-Domain Encoder for Music Source SeparationabstractRecently, the time-domain-based methods (i.e., the method of modeling the raw waveform directly) for audio source separation have shown tremendous potential. In this paper, we propose a model which combines the complexed spectrogram domain feature and time-domain feature by a cross-domain encoder (CDE) and adopts the hierarchic temporal convolutional network (HTCN) for multiple music sources separation. The CDE is designed to enable the network to code the interactive information of the time-domain and complexed spectrogram domain features. HTCN enables it to learn the long-time series dependence effectively. We also designed a feature calibration unit (FCU) to be applied in the HTCN and adopted the multi-stage training strategy during the training stage. The ablation study demonstrates the effectiveness of each designed component in the model. We conducted the experiments on the MUSDB18 dataset. The experimental results indicate that our proposed CDE-HTCN model outperforms the top-of-the-line methods and, compared with the state-of-the-art method, DEMUCS, achieves the improvement of the average SDR score of 0.61 dB. Significantly, the improvement of the SDR score for the$\ bass$source has a sizable margin of 0.91 dB. Ying Hu 0005, Yadong Chen 0003, Wenzhong Yang, Liang He 0003, Hao Huang 0009 |
IEEE Signal Process. Lett. | 4 |
| 2021 | Improved Lightcnn with Attention Modules for Asv Spoofing DetectionabstractWith the advent of many state-of-the-art speech synthesis or conversion techniques, speaker recognition system is confronted with the imperceptible interference caused by spoofed speech. In this paper, we realize that filter bank distributions of cepstral features in the frequency domain cause considerable influence to the system performance through experiments. Therefore, given attention mechanism may help to focus on key information, this paper presents improved light convolutional neural network (LCNN) with attention modules, separately named Squeeze-and-Excitation (SE) block, Convolutional Block Attention Module (CBAM) and Dual Attention Network (DANet). To our knowledge, we are the first to do systematic study of attention mechanisms in the field of speech anti-spoofing. Experimented on ASVspoof 2019 dataset, our proposed single system can get 22% min-tDCF reduction as well as 42% EER reduction over LCNN baseline and its performance can even rank fourth among all participating teams using fusion systems in ASVspoof 2019 competition. Tianyu Liang, Shen Huang, Liang He 0003 |
ICME | 5 |
| 2021 | End-to-End Cross-Lingual Spoken Language Understanding Model with Multilingual Pretraining
Liang He 0003 |
Interspeech | 2 |
| 2020 | THUEE System for NIST SRE19 CTS Challenge
Ruyun Li, Tianyu Liang, Yi Liu 0049, Yangcheng Wu, Can Xu 0003, Xianhong Chen, Weiqiang Zhang 0001, Shouyi Yin, Liang He 0003 |
INTERSPEECH | 12 |
| 2020 | Adaptive Multi-Scale Detection of Acoustic EventsabstractThe goal of acoustic (or sound) events detection (AED or SED) is to predict the temporal position of target events in given audio segments. This task plays a significant role in safety monitoring, acoustic early warning and other scenarios. However, the deficiency of data and diversity of acoustic event sources make the AED task a tough issue, especially for prevalent data-driven methods. In this article, we start from analyzing acoustic events according to their time-frequency domain properties, showing that different acoustic events have different time-frequency scale characteristics. Inspired by the analysis, we propose an adaptive multi-scale detection (AdaMD) method. By taking advantage of hourglass neural network and gated recurrent unit (GRU) module, our AdaMD produces multiple predictions at different temporal and frequency resolutions. An adaptive training algorithm is subsequently adopted to combine multi-scale predictions to enhance the overall capability. Experimental results on Detection and Classification of Acoustic Scenes and Events 2017 (DCASE 2017) Task 2, DCASE 2016 Task 3 and DCASE 2017 Task 3 demonstrate that the AdaMD outperforms published state-of-the-art competitors in terms of the metrics of event error rate (ER) and F1-score. The verification experiment on our collected factory mechanical dataset also proves the noise-resistant capability of the AdaMD, providing the possibility for it to be deployed in the complex environment. Wenhao Ding, Liang He 0003 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Multi-objective Optimization Training of PLDA for Speaker VerificationabstractMost current state-of-the-art text-independent speaker verifi-cation systems take probabilistic linear discriminant analysis (PLDA) as their backend classifiers. The parameters of PL-DA are often estimated by maximizing the objective function, which focuses on increasing the value of log-likelihood function, but ignoring the distinction between speakers. In order to better distinguish speakers, we propose a multi-objective optimization training for PLDA. Experiment results show that the proposed method has more than 10% relative performance improvement in both EER and MinDCF on the NIST SRE14 i-vector challenge dataset, and about 20% relative performance improvement in EER on the MCE18 dataset. Liang He 0003, Xianhong Chen, Can Xu 0003, Jia Liu 0001 |
ICASSP | 1 |
| 2019 | Towards Discriminative Representations and Unbiased Predictions: Class-Specific Angular Softmax for Speech Emotion Recognition
Liang He 0003 |
INTERSPEECH | 2 |
| 2019 | Large Margin Softmax Loss for Speaker VerificationabstractIn neural network based speaker verification, speaker embedding is expected to be discriminative between speakers while the intra-speaker distance should remain small.A variety of loss functions have been proposed to achieve this goal.In this paper, we investigate the large margin softmax loss with different configurations in speaker verification.Ring loss and minimum hyperspherical energy criterion are introduced to further improve the performance.Results on VoxCeleb show that our best system outperforms the baseline approach by 15% in EER, and by 13%, 33% in minDCF08 and minDCF10, respectively. Yi Liu 0049, Liang He 0003, Jia Liu 0001 |
INTERSPEECH | 2 |
| 2019 | Multi-Scale Time-Frequency Attention for Acoustic Event DetectionabstractMost attention-based methods only concentrate along the time axis, which is insufficient for Acoustic Event Detection (AED).Meanwhile, previous methods for AED rarely considered that target events possess distinct temporal and frequential scales.In this work, we propose a Multi-Scale Time-Frequency Attention (MTFA) module for AED.MTFA gathers information at multiple resolutions to generate a time-frequency attention mask which tells the model where to focus along both time and frequency axis.With MTFA, the model could capture the characteristics of target events with different scales.We demonstrate the proposed method on Task 2 of Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 Challenge.Our method achieves competitive results on both development dataset and evaluation dataset. Jingyang Zhang, Wenhao Ding, Jintao Kang, Liang He 0003 |
INTERSPEECH | 4 |
| 2019 | Distance-Dependent Metric LearningabstractIn most existing metric learning methods, the data pairs are equally treated without considering their diversity. In fact, for pairs with different distances, the main purpose of the metric acted on them has some differences. If the pairs have smaller distance, metric should focus more on expanding the negative pairs, which are more easily misjudged. While if the pairs have larger distance, metric should focus more on shrinking the positive pairs. However, most metric learning methods neglect these differences. In this letter, we propose a distance-dependent metric learning (D2ML) method. It partitions data pairs into different clusters according to the ℓ2distance between them. Each cluster is associated with a Mahalanobis metric that learns the pairs' distance. This not only allows us to make each metric more targeted and adapt to the data diversity flexibly, but also avoids the problem of computing the distance between points assigned to different clusters, which happens in some local metric learning methods. D2ML is further extended to D3ML to embrace the nonlinear capacity of neural network. Experiments on UCI datasets and speaker recognition i-vector machine learning challenge show that the proposed methods are superior to other metric learning methods. Xianhong Chen, Liang He 0003, Can Xu 0003, Jia Liu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2018 | MTGAN: Speaker Verification through Multitasking Triplet Generative Adversarial NetworksabstractIn this paper, we propose an enhanced triplet method that improves the encoding process of embeddings by jointly utilizing generative adversarial mechanism and multitasking optimization. We extend our triplet encoder with Generative Adversarial Networks (GANs) and softmax loss function. GAN is introduced for increasing the generality and diversity of samples, while softmax is for reinforcing features about speakers. For simplification, we term our method Multitasking Triplet Generative Adversarial Networks (MTGAN). Experiment on short utterances demonstrates that MTGAN reduces the verification equal error rate (EER) by 67% (relatively) and 32% (relatively) over conventional i-vector method and state-of-the-art triplet loss method respectively. This effectively indicates that MTGAN outperforms triplet methods in the aspect of expressing the high-level feature of speaker information. Wenhao Ding, Liang He 0003 |
INTERSPEECH | 2 |
| 2018 | Speaker Embedding Extraction with Phonetic InformationabstractSpeaker embeddings achieve promising results on many speaker verification tasks.Phonetic information, as an important component of speech, is rarely considered in the extraction of speaker embeddings.In this paper, we introduce phonetic information to the speaker embedding extraction based on the x-vector architecture.Two methods using phonetic vectors and multi-task learning are proposed.On the Fisher dataset, our best system outperforms the original x-vector approach by 20% in EER, and by 15%, 15% in minDCF08 and minDCF10, respectively.Experiments conducted on NIST SRE10 further demonstrate the effectiveness of the proposed methods. Yi Liu 0049, Liang He 0003, Jia Liu 0001, Michael T. Johnson |
INTERSPEECH | 2 |
| 2018 | Defect characterization of amorphous silicon thin film solar cell based on low frequency noise
Linna Hu, Liang He 0003, Xiaofei Jia, Ying Hu 0005, Hongmei Ma, Dandan Guo |
Sci. China Inf. Sci. | 2 |
| 2018 | Semi-supervised minimum redundancy maximum relevance feature selection for audio classification
Xukui Yang 0001, Liang He 0003, Dan Qu 0003, Weiqiang Zhang 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Local Pairwise Linear Discriminant Analysis for Speaker VerificationabstractLinear discriminant analysis-probabilistic linear discriminant analysis (LDA-PLDA) is a standard and effective backend in the field of speaker verification. The object of LDA is to perform dimensionality reduction while minimizing within-class covariance and maximizing between-class covariance. For a target class (or speaker), our task is to make a binary decision about whether a test utterance is from a specific target speaker. Generally, the nontarget test utterances that are close to the target speaker are easily misjudged. Inspired by this idea, we propose a local pairwise linear discriminant analysis (LPLDA) algorithm. This new method focuses on maximizing the local pairwise covariance, which represents the local structure between the target class samples and neighboring nontarget class samples, instead of the between-class covariance, which represents the global structure of the data. Experiments on the NIST SRE 2010, 2014, and 2016 database show that, the proposed LPLDA-PLDA backend has significant performance improvements over the LDA-PLDA backend. Liang He 0003, Xianhong Chen, Can Xu 0003, Jia Liu 0001, Michael T. Johnson |
IEEE Signal Process. Lett. | 1 |
| 2017 | Comparison of multiple features and modeling methods for text-dependent speaker verificationabstractText-dependent speaker verification is becoming popular in the speaker recognition society. However, the conventional i-vector framework which has been successful for speaker identification and other similar tasks works relatively poorly in this task. Researchers have proposed several new methods to improve performance, but it is still unclear that which model is the best choice, especially when the pass-phrases are prompted during enrollment and test. In this paper, we introduce four modeling methods and compare their performance on the newly published RedDots dataset. To further explore the influence of different frame alignments, Viterbi and forward-backward algorithms are both used in the HMM-based models. Several bottleneck features are also investigated. Our experiments show that, by explicitly modeling the lexical content, the HMM-based modeling achieves good results in the fixed-phrase condition. In the prompted-phrase condition, GMM-HMM and i-vector/HMM are not as successful. In both conditions, the forward-backward algorithm brings more benefits to the i-vector/HMM system. Additionally, we also find that even though bottleneck features perform well for text-independent speaker verification, they do not outperform MFCCs on the most challenging Imposter-Correct trials on RedDots. Yi Liu 0049, Liang He 0003, Zhuzi Chen, Jia Liu 0001, Michael T. Johnson |
ASRU | 2 |
| 2017 | Deep neural networks based speaker modeling at different levels of phonetic granularityabstractRecently, a hybrid deep neural network/i-vector framework has been proved effective for speaker verification, where the DNN trained to predict tied-triphone states (senones) is used to produce frame alignments for sufficient statistics extraction. In this work, in order to better understand the impact of different phonetic precision to speaker verification tasks, three levels of phonetic granularity are evaluated when doing frame alignments, which are tied-triphone state, monophone state and monophone. And the distribution of the features associated to a given phonetic unit is further modeled with multiple Gaussians rather than a single Gaussian. We also propose a fast and efficient way to generate phonetic units of different granularity by tying DNN's outputs according to the clustering results based on DNN derived senone embeddings. Experiments are carried out on the NIST SRE 2008 female tasks. Results show that using DNNs with less precise phonetic units and more Gaussians per phonetic unit for speaker modeling generalize better to different speaker verification tasks. Liang He 0003, Weiqiang Zhang 0001, Jia Liu 0001 |
ICASSP | 2 |
| 2016 | THU-EE System Description for NIST LRE 2015
Liang He 0003, Yi Liu 0049, Weiwei Liu 0001, Cai Meng, Jia Liu 0001 |
INTERSPEECH | 1 |
| 2016 | Investigating Various Diarization Algorithms for Speaker in the Wild (SITW) Speaker Recognition Challenge
Yi Liu 0049, Liang He 0003, Jia Liu 0001 |
INTERSPEECH | 3 |
| 2016 | Improving Deep Neural Networks Based Speaker Verification Using Unlabeled Data
Liang He 0003, Weiqiang Zhang 0001, Jia Liu 0001 |
INTERSPEECH | 3 |
| 2015 | Simultaneous utilization of spectral magnitude and phase information to extract supervectors for speaker verification anti-spoofingabstractProtection from spoofing attacks is an essential component of speaker verification systems. This paper proposes a novel approach to detect such attacks by utilizing supervectors derived from spectral magnitude and phase information. Three countermeasures are chosen to represent these important information. To combine different countermeasures, score fusion and an antispoofing supervector (ASSV) are used. Experiments conducted on ASVspoof 2015 show that the combination of magnitude and phase information obtains relative 90% improvement in terms of the equal error rate (EER) compared to the best subsystem in the development set. The two systems can also be fused to further improve the performance. In addition to accuracy improvements, the new supervector framework is extensible and allows for a more flexible interface to the back-end classifier design. Yi Liu 0049, Liang He 0003, Jia Liu 0001, Michael T. Johnson |
INTERSPEECH | 3 |
| 2015 | Investigation of bottleneck features and multilingual deep neural networks for speaker verification
Liang He 0003, Jia Liu 0001 |
INTERSPEECH | 3 |
| 2013 | I-matrix for text-independent speaker recognitionabstractThis paper proposes an i-matrix for text-independent speaker recognition. The framework of the proposed i-matrix is similar to an i-vector. However, the presented method takes short-time cepstral feature matrices as inputs to explore both cepstral feature distribution and temporal information for the recognition task in the phase of statistical modeling. In the i-matrix, the variability of an utterance is constrained by two subspaces U and V, which are estimated by an iterative method on a large database. When U and V are well built, each utterance is represented by an i-matrix. Decision function is a cosine kernel. Experiments were carried out on the tel-tel-English condition of NIST SRE 2008 core task. Compared with an i-vector-LDA, the average EER and MDCF of an i-matrix-LDA showed a relative decrease of 4.82% and 5.12% respectively. Liang He 0003, Jia Liu 0001 |
ICASSP | 1 |
| 2011 | Time-Frequency Cepstral Features and Heteroscedastic Linear Discriminant Analysis for Language RecognitionabstractThe shifted delta cepstrum (SDC) is a widely used feature extraction for language recognition (LRE). With a high context width due to incorporation of multiple frames, SDC outperforms traditional delta and acceleration feature vectors. However, it also introduces correlation into the concatenated feature vector, which increases redundancy and may degrade the performance of backend classifiers. In this paper, we first propose a time-frequency cepstral (TFC) feature vector, which is obtained by performing a temporal discrete cosine transform (DCT) on the cepstrum matrix and selecting the transformed elements in a zigzag scan order. Beyond this, we increase discriminability through a heteroscedastic linear discriminant analysis (HLDA) on the full cepstrum matrix. By utilizing block diagonal matrix constraints, the large HLDA problem is then reduced to several smaller HLDA problems, creating a block diagonal HLDA (BDHLDA) algorithm which has much lower computational complexity. The BDHLDA method is finally extended to the GMM domain, using the simpler TFC features during re-estimation to provide significantly improved computation speed. Experiments on NIST 2003 and 2007 LRE evaluation corpora show that TFC is more effective than SDC, and that the GMM-based BDHLDA results in lower equal error rate (EER) and minimum average cost (Cavg) than either TFC or SDC approaches. Weiqiang Zhang 0001, Liang He 0003, Jia Liu 0001, Michael T. Johnson |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Variant time-frequency cepstral features for speaker recognition
Weiqiang Zhang 0001, Liang He 0003, Jia Liu 0001 |
INTERSPEECH | 3 |