Ta Li

dblp:44/742 · DBLP profile ↗
← Back
28ranked-venue papers
3as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LAES: A local adaptive edge-enhanced spectrogram method for unsupervised anomalous sound detection
Jiyu Lu, Wenbo Guan, Ming Zhang 0035, Ta Li, Yonghong Yan 0002
Signal Process.4
2025 Automatic Text Pronunciation Correlation Generation and Application for Contextual Biasing
abstract
Effectively distinguishing the pronunciation correlations between different written texts is a significant issue in linguistic acoustics. Traditionally, such pronunciation correlations are obtained through manually designed pronunciation lexicons. In this paper, we propose a data-driven method to automatically acquire these pronunciation correlations, called automatic text pronunciation correlation (ATPC). The supervision required for this method is consistent with the supervision needed for training end-to-end automatic speech recognition (E2E-ASR) systems, i.e., speech and corresponding text annotations. First, the iteratively-trained timestamp estimator (ITSE) algorithm is employed to align the speech with their corresponding annotated text symbols. Then, a speech encoder is used to convert the speech into speech embeddings. Finally, we compare the speech embeddings distances of different text symbols to obtain ATPC. Experimental results on Mandarin show that ATPC enhances E2E-ASR performance in contextual biasing and holds promise for dialects or languages lacking artificial pronunciation lexicons.
Gaofeng Cheng, Haitian Lu, Chengxu Yang, Xuyang Wang 0002, Ta Li, Yonghong Yan 0002
ICASSP5
2025 Hybrid Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition
abstract
Pseudo-labeling based semi-supervised learning can mitigate the performance degradation resulting from the absence of labeled data in the target domain. In pseudo-labeling, the quality of pseudo-labels is crucial for the final performance. However, most works overlook the potential benefits of using decoder for pseudo-labels within the the mainstream hybrid Connectionist Temporal Classification (CTC) and attention (CTC/attention) based ASR architecture. Therefore, we propose Hybrid Pseudo-Labeling (HPL) to improve the quality of pseudo-labels during online decoding. HPL introduces a second-stage decoding using the decoder to alleviate substitution errors arising from the conditional independence assumption inherent in CTC for error correction. Furthermore, we propose Hybrid Selection to optimally combine results of encoder and decoder. Additionally, we introduce Speed Perturbation Enhancement (SPE) to further enhance the quality of pseudo-labels via speed perturbation. Experiments demonstrate that HPL achieves state-of-the-art performance compared to other mainstream pseudo-labeling methods.
Han Zhu 0004, Chengxu Yang, Gaofeng Cheng, Ta Li
ICASSP6
2025 Rainbow Delay Compensation: A Multi-Agent Reinforcement Learning Framework for Mitigating Observation Delays
abstract
In real-world multi-agent systems (MASs), observation delays are ubiquitous, preventing agents from making decisions based on the environment's true state. An individual agent's local observation typically comprises multiple components from other agents or dynamic entities within the environment. These discrete observation components with varying delay characteristics pose significant challenges for multi-agent reinforcement learning (MARL). In this paper, we first formulate the decentralized stochastic individual delay partially observable Markov decision process (DSID-POMDP) by extending the standard Dec-POMDP. We then propose the Rainbow Delay Compensation (RDC), a MARL training framework for addressing stochastic individual delays, along with recommended implementations for its constituent modules. We implement the DSID-POMDP's observation generation pattern using standard MARL benchmarks, including MPE and SMAC. Experiments demonstrate that baseline MARL methods suffer severe performance degradation under fixed and unfixed delays. The RDC-enhanced approach mitigates this issue, remarkably achieving ideal delay-free performance in certain delay scenarios while maintaining generalizability. Our work provides a novel perspective on multi-agent delayed observation problems and offers an effective solution framework. The source code is available at https://github.com/linkjoker1006/RDC-pymarl.
Songchen Fu, Siang Chen, Shaojing Zhao, Letian Bai, Ta Li, YongHong Yan
NeurIPS6
2025 Graph Neural Network-Enhanced Feature Learning for Unsupervised Anomalous Sound Detection
abstract
Anomalous sound detection (ASD) is crucial in industrial applications due to its non-invasive and real-time capabilities. However, existing ASD methods often rely on autoencoders, which require machine-specific tuning, or large pre-trained models with high computational costs. Additionally, many self-supervised approaches depend on extensive meta-information, increasing deployment complexity. To address these limitations, we propose a lightweight, metadata-free ASD framework that generalizes across different machine types without requiring complex hyperparameter tuning. Our approach extracts high-dimensional features from Log-Mel spectrograms using MobileNetV2, then refines feature representations through relational learning with a Graph Neural Network-based SAGE-GAT model. Unlike conventional methods that treat machine types independently, our approach leverages cross-category feature propagation through local neighbor relationships, capturing discriminative information from nearby samples. Furthermore, an MLP optimized with ArcFace loss enhances feature structuring, while anomaly detection is performed using K-means clustering. Experiments on the DCASE 2024 Task 2 dataset validate the effectiveness of our approach, demonstrating its robustness, efficiency, and suitability for real-world industrial deployment.
Jiyu Lu, Wenbo Guan, Ming Zhang 0035, Ta Li
SMC4
2025 Anomalous sound detection using sound image and CTF-bilateral filter
Jiyu Lu, Ming Zhang 0035, Ta Li
Neurocomputing3
2025 QTypeMix: Enhancing multi-agent cooperative strategies through heterogeneous and homogeneous value decomposition
Songchen Fu, Shaojing Zhao, Ta Li, YongHong Yan
Neural Networks3
2024 One-Epoch Training with Single Test Sample in Test Time for Better Generalization of Cough-Based Covid-19 Detection Model
abstract
The outbreak of COVID-19 has raised researchers’ attention to audio-based rapid disease detection. Most of the previous studies have obtained competitive detection performance. However, these results are usually obtained by testing data from the same source offline. When making cross-dataset testing, the performance may deteriorate dramatically due to inconsistent data distribution between different datasets. In addition, in practical application, the model has to make a prediction for the current test audio without any prior information, which requires good model generalization under limited training data. To address the above issues, we adopt a test-time training framework to achieve a cough-based COVID-19 detection model with better generalizability. In the model development stage, resnet18 serves as the backbone network and a self-supervised learning branch is added as an auxiliary task. In testing stage, the model parameters are first fine-tuned by the self-supervised branch with the single test audio as input, and then the classification head outputs predictions. The proposed method is validated on three open-source datasets using a variety of hyperparameters. In cross-dataset testing, AUC and UAR increase by 3.65% and 3% on average absolutely, respectively. The results show that the proposed framework is applicable to improve the model performance in practical application.
Jiakun Shen, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002, Qingwei Zhao, Ta Li, Yanfen Tang, Shaoxing Zhang
ICASSP6
2024 Expressive paragraph text-to-speech synthesis with multi-step variational autoencoder
Xuyuan Li, Zengqiang Shang, Peiyang Shi, Hua Hua, Ta Li, Pengyuan Zhang
INTERSPEECH5
2024 Contextual Biasing with Confidence-based Homophone Detector for Mandarin End-to-End Speech Recognition
Chengxu Yang, Sanli Tian, Gaofeng Cheng, Sujie Xiao, Ta Li
INTERSPEECH6
2024 Factorized and progressive knowledge distillation for CTC-based ASR models
Sanli Tian, Zehan Li, Zhaobiao Lyv, Gaofeng Cheng, Ta Li, Qingwei Zhao
Speech Commun.6
2024 Unsupervised Domain Adaptation on End-to-End Multi-Talker Overlapped Speech Recognition
abstract
Serialized Output Training (SOT) has emerged as the mainstream approach for addressing the multi-talker overlapped speech recognition challenge due to its simplicity. However, SOT encounters cross-domain performance degradation which hinders its application. Meanwhile, traditional domain adaption methods may harm the accuracy of speaker change point prediction evaluated by UD-CER, which is an important metric in SOT. To solve these issues, we propose Pseudo-Labeling based SOT (PL-SOT) for domain adaptation by treating speaker change token ($< $sc$>$) specially during training to increase the accuracy of speaker change point prediction. Firstly, we improve CTC loss by proposingWeakening and Enhancing CTC(WE-CTC) loss to weaken the learning of error-prone labels surrounding$<$sc$>$while enhance the emission probability of$< $sc$>$through modifying posteriors of the pseudo-labels. Secondly, we introduceWeighted Confidence Filter(WCF) that assigns higher scores of$<$sc$>$to exclude low-quality pseudo-labels without hurting the$< $sc$>$prediction. Experimental results show that PL-SOT achieves 17.7%/12.8% average relative reduction of CER/UD-CER, with AliMeeting as source domain and AISHELL-4 along with MagicData-RAMC as target domain.
Han Zhu 0004, Sanli Tian, Qingwei Zhao, Ta Li
IEEE Signal Process. Lett.5
2023 A Verifiable Privacy-Preserving Outsourced Prediction Scheme Based on Blockchain in Smart Healthcare
abstract
The swift progression of the Internet of Things and the extensive integration of machine learning have spurred the growth of intelligent healthcare. Many intelligent healthcare devices, limited by their own computing and storage resources, require outsourcing data analysis tasks to cloud platforms for efficient and accurate results. Unfortunately, malicious cloud services lead to privacy breaches in outsourced data and untrustworthiness in learning models. To address these challenges, this paper proposes a verifiable privacy-preserving outsourced prediction scheme based on blockchain in smart healthcare (VPOL). Specifically, by incorporating blockchain technology into VPOL, we build a robust and scalable framework to prevent falsification of outsourced data and learning models in a decentralized and transparent manner. Then, we design a training committee approach to ensure the reliability of outsourced prediction and employ homomorphic encryption and commitment scheme to protect the privacy and integrity of the data. Finally, theoretical analysis proves the effectiveness and security of VPOL. Sufficient experiments demonstrate that VPOL achieves the approximate accuracy of the plaintext.
Ta Li, Youliang Tian, Jinbo Xiong
HealthCom1
2023 SFA: Searching faster architectures for end-to-end automatic speech recognition models
Ta Li, Pengyuan Zhang, Yonghong Yan 0002
Comput. Speech Lang.2
2023 FVP-EOC: Fair, Verifiable, and Privacy-Preserving Edge Outsourcing Computing in 5G-Enabled IIoT
abstract
The 5G-enabled Industrial Internet of Things tilts the data processing model from the cloud to the edge. Users are more inclined to get feedback and data analysis of outsourcing computing results timely from the edge. However, the existing solutions undermine the fairness of multitask outsourcing in edge environment and cannot guarantee the correctness of results. To tackle these challenges, in this article, we propose a fair, verifiable, and privacy-preserving edge outsourcing computing scheme based on blockchain (FVP-EOC). Initially, we propose a task bidding method in the same round of task outsourcing, which improves the utilization of resources and the fairness of the FVP-EOC by dividing tasks into blocks. Furthermore, we design a result verification algorithm and a consensus algorithm to ensure the correctness of the results without a trusted third party. Finally, theoretical analysis and ample simulations indicate that the FVP-EOC is secure and verifiable and ensure the benefits of all the participants in edge outsourcing computing.
Ta Li, Youliang Tian, Jinbo Xiong, Md. Zakirul Alam Bhuiyan
IEEE Trans. Ind. Informatics1
2022 Improving Streaming End-to-End ASR on Transformer-based Causal Models with Encoder States Revision Strategies
abstract
There is often a trade-off between performance and latency in streaming automatic speech recognition (ASR).Traditional methods such as look-ahead and chunk-based methods, usually require information from future frames to advance recognition accuracy, which incurs inevitable latency even if the computation is fast enough.A causal model that computes without any future frames can avoid this latency, but its performance is significantly worse than traditional methods.In this paper, we propose corresponding revision strategies to improve the causal model.Firstly, we introduce a real-time encoder states revision strategy to modify previous states.Encoder forward computation starts once the data is received and revises the previous encoder states after several frames, which is no need to wait for any right context.Furthermore, a CTC spike position alignment decoding algorithm is designed to reduce time costs brought by the proposed revision strategy.Experiments are all conducted on Librispeech datasets.Fine-tuning on the CTC-based wav2vec2.0model, our best method can achieve 3.7/9.2WERs on test-clean/other sets and brings 45% relative improvement for causal models, which is also competitive with the chunkbased methods and the knowledge distillation methods.
Zehan Li, Haoran Miao, Keqi Deng, Gaofeng Cheng, Sanli Tian, Ta Li, Yonghong Yan 0002
INTERSPEECH6
2022 NAS-SCAE: Searching Compact Attention-based Encoders For End-to-end Automatic Speech Recognition
Ta Li, Pengyuan Zhang, Yonghong Yan 0002
INTERSPEECH2
2022 Knowledge Distillation For CTC-based Speech Recognition Via Consistent Acoustic Representation Learning
Sanli Tian, Keqi Deng, Zehan Li, Lingxuan Ye, Gaofeng Cheng, Ta Li, Yonghong Yan 0002
INTERSPEECH6
2022 Self-Supervised Pre-Training for Attention-Based Encoder-Decoder ASR Model
abstract
End-to-end (E2E) models, including the attention-based encoder-decoder (AED) models, have achieved promising performance on the automatic speech recognition (ASR) task. However, the supervised training process of the E2E model needs a large amount of speech-text paired data. In contrast, self-supervised pre-training can pre-train the model on the unlabeled data and then fine-tune it on the limited labeled data to realize better performance. Most of the previous self-supervised pre-training methods focus on learning hidden representations from speech but ignore how to utilize the unpaired text. As a result, previous works often pre-train an acoustic encoder and then fine-tune it as a classification based ASR model, such as Connectionist Temporal Classification (CTC) based model, rather than an AED model. In this paper, we propose a self-supervised pre-training method for the AED model (SP-AED). The SP-AED method contains acoustic pre-training for the encoder, linguistic pre-training for the decoder, and an adaptive combination fine-tuning for the whole system. We first design a linguistic pre-training method for decoder by utilizing the text-only data. The decoder will be pre-trained as a noise-condition language model to learn the prior distribution of the text. Then, we pre-train the AED encoder with the wav2vec2.0 method with some modifications. Finally, we combine the pre-trained encoder and decoder and fine-tune them on the limited labeled data. We design an adaptive combination method during fine-tuning by modifying the decoder’s input and output to prevent catastrophic forgetting. Experiments prove that compared with the random initialized models, the SP-AED pre-trained models can realize up to 17% relative improvement. And with similar model size or computational cost, we can get comparable results to other classification-based models on both English and Chinese corpus.
Changfeng Gao, Gaofeng Cheng, Ta Li, Pengyuan Zhang, Yonghong Yan 0002
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 A Blockchain-Based Machine Learning Framework for Edge Services in IIoT
abstract
Edge services provide an effective and superior means of real-time transmissions and rapid processing of information in the Industrial Internet of Things (IIoT). However, the continuous increase of the number of smart devices results in privacy leakage and insufficient model accuracy of edge services. To tackle these challenges, in this article, we propose a blockchain-based machine learning framework for edge services (BML-ES) in IIoT. Specifically, we construct novel smart contracts to encourage multiparty participation of edge services to improve the efficiency of data processing. Moreover, we propose an aggregation strategy to verify and aggregate model parameters to ensure the accuracy of decision tree models. Finally, based on the SM2 public key cryptosystem, we protect data security and prevent data privacy leakage in edge services. Theoretical analysis and simulation experiments indicate that the BML-ES framework is secure, effective, and efficient, and is better suitable to improve the accuracy of edge services in IIoT.
Youliang Tian, Ta Li, Jinbo Xiong, Md. Zakirul Alam Bhuiyan, Jianfeng Ma 0001, Changgen Peng
IEEE Trans. Ind. Informatics2
2021 RNN-T Based Open-Vocabulary Keyword Spotting in Mandarin with Multi-Level Detection
abstract
Despite the recent prevalence of keyword spotting (KWS) in smart-home, open-vocabulary KWS remains a keen but unmet need among the users. In this paper, we propose an RNN Transducer (RNN-T) based keyword spotting system with a constrained attention mechanism biasing module that biases the RNN-T model towards a specific keyword of interest. The atonal syllables are adopted as the modeling units, which addresses the out-of-vocabulary (OOV) problem. A multi-level detection is applied to the posterior probabilities for the judgement. Evaluating on the AISHELL-2 dataset shows our proposed method outperforms the RNN-T-based approach by 2.70% in false reject rate (FRR) at 1 false alarm (FA) per hour. We further provide insights into the role of each stage of the detection cascade, where most negative samples are filtered out by the first stage with high computational efficiency.
Zuozhen Liu, Ta Li, Pengyuan Zhang
ICASSP2
2021 Keyword Search Using Attention-Based End-to-End ASR and Frame-Synchronous Phoneme Alignments
abstract
Attention-based end-to-end (E2E) automatic speech recognition (ASR) architectures are now the state-of-the-art in terms of recognition performance. However, despite their effectiveness, they have not been widely applied in keyword search (KWS) tasks yet. In this paper, we propose the Att-E2E-KWS architecture, an attention-based E2E ASR framework for KWS that can afford accurate and reliable keyword retrieval results. First, we design a basic framework to carry out KWS based on attention-based E2E ASR. We adopt the connectionist temporal classification and attention (CTC/Att) joint E2E ASR architecture and exploit the spike posterior property of CTC to provide the keywords time stamps. Second, we introduce the frame-synchronous phonemes modeling and use the dynamic programming (DP) algorithm to provide alignments between E2E grapheme outputs and phoneme outputs. We call this alignment procedure dynamic time alignment (DTA), which can provide the proposed Att-E2E-KWS system with more accurate time stamps and reliable confidence scores. Third, we use the Transformer, a self-attention-based encoder-decoder neural network, in place of conventional recurrent neural networks in order to yield more parallelizable models and increased training speed. We conduct comprehensive experiments on English and Mandarin Chinese. To the best of our knowledge, this is the first practical Att-E2E-KWS framework, and experimental results on Switchboard and HKUST corpora show that our proposed Att-E2E-KWS systems significantly outperform the CTC E2E ASR based KWS baselines.
Runyan Yang, Gaofeng Cheng, Haoran Miao, Ta Li, Pengyuan Zhang, Yonghong Yan 0002
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Online Hybrid CTC/Attention Architecture for End-to-End Speech Recognition
Haoran Miao, Gaofeng Cheng, Pengyuan Zhang, Ta Li, Yonghong Yan 0002
INTERSPEECH4
2016 An unsupervised vocabulary selection technique for Chinese automatic speech recognition
abstract
The vocabulary is a vital component of automatic speech recognition(ASR) systems. For a specific Chinese speech recognition task, using a large general vocabulary not only leads to a much longer time to decode, but also hurts the recognition accuracy. In this paper, we proposed an unsupervised algorithm to select task-specific words from a large general vocabulary. The out-of-vocabulary(OOV) rate is a measure of vocabularies, and it is related to the recognition accuracy. However, it is hard to compute OOV rate for a Chinese vocabulary, since OOVs are often segmented into single Chinese characters and most Chinese vocabularies contain all the single Chinese characters. To deal with this problem, we proposed a novel method to estimate the OOV rate of Chinese vocabularies. In experiments, we found that our estimated OOV rate is related to the character error rate(CER) of recognition. Our proposed vocabulary selection method provided both the lowest OOV rate and CER on two Chinese conversational telephone speech(CTS) evaluation sets compared to the general vocabulary and frequency based vocabulary selection method. In addition, our proposed method significantly reduced the size of the language model(LM) and the corresponding weighted finite state transducer(WFST) network, which led to a more efficient decoding.
Pengyuan Zhang, Ta Li, Yonghong Yan 0002
SLT3
2013 Prefix tree based n-best list re-scoring for recurrent neural network language model used in speech recognition system
Yujing Si, Ta Li, Jielin Pan, Yonghong Yan 0002
INTERSPEECH3
2009 Simultaneous Synchronization of Text and Speech for Broadcast News Subtitling
Jie Gao 0020, Qingwei Zhao, Ta Li, Yonghong Yan 0002
ISNN (3)3
2009 Improving Voice Search Using Forward-Backward LVCSR System Combination
Ta Li, Changchun Bao, Weiqun Xu, Jielin Pan, Yonghong Yan 0002
ISNN (4)1
2008 Nonnative speech recognition based on state-candidate bilingual model modification
Ta Li, Jielin Pan, Yonghong Yan 0002
INTERSPEECH2