Sanli Tian

dblp:324/3030 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2024
0000-0002-5560-7673ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2024 Contextual Biasing with Confidence-based Homophone Detector for Mandarin End-to-End Speech Recognition
Chengxu Yang, Sanli Tian, Gaofeng Cheng, Sujie Xiao, Ta Li
INTERSPEECH3
2024 Factorized and progressive knowledge distillation for CTC-based ASR models
Sanli Tian, Zehan Li, Zhaobiao Lyv, Gaofeng Cheng, Ta Li, Qingwei Zhao
Speech Commun.1
2024 Unsupervised Domain Adaptation on End-to-End Multi-Talker Overlapped Speech Recognition
abstract
Serialized Output Training (SOT) has emerged as the mainstream approach for addressing the multi-talker overlapped speech recognition challenge due to its simplicity. However, SOT encounters cross-domain performance degradation which hinders its application. Meanwhile, traditional domain adaption methods may harm the accuracy of speaker change point prediction evaluated by UD-CER, which is an important metric in SOT. To solve these issues, we propose Pseudo-Labeling based SOT (PL-SOT) for domain adaptation by treating speaker change token ($< $sc$>$) specially during training to increase the accuracy of speaker change point prediction. Firstly, we improve CTC loss by proposingWeakening and Enhancing CTC(WE-CTC) loss to weaken the learning of error-prone labels surrounding$<$sc$>$while enhance the emission probability of$< $sc$>$through modifying posteriors of the pseudo-labels. Secondly, we introduceWeighted Confidence Filter(WCF) that assigns higher scores of$<$sc$>$to exclude low-quality pseudo-labels without hurting the$< $sc$>$prediction. Experimental results show that PL-SOT achieves 17.7%/12.8% average relative reduction of CER/UD-CER, with AliMeeting as source domain and AISHELL-4 along with MagicData-RAMC as target domain.
Han Zhu 0004, Sanli Tian, Qingwei Zhao, Ta Li
IEEE Signal Process. Lett.3
2022 Improving Streaming End-to-End ASR on Transformer-based Causal Models with Encoder States Revision Strategies
abstract
There is often a trade-off between performance and latency in streaming automatic speech recognition (ASR).Traditional methods such as look-ahead and chunk-based methods, usually require information from future frames to advance recognition accuracy, which incurs inevitable latency even if the computation is fast enough.A causal model that computes without any future frames can avoid this latency, but its performance is significantly worse than traditional methods.In this paper, we propose corresponding revision strategies to improve the causal model.Firstly, we introduce a real-time encoder states revision strategy to modify previous states.Encoder forward computation starts once the data is received and revises the previous encoder states after several frames, which is no need to wait for any right context.Furthermore, a CTC spike position alignment decoding algorithm is designed to reduce time costs brought by the proposed revision strategy.Experiments are all conducted on Librispeech datasets.Fine-tuning on the CTC-based wav2vec2.0model, our best method can achieve 3.7/9.2WERs on test-clean/other sets and brings 45% relative improvement for causal models, which is also competitive with the chunkbased methods and the knowledge distillation methods.
Zehan Li, Haoran Miao, Keqi Deng, Gaofeng Cheng, Sanli Tian, Ta Li, Yonghong Yan 0002
INTERSPEECH5
2022 Knowledge Distillation For CTC-based Speech Recognition Via Consistent Acoustic Representation Learning
Sanli Tian, Keqi Deng, Zehan Li, Lingxuan Ye, Gaofeng Cheng, Ta Li, Yonghong Yan 0002
INTERSPEECH1
2022 Improving Recognition of Out-of-vocabulary Words in E2E Code-switching ASR by Fusing Speech Generation Methods
Lingxuan Ye, Gaofeng Cheng, Runyan Yang, Zehui Yang, Sanli Tian, Pengyuan Zhang, Yonghong Yan 0002
INTERSPEECH5
2022 VigilanceNet: Decouple Intra- and Inter-Modality Learning for Multimodal Vigilance Estimation in RSVP-Based BCI
abstract
Recently, brain-computer interface (BCI) technology has made impressive progress and has been developed for many applications. Thereinto, the BCI system based on rapid serial visual presentation (RSVP) is a promising information detection technology. However, the use of RSVP is closely related to the user's performance, which can be influenced by their vigilance levels. Therefore it is crucial to detect vigilance levels in RSVP-based BCI. In this paper, we conducted a long-term RSVP target detection experiment to collect electroencephalography (EEG) and electrooculogram (EOG) data at different vigilance levels. In addition, to estimate vigilance levels in RSVP-based BCI, we propose a multimodal method named VigilanceNet using EEG and EOG. Firstly, we define the multiplicative relationships in conventional EOG features that can better describe the relationships between EOG features, and design an outer product embedding module to extract the multiplicative relationships. Secondly, we propose to decouple the learning of intra- and inter-modality to improve multimodal learning. Specifically, for intra-modality, we introduce an intra-modality representation learning (intra-RL) method to obtain effective representations of each modality by letting each modality independently predict vigilance levels during the multimodal training process. For inter-modality, we employ the cross-modal Transformer based on cross-attention to capture the complementary information between EEG and EOG, which only pays attention to the inter-modality relations. Extensive experiments and ablation studies are conducted on the RSVP and SEED-VIG public datasets. The results demonstrate the effectiveness of the method in terms of regression error and correlation.
Wei Wei 0046, Changde Du, Shuang Qiu 0002, Sanli Tian, Huiguang He
ACM Multimedia5