Xian Shi

dblp:15/10772 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
abstract
Forced alignment (FA) predicts start and end timestamps for words or characters in speech, but existing methods are language-specific and prone to cumulative temporal shifts.The multilingual speech understanding and longsequence processing abilities of speech large language models (SLLMs) make them promising for FA in multilingual, crosslingual, and long-form speech settings.However, directly applying the next-token prediction paradigm of SLLMs to FA results in hallucinations and slow inference.To bridge the gap, we propose LLM-ForcedAligner, reformulating FA as a slot-filling paradigm: timestamps are treated as discrete indices, and special timestamp tokens are inserted as slots into the transcript.Conditioned on the speech embeddings and the transcript with slots, the SLLM directly predicts the time indices at slots.During training, causal attention masking with non-shifted input and label sequences allows each slot to predict its own timestamp index based on itself and preceding context, with loss computed only at slot positions.Dynamic slot insertion enables FA at arbitrary positions.Moreover, non-autoregressive inference is supported, avoiding hallucinations and improving speed.Experiments across multilingual, crosslingual, and long-form speech scenarios show that LLM-ForcedAligner achieves a 69%~78% relative reduction in accumulated averaging shift compared with prior methods.Checkpoint and inference code are available at https://huggingface.co/Qwen/ Qwen3-ForcedAligner-0.6B.
Bingshen Mu, Xian Shi, Hexin Liu, Jin Xu 0010, Lei Xie 0001
ACL (1)2
2026 Error Exponents of Probabilistic Entanglement Transformation under Approximately Non-Entangling Instruments
Xian Shi
ISIT1
2024 SeACo-Paraformer: A Non-Autoregressive ASR System with Flexible and Effective Hotword Customization Ability
abstract
Hotword customization is one of the concerned issues remained in ASR field - it is of value to enable users of ASR systems to customize names of entities, persons and other phrases to obtain better experience. The past few years have seen effective modeling strategies for ASR contextualization developed, but they still exhibit space for improvement about training stability and the invisible activation process. In this paper we propose Semantic-Augmented Contextual-Paraformer (SeACo-Paraformer) a novel NAR based ASR system with flexible and effective hotword customization ability. It possesses the advantages of AED-based model’s accuracy, NAR model’s efficiency, and explicit customization capacity of superior performance. Through extensive experiments with 50,000 hours of industrial big data, our proposed model outperforms strong baselines in customization. Besides, we explore an efficient way to filter large-scale incoming hotwords for further improvement. The industrial models compared, source codes and two hotword test sets are all open source.
Xian Shi, Yexin Yang, Yanni Chen, Zhifu Gao, Shiliang Zhang
ICASSP1
2024 SlideSpeech: A Large Scale Slide-Enriched Audio-Visual Corpus
abstract
Multi-Modal automatic speech recognition (ASR) techniques aim to leverage additional modalities to improve the performance of speech recognition systems. While existing approaches primarily focus on video or contextual information, the utilization of extra supplementary textual information has been overlooked. Recognizing the abundance of online conference videos with slides, which provide rich domain-specific information in the form of text and images, we release SlideSpeech, a large-scale audio-visual corpus enriched with slides. The corpus contains 1,705 videos, 1,000+ hours, with 473 hours of high-quality transcribed speech. Moreover, the corpus contains a significant amount of real-time synchronized slides. In this work, we present the pipeline for constructing the corpus and propose baseline methods for utilizing text information in the visual slide context. Through the application of keyword extraction and contextual ASR methods in the benchmark system, we demonstrate the potential of improving speech recognition performance by incorporating textual information from supplementary video slides.
Haoxu Wang, Fan Yu 0002, Xian Shi, Yuezhang Wang, Shiliang Zhang, Ming Li 0026
ICASSP3
2024 LCB-Net: Long-Context Biasing for Audio-Visual Speech Recognition
abstract
The growing prevalence of online conferences and courses presents a new challenge in improving automatic speech recognition (ASR) with enriched textual information from video slides. In contrast to rare phrase lists, the slides within videos are synchronized in real-time with the speech, enabling the extraction of long contextual bias. Therefore, we propose a novel long-context biasing network (LCB-net) for audio-visual speech recognition (AVSR) to leverage the long-context information available in videos effectively. Specifically, we adopt a bi-encoder architecture to simultaneously model audio and long-context biasing. Besides, we also propose a biasing prediction module that utilizes binary cross entropy (BCE) loss to explicitly determine biased phrases in the long-context biasing. Furthermore, we introduce a dynamic contextual phrases simulation to enhance the generalization and robustness of our LCB-net. Experiments on the SlideSpeech, a large-scale audio-visual corpus enriched with slides, reveal that our proposed LCB-net outperforms general ASR model by 9.4%/9.1%/10.9% relative WER/U-WER/B-WER reduction on test set, which enjoys high unbiased and biased performance. Moreover, we also evaluate our model on LibriSpeech corpus, leading to 23.8%/19.2%/35.4% relative WER/U-WER/B-WER reduction over the ASR model.
Fan Yu 0002, Haoxu Wang, Xian Shi, Shiliang Zhang
ICASSP3
2023 BAT: Boundary aware transducer for memory-efficient and low-latency ASR
Keyu An, Xian Shi, Shiliang Zhang
INTERSPEECH2
2023 FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Zhifu Gao, Jiaming Wang 0004, Haoneng Luo, Xian Shi, Mengzhe Chen, Yabin Li, Lingyun Zuo, Zhihao Du, Shiliang Zhang
INTERSPEECH5
2023 Accurate and Reliable Confidence Estimation Based on Non-Autoregressive End-to-End Speech Recognition System
Xian Shi, Haoneng Luo, Zhifu Gao, Shiliang Zhang, Zhijie Yan
INTERSPEECH1
2022 Open-Set Semi-Supervised Learning for 3D Point Cloud Understanding
abstract
Semantic understanding of 3D point cloud relies on learning models with massively annotated data, which, in many cases, are expensive or difficult to collect. This has led to an emerging research interest in semi-supervised learning (SSL) for 3D point cloud. It is commonly assumed in SSL that the unlabeled data are drawn from the same distribution as that of the labeled ones; This assumption, however, rarely holds true in realistic environments. Blindly using out-of-distribution (OOD) unlabeled data could harm SSL performance. In this work, we propose to selectively utilize unlabeled data through sample weighting, so that only conducive unlabeled data would be prioritized. To estimate the weights, we adopt a bi-level optimization framework which iteratively optimizes a meta-objective on a held-out validation set and a task-objective on a training set. Faced with the instability of efficient bi-level optimizers, we further propose three regularization techniques to enhance the training stability. Extensive experiments on 3D point cloud classification and segmentation tasks verify the effectiveness of our proposed method. We also demonstrate the feasibility of a more efficient training strategy. Our code is released on Github1.
Xian Shi, Xun Xu 0002, Wanyue Zhang, Xiatian Zhu, Chuan-Sheng Foo, Kui Jia
ICPR1
2022 Linguistic-Acoustic Similarity Based Accent Shift for Accent Recognition
abstract
General accent recognition (AR) models tend to directly extract low-level information from spectrums, which always significantly overfit on speakers or channels. Considering accent can be regarded as a series of shifts relative to native pronunciation, distinguishing accents will be an easier task with accent shift as input. But due to the lack of native utterance as an anchor, estimating the accent shift is difficult. In this paper, we propose linguistic-acoustic similarity based accent shift (LASAS) for AR tasks. For an accent speech utterance, after mapping the corresponding text vector to multiple accent-associated spaces as anchors, its accent shift could be estimated by the similarities between the acoustic embedding and those anchors. Then, we concatenate the accent shift with a dimension-reduced text vector to obtain a linguistic-acoustic bimodal representation. Compared with pure acoustic embedding, the bimodal representation is richer and more clear by taking full advantage of both linguistic and acoustic information, which can effectively improve AR performance. Experiments on Accented English Speech Recognition Challenge (AESRC) dataset show that our method achieves 77.42% accuracy on Test set, obtaining a 6.94% relative improvement over a competitive system in the challenge.
Qijie Shao, Jinghao Yan, Jian Kang 0006, Xian Shi, Pengfei Hu 0004, Lei Xie 0001
INTERSPEECH5
2022 Rare variant association tests for ancestry-matched case-control data based on conditional logistic regression
abstract
With the increasing volume of human sequencing data available, analysis incorporating external controls becomes a popular and cost-effective approach to boost statistical power in disease association studies. To prevent spurious association due to population stratification, it is important to match the ancestry backgrounds of cases and controls. However, rare variant association tests based on a standard logistic regression model are conservative when all ancestry-matched strata have the same case-control ratio and might become anti-conservative when case-control ratio varies across strata. Under the conditional logistic regression (CLR) model, we propose a weighted burden test (CLR-Burden), a variance component test (CLR-SKAT) and a hybrid test (CLR-MiST). We show that the CLR model coupled with ancestry matching is a general approach to control for population stratification, regardless of the spatial distribution of disease risks. Through extensive simulation studies, we demonstrate that the CLR-based tests robustly control type 1 errors under different matching schemes and are more powerful than the standard Burden, SKAT and MiST tests. Furthermore, because CLR-based tests allow for different case-control ratios across strata, a full-matching scheme can be employed to efficiently utilize all available cases and controls to accelerate the discovery of disease associated genes.
Jingjing Lyu, Xian Shi, Zengmiao Wang, Minghua Deng, Baoluo Sun, Chaolong Wang
Briefings Bioinform.3
2021 The Accented English Speech Recognition Challenge 2020: Open Datasets, Tracks, Baselines, Results and Methods
abstract
The variety of accents has posed a big challenge to speech recognition. The Accented English Speech Recognition Challenge (AESRC2020) is designed for providing a common testbed and promoting accent-related research. Two tracks are set in the challenge – English accent recognition (track 1) and accented English speech recognition (track 2). A set of 160 hours of accented English speech collected from 8 countries is released with labels as the training set. Another 20 hours of speech without labels is later released as the test set, including two unseen accents from another two countries used to test the model generalization ability in track 2. We also provide baseline systems for the participants. This paper first reviews the released dataset, track setups, baselines and then summarizes the challenge results and major techniques used in the submissions.
Xian Shi, Fan Yu 0002, Yizhou Lu, Yuhao Liang, Qiangze Feng, Daliang Wang, Yanmin Qian, Lei Xie 0001
ICASSP1
2021 Cascade RNN-Transducer: Syllable Based Streaming On-Device Mandarin Speech Recognition with a Syllable-To-Character Converter
abstract
End-to-end models are favored in automatic speech recognition (ASR) because of its simplified system structure and superior performance. Among these models, recurrent neural network transducer (RNN-T) has achieved significant progress in streaming on-device speech recognition because of its high-accuracy and low-latency. RNN-T adopts a prediction network to enhance language information, but its language modeling ability is limited because it still needs paired speech-text data to train. Further strengthening the language modeling ability through extra text data, such as shallow fusion with an external language model, only brings a small performance gain. In view of the fact that Mandarin Chinese is a character-based language and each character is pronounced as a tonal syllable, this paper proposes a novel cascade RNN-T approach to improve the language modeling ability of RNN-T. Our approach firstly uses an RNN-T to transform acoustic feature into syllable sequence, and then converts the syllable sequence into character sequence through an RNN-T-based syllable-to-character converter. Thus a rich text repository can be easily used to strengthen the language model ability. By introducing several important tricks, the cascade RNN-T approach surpasses the character-based RNN-T by a large margin on several Mandarin test sets, with much higher recognition quality and similar latency.
Zhuoyuan Yao, Xian Shi, Lei Xie 0001
SLT3
2015 In-vehicle speech recognition and tutorial keywords spotting for novice drivers' performance evaluation
abstract
Novice young drivers are more frequently involved in traffic accidents, and studies have shown that effective supervised driver training is the key in reducing young drivers' risks. Using our previously developed Mobile-UTDrive in-vehicle data acquisition platform, two 16-age novice drivers participated in naturalistic drive training data collection. This paper focuses on analysis of novice driver training signals from an audio processing perspective. Specifically, analysis of supervised driver instruction audio and resulting CAN-Bus maneuver operation is performed. Following a procedure which consists of noise suppression, speech recognition and keyword spotting, five tutorial keywords - Brake, Gas, Left, Right and Stop - are spotted at an overall accuracy rate of 40% versus all spontaneous continuous speech. The time stamps of these keywords are then used as indications of driving maneuvers. As examples of driving performance evaluation, the case of making Left-Turn maneuvers for the two novice drivers are assessed and compared, and the increase of driving skills over experiences are analyzed.
Xian Shi, Amardeep Sathyanarayana, Navid Shokouhi, John H. L. Hansen
Intelligent Vehicles Symposium2
2012 Optimized audit evidence gathering method based on data matching using length-filtering
abstract
Auditing is a continuous process that involves collection of audit evidences. Electronic audit evidence (EAE) is a main form of audit evidence in an informationization environment. Quick and accurate electronic audit evidence in an informationization environment has become an important decision problem for auditors in today's economy. This paper provides optimized audit evidence gathering method based on data matching using length-filtering to assist audit evidence gathering in informationization environments. Finally, based on the audit risk assessment methods proposed, some experiments and analysis of this optimized audit evidence gathering method are given to illustrate the effective.
Wei Chen 0135, Xian Shi, Wally J. Smieliauskas
SMC2