Keyu An

dblp:254/2051 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
7since 2021 · last 2023
0000-0003-0040-0883ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2023 BAT: Boundary aware transducer for memory-efficient and low-latency ASR
Keyu An, Xian Shi, Shiliang Zhang
INTERSPEECH1
2022 CUSIDE: Chunking, Simulating Future Context and Decoding for Streaming ASR
Keyu An, Huahuan Zheng, Zhijian Ou, Hongyu Xiang, Guanglu Wan
INTERSPEECH1
2022 An Empirical Study of Language Model Integration for Transducer based Speech Recognition
abstract
Utilizing text-only data with an external language model (ELM) in end-to-end RNN-Transducer (RNN-T) for speech recognition is challenging.Recently, a class of methods such as density ratio (DR) and internal language model estimation (ILME) have been developed, outperforming the classic shallow fusion (SF) method.The basic idea behind these methods is that RNN-T posterior should first subtract the implicitly learned internal language model (ILM) prior, in order to integrate the ELM.While recent studies suggest that RNN-T only learns some low-order language model information, the DR method uses a well-trained neural language model with full context, which may be inappropriate for the estimation of ILM and deteriorate the integration performance.Based on the DR method, we propose a loworder density ratio method (LODR) by replacing the estimation with a low-order weak language model.Extensive empirical experiments are conducted on both in-domain and cross-domain scenarios on English LibriSpeech & Tedlium-2 and Chinese WenetSpeech & AISHELL-1 datasets.It is shown that LODR consistently outperforms SF in all tasks, while performing generally close to ILME and better than DR in most tests.
Huahuan Zheng, Keyu An, Zhijian Ou, Guanglu Wan
INTERSPEECH2
2021 Multilingual and Crosslingual Speech Recognition Using Phonological-Vector Based Phone Embeddings
abstract
The use of phonological features (PFs) potentially allows language-specific phones to remain linked in training, which is highly desirable for information sharing for multilingual and crosslingual speech recognition methods for low-resourced languages. A drawback suffered by previous methods in using phonological features is that the acoustic-to-PF extraction in a bottom-up way is itself difficult. In this paper, we propose to join phonology driven phone embedding (top-down) and deep neural network (DNN) based acoustic feature extraction (bottom-up) to calculate phone probabilities. The new method is called JoinAP (Joining of Acoustics and Phonology). Remarkably, no inversion from acoustics to phonological features is required for speech recognition. For each phone in the IPA (International Phonetic Alphabet) table, we encode its phonological features to a phonological-vector, and then apply linear or nonlinear transformation of the phonological-vector to obtain the phone embedding. A series of multilingual and crosslingual (both zero-shot and few-shot) speech recognition experiments are conducted on the CommonVoice dataset (German, French, Spanish and Italian) and the AISHLL-1 dataset (Mandarin), and demonstrate the superiority of JoinAP with nonlinear phone embeddings over both JoinAP with linear phone embeddings and the traditional method with flat phone embeddings.
Chengrui Zhu, Keyu An, Huahuan Zheng, Zhijian Ou
ASRU2
2021 Deformable TDNN with Adaptive Receptive Fields for Speech Recognition
abstract
Time Delay Neural Networks (TDNNs) are widely used in both DNN-HMM based hybrid speech recognition systems and recent end-to-end systems.Nevertheless, the receptive fields of TDNNs are limited and fixed, which is not desirable for tasks like speech recognition, where the temporal dynamics of speech are varied and affected by many factors.In this paper, we propose to use deformable TDNNs for adaptive temporal dynamics modeling in end-to-end speech recognition.Inspired by deformable ConvNets, deformable TDNNs augment the temporal sampling locations with additional offsets and learn the offsets automatically based on the ASR criterion, without additional supervision.Experiments show that deformable TDNNs obtain state-of-the-art results on WSJ benchmarks (1.42%/3.45%WER on WSJ eval92/dev93 respectively), outperforming standard TDNNs significantly.Furthermore, we propose the latency control mechanism for deformable TDNNs, which enables deformable TDNNs to do streaming ASR without accuracy degradation.
Keyu An, Zhijian Ou
Interspeech1
2021 The SLT 2021 Children Speech Recognition Challenge: Open Datasets, Rules and Baselines
abstract
Automatic speech recognition (ASR) has been significantly advanced with the use of deep learning and big data. How-ever improving robustness, including achieving equally good performance on diverse speakers and accents, is still a challenging problem. In particular, the performance of children speech recognition (CSR) still lags behind due to 1) the speech and language characteristics of children's voice are substantially different from those of adults and 2) sizable open dataset for children speech is still not available in the research community. To address these problems, we launch the Children Speech Recognition Challenge (CSRC), as a flagship satellite event of IEEE SLT 2021 workshop. The challenge will release about 400 hours of Mandarin speech data for registered teams and set up two challenge tracks and provide a common testbed to benchmark the CSR performance. In this paper, we introduce the datasets, rules, evaluation method as well as baselines.
Fan Yu 0002, Zhuoyuan Yao, Keyu An, Lei Xie 0001, Zhijian Ou, Xiulin Li, Guanqiong Miao
SLT4
2021 Efficient Neural Architecture Search for End-to-End Speech Recognition Via Straight-Through Gradients
abstract
Neural Architecture Search (NAS), the process of automating architecture engineering, is an appealing next step to advancing end-to-end Automatic Speech Recognition (ASR), replacing expert-designed networks with learned, task-specific architectures. In contrast to early computational-demanding NAS methods, recent gradient-based NAS methods, e.g., DARTS (Differentiable ARchiTecture Search), SNAS (Stochastic NAS) and ProxylessNAS, significantly improve the NAS efficiency. In this paper, we make two contributions. First, we rigorously develop an efficient NAS method via Straight-Through (ST) gradients, called ST-NAS. Basically, ST-NAS uses the loss from SNAS but uses ST to back-propagate gradients through discrete variables to optimize the loss, which is not revealed in ProxylessNAS. Using ST gradients to support sub-graph sampling is a core element to achieve efficient NAS beyond DARTS and SNAS. Second, we successfully apply ST-NAS to end-to-end ASR. Experiments over the widely benchmarked 80-hour WSJ and 300-hour Switchboard datasets show that the ST-NAS induced architectures significantly outperform the human-designed architecture across the two datasets. Strengths of ST-NAS such as architecture transferability and low computation cost in memory and time are also reported.
Huahuan Zheng, Keyu An, Zhijian Ou
SLT2
2020 Sequential Deformation for Accurate Scene Text Detection
Shanyu Xiao, Liangrui Peng, Ruijie Yan, Keyu An, Jaesik Min
ECCV (29)4
2020 CAT: A CTC-CRF Based ASR Toolkit Bridging the Hybrid and the End-to-End Approaches Towards Data Efficiency and Low Latency
abstract
In this paper, we present a new open source toolkit for speech recognition, named CAT (CTC-CRF based ASR Toolkit).CAT inherits the data-efficiency of the hybrid approach and the simplicity of the E2E approach, providing a full-fledged implementation of CTC-CRFs and complete training and testing scripts for a number of English and Chinese benchmarks.Experiments show CAT obtains state-of-the-art results, which are comparable to the fine-tuned hybrid models in Kaldi but with a much simpler training pipeline.Compared to existing nonmodularized E2E models, CAT performs better on limited-scale datasets, demonstrating its data efficiency.Furthermore, we propose a new method called contextualized soft forgetting, which enables CAT to do streaming ASR without accuracy degradation.We hope CAT, especially the CTC-CRF based framework and software, will be of broad interest to the community, and can be further explored and improved.
Keyu An, Hongyu Xiang, Zhijian Ou
INTERSPEECH1