VLDB 2026 Research / reviewers in the wild / expert
Hang Lv 0001
dblp:35/9853-1
· DBLP profile ↗
15ranked-venue papers
3as first author
8since 2021 · last 2024
0000-0003-3761-1684ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Conversational Speech Recognition by Learning Audio-Textual Cross-Modal Contextual RepresentationabstractAutomatic Speech Recognition (ASR) in conversational settings presents unique challenges, including extracting relevant contextual information from previous conversational turns. Due to irrelevant content, error propagation, and redundancy, existing methods struggle to extract longer and more effective contexts. To address this issue, we introduce a novel conversational ASR system, extending the Conformer encoder-decoder model with cross-modal conversational representation. Our approach leverages a cross-modal extractor that combines pre-trained speech and text models through a specialized encoder and a modal-level mask input. This enables the extraction of richer historical speech context without explicit error propagation. We also incorporate conditional latent variational modules to learn conversational-level attributes such as role preference and topic coherence. By introducing both cross-modal and conversational representations into the decoder, our model retains longer context without information loss, achieving relative accuracy improvements of 8.8% and 23% on Mandarin conversation datasets HKUST and MagicData-RAMC, respectively, compared to the standard Conformer model. Hang Lv 0001, Lei Xie 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | WENETSPEECH: A 10000+ Hours Multi-Domain Mandarin Corpus for Speech RecognitionabstractIn this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours in total. We collect the data from YouTube and Podcast, which covers a variety of speaking styles, scenarios, domains, topics and noisy conditions. An optical character recognition (OCR) method is introduced to generate the audio/text segmentation candidates for the YouTube data on the corresponding video subtitles, while a high-quality ASR transcription system is used to generate audio/text pair candidates for the Podcast data. Then we propose a novel end-to-end label error detection approach to further validate and filter the candidates. We also provide three manually labelled high-quality test sets along with WenetSpeech for evaluation – Dev for cross-validation purpose in training, Test_Net, collected from Internet for matched test, and Test_Meeting, recorded from real meetings for more challenging mismatched test. Baseline systems trained with WenetSpeech are provided for three popular speech recognition toolkits, namely Kaldi, ESPnet, and WeNet, and recognition results on the three test sets are also provided as benchmarks. To the best of our knowledge, WenetSpeech is the current largest open-source Mandarin speech corpus with transcriptions, which benefits research on production-level speech recognition. Hang Lv 0001, Qijie Shao, Chao Yang 0031, Lei Xie 0001, Hui Bu, Chenchen Zeng, Di Wu 0061, Zhendong Peng |
ICASSP | 2 |
| 2022 | Minimizing Sequential Confusion Error in Speech Command RecognitionabstractSpeech command recognition (SCR) has been commonly used on resource constrained devices to achieve hands-free user experience.However, in real applications, confusion among commands with similar pronunciations often happens due to the limited capacity of small models deployed on edge devices, which drastically affects the user experience.In this paper, inspired by the advances of discriminative training in speech recognition, we propose a novel minimize sequential confusion error (MSCE) training criterion particularly for SCR, aiming to alleviate the command confusion problem.Specifically, we aim to improve the ability of discriminating the target command from other commands on the basis of MCE discriminative criteria.We define the likelihood of different commands through connectionist temporal classification (CTC).During training, we propose several strategies to use prior knowledge creating a confusing sequence set for similar-sounding command instead of creating the whole non-target command set, which can better save the training resources and effectively reduce command confusion errors.Specifically, we design and compare three different strategies for confusing set construction.By using our proposed method, we can relatively reduce the False Reject Rate (FRR) by 33.7% at 0.01 False Alarm Rate (FAR) and confusion errors by 18.28% on our collected speech command set. Zhanheng Yang, Hang Lv 0001, Lei Xie 0001 |
INTERSPEECH | 2 |
| 2022 | WeNet 2.0: More Productive End-to-End Speech Recognition ToolkitabstractRecently, we made available WeNet [1], a production-oriented end-to-end speech recognition toolkit, which introduces a unified two-pass (U2) framework and a built-in runtime to address the streaming and non-streaming decoding modes in a single model.To further improve ASR performance and facilitate various production requirements, in this paper, we present WeNet 2.0 with four important updates.(1) We propose U2++, a unified two-pass framework with bidirectional attention decoders, which includes the future contextual information by a right-toleft attention decoder to improve the representative ability of the shared encoder and the performance during the rescoring stage.(2) We introduce an n-gram based language model and a WFSTbased decoder into WeNet 2.0, promoting the use of rich text data in production scenarios.(3) We design a unified contextual biasing framework, which leverages user-specific context (e.g., contact lists) to provide rapid adaptation ability for production and improves ASR accuracy in both with-LM and without-LM scenarios.(4) We design a unified IO to support large-scale data for effective model training.In summary, the brand-new WeNet 2.0 achieves up to 10% relative recognition performance improvement over the original WeNet on various corpora and makes available several important production-oriented features. Di Wu 0061, Zhendong Peng, Xingchen Song, Zhuoyuan Yao, Hang Lv 0001, Lei Xie 0001, Chao Yang 0031, Fuping Pan, Jianwei Niu 0002 |
INTERSPEECH | 6 |
| 2022 | W-Infer-polation: Approximate reasoning via integrating weighted fuzzy rule inference and interpolation
Hang Lv 0001, Changjing Shang, Qiang Shen 0001 |
Knowl. Based Syst. | 1 |
| 2021 | An Asynchronous WFST-Based Decoder for Automatic Speech RecognitionabstractWe introduce asynchronous dynamic decoder, which adopts an efficient A* algorithm to incorporate big language models in the one-pass decoding for large vocabulary continuous speech recognition. Unlike standard one-pass decoding with on-the-fly composition decoder which might induce a significant computation overhead, the asynchronous dynamic decoder has a novel design where it has two fronts, with one performing "exploration" and the other "backfill". The computation of the two fronts alternates in the decoding process, resulting in more effective pruning than the standard one-pass decoding with an on-the-fly composition decoder. Experiments show that the proposed decoder works notably faster than the standard one-pass decoding with on-the-fly composition decoder, while the acceleration will be more obvious with the increment of data complexity. Hang Lv 0001, Zhehuai Chen, Hainan Xu, Daniel Povey, Lei Xie 0001, Sanjeev Khudanpur |
ICASSP | 1 |
| 2021 | Wake Word Detection with Streaming TransformersabstractModern wake word detection systems usually rely on neural networks for acoustic modeling. Transformers has recently shown superior performance over LSTM and convolutional networks in various sequence modeling tasks with their better temporal modeling power. However it is not clear whether this advantage still holds for short-range temporal modeling like wake word detection. Besides, the vanilla Transformer is not directly applicable to the task due to its non-streaming nature and the quadratic time and space complexity. In this paper we explore the performance of several variants of chunk-wise streaming Transformers tailored for wake word detection in a recently proposed LF-MMI system, including looking-ahead to the next chunk, gradient stopping, different positional embedding methods and adding same-layer dependency between chunks. Our experiments on the Mobvoi wake word dataset demonstrate that our proposed Transformer model outperforms the baseline convolution network by 25% on average in false rejection rate at the same false alarm rate with a comparable model size, while still maintaining linear complexity w.r.t. the sequence length. Yiming Wang 0006, Hang Lv 0001, Daniel Povey, Lei Xie 0001, Sanjeev Khudanpur |
ICASSP | 2 |
| 2021 | LET-Decoder: A WFST-Based Lazy-Evaluation Token-Group Decoder With Exact Lattice GenerationabstractWe propose a novel lazy-evaluation token-group decoding algorithm with on-the-fly composition of weighted finite-state transducers (WFSTs) for large vocabulary continuous speech recognition. In the standard on-the-fly composition decoder, a base WFST and one or more incremental WFSTs are composed during decoding, and then token passing algorithm is employed to generate the lattice on the composed search space, resulting in substantial computation overhead. To improve speed, the proposed algorithm adopts 1) a token-group method, which groups tokens with the same state in the base WFST on each frame and limits the capacity of the group and 2) a lazy-evaluation method, which does not expand a token group and its source token groups until it processes a word label during decoding. Experiments show that the proposed decoder works notably up to 3 times faster than the standard on-the-fly composition decoder. Hang Lv 0001, Daniel Povey, Mahsa Yarmohammadi, Ke Li 0018, Yiming Wang 0006, Lei Xie 0001, Sanjeev Khudanpur |
IEEE Signal Process. Lett. | 1 |
| 2020 | Wake Word Detection with Alignment-Free Lattice-Free MMIabstractAlways-on spoken language interfaces, e.g. personal digital assistants, rely on a wake word to start processing spoken input. We present novel methods to train a hybrid DNN/HMM wake word detection system from partially labeled training data, and to use it in on-line applications: (i) we remove the prerequisite of frame-level alignments in the LF-MMI training algorithm, permitting the use of un-transcribed training examples that are annotated only for the presence/absence of the wake word; (ii) we show that the classical keyword/filler model must be supplemented with an explicit non-speech (silence) model for good performance; (iii) we present an FST-based decoder to perform online detection. We evaluate our methods on two real data sets, showing 50%--90% reduction in false rejection rates at pre-specified false alarm rates over the best previously published figures, and re-validate them on a third (large) data set. Yiming Wang 0006, Hang Lv 0001, Daniel Povey, Lei Xie 0001, Sanjeev Khudanpur |
INTERSPEECH | 2 |
| 2019 | Incremental Lattice Determinization for WFST DecodersabstractWe introduce a lattice determinization algorithm that can operate incrementally. That is, a word-level lattice can be generated for a partial utterance and then, once we have processed more audio, we can obtain a word-level lattice for the extended utterance without redoing all the work of lattice determinization. This is relevant for ASR decoders such as those used in Kaldi, which first generate a state-level lattice and then convert it to a word-level lattice using a determinization algorithm in a special semiring. Our incremental determinization algorithm is useful when word-level lattices are needed prior to the end of the utterance, and also reduces the latency due to determinization at the end of the utterance. Zhehuai Chen, Mahsa Yarmohammadi, Hainan Xu, Hang Lv 0001, Lei Xie 0001, Daniel Povey, Sanjeev Khudanpur |
ASRU | 4 |
| 2019 | Espresso: A Fast End-to-End Neural Speech Recognition ToolkitabstractWe present Espresso, an open-source, modular, extensible end-to-end neural automatic speech recognition (ASR) toolkit based on the deep learning library PyTorch and the popular neural machine translation toolkit FAIRSEQ. ESRESSO supports distributed training across GPUs and computing nodes, and features various decoding approaches commonly employed in ASR, including look-ahead word-based language model fusion, for which a fast, parallelized decoder is implemented. Espresso achieves state-of-the-art ASR performance on the WSJ, LibriSpeech, and Switchboard data sets among other end-to-end systems without data augmentation, and is 4-11x faster for decoding than similar systems (e.g. ESPNET). Yiming Wang 0006, Sanjeev Khudanpur, Tongfei Chen, Hainan Xu, Shuoyang Ding, Hang Lv 0001, Yiwen Shao, Nanyun Peng 0001, Lei Xie 0001, Shinji Watanabe 0001 |
ASRU | 6 |
| 2018 | Acoustic Modeling from Frequency Domain Representations of Speech
Pegah Ghahremani, Hossein Hadian, Hang Lv 0001, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 3 |
| 2016 | Approximate search of audio queries by using DTW with phone time boundary and data augmentationabstractDynamic Time Warping (DTW) is widely used in language independent query-by-example (QbE) spoken term detection (STD) tasks due to its high performance. However, there are two limitations of DTW based template matching, 1) it is not straightforward to perform approximate match of audio queries; 2) DTW is sensitive to the mismatch of signal conditions between the query and the speech search data. To allow approximate search, we propose a partial template matching strategy using phone time boundary information generated by a phone recognizer. To have more invariant representation of audio signals, we use bottleneck features (BNF) as the input of DTW. The BNF network is trained from augmented data, which is generated by adding reverberation and additive noises to the clean training data. Experimental results on QUESST 2015 task shows the effectiveness of the proposed methods for QbE-STD when the queries and search data are both distorted by reverberation and noises. Haihua Xu 0001, Jingyong Hou, Van Tung Pham, Cheung-Chi Leung, Lei Wang 0020, Van Hai Do, Hang Lv 0001, Lei Xie 0001, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 8 |
| 2016 | Toward High-Performance Language-Independent Query-by-Example Spoken Term Detection for MediaEval 2015: Post-Evaluation Analysis
Cheung-Chi Leung, Lei Wang 0020, Haihua Xu 0001, Jingyong Hou, Van Tung Pham, Hang Lv 0001, Lei Xie 0001, Chongjia Ni, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2015 | Language independent query-by-example spoken term detection using N-best phone sequences and partial matchingabstractIn this paper, we propose a partial sequence matching based symbolic search (SS) method for the task of language independent query-by-example spoken term detection. One main drawback of conventional SS approach is the high miss rate for long queries. This is due to high variations in symbol representation of query and search audios, especially in language independent scenario. The successful matching of a query with its instances in search audio becomes exponentially more difficult as the query grows longer. To reduce miss rate, we propose a partial matching strategy, in which all partial phone sequences of a query are used to search for query instances. The partial matching is also suitable for real life applications where exact match is usually not necessary and word prefix, suffix, and order should not affect the search result. When applied to the QUESST 2014 task, results show the partial matching of phone sequences is able to reduce miss rate of long queries significantly compared with conventional full matching method. In addition, for the most challenging inexact matching queries (type 3), it also shows clear advantage over DTW-based methods. Haihua Xu 0001, Lei Xie 0001, Cheung-Chi Leung, Hongjie Chen 0001, Jia Yu 0002, Hang Lv 0001, Lei Wang 0020, Su Jun Leow, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 8 |