VLDB 2026 Research / reviewers in the wild / expert
Jian Luo 0007
dblp:34/3128-7
· DBLP profile ↗
11ranked-venue papers
9as first author
10since 2021 · last 2023
0000-0002-9756-3066ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Dynamic Alignment Mask CTC: Improved Mask CTC With Aligned Cross EntropyabstractBecause of predicting all the target tokens in parallel, the non-autoregressive models greatly improve the decoding efficiency of speech recognition compared with traditional autoregressive models. In this work, we present dynamic alignment Mask CTC, introducing two methods: (1) Aligned Cross Entropy (AXE), finding the monotonic alignment that minimizes the cross-entropy loss through dynamic programming, (2) Dynamic Rectification, creating new training samples by replacing some masks with model predicted tokens. The AXE ignores the absolute position alignment between prediction and ground truth sentence and focuses on tokens matching in relative order. The dynamic rectification method makes the model capable of simulating the non-mask but possible wrong tokens, even if they have high confidence. Our experiments on WSJ dataset demonstrated that not only AXE loss but also the rectification method could improve the WER performance of Mask CTC. Xulong Zhang 0001, Haobin Tang, Jianzong Wang, Ning Cheng 0001, Jian Luo 0007, Jing Xiao 0006 |
ICASSP | 5 |
| 2022 | Speech Augmentation Based Unsupervised Learning for Keyword SpottingabstractIn this paper, we investigated a speech augmentation based unsupervised learning approach for keyword spotting (KWS) task. KWS is a useful speech application, yet also heavily depends on the labeled data. We designed a CNN-Attention architecture to conduct the KWS task. CNN layers focus on the local acoustic features, and attention layers model the long-time dependency. To improve the robustness of KWS model, we also proposed an unsupervised learning method. The unsupervised loss is based on the similarity between the original and augmented speech features, as well as the audio reconstructing information. Two speech augmentation methods are explored in the unsupervised learning: speed and intensity. The experiments on Google Speech Commands V2 Dataset demonstrated that our CNN-Attention model has competitive results. Moreover, the augmentation based unsupervised learning could further improve the classification accuracy of KWS task. In our experiments, with augmentation based unsupervised learning, our KWS model achieves better performance than other unsupervised methods, such as CPC, APC, and MPC. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Haobin Tang, Jing Xiao 0006 |
IJCNN | 1 |
| 2022 | Adaptive Activation Network for Low Resource Multilingual Speech RecognitionabstractLow resource automatic speech recognition (ASR) is a useful but thorny task, since deep learning ASR models usually need huge amounts of training data. The existing models mostly established a bottleneck (BN) layer by pre-training on a large source language, and transferring to the low resource target language. In this work, we introduced an adaptive activation network to the upper layers of ASR model, and applied different activation functions to different languages. We also proposed two approaches to train the model: (1) cross-lingual learning, replacing the activation function from source language to target language, (2) multilingual learning, jointly training the Connectionist Temporal Classification (CTC) loss of each language and the relevance of different languages. Our experiments on IARPA Babel datasets demonstrated that our approaches outperform the from-scratch training and traditional bottleneck feature based methods. In addition, combining the cross-lingual learning and multilingual learning together could further improve the performance of multilingual speech recognition. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Zhenpeng Zheng, Jing Xiao 0006 |
IJCNN | 1 |
| 2022 | Tiny-Sepformer: A Tiny Time-Domain Transformer Network For Speech SeparationabstractTime-domain Transformer neural networks have proven their superiority in speech separation tasks.However, these models usually have a large number of network parameters, thus often encountering the problem of GPU memory explosion.In this paper, we proposed Tiny-Sepformer, a tiny version of Transformer network for speech separation.We present two techniques to reduce the model parameters and memory consumption: (1) Convolution-Attention (CA) block, spliting the vanilla Transformer to two paths, multi-head attention and 1D depthwise separable convolution, (2) parameter sharing, sharing the layer parameters within the CA block.In our experiments, Tiny-Sepformer could greatly reduce the model size, and achieves comparable separation performance with vanilla Sepformer on WSJ0-2/3Mix datasets. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Edward Xiao, Xulong Zhang 0001, Jing Xiao 0006 |
INTERSPEECH | 1 |
| 2021 | Unidirectional Memory-Self-Attention Transducer for Online Speech RecognitionabstractSelf-attention models have been successfully applied in end-to-end speech recognition systems, which greatly improve the performance of recognition accuracy. However, such attention-based models cannot be used in online speech recognition, because these models usually have to utilize a whole acoustic sequences as inputs. A common method is restricting the field of attention sights by a fixed left and right window, which makes the computation costs manageable yet also introduces performance degradation. In this paper, we propose Memory-Self-Attention (MSA), which adds history information into the Restricted-Self-Attention unit. MSA only needs localtime features as inputs, and efficiently models long temporal contexts by attending memory states. Meanwhile, recurrent neural network transducer (RNN-T) has proved to be a great approach for online ASR tasks, because the alignments of RNN-T are local and monotonic. We propose a novel network structure, called Memory-Self-Attention (MSA) Transducer. Both encoder and decoder of the MSA Transducer contain the proposed MSA unit. The experiments demonstrate that our proposed models improve WER results than Restricted-Self-Attention models by 13.5% on WSJ and 7.1% on SWBD datasets relatively, and without much computation costs increase. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 1 |
| 2021 | Cross-Language Transfer Learning and Domain Adaptation for End-to-End Automatic Speech RecognitionabstractIn this paper, we demonstrate the efficacy of transfer learning and continuous learning for various automatic speech recognition (ASR) tasks using end-to-end models trained with CTC loss. We start with a large pre-trained English ASR model and show that transfer learning can be effectively and easily performed on: (1) different English accents, (2) different languages (from English to German, Spanish, Russian, or from Mandarin to Cantonese) and (3) application-specific domains. Our extensive set of experiments demonstrate that in all three cases, transfer learning from a good base model has higher accuracy than a model trained from scratch. Our results indicate that, for fine-tuning, larger pre-trained models are better than small pre-trained models, even if the dataset for fine-tuning is small. We also show that transfer learning significantly speeds up convergence, which could result in significant cost savings when training with large datasets. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Edward Xiao, Jing Xiao 0006, Georg Kucsko, Patrick K. O'Neill, Jagadeesh Balam, Slyne Deng, Adriana Flores, Boris Ginsburg, Jocelyn Huang, Oleksii Kuchaiev, Vitaly Lavrukhin, Jason Li 0007 |
ICME | 1 |
| 2021 | Loss Prediction: End-to-End Active Learning Approach For Speech RecognitionabstractEnd-to-end speech recognition systems usually require huge amounts of labeling resource, while annotating the speech data is complicated and expensive. Active learning is the solution by selecting the most valuable samples for annotation. In this paper, we proposed to use a predicted loss that estimates the uncertainty of the sample. The CTC (Connectionist Temporal Classification) and attention loss are informative for speech recognition since they are computed based on all decoding paths and alignments. We defined an end-to-end active learning pipeline, training an ASR/LP (Automatic Speech Recognition/Loss Prediction) joint model. The proposed approach was validated on an English and a Chinese speech recognition task. The experiments show that our approach achieves competitive results, outperforming random selection, least confidence, and estimated loss method. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 1 |
| 2021 | Dropout Regularization for Self-Supervised Learning of Transformer Encoder Speech RepresentationabstractPredicting the altered acoustic frames is an effective way of self-supervised learning for speech representation.However, it is challenging to prevent the pretrained model from overfitting.In this paper, we proposed to introduce two dropout regularization methods into the pretraining of transformer encoder: (1) attention dropout, (2) layer dropout.Both of the two dropout methods encourage the model to utilize global speech information, and avoid just copying local spectrum features when reconstructing the masked frames.We evaluated the proposed methods on phoneme classification and speaker recognition tasks.The experiments demonstrate that our dropout approaches achieve competitive results, and improve the performance of classification accuracy on downstream tasks. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
Interspeech | 1 |
| 2021 | Multi-Quartznet: Multi-Resolution Convolution for Speech Recognition with Multi-Layer Feature FusionabstractIn this paper, we propose an end-to-end speech recognition network based on Nvidia's previous QuartzNet [1] model. We try to promote the model performance, and design three components: (1) Multi-Resolution Convolution Module, re-places the original 1D time-channel separable convolution with multi-stream convolutions. Each stream has a unique dilated stride on convolutional operations. (2) Channel-Wise Attention Module, calculates the attention weight of each convolutional stream by spatial channel-wise pooling. (3) Multi-Layer Feature Fusion Module, reweights each convolutional block by global multi-layer feature maps. Our experiments demonstrate that Multi-QuartzNet model achieves CER 6.77% on AISHELL-1 data set, which outperforms original QuartzNet and is close to state-of-art result. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Guilin Jiang, Jing Xiao 0006 |
SLT | 1 |
| 2021 | End-To-End Silent Speech Recognition with Acoustic SensingabstractSilent speech interfaces (SSI) has been an exciting area of recent interest. In this paper, we present a non-invasive silent speech interface that uses inaudible acoustic signals to capture people's lip movements when they speak. We exploit the speaker and microphone of the smartphone to emit signals and listen to their reflections, respectively. The extracted phase features of these reflections are fed into the deep learning networks to recognize speech. And we also propose an end-to-end recognition framework, which combines the CNN and attention-based encoder-decoder network. Evaluation results on a limited vocabulary (54 sentences) yield word error rates of 8.4% in speaker-independent and environment-independent settings, and 8.1% for unseen sentence testing. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Guilin Jiang, Jing Xiao 0006 |
SLT | 1 |
| 2020 | MLNET: An Adaptive Multiple Receptive-Field Attention Neural Network for Voice Activity DetectionabstractVoice activity detection (VAD) makes a distinction between speech and non-speech and its performance is of crucial importance for speech based services.Recently, deep neural network (DNN)-based VADs have achieved better performance than conventional signal processing methods.The existed DNNbased models always handcrafted a fixed window to make use of the contextual speech information to improve the performance of VAD.However, the fixed window of contextual speech information can't handle various unpredictable noise environments and highlight the critical speech information to VAD task.In order to solve this problem, this paper proposed an adaptive multiple receptive-field attention neural network, called MLNET, to finish VAD task.The MLNET leveraged multi-branches to extract multiple contextual speech information and investigated an effective attention block to weight the most crucial parts of the context for final classification.Experiments in real-world scenarios demonstrated that the proposed MLNET-based model outperformed other baselines. Zhenpeng Zheng, Jianzong Wang, Ning Cheng 0001, Jian Luo 0007, Jing Xiao 0006 |
INTERSPEECH | 4 |