EDBT 2026 Demo / reviewers in the wild / expert
Zhengkun Tian
dblp:244/9446
· DBLP profile ↗
25ranked-venue papers
6as first author
17since 2021 · last 2024
0000-0002-0469-3049ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 18 · 3 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MSR-86K: An Evolving, Multilingual Corpus with 86, 300 Hours of Transcribed Audio for Speech Recognition Research
Yongbin You, Xuezhi Wang 0008, Zhengkun Tian, Guanglu Wan |
INTERSPEECH | 4 |
| 2024 | SceneFake: An initial dataset and benchmarks for scene fake audio detection
Jiangyan Yi, Chenglong Wang 0001, Jianhua Tao 0001, Chuyuan Zhang, Cunhang Fan, Zhengkun Tian, Haoxin Ma, Ruibo Fu |
Pattern Recognit. | 6 |
| 2023 | Peak-First CTC: Reducing the Peak Latency of CTC Models by Applying Peak-First RegularizationabstractThe CTC model has been widely applied to many application scenarios because of its simple structure, excellent performance, and fast inference speed. There are many peaks in the probability distribution predicted by the CTC models, and each peak represents a non-blank token. The recognition latency of CTC models can be reduced by encouraging the model to predict peaks earlier. Existing methods to reduce latency require modifying the transition relationship between tokens in the forward-backward algorithm, and the gradient calculation. Some of these methods even depend on the forced alignment results provided by other pretrained models. The above methods are complex to implement. To reduce the peak latency, we propose a simple and novel method named peak-first regularization, which utilizes a frame-wise knowledge distillation function to force the probability distribution of the CTC model to shift left along the time axis instead of directly modifying the calculation process of CTC loss and gradients. All the experiments are conducted on a Chinese Mandarin dataset AISHELL-1. We have verified the effectiveness of the proposed regularization on both streaming and non-streaming CTC models respectively. The results show that the proposed method can reduce the average peak latency by about 100 to 200 milliseconds with almost no degradation of recognition accuracy. Zhengkun Tian, Hongyu Xiang, Feifei Lin, Guanglu Wan |
ICASSP | 1 |
| 2023 | Transfer knowledge for punctuation prediction via adversarial training
Jiangyan Yi, Jianhua Tao 0001, Ye Bai 0001, Zhengkun Tian, Cunhang Fan |
Speech Commun. | 4 |
| 2022 | End-to-End Network Based on Transformer for Automatic Detection of Covid-19abstractThe novel coronavirus disease (COVID-19) was declared a pandemic by the World Health Organization. The cumulative number of deaths is more than 4.8 million. Epidemiology experts concur that mass testing is essential for isolating infected individuals, contact tracing, and slowing the progression of the virus. In recent months, some machine learning methods have been proposed utilizing audio cues for COVID-19 detection. However, many works are based on hand-crafted features and deep features to detect COVID-19. There is no evidence that these features are optimal for COVID-19 detection. Therefore, we proposed an end-to-end network based on transformer for automatic detection of COVID-19. It directly learns features from the raw waveform for end-to-end learning, rather than extracting features in advance. We propose a feature extraction module to automatically extract features. And we use the transformer architectures to model the dependencies between the extracted features. It is the first end-to-end learning based on raw waveform for COVID-19 detection. Experiments on COUGHVID dataset show that our method has achieved competitive results. Cong Cai, Bin Liu 0041, Jianhua Tao 0001, Zhengkun Tian |
ICASSP | 4 |
| 2022 | ADD 2022: the first Audio Deep Synthesis Detection ChallengeabstractAudio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three tracks: low-quality fake audio detection (LF), partially fake audio detection (PF) and audio fake game (FG). The LF track focuses on dealing with bona fide and fully fake utterances with various real-world noises etc. The PF track aims to distinguish the partially fake audio from the real. The FG track is a rivalry game, which includes two tasks: an audio generation task and an audio fake detection task. In this paper, we describe the datasets, evaluation metrics, and protocols. We also report major findings that reflect the recent advances in audio deepfake detection tasks. Jiangyan Yi, Ruibo Fu, Jianhua Tao 0001, Shuai Nie 0001, Haoxin Ma, Chenglong Wang 0001, Tao Wang 0074, Zhengkun Tian, Ye Bai 0001, Cunhang Fan, Shan Liang 0007, Shuai Zhang 0014, Xinrui Yan, Zhengqi Wen, Haizhou Li 0001 |
ICASSP | 8 |
| 2022 | reducing multilingual context confusion for end-to-end code-switching automatic speech recognition
Shuai Zhang 0014, Jiangyan Yi, Zhengkun Tian, Jianhua Tao 0001, Yu Ting Yeung, Liqun Deng |
INTERSPEECH | 3 |
| 2022 | Hybrid Autoregressive and Non-Autoregressive Transformer Models for Speech RecognitionabstractThe autoregressive (AR) models, such as attention-based encoder-decoder models and RNN-Transducer, have achieved great success in speech recognition. They predict the output sequence conditioned on the previous tokens and acoustic encoded states, which is inefficient on GPUs. The non-autoregressive (NAR) models can get rid of the temporal dependency between the output tokens and predict the entire output tokens in one inference step. However, the NAR model still faces two major problems. Firstly, there is still a great gap in performance between the NAR models and the advanced AR models. Secondly, it’s difficult for most of the NAR models to train and converge. We propose a hybrid autoregressive and non-autoregressive transformer (HANAT) model, which integrates AR and NAR models deeply by sharing parameters. We assume that the AR model will assist the NAR model to learn some linguistic dependencies and accelerate the convergence. Furthermore, the two-stage hybrid inference is applied to improve the model performance. All the experiments are conducted on a mandarin dataset ASIEHLL-1 and a english dataset librispeech-960 h. The results show that the HANAT can achieve a competitive performance with the AR model and outperform many complicated NAR models. Besides, the RTF is only 1/5 of the AR model. Zhengkun Tian, Jiangyan Yi, Jianhua Tao 0001, Shuai Zhang 0014, Zhengqi Wen |
IEEE Signal Process. Lett. | 1 |
| 2021 | A Large-Scale Chinese Multimodal NER Dataset with Speech CluesabstractDianbo Sui, Zhengkun Tian, Yubo Chen, Kang Liu, Jun Zhao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Dianbo Sui, Zhengkun Tian, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | Decoupling Pronunciation and Language for End-to-End Code-Switching Automatic Speech RecognitionabstractDespite the recent significant advances witnessed in end-to-end (E2E) ASR system for code-switching, hunger for audio-text paired data limits the further improvement of the models’ performance. In this paper, we propose a decoupled transformer model to use mono-lingual paired data and unpaired text data to alleviate the problem of code-switching data shortage. The model is decoupled into two parts: audio-to-phoneme (A2P) network and phoneme-to-text (P2T) network. The A2P network can learn acoustic pattern scenarios using large-scale monolingual paired data. Meanwhile, it generates multiple phoneme sequence candidates for single audio data in real time during the training process. Then the generated phoneme-text paired data is used to train the P2T network. This network can be pre-trained with large amounts of external unpaired text data. By using monolingual data and unpaired text data, the decoupled transformer model reduces the high dependency on code-switching paired training data of E2E model to a certain extent. Finally, the two networks are optimized jointly through attention fusion. We evaluate the proposed method on the public Mandarin-English code-switching dataset. Compared with our transformer baseline, the proposed method achieves 18.14% relative mix error rate reduction. Shuai Zhang 0014, Jiangyan Yi, Zhengkun Tian, Ye Bai 0001, Jianhua Tao 0001, Zhengqi Wen |
ICASSP | 3 |
| 2021 | End-to-End Spelling Correction Conditioned on Acoustic Feature for Code-Switching Speech Recognition
Shuai Zhang 0014, Jiangyan Yi, Zhengkun Tian, Ye Bai 0001, Jianhua Tao 0001, Xuefei Liu, Zhengqi Wen |
Interspeech | 3 |
| 2021 | Continual Learning for Fake Audio DetectionabstractFake audio attack becomes a major threat to the speaker verification system. Although current detection approaches have achieved promising results on dataset-specific scenarios, they encounter difficulties on unseen spoofing data. Fine-tuning and retraining from scratch have been applied to incorporate new data. However, fine-tuning leads to performance degradation on previous data. Retraining takes a lot of time and computation resources. Besides, previous data are unavailable due to privacy in some situations. To solve the above problems, this paper proposes detecting fake without forgetting, a continual-learning-based method, to make the model learn new spoofing attacks incrementally. A knowledge distillation loss is introduced to loss function to preserve the memory of original model. Supposing the distribution of genuine voice is consistent among different scenarios, an extra embedding similarity loss is used as another constraint to further do a positive sample alignment. Experiments are conducted on the ASVspoof2019 dataset. The results show that our proposed method outperforms fine-tuning by the relative reduction of average equal error rate up to 81.62%. Haoxin Ma, Jiangyan Yi, Jianhua Tao 0001, Ye Bai 0001, Zhengkun Tian, Chenglong Wang 0001 |
Interspeech | 5 |
| 2021 | FSR: Accelerating the Inference Process of Transducer-Based Models by Applying Fast-Skip RegularizationabstractTransducer-based models, such as RNN-Transducer and transformer-transducer, have achieved great success in speech recognition. A typical transducer model decodes the output sequence conditioned on the current acoustic state and previously predicted tokens step by step. Statistically, The number of blank tokens in the prediction results accounts for nearly 90\% of all tokens. It takes a lot of computation and time to predict the blank tokens, but only the non-blank tokens will appear in the final output sequence. Therefore, we propose a method named fast-skip regularization, which tries to align the blank position predicted by a transducer with that predicted by a CTC model. During the inference, the transducer model can predict the blank tokens in advance by a simple CTC project layer without many complicated forward calculations of the transducer decoder and then skip them, which will reduce the computation and improve the inference speed greatly. All experiments are conducted on a public Chinese mandarin dataset AISHELL-1. The results show that the fast-skip regularization can indeed help the transducer model learn the blank position alignments. Besides, the inference with fast-skip can be speeded up nearly 4 times with only a little performance degradation. Zhengkun Tian, Jiangyan Yi, Ye Bai 0001, Jianhua Tao 0001, Shuai Zhang 0014, Zhengqi Wen |
Interspeech | 1 |
| 2021 | Half-Truth: A Partially Fake Audio Detection DatasetabstractDiverse promising datasets have been designed to further the development of fake audio detection, such as ASVspoof databases.However, previous datasets ignore an attacking situation, in which the hacker hides some small fake clips in real speech audio.This poses a serious threat since that it is difficult to distinguish the small fake clip from the whole speech utterance.Therefore, this paper develops such a dataset for half-truth audio detection (HAD).Partially fake audio in the HAD dataset involves only changing a few words in an utterance.The audio of the words is generated with the very latest state-of-the-art speech synthesis technology.We can not only detect fake uttrances but also localize manipulated regions in a speech using this dataset.Some benchmark results are presented on this dataset.The results show that partially fake audio presents much more challenging than fully fake audio for fake audio detection.The HAD dataset is publicly available 1 . Jiangyan Yi, Ye Bai 0001, Jianhua Tao 0001, Haoxin Ma, Zhengkun Tian, Chenglong Wang 0001, Tao Wang 0074, Ruibo Fu |
Interspeech | 5 |
| 2021 | Fast End-to-End Speech Recognition Via Non-Autoregressive Models and Cross-Modal Knowledge Transferring From BERTabstractAttention-based encoder-decoder (AED) models have achieved promising performance in speech recognition. However, because the decoder predicts text tokens (such as characters or words) in an autoregressive manner, it is difficult for an AED model to predict all tokens in parallel. This makes the inference speed relatively slow. In contrast, we propose an end-to-end non-autoregressive speech recognition model called LASO (Listen Attentively, and Spell Once). The model aggregates encoded speech features into the hidden representations corresponding to each token with attention mechanisms. Thus, the model can capture the token relations by self-attention on the aggregated hidden representations from the whole speech signal rather than autoregressive modeling on tokens. Without explicitly autoregressive language modeling, this model predicts all tokens in the sequence in parallel so that the inference is efficient. Moreover, we propose a cross-modal transfer learning method to use a text-modal language model to improve the performance of speech-modal LASO by aligning token semantics. We conduct experiments on two scales of public Chinese speech datasets AISHELL-1 and AISHELL-2. Experimental results show that our proposed model achieves a speedup of about 50× and competitive performance, compared with the autoregressive transformer models. And the cross-modal knowledge transferring from the text-modal model can improve the performance of the speech-modal model. Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Zhengqi Wen, Shuai Zhang 0014 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Integrating Knowledge Into End-to-End Speech Recognition From External Text-Only DataabstractAttention-based encoder-decoder (AED) models have achieved promising performance in speech recognition. However, because of the end-to-end training, an AED model is usually trained with speech-text paired data. It is challenging to incorporate external text-only data into AED models. Another issue of the AED model is that it does not use the right context of a text token while predicting the token. To alleviate the above two issues, we propose a unified method called LST (Learn Spelling from Teachers) to integrate knowledge into an AED model from the external text-only data and leverage the whole context in a sentence. The method is divided into two stages. First, in the representation stage, a language model is trained on the text. It can be seen as that the knowledge in the text is compressed into the LM. Then, at the transferring stage, the knowledge is transferred to the AED model via teacher-student learning. To further use the whole context of the text sentence, we propose an LM called causal cloze completer (COR), which estimates the probability of a token, given both the left context and the right context of it. Therefore, with LST training, the AED model can leverage the whole context in the sentence. Different from fusion based methods, which use LM during decoding, the proposed method does not increase any extra complexity at the inference stage. We conduct experiments on two scales of public Chinese datasets AISHELL-1 and AISHELL-2. The experimental results demonstrate the effectiveness of leveraging external text-only data and the whole context in a sentence with our proposed method, compared with baseline hybrid systems and AED model based systems. Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen, Zhengkun Tian, Shuai Zhang 0014 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Gated Recurrent Fusion With Joint Training Framework for Robust End-to-End Speech RecognitionabstractThe joint training framework for speech enhancement and recognition methods have obtained quite good performances for robust end-to-end automatic speech recognition (ASR). However, these methods only utilize the enhanced feature as the input of the speech recognition component, which are affected by the speech distortion problem. In order to address this problem, this paper proposes a gated recurrent fusion (GRF) method with joint training framework for robust end-to-end ASR. The GRF algorithm is used to dynamically combine the noisy and enhanced features. Therefore, the GRF can not only remove the noise signals from the enhanced features, but also learn the raw fine structures from the noisy features so that it can alleviate the speech distortion. The proposed method consists of speech enhancement, GRF and speech recognition. Firstly, the mask based speech enhancement network is applied to enhance the input speech. Secondly, the GRF is applied to address the speech distortion problem. Thirdly, to improve the performance of ASR, the state-of-the-art speech transformer algorithm is used as the speech recognition component. Finally, the joint training framework is utilized to optimize these three components, simultaneously. Our experiments are conducted on an open-source Mandarin speech corpus called AISHELL-1. Experimental results show that the proposed method achieves the relative character error rate (CER) reduction of 10.04% over the conventional joint enhancement and transformer method only using the enhanced features. Especially for the low signal-to-noise ratio (0 dB), our proposed method can achieves better performances with 12.67% CER reduction, which suggests the potential of our proposed method. Cunhang Fan, Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Bin Liu 0041, Zhengqi Wen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Synchronous Transformers for end-to-end Speech RecognitionabstractFor most of the attention-based sequence-to-sequence models, the decoder predicts the output sequence conditioned on the entire input sequence processed by the encoder. The asynchronous problem between the encoding and decoding makes these models difficult to be applied for online speech recognition. In this paper, we propose a model named synchronous transformer to address this problem, which can predict the output sequence chunk by chunk. Once a fixed-length chunk of the input sequence is processed by the encoder, the decoder begins to predict symbols immediately. During training, a forward-backward algorithm is introduced to optimize all the possible alignment paths. Our model is evaluated on a Mandarin dataset AISHELL-1. The experiments show that the synchronous transformer is able to perform encoding and decoding synchronously, and achieves a character error rate of 8.91% on the test set. Zhengkun Tian, Jiangyan Yi, Ye Bai 0001, Jianhua Tao 0001, Shuai Zhang 0014, Zhengqi Wen |
ICASSP | 1 |
| 2020 | Listen Attentively, and Spell Once: Whole Sentence Generation via a Non-Autoregressive Architecture for Low-Latency Speech RecognitionabstractAlthough attention based end-to-end models have achieved promising performance in speech recognition, the multi-pass forward computation in beam-search increases inference time cost, which limits their practical applications.To address this issue, we propose a non-autoregressive end-to-end speech recognition system called LASO (listen attentively, and spell once).Because of the non-autoregressive property, LASO predicts a textual token in the sequence without the dependence on other tokens.Without beam-search, the one-pass propagation much reduces inference time cost of LASO.And because the model is based on the attention based feedforward structure, the computation can be implemented in parallel efficiently.We conduct experiments on publicly available Chinese dataset AISHELL-1.LASO achieves a character error rate of 6.4%, which outperforms the state-of-the-art autoregressive transformer model (6.7%).The average inference latency is 21 ms, which is 1/50 of the autoregressive transformer model. Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Zhengqi Wen, Shuai Zhang 0014 |
INTERSPEECH | 4 |
| 2020 | Spike-Triggered Non-Autoregressive Transformer for End-to-End Speech RecognitionabstractNon-autoregressive transformer models have achieved extremely fast inference speed and comparable performance with autoregressive sequence-to-sequence models in neural machine translation.Most of the non-autoregressive transformers decode the target sequence from a predefined-length mask sequence.If the predefined length is too long, it will cause a lot of redundant calculations.If the predefined length is shorter than the length of the target sequence, it will hurt the performance of the model.To address this problem and improve the inference speed, we propose a spike-triggered non-autoregressive transformer model for end-to-end speech recognition, which introduces a CTC module to predict the length of the target sequence and accelerate the convergence.All the experiments are conducted on a public Chinese mandarin dataset AISHELL-1.The results show that the proposed model can accurately predict the length of the target sequence and achieve a competitive performance with the advanced transformers.What's more, the model even achieves a real-time factor of 0.0056, which exceeds all mainstream speech recognition models. Zhengkun Tian, Jiangyan Yi, Jianhua Tao 0001, Ye Bai 0001, Shuai Zhang 0014, Zhengqi Wen |
INTERSPEECH | 1 |
| 2020 | Focal Loss for Punctuation Prediction
Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Ye Bai 0001, Cunhang Fan |
INTERSPEECH | 3 |
| 2020 | Deep imitator: Handwriting calligraphy imitation via deep attention networks
Bocheng Zhao, Jianhua Tao 0001, Zhengkun Tian, Cunhang Fan, Ye Bai 0001 |
Pattern Recognit. | 4 |
| 2019 | Learn Spelling from Teachers: Transferring Knowledge from Language Models to Sequence-to-Sequence Speech RecognitionabstractIntegrating an external language model into a sequence-tosequence speech recognition system is non-trivial.Previous works utilize linear interpolation or a fusion network to integrate external language models.However, these approaches introduce external components, and increase decoding computation.In this paper, we instead propose a knowledge distillation based training approach to integrating external language models into a sequence-to-sequence model.A recurrent neural network language model, which is trained on large scale external text, generates soft labels to guide the sequence-to-sequence model training.Thus, the language model plays the role of the teacher.This approach does not add any external component to the sequence-to-sequence model during testing.And this approach is flexible to be combined with shallow fusion technique together for decoding.The experiments are conducted on public Chinese datasets AISHELL-1 and CLMAD.Our approach achieves a character error rate of 9.3%, which is relatively reduced by 18.42% compared with the vanilla sequenceto-sequence model. Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Zhengqi Wen |
INTERSPEECH | 4 |
| 2019 | A Time Delay Neural Network with Shared Weight Self-Attention for Small-Footprint Keyword Spotting
Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen, Zhengkun Tian, Chenghao Zhao, Cunhang Fan |
INTERSPEECH | 5 |
| 2019 | Self-Attention Transducers for End-to-End Speech RecognitionabstractRecurrent neural network transducers (RNN-T) have been successfully applied in end-to-end speech recognition. However, the recurrent structure makes it difficult for parallelization . In this paper, we propose a self-attention transducer (SA-T) for speech recognition. RNNs are replaced with self-attention blocks, which are powerful to model long-term dependencies inside sequences and able to be efficiently parallelized. Furthermore, a path-aware regularization is proposed to assist SA-T to learn alignments and improve the performance. Additionally, a chunk-flow mechanism is utilized to achieve online decoding. All experiments are conducted on a Mandarin Chinese dataset AISHELL-1. The results demonstrate that our proposed approach achieves a 21.3% relative reduction in character error rate compared with the baseline RNN-T. In addition, the SA-T with chunk-flow mechanism can perform online decoding with only a little degradation of the performance. Zhengkun Tian, Jiangyan Yi, Jianhua Tao 0001, Ye Bai 0001, Zhengqi Wen |
INTERSPEECH | 1 |