EDBT 2026 Demo / reviewers in the wild / expert
Zhikai Zhou
dblp:119/1899
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Research on Visual Recognition and Localization Methods for Different Feature PartsabstractAimed at the problems of complex industrial site environment, difficult identification of assembly features and low positioning accuracy, a new visual identification and positioning method is proposed, which can well identify and locate parts and products with different characteristics. First, combined filtering is used in preprocessing to repair low-quality images. Then, edge detection and contour finding are performed on the preprocessed image. Finally, the internal and external parameters of the camera are calculated through the camera calibration, and the acquired pixel coordinates are converted into world coordinates. The results show that the combined filtering algorithm has good noise reduction effect and can remove image surface noise, which improves the recognition accuracy for the identification and positioning of different feature parts by the subsequent contour search method, and the algorithm can meet the requirements of the actual positioning accuracy of the assembly. Zhikai Zhou, Wenxian Tang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2022 | LatticeBART: Lattice-to-Lattice Pre-Training for Speech RecognitionabstractTo improve automatic speech recognition, increasing work has attempted to further fix the output of ASR systems with advanced sequence models. However, the output of ASR systems differs significantly from the input form of standard sequence models. To encompass richer information, the output of ASR systems is often a compact lattice structure containing multiple sentences. This mismatch in input form significantly limits sequence models’ ability. On the one hand, the widely used pre-trained models cannot directly input lattice structures and are therefore difficult to use for this task. On the other hand, the sparsity of the supervised training data forces the model to have the ability to learn from limited data. To address these problems, we propose LatticeBART, a model that decodes the sequence from the lattice in an end-to-end fashion and can use the pre-trained language models’ prior. In addition, this paper proposes the lattice-to-lattice pre-training method, which can be used when annotated data is missing, using easily generated lattice with the ASR system for training. The experimental results show that our model can effectively improve the output quality of the ASR system. Lingfeng Dai, Lu Chen 0002, Zhikai Zhou, Kai Yu 0004 |
ICASSP | 3 |
| 2022 | The Sjtu System For Multimodal Information Based Speech Processing Challenge 2021abstractThis paper describes the SJTU system for ICASSP Multi-modal Information based Speech Processing Challenge (MISP) 2021. To solve the speech recognition problem in real complex environments where time-synchronized near- and far-field signals are available for training an enhancement frontend. We build a joint system with speech enhancement frontend and speech recognition backend. These two modules are optimized jointly by both ASR and enhancement criteria. Audio-visual fusion is explored to further boost the ASR performance. ROVER and test time augmentation techniques are used to combine recognition results from multiple systems. The final system achieves Chinese character error rates (CCER) of 34.9% on dev set and 34.0% on test set, which achieved third place in the MISP challenge. The absolute CCER reduction compared with the official baseline system is 26.9% on dev set and 28.7% on test set. Wei Wang 0010, Xun Gong 0005, Zhikai Zhou, Chenda Li, Wangyou Zhang, Bing Han 0008, Yanmin Qian |
ICASSP | 4 |
| 2022 | Punctuation Prediction for Streaming On-Device Speech RecognitionabstractPunctuation prediction is essential for automatic speech recognition (ASR). Although many works have been proposed for punctuation prediction, the on-device scenarios are rarely discussed with an end-to-end ASR. The punctuation prediction task is often treated as a post-processing of ASR outputs, but the mismatch between natural language in training input and ASR hypotheses in testing is ignored. Besides, language models built with deep neural networks are too large for edge devices. In this paper, we discuss one-pass models for both ASR and punctuation prediction to replace the conventional two-pass post-processing pipeline. Then the joint ASR-punctuation model is proposed to utilize multi-task learning to decouple the recognition and punctuation on the ASR decoder. Experimental results show that the proposed joint model not only outperforms the traditional post-processing method with limited extra parameters, but also achieves better accuracy in comparison to the direct ASR modeling on transcripts with punctuation. Zhikai Zhou, Tian Tan 0002, Yanmin Qian |
ICASSP | 1 |
| 2022 | Exploring Effective Data Utilization for Low-Resource Speech RecognitionabstractAutomatic speech recognition (ASR) has suffered great performance degradation when facing low-resource languages with limited training data. In this work, we propose a series of training strategies to exploring more effective data utilization for low-resource speech recognition. In low-resource scenarios, multilingual pretraining is of great help for the above purpose. We exploit relationships among different languages for better pretraining. Then, the knowledge extracted from the language classifier is utilized for data weighing on training samples, making the model more biased towards the target low-resource language. Moreover, dynamic curriculum learning as a warm-up strategy and length perturbation as data augmentation are also designed. All these three methods form a newly improved training strategy for low-resource speech recognition. Meanwhile, we evaluate the proposed strategies using rich-resource languages for pretraining (PT) and finetuning (FT) the model on the target language with limited data. The experimental results show that on the CommonVoice dataset, compared with the commonly used multilingual PT+FT method, the proposed strategies achieve a relative 15-25% reduction in word error rate on different target languages, which shows the significant effects of the proposed data utilization strategy. Zhikai Zhou, Wei Wang 0010, Wangyou Zhang, Yanmin Qian |
ICASSP | 1 |
| 2022 | Knowledge Transfer and Distillation from Autoregressive to Non-Autoregessive Speech Recognition
Xun Gong 0005, Zhikai Zhou, Yanmin Qian |
INTERSPEECH | 2 |
| 2022 | Optimizing Data Usage for Low-Resource Speech RecognitionabstractAutomatic speech recognition has made huge progress recently. However, the current modeling strategy still suffers a large performance degradation when facing the low-resource languages with limited training data. In this paper, we propose a series of methods to optimize the data usage for low-resource speech recognition. Multilingual speech recognition helps a lot in low-resource scenarios. The correlation and similarity between languages are further exploited for multilingual pretraining in our work. We utilize the posterior of the target language extracted from a language classifier to perform data weighing on training samples, which assists the model in being more biased towards the target language during pretraining. Furthermore, dynamic curriculum learning for data allocation and length perturbation for data augmentation are also designed. All these three methods form the new strategy on optimized data usage for low-resource languages. We evaluate the proposed method using rich resource languages for pretraining (PT) and finetuning (FT) the model on the target language with limited data. Experimental results show that the proposed data usage method obtains a 15 to 25% relative word error rate reduction for different target languages compared with the commonly adopted multilingual PT+FT method on CommonVoice dataset. The same improvement and conclusion are also observed on Babel dataset with conversational telephone speech, and$\sim$40% relative character error rate reduction can be obtained for the target low-resource language. Yanmin Qian, Zhikai Zhou |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Towards Data Selection on TTS Data for Children's Speech RecognitionabstractAlthough great progress has been made on automatic speech recognition (ASR) systems, children’s speech recognition still remains a challenging task. General ASR systems for children’s speech suffer from the lack of corpora and mismatch between children’s and adults’ speech. Efforts have been made to reduce such mismatch by applying normalization methods to generate modified adults’ speech for ASR training. However, modified adults’ data can reflect the characteristics of children’s speech to a very limited extent. In this work, we adopt text-to-speech data augmentation to improve the performance of children’s speech recognition system. We find that the children’s TTS model generates speech with inconsistent quality due to children’s substandard pronunciations of phonemes, and the ASR system suffers when trained with these additional synthesized data. To solve this problem, we propose data selection strategies on the TTS augmented data, and the effectiveness of the synthesized data can be substantially boosted for children’s ASR modeling. We show that the speaker embedding similarity based data selection strategy can obtain the best position: relative 14.0% and 14.7% CER reduction for child conversation and child reading test set respectively compared to the baseline model trained on real data. Wei Wang 0010, Zhikai Zhou, Yizhou Lu, Chenpeng Du, Yanmin Qian |
ICASSP | 2 |
| 2021 | Layer-Wise Fast Adaptation for End-to-End Multi-Accent Speech RecognitionabstractAccent variability has posed a huge challenge to automatic speech recognition~(ASR) modeling. Although one-hot accent vector based adaptation systems are commonly used, they require prior knowledge about the target accent and cannot handle unseen accents. Furthermore, simply concatenating accent embeddings does not make good use of accent knowledge, which has limited improvements. In this work, we aim to tackle these problems with a novel layer-wise adaptation structure injected into the E2E ASR model encoder. The adapter layer encodes an arbitrary accent in the accent space and assists the ASR model in recognizing accented speech. Given an utterance, the adaptation structure extracts the corresponding accent information and transforms the input acoustic feature into an accent-related feature through the linear combination of all accent bases. We further explore the injection position of the adaptation layer, the number of accent bases, and different types of accent bases to achieve better accent adaptation. Experimental results show that the proposed adaptation structure brings 12\% and 10\% relative word error rate~(WER) reduction on the AESRC2020 accent dataset and the Librispeech dataset, respectively, compared to the baseline. Xun Gong 0005, Yizhou Lu, Zhikai Zhou, Yanmin Qian |
Interspeech | 3 |
| 2021 | The SJTU System for Short-Duration Speaker Verification Challenge 2021abstractThis paper presents the SJTU system for both text-dependent and text-independent tasks in short-duration speaker verification (SdSV) challenge 2021.In this challenge, we explored different strong embedding extractors to extract robust speaker embedding.For text-independent task, language-dependent adaptive snorm is explored to improve the system performance under the cross-lingual verification condition.For text-dependent task, we mainly focus on the in-domain fine-tuning strategies based on the model pre-trained on large-scale out-of-domain data.In order to improve the distinction between different speakers uttering the same phrase, we proposed several novel phrase-aware fine-tuning strategies and phrase-aware neural PLDA.With such strategies, the system performance is further improved.Finally, we fused the scores of different systems, and our fusion systems achieved 0.0473 in Task1 (rank 3) and 0.0581 in Task2 (rank 8) on the primary evaluation metric. Bing Han 0008, Zhengyang Chen, Zhikai Zhou, Yanmin Qian |
Interspeech | 3 |