Zeyu Zhao 0004

dblp:183/9558-4 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0002-4070-2694ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Less Redraw, More Explore: Suggestion and Completion for Sketch-to-Image
abstract
Sketch-to-image systems let users transform simple line drawings into realistic images, but current workflows force users into tedious redraw-regenerate cycles that slow creative exploration. We introduce two complementary interaction techniques that reduce iteration friction: AutoSketch, which extends partial sketches through AI-driven completions (pre-generation support), and BackSketch, which transforms generated images back into editable sketches at multiple abstraction levels (post-generation support). In a study with 30 participants, the results indicate that both techniques can improve exploration and expressiveness compared to a baseline sketch-to-image system, while AutoSketch also can increase users’ sense of agency and co-creation with the AI. We contribute new evidence that shifting support before or after generation opens distinct pathways for balancing user control and system initiative. Together, our results establish pre- and post-generation assistance as a design space for co-creative sketch-to-image systems.
Zeyu Zhao 0004, Connor Rees, Gavin Bailey, Matt Jones 0001, Simon Robinson 0001, Jennifer Pearson 0001
CHI1
2025 Regarding the Existence of the Internal Language Model in CTC-Based E2E ASR
abstract
Some End-to-End (E2E) Automatic Speech Recognition (ASR) models, such as Attention-based Encoder-Decoder (AED) and Recurrent Neural Network Transducer (RNN-T) are known to have components that effectively act as internal language models (ILM), implicitly modelling the prior probability of the output sequence. However, the existence of an ILM in pure Connectionist Temporal Classification (CTC) ASR systems remains debated. In this paper, we investigate the existence and strength of an ILM in CTC systems. Since CTC posterior probabilities cannot be analytically factorised, we propose a novel empirical method to probe the ILM. After validating our method on a hybrid DNN model with various external language models, we apply it to CTC models trained under different conditions, examining the effects of training data, modelling units, and training or pre-training methods. Our results show no strong evidence of an ILM in CTC-based ASR systems, even with the largest training dataset in our experiments. However, we make the surprising finding that when a CTC encoder is jointly trained with an AED loss, an ILM emerges, even when only the CTC component is used in decoding.
Zeyu Zhao 0004, Peter Bell 0001
ICASSP1
2024 Advancing CTC Models for Better Speech Alignment: A Topological Approach
abstract
Automatic Speech Recognition (ASR) systems often face challenges in alignment quality, particularly with the Connectionist Temporal Classification (CTC) approach, which frequently results in a high number of blank frames, known as the “peaky” issue. In this study, we explore the impact of modifying ASR model topologies on alignment quality without compromising Word Error Rate (WER) performance. Our findings demonstrate that introducing additional states to the CTC topology significantly improves alignment quality and mitigates the peaky issue. Conversely, increasing the minimum traversal frame can degrade alignment quality in our specific settings. These insights emphasise the critical importance of topology design in balancing alignment accuracy and recognition performance in ASR systems.
Zeyu Zhao 0004, Peter Bell 0001
SLT1
2024 Open-Source Conversational AI with SpeechBrain 1.0
abstract
SpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more. It promotes transparency and replicability by releasing both the pre-trained models and the complete recipes of code and algorithms required for training them. This paper presents SpeechBrain 1.0, a significant milestone in the evolution of the toolkit, which now has over 200 recipes for speech, audio, and language processing tasks, and more than 100 models available on Hugging Face. SpeechBrain 1.0 introduces new technologies to support diverse learning modalities, Large Language Model (LLM) integration, and advanced decoding strategies, along with novel models, tasks, and modalities. It also includes a new benchmark repository, offering researchers a unified platform for evaluating models across diverse tasks.
Mirco Ravanelli, Titouan Parcollet, Adel Moumen, Sylvain de Langen, Cem Subakan, Peter Plantinga, Yingzhi Wang 0002, Pooneh Mousavi, Luca Della Libera, Artem Ploujnikov, Francesco Paissan, Davide Borra, Mohamed Salah Zaïem, Zeyu Zhao 0004, Shucong Zhang, Georgios Karakasidis, Sung-Lin Yeh, Pierre Champion, Aku Rouhe, Rudolf Braun, Florian Mai, Juan Zuluaga-Gomez, Seyed Mahed Mousavi, Andreas Nautsch, Xuechen Liu 0001, Sangeet Sagar, Jarod Duret, Salima Mdhaffar, Gaëlle Laperrière, Mickael Rouvier, Renato De Mori, Yannick Estève
J. Mach. Learn. Res.14
2023 ASR and Emotional Speech: A Word-Level Investigation of the Mutual Impact of Speech and Emotion Recognition
abstract
In Speech Emotion Recognition (SER), textual data is often used alongside audio signals to address their inherent variability. However, the reliance on human annotated text in most research hinders the development of practical SER systems. To overcome this challenge, we investigate how Automatic Speech Recognition (ASR) performs on emotional speech by analyzing the ASR performance on emotion corpora and examining the distribution of word errors and confidence scores in ASR transcripts to gain insight into how emotion affects ASR. We utilize four ASR systems, namely Kaldi ASR, wav2vec, Conformer, and Whisper, and three corpora: IEMOCAP, MOSI, and MELD to ensure generalizability. Additionally, we conduct text-based SER on ASR transcripts with increasing word error rates to investigate how ASR affects SER. The objective of this study is to uncover the relationship and mutual impact of ASR and SER, in order to facilitate ASR adaptation to emotional speech and the use of SER in real world.
Yuanchao Li, Zeyu Zhao 0004, Ondrej Klejch, Peter Bell 0001, Catherine Lai
INTERSPEECH2
2023 Regarding Topology and Variant Frame Rates for Differentiable WFST-based End-to-End ASR
abstract
End-to-end (E2E) Automatic Speech Recognition (ASR) has gained popularity in recent years, with most research focusing on designing novel neural network architectures, speech rep resentations, and loss functions. However, the importance of topology in E2E ASR has been largely neglected. There are many aspects of topology to consider; in this paper, we focus on the relationship between topologies’ minimum traversal time and output frame rate, the number of distinct states for each output unit, and the flexibility of alignments admitted. We ex amine several different topologies on two datasets: WSJ and Librispeech. Our experiments reveal that different frame rates have varying optimal topologies and that the commonly used Connectionist Temporal Classification (CTC) topology is not always optimal. Our findings suggest that the choice of topol ogy is an important consideration in the design of E2E ASR systems.
Zeyu Zhao 0004, Peter Bell 0001
INTERSPEECH1
2022 Investigating Sequence-Level Normalisation For CTC-Like End-to-End ASR
abstract
End-to-end Automatic Speech Recognition (E2E ASR) significantly simplifies the training process of an ASR model. Connectionist Temporal Classification (CTC) is one of the most popular methods for E2E ASR training. Implicitly, CTC has a unique topology which is very useful for sequence modelling. However, we find that by changing to another topology, we can make it even more effective. In this paper, we propose a new CTC-like method, for E2E ASR training, by modifying the topology of original CTC, so that the well-known abuse of the blank label in CTC can be resolved theoretically. As we change the topology, a normalisation term is necessary, which makes the form of the final loss function similar to Maximum Mutual Information (MMI); we hence name our method MMI-CTC. In addition to maximising the posterior probability of the target sequence, the normalisation enables models to explicitly minimise the probability of competing hypothesis at the word sequence level. Our experimental results show that MMI-CTC is more efficient than CTC, and that the normalisation is essential for sequence training.
Zeyu Zhao 0004, Peter Bell 0001
ICASSP1
2021 End-to-end keyword search system based on attention mechanism and energy scorer for low resource languages
Zeyu Zhao 0004, Weiqiang Zhang 0001
Neural Networks1
2020 End-to-End Keyword Search Based on Attention and Energy Scorer for Low Resource Languages
Zeyu Zhao 0004, Weiqiang Zhang 0001
INTERSPEECH1