EDBT 2026 Demo / reviewers in the wild / expert
Julian Chan
dblp:158/4295
· DBLP profile ↗
11ranked-venue papers
0as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 8 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised DisentanglementabstractThe imitation of voice, targeted on specific speech attributes such as timbre and speaking style, is crucial in speech generation. However, existing methods rely heavily on annotated data, and struggle with effectively disentangling timbre and style, leading to challenges in achieving controllable generation, especially in zero-shot scenarios. To address these issues, we propose Vevo, a versatile zero-shot voice imitation framework with controllable timbre and style. Vevo operates in two core stages: (1) Content-Style Modeling: Given either text or speech's content tokens as input, we utilize an autoregressive transformer to generate the content-style tokens, which is prompted by a style reference; (2) Acoustic Modeling: Given the content-style tokens as input, we employ a flow-matching transformer to produce acoustic representations, which is prompted by a timbre reference. To obtain the content and content-style tokens of speech, we design a fully self-supervised approach that progressively decouples the timbre, style, and linguistic content of speech. Specifically, we adopt VQ-VAE as the tokenizer for the continuous hidden features of HuBERT. We treat the vocabulary size of the VQ-VAE codebook as the information bottleneck, and adjust it carefully to obtain the disentangled speech representations. Solely self-supervised trained on 60K hours of audiobook speech data, without any fine-tuning on style-specific corpora, Vevo matches or surpasses existing methods in accent and emotion conversion tasks. Additionally, Vevo’s effectiveness in zero-shot voice conversion and text-to-speech tasks further demonstrates its strong generalization and versatility. Audio samples are available at https://versavoice.github.io/. Xueyao Zhang, Kainan Peng, Vimal Manohar, Yingru Liu, Jeff Hwang, Dangna Li, Julian Chan, Zhizheng Wu 0001, Mingbo Ma |
ICLR | 10 |
| 2021 | On Lattice-Free Boosted MMI Training of HMM and CTC-Based Full-Context ASR ModelsabstractHybrid automatic speech recognition (ASR) models are typically sequentially trained with CTC or LF-MMI criteria. However, they have vastly different legacies and are usually implemented in different frameworks. In this paper, by decoupling the concepts of modeling units and label topologies and building proper numerator/denominator graphs accordingly, we establish a generalized framework for hybrid acoustic modeling (AM). In this framework, we show that LF-MMI is a powerful training criterion applicable to both limited-context and full-context models, for wordpiece/mono-char/bi-char/chenone units, with both HMM/CTC topologies. From this framework, we propose three novel training schemes: chenone(ch)/wordpiece(wp)-CTC-bMMI, and wordpiece(wp)-HMM-bMMI with different advantages in training performance, decoding efficiency and decoding time-stamp accuracy. The advantages of different training schemes are evaluated comprehensively on Librispeech, and wp-CTC-bMMI and ch-CTC-bMMI are evaluated on two real world ASR tasks to show their effectiveness. Besides, we also show bi-char(bc) HMM-MMI models can serve as better alignment models than traditional non-neural GMM-HMMs. Xiaohui Zhang 0007, Vimal Manohar, Frank Zhang 0001, Yangyang Shi, Nayan Singhal, Julian Chan, Fuchun Peng, Yatharth Saraf, Mike Seltzer |
ASRU | 7 |
| 2021 | Emformer: Efficient Memory Transformer Based Acoustic Model for Low Latency Streaming Speech RecognitionabstractThis paper proposes an efficient memory transformer Emformer for low latency streaming speech recognition. In Emformer, the long-range history context is distilled into an augmented memory bank to reduce self-attention’s computation complexity. A cache mechanism saves the computation for the key and value in self-attention for the left context. Emformer applies a parallelized block processing in training to support low latency models. We carry out experiments on benchmark LibriSpeech data. Under average latency of 960 ms, Emformer gets WER 2.50% on test-clean and 5.62% on test-other. Comparing with a strong baseline augmented memory transformer (AM-TRF), Emformer gets 4.6 folds training speedup and 18% relative real-time factor (RTF) reduction in decoding with relative WER reduction 17% on test-clean and 9% on test-other. For a low latency scenario with an average latency of 80 ms, Emformer achieves WER 3.01% on test-clean and 7.09% on test-other. Comparing with the LSTM baseline with the same latency and model size, Emformer gets relative WER reduction 9% and 16% on test-clean and test-other, respectively. Yangyang Shi, Yongqiang Wang 0005, Chunyang Wu, Ching-Feng Yeh, Julian Chan, Frank Zhang 0001, Mike Seltzer |
ICASSP | 5 |
| 2021 | Transformer in Action: A Comparative Study of Transformer-Based Acoustic Models for Large Scale Speech Recognition ApplicationsabstractTransformer-based acoustic models have shown promising results very recently. In this paper, we summarize the application of transformer and its streamable variant, Emformer based acoustic model [1] for large scale speech recognition applications. We compare the transformer based acoustic models with their LSTM counterparts on industrial scale tasks. Specifically, we compare Emformer with latency-controlled BLSTM (LCBLSTM) on medium latency tasks and LSTM on low latency tasks. On a low latency voice assistant task, Emformer gets 24% to 26% relative word error rate reductions (WERRs). For medium latency scenarios, comparing with LCBLSTM with similar model size and latency, Emformer gets significant WERR across four languages in video captioning datasets with 2-3 times inference real-time factors reduction. Yongqiang Wang 0005, Yangyang Shi, Frank Zhang 0001, Chunyang Wu, Julian Chan, Ching-Feng Yeh, Alex Xiao |
ICASSP | 5 |
| 2021 | Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow FusionabstractHow to leverage dynamic contextual information in end-toend speech recognition has remained an active research area.Previous solutions to this problem were either designed for specialized use cases that did not generalize well to open-domain scenarios, did not scale to large biasing lists, or underperformed on rare long-tail words.We address these limitations by proposing a novel solution that combines shallow fusion, trie-based deep biasing, and neural network language model contextualization.These techniques result in significant 19.5% relative Word Error Rate improvement over existing contextual biasing approaches and 5.4%-9.3%improvement compared to a strong hybrid baseline on both open-domain and constrained contextualization tasks, where the targets consist of mostly rare long-tail words.Our final system remains lightweight and modular, allowing for quick modification without model re-training. Mahaveer Jain, Gil Keren, Suyoun Kim, Yangyang Shi, Jay Mahadeokar, Julian Chan, Yuan Shangguan, Christian Fügen, Ozlem Kalinli, Yatharth Saraf, Michael L. Seltzer |
Interspeech | 7 |
| 2021 | Dynamic Encoder Transducer: A Flexible Solution for Trading Off Accuracy for LatencyabstractWe propose a dynamic encoder transducer (DET) for on-device speech recognition. One DET model scales to multiple devices with different computation capacities without retraining or finetuning. To trading off accuracy and latency, DET assigns different encoders to decode different parts of an utterance. We apply and compare the layer dropout and the collaborative learning for DET training. The layer dropout method that randomly drops out encoder layers in the training phase, can do on-demand layer dropout in decoding. Collaborative learning jointly trains multiple encoders with different depths in one single model. Experiment results on Librispeech and in-house data show that DET provides a flexible accuracy and latency trade-off. Results on Librispeech show that the full-size encoder in DET relatively reduces the word error rate of the same size baseline by over 8%. The lightweight encoder in DET trained with collaborative learning reduces the model size by 25% but still gets similar WER as the full-size baseline. DET gets similar accuracy as a baseline model with better latency on a large in-house data set by assigning a lightweight encoder for the beginning part of one utterance and a full-size encoder for the rest. Yangyang Shi, Varun Nagaraja, Chunyang Wu, Jay Mahadeokar, Rohit Prabhavalkar, Alex Xiao, Ching-Feng Yeh, Julian Chan, Christian Fügen, Ozlem Kalinli, Michael L. Seltzer |
Interspeech | 9 |
| 2021 | Deep Shallow Fusion for RNN-T PersonalizationabstractEnd-to-end models in general, and Recurrent Neural Network Transducer (RNN-T) in particular, have gained significant traction in the automatic speech recognition community in the last few years due to their simplicity, compactness, and excellent performance on generic transcription tasks. However, these models are more challenging to personalize compared to traditional hybrid systems due to the lack of external language models and difficulties in recognizing rare long-tail words, specifically entity names. In this work, we present novel techniques to improve RNN-T's ability to model rare WordPieces, infuse extra information into the encoder, enable the use of alternative graphemic pronunciations, and perform deep fusion with personalized language models for more robust biasing. We show that these combined techniques result in 15.4%-34.5% relative Word Error Rate improvement compared to a strong RNN-T baseline which uses shallow fusion and text-to-speech augmentation. Our work helps push the boundary of RNN-T personalization and close the gap with hybrid systems on use cases where biasing and entity recognition are crucial. Gil Keren, Julian Chan, Jay Mahadeokar, Christian Fügen, Michael L. Seltzer |
SLT | 3 |
| 2021 | Streaming Attention-Based Models with Augmented Memory for End-To-End Speech RecognitionabstractAttention-based models have been gaining popularity recently for their strong performance demonstrated in fields such as machine translation [1] and automatic speech recognition [2]. One major challenge of attention-based models is the need of access to the full sequence and the quadratically growing computational cost concerning the sequence length. These characteristics pose challenges, especially for low-latency scenarios, where the system is often required to be streaming. In this paper, we build a compact and streaming speech recognition system on top of the end-to-end neural transducer architecture [3] with attention-based modules augmented with convolution [2]. The proposed system equips the end-to-end models with the streaming capability and reduces the large footprint from the streaming attention-based model using augmented memory [4], [5]. On the LibriSpeech [6] dataset, our proposed system achieves word error rates 2.7% on test-clean and 5.8% on test-other, to our best knowledge the lowest among streaming approaches reported so far. Ching-Feng Yeh, Yongqiang Wang 0005, Yangyang Shi, Chunyang Wu, Frank Zhang 0001, Julian Chan, Michael L. Seltzer |
SLT | 6 |
| 2021 | Benchmarking LF-MMI, CTC And RNN-T Criteria For Streaming ASRabstractIn this work, to measure the accuracy and efficiency for a latency-controlled streaming automatic speech recognition (ASR) application, we perform comprehensive evaluations on three popular training criteria: LF-MMI, CTC and RNN-T. In transcribing social media videos of 7 languages with training data 3K - 14K hours, we conduct large-scale controlled experimentation across each criterion using identical datasets and encoder model architecture. We find that RNN-T has consistent wins in ASR accuracy, while CTC models excel at inference efficiency. Moreover, we selectively examine various modeling strategies for different training criteria, including modeling units, encoder architectures, pre-training, etc. Given such large-scale real-world streaming ASR application, to our best knowledge, we present the first comprehensive benchmark on these three widely used training criteria across a great many languages. Xiaohui Zhang 0007, Frank Zhang 0001, Chunxi Liu, Kjell Schubert, Julian Chan, Pradyot Prakash, Ching-Feng Yeh, Fuchun Peng, Yatharth Saraf, Geoffrey Zweig |
SLT | 5 |
| 2014 | Manipulating stance and involvement using collaborative tasks: an exploratory comparisonabstractThe ATAROS project aims to identify acoustic signals of stance-taking in order to inform the development of automatic stance recognition in natural speech. Due to the typically low frequency of stance-taking in existing corpora that have been used to investigate related phenomena such as subjectivity, we are creating an audio corpus of unscripted conversations between dyads as they complete collaborative tasks designed to elicit a high density of stance-taking at increasing levels of involvement. To validate our experimental design and provide a preliminary assessment of the corpus, we examine a fully transcribed and time-aligned portion to compare the speaking styles in two tasks, one expected to elicit low involvement and weak stances, the other high involvement and strong stances. We find that although overall measures such as task duration and total word count do not indicate consistent differences across tasks, speakers do display significant differences in speaking style. Factors such as increases in speaking rate, turn length, and disfluencies from weakto strong-stance tasks are consistent with increased involvement by the participants and provide evidence in support of the experimental design. Valerie Freeman, Julian Chan, Gina-Anne Levow, Richard A. Wright, Mari Ostendorf, Victoria Zayats |
INTERSPEECH | 2 |
| 2014 | Recognition of stance strength and polarity in spontaneous speechabstractFrom activities as simple as scheduling a meeting to those as complex as balancing a national budget, people take stances in negotiations and decision making. While the related areas of subjectivity and sentiment analysis have received significant attention, work has focused almost exclusively on text, and much stance-taking activity is carried out verbally. This paper investigates automatic recognition of stance-taking in spontaneous speech. It first presents a new annotated corpus of spontaneous, conversational speech designed to elicit high densities of stance-taking at different strengths. Speaker spurts are annotated both for strength of stance-taking behavior and polarity of stance. Based on this annotated corpus, we develop classifiers for automatic recognition of stance-taking behavior in speech. We employ a range of lexical, speaking style, and prosodic features in a boosting framework. The classifiers achieve strong accuracies on both binary detection of stance and four-way recognition of stance strength, well above most common class assignment. Finally, we classify the polarity of stance-taking spurts, obtaining accuracies around 80%. The best classifiers rely primarily on word unigram features, with speaking style and prosodic features yielding lower accuracies but still well above common class assignment. Gina-Anne Levow, Valerie Freeman, Alena Hrynkevich, Mari Ostendorf, Richard A. Wright, Julian Chan, Yi Luan, Trang Tran 0001 |
SLT | 6 |