VLDB 2026 Research / reviewers in the wild / expert
Shaojin Ding
dblp:226/1807
· DBLP profile ↗
25ranked-venue papers
14as first author
16since 2021 · last 2024
0000-0002-2108-3111ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 11 first-author · 13 since 2021Artificial intelligence and machine learning · 18 · 12 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | USM-Lite: Quantization and Sparsity Aware Fine-Tuning for Speech Recognition with Universal Speech ModelsabstractEnd-to-end automatic speech recognition (ASR) models have seen revolutionary quality gains with the recent development of large-scale universal speech models (USM). However, deploying these massive USMs is extremely expensive due to the enormous memory usage and computational cost. Therefore, model compression is an important research topic to fit USM-based ASR under budget in real-world scenarios. In this study, we propose a USM fine-tuning approach for ASR, with a low-bit quantization and N:M structured sparsity aware paradigm on the model weights, reducing the model complexity from parameter precision and matrix topology perspectives. We conducted extensive experiments with a 2-billion parameter USM on a large-scale voice search dataset to evaluate our proposed method. A series of ablation studies validate the effectiveness of up to int4 quantization and 2:4 sparsity. However, a single compression technique fails to recover the performance well under extreme setups including int2 quantization and 1:4 sparsity. By contrast, our proposed method can compress the model to have 9.4% of the size, at the cost of only 7.3% relative word error rate (WER) regressions. We also provided in-depth analyses on the results and discussions on the limitations and potential solutions, which would be valuable for future studies. Shaojin Ding, David Qiu, David Rim, Yanzhang He, Oleg Rybakov, Bo Li 0028, Rohit Prabhavalkar, Tara N. Sainath, Zhonglin Han, Amir Yazdanbakhsh, Shivani Agrawal |
ICASSP | 1 |
| 2024 | Rand: Robustness Aware Norm Decay for Quantized Neural NetworksabstractWith the rapid increase in the size of neural networks, model compression has become an important area of research. Quantization is an effective technique for decreasing the size, memory access, and compute load of large models. In this paper, we first benchmark the impact of techniques such as straight through estimator, pseudo-quantization noise (PQN), learnable scale parameter, clipping, etc. on 4-bit seq 2 seq models across a suite of speech recognition datasets ranging from 1,000 hours to 1 million hours, as well as one machine translation dataset to illustrate its applicability outside of speech. It is commonly believed that having dedicated learnable scale parameters per quantization group is critical to model accuracy. Instead, we propose to construct the scale as a Lpnorm of a few of the largest outliers within the quantization group and regularizing that norm within the end-to-end optimization. This outperforms popular learnable scale and clipping methods without the need to introduce extra parameters. PQN-QAT shows a larger improvement under the proposed method, and it opens up the potential to exploit some of its other benefits: 1) training a single model that performs well in mixed precision and 2) improved generalization on long form speech. David Qiu, David Rim, Shaojin Ding, Oleg Rybakov, Yanzhang He |
SLT | 3 |
| 2023 | Efficient Cascaded Streaming ASR System Via Frame Rate ReductionabstractIn this paper, we explore various frame rate reduction schemes on the two-pass cascaded encoder model to improve its efficiency without scarifying the transcription quality. We conduct extensive studies on frame rate reduction strategies, left and right context window length, trade-offs in quality, latency, computation and power consumption, and performance in short-and long-form datasets. With the proposed schemes, we can lower the 2nd pass frame rate to $120 \mathrm{~ms}$, half of the 1st pass’s. This achieves $20 \%$ RTF reduction / $13 \%$ power saving / $19 \%$ lower final latency, without impact on the word-error-rate nor partial results’ latency. If allowing partial latency increase, we can further reduce the frame rate to $180 \mathrm{~ms}$ or even $240 \mathrm{~ms}$ from the 1st pass, and obtain $45 \%$ RTF / 35% power savings, with a similar or even better (on the short-form testset) recognition accuracy. Xingyu Cai, David Qiu, Shaojin Ding, Dongseong Hwang, Antoine Bruguier, Rohit Prabhavalkar, Tara N. Sainath, Yanzhang He |
ASRU | 3 |
| 2023 | The Role of Feature Correlation on Quantized Neural NetworksabstractWith the growing need for large models in speech recognition, quantization has become a valuable technique to reduce their compute and memory transfer costs. Quantized models tends to show worse accuray compared to their float counterparts, due to the noise introduced from quantization. One pathological case is when the noise caused by a group of weights are constructively combined. In contrast to recent advances in post-training quantization, this paper proposes to discourage that constructive combination during the quantization aware training process. This is accomplished by regularizing three different functions of the layer-wise input feature correlation matrix. We show that these methods can improve the 2-bit performance of the Conformer model on LibriSpeech. Additionally, they serve as an informative measure for the ease of quantization in the different types of layers in the Conformer, which we show to help aid selecting layer precision for mixed precision models. David Qiu, Shaojin Ding, Yanzhang He |
ASRU | 2 |
| 2023 | Sharing Low Rank Conformer Weights for Tiny Always-On Ambient Speech Recognition ModelsabstractContinued improvements in machine learning techniques offer exciting new opportunities through the use of larger models and larger training datasets. However, there is a growing need to offer these new capabilities on-board low-powered devices such as smart-phones, wearables and other embedded environments where only low memory is available. Towards this, we consider methods to reduce the model size of Conformer-based speech recognition models which typically require models with greater than 100M parameters down to just 5M parameters while minimizing impact on model quality. Such a model allows us to achieve always-on ambient speech recognition on edge devices with low-memory neural processors. We propose model weight reuse at different levels within our model architecture: (i) repeating full conformer block layers, (ii) sharing specific conformer modules across layers, (iii) sharing sub-components per conformer module, and (iv) sharing decomposed sub-component weights after low-rank decomposition. By sharing weights at different levels of our model, we can retain the full model in-memory while increasing the number of virtual trans-formations applied to the input. Through a series of ablation studies and evaluations, we find that with weight sharing and a low-rank architecture, we can achieve a WER of 2.84 and 2.94 for Librispeech dev-clean and test-clean respectively with a 5M parameter model. Steven M. Hernandez, Ding Zhao, Shaojin Ding, Antoine Bruguier, Rohit Prabhavalkar, Tara N. Sainath, Yanzhang He, Ian McGraw |
ICASSP | 3 |
| 2023 | Conditional Conformer: Improving Speaker Modulation For Single And Multi-User Speech EnhancementabstractRecently, Feature-wise Linear Modulation (FiLM) has been shown to outperform other approaches to incorporate speaker embedding into speech separation and VoiceFilter models. We propose an improved method of incorporating such embeddings into a Voice- Filter frontend for automatic speech recognition (ASR) and text- independent speaker verification (TI-SV). We extend the widely- used Conformer architecture to construct a FiLM Block with additional feature processing before and after the FiLM layers. Apart from its application to single-user VoiceFilter, we show that our system can be easily extended to multi-user VoiceFilter models via element-wise max pooling of the speaker embeddings in a projected space. The final architecture, which we call Conditional Conformer, tightly integrates the speaker embeddings into a Conformer backbone. We improve TI-SV equal error rates by as much as 56% over prior multi-user VoiceFilter models, and our element-wise max pooling reduces relative WER compared to an attention mechanism by as much as 10%. Tom O'Malley, Shaojin Ding, Arun Narayanan, Rajeev Rikhye, Qiao Liang 0001, Yanzhang He, Ian McGraw |
ICASSP | 2 |
| 2023 | Multi-Output RNN-T Joint Networks for Multi-Task Learning of ASR and Auxiliary TasksabstractWe propose a multi-output joint network architecture for RNN-T transducer, for multi-task modeling of ASR and auxiliary tasks that rely on ASR outputs. Each output of the joint network predicts tar-get labels with disjoint vocabularies for each task, while sharing the same audio features by the encoder and language model features by the prediction network. Each task is trained with an RNN-T loss that marginalizes over all possible paths, and we allow multiple tasks to share the blank logit so that they are synchronized. We demonstrate our method on two auxiliary tasks, namely capitalization and pause prediction, and discuss different considerations for modeling and inference procedures. For capitalization, we successfully distill capitalization labels from a standalone text normalization model, and achieve competitive Uppercase Error Rate (UER) while offering streaming capability and improved inference efficiency. In addition, our model has similar capitalization accuracy compared to a mixed-case ASR model, but obtains improved WERs if integrated with external language models. For pause prediction, we achieve the same performance as the previous two-step approach while providing a simpler training recipe without affecting ASR accuracy. Ding Zhao, Shaojin Ding, Hao Zhang 0010, Shuo-Yiin Chang, David Rybach, Tara N. Sainath, Yanzhang He, Ian McGraw, Shankar Kumar |
ICASSP | 3 |
| 2023 | 2-bit Conformer quantization for automatic speech recognition
Oleg Rybakov, Phoenix Meadowlark, Shaojin Ding, David Qiu, David Rim, Yanzhang He |
INTERSPEECH | 3 |
| 2022 | Towards Lifelong Learning of Multilingual Text-to-Speech SynthesisabstractThis work presents a lifelong learning approach to train a multilingual Text-To-Speech (TTS) system, where each language was seen as an individual task and was learned sequentially and continually. It does not require pooled data from all languages altogether, and thus alleviates the storage and computation burden. One of the challenges of lifelong learning methods is "catastrophic forgetting": in TTS scenario it means that model performance quickly degrades on previous languages when adapted to a new language. We approach this problem via a data-replay-based lifelong learning method. We formulate the replay process as a supervised learning problem, and propose a simple yet effective dual-sampler framework to tackle the heavily language-imbalanced training samples. Through objective and subjective evaluations, we show that this supervised learning formulation outperforms other gradient-based and regularization-based lifelong learning methods, achieving 43% Mel-Cepstral Distortion reduction compared to a fine-tuning baseline. Mu Yang, Shaojin Ding, Tianlong Chen 0001, Zhangyang Wang |
ICASSP | 2 |
| 2022 | Audio Lottery: Speech Recognition Made Ultra-Lightweight, Noise-Robust, and Transferable
Shaojin Ding, Tianlong Chen 0001, Zhangyang Wang |
ICLR | 1 |
| 2022 | 4-bit Conformer with Native Quantization Aware Training for Speech RecognitionabstractReducing the latency and model size has always been a significant research problem for live Automatic Speech Recognition (ASR) application scenarios. Along this direction, model quantization has become an increasingly popular approach to compress neural networks and reduce computation cost. Most of the existing practical ASR systems apply post-training 8-bit quantization. To achieve a higher compression rate without introducing additional performance regression, in this study, we propose to develop 4-bit ASR models with native quantization aware training, which leverages native integer operations to effectively optimize both training and inference. We conducted two experiments on state-of-the-art Conformer-based ASR models to evaluate our proposed quantization technique. First, we explored the impact of different precisions for both weight and activation quantization on the LibriSpeech dataset, and obtained a lossless 4-bit Conformer model with 7.7x size reduction compared to the float32 model. Following this, we for the first time investigated and revealed the viability of 4-bit quantization on a practical ASR system that is trained with large-scale datasets, and produced a lossless Conformer ASR model with mixed 4-bit and 8-bit weights that has 5x size reduction compared to the float32 model. Shaojin Ding, Phoenix Meadowlark, Yanzhang He, Lukasz Lew, Shivani Agrawal, Oleg Rybakov |
INTERSPEECH | 1 |
| 2022 | Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech RecognitionabstractPersonalization of on-device speech recognition (ASR) has seen explosive growth in recent years, largely due to the increasing popularity of personal assistant features on mobile devices and smart home speakers. In this work, we present Personal VAD 2.0, a personalized voice activity detector that detects the voice activity of a target speaker, as part of a streaming on-device ASR system. Although previous proof-of-concept studies have validated the effectiveness of Personal VAD, there are still several critical challenges to address before this model can be used in production: first, the quality must be satisfactory in both enrollment and enrollment-less scenarios; second, it should operate in a streaming fashion; and finally, the model size should be small enough to fit a limited latency and CPU/Memory budget. To meet the multi-faceted requirements, we propose a series of novel designs: 1) advanced speaker embedding modulation methods; 2) a new training paradigm to generalize to enrollment-less conditions; 3) architecture and runtime optimizations for latency and resource restrictions. Extensive experiments on a realistic speech recognition system demonstrated the state-of-the-art performance of our proposed method. Shaojin Ding, Rajeev Rikhye, Qiao Liang 0001, Yanzhang He, Arun Narayanan, Tom O'Malley, Ian McGraw |
INTERSPEECH | 1 |
| 2022 | A Unified Cascaded Encoder ASR Model for Dynamic Model SizesabstractIn this paper, we propose a dynamic cascaded encoder Automatic Speech Recognition (ASR) model, which unifies models for different deployment scenarios. Moreover, the model can significantly reduce model size and power consumption without loss of quality. Namely, with the dynamic cascaded encoder model, we explore three techniques to maximally boost the performance of each model size: 1) Use separate decoders for each sub-model while sharing the encoders; 2) Use funnel-pooling to improve the encoder efficiency; 3) Balance the size of causal and non-causal encoders to improve quality and fit deployment constraints. Overall, the proposed large-medium model has 30% smaller size and reduces power consumption by 33%, compared to the baseline cascaded encoder model. The triple-size model that unifies the large, medium, and small models achieves 37% total size reduction with minimal quality loss, while substantially reducing the engineering efforts of having separate models. Shaojin Ding, Ding Zhao, Tara N. Sainath, Yanzhang He, Robert David 0002, Rami Botros, Xin Wang 0116, Rina Panigrahy, Qiao Liang 0001, Dongseong Hwang, Ian McGraw, Rohit Prabhavalkar, Trevor Strohman |
INTERSPEECH | 1 |
| 2022 | Accentron: Foreign accent conversion to arbitrary non-native speakers using zero-shot learning
Shaojin Ding, Guanlong Zhao, Ricardo Gutierrez-Osuna |
Comput. Speech Lang. | 1 |
| 2021 | Textual Echo CancellationabstractIn this paper, we propose Textual Echo Cancellation (TEC) - a framework for cancelling the text-to-speech (TTS) playback echo11TTS playback denotes the synthesized TTS voice, and TTS playback echo denotes the reverberated TTS voice captured by the microphone. from overlapping speech recordings. Such a system can largely im-prove speech recognition performance and user experience for in-telligent devices such as smart speakers, as the user can talk to the device while the device is still playing the TTS signal responding to the previous query. We implement this system by using a novel sequence-to-sequence model with multi-source attention that takes both the microphone mixture signal and source text of the TTS play-back as inputs, and predicts the enhanced audio. Experiments show that the textual information of the TTS playback is critical to en-hancement performance. Besides, the text sequence is much smaller in size compared with the raw acoustic signal of the TTS playback, and can be immediately transmitted to the device or ASR server even before the playback is synthesized. Therefore, our proposed approach effectively reduces Internet communication and latency compared with alternative approaches such as acoustic echo cancellation (AEC). Shaojin Ding, Ye Jia |
ASRU | 1 |
| 2021 | Converting Foreign Accent Speech Without a ReferenceabstractForeign accent conversion (FAC) is the problem of generating a synthetic voice that has the voice identity of a second-language (L2) learner and the pronunciation patterns of a native (L1) speaker. This synthetic voice has been referred to as a “golden-speaker” in the pronunciation-training literature. FAC is generally achieved by building a voice-conversion model that maps utterances from a source (L1) speaker onto the target (L2) speaker. As such, FAC requires that a reference utterance from the L1 speaker be available at synthesis time. This greatly restricts the application scope of the FAC system. In this work, we propose a “reference-free” FAC system that eliminates the need for reference L1 utterances at synthesis time, and transforms L2 utterances directly. The system is trained in two steps. First, a conventional FAC procedure is used to create a golden-speaker using utterances from a reference L1 speaker (which are then discarded) and the L2 speaker. Second, a pronunciation-correction model is trained to convert L2 utterances to match the golden-speaker utterances obtained in the first step. At synthesis time, the pronunciation-correction model directly transforms a novel L2 utterance into its golden-speaker counterpart. Our results show that the system reduces foreign accents in novel L2 utterances, achieving a 20.5% relative reduction in word-error-rate of an American English automatic speech recognizer and a 19% reduction in perceptual ratings of foreign accentedness obtained through listening tests. Over 73% of the listeners also rated golden-speaker utterances as having the same voice identity as the original L2 utterances. Guanlong Zhao, Shaojin Ding, Ricardo Gutierrez-Osuna |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | AutoSpeech: Neural Architecture Search for Speaker RecognitionabstractSpeaker recognition systems based on Convolutional Neural Networks (CNNs) are often built with off-the-shelf backbones such as VGG-Net or ResNet.However, these backbones were originally proposed for image classification, and therefore may not be naturally fit for speaker recognition.Due to the prohibitive complexity of manually exploring the design space, we propose the first neural architecture search approach for the speaker recognition tasks, named as AutoSpeech.Our algorithm first identifies the optimal operation combination in a neural cell and then derives a CNN model by stacking the neural cell for multiple times.The final speaker recognition model can be obtained by training the derived CNN model through the standard scheme.To evaluate the proposed approach, we conduct experiments on both speaker identification and speaker verification tasks using the VoxCeleb1 dataset.Results demonstrate that the derived CNN architectures from the proposed approach significantly outperform current speaker recognition systems based on VGG-M, ResNet-18, and ResNet-34 backbones, while enjoying lower model complexity. Shaojin Ding, Tianlong Chen 0001, Xinyu Gong, Weiwei Zha, Zhangyang Wang |
INTERSPEECH | 1 |
| 2020 | Improving the Speaker Identity of Non-Parallel Many-to-Many Voice Conversion with Adversarial Speaker Recognition
Shaojin Ding, Guanlong Zhao, Ricardo Gutierrez-Osuna |
INTERSPEECH | 1 |
| 2020 | Learning Structured Sparse Representations for Voice ConversionabstractSparse-coding techniques for voice conversion assume that an utterance can be decomposed into a sparse code that only carries linguistic contents, and a dictionary of atoms that captures the speakers' characteristics. However, conventional dictionary-construction and sparse-coding algorithms rarely meet this assumption. The result is that the sparse code is no longer speaker-independent, which leads to lower voice-conversion performance. In this paper, we propose a Cluster-Structured Sparse Representation (CSSR) that improves the speaker independence of the representations. CSSR consists of two complementary components: a Cluster-Structured Dictionary Learning module that groups atoms in the dictionary into clusters, and a Cluster-Selective Objective Function that encourages each speech frame to be represented by atoms from a small number of clusters. We conducted four experiments on the CMU ARCTIC corpus to evaluate the proposed method. In a first ablation study, results show that each of the two CSSR components enhances speaker independence, and that combining both components leads to further improvements. In a second experiment, we find that CSSR uses increasingly larger dictionaries more efficiently than phoneme-based representations by allowing finer-grained decompositions of speech sounds. In a third experiment, results from objective and subjective measurements show that CSSR outperforms prior voice-conversion methods, improving the acoustic quality of the synthesized speech while retaining the target speaker's voice identity. Finally, we show that the CSSR captures latent (i.e., phonetic) information in the speech signal. Shaojin Ding, Guanlong Zhao, Christopher Liberatore, Ricardo Gutierrez-Osuna |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | ABD-Net: Attentive but Diverse Person Re-IdentificationabstractAttention mechanisms have been found effective for person re-identification (Re-ID). However, the learned "attentive'' features are often not naturally uncorrelated or "diverse'', which compromises the retrieval performance based on the Euclidean distance. We advocate the complementary powers of attention and diversity for Re-ID, by proposing an Attentive but Diverse Network (ABD-Net). ABD-Net seamlessly integrates attention modules and diversity regularizations throughout the entire network to learn features that are representative, robust, and more discriminative. Specifically, we introduce a pair of complementary attention modules, focusing on channel aggregation and position awareness, respectively. Then, we plug in a novel orthogonality constraint that efficiently enforces diversity on both hidden activations and weights. Through an extensive set of ablation study, we verify that the attentive and diverse terms each contributes to the performance boosts of ABD-Net. It consistently outperforms existing state-of-the-art methods on there popular person Re-ID benchmarks. Tianlong Chen 0001, Shaojin Ding, Ye Yuan 0012, Wuyang Chen 0001, Zhou Ren, Zhangyang Wang |
ICCV | 2 |
| 2019 | Group Latent Embedding for Vector Quantized Variational Autoencoder in Non-Parallel Voice Conversion
Shaojin Ding, Ricardo Gutierrez-Osuna |
INTERSPEECH | 1 |
| 2019 | Foreign Accent Conversion by Synthesizing Speech from Phonetic Posteriorgrams
Guanlong Zhao, Shaojin Ding, Ricardo Gutierrez-Osuna |
INTERSPEECH | 2 |
| 2019 | Golden speaker builder - An interactive tool for pronunciation training
Shaojin Ding, Christopher Liberatore, Sinem Sonsaat, Ivana Lucic, Alif Silpachai, Guanlong Zhao, Evgeny Chukharev-Hudilainen, John Levis, Ricardo Gutierrez-Osuna |
Speech Commun. | 1 |
| 2018 | Learning Structured Dictionaries for Exemplar-based Voice Conversion
Shaojin Ding, Christopher Liberatore, Ricardo Gutierrez-Osuna |
INTERSPEECH | 1 |
| 2018 | Improving Sparse Representations in Exemplar-Based Voice Conversion with a Phoneme-Selective Objective Function
Shaojin Ding, Guanlong Zhao, Christopher Liberatore, Ricardo Gutierrez-Osuna |
INTERSPEECH | 1 |