EDBT 2026 Demo / reviewers in the wild / expert
Oleg Rybakov
dblp:238/2534
· DBLP profile ↗
15ranked-venue papers
5as first author
12since 2021 · last 2025
0009-0007-0021-047XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SimulTron: On-Device Simultaneous Speech to Speech TranslationabstractSimultaneous speech-to-speech translation (S2ST) holds the promise of breaking down communication barriers and enabling fluid conversations across languages. However, achieving accurate, real-time translation through mobile devices remains a major challenge. We introduce SimulTron, a novel S2ST architecture designed to tackle this task. SimulTron is a lightweight direct S2ST model that uses the strengths of the Translatotron framework while incorporating key modifications for streaming operation, and an adjustable fixed delay. Our experiments show that SimulTron surpasses Translatotron 2 in offline evaluations. Furthermore, real-time evaluations reveal that SimulTron improves upon the performance achieved by Translatotron 1. Additionally, SimulTron achieves superior BLEU scores and latency compared to previous real-time S2ST method on the MuST-C dataset. Significantly, we have successfully deployed SimulTron on a Pixel 7 Pro device, show its potential for simultaneous S2ST on-device. Alex Agranovich, Eliya Nachmani, Oleg Rybakov, Yifan Ding 0004, Ye Jia, Nadav Bar, Heiga Zen, Michelle Tadmor Ramanovich |
ICASSP | 3 |
| 2024 | USM-Lite: Quantization and Sparsity Aware Fine-Tuning for Speech Recognition with Universal Speech ModelsabstractEnd-to-end automatic speech recognition (ASR) models have seen revolutionary quality gains with the recent development of large-scale universal speech models (USM). However, deploying these massive USMs is extremely expensive due to the enormous memory usage and computational cost. Therefore, model compression is an important research topic to fit USM-based ASR under budget in real-world scenarios. In this study, we propose a USM fine-tuning approach for ASR, with a low-bit quantization and N:M structured sparsity aware paradigm on the model weights, reducing the model complexity from parameter precision and matrix topology perspectives. We conducted extensive experiments with a 2-billion parameter USM on a large-scale voice search dataset to evaluate our proposed method. A series of ablation studies validate the effectiveness of up to int4 quantization and 2:4 sparsity. However, a single compression technique fails to recover the performance well under extreme setups including int2 quantization and 1:4 sparsity. By contrast, our proposed method can compress the model to have 9.4% of the size, at the cost of only 7.3% relative word error rate (WER) regressions. We also provided in-depth analyses on the results and discussions on the limitations and potential solutions, which would be valuable for future studies. Shaojin Ding, David Qiu, David Rim, Yanzhang He, Oleg Rybakov, Bo Li 0028, Rohit Prabhavalkar, Tara N. Sainath, Zhonglin Han, Amir Yazdanbakhsh, Shivani Agrawal |
ICASSP | 5 |
| 2024 | FedAQT: Accurate Quantized Training with Federated LearningabstractFederated learning (FL) has been widely used to train neural networks with the decentralized training procedure where data is only accessed on clients’ devices for privacy preservation. However, the limited computation resources on clients’ devices prevent FL of large models. To overcome the constraint, one possible method is to reduce the computation memory usage with quantized neural networks such as quantization aware training on a centralized server. However, directly applying the quantization aware methods does not reduce the memory consumption on the clients’ devices of FL because the full-precision model is still used in the forward propagation of the model computation. To enable FL of the Conformer based ASR models, we propose FedAQT, an accurate quantized training framework under FL by training with quantized variables directly on clients’ devices. We empirically show that our method can achieve comparable WER with only 60% memory of the full-precision model. Renkun Ni, Yonghui Xiao, Phoenix Meadowlark, Oleg Rybakov, Tom Goldstein, Ananda Theertha Suresh, Ignacio López-Moreno, Mingqing Chen, Rajiv Mathews |
ICASSP | 4 |
| 2024 | USM RNN-T model weights binarization
Oleg Rybakov, Dmitriy Serdyuk, Chengjian Zheng |
INTERSPEECH | 1 |
| 2024 | Rand: Robustness Aware Norm Decay for Quantized Neural NetworksabstractWith the rapid increase in the size of neural networks, model compression has become an important area of research. Quantization is an effective technique for decreasing the size, memory access, and compute load of large models. In this paper, we first benchmark the impact of techniques such as straight through estimator, pseudo-quantization noise (PQN), learnable scale parameter, clipping, etc. on 4-bit seq 2 seq models across a suite of speech recognition datasets ranging from 1,000 hours to 1 million hours, as well as one machine translation dataset to illustrate its applicability outside of speech. It is commonly believed that having dedicated learnable scale parameters per quantization group is critical to model accuracy. Instead, we propose to construct the scale as a Lpnorm of a few of the largest outliers within the quantization group and regularizing that norm within the end-to-end optimization. This outperforms popular learnable scale and clipping methods without the need to introduce extra parameters. PQN-QAT shows a larger improvement under the proposed method, and it opens up the potential to exploit some of its other benefits: 1) training a single model that performs well in mixed precision and 2) improved generalization on long form speech. David Qiu, David Rim, Shaojin Ding, Oleg Rybakov, Yanzhang He |
SLT | 4 |
| 2023 | STEP: Learning N: M Structured Sparsity Masks from Scratch with PreconditionabstractRecent innovations on hardware (e.g. Nvidia A100) have motivated learning N:M structured sparsity masks from scratch for fast model inference. However, state-of-the-art learning recipes in this regime (e.g. SR-STE) are proposed for non-adaptive optimizers like momentum SGD, while incurring non-trivial accuracy drop for Adam-trained models like attention-based LLMs. In this paper, we first demonstrate such gap origins from poorly estimated second moment (i.e. variance) in Adam states given by the masked weights. We conjecture that learning N:M masks with Adam should take the critical regime of variance estimation into account. In light of this, we propose STEP, an Adam-aware recipe that learns N:M masks with two phases: first, STEP calculates a reliable variance estimate (precondition phase) and subsequently, the variance remains fixed and is used as a precondition to learn N:M masks (mask-learning phase). STEP automatically identifies the switching point of two phases by dynamically sampling variance changes over the training trajectory and testing the sample concentration. Empirically, we evaluate STEP and other baselines such as ASP and SR-STE on multiple tasks including CIFAR classification, machine translation and LLM fine-tuning (BERT-Base, GPT-2). We show STEP mitigates the accuracy drop of baseline recipes and is robust to aggressive structured sparsity ratios. Yucheng Lu 0003, Shivani Agrawal, Suvinay Subramanian, Oleg Rybakov, Christopher De Sa, Amir Yazdanbakhsh |
ICML | 4 |
| 2023 | Streaming Parrotron for on-device speech-to-speech conversion
Oleg Rybakov, Fadi Biadsy, Liyang Jiang, Phoenix Meadowlark, Shivani Agrawal |
INTERSPEECH | 1 |
| 2023 | 2-bit Conformer quantization for automatic speech recognition
Oleg Rybakov, Phoenix Meadowlark, Shaojin Ding, David Qiu, David Rim, Yanzhang He |
INTERSPEECH | 1 |
| 2023 | Real time spectrogram inversion on mobile phone
Oleg Rybakov, Marco Tagliasacchi, Yunpeng Li 0008, Liyang Jiang, Fadi Biadsy |
INTERSPEECH | 1 |
| 2022 | A Scalable Model Specialization Framework for Training and Inference using Submodels and its Application to Speech Model PersonalizationabstractModel fine-tuning and adaptation have become a common approach for model specialization for downstream tasks or domains. Fine-tuning the entire model or a subset of the parameters using light-weight adaptation has shown considerable success across different specialization tasks. Fine-tuning a model for a large number of domains typically requires starting a new training job for every domain posing scaling limitations. Once these models are trained, deploying them also poses significant scalability challenges for inference for real-time applications. In this paper, building upon prior light-weight adaptation techniques, we propose a modular framework that enables us to substantially improve scalability for model training and inference. We introduce Submodels that can be quickly and dynamically loaded for on-the-fly inference. We also propose multiple approaches for training those Submodels in parallel using an embedding space in the same training job. We test our framework on an extreme use-case which is speech model personalization for atypical speech, requiring a Submodel for each user. We obtain 128x Submodel throughput with a fixed computation budget without a loss of accuracy. We also show that learning a speaker-embedding space can scale further and reduce the amount of personalization training data required per speaker. Fadi Biadsy, Youzheng Chen, Oleg Rybakov, Andrew Rosenberg, Pedro J. Moreno 0001 |
INTERSPEECH | 4 |
| 2022 | 4-bit Conformer with Native Quantization Aware Training for Speech RecognitionabstractReducing the latency and model size has always been a significant research problem for live Automatic Speech Recognition (ASR) application scenarios. Along this direction, model quantization has become an increasingly popular approach to compress neural networks and reduce computation cost. Most of the existing practical ASR systems apply post-training 8-bit quantization. To achieve a higher compression rate without introducing additional performance regression, in this study, we propose to develop 4-bit ASR models with native quantization aware training, which leverages native integer operations to effectively optimize both training and inference. We conducted two experiments on state-of-the-art Conformer-based ASR models to evaluate our proposed quantization technique. First, we explored the impact of different precisions for both weight and activation quantization on the LibriSpeech dataset, and obtained a lossless 4-bit Conformer model with 7.7x size reduction compared to the float32 model. Following this, we for the first time investigated and revealed the viability of 4-bit quantization on a practical ASR system that is trained with large-scale datasets, and produced a lossless Conformer ASR model with mixed 4-bit and 8-bit weights that has 5x size reduction compared to the float32 model. Shaojin Ding, Phoenix Meadowlark, Yanzhang He, Lukasz Lew, Shivani Agrawal, Oleg Rybakov |
INTERSPEECH | 6 |
| 2021 | Real-Time Speech Frequency Bandwidth ExtensionabstractIn this paper we propose a lightweight model for frequency bandwidth extension of speech signals, increasing the sampling frequency from 8kHz to 16kHz while restoring the high frequency content to a level almost indistinguishable from the 16kHz ground truth. The model architecture is based on SEANet (Sound EnhAncement Network), a wave-to-wave fully convolutional model, which uses a combination of feature losses and adversarial losses to reconstruct an enhanced version of the input speech. In addition, we propose a variant of SEANet that can be deployed on-device in streaming mode, achieving an architectural latency of 16ms. When profiled on a single core of a mobile CPU, processing one 16ms frame takes only 1.5ms. The low latency makes it viable for bi-directional voice communication systems. Yunpeng Li 0008, Marco Tagliasacchi, Oleg Rybakov, Victor Ungureanu, Dominik Roblek |
ICASSP | 3 |
| 2020 | Streaming Keyword Spotting on Mobile DevicesabstractIn this work we explore the latency and accuracy of keyword spotting (KWS) models in streaming and non-streaming modes on mobile phones. NN model conversion from non-streaming mode (model receives the whole input sequence and then returns the classification result) to streaming mode (model receives portion of the input sequence and classifies it incrementally) may require manual model rewriting. We address this by designing a Tensorflow/Keras based library which allows automatic conversion of non-streaming models to streaming ones with minimum effort. With this library we benchmark multiple KWS models in both streaming and non-streaming modes on mobile phones and demonstrate different tradeoffs between latency and accuracy. We also explore novel KWS models with multi-head attention which reduce the classification error over the state-of-art by 10% on Google speech commands data sets V2. The streaming library with all experiments is open-sourced. Oleg Rybakov, Natasha Kononenko, Niranjan Subrahmanya, Mirkó Visontai, Stella Laurenzo |
INTERSPEECH | 1 |
| 2019 | Low-Bit Quantization and Quantization-Aware Training for Small-Footprint Keyword SpottingabstractIn this paper, we investigate novel quantization approaches to reduce memory and computational footprint of deep neural network (DNN) based keyword spotters (KWS). We propose a new method for KWS offline and online quantization, which we call dynamic quantization, where we quantize DNN weight matrices column-wise, using each column's exact individual min-max range, and the DNN layers' inputs and outputs are quantized for every input audio frame individually, using the exact min-max range of each input and output vector. We further apply a new quantization-aware training approach that allows us to incorporate quantization errors into KWS model during training. Together, these approaches allow us to significantly improve the performance of KWS in 4-bit and 8-bit quantized precision, achieving the end-to-end accuracy close to that of full precision models while reducing the models' on-device memory footprint by up to 80%. Yuriy Mishchenko, Yusuf Goren, Ming Sun 0007, Chris Beauchene, Spyridon Matsoukas, Oleg Rybakov, Shiv Vitaladevuni |
ICMLA | 6 |
| 2019 | Two Tiered Distributed Training Algorithm for Acoustic Modeling
Pranav Ladkat, Oleg Rybakov, Radhika Arava, Sree Hari Krishnan Parthasarathi, I-Fan Chen, Nikko Strom |
INTERSPEECH | 2 |