VLDB 2026 Research / reviewers in the wild / expert
Rithesh Kumar
dblp:192/1862
· DBLP profile ↗
6ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0000-6881-2783ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Speech recognition and synthesis · 50% Generative modeling · 40% Deep learning architectures and training · 5% | |
| Computer graphics and multimedia
5 papers |
Audio and music processing · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
generative adversarial network |
1.6 | 3 | 2023 | High-Fidelity Audio Compression with Improved RVQGAN · NeurIPS 2023 Chunked Autoregressive GAN for Conditional Waveform Synthesis · ICLR 2022 MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis · NeurIPS 2019 |
Audio and music processing
audio coding |
1.0 | 2 | 2023 | High-Fidelity Audio Compression with Improved RVQGAN · NeurIPS 2023 MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis · NeurIPS 2019 |
Audio and music processing › audio coding
neural audio codec |
1.0 | 2 | 2023 | High-Fidelity Audio Compression with Improved RVQGAN · NeurIPS 2023 MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis · NeurIPS 2019 |
Natural language and speech › Speech recognition and synthesis › speech synthesis
diffusion-based speech synthesis |
0.9 | 1 | 2025 | DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis · ICML 2025 |
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis |
0.9 | 1 | 2025 | DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis · ICML 2025 |
Natural language and speech › Speech recognition and synthesis › speech synthesis
zero-shot speech synthesis |
0.9 | 1 | 2025 | DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis · ICML 2025 |
Machine learning › Generative modeling
audio generation |
0.7 | 1 | 2023 | High-Fidelity Audio Compression with Improved RVQGAN · NeurIPS 2023 |
Audio and music processing
speech synthesis |
0.4 | 1 | 2019 | MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis · NeurIPS 2019 |
Audio and music processing › sound synthesis
music synthesis |
0.3 | 2 | 2023 | High-Fidelity Audio Compression with Improved RVQGAN · NeurIPS 2023 MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis · NeurIPS 2019 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2017 | SampleRNN: An Unconditional End-to-End Neural Audio Generation Model · ICLR (Poster) 2017 |
Audio and music processing › sound synthesis
neural audio synthesis |
0.3 | 1 | 2017 | SampleRNN: An Unconditional End-to-End Neural Audio Generation Model · ICLR (Poster) 2017 |
Audio and music processing
sound synthesis |
0.3 | 1 | 2017 | SampleRNN: An Unconditional End-to-End Neural Audio Generation Model · ICLR (Poster) 2017 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
model distillation |
0.3 | 1 | 2025 | DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis · ICML 2025 |
Interaction techniques and input
voice interaction |
0.3 | 1 | 2025 | SpeakEasy: Enhancing Text-to-Speech Interactions for Expressive Content Creation · CHI 2025 |
Natural language and speech › Speech recognition and synthesis
speech synthesis |
0.2 | 1 | 2023 | High-Fidelity Audio Compression with Improved RVQGAN · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
vector quantization · 2.1adversarial loss · 2.1system design · 1.7residual vector quantization · 1.3chunked generation · 1.1autoregressive modeling · 1.1speaker verification · 0.9knowledge distillation · 0.9diffusion model · 0.9connectionist temporal classification · 0.9mel-spectrogram inversion · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SpeakEasy: Enhancing Text-to-Speech Interactions for Expressive Content CreationabstractCHI ’25, Yokohama, Japan Stephen Brade, Sam Anderson, Rithesh Kumar, Zeyu Jin, Anh Truong |
CHI | 3 |
| 2025 | DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech SynthesisabstractDiffusion models have demonstrated significant potential in speech synthesis tasks, including text-to-speech (TTS) and voice cloning. However, their iterative denoising processes are computationally intensive, and previous distillation attempts have shown consistent quality degradation. Moreover, existing TTS approaches are limited by non-differentiable components or iterative sampling that prevent true end-to-end optimization with perceptual metrics. We introduce DMOSpeech, a distilled diffusion-based TTS model that uniquely achieves both faster inference and superior performance compared to its teacher model. By enabling direct gradient pathways to all model components, we demonstrate the first successful end-to-end optimization of differentiable metrics in TTS, incorporating Connectionist Temporal Classification (CTC) loss and Speaker Verification (SV) loss. Our comprehensive experiments, validated through extensive human evaluation, show significant improvements in naturalness, intelligibility, and speaker similarity while reducing inference time by orders of magnitude. This work establishes a new framework for aligning speech synthesis with human auditory preferences through direct metric optimization. The audio samples are available at https://dmospeech.github.io/demo Yinghao Aaron Li, Rithesh Kumar, Zeyu Jin |
ICML | 2 |
| 2023 | High-Fidelity Audio Compression with Improved RVQGANabstractLanguage models have been successfully used to model natural signals, such as images, speech, and music. A key component of these models is a high quality neural compression model that can compress high-dimensional natural signals into lower dimensional discrete tokens. To that end, we introduce a high-fidelity universal neural audio compression algorithm that achieves ~90x compression of 44.1 KHz audio into tokens at just 8kbps bandwidth. We achieve this by combining advances in high-fidelity audio generation with better vector quantization techniques from the image domain, along with improved adversarial and reconstruction losses. We compress all domains (speech, environment, music, etc.) with a single universal model, making it widely applicable to generative modeling of all audio. We compare with competing audio compression algorithms, and find our method outperforms them significantly. We provide thorough ablations for every design choice, as well as open-source code and trained model weights. We hope our work can lay the foundation for the next generation of high-fidelity audio modeling. Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar |
NeurIPS | 1 |
| 2022 | Chunked Autoregressive GAN for Conditional Waveform Synthesis
Max Morrison, Rithesh Kumar, Prem Seetharaman, Aaron C. Courville, Yoshua Bengio |
ICLR | 2 |
| 2019 | MelGAN: Generative Adversarial Networks for Conditional Waveform SynthesisabstractPrevious works (Donahue et al., 2018a; Engel et al., 2019a) have found that generating coherent raw audio waveforms with GANs is challenging. In this paper, we show that it is possible to train GANs reliably to generate high quality coherent waveforms by introducing a set of architectural changes and simple training techniques. Subjective evaluation metric (Mean Opinion Score, or MOS) shows the effectiveness of the proposed approach for high quality mel-spectrogram inversion. To establish the generality of the proposed techniques, we show qualitative results of our model in speech synthesis, music domain translation and unconditional music synthesis. We evaluate the various components of the model through ablation studies and suggest a set of guidelines to design general purpose discriminators and generators for conditional sequence synthesis tasks. Our model is non-autoregressive, fully convolutional, with significantly fewer parameters than competing models and generalizes to unseen speakers for mel-spectrogram inversion. Our pytorch implementation runs at more than 100x faster than realtime on GTX 1080Ti GPU and more than 2x faster than real-time on CPU, without any hardware specific optimization tricks. Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, Aaron C. Courville |
NeurIPS | 2 |
| 2017 | SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
Soroush Mehri, Ishaan Gulrajani, Rithesh Kumar, Jose Sotelo, Aaron C. Courville, Yoshua Bengio |
ICLR (Poster) | 4 |