Rithesh Kumar

dblp:192/1862 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0000-6881-2783ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Speech recognition and synthesis · 50% Generative modeling · 40% Deep learning architectures and training · 5%
Computer graphics and multimedia
5 papers
Audio and music processing · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
generative adversarial network
1.632023
High-Fidelity Audio Compression with Improved RVQGAN · NeurIPS 2023
Chunked Autoregressive GAN for Conditional Waveform Synthesis · ICLR 2022
MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis · NeurIPS 2019
Audio and music processing
audio coding
1.022023
High-Fidelity Audio Compression with Improved RVQGAN · NeurIPS 2023
MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis · NeurIPS 2019
Audio and music processing › audio coding
neural audio codec
1.022023
High-Fidelity Audio Compression with Improved RVQGAN · NeurIPS 2023
MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis · NeurIPS 2019
Natural language and speech › Speech recognition and synthesis › speech synthesis
diffusion-based speech synthesis
0.912025
DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis · ICML 2025
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis
0.912025
DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis · ICML 2025
Natural language and speech › Speech recognition and synthesis › speech synthesis
zero-shot speech synthesis
0.912025
DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis · ICML 2025
Machine learning › Generative modeling
audio generation
0.712023
High-Fidelity Audio Compression with Improved RVQGAN · NeurIPS 2023
Audio and music processing
speech synthesis
0.412019
MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis · NeurIPS 2019
Audio and music processing › sound synthesis
music synthesis
0.322023
High-Fidelity Audio Compression with Improved RVQGAN · NeurIPS 2023
MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis · NeurIPS 2019
Machine learning › Deep learning architectures and training
recurrent neural network
0.312017
SampleRNN: An Unconditional End-to-End Neural Audio Generation Model · ICLR (Poster) 2017
Audio and music processing › sound synthesis
neural audio synthesis
0.312017
SampleRNN: An Unconditional End-to-End Neural Audio Generation Model · ICLR (Poster) 2017
Audio and music processing
sound synthesis
0.312017
SampleRNN: An Unconditional End-to-End Neural Audio Generation Model · ICLR (Poster) 2017
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
model distillation
0.312025
DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis · ICML 2025
Interaction techniques and input
voice interaction
0.312025
SpeakEasy: Enhancing Text-to-Speech Interactions for Expressive Content Creation · CHI 2025
Natural language and speech › Speech recognition and synthesis
speech synthesis
0.212023
High-Fidelity Audio Compression with Improved RVQGAN · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

vector quantization · 2.1adversarial loss · 2.1system design · 1.7residual vector quantization · 1.3chunked generation · 1.1autoregressive modeling · 1.1speaker verification · 0.9knowledge distillation · 0.9diffusion model · 0.9connectionist temporal classification · 0.9mel-spectrogram inversion · 0.8
YearPublicationVenuePosition
2025 SpeakEasy: Enhancing Text-to-Speech Interactions for Expressive Content Creation
abstract
CHI ’25, Yokohama, Japan
Stephen Brade, Sam Anderson, Rithesh Kumar, Zeyu Jin, Anh Truong
CHI3
2025 DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis
abstract
Diffusion models have demonstrated significant potential in speech synthesis tasks, including text-to-speech (TTS) and voice cloning. However, their iterative denoising processes are computationally intensive, and previous distillation attempts have shown consistent quality degradation. Moreover, existing TTS approaches are limited by non-differentiable components or iterative sampling that prevent true end-to-end optimization with perceptual metrics. We introduce DMOSpeech, a distilled diffusion-based TTS model that uniquely achieves both faster inference and superior performance compared to its teacher model. By enabling direct gradient pathways to all model components, we demonstrate the first successful end-to-end optimization of differentiable metrics in TTS, incorporating Connectionist Temporal Classification (CTC) loss and Speaker Verification (SV) loss. Our comprehensive experiments, validated through extensive human evaluation, show significant improvements in naturalness, intelligibility, and speaker similarity while reducing inference time by orders of magnitude. This work establishes a new framework for aligning speech synthesis with human auditory preferences through direct metric optimization. The audio samples are available at https://dmospeech.github.io/demo
Yinghao Aaron Li, Rithesh Kumar, Zeyu Jin
ICML2
2023 High-Fidelity Audio Compression with Improved RVQGAN
abstract
Language models have been successfully used to model natural signals, such as images, speech, and music. A key component of these models is a high quality neural compression model that can compress high-dimensional natural signals into lower dimensional discrete tokens. To that end, we introduce a high-fidelity universal neural audio compression algorithm that achieves ~90x compression of 44.1 KHz audio into tokens at just 8kbps bandwidth. We achieve this by combining advances in high-fidelity audio generation with better vector quantization techniques from the image domain, along with improved adversarial and reconstruction losses. We compress all domains (speech, environment, music, etc.) with a single universal model, making it widely applicable to generative modeling of all audio. We compare with competing audio compression algorithms, and find our method outperforms them significantly. We provide thorough ablations for every design choice, as well as open-source code and trained model weights. We hope our work can lay the foundation for the next generation of high-fidelity audio modeling.
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar
NeurIPS1
2022 Chunked Autoregressive GAN for Conditional Waveform Synthesis
Max Morrison, Rithesh Kumar, Prem Seetharaman, Aaron C. Courville, Yoshua Bengio
ICLR2
2019 MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
abstract
Previous works (Donahue et al., 2018a; Engel et al., 2019a) have found that generating coherent raw audio waveforms with GANs is challenging. In this paper, we show that it is possible to train GANs reliably to generate high quality coherent waveforms by introducing a set of architectural changes and simple training techniques. Subjective evaluation metric (Mean Opinion Score, or MOS) shows the effectiveness of the proposed approach for high quality mel-spectrogram inversion. To establish the generality of the proposed techniques, we show qualitative results of our model in speech synthesis, music domain translation and unconditional music synthesis. We evaluate the various components of the model through ablation studies and suggest a set of guidelines to design general purpose discriminators and generators for conditional sequence synthesis tasks. Our model is non-autoregressive, fully convolutional, with significantly fewer parameters than competing models and generalizes to unseen speakers for mel-spectrogram inversion. Our pytorch implementation runs at more than 100x faster than realtime on GTX 1080Ti GPU and more than 2x faster than real-time on CPU, without any hardware specific optimization tricks.
Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, Aaron C. Courville
NeurIPS2
2017 SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
Soroush Mehri, Ishaan Gulrajani, Rithesh Kumar, Jose Sotelo, Aaron C. Courville, Yoshua Bengio
ICLR (Poster)4