Ravi Gadde

dblp:230/3603 · also Ravi Teja Gadde · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021
YearPublicationVenuePosition
2025 Audio-Visual Representation Learning For Lip-Sync Estimation Through Ranking Augmented Contrastive Training
abstract
In many applications, particularly in media production and content localization, it is crucial to detect and evaluate varying degrees of audio-visual synchronization, such as selecting high-quality dubbed audio over poorly synchronized tracks. Traditional contrastively pre-trained LipSync models are designed to distinguish perfectly synced audio from unsynced audio. However, these models fall short when it comes to detecting partial synchronization, such as in dubbed audio, because their training objective is focused on pulling synced lip-motion and audio closer together while pushing everything else apart. This approach limits their ability to accurately gauge varying levels of sync, leading to challenges in scenarios that require a more nuanced understanding of synchronization quality. To address this limitation, we propose a novel deep metric learning approach, the Ranking Supervised Multi-Similarity (RSMS) loss formulation, which introduces a ranking prior as a supervision signal. Our method integrates hard-sample mining to enforce this ranking, allowing the model to better differentiate between partial-syncs and completely unsynced audios. Furthermore, we demonstrate the effectiveness of using “Dubbed Audio” as a train-time example of partial-syncs, leading to improved performance in lip-sync models.
Bhavin Jawade, Ravi Gadde, Christophe Bejjani, Yinghong Lan
ICASSP2
2024 Diff2Lip: Audio Conditioned Diffusion Models for Lip-Synchronization
abstract
The task of lip synchronization (lip-sync) seeks to match the lips of human faces with different audio. It has various applications in the film industry as well as for creating virtual avatars and for video conferencing. This is a challenging problem as one needs to simultaneously introduce detailed, realistic lip movements while preserving the identity, pose, emotions, and image quality. Many of the previous methods trying to solve this problem suffer from image quality degradation due to a lack of complete contextual information. In this paper, we present Diff2Lip, an audio-conditioned diffusion-based model which is able to do lip synchronization in-the-wild while preserving these qualities. We train our model on Voxceleb2, a video dataset containing in-the-wild talking face videos. Extensive studies show that our method outperforms popular methods like Wav2Lip and PC-AVS in Fréchet inception distance (FID) metric and Mean Opinion Scores (MOS) of the users. We show results on both reconstruction (same audio-video inputs) as well as cross (different audio-video inputs) settings on Voxceleb2 and LRW datasets. Video results are available at https://soumik-kanad.github.io/diff2lip.
Soumik Mukhopadhyay 0001, Saksham Suri, Ravi Gadde, Abhinav Shrivastava
WACV3
2023 SIDGAN: High-Resolution Dubbed Video Generation via Shift-Invariant Learning
abstract
Dubbed video generation aims to accurately synchronize mouth movements of a given facial video with driving audio while preserving identity and scene-specific visual dynamics, such as head pose and lighting. Despite the accurate lip generation of previous approaches that adopts a pre-trained audio-video synchronization metric as an objective function, called Sync-Loss, extending it to high-resolution videos was challenging due to shift biases in the loss landscape that inhibit tandem optimization of Sync-Loss and visual quality, leading to a loss of detail.To address this issue, we introduce shift-invariant learning, which generates photo-realistic high-resolution videos with accurate Lip-Sync. Further, we employ a pyramid network with coarse-to-fine image generation to improve stability and lip syncronization. Our model outperforms state-of-the-art methods on multiple benchmark datasets, including AVSpeech, HDTF, and LRW, in terms of photo-realism, identity preservation, and Lip-Sync accuracy.
Urwa Muaz, Wondong Jang, Rohun Tripathi, Santhosh Mani, Wenbin Ouyang, Ravi Gadde, Baris Gecer, Sergio Elizondo, Reza Madad, Naveen Nair
ICCV6
2022 Domain Prompts: Towards memory and compute efficient domain adaptation of ASR systems
abstract
Automatic Speech Recognition (ASR) systems have found their use in numerous industrial applications in very diverse domains creating a need to adapt to new domains with small memory and deployment overhead. In this work, we introduce domain prompts, a methodology that involves training a small number of domain embedding parameters to prime a Transformer-based Language Model (LM) to a particular domain. Using this domain-adapted LM for rescoring ASR hypotheses can achieve 7-13% WER reduction for a new domain with just 1000 unlabeled textual domain-specific sentences and a handful of additional parameters. Our method can match or even beat the performance of models fully fine-tuned towards a particular domain with only 0.02% of the parameters. Given the parameter efficiency and negligible deployment overhead, our experiments showcase that such a method is an ideal choice for on-the-fly adaptation of LMs used in ASR systems to progressively scale it to new domains.
Saket Dingliwal, Ashish Shenoy, Sravan Babu Bodapati, Ankur Gandhe, Ravi Gadde, Katrin Kirchhoff
INTERSPEECH5
2022 RefTextLAS: Reference Text Biased Listen, Attend, and Spell Model For Accurate Reading Evaluation
Phani S. Nidadavolu, Nick Jutila, Ravi Gadde, Aswarth Abhilash Dara, Joseph Savold, Sapan Patel, Aaron Hoff, Veerdhawal Pande, Kevin Crews, Ankur Gandhe, Ariya Rastrow, Roland Maas
INTERSPEECH4
2019 Jasper: An End-to-End Convolutional Neural Acoustic Model
abstract
In this paper we report state-of-the-art results on LibriSpeech among end-to-end speech recognition models without any external training data.Our model, Jasper, uses only 1D convolutions, batch normalization, ReLU, dropout, and residual connections.To improve training, we further introduce a new layer-wise optimizer called NovoGrad.Through experiments, we demonstrate that the proposed deep architecture performs as well or better than more complex choices.Our deepest Jasper variant uses 54 convolutional layers.With this architecture, we achieve 2.95% WER using a beam-search decoder with an external neural language model and 3.86% WER with a greedy decoder on LibriSpeech test-clean.We also report competitive results on Wall Street Journal and the Hub5'00 conversational evaluation datasets.
Jason Li 0007, Vitaly Lavrukhin, Boris Ginsburg, Ryan Leary, Oleksii Kuchaiev, Jonathan M. Cohen, Ravi Gadde
INTERSPEECH8