VLDB 2026 Research / reviewers in the wild / expert
Madhav Agarwal
dblp:273/4306
· DBLP profile ↗
6ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0001-8267-1024ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Generative modeling · 70% Trustworthy machine learning · 30% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Understanding the Generalization of Pretrained Diffusion Models on Out-of-Distribution Data · AAAI 2024 |
Machine learning › Generative modeling › image generation › data-efficient image generation
few-shot image generation |
0.8 | 1 | 2024 | Understanding the Generalization of Pretrained Diffusion Models on Out-of-Distribution Data · AAAI 2024 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.8 | 1 | 2024 | Understanding the Generalization of Pretrained Diffusion Models on Out-of-Distribution Data · AAAI 2024 |
Machine learning › Generative modeling
generative adversarial network |
0.2 | 1 | 2024 | Understanding the Generalization of Pretrained Diffusion Models on Out-of-Distribution Data · AAAI 2024 |
Methods — techniques the papers use, named apart from their topics
latent inversion · 0.8geodesic traversal · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian SplattingabstractSpeech-driven talking heads have recently emerged and enable interactive avatars. However, real-world applications are limited, as current methods achieve high visual fidelity but slow or fast yet temporally unstable. Diffusion methods provide realistic image generation, yet struggle with one-shot settings. Gaussian Splatting approaches are real-time, yet inaccuracies in facial tracking, or inconsistent Gaussian mappings, lead to unstable outputs and video artifacts that are detrimental to realistic use cases. We address this problem by mapping Gaussian Splatting using 3D Morphable Models to generate person-specific avatars. We introduce transformer-based prediction of model parameters, directly from audio, to drive temporal consistency. From monocular video and independent audio speech inputs, our method enables generation of real-time talking head videos where we report competitive quantitative and qualitative performance. Madhav Agarwal, Mingtian Zhang, Laura Sevilla-Lara, Steven McDonagh 0001 |
WACV | 1 |
| 2024 | Understanding the Generalization of Pretrained Diffusion Models on Out-of-Distribution DataabstractThis work tackles the important task of understanding out-of-distribution behavior in two prominent types of generative models, i.e., GANs and Diffusion models. Understanding this behavior is crucial in understanding their broader utility and risks as these systems are increasingly deployed in our daily lives. Our first contribution is demonstrating that diffusion spaces outperform GANs' latent spaces in inverting high-quality OOD images. We also provide a theoretical analysis attributing this to the lack of prior holes in diffusion spaces. Our second significant contribution is to provide a theoretical hypothesis that diffusion spaces can be projected onto a bounded hypersphere, enabling image manipulation through geodesic traversal between inverted images. Our analysis shows that different geodesics share common attributes for the same manipulation, which we leverage to perform various image manipulations. We conduct thorough empirical evaluations to support and validate our claims. Finally, our third and final contribution introduces a novel approach to the few-shot sampling for out-of-distribution data by inverting a few images to sample from the cluster formed by the inverted latents. The proposed technique achieves state-of-the-art results for the few-shot generation task in terms of image quality. Our research underscores the promise of diffusion spaces in out-of-distribution imaging and offers avenues for further exploration. Please find more details about our project at \url{http://cvit.iiit.ac.in/research/projects/cvit-projects/diffusionOOD} Sai Niranjan Ramachandran, Rudrabha Mukhopadhyay, Madhav Agarwal, C. V. Jawahar, Vinay P. Namboodiri |
AAAI | 3 |
| 2023 | Audio-Visual Face ReenactmentabstractThis work proposes a novel method to generate realistic talking head videos using audio and visual streams. We animate a source image by transferring head motion from a driving video using a dense motion field generated using learnable keypoints. We improve the quality of lip sync using audio as an additional input, helping the network to attend to the mouth region. We use additional priors using face segmentation and face mesh to improve the structure of the reconstructed faces. Finally, we improve the visual quality of the generations by incorporating a carefully designed identity-aware generator module. The identity-aware generator takes the source image and the warped motion features as input to generate a high-quality output with fine-grained details. Our method produces state-of-the-art results and generalizes well to unseen faces, languages, and voices. We comprehensively evaluate our approach using multiple metrics and outperforming the current techniques both qualitative and quantitatively. Our work opens up several applications, including enabling low bandwidth video calls. We release a demo video and additional information at http://cvit.iiit.ac.in/research/projects/cvit-projects/avfr. Madhav Agarwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. Jawahar |
WACV | 1 |
| 2023 | Dataset agnostic document object detection
Ajoy Mondal, Madhav Agarwal, C. V. Jawahar |
Pattern Recognit. | 2 |
| 2022 | Compressing Video Calls using Synthetic Talking Heads
Madhav Agarwal, Anchit Gupta, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. Jawahar |
BMVC | 1 |
| 2020 | CDeC-Net: Composite Deformable Cascade Network for Table Detection in Document ImagesabstractLocalizing page elements/objects such as tables, figures, equations, etc. is the primary step in extracting information from document images. We propose a novel end-to-end trainable deep network, (cnec-xet) for detecting tables present in the documents. The proposed network consists of a multistage extension of Mask R-CNN with a dual backbone having deformable convolution for detecting tables varying in scale with high detection accuracy at higher IoU threshold. We empirically evaluate CDeC-Net on the publicly available benchmark datasets with extensive experiments. Our solution has three important properties: (i) a single trained model CDeC-Net‡that performs well across all the popular benchmark datasets; (ii) we report excellent performances across multiple, including higher, thresholds of IoU; (iii) by following the same protocol of the recent papers for each of the benchmarks, we consistently demonstrate the superior quantitative performance. Our code and models are publicly available at https://github.com/mdv3101/CDeCNet for enabling reproducibility of the results. Madhav Agarwal, Ajoy Mondal, C. V. Jawahar |
ICPR | 1 |