VLDB 2026 Research / reviewers in the wild / expert
Tanmay Shah
dblp:344/1917
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
3D vision · 60% Generative modeling · 40% | |
| Computer graphics and multimedia
1 paper |
Rendering · 50% Computational photography and imaging · 25% Geometric modeling and processing · 25% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
novel view synthesis |
1.4 | 2 | 2024 | ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image · CVPR 2024 Preface: A Data-driven Volumetric Prior for Few-shot Ultra High-resolution Face Synthesis · ICCV 2023 |
Machine learning › Generative modeling › diffusion model
3d-aware diffusion |
0.8 | 1 | 2024 | ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image · CVPR 2024 |
Computer vision › 3D vision
3d scene reconstruction |
0.8 | 1 | 2024 | ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image · CVPR 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image · CVPR 2024 |
Computer vision › 3D vision › novel view synthesis
single-image novel view synthesis |
0.8 | 1 | 2024 | ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image · CVPR 2024 |
Geometric modeling and processing
3d face modeling |
0.8 | 1 | 2024 | Cafca: High-quality Novel View Synthesis of Expressive Faces from Casual Few-shot Captures · SIGGRAPH Asia 2024 |
Computational photography and imaging
3d vision |
0.8 | 1 | 2024 | Cafca: High-quality Novel View Synthesis of Expressive Faces from Casual Few-shot Captures · SIGGRAPH Asia 2024 |
Rendering
neural radiance fields |
0.8 | 1 | 2024 | Cafca: High-quality Novel View Synthesis of Expressive Faces from Casual Few-shot Captures · SIGGRAPH Asia 2024 |
Rendering
novel view synthesis |
0.8 | 1 | 2024 | Cafca: High-quality Novel View Synthesis of Expressive Faces from Casual Few-shot Captures · SIGGRAPH Asia 2024 |
Machine learning › Generative modeling
face synthesis |
0.7 | 1 | 2023 | Preface: A Data-driven Volumetric Prior for Few-shot Ultra High-resolution Face Synthesis · ICCV 2023 |
Computer vision › 3D vision
neural radiance field |
0.7 | 1 | 2023 | Preface: A Data-driven Volumetric Prior for Few-shot Ultra High-resolution Face Synthesis · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
fine-tuning · 1.5conditional NeRF · 1.53d morphable face model · 1.5score distillation sampling · 0.8camera conditioning · 0.8landmark-based 3d alignment · 0.7identity-conditioned NeRF · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single ImageabstractWe introduce a 3D-aware diffusion model, ZeroNVS, for single-image novel view synthesis for in-the-wild scenes. While existing methods are designed for single objects with masked backgrounds, we propose new techniques to address challenges introduced by in-the-wild multi-object scenes with complex backgrounds. Specifically, we train a generative prior on a mixture of data sources that capture object-centric, indoor, and outdoor scenes. To address issues from data mixture such as depth-scale ambiguity, we propose a novel camera conditioning parameterization and normalization scheme. Further, we observe that Score Distillation Sampling (SDS) tends to truncate the distribution of complex backgrounds during distillation of 360-degree scenes, and propose “SDS anchoring” to improve the diversity of synthesized novel views. Our model sets a new state-of-the-art result in LPIPS on the DTU dataset in the zero-shot setting, even outperforming methods specifically trained on DTU. We further adapt the challenging Mip-NeRF 360 dataset as a new benchmark for single-image novel view synthesis, and demonstrate strong performance in this setting. Code and models are available at this url. Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann, Hong-Xing Yu, Eric R. Chan, Dmitry Lagun, Li Fei-Fei 0001, Deqing Sun, Jiajun Wu 0001 |
CVPR | 3 |
| 2024 | Cafca: High-quality Novel View Synthesis of Expressive Faces from Casual Few-shot CapturesabstractVolumetric modeling and neural radiance field representations have revolutionized 3D face capture and photorealistic novel view synthesis. However, these methods often require hundreds of multi-view input images and are thus inapplicable to cases with less than a handful of inputs. We present a novel volumetric prior on human faces that allows for high-fidelity expressive face modeling from as few as three input views captured in the wild. Our key insight is that an implicit prior trained on synthetic data alone can generalize to extremely challenging real-world identities and expressions and render novel views with fine idiosyncratic details like wrinkles and eyelashes. We leverage a 3D Morphable Face Model to synthesize a large training set, rendering each identity with different expressions, hair, clothing, and other assets. We then train a conditional Neural Radiance Field prior on this synthetic dataset and, at inference time, fine-tune the model on a very sparse set of real images of a single subject. On average, the fine-tuning requires only three inputs to cross the synthetic-to-real domain gap. The resulting personalized 3D model reconstructs strong idiosyncratic facial expressions and outperforms the state-of-the-art in high-quality novel view synthesis of faces from sparse inputs in terms of perceptual and photo-metric quality. Marcel C. Bühler, Gengyan Li 0001, Erroll Wood, Leonhard Helminger, Xu Chen 0025, Tanmay Shah, Daoye Wang, Stephan J. Garbin, Sergio Orts, Otmar Hilliges, Dmitry Lagun, Jérémy Riviere, Paulo F. U. Gotardo, Thabo Beeler, Abhimitra Meka, Kripasindhu Sarkar |
SIGGRAPH Asia | 6 |
| 2024 | TEGLO: High Fidelity Canonical Texture Mapping from Single-View ImagesabstractRecent work in Neural Fields (NFs) learn 3D representations from class-specific single view image collections. However, they are unable to reconstruct the input data preserving high-frequency details. Further, these methods do not disentangle appearance from geometry and hence are not suitable for tasks such as texture transfer and editing. In this work, we propose TEGLO (Textured EG3D-GLO) for learning 3D representations from single view in-the-wild image collections for a given class of objects. We accomplish this by training a conditional Neural Radiance Field (NeRF) without any explicit 3D supervision. We equip our method with editing capabilities by creating a dense correspondence mapping to a 2D canonical space. We demonstrate that such mapping enables texture transfer and texture editing without requiring meshes with shared topology. Our key insight is that by mapping the input image pixels onto the texture space we can achieve near perfect reconstruction (≥ 74 dB PSNR at 10242resolution). Our formulation allows for high quality 3D consistent novel view synthesis with high-frequency details even at megapixel image resolutions. Project Page: teglo-nerf.github.io Vishal Vinod, Tanmay Shah, Dmitry Lagun |
WACV | 2 |
| 2023 | Preface: A Data-driven Volumetric Prior for Few-shot Ultra High-resolution Face SynthesisabstractNeRFs have enabled highly realistic synthesis of human faces including complex appearance and reflectance effects of hair and skin. These methods typically require a large number of multi-view input images, making the process hardware intensive and cumbersome, limiting applicability to unconstrained settings. We propose a novel volumetric human face prior that enables the synthesis of ultra high-resolution novel views of subjects that are not part of the prior’s training distribution. This prior model consists of an identity-conditioned NeRF, trained on a dataset of low-resolution multi-view images of diverse humans with known camera calibration. A simple sparse landmark-based 3D alignment of the training dataset allows our model to learn a smooth latent space of geometry and appearance despite a limited number of training identities. A high-quality volumetric representation of a novel subject can be obtained by model fitting to 2 or 3 camera views of arbitrary resolution. Importantly, our method requires as few as two views of casually captured images as input at inference time. Marcel C. Bühler, Kripasindhu Sarkar, Tanmay Shah, Gengyan Li 0001, Daoye Wang, Leonhard Helminger, Sergio Orts, Dmitry Lagun, Otmar Hilliges, Thabo Beeler, Abhimitra Meka |
ICCV | 3 |