VLDB 2026 Research / reviewers in the wild / expert
Ayush Saraf
dblp:272/7359
· DBLP profile ↗
9ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Textured Gaussians for Enhanced 3D Scene Appearance Modelingabstract3D Gaussian Splatting (3DGS) has emerged as the state-of-the-art 3D reconstruction technique, offering high-quality results with fast training and rendering. However, its expressivity is limited as pixels covered by the same Gaussian share identical colors aside from a Gaussian falloff scaling factor, and individual Gaussians can only represent simple ellipsoids geometrically. To overcome these limitations, we integrate texture and alpha mapping from traditional graphics with 3DGS. Our approach augments each Gaussian with alpha, RGB, or RGBA texture maps to model spatially varying color and opacity across each Gaussian’s extent. This allows Gaussians to represent richer texture patterns and geometric structures beyond single-color ellipsoids. Notably, alpha-only texture maps significantly improve Gaussian expressivity, while further augmenting with RGB texture maps achieve maximum expressivity. We validate our method on a wide variety of standard benchmark datasets and our own custom captures at both the object and scene levels, and demonstrate image quality improvements over existing methods while using a similar or lower number of Gaussians. Brian Chao, Hung-Yu Tseng, Lorenzo Porzi, Chen Gao 0003, Tuotuo Li, Qinbo Li, Ayush Saraf, Jia-Bin Huang 0001, Johannes Kopf 0001, Gordon Wetzstein, Changil Kim 0001 |
CVPR | 7 |
| 2024 | LTM: Lightweight Textured Mesh Extraction and Refinement of Large Unbounded Scenes for Efficient Storage and Real-Time RenderingabstractAdvancements in neural signed distance fields (SDFs) have enabled modeling 3D surface geometry from a set of 2D images of real-world scenes. Baking neural SDFs can extract explicit mesh with appearance baked into texture maps as neural features. The baked meshes still have a large memory footprint and require a powerful GPU for real-time rendering. Neural optimization of such large meshes with differentiable rendering pose significant challenges. We propose a method to produce optimized meshes for large unbounded scenes with low triangle budget and high fidelity of geometry and appearance. We achieve this by combining advancements in baking neural SDFs with classical mesh simplification techniques and proposing a joint appearance-geometry refinement step. The visual quality is comparable to or better than state-of-the-art neural meshing and baking methods with high geometric accuracy despite significant reduction in triangle count, making the produced meshes efficient for storage, transmission, and rendering on mobile hardware. We validate the effectiveness of the proposed method on large unbounded scenes from mip-NeRF 360, Tanks & Temples, and Deep Blending datasets, achieving at-par rendering quality with 73 x reduced triangles and 11 x reduction in memory footprint. Rajvi Shah, Qinbo Li, Yipeng Wang 0018, Ayush Saraf, Changil Kim 0001, Jia-Bin Huang 0001, Dinesh Manocha, Suhib Alsisan, Johannes Kopf 0001 |
CVPR | 5 |
| 2023 | Robust Dynamic Radiance FieldsabstractDynamic radiance field reconstruction methods aim to model the time-varying structure and appearance of a dynamic scene. Existing methods, however, assume that accurate camera poses can be reliably estimated by Structure from Motion (SfM) algorithms. These methods, thus, are unreliable as SfM algorithms often fail or produce erroneous poses on challenging videos with highly dynamic objects, poorly textured surfaces, and rotating camera motion. We address this robustness issue by jointly estimating the static and dynamic radiance fields along with the camera parameters (poses and focal length). We demonstrate the robustness of our approach via extensive quantitative and qualitative experiments. Our results show favorable performance over the state-of-the-art dynamic view synthesis methods. Yu-Lun Liu 0001, Chen Gao 0003, Andreas Meuleman, Hung-Yu Tseng, Ayush Saraf, Changil Kim 0001, Yung-Yu Chuang, Johannes Kopf 0001, Jia-Bin Huang 0001 |
CVPR | 5 |
| 2023 | OmnimatteRF: Robust Omnimatte with 3D Background ModelingabstractVideo matting has broad applications, from adding interesting effects to casually captured movies to assisting video production professionals. Matting with associated effects such as shadows and reflections has also attracted increasing research activity, and methods like Omnimatte have been proposed to separate dynamic foreground objects of interest into their own layers. However, prior works represent video backgrounds as 2D image layers, limiting their capacity to express more complicated scenes, thus hindering application to real-world videos. In this paper, we propose a novel video matting method, OmnimatteRF, that combines dynamic 2D foreground layers and a 3D background model. The 2D layers preserve the details of the subjects, while the 3D background robustly reconstructs scenes in real-world videos. Extensive experiments demonstrate that our method reconstructs scenes with better quality on various videos. Geng Lin, Chen Gao 0003, Jia-Bin Huang 0001, Changil Kim 0001, Yipeng Wang 0018, Matthias Zwicker, Ayush Saraf |
ICCV | 7 |
| 2022 | Boosting View Synthesis with Residual TransferabstractVolumetric view synthesis methods with neural representations, such as NeRF and NeX, have recently demonstrated high-quality novel view synthesis. However, optimizing these representations is slow, and even fully trained models cannot reproduce all fine details in the input views. We present a simple but effective technique to boost the rendering quality, which can be easily integrated with most view synthesis methods. The core idea is to transfer color resid-uals (the difference between the input images and their re-construction) from training views to novel views. We blend the residuals from multiple views using a heuristic weighting scheme depending on ray visibility and angular differ-ences. We integrate our technique with several state-of-the-art view synthesis methods and evaluate the Real Forward-facing and the Shiny datasets. Our results show that at about 1/10th the number of training iterations, we achieve the same rendering quality as fully converged NeRF and NeX models, and when applied to fully converged models, we significantly improve their rendering quality. Xuejian Rong, Jia-Bin Huang 0001, Ayush Saraf, Changil Kim 0001, Johannes Kopf 0001 |
CVPR | 3 |
| 2021 | AMICO: Amodal Instance Composition
Peiye Zhuang, Denis Demandolx, Ayush Saraf, Xuejian Rong, Changil Kim 0001, Jia-Bin Huang 0001 |
BMVC | 3 |
| 2021 | Dynamic View Synthesis from Dynamic Monocular VideoabstractWe present an algorithm for generating novel views at arbitrary viewpoints and any input time step given a monocular video of a dynamic scene. Our work builds upon recent advances in neural implicit representation and uses continuous and differentiable functions for modeling the time-varying structure and the appearance of the scene. We jointly train a time-invariant static NeRF and a time-varying dynamic NeRF, and learn how to blend the results in an unsupervised manner. However, learning this implicit function from a single video is highly ill-posed (with infinitely many solutions that match the input video). To resolve the ambiguity, we introduce regularization losses to encourage a more physically plausible solution. We show extensive quantitative and qualitative results of dynamic view synthesis from casually captured videos. Chen Gao 0003, Ayush Saraf, Johannes Kopf 0001, Jia-Bin Huang 0001 |
ICCV | 2 |
| 2020 | Flow-edge Guided Video Completion
Chen Gao 0003, Ayush Saraf, Jia-Bin Huang 0001, Johannes Kopf 0001 |
ECCV (12) | 2 |
| 2020 | One shot 3D photographyabstract3D photography is a new medium that allows viewers to more fully experience a captured moment. In this work, we refer to a 3D photo as one that displays parallax induced by moving the viewpoint (as opposed to a stereo pair with a fixed viewpoint). 3D photos are static in time, like traditional photos, but are displayed with interactive parallax on mobile or desktop screens, as well as on Virtual Reality devices, where viewing it also includes stereo. We present an end-to-end system for creating and viewing 3D photos, and the algorithmic and design choices therein. Our 3D photos are captured in a single shot and processed directly on a mobile device. The method starts by estimating depth from the 2D input image using a new monocular depth estimation network that is optimized for mobile devices. It performs competitively to the state-of-the-art, but has lower latency and peak memory consumption and uses an order of magnitude fewer parameters. The resulting depth is lifted to a layered depth image, and new geometry is synthesized in parallax regions. We synthesize color texture and structures in the parallax regions as well, using an inpainting network, also optimized for mobile devices, on the LDI directly. Finally, we convert the result into a mesh-based representation that can be efficiently transmitted and rendered even on low-end devices and over poor network connections. Altogether, the processing takes just a few seconds on a mobile device, and the result can be instantly viewed and shared. We perform extensive quantitative evaluation to validate our system and compare its new components against the current state-of-the-art. Johannes Kopf 0001, Kevin Matzen, Suhib Alsisan, Ocean Quigley, Francis Ge, Yangming Chong, Josh Patterson, Jan-Michael Frahm, Matthew Yu, Peizhao Zhang, Peter Vajda, Ayush Saraf, Michael F. Cohen |
ACM Trans. Graph. | 14 |