Jesus Zarzar

dblp:237/9581 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-4966-9621ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 UnCommon Objects in 3D
abstract
We introduce Uncommon Objects in 3D (uCO3D), a new object-centric dataset for 3D deep learning and 3D generative AI. uCO3D is the largest publicly-available collection of high-resolution videos of objects with 3D annotations that ensures full-360° coverage. uCO3D is significantly more diverse than MVImgNet and CO3Dv2, covering more than 1,000 object categories. It is also of higher quality, due to extensive quality checks of both the collected videos and the 3D annotations. Similar to analogous datasets, uCO3D contains annotations for 3D camera poses, depth maps and sparse point clouds. In addition, each object is equipped with a caption and a 3D Gaussian Splat reconstruction. We train several large 3D models on MVImgNet, CO3Dv2, and uCO3D and obtain superior results using the latter, showing that uCO3D is better for learning applications.
Piyush Tayal, Jesus Zarzar, Tom Monnier, Konstantinos Tertikas, Jiali Duan, Antoine Toisoul, Jason Y. Zhang 0001, Natalia Neverova, Andrea Vedaldi, Roman Shapovalov, David Novotný
CVPR4
2025 Twinner: Shining Light on Digital Twins in a Few Snaps
abstract
We present Twinner, the first large reconstruction model capable of recovering a scene’s illumination as well as an object’s geometry and material properties from only a few posed images. Twinner is based on the Large Reconstruction Model and innovates in three key ways: 1)We introduce a memory-efficient voxel-grid transformer whose memory scales only quadratically with the size of the voxel grid. 2) To address the scarcity of high-quality ground-truth PBR- shaded models, we introduce a large, fully synthetic dataset of procedurally generated PBR-textured objects lit with varied illumination. 3) To bridge the synthetic-to-real gap, we finetune the model on real-world datasets using a differentiable physically based shading model, eliminating the need for ground-truth illumination or material properties, which are challenging to obtain in real-world scenarios. We demonstrate the efficacy of our model on the real-world StanfordORB benchmark, where, given a few input views, we achieve reconstruction quality significantly superior to existing feed-forward reconstruction networks and comparable to slower per-scene optimization methods.
Jesus Zarzar, Tom Monnier, Roman Shapovalov, Andrea Vedaldi, David Novotný
CVPR1
2025 A compact stochastic representation for Monte Carlo Path Traced images
abstract
We present a compact, learning-based representation that captures the full Monte Carlo sampling distribution of a rendered image. Our approach enables rendering at arbitrary samples per pixel (SPP) during inference without requiring expensive path tracing operations. This is achieved by fitting parametric distributions to per-pixel radiance values, which can be efficiently estimated, stored, and sampled. Our method proceeds in three stages. First, we map radiance samples into radial log space, which encourages Gaussian-like distributions while preserving angular relationships. Second, we fit each pixel’s distribution using 3D Gaussian Mixture Models (GMMs), trained online with minimal memory overhead, making the approach compatible with standard path tracers. For inference, we introduce an optimized sampling scheme whose complexity is independent of the target SPP, enabling fast synthesis of high-SPP images. Additionally, we demonstrate that the learned representations can be heavily compressed using quantization and codebook techniques with negligible quality loss. Experiments show that GMMs strike an effective balance between expressiveness and sparsity. Compared to alternative models, our method better captures pixel-wise Monte Carlo distributions. Lastly, we illustrate the versatility of our representation with applications such as firefly rejection and ray-distribution-driven denoising.
Matthias Treder, Pavlos Makridis, Alexis Lechat, Jesus Zarzar, Marina Villanueva Barreiro, Roc R. Currius
SIGGRAPH Asia4
2024 TrackNeRF: Bundle Adjusting NeRF from Sparse and Noisy Views via Feature Tracks
Jinjie Mai, Wenxuan Zhu, Sara Rojas 0001, Jesus Zarzar, Abdullah Hamdi, Guocheng Qian, Bing Li 0024, Silvio Giancola, Bernard Ghanem
ECCV (12)4
2024 SplitNeRF: Split Sum Approximation Neural Field for Joint Geometry, Illumination, and Material Estimation
abstract
We present a novel approach for digitizing real-world objects by estimating their geometry, material properties, and environmental lighting from a set of posed images with fixed lighting. Our method incorporates into Neural Radiance Field (NeRF) pipelines the split sum approximation used with image-based lighting for real-time physically based rendering. We propose modeling the scene's lighting with a single scene-specific MLP representing pre-integrated image-based lighting at arbitrary resolutions. We accurately model pre-integrated lighting by exploiting a novel regularizer based on efficient Monte Carlo sampling. Additionally, we propose a new method of supervising self-occlusion predictions by exploiting a similar regularizer based on Monte Carlo sampling. Experimental results demonstrate the efficiency and effectiveness of our approach in estimating scene geometry, material properties, and lighting. Our method attains state-of-the-art relighting quality after only ${\sim}1$ hour of training in a single NVIDIA A100 GPU.
Jesus Zarzar, Bernard Ghanem
NeurIPS1
2023 Re-ReND: Real-time Rendering of NeRFs across Devices
abstract
This paper proposes a novel approach for rendering a pre-trained Neural Radiance Field (NeRF) in real-time on resource-constrained devices. We introduce Re-ReND, a method enabling Real-time Rendering of NeRFs across Devices. Re-ReND is designed to achieve real-time performance by converting the NeRF into a representation that can be efficiently processed by standard graphics pipelines. The proposed method distills the NeRF by extracting the learned density into a mesh, while the learned color information is factorized into a set of matrices that represent the scene’s light field. Factorization implies the field is queried via inexpensive MLP-free matrix multiplications, while using a light field allows rendering a pixel by querying the field a single time—as opposed to hundreds of queries when employing a radiance field. Since the proposed representation can be implemented using a fragment shader, it can be directly integrated with standard rasterization frameworks. Our flexible implementation can render a NeRF in real-time with low memory requirements and on a wide range of resource-constrained devices, including mobiles and AR/VR headsets. Notably, we find that Re-ReND can achieve over a 2.6-fold increase in rendering speed versus the state-of-the-art without perceptible losses in quality.
Sara Rojas 0001, Jesus Zarzar, Juan C. Pérez, Artsiom Sanakoyeu, Ali K. Thabet, Albert Pumarola, Bernard Ghanem
ICCV2
2019 Leveraging Shape Completion for 3D Siamese Tracking
abstract
Point clouds are challenging to process due to their sparsity, therefore autonomous vehicles rely more on appearance attributes than pure geometric features. However, 3D LIDAR perception can provide crucial information for urban navigation in challenging light or weather conditions. In this paper, we investigate the versatility of Shape Completion for 3D Object Tracking in LIDAR point clouds. We design a Siamese tracker that encodes model and candidate shapes into a compact latent representation. We regularize the encoding by enforcing the latent representation to decode into an object model shape. We observe that 3D object tracking and 3D shape completion complement each other. Learning a more meaningful latent representation shows better discriminatory capabilities, leading to improved tracking performance. We test our method on the KITTI Tracking set using car 3D bounding boxes. Our model reaches a 76.94% Success rate and 81.38% Precision for 3D Object Tracking, with the shape completion regularization leading to an improvement of 3% in both metrics.
Silvio Giancola, Jesus Zarzar, Bernard Ghanem
CVPR2