VLDB 2026 Research / reviewers in the wild / expert
Benjamin Attal
dblp:251/1743
· DBLP profile ↗
11ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-0132-5232ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Neural Inverse Rendering from Propagating LightabstractWe present the first system for physically based, neural inverse rendering from multi-viewpoint videos of propagating light. Our approach relies on a time-resolved extension of neural radiance caching — a technique that accelerates inverse rendering by storing infinite-bounce radiance arriving at any point from any direction. The resulting model accurately accounts for direct and indirect light transport effects and, when applied to captured measurements from a flash lidar system, enables state-of-the-art 3D reconstruction in the presence of strong indirect light. Further, we demonstrate view synthesis of propagating light, automatic decomposition of captured measurements into direct and indirect components, as well as novel capabilities such as multi-view time-resolved relighting of captured scenes. Anagh Malik, Benjamin Attal, Andrew Xie, Matthew O'Toole, David B. Lindell |
CVPR | 2 |
| 2025 | Towards Mixed-State Coded Diffraction ImagingabstractCoherent diffraction imaging (CDI) is a computational technique for reconstructing a complex-valued optical field from an intensity measurement. The approach is to illuminate an object with a coherent beam of light to form a diffraction pattern, and use a phase retrieval algorithm to reconstruct the object's complex transmittance from the measurement. However, as the name implies, conventional CDI assumes highly coherent illumination. Recent works therefore extend CDI to account for partial coherence and imperfect detection, by modeling light as an incoherent mixture of multiple fields (e.g., multiple wavelengths) and recovering each field simultaneously. In this work, we make strides towards the practical implementation and usage of multi-wavelength diffraction imaging. In particular, we provide novel analysis of the noise characteristics of multi-wavelength diffraction imaging, and show that it is preferable to coherent diffraction imaging under high signal-independent noise. Additionally, we present a compact coded diffraction imaging system and corresponding phase retrieval algorithms to robustly and simultaneously recover complex fields representing multiple wavelengths. Using a novel mixed-norm color prior, our prototype system reconstructs a larger number of multi-wavelength fields from fewer measurements than existing methods, and supports applications such as micron-scale optical path difference measurement via synthetic wavelength holography. Benjamin Attal, Matthew O'Toole |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Flash Cache: Reducing Bias in Radiance Cache Based Inverse Rendering
Benjamin Attal, Dor Verbin, Ben Mildenhall, Peter Hedman, Jonathan T. Barron, Matthew O'Toole, Pratul P. Srinivasan |
ECCV (28) | 1 |
| 2024 | Flowed Time of Flight Radiance Fields
Mikhail Okunev, Marc Mapeke, Benjamin Attal, Christian Richardt, Matthew O'Toole, James Tompkin 0001 |
ECCV (62) | 3 |
| 2024 | NeRF-Casting: Improved View-Dependent Appearance with Consistent Reflections
Dor Verbin, Pratul P. Srinivasan, Peter Hedman, Ben Mildenhall, Benjamin Attal, Richard Szeliski, Jonathan T. Barron |
SIGGRAPH Asia | 5 |
| 2023 | HyperReel: High-Fidelity 6-DoF Video with Ray-Conditioned SamplingabstractVolumetric scene representations enable photorealistic view synthesis for static scenes and form the basis of several existing 6-DoF video techniques. However, the volume rendering procedures that drive these representations necessitate careful trade-offs in terms of quality, rendering speed, and memory efficiency. In particular, existing methods fail to simultaneously achieve real-time performance, small memory footprint, and high-quality rendering for challenging real-world scenes. To address these issues, we present HyperReel―a novel 6-DoF video representation. The two core components of HyperReel are: (1) a ray-conditioned sample prediction network that enables high-fidelity, high frame rate rendering at high resolutions and (2) a compact and memory-efficient dynamic volume representation. Our 6-DoF video pipeline achieves the best performance compared to prior and contemporary approaches in terms of visual quality with small memory requirements, while also rendering at up to 18 frames-per-second at megapixel resolution without any custom CUDA code. Benjamin Attal, Jia-Bin Huang 0001, Christian Richardt, Michael Zollhöfer, Johannes Kopf 0001, Matthew O'Toole, Changil Kim 0001 |
CVPR | 1 |
| 2023 | Neural Fields for Structured LightingabstractWe present an image formation model and optimization procedure that combines the advantages of neural radiance fields and structured light imaging. Existing depth-supervised neural models rely on depth sensors to accurately capture the scene’s geometry. However, the depth maps recovered by these sensors can be prone to error, or even fail outright. Instead of depending on the fidelity of processed depth maps from a structured light system, a more principled approach is to explicitly model the raw structured light images themselves. Our proposed approach enables the estimation of high-fidelity depth maps, including for objects with complex material properties (e.g., partially-transparent surfaces). Besides computing depth, the raw structured light images also confer other useful radiometric cues, which enable predicting surface normals and decomposing scene appearance in terms of a direct, indirect, and ambient component. We evaluate our framework quantitatively and qualitatively on a range of real and synthetic scenes, and decompose scenes into their constituent components for novel views. Aarrushi Shandilya, Benjamin Attal, Christian Richardt, James Tompkin 0001, Matthew O'Toole |
ICCV | 2 |
| 2022 | Learning Neural Light Fields with Ray-Space EmbeddingabstractNeural radiance fields (NeRFs) produce state-of-the-art view synthesis results, but are slow to render, requiring hundreds of network evaluations per pixel to approximate a volume rendering integral. Baking NeRFs into explicit data structures enables efficient rendering, but results in large memory footprints and, in some cases, quality reduction. Additionally, volumetric representations for view synthesis often struggle to represent challenging view dependent effects such as distorted reflections and refractions. We present a novel neural light field representation that, in contrast to prior work, is fast, memory efficient, and excels at modeling complicated view dependence. Our method supports rendering with a single network evaluation per pixel for small baseline light fields and with only a few evaluations per pixel for light fields with larger baselines. At the core of our approach is a ray-space embedding network that maps 4D ray-space into an intermediate, interpolable latent space. Our method achieves state-of-the-art quality on dense forward-facing datasets such as the Stanford Light Field dataset. In addition, for forward-facing scenes with sparser inputs we achieve results that are competitive with NeRF-based approaches while providing a better speed/quality/memory trade-off with far fewer network evaluations. Benjamin Attal, Jia-Bin Huang 0001, Michael Zollhöfer, Johannes Kopf 0001, Changil Kim 0001 |
CVPR | 1 |
| 2021 | TöRF: Time-of-Flight Radiance Fields for Dynamic Scene View SynthesisabstractNeural networks can represent and accurately reconstruct radiance fields for static 3D scenes (e.g., NeRF). Several works extend these to dynamic scenes captured with monocular video, with promising performance. However, the monocular setting is known to be an under-constrained problem, and so methods rely on data-driven priors for reconstructing dynamic content. We replace these priors with measurements from a time-of-flight (ToF) camera, and introduce a neural representation based on an image formation model for continuous-wave ToF cameras. Instead of working with processed depth maps, we model the raw ToF sensor measurements to improve reconstruction quality and avoid issues with low reflectance regions, multi-path interference, and a sensor's limited unambiguous depth range. We show that this approach improves robustness of dynamic scene reconstruction to erroneous calibration and large motions, and discuss the benefits and limitations of integrating RGB+ToF sensors now available on modern smartphones. Benjamin Attal, Eliot Laidlaw, Aaron Gokaslan, Changil Kim 0001, Christian Richardt, James Tompkin 0001, Matthew O'Toole |
NeurIPS | 1 |
| 2020 | MatryODShka: Real-time 6DoF Video View Synthesis Using Multi-sphere Images
Benjamin Attal, Selena Ling, Aaron Gokaslan, Christian Richardt, James Tompkin 0001 |
ECCV (1) | 1 |
| 2019 | Portal-ble: Intuitive Free-hand Manipulation in Unbounded Smartphone-based Augmented RealityabstractSmartphone augmented reality (AR) lets users interact with physical and virtual spaces simultaneously. With 3D hand tracking, smartphones become apparatus to grab and move virtual objects directly. Based on design considerations for interaction, mobility, and object appearance and physics, we implemented a prototype for portable 3D hand tracking using a smartphone, a Leap Motion controller, and a computation unit. Following an experience prototyping procedure, 12 researchers used the prototype to help explore usability issues and define the design space. We identified issues in perception (moving to the object, reaching for the object), manipulation (successfully grabbing and orienting the object), and behavioral understanding (knowing how to use the smartphone as a viewport). To overcome these issues, we designed object-based feedback and accommodation mechanisms and studied their perceptual and behavioral effects via two tasks: picking up distant objects, and assembling a virtual house from blocks. Our mechanisms enabled significantly faster and more successful user interaction than the initial prototype in picking up and manipulating stationary and moving objects, with a lower cognitive load and greater user preference. The resulting system---Portal-ble---improves user intuition and aids free-hand interactions in mobile situations. Jiaju Ma, Benjamin Attal, Haoming Lai, James Tompkin 0001, John F. Hughes, Jeff Huang 0002 |
UIST | 4 |