Matthew O'Toole

dblp:76/1129 · DBLP profile ↗
← Back
50ranked-venue papers
7as first author
29since 2021 · last 2026
0000-0002-0740-9349ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 6 first-author · 26 since 2021Artificial intelligence and machine learning · 34 · 3 first-author · 21 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 RadarSim: Simulating Single-Chip Radar Via Multimodal Neural Fields
abstract
Radars are an ideal complement to cameras: both are inexpensive, solid-state sensors, with cameras offering fine angular resolution, while radars provide metric depth and robustness under adverse weather. However, radar data is more difficult to interpret than camera images and varies significantly between sensors, necessitating increased reliance on simulation for prototyping sensors and processing pipelines. Recent work treating radar reconstruction as a novel view synthesis problem has shown great promise in reconstructing radar-relevant geometry and simulating lowlevel radar data. However, such methods are constrained by the low spatial resolution of the underlying radar. To address this, we propose a unified differentiable renderer, RadarSim, which leverages the high angular resolution of RGB cameras to generate Doppler radar range images from a camera-initialized neural field. Using a novel data set of calibrated radar camera recordings from a custom handheld rig, we demonstrate that RadarSim produces sharper geometry and Doppler range frames than radar-only reconstructions.
Chuhan Chen, Tianshu Huang, Akarsh Prabhakara, Chaithanya Kumar Mummadi, Zhongxiao Cong, Anthony Rowe 0001, Matthew O'Toole, Deva Ramanan
3DV7
2025 Time of the Flight of the Gaussians: Optimizing Depth Indirectly in Dynamic Radiance Fields
abstract
We present a method to reconstruct dynamic scenes from monocular continuous-wave time-of-flight (C-ToF) cameras using raw sensor samples that achieves similar or better accuracy than neural volumetric approaches and is 100× faster. Quickly achieving high-fidelity dynamic 3D reconstruction from a single viewpoint is a significant challenge in computer vision. In C-ToF radiance field reconstruction, the property of interest—depth—is not directly measured, causing an additional challenge. This problem has a large and underappreciated impact upon the optimization when using a fast primitive-based scene representation like 3D Gaussian splatting, which is commonly used with multi-view data to produce satisfactory results and is brittle in its optimization otherwise. We incorporate two heuristics into the optimization to improve the accuracy of scene geometry represented by Gaussians. Experimental results show that our approach produces accurate reconstructions under constrained C-ToF sensing conditions, including for fast motions like swinging baseball bats. https://visual.cs.brown.edu/gftorf
Runfeng Li, Mikhail Okunev, Anh Ha Duong, Christian Richardt, Matthew O'Toole, James Tompkin 0001
CVPR6
2025 Neural Inverse Rendering from Propagating Light
abstract
We present the first system for physically based, neural inverse rendering from multi-viewpoint videos of propagating light. Our approach relies on a time-resolved extension of neural radiance caching — a technique that accelerates inverse rendering by storing infinite-bounce radiance arriving at any point from any direction. The resulting model accurately accounts for direct and indirect light transport effects and, when applied to captured measurements from a flash lidar system, enables state-of-the-art 3D reconstruction in the presence of strong indirect light. Further, we demonstrate view synthesis of propagating light, automatic decomposition of captured measurements into direct and indirect components, as well as novel capabilities such as multi-view time-resolved relighting of captured scenes.
Anagh Malik, Benjamin Attal, Andrew Xie, Matthew O'Toole, David B. Lindell
CVPR4
2025 Towards Foundational Models for Single-Chip Radar
abstract
mmWave radars are compact, inexpensive, and durable sensors that are robust to occlusions and work regardless of environmental conditions, such as weather and darkness. However, this comes at the cost of poor angular resolution, especially for inexpensive single-chip radars, which are typically used in automotive and indoor sensing applications. Although many have proposed learning-based methods to mitigate this weakness, no standardized foundational models or large datasets for the mmWave radar have emerged, and practitioners have largely trained task-specific models from scratch using relatively small datasets. In this paper, we collect (to our knowledge) the largest available raw radar dataset with 1M samples (29 hours) and train a foundational model for 4D single-chip radar, which can predict 3D occupancy and semantic segmentation with quality that is typically only possible with much higher resolution sensors. We demonstrate that our Generalizable Radar Transformer (GRT) generalizes across diverse settings, can be fine-tuned for different tasks, and shows logarithmic data scaling of 20\% per $10\times$ data. We also run extensive ablations on common design decisions, and find that using raw radar data significantly outperforms widely-used lossy representations, equivalent to a $10\times$ increase in training data. Finally, we roughly estimate that $\approx$100M samples (3000 hours) of data are required to fully exploit the potential of GRT.
Tianshu Huang, Akarsh Prabhakara, Chuhan Chen, Jay Karhade, Deva Ramanan, Matthew O'Toole, Anthony Rowe 0001
ICCV6
2025 Spatially-Varying Autofocus
Yingsi Qin, Aswin C. Sankaranarayanan, Matthew O'Toole
ICCV3
2025 Automated design of compound lenses with discrete-continuous optimization
abstract
We introduce a method that automatically and jointly updates both continuous and discrete parameters of a compound lens design, to improve its performance in terms of sharpness, speed, or both. Previous methods for compound lens design use gradient-based optimization to update continuous parameters (e.g., curvature of individual lens elements) of a given lens topology, requiring extensive expert intervention to realize topology changes. By contrast, our method can additionally optimize discrete parameters such as number and type (e.g., singlet or doublet) of lens elements. Our method achieves this capability by combining gradient-based optimization with a tailored Markov chain Monte Carlo sampling algorithm, using transdimensional mutation and paraxial projection operations for efficient global exploration. We show experimentally on a variety of lens design tasks that our method effectively explores an expanded design space of compound lenses, producing better designs than previous methods and pushing the envelope of speed-sharpness tradeoffs achievable by automated lens design.
Arjun Teh, Delio Vicini, Bernd Bickel, Ioannis Gkioulekas, Matthew O'Toole
SIGGRAPH Asia5
2025 Towards Mixed-State Coded Diffraction Imaging
abstract
Coherent diffraction imaging (CDI) is a computational technique for reconstructing a complex-valued optical field from an intensity measurement. The approach is to illuminate an object with a coherent beam of light to form a diffraction pattern, and use a phase retrieval algorithm to reconstruct the object's complex transmittance from the measurement. However, as the name implies, conventional CDI assumes highly coherent illumination. Recent works therefore extend CDI to account for partial coherence and imperfect detection, by modeling light as an incoherent mixture of multiple fields (e.g., multiple wavelengths) and recovering each field simultaneously. In this work, we make strides towards the practical implementation and usage of multi-wavelength diffraction imaging. In particular, we provide novel analysis of the noise characteristics of multi-wavelength diffraction imaging, and show that it is preferable to coherent diffraction imaging under high signal-independent noise. Additionally, we present a compact coded diffraction imaging system and corresponding phase retrieval algorithms to robustly and simultaneously recover complex fields representing multiple wavelengths. Using a novel mixed-norm color prior, our prototype system reconstructs a larger number of multi-wavelength fields from fewer measurements than existing methods, and supports applications such as micron-scale optical path difference measurement via synthetic wavelength holography.
Benjamin Attal, Matthew O'Toole
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Dual-Shutter Optical Vibration Sensing
abstract
Visual vibrometry is a highly useful tool for remote capture of audio, as well as the physical properties of materials, human heart rate, and more. While visually-observable vibrations can be captured directly with a high-speed camera, minute imperceptible object vibrations can be optically amplified by imaging the displacement of a speckle pattern created by shining a laser beam on the vibrating surface. In this paper, we propose a novel method for sensing vibrations at high speeds (up to 63 kHz), for multiple scene sources at once, using sensors rated for only 130 Hz operation. Our method relies on simultaneously capturing the scene with two cameras equipped with rolling and global shutter sensors, respectively. The rolling shutter camera captures distorted speckle images that encode the high-speed object vibrations. The global shutter camera captures undistorted reference images of the speckle pattern, helping to decode the source vibrations. We demonstrate our method by capturing vibration caused by audio sources (e.g., speakers, human voice, and musical instruments) and analyzing the vibration modes of a tuning fork.
Mark Sheinin, Dorian Chan, Matthew O'Toole, Srinivasa G. Narasimhan
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Coherence as Texture - Passive Textureless 3D Reconstruction by Self-Interference
abstract
Passive depth estimation based on stereo or defocus relies on the presence of the texture on an object to resolve its depth. Hence, recovering the depth of a textureless object-for example, a large white wall-is not just hard but perhaps even impossible. Or is it? We show that spatial coherence, a property of natural light sources, can be used to resolve the depth of a scene point even when it is textureless. Our approach relies on the idea that natural light scattered off a scene point is locally coherent with itself, while incoherent with the light scattered from other surface points; we use this insight to design an optical setup that uses self-interference as a texture feature for estimating depth. Our lab prototype is capable of resolving depths of textureless objects in sunlight as well as indoor lights.
Wei-Yu Chen, Aswin C. Sankaranarayanan, Anat Levin, Matthew O'Toole
CVPR4
2024 Flash Cache: Reducing Bias in Radiance Cache Based Inverse Rendering
Benjamin Attal, Dor Verbin, Ben Mildenhall, Peter Hedman, Jonathan T. Barron, Matthew O'Toole, Pratul P. Srinivasan
ECCV (28)6
2024 Holodepth: Programmable Depth-Varying Projection via Computer-Generated Holography
Dorian Chan, Matthew O'Toole, Sizhuo Ma, Jian Wang 0100
ECCV (61)2
2024 Flowed Time of Flight Radiance Fields
Mikhail Okunev, Marc Mapeke, Benjamin Attal, Christian Richardt, Matthew O'Toole, James Tompkin 0001
ECCV (62)5
2023 HyperReel: High-Fidelity 6-DoF Video with Ray-Conditioned Sampling
abstract
Volumetric scene representations enable photorealistic view synthesis for static scenes and form the basis of several existing 6-DoF video techniques. However, the volume rendering procedures that drive these representations necessitate careful trade-offs in terms of quality, rendering speed, and memory efficiency. In particular, existing methods fail to simultaneously achieve real-time performance, small memory footprint, and high-quality rendering for challenging real-world scenes. To address these issues, we present HyperReel―a novel 6-DoF video representation. The two core components of HyperReel are: (1) a ray-conditioned sample prediction network that enables high-fidelity, high frame rate rendering at high resolutions and (2) a compact and memory-efficient dynamic volume representation. Our 6-DoF video pipeline achieves the best performance compared to prior and contemporary approaches in terms of visual quality with small memory requirements, while also rendering at up to 18 frames-per-second at megapixel resolution without any custom CUDA code.
Benjamin Attal, Jia-Bin Huang 0001, Christian Richardt, Michael Zollhöfer, Johannes Kopf 0001, Matthew O'Toole, Changil Kim 0001
CVPR6
2023 Implicit Neural Head Synthesis via Controllable Local Deformation Fields
abstract
High-quality reconstruction of controllable 3D head avatars from 2D videos is highly desirable for virtual human applications in movies, games, and telepresence. Neural implicit fields provide a powerful representation to model 3D head avatars with personalized shape, expressions, and facial parts, e.g., hair and mouth interior, that go beyond the linear 3D morphable model (3DMM). However, existing methods do not model faces with fine-scale facial features, or local control of facial parts that extrapolate asymmetric expressions from monocular videos. Further, most condition only on 3DMM parameters with poor(er) locality, and resolve local features with a global neural field. We build on part-based implicit shape models that decompose a global deformation field into local ones. Our novel formulation models multiple implicit deformation fields with local semantic rig-like control via 3DMM-based parameters, and representative facial landmarks. Further, we propose a local controlloss and attention mask mechanism that promote sparsity of each learned deformation field. Our formulation renders sharper locally controllable nonlinear deformations than previous implicit monocular approaches, especially mouth interior, asymmetric expressions, and facial details. Project page: https://imaging.cs.cmu.edullocal_deformation_fieldsl
Chuhan Chen, Matthew O'Toole, Gaurav Bharaj, Pablo Garrido 0001
CVPR2
2023 Azimuth Super-Resolution for FMCW Radar in Autonomous Driving
abstract
We tackle the task of Azimuth (angular dimension) super-resolution for Frequency Modulated Continuous Wave (FMCW) multiple-input multiple-output (MIMO) radar. FMCW MIMO radar is widely used in autonomous driving alongside Lidar and RGB cameras. However, compared to Lidar, MIMO radar is usually of low resolution due to hardware size restrictions. For example, achieving 1° azimuth resolution requires at least 100 receivers, but a single MIMO device usually supports at most 12 receivers. Having limitations on the number of receivers is problematic since a high-resolution measurement of azimuth angle is essential for estimating the location and velocity of objects. To improve the azimuth resolution of MIMO radar, we propose a light, yet efficient, Analog-to-Digital super-resolution model (ADC-SR) that predicts or hallucinates additional radar signals using signals from only a few receivers. Compared with the baseline models that are applied to processed radar Range-Azimuth-Doppler (RAD) maps, we show that our ADC-SR method that processes raw ADC signals achieves comparable performance with 98% (50 times) fewer parameters. We also propose a hybrid super-resolution model (Hybrid-SR) combining our ADC-SR with a standard RAD super-resolution model, and show that performance can be improved by a large margin. Experiments on our Pitt-Radar dataset and the RADIal dataset validate the importance of leveraging raw radar ADC signals. To assess the value of our super-resolution model for autonomous driving, we also perform object detection on the results of our super-resolution model and find that our super-resolution model improves detection performance by around 4% in mAP. The Pitt-Radar and the code will be released at the link.
Yu-Jhe Li, Shawn Hunt, Jinhyung Park, Matthew O'Toole, Kris Makoto Kitani
CVPR4
2023 Analyzing Physical Impacts Using Transient Surface Wave Imaging
abstract
The subtle vibrations on an object's surface contain information about the object's physical properties and its interaction with the environment. Prior works imaged surface vibration to recover the object's material properties via modal analysis, which discards the transient vibrations propagating immediately after the object is disturbed. Conversely, prior works that captured transient vibrations focused on recovering localized signals (e.g., recording nearby sound sources), neglecting the spatiotemporal relationship between vibrations at different object points. In this paper, we extract information from the transient surface vibrations simultaneously measured at a sparse set of object points using the dual-shutter camera described by Sheinin et al. [37]. We model the geometry of an elastic wave generated at the moment an object's surface is disturbed (e.g., a knock or a footstep) and use the model to localize the disturbance source for various materials (e.g., wood, plastic, tile). We also show that transient object vibrations contain additional cues about the impact force and the impacting object's material properties. We demonstrate our approach in applications like localizing the strikes of a ping-pong ball on a table mid-play and recovering the footsteps' locations by imaging the floor vibrations they create.
Mark Sheinin, Dorian Chan, Mark Rau, Matthew O'Toole, Srinivasa G. Narasimhan
CVPR5
2023 ST-MVDNet++: Improve Vehicle Detection with Lidar-Radar Geometrical Augmentation via Self-Training
abstract
We aim to improve the performance of the vehicle detection model with Lidar-Radar fusion and data augmentation. The recent works for Lidar-Radar fusion such as MVDNet or ST-MVDNet, have been proposed to have effective performance in detecting vehicles, and address the issue regarding missing modality. However, there are few works applying some global data augmentations such as rotation, translation, and scaling which are common for Lidar-only model. In order to further improve the previous Lidar-Radar fusion model, we propose a model named ST-MVDNet++ by leveraging the self-training teacher-student framework with integrating more common data augmentations such as global rotation, translation, and scaling. To ensure the data augmentations are consistent and matched across Lidar and Radar, we apply the augmentations on bird-eye-view coordinates. We also introduce the student-only augmentation for robust training of the student model with the consistency loss from teacher model. We demonstrate that our leveraging of global consistent Lidar-Radar augmentation improve the previous works by 1 ∼ 2% in all of the experimental settings.
Yu-Jhe Li, Matthew O'Toole, Kris Makoto Kitani
ICASSP2
2023 SpinCam: High-Speed Imaging via a Rotating Point-Spread Function
abstract
High-speed cameras are an indispensable tool used for the slow-motion analysis of scenes. However, the fixed bandwidth of any imaging system quickly becomes a bottleneck, resulting in a fundamental trade-off between the camera’s spatial and temporal resolutions. In recent years, compressive high-speed imaging systems have been proposed to circumvent these issues by optically encoding the signal and using a reconstruction procedure to recover a video. Our work proposes a novel approach for compressive high-speed imaging based on temporally coding the camera’s point-spread function (PSF). By mechanically spinning a diffraction grating in front of a camera, the sensor integrates an image blurred by a PSF that continuously rotates over time. We also propose a deconvolution-based reconstruction algorithm to reconstruct videos from these measurements. Our method achieves superior light efficiency and handles a wider scene class than prior methods. Also, our mechanical design yields flexible temporal resolution that can be easily increased, potentially allowing capture at 192 kHz—far higher than prior works. We demonstrate a prototype for various applications, including motion capture and particle image velocimetry (PIV).
Dorian Chan, Mark Sheinin, Matthew O'Toole
ICCV3
2023 Neural Fields for Structured Lighting
abstract
We present an image formation model and optimization procedure that combines the advantages of neural radiance fields and structured light imaging. Existing depth-supervised neural models rely on depth sensors to accurately capture the scene’s geometry. However, the depth maps recovered by these sensors can be prone to error, or even fail outright. Instead of depending on the fidelity of processed depth maps from a structured light system, a more principled approach is to explicitly model the raw structured light images themselves. Our proposed approach enables the estimation of high-fidelity depth maps, including for objects with complex material properties (e.g., partially-transparent surfaces). Besides computing depth, the raw structured light images also confer other useful radiometric cues, which enable predicting surface normals and decomposing scene appearance in terms of a direct, indirect, and ambient component. We evaluate our framework quantitatively and qualitatively on a range of real and synthetic scenes, and decompose scenes into their constituent components for novel views.
Aarrushi Shandilya, Benjamin Attal, Christian Richardt, James Tompkin 0001, Matthew O'Toole
ICCV5
2023 Light-Efficient Holographic Illumination for Continuous-Wave Time-of-Flight Imaging
abstract
Time-of-flight (TOF) cameras have seen widespread adoption in recent years across the entire spectrum of commodity devices. However, these devices are fundamentally limited by their dynamic range, struggling with saturation from nearby, brighter objects and noisy depth from farther, darker objects. In this work, we explore overcoming these limitations in the context of continuous-wave time-of-flight (CWTOF) devices, by using a holographic light source capable of redistributing light according to arbitrary patterns. In particular, we propose using such a system to move light from overexposed to underexposed regions of a scene, such that the entire scene is well exposed. Such a methodology can be easily integrated with existing illumination schemes for TOF. Our proof-of-concept prototype is constructed from off-the-shelf optical components, and demonstrated on a number of lab scenes.
Dorian Chan, Matthew O'Toole
SIGGRAPH Asia2
2023 Split-Lohmann Multifocal Displays
abstract
This work provides the design of a multifocal display that can create a dense stack of focal planes in a single shot. We achieve this using a novel computational lens that provides spatial selectivity in its focal length, i.e, the lens appears to have different focal lengths across points on a display behind it. This enables a multifocal display via an appropriate selection of the spatially-varying focal length, thereby avoiding time multiplexing techniques that are associated with traditional focus tunable lenses. The idea central to this design is a modification of a Lohmann lens, a focus tunable lens created with two cubic phase plates that translate relative to each other. Using optical relays and a phase spatial light modulator, we replace the physical translation of the cubic plates with an optical one, while simultaneously allowing for different pixels on the display to undergo different amounts of translations and, consequently, different focal lengths. We refer to this design as a Split-Lohmann multifocal display. Split-Lohmann displays provide a large étendue as well as high spatial and depth resolutions; the absence of time multiplexing and the extremely light computational footprint for content processing makes it suitable for video and interactive experiences. Using a lab prototype, we show results over a wide range of static, dynamic, and interactive 3D scenes, showcasing high visual quality over a large working range.
Yingsi Qin, Wei-Yu Chen, Matthew O'Toole, Aswin C. Sankaranarayanan
ACM Trans. Graph.3
2022 Holocurtains: Programming Light Curtains via Binary Holography
abstract
Light curtain systems are designed for detecting the presence of objects within a user-defined 3D region of space, which has many applications across vision and robotics. However, the shape of light curtains have so far been limited to ruled surfaces, i.e., surfaces composed of straight lines. In this work, we propose Holocurtains: a light-efficient approach to producing light curtains of arbitrary shape. The key idea is to synchronize a rolling-shutter camera with a 2D holographic projector, which steers (rather than block) light to generate bright structured light patterns. Our prototype projector uses a binary digital micromirror device (DMD) to generate the holographic interference patterns at high speeds. Our system produces 3D light curtains that cannot be achieved with traditional light curtain setups and thus enables all-new applications, including the ability to simultaneously capture multiple light curtains in a single frame, detect subtle changes in scene geometry, and transform any 3D surface into an optical touch interface.
Dorian Chan, Srinivasa G. Narasimhan, Matthew O'Toole
CVPR3
2022 Modality-Agnostic Learning for Radar-Lidar Fusion in Vehicle Detection
abstract
Fusion of multiple sensor modalities such as camera, Lidar, and Radar, which are commonly found on autonomous vehicles, not only allows for accurate detection but also robustifies perception against adverse weather conditions and individual sensor failures. Due to inherent sensor characteristics, Radar performs well under extreme weather conditions (snow, rain, fog) that significantly degrade camera and Lidar. Recently, a few works have developed vehicle detection methods fusing Lidar and Radar signals, i.e., MVD-Net. However, these models are typically developed under the assumption that the models always have access to two error-free sensor streams. If one of the sensors is unavailable or missing, the model may fail catastrophically. To mitigate this problem, we propose the Self-Training Multimodal Vehicle Detection Network (ST-MVDNet) which leverages a Teacher-Student mutual learning framework and a simulated sensor noise model used in strong data augmentation for Lidar and Radar. We show that by (1) enforcing output consistency between a Teacher network and a Student network and by (2) introducing missing modalities (strong augmentations) during training, our learned model breaks away from the error-free sensor assumption. This consistency enforcement enables the Student model to handle missing data properly and improve the Teacher model by updating it with the Student model's exponential moving average. Our experiments demonstrate that our proposed learning framework for multi-modal detection is able to better handle missing sensor data during inference. Furthermore, our method achieves new state-of-the-art performance (5% gain) on the Oxford Radar Robotcar dataset under various evaluation settings.
Yu-Jhe Li, Jinhyung Park, Matthew O'Toole, Kris Makoto Kitani
CVPR3
2022 Dual-Shutter Optical Vibration Sensing
abstract
Visual vibrometry is a highly useful tool for remote capture of audio, as well as the physical properties of materials, human heart rate, and more. While visually-observable vibrations can be captured directly with a high-speed camera, minute imperceptible object vibrations can be optically amplified by imaging the displacement of a speckle pattern, created by shining a laser beam on the vibrating surface. In this paper, we propose a novel method for sensing vibrations at high speeds (up to 63kHz), for multiple scene sources at once, using sensors rated for only 130Hz operation. Our method relies on simultaneously capturing the scene with two cameras equipped with rolling and global shutter sensors, respectively. The rolling shutter camera captures distorted speckle images that encode the high-speed object vibrations. The global shutter camera captures undistorted reference images of the speckle pattern, helping to decode the source vibrations. We demonstrate our method by capturing vibration caused by audio sources (e.g. speakers, human voice, and musical instruments) and analyzing the vibration modes of a tuning fork.
Mark Sheinin, Dorian Chan, Matthew O'Toole, Srinivasa G. Narasimhan
CVPR3
2022 Adjoint nonlinear ray tracing
abstract
Reconstructing and designing media with continuously-varying refractive index fields remains a challenging problem in computer graphics. A core difficulty in trying to tackle this inverse problem is that light travels inside such media along curves, rather than straight lines. Existing techniques for this problem make strong assumptions on the shape of the ray inside the medium, and thus limit themselves to media where the ray deflection is relatively small. More recently, differentiable rendering techniques have relaxed this limitation, by making it possible to differentiably simulate curved light paths. However, the automatic differentiation algorithms underlying these techniques use large amounts of memory, restricting existing differentiable rendering techniques to relatively small media and low spatial resolutions. We present a method for optimizing refractive index fields that both accounts for curved light paths and has a small, constant memory footprint. We use the adjoint state method to derive a set of equations for computing derivatives with respect to the refractive index field of optimization objectives that are subject to nonlinear ray tracing constraints. We additionally introduce discretization schemes to numerically evaluate these equations, without the need to store nonlinear ray trajectories in memory, significantly reducing the memory requirements of our algorithm. We use our technique to optimize high-resolution refractive index fields for a variety of applications, including creating different types of displays (multiview, lightfield, caustic), designing gradient-index optics, and reconstructing gas flows.
Arjun Teh, Matthew O'Toole, Ioannis Gkioulekas
ACM Trans. Graph.2
2021 Reference Wave Design for Wavefront Sensing
abstract
One of the classical results in wavefront sensing is phase-shifting point diffraction interferometry (PS-PDI), where the phase of a wavefront is measured by interfering it with a planar reference created from the incident wave itself. The limiting drawback of this approach is that the planar reference, often created by passing light through a narrow pinhole, is dim and noise sensitive. We address this limitation with a novel approach called ReWave that uses a non-planar reference that is designed to be brighter. The reference wave is designed in a specific way that would still allow for analytic phase recovery, exploiting ideas of sparse phase retrieval algorithms. ReWave requires only four image intensity measurements and is significantly more robust to noise compared to PS-PDI. We validate the robustness and applicability of our approach using a suite of simulated and real results.
Wei-Yu Chen, Anat Levin, Matthew O'Toole, Aswin C. Sankaranarayanan
ICCP3
2021 Deconvolving Diffraction for Fast Imaging of Sparse Scenes
abstract
Most computer vision techniques rely on cameras which uniformly sample the 2D image plane. However, there exists a class of applications for which the standard uniform 2D sampling of the image plane is sub-optimal. This class consists of applications where the scene points of interest occupy the image plane sparsely (e.g., marker-based motion capture), and thus most pixels of the 2D camera sensor would be wasted. Recently, diffractive optics were used in conjunction with sparse (e.g., line) sensors to achieve high-speed capture of such sparse scenes. One such approach, called “Diffraction Line Imaging”, relies on the use of diffraction gratings to spread the point-spread-function (PSF) of scene points from a point to a color-coded shape (e.g., a horizontal line) whose intersection with a line sensor enables point positioning. In this paper, we extend this approach for arbitrary diffractive optical elements and arbitrary sampling of the sensor plane using a convolution-based image formation model. Sparse scenes are then recovered by formulating a convolutional coding inverse problem that can resolve mixtures of diffraction PSFs without the use of multiple sensors, extending the application of diffraction-based imaging to a new class of significantly denser scenes. For the case of a single-axis diffraction grating, we provide an approach to determine the minimal required sensor sub-sampling for accurate scene recovery. Compared to methods that use a speckle PSF from a narrow-band source or a diffuser-based PSF with a rolling shutter sensor, our approach uses spectrally-coded PSFs from broad-band sources and allows arbitrary sensor sampling, respectively. We demonstrate that the presented combination of the imaging approach and scene recovery method is well suited for high-speed marker based motion capture and particle image velocimetry (PIV) over long periods.
Mark Sheinin, Matthew O'Toole, Srinivasa G. Narasimhan
ICCP2
2021 Multi-Echo LiDAR for 3D Object Detection
abstract
LiDAR sensors can be used to obtain a wide range of measurement signals other than a simple 3D point cloud, and those signals can be leveraged to improve perception tasks like 3D object detection. A single laser pulse can be partially reflected by multiple objects along its path, resulting in multiple measurements called echoes. Multi-echo measurement can provide information about object contours and semi-transparent surfaces which can be used to better identify and locate objects. LiDAR can also measure surface reflectance (intensity of laser pulse return), as well as ambient light of the scene (sunlight reflected by objects). These signals are already available in commercial LiDAR devices but have not been used in most LiDAR-based detection models. We present a 3D object detection model which leverages the full spectrum of measurement signals provided by LiDAR. First, we propose a multi-signal fusion (MSF) module to combine (1) the reflectance and ambient features extracted with a 2D CNN, and (2) point cloud features extracted using a 3D graph neural network (GNN). Second, we propose a multi-echo aggregation (MEA) module to combine the information encoded in different sets of echo points. Compared with traditional single echo point cloud methods, our proposed Multi-Signal LiDAR Detector (MSLiD) extracts richer context information from a wider range of sensing measurements and achieves more accurate 3D object detection. Experiments show that by incorporating the multi-modality of LiDAR, our method outperforms the state-of-the-art by up to relatively 9.1%.
Yunze Man, Xinshuo Weng, Prasanna Kumar Sivakumar, Matthew O'Toole, Kris Makoto Kitani
ICCV4
2021 TöRF: Time-of-Flight Radiance Fields for Dynamic Scene View Synthesis
abstract
Neural networks can represent and accurately reconstruct radiance fields for static 3D scenes (e.g., NeRF). Several works extend these to dynamic scenes captured with monocular video, with promising performance. However, the monocular setting is known to be an under-constrained problem, and so methods rely on data-driven priors for reconstructing dynamic content. We replace these priors with measurements from a time-of-flight (ToF) camera, and introduce a neural representation based on an image formation model for continuous-wave ToF cameras. Instead of working with processed depth maps, we model the raw ToF sensor measurements to improve reconstruction quality and avoid issues with low reflectance regions, multi-path interference, and a sensor's limited unambiguous depth range. We show that this approach improves robustness of dynamic scene reconstruction to erroneous calibration and large motions, and discuss the benefits and limitations of integrating RGB+ToF sensors now available on modern smartphones.
Benjamin Attal, Eliot Laidlaw, Aaron Gokaslan, Changil Kim 0001, Christian Richardt, James Tompkin 0001, Matthew O'Toole
NeurIPS7
2020 Optical Non-Line-of-Sight Physics-Based 3D Human Pose Estimation
abstract
We describe a method for 3D human pose estimation from transient images (i.e., a 3D spatio-temporal histogram of photons) acquired by an optical non-line-of-sight (NLOS) imaging system. Our method can perceive 3D human pose by 'looking around corners' through the use of light indirectly reflected by the environment. We bring together a diverse set of technologies from NLOS imaging, human pose estimation and deep reinforcement learning to construct an end-to-end data processing pipeline that converts a raw stream of photon measurements into a full 3D human pose sequence estimate. Our contributions are the design of data representation process which includes (1) a learnable inverse point spread function (PSF) to convert raw transient images into a deep feature vector; (2) a neural humanoid control policy conditioned on the transient image feature and learned from interactions with a physics simulator; and (3) a data synthesis and augmentation strategy based on depth data that can be transferred to a real-world NLOS imaging system. Our preliminary experiments suggest that our method is able to generalize to real-world NLOS measurement to estimate physically-valid 3D human poses.
Mariko Isogawa, Ye Yuan 0007, Matthew O'Toole, Kris Makoto Kitani
CVPR3
2020 Efficient Non-Line-of-Sight Imaging from Transient Sinograms
Mariko Isogawa, Dorian Chan, Ye Yuan 0007, Kris Makoto Kitani, Matthew O'Toole
ECCV (7)5
2020 Diffraction Line Imaging
Mark Sheinin, N. Dinesh Reddy, Matthew O'Toole, Srinivasa G. Narasimhan
ECCV (2)3
2019 Non-line-of-sight Imaging with Partial Occluders and Surface Normals
abstract
Imaging objects obscured by occluders is a significant challenge for many applications. A camera that could “see around corners” could help improve navigation and mapping capabilities of autonomous vehicles or make search and rescue missions more effective. Time-resolved single-photon imaging systems have recently been demonstrated to record optical information of a scene that can lead to an estimation of the shape and reflectance of objects hidden from the line of sight of a camera. However, existing non-line-of-sight (NLOS) reconstruction algorithms have been constrained in the types of light transport effects they model for the hidden scene parts. We introduce a factored NLOS light transport representation that accounts for partial occlusions and surface normals. Based on this model, we develop a factorization approach for inverse time-resolved light transport and demonstrate high-fidelity NLOS reconstructions for challenging scenes both in simulation and with an experimental NLOS imaging system.
Felix Heide, Matthew O'Toole, Kai Zang, David B. Lindell, Steven Diamond, Gordon Wetzstein
ACM Trans. Graph.2
2019 Wave-based non-line-of-sight imaging using fast f-k migration
abstract
Imaging objects outside a camera's direct line of sight has important applications in robotic vision, remote sensing, and many other domains. Time-of-flight-based non-line-of-sight (NLOS) imaging systems have recently demonstrated impressive results, but several challenges remain. Image formation and inversion models have been slow or limited by the types of hidden surfaces that can be imaged. Moreover, non-planar sampling surfaces and non-confocal scanning methods have not been supported by efficient NLOS algorithms. With this work, we introduce a wave-based image formation model for the problem of NLOS imaging. Inspired by inverse methods used in seismology, we adapt a frequency-domain method, f-k migration, for solving the inverse NLOS problem. Unlike existing NLOS algorithms, f-k migration is both fast and memory efficient, it is robust to specular and other complex reflectance properties, and we show how it can be used with non-confocally scanned measurements as well as for non-planar sampling surfaces. f-k migration is more robust to measurement noise than alternative methods, generally produces better quality reconstructions, and is easy to implement. We experimentally validate our algorithms with a new NLOS imaging system that records room-sized scenes outdoors under indirect sunlight, and scans persons wearing retroreflective clothing at interactive rates.
David B. Lindell, Gordon Wetzstein, Matthew O'Toole
ACM Trans. Graph.3
2018 Tracking Multiple Objects Outside the Line of Sight Using Speckle Imaging
abstract
This paper presents techniques for tracking non-line-of-sight (NLOS) objects using speckle imaging. We develop a novel speckle formation and motion model where both the sensor and the source view objects only indirectly via a diffuse wall. We show that this NLOS imaging scenario is analogous to direct LOS imaging with the wall acting as a virtual, bare (lens-less) sensor. This enables tracking of a single, rigidly moving NLOS object using existing speckle-based motion estimation techniques. However, when imaging multiple NLOS objects, the speckle components due to different objects are superimposed on the virtual bare sensor image, and cannot be analyzed separately for recovering the motion of individual objects. We develop a novel clustering algorithm based on the statistical and geometrical properties of speckle images, which enables identifying the motion trajectories of multiple, independently moving NLOS objects. We demonstrate, for the first time, tracking individual trajectories of multiple objects around a corner with extreme precision (<; 10 microns) using only off-the-shelf imaging components.
Brandon M. Smith 0001, Matthew O'Toole, Mohit Gupta 0001
CVPR2
2018 Towards transient imaging at interactive rates with single-photon detectors
abstract
Active imaging at the picosecond timescale reveals transient light transport effects otherwise not accessible by computer vision and image processing algorithms. For example, analyzing the time of flight of short laser pulses emitted into a scene and scattered back to a detector allows for depth imaging, which is crucial for autonomous driving and many other applications. Moreover, analyzing or removing global light transport effects from photographs becomes feasible. While several transient imaging systems have recently been proposed using various imaging technologies, none is capable of acquiring transient images at interactive framerates. In this paper, we present an imaging system that records transient images at up to 25 Hz. We show several transient video clips recorded with this system and demonstrate transient imaging applications, including direct-global light transport separation and enhanced depth imaging.
David B. Lindell, Matthew O'Toole, Gordon Wetzstein
ICCP2
2018 Single-photon 3D imaging with deep sensor fusion
abstract
Sensors which capture 3D scene information provide useful data for tasks in vehicle navigation, gesture recognition, human pose estimation, and geometric reconstruction. Active illumination time-of-flight sensors in particular have become widely used to estimate a 3D representation of a scene. However, the maximum range, density of acquired spatial samples, and overall acquisition time of these sensors is fundamentally limited by the minimum signal required to estimate depth reliably. In this paper, we propose a data-driven method for photon-efficient 3D imaging which leverages sensor fusion and computational reconstruction to rapidly and robustly estimate a dense depth map from low photon counts. Our sensor fusion approach uses measurements of single photon arrival times from a low-resolution single-photon detector array and an intensity image from a conventional high-resolution camera. Using a multi-scale deep convolutional network, we jointly process the raw measurements from both sensors and output a high-resolution depth map. To demonstrate the efficacy of our approach, we implement a hardware prototype and show results using captured data. At low signal-to-background levels, our depth reconstruction algorithm with sensor fusion outperforms other methods for depth estimation from noisy measurements of photon arrival times.
David B. Lindell, Matthew O'Toole, Gordon Wetzstein
ACM Trans. Graph.2
2017 Reconstructing Transient Images from Single-Photon Sensors
abstract
Computer vision algorithms build on 2D images or 3D videos that capture dynamic events at the millisecond time scale. However, capturing and analyzing “transient images” at the picosecond scale-i.e., at one trillion frames per second-reveals unprecedented information about a scene and light transport within. This is not only crucial for time-of-flight range imaging, but it also helps further our understanding of light transport phenomena at a more fundamental level and potentially allows to revisit many assumptions made in different computer vision algorithms. In this work, we design and evaluate an imaging system that builds on single photon avalanche diode (SPAD) sensors to capture multi-path responses with picosecond-scale active illumination. We develop inverse methods that use modern approaches to deconvolve and denoise measurements in the presence of Poisson noise, and compute transient images at a higher quality than previously reported. The small form factor, fast acquisition rates, and relatively low cost of our system potentially makes transient imaging more practical for a range of applications.
Matthew O'Toole, Felix Heide, David B. Lindell, Kai Zang, Steven Diamond, Gordon Wetzstein
CVPR1
2016 3D Shape and Indirect Appearance by Structured Light Transport
abstract
We consider the problem of deliberately manipulating the direct and indirect light flowing through a time-varying, general scene in order to simplify its visual analysis. Our approach rests on a crucial link between stereo geometry and light transport: while direct light always obeys the epipolar geometry of a projector-camera pair, indirect light overwhelmingly does not. We show that it is possible to turn this observation into an imaging method that analyzes light transport in real time in the optical domain, prior to acquisition. This yields three key abilities that we demonstrate in an experimental camera prototype: (1) producing a live indirect-only video stream for any scene, regardless of geometric or photometric complexity; (2) capturing images that make existing structured-light shape recovery algorithms robust to indirect transport; and (3) turning them into one-shot methods for dynamic 3D shape capture.
Matthew O'Toole, John Mather, Kiriakos N. Kutulakos
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Defocus deblurring and superresolution for time-of-flight depth cameras
abstract
Continuous-wave time-of-flight (ToF) cameras show great promise as low-cost depth image sensors in mobile applications. However, they also suffer from several challenges, including limited illumination intensity, which mandates the use of large numerical aperture lenses, and thus results in a shallow depth of field, making it difficult to capture scenes with large variations in depth. Another shortcoming is the limited spatial resolution of currently available ToF sensors. In this paper we analyze the image formation model for blurred ToF images. By directly working with raw sensor measurements but regularizing the recovered depth and amplitude images, we are able to simultaneously deblur and super-resolve the output of ToF cameras. Our method outperforms existing methods on both synthetic and real datasets. In the future our algorithm should extend easily to cameras that do not follow the cosine model of continuous-wave sensors, as well as to multi-frequency or multi-phase imaging employed in more recent ToF cameras.
Lei Xiao 0014, Felix Heide, Matthew O'Toole, Andreas Kolb 0001, Matthias B. Hullin, Kiriakos N. Kutulakos, Wolfgang Heidrich
CVPR3
2015 Homogeneous codes for energy-efficient illumination and imaging
abstract
Programmable coding of light between a source and a sensor has led to several important results in computational illumination, imaging and display. Little is known, however, about how to utilize energy most effectively, especially for applications in live imaging. In this paper, we derive a novel framework to maximize energy efficiency by "homogeneous matrix factorization" that respects the physical constraints of many coding mechanisms (DMDs/LCDs, lasers, etc. ). We demonstrate energy-efficient imaging using two prototypes based on DMD and laser illumination. For our DMD-based prototype, we use fast local optimization to derive codes that yield brighter images with fewer artifacts in many transport probing tasks. Our second prototype uses a novel combination of a low-power laser projector and a rolling shutter camera. We use this prototype to demonstrate never-seen-before capabilities such as (1) capturing live structured-light video of very bright scenes---even a light bulb that has been turned on; (2) capturing epipolar-only and indirect-only live video with optimal energy efficiency; (3) using a low-power projector to reconstruct 3D objects in challenging conditions such as strong indirect light, strong ambient light, and smoke; and (4) recording live video from a projector's---rather than the camera's---point of view.
Matthew O'Toole, Supreeth Achar, Srinivasa G. Narasimhan, Kiriakos N. Kutulakos
ACM Trans. Graph.1
2014 3D Shape and Indirect Appearance by Structured Light Transport
abstract
We consider the problem of deliberately manipulating the direct and indirect light flowing through a time-varying, fully-general scene in order to simplify its visual analysis. Our approach rests on a crucial link between stereo geometry and light transport: while direct light always obeys the epipolar geometry of a projector-camera pair, indirect light overwhelmingly does not. We show that it is possible to turn this observation into an imaging method that analyzes light transport in real time in the optical domain, prior to acquisition. This yields three key abilities that we demonstrate in an experimental camera prototype: (1) producing a live indirect-only video stream for any scene, regardless of geometric or photometric complexity, (2) capturing images that make existing structured-light shape recovery algorithms robust to indirect transport, and (3) turning them into one-shot methods for dynamic 3D shape capture.
Matthew O'Toole, John Mather, Kiriakos N. Kutulakos
CVPR1
2014 Decomposing Global Light Transport Using Time of Flight Imaging
Di Wu 0006, Andreas Velten, Matthew O'Toole, Belén Masiá, Amit K. Agrawal, Qionghai Dai, Ramesh Raskar
Int. J. Comput. Vis.3
2014 Temporal frequency probing for 5D transient analysis of global light transport
abstract
We analyze light propagation in an unknown scene using projectors and cameras that operate at transient timescales. In this new photography regime, the projector emits a spatio-temporal 3D signal and the camera receives a transformed version of it, determined by the set of all light transport paths through the scene and the time delays they induce. The underlying 3D-to-3D transformation encodes scene geometry and global transport in great detail, but individual transport components ( e.g ., direct reflections, inter-reflections, caustics, etc .) are coupled nontrivially in both space and time. To overcome this complexity, we observe that transient light transport is always separable in the temporal frequency domain . This makes it possible to analyze transient transport one temporal frequency at a time by trivially adapting techniques from conventional projector-to-camera transport. We use this idea in a prototype that offers three never-seen-before abilities: (1) acquiring time-of-flight depth images that are robust to general indirect transport, such as interreflections and caustics; (2) distinguishing between direct views of objects and their mirror reflection; and (3) using a photonic mixer device to capture sharp, evolving wavefronts of "light-in-flight".
Matthew O'Toole, Felix Heide, Lei Xiao 0014, Matthias B. Hullin, Wolfgang Heidrich, Kiriakos N. Kutulakos
ACM Trans. Graph.1
2012 Decomposing global light transport using time of flight imaging
abstract
Global light transport is composed of direct and indirect components. In this paper, we take the first steps toward analyzing light transport using high temporal resolution information via time of flight (ToF) images. The time profile at each pixel encodes complex interactions between the incident light and the scene geometry with spatially-varying material properties. We exploit the time profile to decompose light transport into its constituent direct, subsurface scattering, and interreflection components. We show that the time profile is well modelled using a Gaussian function for the direct and interreflection components, and a decaying exponential function for the subsurface scattering component. We use our direct, subsurface scattering, and interreflection separation algorithm for four computer vision applications: recovering projective depth maps, identifying subsurface scattering objects, measuring parameters of analytical subsurface scattering models, and performing edge detection using ToF images.
Di Wu 0006, Matthew O'Toole, Andreas Velten, Amit K. Agrawal, Ramesh Raskar
CVPR2
2012 Frequency Analysis of Transient Light Transport with Applications in Bare Sensor Imaging
Di Wu 0006, Gordon Wetzstein, Christopher Barsi, Thomas Willwacher, Matthew O'Toole, Nikhil Naik 0003, Qionghai Dai, Kiriakos N. Kutulakos, Ramesh Raskar
ECCV (1)5
2012 Primal-dual coding to probe light transport
abstract
We present primal-dual coding , a photography technique that enables direct fine-grain control over which light paths contribute to a photo. We achieve this by projecting a sequence of patterns onto the scene while the sensor is exposed to light. At the same time, a second sequence of patterns, derived from the first and applied in lockstep, modulates the light received at individual sensor pixels. We show that photography in this regime is equivalent to a matrix probing operation in which the elements of the scene's transport matrix are individually re-scaled and then mapped to the photo. This makes it possible to directly acquire photos in which specific light transport paths have been blocked, attenuated or enhanced. We show captured photos for several scenes with challenging light transport effects, including specular inter-reflections, caustics, diffuse inter-reflections and volumetric scattering. A key feature of primal-dual coding is that it operates almost exclusively in the optical domain: our results consist of directly-acquired, unprocessed RAW photos or differences between them.
Matthew O'Toole, Ramesh Raskar, Kiriakos N. Kutulakos
ACM Trans. Graph.1
2010 A Basis Illumination Approach to BRDF Measurement
Abhijeet Ghosh, Wolfgang Heidrich, Shruthi Achutha, Matthew O'Toole
Int. J. Comput. Vis.4
2010 Optical computing for fast light transport analysis
abstract
We present a general framework for analyzing the transport matrix of a real-world scene at full resolution, without capturing many photos. The key idea is to use projectors and cameras to directly acquire eigenvectors and the Krylov subspace of the unknown transport matrix. To do this, we implement Krylov subspace methods partially in optics, by treating the scene as a "black box subroutine" that enables optical computation of arbitrary matrix-vector products. We describe two methods--- optical Arnoldi to acquire a low-rank approximation of the transport matrix for relighting; and optical GMRES to invert light transport. Our experiments suggest that good quality relighting and transport inversion are possible from a few dozen low-dynamic range photos, even for scenes with complex shadows, caustics, and other challenging lighting effects.
Matthew O'Toole, Kiriakos N. Kutulakos
ACM Trans. Graph.1
2007 BRDF Acquisition with Basis Illumination
abstract
Realistic descriptions of surface reflectance have long been a topic of interest in both computer vision and computer graphics research. In this paper, we describe a novel and fast approach for the acquisition of bidirectional reflectance distribution functions (BRDFs). We develop a novel theory for directly measuring BRDFs in a basis representation by projecting incident light as a sequence of basis functions from a spherical zone of directions. We derive an orthonormal basis over spherical zones that is ideally suited for this task. BRDF values outside the zonal directions are extrapolated by re-projecting the zonal measurements into a spherical harmonics basis, or by fitting analytical reflection models to the data. We verify this approach with a compact optical setup that requires no moving parts and only a small number of image measurements. Using this approach, a BRDF can be measured in just a few minutes.
Abhijeet Ghosh, Shruthi Achutha, Wolfgang Heidrich, Matthew O'Toole
ICCV4