VLDB 2026 Research / reviewers in the wild / expert
Mohit Gupta 0001
dblp:g/MohitGupta
· DBLP profile ↗
78ranked-venue papers
13as first author
33since 2021 · last 2026
0000-0002-2323-7700ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 69 · 11 first-author · 29 since 2021Artificial intelligence and machine learning · 55 · 10 first-author · 24 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Radiance Fields from PhotonsabstractNeural radiance fields, or NeRFs, have become the de facto approach for high-quality view synthesis from a collection of images captured from multiple viewpoints. However, many issues remain when capturing images in-the-wild under challenging conditions, such as in low light, high dynamic range, or with rapid motion, leading to smeared reconstructions with noticeable artifacts. In this work, we introduce quanta radiance fields , a novel class of neural radiance fields that are trained at the granularity of individual photons using single-photon cameras (SPCs). We develop theory and practical computational techniques for building radiance fields and estimating dense camera poses from unconventional, stochastic, and high-speed binary frame sequences captured by SPCs. We demonstrate, both via simulations and a SPC hardware prototype, high-fidelity reconstructions under high-speed motion, in low light, and for extreme dynamic range settings. Sacha Jungerman, Aryan Garg, Mohit Gupta 0001 |
ACM Trans. Graph. | 3 |
| 2025 | Robust 3D Object Detection Using Probabilistic Point Clouds From Single-Photon LidarsabstractLiDAR-based 3D sensors provide point clouds, a canonical 3D representation used in various scene understanding tasks. Modern LiDARs face key challenges in several real-world scenarios, such as long-distance or low-albedo objects, producing sparse or erroneous point clouds. These errors, which are rooted in the noisy raw LiDAR measurements, get propagated to downstream perception models, resulting in potentially severe loss of accuracy. This is because conventional 3D processing pipelines do not retain any uncertainty information from the raw measurements when constructing point clouds. We propose Probabilistic Point Clouds (PPC), a novel 3D scene representation where each point is augmented with a probability attribute that encapsulates the measurement uncertainty (or confidence) in the raw data. We further introduce inference approaches that leverage PPC for robust 3D object detection; these methods are versatile and can be used as computationally lightweight drop-in modules in 3D inference pipelines. We demonstrate, via both simulations and real captures, that PPC-based 3D inference methods outperform several baselines using LiDAR as well as camera-LiDAR fusion models, across challenging indoor and outdoor scenarios involving small, distant, and low-albedo objects, as well as strong ambient light. Our project webpage is at https://bhavyagoyal.github.io/ppc . Bhavya Goyal, Felipe Gutierrez-Barragan, Andreas Velten, Yin Li 0003, Mohit Gupta 0001 |
ICCV | 6 |
| 2025 | Recovering Parametric Scenes from Very Few Time-of-Flight PixelsabstractWe aim to recover the geometry of 3D parametric scenes using very few depth measurements from low-cost, commercially available time-of-flight sensors. These sensors offer very low spatial resolution (i.e., a single pixel), but image a wide field-of-view per pixel and capture detailed time-of-flight data in the form of time-resolved photon counts. This time-of-flight data encodes rich scene information and thus enables recovery of simple scenes from sparse measurements. We investigate the feasibility of using a distributed set of few measurements (e.g., as few as 15 pixels) to recover the geometry of simple parametric scenes with a strong prior, such as estimating the 6D pose of a known object. To achieve this, we design a method that utilizes both feed-forward prediction to infer scene parameters, and differentiable rendering within an analysis-by-synthesis framework to refine the scene parameter estimate. We develop hardware prototypes and demonstrate that our method effectively recovers object pose given an untextured 3D model in both simulations and controlled real-world captures, and show promising initial results for other parametric scenes. We additionally conduct experiments to explore the limits and capabilities of our imaging solution. Carter Sifferman, Yiquan Li, Fangzhou Mu, Michael Gleicher, Mohit Gupta 0001, Yin Li 0003 |
ICCV | 6 |
| 2025 | Quanta Neural Networks: From Photons to Perception
Varun Sundar, Sacha Jungerman, Mohit Gupta 0001 |
ICCV | 4 |
| 2025 | Instant Video Models: Universal Adapters for Stabilizing Image-Based NetworksabstractWhen applied sequentially to video, frame-based networks often exhibit temporal inconsistency—for example, outputs that flicker between frames. This problem is amplified when the network inputs contain time-varying corruptions. In this work, we introduce a general approach for adapting frame-based models for stable and robust inference on video. We describe a class of stability adapters that can be inserted into virtually any architecture and a resource-efficient training process that can be performed with a frozen base network. We introduce a unified conceptual framework for describing temporal stability and corruption robustness, centered on a proposed accuracy-stability-robustness loss. By analyzing the theoretical properties of this loss, we identify the conditions where it produces well-behaved stabilizer training. Our experiments validate our approach on several vision tasks including denoising (NAFNet), image enhancement (HDRNet), monocular depth (Depth Anything v2), and semantic segmentation (DeepLabv3+). Our method improves temporal stability and robustness against a range of image corruptions (including compression artifacts, noise, and adverse weather), while preserving or improving the quality of predictions. Matthew Dutson, Nathan Labiosa, Yin Li 0003, Mohit Gupta 0001 |
NeurIPS | 4 |
| 2025 | Privacy-Enabled Parallax DisplayabstractPrivacy filters for displays are designed to obfuscate or hide visual content from unintended observers, while making the displayed information visible only to selected viewers. Existing privacy filters suffer from either wide viewing field (low selectivity) or limited user positioning. To solve this dilemma, we propose a display technology that allows a narrow but adaptive viewing field that can be directed to arbitrary user location. While conventional parallax barriers provide such capability of modulating the light field according to the user location, it suffers from repeated views. Our key observation is that this view repetition originates from the periodicity of barrier patterns, and we propose a privacy-enabled parallax display based on randomized barrier design. In addition to randomizing the locations of 1D slits, we also propose breaking down the slits into pinholes and randomizing their 2D locations, which results in privacy-preservation along the vertical direction as well. We build a hardware prototype using two off-the-shelf liquid-crystal displays. Experiments show that the proposed randomized parallax barrier can direct to the user a narrow viewing field of about ±6°, providing a significantly improved privacy protection as compared to traditional privacy screens. Sizhuo Ma, Karl Bayer, Gurunandan Krishnan, Mohit Gupta 0001, Shree K. Nayar |
VR | 4 |
| 2025 | Streaming Quanta Sensors for Online, High-Performance Imaging and VisionabstractRecently quanta image sensors (QIS) - ultra-fast, zero-read-noise binary image sensors- have demonstrated remarkable imaging capabilities in many challenging scenarios. Despite their potential, the adoption of these sensors is severely hampered by (a) high data rates and (b) the need for new computational pipelines to handle the unconventional raw data. We introduce a simple, low-bandwidth computational pipeline to address these challenges. Our approach is based on a novel streaming representation with a small memory footprint, efficiently capturing intensity information at multiple temporal scales. Updating the representation requires only 16 floating-point operations/pixel, which can be efficiently computed online at the native frame rate of the binary frames. We use a neural network operating on this representation to reconstruct videos in real-time (10-30 fps). We illustrate why such representation is well-suited for these emerging sensors, and how it offers low latency and high frame rate while retaining flexibility for downstream computer vision. Our approach results in significant data bandwidth reductions () and real-time image reconstruction and computer vision -)reduction in computation than existing state-of-the-art approach[1], while maintaining comparable quality. To the best of our knowledge, our approach is the first to achieve online, real-time image reconstruction on QIS. Matthew Dutson, Vivek Boominathan, Mohit Gupta 0001, Ashok Veeraraghavan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Towards 3D Vision with Low-Cost Single-Photon CamerasabstractWe present a method for reconstructing 3D shape of arbitrary Lambertian objects based on measurements by miniature, energy-efficient, low-cost single-photon cameras. These cameras, operating as time resolved image sensors, illuminate the scene with a very fast pulse of diffuse light and record the shape of that pulse as it returns back from the scene at a high temporal resolution. We propose to model this image formation process, account for its non-idealities, and adapt neural rendering to reconstruct 3D geometry from a set of spatially distributed sensors with known poses. We show that our approach can successfully recover complex 3D shapes from simulated data. We further demonstrate 3D object reconstruction from real-world captures, utilizing measurements from a commodity proximity sensor. Our work draws a connection between image-based modeling and active range scanning, and offers a step towards 3D vision with single-photon cameras. Our project webpage is at https://cpsiff.github.io/towards_3d_vision/. Fangzhou Mu, Carter Sifferman, Sacha Jungerman, Yiquan Li, Mark Han, Michael Gleicher, Mohit Gupta 0001, Yin Li 0003 |
CVPR | 7 |
| 2024 | Generalized Event CamerasabstractEvent cameras capture the world at high time resolution and with minimal bandwidth requirements. However, event streams, which only encode changes in brightness, do not contain sufficient scene information to support a wide variety of downstream tasks. In this work, we design generalized event cameras that inherently preserve scene intensity in a bandwidth-efficient manner. We generalize event cameras in terms of when an event is generated and what information is transmitted. To implement our designs, we turn to single-photon sensors that provide digital access to individual photon detections; this modality gives us the flexibility to realize a rich space of generalized event cameras. Our single-photon event cameras are capable of high-speed, high-fidelity imaging at low readout rates. Consequently, these event cameras can support plug-and-play downstream inference, without capturing new event datasets or designing specialized event-vision models. As a practical implication, our designs, which involve lightweight and near-sensor-compatible computations, provide a way to use single-photon sensors without exorbitant bandwidth costs. Varun Sundar, Matthew Dutson, Andrei Ardelean, Claudio Bruschini, Edoardo Charbon, Mohit Gupta 0001 |
CVPR | 6 |
| 2024 | Photon Inhibition for Energy-Efficient Single-Photon Imaging
Lucas J. Koerner, Atul Ingle, Mohit Gupta 0001 |
ECCV (76) | 4 |
| 2024 | Light-in-Flight for a World-in-Motion
Jongho Lee 0004, Ryan J. Suess, Mohit Gupta 0001 |
ECCV (34) | 3 |
| 2024 | Light Codes for Fast Two-Way Human-Centric Visual CommunicationabstractVisual codes, such as QR codes, are widely used in several applications for conveying information to users. However, user interactions based on spatial codes (e.g., displaying codes on phone screens for exchanging contact information) are often tedious, time consuming, and prone to errors due to image corruptions such as noise, blur, saturation, and perspective distortions. We propose Light Codes (LICO), a novel method for fast and fluid exchange of information among users. Light codes are based on transmitting and receiving temporal codes (instead of spatial) using compact and low-cost transceiver devices. The resulting approach enables seamless and near instantaneous exchange of short messages among users with minimal physical and cognitive effort. We design novel coding techniques, hardware prototypes, and applications that are optimized for human-centric communication, and facilitate fast and fluid user-to-user interactions in various challenging conditions, including a range of distances, motion, and ambient illumination. We evaluate the performance of the proposed methods both via quantitative analysis and user study based comparisons with several existing approaches including display-camera links, Bluetooth, and near-field communication, which show strong preference toward Light Codes in various real-world application scenarios. Mohit Gupta 0001, Jian Wang 0100, Karl Bayer, Shree K. Nayar |
ACM Trans. Graph. | 1 |
| 2023 | Eventful Transformers: Leveraging Temporal Redundancy in Vision TransformersabstractVision Transformers achieve impressive accuracy across a range of visual recognition tasks. Unfortunately, their accuracy frequently comes with high computational costs. This is a particular issue in video recognition, where models are often applied repeatedly across frames or temporal chunks. In this work, we exploit temporal redundancy between subsequent inputs to reduce the cost of Transformers for video processing. We describe a method for identifying and re-processing only those tokens that have changed significantly over time. Our proposed family of models, Eventful Transformers, can be converted from existing Transformers (often without any re-training) and give adaptive control over the compute cost at runtime. We evaluate our method on large-scale datasets for video object detection (ImageNet VID) and action recognition (EPIC-Kitchens 100). Our approach leads to significant computational savings (on the order of 2-4x) with only minor reductions in accuracy. Matthew Dutson, Yin Li 0003, Mohit Gupta 0001 |
ICCV | 3 |
| 2023 | Eulerian Single-Photon VisionabstractSingle-photon sensors measure light signals at the finest possible resolution — individual photons. These sensors introduce two major challenges in the form of strong Poisson noise and extremely large data acquisition rates, which are also inherited by downstream computer vision tasks. Previous work has largely focused on solving the image reconstruction problem first and then using off-the-shelf methods for downstream tasks, but the most general solutions that account for motion are costly and not scalable to large data volumes produced by single-photon sensors.This work forgoes the image reconstruction problem. Instead, we demonstrate computationally light-weight phase-based algorithms for the tasks of edge detection and motion estimation. These methods directly process the raw single-photon data as a 3D volume with a bank of velocity-tuned filters, achieving speed-ups of more than two orders of magnitude compared to explicit reconstruction-based methods.Project webpage: https://wisionlab.com/project/eulerian-single-photon-vision/ Mohit Gupta 0001 |
ICCV | 2 |
| 2023 | Learned Compressive Representations for Single-Photon 3D ImagingabstractSingle-photon 3D cameras can record the time-of-arrival of billions of photons per second with picosecond accuracy. One common approach to summarize the photon data stream is to build a per-pixel timestamp histogram, resulting in a 3D histogram tensor that encodes distances along the time axis. As the spatio-temporal resolution of the histogram tensor increases, the in-pixel memory requirements and output data rates can quickly become impractical. To overcome this limitation, we propose a family of linear compressive representations of histogram tensors that can be computed efficiently, in an online fashion, as a matrix operation. We design practical lightweight compressive representations that are amenable to an in-pixel implementation and consider the spatio-temporal information of each timestamp. Furthermore, we implement our proposed framework as the first layer of a neural network, which enables the joint end-to-end optimization of the compressive representations and a downstream SPAD data processing model. We find that a well-designed compressive representation can reduce in-sensor memory and data rates up to 2 orders of magnitude without significantly reducing 3D imaging quality. Finally, we analyze the power consumption implications through an on-chip implementation. Felipe Gutierrez-Barragan, Fangzhou Mu, Andrei Ardelean, Atul Ingle, Claudio Bruschini, Edoardo Charbon, Yin Li 0003, Mohit Gupta 0001, Andreas Velten |
ICCV | 8 |
| 2023 | Panoramas from PhotonsabstractScene reconstruction in the presence of high-speed motion and low illumination is important in many applications such as augmented and virtual reality, drone navigation, and autonomous robotics. Traditional motion estimation techniques fail in such conditions, suffering from too much blur in the presence of high-speed motion and strong noise in low-light conditions. Single-photon cameras have recently emerged as a promising technology capable of capturing hundreds of thousands of photon frames per second thanks to their high speed and extreme sensitivity. Unfortunately, traditional computer vision techniques are not well suited for dealing with the binary-valued photon data captured by these cameras because these are corrupted by extreme Poisson noise. Here we present a method capable of estimating extreme scene motion under challenging conditions, such as low light or high dynamic range, from a sequence of high-speed image frames such as those captured by a single-photon camera. Our method relies on iteratively improving a motion estimate by grouping and aggregating frames after-the-fact, in a stratified manner. We demonstrate the creation of high-quality panoramas under fast motion and extremely low light, and super-resolution results using a custom single-photon camera prototype. For code and supplemental material see our project webpage. Sacha Jungerman, Atul Ingle, Mohit Gupta 0001 |
ICCV | 3 |
| 2023 | Computational 3D Imaging with Position SensorsabstractUnderlying many structured light systems, especially those based on laser scanning, is a simple vision task: tracking a light spot. To accomplish this, scanners use conventional CMOS sensors to capture, transmit, and process millions of pixel measurements. This approach, while capable of achieving high-fidelity 3D scans, is wasteful in terms of (often scarce) sensing and computational resources. We present a structured light system based on position sensing diodes (PSDs), an unconventional sensing modality that directly measures the centroid of the spatial distribution of incident light, thus enabling high-resolution 3D laser scanning with a minimal amount of sensor data. We develop theory and computational algorithms for PSD-based structured light under a variety of light transport effects. We demonstrate the benefits of the proposed techniques using a hardware prototype on several real-world scenes, including optically-challenging objects with long-range interreflections and scattering. Jeremy Klotz, Mohit Gupta 0001, Aswin C. Sankaranarayanan |
ICCV | 2 |
| 2023 | SoDaCam: Software-defined Cameras via Single-Photon ImagingabstractReinterpretable cameras are defined by their post-processing capabilities that exceed traditional imaging. We present "SoDaCam" that provides reinterpretable cameras at the granularity of photons, from photon-cubes acquired by single-photon devices. Photon-cubes represent the spatio-temporal detections of photons as a sequence of binary frames, at frame-rates as high as 100 kHz. We show that simple transformations of the photon-cube, or photon-cube projections, provide the functionality of numerous imaging systems including: exposure bracketing, flutter shutter cameras, video compressive systems, event cameras, and even cameras that move during exposure. Our photon-cube projections offer the flexibility of being software-defined constructs that are only limited by what is computable, and shot-noise. We exploit this flexibility to provide new capabilities for the emulated cameras. As an added benefit, our projections provide camera-dependent compression of photon-cubes, which we demonstrate using an implementation of our projections on a novel compute architecture that is designed for single-photon imaging. Varun Sundar, Andrei Ardelean, Tristan Swedish, Claudio Bruschini, Edoardo Charbon, Mohit Gupta 0001 |
ICCV | 6 |
| 2023 | QfaR: Location-Guided Scanning of Visual Codes from Long DistancesabstractVisual codes such as QR codes provide a low-cost and convenient communication channel between physical objects and mobile devices, but typically operate when the code and the device are in close physical proximity. We propose a system, called QfaR, which enables mobile devices to scan visual codes across long distances even where the image resolution of the visual codes is extremely low. QfaR is based on location-guided code scanning, where we utilize a crowd-sourced database of physical locations of codes. Our key observation is that if the approximate location of the codes and the user is known, the space of possible codes can be dramatically pruned down. Then, even if every "single bit" from the low-resolution code cannot be recovered, QfaR can still identify the visual code from the pruned list with high probability. By applying computer vision techniques, QfaR is also robust against challenging imaging conditions, such as tilt, motion blur, etc. Experimental results with common iOS and Android devices show that QfaR can significantly enhance distances at which codes can be scanned, e.g., 3.6cm-sized codes can be scanned at a distance of 7.5 meters, and 0.5m-sized codes at about 100 meters. QfaR has many potential applications, and beyond our diverse experiments, we also conduct a simple case study on its use for efficiently scanning QR code-based badges to estimate event attendance. Sizhuo Ma, Jian Wang 0100, Wenzheng Chen, Suman Banerjee 0001, Mohit Gupta 0001, Shree K. Nayar |
MobiCom | 5 |
| 2023 | Spike-Based Anytime PerceptionabstractIn many emerging computer vision applications, it is critical to adhere to stringent latency and power constraints. The current neural network paradigm of frame-based, floating-point inference is often ill-suited to these resource-constrained applications. Spike-based perception – enabled by spiking neural networks (SNNs) – is one promising alternative. Unlike conventional neural networks (ANNs), spiking networks exhibit smooth tradeoffs between latency, power, and accuracy. SNNs are the archetype of an "anytime algorithm" whose accuracy improves smoothly over time. This property allows SNNs to adapt their computational investment in response to changing resource constraints. Unfortunately, mainstream algorithms for training SNNs (i.e., those based on ANN-to-SNN conversion) tend to produce models that are inefficient in practice. To mitigate this problem, we propose a set of principled optimizations that reduce latency and power consumption by 1–2 orders of magnitude in converted SNNs. These optimizations leverage a set of novel efficiency metrics designed for anytime algorithms. We also develop a state-of-the-art simulator, SaRNN, which can simulate SNNs using commodity GPU hardware and neuromorphic platforms. We hope that the proposed optimizations, metrics, and tools will facilitate the future development of spike-based vision systems. Matthew Dutson, Yin Li 0003, Mohit Gupta 0001 |
WACV | 3 |
| 2023 | Burst Vision Using Single-Photon CamerasabstractSingle-photon avalanche diodes (SPADs) are novel image sensors that record the arrival of individual photons at extremely high temporal resolution. In the past, they were only available as single pixels or small-format arrays, for various active imaging applications such as LiDAR and microscopy. Recently, high-resolution SPAD arrays up to 3.2 megapixel have been realized, which for the first time may be able to capture sufficient spatial details for general computer vision tasks, purely as a passive sensor. However, existing vision algorithms are not directly applicable on the binary data captured by SPADs. In this paper, we propose developing quanta vision algorithms based on burst processing for extracting scene information from SPAD photon streams. With extensive real-world data, we demonstrate that current SPAD arrays, along with burst processing as an example plug-and-play algorithm, are capable of a wide range of downstream vision tasks in extremely challenging imaging conditions including fast motion, low light (< 5 lux) and high dynamic range. To our knowledge, this is the first attempt to demonstrate the capabilities of SPAD sensors for a wide gamut of real-world computer vision tasks including object detection, pose estimation, SLAM, and text recognition. We hope this work will inspire future research into developing computer vision algorithms in extreme scenarios using single-photon cameras. Sizhuo Ma, Paul Mos, Edoardo Charbon, Mohit Gupta 0001 |
WACV | 4 |
| 2023 | Mitigating AC and DC Interference in Multi-ToF-Camera EnvironmentsabstractMulti-camera interference (MCI) is an important challenge faced by continuous-wave time-of-flight (C-ToF) cameras. In the presence of other cameras, a C-ToF camera may receive light from other cameras' sources, resulting in potentially large depth errors. We propose stochastic exposure coding (SEC), a novel approach to mitigate MCI. In SEC, the camera integration time is divided into multiple time slots. Each camera is turned on during a slot with an optimal probability to avoid interference while maintaining high signal-to-noise ratio (SNR). The proposed approach has the following benefits. First, SEC can filter out both the AC and DC components of interfering signals effectively, which simultaneously achieves high SNR and mitigates depth errors. Second, time-slotting in SEC enables 3D imaging without saturation in the high photon flux regime. Third, the energy savings due to camera turning on during only a fraction of integration time can be utilized to amplify the source peak power, which increases the robustness of SEC to ambient light. Lastly, SEC can be implemented without modifying the C-ToF camera's coding functions, and thus, can be used with a wide range of cameras with minimal changes. We demonstrate the performance benefits of SEC with thorough theoretical analysis, simulations and real experiments, across a wide range of imaging scenarios. Jongho Lee 0004, Mohit Gupta 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Seeing Photons in ColorabstractMegapixel single-photon avalanche diode (SPAD) arrays have been developed recently, opening up the possibility of deploying SPADs as generalpurpose passive cameras for photography and computer vision. However, most previous work on SPADs has been limited to monochrome imaging. We propose a computational photography technique that reconstructs high-quality color images from mosaicked binary frames captured by a SPAD array, even for high-dyanamic-range (HDR) scenes with complex and rapid motion. Inspired by conventional burst photography approaches, we design algorithms that jointly denoise and demosaick single-photon image sequences. Based on the observation that motion effectively increases the color sample rate, we design a blue-noise pseudorandom RGBW color filter array for SPADs, which is tailored for imaging dark, dynamic scenes. Results on simulated data, as well as real data captured with a fabricated color SPAD hardware prototype shows that the proposed method can reconstruct high-quality images with minimal color artifacts even for challenging low-light, HDR and fast-moving scenes. We hope that this paper, by adding color to computational single-photon imaging, spurs rapid adoption of SPADs for real-world passive imaging applications. Sizhuo Ma, Varun Sundar, Paul Mos, Claudio Bruschini, Edoardo Charbon, Mohit Gupta 0001 |
ACM Trans. Graph. | 6 |
| 2022 | Compressive Single-Photon 3D CamerasabstractSingle-photon avalanche diodes (SPADs) are an emerging pixel technology for time-of-flight (ToF) 3D cameras that can capture the time-of-arrival of individual photons at picosecond resolution. To estimate depths, current SPAD-based 3D cameras measure the round-trip time of a laser pulse by building a per-pixel histogram of photon times-tamps. As the spatial and timestamp resolution of SPAD-based cameras increase, their output data rates far exceed the capacity of existing data transfer technologies. One major reason for SPAD's bandwidth-intensive operation is the tight coupling that exists between depth resolution and histogram resolution. To weaken this coupling, we propose compressive single-photon histograms (CSPH). CSPHs are a per-pixel compressive representation of the high-resolution histogram, that is built on-the-fly, as each photon is detected. They are based on a family of linear coding schemes that can be expressed as a simple matrix operation. We design different CSPH coding schemes for 3D imaging and evaluate them under different signal and background levels, laser waveforms, and illumination setups. Our results show that a well-designed CSPH can consistently reduce data rates by 1–2 orders of magnitude without compromising depth precision. Felipe Gutierrez-Barragan, Atul Ingle, Trevor Seets, Mohit Gupta 0001, Andreas Velten |
CVPR | 4 |
| 2022 | Single-Photon Structured LightabstractWe present a novel structured light technique that uses Single Photon Avalanche Diode (SPAD) arrays to enable 3D scanning at high-frame rates and low-light levels. This technique, called “Single-Photon Structured Light”, works by sensing binary images that indicates the presence or absence of photon arrivals during each exposure; the SPAD array is used in conjunction with a high-speed binary projector, with both devices operated at speeds as high as 20 kHz. The binary images that we acquire are heavily influenced by photon noise and are easily corrupted by ambient sources of light. To address this, we develop novel temporal sequences using error correction codes that are designed to be robust to short-range effects like projector and camera defocus as well as resolution mismatch between the two devices. Our lab prototype is capable of 3D imaging in challenging scenarios involving objects with extremely low albedo or undergoing fast motion, as well as scenes under strong ambient illumination. Varun Sundar, Sizhuo Ma, Aswin C. Sankaranarayanan, Mohit Gupta 0001 |
CVPR | 4 |
| 2022 | Event Neural Networks
Matthew Dutson, Yin Li 0003, Mohit Gupta 0001 |
ECCV (11) | 3 |
| 2022 | 3D Scene Inference from Transient Histograms
Sacha Jungerman, Atul Ingle, Yin Li 0003, Mohit Gupta 0001 |
ECCV (7) | 4 |
| 2022 | Robust Scene Inference under Noise-Blur Dual CorruptionsabstractScene inference under low-light is a challenging problem due to severe noise in the captured images. One way to reduce noise is to use longer exposure during the capture. However, in the presence of motion (scene or camera motion), longer exposures lead to motion blur, resulting in loss of image information. This creates a trade-off between these two kinds of image degradations: motion blur (due to long exposure) vs. noise (due to short exposure), also referred as a dual image corruption pair in this paper. With the rise of cameras capable of capturing multiple exposures of the same scene simultaneously, it is possible to overcome this trade-off. Our key observation is that although the amount and nature of degradation varies for these different image captures, the semantic content remains the same across all images. To this end, we propose a method to leverage these multi exposure captures for robust inference under low-light and motion. Our method builds on a feature consistency loss to encourage similar results from these individual captures, and uses the ensemble of their final predictions for robust visual recognition. We demonstrate the effectiveness of our approach on simulated images as well as real captures with multiple exposures, and across the tasks of object detection and image classification. Project: https://wisionlab.com/project/noiseblurdual Bhavya Goyal, Jean-François Lalonde, Yin Li 0003, Mohit Gupta 0001 |
ICCP | 4 |
| 2022 | Single-Photon Camera Guided Extreme Dynamic Range ImagingabstractReconstruction of high-resolution extreme dynamic range images from a small number of low dynamic range (LDR) images is crucial for many computer vision applications. Current high dynamic range (HDR) cameras based on CMOS image sensor technology rely on multi-exposure bracketing which suffers from motion artifacts and signal-to-noise (SNR) dip artifacts in extreme dynamic range scenes. Recently, single-photon cameras (SPCs) have been shown to achieve orders of magnitude higher dynamic range for passive imaging than conventional CMOS sensors. SPCs are becoming increasingly available commercially, even in some consumer devices. Unfortunately, current SPCs suffer from low spatial resolution. To overcome the limitations of CMOS and SPC sensors, we propose a learning-based CMOS-SPC fusion method to recover high-resolution extreme dynamic range images. We compare the performance of our method against various traditional and state-of-the-art baselines using both synthetic and experimental data. Our method outperforms these baselines, both in terms of visual quality and quantitative metrics. Yuhao Liu 0012, Felipe Gutierrez-Barragan, Atul Ingle, Mohit Gupta 0001, Andreas Velten |
WACV | 4 |
| 2021 | Passive Inter-Photon ImagingabstractDigital camera pixels measure image intensities by converting incident light energy into an analog electrical current, and then digitizing it into a fixed-width binary representation. This direct measurement method, while conceptually simple, suffers from limited dynamic range and poor performance under extreme illumination — electronic noise dominates under low illumination, and pixel full-well capacity results in saturation under bright illumination. We propose a novel intensity cue based on measuring inter-photon timing, defined as the time delay between detection of successive photons. Based on the statistics of inter-photon times measured by a time-resolved single-photon sensor, we develop theory and algorithms for a scene brightness estimator which works over extreme dynamic range; we experimentally demonstrate imaging scenes with a dynamic range of over ten million to one. The proposed techniques, aided by the emergence of single-photon sensors such as single-photon avalanche diodes (SPADs) with picosecond timing resolution, will have implications for a wide range of imaging applications: robotics, consumer photography, astronomy, microscopy and biomedical imaging. Atul Ingle, Trevor Seets, Mauro Buttafava, Alberto Tosi, Mohit Gupta 0001, Andreas Velten |
CVPR | 6 |
| 2021 | Blocks-World CamerasabstractFor several vision and robotics applications, 3D geometry of man-made environments such as indoor scenes can be represented with a small number of dominant planes. However, conventional 3D vision techniques typically first acquire dense 3D point clouds before estimating the compact piece-wise planar representations (e.g., by plane-fitting). This approach is costly, both in terms of acquisition and computational requirements, and potentially unreliable due to noisy point clouds. We propose Blocks-World Cameras, a class of imaging systems which directly recover dominant planes of piece-wise planar scenes (Blocks-World), without requiring point clouds. The Blocks-World Cameras are based on a structured-light system projecting a single pattern with a sparse set of cross-shaped features. We develop a novel geometric algorithm for recovering scene planes without explicit correspondence matching, thereby avoiding computationally intensive search or optimization routines. The proposed approach has low device and computational complexity, and requires capturing only one or two images. We demonstrate highly efficient and precise planar-scene sensing with simulations and real experiments, across various imaging conditions, including defocus blur, large lighting variations, ambient illumination, and scene clutter. Jongho Lee 0004, Mohit Gupta 0001 |
CVPR | 2 |
| 2021 | Invisible Perturbations: Physical Adversarial Examples Exploiting the Rolling Shutter EffectabstractPhysical adversarial examples for camera-based computer vision have so far been achieved through visible artifacts — a sticker on a Stop sign, colorful borders around eyeglasses or a 3D printed object with a colorful texture. An implicit assumption here is that the perturbations must be visible so that a camera can sense them. By contrast, we contribute a procedure to generate, for the first time, physical adversarial examples that are invisible to human eyes. Rather than modifying the victim object with visible artifacts, we modify light that illuminates the object. We demonstrate how an attacker can craft a modulated light signal that adversarially illuminates a scene and causes targeted misclassifications on a state-of-the-art ImageNet deep learning model. Concretely, we exploit the radiometric rolling shutter effect in commodity cameras to create precise striping patterns that appear on images. To human eyes, it appears like the object is illuminated, but the camera creates an image with stripes that will cause ML models to output the attacker-desired classification. We conduct a range of simulation and physical experiments with LEDs, demonstrating targeted attack rates up to 84%. Athena Sayles, Ashish Hooda, Mohit Gupta 0001, Rahul Chatterjee 0001, Earlence Fernandes |
CVPR | 3 |
| 2021 | Photon-Starved Scene Inference using Single Photon CamerasabstractScene understanding under low-light conditions is a challenging problem. This is due to the small number of photons captured by the camera and the resulting low signal-to-noise ratio (SNR). Single-photon cameras (SPCs) are an emerging sensing modality that are capable of capturing images with high sensitivity. Despite having minimal read-noise, images captured by SPCs in photon-starved conditions still suffer from strong shot noise, preventing reliable scene inference. We propose photon scale-space – a collection of high-SNR images spanning a wide range of photons-per-pixel (PPP) levels (but same scene content) as guides to train inference model on low photon flux images. We develop training techniques that push images with different illumination levels closer to each other in feature representation space. The key idea is that having a spectrum of different brightness levels during training enables effective guidance, and increases robustness to shot noise even in extreme noise cases. Based on the proposed approach, we demonstrate, via simulations and real experiments with a SPAD camera, high-performance on various inference tasks such as image classification and monocular depth estimation under ultra low-light, down to < 1 PPP. Project Page: https://wisionlab.cs.wisc.edu/project/photon-net Bhavya Goyal, Mohit Gupta 0001 |
ICCV | 2 |
| 2020 | Smart Time-Multiplexing of Quads Solves the Multicamera Interference ProblemabstractTime-of-flight (ToF) cameras are becoming increasingly popular for 3D imaging. Their optimal usage has been studied from the several aspects. One of the open research problems is the possibility of a multicamera interference problem when two or more ToF cameras are operating simultaneously. In this work we present an efficient method to synchronize multiple operating ToF cameras. Our method is based on the time-division multiplexing, but unlike traditional time multiplexing, it does not decrease the effective camera frame rate. Additionally, for unsynchronized cameras, we provide a robust method to extract from their corresponding video streams, frames which are not subject to multicamera interference problem. We demonstrate our approach through a series of experiments and with a different level of support available for triggering, ranging from a hardware triggering to purely random software triggering. Tomislav Pribanic, Tomislav Petkovic, David Bojanic, Kristijan Bartol, Mohit Gupta 0001 |
3DV | 5 |
| 2020 | Inertial Safety from Structured Light
Sizhuo Ma, Mohit Gupta 0001 |
ECCV (23) | 2 |
| 2020 | Differential Scene Flow from Light Field Gradients
Sizhuo Ma, Brandon M. Smith 0001, Mohit Gupta 0001 |
Int. J. Comput. Vis. | 3 |
| 2020 | Quanta burst photographyabstractSingle-photon avalanche diodes (SPADs) are an emerging sensor technology capable of detecting individual incident photons, and capturing their time-of-arrival with high timing precision. While these sensors were limited to singlepixel or low-resolution devices in the past, recently, large (up to 1 MPixel) SPAD arrays have been developed. These single-photon cameras (SPCs) are capable of capturing high-speed sequences of binary single-photon images with no read noise. We present quanta burst photography, a computational photography technique that leverages SPCs as passive imaging devices for photography in challenging conditions, including ultra low-light and fast motion. Inspired by recent success of conventional burst photography, we design algorithms that align and merge binary sequences captured by SPCs into intensity images with minimal motion blur and artifacts, high signal-to-noise ratio (SNR), and high dynamic range. We theoretically analyze the SNR and dynamic range of quanta burst photography, and identify the imaging regimes where it provides significant benefits. We demonstrate, via a recently developed SPAD array, that the proposed method is able to generate high-quality images for scenes with challenging lighting, complex geometries, high dynamic range and moving objects. With the ongoing development of SPAD arrays, we envision quanta burst photography finding applications in both consumer and scientific photography. Sizhuo Ma, Arin C. Ulku, Claudio Bruschini, Edoardo Charbon, Mohit Gupta 0001 |
ACM Trans. Graph. | 6 |
| 2019 | Photon-Flooded Single-Photon 3D CamerasabstractSingle-photon avalanche diodes (SPADs) are starting to play a pivotal role in the development of photon-efficient, long-range LiDAR systems. However, due to non-linearities in their image formation model, a high photon flux (e.g., due to strong sunlight) leads to distortion of the incident temporal waveform, and potentially, large depth errors. Operating SPADs in low flux regimes can mitigate these distortions, but, often requires attenuating the signal and thus, results in low signal-to-noise ratio. In this paper, we address the following basic question: what is the optimal photon flux that a SPAD-based LiDAR should be operated in? We derive a closed form expression for the optimal flux, which is quasi-depth-invariant, and depends on the ambient light strength. The optimal flux is lower than what a SPAD typically measures in real world scenarios, but surprisingly, considerably higher than what is conventionally suggested for avoiding distortions. We propose a simple, adaptive approach for achieving the optimal flux by attenuating incident flux based on an estimate of ambient light strength. Using extensive simulations and a hardware prototype, we show that the optimal flux criterion holds for several depth estimators, under a wide range of illumination conditions. Anant Gupta, Atul Ingle, Andreas Velten, Mohit Gupta 0001 |
CVPR | 4 |
| 2019 | Practical Coding Function Design for Time-Of-Flight ImagingabstractThe depth resolution of a continuous-wave time-of-flight (CW-ToF) imaging system is determined by its coding functions. Recently, there has been growing interest in the design of new high-performance CW-ToF coding functions. However, these functions are typically designed in a hardware agnostic manner, i.e., without considering the practical device limitations, such as bandwidth, source power, digital (binary) function generation. Therefore, despite theoretical improvements, practical implementation of these functions remains a challenge. We present a constrained optimization approach for designing practical coding functions that adhere to hardware constraints. The optimization problem is non-convex with a large search space and no known globally optimal solutions. To make the problem tractable, we design an iterative, alternating least-squares algorithm, along with convex relaxation of the constraints. Using this approach, we design high-performance coding functions that can be implemented on existing hardware with minimal modifications. We demonstrate the performance benefits of the resulting functions via extensive simulations and a hardware prototype. Felipe Gutierrez-Barragan, Syed Azer Reza, Andreas Velten, Mohit Gupta 0001 |
CVPR | 4 |
| 2019 | High Flux Passive Imaging With Single-Photon SensorsabstractSingle-photon avalanche diodes (SPADs) are an emerging technology with a unique capability of capturing individual photons with high timing precision. SPADs are being used in several active imaging systems (e.g., fluorescence lifetime microscopy and LiDAR), albeit mostly limited to low photon flux settings. We propose passive free-running SPAD (PF-SPAD) imaging, an imaging modality that uses SPADs for capturing 2D intensity images with unprecedented dynamic range under ambient lighting, without any active light source. Our key observation is that the precise inter-photon timing measured by a SPAD can be used for estimating scene brightness under ambient lighting conditions, even for very bright scenes. We develop a theoretical model for PF-SPAD imaging, and derive a scene brightness estimator based on the average time of darkness between successive photons detected by a PF-SPAD pixel. Our key insight is that due to the stochastic nature of photon arrivals, this estimator does not suffer from a hard saturation limit. Coupled with high sensitivity at low flux, this enables a PF-SPAD pixel to measure a wide range of scene brightnesses, from very low to very high, thereby achieving extreme dynamic range. We demonstrate an improvement of over 2 orders of magnitude over conventional sensors by imaging scenes spanning a dynamic range of 10^6:1. Atul Ingle, Andreas Velten, Mohit Gupta 0001 |
CVPR | 3 |
| 2019 | Asynchronous Single-Photon 3D ImagingabstractSingle-photon avalanche diodes (SPADs) are becoming popular in time-of-flight depth-ranging due to their unique ability to capture individual photons with picosecond timing resolution. However, ambient light (e.g., sunlight) incident on a SPAD-based 3D camera leads to severe non-linear distortions (pileup) in the measured waveform, resulting in large depth errors. We propose asynchronous single-photon 3D imaging, a family of acquisition schemes to mitigate pileup during data acquisition itself. Asynchronous acquisition temporally misaligns SPAD measurement windows and the laser cycles through deterministically predefined or randomized offsets. Our key insight is that pileup distortions can be “averaged out” by choosing a sequence of offsets that span the entire depth range. We develop a generalized image formation model and perform theoretical analysis to explore the space of asynchronous acquisition schemes and design high-performance schemes. Our simulations and experiments demonstrate an improvement in depth accuracy of up to an order of magnitude as compared to the state-of-the-art, across a wide range of imaging scenarios, including those with high ambient flux. Anant Gupta, Atul Ingle, Mohit Gupta 0001 |
ICCV | 3 |
| 2019 | Stochastic Exposure Coding for Handling Multi-ToF-Camera InterferenceabstractAs continuous-wave time-of-flight (C-ToF) cameras become popular in 3D imaging applications, they need to contend with the problem of multi-camera interference (MCI). In a multi-camera environment, a ToF camera may receive light from the sources of other cameras, resulting in large depth errors. In this paper, we propose stochastic exposure coding (SEC), a novel approach for mitigating. SEC involves dividing a camera's integration time into multiple slots, and switching the camera off and on stochastically during each slot. This approach has two benefits. First, by appropriately choosing the on probability for each slot, the camera can effectively filter out both the AC and DC components of interfering signals, thereby mitigating depth errors while also maintaining high signal-to-noise ratio. This enables high accuracy depth recovery with low power consumption. Second, this approach can be implemented without modifying the C-ToF camera's coding functions, and thus, can be used with a wide range of cameras with minimal changes. We demonstrate the performance benefits of SEC with theoretical analysis, simulations and real experiments, across a wide range of imaging scenarios. Jongho Lee 0004, Mohit Gupta 0001 |
ICCV | 2 |
| 2019 | Micro-Baseline Structured LightabstractWe propose Micro-baseline Structured Light (MSL), a novel 3D imaging approach designed for small form-factor devices such as cell-phones and miniature robots. MSL operates with small projector-camera baseline and low-cost projection hardware, and can recover scene depths with computationally lightweight algorithms. The main observation is that a small baseline leads to small disparities, enabling a first-order approximation of the non-linear SL image formation model. This leads to the key theoretical result of the paper: the MSL equation, a linearized version of SL image formation. MSL equation is under-constrained due to two unknowns (depth and albedo) at each pixel, but can be efficiently solved using a local least squares approach. We analyze the performance of MSL in terms of various system parameters such as projected pattern and baseline, and provide guidelines for optimizing performance. Armed with these insights, we build a prototype to experimentally examine the theory and its practicality. Vishwanath Saragadam, Raja Venkata, Jian Wang 0100, Shree K. Nayar, Mohit Gupta 0001 |
ICCV | 5 |
| 2019 | Coding Scheme Optimization for Fast Fluorescence Lifetime ImagingabstractFluorescence lifetime imaging (FLIM) is used for measuring material properties in a wide range of applications, including biology, medical imaging, chemistry, and material science. In frequency-domain FLIM (FD-FLIM), the object of interest is illuminated with a temporally modulated light source. The fluorescence lifetime is measured by computing the correlations of the emitted light with a demodulation function at the sensor. The signal-to-noise ratio (SNR) and the acquisition time of a FD-FLIM system is determined by the coding scheme (modulation and demodulation functions). In this article, we develop theory and algorithms for designing high-performance FD-FLIM coding schemes that can achieve high SNR and short acquisition time, given a fixed source power budget. Based on a geometric analysis of the image formation and noise model, we propose a novel surrogate objective for the performance of a given coding scheme. The surrogate objective is extremely fast to compute, and can be used to efficiently explore the entire space of coding schemes. Based on this objective, we design novel, high-performance coding schemes that achieve up to an order of magnitude shorter acquisition time as compared to existing approaches. We demonstrate the performance advantage of the proposed schemes in a variety of imaging conditions, using a modular hardware prototype that can implement various coding schemes. Jongho Lee 0004, Jenu Varghese Chacko, Bing Dai, Syed Azer Reza, Abdul Kader Sagar, Kevin W. Eliceiri, Andreas Velten, Mohit Gupta 0001 |
ACM Trans. Graph. | 8 |
| 2018 | Tracking Multiple Objects Outside the Line of Sight Using Speckle ImagingabstractThis paper presents techniques for tracking non-line-of-sight (NLOS) objects using speckle imaging. We develop a novel speckle formation and motion model where both the sensor and the source view objects only indirectly via a diffuse wall. We show that this NLOS imaging scenario is analogous to direct LOS imaging with the wall acting as a virtual, bare (lens-less) sensor. This enables tracking of a single, rigidly moving NLOS object using existing speckle-based motion estimation techniques. However, when imaging multiple NLOS objects, the speckle components due to different objects are superimposed on the virtual bare sensor image, and cannot be analyzed separately for recovering the motion of individual objects. We develop a novel clustering algorithm based on the statistical and geometrical properties of speckle images, which enables identifying the motion trajectories of multiple, independently moving NLOS objects. We demonstrate, for the first time, tracking individual trajectories of multiple objects around a corner with extreme precision (<; 10 microns) using only off-the-shelf imaging components. Brandon M. Smith 0001, Matthew O'Toole, Mohit Gupta 0001 |
CVPR | 3 |
| 2018 | Trapping Light for Time of FlightabstractWe propose a novel imaging method for near-complete, surround, 3D reconstruction of geometrically complex objects, in a single scan. The key idea is to augment a time-of-flight (ToF) based 3D sensor with a multi-mirror system, called a light-trap. The shape of the trap is chosen so that light rays entering it bounce multiple times inside the trap, thereby visiting every position inside the trap multiple times from various directions. We show via simulations that this enables light rays to reach more than 99.9% of the surface of objects placed inside the trap, even those with strong occlusions, for example, lattice-shaped objects. The ToF sensor provides the path length for each light ray, which, along with the known shape of the trap, is used to reconstruct the complete paths of all the rays. This enables performing dense, surround 3D reconstructions of objects with highly complex 3D shapes, in a single scan. We have developed a proof-of-concept hardware prototype consisting of a pulsed ToF sensor, and a light trap built with planar mirrors. We demonstrate the effectiveness of the light trap based 3D reconstruction method on a variety of objects with a broad range of geometry and reflectance properties. Ruilin Xu 0001, Mohit Gupta 0001, Shree K. Nayar |
CVPR | 2 |
| 2018 | A Geometric Perspective on Structured Light Coding
Mohit Gupta 0001, Nikhil Nakhate |
ECCV (16) | 1 |
| 2018 | 3D Scene Flow from 4D Light Field Gradients
Sizhuo Ma, Brandon M. Smith 0001, Mohit Gupta 0001 |
ECCV (8) | 3 |
| 2018 | SH-ToF: Micro resolution time-of-flight imaging with superheterodyne interferometryabstractThree dimensional imaging techniques have been widely used in both industry and academia. Time-of-flight (ToF) sensors offer a promising method of 3D imaging due to compact size and low complexity. However, state-of-the-art ToF sensors only have depth resolutions of centimeters due to limitations in the modulation frequencies that can be used. In this paper, we propose a technique to generate modulation frequencies as high as 1 THz using optical superheterodyne interferometry. Our proposed system provides great flexibility in imaging range and resolution. We experimentally demonstrate an increase in depth resolution by an order of magnitude relative to currently available commercial ToF cameras. Fengqiang Li, Florian Willomitzer, Prasanna Rangarajan, Mohit Gupta 0001, Andreas Velten, Oliver Cossairt |
ICCP | 4 |
| 2018 | What Are Optimal Coding Functions for Time-of-Flight Imaging?abstractThe depth resolution achieved by a continuous wave time-of-flight (C-ToF) imaging system is determined by the coding (modulation and demodulation) functions that it uses. Almost all current C-ToF systems use sinusoid or square coding functions, resulting in a limited depth resolution. In this article, we present a mathematical framework for exploring and characterizing the space of C-ToF coding functions in a geometrically intuitive space. Using this framework, we design families of novel coding functions that are based on Hamiltonian cycles on hypercube graphs. Given a fixed total source power and acquisition time, the new Hamiltonian coding scheme can achieve up to an order of magnitude higher resolution as compared to the current state-of-the-art methods, especially in low signal-to-noise ratio (SNR) settings. We also develop a comprehensive physically-motivated simulator for C-ToF cameras that can be used to evaluate various coding schemes prior to a real hardware implementation. Since most off-the-shelf C-ToF sensors use sinusoid or square functions, we develop a hardware prototype that can implement a wide range of coding functions. Using this prototype and our software simulator, we demonstrate the performance advantages of the proposed Hamiltonian coding functions in a wide range of imaging settings. Mohit Gupta 0001, Andreas Velten, Shree K. Nayar, Eric Breitbach |
ACM Trans. Graph. | 1 |
| 2017 | CoLux: multi-object 3D micro-motion analysis using speckle imagingabstractWe present CoLux, a novel system for measuring micro 3D motion of multiple independently moving objects at macroscopic standoff distances. CoLux is based on speckle imaging, where the scene is illuminated with a coherent light source and imaged with a camera. Coherent light, on interacting with optically rough surfaces, creates a high-frequency speckle pattern in the captured images. The motion of objects results in movement of speckle, which can be measured to estimate the object motion. Speckle imaging is widely used for micro-motion estimation in several applications, including industrial inspection, scientific imaging, and user interfaces (e.g., optical mice). However, current speckle imaging methods are largely limited to measuring 2D motion (parallel to the sensor image plane) of a single rigid object. We develop a novel theoretical model for speckle movement due to multi-object motion, and present a simple technique based on global scale-space speckle motion analysis for measuring small (5--50 microns) compound motion of multiple objects, along all three axes. Using these tools, we develop a method for measuring 3D micro-motion histograms of multiple independently moving objects, without tracking the individual motion trajectories. In order to demonstrate the capabilities of CoLux, we develop a hardware prototype and a proof-of-concept subtle hand gesture recognition system with a broad range of potential applications in user interfaces and interactive computer graphics. Brandon M. Smith 0001, Pratham Desai, Vishal Agarwal, Mohit Gupta 0001 |
ACM Trans. Graph. | 4 |
| 2016 | Dual Structured Light 3D Using a 1D Sensor
Jian Wang 0100, Aswin C. Sankaranarayanan, Mohit Gupta 0001, Srinivasa G. Narasimhan |
ECCV (6) | 3 |
| 2016 | DisCo: Display-Camera Communication Using Rolling Shutter SensorsabstractWe present DisCo, a novel display-camera communication system. DisCo enables displays and cameras to communicate with each other while also displaying and capturing images for human consumption. Messages are transmitted by temporally modulating the display brightness at high frequencies so that they are imperceptible to humans. Messages are received by a rolling shutter camera that converts the temporally modulated incident light into a spatial flicker pattern. In the captured image, the flicker pattern is superimposed on the pattern shown on the display. The flicker and the display pattern are separated by capturing two images with different exposures. The proposed system performs robustly in challenging real-world situations such as occlusion, variable display size, defocus blur, perspective distortion, and camera rotation. Unlike several existing visible light communication methods, DisCo works with off-the-shelf image sensors. It is compatible with a variety of sources (including displays, single LEDs), as well as reflective surfaces illuminated with light sources. We have built hardware prototypes that demonstrate DisCo’s performance in several scenarios. Because of its robustness, speed, ease of use, and generality, DisCo can be widely deployed in several applications, such as advertising, pairing of displays with cell phones, tagging objects in stores and museums, and indoor navigation. Kensei Jo, Mohit Gupta 0001, Shree K. Nayar |
ACM Trans. Graph. | 2 |
| 2015 | MC3D: Motion Contrast 3D ScanningabstractStructured light 3D scanning systems are fundamentally constrained by limited sensor bandwidth and light source power, hindering their performance in real-world applications where depth information is essential, such as industrial automation, autonomous transportation, robotic surgery, and entertainment. We present a novel structured light technique called Motion Contrast 3D scanning (MC3D) that maximizes bandwidth and light source power to avoid performance trade-offs. The technique utilizes motion contrast cameras that sense temporal gradients asynchronously, i.e., independently for each pixel, a property that minimizes redundant sampling. This allows laser scanning resolution with single-shot speed, even in the presence of strong ambient illumination, significant inter-reflections, and highly reflective surfaces. The proposed approach will allow 3D vision systems to be deployed in challenging and hitherto inaccessible real-world scenarios requiring high performance using limited power and bandwidth. Nathan Matsuda, Oliver Cossairt, Mohit Gupta 0001 |
ICCP | 3 |
| 2015 | LiSens- A Scalable Architecture for Video Compressive SensingabstractThe measurement rate of cameras that take spatially multiplexed measurements by using spatial light modulators (SLM) is often limited by the switching speed of the SLMs. This is especially true for single-pixel cameras where the photodetector operates at a rate that is many orders-of-magnitude greater than the SLM. We study the factors that determine the measurement rate for such spatial multiplexing cameras (SMC) and show that increasing the number of pixels in the device improves the measurement rate, but there is an optimum number of pixels (typically, few thousands) beyond which the measurement rate does not increase. This motivates the design of LiSens, a novel imaging architecture, that replaces the photodetector in the single-pixel camera with a 1D linear array or a line-sensor. We illustrate the optical architecture underlying LiSens, build a prototype, and demonstrate results of a range of indoor and outdoor scenes. LiSens delivers on the promise of SMCs: imaging at a megapixel resolution, at video rate, using an inexpensive low-resolution sensor. Jian Wang 0100, Mohit Gupta 0001, Aswin C. Sankaranarayanan |
ICCP | 2 |
| 2015 | SpeDo: 6 DOF Ego-Motion Sensor Using Speckle Defocus ImagingabstractSensors that measure their motion with respect to the surrounding environment (ego-motion sensors) can be broadly classified into two categories. First is inertial sensors such as accelerometers. In order to estimate position and velocity, these sensors integrate the measured acceleration, which often results in accumulation of large errors over time. Second, camera-based approaches such as SLAM that can measure position directly, but their performance depends on the surrounding scene's properties. These approaches cannot function reliably if the scene has low frequency textures or small depth variations. We present a novel ego-motion sensor called SpeDo that addresses these fundamental limitations. SpeDo is based on using coherent light sources and cameras with large defocus. Coherent light, on interacting with a scene, creates a high frequency interferometric pattern in the captured images, called speckle. We develop a theoretical model for speckle flow (motion of speckle as a function of sensor motion), and show that it is quasi-invariant to surrounding scene's properties. As a result, SpeDo can measure ego-motion (not derivative of motion) simply by estimating optical flow at a few image locations. We have built a low-cost and compact hardware prototype of SpeDo and demonstrated high precision 6 DOF ego-motion estimation for complex trajectories in scenarios where the scene properties are challenging (e.g., repeating or no texture) as well as unknown. Kensei Jo, Mohit Gupta 0001, Shree K. Nayar |
ICCV | 2 |
| 2015 | Phasor Imaging: A Generalization of Correlation-Based Time-of-Flight ImagingabstractIn correlation-based time-of-flight (C-ToF) imaging systems, light sources with temporally varying intensities illuminate the scene. Due to global illumination, the temporally varying radiance received at the sensor is a combination of light received along multiple paths. Recovering scene properties (e.g., scene depths) from the received radiance requires separating these contributions, which is challenging due to the complexity of global illumination and the additional temporal dimension of the radiance. We propose phasor imaging, a framework for performing fast inverse light transport analysis using C-ToF sensors. Phasor imaging is based on the idea that, by representing light transport quantities as phasors and light transport events as phasor transformations, light transport analysis can be simplified in the temporal frequency domain. We study the effect of temporal illumination frequencies on light transport and show that, for a broad range of scenes, global radiance (inter-reflections and volumetric scattering) vanishes for frequencies higher than a scene-dependent threshold. We use this observation for developing two novel scene recovery techniques. First, we present micro-ToF imaging, a ToF-based shape recovery technique that is robust to errors due to inter-reflections (multipath interference) and volumetric scattering. Second, we present a technique for separating the direct and global components of radiance. Both techniques require capturing as few as 3--4 images and minimal computations. We demonstrate the validity of the presented techniques via simulations and experiments performed with our hardware prototype. Mohit Gupta 0001, Shree K. Nayar, Matthias B. Hullin, Jaime Martín |
ACM Trans. Graph. | 1 |
| 2014 | Recovering Scene Geometry under Wavy Fluid via Distortion and Defocus Analysis
Mohit Gupta 0001, Jin-Li Suo, Qionghai Dai |
ECCV (5) | 3 |
| 2014 | Digital refocusing with incoherent holographyabstractLight field cameras allow us to digitally refocus a photograph after the time of capture. However, recording a light field requires either a significant loss in spatial resolution [11, 21, 10] or a large number of images to be captured [12]. In this paper, we propose incoherent holography for digital refocusing without loss of spatial resolution from only 3 captured images. The main idea is to capture 2D coherent holograms of the scene instead of the 4D light fields. The key properties of coherent light propagation are that the coherent spread function (hologram of a single point source) encodes scene depths and has a broadband spatial frequency response. These properties enable digital refocusing with 2D coherent holograms, which can be captured on sensors without loss of spatial resolution. Incoherent holography does not require illuminating the scene with high power coherent laser, making it possible to acquire holograms even for passively illuminated scenes. We provide an in-depth performance comparison between light field and incoherent holographic cameras in terms of the signal-to-noise-ratio (SNR). We show that given the same sensing resources, an incoherent holography camera outperforms light field cameras in most real world settings. We demonstrate a prototype incoherent holography camera capable of performing digital refocusing from only 3 acquired images. We show results on a variety of scenes that verify the accuracy of our theoretical analysis. Oliver Cossairt, Nathan Matsuda, Mohit Gupta 0001 |
ICCP | 3 |
| 2014 | Efficient Space-Time Sampling with Pixel-Wise Coded Exposure for High-Speed ImagingabstractCameras face a fundamental trade-off between spatial and temporal resolution. Digital still cameras can capture images with high spatial resolution, but most high-speed video cameras have relatively low spatial resolution. It is hard to overcome this trade-off without incurring a significant increase in hardware costs. In this paper, we propose techniques for sampling, representing, and reconstructing the space-time volume to overcome this trade-off. Our approach has two important distinctions compared to previous works: 1) We achieve sparse representation of videos by learning an overcomplete dictionary on video patches, and 2) we adhere to practical hardware constraints on sampling schemes imposed by architectures of current image sensors, which means that our sampling function can be implemented on CMOS image sensors with modified control units in the future. We evaluate components of our approach, sampling function and sparse representation, by comparing them to several existing approaches. We also implement a prototype imaging system with pixel-wise coded exposure control using a liquid crystal on silicon device. System characteristics such as field of view and modulation transfer function are evaluated for our imaging system. Both simulations and experiments on a wide range of scenes show that our method can effectively reconstruct a video from a single coded image while maintaining high spatial resolution. Dengyu Liu, Jinwei Gu, Yasunobu Hitomi, Mohit Gupta 0001, Tomoo Mitsunaga, Shree K. Nayar |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2013 | Fibonacci Exposure Bracketing for High Dynamic Range ImagingabstractExposure bracketing for high dynamic range (HDR) imaging involves capturing several images of the scene at different exposures. If either the camera or the scene moves during capture, the captured images must be registered. Large exposure differences between bracketed images lead to inaccurate registration, resulting in artifacts such as ghosting (multiple copies of scene objects) and blur. We present two techniques, one for image capture (Fibonacci exposure bracketing) and one for image registration (generalized registration), to prevent such motion-related artifacts. Fibonacci bracketing involves capturing a sequence of images such that each exposure time is the sum of the previous N(N > 1) exposures. Generalized registration involves estimating motion between sums of contiguous sets of frames, instead of between individual frames. Together, the two techniques ensure that motion is always estimated between frames of the same total exposure time. This results in HDR images and videos which have both a large dynamic range and minimal motion-related artifacts. We show, by results for several real-world indoor and outdoor scenes, that the proposed approach significantly outperforms several existing bracketing schemes. Mohit Gupta 0001, Daisuke Iso, Shree K. Nayar |
ICCV | 1 |
| 2013 | Structured Light in SunlightabstractStrong ambient illumination severely degrades the performance of structured light based techniques. This is especially true in outdoor scenarios, where the structured light sources have to compete with sunlight, whose power is often 2-5 orders of magnitude larger than the projected light. In this paper, we propose the concept of light concentration to overcome strong ambient illumination. Our key observation is that given a fixed light (power) budget, it is always better to allocate it sequentially in several portions of the scene, as compared to spreading it over the entire scene at once. For a desired level of accuracy, we show that by distributing light appropriately, the proposed approach requires 1-2 orders lower acquisition time than existing approaches. Our approach is illumination-adaptive as the optimal light distribution is determined based on a measurement of the ambient illumination level. Since current light sources have a fixed light distribution, we have built a prototype light source that supports flexible light distribution by controlling the scanning speed of a laser scanner. We show several high quality 3D scanning results in a wide range of outdoor scenarios. The proposed approach will benefit 3D vision systems that need to operate outdoors under extreme ambient illumination levels on a limited time and power budget. Mohit Gupta 0001, Qi Yin, Shree K. Nayar |
ICCV | 1 |
| 2013 | A Practical Approach to 3D Scanning in the Presence of Interreflections, Subsurface Scattering and Defocus
Mohit Gupta 0001, Amit K. Agrawal, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
Int. J. Comput. Vis. | 1 |
| 2013 | When Does Computational Imaging Improve Performance?abstractA number of computational imaging techniques are introduced to improve image quality by increasing light throughput. These techniques use optical coding to measure a stronger signal level. However, the performance of these techniques is limited by the decoding step, which amplifies noise. Although it is well understood that optical coding can increase performance at low light levels, little is known about the quantitative performance advantage of computational imaging in general settings. In this paper, we derive the performance bounds for various computational imaging techniques. We then discuss the implications of these bounds for several real-world scenarios (e.g., illumination conditions, scene properties, and sensor noise characteristics). Our results show that computational imaging techniques do not provide a significant performance advantage when imaging with illumination that is brighter than typical daylight. These results can be readily used by practitioners to design the most suitable imaging systems given the application at hand. Oliver Cossairt, Mohit Gupta 0001, Shree K. Nayar |
IEEE Trans. Image Process. | 2 |
| 2012 | Micro Phase ShiftingabstractWe consider the problem of shape recovery for real world scenes, where a variety of global illumination (inter-reflections, subsurface scattering, etc.) and illumination defocus effects are present. These effects introduce systematic and often significant errors in the recovered shape. We introduce a structured light technique called Micro Phase Shifting, which overcomes these problems. The key idea is to project sinusoidal patterns with frequencies limited to a narrow, high-frequency band. These patterns produce a set of images over which global illumination and defocus effects remain constant for each point in the scene. This enables high quality reconstructions of scenes which have traditionally been considered hard, using only a small number of images. We also derive theoretical lower bounds on the number of input images needed for phase shifting and show that Micro PS achieves the bound. Mohit Gupta 0001, Shree K. Nayar |
CVPR | 1 |
| 2012 | Diffuse structured lightabstractToday, structured light systems are widely used in applications such as robotic assembly, visual inspection, surgery, entertainment, games and digitization of cultural heritage. Current structured light methods are faced with two serious limitations. First, they are unable to cope with scene regions that produce strong highlights due to specular reflection. Second, they cannot recover useful information for regions that lie within shadows. We observe that many structured light methods use illumination patterns that have translational symmetry, i.e., two-dimensional patterns that vary only along one of the two dimensions. We show that, for this class of patterns, diffusion of the patterns along the axis of translation can mitigate the adverse effects of specularities and shadows. We show results for two applications - 3D scanning using phase shifting of sinusoidal patterns and separation of direct and global components of light transport using high-frequency binary stripes. Shree K. Nayar, Mohit Gupta 0001 |
ICCP | 2 |
| 2012 | A Combined Theory of Defocused Illumination and Global Light Transport
Mohit Gupta 0001, Yuandong Tian, Srinivasa G. Narasimhan, Li Zhang 0003 |
Int. J. Comput. Vis. | 1 |
| 2011 | Structured light 3D scanning in the presence of global illuminationabstractGlobal illumination effects such as inter-reflections, diffusion and sub-surface scattering severely degrade the performance of structured light-based 3D scanning. In this paper, we analyze the errors caused by global illumination in structured light-based shape recovery. Based on this analysis, we design structured light patterns that are resilient to individual global illumination effects using simple logical operations and tools from combinatorial mathematics. Scenes exhibiting multiple phenomena are handled by combining results from a small ensemble of such patterns. This combination also allows us to detect any residual errors that are corrected by acquiring a few additional images. Our techniques do not require explicit separation of the direct and global components of scene radiance and hence work even in scenarios where the separation fails or the direct component is too low. Our methods can be readily incorporated into existing scanning systems without significant overhead in terms of capture time or hardware. We show results on a variety of scenes with complex shape and material properties and challenging global illumination effects. Mohit Gupta 0001, Amit K. Agrawal, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
CVPR | 1 |
| 2011 | Multiplexed illumination for scene recovery in the presence of global illuminationabstractGlobal illumination effects such as inter-reflections and subsurface scattering result in systematic, and often significant errors in scene recovery using active illumination. Recently, it was shown that the direct and global components could be separated efficiently for a scene illuminated with a single light source. In this paper, we study the problem of direct-global separation for multiple light sources. We derive a theoretical lower bound for the number of required images, and propose a multiplexed illumination scheme which achieves this lower bound. We analyze the signal-to-noise ratio (SNR) characteristics of the proposed illumination multiplexing method in the context of direct-global separation. We apply our method to several scene recovery techniques requiring multiple light sources, including shape from shading, structured light 3D scanning, photometric stereo, and reflectance estimation. Both simulation and experimental results show that the proposed method can accurately recover scene information with fewer images compared to sequentially separating direct-global components for each light source. Jinwei Gu, Toshihiro Kobayashi, Mohit Gupta 0001, Shree K. Nayar |
ICCV | 3 |
| 2011 | Video from a single coded exposure photograph using a learned over-complete dictionaryabstractCameras face a fundamental tradeoff between the spatial and temporal resolution - digital still cameras can capture images with high spatial resolution, but most high-speed video cameras suffer from low spatial resolution. It is hard to overcome this tradeoff without incurring a significant increase in hardware costs. In this paper, we propose techniques for sampling, representing and reconstructing the space-time volume in order to overcome this tradeoff. Our approach has two important distinctions compared to previous works: (1) we achieve sparse representation of videos by learning an over-complete dictionary on video patches, and (2) we adhere to practical constraints on sampling scheme which is imposed by architectures of present image sensor devices. Consequently, our sampling scheme can be implemented on image sensors by making a straightforward modification to the control unit. To demonstrate the power of our approach, we have implemented a prototype imaging system with per-pixel coded exposure control using a liquid crystal on silicon (LCoS) device. Using both simulations and experiments on a wide range of scenes, we show that our method can effectively reconstruct a video from a single image maintaining high spatial resolution. Yasunobu Hitomi, Jinwei Gu, Mohit Gupta 0001, Tomoo Mitsunaga, Shree K. Nayar |
ICCV | 3 |
| 2010 | Optimal coded sampling for temporal super-resolutionabstractConventional low frame rate cameras result in blur and/or aliasing in images while capturing fast dynamic events. Multiple low speed cameras have been used previously with staggered sampling to increase the temporal resolution. However, previous approaches are inefficient: they either use small integration time for each camera which does not provide light benefit, or use large integration time in a way that requires solving a big ill-posed linear system. We propose coded sampling that address these issues: using N cameras it allows N times temporal superresolution while allowing ~N/2 times more light compared to an equivalent high speed camera. In addition, it results in a well-posed linear system which can be solved independently for each frame, avoiding reconstruction artifacts and significantly reducing the computational time and memory. Our proposed sampling uses optimal multiplexing code considering additive Gaussian noise to achieve the maximum possible SNR in the recovered video. We show how to implement coded sampling on off-the-shelf machine vision cameras. We also propose a new class of invertible codes that allow continuous blur in captured frames, leading to an easier hardware implementation. Amit K. Agrawal, Mohit Gupta 0001, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
CVPR | 2 |
| 2010 | Flexible Voxels for Motion-Aware Videography
Mohit Gupta 0001, Amit K. Agrawal, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
ECCV (1) | 1 |
| 2009 | (De) focusing on global light transport for active scene recoveryabstractMost active scene recovery techniques assume that a scene point is illuminated only directly by the illumination source. Consequently, global illumination effects due to inter-reflections, sub-surface scattering and volumetric scattering introduce strong biases in the recovered scene shape. Our goal is to recover scene properties in the presence of global illumination. To this end, we study the interplay between global illumination and the depth cue of illumination defocus. By expressing both these effects as low pass filters, we derive an approximate invariant that can be used to separate them without explicitly modeling the light transport. This is directly useful in any scenario where limited depth-of-field devices (such as projectors) are used to illuminate scenes with global light transport and significant depth variations. We show two applications: (a) accurate depth recovery in the presence of global illumination, and (b) factoring out the effects of defocus for correct direct-global separation in large depth scenes. We demonstrate our approach using scenes with complex shapes, reflectances, textures and translucencies. Mohit Gupta 0001, Yuandong Tian, Srinivasa G. Narasimhan, Li Zhang 0003 |
CVPR | 1 |
| 2008 | On controlling light transport in poor visibility environmentsabstractPoor visibility conditions due to murky water, bad weather, dust and smoke severely impede the performance of vision systems. Passive methods have been used to restore scene contrast under moderate visibility by digital post-processing. However, these methods are ineffective when the quality of acquired images is poor to begin with. In this work, we design active lighting and sensing systems for controlling light transport before image formation, and hence obtain higher quality data. First, we present a technique of polarized light striping based on combining polarization imaging and structured light striping. We show that this technique out-performs different existing illumination and sensing methodologies. Second, we present a numerical approach for computing the optimal relative sensor-source position, which results in the best quality image. Our analysis accounts for the limits imposed by sensor noise. Mohit Gupta 0001, Srinivasa G. Narasimhan, Yoav Y. Schechner |
CVPR | 1 |
| 2008 | High Resolution Tracking of Non-Rigid Motion of Densely Sampled 3D Data Using Harmonic Maps
Yang Wang 0001, Mohit Gupta 0001, Song Zhang 0002, Xianfeng Gu, Dimitris Samaras, Peisen Huang |
Int. J. Comput. Vis. | 2 |
| 2006 | Acquiring scattering properties of participating media by dilutionabstractThe visual world around us displays a rich set of volumetric effects due to participating media. The appearance of these media is governed by several physical properties such as particle densities, shapes and sizes, which must be input (directly or indirectly) to a rendering algorithm to generate realistic images. While there has been significant progress in developing rendering techniques (for instance, volumetric Monte Carlo methods and analytic approximations), there are very few methods that measure or estimate these properties for media that are of relevance to computer graphics. In this paper, we present a simple device and technique for robustly estimating the properties of a broad class of participating media that can be either (a) diluted in water such as juices, beverages, paints and cleaning supplies, or (b) dissolved in water such as powders and sugar/salt crystals, or (c) suspended in water such as impurities. The key idea is to dilute the concentrations of the media so that single scattering effects dominate and multiple scattering becomes negligible, leading to a simple and robust estimation algorithm. Furthermore, unlike previous approaches that require complicated or separate measurement setups for different types or properties of media, our method and setup can be used to measure media with a complete range of absorption and scattering properties from a single HDR photograph. Once the parameters of the diluted medium are estimated, a volumetric Monte Carlo technique may be used to create renderings of any medium concentration and with multiple scattering. We have measured the scattering parameters of forty commonly found materials, that can be immediately used by the computer graphics community. We can also create realistic images of combinations or mixtures of the original measured materials, thus giving the user a wide flexibility in making realistic images of participating media. Srinivasa G. Narasimhan, Mohit Gupta 0001, Craig Donner, Ravi Ramamoorthi, Shree K. Nayar, Henrik Wann Jensen |
ACM Trans. Graph. | 2 |
| 2005 | Face Modeling and Analysis in Stony Brook UniversityabstractIn this paper, we present our latest work on facial expression analysis, synthesis and face recognition. The advent of new technologies that allow the capture of massive amounts of high resolution, high frame rate face data, leads us to propose data-driven face models that accurately describe the appearance of faces under unknown pose and illumination conditions as well as to track subtle geometry changes that occur during expressions. In this paper, we also demonstrate our results for expression transfer among different subjects. We reduce the dimensionality of our data onto a lower dimensional space manifold and then decompose it into style and content parameters. This allows us to transfer subtle expression information (in the form of a style vector) between individuals to synthesize new expressions, as well as smoothly morph geometry and motion. Finally, we demonstrate the accuracy of our face modeling methods through an integrated example of image-driven re-targeting and relighting of facial expressions, where transfer of expression and illumination information between different individuals is possible. Dimitris Samaras, Yang Wang 0001, Lei Zhang 0002, Mohit Gupta 0001 |
CVPR (2) | 5 |
| 2005 | High Resolution Tracking of Non-Rigid 3D Motion of Densely Sampled Data Using Harmonic MapsabstractWe present a novel fully automatic method for high resolution, nonrigid dense 3D point tracking. High quality dense point clouds of nonrigid geometry moving at video speeds are acquired using a phase-shifting structured light ranging technique. To use such data for the temporal study of subtle motions such as those seen in facial expressions, an efficient nonrigid 3D motion tracking algorithm is needed to establish inter-frame correspondences. The novelty of this paper is the development of an algorithmic framework for 3D tracking that unifies tracking of intensity and geometric features, using harmonic maps with added feature correspondence constraints. While the previous uses of harmonic maps provided only global alignment, the proposed introduction of interior feature constraints guarantees that nonrigid deformations are accurately tracked as well. The harmonic map between two topological disks is a diffeomorphism with minimal stretching energy and bounded angle distortion. The map is stable, insensitive to resolution changes and is robust to noise. Due to the strong implicit and explicit smoothness constraints imposed by the algorithm and the high-resolution data, the resulting registration/deformation field is smooth, continuous and gives dense one-to-one inter-frame correspondences. Our method is validated through a series of experiments demonstrating its accuracy and efficiency. Yang Wang 0001, Mohit Gupta 0001, Song Zhang 0002, Xianfeng Gu, Dimitris Samaras, Peisen Huang |
ICCV | 2 |