Oliver Cossairt

dblp:37/3504 · also Oliver S. Cossairt · DBLP profile ↗
← Back
49ranked-venue papers
6as first author
22since 2021 · last 2025
0000-0003-0501-0163ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 43 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 9 · 6 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Holospeed: High-Speed Holographic Displays for Dynamic Content
abstract
Holographic displays are plagued by speckle - noise-like artifacts caused by the coherent interference of laser light. To mitigate this challenge, state-of-the-art systems use time multiplexing on fast spatial light modulators (SLMs) to effectively temporally smooth out these effects. In our work, we observe that such an approach struggles in practice in the context of dynamic content, manifesting motion blur and stroboscopic artifacts thanks to a fundamental mismatch between expected and displayed motion. To tackle this challenge, we propose a paradigm of holographic high-speed display, where we use the underlying fast SLM to reproduce target content that changes at the same framerate. Approaches built using this paradigm mitigate motion blur and strobing, and simultaneously minimize speckle and maximize contrast with the right loss functions. We demonstrate such a methodology in both simulation and a real system.
Dorian Chan, Oliver Cossairt, Nathan Matsuda, Grace Kuo
ICCP2
2025 Artifact-Resilient Real-Time Holography
abstract
Holographic near-eye displays promise unparalleled depth cues, high-resolution imagery, and realistic three-dimensional parallax at a compact form factor, making them promising candidates for emerging augmented and virtual reality systems. However, existing holographic display methods often assume ideal viewing conditions and overlook real-world factors such as eye floaters and eyelashes—obstructions that can severely degrade perceived image quality. In this work, we propose a new metric that quantifies hologram resilience to artifacts and apply it to computer generated holography (CGH) optimization. We call this Artifact Resilient Holography (ARH). We begin by introducing a simulation method that models the effects of pre- and post-pupil obstructions on holographic displays. Our analysis reveals that eyebox regions dominated by low frequencies—produced especially by the smooth-phase holograms broadly adopted in recent holography work—are vulnerable to visual degradation from dynamic obstructions such as floaters and eyelashes. In contrast, random phase holograms spread energy more uniformly across the eyebox spectrum, enabling them to diffract around obstructions without producing prominent artifacts. By characterizing a random phase eyebox using the Rayleigh Distribution, we derive a differentiable metric in the eyebox domain. We then apply this metric to train a real-time neural network-based phase generator, enabling it to produce artifact-resilient 3D holograms that preserve visual fidelity across a range of practical viewing conditions—enhancing both robustness and user interactivity.
Victor Chu, Oscar Pueyo-Ciutad, Ethan Tseng, Florian Andreas Schiffers, Grace Kuo, Nathan Matsuda, Albert Redo-Sanchez, Douglas Lanman, Oliver Cossairt, Felix Heide
ACM Trans. Graph.9
2025 HoloChrome: Polychromatic Illumination for Speckle Reduction in Holographic Near-Eye Displays
abstract
Holographic displays hold the promise of providing authentic depth cues, resulting in enhanced immersive visual experiences for near-eye applications. However, current holographic displays are hindered by speckle noise, which limits accurate reproduction of color and texture in displayed images. We present HoloChrome, a polychromatic holographic display framework designed to mitigate these limitations. HoloChrome utilizes an ultrafast, wavelength-adjustable laser and a dual-Spatial Light Modulator (SLM) architecture, enabling the multiplexing of a large set of discrete wavelengths across the visible spectrum. By leveraging spatial separation in our dual-SLM setup, we independently manipulate speckle patterns across multiple wavelengths. This novel approach effectively reduces speckle noise through incoherent averaging achieved by wavelength multiplexing, specifically by using a single SLM pattern to modulate multiple wavelengths simultaneously on one or more SLM devices. Our method is complementary to existing speckle reduction techniques, offering a new pathway to address this challenge. Furthermore, the use of polychromatic illumination broadens the achievable color gamut compared to traditional three-color primary holographic displays. Our simulations and tabletop experiments validate that HoloChrome significantly reduces speckle noise and expands the color gamut. These advancements enhance the performance of holographic near-eye displays, moving us closer to practical, immersive next-generation visual experiences.
Florian Andreas Schiffers, Grace Kuo, Nathan Matsuda, Douglas Lanman, Oliver Cossairt
ACM Trans. Graph.5
2024 A Joint Intensity-Neuromorphic Event Imaging System With Bandwidth-Limited Communication Channel
abstract
We present a novel adaptive multimodal intensity-event algorithm to optimize an overall objective of object tracking under bit rate constraints for a host-chip architecture. The chip is a computationally resource-constrained device acquiring high-resolution intensity frames and events, while the host is capable of performing computationally expensive tasks. We develop a joint intensity-neuromorphic event rate-distortion compression framework with a quadtree (QT)-based compression of intensity and events scheme. The goal of this compression framework is to optimally allocate bits to the intensity frames and neuromorphic events based on the minimum distortion at a given communication channel capacity. The data acquisition on the chip is driven by the presence of objects of interest in the scene as detected by an object detector. The most informative intensity and event data are communicated to the host under rate constraints so that the best possible tracking performance is obtained. The detection and tracking of objects in the scene are done on the distorted data at the host. Intensity and events are jointly used in a fusion framework to enhance the quality of the distorted images, in order to improve the object detection and tracking performance. The performance assessment of the overall system is done in terms of the multiple object tracking accuracy (MOTA) score. Compared with using intensity modality only, there is an improvement in MOTA using both these modalities in different scenarios.
Srutarshi Banerjee, Henry Chopp, Zihao W. Wang, Oliver Cossairt, Aggelos K. Katsaggelos
IEEE Trans. Neural Networks Learn. Syst.6
2023 Thermal Spread Functions (TSF): Physics-Guided Material Classification
abstract
Robust and non-destructive material classification is a challenging but crucial first-step in numerous vision applications. We propose a physics-guided material classification framework that relies on thermal properties of the object. Our key observation is that the rate of heating and cooling of an object depends on the unique intrinsic properties of the material, namely the emissivity and diffusivity. We leverage this observation by gently heating the objects in the scene with a low-power laser for a fixed duration and then turning it off, while a thermal camera captures measurements during the heating and cooling process. We then take this spatial and temporal “thermal spread function” (TSF) to solve an inverse heat equation using the finite-differences approach, resulting in a spatially varying estimate of diffusivity and emissivity. These tuples are then used to train a classifier that produces a fine-grained material label at each spatial pixel. Our approach is extremely simple requiring only a small light source (low power laser) and a thermal camera, and produces robust classification results with 86% accuracy over 16 classes11Code: https://github.com/aniketdashpute/TSF.
Aniket Dashpute, Vishwanath Saragadam, Emma Alexander, Florian Willomitzer, Aggelos K. Katsaggelos, Ashok Veeraraghavan, Oliver Cossairt
CVPR7
2023 Stochastic Light Field Holography
abstract
The Visual Turing Test is the ultimate goal to evaluate the realism of holographic displays. Previous studies have focused on addressing challenges such as limited étendue and image quality over a large focal volume, but they have not investigated the effect of pupil sampling on the viewing experience in full 3D holograms. In this work, we tackle this problem with a novel hologram generation algorithm motivated by matching the projection operators of incoherent (Light Field) and coherent (Wigner Function) light transport. To this end, we supervise hologram computation using synthesized photographs, which are rendered on-the-fly using Light Field refocusing from stochastically sampled pupil states during optimization. The proposed method produces holograms with correct parallax and focus cues, which are important for passing the Visual Turing Test. We validate that our approach compares favorably to state-of-the-art CGH algorithms that use Light Field and Focal Stack supervision. Our experiments demonstrate that our algorithm improves the viewing experience when evaluated under a large variety of different pupil states.
Florian Schiffers, Praneeth Chakravarthula, Nathan Matsuda, Grace Kuo, Ethan Tseng, Douglas Lanman, Felix Heide, Oliver Cossairt
ICCP8
2023 Spiking GLOM: Bio-Inspired Architecture for Next-Generation Object Recognition
abstract
Today, artificial neural networks (ANNs) have demonstrated extraordinary abilities in many cognition tasks. Nevertheless, the limitations of many ANN-based techniques are evident, such as the low energy efficiency and the lack of interpretability. To alleviate these problems, researchers have directed their attention to bio-inspired models, including energy-efficient Spiking Neural Networks (SNNs) and the GLOM model representing part-whole hierarchies in neural networks. In this paper, we propose a novel bio-inspired solution to next-generation object recognition. Specifically, we propose an energy-efficient and interpretable model – Spiking GLOM by introducing spiking neurons and neuronal dynamics into the GLOM model. Moreover, we evaluate our model and its variants on CIFAR-10. Extensive experiments demonstrate the effectiveness of our proposed models for object recognition and show the superiority of our models in energy efficiency and interpretability.
Srutarshi Banerjee, Henry Chopp, Aggelos K. Katsaggelos, Oliver Cossairt
ICIP5
2023 Simultaneous Color Computer Generated Holography
abstract
Computer generated holography has long been touted as the future of augmented and virtual reality (AR/VR) displays, but has yet to be realized in practice. Previous high-quality, color holographic displays have made either a 3 × sacrifice on frame rate by using a sequential color illumination scheme or used more than one spatial light modulator (SLM) and/or bulky, complex optical setups. The reduced frame rate of sequential color introduces distracting judder and color fringing in the presence of head motion while the form factor of current simultaneous color systems is incompatible with a head-mounted display. In this work, we propose a framework for simultaneous color holography that allows the use of the full SLM frame rate while maintaining a compact and simple optical setup. Simultaneous color holograms are optimized through the use of a perceptual loss function, a physics-based neural network wavefront propagator, and a camera-calibrated forward model. We measurably improve hologram quality compared to other simultaneous color methods and move one step closer to the realization of color holographic displays for AR/VR.
Eric Markley, Nathan Matsuda, Florian Schiffers, Oliver Cossairt, Grace Kuo
SIGGRAPH Asia4
2023 Multisource Holography
abstract
Holographic displays promise several benefits including high quality 3D imagery, accurate accommodation cues, and compact form-factors. However, holography relies on coherent illumination which can create undesirable speckle noise in the final image. Although smooth phase holograms can be speckle-free, their non-uniform eyebox makes them impractical, and speckle mitigation with partially coherent sources also reduces resolution. Averaging sequential frames for speckle reduction requires high speed modulators and consumes temporal bandwidth that may be needed elsewhere in the system. In this work, we propose multisource holography, a novel architecture that uses an array of sources to suppress speckle in a single frame without sacrificing resolution. By using two spatial light modulators, arranged sequentially, each source in the array can be controlled almost independently to create a version of the target content with different speckle. Speckle is then suppressed when the contributions from the multiple sources are averaged at the image plane. We introduce an algorithm to calculate multisource holograms, analyze the design space, and demonstrate up to a 10 dB increase in peak signal-to-noise ratio compared to an equivalent single source system. Finally, we validate the concept with a benchtop experimental prototype by producing both 2D images and focal stacks with natural defocus cues.
Grace Kuo, Florian Schiffers, Douglas Lanman, Oliver Cossairt, Nathan Matsuda
ACM Trans. Graph.4
2022 Investigating the Potential of Auxiliary-Classifier Gans for Image Classification in Low Data Regimes
abstract
Generative Adversarial Networks (GANs) have shown promise in augmenting datasets and boosting convolutional neural network (CNN) performance on image classification tasks. But they introduce more hyperparameters to tune as well as the need for additional time and computational power to train, supplementary to the CNN. In this work, we examine the potential for Auxiliary-Classifier GANs (AC-GANs) as a ’one-stop-shop’ architecture for image classification, particularly in low data regimes. Additionally, we explore modifications to the typical AC-GAN framework, changing the generator’s latent space sampling scheme and employing a Wasserstein loss with gradient penalty to stabilize the simultaneous training of image synthesis and classification. Through experiments on images of varying resolutions and complexity, we demonstrate that AC-GANs show promise in image classification, achieving competitive performance with standard CNNs. These methods can be employed as an ’all-in-one’ framework with particular utility in the absence of large amounts of training data.
Amil Dravid, Florian Schiffers, Yunan Wu, Oliver Cossairt, Aggelos K. Katsaggelos
ICASSP4
2022 Event - Driven Tactile Learning with Location Spiking Neurons
abstract
The sense of touch is essential for a variety of daily tasks. New advances in event-based tactile sensors and Spiking Neural Networks (SNNs) spur the research in event-driven tactile learning. However, SNN -enabled event-driven tactile learning is still in its infancy due to the limited representative abilities of existing spiking neurons and high spatio-temporal complexity in the data. In this paper, to improve the representative capabilities of existing spiking neurons, we propose a novel neuron model called “location spiking neuron”, which enables us to extract features of event-based data in a novel way. Moreover, based on the classical Time Spike Response Model (TSRM), we develop a specific location spiking neuron model - Location Spike Response Model (LSRM) that serves as a new building block of SNNs11The TSRM is the classical SRM in the literature. We add the character “T” to highlight its difference with the LSRM.• Furthermore, we propose a hybrid model which combines an SNN with TSRM neurons and an SNN with LSRM neurons to capture the complex spatio-temporal dependencies in the data. Extensive experiments demonstrate the significant improvements of our models over other works on event-driven tactile learning and show the superior energy efficiency of our models and location spiking neurons, which may unlock their potential on neuromorphic hardware.
Srutarshi Banerjee, Henry Chopp, Aggelos K. Katsaggelos, Oliver Cossairt
IJCNN5
2022 Guided Event Filtering: Synergy Between Intensity Images and Neuromorphic Events for High Performance Imaging
abstract
Many visual and robotics tasks in real-world scenarios rely on robust handling of high speed motion and high dynamic range (HDR) with effectively high spatial resolution and low noise. Such stringent requirements, however, cannot be directly satisfied by a single imager or imaging modality, rather by multi-modal sensors with complementary advantages. In this paper, we address high performance imaging by exploring the synergy between traditional frame-based sensors with high spatial resolution and low sensor noise, and emerging event-based sensors with high speed and high dynamic range. We introduce a novel computational framework, termed Guided Event Filtering (GEF), to process these two streams of input data and output a stream of super-resolved yet noise-reduced events. To generate high quality events, GEF first registers the captured noisy events onto the guidance image plane according to our flow model. it then performs joint image filtering that inherits the mutual structure from both inputs. Lastly, GEF re-distributes the filtered event frame in the space-time volume while preserving the statistical characteristics of the original events. When the guidance images under-perform, GEF incorporates an event self-guiding mechanism that resorts to neighbor events for guidance. We demonstrate the benefits of GEF by applying the output high quality events to existing event-based algorithms across diverse application categories, including high speed object tracking, depth estimation, high frame-rate video synthesis, and super resolution/HDR/color image restoration.
Peiqi Duan 0002, Zihao W. Wang, Boxin Shi, Oliver Cossairt, Tiejun Huang 0001, Aggelos K. Katsaggelos
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Pupil-Aware Holography
abstract
Holographic displays promise to deliver unprecedented display capabilities in augmented reality applications, featuring a wide field of view, wide color gamut, spatial resolution, and depth cues all in a compact form factor. While emerging holographic display approaches have been successful in achieving large étendue and high image quality as seen by a camera, the large étendue also reveals a problem that makes existing displays impractical: the sampling of the holographic field by the eye pupil. Existing methods have not investigated this issue due to the lack of displays with large enough étendue, and, as such, they suffer from severe artifacts with varying eye pupil size and location. We show that the holographic field as sampled by the eye pupil is highly varying for existing display setups, and we propose pupil-aware holography that maximizes the perceptual image quality irrespective of the size, location, and orientation of the eye pupil in a near-eye holographic display. We validate the proposed approach both in simulations and on a prototype holographic display and show that our method eliminates severe artifacts and significantly outperforms existing approaches.
Praneeth Chakravarthula, Seung-Hwan Baek, Florian Schiffers, Ethan Tseng, Grace Kuo, Andrew Maimone, Nathan Matsuda, Oliver Cossairt, Douglas Lanman, Felix Heide
ACM Trans. Graph.8
2021 Depth from Defocus as a Special Case of the Transport of Intensity Equation
abstract
The Transport of Intensity Equation (TIE) in microscopy and the Depth from Differential Defocus (DfDD) method in photography both describe the effect of a small change in defocus on image intensity. They are based on different assumptions and may appear to contradict each other. Using the Wigner Distribution Function, we show that DfDD can be interpreted as a special case of the TIE, well-suited to applications where the generalized phase measurements recovered by the TIE are connected to depth rather than phase, such as photography and fluorescence microscopy. The level of spatial coherence is identified as the driving factor in the trade-off between the usefulness of each technique. Specifically, the generalized phase corresponds to the sample's phase under high-coherence illumination and reveals scene depth in low-coherence settings. When coherence varies spatially, as in multi-modal phase and fluorescence microscopy, we show that complementary information is available in different regions of the image.
Emma Alexander, Leyla A. Kabuli, Oliver Cossairt, Laura Waller
ICCP3
2021 SeLFVi: Self-supervised Light-Field Video Reconstruction from Stereo Video
abstract
Light-field imaging is appealing to the mobile devices market because of its capability for intuitive post-capture processing. Acquiring light field (LF) data with high angular, spatial and temporal resolution poses significant challenges, especially with space constraints preventing bulky optics. At the same time, stereo video capture, now available on many consumer devices, can be interpreted as a sparse LF-capture. We explore the application of small baseline stereo videos for reconstructing high fidelity LF videos.We propose a self-supervised learning-based algorithm for LF video reconstruction from stereo video. The self- supervised LF video reconstruction is guided via the geometric information from the individual stereo pairs and the temporal information from the video sequence. LF estimation is further regularized by a low-rank constraint based on layered LF displays. The proposed self-supervised algorithm facilitates advantages such as post-training finetuning on test sequences and variable angular view interpolation and extrapolation. Quantitatively the reconstructed LF videos show higher fidelity than previously proposed unsupervised approaches. We demonstrate our results via LF videos generated from publicly available stereo videos acquired from commercially available stereoscopic cameras. Finally, we demonstrate that our reconstructed LF videos allow applications such as post-capture focus control and region-of-interest (RoI) based focus tracking for videos.
Prasan A. Shedligeri, Florian Schiffers, Sushobhan Ghosh, Oliver Cossairt, Kaushik Mitra
ICCV4
2021 Lossy Event Compression Based On Image-Derived Quad Trees And Poisson Disk Sampling
abstract
Event cameras have provided new opportunities for tackling visual tasks under challenging scenarios over conventional RGB cameras. However, not much focus has been given on event compression algorithms. The main challenge for compressing events is its unique asynchronous form. To address this problem, we propose a novel event compression algorithm based on a quad tree (QT) segmentation map derived from the adjacent intensity images. The QT informs 2D spatial priority within the 3D space-time volume. In the event encoding step, events are first aggregated over time to form polarity-based event histograms. The histograms are then variably sampled via Poisson Disk Sampling prioritized by the QT based segmentation map. Next, differential encoding and run length encoding are employed for encoding the spatial and polarity information of the sampled events, respectively, followed by Huffman encoding to produce the final encoded events. Our algorithm achieves greater than $6 \times$ higher compression compared to the state of the art.
Srutarshi Banerjee, Zihao W. Wang, Henry Chopp, Oliver Cossairt, Aggelos K. Katsaggelos
ICIP4
2021 A Data Fusion Method For The Delayering Of X-Ray Fluorescence Images Of Painted Works Of Art
abstract
In this manuscript, we address the problem of studying layer structure in X-ray Fluorescence (XRF) elemental maps of paintings through the incorporation of reflectance imaging spectral data in the visible or near IR range. We propose a conceptually flexible approach, which involves an initial clustering step for the visible hyperspectral reflectance data (RIS) and the formation of a synthetic surface XRF image. Considering the difference of the full and synthetic surface XRF images, surface and subsurface correlated features are then identified. Results are demonstrated on real and simulated data.
Lionel Fiske, Aggelos K. Katsaggelos, Maurice C. G. Aalders, Matthias Alfeld, Marc Walton, Oliver Cossairt
ICIP6
2021 A Two-Stage Framework for Compound Figure Separation
abstract
Scientific literature contains large volumes of complex, unstructured figures that are compound in nature (i.e. composed of multiple images, graphs, and drawings). Separation of these compound figures is critical for information retrieval from these figures. In this paper, we propose a new strategy for compound Figure separation, which decomposes the compound figures into constituent subfigures while preserving the association between the subfigures and their respective caption components. We propose a two-stage framework to address the proposed compound Figure separation problem. In the first stage, the subfigure label detection module detects all subfigure labels. Then, in the subfigure detection module, the detected subfigure labels help to detect the subfigures by optimizing the feature selection process and providing the global layout information as extra features. Extensive experiments are conducted to validate the effectiveness and superiority of the proposed framework, which improves the precision by 9% compared to the benchmark.
Weixin Jiang, Eric Schwenker, Trevor Spreadbury, Nicola J. Ferrier, Maria K. Y. Chan, Oliver Cossairt
ICIP6
2021 Human Vision-Like Robust Object Recognition
abstract
Previous research always solely utilizes Artificial Neural Networks (ANNs) or Spiking Neural Networks (SNNs) for object recognition. However, evidence in neuroscience suggests that the visual processing in human vision is performed hierarchically in the combination of analog and digital processing. To construct a more human vision-like object recognition system, we propose a general hierarchical ANN-SNN model. We evaluate our model and its variants on two popular datasets to show its effectiveness, robustness, efficiency, and generality. Extensive experiments clearly demonstrate the superiority of our proposed models for robust object recognition.
Srutarshi Banerjee, Henry Chopp, Aggelos K. Katsaggelos, Oliver Cossairt
ICIP6
2021 Skinscan: Low-Cost 3D-Scanning for Dermatologic Diagnosis and Documentation
abstract
The utilization of computational photography becomes increasingly essential in the medical field. Today, imaging techniques for dermatology range from two-dimensional (2D) color imagery with a mobile device to professional clinical imaging systems measuring additional detailed three-dimensional (3D) data. The latter are commonly expensive and not accessible to a broad audience. In this work, we propose a novel system and software framework that relies only on low-cost (and even mobile) commodity devices present in every household to measure detailed 3D information of the human skin with a 3D-gradient-illumination-based method. We believe that our system has great potential for early-stage diagnosis and monitoring of skin diseases, especially in vastly populated or underdeveloped areas.
Merlin A. Nau, Florian Schiffers, Andreas K. Maier, Jack Tumblin, Marc Walton, Aggelos K. Katsaggelos, Florian Willomitzer, Oliver Cossairt
ICIP10
2021 Improving Acquisition Speed of X-Ray Ptychography Through Spatial Undersampling and Regularization
abstract
X-ray ptychography is one of the versatile techniques for nanometer resolution imaging. The magnitude of the diffraction patterns is recorded on a detector, and the phase of the diffraction patterns is estimated using phase retrieval techniques. Most phase retrieval algorithms make the solution well-posed by relying on the constraints imposed by the overlapping region between neighboring diffraction pattern samples. As the overlap between neighboring diffraction patterns reduces, the problem becomes ill-posed, and the object cannot be recovered. To avoid the ill-posedness, we investigate the effect of regularizing the phase retrieval algorithm with image priors for various overlap ratios between the neighboring diffraction patterns. We show that the object can be faithfully reconstructed at low overlap ratios by regularizing the phase retrieval algorithm with image priors such as Total-Variation prior and Structure Tensor Prior. We also show the effectiveness of our proposed algorithm on real data acquired from an IC chip with a coherent X-ray beam.
Prasan A. Shedligeri, Florian Schiffers, Semih Barutcu, Pablo Ruiz 0002, Aggelos K. Katsaggelos, Oliver Cossairt
ICIP6
2021 Exploiting Wavelength Diversity for High Resolution Time-of-Flight 3D Imaging
abstract
The poor lateral and depth resolution of state-of-the-art 3D sensors based on the time-of-flight (ToF) principle has limited widespread adoption to a few niche applications. In this work, we introduce a novel sensor concept that provides ToF-based 3D measurements of real world objects and surfaces with depth precision up to 35 μm and point cloud densities commensurate with the native sensor resolution of standard CMOS/CCD detectors (up to several megapixels). Such capabilities are realized by combining the best attributes of continuous wave ToF sensing, multi-wavelength interferometry, and heterodyne interferometry into a single approach. We describe multiple embodiments of the approach, each featuring a different sensing modality and associated tradeoffs.
Fengqiang Li, Florian Willomitzer, Muralidhar Madabhushi Balaji, Prasanna Rangarajan, Oliver Cossairt
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 Joint Filtering of Intensity Images and Neuromorphic Events for High-Resolution Noise-Robust Imaging
abstract
We present a novel computational imaging system with high resolution and low noise. Our system consists of a traditional video camera which captures high-resolution intensity images, and an event camera which encodes high-speed motion as a stream of asynchronous binary events. To process the hybrid input, we propose a unifying framework that first bridges the two sensing modalities via a noise-robust motion compensation model, and then performs joint image filtering. The filtered output represents the temporal gradient of the captured space-time volume, which can be viewed as motion-compensated event frames with high resolution and low noise. Therefore, the output can be widely applied to many existing event-based algorithms that are highly dependent on spatial resolution and noise robustness. In experimental results performed on both publicly available datasets as well as our contributing RGB-DAVIS dataset, we show systematic performance improvement in applications such as high frame-rate video synthesis, feature/corner detection and tracking, as well as high dynamic range image reconstruction.
Zihao W. Wang, Peiqi Duan 0002, Oliver Cossairt, Aggelos K. Katsaggelos, Tiejun Huang 0001, Boxin Shi
CVPR3
2020 WISHED: Wavefront imaging sensor with high resolution and depth ranging
abstract
Phase-retrieval based wavefront sensors have been shown to reconstruct the complex field from an object with a high spatial resolution. Although the reconstructed complex field encodes the depth information of the object, it is impractical to be used as a depth sensor for macroscopic objects, since the unambiguous depth imaging range is limited by the optical wavelength. To improve the depth range of imaging and handle depth discontinuities, we propose a novel three-dimensional sensor by leveraging wavelength diversity and wavefront sensing. Complex fields at two optical wavelengths are recorded, and a synthetic wavelength can be generated by correlating those wavefronts. The proposed system achieves high lateral and depth resolutions. Our experimental prototype shows an unambiguous range of more than 1,000 x larger compared with the optical wavelengths, while the depth precision is up to 9µm for smooth objects and up to 69µm for rough objects. We experimentally demonstrate 3D reconstructions for transparent, translucent, and opaque objects with smooth and rough surfaces.
Fengqiang Li, Florian Willomitzer, Ashok Veeraraghavan, Oliver Cossairt
ICCP5
2020 Adaptive Image Sampling Using Deep Learning and Its Application on X-Ray Fluorescence Image Reconstruction
abstract
This paper presents an adaptive image sampling algorithm based on Deep Learning (DL). It consists of an adaptive sampling mask generation network which is jointly trained with an image inpainting network. The sampling rate is controlled by the mask generation network, and a binarization strategy is investigated to make the sampling mask binary. In addition to the image sampling and reconstruction process, we show how it can be extended and used to speed up raster scanning such as the X-Ray fluorescence (XRF) image scanning process. Recently XRF laboratory-based systems have evolved into lightweight and portable instruments thanks to technological advancements in both X-Ray generation and detection. However, the scanning time of an XRF image is usually long due to the long exposure requirements (e.g., 100 μs - 1 ms per point). We propose an XRF image in painting approach to address the long scanning times, thus speeding up the scanning process, while being able to reconstruct a high quality XRF image. The proposed adaptive image sampling algorithm is applied to the RGB image of the scanning target to generate the sampling mask. The XRF scanner is then driven according to the sampling mask to scan a subset of the total image pixels. Finally, we inpaint the scanned XRF image by fusing the RGB image to reconstruct the full scan XRF image. The experiments show that the proposed adaptive sampling algorithm is able to effectively sample the image and achieve a better reconstruction accuracy than that of existing methods.
Qiqin Dai, Henry Chopp, Emeline Pouyet, Oliver Cossairt, Marc Walton, Aggelos K. Katsaggelos
IEEE Trans. Multim.4
2019 Multi-frame Super-resolution for Time-of-flight Imaging
abstract
Recently, time-of-flight (ToF) sensors have emerged as a promising three-dimensional sensing technology that can be manufactured inexpensively in a compact size. However, current state-of-the-art ToF sensors suffer from low spatial resolution due to physical limitations in the fabrication process. In this paper, we analyze the ToF sensor's output as a complex value coupling the depth and intensity information in a phasor representation. Based on this analysis, we introduce a novel multi-frame superresolution technique that can improve both spatial resolution in intensity and depth images simultaneously. We believe our proposed method can benefit numerous applications where high resolution depth sensing is desirable, such as precision automated navigation and collision avoidance.
Fengqiang Li, Pablo Ruiz 0002, Oliver Cossairt, Aggelos K. Katsaggelos
ICASSP3
2019 Pigment Unmixing of Hyperspectral Images of Paintings Using Deep Neural Networks
abstract
In this paper, the problem of automatic nonlinear unmixing of hyperspectral reflectance data using works of art as test cases is described. We use a deep neural network to decompose a given spectrum quantitatively to the abundance values of pure pigments. We show that adding another step to identify the constituent pigments of a given spectrum leads to more accurate unmixing results. Towards this, we use another deep neural network to identify pigments first and integrate this information to different layers of the network used for pigment unmixing. As a test set, the hyperspectral images of a set of mock-up paintings consisting of a broad palette of pigment mixtures, and pure pigment exemplars, were measured. The results of the algorithm on the mock-up test set are reported and analyzed.
Neda Rohani, Emeline Pouyet, Marc Walton, Oliver Cossairt, Aggelos K. Katsaggelos
ICASSP4
2018 3D Image Reconstruction from Multi-Focus Microscope: Axial Super-Resolution and Multiple-Frame Processing
abstract
Multi-focus microscope (MFM) provides a way to obtain 3D information by simultaneously capturing multiple focal planes. The naive method for MFM reconstruction is to stack the sub-images with alignment. However, the resolution in the z-axis in this method is limited by the number of acquired focal planes. In this work we build on a recent reconstruction algorithm for MFM, using information from multiple frames to improve the reconstruction quality. We propose two multiple-frame MFM image reconstruction algorithms: batch and recursive approaches. In the batch approach, we take multiple MFM frames and jointly estimate the 3D image and the motion for each frame. In the recursive approach, we utilize the reconstructed image from the previous frame. Experimental results show that the proposed algorithms produce a sequence of 3D object reconstruction with high quality that enable reconstruction of dynamic extended objects.
Seunghwan Yoo, Pablo Ruiz 0002, Xiang Huang 0006, Kuan He, Nicola J. Ferrier, Mark Hereld, Alan Selewa, Matthew Daddysman, Norbert Scherer, Oliver Cossairt, Aggelos K. Katsaggelos
ICASSP10
2018 ADP: Automatic differentiation ptychography
abstract
Ptychography is an imaging technique which aims to recover the complex-valued exit wavefront of an object from a set of its diffraction pattern magnitudes. Ptychography is one of the most popular techniques for sub-30 nanometer imaging as it does not suffer from the limitations of typical lens based imaging techniques. The object can be reconstructed from the captured diffraction patterns using iterative phase retrieval algorithms. Over time many algorithms have been proposed for iterative reconstruction of the object based on manually derived update rules. In this paper, we adapt automatic differentiation framework to solve practical and complex ptychographic phase retrieval problems and demonstrate its advantages in terms of speed, accuracy, adaptability and generalizability across different scanning techniques.
Sushobhan Ghosh, Youssef S. G. Nashed, Oliver Cossairt, Aggelos K. Katsaggelos
ICCP3
2018 SH-ToF: Micro resolution time-of-flight imaging with superheterodyne interferometry
abstract
Three dimensional imaging techniques have been widely used in both industry and academia. Time-of-flight (ToF) sensors offer a promising method of 3D imaging due to compact size and low complexity. However, state-of-the-art ToF sensors only have depth resolutions of centimeters due to limitations in the modulation frequencies that can be used. In this paper, we propose a technique to generate modulation frequencies as high as 1 THz using optical superheterodyne interferometry. Our proposed system provides great flexibility in imaging range and resolution. We experimentally demonstrate an increase in depth resolution by an order of magnitude relative to currently available commercial ToF cameras.
Fengqiang Li, Florian Willomitzer, Prasanna Rangarajan, Mohit Gupta 0001, Andreas Velten, Oliver Cossairt
ICCP6
2018 An Interior Point Method for Nonnegative Sparse Signal Reconstruction
abstract
We present a primal-dual interior point method (IPM) with a novel preconditioner to solve the ℓ1-norm regularized least square problem for nonnegative sparse signal reconstruction. IPM is a second-order method that uses both gradient and Hessian information to compute effective search directions and achieve super-linear convergence rates. It therefore requires many fewer iterations than first-order methods such as iterative shrinkage/thresholding algorithms (ISTA) that only achieve sub-linear convergence rates. However, each iteration of IPM is more expensive than in ISTA because it needs to evaluate an inverse of a Hessian matrix to compute the Newton direction. We propose to approximate each Hessian matrix by a diagonal matrix plus a rank-one matrix. This approximation matrix is easily invertible using the Sherman-Morrison formula, and is used as a novel preconditioner in a preconditioned conjugate gradient method to compute a truncated Newton direction. We demonstrate the efficiency of our algorithm in compressive 3D volumetric image reconstruction. Numerical experiments show favorable results of our method in comparison with previous interior point based and iterative shrinkage/thresholding based algorithms.
Xiang Huang 0006, Kuan He, Seunghwan Yoo, Oliver Cossairt, Aggelos K. Katsaggelos, Nicola J. Ferrier, Mark Hereld
ICIP4
2018 Bayesian Approach for Automatic Joint Parameter Estimation in 3D Image Reconstruction from Multi-Focus Microscope
abstract
We present a Bayesian approach for 3D image reconstruction of an extended object imaged with multi-focus microscopy (MFM). MFM simultaneously captures multiple sub-images of different focal planes to provide 3D information of the sample. The naive method to reconstruct the object is to stack the sub-images along the z-axis, but the result suffers from poor resolution in the z-axis. The maximum a posteriori framework provides a way to reconstruct a 3D image according to its observation model and prior knowledge. It jointly estimates the 3D image and the model parameters. Experimental results with synthetic and real experimental data show that it enables the high-quality 3D reconstruction of an extended object from MFM.
Seunghwan Yoo, Pablo Ruiz 0002, Xiang Huang 0006, Kuan He, Itay Gdor, Alan Selewa, Matthew Daddysman, Nicola J. Ferrier, Mark Hereld, Norbert Scherer, Oliver Cossairt, Aggelos K. Katsaggelos
ICIP12
2017 Shape-from-Shifting: Uncalibrated Photometric Stereo with a Mobile Device
abstract
Surface shape scanning techniques, such as laser scanning and photometric stereo, are widespread analytical tools used in the field of cultural heritage. Compared to regular 2D RGB photos, 3D surface scans provide higher fidelity of an object's surface shape which assist conservators, art historians, and archaeologists in understanding how these artworks and artifacts are made and to digitally document them for purposes of conservation. However, current state-of-the-art 3D surface scanning tools used in art conservation are often expensive and bulky-such as light dome structures that are often over 1 m in diameter. In this paper, we introduce mobile shape-from-shifting (SfS): a simple, low-cost and streamlined photometric stereo framework for scanning planar surfaces with a consumer mobile device coupled to a low-cost add-on component. Our free-form mobile SfS framework relaxes the rigorous hardware and other complex requirements inherent to conventional 3D scanning tools. This is achieved by taking a sequence of photos with the on-board camera and flash of a mobile device. The sequence of captures are used to reconstruct high quality normal maps using nearlight photometric stereo algorithms, which are of comparable quality to conventional photometric stereo. We demonstrate 3D surface reconstructions with SfS on different materials and scales. Moreover, the mobile SfS technique can be used "in the wild" so that 3D scans may be performed in their natural environment, eliminating the need for transport to a laboratory setting. With the elegant design and low cost, we believe our Mobile SfS can greatly benefit the conservation community by providing a userfriendly and cost-effective solution for 3D surface scanning.
Chia-Kai Yeh, Fengqiang Li, Gianluca Pastorelli, Marc Walton, Aggelos K. Katsaggelos, Oliver Cossairt
eScience6
2017 Linear systems approach to identifying performance bounds in indirect imaging
abstract
Light scattering on diffuse rough surfaces was long assumed to destroy geometry and photometry information about hidden (non line of sight) objects making `looking around the corner' (LATC) and `non line of sight' (NLOS) imaging impractical. Recent work pioneered by Kirmani et al. [1], Velten et al. [2] demonstrated that transient information (time of flight information) from these scattered third bounce photons can be exploited to solve LATC and NLOS imaging. In this paper, we quantify the geometric and photometric reconstruction limits of LATC and NLOS imaging for the first time using a classical linear systems approach. The relationship between the albedo of the voxels in a hidden volume to the third bounce measurements at the sensor is a linear system that is determined by the geometry and the illumination source. We study this linear system and employ empirical techniques to find the limits of the information contained in the third bounce photons as a function of various system parameters.
Adithya Kumar Pediredla, Nathan Matsuda, Oliver Cossairt, Ashok Veeraraghavan
ICASSP3
2017 Coherent inverse scattering via transmission matrices: Efficient phase retrieval algorithms and a public dataset
abstract
A transmission matrix describes the input-output relationship of a complex wavefront as it passes through/reflects off a multiple-scattering medium, such as frosted glass or a painted wall. Knowing a medium's transmission matrix enables one to image through the medium, send signals through the medium, or even use the medium as a lens. The double phase retrieval method is a recently proposed technique to learn a medium's transmission matrix that avoids difficult-to-capture interferometric measurements. Unfortunately, to perform high resolution imaging, existing double phase retrieval methods require (1) a large number of measurements and (2) an unreasonable amount of computation. In this work we focus on the latter of these two problems and reduce computation times with two distinct methods: First, we develop a new phase retrieval algorithm that is significantly faster than existing methods, especially when used with an amplitude-only spatial light modulator (SLM). Second, we calibrate the system using a phase-only SLM, rather than an amplitude-only SLM which was used in previous double phase retrieval experiments. This seemingly trivial change enables us to use a far faster class of phase retrieval algorithms. As a result of these advances, we achieve a 100x reduction in computation times, thereby allowing us to image through scattering media at state-of-the-art resolutions. In addition to these advances, we also release the first publicly available transmission matrix dataset. This contribution will enable phase retrieval researchers to apply their algorithms to real data. Of particular interest to this community, our measurement vectors are naturally i.i.d. subgaussian, i.e., no coded diffraction pattern is required.
Christopher A. Metzler, Sudarshan Nagesh, Richard G. Baraniuk, Oliver Cossairt, Ashok Veeraraghavan
ICCP5
2017 Reconstructing rooms using photon echoes: A plane based model and reconstruction algorithm for looking around the corner
abstract
Can we reconstruct the entire internal shape of a room if all we can directly observe is a small portion of one internal wall, presumably through a window in the room? While conventional wisdom may indicate that this is not possible, motivated by recent work on `looking around corners', we show that one can exploit light echoes to reconstruct the internal shape of hidden rooms. Existing techniques for looking around the corner using transient images model the hidden volume using voxels and try to explain the captured transient response as the sum of the transient responses obtained from individual voxels. Such a technique inherently suffers from challenges with regards to low signal to background ratios (SBR) and has difficulty scaling to larger volumes. In contrast, in this paper, we argue for using a plane-based model for the hidden surfaces. We demonstrate that such a plane-based model results in much higher SBR while simultaneously being amenable to larger spatial scales. We build an experimental prototype composed of a pulsed laser source and a single-photon avalanche detector (SPAD) that can achieve a time resolution of about 30ps and demonstrate high-fidelity reconstructions both of individual planes in a hidden volume and for reconstructing entire polygonal rooms composed of multiple planar walls.
Adithya Kumar Pediredla, Mauro Buttafava, Alberto Tosi, Oliver Cossairt, Ashok Veeraraghavan
ICCP4
2017 Ptychnet: CNN based fourier ptychography
abstract
Fourier ptychography is an imaging technique that overcomes the diffraction limit of conventional cameras with applications in microscopy and long range imaging. Diffraction blur causes resolution loss in both cases. In Fourier ptychography, a coherent light source illuminates an object, which is then imaged from multiple viewpoints. The reconstruction of the object from these set of recordings can be obtained by an iterative phase retrieval algorithm. However, the retrieval process is slow and does not work well under certain conditions. In this paper, we propose a new reconstruction algorithm that is based on convolutional neural networks and demonstrate its advantages in terms of speed and performance.
Armin Kappeler, Sushobhan Ghosh, Jason Holloway, Oliver Cossairt, Aggelos K. Katsaggelos
ICIP4
2016 Compressive reconstruction for 3D incoherent holographic microscopy
abstract
Incoherent holography has recently attracted significant research interest due to its flexibility for a wide variety of light sources. In this paper, we use compressive sensing to reconstruct a three-dimensional volumetric object from its two-dimensional Fresnel incoherent correlation hologram. We show how compressed sensing enables reconstruction without out-of-focus artifacts, when compared to conventional back-propagation recovery. Finally, we analyze the reconstruction guarantees of the proposed approach both numerically and theoretically and compare that with coherent holography.
Oliver Cossairt, Kuan He, Ruibo Shang, Nathan Matsuda, Xiang Huang 0006, Aggelos K. Katsaggelos, Leonidas Spinoulas, Seunghwan Yoo
ICIP1
2015 MC3D: Motion Contrast 3D Scanning
abstract
Structured light 3D scanning systems are fundamentally constrained by limited sensor bandwidth and light source power, hindering their performance in real-world applications where depth information is essential, such as industrial automation, autonomous transportation, robotic surgery, and entertainment. We present a novel structured light technique called Motion Contrast 3D scanning (MC3D) that maximizes bandwidth and light source power to avoid performance trade-offs. The technique utilizes motion contrast cameras that sense temporal gradients asynchronously, i.e., independently for each pixel, a property that minimizes redundant sampling. This allows laser scanning resolution with single-shot speed, even in the presence of strong ambient illumination, significant inter-reflections, and highly reflective surfaces. The proposed approach will allow 3D vision systems to be deployed in challenging and hitherto inaccessible real-world scenarios requiring high performance using limited power and bandwidth.
Nathan Matsuda, Oliver Cossairt, Mohit Gupta 0001
ICCP2
2015 Sampling optimization for on-chip compressive video
abstract
In this paper, we consider the problem of on-chip temporal compressive sensing for video reconstruction at high frame-rates without the need of any additional optical components. We devise an optimization scheme in order to achieve adequate spatio-temporal sampling of subsequent frames under maximal capturing speed, based on the bandwidth constraints of a sensor. We test this optimization strategy on a commercially available camera and propose a set of reconstruction steps that can achieve reasonable performance but, at the same time, accommodate high-resolution video reconstruction under realistic time requirements. Our analysis constitutes a set of first steps bringing high-speed compressive video capture within the realm of commercial availability.
Leonidas Spinoulas, Oliver Cossairt, Aggelos K. Katsaggelos
ICIP2
2014 Digital refocusing with incoherent holography
abstract
Light field cameras allow us to digitally refocus a photograph after the time of capture. However, recording a light field requires either a significant loss in spatial resolution [11, 21, 10] or a large number of images to be captured [12]. In this paper, we propose incoherent holography for digital refocusing without loss of spatial resolution from only 3 captured images. The main idea is to capture 2D coherent holograms of the scene instead of the 4D light fields. The key properties of coherent light propagation are that the coherent spread function (hologram of a single point source) encodes scene depths and has a broadband spatial frequency response. These properties enable digital refocusing with 2D coherent holograms, which can be captured on sensors without loss of spatial resolution. Incoherent holography does not require illuminating the scene with high power coherent laser, making it possible to acquire holograms even for passively illuminated scenes. We provide an in-depth performance comparison between light field and incoherent holographic cameras in terms of the signal-to-noise-ratio (SNR). We show that given the same sensing resources, an incoherent holography camera outperforms light field cameras in most real world settings. We demonstrate a prototype incoherent holography camera capable of performing digital refocusing from only 3 acquired images. We show results on a variety of scenes that verify the accuracy of our theoretical analysis.
Oliver Cossairt, Nathan Matsuda, Mohit Gupta 0001
ICCP1
2014 Can we beat Hadamard multiplexing? Data driven design and analysis for computational imaging systems
abstract
Computational Imaging (CI) systems that exploit optical multiplexing and algorithmic demultiplexing have been shown to improve imaging performance in tasks such as motion deblurring, extended depth of field, light field and hyper-spectral imaging. Design and performance analysis of many of these approaches tend to ignore the role of image priors. It is well known that utilizing statistical image priors significantly improves demultiplexing performance. In this paper, we extend the Gaussian Mixture Model as a data-driven image prior (proposed by Mitra et. al [21]) to under-determined linear systems and study compressive CI methods such as light-field and hyper-spectral imaging. Further, we derive a novel algorithm for optimizing multiplexing matrices that simultaneously accounts for (a) sensor noise (b) image priors and (c) CI design constraints. We use our algorithm to design data-optimal multiplexing matrices for a variety of existing CI designs, and we use these matrices to analyze the performance of CI systems as a function of noise level. Our analysis gives new insight into the optimal performance of CI systems, and how this relates to the performance of classical multiplexing designs such as Hadamard matrices.
Kaushik Mitra, Oliver Cossairt, Ashok Veeraraghavan
ICCP2
2014 A Framework for Analysis of Computational Imaging Systems: Role of Signal Prior, Sensor Noise and Multiplexing
abstract
Over the last decade, a number of computational imaging (CI) systems have been proposed for tasks such as motion deblurring, defocus deblurring and multispectral imaging. These techniques increase the amount of light reaching the sensor via multiplexing and then undo the deleterious effects of multiplexing by appropriate reconstruction algorithms. Given the widespread appeal and the considerable enthusiasm generated by these techniques, a detailed performance analysis of the benefits conferred by this approach is important. Unfortunately, a detailed analysis of CI has proven to be a challenging problem because performance depends equally on three components: (1) the optical multiplexing, (2) the noise characteristics of the sensor, and (3) the reconstruction algorithm which typically uses signal priors. A few recent papers [12], [30], [49] have performed analysis taking multiplexing and noise characteristics into account. However, analysis of CI systems under state-of-the-art reconstruction algorithms, most of which exploit signal prior models, has proven to be unwieldy. In this paper, we present a comprehensive analysis framework incorporating all three components. In order to perform this analysis, we model the signal priors using a Gaussian Mixture Model (GMM). A GMM prior confers two unique characteristics. First, GMM satisfies the universal approximation property which says that any prior density function can be approximated to any fidelity using a GMM with appropriate number of mixtures. Second, a GMM prior lends itself to analytical tractability allowing us to derive simple expressions for the `minimum mean square error' (MMSE) which we use as a metric to characterize the performance of CI systems. We use our framework to analyze several previously proposed CI techniques (focal sweep, flutter shutter, parabolic exposure, etc.), giving conclusive answer to the question: `How much performance gain is due to use of a signal prior and how much is due to multiplexing? Our analysis also clearly shows that multiplexing provides significant performance gains above and beyond the gains obtained due to use of signal priors.
Kaushik Mitra, Oliver Cossairt, Ashok Veeraraghavan
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Focal sweep videography with deformable optics
abstract
A number of cameras have been introduced that sweep the focal plane using mechanical motion. However, mechanical motion makes video capture impractical and is unsuitable for long focal length cameras. In this paper, we present a focal sweep telephoto camera that uses a variable focus lens to sweep the focal plane. Our camera requires no mechanical motion and is capable of sweeping the focal plane periodically at high speeds. We use our prototype camera to capture EDOF videos at 20fps, and demonstrate space-time refocusing for scenes with a wide depth range. In addition, we capture periodic focal stacks, and show how they can be used for several interesting applications such as video refocusing and trajectory estimation of moving objects.
Daniel Miau, Oliver Cossairt, Shree K. Nayar
ICCP2
2013 When Does Computational Imaging Improve Performance?
abstract
A number of computational imaging techniques are introduced to improve image quality by increasing light throughput. These techniques use optical coding to measure a stronger signal level. However, the performance of these techniques is limited by the decoding step, which amplifies noise. Although it is well understood that optical coding can increase performance at low light levels, little is known about the quantitative performance advantage of computational imaging in general settings. In this paper, we derive the performance bounds for various computational imaging techniques. We then discuss the implications of these bounds for several real-world scenarios (e.g., illumination conditions, scene properties, and sensor noise characteristics). Our results show that computational imaging techniques do not provide a significant performance advantage when imaging with illumination that is brighter than typical daylight. These results can be readily used by practitioners to design the most suitable imaging systems given the application at hand.
Oliver Cossairt, Mohit Gupta 0001, Shree K. Nayar
IEEE Trans. Image Process.1
2011 Gigapixel Computational Imaging
abstract
Today, consumer cameras produce photographs with tens of millions of pixels. The recent trend in image sensor resolution seems to suggest that we will soon have cameras with billions of pixels. However, the resolution of any camera is fundamentally limited by geometric aberrations. We derive a scaling law that shows that, by using computations to correct for aberrations, we can create cameras with unprecedented resolution that have low lens complexity and compact form factor. In this paper, we present an architecture for gigapixel imaging that is compact and utilizes a simple optical design. The architecture consists of a ball lens shared by several small planar sensors, and a post-capture image processing stage. Several variants of this architecture are shown for capturing a contiguous hemispherical field of view as well as a complete spherical field of view. We demonstrate the effectiveness of our architecture by showing example images captured with two proof-of-concept gigapixel cameras.
Oliver Cossairt, Daniel Miau, Shree K. Nayar
ICCP1
2010 Depth from Diffusion
abstract
An optical diffuser is an element that scatters light and is commonly used to soften or shape illumination. In this paper, we propose a novel depth estimation method that places a diffuser in the scene prior to image capture. We call this approach depth-from-diffusion (DFDiff). We show that DFDiff is analogous to conventional depth-from-defocus (DFD), where the scatter angle of the diffuser determines the effective aperture of the system. The main benefit of DFDiff is that while DFD requires very large apertures to improve depth sensitivity, DFDiff only requires an increase in the diffusion angle-a much less expensive proposition. We perform a detailed analysis of the image formation properties of a DFDiff system, and show a variety of examples demonstrating greater precision in depth estimation when using DFDiff.
Changyin Zhou, Oliver Cossairt, Shree K. Nayar
CVPR2
2010 Diffusion coded photography for extended depth of field
abstract
In recent years, several cameras have been introduced which extend depth of field (DOF) by producing a depth-invariant point spread function (PSF). These cameras extend DOF by deblurring a captured image with a single spatially-invariant PSF. For these cameras, the quality of recovered images depends both on the magnitude of the PSF spectrum (MTF) of the camera, and the similarity between PSFs at different depths. While researchers have compared the MTFs of different extended DOF cameras, relatively little attention has been paid to evaluating their depth invariances. In this paper, we compare the depth invariance of several cameras, and introduce a new camera that improves in this regard over existing designs, while still maintaining a good MTF. Our technique utilizes a novel optical element placed in the pupil plane of an imaging system. Whereas previous approaches use optical elements characterized by their amplitude or phase profile, our approach utilizes one whose behavior is characterized by its scattering properties. Such an element is commonly referred to as an optical diffuser, and thus we refer to our new approach as diffusion coding . We show that diffusion coding can be analyzed in a simple and intuitive way by modeling the effect of a diffuser as a kernel in light field space. We provide detailed analysis of diffusion coded cameras and show results from an implementation using a custom designed diffuser.
Oliver Cossairt, Changyin Zhou, Shree K. Nayar
ACM Trans. Graph.1
2008 Light field transfer: global illumination between real and synthetic objects
abstract
We present a novel image-based method for compositing real and synthetic objects in the same scene with a high degree of visual realism. Ours is the first technique to allow global illumination and near-field lighting effects between both real and synthetic objects at interactive rates, without needing a geometric and material model of the real scene. We achieve this by using a light field interface between real and synthetic components---thus, indirect illumination can be simulated using only two 4D light fields, one captured from and one projected onto the real scene. Multiple bounces of interreflections are obtained simply by iterating this approach. The interactivity of our technique enables its use with time-varying scenes, including dynamic objects. This is in sharp contrast to the alternative approach of using 6D or 8D light transport functions of real objects, which are very expensive in terms of acquisition and storage and hence not suitable for real-time applications. In our method, 4D radiance fields are simultaneously captured and projected by using a lens array, video camera, and digital projector. The method supports full global illumination with restricted object placement, and accommodates moderately specular materials. We implement a complete system and show several example scene compositions that demonstrate global illumination effects between dynamic real and synthetic objects. Our implementation requires a single point light source and dark background.
Oliver Cossairt, Shree K. Nayar, Ravi Ramamoorthi
ACM Trans. Graph.1