Shree K. Nayar

dblp:n/ShreeKNayar · DBLP profile ↗
← Back
220ranked-venue papers
42as first author
12since 2021 · last 2025
0000-0002-6452-6998ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 161 · 31 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 151 · 23 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 11 · 3 since 2021Systems, architecture and hardware · 7 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 3 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Hierarchical Material Recognition from Local Appearance
Matthew Beveridge, Shree K. Nayar
ICCV2
2025 Privacy-Enabled Parallax Display
abstract
Privacy filters for displays are designed to obfuscate or hide visual content from unintended observers, while making the displayed information visible only to selected viewers. Existing privacy filters suffer from either wide viewing field (low selectivity) or limited user positioning. To solve this dilemma, we propose a display technology that allows a narrow but adaptive viewing field that can be directed to arbitrary user location. While conventional parallax barriers provide such capability of modulating the light field according to the user location, it suffers from repeated views. Our key observation is that this view repetition originates from the periodicity of barrier patterns, and we propose a privacy-enabled parallax display based on randomized barrier design. In addition to randomizing the locations of 1D slits, we also propose breaking down the slits into pinholes and randomizing their 2D locations, which results in privacy-preservation along the vertical direction as well. We build a hardware prototype using two off-the-shelf liquid-crystal displays. Experiments show that the proposed randomized parallax barrier can direct to the user a narrow viewing field of about ±6°, providing a significantly improved privacy protection as compared to traditional privacy screens.
Sizhuo Ma, Karl Bayer, Gurunandan Krishnan, Mohit Gupta 0001, Shree K. Nayar
VR5
2024 Minimalist Vision with Freeform Pixels
Jeremy Klotz, Shree K. Nayar
ECCV (64)2
2024 Light Codes for Fast Two-Way Human-Centric Visual Communication
abstract
Visual codes, such as QR codes, are widely used in several applications for conveying information to users. However, user interactions based on spatial codes (e.g., displaying codes on phone screens for exchanging contact information) are often tedious, time consuming, and prone to errors due to image corruptions such as noise, blur, saturation, and perspective distortions. We propose Light Codes (LICO), a novel method for fast and fluid exchange of information among users. Light codes are based on transmitting and receiving temporal codes (instead of spatial) using compact and low-cost transceiver devices. The resulting approach enables seamless and near instantaneous exchange of short messages among users with minimal physical and cognitive effort. We design novel coding techniques, hardware prototypes, and applications that are optimized for human-centric communication, and facilitate fast and fluid user-to-user interactions in various challenging conditions, including a range of distances, motion, and ambient illumination. We evaluate the performance of the proposed methods both via quantitative analysis and user study based comparisons with several existing approaches including display-camera links, Bluetooth, and near-field communication, which show strong preference toward Light Codes in various real-world application scenarios.
Mohit Gupta 0001, Jian Wang 0100, Karl Bayer, Shree K. Nayar
ACM Trans. Graph.4
2024 Cricket: A Self-Powered Chirping Pixel
abstract
We present a sensor that can measure light and wirelessly communicate the measurement, without the need for an external power source or a battery. Our sensor, called cricket, harvests energy from incident light. It is asleep for most of the time and transmits a short and strong radio frequency chirp when its harvested energy reaches a specific level. The carrier frequency of each cricket is fixed and reveals its identity, and the duration between consecutive chirps is a measure of the incident light level. We have characterized the radiometric response function, signal-to-noise ratio and dynamic range of cricket. We have experimentally verified that cricket can be miniaturized at the expense of increasing the duration between chirps. We show that a cube with a cricket on each of its sides can be used to estimate the centroid of any complex illumination, which has value in applications such as solar tracking. We also demonstrate the use of crickets for creating untethered sensor arrays that can produce video and control lighting for energy conservation. Finally, we modified cricket's circuit to develop battery-free electronic sunglasses that can instantly adapt to environmental illumination.
Shree K. Nayar, Jeremy Klotz, Nikhil Nanda, Mikhail Fridberg
ACM Trans. Graph.1
2024 Perspective-Aligned AR Mirror with Under-Display Camera
abstract
Augmented reality (AR) mirrors are novel displays that have great potential for commercial applications such as virtual apparel try-on. Typically the camera is placed beside the display, leading to distorted perspectives during user interaction. In this paper, we present a novel approach to address this problem by placing the camera behind a transparent display, thereby providing users with a perspective-aligned experience. Simply placing the camera behind the display can compromise image quality due to optical effects. We meticulously analyze the image formation process, and present an image restoration algorithm that benefits from physics-based data synthesis and network design. Our method significantly improves image quality and outperforms existing methods especially on the underexplored wire and backscatter artifacts. We then carefully design a full AR mirror system including display and camera selection, real-time processing pipeline, and mechanical design. Our user study demonstrates that the system is exceptionally well-received by users, highlighting its advantages over existing camera configurations not only as an AR mirror, but also for video conferencing. Our work represents a step forward in the development of AR mirrors, with potential applications in retail, cosmetics, fashion, etc. The image restoration dataset and code are available at https://perspective-armirror.github.io/.
Jian Wang 0100, Sizhuo Ma, Karl Bayer, Yi Zhang 0108, Peihao Wang, Bing Zhou 0001, Shree K. Nayar, Gurunandan Krishnan
ACM Trans. Graph.7
2023 AO-Finger: Hands-free Fine-grained Finger Gesture Recognition via Acoustic-Optic Sensor Fusing
abstract
Finger gesture recognition is gaining great research interest for wearable device interactions such as smartwatches and AR/VR headsets. In this paper, we propose a hands-free fine-grained finger gesture recognition system AO-Finger based on acoustic-optic sensor fusing. Specifically, we design a wristband with a modified stethoscope microphone and two high-speed optic motion sensors to capture signals generated from finger movements. We propose a set of natural, inconspicuous and effortless micro finger gestures that can be reliably detected from the complementary signals from both sensors. We design a multi-modal CNN-Transformer model for fast gesture recognition (flick/pinch/tap), and a finger swipe contact detection model to enable fine-grained swipe gesture tracking. We built a prototype which achieves an overall accuracy of 94.83% in detecting fast gestures and enables fine-grained continuous swipe gestures tracking. AO-Finger is practical for use as a wearable device and ready to be integrated into existing wrist-worn devices such as smartwatches.
Chenhan Xu, Bing Zhou 0001, Gurunandan Krishnan, Shree K. Nayar
CHI4
2023 Personalized Dereverberation of Speech
Ruilin Xu 0001, Gurunandan Krishnan, Changxi Zheng, Shree K. Nayar
INTERSPEECH4
2023 QfaR: Location-Guided Scanning of Visual Codes from Long Distances
abstract
Visual codes such as QR codes provide a low-cost and convenient communication channel between physical objects and mobile devices, but typically operate when the code and the device are in close physical proximity. We propose a system, called QfaR, which enables mobile devices to scan visual codes across long distances even where the image resolution of the visual codes is extremely low. QfaR is based on location-guided code scanning, where we utilize a crowd-sourced database of physical locations of codes. Our key observation is that if the approximate location of the codes and the user is known, the space of possible codes can be dramatically pruned down. Then, even if every "single bit" from the low-resolution code cannot be recovered, QfaR can still identify the visual code from the pruned list with high probability. By applying computer vision techniques, QfaR is also robust against challenging imaging conditions, such as tilt, motion blur, etc. Experimental results with common iOS and Android devices show that QfaR can significantly enhance distances at which codes can be scanned, e.g., 3.6cm-sized codes can be scanned at a distance of 7.5 meters, and 0.5m-sized codes at about 100 meters. QfaR has many potential applications, and beyond our diverse experiments, we also conduct a simple case study on its use for efficiently scanning QR code-based badges to estimate event attendance.
Sizhuo Ma, Jian Wang 0100, Wenzheng Chen, Suman Banerjee 0001, Mohit Gupta 0001, Shree K. Nayar
MobiCom6
2022 Seeing Far in the Dark with Patterned Flash
Zhanghao Sun, Jian Wang 0100, Shree K. Nayar
ECCV (6)4
2021 BackTrack: 2D Back-of-device Interaction Through Front Touchscreen
abstract
We present BackTrack, a trackpad placed on the back of a smartphone to track fine-grained finger motions. Our system has a small form factor, with all the circuits encapsulated in a thin layer attached to a phone case. It can be used with any off-the-shelf smartphone, requiring no power supply or modification of the operating systems. BackTrack simply extends the finger tracking area of the front screen, without interrupting the use of the front screen. It also provides a switch to prevent unintentional touch on the trackpad. All these features are enabled by a battery-free capacitive circuit, part of which is a transparent, thin-film conductor coated on a thin glass and attached to the front screen. To ensure accurate and robust tracking, the capacitive circuits are carefully designed. Our design is based on a circuit model of capacitive touchscreens, justified through both physics-based finite-element simulation and controlled laboratory experiments. We conduct user studies to evaluate the performance of using BackTrack. We also demonstrate its use in a number of smartphone applications.
Chang Xiao 0003, Karl Bayer, Changxi Zheng, Shree K. Nayar
CHI4
2021 Seeing in Extra Darkness Using a Deep-Red Flash
abstract
We propose a new flash technique for low-light imaging, using deep-red light as an illuminating source. Our main observation is that in a dim environment, the human eye mainly uses rods for the perception of light, which are not sensitive to wavelengths longer than 620nm, yet the camera sensor still has a spectral response. We propose a novel modulation strategy when training a modern CNN model for guided image filtering, fusing a noisy RGB frame and a flash frame. This fusion network is further extended for video reconstruction. We have built a prototype with minor hardware adjustments and tested the new flash technique on a variety of static and dynamic scenes. The experimental results demonstrate that our method produces compelling reconstructions, even in extra dim conditions.
Jinhui Xiong, Jian Wang 0100, Wolfgang Heidrich, Shree K. Nayar
CVPR4
2019 Micro-Baseline Structured Light
abstract
We propose Micro-baseline Structured Light (MSL), a novel 3D imaging approach designed for small form-factor devices such as cell-phones and miniature robots. MSL operates with small projector-camera baseline and low-cost projection hardware, and can recover scene depths with computationally lightweight algorithms. The main observation is that a small baseline leads to small disparities, enabling a first-order approximation of the non-linear SL image formation model. This leads to the key theoretical result of the paper: the MSL equation, a linearized version of SL image formation. MSL equation is under-constrained due to two unknowns (depth and albedo) at each pixel, but can be efficiently solved using a local least squares approach. We analyze the performance of MSL in terms of various system parameters such as projected pattern and baseline, and provide guidelines for optimizing performance. Armed with these insights, we build a prototype to experimentally examine the theory and its practicality.
Vishwanath Saragadam, Raja Venkata, Jian Wang 0100, Shree K. Nayar, Mohit Gupta 0001
ICCV4
2019 Audiovisual Zooming: What You See Is What You Hear
abstract
When capturing videos on a mobile platform, often the target of interest is contaminated by the surrounding environment. To alleviate the visual irrelevance, camera panning and zooming provide the means to isolate a desired field of view (FOV). However, the captured audio is still contaminated by signals outside the FOV. This effect is unnatural---for human perception, visual and auditory cues must go hand-in-hand. We present the concept ofAudiovisual Zooming, whereby an auditory FOV is formed to match the visual. Our framework is built around the classic idea of beamforming, a computational approach to enhancing sound from a single direction using a microphone array. Yet, beamforming on its own can not incorporate the auditory FOV, as the FOV may include an arbitrary number of directional sources. We formulate our audiovisual zooming as a generalized eigenvalue problem and propose an algorithm for efficient computation on mobile platforms. To inform the algorithmic and physical implementation, we offer a theoretical analysis of our algorithmic components as well as numerical studies for understanding various design choices of microphone arrays. Finally, we demonstrate audiovisual zooming on two different mobile platforms: a mobile smartphone and a 360$^\circ $ spherical imaging system for video conference settings.
Arun Asokan Nair, Austin Reiter, Changxi Zheng, Shree K. Nayar
ACM Multimedia4
2019 Vidgets: modular mechanical widgets for mobile devices
abstract
We present Vidgets , a family of mechanical widgets, specifically push buttons and rotary knobs that augment mobile devices with tangible user interfaces. When these widgets are attached to a mobile device and a user interacts with them, the widgets' nonlinear mechanical response shifts the device slightly and quickly, and this subtle motion can be detected by the accelerometer commonly equipped on mobile devices. We propose a physics-based model to understand the nonlinear mechanical response of widgets. This understanding enables us to design tactile force profiles of these widgets so that the resulting accelerometer signals become easy to recognize. We then develop a lightweight signal processing algorithm that analyzes the accelerometer signals and recognizes how the user interacts with the widgets in real time. Vidgets widgets are low-cost, compact, reconfigurable, and power efficient. They can form a diverse set of physical interfaces that enrich users' interactions with mobile devices in various practical scenarios. We demonstrate their use in three applications: photo capture with single-handed zoom, control of mobile games, and making a playable mobile music instrument.
Chang Xiao 0003, Karl Bayer, Changxi Zheng, Shree K. Nayar
ACM Trans. Graph.4
2018 The Minimalist Camera
Parita Pooj, Michael D. Grossberg, Peter N. Belhumeur, Shree K. Nayar
BMVC4
2018 The RAD: Making Racing Games Equivalently Accessible to People Who Are Blind
abstract
We introduce the racing auditory display (RAD), an audio-based user interface that allows players who are blind to play the same types of racing games that sighted players can play with an efficiency and sense of control that are similar to what sighted players have. The RAD works with a standard pair of headphones and comprises two novel sonification techniques: the sound slider for understanding a car's speed and trajectory on a racetrack and the turn indicator system for alerting players of the direction, sharpness, length, and timing of upcoming turns. In a user study with 15 participants (3 blind; the rest blindfolded and analyzed separately), we found that players preferred the RAD's interface over that of Mach 1, a popular blind-accessible racing game. We also found that the RAD allows an avid gamer who is blind to race as well on a complex racetrack as casual sighted players can, without a significant difference between lap times or driving paths.
Brian A. Smith 0001, Shree K. Nayar
CHI2
2018 Trapping Light for Time of Flight
abstract
We propose a novel imaging method for near-complete, surround, 3D reconstruction of geometrically complex objects, in a single scan. The key idea is to augment a time-of-flight (ToF) based 3D sensor with a multi-mirror system, called a light-trap. The shape of the trap is chosen so that light rays entering it bounce multiple times inside the trap, thereby visiting every position inside the trap multiple times from various directions. We show via simulations that this enables light rays to reach more than 99.9% of the surface of objects placed inside the trap, even those with strong occlusions, for example, lattice-shaped objects. The ToF sensor provides the path length for each light ray, which, along with the known shape of the trap, is used to reconstruct the complete paths of all the rays. This enables performing dense, surround 3D reconstructions of objects with highly complex 3D shapes, in a single scan. We have developed a proof-of-concept hardware prototype consisting of a pulsed ToF sensor, and a light trap built with planar mirrors. We demonstrate the effectiveness of the light trap based 3D reconstruction method on a variety of objects with a broad range of geometry and reflectance properties.
Ruilin Xu 0001, Mohit Gupta 0001, Shree K. Nayar
CVPR3
2018 What Are Optimal Coding Functions for Time-of-Flight Imaging?
abstract
The depth resolution achieved by a continuous wave time-of-flight (C-ToF) imaging system is determined by the coding (modulation and demodulation) functions that it uses. Almost all current C-ToF systems use sinusoid or square coding functions, resulting in a limited depth resolution. In this article, we present a mathematical framework for exploring and characterizing the space of C-ToF coding functions in a geometrically intuitive space. Using this framework, we design families of novel coding functions that are based on Hamiltonian cycles on hypercube graphs. Given a fixed total source power and acquisition time, the new Hamiltonian coding scheme can achieve up to an order of magnitude higher resolution as compared to the current state-of-the-art methods, especially in low signal-to-noise ratio (SNR) settings. We also develop a comprehensive physically-motivated simulator for C-ToF cameras that can be used to evaluate various coding schemes prior to a real hardware implementation. Since most off-the-shelf C-ToF sensors use sinusoid or square functions, we develop a hardware prototype that can implement a wide range of coding functions. Using this prototype and our software simulator, we demonstrate the performance advantages of the proposed Hamiltonian coding functions in a wide range of imaging settings.
Mohit Gupta 0001, Andreas Velten, Shree K. Nayar, Eric Breitbach
ACM Trans. Graph.3
2017 AirCode: Unobtrusive Physical Tags for Digital Fabrication
abstract
We present AirCode, a technique that allows the user to tag physically fabricated objects with given information. An AirCode tag consists of a group of carefully designed air pockets placed beneath the object surface. These air pockets are easily produced during the fabrication process of the object, without any additional material or postprocessing. Meanwhile, the air pockets affect only the scattering light transport under the surface, and thus are hard to notice to our naked eyes. But, by using a computational imaging method, the tags become detectable. We present a tool that automates the design of air pockets for the user to encode information. AirCode system also allows the user to retrieve the information from captured images via a robust decoding algorithm. We demonstrate our tagging technique with applications for metadata embedding, robotic grasping, as well as conveying object affordances.
Dingzeyu Li, Avinash S. Nair, Shree K. Nayar, Changxi Zheng
UIST3
2016 Towards flexible sheet cameras: Deformable lens arrays with intrinsic optical adaptation
abstract
We propose a framework for developing a new class of imaging systems that are thin and flexible. Such an imaging sheet can be flexed at will and wrapped around everyday objects to capture unconventional fields of view. Our approach is to use a lens array attached to a sheet with a 2D grid of pixels. A major challenge with this type of a system is that its sampling of the scene varies with the curvature of the sheet. To avoid undesirable aliasing effects due to under-sampling in high curvature regions of the sheet, we design a deformable lens array with adaptive optical properties. We show that the material and geometric properties of the lens array can be optimized so that the object-side point spread function corresponding to each pixel widens with the curvature of the sheet at that pixel. This intrinsic adaptation of focal length is passive (without the use of actuators or other control mechanisms), and enables a sheet camera to capture images without aliasing, irrespective of its shape. We have designed a 33×33 lens array, fabricated it using silicone rubber, and conducted several experiments to verify its optical adaptation characteristics. We conclude with a discussion on the advantages of our proposed approach as well as future work.
Daniel C. Sims, Yonghao Yue, Shree K. Nayar
ICCP3
2016 Mining Controller Inputs to Understand Gameplay
abstract
Today's game analytics systems are powered by event logs, which reveal information about what players are doing but offer little insight about the types of gameplay that games foster. Moreover, the concept of gameplay itself is difficult to define and quantify. In this paper, we show that analyzing players' controller inputs using probabilistic topic models allows game developers to describe the types of gameplay -- or action -- in games in a quantitative way. More specifically, developers can discover the types of action that a game fosters and the extent that each game level fosters each type of action, all in an unsupervised manner. They can use this information to verify that their levels feature the appropriate style of gameplay and to recommend levels with gameplay that is similar to levels that players like. We begin with latent Dirichlet allocation (LDA), the simplest topic model, then develop the player-gameplay action (PGA) model to make the same types of discoveries about gameplay in a way that is independent of each player's play style. We train a player recognition system on the PGA model's output to verify that its discoveries about gameplay are in fact independent of each player's play style. The system recognizes players with over 90% accuracy in about 20 seconds of playtime.
Brian A. Smith 0001, Shree K. Nayar
UIST2
2016 DisCo: Display-Camera Communication Using Rolling Shutter Sensors
abstract
We present DisCo, a novel display-camera communication system. DisCo enables displays and cameras to communicate with each other while also displaying and capturing images for human consumption. Messages are transmitted by temporally modulating the display brightness at high frequencies so that they are imperceptible to humans. Messages are received by a rolling shutter camera that converts the temporally modulated incident light into a spatial flicker pattern. In the captured image, the flicker pattern is superimposed on the pattern shown on the display. The flicker and the display pattern are separated by capturing two images with different exposures. The proposed system performs robustly in challenging real-world situations such as occlusion, variable display size, defocus blur, perspective distortion, and camera rotation. Unlike several existing visible light communication methods, DisCo works with off-the-shelf image sensors. It is compatible with a variety of sources (including displays, single LEDs), as well as reflective surfaces illuminated with light sources. We have built hardware prototypes that demonstrate DisCo’s performance in several scenarios. Because of its robustness, speed, ease of use, and generality, DisCo can be widely deployed in several applications, such as advertising, pairing of displays with cell phones, tagging objects in stores and museums, and indoor navigation.
Kensei Jo, Mohit Gupta 0001, Shree K. Nayar
ACM Trans. Graph.3
2015 Towards Self-Powered Cameras
abstract
We propose a simple pixel design, where the pixel's photodiode can be used to not only measure the incident light level, but also to convert the incident light into electrical energy. A sensor architecture is proposed where, during each image capture cycle, the pixels are used first to record and read out the image and then used to harvest energy and charge the sensors' power supply. We have conducted several experiments using off-the-shelf discrete components to validate the practical feasibility of our approach. We first developed a single pixel based on our design and used it to physically scan images of scenes. Next, we developed a fully self-powered camera that produces 30×40 images. The camera uses a supercap rather than an external source as its power supply. For a scene that is around 300 lux in brightness, the voltage across the supercap remains well above the minimum needed for the camera to indefinitely produce an image per second. For scenarios where scene brightness may vary dramatically, we present an adaptive algorithm that adjusts the framerate of the camera based on the voltage of the supercap and the brightness of the scene. Finally, we analyze the light gathering and harvesting properties of our design and explain why we believe it could lead to a fully self-powered solid-state image sensor that produces a useful resolution and framerate.
Shree K. Nayar, Daniel C. Sims, Mikhail Fridberg
ICCP1
2015 SpeDo: 6 DOF Ego-Motion Sensor Using Speckle Defocus Imaging
abstract
Sensors that measure their motion with respect to the surrounding environment (ego-motion sensors) can be broadly classified into two categories. First is inertial sensors such as accelerometers. In order to estimate position and velocity, these sensors integrate the measured acceleration, which often results in accumulation of large errors over time. Second, camera-based approaches such as SLAM that can measure position directly, but their performance depends on the surrounding scene's properties. These approaches cannot function reliably if the scene has low frequency textures or small depth variations. We present a novel ego-motion sensor called SpeDo that addresses these fundamental limitations. SpeDo is based on using coherent light sources and cameras with large defocus. Coherent light, on interacting with a scene, creates a high frequency interferometric pattern in the captured images, called speckle. We develop a theoretical model for speckle flow (motion of speckle as a function of sensor motion), and show that it is quasi-invariant to surrounding scene's properties. As a result, SpeDo can measure ego-motion (not derivative of motion) simply by estimating optical flow at a few image locations. We have built a low-cost and compact hardware prototype of SpeDo and demonstrated high precision 6 DOF ego-motion estimation for complex trajectories in scenarios where the scene properties are challenging (e.g., repeating or no texture) as well as unknown.
Kensei Jo, Mohit Gupta 0001, Shree K. Nayar
ICCV3
2015 Extended Depth of Field Catadioptric Imaging Using Focal Sweep
abstract
Catadioptric imaging systems use curved mirrors to capture wide fields of view. However, due to the curvature of the mirror, these systems tend to have very limited depth of field (DOF), with the point spread function (PSF) varying dramatically over the field of view and as a function of scene depth. In recent years, focal sweep has been used extensively to extend the DOF of conventional imaging systems. It has been shown that focal sweep produces an integrated point spread function (IPSF) that is nearly space-invariant and depth-invariant, enabling the recovery of an extended depth of field (EDOF) image by deconvolving the captured focal sweep image with a single IPSF. In this paper, we use focal sweep to extend the DOF of a catadioptric imaging system. We show that while the IPSF is spatially varying when a curved mirror is used, it remains quasi depth-invariant over the wide field of view of the imaging system. We have developed a focal sweep system where mirrors of different shapes can be used to capture wide field of view EDOF images. In particular, we show experimental results using spherical and paraboloidal mirrors.
Ryunosuke Yokoya, Shree K. Nayar
ICCV2
2015 Phasor Imaging: A Generalization of Correlation-Based Time-of-Flight Imaging
abstract
In correlation-based time-of-flight (C-ToF) imaging systems, light sources with temporally varying intensities illuminate the scene. Due to global illumination, the temporally varying radiance received at the sensor is a combination of light received along multiple paths. Recovering scene properties (e.g., scene depths) from the received radiance requires separating these contributions, which is challenging due to the complexity of global illumination and the additional temporal dimension of the radiance. We propose phasor imaging, a framework for performing fast inverse light transport analysis using C-ToF sensors. Phasor imaging is based on the idea that, by representing light transport quantities as phasors and light transport events as phasor transformations, light transport analysis can be simplified in the temporal frequency domain. We study the effect of temporal illumination frequencies on light transport and show that, for a broad range of scenes, global radiance (inter-reflections and volumetric scattering) vanishes for frequencies higher than a scene-dependent threshold. We use this observation for developing two novel scene recovery techniques. First, we present micro-ToF imaging, a ToF-based shape recovery technique that is robust to errors due to inter-reflections (multipath interference) and volumetric scattering. Second, we present a technique for separating the direct and global components of radiance. Both techniques require capturing as few as 3--4 images and minimal computations. We demonstrate the validity of the presented techniques via simulations and experiments performed with our hardware prototype.
Mohit Gupta 0001, Shree K. Nayar, Matthias B. Hullin, Jaime Martín
ACM Trans. Graph.2
2014 Efficient Space-Time Sampling with Pixel-Wise Coded Exposure for High-Speed Imaging
abstract
Cameras face a fundamental trade-off between spatial and temporal resolution. Digital still cameras can capture images with high spatial resolution, but most high-speed video cameras have relatively low spatial resolution. It is hard to overcome this trade-off without incurring a significant increase in hardware costs. In this paper, we propose techniques for sampling, representing, and reconstructing the space-time volume to overcome this trade-off. Our approach has two important distinctions compared to previous works: 1) We achieve sparse representation of videos by learning an overcomplete dictionary on video patches, and 2) we adhere to practical hardware constraints on sampling schemes imposed by architectures of current image sensors, which means that our sampling function can be implemented on CMOS image sensors with modified control units in the future. We evaluate components of our approach, sampling function and sparse representation, by comparing them to several existing approaches. We also implement a prototype imaging system with pixel-wise coded exposure control using a liquid crystal on silicon device. System characteristics such as field of view and modulation transfer function are evaluated for our imaging system. Both simulations and experiments on a wide range of scenes show that our method can effectively reconstruct a video from a single coded image while maintaining high spatial resolution.
Dengyu Liu, Jinwei Gu, Yasunobu Hitomi, Mohit Gupta 0001, Tomoo Mitsunaga, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.6
2013 Focal sweep videography with deformable optics
abstract
A number of cameras have been introduced that sweep the focal plane using mechanical motion. However, mechanical motion makes video capture impractical and is unsuitable for long focal length cameras. In this paper, we present a focal sweep telephoto camera that uses a variable focus lens to sweep the focal plane. Our camera requires no mechanical motion and is capable of sweeping the focal plane periodically at high speeds. We use our prototype camera to capture EDOF videos at 20fps, and demonstrate space-time refocusing for scenes with a wide depth range. In addition, we capture periodic focal stacks, and show how they can be used for several interesting applications such as video refocusing and trajectory estimation of moving objects.
Daniel Miau, Oliver Cossairt, Shree K. Nayar
ICCP3
2013 Fibonacci Exposure Bracketing for High Dynamic Range Imaging
abstract
Exposure bracketing for high dynamic range (HDR) imaging involves capturing several images of the scene at different exposures. If either the camera or the scene moves during capture, the captured images must be registered. Large exposure differences between bracketed images lead to inaccurate registration, resulting in artifacts such as ghosting (multiple copies of scene objects) and blur. We present two techniques, one for image capture (Fibonacci exposure bracketing) and one for image registration (generalized registration), to prevent such motion-related artifacts. Fibonacci bracketing involves capturing a sequence of images such that each exposure time is the sum of the previous N(N > 1) exposures. Generalized registration involves estimating motion between sums of contiguous sets of frames, instead of between individual frames. Together, the two techniques ensure that motion is always estimated between frames of the same total exposure time. This results in HDR images and videos which have both a large dynamic range and minimal motion-related artifacts. We show, by results for several real-world indoor and outdoor scenes, that the proposed approach significantly outperforms several existing bracketing schemes.
Mohit Gupta 0001, Daisuke Iso, Shree K. Nayar
ICCV3
2013 Structured Light in Sunlight
abstract
Strong ambient illumination severely degrades the performance of structured light based techniques. This is especially true in outdoor scenarios, where the structured light sources have to compete with sunlight, whose power is often 2-5 orders of magnitude larger than the projected light. In this paper, we propose the concept of light concentration to overcome strong ambient illumination. Our key observation is that given a fixed light (power) budget, it is always better to allocate it sequentially in several portions of the scene, as compared to spreading it over the entire scene at once. For a desired level of accuracy, we show that by distributing light appropriately, the proposed approach requires 1-2 orders lower acquisition time than existing approaches. Our approach is illumination-adaptive as the optimal light distribution is determined based on a measurement of the ambient illumination level. Since current light sources have a fixed light distribution, we have built a prototype light source that supports flexible light distribution by controlling the scanning speed of a laser scanner. We show several high quality 3D scanning results in a wide range of outdoor scenarios. The proposed approach will benefit 3D vision systems that need to operate outdoors under extreme ambient illumination levels on a limited time and power budget.
Mohit Gupta 0001, Qi Yin, Shree K. Nayar
ICCV3
2013 Gaze locking: passive eye contact detection for human-object interaction
abstract
Eye contact plays a crucial role in our everyday social interactions. The ability of a device to reliably detect when a person is looking at it can lead to powerful human-object interfaces. Today, most gaze-based interactive systems rely on gaze tracking technology. Unfortunately, current gaze tracking techniques require active infrared illumination, calibration, or are sensitive to distance and pose. In this work, we propose a different solution-a passive, appearance-based approach for sensing eye contact in an image. By focusing on gaze *locking* rather than gaze tracking, we exploit the special appearance of direct eye gaze, achieving a Matthews correlation coefficient (MCC) of over 0.83 at long distances (up to 18 m) and large pose variations (up to ±30° of head yaw rotation) using a very basic classifier and without calibration. To train our detector, we also created a large publicly available gaze data set: 5,880 images of 56 people over varying gaze directions and head poses. We demonstrate how our method facilitates human-object interaction, user analytics, image filtering, and gaze-triggered photography.
Brian A. Smith 0001, Qi Yin, Steven K. Feiner, Shree K. Nayar
UIST4
2013 Compressive Structured Light for Recovering Inhomogeneous Participating Media
abstract
We propose a new method named compressive structured light for recovering inhomogeneous participating media. Whereas conventional structured light methods emit coded light patterns onto the surface of an opaque object to establish correspondence for triangulation, compressive structured light projects patterns into a volume of participating medium to produce images which are integral measurements of the volume density along the line of sight. For a typical participating medium encountered in the real world, the integral nature of the acquired images enables the use of compressive sensing techniques that can recover the entire volume density from only a few measurements. This makes the acquisition process more efficient and enables reconstruction of dynamic volumetric phenomena. Moreover, our method requires the projection of multiplexed coded illumination, which has the added advantage of increasing the signal-to-noise ratio of the acquisition. Finally, we propose an iterative algorithm to correct for the attenuation of the participating medium during the reconstruction process. We show the effectiveness of our method with simulations as well as experiments on the volumetric recovery of multiple translucent layers, 3D point clouds etched in glass, and the dynamic process of milk drops dissolving in water.
Jinwei Gu, Shree K. Nayar, Eitan Grinspun, Peter N. Belhumeur, Ravi Ramamoorthi
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 When Does Computational Imaging Improve Performance?
abstract
A number of computational imaging techniques are introduced to improve image quality by increasing light throughput. These techniques use optical coding to measure a stronger signal level. However, the performance of these techniques is limited by the decoding step, which amplifies noise. Although it is well understood that optical coding can increase performance at low light levels, little is known about the quantitative performance advantage of computational imaging in general settings. In this paper, we derive the performance bounds for various computational imaging techniques. We then discuss the implications of these bounds for several real-world scenarios (e.g., illumination conditions, scene properties, and sensor noise characteristics). Our results show that computational imaging techniques do not provide a significant performance advantage when imaging with illumination that is brighter than typical daylight. These results can be readily used by practitioners to design the most suitable imaging systems given the application at hand.
Oliver Cossairt, Mohit Gupta 0001, Shree K. Nayar
IEEE Trans. Image Process.3
2013 PiCam: an ultra-thin high performance monolithic camera array
abstract
We presentPiCam(Pelican Imaging Camera-Array), an ultra-thin high performance monolithic camera array, that captures light fields and synthesizes high resolution images along with a range image (scene depth) through integrated parallax detection and superresolution. The camera is passive, supporting both stills and video, low light capable, and small enough to be included in the next generation of mobile devices including smartphones. Prior works [Rander et al. 1997; Yang et al. 2002; Zhang and Chen 2004; Tanida et al. 2001; Tanida et al. 2003; Duparré et al. 2004] in camera arrays have explored multiple facets of light field capture - from viewpoint synthesis, synthetic refocus, computing range images, high speed video, and micro-optical aspects of system miniaturization. However, none of these have addressed the modifications needed to achieve the strict form factor and image quality required to make array cameras practical for mobile devices. In our approach, we customize many aspects of the camera array including lenses, pixels, sensors, and software algorithms to achieve imaging performance and form factor comparable to existing mobile phone cameras. Our contributions to the post-processing of images from camera arrays include a cost function for parallax detection that integrates across multiple color channels, and a regularized image restoration (superresolution) process that takes into account all the system degradations and adapts to a range of practical imaging conditions. The registration uncertainty from the parallax detection process is integrated into a Maximum-a-Posteriori formulation that synthesizes an estimate of the high resolution image and scene depth. We conclude with some examples of our array capabilities such as postcapture (still) refocus, video refocus, view synthesis to demonstrate motion parallax, 3D range images, and briefly address future work.
Kartik Venkataraman, Dan Lelescu, Jacques Duparré, Andrew McMahon, Gabriel Molina, Priyam Chatterjee, Robert Mullis, Shree K. Nayar
ACM Trans. Graph.8
2012 Micro Phase Shifting
abstract
We consider the problem of shape recovery for real world scenes, where a variety of global illumination (inter-reflections, subsurface scattering, etc.) and illumination defocus effects are present. These effects introduce systematic and often significant errors in the recovered shape. We introduce a structured light technique called Micro Phase Shifting, which overcomes these problems. The key idea is to project sinusoidal patterns with frequencies limited to a narrow, high-frequency band. These patterns produce a set of images over which global illumination and defocus effects remain constant for each point in the scene. This enables high quality reconstructions of scenes which have traditionally been considered hard, using only a small number of images. We also derive theoretical lower bounds on the number of input images needed for phase shifting and show that Micro PS achieves the bound.
Mohit Gupta 0001, Shree K. Nayar
CVPR2
2012 Diffuse structured light
abstract
Today, structured light systems are widely used in applications such as robotic assembly, visual inspection, surgery, entertainment, games and digitization of cultural heritage. Current structured light methods are faced with two serious limitations. First, they are unable to cope with scene regions that produce strong highlights due to specular reflection. Second, they cannot recover useful information for regions that lie within shadows. We observe that many structured light methods use illumination patterns that have translational symmetry, i.e., two-dimensional patterns that vary only along one of the two dimensions. We show that, for this class of patterns, diffusion of the patterns along the axis of translation can mitigate the adverse effects of specularities and shadows. We show results for two applications - 3D scanning using phase shifting of sinusoidal patterns and separation of direct and global components of light transport using high-frequency binary stripes.
Shree K. Nayar, Mohit Gupta 0001
ICCP1
2011 Gigapixel Computational Imaging
abstract
Today, consumer cameras produce photographs with tens of millions of pixels. The recent trend in image sensor resolution seems to suggest that we will soon have cameras with billions of pixels. However, the resolution of any camera is fundamentally limited by geometric aberrations. We derive a scaling law that shows that, by using computations to correct for aberrations, we can create cameras with unprecedented resolution that have low lens complexity and compact form factor. In this paper, we present an architecture for gigapixel imaging that is compact and utilizes a simple optical design. The architecture consists of a ball lens shared by several small planar sensors, and a post-capture image processing stage. Several variants of this architecture are shown for capturing a contiguous hemispherical field of view as well as a complete spherical field of view. We demonstrate the effectiveness of our architecture by showing example images captured with two proof-of-concept gigapixel cameras.
Oliver Cossairt, Daniel Miau, Shree K. Nayar
ICCP3
2011 Multiplexed illumination for scene recovery in the presence of global illumination
abstract
Global illumination effects such as inter-reflections and subsurface scattering result in systematic, and often significant errors in scene recovery using active illumination. Recently, it was shown that the direct and global components could be separated efficiently for a scene illuminated with a single light source. In this paper, we study the problem of direct-global separation for multiple light sources. We derive a theoretical lower bound for the number of required images, and propose a multiplexed illumination scheme which achieves this lower bound. We analyze the signal-to-noise ratio (SNR) characteristics of the proposed illumination multiplexing method in the context of direct-global separation. We apply our method to several scene recovery techniques requiring multiple light sources, including shape from shading, structured light 3D scanning, photometric stereo, and reflectance estimation. Both simulation and experimental results show that the proposed method can accurately recover scene information with fewer images compared to sequentially separating direct-global components for each light source.
Jinwei Gu, Toshihiro Kobayashi, Mohit Gupta 0001, Shree K. Nayar
ICCV4
2011 Video from a single coded exposure photograph using a learned over-complete dictionary
abstract
Cameras face a fundamental tradeoff between the spatial and temporal resolution - digital still cameras can capture images with high spatial resolution, but most high-speed video cameras suffer from low spatial resolution. It is hard to overcome this tradeoff without incurring a significant increase in hardware costs. In this paper, we propose techniques for sampling, representing and reconstructing the space-time volume in order to overcome this tradeoff. Our approach has two important distinctions compared to previous works: (1) we achieve sparse representation of videos by learning an over-complete dictionary on video patches, and (2) we adhere to practical constraints on sampling scheme which is imposed by architectures of present image sensor devices. Consequently, our sampling scheme can be implemented on image sensors by making a straightforward modification to the control unit. To demonstrate the power of our approach, we have implemented a prototype imaging system with per-pixel coded exposure control using a liquid crystal on silicon (LCoS) device. Using both simulations and experiments on a wide range of scenes, we show that our method can effectively reconstruct a video from a single image maintaining high spatial resolution.
Yasunobu Hitomi, Jinwei Gu, Mohit Gupta 0001, Tomoo Mitsunaga, Shree K. Nayar
ICCV5
2011 Coded Aperture Pairs for Depth from Defocus and Defocus Deblurring
Changyin Zhou, Stephen Lin 0001, Shree K. Nayar
Int. J. Comput. Vis.3
2011 Describable Visual Attributes for Face Verification and Image Search
abstract
We introduce the use of describable visual attributes for face verification and image search. Describable visual attributes are labels that can be given to an image to describe its appearance. This paper focuses on images of faces and the attributes used to describe them, although the concepts also apply to other domains. Examples of face attributes include gender, age, jaw shape, nose size, etc. The advantages of an attribute-based representation for vision tasks are manifold: They can be composed to create descriptions at various levels of specificity; they are generalizable, as they can be learned once and then applied to recognize new objects or categories without any further training; and they are efficient, possibly requiring exponentially fewer attributes (and training data) than explicitly naming each category. We show how one can create and label large data sets of real-world images to train classifiers which measure the presence, absence, or degree to which an attribute is expressed in images. These classifiers can then automatically label new images. We demonstrate the current effectiveness--and explore the future potential--of using attributes for face verification and image search via human and computational experiments. Finally, we introduce two new face data sets, named FaceTracer and PubFig, with labeled attributes and identities, respectively.
Neeraj Kumar 0006, Alexander C. Berg, Peter N. Belhumeur, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.4
2011 Flexible Depth of Field Photography
abstract
The range of scene depths that appear focused in an image is known as the depth of field (DOF). Conventional cameras are limited by a fundamental trade-off between depth of field and signal-to-noise ratio (SNR). For a dark scene, the aperture of the lens must be opened up to maintain SNR, which causes the DOF to reduce. Also, today's cameras have DOFs that correspond to a single slab that is perpendicular to the optical axis. In this paper, we present an imaging system that enables one to control the DOF in new and powerful ways. Our approach is to vary the position and/or orientation of the image detector during the integration time of a single photograph. Even when the detector motion is very small (tens of microns), a large range of scene depths (several meters) is captured, both in and out of focus. Our prototype camera uses a micro-actuator to translate the detector along the optical axis during image integration. Using this device, we demonstrate four applications of flexible DOF. First, we describe extended DOF where a large depth range is captured with a very wide aperture (low noise) but with nearly depth-independent defocus blur. Deconvolving a captured image with a single blur kernel gives an image with extended DOF and high SNR. Next, we show the capture of images with discontinuous DOFs. For instance, near and far objects can be imaged with sharpness, while objects in between are severely blurred. Third, we show that our camera can capture images with tilted DOFs (Scheimpflug imaging) without tilting the image detector. Finally, we demonstrate how our camera can be used to realize nonplanar DOFs. We believe flexible DOF imaging can open a new creative dimension in photography and lead to new capabilities in scientific imaging, vision, and graphics.
Sujit Kuthirummal, Hajime Nagahara, Changyin Zhou, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.4
2011 Computational Cameras: Convergence of Optics and Processing
abstract
A computational camera uses a combination of optics and processing to produce images that cannot be captured with traditional cameras. In the last decade, computational imaging has emerged as a vibrant field of research. A wide variety of computational cameras has been demonstrated to encode more useful visual information in the captured images, as compared with conventional cameras. In this paper, we survey computational cameras from two perspectives. First, we present a taxonomy of computational camera designs according to the coding approaches, including object side coding, pupil plane coding, sensor side coding, illumination coding, camera arrays and clusters, and unconventional imaging systems. Second, we use the abstract notion of light field representation as a general tool to describe computational camera designs, where each camera can be formulated as a projection of a high-dimensional light field to a 2-D image sensor. We show how individual optical devices transform light fields and use these transforms to illustrate how different computational camera designs (collections of optical devices) capture and encode useful visual information.
Changyin Zhou, Shree K. Nayar
IEEE Trans. Image Process.2
2010 Depth from Diffusion
abstract
An optical diffuser is an element that scatters light and is commonly used to soften or shape illumination. In this paper, we propose a novel depth estimation method that places a diffuser in the scene prior to image capture. We call this approach depth-from-diffusion (DFDiff). We show that DFDiff is analogous to conventional depth-from-defocus (DFD), where the scatter angle of the diffuser determines the effective aperture of the system. The main benefit of DFDiff is that while DFD requires very large apertures to improve depth sensitivity, DFDiff only requires an increase in the diffusion angle-a much less expensive proposition. We perform a detailed analysis of the image formation properties of a DFDiff system, and show a variety of examples demonstrating greater precision in depth estimation when using DFDiff.
Changyin Zhou, Oliver Cossairt, Shree K. Nayar
CVPR3
2010 Programmable Aperture Camera Using LCoS
Hajime Nagahara, Changyin Zhou, Hiroshi Ishiguro, Shree K. Nayar
ECCV (6)5
2010 Generalized Assorted Pixel Camera: Postcapture Control of Resolution, Dynamic Range, and Spectrum
abstract
We propose the concept of a generalized assorted pixel (GAP) camera, which enables the user to capture a single image of a scene and, after the fact, control the tradeoff between spatial resolution, dynamic range and spectral detail. The GAP camera uses a complex array (or mosaic) of color filters. A major problem with using such an array is that the captured image is severely under-sampled for at least some of the filter types. This leads to reconstructed images with strong aliasing. We make four contributions in this paper: 1) we present a comprehensive optimization method to arrive at the spatial and spectral layout of the color filter array of a GAP camera. 2) We develop a novel algorithm for reconstructing the under-sampled channels of the image while minimizing aliasing artifacts. 3) We demonstrate how the user can capture a single image and then control the tradeoff of spatial resolution to generate a variety of images, including monochrome, high dynamic range (HDR) monochrome, RGB, HDR RGB, and multispectral images. 4) Finally, the performance of our GAP camera has been verified using extensive simulations that use multispectral images of real world scenes. A large database of these multispectral images has been made available at http://www1.cs.columbia.edu/CAVE/projects/gap_camera/ for use by the research community.
Fumihito Yasuma, Tomoo Mitsunaga, Daisuke Iso, Shree K. Nayar
IEEE Trans. Image Process.4
2010 Diffusion coded photography for extended depth of field
abstract
In recent years, several cameras have been introduced which extend depth of field (DOF) by producing a depth-invariant point spread function (PSF). These cameras extend DOF by deblurring a captured image with a single spatially-invariant PSF. For these cameras, the quality of recovered images depends both on the magnitude of the PSF spectrum (MTF) of the camera, and the similarity between PSFs at different depths. While researchers have compared the MTFs of different extended DOF cameras, relatively little attention has been paid to evaluating their depth invariances. In this paper, we compare the depth invariance of several cameras, and introduce a new camera that improves in this regard over existing designs, while still maintaining a good MTF. Our technique utilizes a novel optical element placed in the pupil plane of an imaging system. Whereas previous approaches use optical elements characterized by their amplitude or phase profile, our approach utilizes one whose behavior is characterized by its scattering properties. Such an element is commonly referred to as an optical diffuser, and thus we refer to our new approach as diffusion coding . We show that diffusion coding can be analyzed in a simple and intuitive way by modeling the effect of a diffuser as a kernel in light field space. We provide detailed analysis of diffusion coded cameras and show results from an implementation using a custom designed diffuser.
Oliver Cossairt, Changyin Zhou, Shree K. Nayar
ACM Trans. Graph.3
2009 Multiple view image denoising
abstract
We present a novel multi-view denoising algorithm. Our algorithm takes noisy images taken from different viewpoints as input and groups similar patches in the input images using depth estimation. We model intensity-dependent noise in low-light conditions and use the principal component analysis and tensor analysis to remove such noise. The dimensionalities for both PCA and tensor analysis are automatically computed in a way that is adaptive to the complexity of image structures in the patches. Our method is based on a probabilistic formulation that marginalizes depth maps as hidden variables and therefore does not require perfect depth estimation. We validate our algorithm on both synthetic and real images with different content. Our algorithm compares favorably against several state-of-the-art denoising algorithms.
Li Zhang 0003, Sundeep Vaddadi, Hailin Jin, Shree K. Nayar
CVPR4
2009 Attribute and simile classifiers for face verification
abstract
We present two novel methods for face verification. Our first method - “attribute” classifiers - uses binary classifiers trained to recognize the presence or absence of describable aspects of visual appearance (e.g., gender, race, and age). Our second method - “simile” classifiers - removes the manual labeling required for attribute classification and instead learns the similarity of faces, or regions of faces, to specific reference people. Neither method requires costly, often brittle, alignment between image pairs; yet, both methods produce compact visual descriptions, and work on real-world images. Furthermore, both the attribute and simile classifiers improve on the current state-of-the-art for the LFW data set, reducing the error rates compared to the current best by 23.92% and 26.34%, respectively, and 31.68% when combined. For further testing across pose, illumination, and expression, we introduce a new data set - termed PubFig - of real-world images of public figures (celebrities and politicians) acquired from the internet. This data set is both larger (60,000 images) and deeper (300 images per individual) than existing data sets of its kind. Finally, we present an evaluation of human performance.
Neeraj Kumar 0006, Alexander C. Berg, Peter N. Belhumeur, Shree K. Nayar
ICCV4
2009 Coded aperture pairs for depth from defocus
abstract
The classical approach to depth from defocus uses two images taken with circular apertures of different sizes. We show in this paper that the use of a circular aperture severely restricts the accuracy of depth from defocus. We derive a criterion for evaluating a pair of apertures with respect to the precision of depth recovery. This criterion is optimized using a genetic algorithm and gradient descent search to arrive at a pair of high resolution apertures. The two coded apertures are found to complement each other in the scene frequencies they preserve. This property enables them to not only recover depth with greater fidelity but also obtain a high quality all-focused image from the two captured images. Extensive simulations as well as experiments on a variety of scenes demonstrate the benefits of using the coded apertures over conventional circular apertures.
Changyin Zhou, Stephen Lin 0001, Shree K. Nayar
ICCV3
2009 An empirical BSSRDF model
abstract
We present a new model of the homogeneous BSSRDF based on large-scale simulations. Our model captures the appearance of materials that are not accurately represented using existing single scattering models or multiple isotropic scattering models (e.g. the diffusion approximation). We use an analytic function to model the 2D hemispherical distribution of exitant light at a point on the surface, and a table of parameter values of this function computed at uniformly sampled locations over the remaining dimensions of the BSSRDF domain. This analytic function is expressed in elliptic coordinates and has six parameters which vary smoothly with surface position, incident angle, and the underlying optical properties of the material (albedo, mean free path length, phase function and the relative index of refraction). Our model agrees well with measured data, and is compact, requiring only 250MB to represent the full spatial- and angular-distribution of light across a wide spectrum of materials. In practice, rendering a single material requires only about 100KB to represent the BSSRDF.
Craig Donner, Jason Lawrence, Ravi Ramamoorthi, Toshiya Hachisuka, Henrik Wann Jensen, Shree K. Nayar
ACM Trans. Graph.6
2009 Removing image artifacts due to dirty camera lenses and thin occluders
abstract
Dirt on camera lenses, and occlusions from thin objects such as fences, are two important types of artifacts in digital imaging systems. These artifacts are not only an annoyance for photographers, but also a hindrance to computer vision and digital forensics. In this paper, we show that both effects can be described by a single image formation model, wherein an intermediate layer (of dust, dirt or thin occluders) both attenuates the incoming light and scatters stray light towards the camera. Because of camera defocus, these artifacts are low-frequency and either additive or multiplicative, which gives us the power to recover the original scene radiance pointwise. We develop a number of physics-based methods to remove these effects from digital photographs and videos. For dirty camera lenses, we propose two methods to estimate the attenuation and the scattering of the lens dirt and remove the artifacts -- either by taking several pictures of a structured calibration pattern beforehand, or by leveraging natural image statistics for post-processing existing images. For artifacts from thin occluders, we propose a simple yet effective iterative method that recovers the original scene from multiple apertures. The method requires two images if the depths of the scene and the occluder layer are known, or three images if the depths are unknown. The effectiveness of our proposed methods are demonstrated by both simulated and real experimental results.
Jinwei Gu, Ravi Ramamoorthi, Peter N. Belhumeur, Shree K. Nayar
ACM Trans. Graph.4
2008 Compressive Structured Light for Recovering Inhomogeneous Participating Media
Jinwei Gu, Shree K. Nayar, Eitan Grinspun, Peter N. Belhumeur, Ravi Ramamoorthi
ECCV (4)2
2008 FaceTracer: A Search Engine for Large Collections of Images with Faces
Neeraj Kumar 0006, Peter N. Belhumeur, Shree K. Nayar
ECCV (4)3
2008 What Is a Good Nearest Neighbors Algorithm for Finding Similar Patches in Images?
Neeraj Kumar 0006, Li Zhang 0003, Shree K. Nayar
ECCV (2)3
2008 Priors for Large Photo Collections and What They Reveal about Cameras
Sujit Kuthirummal, Aseem Agarwala, Dan B. Goldman, Shree K. Nayar
ECCV (4)4
2008 Flexible Depth of Field Photography
Hajime Nagahara, Sujit Kuthirummal, Changyin Zhou, Shree K. Nayar
ECCV (4)4
2008 Creating a Speech Enabled Avatar from a Single Photograph
abstract
This paper presents a complete framework for creating a speech-enabled avatar from a single image of a person. Our approach uses a generic facial motion model which represents deformations of a prototype face during speech. We have developed an HMM-based facial animation algorithm which takes into account both lexical stress and coarticulation. This algorithm produces realistic animations of the prototype facial surface from either text or speech. The generic facial motion model can be transformed to a novel face geometry using a set of corresponding points between the prototype face surface and the novel face. Given a face photograph, a small number of manually selected features in the photograph are used to deform the prototype face surface. The deformed surface is then used to animate the face in the photograph. We show several examples of avatars that are driven by text and speech inputs.
Dmitri Bitouk, Shree K. Nayar
VR2
2008 Capturing Images with Sparse Informational Pixels using Projected 3D Tags
abstract
In this paper, we propose a novel imaging system that enables the capture of photos and videos with sparse informational pixels. Our system is based on the projection and detection of 3D optical tags. We use an infrared (IR) projector to project temporally-coded (blinking) dots onto selected points in a scene. These tags are invisible to the human eye, but appear as clearly visible time-varying codes to an IR photosensor. As a proof of concept, we have built a prototype camera system (consisting of co-located visible and IR sensors) to simultaneously capture visible and IR images. When a user takes an image of a tagged scene using such a camera system, all the scene tags that are visible from the system's viewpoint are detected. In addition, tags that lie in the field of view but are occluded, and ones that lie just outside the field of view, are also automatically generated for the image. Associated with each tagged pixel is its 3D location and the identity of the object that the tag falls on. Our system can interface with conventional image recognition methods for efficient scene authoring, enabling objects in an image to be robustly identified using cheap cameras, minimal computations, and no domain knowledge. We demonstrate several applications of our system, including, photo-browsing, e-commerce, augmented reality, and objection localization.
Li Zhang 0003, Neesha Subramaniam, Robert Lin, Ramesh Raskar, Shree K. Nayar
VR5
2008 Cata-Fisheye Camera for Panoramic Imaging
abstract
We present a novel panoramic imaging system which uses a curved mirror as a simple optical attachment to a fish- eye lens. When compared to existing panoramic cameras, our "'cata-fisheye" camera has a simple, compact and inexpensive design, and yet yields high optical performance. It captures the desired panoramic field of view in two parts. The upper part is obtained directly by the fisheye lens and the lower part after reflection by the curved mirror. These two parts of the field of view have a small overlap that is used to stitch them into a single seamless panorama. The cata-fisheye concept allows us to design cameras with a wide range of fields of view by simply varying the parameters and position of the curved mirror. We provide an automatic method for the one-time calibration needed to stitch the two parts of the panoramic field of view. We have done a complete performance evaluation of our concept with respect to (i) the optical quality of the captured images, (ii) the working range of the camera over which the parallax is negligible, and (iii) the spatial resolution of the computed panorama. Finally, we have built a prototype cata-fisheye video camera with a spherical mirror that can capture high resolution panoramic images (3600times550pixels) with a 360deg (horizontal) x 55deg (vertical) field of view.
Gurunandan Krishnan, Shree K. Nayar
WACV2
2008 Face swapping: automatically replacing faces in photographs
abstract
In this paper, we present a complete system for automatic face replacement in images. Our system uses a large library of face images created automatically by downloading images from the internet, extracting faces using face detection software, and aligning each extracted face to a common coordinate system. This library is constructed off-line, once, and can be efficiently accessed during face replacement. Our replacement algorithm has three main stages. First, given an input image, we detect all faces that are present, align them to the coordinate system used by our face library, and select candidate face images from our face library that are similar to the input face in appearance and pose. Second, we adjust the pose, lighting, and color of the candidate face images to match the appearance of those in the input image, and seamlessly blend in the results. Third, we rank the blended candidate replacements by computing a match distance over the overlap region. Our approach requires no 3D model, is fully automatic, and generates highly plausible results across a wide range of skin tones, lighting conditions, and viewpoints. We show how our approach can be used for a variety of applications including face de-identification and the creation of appealing group photographs from a set of images. We conclude with a user study that validates the high quality of our replacement results, and a discussion on the current limitations of our system.
Dmitri Bitouk, Neeraj Kumar 0006, Samreen Dhillon, Peter N. Belhumeur, Shree K. Nayar
ACM Trans. Graph.5
2008 Light field transfer: global illumination between real and synthetic objects
abstract
We present a novel image-based method for compositing real and synthetic objects in the same scene with a high degree of visual realism. Ours is the first technique to allow global illumination and near-field lighting effects between both real and synthetic objects at interactive rates, without needing a geometric and material model of the real scene. We achieve this by using a light field interface between real and synthetic components---thus, indirect illumination can be simulated using only two 4D light fields, one captured from and one projected onto the real scene. Multiple bounces of interreflections are obtained simply by iterating this approach. The interactivity of our technique enables its use with time-varying scenes, including dynamic objects. This is in sharp contrast to the alternative approach of using 6D or 8D light transport functions of real objects, which are very expensive in terms of acquisition and storage and hence not suitable for real-time applications. In our method, 4D radiance fields are simultaneously captured and projected by using a lens array, video camera, and digital projector. The method supports full global illumination with restricted object placement, and accommodates moderately specular materials. We implement a complete system and show several example scene compositions that demonstrate global illumination effects between dynamic real and synthetic objects. Our implementation requires a single point light source and dark background.
Oliver Cossairt, Shree K. Nayar, Ravi Ramamoorthi
ACM Trans. Graph.2
2007 Flexible Mirror Imaging
abstract
The field of view of a traditional camera has a fixed shape. This severely restricts how scene elements can be composed into an image. We present a novel imaging system that uses a flexible mirror in conjunction with a camera to overcome this limitation. By deforming the mirror, our system can produce fields of view with a wide range of shapes and sizes. A captured image is typically a multi-perspective view of the scene with spatially varying resolution. As a result, scene objects appear distorted. To minimize these distortions, we have developed an efficient algorithm that maps a captured image to one with almost uniform resolution. To determine this mapping we need to know the shape of the mirror. For this, we have developed a simple calibration method that automatically estimates the mirror shape from its boundary, which is visible in the captured image. We present a number of examples that demonstrate that a flexible field of view imaging system can be used to compose scenes in ways that have not been possible before. This flexibility can be exploited in applications such as video surveillance and monitoring.
Sujit Kuthirummal, Shree K. Nayar
ICCV2
2007 Multispectral Imaging Using Multiplexed Illumination
abstract
Many vision tasks such as scene segmentation, or the recognition of materials within a scene, become considerably easier when it is possible to measure the spectral reflectance of scene surfaces. In this paper, we present an efficient and robust approach for recovering spectral reflectance in a scene that combines the advantages of using multiple spectral sources and a multispectral camera. We have implemented a system based on this approach using a cluster of light sources with different spectra to illuminate the scene and a conventional RGB camera to acquire images. Rather than sequentially activating the sources, we have developed a novel technique to determine the optimal multiplexing sequence of spectral sources so as to minimize the number of acquired images. We use our recovered spectral measurements to recover the continuous spectral reflectance for each scene point by using a linear model for spectral reflectance. Our imaging system can produce multispectral videos of scenes at 30fps. We demonstrate the effectiveness of our system through extensive evaluation. As a demonstration, we present the results of applying data recovered by our system to material segmentation and spectral relighting.
Jong-Il Park, Moon-Hyun Lee, Michael D. Grossberg, Shree K. Nayar
ICCV4
2007 A VQ-Based Demosaicing by Self-Similarity
abstract
In this paper, we propose a learning-based demosaicing and a restoration error detection. A Vector Quantization (VQ)-based method is utilized for learning. We take advantage of a self-similarity in an image for a codebook generation in VQ. The mosaic image is interpolated via a traditional method, and applied scaling, blurring, phase-shifting and resampling are used to create a training data for the codebook. The characteristics of the training data are similar to those of an ideal image. Using such training data and approximation of an ideal codevector by a locally linear embedding (LLE)-based method increases the probability of finding a suitable codevector from the codebook. Even if we cannot find a good codevector in an ill-conditioned case, the error detection finds poorly estimated pixel values and replaces them with better restoration results by another demosaicing method.
Yoshikuni Nomura, Shree K. Nayar
ICIP (3)2
2007 Visual Chatter in the Real World
Shree K. Nayar, Gurunandan Krishnan, Michael D. Grossberg, Ramesh Raskar
ISRR1
2007 Material Based Splashing of Water Drops
Kshitiz Garg, Gurunandan Krishnan, Shree K. Nayar
Rendering Techniques3
2007 Dirty Glass: Rendering Contamination on Transparent Surfaces
Jinwei Gu, Ravi Ramamoorthi, Peter N. Belhumeur, Shree K. Nayar
Rendering Techniques4
2007 Scene Collages and Flexible Camera Arrays
Yoshikuni Nomura, Li Zhang 0003, Shree K. Nayar
Rendering Techniques3
2007 Vision and Rain
Kshitiz Garg, Shree K. Nayar
Int. J. Comput. Vis.2
2007 Multiplexing for Optimal Lighting
abstract
Imaging of objects under variable lighting directions is an important and frequent practice in computer vision, machine vision, and image-based rendering. Methods for such imaging have traditionally used only a single light source per acquired image. They may result in images that are too dark and noisy, e.g., due to the need to avoid saturation of highlights. We introduce an approach that can significantly improve the quality of such images, in which multiple light sources illuminate the object simultaneously from different directions. These illumination-multiplexed frames are then computationally demultiplexed. The approach is useful for imaging dim objects, as well as objects having a specular reflection component. We give the optimal scheme by which lighting should be multiplexed to obtain the highest quality output, for signal-independent noise. The scheme is based on Hadamard codes. The consequences of imperfections such as stray light, saturation, and noisy illumination sources are then studied. In addition, the paper analyzes the implications of shot noise, which is signal-dependent, to Hadamard multiplexing. The approach facilitates practical lighting setups having high directional resolution. This is shown by a setup we devise, which is flexible, scalable, and programmable. We used it to demonstrate the benefit of multiplexing in experiments.
Yoav Y. Schechner, Shree K. Nayar, Peter N. Belhumeur
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Active refocusing of images and videos
abstract
We present a system for refocusing images and videos of dynamic scenes using a novel, single-view depth estimation method. Our method for obtaining depth is based on the defocus of a sparse set of dots projected onto the scene. In contrast to other active illumination techniques, the projected pattern of dots can be removed from each captured image and its brightness easily controlled in order to avoid under- or over-exposure. The depths corresponding to the projected dots and a color segmentation of the image are used to compute an approximate depth map of the scene with clean region boundaries. The depth map is used to refocus the acquired image after the dots are removed, simulating realistic depth of field effects. Experiments on a wide variety of scenes, including close-ups and live action, demonstrate the effectiveness of our method.
Francesc Moreno-Noguer, Peter N. Belhumeur, Shree K. Nayar
ACM Trans. Graph.3
2007 Prakash: lighting aware motion capture using photosensing markers and multiplexed illuminators
abstract
In this paper, we present a high speed optical motion capture method that can measure three dimensional motion, orientation, and incident illumination at tagged points in a scene. We use tracking tags that work in natural lighting conditions and can be imperceptibly embedded in attire or other objects. Our system supports an unlimited number of tags in a scene, with each tag uniquely identified to eliminate marker reacquisition issues. Our tags also provide incident illumination data which can be used to match scene lighting when inserting synthetic elements. The technique is therefore ideal for on-set motion capture or real-time broadcasting of virtual sets. Unlike previous methods that employ high speed cameras or scanning lasers, we capture the scene appearance using the simplest possible optical devices - a light-emitting diode (LED) with a passive binary mask used as the transmitter and a photosensor used as the receiver. We strategically place a set of optical transmitters to spatio-temporally encode the volume of interest. Photosensors attached to scene points demultiplex the coded optical signals from multiple transmitters, allowing us to compute not only receiver location and orientation but also their incident illumination and the reflectance of the surfaces to which the photosensors are attached. We use our untethered tag system, called Prakash, to demonstrate methods of adding special effects to captured videos that cannot be accomplished using pure vision techniques that rely on camera images.
Ramesh Raskar, Hideaki Nii, Bert de Decker, Yuki Hashimoto, Jay Summet, Dylan Moore 0002, Jonathan Westhues, Paul H. Dietz, John Barnwell, Shree K. Nayar, Masahiko Inami, Philippe Bekaert, Michael Noland, Vlad Branzoi, Erich Bruns
ACM Trans. Graph.11
2007 Time-Varying BRDFs
abstract
The properties of virtually all real-world materials change with time, causing their bidirectional reflectance distribution functions (BRDFs) to be time varying. However, none of the existing BRDF models and databases take time variation into consideration; they represent the appearance of a material at a single time instance. In this paper, we address the acquisition, analysis, modeling, and rendering of a wide range of time-varying BRDFs (TVBRDFs). We have developed an acquisition system that is capable of sampling a material's BRDF at multiple time instances, with each time sample acquired within 36 sec. We have used this acquisition system to measure the BRDFs of a wide range of time-varying phenomena, which include the drying of various types of paints (watercolor, spray, and oil), the drying of wet rough surfaces (cement, plaster, and fabrics), the accumulation of dusts (household and joint compound) on surfaces, and the melting of materials (chocolate). Analytic BRDF functions are fit to these measurements and the model parameters' variations with time are analyzed. Each category exhibits interesting and sometimes nonintuitive parameter trends. These parameter trends are then used to develop analytic TVBRDF models. The analytic TVBRDF models enable us to apply effects such as paint drying and dust accumulation to arbitrary surfaces and novel materials.
Kalyan Sunkavalli, Ravi Ramamoorthi, Peter N. Belhumeur, Shree K. Nayar
IEEE Trans. Vis. Comput. Graph.5
2006 Lensless Imaging with a Controllable Aperture
abstract
In this paper we propose a novel, highly flexible camera. The camera consists of an image detector and a special aperture, but no lens. The aperture is a set of parallel light attenuating layers whose transmittances are controllable in space and time. By applying different transmittance patterns to this aperture, it is possible to modulate the incoming light in useful ways and capture images that are impossible to capture with conventional lens-based cameras. For example, the camera can pan and tilt its field of view without the use of any moving parts. It can also capture disjoint regions of interest in the scene without having to capture the regions in between them. In addition, the camera can be used as a computational sensor, where the detector measures the end result of computations performed by the attenuating layers on the scene radiance values. These and other imaging functionalities can be implemented with the same physical camera and the functionalities can be switched from one video frame to the next via software. We have built a prototype camera based on this approach using a bare image detector and a liquid crystal modulator for the aperture. We discuss in detail the merits and limitations of lensless imaging using controllable apertures.
Assaf Zomet, Shree K. Nayar
CVPR (1)2
2006 Spatio-Angular Resolution Tradeoffs in Integral Photography
Todor G. Georgiev, Ke Colin Zheng, Brian Curless, David Salesin, Shree K. Nayar, Chintan Intwala
Rendering Techniques5
2006 Visual Chatter in the Real World
Shree K. Nayar, Gurunandan Krishnan
Rendering Techniques1
2006 Programmable Imaging: Towards a Flexible Camera
Shree K. Nayar, Vlad Branzoi, Terrance E. Boult
Int. J. Comput. Vis.1
2006 Corneal Imaging System: Environment from Eyes
Ko Nishino, Shree K. Nayar
Int. J. Comput. Vis.2
2006 Non-Single Viewpoint Catadioptric Cameras: Geometry and Analysis
Rahul Swaminathan, Michael D. Grossberg, Shree K. Nayar
Int. J. Comput. Vis.3
2006 Photorealistic rendering of rain streaks
abstract
Photorealistic rendering of rain streaks with lighting and viewpoint effects is a challenging problem. Raindrops undergo rapid shape distortions as they fall, a phenomenon referred to as oscillations. Due to these oscillations, the reflection of light by, and the refraction of light through, a falling raindrop produce complex brightness patterns within a single motion-blurred rain streak captured by a camera or observed by a human. The brightness pattern of a rain streak typically includes speckles, multiple smeared highlights and curved brightness contours. In this work, we propose a new model for rain streak appearance that captures the complex interactions between the lighting direction, the viewing direction and the oscillating shape of the drop. Our model builds upon a raindrop oscillation model that has been developed in atmospheric sciences. We have measured rain streak appearances under a wide range of lighting and viewing conditions and empirically determined the oscillation parameters that are dominant in raindrops. Using these parameters, we have rendered thousands of rain streaks to create a database that captures the variations in streak appearance with respect to lighting and viewing directions. We have developed an efficient image-based rendering algorithm that uses our streak database to add rain to a single image or a captured video with moving objects and sources. The rendering algorithm is very simple to use as it only requires a coarse depth map of the scene and the locations and properties of the light sources. We have rendered rain in a wide range of scenarios and the results show that our physically-based rain streak model greatly enhances the visual realism of rendered rain.
Kshitiz Garg, Shree K. Nayar
ACM Trans. Graph.2
2006 Time-varying surface appearance: acquisition, modeling and rendering
abstract
For computer graphics rendering, we generally assume that the appearance of surfaces remains static over time. Yet, there are a number of natural processes that cause surface appearance to vary dramatically, such as burning of wood, wetting and drying of rock and fabric, decay of fruit skins, and corrosion and rusting of steel and copper. In this paper, we take a significant step towards measuring, modeling, and rendering time-varying surface appearance. We describe the acquisition of the first time-varying database of 26 samples, encompassing a variety of natural processes including burning, drying, decay, and corrosion. Our main technical contribution is a Space-Time Appearance Factorization (STAF). This model factors space and time-varying effects. We derive an overall temporal appearance variation characteristic curve of the specific process, as well as space-dependent textures, rates, and offsets. This overall temporal curve controls different spatial locations evolve at the different rates, causing spatial patterns on the surface over time. We show that the model accurately represents a variety of phenomena. Moreover, it enables a number of novel rendering applications, such as transfer of the time-varying effect to a new static surface, control to accelerate time evolution in certain areas, extrapolation beyond the acquired sequence, and texture synthesis of time-varying appearance.
Jinwei Gu, Chien-I Tu, Ravi Ramamoorthi, Peter N. Belhumeur, Wojciech Matusik, Shree K. Nayar
ACM Trans. Graph.6
2006 Multiview radial catadioptric imaging for scene capture
abstract
In this paper, we present a class of imaging systems, called radial imaging systems , that capture a scene from a large number of view-points within a single image, using a camera and a curved mirror. These systems can recover scene properties such as geometry, reflectance, and texture. We derive analytic expressions that describe the properties of a complete family of radial imaging systems, including their loci of viewpoints, fields of view, and resolution characteristics. We have built radial imaging systems that, from a single image, recover the frontal 3D structure of an object, generate the complete texture map of a convex object, and estimate the parameters of an analytic BRDF model for an isotropic material. In addition, one of our systems can recover the complete geometry of a convex object by capturing only two images. These results show that radial imaging systems are simple, effective, and convenient devices for a wide range of applications in computer graphics and computer vision.
Sujit Kuthirummal, Shree K. Nayar
ACM Trans. Graph.2
2006 Acquiring scattering properties of participating media by dilution
abstract
The visual world around us displays a rich set of volumetric effects due to participating media. The appearance of these media is governed by several physical properties such as particle densities, shapes and sizes, which must be input (directly or indirectly) to a rendering algorithm to generate realistic images. While there has been significant progress in developing rendering techniques (for instance, volumetric Monte Carlo methods and analytic approximations), there are very few methods that measure or estimate these properties for media that are of relevance to computer graphics. In this paper, we present a simple device and technique for robustly estimating the properties of a broad class of participating media that can be either (a) diluted in water such as juices, beverages, paints and cleaning supplies, or (b) dissolved in water such as powders and sugar/salt crystals, or (c) suspended in water such as impurities. The key idea is to dilute the concentrations of the media so that single scattering effects dominate and multiple scattering becomes negligible, leading to a simple and robust estimation algorithm. Furthermore, unlike previous approaches that require complicated or separate measurement setups for different types or properties of media, our method and setup can be used to measure media with a complete range of absorption and scattering properties from a single HDR photograph. Once the parameters of the diluted medium are estimated, a volumetric Monte Carlo technique may be used to create renderings of any medium concentration and with multiple scattering. We have measured the scattering parameters of forty commonly found materials, that can be immediately used by the computer graphics community. We can also create realistic images of combinations or mixtures of the original measured materials, thus giving the user a wide flexibility in making realistic images of participating media.
Srinivasa G. Narasimhan, Mohit Gupta 0001, Craig Donner, Ravi Ramamoorthi, Shree K. Nayar, Henrik Wann Jensen
ACM Trans. Graph.5
2006 Fast separation of direct and global components of a scene using high frequency illumination
abstract
We present fast methods for separating the direct and global illumination components of a scene measured by a camera and illuminated by a light source. In theory, the separation can be done with just two images taken with a high frequency binary illumination pattern and its complement. In practice, a larger number of images are used to overcome the optical and resolution limitations of the camera and the source. The approach does not require the material properties of objects and media in the scene to be known. However, we require that the illumination frequency is high enough to adequately sample the global components received by scene points. We present separation results for scenes that include complex interreflections, subsurface scattering and volumetric scattering. Several variants of the separation approach are also described. When a sinusoidal illumination pattern is used with different phase shifts, the separation can be done using just three images. When the computed images are of lower resolution than the source and the camera, smoothness constraints are used to perform the separation using a single image. Finally, in the case of a static scene that is lit by a simple point source, such as the sun, a moving occluder and a video camera can be used to do the separation. We also show several simple examples of how novel images of a scene can be computed from the separation results.
Shree K. Nayar, Gurunandan Krishnan, Michael D. Grossberg, Ramesh Raskar
ACM Trans. Graph.1
2006 Projection defocus analysis for scene capture and image display
abstract
In order to produce bright images, projectors have large apertures and hence narrow depths of field. In this paper, we present methods for robust scene capture and enhanced image display based on projection defocus analysis. We model a projector's defocus using a linear system. This model is used to develop a novel temporal defocus analysis method to recover depth at each camera pixel by estimating the parameters of its projection defocus kemel in frequency domain. Compared to most depth recovery methods, our approach is more accurate near depth discontinuities. Furthermore, by using a coaxial projector-camera system, we ensure that depth is computed at all camera pixels, without any missing parts. We show that the recovered scene geometry can be used for refocus synthesis and for depth-based image composition. Using the same projector defocus model and estimation technique, we also propose a defocus compensation method that filters a projection image in a spatially-varying, depth-dependent manner to minimize its defocus blur after it is projected onto the scene. This method effectively increases the depth of field of a projector without modifying its optics. Finally, we present an algorithm that exploits projector defocus to reduce the strong pixelation artifacts produced by digital projectors, while preserving the quality of the projected image. We have experimentally verified each of our methods using real scenes.
Li Zhang 0003, Shree K. Nayar
ACM Trans. Graph.2
2005 A Projector-Camera System with Real-Time Photometric Adaptation for Dynamic Environments
abstract
Projection systems can be used to implement augmented reality, as well as to create both displays and interfaces on ordinary surfaces. Ordinary surfaces have varying reflectance, color, and geometry. These variations can be accounted for by integrating a camera into the projection system and applying methods from computer vision. The methods currently applied are fundamentally limited since they assume the camera, projector, and scene are static. In this paper, we describe a technique for photometrically adaptive projection that makes it possible to handle a dynamic environment. We begin by presenting a co-axial projector-camera system whose geometric correspondence is independent of changes in the environment. To handle photometric changes, our method uses the errors between the desired and measured appearance of the projected image. A key novel aspect of our algorithm is that we combine a physics-based model with dynamic feedback to achieve real time adaptation to the changing environment. We verify our algorithm through a wide variety of experiments. We show that it is accurate and runs in real-time. Our algorithm can be applied broadly to assist HCI, visualization, shape recovery, and entertainment applications.
Kensaku Fujii, Michael D. Grossberg, Shree K. Nayar
CVPR (1)3
2005 A Projector-Camera System with Real-Time Photometric Adaptation for Dynamic Environments
abstract
Projection systems can be used to implement augmented reality, as well as to create both displays and interfaces on ordinary surfaces. Ordinary surfaces have varying reflectance, color, and geometry. Current methods use a camera to account for these variations, but are fundamentally limited since they assume the camera, projector, and scene are static. In this article, we describe a technique for photometrically adaptive projection that makes it possible to handle a dynamic environment. We begin by presenting a co-axial projector-camera system. It consists of a camera and beam splitter, which attaches to an off-the-shelf projector. The co-axial design makes geometric calibration scene-independent. To handle photometric changes, our method uses the errors between the desired and measured appearance of the projected image. A key novel aspect of our algorithm is that we combine a physics-based model with dynamic feedback to achieve real time adaptation to the changing environment. We verify our algorithm through a wide variety of experiments. We show that it is accurate and runs in real-time. Our algorithm moves beyond the limits of a static environment to make real-time color compensation in a dynamic environment possible. It can be applied broadly to assist HCI, visualization, shape recovery, and entertainment applications.
Kensaku Fujii, Michael D. Grossberg, Shree K. Nayar
CVPR (2)3
2005 When Does a Camera See Rain?
abstract
Rain produces sharp intensity fluctuations in images and videos, which degrade the performance of outdoor vision systems. These intensity fluctuations depend on various factors, such as the camera parameters, the properties of rain, and the brightness of the scene. We show that the properties of rain - its small drop size, high velocity, and low density - make its visibility strongly dependent on camera parameters such as exposure time and depth of field. We show that these parameters can be selected so as to reduce or even remove the effects of rain without altering the appearance of the scene. Conversely, the parameters of a camera can also be set to enhance the visual effects of rain. This can be used to develop an inexpensive and portable camera-based rain gauge that provides instantaneous rain rate measurements. The proposed methods serve to make vision algorithms more robust to rain without any necessity for post-processing. In addition, they can be used to control the visual effects of rain during the filming of movies
Kshitiz Garg, Shree K. Nayar
ICCV2
2005 Structured Light in Scattering Media
abstract
Virtually all structured light methods assume that the scene and the sources are immersed in pure air and that light is neither scattered nor absorbed. Recently, however, structured lighting has found growing application in underwater and aerial imaging, where scattering effects cannot be ignored. In this paper, we present a comprehensive analysis of two representative methods - light stripe range scanning and photometric stereo - in the presence of scattering. For both methods, we derive physical models for the appearances of a surface immersed in a scattering medium. Based on these models, we present results on (a) the condition for object detectability in light striping and (b) the number of sources required for photometric stereo. In both cases, we demonstrate that while traditional methods fail when scattering is significant, our methods accurately recover the scene (depths, normals, albedos) as well as the properties of the medium. These results are in turn used to restore the appearances of scenes as if they were captured in clear air. Although we have focused on light striping and photometric stereo, our approach can also be extended to other methods such as grid coding, gated and active polarization imaging.
Srinivasa G. Narasimhan, Shree K. Nayar, Sanjeev J. Koppal
ICCV2
2005 Using Eye Reflections for Face Recognition Under Varying Illumination
abstract
Face recognition under varying illumination remains a challenging problem. Much progress has been made toward a solution through methods that require multiple gallery images of each subject under varying illumination. Yet for many applications, this requirement is too severe. In this paper, we propose a novel method that requires only a single gallery image per subject taken under unknown lighting. The method builds upon two contributions. We first estimate the lighting from its reflection in the eyes. This allows us to explicitly recover the illumination in the single gallery images as well as the probe image. Next, we exploit the local linearity of face appearance variation across different people. We represent the gallery images as locally linear montages of images of many different faces taken under the same lighting (bootstrap images). Then, we transfer the estimated combination of bootstrap images to synthesize each subject's face under tile probe lighting to accomplish recognition. Finally, we show through tests on the CMU PIE database that we can achieve better recognition results using our lighting estimation method and locally linear montages than the current state-of-the-art.
Ko Nishino, Peter N. Belhumeur, Shree K. Nayar
ICCV3
2005 The Raxel Imaging Model and Ray-Based Calibration
Michael D. Grossberg, Shree K. Nayar
Int. J. Comput. Vis.2
2005 Video Super-Resolution Using Controlled Subpixel Detector Shifts
abstract
Video cameras must produce images at a reasonable frame-rate and with a reasonable depth of field. These requirements impose fundamental physical limits on the spatial resolution of the image detector. As a result, current cameras produce videos with a very low resolution. The resolution of videos can be computationally enhanced by moving the camera and applying super-resolution reconstruction algorithms. However, a moving camera introduces motion blur, which limits super-resolution quality. We analyze this effect and derive a theoretical result showing that motion blur has a substantial degrading effect on the performance of super-resolution. The conclusion is that, in order to achieve the highest resolution, motion blur should be avoided. Motion blur can be minimized by sampling the space-time volume of the video in a specific manner. We have developed a novel camera, called the "jitter camera," that achieves this sampling. By applying an adaptive super-resolution algorithm to the video produced by the jitter camera, we show that resolution can be notably enhanced for stationary or slowly moving objects, while it is improved slightly or left unchanged for objects with fast and complex motions. The end result is a video that has a significantly higher resolution than the captured one.
Moshe Ben-Ezra, Assaf Zomet, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 Enhancing Resolution Along Multiple Imaging Dimensions Using Assorted Pixels
abstract
Multisampled imaging is a general framework for using pixels on an image detector to simultaneously sample multiple dimensions of imaging (space, time, spectrum, brightness, polarization, etc.). The mosaic of red, green, and blue spectral filters found in most solid-state color cameras is one example of multisampled imaging. We briefly describe how multisampling can be used to explore other dimensions of imaging. Once such an image is captured, smooth reconstructions along the individual dimensions can be obtained using standard interpolation algorithms. Typically, this results in a substantial reduction of resolution (and, hence, image quality). One can extract significantly greater resolution in each dimension by noting that the light fields associated with real scenes have enormous redundancies within them, causing different dimensions to be highly correlated. Hence, multisampled images can be better interpolated using local structural models that are learned offline from a diverse set of training images. The specific type of structural models we use are based on polynomial functions of measured image intensities. They are very effective as well as computationally efficient. We demonstrate the benefits of structural interpolation using three specific applications. These are 1) traditional color imaging with a mosaic of color filters, 2) high dynamic range monochrome imaging using a mosaic of exposure filters, and 3) high dynamic range color imaging using a mosaic of overlapping color and exposure filters.
Srinivasa G. Narasimhan, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Clustered Blockwise PCA for Representing Visual Data
abstract
Principal Component Analysis (PCA) is extensively used in computer vision and image processing. Since it provides the optimal linear subspace in a least-square sense, it has been used for dimensionality reduction and subspace analysis in various domains. However, its scalability is very limited because of its inherent computational complexity. We introduce a new framework for applying PCA to visual data which takes advantage of the spatio-temporal correlation and localized frequency variations that are typically found in such data. Instead of applying PCA to the whole volume of data (complete set of images), we partition the volume into a set of blocks and apply PCA to each block. Then, we group the subspaces corresponding to the blocks and merge them together. As a result, we not only achieve greater efficiency in the resulting representation of the visual data, but also successfully scale PCA to handle large data sets. We present a thorough analysis of the computational complexity and storage benefits of our approach. We apply our algorithm to several types of videos. We show that, in addition to its storage and speed benefits, the algorithm results in a useful representation of the visual data.
Ko Nishino, Shree K. Nayar, Tony Jebara
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Generalized Mosaicing: Polarization Panorama
abstract
We present an approach to image the polarization state of object points in a wide field of view, while enhancing the radiometric dynamic range of maging systems by generalizing image mosaicing. The approach is biologically-inspired, as it emulates spatially varying polarization sensitivity of some animals. In our method, a spatially varying polarization and attenuation filter is rigidly attached to a camera. As the system moves, it senses each scene point multiple times, each time filtering it through a different filter polarizing angle, polarizance, and transmittance. Polarization is an additional dimension of the generalized mosaicing paradigm, which has recently yielded high dynamic range images and multispectral images in a wide field of view using other kinds of filters. The image acquisition is as easy as in traditional image mosaics. The computational algorithm can easily handle nonideal polarization filters (partial polarizers), variable exposures, and saturation in a single framework. The resulting mosaic represents the polarization state at each scene point. Using data acquired by this method, we demonstrate attenuation and enhancement of specular reflections and semireflection separation in an image mosaic.
Yoav Y. Schechner, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Removing photography artifacts using gradient projection and flash-exposure sampling
abstract
Flash images are known to suffer from several problems: saturation of nearby objects, poor illumination of distant objects, reflections of objects strongly lit by the flash and strong highlights due to the reflection of flash itself by glossy surfaces. We propose to use a flash and no-flash (ambient) image pair to produce better flash images. We present a novel gradient projection scheme based on a gradient coherence model that allows removal of reflections and highlights from flash images. We also present a brightness-ratio based algorithm that allows us to compensate for the falloff in the flash image brightness due to depth. In several practical scenarios, the quality of flash/no-flash images may be limited in terms of dynamic range. In such cases, we advocate using several images taken under different flash intensities and exposures. We analyze the flash intensity-exposure space and propose a method for adaptively sampling this space so as to minimize the number of captured images for any given scene. We present several experimental results that demonstrate the ability of our algorithms to produce improved flash images.
Amit K. Agrawal, Ramesh Raskar, Shree K. Nayar, Yuanzhen Li
ACM Trans. Graph.3
2005 A practical analytic single scattering model for real time rendering
abstract
We consider real-time rendering of scenes in participating media, capturing the effects of light scattering in fog, mist and haze. While a number of sophisticated approaches based on Monte Carlo and finite element simulation have been developed, those methods do not work at interactive rates. The most common real-time methods are essentially simple variants of the OpenGL fog model. While easy to use and specify, that model excludes many important qualitative effects like glows around light sources, the impact of volumetric scattering on the appearance of surfaces such as the diffusing of glossy highlights, and the appearance under complex lighting such as environment maps. In this paper, we present an alternative physically based approach that captures these effects while maintaining real time performance and the ease-of-use of the OpenGL fog model. Our method is based on an explicit analytic integration of the single scattering light transport equations for an isotropic point light source in a homogeneous participating medium. We can implement the model in modern programmable graphics hardware using a few small numerical lookup tables stored as texture maps. Our model can also be easily adapted to generate the appearances of materials with arbitrary BRDFs, environment map lighting, and precomputed radiance transfer methods, in the presence of participating media. Hence, our techniques can be widely used in real-time rendering.
Ravi Ramamoorthi, Srinivasa G. Narasimhan, Shree K. Nayar
ACM Trans. Graph.4
2004 Jitter Camera: High Resolution Video from a Low Resolution Detector
Moshe Ben-Ezra, Assaf Zomet, Shree K. Nayar
CVPR (2)3
2004 Detection and Removal of Rain from Videos
Kshitiz Garg, Shree K. Nayar
CVPR (1)2
2004 Making One Object Look Like Another: Controlling Appearance Using a Projector-Camera System
Michael D. Grossberg, Harish Peri, Shree K. Nayar, Peter N. Belhumeur
CVPR (1)3
2004 Programmable Imaging Using a Digital Micromirror Array
Shree K. Nayar, Vlad Branzoi, Terrance E. Boult
CVPR (1)1
2004 The World in an Eye
Ko Nishino, Shree K. Nayar
CVPR (1)2
2004 Uncontrolled Modulation Imaging
Yoav Y. Schechner, Shree K. Nayar
CVPR (2)2
2004 Motion-Based Motion Deblurring
abstract
Motion blur due to camera motion can significantly degrade the quality of an image. Since the path of the camera motion can be arbitrary, deblurring of motion blurred images is a hard problem. Previous methods to deal with this problem have included blind restoration of motion blurred images, optical correction using stabilized lenses, and special cmos sensors that limit the exposure time in the presence of motion. In this paper, we exploit the fundamental trade off between spatial resolution and temporal resolution to construct a hybrid camera that can measure its own motion during image integration. The acquired motion information is used to compute a point spread function (psf) that represents the path of the camera during integration. This psf is then used to deblur the image. To verify the feasibility of hybrid imaging for motion deblurring, we have implemented a prototype hybrid camera. This prototype system was evaluated in different indoor and outdoor scenes using long exposures and complex camera motion paths. The results show that, with minimal resources, hybrid imaging outperforms previous approaches to the motion blur problem. We conclude with a brief discussion on how our ideas can be extended beyond the case of global camera motion to the case where individual objects in the scene move with different velocities.
Moshe Ben-Ezra, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Modeling the Space of Camera Response Functions
abstract
Many vision applications require precise measurement of scene radiance. The function relating scene radiance to image intensity of an imaging system is called the camera response. We analyze the properties that all camera responses share. This allows us to find the constraints that any response function must satisfy. These constraints determine the theoretical space of all possible camera responses. We have collected a diverse database of real-world camera response functions (DoRF). Using this database, we show that real-world responses occupy a small part of the theoretical space of all possible responses. We combine the constraints from our theoretical space with the data from DoRF to create a low-parameter empirical model of response (EMoR). This response model allows us to accurately interpolate the complete response function of a camera from a small number of measurements obtained using a standard chart. We also show that the model can be used to accurately estimate the camera response from images of an arbitrary scene taken using different exposures. The DoRF database and the EMoR model can be downloaded at http://www.cs.columbia.edu/CAVE.
Michael D. Grossberg, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Multiresolution Histograms and Their Use for Recognition
abstract
The histogram of image intensities is used extensively for recognition and for retrieval of images and video from visual databases. A single image histogram, however, suffers from the inability to encode spatial image variation. An obvious way to extend this feature is to compute the histograms of multiple resolutions of an image to form a multiresolution histogram. The multiresolution histogram shares many desirable properties with the plain histogram including that they are both fast to compute, space efficient, invariant to rigid motions, and robust to noise. In addition, the multiresolution histogram directly encodes spatial information. We describe a simple yet novel matching algorithm based on the multiresolution histogram that uses the differences between histograms of consecutive image resolutions. We evaluate it against five widely used image features. We show that with our simple feature we achieve or exceed the performance obtained with more complicated features. Further, we show our algorithm to be the most efficient and robust.
Efstathios Hadjidemetriou, Michael D. Grossberg, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Lighting sensitive display
abstract
Although display devices have been used for decades, they have functioned without taking into account the illumination of their environment. We present the concept of a lighting sensitive display (LSD)---a display that measures the incident illumination and modifies its content accordingly. An ideal LSD would be able to measure the 4D illumination field incident upon it and generate a 4D light field in response to the illumination. However, current sensing and display technologies do not allow for such an ideal implementation. Our initial LSD prototype uses a 2D measurement of the illumination field and produces a 2D image in response to it. In particular, it renders a 3D scene such that it always appears to be lit by the real environment that the display resides in. The current system is designed to perform best when the light sources in the environment are distant from the display, and a single user in a known location views the display. The displayed scene is represented by compressing a very large set of images (acquired or rendered) of the scene that correspond to different lighting conditions. The compression algorithm is a lossy one that exploits not only image correlations over the illumination dimensions but also coherences over the spatial dimensions of the image. This results in a highly compressed representation of the original image set. This representation enables us to achieve high quality relighting of the scene in real time. Our prototype LSD can render 640 × 480 images of scenes under complex and varying illuminations at 15 frames per second using a 2 GHz processor. We conclude with a discussion on the limitations of the current implementation and potential areas for future research.
Shree K. Nayar, Peter N. Belhumeur, Terrance E. Boult
ACM Trans. Graph.1
2004 Eyes for relighting
abstract
The combination of the cornea of an eye and a camera viewing the eye form a catadioptric (mirror + lens) imaging system with a very wide field of view. We present a detailed analysis of the characteristics of this corneal imaging system. Anatomical studies have shown that the shape of a normal cornea (without major defects) can be approximated with an ellipsoid of fixed eccentricity and size. Using this shape model, we can determine the geometric parameters of the corneal imaging system from the image. Then, an environment map of the scene with a large field of view can be computed from the image. The environment map represents the illumination of the scene with respect to the eye. This use of an eye as a natural light probe is advantageous in many relighting scenarios. For instance, it enables us to insert virtual objects into an image such that they appear consistent with the illumination of the scene. The eye is a particularly useful probe when relighting faces. It allows us to reconstruct the geometry of a face by simply waving a light source in front of the face. Finally, in the case of an already captured image, eyes could be the only direct means for obtaining illumination information. We show how illumination computed from eyes can be used to replace a face in an image with another one. We believe that the eye not only serves as a useful tool for relighting but also makes relighting possible in situations where current approaches are hard to use.
Ko Nishino, Shree K. Nayar
ACM Trans. Graph.2
2003 Motion Deblurring Using Hybrid Imaging
abstract
Motion blur due to camera motion can significantly degrade the quality of an image. Since the path of the camera motion can be arbitrary, deblurring of motion blurred images is a hard problem. Previous methods to deal with this problem have included blind restoration of motion blurred images, optical correction using stabilized lenses, and special CMOS sensors that limit the exposure time in the presence of motion. In this paper, we exploit the fundamental tradeoff between spatial resolution and temporal resolution to construct a hybrid camera that can measure its own motion during image integration. The acquired motion information is used to compute a point spread function (PSF) that represents the path of the camera during integration. This PSF is then used to deblur the image. To verify the feasibility of hybrid imaging for motion deblurring, we have implemented a prototype hybrid camera. This prototype system was evaluated in different indoor and outdoor scenes using long exposures and complex camera motion paths. The results show that, with minimal resources, hybrid imaging outperforms previous approaches to the motion blur problem.
Moshe Ben-Ezra, Shree K. Nayar
CVPR (1)2
2003 What is the Space of Camera Response Functions?
abstract
Many vision applications require precise measurement of scene radiance. The function relating scene radiance to image brightness is called the camera response. We analyze the properties that all camera responses share. This allows us to find the constraints that any response function must satisfy. These constraints determine the theoretical space of all possible camera responses. We have collected a diverse database of real-world camera response functions (DoRF). Using this database we show that real-world responses occupy a small part of the theoretical space of all possible responses. We combine the constraints from our theoretical space with the data from DoRF to create a low-parameter Empirical Model of Response (EMoR). This response model allows us to accurately interpolate the complete response function of a camera from a small number of measurements obtained using a standard chart. We also show that the model can be used to accurately estimate the camera response from images of an arbitrary scene taken using different exposures. The DoRF database and the EMoR model can be downloaded at http://www.cs.columbia.edu/CAVE.
Michael D. Grossberg, Shree K. Nayar
CVPR (2)2
2003 Shedding Light on the Weather
abstract
Virtually all methods in image processing and computer vision, for removing weather effects from images, assume single scattering of light by particles in the atmosphere. In reality, multiple scattering effects are significant. A common manifestation of multiple scattering is the appearance of glows around light sources in bad weather. Modeling multiple scattering is critical to understanding the complex effects of weather on images, and hence essential for improving the performance of outdoor vision systems. We develop a new physics-based model for the multiple scattering of light rays as they travel from a source to an observer. This model is valid for various weather conditions including fog, haze, mist and rain. Our model enables us to recover from a single image the shapes and depths of sources in the scene. In addition, the weather condition and the visibility of the atmosphere can be estimated. These quantities can, in turn, be used to remove the glows of sources to obtain a clear picture of the scene. Based on these results, we demonstrate that a camera observing a distant source can serve as a "visual weather meter". The model and techniques described in this paper can also be used to analyze scattering in other media, such as fluids and tissues. Therefore, in addition to vision in bad weather, our work has implications for medical and underwater imaging.
Srinivasa G. Narasimhan, Shree K. Nayar
CVPR (1)2
2003 A Perspective on Distortions
abstract
A framework for analyzing distortions in non-single viewpoint imaging systems is presented. Such systems possess loci of viewpoints called caustics. In general, perspective (or undistorted) views cannot be computed from images acquired with such systems without knowing scene structure. Views computed without scene structure will exhibit distortions, which we call caustic distortions. We first introduce a taxonomy of distortions based on the geometry of imaging systems. Then, we derive a metric to quantify caustic distortions. We present an algorithm to compute minimally distorted views using simple priors on scene structure. These priors are defined as parameterized primitives such as spheres, planes and cylinders with simple uncertainty models for the parameters. To validate our method, we conducted extensive experiments on rendered and real images. In all cases our method produces nearly undistorted views even though the acquired images were strongly distorted. We also provide an approximation of the above method that warps the entire captured image into a quasi single viewpoint representation that can be used by any "viewer" to compute near-perspective views in real-time.
Rahul Swaminathan, Michael D. Grossberg, Shree K. Nayar
CVPR (2)3
2003 What Does Motion Reveal About Transparency?
abstract
The perception of transparent objects from images is known to be a very hard problem in vision. Given a single image, it is difficult to even detect the presence of transparent objects in the scene. In this paper, we explore what can be said about transparent objects by a moving observer. We show how features that are imaged through a transparent object behave differently from those that are rigidly attached to the scene. We present a novel model-based approach to recover the shapes and the poses of transparent objects from known motion. The objects can be complex in that they may be composed of multiple layers with different refractive indices. We have conducted numerous simulations to verify the practical feasibility of our algorithm. We have applied it to real scenes that include transparent objects and recovered the shapes of the objects with high accuracy.
Moshe Ben-Ezra, Shree K. Nayar
ICCV2
2003 A Class of Photometric Invariants: Separating Material from Shape and Illumination
abstract
We derive a new class of photometric invariants that can be used for a variety of vision tasks including lighting invariant material segmentation, change detection and tracking, as well as material invariant shape recognition. The key idea is the formulation of a scene radiance model for the class of "separable" BRDFs, that can be decomposed into material related terms and object shape and lighting related terms. All the proposed invariants are simple rational functions of the appearance parameters (say, material or shape and lighting). The invariants in this class differ from one another in the number and type of image measurements they require. Most of the invariants in this class need changes in illumination or object position between image acquisitions. The invariants can handle large changes in lighting which pose problems for most existing vision algorithms. We demonstrate the power of these invariants using scenes with complex shapes, materials, textures, shadows and specularities.
Srinivasa G. Narasimhan, Visvanathan Ramesh, Shree K. Nayar
ICCV3
2003 Adaptive Dynamic Range Imaging: Optical Control of Pixel Exposures Over Space and Time
abstract
This paper presents a new approach to imaging that significantly enhances the dynamic range of a camera. The key idea is to adapt the exposure of each pixel on the image detector, based on the radiance value of the corresponding scene point. This adaptation is done in the optical domain, that is, during image formation. In practice, this is achieved using a spatial light modulator whose transmittance can be varied with high resolution over space and time. A real-time control algorithm is developed that uses acquired images to automatically adjust the transmittance function of the spatial modulator. Each captured image and its corresponding transmittance function are used to compute a very high dynamic range image that is linear in scene radiance. We have implemented a video-rate adaptive dynamic range camera that consists of a color CCD detector and a controllable liquid crystal light modulator. Experiments have been conducted in scenarios with complex and harsh lighting conditions. The results indicate that adaptive imaging can have a significant impact on vision applications such as monitoring, tracking, recognition, and navigation.
Shree K. Nayar, Vlad Branzoi
ICCV1
2003 A Theory of Multiplexed Illumination
abstract
Imaging of objects under variable lighting directions is an important and frequent practice in computer vision and image-based rendering. We introduce an approach that significantly improves the quality of such images. Traditional methods for acquiring images under variable illumination directions use only a single light source per acquired image. In contrast, our approach is based on a multiplexing principle, in which multiple light sources illuminate the object simultaneously from different directions. Thus, the object irradiance is much higher. The acquired images are then computationally demultiplexed. The number of image acquisitions is the same as in the single-source method. The approach is useful for imaging dim object areas. We give the optimal code by which the illumination should be multiplexed to obtain the highest quality output. For n images corresponding to n light sources, the noise is reduced by /spl radic/(n)/2 relative to the signal. This noise reduction translates to a faster acquisition time or an increase in density of illumination direction samples. It also enables one to use lighting with high directional resolution using practical setups, as we demonstrate in our experiments.
Yoav Y. Schechner, Shree K. Nayar, Peter N. Belhumeur
ICCV2
2003 Seeing Through Bad Weather
Shree K. Nayar, Srinivasa G. Narasimhan
ISRR1
2003 Generalized Mosaicing: High Dynamic Range in a Wide Field of View
Yoav Y. Schechner, Shree K. Nayar
Int. J. Comput. Vis.2
2003 Determining the Camera Response from Images: What Is Knowable?
abstract
An image acquired by a camera consists of measured intensity values which are related to scene radiance by a function called the camera response function. Knowledge of this response is necessary for computer vision algorithms which depend on scene radiance. One way the response has been determined is by establishing a mapping of intensity values between images taken with different exposures. We call this mapping the intensity mapping function. In this paper, we address two basic questions. What information from a pair of images taken at different exposures is needed to determine the intensity mapping function? Given this function, can the response of the camera and the exposures of the images be determined? We completely determine the ambiguities associated with the recovery of the response and the ratios of the exposures. We show all methods that have been used to recover the response break these ambiguities by making assumptions on the exposures or on the form of the response. We also show when the ratio of exposures can be recovered directly from the intensity mapping, without recovering the response. We show that the intensity mapping between images is determined solely by the intensity histograms of the images. We describe how this allows determination of the intensity mapping between images without registration. This makes it possible to determine the intensity mapping in sequences with some motion of both the camera and objects in the scene.
Michael D. Grossberg, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Contrast Restoration of Weather Degraded Images
abstract
Images of outdoor scenes captured in bad weather suffer from poor contrast. Under bad weather conditions, the light reaching a camera is severely scattered by the atmosphere. The resulting decay in contrast varies across the scene and is exponential in the depths of scene points. Therefore, traditional space invariant image processing techniques are not sufficient to remove weather effects from images. We present a physics-based model that describes the appearances of scenes in uniform bad weather conditions. Changes in intensities of scene points under different weather conditions provide simple constraints to detect depth discontinuities in the scene and also to compute scene structure. Then, a fast algorithm to restore scene contrast is presented. In contrast to previous techniques, our weather removal algorithm does not require any a priori scene structure, distributions of scene reflectances, or detailed knowledge about the particular weather condition. All the methods described in this paper are effective under a wide range of weather conditions including haze, mist, fog, and conditions arising due to other aerosols. Further, our methods can be applied to gray scale, RGB color, multispectral and even IR images. We also extend our techniques to restore contrast of scenes with moving objects, captured using a video camera.
Srinivasa G. Narasimhan, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 What Can Be Known about the Radiometric Response from Images?
Michael D. Grossberg, Shree K. Nayar
ECCV (4)2
2002 Resolution Selection Using Generalized Entropies of Multiresolution Histograms
Efstathios Hadjidemetriou, Michael D. Grossberg, Shree K. Nayar
ECCV (1)3
2002 All the Images of an Outdoor Scene
Srinivasa G. Narasimhan, Shree K. Nayar
ECCV (3)3
2002 Assorted Pixels: Multi-sampled Imaging with Structural Models
Shree K. Nayar, Srinivasa G. Narasimhan
ECCV (4)1
2002 On the Motion and Appearance of Specularities in Image Sequences
Rahul Swaminathan, Sing Bing Kang, Richard Szeliski, Antonio Criminisi, Shree K. Nayar
ECCV (1)5
2002 Vision and the Atmosphere
Srinivasa G. Narasimhan, Shree K. Nayar
Int. J. Comput. Vis.2
2002 Rectified Catadioptric Stereo Sensors
abstract
It has been shown elsewhere how mirrors can be used to capture stereo images with a single camera, an approach termed catadioptric stereo. We present novel catadioptric sensors that use mirrors to produce rectified stereo images. The scanline correspondence of these images benefits real-time stereo by avoiding the computational cost and image degradation due to resampling when rectification is performed after image capture. First, we develop a theory which determines the number of mirrors that must be used and the constraints on those mirrors that must be satisfied to obtain rectified stereo images with a single camera. Then, we discuss in detail the use of both one and three mirrors. In addition, we show how the mirrors should be placed in order to minimize sensor size for a given baseline, an important design consideration. In order to understand the feasibility of building these sensors, we analyze rectification errors due to misplacement of the camera with respect to the mirrors.
Joshua Gluckman, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Generalized Mosaicing: Wide Field of View Multispectral Imaging
abstract
We present an approach to significantly enhance the spectral resolution of imaging systems by generalizing image mosaicing. A filter transmitting spatially varying spectral bands is rigidly attached to a camera. As the system moves, it senses each scene point multiple times, each time in a different spectral band. This is an additional dimension of the generalized mosaic paradigm, which has demonstrated yielding high radiometric dynamic range images in a wide field of view, using a spatially varying density filter. The resulting mosaic represents the spectrum at each scene point. The image acquisition is as easy as in traditional image mosaics. We derive an efficient scene sampling rate, and use a registration method that accommodates the spatially varying properties of the filter. Using the data acquired by this method, we demonstrate scene rendering under different simulated illumination spectra. We are also able to infer information about the scene illumination. The approach was tested using a standard 8-bit black/white video camera and a fixed spatially varying spectral (interference) filter.
Yoav Y. Schechner, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Rectifying Transformations That Minimize Resampling Effects
abstract
Image rectification is the process of warping a pair of stereo images in order to align the epipolar lines with the scan-lines of the images. Once a pair of images is rectified, stereo matching can be implemented in an efficient manner. Given the epipolar geometry, it is straightforward to define a rectifying transformation, however, many transformations will lead to unwanted image distortions. In this paper, we present a novel method for stereo rectification that determines the transformation that minimizes the effects of resampling that can impede stereo matching. The effects we seek to minimize are the loss of pixels due to under-sampling and the creation of new pixels due to over-sampling. To minimize these effects we parameterize the family of rectification transformations and solve for the one that minimizes the change in local area integrated over the area of the images.
Joshua Gluckman, Shree K. Nayar
CVPR (1)2
2001 Spatial Information in Multiresolution Histograms
abstract
Intensity histograms have been used extensively for recognition and for retrieval of images and video from visual databases. Intensity histograms of images at different individual resolutions have also been used for indexing. They suffer, however, from the inability to encode spatial image information. Spatial information can be incorporated into histograms simply by taking histograms of an image at multiple resolutions together to form a multiresolution histogram. Multiresolution histograms can also be computed, stored, and matched efficiently. The authors analyze and quantify the relation and sensitivity of the multiresolution histogram to spatial image information as well as to properties of shapes and textures in an image. They verify the analytical results experimentally and demonstrate the ability of multiresolution histograms to discriminate between images, as well as their robustness to noise.
Efstathios Hadjidemetriou, Michael D. Grossberg, Shree K. Nayar
CVPR (1)3
2001 Removing Weather Effects from Monochrome Images
abstract
Images of outdoor scenes captured in bad weather suffer from poor contrast. Under bad weather conditions, the light reaching a camera is severely scattered by the atmosphere. The resulting decay in contrast varies across the scene and is exponential in the depths of scene points. Therefore, traditional space invariant image processing techniques are not sufficient to remove weather effects from images. In this paper, we present a fast physics-based method to compute scene structure and hence restore contrast of the scene from two or more images taken in bad weather In contrast to previous techniques, our method does not require any a priori weather-specific or scene information, and is effective under a wide range of weather conditions including haze, mist, fog and other aerosols. Further, our method can be applied to gray-scale, RGB color, multi-spectral and even IR images. We also extend the technique to restore contrast of scenes with moving objects, captured using a video camera.
Srinivasa G. Narasimhan, Shree K. Nayar
CVPR (2)2
2001 Instant Dehazing of Images Using Polarization
abstract
We present an approach to easily remove the effects of haze from images. It is based on the fact that usually airlight scattered by atmospheric particles is partially polarized. Polarization filtering alone cannot remove the haze effects, except in restricted situations. Our method, however, works under a wide range of atmospheric and viewing conditions. We analyze the image formation process, taking into account polarization effects of atmospheric scattering. We then invert the process to enable the removal of haze from images. The method can be used with as few as two images taken through a polarizer at different orientations. This method works instantly, without relying on changes of weather conditions. We present experimental results of complete dehazing in far from ideal conditions for polarization filtering. We obtain a great improvement of scene contrast and correction of color. As a by product, the method also yields a range (depth) map of the scene, and information about properties of the atmospheric particles.
Yoav Y. Schechner, Srinivasa G. Narasimhan, Shree K. Nayar
CVPR (1)3
2001 A General Imaging Model and a Method for Finding its Parameters
abstract
Linear perspective projection has served as the dominant imaging model in computer vision. Recent developments in image sensing make the perspective model highly restrictive. This paper presents a general imaging model that can be used to represent an arbitrary imaging system. It is observed that all imaging systems perform a mapping from incoming scene rays to photo-sensitive elements on the image detector. This mapping can be conveniently described using a set of virtual sensing elements called raxels. Raxels include geometric, radiometric and optical properties. We present a novel calibration method that uses structured light patterns to extract the raxel parameters of an arbitrary imaging system. Experimental results for perspective as well as ion-perspective imaging systems are included.
Michael D. Grossberg, Shree K. Nayar
ICCV2
2001 Finding "Anomalies" in an Arbitrary Image
abstract
A fast and general method to extract "anomalies" in an arbitrary image is proposed. The basic idea is to compute a probability density for sub-regions in an image, conditioned upon the areas surrounding the sub-regions. Linear estimation and Independent Component Analysis (ICA) are combined to obtain the probability estimates. Pseudo non-parametric correlation is used to group sets of similar surrounding patterns, from which a probability for the occurrence of a given sub-region is derived. A carefully designed multi-dimensional histogram, based on compressed vector representations, enables efficient and high-resolution extraction of anomalies from the image. Our current (unoptimized) implementation performs anomaly extraction in about 30 seconds for a 640/spl times/480 image using a 700 MHz PC. Experimental results are included that demonstrate the performance of the proposed method.
Toshifumi Honda, Shree K. Nayar
ICCV2
2001 Generalized Mosaicing
Yoav Y. Schechner, Shree K. Nayar
ICCV2
2001 Caustics of Catadioptric Cameras
abstract
Conventional vision systems and algorithms assume the camera to have a single viewpoint. However, sensors need not always maintain a single viewpoint. For instance, an incorrectly aligned system could cause non-single viewpoints. Also, systems could be designed to specifically deviate from a single viewpoint to trade-off image characteristics such as resolution and field of view. In these cases, the locus of viewpoints forms what is called a caustic. In this paper, we present an in-depth analysis of caustics of catadioptric cameras with conic reflectors. Properties of caustics with respect to field of view and resolution are presented. Finally, we present ways to calibrate conic catadioptric systems and estimate their caustics from known camera motion.
Rahul Swaminathan, Michael D. Grossberg, Shree K. Nayar
ICCV3
2001 Catadioptric Stereo Using Planar Mirrors
Joshua Gluckman, Shree K. Nayar
Int. J. Comput. Vis.2
2001 Histogram Preserving Image Transformations
Efstathios Hadjidemetriou, Michael D. Grossberg, Shree K. Nayar
Int. J. Comput. Vis.3
2000 Rectified Catadioptric Stereo Sensors
abstract
It has been previously shown how mirrors can be used to capture stereo images with a single camera, an approach termed catadioptric stereo. In this paper we present novel catadioptric sensors which use mirrors to produce rectified stereo images. The scan-line correspondence of these images benefits real-time stereo by avoiding the computational cost and image degradation due to resampling when rectification is performed after image capture. First, we develop a theory which determines the number of mirrors that must be used and the constraints on those mirrors that must be satisfied to obtain rectified stereo images with a single camera. Then we discuss in detail the use of both one and three mirrors. In addition, we show how the mirrors should be placed in order to minimize sensor size for a given baseline, an important design consideration.
Joshua Gluckman, Shree K. Nayar
CVPR2
2000 Histogram Preserving Image Transformations
abstract
Histograms are used to analyze and classify images. They have been found experimentally to have low sensitivity to certain types of image morphisms, for example, viewpoint changes and object deformations. However the precise effect of these image morphisms on the histogram has not been studied. In this work we derive the complete class of local transformations that preserve the histogram or simply scale its magnitude. To achieve this the transformations are represented as solutions to families of vector fields acting on the image. It is then shown that weak perspective projection and paraperspective projection belong to this class and simply scale the histogram. The results on weak perspective projection, together with the effect of illumination, are used to compute the histogram of the projection of 3D polyhedral objects. We verify the analytical results with several examples. Moreover we present and test a system that recognizes and approximates the poses of 3D polyhedral objects independent of viewpoint.
Efstathios Hadjidemetriou, Michael D. Grossberg, Shree K. Nayar
CVPR3
2000 Chromatic Framework for Vision in Bad Weather
abstract
Conventional vision systems are designed to perform in clear weather. However, any outdoor vision system is incomplete without mechanisms that guarantee satisfactory performance under poor weather conditions. It is known that the atmosphere can significantly alter light energy reaching an observer. Therefore, atmospheric scattering models must be used to make vision systems robust in bad weather. In this paper, we develop a geometric framework for analyzing the chromatic effects of atmospheric scattering. First, we study a simple color model for atmospheric scattering and verify it for fog and haze. Then, based on the physics of scattering, we derive several geometric constraints on scene color changes, caused by varying atmospheric conditions. Finally, using these constraints we develop algorithms for computing fog or haze color depth segmentation, extracting three dimensional structure, and recovering "true" scene colors, from two or more images taken under different but unknown weather conditions.
Srinivasa G. Narasimhan, Shree K. Nayar
CVPR2
2000 360 x 360 Mosaics
abstract
Current mosaicing methods use narrow field of view cameras to acquire image data. This poses problems when computing a complete spherical mosaic. First, a large number of images are needed to capture a sphere. Second, errors in mosaicing make it difficult to complete the spherical mosaic without seams. Third, with a hand-held camera it is hard for the user to ensure complete coverage of the sphere. This paper presents two approaches to spherical mosaicing. The first is to rotate a 360 degree camera about a single axis to capture a sequence of 360 degree strips. The unknown rotations between the strips are estimated and the strips are blended together to obtain a spherical mosaic. The second approach seeks to significantly enhance the resolution of the computed mosaic by capturing 360 degree slices rather than strips. A variety of slice cameras are proposed that map a thin 360 degree sheet of rays onto a large image area. This results in the capture of high resolution slices despite the use of a low resolution video camera. A slice camera is rotated using a motorized turntable to obtain regular as well as stereoscopic spherical mosaics.
Shree K. Nayar, Amruta Karmarkar
CVPR1
2000 High Dynamic Range Imaging: Spatially Varying Pixel Exposures
abstract
While real scenes produce a wide range of brightness variations, vision systems use low dynamic range image detectors that typically provide 8 bits of brightness data at each pixel. The resulting low quality images greatly limit what vision can accomplish today. This paper proposes a very simple method for significantly enhancing the dynamic range of virtually any imaging system. The basic principle is to simultaneously sample the spatial and exposure dimensions of image irradiance. One of several ways to achieve this is by placing an optical mask adjacent to a conventional image detector array. The mask has a pattern with spatially varying transmittance, thereby giving adjacent pixels on the detector different exposures to the scene. The captured image is mapped to a high dynamic range image using an efficient image reconstruction algorithm. The end result is an imaging system that can measure a very wide range of scene radiance and produce a substantially larger number of brightness levels, with a slight reduction in spatial resolution. We conclude with several examples of high dynamic range images computed using spatially varying pixel exposures.
Shree K. Nayar, Tomoo Mitsunaga
CVPR1
2000 Nonmetric Calibration of Wide-Angle Lenses and Polycameras
abstract
Images taken with wide-angle cameras tend to have severe distortions which pull points towards the optical center. This paper proposes a simple method for recovering the distortion parameters without the use of any calibration objects. Since distortions cause straight lines in the scene to appear as curves in the image, our algorithm seeks to find the distortion parameters that map the image curves to straight lines. The user selects a small set of points along the image curves. Recovery of the distortion parameters is formulated as the minimization of an objective function which is designed to explicitly account for noise in the selected image points. Experimental results are presented for synthetic data as well as real images. We also present the idea of a polycamera which is defined as a tightly packed camera cluster. Possible configurations are proposed to capture very large fields of view. Such camera clusters tend to have a nonsingle viewpoint. We therefore provide analysis of what we call the minimum working distance for such clusters. Finally, we present results for a polycamera consisting of four wide-angle sensors having a minimum working distance of about 4 m. On undistorting the acquired images using our proposed technique, we create real-time high resolution panoramas.
Rahul Swaminathan, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.2
1999 Global Measures of Coherence for Edge Detector Evaluation
abstract
We propose a class of benchmarks for edge detector evaluation that require no ground truth. Each benchmark consists of a large number of images of a carefully designed scene for which we enforce a constraint on the edges, for example, that they are co-linear. We sample the space of edge appearances as densely as possible by capturing the images under widely varying imaging conditions. Not only do we change the viewing geometry and the illumination direction, but we also vary the camera parameters and the physical properties of the objects in the scene. We show that the degrees to which the constraints hold in the output edge-maps can be used as highly discriminating measures of edge detector performance. The code, images, and results which form our benchmarks are all available from the website http://www.cs.columbia.edu/CAVE/. The code and images enable a user to compare any new detector against several previous ones with minimal effort.
Simon Baker, Shree K. Nayar
CVPR2
1999 Planar Catadioptric Stereo: Geometry and Calibration
abstract
By using mirror reflections of a scene, stereo images can be captured with a single camera (catadioptric stereo). Single camera stereo provides both geometric and radiometric advantages over traditional two camera stereo. In this paper we discuss the geometry and calibration of catadioptric stereo with two planar mirrors and show how the relative orientation, the epipolar geometry and the estimation of the focal length are constrained by planar motion. In addition, we have implemented a real-time system which demonstrates the viability of stereo with mirrors as an alternative to traditional two camera stereo.
Joshua Gluckman, Shree K. Nayar
CVPR2
1999 Radiometric Self Calibration
abstract
A simple algorithm is described that computes the radiometric response function of an imaging system, from images of an arbitrary scene taken using different exposures. The exposure is varied by changing either the aperture setting or the shutter speed. The algorithm does not require precise estimates of the exposures used. Rough estimates of the ratios of the exposures (e.g. F-number settings on an inexpensive lens) are sufficient for accurate recovery of the response function as well as the actual exposure ratios. The computed response function is used to fuse the multiple images into a single high dynamic range radiance image. Robustness is tested using a variety of scenes and cameras as well as noisy synthetic images generated using 100 randomly selected response curves. Automatic rejection of image areas that have large vignetting effects or temporal scene variations make the algorithm applicable to not just photographic but also video cameras.
Tomoo Mitsunaga, Shree K. Nayar
CVPR2
1999 Folded Catadioptric Cameras
abstract
A framework is developed for the design and analysis of single-viewpoint catadioptric cameras that use two or more mirrors. The use of multiple mirrors permits folding of the optics which leads to more compact camera designs than ones that use a single mirror. A dictionary of camera designs that use two conic mirrors is presented. We show that any folded system that uses conic mirrors has a geometrically equivalent system that uses a single conic mirror. This result makes it easy to determine the scene-to-image mapping of a conic folded system. In addition, we discuss the optical benefits of using folded systems. As an example, we choose a camera design from our dictionary and optimize its parameters via optical simulations. This design is used to construct a compact video camera that provides a hemispherical field of view.
Shree K. Nayar, Venkata Peri
CVPR1
1999 Non-Metric Calibration of Wide-Angle Lenses and Polycameras
abstract
Images taken with wide-angle cameras tend to have severe distortions which pull points towards the optical center. This paper proposes a method for recovering the distortion parameters without the use of any calibration objects. The distortions cause straight lines in the scene to appear as curves in the image. Our algorithm seeks to find the distortion parameters that would map the image curves to straight lines. The user selects a small set of points along the image curves. Recovery of the parameters is formulated as the minimization of an objective function which is designed to explicitly account for noise in the selected image points. Experimental results are presented for synthetic data with different noise levels as well as for real images. Once calibrated, the image streams from these cameras can be undistorted in real time using look up tables. We also present an application of this calibration method for wide-angle camera clusters, which we call polycameras. We apply our distortion correction technique to a polycamera with four wide-angle cameras to create a high resolution 360 degree panorama in real-time.
Rahul Swaminathan, Shree K. Nayar
CVPR2
1999 Correlation Model for 3D Texture
abstract
While an exact definition of texture is somewhat elusive, texture can be qualitatively described as a distribution of color, albedo or local normal on a surface. In the literature, the word texture is often used to describe a color or albedo variation on a smooth surface. We refer to such texture as 2D texture. In real world scenes, texture is often due to surface height variations and can be termed 3D texture. Because of local foreshortening and masking, oblique views of 3D texture are not simple transformations of the frontal view. Consequently, texture representations such as the correlation function or power spectrum are also affected by local foreshortening and masking. This work presents a correlation model for a particular class of 3D textures. The model characterizes the spatial relationship among neighboring pixels in an image of 3D texture and the change of this spatial relationship with viewing direction.
Kristin J. Dana, Shree K. Nayar
ICCV2
1999 Vision in Bad Weather
abstract
Current vision systems are designed to perform in clear weather. Needless to say, in any outdoor application, there is no escape from "bad" weather. Ultimately, computer vision systems must include mechanisms that enable them to function (even if somewhat less reliably) in the presence of haze, fog, rain, hail and snow. We begin by studying the visual manifestations of different weather conditions. For this, we draw on what is already known about atmospheric optics. Next, we identify effects caused by bad weather that can be turned to our advantage. Since the atmosphere modulates the information carried from a scene point to the observer it can be viewed as a mechanism of visual information coding. Based on this observation, we develop models and methods for recovering pertinent scene properties, such as three-dimensional structure, from images taken under poor weather conditions.
Shree K. Nayar, Srinivasa G. Narasimhan
ICCV1
1999 A Theory of Single-Viewpoint Catadioptric Image Formation
Simon Baker, Shree K. Nayar
Int. J. Comput. Vis.2
1999 Bidirectional Reflection Distribution Function of Thoroughly Pitted Surfaces
Jan J. Koenderink, Andrea J. van Doorn, Kristin J. Dana, Shree K. Nayar
Int. J. Comput. Vis.4
1999 Reflectance and Texture of Real-World Surfaces
abstract
In this work, we investigate the visual appearance of real-world surfaces and the dependence of appearance on the geometry of imaging conditions. We discuss a new texture representation called the BTF (bidirectional texture function) which captures the variation in texture with illumination and viewing direction. We present a BTF database with image textures from over 60 different samples, each observed with over 200 different combinations of viewing and illumination directions. We describe the methods involved in collecting the database as well as the importqance and uniqueness of this database for computer graphics. A related quantity to the BTF is the familiar BRDF (bidirectional reflectance distribution function). The measurement methods involved in the BTF database are conducive to simultaneous measurement of the BRDF. Accordingly, we also present a BRDF database with reflectance measurements for over 60 different samples, each observed with over 200 different combinations of viewing and illumination directions. Both of these unique databases are publicly available and have important implications for computer graphics.
Kristin J. Dana, Bram van Ginneken, Shree K. Nayar, Jan J. Koenderink
ACM Trans. Graph.3
1998 Omnidirectional Vision
Shree K. Nayar
BMVC1
1998 Histogram Model for 3D Textures
abstract
Image texture can arise not only from surface albedo variations (2D texture) but also from surface height variations (3D texture). Since the appearance of 3D texture depends on the illumination and viewing direction in a complicated manner, such image texture can be called a bidirectional texture function. A fundamental representation of image texture is the histogram of pixel intensities. Since the histogram of 3D texture also depends on the illumination and viewing directions in a complex fashion, we refer to it as a bidirectional histogram. In this work, we present a concise analytical model for the bidirectional histogram of Lambertian, isotropic, randomly rough surfaces, which are common in real-world scenes. We demonstrate the accuracy of the histogram model by fitting to several samples from the Columbia-Utrecht texture database. The parameters obtained from the model fits are roughness measures which can be used in texture recognition schemes. In addition, the model has potential application in estimating illumination direction in scenes where surfaces of known tilt and roughness are visible. We demonstrate the usefulness of our model by employing it in a novel 3D texture synthesis procedure.
Kristin J. Dana, Shree K. Nayar
CVPR2
1998 A Theory of Catadioptric Image Formation
abstract
Conventional video cameras have limited fields of view which make them restrictive for certain applications in computational vision. A catadioptric sensor uses a combination of lenses and mirrors placed in a carefully arranged configuration to capture a much wider field of view. When designing a catadioptric sensor, the shape of the mirror(s) should ideally be selected to ensure that the complete catadioptric system has a single effective viewpoint. In this paper, we derive the complete class of single-lens single-mirror catadioptric sensors which have a single viewpoint and an expression for the spatial resolution of a catadioptric sensor in terms of the resolution of the camera used to construct it. We also include a preliminary analysis of the defocus blur caused by the use of a curved mirror.
Simon Baker, Shree K. Nayar
ICCV2
1998 Ego-Motion and Omnidirectional Cameras
abstract
Recent research in image sensors has produced cameras with very large fields of view. An area of computer vision research which will benefit from this technology is the computation of camera motion (ego-motion) from a sequence of images. Traditional cameras suffer from the problem that the direction of translation may lie outside of the field of view, making the computation of camera motion sensitive to noise. In this paper, we present a method for the recovery of ego-motion using omnidirectional cameras. Noting the relationship between spherical projection and wide-angle imaging devices, we propose mapping the image velocity vectors to a sphere, using the Jacobian of the transformation between the projection model of the camera and spherical projection. Once the velocity vectors are mapped to a sphere, we show how existing ego-motion algorithms can be applied and present some experimental results. These results demonstrate the ability to compute egomotion with omnidirectional cameras. 1
Joshua Gluckman, Shree K. Nayar
ICCV2
1998 Stereo with Mirrors
abstract
In this paper, we propose the use of mirrors and a single camera for computational stereo. When compared to conventional stereo systems that use two cameras, our method has a number of significant advantages such as wide field of view, single viewpoint projection, identical camera parameters and ease of calibration. We propose four stereo systems that use a single camera pointed towards planar, ellipsoidal, hyperboloidal, and paraboloidal mirrors. In each case, we present a derivation of the epipolar constraints. Next, we attempt to understand what can be seen by each system, and formalize the notion of field of view. We conclude with two experiments to obtain 3-D structure. In the first we use a pair of planar mirrors, and in the second a pair of paraboloidal mirrors. The results of our experiments demonstrate the viability of stereo using mirrors.
Sameer A. Nene, Shree K. Nayar
ICCV2
1998 Catadioptric video sensors
abstract
Conventional video cameras have limited fields of view which make them restrictive in a variety of applications. A catadioptric sensor uses a combination of lenses and mirrors placed in a carefully arranged configuration to capture a much wider field of view. At Columbia University, we have developed a wide range of catadioptric sensors. Some of these sensors have been designed to produce unusually large fields of view. Others have been constructed for the purpose of depth computation. All our sensors perform in real time using just a PC.
Shree K. Nayar, Joshua Gluckman, Rahul Swaminathan, Simon Lok, Terrance E. Boult
WACV1
1998 Parametric Feature Detection
Simon Baker, Shree K. Nayar, Hiroshi Murase
Int. J. Comput. Vis.2
1998 Stereo and Specular Reflection
Dinkar N. Bhat, Shree K. Nayar
Int. J. Comput. Vis.2
1998 Rational Filters for Passive Depth from Defocus
Masahiro Watanabe, Shree K. Nayar
Int. J. Comput. Vis.2
1998 Improved Diffuse Reflection Models for Computer Vision
Lawrence B. Wolff, Shree K. Nayar, Michael Oren
Int. J. Comput. Vis.2
1998 Ordinal Measures for Image Correspondence
abstract
We present ordinal measures of association for image correspondence in the context of stereo. Linear correspondence measures like correlation and the sum of squared difference between intensity distributions are known to be fragile. Ordinal measures which are based on relative ordering of intensity values in windows-rank permutations-have demonstrable robustness. By using distance metrics between two rank permutations, ordinal measures are defined. These measures are independent of absolute intensity scale and invariant to monotone transformations of intensity values like gamma variation between images. We have developed simple algorithms for their efficient implementation. Experiments suggest the superiority of ordinal measures over existing techniques under nonideal conditions. These measures serve as a general tool for image matching that are applicable to other vision problems such as motion estimation and texture-based image retrieval.
Dinkar N. Bhat, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 Motion estimation using ordinal measures
abstract
We present a method for motion estimation using ordinal measures. Ordinal measures are based on relative ordering of intensity values in an image region called rank permutation. While popular measures like the sum-of-squared-difference (SSD) and normalized correlation (NCC) rely on linearity between corresponding intensity values, ordinal measures only require them to be monotonically related so that rank permutations between corresponding regions are presented. This property turns out to be useful for motion estimation in tagged magnetic resonance images. We study the imaging equation involved in two methods of tagging and observe temporal monotonicity in intensity under certain conditions though the tags themselves fade. We compare our method to SSD and NCC in a rotating ring phantom image sequence. We present an experiment on a real heart image sequence which suggests the suitability of our method.
Dinkar N. Bhat, Shree K. Nayar, Alok Gupta
CVPR2
1997 Reflectance and Texture of Real-World Surfaces Authors
abstract
In this work, we investigate the visual appearance of real-world surfaces and the dependence of appearance on imaging conditions. We present a BRDF (bidirectional reflectance distribution function) database with reflectance measurements for over 60 different samples, each observed with over 200 different combinations of viewing and source directions. We fit the BRDF measurements to two recent models to obtain a BRDF parameter database. These BRDF parameters can be directly used for both image analysis and image synthesis. Finally, we present a BTF (bidirectional texture function) database with image textures from over 60 different samples, each observed with over 200 different combinations of viewing and source directions. Each of these unique databases has important implications for a variety of vision algorithms and each is made publicly available.
Kristin J. Dana, Shree K. Nayar, Bram van Ginneken, Jan J. Koenderink
CVPR2
1997 Catadioptric Omnidirectional Camera
abstract
Conventional video cameras have limited fields of view that make them restrictive in a variety of vision applications. There are several ways to enhance the field of view of an imaging system. However, the entire imaging system must have a single effective viewpoint to enable the generation of pure perspective images from a sensed image. A new camera with a hemispherical field of view is presented. Two such cameras can be placed back-to-back, without violating the single viewpoint constraint, to arrive at a truly omnidirectional sensor. Results are presented on the software generation of pure perspective images from an omnidirectional image, given any user-selected viewing direction and magnification. The paper concludes with a discussion on the spatial resolution of the proposed camera.
Shree K. Nayar
CVPR1
1997 Are Textureless Scenes Recoverable?
abstract
It is widely accepted that textureless surfaces cannot be recovered using passive sensing techniques. The problem is approached by viewing image formation as a fully three-dimensional mapping. It is shown that the lens encodes structural information of the scene within a compact three-dimensional space behind it. After analyzing the information content of this space and by using its properties we derive necessary and sufficient conditions for the recovery of textureless scenes. Based on these conditions, a simple procedure for recovering textureless scenes is described. We experimentally demonstrate the recovery of three textureless surfaces, namely, a line, a plane, and a paraboloid. Since textureless surfaces represent the worst case recovery scenario, all the results and the recovery procedure are naturally applicable to scenes with texture.
Hari Sundaram, Shree K. Nayar
CVPR2
1997 Separation of Reflection Components Using Color and Polarization
Shree K. Nayar, Xi-Sheng Fang, Terrance E. Boult
Int. J. Comput. Vis.1
1997 A Theory of Specular Surface Geometry
Michael Oren, Shree K. Nayar
Int. J. Comput. Vis.2
1997 Telecentric Optics for Focus Analysis
abstract
Magnification variations due to changes in focus setting pose problems for vision techniques, such as, depth from focus and defocus. The magnification of a conventional lens can be made invariant to defocus by simply adding an aperture at an analytically derived location. The resulting optical configuration is called "telecentric". It is shown that most commercially available lenses can be turned into telecentric ones. The procedure for calculating the position of the additional aperture and a detailed analysis of the photometric and geometric properties of telecentric lenses are presented. Experiments are reported that use a phase-based shift detection algorithm to demonstrate the magnification invariance of telecentric lenses.
Masahiro Watanabe, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 Detection of 3D objects in cluttered scenes using hierarchical eigenspace
abstract
This paper proposes a novel method to detect three-dimensional objects in arbitrary poses and sizes from a complex image and to simultaneously measure their poses and sizes using appearance matching. In the learning stage, for a sample object to be learned, a set of images is obtained by varying pose and size. This large image set is compactly represented by a manifold in compressed subspace spanned by eigenvectors of the image set. This representation is called the parametric eigenspace representation. In the object detection stage, a partial region in an input image is projected to the eigenspace, and the location of the projection relative to the manifold determines whether this region belongs to the object, and what its pose is in the scene. This process is sequentially applied to the entire image at different resolutions. Experimental results show that this method accurately detects the target objects.
Hiroshi Murase, Shree K. Nayar
Pattern Recognit. Lett.2
1996 Pattern Rejection
abstract
The efficiency of pattern recognition is particularly crucial in two scenarios; whenever there are a large number of classes to discriminate, and, whenever recognition must be performed a large number of times. We propose a single technique, namely, pattern rejection, that greatly enhances efficiency in both cases. A rejector is a generalization of a classifier, that quickly eliminates a large fraction of the candidate classes or inputs. This allows a recognition algorithm to dedicate its efforts to a much smaller number of possibilities. Importantly, a collection of rejectors may be combined to form a composite rejector, which is shown to be far more effective than any of its individual components. A simple algorithm is proposed for the construction of each of the component rejectors. Its generality is established through close relationships with the Karhunen-Loeve expansion and Fisher's discriminant analysis. Composite rejectors were constructed for two representative applications, namely, appearance matching based object recognition and local feature detection. The results demonstrate substantial efficiency improvements over existing approaches, most notably Fisher's discriminant analysis.
Simon Baker, Shree K. Nayar
CVPR2
1996 Ordinal Measures for Visual Correspondence
abstract
We present ordinal measures for establishing image correspondence. Linear correspondence measures like correlation and the sum of squared differences are known to be fragile. Ordinal measures, which are based on relative ordering of intensity values in windows, have demonstrable robustness to depth discontinuities, occlusion and noise. The relative ordering of intensity values in each window is represented by a rank permutation which is obtained by sorting the corresponding intensity data. By using a novel distance metric between the rank permutations, we arrive at ordinal correlation coefficients. These coefficients are independent of absolute intensity scale, i.e. they are normalized measures. Further, since rank permutations are invariant to monotone transformations of the intensity values, the coefficients are unaffected by nonlinear effects like gamma variation between images. We have developed a simple algorithm for their efficient implementation. Experiments suggest the superiority of ordinal measures over existing techniques under non-ideal conditions. Though we present ordinal measures in the context of stereo, they serve as a general tool for image matching that is applicable to other vision problems such as motion estimation and image registration.
Dinkar N. Bhat, Shree K. Nayar
CVPR2
1996 Parametric Feature Detection
abstract
We propose an algorithm to automatically construct feature detectors for arbitrary parametric features. To obtain a high level of robustness we advocate the use of realistic multi-parameter feature models and incorporate optical and sensing effects. Each feature is represented as a densely sampled parametric manifold in a low dimensional subspace of a Hilbert space. During detection, the brightness distribution around each image pixel is projected into the subspace. If the projection lies sufficiently close to the feature manifold, the feature is detected and the location of the closest manifold point yields the feature parameters. The concepts of parameter reduction by normalization, dimension reduction, pattern rejection, and heuristic search are all employed to achieve the required efficiency. By applying the algorithm to appropriate parametric feature models, detectors have been constructed for five features, namely, step edge, roof edge, line, corner, and circular disc. Detailed experiments are reported on the robustness of detection and the accuracy of parameter estimation.
Shree K. Nayar, Simon Baker, Hiroshi Murase
CVPR1
1996 Closest Point Search in High Dimensions
abstract
The problem of finding the closest point in high-dimensional spaces is common in computational vision. Unfortunately, the complexity of most existing search algorithms, such as k-d tree and R-tree, grows exponentially with dimension, making them impractical for dimensionality above 15. In nearly all applications, the closest point is of interest only if it lies within a user specified distance /spl epsiv/. We present a simple and practical algorithm to efficiently search for the nearest neighbor within Euclidean distance /spl epsiv/. Our algorithm uses a projection search technique along with a novel data structure to dramatically improve performance in high dimensions. A complexity analysis is presented which can help determine /spl epsiv/ in structured problems. Benchmarks clearly show the superiority of the proposed algorithm for high dimensional search problems frequently encountered in machine vision, such as real-time object recognition.
Sameer A. Nene, Shree K. Nayar
CVPR2
1996 Minimal operator set for passive depth from defocus
abstract
A fundamental problem in depth from defocus is the measurement of relative defocus between images. We propose a class of broadband operators that, when used together, provide invariance to scene texture and produce accurate and dense depth maps. Since the operators are broadband, a small number of them are sufficient for depth estimation of scenes with complex textural properties. Experiments are conducted on both synthetic and real scenes to evaluate the performance of the proposed operators. The depth detection gain error is less than 1%, irrespective of texture frequency. Depth accuracy is found to be 0.5/spl sim/1.2% of the distance of the object from the imaging optics.
Masahiro Watanabe, Shree K. Nayar
CVPR2
1996 Telecentric Optics for Computational Vision
Masahiro Watanabe, Shree K. Nayar
ECCV (2)2
1996 Algorithms for pattern rejection
abstract
The efficiency of pattern recognition is particularly crucial in two situations; whenever there are a large number of classes to discriminate, and, whenever recognition must be performed a large number of times. We develop a number of algorithms to cope with the demands of these difficult conditions. The algorithms achieve high efficiency by using pattern rejectors. A pattern rejector is a generalization of a classifier that quickly eliminates a large fraction of the candidate classes or inputs. After applying a rejector the recognition algorithms can concentrate their computational efforts on verifying the small number of remaining possibilities. The generality of our algorithms is established through a close relationship with the Karhunen-Loeve expansion. We experimented on two representative applications, namely, object recognition and feature detection. The results demonstrate substantial efficiency improvements over existing approaches, most notably Fisher's discriminant analysis (1939).
Simon Baker, Shree K. Nayar
ICPR2
1996 Learning by a generation approach to appearance-based object recognition
abstract
We propose a methodology for the generation of learning samples in appearance-based object recognition. In many practical situations, it is not easy to obtain a large number of learning samples. The proposed method learns object models from a large number of generated samples derived from a small number of actually observed images. The learning algorithm has two steps: 1) generation of a large number of images by image interpolation, or image deformation, and 2) compression of the large sample sets using parametric eigenspace representation. We compare our method with the previous methods that interpolate sample points in eigenspace, and show the performance of our method to be superior. Experiments were conducted for 432 image samples for 4 objects to demonstrate the effectiveness of the method.
Hiroshi Murase, Shree K. Nayar
ICPR2
1996 Dimensionality of illumination in appearance matching
abstract
Appearance matching was recently demonstrated as a robust and efficient approach to 3D object recognition and pose estimation. Each object is represented as a continuous appearance manifold in a low-dimensional subspace parametrized by object pose and illumination direction. Here, the structural properties of appearance manifolds are analyzed with the aim of making appearance representation efficient in off-line computation, storage requirements, and online recognition time. In particular, the effect of illumination on the structure of the appearance manifold is studied. It is shown that for an ideal diffused surface of arbitrary texture, the appearance manifold is linear and three dimensional. This enables the construction of the entire illumination manifold from just three images of the object taken using linearly independent light sources. This result is shown to hold even for illumination by multiple light sources and for concave surfaces that exhibit inter-reflections. Finally, a simple but efficient algorithm is presented that uses just three manifold points for recognizing images taken under novel illuminations.
Shree K. Nayar, Hiroshi Murase
ICRA1
1996 Real-time 100 object recognition system
abstract
A real-time vision system is described that can recognize 100 complex three-dimensional objects. In contrast to traditional strategies that rely on object geometry and local image features, the present system is founded on the concept of appearance matching. Appearance manifolds of the 100 objects were automatically learned using a computer-controlled turntable. The entire learning process was completed in 1 day. A recognition loop has been implemented that performs scene change detection, image segmentation, region normalizations, and appearance matching, in less than 1 second. The hardware used by the recognition system includes no more than a CCD color camera and a workstation. The real-time capability and interactive nature of the system have allowed numerous observers to test its performance. To quantify performance, we have conducted controlled experiments on recognition and pose estimation. The recognition rate was found to be 100% and object pose was estimated with a mean absolute error of 2.02 degrees and standard deviation of 1.67 degrees.
Shree K. Nayar, Sameer A. Nene, Hiroshi Murase
ICRA1
1996 Transparent grippers for robot vision
abstract
While existing grippers execute manipulation tasks, they occlude parts of the objects they grasp as well as parts of the workspace from vision sensors. We present the concept of a transparent gripper that enables vision sensors to image an object without occlusion while it is being manipulated. The physics of refraction, total internal reflection, lens effects, dispersion and transmittance are analyzed to determine the geometry and material properties of a practical transparent gripper. Algorithms are presented that compensate for image shifts caused by refraction. The experiments demonstrate the proposed gripper to be an effective solution to a variety of problems, including 3D object model generation.
Anton Nikolaev, Shree K. Nayar
ICRA2
1996 Reflectance based object recognition
Shree K. Nayar, Ruud M. Bolle
Int. J. Comput. Vis.1
1996 Real-Time Focus Range Sensor
abstract
Structures of dynamic scenes can only be recovered using a real-time range sensor. Depth from defocus offers an effective solution to fast and dense range estimation. However, accurate depth estimation requires theoretical and practical solutions to a variety of problems including recovery of textureless surfaces, precise blur estimation, and magnification variations caused by defocusing. Both textured and textureless surfaces are recovered using an illumination pattern that is projected via the same optical path used to acquire images. The illumination pattern is optimized to maximize accuracy and spatial resolution in computed depth. The relative blurring in two images is computed using a narrow-band linear operator that is designed by considering all the optical, sensing, and computational elements of the depth from defocus system. Defocus invariant magnification is achieved by the use of an additional aperture in the imaging optics. A prototype focus range sensor has been developed that has a workspace of 1 cubic foot and produces up to 512/spl times/480 depth estimates at 30 Hz with an average RMS error of 0.2%. Several experimental results are included to demonstrate the performance of the sensor.
Shree K. Nayar, Masahiro Watanabe, Minori Noguchi
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 Automatic generation of RBF networks using wavelets
abstract
Learning can be viewed as mapping from an input space to an output space. Examples of these mappings are used to construct a continuous function that approximates the given data and generalizes for intermediate instances. Radial-basis function (RBF) networks are used to formulate this approximating function. A novel method is introduced that automatically constructs a generalized radial-basis function (GRBF) network for a given mapping and error bound. This network is shown to be the smallest network within the error bound for the given mapping. The integral wavelet transform is used to determine the parameters of the network. Simple one-dimensional examples are used to demonstrate how the network constructed using the transform is superior to that constructed using standard ad hoc optimization techniques. The paper concludes with the automatic generation of GRBF networks for a multi-dimensional problem, namely, real-time 3D object recognition and pose estimation. The results of this application are favorable.
Sayan Mukherjee 0001, Shree K. Nayar
Pattern Recognit.2
1996 Subspace methods for robot vision
abstract
In contrast to the traditional approach, visual recognition is formulated as one of matching appearance rather than shape. For any given robot vision task, all possible appearance variations define its visual workspace. A set of images is obtained by coarsely sampling the workspace. The image set is compressed to obtain a low-dimensional subspace, called the eigenspace, in which the visual workspace is represented as a continuous appearance manifold. Given an unknown input image, the recognition system first projects the image to eigenspace. The parameters of the vision task are recognized based on the exact location of the projection on the appearance manifold. An efficient algorithm for finding the closest manifold point is described. The proposed appearance representation has several applications in robot vision. As examples, a precise visual positioning system, a real-time visual tracking system, and a real-time temporal inspection system are described.
Shree K. Nayar, Sameer A. Nene, Hiroshi Murase
IEEE Trans. Robotics Autom.1
1995 Stereo in the Presence of Specular Reflection
abstract
The problem of accurate depth estimation using stereo in the presence of specular reflection is addressed. Specular reflection, a fundamental and ubiquitous reflection mechanism, is viewpoint dependent and can cause large intensity differences at corresponding points, resulting in significant depth errors. We analyze the physics of specular reflection and the geometry of stereopsis which led us to a relationship between stereo vergence, surface roughness, and the likelihood of a correct match. Given a lower bound on surface roughness, an optimal binocular stereo configuration can be determined which maximizes precision in depth estimation despite specular reflection. However, surface roughness is difficult to estimate in unstructured environments. Therefore, trinocular configurations, independent of surface roughness, are determined such that at each scene point visible to all sensors, at least one stereo pair can compute produce depth. We have developed a simple algorithm to reconstruct depth from the multiple stereo pairs.>
Dinkar N. Bhat, Shree K. Nayar
ICCV2
1995 Automatic Generation of GRBF Networks for Visual Learning
abstract
Learning can often be viewed as the problem of mapping from an input space to an output space. Examples of these mappings are used to construct a continuous function that approximates given data and generalizes for intermediate instances. Generalized Radial Basis Function (GRBF) networks are used to formulate this approximating function. A novel method is introduced to construct an optimal GRBF network for a given mapping and error bound using the integral wavelet transform. Simple one-dimensional examples are used to demonstrate how the optimal network is superior to one constructed using standard ad hoc optimization techniques. The paper concludes with an application of optimal GRBF networks to object recognition and pose estimation. The results of this application are favorable.>
Sayan Mukherjee 0001, Shree K. Nayar
ICCV2
1995 Real-Time Focus Range Sensor
abstract
Structures of dynamic scenes can only be recovered using a real-time range sensor. Depth-from-defocus offers a direct solution to fast and dense range estimation. It is computationally efficient as it circumvents the correspondence problem faced by stereo and feature tracking in structure-from-motion. However, accurate depth estimation requires theoretical and practical solutions to a variety of problems including the recovery of textureless surfaces, precise blur estimation, and magnification variations caused by defocusing. Both textured and textureless surfaces are recovered using an illumination pattern that is projected via the same optical path used to acquire images. The illumination pattern is optimized to ensure maximum accuracy and spatial resolution in the computed depth. The relative blurring in two images is computed using a narrow-band linear operator that is designed by considering all the optical, sensing and computational elements of the depth-from-defocus system. Defocus-invariant magnification is achieved by the use of an additional aperture in the imaging optics. A prototype focus range sensor has been developed that produces up to 512/spl times/480 depth estimates at 30 Hz with an accuracy better than 0.3%. Several experimental results are included to demonstrate the performance of the sensor.>
Shree K. Nayar, Masahiro Watanabe, Minori Noguchi
ICCV1
1995 A Theory of Specular Surface Geometry
abstract
A theoretical framework is introduced for the perception of specular surface geometry. When an observer moves in three-dimensional space, real scene features, such as surface markings, remain stationary with respect to the surfaces they belong to. In contrast, a virtual feature, which is the specular reflection of a real feature, travels on the surface. Based on the notion of caustics, a novel feature classification algorithm is developed that distinguishes real and virtual features from their image trajectories that result from observer motion. Next, using support functions of curves, a closed-form relation is derived between the image trajectory of a virtual feature and the geometry of the specular surface it travels on. It is shown that in the 2D case where camera motion and the surface profile are coplanar, the profile is uniquely recovered by tracking just two unknown virtual features. Finally, these results are generalized to the case of arbitrary 3D surface profiles that are travelled by virtual features when camera motion is not confined to a plane. An algorithm is developed that uniquely recovers 3D surface profiles using a single virtual feature tracked from the occluding boundary of the object. All theoretical derivations and proposed algorithms are substantiated by experiments.>
Michael Oren, Shree K. Nayar
ICCV2
1995 Visual learning and recognition of 3-d objects from appearance
Hiroshi Murase, Shree K. Nayar
Int. J. Comput. Vis.2
1995 Generalization of the Lambertian model and implications for machine vision
Michael Oren, Shree K. Nayar
Int. J. Comput. Vis.2
1994 Illumination planning for object recognition in structured environments
abstract
This paper addresses the problem of illumination planning for robust object recognition in structured environments. Given a set of objects, the goal is to determine the illumination for which the objects are most distinguishable in appearance from each other. For each object, a large number of images is automatically obtained by varying pose and illumination. Images of all objects, together, constitute the planning image set. The planning set is compressed using the Karhunen-Loeve transform to obtain a low-dimensional subspace. For any given illumination, objects are represented as parametrized manifolds in the subspace. The minimum distance between the manifolds of too objects represents the similarity between the objects in the correlation sense. The optimal illumination is therefore one that maximizes the shortest distance between object manifolds. Results produced by the illumination planner heave been used to enhance the performance of an object recognition system.>
Hiroshi Murase, Shree K. Nayar
CVPR2
1994 Seeing Beyond Lambert's Law
Michael Oren, Shree K. Nayar
ECCV (2)2
1994 Microscopic shape from focus using active illumination
abstract
Shape from focus relies on surface texture for the computation of depth. In many real-world applications, surfaces can be smoothly shaded and lacking in detectable texture. In such cases, shape from focus generates inaccurate and sparse depth maps. This paper presents a novel extension to the original shape from focus method. A strong texture is forced on imaged surfaces by the use of active illumination. The exact pattern of the projected illumination is determined through a careful Fourier analysis of all the optical effects involved in focus analysis. When the focus operator used is a 2D Laplacian, the optimal illumination pattern is found to be a checkerboard whose pitch is the same size as the distance between adjacent elements in the discrete Laplacian kernel. This analysis also reveals the exact number of images required for accurate shape recovery. These results are experimentally verified using an optical microscope. Surfaces lacking in texture, such as, three-dimensional structures on silicon substrates and solder joints on circuit boards were used in the experiments. The results show that the derived illumination pattern is in fact optimal, facilitating accurate shape recovery of complex and pertinent industrial samples.
Minori Noguchi, Shree K. Nayar
ICPR (1)2
1994 Learning, Positioning, and Tracking Visual Appearance
abstract
The problem of vision-based robot positioning and tracking is addressed. A general learning algorithm is presented for determining the mapping between robot position and object appearance. The robot is first moved through several displacements with respect to its desired position, and a large set of object images is acquired. This image set is compressed using principal component analysis to obtain a four-dimensional subspace. Variations in object images due to robot displacements are represented as a compact parametrized manifold in the subspace. While positioning or tracking, errors in end-effector coordinates are efficiently computed from a single brightness image using the parametric manifold representation. The learning component enables accurate visual control without any prior hand-eye calibration. Several experiments have been conducted to demonstrate the practical feasibility of the proposed positioning/tracking approach and its relevance to industrial applications.>
Shree K. Nayar, Hiroshi Murase, Sameer A. Nene
ICRA1
1994 Generalization of Lambert's reflectance model
abstract
Lambert's model for body reflection is widely used in computer graphics. It is used extensively by rendering techniques such as radiosity and ray tracing. For several real-world objects, however, Lambert's model can prove to be a very inaccurate approximation to the body reflectance. While the brightness of a Lambertian surface is independent of viewing direction, that of a rough surface increases as the viewing direction approaches the light source direction. In this paper, a comprehensive model is developed that predicts body reflectance from rough surfaces. The surface is modeled as a collection of Lambertian facets. It is shown that such a surface is inherently non-Lambertian due to the foreshortening of the surface facets. Further, the model accounts for complex geometric and radiometric phenomena such as masking, shadowing, and interreflections between facets. Several experiments have been conducted on samples of rough diffuse surfaces, such as, plaster, sand, clay, and cloth. All these surface demonstrate significant deviation from Lambertian behavior. The reflectance measurements obtained are in strong agreement with the reflectance predicted by the model.
Michael Oren, Shree K. Nayar
SIGGRAPH2
1994 Illumination Planning for Object Recognition Using Parametric Eigenspaces
abstract
Presents a novel approach to the problem of illumination planning for robust object recognition in structured environments. Given a set of objects, the goal is to determine the illumination for which the objects are most distinguishable in appearance from each other. Correlation is used as a measure of similarity between objects. For each object, a large number of images is automatically obtained by varying the pose and the illumination direction. Images of all objects together constitute the planning image set. The planning set is compressed using the Karhunen-Loeve transform to obtain a low-dimensional subspace, called the eigenspace. For each illumination direction, objects are represented as parametrized manifolds in the eigenspace. The minimum distance between the manifolds of two objects represents the similarity between the objects in the correlation sense. The optimal source direction is therefore the one that maximizes the shortest distance between the object manifolds. Several experiments have been conducted using real objects. The results produced by the illumination planner have been used to enhance the performance of an object recognition system.>
Hiroshi Murase, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.2
1994 Shape from Focus
abstract
The shape from focus method presented here uses different focus levels to obtain a sequence of object images. The sum-modified-Laplacian (SML) operator is developed to provide local measures of the quality of image focus. The operator is applied to the image sequence to determine a set of focus measures at each image point. A depth estimation algorithm interpolates a small number of focus measure values to obtain accurate depth estimates. A fully automated shape from focus system has been implemented using an optical microscope and tested on a variety of industrial samples. Experimental results are presented that demonstrate the accuracy and robustness of the proposed method. These results suggest shape from focus to be an effective approach for a variety of challenging visual inspection tasks.>
Shree K. Nayar, Yasuo Nakagawa
IEEE Trans. Pattern Anal. Mach. Intell.1
1993 Learning Object Models from Appearance
Hiroshi Murase, Shree K. Nayar
AAAI2
1993 Removal of specularities using color and polarization
abstract
An algorithm for separating the specular and diffuse components of reflection from images is presented. The method uses color and polarization simultaneously to obtain strong constraints on the reflection components at each image point. Polarization is used to locally determine the color of the specular component, constraining the diffuse color at a pixel to a one-dimensional linear subspace. This subspace is used to find neighboring pixels whose color is consistent with the pixel. Diffuse color information from consistent neighbors is used to determine the diffuse color of the pixel. In contrast to previous separation algorithms, the proposed method can handle highlights that have a varying diffuse component, as well as highlights that include regions with different reflectance and material properties. Experimental results obtained by applying the algorithm to complex scenes with textured objects and strong interreflections are presented.>
Shree K. Nayar, Xi-Sheng Fang, Terrance E. Boult
CVPR1
1993 Diffuse reflectance from rough surfaces
abstract
A comprehensive model that predicts reflectance from rough diffuse surfaces is presented. It is shown that diffuse reflectance from rough surfaces increases as the viewing direction approaches the source direction. This is in contrast to Lambertian surfaces, where radiance is independent of the viewing direction. The new model is a generalization of the Lambertian model, and has significant implications for machine vision, graphics, and remote sensing.>
Michael Oren, Shree K. Nayar
CVPR2
1993 Reflectance ratio: A photometric invariant for object recognition
abstract
Neighboring points on a smoothly curved surface have similar surface orientations and illumination conditions. Hence, their brightness values can be used to compute the ratio of their reflectance coefficients. Based on this observation, an efficient algorithm is developed that estimates a reflectance ratio for each region in an image with respect to its background. The region reflectance ratio represents a physical property of a region that is invariant to the illumination conditions. The ratio invariant is used to recognize objects from a single brightness image of a scene. Experimental results demonstrate the power of using reflectance and geometric properties of objects simultaneously.>
Shree K. Nayar, Ruud M. Bolle
ICCV1
1993 Computing reflectance ratios from an image
Shree K. Nayar, Ruud M. Bolle
Pattern Recognit.1
1992 Extracting 3-D structure and focused images using an optical microscope
abstract
Describes a novel method for recovering three-dimensional shapes and focused images of microscopic objects. A small sequence of object images is obtained by translating the object through the focused plane of the microscope. The sum-modified-Laplacian (SML) operator is developed to compute local measures of the quality of image focus. The SML operator is applied to the image sequence, and the focus measures obtained at each image point are used to compute local depth estimates. Two algorithms for depth estimation are presented. The first algorithm simply looks for the focus level that maximizes the focus measure at each image point. The second algorithm uses a model to interpolate the focus measures to obtain more accurate depth estimates. The algorithms were implemented and tested on a variety of biological specimens. The authors present a brief description of a fully automated shape-from-focus system that is currently being used to inspect both medical and industrial samples.>
Ushir B. Shah, Shree K. Nayar
CBMS2
1992 Shape from focus system
abstract
A shape-from-focus method that uses different focus levels to obtain a sequence of object images is described. A sum-modified-Laplacian operator is developed to provide local measures of the quality of image focus. The operator is applied to the sequence of images of the object to determine a set of focus measures as each image point. A model is developed to describe the variation of focus measure values due to defocusing. This model is used by a depth estimation algorithm to interpolate focus measure values and obtain accurate depth estimates. A fully automated system that has been implemented using an optical microscope and tested on a variety of industrial samples is described.>
Shree K. Nayar
CVPR1
1992 Shape recovery methods for visual inspection
abstract
The advancement of three-dimensional machine vision is closely related to the development of robust and efficient shape recovery methods. The author addresses the recovery problem associated with three different classes of surfaces: (a) specular surfaces; (b) surfaces with varying reflectance; and (c) rough and textured surfaces. Three realtime machine vision systems have been developed based on these results. Experimental results demonstrate that the proposed methods and systems are applicable to a variety of visual inspection problems.>
Shree K. Nayar
WACV1
1991 Recovering shape in the presence of interreflections
abstract
An algorithm for recovering the shape and reflectance of Lambertian surfaces in the presence of interreflections is presented. The surfaces may be of arbitrary but continuous shape, and with possibly varying and unknown reflectance. The actual shape and reflectance are recovered from the pseudoshape and pseudoreflectance estimated by a local shape-from-intensity method (e.g., photometric stereo). Thus, the algorithm enhances the performance and the utility of existing shape-from-intensity methods. From the results reported, two observations can be made that are pertinent to machine vision: interreflections can cause vision algorithms to produce unacceptably erroneous results and hence should not be ignored; and at least some interreflection problems are tractable and solvable.>
Shree K. Nayar, Katsushi Ikeuchi, Takeo Kanade
ICRA1
1991 Shape from interreflections
Shree K. Nayar, Katsushi Ikeuchi, Takeo Kanade
Int. J. Comput. Vis.1
1991 Surface Reflection: Physical and Geometrical Perspectives
abstract
Reflectance models based on physical optics and geometrical optics are studied. Specifically, the authors consider the Beckmann-Spizzichino (physical optics) model and the Torrance-Sparrow (geometrical optics) model. These two models were chosen because they have been reported to fit experimental data well. Each model is described in detail, and the conditions that determine the validity of the model are clearly stated. By studying reflectance curves predicted by the two models, the authors propose a reflectance framework comprising three components: the diffuse lobe, the specular lobe, and the specular spike. The effects of surface roughness on the three primary components are analyzed in detail.>
Shree K. Nayar, Katsushi Ikeuchi, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.1
1990 Shape from interreflections
abstract
An iterative algorithm is presented that simultaneously recovers the actual shape and the actual reflectance from the pseudo estimates. The recovery algorithm works on Lambertian surfaces of arbitrary shape with possibly varying and unknown reflectance. The general behavior of the algorithm and its convergence properties are discussed. Both simulation and experimental results are included to demonstrate the accuracy and stability of the algorithm.>
Shree K. Nayar, Katsushi Ikeuchi, Takeo Kanade
ICCV1
1990 Shape from focus: an effective approach for rough surfaces
abstract
Two algorithms for depth estimation are presented. The first algorithm simply looks for the focus level that maximizes the focus measure at each image point. The second algorithm uses a Gaussian model to interpolate the focus measures to obtain more accurate depth estimates. The algorithms were implemented and tested on surfaces of different roughness and reflectance properties. The results indicate that the shape-from-focus method may be applied to a variety of industrial vision problems.>
Shree K. Nayar, Yasuo Nakagawa
ICRA1
1990 Determining shape and reflectance of hybrid surfaces by photometric sampling
abstract
A method is presented for determining the shapes of hybrid surfaces without prior knowledge of the relative strengths of the Lambertian and specular components of reflection. The object surface is illuminated using extended light sources and is viewed from a single direction. Surface illumination using extended sources makes it possible to ensure the detection of both Lambertian and specular reflections. Uniformly distributed source directions are used to obtain an image sequence of the object. This method of obtaining photometric measurements is called photometric sampling. An extraction algorithm uses the set of image intensity values measured at each surface point to compute orientation as well as relative strengths of the Lambertian and specular reflection components. The simultaneous recovery of shape and reflectance parameters enables the method to adapt to variations in reflectance properties from one scene point to another. Experiments were conducted on Lambertian surfaces, specular surfaces, and hybrid surfaces.>
Shree K. Nayar, Katsushi Ikeuchi, Takeo Kanade
IEEE Trans. Robotics Autom.1
1990 Specular surface inspection using structured highlight and Gaussian images
abstract
The structured highlight inspection method uses an array of point sources to illuminate a specular object surface. The point sources are scanned, and highlights on the object surface resulting from each source are used to derive local surface orientation information. The extended Gaussian image (EGI) is obtained by placing at each point on a Gaussian sphere a mass proportional to the area of elements on the object surface that have a specific orientation. The EGI summarizes shape properties of the object surface and can be efficiently calculated from structured highlight data without surface reconstruction. Features of the estimated EGI including areas, moments, principal axes, homogeneity measures, and polygonality can be used as the basis for classification and inspection. The structured highlight inspection system (SHINY) has been implemented using a hemisphere of 127 point sources. The SHINY system uses a binary coding scheme to make the scanning of point sources efficient. Experiments have used the SHINY system and EGI features for the inspection and classification of surface-mounted-solder joints.>
Shree K. Nayar, Arthur C. Sanderson, Lee E. Weiss, David A. Simon
IEEE Trans. Robotics Autom.1
1989 Shape and reflectance from an image sequence generated using extended sources
abstract
The authors present a method for determining the shapes of surfaces whose reflectance properties may vary from Lambertian to specular, without prior knowledge of the relative strengths of the Lambertian and specular components of reflection. The object surface is illuminated using extended light sources and is viewed from a single direction. Surface illumination using extended sources makes it possible to ensure the detection of both Lambertian and specular reflections. Multiple source directions are used to obtain an image sequence of the object. An extraction algorithm uses the set of image intensity values measured at each surface point to compute orientation as well as relative strengths of the Lambertian and specular reflection components. The proposed method is called photometric sampling, as it uses samples of photometric function that relates image intensity to surface orientation, reflectance, and light source characteristics. Experiments were conducted on Lambertian surfaces, specular surfaces, and hybrid surfaces, whose reflectance models are composed of both Lambertian and specular components. The results show high accuracy in measured orientations and estimated reflectance parameters.>
Shree K. Nayar, Katsushi Ikeuchi, Takeo Kanade
ICRA1
1988 Structured Highlight Inspection of Specular Surfaces
abstract
An approach to illumination and imaging of specular surfaces that yields three-dimensional shape information is described. The structured highlight approach uses a scanned array of point sources and images of the resulting reflected highlights to compute local surface height and orientation. A prototype structured highlight inspection system, called SHINY, has been implemented. SHINY demonstrates the determination of surface shape for several test objects including solder joints. The current SHINY system makes the distant-source assumption and requires only one camera. A stereo structured highlight system using two cameras is proposed to determine surface-element orientation for objects in a much larger field of view. Analysis and description of the algorithms are included. The proposed structured highlight techniques are promising for many industrial tasks.>
Arthur C. Sanderson, Lee E. Weiss, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.3