VLDB 2026 Research / reviewers in the wild / expert
Srinivasa G. Narasimhan
dblp:57/2011
· DBLP profile ↗
117ranked-venue papers
11as first author
29since 2021 · last 2025
0000-0003-0389-1921ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 93 · 8 first-author · 24 since 2021Artificial intelligence and machine learning · 88 · 10 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Incorporating Dense Metric Depth into Neural 3D Representations for View Synthesis and RelightingabstractCapturing photo realistic appearance and geometry of scenes is a fundamental problem in computer vision and graphics with a set of mature tools and solutions for content creation [12], [67], large scale scene mapping [5], augmented reality and cinematography [6], [80], [97]. Enthusiast level 3D photogrammetry, especially for small or tabletop scenes, has been supercharged by more capable smartphone cameras and new toolboxes like RealityCapture and NeRF-Studio. A subset of these solutions are geared towards view synthesis where the focus is on photo-realistic view interpolation rather than recovery of accurate scene geometry. These solutions take the “shape-radiance ambiguity”[58] into stride by decoupling the scene transmissivity (related to geometry) from the scene appearance prediction. But without diverse training views, several neural scene representations (e.g. [39], [73], [76]) are prone to poor shape reconstructions while estimating accurate appearance. Arkadeep Narayan Chaudhury, Igor Vasiljevic, Sergey Zakharov, Vitor Campagnolo Guizilini, Rares Ambrus, Srinivasa G. Narasimhan, Christopher G. Atkeson |
3DV | 6 |
| 2025 | AerialMegaDepth: Learning Aerial-Ground Reconstruction and View SynthesisabstractWe explore the task of geometric reconstruction of images captured from a mixture of ground and aerial views. Current state-of-the-art learning-based approaches fail to handle the extreme viewpoint variation between aerial-ground image pairs. Our hypothesis is that the lack of high-quality, co-registered aerial-ground datasets for training is a key reason for this failure. Such data is difficult to assemble precisely because it is difficult to reconstruct in a scalable way. To overcome this challenge, we propose a scalable framework combining pseudo-synthetic renderings from 3D city-wide meshes (e.g., Google Earth) with real, ground-level crowd-sourced images (e.g., MegaDepth [29]). The pseudo-synthetic data simulates a wide range of aerial viewpoints, while the real, crowd-sourced images help improve visual fidelity for ground-level images where mesh-based renderings lack sufficient detail, effectively bridging the domain gap between real images and pseudo-synthetic renderings. Using this hybrid dataset, we fine-tune several state-of-the-art algorithms and achieve significant improvements on real-world, zero-shot aerial-ground tasks. For example, we observe that baseline DUSt3R [64] localizes fewer than 5% of aerial-ground pairs within 5 degrees of camera rotation error, while fine-tuning with our data raises accuracy to nearly 56%, addressing a major failure point in handling large viewpoint changes. Beyond camera estimation and scene reconstruction, our dataset also improves performance on downstream tasks like novel-view synthesis in challenging aerial-ground scenarios, demonstrating the practical value of our approach in real-world applications. Khiem Vuong, Anurag Ghosh, Deva Ramanan, Srinivasa G. Narasimhan, Shubham Tulsiani |
CVPR | 4 |
| 2025 | Resolving Shape Ambiguities using Heat Conduction and ShadingabstractShape from shading using a single image of a Lambertian surface is inherently ambiguous. When the light source direction is known, the surface normal estimation has a cone-ambiguity, which worsens when the source is unknown. Recently, shape from heat conduction has emerged as an approach that leverages heat transport equations to estimate the Shape Laplacian operator, an intrinsic measure of shape. However, deriving surface normals from the Laplacian operator encounters a local binary convex/concave ambiguity. Our contribution introduces a novel theory to resolve these local shape ambiguities (excluding a few degeneracies) without relying on priors like smoothness, by combining the cues from shading and heat conduction. Our method ensures the mathematical constraints of both shading and the Laplacian are satisfied simultaneously, even with an unknown light source. We validate our theory through simulations of complex shapes and analyze its performance in the presence of noise. Index Terms-Shape Reconstruction, Heat Conduction, Concave/convex Ambiguity, Thermal Video Akihiko Oharazawa, Sriram Narayanan, Manikandasriram Srinivasan Ramanagopal, Srinivasa G. Narasimhan |
ICCP | 4 |
| 2025 | ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones
Anurag Ghosh, Robert Tamburo, Khiem Vuong, Juan R. Alvarez-Padilla, Hailiang Zhu, Michael Cardei, Nicholas Dunn, Christoph Mertz, Srinivasa G. Narasimhan |
ICCV | 10 |
| 2025 | Object-level Visual Prompts for Compositional Image GenerationabstractWe introduce a method for composing object-level visual prompts within a text-to-image diffusion model. Our approach addresses the task of generating semantically coherent compositions across diverse scenes and styles, similar to the versatility and expressiveness offered by text prompts. A key challenge in this task is to preserve the identity of the objects depicted in the input visual prompts, while also generating diverse compositions across different images. To address this challenge, we introduce a new KV-mixed cross-attention mechanism, in which keys and values are learned from distinct visual representations. The keys are derived from an encoder with a small bottleneck for layout control, whereas the values come from a larger bottleneck encoder that captures fine-grained appearance details. By mixing keys and values from these complementary sources, our model preserves the identity of the visual prompts while supporting flexible variations in object arrangement, pose, and composition. During inference, we further propose object-level compositional guidance to improve the method’s identity preservation and layout correctness. Results show that our technique produces diverse scene compositions that preserve the unique characteristics of each visual prompt, expanding the creative potential of text-to-image generation. Gaurav Parmar, Or Patashnik, Kuan-Chieh Wang, Daniil Ostashev, Srinivasa G. Narasimhan, Jun-Yan Zhu, Daniel Cohen-Or, Kfir Aberman |
SIGGRAPH Asia | 5 |
| 2025 | Instance-Warp: Saliency Guided Image Warping for Unsupervised Domain Adaptation
Anurag Ghosh, Srinivasa G. Narasimhan |
WACV | 3 |
| 2025 | Dual-Shutter Optical Vibration SensingabstractVisual vibrometry is a highly useful tool for remote capture of audio, as well as the physical properties of materials, human heart rate, and more. While visually-observable vibrations can be captured directly with a high-speed camera, minute imperceptible object vibrations can be optically amplified by imaging the displacement of a speckle pattern created by shining a laser beam on the vibrating surface. In this paper, we propose a novel method for sensing vibrations at high speeds (up to 63 kHz), for multiple scene sources at once, using sensors rated for only 130 Hz operation. Our method relies on simultaneously capturing the scene with two cameras equipped with rolling and global shutter sensors, respectively. The rolling shutter camera captures distorted speckle images that encode the high-speed object vibrations. The global shutter camera captures undistorted reference images of the speckle pattern, helping to decode the source vibrations. We demonstrate our method by capturing vibration caused by audio sources (e.g., speakers, human voice, and musical instruments) and analyzing the vibration modes of a tuning fork. Mark Sheinin, Dorian Chan, Matthew O'Toole, Srinivasa G. Narasimhan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | A Theory of Joint Light and Heat Transport for Lambertian ScenesabstractWe present a novel theory that establishes the relation-ship between light transport in visible and thermal infrared, and heat transport in solids. We show that heat generated due to light absorption can be estimated by modeling heat transport using a thermal camera. For situations where heat conduction is negligible, we analytically solve the heat transport equation to derive a simple expression relating the change in thermal image intensity to the absorbed light intensity and heat capacity of the material. Next, we prove that intrinsic image decomposition for Lambertian scenes becomes a well-posed problem if one has access to the ab-sorbed light. Our theory generalizes to arbitrary shapes and unstructured illumination. Our theory is based on ap-plying energy conservation principle at each pixel indepen-dently. We validate our theory using real-world experi-ments on diffuse objects made of different materials that ex-hibit both direct and global components (inter-reflections) of light transport under unknown complex lighting. Manikandasriram Srinivasan Ramanagopal, Sriram Narayanan, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan |
CVPR | 4 |
| 2024 | Projecting Trackable Thermal Patterns for Dynamic Computer VisionabstractAdding artificial patterns to objects, like QR codes, can ease tasks such as object tracking, robot navigation, and conveying information (e.g., a label or a website link). However, these patterns require a physical application and they alter the object's appearance. Conversely, projected patterns can temporarily change the object's appearance, aiding tasks like 3D scanning and retrieving object textures and shading. However, projected patterns impede dynamic tasks like object tracking because they do not ‘stick’ to the object's surface. Or do they? This paper introduces a novel approach combining the advantages of projected and persistent physical patterns. Our system projects heat patterns using a laser beam (similar in spirit to a LIDAR), which a thermal camera observes and tracks. Such thermal patterns enable tracking poorly-textured objects whose tracking is highly challenging with standard cameras while not affecting the object's appearance or physical properties. To avail these thermal patterns in existing vision frameworks, we train a network to reverse heat diffusion's effects and remove inconsistent pattern points between different thermal frames. We prototyped and tested this approach on dynamic vision tasks like structure from motion, optical flow, and object tracking of everyday textureless objects. Mark Sheinin, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan |
CVPR | 3 |
| 2024 | WALT3D: Generating Realistic Training Data from Time-Lapse Imagery for Reconstructing Dynamic Objects Under OcclusionabstractCurrent methods for 2D and 3D object understanding struggle with severe occlusions in busy urban environments, partly due to the lack of large-scale labeled ground-truth annotations for learning occlusion. In this work, we introduce a novel framework for automatically generating a large, realistic dataset of dynamic objects under occlusions using freely available time-lapse imagery. By leveraging off-the-shelf2D (bounding box, segmentation, keypoint) and 3D (pose, shape) predictions as pseudo-groundtruth, unoccluded 3D objects are identified automatically and composited into the background in a clip-art style, ensuring realistic appearances and physically accurate occlusion configurations. The resulting clip-art image with pseudogroundtruth enables efficient training of object reconstruction methods that are robust to occlusions. Our method demonstrates significant improvements in both 2D and 3D reconstruction, particularly in scenarios with heavily occluded objects like vehicles and people in urban scenes. Khiem Vuong, N. Dinesh Reddy, Robert Tamburo, Srinivasa G. Narasimhan |
CVPR | 4 |
| 2024 | Shape from Heat Conduction
Sriram Narayanan, Manikandasriram Srinivasan Ramanagopal, Mark Sheinin, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan |
ECCV (38) | 5 |
| 2024 | Toward Planet-Wide Traffic Camera CalibrationabstractDespite the widespread deployment of outdoor cameras, their potential for automated analysis remains largely untapped due, in part, to calibration challenges. The absence of precise camera calibration data, including intrinsic and extrinsic parameters, hinders accurate real-world distance measurements from captured videos. To address this, we present a scalable framework that utilizes street-level imagery to reconstruct a metric 3D model, facilitating precise calibration of in-the-wild traffic cameras. Notably, our framework achieves 3D scene reconstruction and accurate localization of over 100 global traffic cameras and is scalable to any camera with sufficient street-level imagery. For evaluation, we introduce a dataset of 20 fully calibrated traffic cameras, demonstrating our method’s significant enhancements over existing automatic calibration techniques. Furthermore, we highlight our approach’s utility in traffic analysis by extracting insights via 3D vehicle reconstruction and speed measurement, thereby opening up the potential of using outdoor cameras for automated analysis. Code and dataset will be available on the project website. Khiem Vuong, Robert Tamburo, Srinivasa G. Narasimhan |
WACV | 3 |
| 2024 | TPSeNCE: Towards Artifact-Free Realistic Rain Generation for Deraining and Object Detection in RainabstractRain generation algorithms have the potential to improve the generalization of deraining methods and scene understanding in rainy conditions. However, in practice, they produce artifacts and distortions and struggle to control the amount of rain generated due to a lack of proper constraints. In this paper, we propose an unpaired image-to-image translation framework for generating realistic rainy images. We first introduce a Triangular Probability Similarity (TPS) constraint to guide the generated images toward clear and rainy images in the discriminator manifold, thereby minimizing artifacts and distortions during rain generation. Unlike conventional contrastive learning approaches, which indiscriminately push negative samples away from the anchors, we propose a Semantic Noise Contrastive Estimation (SeNCE) strategy and reassess the pushing force of negative samples based on the semantic similarity between the clear and the rainy images and the feature similarity between the anchor and the negative samples. Experiments demonstrate realistic rain generation with minimal artifacts and distortions, which benefits image deraining and object detection in rain. Furthermore, the method can be used to generate realistic snowy and night images, underscoring its potential for broader applicability. Code is available at https://github.com/ShenZheng2000/TPSeNCE. Changjie Lu, Srinivasa G. Narasimhan |
WACV | 3 |
| 2024 | Virtual home staging and relighting from a single panorama under natural illuminationabstractAbstract Virtual staging technique can digitally showcase a variety of real-world scenes. However, relighting indoor scenes from a single image is challenging due to unknown scene geometry, material properties, and outdoor spatially-varying lighting. In this study, we use the High Dynamic Range (HDR) technique to capture an indoor panorama and its paired outdoor hemispherical photograph, and we develop a novel inverse rendering approach for scene relighting and editing. Our method consists of four key components: (1) panoramic furniture detection and removal, (2) automatic floor layout design, (3) global rendering with scene geometry, new furniture objects, and the real-time outdoor photograph, and (4) virtual staging with new camera position, outdoor illumination, scene texture, and electrical light. The results demonstrate that a single indoor panorama can be used to generate high-quality virtual scenes under new environmental conditions. Additionally, we contribute a new calibrated HDR (Cali-HDR) dataset that consists of 137 paired indoor and outdoor photographs. The animation for virtual rendered scenes is available here . Guanzhou Ji, Azadeh O. Sawyer, Srinivasa G. Narasimhan |
Mach. Vis. Appl. | 3 |
| 2023 | Learned Two-Plane Perspective Prior based Image Resampling for Efficient Object DetectionabstractReal-time efficient perception is critical for autonomous navigation and city scale sensing. Orthogonal to architectural improvements, streaming perception approaches have exploited adaptive sampling improving real-time detection performance. In this work, we propose a learnable geometry-guided prior that incorporates rough geometry of the 3D scene (a ground plane and a plane above) to resample images for efficient object detection. This significantly improves small and far-away object detection performance while also being more efficient both in terms of latency and memory. For autonomous navigation, using the same detector and scale, our approach improves detection rate by +4.1 APsor +39% and in real-time performance by +5.3 sAPs or +63% for small objects over state-of-the-art (SOTA). For fixed traffic cameras, our approach detects small objects at image scales other methods cannot. At the same scale, our approach improves detection of small objects by 195% (+12.5 APS) over naive-downsampling and 63% (+4.2 APS) over SOTA. Anurag Ghosh, N. Dinesh Reddy, Christoph Mertz, Srinivasa G. Narasimhan |
CVPR | 4 |
| 2023 | Megahertz Light Steering Without Moving PartsabstractWe introduce a light steering technology that operates at megahertz frequencies, has no moving parts, and costs less than a hundred dollars. Our technology can benefit many projector and imaging systems that critically rely on high-speed, reliable, low-cost, and wavelength-independent light steering, including laser scanning projectors, LiDAR sensors, and fluorescence microscopes. Our technology uses ultrasound waves to generate a spatiotemporally-varying refractive index field inside a compressible medium, such as water, turning the medium into a dynamic traveling lens. By controlling the electrical input of the ultrasound transducers that generate the waves, we can change the lens, and thus steer light, at the speed of sound (1.5 km/s in water). We build a physical prototype of this technology, use it to realize different scanning techniques at megahertz rates (three orders of magnitude faster than commercial alternatives such as galvo mirror scanners), and demonstrate proof-of-concept projector and LiDAR applications. To encourage further innovation towards this new technology, we derive theory for its fundamental limits and develop a physically-accurate simulator for virtual design. Our technology offers a promising solution for achieving high-speed and low-cost light steering in a variety of applications. Adithya Kumar Pediredla, Srinivasa G. Narasimhan, Maysamreza Chamanzar, Ioannis Gkioulekas |
CVPR | 2 |
| 2023 | Analyzing Physical Impacts Using Transient Surface Wave ImagingabstractThe subtle vibrations on an object's surface contain information about the object's physical properties and its interaction with the environment. Prior works imaged surface vibration to recover the object's material properties via modal analysis, which discards the transient vibrations propagating immediately after the object is disturbed. Conversely, prior works that captured transient vibrations focused on recovering localized signals (e.g., recording nearby sound sources), neglecting the spatiotemporal relationship between vibrations at different object points. In this paper, we extract information from the transient surface vibrations simultaneously measured at a sparse set of object points using the dual-shutter camera described by Sheinin et al. [37]. We model the geometry of an elastic wave generated at the moment an object's surface is disturbed (e.g., a knock or a footstep) and use the model to localize the disturbance source for various materials (e.g., wood, plastic, tile). We also show that transient object vibrations contain additional cues about the impact force and the impacting object's material properties. We demonstrate our approach in applications like localizing the strikes of a ping-pong ball on a table mid-play and recovering the footsteps' locations by imaging the floor vibrations they create. Mark Sheinin, Dorian Chan, Mark Rau, Matthew O'Toole, Srinivasa G. Narasimhan |
CVPR | 6 |
| 2022 | Holocurtains: Programming Light Curtains via Binary HolographyabstractLight curtain systems are designed for detecting the presence of objects within a user-defined 3D region of space, which has many applications across vision and robotics. However, the shape of light curtains have so far been limited to ruled surfaces, i.e., surfaces composed of straight lines. In this work, we propose Holocurtains: a light-efficient approach to producing light curtains of arbitrary shape. The key idea is to synchronize a rolling-shutter camera with a 2D holographic projector, which steers (rather than block) light to generate bright structured light patterns. Our prototype projector uses a binary digital micromirror device (DMD) to generate the holographic interference patterns at high speeds. Our system produces 3D light curtains that cannot be achieved with traditional light curtain setups and thus enables all-new applications, including the ability to simultaneously capture multiple light curtains in a single frame, detect subtle changes in scene geometry, and transform any 3D surface into an optical touch interface. Dorian Chan, Srinivasa G. Narasimhan, Matthew O'Toole |
CVPR | 2 |
| 2022 | WALT: Watch And Learn 2D amodal representation from Time-lapse imageryabstractCurrent methods for object detection, segmentation, and tracking fail in the presence of severe occlusions in busy urban environments. Labeled real data of occlusions is scarce (even in large datasets) and synthetic data leaves a domain gap, making it hard to explicitly model and learn occlusions. In this work, we present the best of both the real and synthetic worlds for automatic occlusion supervision using a large readily available source of data: time-lapse imagery from stationary webcams observing street intersections over weeks, months, or even years. We introduce a new dataset, Watch and Learn Time-lapse (WALT), consisting of 12 (4K and 1080p) cameras capturing urban environments over a year. We exploit this real data in a novel way to automatically mine a large set of unoccluded objects and then composite them in the same views to generate occlusions. This longitudinal self-supervision is strong enough for an amodal network to learn object-occluder-occluded layer representations. We show how to speed up the discovery of unoccluded objects and relate the confidence in this discovery to the rate and accuracy of training occluded objects. After watching and automatically learning for several days, this approach shows significant performance improvement in detecting and segmenting occluded people and vehicles, over human-supervised amodal approaches. N. Dinesh Reddy, Robert Tamburo, Srinivasa G. Narasimhan |
CVPR | 3 |
| 2022 | Dual-Shutter Optical Vibration SensingabstractVisual vibrometry is a highly useful tool for remote capture of audio, as well as the physical properties of materials, human heart rate, and more. While visually-observable vibrations can be captured directly with a high-speed camera, minute imperceptible object vibrations can be optically amplified by imaging the displacement of a speckle pattern, created by shining a laser beam on the vibrating surface. In this paper, we propose a novel method for sensing vibrations at high speeds (up to 63kHz), for multiple scene sources at once, using sensors rated for only 130Hz operation. Our method relies on simultaneously capturing the scene with two cameras equipped with rolling and global shutter sensors, respectively. The rolling shutter camera captures distorted speckle images that encode the high-speed object vibrations. The global shutter camera captures undistorted reference images of the speckle pattern, helping to decode the source vibrations. We demonstrate our method by capturing vibration caused by audio sources (e.g. speakers, human voice, and musical instruments) and analyzing the vibration modes of a tuning fork. Mark Sheinin, Dorian Chan, Matthew O'Toole, Srinivasa G. Narasimhan |
CVPR | 4 |
| 2022 | Learning Continuous Implicit Representation for Near-Periodic Patterns
Bowei Chen 0004, Tiancheng Zhi, Martial Hebert, Srinivasa G. Narasimhan |
ECCV (15) | 4 |
| 2022 | Spatiotemporal Bundle Adjustment for Dynamic 3D Human Reconstruction in the WildabstractBundle adjustment jointly optimizes camera intrinsics and extrinsics and 3D point triangulation to reconstruct a static scene. The triangulation constraint, however, is invalid for moving points captured in multiple unsynchronized videos and bundle adjustment is not designed to estimate the temporal alignment between cameras. We present a spatiotemporal bundle adjustment framework that jointly optimizes four coupled sub-problems: estimating camera intrinsics and extrinsics, triangulating static 3D points, as well as sub-frame temporal alignment between cameras and computing 3D trajectories of dynamic points. Key to our joint optimization is the careful integration of physics-based motion priors within the reconstruction pipeline, validated on a large motion capture corpus of human subjects. We devise an incremental reconstruction and alignment algorithm to strictly enforce the motion prior during the spatiotemporal bundle adjustment. This algorithm is further made more efficient by a divide and conquer scheme while still maintaining high accuracy. We apply this algorithm to reconstruct 3D motion trajectories of human bodies in dynamic events captured by multiple uncalibrated and unsynchronized video cameras in the wild. To make the reconstruction visually more interpretable, we fit a statistical 3D human body model to the asynchronous video streams. Compared to the baseline, the fitting significantly benefits from the proposed spatiotemporal bundle adjustment procedure. Because the videos are aligned with sub-frame precision, we reconstruct 3D motion at much higher temporal resolution than the input videos. Website: http://www.cs.cmu.edu/~ILIM/projects/IM/STBA. Minh Vo, Yaser Sheikh, Srinivasa G. Narasimhan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Semantically supervised appearance decomposition for virtual staging from a single panoramaabstractWe describe a novel approach to decompose a single panorama of an empty indoor environment into four appearance components: specular, direct sunlight, diffuse and diffuse ambient without direct sunlight. Our system is weakly supervised by automatically generated semantic maps (with floor, wall, ceiling, lamp, window and door labels) that have shown success on perspective views and are trained for panoramas using transfer learning without any further annotations. A GAN-based approach supervised by coarse information obtained from the semantic map extracts specular reflection and direct sunlight regions on the floor and walls. These lighting effects are removed via a similar GAN-based approach and a semantic-aware inpainting step. The appearance decomposition enables multiple applications including sun direction estimation, virtual furniture insertion, floor material replacement, and sun direction change, providing an effective tool for virtual home staging. We demonstrate the effectiveness of our approach on a large and recently released dataset of panoramas of empty homes. Tiancheng Zhi, Bowei Chen 0004, Ivaylo Boyadzhiev, Sing Bing Kang, Martial Hebert, Srinivasa G. Narasimhan |
ACM Trans. Graph. | 6 |
| 2021 | Exploiting & Refining Depth Distributions With Triangulation Light CurtainsabstractActive sensing through the use of Adaptive Depth Sensors is a nascent field, with potential in areas such as Advanced driver-assistance systems (ADAS). They do however require dynamically driving a laser / light-source to a specific location to capture information, with one such class of sensor being the Triangulation Light Curtains (LC). In this work, we introduce a novel approach that exploits prior depth distributions from RGB cameras to drive a Light Curtain’s laser line to regions of uncertainty to get new measurements. These measurements are utilized such that depth uncertainty is reduced and errors get corrected recursively. We show real-world experiments that validate our approach in outdoor and driving settings, and demonstrate qualitative and quantitative improvements in depth RMSE when RGB cameras are used in tandem with a Light Curtain. Yaadhav Raaj, Siddharth Ancha, Robert Tamburo, David Held, Srinivasa G. Narasimhan |
CVPR | 5 |
| 2021 | TesseTrack: End-to-End Learnable Multi-Person Articulated 3D Pose TrackingabstractWe consider the task of 3D pose estimation and tracking of multiple people seen in an arbitrary number of camera feeds. We propose TesseTrack1, a novel top-down approach that simultaneously reasons about multiple individuals’ 3D body joint reconstructions and associations in space and time in a single end-to-end learnable framework. At the core of our approach is a novel spatio-temporal formulation that operates in a common voxelized feature space aggregated from single- or multiple camera views. After a person detection step, a 4D CNN produces short-term person-specific representations which are then linked across time by a differentiable matcher. The linked descriptions are then merged and deconvolved into 3D poses. This joint spatio-temporal formulation contrasts with previous piecewise strategies that treat 2D pose estimation, 2D-to-3D lifting, and 3D pose tracking as independent sub-problems that are error-prone when solved in isolation. Furthermore, unlike previous methods, TesseTrack is robust to changes in the number of camera views and achieves very good results even if a single view is available at inference time. Quantitative evaluation of 3D pose reconstruction accuracy on standard benchmarks shows significant improvements over the state of the art. Evaluation of multi-person articulated 3D pose tracking in our novel evaluation framework demonstrates the superiority of TesseTrack over strong baselines. N. Dinesh Reddy, Laurent Guigues, Leonid Pishchulin, Jayan Eledath, Srinivasa G. Narasimhan |
CVPR | 5 |
| 2021 | Deconvolving Diffraction for Fast Imaging of Sparse ScenesabstractMost computer vision techniques rely on cameras which uniformly sample the 2D image plane. However, there exists a class of applications for which the standard uniform 2D sampling of the image plane is sub-optimal. This class consists of applications where the scene points of interest occupy the image plane sparsely (e.g., marker-based motion capture), and thus most pixels of the 2D camera sensor would be wasted. Recently, diffractive optics were used in conjunction with sparse (e.g., line) sensors to achieve high-speed capture of such sparse scenes. One such approach, called “Diffraction Line Imaging”, relies on the use of diffraction gratings to spread the point-spread-function (PSF) of scene points from a point to a color-coded shape (e.g., a horizontal line) whose intersection with a line sensor enables point positioning. In this paper, we extend this approach for arbitrary diffractive optical elements and arbitrary sampling of the sensor plane using a convolution-based image formation model. Sparse scenes are then recovered by formulating a convolutional coding inverse problem that can resolve mixtures of diffraction PSFs without the use of multiple sensors, extending the application of diffraction-based imaging to a new class of significantly denser scenes. For the case of a single-axis diffraction grating, we provide an approach to determine the minimal required sensor sub-sampling for accurate scene recovery. Compared to methods that use a speckle PSF from a narrow-band source or a diffuser-based PSF with a rolling shutter sensor, our approach uses spectrally-coded PSFs from broad-band sources and allows arbitrary sensor sampling, respectively. We demonstrate that the presented combination of the imaging approach and scene recovery method is well suited for high-speed marker based motion capture and particle image velocimetry (PIV) over long periods. Mark Sheinin, Matthew O'Toole, Srinivasa G. Narasimhan |
ICCP | 3 |
| 2021 | Traffic4D: Single View Reconstruction of Repetitious Activity Using Longitudinal Self-SupervisionabstractReconstructing 4D vehicular activity (3D space and time) from cameras is useful for autonomous vehicles, commuters and local authorities to plan for smarter and safer cities. Traffic is inherently repetitious over long periods, yet current deep learning-based 3D reconstruction methods have not considered such repetitions and have difficulty generalizing to new intersection-installed cameras. We present a novel approach exploiting longitudinal (long-term) repetitious motion as self-supervision to reconstruct 3D vehicular activity from a video captured by a single fixed camera. Starting from off-the-shelf 2D keypoint detections, our algorithm optimizes 3D vehicle shapes and poses, and then clusters their trajectories in 3D space. The 2D keypoints and trajectory clusters accumulated over long-term are later used to improve the 2D and 3D keypoints via self-supervision without any human annotation. Our method improves reconstruction accuracy over state of the art on scenes with a significant visual difference from the keypoint detector's training data, and has many applications including velocity estimation, anomaly detection and vehicle counting. We demonstrate results on traffic videos captured at multiple city intersections, collected using our smartphones, YouTube, and other public datasets. N. Dinesh Reddy, Srinivasa G. Narasimhan |
IV | 4 |
| 2021 | Self-Supervised Multi-View Person Association and its ApplicationsabstractReliable markerless motion tracking of people participating in a complex group activity from multiple moving cameras is challenging due to frequent occlusions, strong viewpoint and appearance variations, and asynchronous video streams. To solve this problem, reliable association of the same person across distant viewpoints and temporal instances is essential. We present a self-supervised framework to adapt a generic person appearance descriptor to the unlabeled videos by exploiting motion tracking, mutual exclusion constraints, and multi-view geometry. The adapted discriminative descriptor is used in a tracking-by-clustering formulation. We validate the effectiveness of our descriptor learning on WILDTRACK T. Chavdarova et al., "WILDTRACK: A multi-camera HD dataset for dense unscripted pedestrian detection," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2018, pp. 5030-5039. and three new complex social scenes captured by multiple cameras with up to 60 people "in the wild". We report significant improvement in association accuracy (up to 18 percent) and stable and coherent 3D human skeleton tracking (5 to 10 times) over the baseline. Using the reconstructed 3D skeletons, we cut the input videos into a multi-angle video where the image of a specified person is shown from the best visible front-facing camera. Our algorithm detects inter-human occlusion to determine the camera switching moment while still maintaining the flow of the action well. Website: http://www.cs.cmu.edu/~ILIM/projects/IM/Association4Tracking. Minh Vo, Ersin Yumer, Kalyan Sunkavalli, Sunil Hadap, Yaser Sheikh, Srinivasa G. Narasimhan |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Programmable Non-Epipolar Indirect Light Transport: Capture and AnalysisabstractThe decomposition of light transport into direct and global components, diffuse and specular interreflections, and subsurface scattering allows for new visualizations of light in everyday scenes. In particular, indirect light contains a myriad of information about the complex appearance of materials useful for computer vision and inverse rendering applications. In this paper, we present a new imaging technique that captures and analyzes components of indirect light via light transport using a synchronized projector-camera system. The rectified system illuminates the scene with epipolar planes corresponding to projector rows, and we vary two key parameters to capture plane-to-ray light transport between projector row and camera pixel: (1) the offset between projector row and camera row in the rolling shutter (implemented as synchronization delay), and (2) the exposure of the camera row. We describe how this synchronized rolling shutter performs illumination multiplexing, and develop a nonlinear optimization algorithm to demultiplex the resulting 3D light transport operator. Using our system, we are able to capture live short and long-range non-epipolar indirect light transport, disambiguate subsurface scattering, diffuse and specular interreflections, and distinguish materials according to their subsurface scattering properties. In particular, we show the utility of indirect imaging for capturing and analyzing the hidden structure of veins in human skin. Hiroyuki Kubo, Suren Jayasuriya, Takafumi Iwaguchi, Takuya Funatomi, Yasuhiro Mukaigawa, Srinivasa G. Narasimhan |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2020 | 4D Visualization of Dynamic Events From Unconstrained Multi-View VideosabstractWe present a data-driven approach for 4D space-time visualization of dynamic events from videos captured by hand-held multiple cameras. Key to our approach is the use of self-supervised neural networks specific to the scene to compose static and dynamic aspects of an event. Though captured from discrete viewpoints, this model enables us to move around the space-time of the event continuously. This model allows us to create virtual cameras that facilitate: (1) freezing the time and exploring views; (2) freezing a view and moving through time; and (3) simultaneously changing both time and view. We can also edit the videos and reveal occluded objects for a given view if it is visible in any of the other views. We validate our approach on challenging in-the-wild events captured using up to 15 mobile cameras. Aayush Bansal, Minh Vo, Yaser Sheikh, Deva Ramanan, Srinivasa G. Narasimhan |
CVPR | 5 |
| 2020 | Active Perception Using Light Curtains for Autonomous Driving
Siddharth Ancha, Yaadhav Raaj, Peiyun Hu, Srinivasa G. Narasimhan, David Held |
ECCV (5) | 4 |
| 2020 | Diffraction Line Imaging
Mark Sheinin, N. Dinesh Reddy, Matthew O'Toole, Srinivasa G. Narasimhan |
ECCV (2) | 4 |
| 2020 | TexMesh: Reconstructing Detailed Human Texture and Geometry from RGB-D Video
Tiancheng Zhi, Christoph Lassner, Tony Tung, Carsten Stoll, Srinivasa G. Narasimhan, Minh Vo |
ECCV (10) | 5 |
| 2020 | High Resolution Diffuse Optical Tomography using Short Range Indirect Subsurface ImagingabstractDiffuse optical tomography (DOT) is an approach to recover subsurface structures beneath the skin by measuring light propagation beneath the surface. The method is based on optimizing the difference between the images collected and a forward model that accurately represents diffuse photon propagation within a heterogeneous scattering medium. However, to date, most works have used a few source-detector pairs and recover the medium at only a very low resolution. And increasing the resolution requires prohibitive computations/storage. In this work, we present a fast imaging and algorithm for high resolution diffuse optical tomography with a line imaging and illumination system. Key to our approach is a convolution approximation of the forward heterogeneous scattering model that can be inverted to produce deeper than ever before structured beneath the surface. We show that our proposed method can detect reasonably accurate boundaries and relative depth of heterogeneous structures up to a depth of 8 mm below highly scattering medium such as milk. This work can extend the potential of DOT to recover more intricate structures (vessels, tissue, tumors, etc.) beneath the skin for diagnosing many dermatological and cardio-vascular conditions. Chao Liu 0064, Akash K. Maity, Artur Dubrawski, Ashutosh Sabharwal, Srinivasa G. Narasimhan |
ICCP | 5 |
| 2020 | Path tracing estimators for refractive radiative transferabstractRendering radiative transfer through media with a heterogeneous refractive index is challenging because the continuous refractive index variations result in light traveling along curved paths. Existing algorithms are based on photon mapping techniques, and thus are biased and result in strong artifacts. On the other hand, existing unbiased methods such as path tracing and bidirectional path tracing cannot be used in their current form to simulate media with a heterogeneous refractive index. We change this state of affairs by deriving unbiased path tracing estimators for this problem. Starting from the refractive radiative transfer equation (RRTE), we derive a path-integral formulation, which we use to generalize path tracing with next-event estimation and bidirectional path tracing to the heterogeneous refractive index setting. We then develop an optimization approach based on fast analytic derivative computations to produce the point-to-point connections required by these path tracing algorithms. We propose several acceleration techniques to handle complex scenes (surfaces and volumes) that include participating media with heterogeneous refractive fields. We use our algorithms to simulate a variety of scenes combining heterogeneous refraction and scattering, as well as tissue imaging techniques based on ultrasonic virtual waveguides and lenses. Our algorithms and publicly-available implementation can be used to characterize imaging systems such as refractive index microscopy, schlieren imaging, and acousto-optic imaging, and can facilitate the development of inverse rendering techniques for related applications. Adithya Kumar Pediredla, Yasin Karimi Chalmiani, Matteo Giuseppe Scopelliti, Maysamreza Chamanzar, Srinivasa G. Narasimhan, Ioannis Gkioulekas |
ACM Trans. Graph. | 5 |
| 2019 | Neural RGB(r)D Sensing: Depth and Uncertainty From a Video CameraabstractDepth sensing is crucial for 3D reconstruction and scene understanding. Active depth sensors provide dense metric measurements, but often suffer from limitations such as restricted operating ranges, low spatial resolution, sensor interference, and high power consumption. In this paper, we propose a deep learning (DL) method to estimate per-pixel depth and its uncertainty continuously from a monocular video stream, with the goal of effectively turning an RGB camera into an RGB-D camera. Unlike prior DL-based methods, we estimate a depth probability distribution for each pixel rather than a single depth value, leading to an estimate of a 3D depth probability volume for each input frame. These depth probability volumes are accumulated over time under a Bayesian filtering framework as more incoming frames are processed sequentially, which effectively reduces depth uncertainty and improves accuracy, robustness, and temporal stability. Compared to prior work, the proposed approach achieves more accurate and stable results, and generalizes better to new datasets. Experimental results also show the output of our approach can be directly fed into classical RGB-D based 3D scanning methods for 3D scene reconstruction. Chao Liu 0064, Jinwei Gu, Srinivasa G. Narasimhan, Jan Kautz |
CVPR | 4 |
| 2019 | Occlusion-Net: 2D/3D Occluded Keypoint Localization Using Graph NetworksabstractWe present Occlusion-Net, a framework to predict 2D and 3D locations of occluded keypoints for objects, in a largely self-supervised manner. We use an off-the-shelf detector as input (like MaskRCNN) that is trained only on visible key point annotations. This is the only supervision used in this work. A graph encoder network then explicitly classifies invisible edges and a graph decoder network corrects the occluded keypoint locations from the initial detector. Central to this work is a trifocal tensor loss that provides indirect self-supervision for occluded keypoint locations that are visible in other views of the object. The 2D keypoints are then passed into a 3D graph network that estimates the 3D shape and camera pose using the self-supervised re-projection loss. At test time, our approach successfully localizes keypoints in a single view under a diverse set of severe occlusion settings. We demonstrate and evaluate our approach on synthetic CAD data as well as a large image set capturing vehicles at many busy city intersections. As an interesting aside, we compare the accuracy of human labels of invisible keypoints against those obtained from geometric trifocal-tensor loss. N. Dinesh Reddy, Minh Vo, Srinivasa G. Narasimhan |
CVPR | 3 |
| 2019 | A Theory of Fermat Paths for Non-Line-Of-Sight Shape ReconstructionabstractWe present a novel theory of Fermat paths of light between a known visible scene and an unknown object not in the line of sight of a transient camera. These light paths either obey specular reflection or are reflected by the object's boundary, and hence encode the shape of the hidden object. We prove that Fermat paths correspond to discontinuities in the transient measurements. We then derive a novel constraint that relates the spatial derivatives of the path lengths at these discontinuities to the surface normal. Based on this theory, we present an algorithm, called Fermat Flow, to estimate the shape of the non-line-of-sight object. Our method allows, for the first time, accurate shape recovery of complex objects, ranging from diffuse to specular, that are hidden around the corner as well as hidden behind a diffuser. Finally, our approach is agnostic to the particular technology used for transient imaging. As such, we demonstrate mm-scale shape recovery from pico-second scale transients using a SPAD and ultrafast laser, as well as micron-scale reconstruction from femto-second scale transients using interferometry. We believe our work is a significant advance over the state-of-the-art in non-line-of-sight imaging. Shumian Xin, Sotiris Nousias, Kiriakos N. Kutulakos, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan, Ioannis Gkioulekas |
CVPR | 5 |
| 2019 | Multispectral Imaging for Fine-Grained Recognition of Powders on Complex BackgroundsabstractHundreds of materials, such as drugs, explosives, makeup, food additives, are in the form of powder. Recognizing such powders is important for security checks, criminal identification, drug control, and quality assessment. However, powder recognition has drawn little attention in the computer vision community. Powders are hard to distinguish: they are amorphous, appear matte, have little color or texture variation and blend with surfaces they are deposited on in complex ways. To address these challenges, we present the first comprehensive dataset and approach for powder recognition using multi-spectral imaging. By using Shortwave Infrared (SWIR) multi-spectral imaging together with visible light (RGB) and Near Infrared (NIR), powders can be discriminated with reasonable accuracy. We present a method to select discriminative spectral bands to significantly reduce acquisition time while improving recognition accuracy. We propose a blending model to synthesize images of powders of various thickness deposited on a wide range of surfaces. Incorporating band selection and image synthesis, we conduct fine-grained recognition of 100 powders on complex backgrounds, and achieve 60%~70% accuracy on recognition with known powder location, and over 40% mean IoU without known location. Tiancheng Zhi, Bernardo Rodrigues Pires, Martial Hebert, Srinivasa G. Narasimhan |
CVPR | 4 |
| 2019 | Quantifying the Benefits of Dynamic Partial Reconfiguration for Embedded Vision ApplicationsabstractDynamic partial reconfiguration (DPR) allows parts of an FPGA to be reprogrammed at runtime (i.e., repurposed). Though DPR has been supported by commercial devices and tools for more than a decade, it has been underutilized, perhaps, due to a shortage of demonstrated use-cases and quantified benefits over static FPGA mapping (without DPR). In this paper, we quantify the benefits of dynamic FPGA mapping (with DPR) over traditional static FPGA mapping for two vision applications deployed on systems with area/device cost, power or energy constraints (i.e., smart car and smart robot). In both applications, the FPGA needs to accelerate multiple tasks at 60 fps. However, all tasks are not required at the same time. In this work, instead of mapping all tasks statically on a large FPGA, the set of tasks needed at a given time is (1) repurposed on a smaller FPGA and (2) still meets the functional and performance requirements (i.e., 60 fps). In the two application examples, we show that dynamic mapping on smaller FPGAs reduces logic resource utilization by up to 3.2x, device cost by up to 10x, and power and energy consumption by up to 30% in comparison with static mapping on larger FPGAs. These benefits are crucial for applications deployed on systems where reducing area/device cost, power and energy is as important as meeting performance requirement. Marie Nguyen, Robert Tamburo, Srinivasa G. Narasimhan, James C. Hoe |
FPL | 3 |
| 2019 | Episcan360: Active Epipolar Imaging for Live Omni-directional StereoabstractActive epipolar imaging simultaneously illuminates and images a scene along epipolar planes. The recent devices based on this principle, such as Episcan [1] and EpiToF [2], significantly reduce the effects of indirect light transport and ambient light, resulting in live capture of high resolution 3D at longer ranges indoors and outdoors. However, these devices are designed for narrow fields of view where lens distortion can be ignored and epipolar plane/line constraints are satisfied. In this work, we extend active epipolar imaging to obtain live-omnidirectional stereo for the first time. Instead of using a 2D sensor/projector, we use a 1D sensor and 1D light sheet source that are placed in a rectified configuration. We observe that when the lens center axis and the sensor are aligned, the distortion is mostly along that axis and epipolar plane constraint remains satisfied. This allows us to use small wide-angle lenses resulting in a compact 1D Episcan. This 1D Episcan is then spun quickly to obtain an near-spherical field of view. Based on this design, we demonstrate a custom-built hand-held working prototype for live capture of omni-directional active stereo images. D. W. Wilson Hamilton, Jaime Bourne, Jeffrey D. McMahill, Joan Campoy, Herman Herman, Srinivasa G. Narasimhan |
ICCP | 6 |
| 2019 | STORM: Super-resolving Transients by OveRsampled MeasurementsabstractImage sensors that can measure the time of travel of photons are gaining importance in a myriad of applications such as LIDAR, non-line of sight imaging, light-in-flight imaging, and imaging through scattering media. While the price of these sensors is dramatically shrinking, there remains a trade-off between spatial resolution and temporal resolution. While single-pixel detectors using the single photon avalanche diode (SPAD) technology can achieve 10-30 ps time resolution, the current generation array detectors can only produce an order of magnitude lower temporal resolution due to space-related fabrication constraints. Moreover, this limit is due to bandwidth, read-out and circuit-area constraints on the detector array and therefore unlikely to dramatically change in the next few years.In this paper, we demonstrate a computational imaging approach that utilizes multiple measurements with calibrated sub-temporal resolution delays on the illumination pulse and super-resolution post-processing algorithms that together can achieve an order of magnitude improvement in the time resolution of the acquired transients. We build an experimental prototype, using a 32 × 32 SPAD detector array with 400ps time resolution and demonstrate recovery of transients with ≈ 50ps time resolution, an 8× improvement in time resolution resulting in a 5× improvement in depth reconstruction error. Ankit Raghuram, Adithya Kumar Pediredla, Srinivasa G. Narasimhan, Ioannis Gkioulekas, Ashok Veeraraghavan |
ICCP | 3 |
| 2019 | Agile Depth Sensing Using Triangulation Light CurtainsabstractDepth sensors like LIDARs and Kinect use a fixed depth acquisition strategy that is independent of the scene of interest. Due to the low spatial and temporal resolution of these sensors, this strategy can undersample parts of the scene that are important (small or fast moving objects), or oversample areas that are not informative for the task at hand (a fixed planar wall). In this paper, we present an approach and system to dynamically and adaptively sample the depths of a scene using the principle of triangulation light curtains. The approach directly detects the presence or absence of objects at specified 3D lines. These 3D lines can be sampled sparsely, non-uniformly, or densely only at specified regions. The depth sampling can be varied in real-time, enabling quick object discovery or detailed exploration of areas of interest. These results are achieved using a novel prototype light curtain system that is based on a 2D rolling shutter camera with higher light efficiency, working range, and faster adaptation than previous work, making it useful broadly for autonomous navigation and exploration. Joseph R. Bartels, Jian Wang 0100, William Whittaker, Srinivasa G. Narasimhan |
ICCV | 4 |
| 2018 | CarFusion: Combining Point Tracking and Part Detection for Dynamic 3D Reconstruction of VehiclesabstractDespite significant research in the area, reconstruction of multiple dynamic rigid objects (eg. vehicles) observed from wide-baseline, uncalibrated and unsynchronized cameras, remains hard. On one hand, feature tracking works well within each view but is hard to correspond across multiple cameras with limited overlap infields of view or due to occlusions. On the other hand, advances in deep learning have resulted in strong detectors that work across different viewpoints but are still not precise enough for triangulation-based reconstruction. In this work, we develop a framework to fuse both the single-view feature tracks and multiview detected part locations to significantly improve the detection, localization and reconstruction of moving vehicles, even in the presence of strong occlusions. We demonstrate our framework at a busy traffic intersection by reconstructing over 62 vehicles passing within a 3-minute window. We evaluate the different components within our framework and compare to alternate approaches such as reconstruction using tracking-by-detection. N. Dinesh Reddy, Minh Vo, Srinivasa G. Narasimhan |
CVPR | 3 |
| 2018 | Deep Material-Aware Cross-Spectral Stereo MatchingabstractCross-spectral imaging provides strong benefits for recognition and detection tasks. Often, multiple cameras are used for cross-spectral imaging, thus requiring image alignment, or disparity estimation in a stereo setting. Increasingly, multi-camera cross-spectral systems are embedded in active RGBD devices (e.g. RGB-NIR cameras in Kinect and iPhone X). Hence, stereo matching also provides an opportunity to obtain depth without an active projector source. However, matching images from different spectral bands is challenging because of large appearance variations. We develop a novel deep learning framework to simultaneously transform images across spectral bands and estimate disparity. A material-aware loss function is incorporated within the disparity prediction network to handle regions with unreliable matching such as light sources, glass windshields and glossy surfaces. No depth supervision is required by our method. To evaluate our method, we used a vehicle-mounted RGB-NIR stereo system to collect 13.7 hours of video data across a range of areas in and around a city. Experiments show that our method achieves strong performance and reaches real-time speed. Tiancheng Zhi, Bernardo Rodrigues Pires, Martial Hebert, Srinivasa G. Narasimhan |
CVPR | 4 |
| 2018 | Programmable Triangulation Light Curtains
Jian Wang 0100, Joseph R. Bartels, William Whittaker, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan |
ECCV (3) | 5 |
| 2018 | Acquiring and characterizing plane-to-ray indirect light transportabstractSeparation of light transport into direct and indirect paths has enabled new visualizations of light in everyday scenes. However, indirect light itself contains a variety of components from subsurface scattering to diffuse and specular interreflections, all of which contribute to complex visual appearance. In this paper, we present a new imaging technique that captures and analyzes these components of indirect light via light transport between epipolar planes of illumination and rays of received light. This plane-to-ray light transport is captured using a rectified projector-camera system where we vary the offset between projector and camera rows (implemented as synchronization delay) as well as the exposure of each camera row. The resulting delay-exposure stack of images can capture live short and long-range indirect light transport, disambiguate subsurface scattering, diffuse and specular interreflections, and distinguish materials according to their subsurface scattering properties. Hiroyuki Kubo, Suren Jayasuriya, Takafumi Iwaguchi, Takuya Funatomi, Yasuhiro Mukaigawa, Srinivasa G. Narasimhan |
ICCP | 6 |
| 2018 | Near-light photometric stereo using circularly placed point light sourcesabstractMost photometric stereo approaches assume distant or directional lighting and orthographic imaging. However, when the source is divergent and is near the object and the camera is projective, the image intensity of a Lambertian object is a non-linear function of both the unknown surface normals and the unknown distances of the source to the surface points. The resulting non-linear optimization is non-convex and highly sensitive to the initial guess. In this paper, we propose a two-stage near-light photometric stereo method using circularly placed point light sources (commonly seen in recent consumer imaging devices like NESTcam, Amazon Cloudcam, etc). We represent the scene using a 3D mesh and directly optimize the vertices of the mesh. This reduces the complexity of the relationship between surface normals and depths in the image formation model. In the first stage, we optimize the vertex positions using the differential images induced by small changes in light source position. This procedure yields a strong initial guess for the second stage that refines the estimations using the raw captured images. We propose an accurate calibration approach to estimate the positions of the sources. Our approach performs better on simulations and on real Lambertian scenes with complex shapes than the state-of-the-art method with near-field lighting. Chao Liu 0064, Srinivasa G. Narasimhan, Artur Dubrawski |
ICCP | 2 |
| 2018 | Acquiring short range 4D light transport with synchronized projector camera systemabstractLight interacts with a scene in various ways. For scene understanding, a light transport is useful because it describes a relationship between the incident light ray and the result of the interaction. Our goal is to acquire the 4D light transport between the projector and the camera, focusing on direct and short-range transport that include the effect of the diffuse reflections, subsurface scattering, and inter-reflections. The acquisition of the light transport is challenging since the acquisition of the full 4D light transport requires a large number of measurement. We propose an efficient method to acquire short range light transport, which is dominant in the general scene, using synchronized projector-camera system. We show the transport profile of various materials, including uniform or heterogeneous subsurface scattering. Takafumi Iwaguchi, Hiroyuki Kubo, Takuya Funatomi, Yasuhiro Mukaigawa, Srinivasa G. Narasimhan |
VRST | 5 |
| 2017 | Matting and Depth Recovery of Thin Structures Using a Focal StackabstractThin structures such as fence, grass and vessels are common in photography and scientific imaging. They exhibit complex 3D structures with sharp depth variations/discontinuities and mutual occlusions. In this paper, we develop a method to estimate the occlusion matte and depths of thin structures from a focal image stack, which is obtained either by varying the focus/aperture of the lens or computed from a one-shot light field image. We propose an image formation model that explicitly describes the spatially varying optical blur and mutual occlusions for structures located at different depths. Based on the model, we derive an efficient MCMC inference algorithm that enables direct and analytical computations of the iterative update for the model/images without re-rendering images in the sampling process. Then, the depths of the thin structures are recovered using gradient descent with the differential terms computed using the image formation model. We apply the proposed method to scenes at both macro and micro scales. For macro-scale, we evaluate our method on scenes with complex 3D thin structures such as tree branches and grass. For micro-scale, we apply our method to in-vivo microscopic images of micro-vessels with diameters less than 50 μm. To our knowledge, the proposed method is the first approach to reconstruct the 3D structures of micro-vessels from non-invasive in-vivo image measurements. Chao Liu 0064, Srinivasa G. Narasimhan, Artur Dubrawski |
CVPR | 2 |
| 2017 | The Geometry of First-Returning Photons for Non-Line-of-Sight ImagingabstractNon-line-of-sight (NLOS) imaging utilizes the full 5D light transient measurements to reconstruct scenes beyond the cameras field of view. Mathematically, this requires solving an elliptical tomography problem that unmixes the shape and albedo from spatially-multiplexed measurements of the NLOS scene. In this paper, we propose a new approach for NLOS imaging by studying the properties of first-returning photons from three-bounce light paths. We show that the times of flight of first-returning photons are dependent only on the geometry of the NLOS scene and each observation is almost always generated from a single NLOS scene point. Exploiting these properties, we derive a space carving algorithm for NLOS scenes. In addition, by assuming local planarity, we derive an algorithm to localize NLOS scene points in 3D and estimate their surface normals. Our methods do not require either the full transient measurements or solving the hard elliptical tomography problem. We demonstrate the effectiveness of our methods through simulations as well as real data captured from a SPAD sensor. Chia-Yin Tsai, Kiriakos N. Kutulakos, Srinivasa G. Narasimhan, Aswin C. Sankaranarayanan |
CVPR | 3 |
| 2017 | Epipolar time-of-flight imagingabstractConsumer time-of-flight depth cameras like Kinect and PMD are cheap, compact and produce video-rate depth maps in short-range applications. In this paper we apply energy-efficient epipolar imaging to the ToF domain to significantly expand the versatility of these sensors: we demonstrate live 3D imaging at over 15 m range outdoors in bright sunlight; robustness to global transport effects such as specular and diffuse inter-reflections---the first live demonstration for this ToF technology; interference-free 3D imaging in the presence of many ToF sensors, even when they are all operating at the same optical wavelength and modulation frequency; and blur-free, distortion-free 3D video in the presence of severe camera shake. We believe these achievements can make such cheap ToF devices broadly applicable in consumer and robotics domains. Supreeth Achar, Joseph R. Bartels, William Whittaker, Kiriakos N. Kutulakos, Srinivasa G. Narasimhan |
ACM Trans. Graph. | 5 |
| 2016 | Simultaneous Estimation of Near IR BRDF and Fine-Scale Surface GeometryabstractNear-Infrared (NIR) images of most materials exhibit less texture or albedo variations making them beneficial for vision tasks such as intrinsic image decomposition and structured light depth estimation. Understanding the reflectance properties (BRDF) of materials in the NIR wavelength range can be further useful for many photometric methods including shape from shading and inverse rendering. However, even with less albedo variation, many materials e.g. fabrics, leaves, etc. exhibit complex fine-scale surface detail making it hard to accurately estimate BRDF. In this paper, we present an approach to simultaneously estimate NIR BRDF and fine-scale surface details by imaging materials under different IR lighting and viewing directions. This is achieved by an iterative scheme that alternately estimates surface detail and NIR BRDF of materials. Our setup does not require complicated gantries or calibration and we present the first NIR dataset of 100 materials including a variety of fabrics (knits, weaves, cotton, satin, leather), and organic (skin, leaves, jute, trunk, fur) and inorganic materials (plastic, concrete, carpet). The NIR BRDFs measured from material samples are used with a shape-from-shading algorithm to demonstrate fine-scale reconstruction of objects from a single NIR image. Gyeongmin Choe, Srinivasa G. Narasimhan, In-So Kweon |
CVPR | 2 |
| 2016 | Spatiotemporal Bundle Adjustment for Dynamic 3D ReconstructionabstractBundle adjustment jointly optimizes camera intrinsics and extrinsics and 3D point triangulation to reconstruct a static scene. The triangulation constraint however is invalid for moving points captured in multiple unsynchronized videos and bundle adjustment is not purposed to estimate the temporal alignment between cameras. In this paper, we present a spatiotemporal bundle adjustment approach that jointly optimizes four coupled sub-problems: estimating camera intrinsics and extrinsics, triangulating 3D static points, as well as subframe temporal alignment between cameras and estimating 3D trajectories of dynamic points. Key to our joint optimization is the careful integration of physics-based motion priors within the reconstruction pipeline, validated on a large motion capture corpus. We present an end-to-end pipeline that takes multiple uncalibrated and unsynchronized video streams and produces a dynamic reconstruction of the event. Because the videos are aligned with sub-frame precision, we reconstruct 3D trajectories of unconstrained outdoor activities at much higher temporal resolution than the input videos. Minh Vo, Srinivasa G. Narasimhan, Yaser Sheikh |
CVPR | 2 |
| 2016 | Dual Structured Light 3D Using a 1D Sensor
Jian Wang 0100, Aswin C. Sankaranarayanan, Mohit Gupta 0001, Srinivasa G. Narasimhan |
ECCV (6) | 4 |
| 2016 | Model effectiveness prediction and system adaptation for photometric stereo in murky water
Chourmouzios Tsiotsios, Tae-Kyun Kim 0001, Andrew J. Davison, Srinivasa G. Narasimhan |
Comput. Vis. Image Underst. | 4 |
| 2016 | Texture Illumination Separation for Single-Shot Structured Light ReconstructionabstractActive illumination based methods have a trade-off between acquisition time and resolution of the estimated 3D shapes. Multi-shot approaches can generate dense reconstructions but require stationary scenes. Single-shot methods are applicable to dynamic objects but can only estimate sparse reconstructions and are sensitive to surface texture. We present a single-shot approach to produce dense shape reconstructions of highly textured objects illuminated by one or more projectors. The key to our approach is an image decomposition scheme that can recover the illumination image of different projectors and the texture images of the scene from their mixed appearances. We focus on three cases of mixed appearances: the illumination from one projector onto textured surface, illumination from multiple projectors onto a textureless surface, or their combined effect. Our method can accurately compute per-pixel warps from the illumination patterns and the texture template to the observed image. The texture template is obtained by interleaving the projection sequence with an all-white pattern. The estimated warps are reliable even with infrequent interleaved projection and strong object deformation. Thus, we obtain detailed shape reconstruction and dense motion tracking of the textured surfaces. The proposed method, implemented using a one camera and two projectors system, is validated on synthetic and real data containing subtle non-rigid surface deformations. Minh Vo, Srinivasa G. Narasimhan, Yaser Sheikh |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Real-time visual analysis of microvascular blood flow for critical careabstractMicrocirculatory monitoring plays an important role in diagnosis and treatment of critical care patients. Sidestream Dark Field (SDF) imaging devices have been used to visualize and support interpretation of the micro-vascular blood flow. However, due to subsurface scattering within the tissue that embeds the capillaries, transparency of plasma, imaging noise and lack of features, it is difficult to obtain reliable physiological data from SDF videos. Therefore, thus far microcirculatory videos have been analyzed manually with significant input from expert clinicians. In this paper, we present a framework that automates the analysis process. It includes stages of video stabilization, enhancement, and micro-vessel extraction, in order to automatically estimate statistics of the micro blood flows from SDF videos. Our method has been validated in critical care experiments conducted carefully to record the microcirculatory blood flow in test animal subjects before, during and after induced bleeding episodes, as well as to study the effect of fluid resuscitation. Our method is able to extract microcirculatory measurements that are consistent with clinical intuition and it has a potential to become a useful tool in critical care medicine. Chao Liu 0064, Hernando Gómez, Srinivasa G. Narasimhan, Artur Dubrawski, Michael R. Pinsky, Brian Zuckerbraun |
CVPR | 3 |
| 2015 | Performance Characterization of Reactive Visual SystemsabstractWe consider the class of projector-camera systems that adaptively image and illuminate a dynamic environment. Examples include adaptive front lighting in vehicles, dynamic stage performance lighting, adaptive dynamic range imaging and volumetric displays. A simulator is developed to explore the design space of such Reactive Visual Systems. Simulations are conducted to characterize system performance by analyzing the effects of end-to-end latency, jitter, and prediction algorithm complexity. Key operating points are identified where systems with simple prediction algorithms can outperform systems with more complex prediction algorithms. Based on the lessons learned from simulations, a low latency and low jitter, tight closed-loop reactive visual system is built. For the first time, we measure end-to-end latency, perform jitter analysis, investigate various prediction algorithms and their effect on system performance, compare our system's performance to previous work, and demonstrate dis-illumination of falling snow-like particles and photography of fast moving scenes. Subhagato Dutta, Abhishek Chugh, Robert Tamburo, Anthony Rowe 0001, Srinivasa G. Narasimhan |
ICCP | 5 |
| 2015 | Theory and Practice of Hierarchical Data-driven Descent for Optimal Deformation Estimation
Yuandong Tian, Srinivasa G. Narasimhan |
Int. J. Comput. Vis. | 2 |
| 2015 | Homogeneous codes for energy-efficient illumination and imagingabstractProgrammable coding of light between a source and a sensor has led to several important results in computational illumination, imaging and display. Little is known, however, about how to utilize energy most effectively, especially for applications in live imaging. In this paper, we derive a novel framework to maximize energy efficiency by "homogeneous matrix factorization" that respects the physical constraints of many coding mechanisms (DMDs/LCDs, lasers, etc. ). We demonstrate energy-efficient imaging using two prototypes based on DMD and laser illumination. For our DMD-based prototype, we use fast local optimization to derive codes that yield brighter images with fewer artifacts in many transport probing tasks. Our second prototype uses a novel combination of a low-power laser projector and a rolling shutter camera. We use this prototype to demonstrate never-seen-before capabilities such as (1) capturing live structured-light video of very bright scenes---even a light bulb that has been turned on; (2) capturing epipolar-only and indirect-only live video with optimal energy efficiency; (3) using a low-power projector to reconstruct 3D objects in challenging conditions such as strong indirect light, strong ambient light, and smoke; and (4) recording live video from a projector's---rather than the camera's---point of view. Matthew O'Toole, Supreeth Achar, Srinivasa G. Narasimhan, Kiriakos N. Kutulakos |
ACM Trans. Graph. | 3 |
| 2014 | Multi Focus Structured Light for Recovering Scene Shape and Global Illumination
Supreeth Achar, Srinivasa G. Narasimhan |
ECCV (1) | 2 |
| 2014 | Passive Tomography of Turbulence Strength
Marina Alterman, Yoav Y. Schechner, Minh Vo, Srinivasa G. Narasimhan |
ECCV (4) | 4 |
| 2014 | Programmable Automotive Headlights
Robert Tamburo, Eriko Nurvitadhi, Abhishek Chugh, Anthony Rowe 0001, Takeo Kanade, Srinivasa G. Narasimhan |
ECCV (4) | 7 |
| 2014 | Translucent Radiosity: Efficiently CombiningDiffuse Inter-Reflection andSubsurface ScatteringabstractIt is hard to efficiently model the light transport in scenes with translucent objects for interactive applications. The inter-reflection between objects and their environments and the subsurface scattering through the materials intertwine to produce visual effects like color bleeding, light glows, and soft shading. Monte-Carlo based approaches have demonstrated impressive results but are computationally expensive, and faster approaches model either only inter-reflection or only subsurface scattering. In this paper, we present a simple analytic model that combines diffuse inter-reflection and isotropic subsurface scattering. Our approach extends the classical work in radiosity by including a subsurface scattering matrix that operates in conjunction with the traditional form factor matrix. This subsurface scattering matrix can be constructed using analytic, measurement-based or simulation-based models and can capture both homogeneous and heterogeneous translucencies. Using a fast iterative solution to radiosity, we demonstrate scene relighting and dynamically varying object translucencies at near interactive rates. Yu Sheng, Yulong Shi, Lili Wang 0006, Srinivasa G. Narasimhan |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2013 | Compensating for Motion during Direct-Global SeparationabstractSeparating the direct and global components of radiance can aid shape recovery algorithms and can provide useful information about materials in a scene. Practical methods for finding the direct and global components use multiple images captured under varying illumination patterns and require the scene, light source and camera to remain stationary during the image acquisition process. In this paper, we develop a motion compensation method that relaxes this condition and allows direct-global separation to be performed on video sequences of dynamic scenes captured by moving projector-camera systems. Key to our method is being able to register frames in a video sequence to each other in the presence of time varying, high frequency active illumination patterns. We compare our motion compensated method to alternatives such as single shot separation and frame interleaving as well as ground truth. We present results on challenging video sequences that include various types of motions and deformations in scenes that contain complex materials like fabric, skin, leaves and wax. Supreeth Achar, Stephen Nuske, Srinivasa G. Narasimhan |
ICCV | 3 |
| 2013 | Hierarchical Data-Driven Descent for Efficient Optimal Deformation EstimationabstractReal-world surfaces such as clothing, water and human body deform in complex ways. The image distortions observed are high-dimensional and non-linear, making it hard to estimate these deformations accurately. The recent data-driven descent approach applies Nearest Neighbor estimators iteratively on a particular distribution of training samples to obtain a globally optimal and dense deformation field between a template and a distorted image. In this work, we develop a hierarchical structure for the Nearest Neighbor estimators, each of which can have only a local image support. We demonstrate in both theory and practice that this algorithm has several advantages over the non-hierarchical version: it guarantees global optimality with significantly fewer training samples, is several orders faster, provides a metric to decide whether a given image is ``hard'' (or ``easy'') requiring more (or less) samples, and can handle more complex scenes that include both global motion and local deformation. The proposed algorithm successfully tracks a broad range of non-rigid scenes including water, clothing, and medical images, and compares favorably against several other deformation estimation and tracking approaches that do not provide optimality guarantees. Yuandong Tian, Srinivasa G. Narasimhan |
ICCV | 2 |
| 2013 | A practical analytic model for the radiosity of translucent scenesabstractLight propagation in scenes with translucent objects is hard to model efficiently for interactive applications. The inter-reflections between objects and their environments and the subsurface scattering through the materials intertwine to produce visual effects like color bleeding, light glows and soft shading. Monte-Carlo based approaches have demonstrated impressive results but are computationally expensive, and faster approaches model either only inter-reflections or only subsurface scattering. In this paper, we present a simple analytic model that combines diffuse inter-reflections and isotropic subsurface scattering. Our approach extends the classical work in radiosity by including a subsurface scattering matrix that operates in conjunction with the traditional form-factor matrix. This subsurface scattering matrix can be constructed using analytic, measurement-based or simulation-based models and can capture both homogeneous and heterogeneous translucencies. Using a fast iterative solution to radiosity, we demonstrate scene relighting and dynamically varying object translucencies at near interactive rates. Yu Sheng, Yulong Shi, Lili Wang 0006, Srinivasa G. Narasimhan |
I3D | 4 |
| 2013 | A Practical Approach to 3D Scanning in the Presence of Interreflections, Subsurface Scattering and Defocus
Mohit Gupta 0001, Amit K. Agrawal, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
Int. J. Comput. Vis. | 4 |
| 2013 | Non-polynomial Galerkin projection on deforming meshesabstractThis paper extends Galerkin projection to a large class of non-polynomial functions typically encountered in graphics. We demonstrate the broad applicability of our approach by applying it to two strikingly different problems: fluid simulation and radiosity rendering, both using deforming meshes. Standard Galerkin projection cannot efficiently approximate these phenomena. Our approach, by contrast, enables the compact representation and approximation of these complex non-polynomial systems, including quotients and roots of polynomials. We rely on representing each function to be model-reduced as a composition of tensor products, matrix inversions, and matrix roots. Once a function has been represented in this form, it can be easily model-reduced, and its reduced form can be evaluated with time and memory costs dependent only on the dimension of the reduced space. Matt Stanton, Yu Sheng, Martin Wicke, Federico Perazzi, Amos Yuen, Srinivasa G. Narasimhan, Adrien Treuille |
ACM Trans. Graph. | 6 |
| 2012 | Depth from optical turbulenceabstractTurbulence near hot surfaces such as desert terrains and roads during the summer, causes shimmering, distortion and blurring in images. While recent works have focused on image restoration, this paper explores what information about the scene can be extracted from the distortion caused by turbulence. Based on the physical model of wave propagation, we first study the relationship between the scene depth and the amount of distortion caused by homogenous turbulence. We then extend this relationship to more practical scenarios such as finite extent and height-varying turbulence, and present simple algorithms to estimate depth ordering, depth discontinuity and relative depth, from a sequence of short exposure images. In the case of general non-homogenous turbulence, we show that a statistical property of turbulence can be used to improve long-range structure-from-motion (or stereo). We demonstrate the accuracy of our methods in both laboratory and outdoor settings and conclude that turbulence (when present) can be a strong and useful depth cue. Yuandong Tian, Srinivasa G. Narasimhan, Alan Van Nevel |
CVPR | 2 |
| 2012 | Exploring the Spatial Hierarchy of Mixture Models for Human Pose Estimation
Yuandong Tian, C. Lawrence Zitnick, Srinivasa G. Narasimhan |
ECCV (5) | 3 |
| 2012 | Fast reactive control for illumination through rain and snowabstractDuring low-light conditions, drivers rely mainly on headlights to improve visibility. But in the presence of rain and snow, headlights can paradoxically reduce visibility due to light reflected off of precipitation back towards the driver. Precipitation also scatters light across a wide range of angles that disrupts the vision of drivers in oncoming vehicles. In contrast to recent computer vision methods that digitally remove rain and snow streaks from captured images, we present a system that will directly improve driver visibility by controlling illumination in response to detected precipitation. The motion of precipitation is tracked and only the space around particles is illuminated using fast dynamic control. Using a physics-based simulator, we show how such a system would perform under a variety of weather conditions. We build and evaluate a proof-of-concept system that can avoid water drops generated in the laboratory. Raoul de Charette, Robert Tamburo, Peter C. Barnum, Anthony Rowe 0001, Takeo Kanade, Srinivasa G. Narasimhan |
ICCP | 6 |
| 2012 | A Combined Theory of Defocused Illumination and Global Light Transport
Mohit Gupta 0001, Yuandong Tian, Srinivasa G. Narasimhan, Li Zhang 0003 |
Int. J. Comput. Vis. | 3 |
| 2012 | Exploiting DLP Illumination Dithering for Reconstruction and Photography of High-Speed Scenes
Sanjeev J. Koppal, Shuntaro Yamazaki, Srinivasa G. Narasimhan |
Int. J. Comput. Vis. | 3 |
| 2012 | Estimating the Natural Illumination Conditions from a Single Outdoor Image
Jean-François Lalonde, Alexei A. Efros, Srinivasa G. Narasimhan |
Int. J. Comput. Vis. | 3 |
| 2012 | Globally Optimal Estimation of Nonrigid Image Distortion
Yuandong Tian, Srinivasa G. Narasimhan |
Int. J. Comput. Vis. | 2 |
| 2011 | Structured light 3D scanning in the presence of global illuminationabstractGlobal illumination effects such as inter-reflections, diffusion and sub-surface scattering severely degrade the performance of structured light-based 3D scanning. In this paper, we analyze the errors caused by global illumination in structured light-based shape recovery. Based on this analysis, we design structured light patterns that are resilient to individual global illumination effects using simple logical operations and tools from combinatorial mathematics. Scenes exhibiting multiple phenomena are handled by combining results from a small ensemble of such patterns. This combination also allows us to detect any residual errors that are corrected by acquiring a few additional images. Our techniques do not require explicit separation of the direct and global components of scene radiance and hence work even in scenarios where the separation fails or the direct component is too low. Our methods can be readily incorporated into existing scanning systems without significant overhead in terms of capture time or hardware. We show results on a variety of scenes with complex shape and material properties and challenging global illumination effects. Mohit Gupta 0001, Amit K. Agrawal, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
CVPR | 4 |
| 2011 | Rectification and 3D reconstruction of curved document imagesabstractDistortions in images of documents, such as the pages of books, adversely affect the performance of optical character recognition (OCR) systems. Removing such distortions requires the 3D deformation of the document that is often measured using special and precisely calibrated hardware (stereo, laser range scanning or structured light). In this paper, we introduce a new approach that automatically reconstructs the 3D shape and rectifies a deformed text document from a single image. We first estimate the 2D distortion grid in an image by exploiting the line structure and stroke statistics in text documents. This approach does not rely on more noise-sensitive operations such as image binarization and character segmentation. The regularity in the text pattern is used to constrain the 2D distortion grid to be a perspective projection of a 3D parallelogram mesh. Based on this constraint, we present a new shape-from-texture method that computes the 3D deformation up to a scale factor using SVD. Unlike previous work, this formulation imposes no restrictions on the shape (e.g., a developable surface). The estimated shape is then used to remove both geometric distortions and photometric (shading) effects in the image. We demonstrate our techniques on documents containing a variety of languages, fonts and sizes. Yuandong Tian, Srinivasa G. Narasimhan |
CVPR | 2 |
| 2011 | Yield estimation in vineyards by visual grape detectionabstractThe harvest yield in vineyards can vary significantly from year to year and also spatially within plots due to variations in climate, soil conditions and pests. Fine grained knowledge of crop yields can allow viticulturists to better manage their vineyards. The current industry practice for yield prediction is destructive, expensive and spatially sparse - during the growing season sparse samples are taken and extrapolated to determine overall yield. We present an automated method that uses computer vision to detect and count grape berries. The method could potentially be deployed across large vineyards taking measurements at every vine in a non-destructive manner. Our berry detection uses both shape and visual texture and we can demonstrate detection of green berries against a green leaf background. Berry detections are counted and the eventual harvest yield is predicted. Results are presented for 224 vines (over 450 meters) of two different grape varieties and compared against the actual harvest yield as groundtruth. We calibrate our berry count to yield and find that we can predict yield of individual vineyard rows to within 9.8% of actual crop weight. Stephen Nuske, Supreeth Achar, Terry Bates, Srinivasa G. Narasimhan, Sanjiv Singh |
IROS | 4 |
| 2010 | Optimal coded sampling for temporal super-resolutionabstractConventional low frame rate cameras result in blur and/or aliasing in images while capturing fast dynamic events. Multiple low speed cameras have been used previously with staggered sampling to increase the temporal resolution. However, previous approaches are inefficient: they either use small integration time for each camera which does not provide light benefit, or use large integration time in a way that requires solving a big ill-posed linear system. We propose coded sampling that address these issues: using N cameras it allows N times temporal superresolution while allowing ~N/2 times more light compared to an equivalent high speed camera. In addition, it results in a well-posed linear system which can be solved independently for each frame, avoiding reconstruction artifacts and significantly reducing the computational time and memory. Our proposed sampling uses optimal multiplexing code considering additive Gaussian noise to achieve the maximum possible SNR in the recovered video. We show how to implement coded sampling on off-the-shelf machine vision cameras. We also propose a new class of invertible codes that allow continuous blur in captured frames, leading to an easier hardware implementation. Amit K. Agrawal, Mohit Gupta 0001, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
CVPR | 4 |
| 2010 | A globally optimal data-driven approach for image distortion estimationabstractImage alignment in the presence of non-rigid distortions is a challenging task. Typically, this involves estimating the parameters of a dense deformation field that warps a distorted image back to its undistorted template. Generative approaches based on parameter optimization such as Lucas-Kanade can get trapped within local minima. On the other hand, discriminative approaches like Nearest-Neighbor require a large number of training samples that grows exponentially with the desired accuracy. In this work, we develop a novel data-driven iterative algorithm that combines the best of both generative and discriminative approaches. For this, we introduce the notion of a “pull-back” operation that enables us to predict the parameters of the test image using training samples that are not in its neighborhood (not ϵ-close) in parameter space. We prove that our algorithm converges to the global optimum using a significantly lower number of training samples that grows only logarithmically with the desired accuracy. We analyze the behavior of our algorithm extensively using synthetic data and demonstrate successful results on experiments with complex deformations due to water and clothing. Yuandong Tian, Srinivasa G. Narasimhan |
CVPR | 2 |
| 2010 | Flexible Voxels for Motion-Aware Videography
Mohit Gupta 0001, Amit K. Agrawal, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
ECCV (1) | 4 |
| 2010 | Detecting Ground Shadows in Outdoor Consumer Photographs
Jean-François Lalonde, Alexei A. Efros, Srinivasa G. Narasimhan |
ECCV (2) | 3 |
| 2010 | Analysis of Rain and Snow in Frequency Space
Peter C. Barnum, Srinivasa G. Narasimhan, Takeo Kanade |
Int. J. Comput. Vis. | 2 |
| 2010 | What Do the Sun and the Sky Tell Us About the Camera?
Jean-François Lalonde, Srinivasa G. Narasimhan, Alexei A. Efros |
Int. J. Comput. Vis. | 2 |
| 2010 | A Multi-Image Shape-from-Shading Framework for Near-Lighting Perspective Endoscopes
Srinivasa G. Narasimhan, Branislav Jaramaz |
Int. J. Comput. Vis. | 2 |
| 2010 | A multi-layered display with water dropsabstractWe present a multi-layered display that uses water drops as voxels. Water drops refract most incident light, making them excellent wide-angle lenses. Each 2D layer of our display can exhibit arbitrary visual content, creating a layered-depth (2.5D) display. Our system consists of a single projector-camera system and a set of linear drop generator manifolds that are tightly synchronized and controlled using a computer. Following the principles of fluid mechanics, we are able to accurately generate and control drops so that, at any time instant, no two drops occupy the same projector pixel's line-of-sight. This drop control is combined with an algorithm for space-time division of projector light rays. Our prototype system has up to four layers, with each layer consisting of an row of 50 drops that can be generated at up to 60 Hz. The effective resolution of the display is 50x projector vertical-resolution x number of layers. We show how this water drop display can be used for text, videos, and interactive games. Peter C. Barnum, Srinivasa G. Narasimhan, Takeo Kanade |
ACM Trans. Graph. | 2 |
| 2009 | (De) focusing on global light transport for active scene recoveryabstractMost active scene recovery techniques assume that a scene point is illuminated only directly by the illumination source. Consequently, global illumination effects due to inter-reflections, sub-surface scattering and volumetric scattering introduce strong biases in the recovered scene shape. Our goal is to recover scene properties in the presence of global illumination. To this end, we study the interplay between global illumination and the depth cue of illumination defocus. By expressing both these effects as low pass filters, we derive an approximate invariant that can be used to separate them without explicitly modeling the light transport. This is directly useful in any scenario where limited depth-of-field devices (such as projectors) are used to illuminate scenes with global light transport and significant depth variations. We show two applications: (a) accurate depth recovery in the presence of global illumination, and (b) factoring out the effects of defocus for correct direct-global separation in large depth scenes. We demonstrate our approach using scenes with complex shapes, reflectances, textures and translucencies. Mohit Gupta 0001, Yuandong Tian, Srinivasa G. Narasimhan, Li Zhang 0003 |
CVPR | 3 |
| 2009 | Shadow cameras: Reciprocal views from illumination masksabstractScene appearance from the point of view of a light source is called a reciprocal or dual view. Since there exists a large diversity in illumination, these virtual views may be non-perspective and multi-viewpoint in nature. In this paper, we demonstrate the use of occluding masks to recover these dual views, which we term shadow cameras. We first show how to render a single reciprocal scene view by swapping the camera and light source positions. We extend this technique for multiple views by building a virtual shadow camera array with static masks and a moving source. We also capture non-perspective views such as orthographic, cross-slit and a pushbroom variant, while introducing novel applications such as converting between camera projections and removing refractive and catadioptric distortions. Finally, since a shadow camera is artificial, we can manipulate any of its intrinsic parameters, such as camera skew, to create perspective distortions. Sanjeev J. Koppal, Srinivasa G. Narasimhan |
ICCV | 2 |
| 2009 | Estimating natural illumination from a single outdoor imageabstractGiven a single outdoor image, we present a method for estimating the likely illumination conditions of the scene. In particular, we compute the probability distribution over the sun position and visibility. The method relies on a combination of weak cues that can be extracted from different portions of the image: the sky, the vertical surfaces, and the ground. While no single cue can reliably estimate illumination by itself, each one can reinforce the others to yield a more robust estimate. This is combined with a data-driven prior computed over a dataset of 6 million Internet photos. We present quantitative results on a webcam dataset with annotated sun positions, as well as qualitative results on consumer-grade photographs downloaded from Internet. Based on the estimated illumination, we show how to realistically insert synthetic 3-D objects into the scene. Jean-François Lalonde, Alexei A. Efros, Srinivasa G. Narasimhan |
ICCV | 3 |
| 2009 | Seeing through water: Image restoration using model-based trackingabstractA video sequence of an underwater scene taken from above the water surface suffers from severe distortions due to water fluctuations. In this paper, we simultaneously estimate the shape of the water surface and recover the planar underwater scene without using any calibration patterns, image priors, multiple viewpoints or active illumination. The key idea is to build a compact spatial distortion model of the water surface using the wave equation. Based on this model, we present a novel tracking technique that is designed specifically for water surfaces and addresses two unique challenges—the absence of an object model or template and the presence of complex appearance changes in the scene due to water fluctuation. We show the effectiveness of our approach on both simulated and real scenes, with text and texture. Yuandong Tian, Srinivasa G. Narasimhan |
ICCV | 2 |
| 2009 | The Theory and Practice of Coplanar Shadowgram Imaging for Acquiring Visual Hulls of Intricate Objects
Shuntaro Yamazaki, Srinivasa G. Narasimhan, Simon Baker, Takeo Kanade |
Int. J. Comput. Vis. | 2 |
| 2009 | Appearance Derivatives for Isonormal Clustering of ScenesabstractA new technique is proposed for scene analysis, called "appearance clustering." The key result of this approach is that the scene points can be clustered according to their surface normals, even when the geometry, material, and lighting are all unknown. This is achieved by analyzing an image sequence of a scene as it is illuminated by a smoothly moving distant light source. In such a scenario, the brightness measurements at each pixel form a "continuous appearance profile." When the source path follows an unstructured trajectory (obtained, say, by smoothly hand-waving a light source), the locations of the extrema of the appearance profile provide a strong cue for the scene point's surface normal. Based on this observation, a simple transformation of the appearance profiles and a distance metric are introduced that, together, can be used with any unsupervised clustering algorithm to obtain isonormal clusters of a scene. We support our algorithm empirically with comprehensive simulations of the Torrance-Sparrow and Oren-Nayar analytic BRDFs, as well as experiments with 25 materials obtained from the MERL database of measured BRDFs. The method is also demonstrated on 45 examples from the CURET database, obtaining clusters on scenes with real textures such as artificial grass and ceramic tile, as well as anisotropic materials such as satin and velvet. The results of applying our algorithm to indoor and outdoor scenes containing a variety of complex geometry and materials are shown. As an example application, isonormal clusters are used for lighting-consistent texture transfer. Our algorithm is simple and does not require any complex lighting setup for data collection. Sanjeev J. Koppal, Srinivasa G. Narasimhan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Webcam clip art: appearance and illuminant transfer from time-lapse sequencesabstractWebcams placed all over the world observe and record the visual appearance of a variety of outdoor scenes over long periods of time. The recorded time-lapse image sequences cover a wide range of illumination and weather conditions -- a vast untapped resource for creating visual realism. In this work, we propose to use a large repository of webcams as a "clip art" library from which users may transfer scene appearance (objects, scene backdrops, outdoor illumination) into their own time-lapse sequences or even single photographs. The goal is to combine the recent ideas from data-driven appearance transfer techniques with a general and theoretically-grounded physically-based illumination model. To accomplish this, the paper presents three main research contributions: 1) a new, high-quality outdoor webcam database that has been calibrated radiometrically and geometrically; 2) a novel approach for matching illuminations across different scenes based on the estimation of the properties of natural illuminants (sun, sky, weather and clouds), the camera geometry, and illumination-dependent scene features; 3) a new algorithm for generating physically plausible high dynamic range environment maps for each frame in a webcam sequence. Jean-François Lalonde, Alexei A. Efros, Srinivasa G. Narasimhan |
ACM Trans. Graph. | 3 |
| 2008 | On controlling light transport in poor visibility environmentsabstractPoor visibility conditions due to murky water, bad weather, dust and smoke severely impede the performance of vision systems. Passive methods have been used to restore scene contrast under moderate visibility by digital post-processing. However, these methods are ineffective when the quality of acquired images is poor to begin with. In this work, we design active lighting and sensing systems for controlling light transport before image formation, and hence obtain higher quality data. First, we present a technique of polarized light striping based on combining polarization imaging and structured light striping. We show that this technique out-performs different existing illumination and sensing methodologies. Second, we present a numerical approach for computing the optimal relative sensor-source position, which results in the best quality image. Our analysis accounts for the limits imposed by sensor noise. Mohit Gupta 0001, Srinivasa G. Narasimhan, Yoav Y. Schechner |
CVPR | 2 |
| 2008 | What Does the Sky Tell Us about the Camera?
Jean-François Lalonde, Srinivasa G. Narasimhan, Alexei A. Efros |
ECCV (4) | 2 |
| 2008 | Temporal Dithering of Illumination for Fast Active Vision
Srinivasa G. Narasimhan, Sanjeev J. Koppal, Shuntaro Yamazaki |
ECCV (4) | 1 |
| 2008 | Scattering
Diego Gutierrez, Srinivasa G. Narasimhan, Henrik Wann Jensen, Wojciech Jarosz |
SIGGRAPH ASIA Courses | 2 |
| 2007 | Novel Depth Cues from Uncalibrated Near-field LightingabstractWe present the first method to compute depth cues from images taken solely under uncalibrated near point lighting. A stationary scene is illuminated by a point source that is moved approximately along a line or in a plane. We observe the brightness profile at each pixel and demonstrate how to obtain three novel cues: plane-scene intersections, depth ordering and mirror symmetries. These cues are defined with respect to the line/plane in which the light source moves, and not the camera viewpoint. Plane-Scene Intersections are detected by finding those scene points that are closest to the light source path at some time instance. Depth Ordering for scenes with homogeneous BRDFs is obtained by sorting pixels according to their shortest distances from a plane containing the light source path. Mirror Symmetry pairs for scenes with homogeneous BRDFs are detected by reflecting scene points across a plane in which the light source moves. We show analytic results for Lambertian objects and demonstrate empirical evidence for a variety of other BRDFs. Sanjeev J. Koppal, Srinivasa G. Narasimhan |
ICCV | 2 |
| 2007 | Coplanar Shadowgrams for Acquiring Visual Hulls of Intricate ObjectsabstractAcquiring 3D models of intricate objects (like tree branches, bicycles and insects) is a hard problem due to severe self-occlusions, repeated thin structures and surface discontinuities. In theory, a shape-from-silhouettes (SFS) approach can overcome these difficulties and use many views to reconstruct visual hulls that are close to the actual shapes. In practice, however, SFS is highly sensitive to errors in silhouette contours and the calibration of the imaging system, and therefore not suitable for obtaining reliable shapes with a large number of views. We present a practical approach to SFS using a novel technique called coplanar shadowgram imaging, that allows us to use dozens to even hundreds of views for visual hull reconstruction. Here, a point light source is moved around an object and the shadows (silhouettes) cast onto a single background plane are observed. We characterize this imaging system in terms of image projection, reconstruction ambiguity, epipolar geometry, and shape and source recovery. The coplanarity of the shadowgrams yields novel geometric properties that are not possible in traditional multi-view camera- based imaging systems. These properties allow us to derive a robust and automatic algorithm to recover the visual hull of an object and the 3D positions of light source simultaneously, regardless of the complexity of the object. We demonstrate the acquisition of several intricate shapes with severe occlusions and thin structures, using 50 to 120 views. Shuntaro Yamazaki, Srinivasa G. Narasimhan, Simon Baker, Takeo Kanade |
ICCV | 2 |
| 2006 | Clustering Appearance for Scene AnalysisabstractWe propose a new approach called "appearance clustering" for scene analysis. The key idea in this approach is that the scene points can be clustered according to their surface normals, even when the geometry, material and lighting are all unknown. We achieve this by analyzing an image sequence of a scene as it is illuminated by a smoothly moving distant source. Each pixel thus gives rise to a "continuous appearance profile" that yields information about derivatives of the BRDF w.r.t source direction. This information is directly related to the surface normal of the scene point when the source path follows an unstructured trajectory (obtained, say, by "hand-waving"). Based on this observation, we transform the appearance profiles and propose a metric that can be used with any unsupervised clustering algorithm to obtain iso-normal clusters. We successfully demonstrate appearance clustering for complex indoor and outdoor scenes. In addition, iso-normal clusters serve as excellent priors for scene geometry and can strongly impact any vision algorithm that attempts to estimate material, geometry and/or lighting properties in a scene from images. We demonstrate this impact for applications such as diffuse and specular separation, both calibrated and uncalibrated photometric stereo of non-lambertian scenes, light source estimation and texture transfer. Sanjeev J. Koppal, Srinivasa G. Narasimhan |
CVPR (2) | 2 |
| 2006 | Acquiring scattering properties of participating media by dilutionabstractThe visual world around us displays a rich set of volumetric effects due to participating media. The appearance of these media is governed by several physical properties such as particle densities, shapes and sizes, which must be input (directly or indirectly) to a rendering algorithm to generate realistic images. While there has been significant progress in developing rendering techniques (for instance, volumetric Monte Carlo methods and analytic approximations), there are very few methods that measure or estimate these properties for media that are of relevance to computer graphics. In this paper, we present a simple device and technique for robustly estimating the properties of a broad class of participating media that can be either (a) diluted in water such as juices, beverages, paints and cleaning supplies, or (b) dissolved in water such as powders and sugar/salt crystals, or (c) suspended in water such as impurities. The key idea is to dilute the concentrations of the media so that single scattering effects dominate and multiple scattering becomes negligible, leading to a simple and robust estimation algorithm. Furthermore, unlike previous approaches that require complicated or separate measurement setups for different types or properties of media, our method and setup can be used to measure media with a complete range of absorption and scattering properties from a single HDR photograph. Once the parameters of the diluted medium are estimated, a volumetric Monte Carlo technique may be used to create renderings of any medium concentration and with multiple scattering. We have measured the scattering parameters of forty commonly found materials, that can be immediately used by the computer graphics community. We can also create realistic images of combinations or mixtures of the original measured materials, thus giving the user a wide flexibility in making realistic images of participating media. Srinivasa G. Narasimhan, Mohit Gupta 0001, Craig Donner, Ravi Ramamoorthi, Shree K. Nayar, Henrik Wann Jensen |
ACM Trans. Graph. | 1 |
| 2005 | Structured Light in Scattering MediaabstractVirtually all structured light methods assume that the scene and the sources are immersed in pure air and that light is neither scattered nor absorbed. Recently, however, structured lighting has found growing application in underwater and aerial imaging, where scattering effects cannot be ignored. In this paper, we present a comprehensive analysis of two representative methods - light stripe range scanning and photometric stereo - in the presence of scattering. For both methods, we derive physical models for the appearances of a surface immersed in a scattering medium. Based on these models, we present results on (a) the condition for object detectability in light striping and (b) the number of sources required for photometric stereo. In both cases, we demonstrate that while traditional methods fail when scattering is significant, our methods accurately recover the scene (depths, normals, albedos) as well as the properties of the medium. These results are in turn used to restore the appearances of scenes as if they were captured in clear air. Although we have focused on light striping and photometric stereo, our approach can also be extended to other methods such as grid coding, gated and active polarization imaging. Srinivasa G. Narasimhan, Shree K. Nayar, Sanjeev J. Koppal |
ICCV | 1 |
| 2005 | Enhancing Resolution Along Multiple Imaging Dimensions Using Assorted PixelsabstractMultisampled imaging is a general framework for using pixels on an image detector to simultaneously sample multiple dimensions of imaging (space, time, spectrum, brightness, polarization, etc.). The mosaic of red, green, and blue spectral filters found in most solid-state color cameras is one example of multisampled imaging. We briefly describe how multisampling can be used to explore other dimensions of imaging. Once such an image is captured, smooth reconstructions along the individual dimensions can be obtained using standard interpolation algorithms. Typically, this results in a substantial reduction of resolution (and, hence, image quality). One can extract significantly greater resolution in each dimension by noting that the light fields associated with real scenes have enormous redundancies within them, causing different dimensions to be highly correlated. Hence, multisampled images can be better interpolated using local structural models that are learned offline from a diverse set of training images. The specific type of structural models we use are based on polynomial functions of measured image intensities. They are very effective as well as computationally efficient. We demonstrate the benefits of structural interpolation using three specific applications. These are 1) traditional color imaging with a mosaic of color filters, 2) high dynamic range monochrome imaging using a mosaic of exposure filters, and 3) high dynamic range color imaging using a mosaic of overlapping color and exposure filters. Srinivasa G. Narasimhan, Shree K. Nayar |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | A practical analytic single scattering model for real time renderingabstractWe consider real-time rendering of scenes in participating media, capturing the effects of light scattering in fog, mist and haze. While a number of sophisticated approaches based on Monte Carlo and finite element simulation have been developed, those methods do not work at interactive rates. The most common real-time methods are essentially simple variants of the OpenGL fog model. While easy to use and specify, that model excludes many important qualitative effects like glows around light sources, the impact of volumetric scattering on the appearance of surfaces such as the diffusing of glossy highlights, and the appearance under complex lighting such as environment maps. In this paper, we present an alternative physically based approach that captures these effects while maintaining real time performance and the ease-of-use of the OpenGL fog model. Our method is based on an explicit analytic integration of the single scattering light transport equations for an isotropic point light source in a homogeneous participating medium. We can implement the model in modern programmable graphics hardware using a few small numerical lookup tables stored as texture maps. Our model can also be easily adapted to generate the appearances of materials with arbitrary BRDFs, environment map lighting, and precomputed radiance transfer methods, in the presence of participating media. Hence, our techniques can be widely used in real-time rendering. Ravi Ramamoorthi, Srinivasa G. Narasimhan, Shree K. Nayar |
ACM Trans. Graph. | 3 |
| 2003 | Shedding Light on the WeatherabstractVirtually all methods in image processing and computer vision, for removing weather effects from images, assume single scattering of light by particles in the atmosphere. In reality, multiple scattering effects are significant. A common manifestation of multiple scattering is the appearance of glows around light sources in bad weather. Modeling multiple scattering is critical to understanding the complex effects of weather on images, and hence essential for improving the performance of outdoor vision systems. We develop a new physics-based model for the multiple scattering of light rays as they travel from a source to an observer. This model is valid for various weather conditions including fog, haze, mist and rain. Our model enables us to recover from a single image the shapes and depths of sources in the scene. In addition, the weather condition and the visibility of the atmosphere can be estimated. These quantities can, in turn, be used to remove the glows of sources to obtain a clear picture of the scene. Based on these results, we demonstrate that a camera observing a distant source can serve as a "visual weather meter". The model and techniques described in this paper can also be used to analyze scattering in other media, such as fluids and tissues. Therefore, in addition to vision in bad weather, our work has implications for medical and underwater imaging. Srinivasa G. Narasimhan, Shree K. Nayar |
CVPR (1) | 1 |
| 2003 | A Class of Photometric Invariants: Separating Material from Shape and IlluminationabstractWe derive a new class of photometric invariants that can be used for a variety of vision tasks including lighting invariant material segmentation, change detection and tracking, as well as material invariant shape recognition. The key idea is the formulation of a scene radiance model for the class of "separable" BRDFs, that can be decomposed into material related terms and object shape and lighting related terms. All the proposed invariants are simple rational functions of the appearance parameters (say, material or shape and lighting). The invariants in this class differ from one another in the number and type of image measurements they require. Most of the invariants in this class need changes in illumination or object position between image acquisitions. The invariants can handle large changes in lighting which pose problems for most existing vision algorithms. We demonstrate the power of these invariants using scenes with complex shapes, materials, textures, shadows and specularities. Srinivasa G. Narasimhan, Visvanathan Ramesh, Shree K. Nayar |
ICCV | 1 |
| 2003 | Seeing Through Bad Weather
Shree K. Nayar, Srinivasa G. Narasimhan |
ISRR | 2 |
| 2003 | Contrast Restoration of Weather Degraded ImagesabstractImages of outdoor scenes captured in bad weather suffer from poor contrast. Under bad weather conditions, the light reaching a camera is severely scattered by the atmosphere. The resulting decay in contrast varies across the scene and is exponential in the depths of scene points. Therefore, traditional space invariant image processing techniques are not sufficient to remove weather effects from images. We present a physics-based model that describes the appearances of scenes in uniform bad weather conditions. Changes in intensities of scene points under different weather conditions provide simple constraints to detect depth discontinuities in the scene and also to compute scene structure. Then, a fast algorithm to restore scene contrast is presented. In contrast to previous techniques, our weather removal algorithm does not require any a priori scene structure, distributions of scene reflectances, or detailed knowledge about the particular weather condition. All the methods described in this paper are effective under a wide range of weather conditions including haze, mist, fog, and conditions arising due to other aerosols. Further, our methods can be applied to gray scale, RGB color, multispectral and even IR images. We also extend our techniques to restore contrast of scenes with moving objects, captured using a video camera. Srinivasa G. Narasimhan, Shree K. Nayar |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | All the Images of an Outdoor Scene
Srinivasa G. Narasimhan, Shree K. Nayar |
ECCV (3) | 1 |
| 2002 | Assorted Pixels: Multi-sampled Imaging with Structural Models
Shree K. Nayar, Srinivasa G. Narasimhan |
ECCV (4) | 2 |
| 2002 | Vision and the Atmosphere
Srinivasa G. Narasimhan, Shree K. Nayar |
Int. J. Comput. Vis. | 1 |
| 2001 | Removing Weather Effects from Monochrome ImagesabstractImages of outdoor scenes captured in bad weather suffer from poor contrast. Under bad weather conditions, the light reaching a camera is severely scattered by the atmosphere. The resulting decay in contrast varies across the scene and is exponential in the depths of scene points. Therefore, traditional space invariant image processing techniques are not sufficient to remove weather effects from images. In this paper, we present a fast physics-based method to compute scene structure and hence restore contrast of the scene from two or more images taken in bad weather In contrast to previous techniques, our method does not require any a priori weather-specific or scene information, and is effective under a wide range of weather conditions including haze, mist, fog and other aerosols. Further, our method can be applied to gray-scale, RGB color, multi-spectral and even IR images. We also extend the technique to restore contrast of scenes with moving objects, captured using a video camera. Srinivasa G. Narasimhan, Shree K. Nayar |
CVPR (2) | 1 |
| 2001 | Instant Dehazing of Images Using PolarizationabstractWe present an approach to easily remove the effects of haze from images. It is based on the fact that usually airlight scattered by atmospheric particles is partially polarized. Polarization filtering alone cannot remove the haze effects, except in restricted situations. Our method, however, works under a wide range of atmospheric and viewing conditions. We analyze the image formation process, taking into account polarization effects of atmospheric scattering. We then invert the process to enable the removal of haze from images. The method can be used with as few as two images taken through a polarizer at different orientations. This method works instantly, without relying on changes of weather conditions. We present experimental results of complete dehazing in far from ideal conditions for polarization filtering. We obtain a great improvement of scene contrast and correction of color. As a by product, the method also yields a range (depth) map of the scene, and information about properties of the atmospheric particles. Yoav Y. Schechner, Srinivasa G. Narasimhan, Shree K. Nayar |
CVPR (1) | 2 |
| 2000 | Chromatic Framework for Vision in Bad WeatherabstractConventional vision systems are designed to perform in clear weather. However, any outdoor vision system is incomplete without mechanisms that guarantee satisfactory performance under poor weather conditions. It is known that the atmosphere can significantly alter light energy reaching an observer. Therefore, atmospheric scattering models must be used to make vision systems robust in bad weather. In this paper, we develop a geometric framework for analyzing the chromatic effects of atmospheric scattering. First, we study a simple color model for atmospheric scattering and verify it for fog and haze. Then, based on the physics of scattering, we derive several geometric constraints on scene color changes, caused by varying atmospheric conditions. Finally, using these constraints we develop algorithms for computing fog or haze color depth segmentation, extracting three dimensional structure, and recovering "true" scene colors, from two or more images taken under different but unknown weather conditions. Srinivasa G. Narasimhan, Shree K. Nayar |
CVPR | 1 |
| 1999 | Vision in Bad WeatherabstractCurrent vision systems are designed to perform in clear weather. Needless to say, in any outdoor application, there is no escape from "bad" weather. Ultimately, computer vision systems must include mechanisms that enable them to function (even if somewhat less reliably) in the presence of haze, fog, rain, hail and snow. We begin by studying the visual manifestations of different weather conditions. For this, we draw on what is already known about atmospheric optics. Next, we identify effects caused by bad weather that can be turned to our advantage. Since the atmosphere modulates the information carried from a scene point to the observer it can be viewed as a mechanism of visual information coding. Based on this observation, we develop models and methods for recovering pertinent scene properties, such as three-dimensional structure, from images taken under poor weather conditions. Shree K. Nayar, Srinivasa G. Narasimhan |
ICCV | 2 |