VLDB 2026 Research / reviewers in the wild / expert
Aswin C. Sankaranarayanan
dblp:68/4264
· DBLP profile ↗
86ranked-venue papers
11as first author
28since 2021 · last 2025
0000-0003-0906-4046ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 72 · 9 first-author · 22 since 2021Artificial intelligence and machine learning · 43 · 3 first-author · 17 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PhotonSplat: 3D Scene Reconstruction and Colorization from SPAD SensorsabstractAdvances in 3D reconstruction using neural rendering have enabled high-quality 3D capture. However, they often fail when the input imagery is corrupted by motion blur, due to fast motion of the camera or the objects in the scene. This work advances neural rendering techniques in such scenarios by using single-photon avalanche diode (SPAD) arrays, an emerging sensing technology capable of sensing images at extremely high speeds. However, the use of SPADs presents its own set of unique challenges in the form of binary images, that are driven by stochastic photon arrivals. To address this, we introduce PhotonSplat, a framework designed to reconstruct 3D scenes directly from SPAD binary images, effectively navigating the noise vs. blur trade-off. Our approach incorporates a novel 3D spatial filtering technique to reduce noise in the renderings. The framework also supports both no-reference using generative priors and reference-based colorization from a single blurry image, enabling downstream applications such as segmentation, object detection and appearance editing tasks. Additionally, we extend our method to incorporate dynamic scene representations, making it suitable for scenes with moving objects. We further contribute PhotonScenes, a real-world multi-view dataset captured with the SPAD sensors. Code, data and video results are available at vinayak-vg.github.io/PhotonSplat/. Kuppa Sai Sri Teja, Sreevidya Chintalapati, Mukund Varma T., Haejoon Lee, Aswin C. Sankaranarayanan, Kaushik Mitra |
ICCP | 6 |
| 2025 | Spatially-Varying Autofocus
Yingsi Qin, Aswin C. Sankaranarayanan, Matthew O'Toole |
ICCV | 2 |
| 2025 | RT-X Net: RGB-Thermal Cross Attention Network for Low-Light Image EnhancementabstractIn nighttime conditions, high noise levels and bright illumination sources degrade image quality, making low-light image enhancement challenging. Thermal images provide complementary information, offering richer textures and structural details. We propose RT-X Net, a cross-attention network that fuses RGB and thermal images for nighttime image enhancement. We leverage self-attention networks for feature extraction and a cross-attention mechanism for fusion to effectively integrate information from both modalities. To support research in this domain, we introduce the Visible-Thermal Image Enhancement Evaluation (V-TIEE) dataset, comprising 50 co-located visible and thermal images captured under diverse nighttime conditions. Extensive evaluations on the publicly available LLVIP dataset and our V-TIEE dataset demonstrate that RT-X Net outperforms state-of-the-art methods in low-light image enhancement. The code and the V-TIEE can be found here https://github.com/jhakrraman/rt-xnet. Raman Jha, Adithya Lenka, Manikandasriram Srinivasan Ramanagopal, Aswin C. Sankaranarayanan, Kaushik Mitra |
ICIP | 4 |
| 2025 | Wide-Baseline Light Fields Using Ellipsoidal MirrorsabstractTraditional hand-held light field cameras only observe a small fraction of the cone of light emitted by a scene point. As a consequence, the study of interesting angular effects like iridescence are beyond the scope of such cameras. This paper envisions a new design for sensing light fields with wide baselines, so as to sense a significantly larger fraction of the cone of light emitted by scene points. Our system achieves this by imaging the scene, indirectly, through an ellipsoidal mirror. We show that an ellipsoidal mirror maps a wide cone of light from locations near one of its foci to a narrower cone at its other focus; thus, by placing a conventional light field camera at a focus, we can observe a wide-baseline light field from the scene near the other focus. We show via simulations and a lab prototype that wide-baseline light fields excel in the traditional applications involving changes in focus and perspective. Additionally, the larger cone of light that they observe allows the study of iridescence and thin-film interference. Perhaps surprisingly, the larger cone of light allows us to estimate surface normals of scene points by reasoning about their visibility. Michael DeZeeuw, Aswin C. Sankaranarayanan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | PS$^{2}$2 F: Polarized Spiral Point Spread Function for Single-Shot 3D SensingabstractWe propose a compact snapshot monocular depth estimation technique that relies on an engineered point spread function (PSF). Traditional approaches used in microscopic super-resolution imaging such as the Double-Helix PSF (DHPSF) are ill-suited for scenes that are more complex than a sparse set of point light sources. We show, using the Cramér-Rao lower bound, that separating the two lobes of the DHPSF and thereby capturing two separate images leads to a dramatic increase in depth accuracy. A special property of the phase mask used for generating the DHPSF is that a separation of the phase mask into two halves leads to a spatial separation of the two lobes. We leverage this property to build a compact polarization-based optical setup, where we place two orthogonal linear polarizers on each half of the DHPSF phase mask and then capture the resulting image with a polarization-sensitive camera. Results from simulations and a lab prototype demonstrate that our technique achieves up to $50\%$50% lower depth error compared to state-of-the-art designs including the DHPSF and the Tetrapod PSF, with little to no loss in spatial resolution. Bhargav Ghanekar, Vishwanath Saragadam, Dushyant Mehra, Anna-Karin Gustavsson, Aswin C. Sankaranarayanan, Ashok Veeraraghavan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Computational Imaging for Long-Term Prediction of Solar IrradianceabstractThe occlusion of the sun by clouds is one of the primary sources of uncertainties in solar power generation, and is a factor that affects the wide-spread use of solar power as a primary energy source. Real-time forecasting of cloud movement and, as a result, solar irradiance is necessary to schedule and allocate energy across grid-connected photovoltaic systems. Previous works monitored cloud movement using wide-angle field of view imagery of the sky. However, such images have poor resolution for clouds that appear near the horizon, which reduces their effectiveness for long term prediction of solar occlusion. Specifically, to be able to predict occlusion of the sun over long time periods, clouds that are near the horizon need to be detected, and their velocities estimated precisely. To enable such a system, we design and deploy a catadioptric system that delivers wide-angle imagery with uniform spatial resolution of the sky over its field of view. To enable prediction over a longer time horizon, we design an algorithm that uses carefully selected spatio-temporal slices of the imagery using estimated wind direction and velocity as inputs. Using ray-tracing simulations as well as a real testbed deployed outdoors, we show that the system is capable of predicting solar occlusion as well as irradiance for tens of minutes in the future, which is an order of magnitude improvement over prior work. Leron K. Julian, Haejoon Lee, Soummya Kar, Aswin C. Sankaranarayanan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Coherence as Texture - Passive Textureless 3D Reconstruction by Self-InterferenceabstractPassive depth estimation based on stereo or defocus relies on the presence of the texture on an object to resolve its depth. Hence, recovering the depth of a textureless object-for example, a large white wall-is not just hard but perhaps even impossible. Or is it? We show that spatial coherence, a property of natural light sources, can be used to resolve the depth of a scene point even when it is textureless. Our approach relies on the idea that natural light scattered off a scene point is locally coherent with itself, while incoherent with the light scattered from other surface points; we use this insight to design an optical setup that uses self-interference as a texture feature for estimating depth. Our lab prototype is capable of resolving depths of textureless objects in sunlight as well as indoor lights. Wei-Yu Chen, Aswin C. Sankaranarayanan, Anat Levin, Matthew O'Toole |
CVPR | 2 |
| 2024 | A Theory of Joint Light and Heat Transport for Lambertian ScenesabstractWe present a novel theory that establishes the relation-ship between light transport in visible and thermal infrared, and heat transport in solids. We show that heat generated due to light absorption can be estimated by modeling heat transport using a thermal camera. For situations where heat conduction is negligible, we analytically solve the heat transport equation to derive a simple expression relating the change in thermal image intensity to the absorbed light intensity and heat capacity of the material. Next, we prove that intrinsic image decomposition for Lambertian scenes becomes a well-posed problem if one has access to the ab-sorbed light. Our theory generalizes to arbitrary shapes and unstructured illumination. Our theory is based on ap-plying energy conservation principle at each pixel indepen-dently. We validate our theory using real-world experi-ments on diffuse objects made of different materials that ex-hibit both direct and global components (inter-reflections) of light transport under unknown complex lighting. Manikandasriram Srinivasan Ramanagopal, Sriram Narayanan, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan |
CVPR | 3 |
| 2024 | Projecting Trackable Thermal Patterns for Dynamic Computer VisionabstractAdding artificial patterns to objects, like QR codes, can ease tasks such as object tracking, robot navigation, and conveying information (e.g., a label or a website link). However, these patterns require a physical application and they alter the object's appearance. Conversely, projected patterns can temporarily change the object's appearance, aiding tasks like 3D scanning and retrieving object textures and shading. However, projected patterns impede dynamic tasks like object tracking because they do not ‘stick’ to the object's surface. Or do they? This paper introduces a novel approach combining the advantages of projected and persistent physical patterns. Our system projects heat patterns using a laser beam (similar in spirit to a LIDAR), which a thermal camera observes and tracks. Such thermal patterns enable tracking poorly-textured objects whose tracking is highly challenging with standard cameras while not affecting the object's appearance or physical properties. To avail these thermal patterns in existing vision frameworks, we train a network to reverse heat diffusion's effects and remove inconsistent pattern points between different thermal frames. We prototyped and tested this approach on dynamic vision tasks like structure from motion, optical flow, and object tracking of everyday textureless objects. Mark Sheinin, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan |
CVPR | 2 |
| 2024 | Spectral Subsurface Scattering for Material Classification
Haejoon Lee, Aswin C. Sankaranarayanan |
ECCV (5) | 2 |
| 2024 | Shape from Heat Conduction
Sriram Narayanan, Manikandasriram Srinivasan Ramanagopal, Mark Sheinin, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan |
ECCV (38) | 4 |
| 2024 | Scanning Iridescent ReflectanceabstractAcquisition of high-dimensional spatially-varying bidirectional reflectance distribution functions (SVBRDFs) is extremely challenging for scenes that exhibit high-frequency angular effects-in particular, iridescence-that requires dense sampling across spatial and angular dimensions of reflectance. This is further complicated by the need to illuminate and observe the material along grazing angles, where the effects of Iridescence are often most prominent. This paper proposes an imaging system for acquiring the reflectance associated with such iridescent materials. Our system uses an imaging setup, consisting of an ellipsoidal mirror and a light field camera, that can acquire a dense slice of the SVBRDF across spatial and observation angle axes from a single image, including at grazing observation angles. We use an active illumination setup to sparsely sample the incident illumination angles, thereby providing an acquisition setup that only requires a few (light field) photographs. To compensate for the lack of density in the incident angles, we train a neural model that implicitly represents the SVBRDF, creating a light-weight data-driven reflectance function that enables interpolation of the missing measurements. We show that our imaging system can reconstruct textures that exhibit spatially-varying iridescence stemming from diffraction as well as structural coloration. Michael DeZeeuw, Aswin C. Sankaranarayanan |
ICCP | 2 |
| 2023 | Neural Kaleidoscopic Space SculptingabstractWe introduce a method that recovers full-surround 3D reconstructions from a single kaleidoscopic image using a neural surface representation. Full-surround 3D reconstruction is critical for many applications, such as augmented and virtual reality. A kaleidoscope, which uses a single camera and multiple mirrors, is a convenient way of achieving full-surround coverage, as it redistributes light directions and thus captures multiple viewpoints in a single image. This enables single-shot and dynamic full-surround 3D reconstruction. However, using a kaleidoscopic image for multiview stereo is challenging, as we need to decompose the image into multi-view images by identifying which pixel corresponds to which virtual camera, a process we call labeling. To address this challenge, pur approach avoids the need to explicitly estimate labels, but instead “sculpts” a neural surface representation through the careful use of silhouette, background, foreground, and texture information present in the kaleidoscopic image. We demonstrate the advantages of our method in a range of simulated and real experiments, on both static and dynamic scenes. Byeongjoo Ahn, Michael DeZeeuw, Ioannis Gkioulekas, Aswin C. Sankaranarayanan |
CVPR | 4 |
| 2023 | Programmable Spectral Filter Arrays using Phase Spatial Light ModulatorsabstractComputational imaging has always benefited from tools that modulate light along the many dimensions of its plenoptic function. This paper provides a practical architecture for achieving spatially varying spectral modulation using a liquid crystal phase spatial light modulator (SLM). The use of a phase SLM, however, results in strong optical aberrations due to the unintended phase modulation, thereby precluding spectral modulation at high spatial resolutions. To mitigate this, we provide a careful and systematic analysis of the aberrations arising out of phase SLMs for the purpose of spatially varying spectral modulation; this analysis results in a dual strategy of “good patterns” that minimize the optical aberrations and a deep restoration network that overcomes any residual aberrations. We show a number of unique operating points with our prototype including single- and multi-image hyperspectral imaging, material classification (fewer than two images), and dynamic spectral filtering at video rates. Vishwanath Saragadam, Vijay Rengarajan, Ryuichi Tadano, Tuo Zhuang, Hideki Oyaizu, Jun Murayama, Aswin C. Sankaranarayanan |
ICCP | 7 |
| 2023 | Computational 3D Imaging with Position SensorsabstractUnderlying many structured light systems, especially those based on laser scanning, is a simple vision task: tracking a light spot. To accomplish this, scanners use conventional CMOS sensors to capture, transmit, and process millions of pixel measurements. This approach, while capable of achieving high-fidelity 3D scans, is wasteful in terms of (often scarce) sensing and computational resources. We present a structured light system based on position sensing diodes (PSDs), an unconventional sensing modality that directly measures the centroid of the spatial distribution of incident light, thus enabling high-resolution 3D laser scanning with a minimal amount of sensor data. We develop theory and computational algorithms for PSD-based structured light under a variety of light transport effects. We demonstrate the benefits of the proposed techniques using a hardware prototype on several real-world scenes, including optically-challenging objects with long-range interreflections and scattering. Jeremy Klotz, Mohit Gupta 0001, Aswin C. Sankaranarayanan |
ICCV | 3 |
| 2023 | Designing Phase Masks for Under-Display CamerasabstractDiffractive blur and low light levels are two fundamental challenges in producing high-quality photographs in under-display cameras (UDCs). In this paper, we incorporate phase masks on display panels to tackle both challenges. Our design inserts two phase masks, specifically two microlens arrays, in front of and behind a display panel. The first phase mask concentrates light on the locations where the display is transparent so that more light passes through the display, and the second phase mask reverts the effect of the first phase mask. We further optimize the folding height of each microlens to improve the quality of PSFs and suppress chromatic aberration. We evaluate our design using a physically-accurate simulator based on Fourier optics. The proposed design is able to double the light throughput while improving the invertibility of the PSFs. Lastly, we discuss the effect of our design on the display quality and show that implementation with polarization-dependent phase masks can leave the display quality uncompromised. Anqi Yang, Eunhee Kang, Hyong-Euk Lee, Aswin C. Sankaranarayanan |
ICCV | 4 |
| 2023 | Split-Lohmann Multifocal DisplaysabstractThis work provides the design of a multifocal display that can create a dense stack of focal planes in a single shot. We achieve this using a novel computational lens that provides spatial selectivity in its focal length, i.e, the lens appears to have different focal lengths across points on a display behind it. This enables a multifocal display via an appropriate selection of the spatially-varying focal length, thereby avoiding time multiplexing techniques that are associated with traditional focus tunable lenses. The idea central to this design is a modification of a Lohmann lens, a focus tunable lens created with two cubic phase plates that translate relative to each other. Using optical relays and a phase spatial light modulator, we replace the physical translation of the cubic plates with an optical one, while simultaneously allowing for different pixels on the display to undergo different amounts of translations and, consequently, different focal lengths. We refer to this design as a Split-Lohmann multifocal display. Split-Lohmann displays provide a large étendue as well as high spatial and depth resolutions; the absence of time multiplexing and the extremely light computational footprint for content processing makes it suitable for video and interactive experiences. Using a lab prototype, we show results over a wide range of static, dynamic, and interactive 3D scenes, showcasing high visual quality over a large working range. Yingsi Qin, Wei-Yu Chen, Matthew O'Toole, Aswin C. Sankaranarayanan |
ACM Trans. Graph. | 4 |
| 2022 | Single-Photon Structured LightabstractWe present a novel structured light technique that uses Single Photon Avalanche Diode (SPAD) arrays to enable 3D scanning at high-frame rates and low-light levels. This technique, called “Single-Photon Structured Light”, works by sensing binary images that indicates the presence or absence of photon arrivals during each exposure; the SPAD array is used in conjunction with a high-speed binary projector, with both devices operated at speeds as high as 20 kHz. The binary images that we acquire are heavily influenced by photon noise and are easily corrupted by ambient sources of light. To address this, we develop novel temporal sequences using error correction codes that are designed to be robust to short-range effects like projector and camera defocus as well as resolution mismatch between the two devices. Our lab prototype is capable of 3D imaging in challenging scenarios involving objects with extremely low albedo or undergoing fast motion, as well as scenes under strong ambient illumination. Varun Sundar, Sizhuo Ma, Aswin C. Sankaranarayanan, Mohit Gupta 0001 |
CVPR | 3 |
| 2022 | Computational Imaging using Ultrasonically-Sculpted Virtual LensesabstractUltrasonically-sculpted waveguides provide interesting opportunities for in situ optical imaging in transparent and scattering media. The interference of ultrasonic waves can be designed to form a spatially-varying refractive index pattern in the target medium, acting as a virtual lens to guide light and relay images. The images formed by such lenses are subject to a large amount of spatially-varying blur, which significantly reduces their contrast. To alleviate this issue, the images can be computationally deblurred post experiment to restore the image. First, we demonstrate a brute force deconvolution technique to deblur the relayed images. While effective, this method proves to be computationally intensive due to the spatially-varying blur kernel. The reconfigurability of ultrasonically sculpted waveguides can be leveraged to address this issue by measuring line integrals of the image at multiple angles to form the Radon image of the target, which can be deblurred very efficiently with a simple linear model. We validate this method using simulated and experimental results. The tantalizing notion of using the reconfigurable ultrasonically-sculpted waveguides as part of the imaging systems, demonstrated in this paper, opens new opportunities for hybrid physical-computational optical systems. Hossein Baktash, Yash Belhe, Matteo Giuseppe Scopelliti, Aswin C. Sankaranarayanan, Maysamreza Chamanzar |
ICCP | 5 |
| 2022 | Exponentially-wide etendue displays using a tilting cascadeabstractA fundamental limitation of spatial light modulation (SLM) devices is that their etendue, defined as the product of the display's size and angular range, is bounded by the number of pixel units. Current SLMs are woefully inadequate in meeting the spatial and angular ranges of many real applications. In particular using these SLMs to generate realistic computer-controlled holographic displays would require scaling the number of display units by a few orders of magnitude. In this work, we suggest that rather than excessively increasing the pixel count, etendue can be expanded by augmenting the display units with tilting capabilities. Furthermore, we show that tiltable displays can be realized using a cascade of binary tilt layers, each capable of tilting the light towards one of two orientations. With proper design, the etendue-expansion factor scales exponentially with the number of display layers; hence, a very small number of such layers can effectively realize wide expansion factors. We implement a proof of concept display and demonstrate its applicability for displaying multi-view or holographic content with increased size and angular field-of-view. Sagi Monin, Aswin C. Sankaranarayanan, Anat Levin |
ICCP | 2 |
| 2022 | Analyzing phase masks for wide étendue holographic displaysabstractSpatial light modulator (SLM) technology forms the centerpiece of digital holographic displays. However, an inherent limitation of these devices is that their etendue, defined as the product of the display's eye box and field of view, is bounded by the number of pixel units. As a consequence, current SLMs are far from meeting the required field-of-view and eye box for the human visual system, which would require scaling the number of display units by a few orders of magnitude. Existing strategies for etendue-expansion rely on introducing a diffractive optical element (DOE), a fixed random phase mask whose pitch is much smaller than that of the original display, thereby spreading light over a wider angle. Displayed content is then optimized under perceptual constraints on the generated image. However, since the phase mask is fixed, the number of degrees of freedom does not increase and hence, the expansion in etendue necessarily comes with a loss of image quality. The tradeofts involved with such phase masks are not well understood. This paper studies the space of phase masks that can be attached to an SLM to increase its angular range. It attempts to characterize what trade-offs are involved in etendue-expansion, and whatever specific phase mask designs would support better holograms. Our theoretical results show that etendue expansion comes with a commensurate loss of contrast or resolution, depending on the specifics of the mask that we use. We show that while pseudo random masks support wide-etendue, they involve an inherent loss of contrast. Perhaps surprisingly, simple commonly-available phase masks like lenslet arrays provide near-optimal results that can largely outperform random masks. Sagi Monin, Aswin C. Sankaranarayanan, Anat Levin |
ICCP | 2 |
| 2022 | Exploring mmWave Radar and Camera Fusion for High-Resolution and Long-Range Depth ImagingabstractRobotic geo-fencing and surveillance systems require accurate monitoring of objects if/when they violate perimeter restrictions. In this paper, we seek a solution for depth imaging of such objects of interest at high accuracy (few tens of cm) over extended ranges (up to 300 meters) from a single vantage point, such as a pole mounted platform. Unfortunately, the rich literature in depth imaging using camera, lidar and radar in isolation struggles to meet these tight requirements in real-world conditions. This paper proposes Metamoran, a solution that explores long-range depth imaging of objects of interest by fusing the strengths of two complementary technologies: mmWave radar and camera. Unlike cameras, mmWave radars offer excellent cm-scale depth resolution even at very long ranges. However, their angular resolution is at least 10x worse than camera systems. Fusing these two modalities is natural, but in scenes with high clutter and at long ranges, radar reflections are weak and experience spurious artifacts. Metamoran's core contribution is to leverage image segmentation and monocular depth estimation on camera images to help declutter radar and discover true object reflections. We perform a detailed evaluation of Metamoran's depth imaging capabilities in 400 diverse scenarios. Our evaluation shows that Metamoran estimates the depth of static objects up to 90 m away and moving objects up to 305 m away and with a median error of 28 cm, an improvement of 13 x over a naive radar+camera baseline and 23 x compared to monocular depth estimation. Akarsh Prabhakara, Diana Zhang, Sirajum Munir, Aswin C. Sankaranarayanan, Anthony Rowe 0001, Swarun Kumar |
IROS | 5 |
| 2021 | Reference Wave Design for Wavefront SensingabstractOne of the classical results in wavefront sensing is phase-shifting point diffraction interferometry (PS-PDI), where the phase of a wavefront is measured by interfering it with a planar reference created from the incident wave itself. The limiting drawback of this approach is that the planar reference, often created by passing light through a narrow pinhole, is dim and noise sensitive. We address this limitation with a novel approach called ReWave that uses a non-planar reference that is designed to be brighter. The reference wave is designed in a specific way that would still allow for analytic phase recovery, exploiting ideas of sparse phase retrieval algorithms. ReWave requires only four image intensity measurements and is significantly more robust to noise compared to PS-PDI. We validate the robustness and applicability of our approach using a suite of simulated and real results. Wei-Yu Chen, Anat Levin, Matthew O'Toole, Aswin C. Sankaranarayanan |
ICCP | 4 |
| 2021 | A Simple Framework for 3D Lensless Imaging with Programmable MasksabstractLensless cameras provide a framework to build thin imaging systems by replacing the lens in a conventional camera with an amplitude or phase mask near the sensor. Existing methods for lensless imaging can recover the depth and intensity of the scene, but they require solving computationally-expensive inverse problems. Furthermore, existing methods struggle to recover dense scenes with large depth variations. In this paper, we propose a lensless imaging system that captures a small number of measurements using different patterns on a programmable mask. In this context, we make three contributions. First, we present a fast recovery algorithm to recover textures on a fixed number of depth planes in the scene. Second, we consider the mask design problem, for programmable lensless cameras, and provide a design template for optimizing the mask patterns with the goal of improving depth estimation. Third, we use a refinement network as a post-processing step to identify and remove artifacts in the reconstruction. These modifications are evaluated extensively with experimental results on a lensless camera prototype to showcase the performance benefits of the optimized masks and recovery algorithms over the state of the art. Yucheng Zheng, Aswin C. Sankaranarayanan, Muhammad Salman Asif |
ICCV | 3 |
| 2021 | SliceNets - A Scalable Approach for Object Detection in 3D CT ScansabstractOne of the most promising approaches for automated detection of guns and other prohibited items in aviation baggage screening is the use of 3D computed tomography (CT) scans. However, automated detection, especially with deep neural networks, faces two key challenges: the high dimensionality of individual 3D scans, and the lack of labelled training data. We address these challenges using a novel image-based detection and segmentation technique that we call the slice-and-fuse framework. Our approach relies on slicing the input 3D volumes, generating 2D predictions on each slice using 2D Convolutional Neural Networks (CNNs), and fusing them to obtain a 3D prediction. We develop two distinct detectors based on this slice-and-fuse strategy: the Retinal-SliceNet that uses a unified, single network with end-to-end training, and the U-SliceNet that uses a two-stage paradigm, first generating proposals using a voxel labeling network and, subsequently, refining the proposals by a 3D classification network. The networks are trained using a data augmentation approach that creates a very large training dataset by inserting weapons into 3D CT scans of threat-free bags. We demonstrate that the two SliceNets outperform state-of-the-art methods on a large-scale 3D baggage CT dataset for baggage classification, 3D object detection, and 3D semantic segmentation. Anqi Yang, Vishwanath Saragadam, Duy Dao, Zhuo Hui, Jen-Hao Rick Chang, Aswin C. Sankaranarayanan |
WACV | 7 |
| 2021 | SASSI - Super-Pixelated Adaptive Spatio-Spectral ImagingabstractWe introduce a novel video-rate hyperspectral imager with high spatial, temporal and spectral resolutions. Our key hypothesis is that spectral profiles of pixels within each super-pixel tend to be similar. Hence, a scene-adaptive spatial sampling of a hyperspectral scene, guided by its super-pixel segmented image, is capable of obtaining high-quality reconstructions. To achieve this, we acquire an RGB image of the scene, compute its super-pixels, from which we generate a spatial mask of locations where we measure high-resolution spectrum. The hyperspectral image is subsequently estimated by fusing the RGB image and the spectral measurements using a learnable guided filtering approach. Due to low computational complexity of the superpixel estimation step, our setup can capture hyperspectral images of the scenes with little overhead over traditional snapshot hyperspectral cameras, but with significantly higher spatial and spectral resolutions. We validate the proposed technique with extensive simulations as well as a lab prototype that measures hyperspectral video at a spatial resolution of 600 ×900 pixels, at a spectral resolution of 10 nm over visible wavebands, and achieving a frame rate at 18fps. Vishwanath Saragadam, Michael DeZeeuw, Richard G. Baraniuk, Ashok Veeraraghavan, Aswin C. Sankaranarayanan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Designing Display Pixel Layouts for Under-Panel CamerasabstractUnder-panel cameras provide an intriguing way to maximize the display area for a mobile device. An under-panel camera images a scene via the openings in the display panel; hence, a captured photograph is noisy as well as endowed with a large diffractive blur as the display acts as an aperture on the lens. Unfortunately, the pattern of openings commonly found in current LED displays are not conducive to high-quality deblurring. This paper redesigns the layout of openings in the display to engineer a blur kernel that is robustly invertible in the presence of noise. We first provide a basic analysis using Fourier optics that indicates that the nature of the blur is critically affected by the periodicity of the display openings as well as the shape of the opening at each individual display pixel. Armed with this insight, we provide a suite of modifications to the pixel layout that promote the invertibility of the blur kernels. We evaluate the proposed layouts with photomasks placed in front of a cellphone camera, thereby emulating an under-panel camera. A key takeaway is that optimizing the display layout does indeed produce significant improvements. Anqi Yang, Aswin C. Sankaranarayanan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Kaleidoscopic structured lightabstractFull surround 3D imaging for shape acquisition is essential for generating digital replicas of real-world objects. Surrounding an object we seek to scan with a kaleidoscope, that is, a configuration of multiple planar mirrors, produces an image of the object that encodes information from a combinatorially large number of virtual viewpoints. This information is practically useful for the full surround 3D reconstruction of the object, but cannot be used directly, as we do not know what virtual viewpoint each image pixel corresponds---the pixel label. We introduce a structured light system that combines a projector and a camera with a kaleidoscope. We then prove that we can accurately determine the labels of projector and camera pixels, for arbitrary kaleidoscope configurations, using the projector-camera epipolar geometry. We use this result to show that our system can serve as a multi-view structured light system with hundreds of virtual projectors and cameras. This makes our system capable of scanning complex shapes precisely and with full coverage. We demonstrate the advantages of the kaleidoscopic structured light system by scanning objects that exhibit a large range of shapes and reflectances. Byeongjoo Ahn, Ioannis Gkioulekas, Aswin C. Sankaranarayanan |
ACM Trans. Graph. | 3 |
| 2020 | 3PointTM: Faster Measurement of High-Dimensional Transmission Matrices
Yujun Chen, Ashutosh Sabharwal, Ashok Veeraraghavan, Aswin C. Sankaranarayanan |
ECCV (8) | 5 |
| 2020 | FreeCam3D: Snapshot Structured Light 3D with Freely-Moving Cameras
Vivek Boominathan, Jacob T. Robinson, Hiroshi Kawasaki, Aswin C. Sankaranarayanan, Ashok Veeraraghavan |
ECCV (27) | 6 |
| 2020 | Programmable Spectrometry: Per-pixel Material Classification using Learned Spectral FiltersabstractMany materials have distinct spectral profiles, which facilitates estimation of the material composition of a scene by processing its hyperspectral image (HSI). However, this process is inherently wasteful since high-dimensional HSIs are expensive to acquire and only a set of linear projections of the HSI contribute to the classification task. This paper proposes the concept of programmable spectrometry for per-pixel material classification, where instead of sensing the HSI of the scene and then processing it, we optically compute the spectrally-filtered images. This is achieved using a computational camera with a programmable spectral response. Our approach provides gains both in terms of acquisition speed - since only the relevant measurements are acquired - and in signal-to-noise ratio - since we invariably avoid narrowband filters that are light inefficient. Given ample training data, we use learning techniques to identify the bank of spectral profiles that facilitate material classification. We verify the method in simulations, as well as validate our findings using a lab prototype of the camera. Vishwanath Saragadam, Aswin C. Sankaranarayanan |
ICCP | 2 |
| 2020 | SweepCam - Depth-Aware Lensless Imaging Using Programmable MasksabstractLensless cameras, while extremely useful for imaging in constrained scenarios, struggle with resolving scenes with large depth variations. To resolve this, we propose imaging with a set of mask patterns displayed on a programmable mask, and introduce a computational focusing operator that helps to resolve the depth of scene points. As a result, the proposed imager can resolve dense scenes with large depth variations, allowing for more practical applications of lensless cameras. We also present a fast reconstruction algorithm for scene at multiple depths that reduces reconstruction time by two orders of magnitude. Finally, we build a prototype to show the proposed method improves both image quality and depth resolution of lensless cameras. Shigeki Nakamura, Muhammad Salman Asif, Aswin C. Sankaranarayanan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Towards occlusion-aware multifocal displaysabstractThe human visual system uses numerous cues for depth perception, including disparity, accommodation, motion parallax and occlusion. It is incumbent upon virtual-reality displays to satisfy these cues to provide an immersive user experience. Multifocal displays, one of the classic approaches to satisfy the accommodation cue, place virtual content at multiple focal planes, each at a different depth. However, the content on focal planes close to the eye do not occlude those farther away; this deteriorates the occlusion cue as well as reduces contrast at depth discontinuities due to leakage of the defocus blur. This paper enables occlusion-aware multifocal displays using a novel ConeTilt operator that provides an additional degree of freedom --- tilting the light cone emitted at each pixel of the display panel. We show that, for scenes with relatively simple occlusion configurations, tilting the light cones provides the same effect as physical occlusion. We demonstrate that ConeTilt can be easily implemented by a phase-only spatial light modulator. Using a lab prototype, we show results that demonstrate the presence of occlusion cues and the increased contrast of the display at depth edges. Jen-Hao Rick Chang, Anat Levin, B. V. K. Vijaya Kumar, Aswin C. Sankaranarayanan |
ACM Trans. Graph. | 4 |
| 2019 | Learning to Separate Multiple Illuminants in a Single ImageabstractWe present a method to separate a single image captured under two illuminants, with different spectra, into the two images corresponding to the appearance of the scene under each individual illuminant. We do this by training a deep neural network to predict the per-pixel reflectance chromaticity of the scene, which we use in conjunction with a previous flash/no-flash image-based separation algorithm to produce the final two output images. We design our reflectance chromaticity network and loss functions by incorporating intuitions from the physics of image formation. We show that this leads to significantly better performance than other single image techniques and even approaches the quality of the two image separation method. Zhuo Hui, Ayan Chakrabarti, Kalyan Sunkavalli, Aswin C. Sankaranarayanan |
CVPR | 4 |
| 2019 | Beyond Volumetric Albedo - A Surface Optimization Framework for Non-Line-Of-Sight ImagingabstractNon-line-of-sight (NLOS) imaging is the problem of reconstructing properties of scenes occluded from a sensor, using measurements of light that indirectly travels from the occluded scene to the sensor through intermediate diffuse reflections. We introduce an analysis-by-synthesis framework that can reconstruct complex shape and reflectance of an NLOS object. Our framework deviates from prior work on NLOS reconstruction, by directly optimizing for a surface representation of the NLOS object, in place of commonly employed volumetric representations. At the core of our framework is a new rendering formulation that efficiently computes derivatives of radiometric measurements with respect to NLOS geometry and reflectance, while accurately modeling the underlying light transport physics. By coupling this with stochastic optimization and geometry processing techniques, we are able to reconstruct NLOS surface at a level of detail significantly exceeding what is possible with previous volumetric reconstruction methods. Chia-Yin Tsai, Aswin C. Sankaranarayanan, Ioannis Gkioulekas |
CVPR | 2 |
| 2019 | A Theory of Fermat Paths for Non-Line-Of-Sight Shape ReconstructionabstractWe present a novel theory of Fermat paths of light between a known visible scene and an unknown object not in the line of sight of a transient camera. These light paths either obey specular reflection or are reflected by the object's boundary, and hence encode the shape of the hidden object. We prove that Fermat paths correspond to discontinuities in the transient measurements. We then derive a novel constraint that relates the spatial derivatives of the path lengths at these discontinuities to the surface normal. Based on this theory, we present an algorithm, called Fermat Flow, to estimate the shape of the non-line-of-sight object. Our method allows, for the first time, accurate shape recovery of complex objects, ranging from diffuse to specular, that are hidden around the corner as well as hidden behind a diffuser. Finally, our approach is agnostic to the particular technology used for transient imaging. As such, we demonstrate mm-scale shape recovery from pico-second scale transients using a SPAD and ultrafast laser, as well as micron-scale reconstruction from femto-second scale transients using interferometry. We believe our work is a significant advance over the state-of-the-art in non-line-of-sight imaging. Shumian Xin, Sotiris Nousias, Kiriakos N. Kutulakos, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan, Ioannis Gkioulekas |
CVPR | 4 |
| 2019 | Wavelet Tree Parsing with Freeform LensingabstractWe propose an architecture for adaptive sensing of images by progressively measuring its wavelet coefficients. Our approach, commonly referred to as wavelet tree parsing, adaptively selects the specific wavelet coefficients to be sensed by modeling the children of dominant coefficients to be dominant themselves. A key challenge for practical implementation of this technique is that the wavelet patterns, especially at finer scales, occupy a tiny portion of the field of view and, hence, the resulting measurements have very poor light levels and signal-to-noise ratios (SNR). To address this, we propose a novel imaging architecture that uses a phase-only spatial light modulator as a freeform lens to concentrate a light source and create the wavelet patterns. This ensures that the SNR of measurements remain constant across different spatial scales. Using a lab prototype, we demonstrate successful reconstruction on a wide range of real scenes and show that concentrating illumination enables us to outperform non-adaptive techniques as well as adaptive techniques based on traditional projectors. Vishwanath Saragadam, Aswin C. Sankaranarayanan |
ICCP | 2 |
| 2019 | PhaseCam3D - Learning Phase Masks for Passive Single View Depth EstimationabstractThere is an increasing need for passive 3D scanning in many applications that have stringent energy constraints. In this paper, we present an approach for single frame, single viewpoint, passive 3D imaging using a phase mask at the aperture plane of a camera. Our approach relies on an end-to-end optimization framework to jointly learn the optimal phase mask and the reconstruction algorithm that allows an accurate estimation of range image from captured data. Using our optimization framework, we design a new phase mask that performs significantly better than existing approaches. We build a prototype by inserting a phase mask fabricated using photolithography into the aperture plane of a conventional camera and show compelling performance in 3D imaging. Vivek Boominathan, Huaijin G. Chen, Aswin C. Sankaranarayanan, Ashok Veeraraghavan |
ICCP | 4 |
| 2019 | Convolutional Approximations to the General Non-Line-of-Sight Imaging OperatorabstractNon-line-of-sight (NLOS) imaging aims to reconstruct scenes outside the field of view of an imaging system. A common approach is to measure the so-called light transients, which facilitates reconstructions through ellipsoidal tomography that involves solving a linear least-squares. Unfortunately, the corresponding linear operator is very high-dimensional and lacks structures that facilitate fast solvers, and so, the ensuing optimization is a computationally daunting task. We introduce a computationally tractable framework for solving the ellipsoidal tomography problem. Our main observation is that the Gram of the ellipsoidal tomography operator is convolutional, either exactly under certain idealized imaging conditions, or approximately in practice. This, in turn, allows us to obtain the ellipsoidal tomography solution by using efficient deconvolution procedures to solve a linear least-squares problem involving the Gram operator. The computational tractability of our approach also facilitates the use of various regularizers during the deconvolution procedure. We demonstrate the advantages of our framework in a variety of simulated and real experiments. Byeongjoo Ahn, Akshat Dave, Ashok Veeraraghavan, Ioannis Gkioulekas, Aswin C. Sankaranarayanan |
ICCV | 5 |
| 2019 | Cross-Scale Predictive DictionariesabstractSparse representations using data dictionaries provide an efficient model particularly for signals that do not enjoy alternate analytic sparsifying transformations. However, solving inverse problems with sparsifying dictionaries can be computationally expensive, especially when the dictionary under consideration has a large number of atoms. In this paper, we incorporate additional structure on to dictionary-based sparse representations for visual signals to enable speedups when solving sparse approximation problems. The specific structure that we endow onto sparse models is that of a multi-scale modeling where the sparse representation at each scale is constrained by the sparse representation at coarser scales. We show that this cross-scale predictive model delivers significant speedups, often in the range of , with little loss in accuracy for linear inverse problems associated with images, videos, and light fields. Vishwanath Saragadam, Xin Li 0001, Aswin C. Sankaranarayanan |
IEEE Trans. Image Process. | 3 |
| 2019 | KRISM - Krylov Subspace-based Optical Computing of Hyperspectral ImagesabstractWe present an adaptive imaging technique that optically computes a low-rank approximation of a scene’s hyperspectral image, conceptualized as a matrix. Central to the proposed technique is the optical implementation of two measurement operators: a spectrally coded imager and a spatially coded spectrometer. By iterating between the two operators, we show that the top singular vectors and singular values of a hyperspectral image can be adaptively and optically computed with only a few iterations. We present an optical design that uses pupil plane coding for implementing the two operations and show several compelling results using a lab prototype to demonstrate the effectiveness of the proposed hyperspectral imager. Vishwanath Saragadam, Aswin C. Sankaranarayanan |
ACM Trans. Graph. | 2 |
| 2018 | Illuminant Spectra-Based Source Separation Using Flash PhotographyabstractReal-world lighting often consists of multiple illuminants with different spectra. Separating and manipulating these illuminants in post-process is a challenging problem that requires either significant manual input or calibrated scene geometry and lighting. In this work, we leverage a flash/no-flash image pair to analyze and edit scene illuminants based on their spectral differences. We derive a novel physics-based relationship between color variations in the observed flash/no-flash intensities and the spectra and surface shading corresponding to individual scene illuminants. Our technique uses this constraint to automatically separate an image into constituent images lit by each illuminant. This separation can be used to support applications like white balancing, lighting editing, and RGB photometric stereo, where we demonstrate results that outperform state-of-the-art techniques on a wide range of images. Zhuo Hui, Kalyan Sunkavalli, Sunil Hadap, Aswin C. Sankaranarayanan |
CVPR | 4 |
| 2018 | Programmable Triangulation Light Curtains
Jian Wang 0100, Joseph R. Bartels, William Whittaker, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan |
ECCV (3) | 4 |
| 2018 | Towards multifocal displays with dense focal stacksabstractWe present a virtual reality display that is capable of generating a dense collection of depth/focal planes. This is achieved by driving a focus-tunable lens to sweep a range of focal lengths at a high frequency and, subsequently, tracking the focal length precisely at microsecond time resolutions using an optical module. Precise tracking of the focal length, coupled with a high-speed display, enables our lab prototype to generate 1600 focal planes per second. This enables a novel first-of-its-kind virtual reality multifocal display that is capable of resolving the vergence-accommodation conflict endemic to today's displays. Jen-Hao Rick Chang, B. V. K. Vijaya Kumar, Aswin C. Sankaranarayanan |
ACM Trans. Graph. | 3 |
| 2017 | The Geometry of First-Returning Photons for Non-Line-of-Sight ImagingabstractNon-line-of-sight (NLOS) imaging utilizes the full 5D light transient measurements to reconstruct scenes beyond the cameras field of view. Mathematically, this requires solving an elliptical tomography problem that unmixes the shape and albedo from spatially-multiplexed measurements of the NLOS scene. In this paper, we propose a new approach for NLOS imaging by studying the properties of first-returning photons from three-bounce light paths. We show that the times of flight of first-returning photons are dependent only on the geometry of the NLOS scene and each observation is almost always generated from a single NLOS scene point. Exploiting these properties, we derive a space carving algorithm for NLOS scenes. In addition, by assuming local planarity, we derive an algorithm to localize NLOS scene points in 3D and estimate their surface normals. Our methods do not require either the full transient measurements or solving the hard elliptical tomography problem. We demonstrate the effectiveness of our methods through simulations as well as real data captured from a SPAD sensor. Chia-Yin Tsai, Kiriakos N. Kutulakos, Srinivasa G. Narasimhan, Aswin C. Sankaranarayanan |
CVPR | 4 |
| 2017 | Compressive spectral anomaly detectionabstractWe propose a novel compressive imager for detecting anomalous spectral profiles in a scene. We model the background spectrum as a low-dimensional subspace while assuming the anomalies to form a spatially-sparse set of spectral profiles different from the background. Our core contributions are in the form of a two-stage sensing mechanism. In the first stage, we estimate the subspace for the background spectrum by acquiring spectral measurements at a few randomly-selected pixels. In the second stage, we acquire spatially-multiplexed spectral measurements of the scene. We remove the contributions of the background spectrum from the spatially-multiplexed measurements by projecting onto the complementary subspace of the background spectrum; the resulting measurements are of a sparse matrix that encodes the presence and spectra of anomalies, which can be recovered using a Multiple Measurement Vector formulation. Theoretical analysis and simulations show significant speed up in acquisition time over other anomaly detection techniques. A lab prototype based on a DMD and a visible spectrometer validates our proposed imager. Vishwanath Saragadam, Jian Wang 0100, Xin Li 0001, Aswin C. Sankaranarayanan |
ICCP | 4 |
| 2017 | Reflectance Capture Using Univariate Sampling of BRDFsabstractWe propose the use of a light-weight setup consisting of a collocated camera and light source – commonly found on mobile devices – to reconstruct surface normals and spatially-varying BRDFs of near-planar material samples. A collocated setup provides only a 1-D “univariate” sampling of a 3-D isotropic BRDF. We show that a univariate sampling is sufficient to estimate parameters of commonly used analytical BRDF models. Subsequently, we use a dictionary-based reflectance prior to derive a robust technique for per-pixel normal and BRDF estimation. We demonstrate real-world shape and capture, and its application to material editing and classification, using real data acquired using a mobile phone. Zhuo Hui, Kalyan Sunkavalli, Joon-Young Lee, Sunil Hadap, Jian Wang 0100, Aswin C. Sankaranarayanan |
ICCV | 6 |
| 2017 | Shape and Spatially-Varying Reflectance Estimation from Virtual ExemplarsabstractThis paper addresses the problem of estimating the shape of objects that exhibit spatially-varying reflectance. We assume that multiple images of the object are obtained under a fixed view-point and varying illumination, i.e., the setting of photometric stereo. At the core of our techniques is the assumption that the BRDF at each pixel lies in the non-negative span of a known BRDF dictionary. This assumption enables a per-pixel surface normal and BRDF estimation framework that is computationally tractable and requires no initialization in spite of the underlying problem being non-convex. Our estimation framework first solves for the surface normal at each pixel using a variant of example-based photometric stereo. We design an efficient multi-scale search strategy for estimating the surface normal and subsequently, refine this estimate using a gradient descent procedure. Given the surface normal estimate, we solve for the spatially-varying BRDF by constraining the BRDF at each pixel to be in the span of the BRDF dictionary; here, we use additional priors to further regularize the solution. A hallmark of our approach is that it does not require iterative optimization techniques nor the need for careful initialization, both of which are endemic to most state-of-the-art techniques. We showcase the performance of our technique on a wide range of simulated and real scenes where we outperform competing methods. Zhuo Hui, Aswin C. Sankaranarayanan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Random Features for Sparse Signal ClassificationabstractRandom features is an approach for kernel-based inference on large datasets. In this paper, we derive performance guarantees for random features on signals, like images, that enjoy sparse representations and show that the number of random features required to achieve a desired approximation of the kernel similarity matrix can be significantly smaller for sparse signals. Based on this, we propose a scheme termed compressive random features that first obtains low-dimensional projections of a dataset and, subsequently, derives random features on the low-dimensional projections. This scheme provides significant improvements in signal dimensionality, computational time, and storage costs over traditional random features while enjoying similar theoretical guarantees for achieving inference performance. We support our claims by providing empirical results across many datasets. Jen-Hao Rick Chang, Aswin C. Sankaranarayanan, B. V. K. Vijaya Kumar |
CVPR | 2 |
| 2016 | Dual Structured Light 3D Using a 1D Sensor
Jian Wang 0100, Aswin C. Sankaranarayanan, Mohit Gupta 0001, Srinivasa G. Narasimhan |
ECCV (6) | 2 |
| 2016 | White balance under mixed illumination using flash photographyabstractReal-world illumination is often a complex spatially-varying combination of multiple illuminants. In this work, we present a technique to white-balance images captured in such illumination by leveraging flash photography. Even though this problem is severely ill-posed, we show that using two images — captured with and without flash lighting — leads to a closed form solution for spatially-varying mixed illumination. Our solution is completely automatic and makes no assumptions about the number or nature of the illuminants. We also propose an extension of our scheme to handle practical challenges such as shadows, specularities, as well as the camera and scene motion. We evaluate our technique on datasets captured in both the laboratory and the real-world, and show that it significantly outperforms a number of previous white balance algorithms. Zhuo Hui, Aswin C. Sankaranarayanan, Kalyan Sunkavalli, Sunil Hadap |
ICCP | 2 |
| 2016 | Shape and reflectance from two-bounce light transientsabstractComputer vision and image-based inference have predominantly focused on extracting scene information by assuming that the camera measures direct light transport (i.e., single-bounce light paths). As a consequence, strong multi-bounce effects are treated typically as sources of noise and, in many scenarios, the presence of such effects can result in gross errors in the estimates of shape and reflectance. This paper provides the theoretical and algorithmic foundations for shape and reflectance estimation from two-bounce light transients, i.e., scenarios where photons from a light source interact with the scene exactly twice before reaching the sensor We derive sufficient conditions for exact recovery of shape and reflectance given lengths and intensities associated with two-bounce light paths. We also develop algorithms for recovery of shape and reflectance, and validate these on a range of simulated scenes. Chia-Yin Tsai, Ashok Veeraraghavan, Aswin C. Sankaranarayanan |
ICCP | 3 |
| 2016 | Cross-scale predictive dictionaries for image and video restorationabstractWe propose a novel signal model, based on sparse representations, that captures cross-scale features for visual signals. We show that cross-scale predictive model enables faster solutions to sparse approximation problems. This is achieved by first solving the sparse approximation problem for the downsampled signal and using the support of the solution to constrain the support at the original resolution. The speedups obtained are especially compelling for high-dimensional signals that require large dictionaries to provide precise sparse approximations. We demonstrate speedups in the order of 10-20x for denoising and up to 9x speed-ups for compressive sensing of images and videos. Vishwanath Saragadam, Aswin C. Sankaranarayanan, Xin Li 0001 |
ICIP | 2 |
| 2015 | FPA-CS: Focal plane array-based compressive imaging in short-wave infraredabstractCameras for imaging in short and mid-wave infrared spectra are significantly more expensive than their counterparts in visible imaging. As a result, high-resolution imaging in those spectrum remains beyond the reach of most consumers. Over the last decade, compressive sensing (CS) has emerged as a potential means to realize inexpensive short-wave infrared cameras. One approach for doing this is the single-pixel camera (SPC) where a single detector acquires coded measurements of a high-resolution image. A computational reconstruction algorithm is then used to recover the image from these coded measurements. Unfortunately, the measurement rate of a SPC is insufficient to enable imaging at high spatial and temporal resolutions. We present a focal plane array-based compressive sensing (FPA-CS) architecture that achieves high spatial and temporal resolutions. The idea is to use an array of SPCs that sense in parallel to increase the measurement rate, and consequently, the achievable spatio-temporal resolution of the camera. We develop a proof-of-concept prototype in the short-wave infrared using a sensor with 64× 64 pixels; the prototype provides a 4096× increase in the measurement rate compared to the SPC and achieves a megapixel resolution at video rate using CS techniques. Huaijin G. Chen, Muhammad Salman Asif, Aswin C. Sankaranarayanan, Ashok Veeraraghavan |
CVPR | 3 |
| 2015 | Dynamic sparse state estimation using ℓ1-ℓ1 minimization: Adaptive-rate measurement bounds, algorithms and applicationsabstractWe propose a recursive algorithm for estimating time-varying signals from a few linear measurements. The signals are assumed sparse, with unknown support, and are described by a dynamical model. In each iteration, the algorithm solves an ℓ1-ℓ1minimization problem and estimates the number of measurements that it has to take at the next iteration. These estimates are computed based on recent theoretical results for ℓ1-ℓ1minimization. We also provide sufficient conditions for perfect signal reconstruction at each time instant as a function of an algorithm parameter. The algorithm exhibits high performance in compressive tracking on a real video sequence, as shown in our experimental results. João F. C. Mota, Nikos Deligiannis, Aswin C. Sankaranarayanan, Volkan Cevher, Miguel R. D. Rodrigues |
ICASSP | 3 |
| 2015 | A Dictionary-Based Approach for Estimating Shape and Spatially-Varying ReflectanceabstractWe present a technique for estimating the shape and reflectance of an object in terms of its surface normals and spatially-varying BRDF. We assume that multiple images of the object are obtained under fixed view-point and varying illumination, i.e, the setting of photometric stereo. Assuming that the BRDF at each pixel lies in the non-negative span of a known BRDF dictionary, we derive a per-pixel surface normal and BRDF estimation framework that requires neither iterative optimization techniques nor careful initialization, both of which are endemic to most state-of the-art techniques. We showcase the performance of our technique on a wide range of simulated and real scenes where we outperform competing methods. Zhuo Hui, Aswin C. Sankaranarayanan |
ICCP | 2 |
| 2015 | LiSens- A Scalable Architecture for Video Compressive SensingabstractThe measurement rate of cameras that take spatially multiplexed measurements by using spatial light modulators (SLM) is often limited by the switching speed of the SLMs. This is especially true for single-pixel cameras where the photodetector operates at a rate that is many orders-of-magnitude greater than the SLM. We study the factors that determine the measurement rate for such spatial multiplexing cameras (SMC) and show that increasing the number of pixels in the device improves the measurement rate, but there is an optimum number of pixels (typically, few thousands) beyond which the measurement rate does not increase. This motivates the design of LiSens, a novel imaging architecture, that replaces the photodetector in the single-pixel camera with a 1D linear array or a line-sensor. We illustrate the optical architecture underlying LiSens, build a prototype, and demonstrate results of a range of indoor and outdoor scenes. LiSens delivers on the promise of SMCs: imaging at a megapixel resolution, at video rate, using an inexpensive low-resolution sensor. Jian Wang 0100, Mohit Gupta 0001, Aswin C. Sankaranarayanan |
ICCP | 3 |
| 2015 | Photometric Stereo with Small Angular VariationsabstractMost existing successful photometric stereo setups require large angular variations in illumination directions, which results in acquisition rigs that have large spatial extent. For many applications, especially involving mobile devices, it is important that the device be spatially compact. This naturally implies smaller angular variations in the illumination directions. This paper studies the effect of small angular variations in illumination directions to photometric stereo. We explore both theoretical justification and practical issues in the design of a compact and portable photometric stereo device on which a camera is surrounded by a ring of point light sources. We first derive the relationship between the estimation error of surface normal and the baseline of the point light sources. Armed with this theoretical insight, we develop a small baseline photometric stereo prototype to experimentally examine the theory and its practicality. Jian Wang 0100, Yasuyuki Matsushita, Boxin Shi, Aswin C. Sankaranarayanan |
ICCV | 4 |
| 2015 | What does a single light-ray reveal about a transparent object?abstractWe address the following problem in refractive shape estimation: given a single light-ray correspondence, what shape information of the transparent object is revealed along the path of the light-ray, assuming that the light-ray refracts twice. We answer this question in the form of two depth-normal ambiguities. First, specifying the surface normal at which refraction occurs constrains the depth to a unique value. Second, specifying the depth at which refraction occurs constrains the surface normal to lie on a 1D curve. These two depth-normal ambiguities are fundamental to shape estimation of transparent objects and can be used to derive additional properties. For example, we show that correspondences from three light-rays passing through a point are needed to correctly estimate its surface normal. Another contribution of this work is that we can reduce the number of views required to reconstruct an object by enforcing shape models. We demonstrate this property on real data where we reconstruct shape of an object, with light-rays observed from a single view, by enforcing a locally planar shape model. Chia-Yin Tsai, Ashok Veeraraghavan, Aswin C. Sankaranarayanan |
ICIP | 3 |
| 2015 | Video Compressive Sensing for Spatial Multiplexing Cameras Using Motion-Flow ModelsabstractSpatial multiplexing cameras (SMCs) acquire a (typically static) scene through a series of coded projections using a spatial light modulator (e.g., a digital micromirror device) and a few optical sensors. This approach finds use in imaging applications where full-frame sensors are either too expensive (e.g., for short-wave infrared wavelengths) or unavailable. Existing SMC systems reconstruct static scenes using techniques from compressive sensing (CS). For videos, however, existing acquisition and recovery methods deliver poor quality. In this paper, we propose the CS multiscale video (CS-MUVI) sensing and recovery framework for high-quality video acquisition and recovery using SMCs. Our framework features novel sensing matrices that enable the efficient computation of a low-resolution video preview, while enabling high-resolution video recovery using convex optimization. To further improve the quality of the reconstructed videos, we extract optical-flow estimates from the low-resolution previews and impose them as constraints in the recovery procedure. We demonstrate the efficacy of our CS-MUVI framework for a host of synthetic and real measured SMC video data, and we show that high-quality videos can be recovered at roughly $60\times$ compression. Aswin C. Sankaranarayanan, Christoph Studer, Kevin F. Kelly, Richard G. Baraniuk |
SIAM J. Imaging Sci. | 1 |
| 2014 | LIE operators for compressive sensingabstractWe consider the efficient acquisition, parameter estimation, and recovery of signal ensembles that lie on a low-dimensional manifold in a high-dimensional ambient signal space. Our particular focus is on randomized, compressive acquisition of signals from the manifold generated by the transformation of a base signal by operators from a Lie group. Such manifolds factor prominently in a number of applications, including radar and sonar array processing, camera arrays, and video processing. Leveraging the fact that Lie group manifolds admit a convenient analytical characterization, we develop new theory and algorithms for: (1) estimating the Lie operator parameters from compressive measurements, and (2) recovering the base signal from compressive measurements. We validate our approach with several of numerical simulations, including the reconstruction of an affine-transformed video sequence from compressive measurements. Chinmay Hegde, Aswin C. Sankaranarayanan, Richard G. Baraniuk |
ICASSP | 2 |
| 2014 | DALM-SVD: Accelerated sparse coding through singular value decomposition of the dictionaryabstractSparse coding techniques have seen an increasing range of applications in recent years, especially in the area of image processing. In particular, sparse coding using ℓ1-regularization has been efficiently solved with the Augmented Lagrangian (AL) applied to its dual formulation (DALM). This paper proposes the decomposition of the dictionary matrix in its Singular Value/Vector form in order to simplify and speed-up the implementation of the DALM algorithm. Furthermore, we propose an update rule for the penalty parameter used in AL methods that improves the convergence rate. The SVD of the dictionary matrix is done as a pre-processing step prior to the sparse coding, and thus the method is better suited for applications where the same dictionary is reused for several sparse recovery steps, such as block image processing. Hugo R. Gonçalves, Miguel Correia 0002, Xin Li 0001, Aswin C. Sankaranarayanan, Vítor Grade Tavares |
ICIP | 4 |
| 2014 | Robust Face Recognition From Multi-View VideosabstractMultiview face recognition has become an active research area in the last few years. In this paper, we present an approach for video-based face recognition in camera networks. Our goal is to handle pose variations by exploiting the redundancy in the multiview video data. However, unlike traditional approaches that explicitly estimate the pose of the face, we propose a novel feature for robust face recognition in the presence of diffuse lighting and pose variations. The proposed feature is developed using the spherical harmonic representation of the face texture-mapped onto a sphere; the texture map itself is generated by back-projecting the multiview video data. Video plays an important role in this scenario. First, it provides an automatic and efficient way for feature extraction. Second, the data redundancy renders the recognition algorithm more robust. We measure the similarity between feature sets from different videos using the reproducing kernel Hilbert space. We demonstrate that the proposed approach outperforms traditional algorithms on a multiview video database. Aswin C. Sankaranarayanan, Rama Chellappa |
IEEE Trans. Image Process. | 2 |
| 2014 | Compressive epsilon photography for post-capture control in digital imagingabstractA traditional camera requires the photographer to select the many parameters at capture time. While advances in light field photography have enabled post-capture control of focus and perspective, they suffer from several limitations including lower spatial resolution, need for hardware modifications, and restrictive choice of aperture and focus setting. In this paper, we propose "compressive epsilon photography," a technique for achieving complete post-capture control of focus and aperture in a traditional camera by acquiring a carefully selected set of 8 to 16 images and computationally reconstructing images corresponding to all other focus-aperture settings. We make the following contributions: first, we learn the statistical redundancies in focal-aperture stacks using a Gaussian Mixture Model; second, we derive a greedy sampling strategy for selecting the best focus-aperture settings; and third, we develop an algorithm for reconstructing the entire focal-aperture stack from a few captured images. As a consequence, only a burst of images with carefully selected camera settings are acquired. Post-capture, the user can then select any focal-aperture setting of choice and the corresponding image can be rendered using our algorithm. We show extensive results on several real data sets. Atsushi Ito, Salil Tambe, Kaushik Mitra, Aswin C. Sankaranarayanan, Ashok Veeraraghavan |
ACM Trans. Graph. | 4 |
| 2013 | Greedy feature selection for subspace clustering
Eva L. Dyer, Aswin C. Sankaranarayanan, Richard G. Baraniuk |
J. Mach. Learn. Res. | 2 |
| 2013 | Joint Albedo Estimation and Pose Tracking from VideoabstractThe albedo of a Lambertian object is a surface property that contributes to an object's appearance under changing illumination. As a signature independent of illumination, the albedo is useful for object recognition. Single image-based albedo estimation algorithms suffer due to shadows and non-Lambertian effects of the image. In this paper, we propose a sequential algorithm to estimate the albedo from a sequence of images of a known 3D object in varying poses and illumination conditions. We first show that by knowing/estimating the pose of the object at each frame of a sequence, the object's albedo can be efficiently estimated using a Kalman filter. We then extend this for the case of unknown pose by simultaneously tracking the pose as well as updating the albedo through a Rao-Blackwellized particle filter (RBPF). More specifically, the albedo is marginalized from the posterior distribution and estimated analytically using the Kalman filter, while the pose parameters are estimated using importance sampling and by minimizing the projection error of the face onto its spherical harmonic subspace, which results in an illumination-insensitive pose tracking algorithm. Illustrations and experiments are provided to validate the effectiveness of the approach using various synthetic and real sequences followed by applications to unconstrained, video-based face recognition. Sima Taheri, Aswin C. Sankaranarayanan, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Compressive Acquisition of Linear Dynamical SystemsabstractCompressive sensing (CS) enables the acquisition and recovery of sparse signals and images at sampling rates significantly below the classical Nyquist rate. Despite significant progress in the theory and methods of CS, little headway has been made in compressive video acquisition and recovery. Video CS is complicated by the ephemeral nature of dynamic events, which makes direct extensions of standard CS imaging architectures and signal models difficult. In this paper, we develop a new framework for video CS for dynamic textured scenes that models the evolution of the scene as a linear dynamical system (LDS). This reduces the video recovery problem to first estimating the model parameters of the LDS from compressive measurements and then reconstructing the image frames. We exploit the low-dimensional dynamic parameters (the state sequence) and high-dimensional static parameters (the observation matrix) of the LDS to devise a novel compressive measurement strategy that measures only the time-varying parameters at each instant and accumulates measurements over time to estimate the time-invariant parameters. This enables us to lower the compressive measurement rate considerably. We validate our approach and demonstrate its effectiveness with a range of experiments involving video recovery and scene classification. Aswin C. Sankaranarayanan, Pavan Turaga, Rama Chellappa, Richard G. Baraniuk |
SIAM J. Imaging Sci. | 1 |
| 2012 | Flutter Shutter Video Camera for compressive sensing of videosabstractVideo cameras are invariably bandwidth limited and this results in a trade-off between spatial and temporal resolution. Advances in sensor manufacturing technology have tremendously increased the available spatial resolution of modern cameras while simultaneously lowering the costs of these sensors. In stark contrast, hardware improvements in temporal resolution have been modest. One solution to enhance temporal resolution is to use high bandwidth imaging devices such as high speed sensors and camera arrays. Unfortunately, these solutions are expensive. An alternate solution is motivated by recent advances in computational imaging and compressive sensing. Camera designs based on these principles, typically, modulate the incoming video using spatio-temporal light modulators and capture the modulated video at a lower bandwidth. Reconstruction algorithms, motivated by compressive sensing, are subsequently used to recover the high bandwidth video at high fidelity. Though promising, these methods have been limited since they require complex and expensive light modulators that make the techniques difficult to realize in practice. In this paper, we show that a simple coded exposure modulation is sufficient to reconstruct high speed videos. We propose the Flutter Shutter Video Camera (FSVC) in which each exposure of the sensor is temporally coded using an independent pseudo-random sequence. Such exposure coding is easily achieved in modern sensors and is already a feature of several machine vision cameras. We also develop two algorithms for reconstructing the high speed video; the first based on minimizing the total variation of the spatio-temporal slices of the video and the second based on a data driven dictionary based approximation. We perform evaluation on simulated videos and real data to illustrate the robustness of our system. Jason Holloway, Aswin C. Sankaranarayanan, Ashok Veeraraghavan, Salil Tambe |
ICCP | 2 |
| 2012 | CS-MUVI: Video compressive sensing for spatial-multiplexing camerasabstractCompressive sensing (CS)-based spatial-multiplexing cameras (SMCs) sample a scene through a series of coded projections using a spatial light modulator and a few optical sensor elements. SMC architectures are particularly useful when imaging at wavelengths for which full-frame sensors are too cumbersome or expensive. While existing recovery algorithms for SMCs perform well for static images, they typically fail for time-varying scenes (videos). In this paper, we propose a novel CS multi-scale video (CS-MUVI) sensing and recovery framework for SMCs. Our framework features a co-designed video CS sensing matrix and recovery algorithm that provide an efficiently computable low-resolution video preview. We estimate the scene's optical flow from the video preview and feed it into a convex-optimization algorithm to recover the high-resolution video. We demonstrate the performance and capabilities of the CS-MUVI framework for different scenes. Aswin C. Sankaranarayanan, Christoph Studer, Richard G. Baraniuk |
ICCP | 1 |
| 2011 | SpaRCS: Recovering low-rank and sparse matrices from compressive measurementsabstractWe consider the problem of recovering a matrix $\mathbf{M}$ that is the sum of a low-rank matrix $\mathbf{L}$ and a sparse matrix $\mathbf{S}$ from a small set of linear measurements of the form $\mathbf{y} = \mathcal{A}(\mathbf{M}) = \mathcal{A}({\bf L}+{\bf S})$. This model subsumes three important classes of signal recovery problems: compressive sensing, affine rank minimization, and robust principal component analysis. We propose a natural optimization problem for signal recovery under this model and develop a new greedy algorithm called SpaRCS to solve it. SpaRCS inherits a number of desirable properties from the state-of-the-art CoSaMP and ADMiRA algorithms, including exponential convergence and efficient implementation. Simulation results with video compressive sensing, hyperspectral imaging, and robust matrix completion data sets demonstrate both the accuracy and efficacy of the algorithm. Andrew E. Waters, Aswin C. Sankaranarayanan, Richard G. Baraniuk |
NIPS | 2 |
| 2010 | Specular surface reconstruction from sparse reflection correspondencesabstractWe present a practical approach for surface reconstruction of smooth mirror-like objects using sparse reflection correspondences (RCs). Assuming finite object motion with a fixed camera and un-calibrated environment, we derive the relationship between RC and the surface shape. We show that by locally modeling the surface as a quadric, the relationship between the RCs and unknown surface parameters becomes linear. We develop a simple surface reconstruction algorithm that amounts to solving either an eigenvalue problem or a second order cone program (SOCP). Ours is the first method that allows for reconstruction of mirror surfaces from sparse RCs, obtained from standard algorithms such as SIFT. Our approach overcomes the practical issues in shape from specular flow (SFSF) such as the requirement of dense optical flow and undefined/infinite flow at parabolic points. We also show how to incorporate auxiliary information such as sparse surface normals into our framework. Experiments, both real and synthetic are shown that validate the theory presented. Aswin C. Sankaranarayanan, Ashok Veeraraghavan, Oncel Tuzel, Amit K. Agrawal |
CVPR | 1 |
| 2010 | Compressive Acquisition of Dynamic Scenes
Aswin C. Sankaranarayanan, Pavan Turaga, Richard G. Baraniuk, Rama Chellappa |
ECCV (1) | 1 |
| 2010 | Image Invariants for Smooth Reflective Surfaces
Aswin C. Sankaranarayanan, Ashok Veeraraghavan, Oncel Tuzel, Amit K. Agrawal |
ECCV (2) | 1 |
| 2010 | Online Empirical Evaluation of Tracking AlgorithmsabstractEvaluation of tracking algorithms in the absence of ground truth is a challenging problem. There exist a variety of approaches for this problem, ranging from formal model validation techniques to heuristics that look for mismatches between track properties and the observed data. However, few of these methods scale up to the task of visual tracking, where the models are usually nonlinear and complex and typically lie in a high-dimensional space. Further, scenarios that cause track failures and/or poor tracking performance are also quite diverse for the visual tracking problem. In this paper, we propose an online performance evaluation strategy for tracking systems based on particle filters using a time-reversed Markov chain. The key intuition of our proposed methodology relies on the time-reversible nature of physical motion exhibited by most objects, which in turn should be possessed by a good tracker. In the presence of tracking failures due to occlusion, low SNR, or modeling errors, this reversible nature of the tracker is violated. We use this property for detection of track failures. To evaluate the performance of the tracker at time instant t, we use the posterior of the tracking algorithm to initialize a time-reversed Markov chain. We compute the posterior density of track parameters at the starting time t=0 by filtering back in time to the initial time instant. The distance between the posterior density of the time-reversed chain (at t=0) and the prior density used to initialize the tracking algorithm forms the decision statistic for evaluation. It is observed that when the data are generated by the underlying models, the decision statistic takes a low value. We provide a thorough experimental analysis of the evaluation methodology. Specifically, we demonstrate the effectiveness of our approach for tackling common challenges such as occlusion, pose, and illumination changes and provide the Receiver Operating Characteristic (ROC) curves. Finally, we also show the applicability of the core ideas of the paper to other tracking algorithms such as the Kanade-Lucas-Tomasi (KLT) feature tracker and the mean-shift tracker. Hao Wu 0014, Aswin C. Sankaranarayanan, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Compressive Sensing for Background Subtraction
Volkan Cevher, Aswin C. Sankaranarayanan, Marco F. Duarte, Dikpal Reddy, Richard G. Baraniuk, Rama Chellappa |
ECCV (2) | 2 |
| 2008 | Factorized variational approximations for acoustic multi source localizationabstractEstimation based on received signal strength (RSS) is crucial in sensor networks for sensor localization, target tracking, etc. In this paper, we present a Gaussian approximation of the Chi distribution that is applicable to general RSS source localization problems in sensor networks. Using our Gaussian approximation, we provide a factorized variational Bayes (VB) approximation to the location and power posterior of multiple sources using a sensor network. When the source signal and the sensor noise have uncorrelated Gaussian distributions, we demonstrate that the envelope of the sensor output can be accurately modeled with a multiplicative Gaussian noise model. In turn, our factorized VB approximations decrease the computational complexity and provide computational robustness as the number of targets increases. Simulations are provided to demonstrate the effectiveness of the proposed approximations. Volkan Cevher, Aswin C. Sankaranarayanan, Rama Chellappa |
ICASSP | 2 |
| 2008 | Compressed sensing for multi-view tracking and 3-D voxel reconstructionabstractCompressed sensing (CS) suggests that a signal, sparse in some basis, can be recovered from a small number of random projections. In this paper, we apply the CS theory on sparse background-subtracted silhouettes and show the usefulness of such an approach in various multi-view estimation problems. The sparsity of the silhouette images corresponds to sparsity of object parameters (location, volume etc.) in the scene. We use random projections (compressed measurements) of the silhouette images for directly recovering object parameters in the scene coordinates. To keep the computational requirements of this recovery procedure reasonable, we tessellate the scene into a bunch of non-overlapping lines and perform estimation on each of these lines. Our method is scalable in the number of cameras and utilizes very few measurements for transmission among cameras. We illustrate the usefulness of our approach for multi-view tracking and 3-D voxel reconstruction problems. Dikpal Reddy, Aswin C. Sankaranarayanan, Volkan Cevher, Rama Chellappa |
ICIP | 2 |
| 2008 | Stochastic fusion of multi-view gradient fieldsabstractImage gradients form powerful cues in a host of vision and graphics applications. In this paper, we consider multiple views of a textured planar scene and consider the problem of estimating the scene texture map using these multi-view inputs. Modeling each camera view as a projective transformation of the scene, we show that the problem is equivalent to that of studying the effect of noise (and the projective imaging) on the gradient fields induced by this texture map. We show that these noisy gradient fields can be modeled as complete observers of the scene radiance. Further, the corrupting noise can be shown to be additive and linear, although spatially varying. However, the specific form of the noise term can be exploited to design linear estimators that fuse the gradient fields obtained from each of the individual views. The fused gradient field forms a robust estimate of the scene gradients and can be used for scene reconstruction. Aswin C. Sankaranarayanan, Rama Chellappa |
ICIP | 1 |
| 2008 | Object Detection, Tracking and Recognition for Multiple Smart CamerasabstractVideo cameras are among the most commonly used sensors in a large number of applications, ranging from surveillance to smart rooms for videoconferencing. There is a need to develop algorithms for tasks such as detection, tracking, and recognition of objects, specifically using distributed networks of cameras. The projective nature of imaging sensors provides ample challenges for data association across cameras. We first discuss the nature of these challenges in the context of visual sensor networks. Then, we show how real-world constraints can be favorably exploited in order to tackle these challenges. Examples of real-world constraints are (a) the presence of a world plane, (b) the presence of a three-dimiensional scene model, (c) consistency of motion across cameras, and (d) color and texture properties. In this regard, the main focus of this paper is towards highlighting the efficient use of the geometric constraints induced by the imaging devices to derive distributed algorithms for target detection, tracking, and recognition. Our discussions are supported by several examples drawn from real applications. Lastly, we also describe several potential research problems that remain to be addressed. Aswin C. Sankaranarayanan, Ashok Veeraraghavan, Rama Chellappa |
Proc. IEEE | 1 |
| 2008 | Algorithmic and Architectural Optimizations for Computationally Efficient Particle FilteringabstractIn this paper, we analyze the computational challenges in implementing particle filtering, especially to video sequences. Particle filtering is a technique used for filtering nonlinear dynamical systems driven by non-Gaussian noise processes. It has found widespread applications in detection, navigation, and tracking problems. Although, in general, particle filtering methods yield improved results, it is difficult to achieve real time performance. In this paper, we analyze the computational drawbacks of traditional particle filtering algorithms, and present a method for implementing the particle filter using the Independent Metropolis Hastings sampler, that is highly amenable to pipelined implementations and parallelization. We analyze the implementations of the proposed algorithm, and, in particular, concentrate on implementations that have minimum processing times. It is shown that the design parameters for the fastest implementation can be chosen by solving a set of convex programs. The proposed computational methodology was verified using a cluster of PCs for the application of visual tracking. We demonstrate a linear speed-up of the algorithm using the methodology proposed in the paper. Aswin C. Sankaranarayanan, Ankur Srivastava 0001, Rama Chellappa |
IEEE Trans. Image Process. | 1 |
| 2007 | In Situ Evaluation of Tracking Algorithms Using Time Reversed ChainsabstractAutomatic evaluation of visual tracking algorithms in the absence of ground truth is a very challenging and important problem. In the context of online appearance modeling, there is an additional ambiguity involving the correctness of the appearance model. In this paper, we propose a novel performance evaluation strategy for tracking systems based on particle filter using a time reversed Markov chain. Starting from the latest observation, the time reversed chain is propagated back till the starting time t = 0 of the tracking algorithm. The posterior density of the time reversed chain is also computed. The distance between the posterior density of the time reversed chain (at t = 0) and the prior density used to initialize the tracking algorithm forms the decision statistic for evaluation. It is postulated that when the data is generated true to the underlying models, the decision statistic takes a low value. We empirically demonstrate the performance of the algorithm against various common failure modes in the generic visual tracking problem. Finally, we derive a small frame approximation that allows for very efficient computation of the decision statistic. Hao Wu 0014, Aswin C. Sankaranarayanan, Rama Chellappa |
CVPR | 2 |
| 2007 | Joint Acoustic-Video Fingerprinting of Vehicles, Part IIabstractIn this second paper, we first show how to estimate the wheelbase length of a vehicle using line metrology in video. We then address the vehicle fingerprinting problem using vehicle silhouettes and color invariants. We combine the acoustic metrology and classification results discussed in Part I with the video results to improve estimation performance and robustness. The acoustic video fusion is achieved in a Bayesian framework by assuming conditional independence of the observations of each modality. For the metrology density functions, Laplacian approximations are used for computational efficiency. Experimental results are given using field data. Volkan Cevher, Feng Guo 0006, Aswin C. Sankaranarayanan, Rama Chellappa |
ICASSP (2) | 3 |
| 2007 | Robust Visual Tracking Using the Time-Reversibility ConstraintabstractVisual tracking is a very important front-end to many vision applications. We present a new framework for robust visual tracking in this paper. Instead of just looking forward in the time domain, we incorporate both forward and backward processing of video frames using a novel time-reversibility constraint. This leads to a new minimization criterion that combines the forward and backward similarity functions and the distances of the state vectors between the forward and backward states of the tracker. The new framework reduces the possibility of the tracker getting stuck in local minima and significantly improves the tracking robustness and accuracy. Our approach is general enough to be incorporated into most of the current tracking algorithms. We illustrate the improvements due to the proposed approach for the popular KLT tracker and a search based tracker. The experimental results show that the improved KLT tracker significantly outperforms the original KLT tracker. The time-reversibility constraint used for tracking can be incorporated to improve the performance of optical flow, mean shift tracking and other algorithms. Hao Wu 0014, Rama Chellappa, Aswin C. Sankaranarayanan, Shaohua Kevin Zhou |
ICCV | 3 |
| 2007 | Target Tracking Using a Joint Acoustic Video SystemabstractIn this paper, a multitarget tracking system for collocated video and acoustic sensors is presented. We formulate the tracking problem using a particle filter based on a state-space approach. We first discuss the acoustic state-space formulation whose observations use a sliding window of direction-of-arrival estimates. We then present the video state space that tracks a target's position on the image plane based on online adaptive appearance models. For the joint operation of the filter, we combine the state vectors of the individual modalities and also introduce a time-delay variable to handle the acoustic-video data synchronization issue, caused by acoustic propagation delays. A novel particle filter proposal strategy for joint state-space tracking is introduced, which places the random support of the joint filter where the final posterior is likely to lie. By using the Kullback-Leibler divergence measure, it is shown that the joint operation of the filter decreases the worst case divergence of the individual modalities. The resulting joint tracking filter is quite robust against video and acoustic occlusions due to our proposal strategy. Computer simulations are presented with synthetic and field data to demonstrate the filter's performance Volkan Cevher, Aswin C. Sankaranarayanan, James H. McClellan, Rama Chellappa |
IEEE Trans. Multim. | 2 |
| 2005 | Algorithmic and Architectural Design Methodology for Particle Filters in HardwareabstractIn this paper, we present algorithmic and architectural methodology for building particle filters in hardware. Particle filtering is a new paradigm for filtering in presence of nonGaussian nonlinear state evolution and observation models. This technique has found wide-spread application in tracking, navigation, detection problems especially in a sensing environment. So far most particle filtering implementations are not lucrative for real time problems due to excessive computational complexity involved. In this paper, we re-derive the particle filtering theory to make it more amenable to simplified VLSI implementations. Furthermore, we present and analyze pipelined architectural methodology for designing these computational blocks. Finally, we present an application using the bearing only tracking problem and evaluate the proposed architecture and algorithmic methodology. Aswin C. Sankaranarayanan, Rama Chellappa, Ankur Srivastava 0001 |
ICCD | 1 |
| 2005 | Tracking objects in video using motion and appearance modelsabstractThis paper proposes a visual tracking algorithm that combines motion and appearance in a statistical framework. It is assumed that image observations are generated simultaneously from a background model and a target appearance model. This is different from conventional appearance-based tracking, that does not use motion information. The proposed algorithm attempts to maximize the likelihood ratio of the tracked region, derived from appearance and background models. Incorporation of motion in appearance based tracking provides robust tracking, even when the target violates the appearance model. We show that the proposed algorithm performs well in tracking targets efficiently over long time intervals. Aswin C. Sankaranarayanan, Rama Chellappa, Qinfen Zheng |
ICIP (2) | 1 |