EDBT 2026 Demo / reviewers in the wild / expert
Ashok Veeraraghavan
dblp:84/858 · also Ashok Narayanan Veeraraghavan
· DBLP profile ↗
116ranked-venue papers
11as first author
38since 2021 · last 2026
0000-0001-5043-7460ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 84 · 8 first-author · 21 since 2021Artificial intelligence and machine learning · 67 · 7 first-author · 26 since 2021Systems, architecture and hardware · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The feasibility of passively tracking children's TV viewing and mobile device use in naturalistic settingsabstractResearch on children's technology and digital media (TDM) is hampered by a lack of robust approaches for assessing TDM use. This study assessed the feasibility of passively measuring children's TV screens and mobile devices (TDM) in a naturalistic setting. In the three-day feasibility study, FLASH-TV was set up on one to two TVs the child (5-12 year olds) typically used in the home (n=20). Children's mobile device use was assessed with either the Chronicle App or ScreenTime screenshots. Parents completed three TDM diaries. An exit interview with the parent explored their perceptions of the assessments and the child's TDM use report. Complete data were obtained on 86.7% of days for passive assessment of TV viewing and 84.3% of days for mobile device use. Fifteen parents reviewed complete TDM use reports for their child, with most stating the reports appeared correct for TV (80%) and mobile device (80%). Almost two-thirds had no concerns about having the FLASH-TV installed in their home, while some reported issues about feeling observed. Parents described high burden and frustration with the TDM diaries. Data provided preliminary evidence that passive measurement is feasible for assessing children's TV and mobile device use, with reduced burden for parents. Teresia M. O'Connor, Tatyana Garza, Uzair Alam, Anil Kumar Vadathya, Jennette P. Moreno, Alicia Beltran, Samah Haidar, Nimah Haidar, Sheryl O. Hughes, Debbe Thompson, Salma M. Musaad, Thomas Baranowski, Jason A. Mendoza, Joseph Young, Akane Sano, Ashok Veeraraghavan |
Behav. Inf. Technol. | 16 |
| 2025 | Diffusion Model Based Image Reconstruction in Lensless ImagingabstractLensless imaging systems eliminate the need for lenses by employing an encoding element to multiplex incident light signals, which are then captured directly onto a bare camera sensor. They present a promising alternative to traditional lens-based imaging systems by offering significant advantages in terms of compactness, versatility, and cost. Due to the multiplexed nature of measurements, image reconstruction takes place computationally. However, existing techniques for image reconstruction in lensless imaging fall short of the image quality offered by traditional lens-based imaging. In this work, we consider the application of diffusion models, a class of deep generative models, for image reconstruction in a lensless imaging modality. These models currently achieve state-of-the-art performance in image generation. Specifically, we focus on the PhlatCam lensless system, which consists of a coded phase mask as the encoding element placed close to the camera sensor. We use a ControlNet based diffusion model to improve the perceptual quality of image reconstruction. The performance is measured in terms of peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), and learned perceptual image patch similarity (LPIPS). The proposed method improves the performance in these metrics for synthetic measurements. For real measurements, the improvement in image quality comes at the expense of a small bias in color, which is attributed to the generative nature of the diffusion prior itself. Vivek Boominathan, Ashok Veeraraghavan, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2025 | Fit Pixels, Get Labels: Meta-learned Implicit Networks for Image Segmentation
Kushal Vyas, Ashok Veeraraghavan, Guha Balakrishnan |
MICCAI (3) | 2 |
| 2025 | CogPhys: Assessing Cognitive Load via Multimodal Remote and Contact-based Physiological SensingabstractRemote physiological sensing is an evolving area of research. As systems approach clinical precision, there is increasing focus on complex applications such as cognitive state estimation. Hence, there is a need for large datasets that facilitate research into complex downstream tasks such as remote cognitive load estimation. A first-of-its-kind, our paper introduces an open-source multimodal multi-vital sign dataset consisting of concurrent recordings from RGB, NIR (near-infrared), thermal, and RF (radio-frequency) sensors alongside contact-based physiological signals, such as pulse oximeter and chest bands, providing a benchmark for cognitive state assessment. By adopting a multimodal approach to remote health sensing, our dataset and its associated hardware system excel at modeling the complexities of cognitive load. Here, cognitive load is defined as the mental effort exerted during tasks such as reading, memorizing, and solving math problems. By using the NASA-TLX survey, we set personalized thresholds for defining high/low cognitive levels, enabling a more reliable benchmark. Our benchmarking scheme bridges the gap between existing remote sensing strategies and cognitive load estimation techniques by using vital signs (such as photoplethysmography (PPG) and respiratory waveforms) and physiological signals (blink waveforms) as an intermediary. Through this paper, we focus on replacing the need for intrusive contact-based physiological measurements with more user-friendly remote sensors. Our benchmarking demonstrates that multimodal fusion significantly improves remote vital sign estimation, with our fusion model achieving $<3~BPM$ (beats per minute) error for vital sign estimation. For cognitive load classification, the combination of remote PPG, remote respiratory signals, and blink markers achieves $86.49$% accuracy, approaching the performance of contact-based sensing ($87.5$%) and validating the feasibility of non-intrusive cognitive monitoring. Anirudh Bindiganavale Harish, Peikun Guo, Bhargav Ghanekar, Diya Gupta, Akilesh Rajavenkatanarayanan, Maureen August, Akane Sano, Ashok Veeraraghavan |
NeurIPS | 9 |
| 2025 | CoIR: Compressive Implicit RadarabstractUsing millimeter wave (mmWave) signals for imaging has an important advantage in that they can penetrate through poor environmental conditions such as fog, dust, and smoke that severely degrade optical-based imaging systems. However, mmWave radars, contrary to cameras and LiDARs, suffer from low angular resolution because of small physical apertures and conventional signal processing techniques. Sparse radar imaging, on the other hand, can increase the aperture size while minimizing power consumption and read-out bandwidth. This article presents CoIR, an analysis by synthesis method that leverages the implicit neural network bias in convolutional decoders and compressed sensing to perform high-accuracy sparse radar imaging. The proposed system is data set-agnostic and does not require any auxiliary sensors for training or testing. We introduce a sparse array design that allows for a $5.5\times$5.5× reduction in the number of antenna elements needed compared to conventional MIMO array designs. We demonstrate our system's improved imaging performance over standard mmWave radars and other competitive untrained methods on both simulated and experimental mmWave radar data. Sean M. Farrell, Vivek Boominathan, Nate Raymondi, Ashutosh Sabharwal, Ashok Veeraraghavan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | PS$^{2}$2 F: Polarized Spiral Point Spread Function for Single-Shot 3D SensingabstractWe propose a compact snapshot monocular depth estimation technique that relies on an engineered point spread function (PSF). Traditional approaches used in microscopic super-resolution imaging such as the Double-Helix PSF (DHPSF) are ill-suited for scenes that are more complex than a sparse set of point light sources. We show, using the Cramér-Rao lower bound, that separating the two lobes of the DHPSF and thereby capturing two separate images leads to a dramatic increase in depth accuracy. A special property of the phase mask used for generating the DHPSF is that a separation of the phase mask into two halves leads to a spatial separation of the two lobes. We leverage this property to build a compact polarization-based optical setup, where we place two orthogonal linear polarizers on each half of the DHPSF phase mask and then capture the resulting image with a polarization-sensitive camera. Results from simulations and a lab prototype demonstrate that our technique achieves up to $50\%$50% lower depth error compared to state-of-the-art designs including the DHPSF and the Tetrapod PSF, with little to no loss in spatial resolution. Bhargav Ghanekar, Vishwanath Saragadam, Dushyant Mehra, Anna-Karin Gustavsson, Aswin C. Sankaranarayanan, Ashok Veeraraghavan |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Guest Editorial: Introduction to the Special Section on Computational Photography
Boxin Shi, Ashok Veeraraghavan, Roarke Horstmeyer, Wolfgang Heidrich |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Streaming Quanta Sensors for Online, High-Performance Imaging and VisionabstractRecently quanta image sensors (QIS) - ultra-fast, zero-read-noise binary image sensors- have demonstrated remarkable imaging capabilities in many challenging scenarios. Despite their potential, the adoption of these sensors is severely hampered by (a) high data rates and (b) the need for new computational pipelines to handle the unconventional raw data. We introduce a simple, low-bandwidth computational pipeline to address these challenges. Our approach is based on a novel streaming representation with a small memory footprint, efficiently capturing intensity information at multiple temporal scales. Updating the representation requires only 16 floating-point operations/pixel, which can be efficiently computed online at the native frame rate of the binary frames. We use a neural network operating on this representation to reconstruct videos in real-time (10-30 fps). We illustrate why such representation is well-suited for these emerging sensors, and how it offers low latency and high frame rate while retaining flexibility for downstream computer vision. Our approach results in significant data bandwidth reductions () and real-time image reconstruction and computer vision -)reduction in computation than existing state-of-the-art approach[1], while maintaining comparable quality. To the best of our knowledge, our approach is the first to achieve online, real-time image reconstruction on QIS. Matthew Dutson, Vivek Boominathan, Mohit Gupta 0001, Ashok Veeraraghavan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Downscaling Extreme Precipitation With Wasserstein Regularized DiffusionabstractUnderstanding the risks posed by extreme rainfall events requires analysis of precipitation fields with high resolution (to assess localized hazards) and extensive historical coverage (to capture sufficient examples of rare occurrences). Radar and mesonet networks provide precipitation fields at 1 km resolution but with limited historical and geographical coverage, while gauge-based records and reanalysis products cover decades of time on a global scale, but only at 30–50 km resolution. To help provide high-resolution precipitation estimates over long time scales, this study presents Wasserstein Regularized Diffusion (WassDiff), a diffusion framework to downscale (super-resolve) precipitation fields from low-resolution gauge and reanalysis products. Crucially, unlike related deep generative models, WassDiff integrates a Wasserstein distribution-matching regularizer to the denoising process to reduce empirical biases at extreme intensities. Comprehensive evaluations demonstrate that WassDiff quantitatively outperforms existing state-of-the-art generative downscaling methods at recovering extreme weather phenomena such as tropical storms and cold fronts. Case studies further qualitatively demonstrate WassDiff’s ability to reproduce realistic fine-scale structures and accurate peak intensities of these phenomena. By unlocking decades of high-resolution rainfall information from globally available coarse records, WassDiff offers a practical pathway toward more accurate flood-risk assessments and climate-adaptation planning. Yuhao Liu 0012, James Doss-Gollin, Qiushi Dai, Ashok Veeraraghavan, Guha Balakrishnan |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | NeST: Neural Stress Tensor Tomography by leveraging 3D PhotoelasticityabstractPhotoelasticity enables full-field stress analysis in transparent objects through stress-induced birefringence. Existing techniques are limited to two-dimensional (2D) slices and require destructively slicing the object. Recovering the internal three-dimensional (3D) stress distribution of the entire object is challenging, as it involves solving a tensor tomography problem and handling phase wrapping ambiguities. We introduce NeST, an analysis-by-synthesis approach for reconstructing 3D stress tensor fields as neural implicit representations from polarization measurements. Our key insight is to jointly handle phase unwrapping and tensor tomography using a differentiable forward model based on Jones calculus. Our non-linear model faithfully matches real captures, unlike prior linear approximations. We develop an experimental multi-axis polariscope setup to capture 3D photoelasticity and experimentally demonstrate that NeST reconstructs the internal stress distribution for objects with varying shape and force conditions. Additionally, we showcase novel applications in stress analysis, such as visualizing photoelastic fringes by virtually slicing the object and viewing photoelastic fringes from unseen viewpoints. NeST paves the way for scalable non-destructive 3D photoelastic analysis. Akshat Dave, Aaron Young, Ramesh Raskar, Wolfgang Heidrich, Ashok Veeraraghavan |
ACM Trans. Graph. | 6 |
| 2024 | Passive Snapshot Coded Aperture Dual-Pixel RGB-D ImagingabstractPassive, compact, single-shot 3D sensing is useful in many application areas such as microscopy, medical imaging, surgical navigation, and autonomous driving where form factor, time, and power constraints can exist. Ob-taining RGB-D scene information over a short imaging distance, in an ultra-compact form factor, and in a passive, snapshot manner is challenging. Dual-pixel (DP) sensors are a potential solution to achieve the same. DP sensors collect light rays from two different halves of the lens in two interleaved pixel arrays, thus capturing two slightly different views of the scene, like a stereo camera system. However, imaging with a DP sensor implies that the defocus blur size is directly proportional to the disparity seen between the views. This creates a tradeoff between disparity estimation vs. deblurring accuracy. To improve this tradeoff effect, we propose CADS (Coded Aperture Dual-Pixel Sensing), in which we use a coded aperture in the imaging lens along with a DP sensor. In our approach, we jointly learn an optimal coded pattern and the reconstruction algorithm in an end-to-end optimization setting. Our resulting CADS imaging system demonstrates improvement of> 1.5 dB PSNR in all-in-focus (AIF) estimates and 5-6% in depth estimation quality over naive DP sensing for a wide range of aperture settings. Furthermore, we build the proposed CADS prototypes for DSLR photography settings and in an endoscope and a dermoscope form factor. Our novel coded dual-pixel sensing approach demonstrates accurate RGB-D reconstruction results in simulations and real-world experiments in a passive, snapshot, and compact manner. Bhargav Ghanekar, Salman Siddique Khan, Pranav Sharma, Shreyas Singh, Vivek Boominathan, Kaushik Mitra, Ashok Veeraraghavan |
CVPR | 7 |
| 2024 | WaveMo: Learning Wavefront Modulations to See Through ScatteringabstractImaging through scattering media is a fundamental and pervasive challenge infields ranging from medical diagnos-tics to astronomy. A promising strategy to overcome this challenge is wavefront modulation, which induces measure-ment diversity during image acquisition. Despite its importance, designing optimal wavefront modulations to image through scattering remains under-explored. This paper in-troduces a novel learning-based framework to address the gap. Our approach jointly optimizes wavefront modulations and a computationally lightweight feedforward “proxy” re-construction network. This network is trained to recover scenes obscured by scattering, using measurements that are modified by these modulations. The learned modulations produced by our framework generalize effectively to un-seen scattering scenarios and exhibit remarkable versatility. During deployment, the learned modulations can be decou-pled from the proxy network to augment other more computationally expensive restoration algorithms. Through ex-tensive experiments, we demonstrate our approach signifi-cantly advances the state of the art in imaging through scat-tering media. Our project webpage is at https://wavemo-2024.github.io/. Mingyang Xie, Haiyun Guo, Brandon Yushan Feng, Lingbo Jin, Ashok Veeraraghavan, Christopher A. Metzler |
CVPR | 5 |
| 2024 | DecentNeRFs: Decentralized Neural Radiance Fields from Crowdsourced Images
Zaid Tasneem, Akshat Dave, Abhishek Singh 0005, Kushagra Tiwary, Praneeth Vepakomma, Ashok Veeraraghavan, Ramesh Raskar |
ECCV (59) | 6 |
| 2024 | Message from the ChairsabstractWelcome to the 16th IEEE International Conference on Computational Photography (ICCP 2024), taking place at the Ecole Polytéchnique Fédérale (EPFL) in Lausanne, Switzerland! This is the first time that ICCP takes place in Continental Europe, specifically on the shore of beautiful Lake Geneva. This year's two-and-a-half day conference features 18 accepted papers, 3 keynote talks, 8 invited talks, and 50 posters and/or demos. Sabine Süsstrunk, Ashok Veeraraghavan, Roarke Horstmeyer, Wolfgang Heidrich |
ICCP | 2 |
| 2024 | Temporally Consistent Atmospheric Turbulence Mitigation with Neural RepresentationsabstractAtmospheric turbulence, caused by random fluctuations in the atmosphere's refractive index, introduces complex spatio-temporal distortions in imagery captured at long range. Video Atmospheric Turbulence Mitigation (ATM) aims to restore videos affected by these distortions. However, existing video ATM methods, both supervised and self-supervised, struggle to maintain temporally consistent mitigation across frames, leading to visually incoherent results. This limitation arises from the stochastic nature of atmospheric turbulence, which varies across space and time. Inspired by the observation that atmospheric turbulence induces high-frequency temporal variations, we propose ConVRT, a novel framework for consistent video restoration through turbulence. ConVRT introduces a neural video representation that explicitly decouples spatial and temporal information into a spatial content field and a temporal deformation field, enabling targeted regularization of the network's temporal representation capability. By leveraging the low-pass filtering properties of the regularized temporal representations, ConVRT effectively mitigates turbulence-induced temporal frequency variations and promotes temporal consistency. Furthermore, our training framework seamlessly integrates supervised pre-training on synthetic turbulence data with self-supervised learning on real-world videos, significantly improving the temporally consistent mitigation of ATM methods on diverse real-world data. More information can be found on our project page: https://convrt-2024.github.io/ Haoming Cai, Jingxi Chen, Brandon Yushan Feng, Weiyun Jiang, Mingyang Xie, Kevin Zhang 0003, Cornelia Fermüller, Yiannis Aloimonos, Ashok Veeraraghavan, Christopher A. Metzler |
NeurIPS | 9 |
| 2024 | Learning Transferable Features for Implicit Neural RepresentationsabstractImplicit neural representations (INRs) have demonstrated success in a variety of applications, including inverse problems and neural rendering. An INR is typically trained to capture one signal of interest, resulting in learned neural features that are highly attuned to that signal. Assumed to be less generalizable, we explore the aspect of transferability of such learned neural features for fitting similar signals. We introduce a new INR training framework, STRAINER that learns transferable features for fitting INRs to new signals from a given distribution, faster and with better reconstruction quality. Owing to the sequential layer-wise affine operations in an INR, we propose to learn transferable representations by sharing initial encoder layers across multiple INRs with independent decoder layers. At test time, the learned encoder representations are transferred as initialization for an otherwise randomly initialized INR. We find STRAINER to yield extremely powerful initialization for fitting images from the same domain and allow for a ≈ +10dB gain in signal quality early on compared to an untrained INR itself. STRAINER also provides a simple way to encode data-driven priors in INRs. We evaluate STRAINER on multiple in-domain and out-of-domain signal fitting tasks and inverse problems and further provide detailed analysis and discussion on the transferability of STRAINER’s features. Kushal Vyas, Ahmed Imtiaz Humayun, Aniket Dashpute, Richard G. Baraniuk, Ashok Veeraraghavan, Guha Balakrishnan |
NeurIPS | 5 |
| 2024 | Development of family level assessment of screen use in the home for television (FLASH-TV)
Anil Kumar Vadathya, Thomas Baranowski, Teresia M. O'Connor, Alicia Beltran, Salma M. Musaad, Oriana Perez, Jason A. Mendoza, Sheryl O. Hughes, Ashok Veeraraghavan |
Multim. Tools Appl. | 9 |
| 2024 | DeepTensor: Low-Rank Tensor Decomposition With Deep Network PriorsabstractDeepTensor is a computationally efficient framework for low-rank decomposition of matrices and tensors using deep generative networks. We decompose a tensor as the product of low-rank tensor factors (e.g., a matrix as the outer product of two vectors), where each low-rank tensor is generated by a deep network (DN) that is trained in a self-supervised manner to minimize the mean-square approximation error. Our key observation is that the implicit regularization inherent in DNs enables them to capture nonlinear signal structures (e.g., manifolds) that are out of the reach of classical linear methods like the singular value decomposition (SVD) and principal components analysis (PCA). Furthermore, in contrast to the SVD and PCA, whose performance deteriorates when the tensor's entries deviate from additive white Gaussian noise, we demonstrate that the performance of DeepTensor is robust to a wide range of distributions. We validate that DeepTensor is a robust and computationally efficient drop-in replacement for the SVD, PCA, nonnegative matrix factorization (NMF), and similar decompositions by exploring a range of real-world applications, including hyperspectral image denoising, 3D MRI tomography, and image classification. In particular, DeepTensor offers a 6 dB signal-to-noise ratio improvement over standard denoising methods for signal corrupted by Poisson noise and learns to decompose 3D tensors 60 times faster than a single DN equipped with 3D convolutions. Vishwanath Saragadam, Randall Balestriero, Ashok Veeraraghavan, Richard G. Baraniuk |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Thermal Spread Functions (TSF): Physics-Guided Material ClassificationabstractRobust and non-destructive material classification is a challenging but crucial first-step in numerous vision applications. We propose a physics-guided material classification framework that relies on thermal properties of the object. Our key observation is that the rate of heating and cooling of an object depends on the unique intrinsic properties of the material, namely the emissivity and diffusivity. We leverage this observation by gently heating the objects in the scene with a low-power laser for a fixed duration and then turning it off, while a thermal camera captures measurements during the heating and cooling process. We then take this spatial and temporal “thermal spread function” (TSF) to solve an inverse heat equation using the finite-differences approach, resulting in a spatially varying estimate of diffusivity and emissivity. These tuples are then used to train a classifier that produces a fine-grained material label at each spatial pixel. Our approach is extremely simple requiring only a small light source (low power laser) and a thermal camera, and produces robust classification results with 86% accuracy over 16 classes11Code: https://github.com/aniketdashpute/TSF. Aniket Dashpute, Vishwanath Saragadam, Emma Alexander, Florian Willomitzer, Aggelos K. Katsaggelos, Ashok Veeraraghavan, Oliver Cossairt |
CVPR | 6 |
| 2023 | WIRE: Wavelet Implicit Neural RepresentationsabstractImplicit neural representations (INRs) have recently advanced numerous vision-related areas. INR performance depends strongly on the choice of activation function employed in its MLP network. A wide range of nonlinearities have been explored, but, unfortunately, current INRs designed to have high accuracy also suffer from poor robustness (to signal noise, parameter variation, etc.). Inspired by harmonic analysis, we develop a new, highly accurate and robust INR that does not exhibit this trade off. Our Wavelet Implicit neural REpresentation (WIRE) uses as its activation function the complex Gabor wavelet that is well-known to be optimally concentrated in space-frequency and to have excellent biases for representing images. A wide range of experiments (image denoising, image inpainting, super-resolution, computed tomography reconstruction, image over fitting, and novel view synthesis with neural radiance fields) demonstrate that WIRE defines the new state of the art in INR accuracy, training time, and robustness. Vishwanath Saragadam, Daniel LeJeune, Jasper Tan, Guha Balakrishnan, Ashok Veeraraghavan, Richard G. Baraniuk |
CVPR | 5 |
| 2023 | Role of Transients in Two-Bounce Non-Line-of-Sight ImagingabstractThe goal of non-line-of-sight (NLOS) imaging is to image objects occluded from the camera's field of view using multiply scattered light. Recent works have demonstrated the feasibility of two-bounce (2B) NLOS imaging by scanning a laser and measuring cast shadows of occluded objects in scenes with two relay surfaces. In this work, we study the role of time-of-flight (ToF) measurements, i.e. transients, in 2B-NLOS under multiplexed illumination. Specifically, we study how ToF information can reduce the number of measurements and spatial resolution needed for shape reconstruction. We present our findings with respect to tradeoffs in (1) temporal resolution, (2) spatial resolution, and (3) number of image captures by studying SNR and recoverability as functions of system parameters. This leads to a formal definition of the mathematical constraints for 2B lidar. We believe that our work lays an analytical ground- workfor design of future NLOS imaging systems, especially as ToF sensors become increasingly ubiquitous. Siddharth Somasundaram, Akshat Dave, Connor Henley, Ashok Veeraraghavan, Ramesh Raskar |
CVPR | 4 |
| 2023 | ORCa: Glossy Objects as Radiance-Field CamerasabstractReflections on glossy objects contain valuable and hidden information about the surrounding environment. By converting these objects into cameras, we can unlock exciting applications, including imaging beyond the camera's field-of-view and from seemingly impossible vantage points, e.g. from reflections on the human eye. However, this task is challenging because reflections depend jointly on object geometry, material properties, the 3D environment, and the observer's viewing direction. Our approach converts glossy objects with unknown geometry into radiance-field cameras to image the world from the object's perspective. Our key insight is to convert the object surface into a virtual sensor that captures cast reflections as a 2D projection of the 5D environment radiance field visible to and surrounding the object. We show that recovering the environment radiance fields enables depth and radiance estimation from the object to its surroundings in addition to beyond field-of-view novel-view synthesis, i.e. rendering of novel views that are only directly visible to the glossy object present in the scene, but not the observer. Moreover, using the radiance field we can image around occluders caused by close-by objects in the scene. Our method is trained end-to-end on multi-view images of the object and jointly estimates object geometry, diffuse radiance, and the 5D environment radiance field. For more information, visit our website. Kushagra Tiwary, Akshat Dave, Nikhil Behari, Tzofi Klinghoffer, Ashok Veeraraghavan, Ramesh Raskar |
CVPR | 5 |
| 2023 | Designing Optics and Algorithm for Ultra-Thin, High-Speed Lensless CamerasabstractThere is a growing demand for small, light-weight and low-latency cameras in the robotics and AR/VR community. Mask-based lensless cameras, by design, provide a combined advantage of form-factor, weight and speed. They do so by replacing the classical lens with a thin optical mask and computation. Recent works have explored deep learning based post-processing operations on lensless captures that allow high quality scene reconstruction. However, the ability of deep learning to find the optimal optics for thin lensless cameras has not been explored. In this work, we propose a learning based framework for designing the optics of thin lensless cameras. To highlight the effectiveness of our framework, we learn the optical phase mask for multiple tasks using physics-based neural networks. Specifically, we learn the optimal mask using a weighted loss defined for the following tasks-2D scene reconstructions, optical flow estimation and face detection. We show that mask learned through this framework is better than heuristically designed masks especially for small sensors sizes that allow lower bandwidth and faster readout. Finally, we verify the performance of our learned phase-mask on real data. Salman Siddique Khan, Vivek Boominathan, Ashok Veeraraghavan, Kaushik Mitra |
ICME | 3 |
| 2022 | PANDORA: Polarization-Aided Neural Decomposition of Radiance
Akshat Dave, Yongyi Zhao, Ashok Veeraraghavan |
ECCV (7) | 3 |
| 2022 | MINER: Multiscale Implicit Neural Representation
Vishwanath Saragadam, Jasper Tan, Guha Balakrishnan, Richard G. Baraniuk, Ashok Veeraraghavan |
ECCV (23) | 5 |
| 2022 | Learning Phase Mask for Privacy-Preserving Passive Depth Estimation
Zaid Tasneem, Giovanni Milione, Yi-Hsuan Tsai, Xiang Yu 0002, Ashok Veeraraghavan, Manmohan Krishna Chandraker, Francesco Pittaluga |
ECCV (7) | 5 |
| 2022 | First Arrival Differential LiDARabstractSingle-photon avalanche diode (SPAD) based LiDAR is becoming the de-facto choice for 3D imaging in many emerging applications. However, they suffer from three significant limitations: (a) the additional time-of-arrival dimension results in a data throughput bottleneck, (b) limited spatial resolution due to either low fill-factor (flash LiDAR) or scanning time (scanning-based LiDAR), and (c) coarse depth resolution due to quantization of photon timing by existing SPAD timing circuitries. In this paper, we present a novel, in-pixel computing architecture that we term first arrival differential (FAD) LiDAR, where instead of recording quantized time-of-arrival information at individual pixels, we record a temporal differential measurement between pairs of pixels. FAD captures relative order of photon arrivals at the two pixels (within a cycle or laser period) and creates a one-to-one mapping between this differential measurement and depth differences between the two pixels. We perform detailed system analysis and characterization using Monte Carlo simulation, and experimental emulation using a scanning-based single-photon avalanche diode. FAD pixels can result in a 10–100x reduction in per-pixel data throughput compared to TDC-based pixels. Under the same bandwidth constraints, FAD-LiDAR achieves better depth resolution and/or range than several state-of-the-art TDC-based LiDAR baselines. Mel J. White, Akshat Dave, Shahaboddin Ghajari, Ankit Raghuram, Alyosha C. Molnar, Ashok Veeraraghavan |
ICCP | 7 |
| 2022 | EyeCoD: eye tracking system acceleration via flatcam-based algorithm & accelerator co-designabstractEye tracking has become an essential human-machine interaction modality for providing immersive experience in numerous virtual and augmented reality (VR/AR) applications desiring high throughput (e.g., 240 FPS), small-form, and enhanced visual privacy. However, existing eye tracking systems are still limited by their: (1) large form-factor largely due to the adopted bulky lens-based cameras; (2) high communication cost required between the camera and backend processor; and (3) potentially concerned low visual privacy, thus prohibiting their more extensive applications. To this end, we propose, develop, and validate a lensless FlatCambased eye tracking algorithm and accelerator co-design framework dubbed EyeCoD to enable eye tracking systems with a much reduced form-factor and boosted system efficiency without sacrificing the tracking accuracy, paving the way for next-generation eye tracking solutions. On the system level, we advocate the use of lensless FlatCams instead of lens-based cameras to facilitate the small form-factor need in mobile eye tracking systems, which also leaves rooms for a dedicated sensing-processor co-design to reduce the required camera-processor communication latency. On the algorithm level, EyeCoD integrates a predict-then-focus pipeline that first predicts the region-of-interest (ROI) via segmentation and then only focuses on the ROI parts to estimate gaze directions, greatly reducing redundant computations and data movements. On the hardware level, we further develop a dedicated accelerator that (1) integrates a novel workload orchestration between the aforementioned segmentation and gaze estimation models, (2) leverages intra-channel reuse opportunities for depth-wise layers, (3) utilizes input feature-wise partition to save activation memory size, and (4) develops a sequential-write-parallel-read input buffer to alleviate the bandwidth requirement for the activation global buffer. On-silicon measurement and extensive experiments validate that our EyeCoD consistently reduces both the communication and computation costs, leading to an overall system speedup of 10.95×, 3.21×, and 12.85× over general computing platforms including CPUs and GPUs, and a prior-art eye tracking processor called CIS-GEP, respectively, while maintaining the tracking accuracy. Codes are available at https://github.com/RICE-EIC/EyeCoD. Haoran You, Cheng Wan 0005, Yang Zhao 0013, Zhongzhi Yu, Yonggan Fu, Jiayi Yuan 0001, Shang Wu 0003, Yongan Zhang, Chaojian Li, Vivek Boominathan, Ashok Veeraraghavan, Ziyun Li 0001, Yingyan (Celine) Lin |
ISCA | 12 |
| 2022 | A Differential SPAD Array Architecture in 0.18 μm CMOS for HDR ImagingabstractWe propose a scalable architecture for a differential single-photon avalanche diode (D-SPAD) array which generates measurements based on differential time-of-arrival in lieu of absolute time-of-arrival. This design addresses the throughput bottleneck in conventional sensors that record time-of-arrival statistics (such as histograms or raw arrival-time data) directly. In addition, this design also mitigates saturation at the pixel level and at the counter, making it an ideal candidate for use when the scene being imaged covers a high dynamic range (HDR). The differential nature of the data also obviates the need for large digital circuitry such as high bit depth counters or time to digital converters (TDCs). A prototype test structure of 16 pixels was fabricated in 0.18 $\mu$m CMOS, and we show images reconstructed from this chip that illustrate its capabilities. Mel J. White, Shahaboddin Ghajari, Akshat Dave, Ashok Veeraraghavan, Alyosha C. Molnar |
ISCAS | 5 |
| 2022 | FlatNet: Towards Photorealistic Scene Reconstruction From Lensless MeasurementsabstractLensless imaging has emerged as a potential solution towards realizing ultra-miniature cameras by eschewing the bulky lens in a traditional camera. Without a focusing lens, the lensless cameras rely on computational algorithms to recover the scenes from multiplexed measurements. However, the current iterative-optimization-based reconstruction algorithms produce noisier and perceptually poorer images. In this work, we propose a non-iterative deep learning-based reconstruction approach that results in orders of magnitude improvement in image quality for lensless reconstructions. Our approach, called FlatNet, lays down a framework for reconstructing high-quality photorealistic images from mask-based lensless cameras, where the camera's forward model formulation is known. FlatNet consists of two stages: (1) an inversion stage that maps the measurement into a space of intermediate reconstruction by learning parameters within the forward model formulation, and (2) a perceptual enhancement stage that improves the perceptual quality of this intermediate reconstruction. These stages are trained together in an end-to-end manner. We show high-quality reconstructions by performing extensive experiments on real and challenging scenes using two different types of lensless prototypes: one which uses a separable forward model and another, which uses a more general non-separable cropped-convolution model. Our end-to-end approach is fast, produces photorealistic reconstructions, and is easy to adopt for other mask-based lensless cameras. Salman Siddique Khan, Varun Sundar, Vivek Boominathan, Ashok Veeraraghavan, Kaushik Mitra |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | FaceEngage: Robust Estimation of Gameplay Engagement from User-Contributed (YouTube) VideosabstractMeasuring user engagement in interactive tasks can facilitate numerous applications toward optimizing user experience, ranging from eLearning to gaming. However, a significant challenge is the lack of non-contact engagement estimation methods that are robust in unconstrained environments. We present FaceEngage, a non-intrusive engagement estimator leveraging user facial recordings during actual gameplay in naturalistic conditions. Our contributions are three-fold. First, we show the potential of using front-facing videos as training data to build the engagement estimator. We compile FaceEngage Dataset with over 700 picture-in-picture, realisitic, and user-contributed YouTube gaming videos (i.e., with both full-screen game scenes and time-synchronized user facial recordings in subwindows). Second, we develop FaceEngage system, that captures relevant gamer facial features from front-facing recordings to infer task engagement. We implement two FaceEngage pipelines: an estimator trained on user facial motion features inspired by prior psychological works, and a deep learning-enabled estimator. Lastly, we conduct extensive experiments and conclude: (i) certain user facial motion cues (e.g., blink rates, head movements) are engagement-indicative; (ii) our deep learning-enabled FaceEngage pipeline can automatically extract more informative features, outperforming the facial motion feature-based pipeline; (iii) FaceEngage is robust to various video lengths, users/game genres and interpretable. Despite the challenging nature of realistic videos, FaceEngage attains the accuracy of 83.8 percent and leave-one-user-out precision of 79.9 percent, both of which are superior to our face motion-based model. Xu Chen 0011, Li Niu 0002, Ashok Veeraraghavan, Ashutosh Sabharwal |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Near-Infrared Imaging Photoplethysmography During DrivingabstractImaging photoplethysmography (iPPG) could greatly improve driver safety systems by enabling capabilities ranging from identifying driver fatigue to unobtrusive early heart failure detection. Unfortunately, the driving context poses unique challenges to iPPG, including illumination and motion. First, drastic illumination variations present during driving can overwhelm the small intensity-based iPPG signals. Second, significant driver head motion during driving, as well as camera motion (e.g., vibration) make it challenging to recover iPPG signals. To address these two challenges, we present two innovations. First, we demonstrate that we can reduce most outside light variations using narrow-band near-infrared (NIR) video recordings and obtain reliable heart rate estimates. Second, we present a novel optimization algorithm, which we call AutoSparsePPG, that leverages the quasi-periodicity of iPPG signals and achieves better performance than the state-of-the-art methods. In addition, we release the first publicly available driving dataset that contains both NIR and RGB video recordings of a passenger’s face with simultaneous ground truth pulse oximeter recordings. Ewa Magdalena Nowara, Tim K. Marks, Hassan Mansour, Ashok Veeraraghavan |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | CodedStereo: Learned Phase Masks for Large Depth-of-Field StereoabstractConventional stereo suffers from a fundamental trade-off between imaging volume and signal-to-noise ratio (SNR) – due to the conflicting impact of aperture size on both these variables. Inspired by the extended depth of field cameras, we propose a novel end-to-end learning-based technique to overcome this limitation, by introducing a phase mask at the aperture plane of the cameras in a stereo imaging system. The phase mask creates a depth-dependent yet numerically invertible point spread function, allowing us to recover sharp image texture and stereo correspondence over a significantly extended depth of field (EDOF) than conventional stereo. The phase mask pattern, the EDOF image reconstruction, and the stereo disparity estimation are all trained together using an end-to-end learned deep neural network. We perform theoretical analysis and characterization of the proposed approach and show a 6× increase in volume that can be imaged in simulation. We also build an experimental prototype and validate the approach using real-world results acquired using this prototype system. Shiyu Tan, Shoou-I Yu, Ashok Veeraraghavan |
CVPR | 4 |
| 2021 | SACoD: Sensor Algorithm Co-Design Towards Efficient CNN-powered Intelligent PhlatCamabstractThere has been a booming demand for integrating Convolutional Neural Networks (CNNs) powered functionalities into Internet-of-Thing (IoT) devices to enable ubiquitous intelligent "IoT cameras". However, more extensive applications of such IoT systems are still limited by two challenges. First, some applications, especially medicine-and wearable-related ones, impose stringent requirements on the camera form factor. Second, powerful CNNs often require considerable storage and energy cost, whereas IoT devices often suffer from limited resources. PhlatCam, with its form factor potentially reduced by orders of magnitude, has emerged as a promising solution to the first aforementioned challenge, while the second one remains a bottleneck. Existing compression techniques, which can potentially tackle the second challenge, are far from realizing the full potential in storage and energy reduction, because they mostly focus on the CNN algorithm itself. To this end, this work proposes SACoD, a Sensor Algorithm Co-Design framework to develop more efficient CNN-powered PhlatCam. In particular, the mask coded in the Phlat-Cam sensor and the backend CNN model are jointly optimized in terms of both model parameters and architectures via differential neural architecture search. Extensive experiments including both simulation and physical measurement on manufactured masks show that the proposed SACoD framework achieves aggressive model compression and energy savings while maintaining or even boosting the task accuracy, when benchmarking over two state-of-the-art (SOTA) designs with six datasets across four different vision tasks including classification, segmentation, image translation, and face recognition. Our codes are available at: https://github.com/RICE-EIC/SACoD. Yonggan Fu, Yang Zhang 0001, Yue Wang 0036, Zhihan Lyu, Vivek Boominathan, Ashok Veeraraghavan, Yingyan (Celine) Lin |
ICCV | 6 |
| 2021 | The Benefit of Distraction: Denoising Camera-Based Physiological Measurements using Inverse AttentionabstractAttention networks perform well on diverse computer vision tasks. The core idea is that the signal of interest is stronger in some pixels ("foreground"), and by selectively focusing computation on these pixels, networks can extract subtle information buried in noise and other sources of corruption. Our paper is based on one key observation: in many real-world applications, many sources of corruption, such as illumination and motion, are often shared between the "foreground" and the "background" pixels. Can we utilize this to our advantage? We propose the utility of inverse attention networks, which focus on extracting information about these shared sources of corruption. We show that this helps to effectively suppress shared covariates and amplify signal information, resulting in improved performance. We illustrate this on the task of camera-based physiological measurement where the signal of interest is weak and global illumination variations and motion act as significant shared sources of corruption. We perform experiments on three datasets and show that our approach of inverse attention produces state-of-the-art results, increasing the signal-to-noise ratio by up to 5.8 dB, reducing heart rate and breathing rate estimation errors by as much as 30 %, recovering subtle waveform dynamics, and generalizing from RGB to NIR videos without retraining. Ewa Magdalena Nowara, Daniel McDuff, Ashok Veeraraghavan |
ICCV | 3 |
| 2021 | How to Train Neural Networks for Flare RemovalabstractWhen a camera is pointed at a strong light source, the resulting photograph may contain lens flare artifacts. Flares appear in a wide variety of patterns (halos, streaks, color bleeding, haze, etc.) and this diversity in appearance makes flare removal challenging. Existing analytical solutions make strong assumptions about the artifact’s geometry or brightness, and therefore only work well on a small subset of flares. Machine learning techniques have shown success in removing other types of artifacts, like reflections, but have not been widely applied to flare removal due to the lack of training data. To solve this problem, we explicitly model the optical causes of flare either empirically or using wave optics, and generate semi-synthetic pairs of flare-corrupted and clean images. This enables us to train neural networks to remove lens flare for the first time. Experiments show our data synthesis approach is critical for accurate flare removal, and that models trained with our technique generalize well to real lens flares across different scenes, lighting conditions, and cameras. Qiurui He 0001, Tianfan Xue, Rahul Garg 0002, Jiawen Chen 0001, Ashok Veeraraghavan, Jonathan T. Barron |
ICCV | 6 |
| 2021 | SASSI - Super-Pixelated Adaptive Spatio-Spectral ImagingabstractWe introduce a novel video-rate hyperspectral imager with high spatial, temporal and spectral resolutions. Our key hypothesis is that spectral profiles of pixels within each super-pixel tend to be similar. Hence, a scene-adaptive spatial sampling of a hyperspectral scene, guided by its super-pixel segmented image, is capable of obtaining high-quality reconstructions. To achieve this, we acquire an RGB image of the scene, compute its super-pixels, from which we generate a spatial mask of locations where we measure high-resolution spectrum. The hyperspectral image is subsequently estimated by fusing the RGB image and the spectral measurements using a learnable guided filtering approach. Due to low computational complexity of the superpixel estimation step, our setup can capture hyperspectral images of the scenes with little overhead over traditional snapshot hyperspectral cameras, but with significantly higher spatial and spectral resolutions. We validate the proposed technique with extensive simulations as well as a lab prototype that measures hyperspectral video at a spatial resolution of 600 ×900 pixels, at a spectral resolution of 10 nm over visible wavebands, and achieving a frame rate at 18fps. Vishwanath Saragadam, Michael DeZeeuw, Richard G. Baraniuk, Ashok Veeraraghavan, Aswin C. Sankaranarayanan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | High Resolution, Deep Imaging Using Confocal Time-of-Flight Diffuse Optical TomographyabstractLight scattering by tissue severely limits how deep beneath the surface one can image, and the spatial resolution one can obtain from these images. Diffuse optical tomography (DOT) is one of the most powerful techniques for imaging deep within tissue - well beyond the conventional ∼ 10-15 mean scattering lengths tolerated by ballistic imaging techniques such as confocal and two-photon microscopy. Unfortunately, existing DOT systems are limited, achieving only centimeter-scale resolution. Furthermore, they suffer from slow acquisition times and slow reconstruction speeds making real-time imaging infeasible. We show that time-of-flight diffuse optical tomography (ToF-DOT) and its confocal variant (CToF-DOT), by exploiting the photon travel time information, allow us to achieve millimeter spatial resolution in the highly scattered diffusion regime ( mean free paths). In addition, we demonstrate two additional innovations: focusing on confocal measurements, and multiplexing the illumination sources allow us to significantly reduce the measurement acquisition time. Finally, we rely on a novel convolutional approximation that allows us to develop a fast reconstruction algorithm, achieving a 100× speedup in reconstruction time compared to traditional DOT reconstruction techniques. Together, we believe that these technical advances serve as the first step towards real-time, millimeter resolution, deep tissue imaging using DOT. Yongyi Zhao, Ankit Raghuram, Hyun Keol Kim, Andreas H. Hielscher, Jacob T. Robinson, Ashok Veeraraghavan |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | 3PointTM: Faster Measurement of High-Dimensional Transmission Matrices
Yujun Chen, Ashutosh Sabharwal, Ashok Veeraraghavan, Aswin C. Sankaranarayanan |
ECCV (8) | 4 |
| 2020 | FreeCam3D: Snapshot Structured Light 3D with Freely-Moving Cameras
Vivek Boominathan, Jacob T. Robinson, Hiroshi Kawasaki, Aswin C. Sankaranarayanan, Ashok Veeraraghavan |
ECCV (27) | 7 |
| 2020 | Fast confocal microscopy imaging based on deep learningabstractConfocal microscopy is the de-facto standard technique in bio-imaging for acquiring 3D images in the presence of tissue scattering. However, the point-scanning mechanism inherent in confocal microscopy implies that the capture speed is much too slow for imaging dynamic objects at sufficient spatial resolution and signal to noise ratio(SNR). In this paper, we propose an algorithm for super-resolution confocal microscopy that allows us to capture high-resolution, high SNR confocal images at an order of magnitude faster acquisition speed. The proposed Back-Projection Generative Adversarial Network (BPGAN) consists of a feature extraction step followed by a back-projection feedback module (BPFM) and an associated reconstruction network, these together allow for super-resolution of low-resolution confocal scans. We validate our method using real confocal captures of multiple biological specimens and the results demonstrate that our proposed BPGAN is able to achieve similar quality to high-resolution confocal scans while the imaging speed can be up to 64 times faster. Xiu Li 0001, Jiuyang Dong, Yongbing Zhang 0002, Ashok Veeraraghavan, Xiangyang Ji |
ICCP | 6 |
| 2020 | WISHED: Wavefront imaging sensor with high resolution and depth rangingabstractPhase-retrieval based wavefront sensors have been shown to reconstruct the complex field from an object with a high spatial resolution. Although the reconstructed complex field encodes the depth information of the object, it is impractical to be used as a depth sensor for macroscopic objects, since the unambiguous depth imaging range is limited by the optical wavelength. To improve the depth range of imaging and handle depth discontinuities, we propose a novel three-dimensional sensor by leveraging wavelength diversity and wavefront sensing. Complex fields at two optical wavelengths are recorded, and a synthetic wavelength can be generated by correlating those wavefronts. The proposed system achieves high lateral and depth resolutions. Our experimental prototype shows an unambiguous range of more than 1,000 x larger compared with the optical wavelengths, while the depth precision is up to 9µm for smooth objects and up to 69µm for rough objects. We experimentally demonstrate 3D reconstructions for transparent, translucent, and opaque objects with smooth and rough surfaces. Fengqiang Li, Florian Willomitzer, Ashok Veeraraghavan, Oliver Cossairt |
ICCP | 4 |
| 2020 | CANOPIC: Pre-Digital Privacy-Enhancing Encodings for Computer VisionabstractThe standard pipeline for many vision tasks uses a conventional camera to capture an image that is then passed to a digital processor for information extraction. In some deployments, such as private locations, the captured digital imagery contains sensitive information exposed to digital vulnerabilities such as spyware, Trojans, etc. However, in many applications, the full imagery is unnecessary for the vision task at hand. In this paper we propose an optical and analog system that preprocesses the light from the scene before it reaches the digital imager to destroy sensitive information. We explore analog and optical encodings consisting of easily implementable operations such as convolution, pooling, and quantization. We perform a case study to evaluate how such encodings can destroy face identity information while preserving enough information for face detection. The encoding parameters are learned via an alternating optimization scheme based on adversarial learning with deep neural networks. We name our system CAnOPIC (Camera with Analog and Optical Privacy-Integrating Computations) and show that it has better performance in terms of both privacy and utility than conventional optical privacy-enhancing methods such as blurring and pixelation. Jasper Tan, Salman Siddique Khan, Vivek Boominathan, Jeffrey Byrne, Richard G. Baraniuk, Kaushik Mitra, Ashok Veeraraghavan |
ICME | 7 |
| 2020 | PhlatCam: Designed Phase-Mask Based Thin Lensless CameraabstractWe demonstrate a versatile thin lensless camera with a designed phase-mask placed at sub-2 mm from an imaging CMOS sensor. Using wave optics and phase retrieval methods, we present a general-purpose framework to create phase-masks that achieve desired sharp point-spread-functions (PSFs) for desired camera thicknesses. From a single 2D encoded measurement, we show the reconstruction of high-resolution 2D images, computational refocusing, and 3D imaging. This ability is made possible by our proposed high-performance contour-based PSF. The heuristic contour-based PSF is designed using concepts in signal processing to achieve maximal information transfer to a bit-depth limited sensor. Due to the efficient coding, we can use fast linear methods for high-quality image reconstructions and switch to iterative nonlinear methods for higher fidelity reconstructions and 3D imaging. Vivek Boominathan, Jesse K. Adams, Jacob T. Robinson, Ashok Veeraraghavan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | SNLOS: Non-line-of-sight Scanning through Temporal FocusingabstractOver the last decade, several techniques have been developed for looking around the corner by exploiting the round-trip travel time of photons. Typically, these techniques necessitate the collection of a large number of measurements with varying virtual source and virtual detector locations. This data is then processed by a reconstruction algorithm to estimate the hidden scene. As a consequence, even when the region of interest in the hidden volume is small and limited, the acquisition time needed is large as the entire dataset has to be acquired and then processed.In this paper, we present the first example of scanning based non-line-of-sight imaging technique. The key idea is that if the virtual sources (pulsed sources) on the wall are delayed using a quadratic delay profile (much like the quadratic phase of a focusing lens), then these pulses arrive at the same instant at a single point in the hidden volume – the point being scanned. On the imaging side, applying quadratic delays to the virtual detectors before integration on a single gated detector allows us to ‘focus’ and scan each point in the hidden volume. By changing the quadratic delay profiles, we can focus light at different points in the hidden volume. This provides the first example of scanning based non-line-of-sight imaging, allowing us to focus our measurements only in the region of interest. We derive the theoretical underpinnings of ‘temporal focusing’, show compelling simulations of performance analysis, build a hardware prototype system and demonstrate real results. Adithya Kumar Pediredla, Akshat Dave, Ashok Veeraraghavan |
ICCP | 3 |
| 2019 | STORM: Super-resolving Transients by OveRsampled MeasurementsabstractImage sensors that can measure the time of travel of photons are gaining importance in a myriad of applications such as LIDAR, non-line of sight imaging, light-in-flight imaging, and imaging through scattering media. While the price of these sensors is dramatically shrinking, there remains a trade-off between spatial resolution and temporal resolution. While single-pixel detectors using the single photon avalanche diode (SPAD) technology can achieve 10-30 ps time resolution, the current generation array detectors can only produce an order of magnitude lower temporal resolution due to space-related fabrication constraints. Moreover, this limit is due to bandwidth, read-out and circuit-area constraints on the detector array and therefore unlikely to dramatically change in the next few years.In this paper, we demonstrate a computational imaging approach that utilizes multiple measurements with calibrated sub-temporal resolution delays on the illumination pulse and super-resolution post-processing algorithms that together can achieve an order of magnitude improvement in the time resolution of the acquired transients. We build an experimental prototype, using a 32 × 32 SPAD detector array with 400ps time resolution and demonstrate recovery of transients with ≈ 50ps time resolution, an 8× improvement in time resolution resulting in a 5× improvement in depth reconstruction error. Ankit Raghuram, Adithya Kumar Pediredla, Srinivasa G. Narasimhan, Ioannis Gkioulekas, Ashok Veeraraghavan |
ICCP | 5 |
| 2019 | PhaseCam3D - Learning Phase Masks for Passive Single View Depth EstimationabstractThere is an increasing need for passive 3D scanning in many applications that have stringent energy constraints. In this paper, we present an approach for single frame, single viewpoint, passive 3D imaging using a phase mask at the aperture plane of a camera. Our approach relies on an end-to-end optimization framework to jointly learn the optimal phase mask and the reconstruction algorithm that allows an accurate estimation of range image from captured data. Using our optimization framework, we design a new phase mask that performs significantly better than existing approaches. We build a prototype by inserting a phase mask fabricated using photolithography into the aperture plane of a conventional camera and show compelling performance in 3D imaging. Vivek Boominathan, Huaijin G. Chen, Aswin C. Sankaranarayanan, Ashok Veeraraghavan |
ICCP | 5 |
| 2019 | Convolutional Approximations to the General Non-Line-of-Sight Imaging OperatorabstractNon-line-of-sight (NLOS) imaging aims to reconstruct scenes outside the field of view of an imaging system. A common approach is to measure the so-called light transients, which facilitates reconstructions through ellipsoidal tomography that involves solving a linear least-squares. Unfortunately, the corresponding linear operator is very high-dimensional and lacks structures that facilitate fast solvers, and so, the ensuing optimization is a computationally daunting task. We introduce a computationally tractable framework for solving the ellipsoidal tomography problem. Our main observation is that the Gram of the ellipsoidal tomography operator is convolutional, either exactly under certain idealized imaging conditions, or approximately in practice. This, in turn, allows us to obtain the ellipsoidal tomography solution by using efficient deconvolution procedures to solve a linear least-squares problem involving the Gram operator. The computational tractability of our approach also facilitates the use of various regularizers during the deconvolution procedure. We demonstrate the advantages of our framework in a variety of simulated and real experiments. Byeongjoo Ahn, Akshat Dave, Ashok Veeraraghavan, Ioannis Gkioulekas, Aswin C. Sankaranarayanan |
ICCV | 3 |
| 2019 | Towards Photorealistic Reconstruction of Highly Multiplexed Lensless ImagesabstractRecent advancements in fields like Internet of Things (IoT), augmented reality, etc. have led to an unprecedented demand for miniature cameras with low cost that can be integrated anywhere and can be used for distributed monitoring. Mask-based lensless imaging systems make such inexpensive and compact models realizable. However, reduction in the size and cost of these imagers comes at the expense of their image quality due to the high degree of multiplexing inherent in their design. In this paper, we present a method to obtain image reconstructions from mask-based lensless measurements that are more photorealistic than those currently available in the literature. We particularly focus on FlatCam, a lensless imager consisting of a coded mask placed over a bare CMOS sensor. Existing techniques for reconstructing FlatCam measurements suffer from several drawbacks including lower resolution and dynamic range than lens-based cameras. Our approach overcomes these drawbacks using a fully trainable non-iterative deep learning based model. Our approach is based on two stages: an inversion stage that maps the measurement into the space of intermediate reconstruction and a perceptual enhancement stage that improves this intermediate reconstruction based on perceptual and signal distortion metrics. Our proposed method is fast and produces photo-realistic reconstruction as demonstrated on many real and challenging scenes. Salman Siddique Khan, Adarsh V. R, Vivek Boominathan, Jasper Tan, Ashok Veeraraghavan, Kaushik Mitra |
ICCV | 5 |
| 2019 | Zero-Shot Learning via Category-Specific Visual-Semantic Mapping and Label RefinementabstractZero-Shot Learning (ZSL) aims to classify a test instance from an unseen category based on the training instances from seen categories, in which the gap between seen categories and unseen categories is generally bridged via visual-semantic mapping between the low-level visual feature space and the intermediate semantic space. However, the visual-semantic mapping (i.e., projection) learnt based on seen categories may not generalize well to unseen categories, which is known as the projection domain shift in ZSL. To address this projection domain shift issue, we propose a method named Adaptive Embedding ZSL (AEZSL) to learn an adaptive visual-semantic mapping for each unseen category, followed by progressive label refinement. Moreover, to avoid learning visual-semantic mapping for each unseen category in the large-scale classification task, we additionally propose a deep adaptive embedding model named Deep AEZSL (DAEZSL) sharing the similar idea (i.e., visual-semantic mapping should be category-specific and related to the semantic space) with AEZSL, which only needs to be trained once, but can be applied to arbitrary number of unseen categories. Extensive experiments demonstrate that our proposed methods achieve the state-of-theart results for image classification on three small-scale benchmark datasets and one large-scale benchmark dataset. Li Niu 0002, Jianfei Cai 0001, Ashok Veeraraghavan, Liqing Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Ellipsoidal path connections for time-gated renderingabstractDuring the last decade, we have been witnessing the continued development of new time-of-flight imaging devices, and their increased use in numerous and varied applications. However, physics-based rendering techniques that can accurately simulate these devices are still lacking: while existing algorithms are adequate for certain tasks, such as simulating transient cameras, they are very inefficient for simulating time-gated cameras because of the large number of wasted path samples. We take steps towards addressing these deficiencies, by introducing a procedure for efficiently sampling paths with a predetermined length, and incorporating it within rendering frameworks tailored towards simulating time-gated imaging. We use our open-source implementation of the above to empirically demonstrate improved rendering performance in a variety of applications, including simulating proximity sensors, imaging through occlusions, depth-selective cameras, transient imaging in dynamic scenes, and non-line-of-sight imaging. Adithya Kumar Pediredla, Ashok Veeraraghavan, Ioannis Gkioulekas |
ACM Trans. Graph. | 2 |
| 2018 | Learning From Noisy Web Data With Category-Level SupervisionabstractLearning from web data is increasingly popular due to abundant free web resources. However, the performance gap between webly supervised learning and traditional supervised learning is still very large, due to the label noise of web data as well as the domain shift between web data and test data. To fill this gap, most existing methods propose to purify or augment web data using instance-level supervision, which generally requires heavy annotation. Instead, we propose to address the label noise and domain shift by using more accessible category-level supervision. In particular, we build our deep probabilistic framework upon variational autoencoder (VAE), in which classification network and VAE can jointly leverage category-level hybrid information. Then, we extend our method for domain adaptation followed by our low-rank refinement strategy. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our proposed method. Li Niu 0002, Qingtao Tang, Ashok Veeraraghavan, Ashutosh Sabharwal |
CVPR | 3 |
| 2018 | Webly Supervised Learning Meets Zero-Shot Learning: A Hybrid Approach for Fine-Grained ClassificationabstractFine-grained image classification, which targets at distinguishing subtle distinctions among various subordinate categories, remains a very difficult task due to the high annotation cost of enormous fine-grained categories. To cope with the scarcity of well-labeled training images, existing works mainly follow two research directions: 1) utilize freely available web images without human annotation; 2) only annotate some fine-grained categories and transfer the knowledge to other fine-grained categories, which falls into the scope of zero-shot learning (ZSL). However, the above two directions have their own drawbacks. For the first direction, the labels of web images are very noisy and the data distribution between web images and test images are considerably different. For the second direction, the performance gap between ZSL and traditional supervised learning is still very large. The drawbacks of the above two directions motivate us to design a new framework which can jointly leverage both web data and auxiliary labeled categories to predict the test categories that are not associated with any well-labeled training images. Comprehensive experiments on three benchmark datasets demonstrate the effectiveness of our proposed framework. Li Niu 0002, Ashok Veeraraghavan, Ashutosh Sabharwal |
CVPR | 2 |
| 2018 | Reblur2Deblur: Deblurring videos via self-supervised learningabstractMotion blur is a fundamental problem in computer vision as it impacts image quality and hinders inference. Traditional deblurring algorithms leverage the physics of the image formation model and use hand-crafted priors: they usually produce results that better reflect the underlying scene, but present artifacts. Recent learning-based methods implicitly extract the distribution of natural images directly from the data and use it to synthesize plausible images. Their results are impressive, but they are not always faithful to the content of the latent image. We present an approach that bridges the two. Our method fine-tunes existing deblurring neural networks in a self-supervised fashion by enforcing that the output, when blurred based on the optical flow between subsequent frames, matches the input blurry image. We show that our method significantly improves the performance of existing methods on several datasets both visually and in terms of image quality metrics. Huaijin G. Chen, Jinwei Gu, Orazio Gallo, Ming-Yu Liu 0001, Ashok Veeraraghavan, Jan Kautz |
ICCP | 5 |
| 2018 | prDeep: Robust Phase Retrieval with a Flexible Deep NetworkabstractPhase retrieval algorithms have become an important component in many modern computational imaging systems. For instance, in the context of ptychography and speckle correlation imaging, they enable imaging past the diffraction limit and through scattering media, respectively. Unfortunately, traditional phase retrieval algorithms struggle in the presence of noise. Progress has been made recently on developing more robust algorithms using signal priors, but at the expense of limiting the range of supported measurement models (e.g., to Gaussian or coded diffraction patterns). In this work we leverage the regularization-by-denoising framework and a convolutional neural network denoiser to create prDeep, a new phase retrieval algorithm that is both robust and broadly applicable. We test and validate prDeep in simulation to demonstrate that it is robust to noise and can handle a variety of system models. Christopher A. Metzler, Philip Schniter, Ashok Veeraraghavan, Richard G. Baraniuk |
ICML | 3 |
| 2018 | Deep k-Means: Re-Training and Parameter Sharing with Harder Cluster Assignments for Compressing Deep ConvolutionsabstractThe current trend of pushing CNNs deeper with convolutions has created a pressing demand to achieve higher compression gains on CNNs where convolutions dominate the computation and parameter amount (e.g., GoogLeNet, ResNet and Wide ResNet). Further, the high energy consumption of convolutions limits its deployment on mobile devices. To this end, we proposed a simple yet effective scheme for compressing convolutions though applying k-means clustering on the weights, compression is achieved through weight-sharing, by only recording $K$ cluster centers and weight assignment indexes. We then introduced a novel spectrally relaxed $k$-means regularization, which tends to make hard assignments of convolutional layer weights to $K$ learned cluster centers during re-training. We additionally propose an improved set of metrics to estimate energy consumption of CNN hardware implementations, whose estimation results are verified to be consistent with previously proposed energy estimation tool extrapolated from actual hardware measurements. We finally evaluated Deep $k$-Means across several CNN models in terms of both compression ratio and energy consumption reduction, observing promising results without incurring accuracy loss. The code is available at https://github.com/Sandbox3aster/Deep-K-Means Yue Wang 0036, Zhenyu Wu 0002, Zhangyang Wang, Ashok Veeraraghavan, Yingyan (Celine) Lin |
ICML | 5 |
| 2017 | PPGSecure: Biometric Presentation Attack Detection Using PhotopletysmogramsabstractAuthentication of users by exploiting face as a biometric is gaining widespread traction due to recent advances in face detection and recognition algorithms. While face recognition has made rapid advances in its performance, such face-based authentication systems remain vulnerable to biometric presentation attacks. Biometric presentation attacks are varied and the most common attacks include the presentation of a video or photograph on a display device, the presentation of a printed photograph or the presentation of a face mask resembling the user to be authenticated. In this paper, we present PPGSecure, a novel methodology that relies on camera-based physiology measurements to detect and thwart such biometric presentation attacks. PPGSecure uses a photoplethysmogram (PPG), which is an estimate of vital signs from the small color changes in the video observed due to minor pulsatile variations in the volume of blood flowing to the face. We demonstrate that the temporal frequency spectra of the estimated PPG signal for real live individuals are distinctly different than those of presentation attacks and exploit these differences to detect presentation attacks. We demonstrate that PPGSecure achieves significantly better performance than existing state of the art presentation attack detection methods. Ewa Magdalena Nowara, Ashutosh Sabharwal, Ashok Veeraraghavan |
FG | 3 |
| 2017 | Linear systems approach to identifying performance bounds in indirect imagingabstractLight scattering on diffuse rough surfaces was long assumed to destroy geometry and photometry information about hidden (non line of sight) objects making `looking around the corner' (LATC) and `non line of sight' (NLOS) imaging impractical. Recent work pioneered by Kirmani et al. [1], Velten et al. [2] demonstrated that transient information (time of flight information) from these scattered third bounce photons can be exploited to solve LATC and NLOS imaging. In this paper, we quantify the geometric and photometric reconstruction limits of LATC and NLOS imaging for the first time using a classical linear systems approach. The relationship between the albedo of the voxels in a hidden volume to the third bounce measurements at the sensor is a linear system that is determined by the geometry and the illumination source. We study this linear system and employ empirical techniques to find the limits of the information contained in the third bounce photons as a function of various system parameters. Adithya Kumar Pediredla, Nathan Matsuda, Oliver Cossairt, Ashok Veeraraghavan |
ICASSP | 4 |
| 2017 | Flat focus: depth of field analysis for the FlatCam lensless imaging systemabstractLensless imaging systems, such as the recently proposed FlatCam, offer numerous advantages over lens-based systems such as a thin form-factor, low cost, and higher light throughput. However, little work has been done in analyzing these systems' depth of field characteristics. A depth-dependent calibration step is necessary to obtain the image from the FlatCam measurements, and this calibration determines the system's depth of field. In this paper, we characterize the FlatCam's depth of field properties and show that (a) for scene depths on the order of tens of centimeters, it is possible to perform depth-selective refocusing from a single captured image and (b) for sufficiently large scene depths, calibratin.g for one depth can provide a very large depth of field. Jasper Tan, Vivek Boominathan, Ashok Veeraraghavan, Richard G. Baraniuk |
ICASSP | 3 |
| 2017 | Coherent inverse scattering via transmission matrices: Efficient phase retrieval algorithms and a public datasetabstractA transmission matrix describes the input-output relationship of a complex wavefront as it passes through/reflects off a multiple-scattering medium, such as frosted glass or a painted wall. Knowing a medium's transmission matrix enables one to image through the medium, send signals through the medium, or even use the medium as a lens. The double phase retrieval method is a recently proposed technique to learn a medium's transmission matrix that avoids difficult-to-capture interferometric measurements. Unfortunately, to perform high resolution imaging, existing double phase retrieval methods require (1) a large number of measurements and (2) an unreasonable amount of computation. In this work we focus on the latter of these two problems and reduce computation times with two distinct methods: First, we develop a new phase retrieval algorithm that is significantly faster than existing methods, especially when used with an amplitude-only spatial light modulator (SLM). Second, we calibrate the system using a phase-only SLM, rather than an amplitude-only SLM which was used in previous double phase retrieval experiments. This seemingly trivial change enables us to use a far faster class of phase retrieval algorithms. As a result of these advances, we achieve a 100x reduction in computation times, thereby allowing us to image through scattering media at state-of-the-art resolutions. In addition to these advances, we also release the first publicly available transmission matrix dataset. This contribution will enable phase retrieval researchers to apply their algorithms to real data. Of particular interest to this community, our measurement vectors are naturally i.i.d. subgaussian, i.e., no coded diffraction pattern is required. Christopher A. Metzler, Sudarshan Nagesh, Richard G. Baraniuk, Oliver Cossairt, Ashok Veeraraghavan |
ICCP | 6 |
| 2017 | Reconstructing rooms using photon echoes: A plane based model and reconstruction algorithm for looking around the cornerabstractCan we reconstruct the entire internal shape of a room if all we can directly observe is a small portion of one internal wall, presumably through a window in the room? While conventional wisdom may indicate that this is not possible, motivated by recent work on `looking around corners', we show that one can exploit light echoes to reconstruct the internal shape of hidden rooms. Existing techniques for looking around the corner using transient images model the hidden volume using voxels and try to explain the captured transient response as the sum of the transient responses obtained from individual voxels. Such a technique inherently suffers from challenges with regards to low signal to background ratios (SBR) and has difficulty scaling to larger volumes. In contrast, in this paper, we argue for using a plane-based model for the hidden surfaces. We demonstrate that such a plane-based model results in much higher SBR while simultaneously being amenable to larger spatial scales. We build an experimental prototype composed of a pulsed laser source and a single-photon avalanche detector (SPAD) that can achieve a time resolution of about 30ps and demonstrate high-fidelity reconstructions both of individual planes in a hidden volume and for reconstructing entire polygonal rooms composed of multiple planar walls. Adithya Kumar Pediredla, Mauro Buttafava, Alberto Tosi, Oliver Cossairt, Ashok Veeraraghavan |
ICCP | 5 |
| 2017 | TabletGaze: dataset and analysis for unconstrained appearance-based gaze estimation in mobile tablets
Qiong Huang 0002, Ashok Veeraraghavan, Ashutosh Sabharwal |
Mach. Vis. Appl. | 2 |
| 2016 | ASP Vision: Optically Computing the First Layer of Convolutional Neural Networks Using Angle Sensitive PixelsabstractDeep learning using convolutional neural networks (CNNs) is quickly becoming the state-of-the-art for challenging computer vision applications. However, deep learning's power consumption and bandwidth requirements currently limit its application in embedded and mobile systems with tight energy budgets. In this paper, we explore the energy savings of optically computing the first layer of CNNs. To do so, we utilize bio-inspired Angle Sensitive Pixels (ASPs), custom CMOS diffractive image sensors which act similar to Gabor filter banks in the V1 layer of the human visual cortex. ASPs replace both image sensing and the first layer of a conventional CNN by directly performing optical edge filtering, saving sensing energy, data bandwidth, and CNN FLOPS to compute. Our experimental results (both on synthetic data and a hardware prototype) for a variety of vision tasks such as digit recognition, object recognition, and face identification demonstrate 97% reduction in image sensor power consumption and 90% reduction in data bandwidth from sensor to CPU, while achieving similar performance compared to traditional deep learning pipelines. Huaijin G. Chen, Suren Jayasuriya, Jiyue Yang, Judy Stephen, Sriram Sivaramakrishnan, Ashok Veeraraghavan, Alyosha C. Molnar |
CVPR | 6 |
| 2016 | Shape and reflectance from two-bounce light transientsabstractComputer vision and image-based inference have predominantly focused on extracting scene information by assuming that the camera measures direct light transport (i.e., single-bounce light paths). As a consequence, strong multi-bounce effects are treated typically as sources of noise and, in many scenarios, the presence of such effects can result in gross errors in the estimates of shape and reflectance. This paper provides the theoretical and algorithmic foundations for shape and reflectance estimation from two-bounce light transients, i.e., scenarios where photons from a light source interact with the scene exactly twice before reaching the sensor We derive sufficient conditions for exact recovery of shape and reflectance given lengths and intensities associated with two-bounce light paths. We also develop algorithms for recovery of shape and reflectance, and validate these on a range of simulated scenes. Chia-Yin Tsai, Ashok Veeraraghavan, Aswin C. Sankaranarayanan |
ICCP | 2 |
| 2016 | Focal-sweep for large aperture time-of-flight camerasabstractTime-of-flight (ToF) imaging is an active method that utilizes a temporally modulated light source and a correlation-based (or lock-in) imager that computes the round-trip travel time from source to scene and back. Much like conventional imaging ToF cameras suffer from the trade-off between depth of field (DOF) and light throughput-larger apertures allow for more light collection but results in lower DoF. This trade-off is especially crucial in ToF systems since they require active illumination and have limited power, which limits performance in long-range imaging or imaging in strong ambient illumination (such as outdoors). Motivated by recent work in extended depth of field imaging for photography, we propose a focal sweep-based image acquisition methodology to increase depth-of-field and eliminate defocus blur. Our approach allows for a simple inversion algorithm to recover all-in-focus images. We validate our technique through simulation and experimental results. We demonstrate a proof-of-concept focal sweep time-of-flight acquisition system and show results for a real scene. Sagar Honnungar, Jason Holloway, Adithya Kumar Pediredla, Ashok Veeraraghavan, Kaushik Mitra |
ICIP | 4 |
| 2016 | Spatial Phase-Sweep: Increasing temporal resolution of transient imaging using a light source arrayabstractTransient imaging techniques capture the propagation of an ultra-short pulse of light through a scene, which in effect captures the optical impulse response of the scene. Recently, it has been shown that we can capture transient images using commercial, correlation imager based Time-of-Flight (ToF) systems. But the temporal resolution of these transient images are currently limited by high-speed electronics. In this paper, we propose `Spatial Phase-Sweep' (SPS), a technique that exploits the speed of light to increase the temporal resolution of transient imaging beyond the limit imposed by electronic circuits in these commercial ToF sensors. SPS uses a linear array of light sources with a controlled spatial separation between these sources. The differential positioning of these sources introduce sub nano-second time shifts in the light wavefront, improving the time resolution of captured transients. As a proof of concept, we demonstrate a prototype which improves the temporal resolution of transient imaging by a factor of 10x, without any modification to the underlying electronics. Ryuichi Tadano, Adithya Kumar Pediredla, Kaushik Mitra, Ashok Veeraraghavan |
ICIP | 4 |
| 2016 | Direct face detection and video reconstruction from event camerasabstractEvent cameras are emerging as a new class of cameras, to potentially rival conventional CMOS cameras, because of their high speed operation and low power consumption. Pixels in an event camera operate in parallel and fire asynchronous spikes when individual pixels encounter a change in intensity that is greater than a pre-determined threshold. Such event-based cameras have an immense potential in battery-operated or always-on application scenarios, owing to their low power consumption. These event-based cameras can be used for direct detection from event streams, and we demonstrate this potential using face detection as an example application. We first propose and develop a patch-based model for the event streams acquired from such cameras. We demonstrate the utility and robustness of the patch-based model for event-based video reconstruction and event-based direct face detection. We are able to reconstruct images and videos at over 2,000 fps from the acquired event streams. In addition, we demonstrate the first direct face detection from event streams, highlighting the potential of these event-based cameras for power-efficient vision applications. Souptik Barua, Yoshitaka Miyatani, Ashok Veeraraghavan |
WACV | 3 |
| 2015 | Depth Fields: Extending Light Field Techniques to Time-of-Flight ImagingabstractA variety of techniques such as light field, structured illumination, and time-of-flight (TOF) are commonly used for depth acquisition in consumer imaging, robotics and many other applications. Unfortunately, each technique suffers from its individual limitations preventing robust depth sensing. In this paper, we explore the strengths and weaknesses of combining light field and time-of-flight imaging, particularly the feasibility of an on-chip implementation as a single hybrid depth sensor. We refer to this combination as depth field imaging. Depth fields combine light field advantages such as synthetic aperture refocusing with TOF imaging advantages such as high depth resolution and coded signal processing to resolve multipath interference. We show applications including synthesizing virtual apertures for TOF imaging, improved depth mapping through partial and scattering occluders, and single frequency TOF phase unwrapping. Utilizing space, angle, and temporal coding, depth fields can improve depth sensing in the wild and generate new insights into the dimensions of light's plenoptic function. Suren Jayasuriya, Adithya Kumar Pediredla, Sriram Sivaramakrishnan, Alyosha C. Molnar, Ashok Veeraraghavan |
3DV | 5 |
| 2015 | FPA-CS: Focal plane array-based compressive imaging in short-wave infraredabstractCameras for imaging in short and mid-wave infrared spectra are significantly more expensive than their counterparts in visible imaging. As a result, high-resolution imaging in those spectrum remains beyond the reach of most consumers. Over the last decade, compressive sensing (CS) has emerged as a potential means to realize inexpensive short-wave infrared cameras. One approach for doing this is the single-pixel camera (SPC) where a single detector acquires coded measurements of a high-resolution image. A computational reconstruction algorithm is then used to recover the image from these coded measurements. Unfortunately, the measurement rate of a SPC is insufficient to enable imaging at high spatial and temporal resolutions. We present a focal plane array-based compressive sensing (FPA-CS) architecture that achieves high spatial and temporal resolutions. The idea is to use an array of SPCs that sense in parallel to increase the measurement rate, and consequently, the achievable spatio-temporal resolution of the camera. We develop a proof-of-concept prototype in the short-wave infrared using a sensor with 64× 64 pixels; the prototype provides a 4096× increase in the measurement rate compared to the SPC and achieves a megapixel resolution at video rate using CS techniques. Huaijin G. Chen, Muhammad Salman Asif, Aswin C. Sankaranarayanan, Ashok Veeraraghavan |
CVPR | 4 |
| 2015 | Depth Selective Camera: A Direct, On-Chip, Programmable Technique for Depth Selectivity in PhotographyabstractTime of flight (ToF) cameras use a temporally modulated light source and measure correlation between the reflected light and a sensor modulation pattern, in order to infer scene depth. In this paper, we show that such correlational sensors can also be used to selectively accept or reject light rays from certain scene depths. The basic idea is to carefully select illumination and sensor modulation patterns such that the correlation is non-zero only in the selected depth range - thus light reflected from objects outside this depth range do not affect the correlational measurements. We demonstrate a prototype depth-selective camera and highlight two potential applications: imaging through scattering media and virtual blue screening. This depth-selectivity can be used to reject back-scattering and reflection from media in front of the subjects of interest, thereby significantly enhancing the ability to image through scattering media-critical for applications such as car navigation in fog and rain. Similarly, such depth selectivity can also be utilized as a virtual blue-screen in cinematography by rejecting light reflecting from background, while selectively retaining light contributions from the foreground subject. Ryuichi Tadano, Adithya Kumar Pediredla, Ashok Veeraraghavan |
ICCV | 3 |
| 2015 | What does a single light-ray reveal about a transparent object?abstractWe address the following problem in refractive shape estimation: given a single light-ray correspondence, what shape information of the transparent object is revealed along the path of the light-ray, assuming that the light-ray refracts twice. We answer this question in the form of two depth-normal ambiguities. First, specifying the surface normal at which refraction occurs constrains the depth to a unique value. Second, specifying the depth at which refraction occurs constrains the surface normal to lie on a 1D curve. These two depth-normal ambiguities are fundamental to shape estimation of transparent objects and can be used to derive additional properties. For example, we show that correspondences from three light-rays passing through a point are needed to correctly estimate its surface normal. Another contribution of this work is that we can reduce the number of views required to reconstruct an object by enforcing shape models. We demonstrate this property on real data where we reconstruct shape of an object, with light-rays observed from a single view, by enforcing a locally planar shape model. Chia-Yin Tsai, Ashok Veeraraghavan, Aswin C. Sankaranarayanan |
ICIP | 2 |
| 2015 | Generalized Assorted Camera Arrays: Robust Cross-Channel Registration and ApplicationsabstractOne popular technique for multimodal imaging is generalized assorted pixels (GAP), where an assorted pixel array on the image sensor allows for multimodal capture. Unfortunately, GAP is limited in its applicability because of the need for multimodal filters that are amenable with semiconductor fabrication processes and results in a fixed multimodal imaging configuration. In this paper, we advocate for generalized assorted camera (GAC) arrays for multimodal imaging--i.e., a camera array with filters of different characteristics placed in front of each camera aperture. The GAC provides us with three distinct advantages over GAP: ease of implementation, flexible application-dependent imaging since filters are external and can be changed and depth information that can be used for enabling novel applications (e.g., postcapture refocusing). The primary challenge in GAC arrays is that since the different modalities are obtained from different viewpoints, there is a need for accurate and efficient cross-channel registration. Traditional approaches such as sum-of-squared differences, sum-of-absolute differences, and mutual information all result in multimodal registration errors. Here, we propose a robust cross-channel matching cost function, based on aligning normalized gradients, which allows us to compute cross-channel subpixel correspondences for scenes exhibiting nontrivial geometry. We highlight the promise of GAC arrays with our cross-channel normalized gradient cost for several applications such as low-light imaging, postcapture refocusing, skin perfusion imaging using color + near infrared, and hyperspectral imaging. Jason Holloway, Kaushik Mitra, Sanjeev J. Koppal, Ashok Veeraraghavan |
IEEE Trans. Image Process. | 4 |
| 2014 | Improving resolution and depth-of-field of light field cameras using a hybrid imaging systemabstractCurrent light field (LF) cameras provide low spatial resolution and limited depth-of-field (DOF) control when compared to traditional digital SLR (DSLR) cameras. We show that a hybrid imaging system consisting of a standard LF camera and a high-resolution standard camera enables (a) achieve high-resolution digital refocusing, (b) better DOF control than LF cameras, and (c) render graceful high-resolution viewpoint variations, all of which were previously unachievable. We propose a simple patch-based algorithm to super-resolve the low-resolution views of the light field using the high-resolution patches captured using a high-resolution SLR camera. The algorithm does not require the LF camera and the DSLR to be co-located or for any calibration information regarding the two imaging systems. We build an example prototype using a Lytro camera (380×380 pixel spatial resolution) and a 18 megapixel (MP) Canon DSLR camera to generate a light field with 11 MP resolution (9× super-resolution) and about 1 over 9thof the DOF of the Lytro camera. We show several experimental results on challenging scenes containing occlusions, specularities and complex non-lambertian materials, demonstrating the effectiveness of our approach. Vivek Boominathan, Kaushik Mitra, Ashok Veeraraghavan |
ICCP | 3 |
| 2014 | Can we beat Hadamard multiplexing? Data driven design and analysis for computational imaging systemsabstractComputational Imaging (CI) systems that exploit optical multiplexing and algorithmic demultiplexing have been shown to improve imaging performance in tasks such as motion deblurring, extended depth of field, light field and hyper-spectral imaging. Design and performance analysis of many of these approaches tend to ignore the role of image priors. It is well known that utilizing statistical image priors significantly improves demultiplexing performance. In this paper, we extend the Gaussian Mixture Model as a data-driven image prior (proposed by Mitra et. al [21]) to under-determined linear systems and study compressive CI methods such as light-field and hyper-spectral imaging. Further, we derive a novel algorithm for optimizing multiplexing matrices that simultaneously accounts for (a) sensor noise (b) image priors and (c) CI design constraints. We use our algorithm to design data-optimal multiplexing matrices for a variety of existing CI designs, and we use these matrices to analyze the performance of CI systems as a function of noise level. Our analysis gives new insight into the optimal performance of CI systems, and how this relates to the performance of classical multiplexing designs such as Hadamard matrices. Kaushik Mitra, Oliver Cossairt, Ashok Veeraraghavan |
ICCP | 3 |
| 2014 | The radon image as plenoptic functionabstractWe introduce a novel plenoptic function that can be directly captured or generated after the fact in plenoptic cameras. Whereas previous approaches represent the plenoptic function over a 4D ray space (as radiance or light field), we introduce the representation of the plenoptic function over a 3D plane space. Our approach uses the Radon plenoptic function instead of the traditional 4D plenoptic function to achieve 3D representation - which promises reduced size, making it suitable for use in mobile devices. Moreover, we show that the original 3D luminous density of the scene can be recovered via the inverse Radon transform. Finally, we demonstrate how various 3D views and differently-focused pictures can be rendered directly from this new representation. Todor G. Georgiev, Salil Tambe, Andrew Lumsdaine, Jennifer Gille, Ashok Veeraraghavan |
ICIP | 5 |
| 2014 | Image classification in natural scenes: Are a few selective spectral channels sufficient?abstractA tenet of object classification is that accuracy improves with an increasing number (and variety) of spectral channels available to the classifier. Hyperspectral images provide hundreds of narrowband measurements over a wide spectral range, and offer superior classification performance over color images. However, hyperspectral data is highly redundant. In this paper we suggest that only 6 measurements are needed to obtain classification results comparable to those realized using hyperspectral data. We present classification results for a natural scene using three imaging modalities: 1) using three broadband color filters (RGB) and three narrowband samples, 2) using six narrowband samples, and 3) using six commonly available optical filters. If these results hold for larger datasets of natural images, recently proposed multispectral image sensors [1, 2] can be used to offer material classification results equal to that of hyperspectral data. Jason Holloway, Tanu Priya, Ashok Veeraraghavan, Saurabh Prasad |
ICIP | 3 |
| 2014 | A Framework for Analysis of Computational Imaging Systems: Role of Signal Prior, Sensor Noise and MultiplexingabstractOver the last decade, a number of computational imaging (CI) systems have been proposed for tasks such as motion deblurring, defocus deblurring and multispectral imaging. These techniques increase the amount of light reaching the sensor via multiplexing and then undo the deleterious effects of multiplexing by appropriate reconstruction algorithms. Given the widespread appeal and the considerable enthusiasm generated by these techniques, a detailed performance analysis of the benefits conferred by this approach is important. Unfortunately, a detailed analysis of CI has proven to be a challenging problem because performance depends equally on three components: (1) the optical multiplexing, (2) the noise characteristics of the sensor, and (3) the reconstruction algorithm which typically uses signal priors. A few recent papers [12], [30], [49] have performed analysis taking multiplexing and noise characteristics into account. However, analysis of CI systems under state-of-the-art reconstruction algorithms, most of which exploit signal prior models, has proven to be unwieldy. In this paper, we present a comprehensive analysis framework incorporating all three components. In order to perform this analysis, we model the signal priors using a Gaussian Mixture Model (GMM). A GMM prior confers two unique characteristics. First, GMM satisfies the universal approximation property which says that any prior density function can be approximated to any fidelity using a GMM with appropriate number of mixtures. Second, a GMM prior lends itself to analytical tractability allowing us to derive simple expressions for the `minimum mean square error' (MMSE) which we use as a metric to characterize the performance of CI systems. We use our framework to analyze several previously proposed CI techniques (focal sweep, flutter shutter, parabolic exposure, etc.), giving conclusive answer to the question: `How much performance gain is due to use of a signal prior and how much is due to multiplexing? Our analysis also clearly shows that multiplexing provides significant performance gains above and beyond the gains obtained due to use of signal priors. Kaushik Mitra, Oliver Cossairt, Ashok Veeraraghavan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Compressive epsilon photography for post-capture control in digital imagingabstractA traditional camera requires the photographer to select the many parameters at capture time. While advances in light field photography have enabled post-capture control of focus and perspective, they suffer from several limitations including lower spatial resolution, need for hardware modifications, and restrictive choice of aperture and focus setting. In this paper, we propose "compressive epsilon photography," a technique for achieving complete post-capture control of focus and aperture in a traditional camera by acquiring a carefully selected set of 8 to 16 images and computationally reconstructing images corresponding to all other focus-aperture settings. We make the following contributions: first, we learn the statistical redundancies in focal-aperture stacks using a Gaussian Mixture Model; second, we derive a greedy sampling strategy for selecting the best focus-aperture settings; and third, we develop an algorithm for reconstructing the entire focal-aperture stack from a few captured images. As a consequence, only a burst of images with carefully selected camera settings are acquired. Post-capture, the user can then select any focal-aperture setting of choice and the corresponding image can be rendered using our algorithm. We show extensive results on several real data sets. Atsushi Ito, Salil Tambe, Kaushik Mitra, Aswin C. Sankaranarayanan, Ashok Veeraraghavan |
ACM Trans. Graph. | 5 |
| 2013 | Towards Motion Aware Light Field Video for Dynamic ScenesabstractCurrent Light Field (LF) cameras offer fixed resolution in space, time and angle which is decided a-priori and is independent of the scene. These cameras either trade-off spatial resolution to capture single-shot LF or tradeoff temporal resolution by assuming a static scene to capture high spatial resolution LF. Thus, capturing high spatial resolution LF video for dynamic scenes remains an open and challenging problem. We present the concept, design and implementation of a LF video camera that allows capturing high resolution LF video. The spatial, angular and temporal resolution are not fixed a-priori and we exploit the scene-specific redundancy in space, time and angle. Our reconstruction is motion-aware and offers a continuum of resolution tradeoff with increasing motion in the scene. The key idea is (a) to design efficient multiplexing matrices that allow resolution tradeoffs, (b) use dictionary learning and sparse representations for robust reconstruction, and (c) perform local motion-aware adaptive reconstruction. We perform extensive analysis and characterize the performance of our motion-aware reconstruction algorithm. We show realistic simulations using a graphics simulator as well as real results using a LCoS based programmable camera. We demonstrate novel results such as high resolution digital refocusing for dynamic moving objects. Salil Tambe, Ashok Veeraraghavan, Amit K. Agrawal |
ICCV | 2 |
| 2013 | A Practical Approach to 3D Scanning in the Presence of Interreflections, Subsurface Scattering and Defocus
Mohit Gupta 0001, Amit K. Agrawal, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
Int. J. Comput. Vis. | 3 |
| 2012 | Flutter Shutter Video Camera for compressive sensing of videosabstractVideo cameras are invariably bandwidth limited and this results in a trade-off between spatial and temporal resolution. Advances in sensor manufacturing technology have tremendously increased the available spatial resolution of modern cameras while simultaneously lowering the costs of these sensors. In stark contrast, hardware improvements in temporal resolution have been modest. One solution to enhance temporal resolution is to use high bandwidth imaging devices such as high speed sensors and camera arrays. Unfortunately, these solutions are expensive. An alternate solution is motivated by recent advances in computational imaging and compressive sensing. Camera designs based on these principles, typically, modulate the incoming video using spatio-temporal light modulators and capture the modulated video at a lower bandwidth. Reconstruction algorithms, motivated by compressive sensing, are subsequently used to recover the high bandwidth video at high fidelity. Though promising, these methods have been limited since they require complex and expensive light modulators that make the techniques difficult to realize in practice. In this paper, we show that a simple coded exposure modulation is sufficient to reconstruct high speed videos. We propose the Flutter Shutter Video Camera (FSVC) in which each exposure of the sensor is temporally coded using an independent pseudo-random sequence. Such exposure coding is easily achieved in modern sensors and is already a feature of several machine vision cameras. We also develop two algorithms for reconstructing the high speed video; the first based on minimizing the total variation of the spatio-temporal slices of the video and the second based on a data driven dictionary based approximation. We perform evaluation on simulated videos and real data to illustrate the robustness of our system. Jason Holloway, Aswin C. Sankaranarayanan, Ashok Veeraraghavan, Salil Tambe |
ICCP | 3 |
| 2012 | Variable focus video: Reconstructing depth and video for dynamic scenesabstractTraditional depth from defocus (DFD) algorithms assume that the camera and the scene are static during acquisition time. In this paper, we examine the effects of camera and scene motion on DFD algorithms. We show that, given accurate estimates of optical flow (OF), one can robustly warp the focal stack (FS) images to obtain a virtual static FS and apply traditional DFD algorithms on the static FS. Acquiring accurate OF in the presence of varying focal blur is a challenging task. We show how defocus blur variations cause inherent biases in the estimates of optical flow. We then show how to robustly handle these biases and compute accurate OF estimates in the presence of varying focal blur. This leads to an architecture and an algorithm that converts a traditional 30 fps video camera into a co-located 30 fps image and a range sensor. Further, the ability to extract image and range information allows us to render images with artistic depth-of field effects, both extending and reducing the depth of field of the captured images. We demonstrate experimental results on challenging scenes captured using a camera prototype. Nitesh Shroff, Ashok Veeraraghavan, Yuichi Taguchi, Oncel Tuzel, Amit K. Agrawal, Rama Chellappa |
ICCP | 2 |
| 2011 | Structured light 3D scanning in the presence of global illuminationabstractGlobal illumination effects such as inter-reflections, diffusion and sub-surface scattering severely degrade the performance of structured light-based 3D scanning. In this paper, we analyze the errors caused by global illumination in structured light-based shape recovery. Based on this analysis, we design structured light patterns that are resilient to individual global illumination effects using simple logical operations and tools from combinatorial mathematics. Scenes exhibiting multiple phenomena are handled by combining results from a small ensemble of such patterns. This combination also allows us to detect any residual errors that are corrected by acquiring a few additional images. Our techniques do not require explicit separation of the direct and global components of scene radiance and hence work even in scenarios where the separation fails or the direct component is too low. Our methods can be readily incorporated into existing scanning systems without significant overhead in terms of capture time or hardware. We show results on a variety of scenes with complex shape and material properties and challenging global illumination effects. Mohit Gupta 0001, Amit K. Agrawal, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
CVPR | 3 |
| 2011 | P2C2: Programmable pixel compressive camera for high speed imagingabstractWe describe an imaging architecture for compressive video sensing termed programmable pixel compressive camera (P2C2). P2C2 allows us to capture fast phenomena at frame rates higher than the camera sensor. In P2C2, each pixel has an independent shutter that is modulated at a rate higher than the camera frame-rate. The observed intensity at a pixel is an integration of the incoming light modulated by its specific shutter. We propose a reconstruction algorithm that uses the data from P2C2 along with additional priors about videos to perform temporal super-resolution. We model the spatial redundancy of videos using sparse representations and the temporal redundancy using brightness constancy constraints inferred via optical flow. We show that by modeling such spatio-temporal redundancies in a video volume, one can faithfully recover the underlying high-speed video frames from the observed low speed coded video. The imaging architecture and the reconstruction algorithm allows us to achieve temporal super-resolution without loss in spatial resolution. We implement a prototype of P2C2 using an LCOS modulator and recover several videos at 200 fps using a 25 fps camera. Dikpal Reddy, Ashok Veeraraghavan, Rama Chellappa |
CVPR | 2 |
| 2011 | Finding a needle in a specular haystackabstractProgress in machine vision algorithms has led to widespread adoption of these techniques to automate several industrial assembly tasks. Nevertheless, shiny or specular objects which are common in industrial environments still present a great challenge for vision systems. In this paper, we take a step towards this problem under the context of vision-aided robotic assembly. We show that when the illumination source moves, the specular highlights remain in a region whose radius is inversely proportional to the surface curvature. This allows us to extract regions of the object that have high surface curvature. These points of high curvature can be used as features for specular objects. Further, an inexpensive multi-flash camera (MFC) design can be used to reliably extract these features. We show that one can use multiple views of the object using the MFC in order to triangulate and obtain the 3D location and pose of the shiny objects. Finally, we show a system consisting of a robot arm with an MFC that can perform automated detection and pose estimation of shiny screws within a cluttered bin, achieving position and orientation errors less than 0.5 mm and 0.8° respectively. Nitesh Shroff, Yuichi Taguchi, Oncel Tuzel, Ashok Veeraraghavan, Srikumar Ramalingam, Haruhisa Okuda |
ICRA | 4 |
| 2011 | A Fast Bilinear Structure from Motion Algorithm Using a Video Sequence and Inertial SensorsabstractIn this paper, we study the benefits of the availability of a specific form of additional information—the vertical direction (gravity) and the height of the camera, both of which can be conveniently measured using inertial sensors and a monocular video sequence for 3D urban modeling. We show that in the presence of this information, the SfM equations can be rewritten in a bilinear form. This allows us to derive a fast, robust, and scalable SfM algorithm for large scale applications. The SfM algorithm developed in this paper is experimentally demonstrated to have favorable properties compared to the sparse bundle adjustment algorithm. We provide experimental evidence indicating that the proposed algorithm converges in many cases to solutions with lower error than state-of-art implementations of bundle adjustment. We also demonstrate that for the case of large reconstruction problems, the proposed algorithm takes lesser time to reach its solution compared to bundle adjustment. We also present SfM results using our algorithm on the Google StreetView research data set. Mahesh Ramachandran, Ashok Veeraraghavan, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Statistical Computations on Grassmann and Stiefel Manifolds for Image and Video-Based RecognitionabstractIn this paper, we examine image and video-based recognition applications where the underlying models have a special structure—the linear subspace structure. We discuss how commonly used parametric models for videos and image sets can be described using the unified framework of Grassmann and Stiefel manifolds. We first show that the parameters of linear dynamic models are finite-dimensional linear subspaces of appropriate dimensions. Unordered image sets as samples from a finite-dimensional linear subspace naturally fall under this framework. We show that an inference over subspaces can be naturally cast as an inference problem on the Grassmann manifold. To perform recognition using subspace-based models, we need tools from the Riemannian geometry of the Grassmann manifold. This involves a study of the geometric properties of the space, appropriate definitions of Riemannian metrics, and definition of geodesics. Further, we derive statistical modeling of inter and intraclass variations that respect the geometry of the space. We apply techniques such as intrinsic and extrinsic statistics to enable maximum-likelihood classification. We also provide algorithms for unsupervised clustering derived from the geometry of the manifold. Finally, we demonstrate the improved performance of these methods in a wide variety of vision applications such as activity recognition, video-based face recognition, object recognition from image sets, and activity-based video clustering. Pavan Turaga, Ashok Veeraraghavan, Anuj Srivastava, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Coded Strobing Photography: Compressive Sensing of High Speed Periodic VideosabstractWe show that, via temporal modulation, one can observe and capture a high-speed periodic video well beyond the abilities of a low-frame-rate camera. By strobing the exposure with unique sequences within the integration time of each frame, we take coded projections of dynamic events. From a sequence of such frames, we reconstruct a high-speed video of the high-frequency periodic process. Strobing is used in entertainment, medical imaging, and industrial inspection to generate lower beat frequencies. But this is limited to scenes with a detectable single dominant frequency and requires high-intensity lighting. In this paper, we address the problem of sub-Nyquist sampling of periodic signals and show designs to capture and reconstruct such signals. The key result is that for such signals, the Nyquist rate constraint can be imposed on the strobe rate rather than the sensor rate. The technique is based on intentional aliasing of the frequency components of the periodic signal while the reconstruction algorithm exploits recent advances in sparse representations and compressive sensing. We exploit the sparsity of periodic signals in the Fourier domain to develop reconstruction algorithms that are inspired by compressive sensing. Ashok Veeraraghavan, Dikpal Reddy, Ramesh Raskar |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Optimal coded sampling for temporal super-resolutionabstractConventional low frame rate cameras result in blur and/or aliasing in images while capturing fast dynamic events. Multiple low speed cameras have been used previously with staggered sampling to increase the temporal resolution. However, previous approaches are inefficient: they either use small integration time for each camera which does not provide light benefit, or use large integration time in a way that requires solving a big ill-posed linear system. We propose coded sampling that address these issues: using N cameras it allows N times temporal superresolution while allowing ~N/2 times more light compared to an equivalent high speed camera. In addition, it results in a well-posed linear system which can be solved independently for each frame, avoiding reconstruction artifacts and significantly reducing the computational time and memory. Our proposed sampling uses optimal multiplexing code considering additive Gaussian noise to achieve the maximum possible SNR in the recovered video. We show how to implement coded sampling on off-the-shelf machine vision cameras. We also propose a new class of invertible codes that allow continuous blur in captured frames, leading to an easier hardware implementation. Amit K. Agrawal, Mohit Gupta 0001, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
CVPR | 3 |
| 2010 | Fast directional chamfer matchingabstractWe study the object localization problem in images given a single hand-drawn example or a gallery of shapes as the object model. Although many shape matching algorithms have been proposed for the problem over the decades, chamfer matching remains to be the preferred method when speed and robustness are considered. In this paper, we significantly improve the accuracy of chamfer matching while reducing the computational time from linear to sublinear (shown empirically). Specifically, we incorporate edge orientation information in the matching algorithm such that the resulting cost function is piecewise smooth and the cost variation is tightly bounded. Moreover, we present a sublinear time algorithm for exact computation of the directional chamfer matching score using techniques from 3D distance transforms and directional integral images. In addition, the smooth cost function allows to bound the cost distribution of large neighborhoods and skip the bad hypotheses within. Experiments show that the proposed approach improves the speed of the original chamfer matching upto an order of 45×, and it is much faster than many state of art techniques while the accuracy is comparable. Ming-Yu Liu 0001, Oncel Tuzel, Ashok Veeraraghavan, Rama Chellappa |
CVPR | 3 |
| 2010 | Robust RVM regression using sparse outlier modelabstractKernel regression techniques such as Relevance Vector Machine (RVM) regression, Support Vector Regression and Gaussian processes are widely used for solving many computer vision problems such as age, head pose, 3D human pose and lighting estimation. However, the presence of outliers in the training dataset makes the estimates from these regression techniques unreliable. In this paper, we propose robust versions of the RVM regression that can handle outliers in the training dataset. We decompose the noise term in the RVM formulation into a (sparse) outlier noise term and a Gaussian noise term. We then estimate the outlier noise along with the model parameters. We present two approaches for solving this estimation problem: (1) a Bayesian approach, which essentially follows the RVM framework and (2) an optimization approach based on Basis Pursuit Denoising. In the Bayesian approach, the robust RVM problem essentially becomes a bigger RVM problem with the advantage that it can be solved efficiently by a fast algorithm. Empirical evaluations, and real experiments on image de-noising and age estimation demonstrate the better performance of the robust RVM algorithms over that of the RVM reg ression. Kaushik Mitra, Ashok Veeraraghavan, Rama Chellappa |
CVPR | 2 |
| 2010 | Specular surface reconstruction from sparse reflection correspondencesabstractWe present a practical approach for surface reconstruction of smooth mirror-like objects using sparse reflection correspondences (RCs). Assuming finite object motion with a fixed camera and un-calibrated environment, we derive the relationship between RC and the surface shape. We show that by locally modeling the surface as a quadric, the relationship between the RCs and unknown surface parameters becomes linear. We develop a simple surface reconstruction algorithm that amounts to solving either an eigenvalue problem or a second order cone program (SOCP). Ours is the first method that allows for reconstruction of mirror surfaces from sparse RCs, obtained from standard algorithms such as SIFT. Our approach overcomes the practical issues in shape from specular flow (SFSF) such as the requirement of dense optical flow and undefined/infinite flow at parabolic points. We also show how to incorporate auxiliary information such as sparse surface normals into our framework. Experiments, both real and synthetic are shown that validate the theory presented. Aswin C. Sankaranarayanan, Ashok Veeraraghavan, Oncel Tuzel, Amit K. Agrawal |
CVPR | 2 |
| 2010 | Axial light field for curved mirrors: Reflect your perspective, widen your viewabstractMirrors have been used to enable wide field-of-view (FOV) catadioptric imaging. The mapping between the incoming and reflected light rays depends non-linearly on the mirror shape and has been well-studied using caustics. We analyze this mapping using two-plane light field parameterization, which provides valuable insight into the geometric structure of reflected rays. Using this analysis, we study the problem of generating a single-viewpoint virtual perspective image for catadioptric systems, which is unachievable for several common configurations. Instead of minimizing distortions appearing in a single image, we propose to capture all the rays required to generate a virtual perspective by capturing a light field. We consider rotationally symmetric mirrors and show that a traditional planar light field results in significant aliasing artifacts. We propose axial light field, captured by moving the camera along the mirror rotation axis, for efficient sampling and to remove aliasing artifacts. This allows us to computationally generate wide FOV virtual perspectives using a wider class of mirrors than before, without using scene priors or depth estimation. We analyze the relationship between the axial light field parameters and the FOV/resolution of the resulting virtual perspective. Real results using a spherical mirror demonstrate generating 140° FOV virtual perspective using multiple 30° FOV images. Yuichi Taguchi, Amit K. Agrawal, Srikumar Ramalingam, Ashok Veeraraghavan |
CVPR | 4 |
| 2010 | Increasing depth resolution of electron microscopy of neural circuits using sparse tomographic reconstructionabstractFuture progress in neuroscience hinges on reconstruction of neuronal circuits to the level of individual synapses. Because of the specifics of neuronal architecture, imaging must be done with very high resolution and throughput. While Electron Microscopy (EM) achieves the required resolution in the transverse directions, its depth resolution is a severe limitation. Computed tomography (CT) may be used in conjunction with electron microscopy to improve the depth resolution, but this severely limits the throughput since several tens or hundreds of EM images need to be acquired. Here, we exploit recent advances in signal processing to obtain high depth resolution EM images computationally. First, we show that the brain tissue can be represented as sparse linear combination of local basis functions that are thin membrane-like structures oriented in various directions. We then develop reconstruction techniques inspired by compressive sensing that can reconstruct the brain tissue from very few (typically 5) tomographic views of each section. This enables tracing of neuronal connections across layers and, hence, high throughput reconstruction of neural circuits to the level of individual synapses. Ashok Veeraraghavan, Alexander Genkin, Shiv Vitaladevuni, Louis K. Scheffer, Harald F. Hess, Richard Fetter, Marco Cantoni, Graham Knott, Dmitri B. Chklovskii |
CVPR | 1 |
| 2010 | Flexible Voxels for Motion-Aware Videography
Mohit Gupta 0001, Amit K. Agrawal, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
ECCV (1) | 3 |
| 2010 | Image Invariants for Smooth Reflective Surfaces
Aswin C. Sankaranarayanan, Ashok Veeraraghavan, Oncel Tuzel, Amit K. Agrawal |
ECCV (2) | 2 |
| 2010 | Robust regression using sparse learning for high dimensional parameter estimation problemsabstractAlgorithms such as Least Median of Squares (LMedS) and Random Sample Consensus (RANSAC) have been very successful for low-dimensional robust regression problems. However, the combinatorial nature of these algorithms makes them practically unusable for high-dimensional applications. In this paper, we introduce algorithms that have cubic time complexity in the dimension of the problem, which make them computationally efficient for high-dimensional problems. We formulate the robust regression problem by projecting the dependent variable onto the null space of the independent variables which receives significant contributions only from the outliers. We then identify the outliers using sparse representation/learning based algorithms. Under certain conditions, that follow from the theory of sparse representation, these polynomial algorithms can accurately solve the robust regression problem which is, in general, a combinatorial problem. We present experimental results that demonstrate the efficacy of the proposed algorithms. We also analyze the intrinsic parameter space of robust regression and identify an efficient and accurate class of algorithms for different operating conditions. An application to facial age estimation is presented. Kaushik Mitra, Ashok Veeraraghavan, Rama Chellappa |
ICASSP | 2 |
| 2010 | Streaming Compressive Sensing for high-speed periodic videosabstractThe ability of Compressive Sensing (CS) to recover sparse signals from limited measurements has been recently exploited in computational imaging to acquire high-speed periodic and near-periodic videos using only a low-speed camera with coded exposure and intensive off-line processing. Each low-speed frame integrates a coded sequence of high-speed frames during its exposure time. The high-speed video can be reconstructed from the low-speed coded frames using a sparse recovery algorithm. This paper presents a new streaming CS algorithm specifically tailored to this application. Our streaming approach allows causal on-line acquisition and reconstruction of the video, with a small, controllable, and guaranteed buffer delay and low computational cost. The algorithm adapts to changes in the signal structure and, thus, outperforms the off-line algorithm in realistic signals. Muhammad Salman Asif, Dikpal Reddy, Petros Boufounos, Ashok Veeraraghavan |
ICIP | 4 |
| 2010 | Pose estimation in heavy clutter using a multi-flash cameraabstractWe propose a novel solution to object detection, localization and pose estimation with applications in robot vision. The proposed method is especially applicable when the objects of interest may not be richly textured and are immersed in heavy clutter. We show that a multi-flash camera (MFC) provides accurate separation of depth edges and texture edges in such scenes. Then, we reformulate the problem, as one of finding matches between the depth edges obtained in one or more MFC images to the rendered depth edges that are computed offline using 3D CAD model of the objects. In order to facilitate accurate matching of these binary depth edge maps, we introduce a novel cost function that respects both the position and the local orientation of each edge pixel. This cost function is significantly superior to traditional Chamfer cost and leads to accurate matching even in heavily cluttered scenes where traditional methods are unreliable. We present a sub-linear time algorithm to compute the cost function using techniques from 3D distance transforms and integral images. Finally, we also propose a multi-view based pose-refinement algorithm to improve the estimated pose. We implemented the algorithm on an industrial robot arm and obtained location and angular estimation accuracy of the order of 1 mm and 2° respectively for a variety of parts with minimal texture. Ming-Yu Liu 0001, Oncel Tuzel, Ashok Veeraraghavan, Rama Chellappa, Amit K. Agrawal, Haruhisa Okuda |
ICRA | 3 |
| 2010 | Reinterpretable Imager: Towards Variable Post-Capture Space, Angle and Time Resolution in PhotographyabstractAbstract We describe a novel multiplexing approach to achieve tradeoffs in space, angle and time resolution in photography. We explore the problem of mapping useful subsets of time‐varying 4D lightfields in a single snapshot. Our design is based on using a dynamic mask in the aperture and a static mask close to the sensor. The key idea is to exploit scene‐specific redundancy along spatial, angular and temporal dimensions and to provide a programmable or variable resolution tradeoff among these dimensions. This allows a user to reinterpret the single captured photo as either a high spatial resolution image, a refocusable image stack or a video for different parts of the scene in post‐processing. A lightfield camera or a video camera forces a‐priori choice in space‐angle‐time resolution. We demonstrate a single prototype which provides flexible post‐capture abilities not possible using either a single‐shot lightfield camera or a multi‐frame video camera. We show several novel results including digital refocusing on objects moving in depth and capturing multiple facial expressions in a single photo. Amit K. Agrawal, Ashok Veeraraghavan, Ramesh Raskar |
Comput. Graph. Forum | 2 |
| 2010 | Axial-cones: modeling spherical catadioptric cameras for wide-angle light field renderingabstractCatadioptric imaging systems are commonly used for wide-angle imaging, but lead to multi-perspective images which do not allow algorithms designed for perspective cameras to be used. Efficient use of such systems requires accurate geometric ray modeling as well as fast algorithms. We present accurate geometric modeling of the multi-perspective photo captured with a spherical catadioptric imaging system usingaxial-cone cameras:multiple perspective cameras lying on an axis each with a different viewpoint and a different cone of rays. This modeling avoids geometric approximations and allows several algorithms developed for perspective cameras to be applied to multi-perspective catadioptric cameras. We demonstrate axial-cone modeling in the context of rendering wide-angle light fields, captured using a spherical mirror array. We present several applications such as spherical distortion correction, digital refocusing for artistic depth of field effects in wide-angle scenes, and wide-angle dense depth estimation. Our GPU implementation using axial-cone modeling achieves up to three orders of magnitude speed up over ray tracing for these applications. Yuichi Taguchi, Amit K. Agrawal, Ashok Veeraraghavan, Srikumar Ramalingam, Ramesh Raskar |
ACM Trans. Graph. | 3 |
| 2009 | Unsupervised view and rate invariant clustering of video sequences
Pavan Turaga, Ashok Veeraraghavan, Rama Chellappa |
Comput. Vis. Image Underst. | 2 |
| 2009 | Rate-Invariant Recognition of Humans and Their ActivitiesabstractPattern recognition in video is a challenging task because of the multitude of spatio-temporal variations that occur in different videos capturing the exact same event. While traditional pattern-theoretic approaches account for the spatial changes that occur due to lighting and pose, very little has been done to address the effect of temporal rate changes in the executions of an event. In this paper, we provide a systematic model-based approach to learn the nature of such temporal variations (time warps) while simultaneously allowing for the spatial variations in the descriptors. We illustrate our approach for the problem of action recognition and provide experimental justification for the importance of accounting for rate variations in action recognition. The model is composed of a nominal activity trajectory and a function space capturing the probability distribution of activity-specific time warping transformations. We use the square-root parameterization of time warps to derive geodesics, distance measures, and probability distributions on the space of time warping functions. We then design a Bayesian algorithm which treats the execution rate function as a nuisance variable and integrates it out using Monte Carlo sampling, to generate estimates of class posteriors. This approach allows us to learn the space of time warps for each activity while simultaneously capturing other intra- and interclass variations. Next, we discuss a special case of this approach which assumes a uniform distribution on the space of time warping functions and show how computationally efficient inference algorithms may be derived for this special case. We discuss the relative advantages and disadvantages of both approaches and show their efficacy using experiments on gait-based person identification and activity recognition. Ashok Veeraraghavan, Anuj Srivastava, Amit K. Roy-Chowdhury, Rama Chellappa |
IEEE Trans. Image Process. | 1 |
| 2008 | Statistical analysis on Stiefel and Grassmann manifolds with applications in computer visionabstractMany applications in computer vision and pattern recognition involve drawing inferences on certain manifold-valued parameters. In order to develop accurate inference algorithms on these manifolds we need to a) understand the geometric structure of these manifolds b) derive appropriate distance measures and c) develop probability distribution functions (pdf) and estimation techniques that are consistent with the geometric structure of these manifolds. In this paper, we consider two related manifolds - the Stiefel manifold and the Grassmann manifold, which arise naturally in several vision applications such as spatio-temporal modeling, affine invariant shape analysis, image matching and learning theory. We show how accurate statistical characterization that reflects the geometry of these manifolds allows us to design efficient algorithms that compare favorably to the state of the art in these very different applications. In particular, we describe appropriate distance measures and parametric and non-parametric density estimators on these manifolds. These methods are then used to learn class conditional densities for applications such as activity recognition, video based face recognition and shape classification. Pavan Turaga, Ashok Veeraraghavan, Rama Chellappa |
CVPR | 2 |
| 2008 | Non-refractive modulators for encoding and capturing scene appearance and depthabstractWe analyze the modulation of a light field via non-refracting attenuators. In the most general case, any desired modulation can be achieved with attenuators having four degrees of freedom in ray-space. We motivate the discussion with a universal 4D ray modulator (ray-filter) which can attenuate the intensity of each ray independently. We describe operation of such a fantasy ray-filter in the context of altering the 4D light field incident on a 2D camera sensor. Ray-filters are difficult to realize in practice but we can achieve reversible encoding for light field capture using patterned attenuating mask. Two mask-based designs are analyzed in this framework. The first design closely mimics the angle-dependent ray-sorting possible with the ray filter. The second design exploits frequency-domain modulation to achieve a more efficient encoding. We extend these designs for optimal sampling of light field by matching the modulation function to the specific shape of the band-limit frequency transform of light field. We also show how a hand-held version of an attenuator based light field camera can be built using a medium-format digital camera and an inexpensive mask. Ashok Veeraraghavan, Amit K. Agrawal, Ramesh Raskar, Ankit Mohan, Jack Tumblin |
CVPR | 1 |
| 2008 | Homography based distributed video coding for a network of camerasabstractNetworks of multiple video cameras are being deployed in several scenarios like surveillance, traffic enforcement, human motion analysis and sports telecast. The unmanageable size of the raw video sequences necessitates the use of compression schemes to efficiently encode these videos. Traditional video compression schemes account for spatial and temporal redundancy in a video sequence. In this paper, we present a compression technique that also leverages the information redundancy between video sequences across different cameras with overlapping fields of view. An algorithm based on homography for efficient compression of multiple video sequences is presented that performs significantly better than current schemes. We also derive an efficient distributed version of the algorithm that can be implemented on a large scale network of cameras. This distributed algorithm minimizes both the communication costs and the compression costs simultaneously. Ashok Veeraraghavan, Mahesh Ramachandran, Mareboyana Manohar |
ICIP | 1 |
| 2008 | Shape-and-Behavior Encoded Tracking of Bee DancesabstractBehavior analysis of social insects has garnered impetus in recent years and has led to some advances in fields like control systems, flight navigation etc. Manual labeling of insect motions required for analyzing the behaviors of insects requires significant investment of time and effort. In this paper, we propose certain general principles that help in simultaneous automatic tracking and behavior analysis with applications in tracking bees and recognizing specific behaviors exhibited by them. The state space for tracking is defined using position, orientation and the current behavior of the insect being tracked. The position and orientation are parametrized using a shape model while the behavior is explicitly modeled using a three-tier hierarchical motion model. The first tier (dynamics) models the local motions exhibited and the models built in this tier act as a vocabulary for behavior modeling. The second tier is a Markov motion model built on top of the local motion vocabulary which serves as the behavior model. The third tier of the hierarchy models the switching between behaviors and this is also modeled as a Markov model. We address issues in learning the three-tier behavioral model, in discriminating between models, detecting and in modeling abnormal behaviors. Another important aspect of this work is that it leads to joint tracking and behavior analysis instead of the traditional track and then recognize approach. We apply these principles for tracking bees in a hive while they are executing the waggle dance and the round dance. Ashok Veeraraghavan, Rama Chellappa, Mandyam V. Srinivasan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Object Detection, Tracking and Recognition for Multiple Smart CamerasabstractVideo cameras are among the most commonly used sensors in a large number of applications, ranging from surveillance to smart rooms for videoconferencing. There is a need to develop algorithms for tasks such as detection, tracking, and recognition of objects, specifically using distributed networks of cameras. The projective nature of imaging sensors provides ample challenges for data association across cameras. We first discuss the nature of these challenges in the context of visual sensor networks. Then, we show how real-world constraints can be favorably exploited in order to tackle these challenges. Examples of real-world constraints are (a) the presence of a world plane, (b) the presence of a three-dimiensional scene model, (c) consistency of motion across cameras, and (d) color and texture properties. In this regard, the main focus of this paper is towards highlighting the efficient use of the geometric constraints induced by the imaging devices to derive distributed algorithms for target detection, tracking, and recognition. Our discussions are supported by several examples drawn from real applications. Lastly, we also describe several potential research problems that remain to be addressed. Aswin C. Sankaranarayanan, Ashok Veeraraghavan, Rama Chellappa |
Proc. IEEE | 2 |
| 2008 | Glare aware photography: 4D ray sampling for reducing glare effects of camera lensesabstractGlare arises due to multiple scattering of light inside the camera's body and lens optics and reduces image contrast. While previous approaches have analyzed glare in 2D image space, we show that glare is inherently a 4D ray-space phenomenon. By statistically analyzing the ray-space inside a camera, we can classify and remove glare artifacts. In ray-space, glare behaves as high frequency noise and can be reduced by outlier rejection. While such analysis can be performed by capturing the light field inside the camera, it results in the loss of spatial resolution. Unlike light field cameras, we do not need to reversibly encode the spatial structure of the ray-space, leading to simpler designs. We explore masks for uniform and non-uniform ray sampling and show a practical solution to analyze the 4D statistics without significantly compromising image resolution. Although diffuse scattering of the lens introduces 4D low-frequency glare, we can produce useful solutions in a variety of common scenarios. Our approach handles photography looking into the sun and photos taken without a hood, removes the effect of lens smudges and reduces loss of contrast due to camera body reflections. We show various applications in contrast enhancement and glare manipulation. Ramesh Raskar, Amit K. Agrawal, Cyrus A. Wilson, Ashok Veeraraghavan |
ACM Trans. Graph. | 4 |
| 2007 | From Videos to Verbs: Mining Videos for Activities using a Cascade of Dynamical SystemsabstractClustering video sequences in order to infer and extract activities from a single video stream is an extremely important problem and has significant potential in video indexing, surveillance, activity discovery and event recognition. Clustering a video sequence into activities requires one to simultaneously recognize activity boundaries (activity consistent subsequences) and cluster these activity subsequences. In order to do this, we build a generative model for activities (in video) using a cascade of dynamical systems and show that this model is able to capture and represent a diverse class of activities. We then derive algorithms to learn the model parameters from a video stream and also show how a single video sequence may be clustered into different clusters where each cluster represents an activity. We also propose a novel technique to build affine, view, rate invariance of the activity into the distance metric for clustering. Experiments show that the clusters found by the algorithm correspond to semantically meaningful activities. Pavan Turaga, Ashok Veeraraghavan, Rama Chellappa |
CVPR | 2 |
| 2007 | Fast Bilinear SfM with Side InformationabstractWe study the beneficial effect of side information on the Structure from Motion (SfM) estimation problem. The side information that we consider is measurement of a 'reference vector' and distance from fixed plane perpendicular to that reference vector. Firstly, we show that in the presence of this information, the SfM equations can be rewritten similar to a bilinear form in its unknowns. Secondly, we describe a fast iterative estimation procedure to recover the structure of both stationary scenes and moving objects that capitalizes on this information. We also provide a refinement procedure in order to tackle incomplete or noisy side information. We characterize the algorithm with respect to its reconstruction accuracy, memory requirements and stability. Finally, we describe two classes of commonly occurring real-world scenarios in which this algorithm will be effective: (a) presence of a dominant ground plane in the scene and (b) presence of an inertial measurement unit on board. Experiments using both real data and rigorous simulations show the efficacy of the algorithm. Mahesh Ramachandran, Ashok Veeraraghavan, Rama Chellappa |
ICCV | 2 |
| 2007 | Dappled photography: mask enhanced cameras for heterodyned light fields and coded aperture refocusingabstractWe describe a theoretical framework for reversibly modulating 4D light fields using an attenuating mask in the optical path of a lens based camera. Based on this framework, we present a novel design to reconstruct the 4D light field from a 2D camera image without any additional refractive elements as required by previous light field cameras. The patterned mask attenuates light rays inside the camera instead of bending them, and the attenuation recoverably encodes the rays on the 2D sensor. Our mask-equipped camera focuses just as a traditional camera to capture conventional 2D photos at full sensor resolution, but the raw pixel values also hold a modulated 4D light field. The light field can be recovered by rearranging the tiles of the 2D Fourier transform of sensor values into 4D planes, and computing the inverse Fourier transform. In addition, one can also recover the full resolution image information for the in-focus parts of the scene. We also show how a broadband mask placed at the lens enables us to compute refocused images at full sensor resolution for layered Lambertian scenes. This partial encoding of 4D ray-space data enables editing of image contents by depth, yet does not require computational recovery of the complete 4D light field. Ashok Veeraraghavan, Ramesh Raskar, Amit K. Agrawal, Ankit Mohan, Jack Tumblin |
ACM Trans. Graph. | 1 |
| 2006 | The Function Space of an ActivityabstractAn activity consists of an actor performing a series of actions in a pre-defined temporal order. An action is an individual atomic unit of an activity. Different instances of the same activity may consist of varying relative speeds at which the various actions are executed, in addition to other intra- and inter- person variabilities. Most existing algorithms for activity recognition are not very robust to intra- and inter-personal changes of the same activity, and are extremely sensitive to warping of the temporal axis due to variations in speed profile. In this paper, we provide a systematic approach to learn the nature of such time warps while simultaneously allowing for the variations in descriptors for actions. For each activity we learn an ‘average’ sequence that we denote as the nominal activity trajectory. We also learn a function space of time warpings for each activity separately. The model can be used to learn individualspecific warping patterns so that it may also be used for activity based person identification. The proposed model leads us to algorithms for learning a model for each activity, clustering activity sequences and activity recognition that are robust to temporal, intra- and inter-person variations. We provide experimental results using two datasets. Ashok Veeraraghavan, Amit K. Roy-Chowdhury |
CVPR (1) | 1 |
| 2006 | Motion Based Correspondence for 3D Tracking of Multiple Dim ObjectsabstractTracking multiple objects in a video is a demanding task that is frequently encountered in several systems such as surveillance and motion analysis. Ability to track objects in 3D requires the use of multiple cameras. While tracking multiple objects using multiples video cameras, establishing correspondence between objects in the various cameras is a non-trivial task. Specifically, when the targets are dim or are very far away from the camera, appearance cannot be used in order to establish this correspondence. Here, we propose a technique to establish correspondence across cameras using the motion features extracted from the targets, even when the relative position of the cameras is unknown. Experimental results are provided for the problem of tracking multiple bees in natural flight using two cameras. The reconstructed 3D flight paths of the bees show some interesting flight patterns. Ashok Veeraraghavan, Mandyam V. Srinivasan, Rama Chellappa, Emily Baird, Richard Lamont |
ICASSP (2) | 1 |
| 2005 | Matching Shape Sequences in Video with Applications in Human Movement AnalysisabstractWe present an approach for comparing two sequences of deforming shapes using both parametric models and nonparametric methods. In our approach, Kendall's definition of shape is used for feature extraction. Since the shape feature rests on a non-Euclidean manifold, we propose parametric models like the autoregressive model and autoregressive moving average model on the tangent space and demonstrate the ability of these models to capture the nature of shape deformations using experiments on gait-based human recognition. The nonparametric model is based on Dynamic Time-Warping. We suggest a modification of the Dynamic time-warping algorithm to include the nature of the non-Euclidean space in which the shape deformations take place. We also show the efficacy of this algorithm by its application to gait-based human recognition. We exploit the shape deformations of a person's silhouette as a discriminating feature and provide recognition results using the nonparametric model. Our analysis leads to some interesting observations on the role of shape and kinematics in automated gait-based person authentication. Ashok Veeraraghavan, Amit K. Roy-Chowdhury, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Role of Shape and Kinematics in Human Movement Analysis
Ashok Veeraraghavan, Amit K. Roy-Chowdhury, Rama Chellappa |
CVPR (1) | 1 |