Vivek Boominathan

dblp:120/7119 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0003-4875-3135ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Event Fields: Capturing Light Fields at High Speed, Resolution, and Dynamic Range
abstract
Event cameras, which feature pixels that independently respond to changes in brightness, are becoming increasingly popular in high- speed applications due to their lower latency, reduced bandwidth requirements, and enhanced dynamic range compared to traditional frame- based cameras. Numerous imaging and vision techniques have leveraged event cameras for high- speed scene understanding by capturing high- framerate, high- dynamic range videos, primarily utilizing the temporal advantages inherent to event cameras. Additionally, imaging and vision techniques have utilized the light field—a complementary dimension to temporal information—for enhanced scene understanding.In this work, we propose "Event Fields", a new approach that utilizes innovative optical designs for event cameras to capture light fields at high speed. We develop the underlying mathematical framework for Event Fields and introduce two foundational frameworks to capture them practically: spatial multiplexing to capture temporal derivatives and temporal multiplexing to capture angular derivatives. To realize these, we design two complementary optical setups— one using a kaleidoscope for spatial multiplexing and another using a galvanometer for temporal multiplexing. We evaluate the performance of both designs using a custom-built simulator and real hardware prototypes, showcasing their distinct benefits. Our event fields unlock the full advantages of typical light fields—like post- capture refocusing and depth estimation—now supercharged for high- speed and high- dynamic range scenes. This novel light- sensing paradigm opens doors to new applications in photography, robotics, and AR/VR, and presents fresh challenges in rendering and machine learning.
Ziyuan Qu, Zihao Zou, Vivek Boominathan, Praneeth Chakravarthula, Adithya Kumar Pediredla
CVPR3
2025 Diffusion Model Based Image Reconstruction in Lensless Imaging
abstract
Lensless imaging systems eliminate the need for lenses by employing an encoding element to multiplex incident light signals, which are then captured directly onto a bare camera sensor. They present a promising alternative to traditional lens-based imaging systems by offering significant advantages in terms of compactness, versatility, and cost. Due to the multiplexed nature of measurements, image reconstruction takes place computationally. However, existing techniques for image reconstruction in lensless imaging fall short of the image quality offered by traditional lens-based imaging. In this work, we consider the application of diffusion models, a class of deep generative models, for image reconstruction in a lensless imaging modality. These models currently achieve state-of-the-art performance in image generation. Specifically, we focus on the PhlatCam lensless system, which consists of a coded phase mask as the encoding element placed close to the camera sensor. We use a ControlNet based diffusion model to improve the perceptual quality of image reconstruction. The performance is measured in terms of peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), and learned perceptual image patch similarity (LPIPS). The proposed method improves the performance in these metrics for synthetic measurements. For real measurements, the improvement in image quality comes at the expense of a small bias in color, which is attributed to the generative nature of the diffusion prior itself.
Vivek Boominathan, Ashok Veeraraghavan, Chandra Sekhar Seelamantula
ICASSP2
2025 CoIR: Compressive Implicit Radar
abstract
Using millimeter wave (mmWave) signals for imaging has an important advantage in that they can penetrate through poor environmental conditions such as fog, dust, and smoke that severely degrade optical-based imaging systems. However, mmWave radars, contrary to cameras and LiDARs, suffer from low angular resolution because of small physical apertures and conventional signal processing techniques. Sparse radar imaging, on the other hand, can increase the aperture size while minimizing power consumption and read-out bandwidth. This article presents CoIR, an analysis by synthesis method that leverages the implicit neural network bias in convolutional decoders and compressed sensing to perform high-accuracy sparse radar imaging. The proposed system is data set-agnostic and does not require any auxiliary sensors for training or testing. We introduce a sparse array design that allows for a $5.5\times$5.5× reduction in the number of antenna elements needed compared to conventional MIMO array designs. We demonstrate our system's improved imaging performance over standard mmWave radars and other competitive untrained methods on both simulated and experimental mmWave radar data.
Sean M. Farrell, Vivek Boominathan, Nate Raymondi, Ashutosh Sabharwal, Ashok Veeraraghavan
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Streaming Quanta Sensors for Online, High-Performance Imaging and Vision
abstract
Recently quanta image sensors (QIS) - ultra-fast, zero-read-noise binary image sensors- have demonstrated remarkable imaging capabilities in many challenging scenarios. Despite their potential, the adoption of these sensors is severely hampered by (a) high data rates and (b) the need for new computational pipelines to handle the unconventional raw data. We introduce a simple, low-bandwidth computational pipeline to address these challenges. Our approach is based on a novel streaming representation with a small memory footprint, efficiently capturing intensity information at multiple temporal scales. Updating the representation requires only 16 floating-point operations/pixel, which can be efficiently computed online at the native frame rate of the binary frames. We use a neural network operating on this representation to reconstruct videos in real-time (10-30 fps). We illustrate why such representation is well-suited for these emerging sensors, and how it offers low latency and high frame rate while retaining flexibility for downstream computer vision. Our approach results in significant data bandwidth reductions () and real-time image reconstruction and computer vision -)reduction in computation than existing state-of-the-art approach[1], while maintaining comparable quality. To the best of our knowledge, our approach is the first to achieve online, real-time image reconstruction on QIS.
Matthew Dutson, Vivek Boominathan, Mohit Gupta 0001, Ashok Veeraraghavan
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Passive Snapshot Coded Aperture Dual-Pixel RGB-D Imaging
abstract
Passive, compact, single-shot 3D sensing is useful in many application areas such as microscopy, medical imaging, surgical navigation, and autonomous driving where form factor, time, and power constraints can exist. Ob-taining RGB-D scene information over a short imaging distance, in an ultra-compact form factor, and in a passive, snapshot manner is challenging. Dual-pixel (DP) sensors are a potential solution to achieve the same. DP sensors collect light rays from two different halves of the lens in two interleaved pixel arrays, thus capturing two slightly different views of the scene, like a stereo camera system. However, imaging with a DP sensor implies that the defocus blur size is directly proportional to the disparity seen between the views. This creates a tradeoff between disparity estimation vs. deblurring accuracy. To improve this tradeoff effect, we propose CADS (Coded Aperture Dual-Pixel Sensing), in which we use a coded aperture in the imaging lens along with a DP sensor. In our approach, we jointly learn an optimal coded pattern and the reconstruction algorithm in an end-to-end optimization setting. Our resulting CADS imaging system demonstrates improvement of> 1.5 dB PSNR in all-in-focus (AIF) estimates and 5-6% in depth estimation quality over naive DP sensing for a wide range of aperture settings. Furthermore, we build the proposed CADS prototypes for DSLR photography settings and in an endoscope and a dermoscope form factor. Our novel coded dual-pixel sensing approach demonstrates accurate RGB-D reconstruction results in simulations and real-world experiments in a passive, snapshot, and compact manner.
Bhargav Ghanekar, Salman Siddique Khan, Pranav Sharma, Shreyas Singh, Vivek Boominathan, Kaushik Mitra, Ashok Veeraraghavan
CVPR5
2023 Designing Optics and Algorithm for Ultra-Thin, High-Speed Lensless Cameras
abstract
There is a growing demand for small, light-weight and low-latency cameras in the robotics and AR/VR community. Mask-based lensless cameras, by design, provide a combined advantage of form-factor, weight and speed. They do so by replacing the classical lens with a thin optical mask and computation. Recent works have explored deep learning based post-processing operations on lensless captures that allow high quality scene reconstruction. However, the ability of deep learning to find the optimal optics for thin lensless cameras has not been explored. In this work, we propose a learning based framework for designing the optics of thin lensless cameras. To highlight the effectiveness of our framework, we learn the optical phase mask for multiple tasks using physics-based neural networks. Specifically, we learn the optimal mask using a weighted loss defined for the following tasks-2D scene reconstructions, optical flow estimation and face detection. We show that mask learned through this framework is better than heuristically designed masks especially for small sensors sizes that allow lower bandwidth and faster readout. Finally, we verify the performance of our learned phase-mask on real data.
Salman Siddique Khan, Vivek Boominathan, Ashok Veeraraghavan, Kaushik Mitra
ICME2
2022 EyeCoD: eye tracking system acceleration via flatcam-based algorithm & accelerator co-design
abstract
Eye tracking has become an essential human-machine interaction modality for providing immersive experience in numerous virtual and augmented reality (VR/AR) applications desiring high throughput (e.g., 240 FPS), small-form, and enhanced visual privacy. However, existing eye tracking systems are still limited by their: (1) large form-factor largely due to the adopted bulky lens-based cameras; (2) high communication cost required between the camera and backend processor; and (3) potentially concerned low visual privacy, thus prohibiting their more extensive applications. To this end, we propose, develop, and validate a lensless FlatCambased eye tracking algorithm and accelerator co-design framework dubbed EyeCoD to enable eye tracking systems with a much reduced form-factor and boosted system efficiency without sacrificing the tracking accuracy, paving the way for next-generation eye tracking solutions. On the system level, we advocate the use of lensless FlatCams instead of lens-based cameras to facilitate the small form-factor need in mobile eye tracking systems, which also leaves rooms for a dedicated sensing-processor co-design to reduce the required camera-processor communication latency. On the algorithm level, EyeCoD integrates a predict-then-focus pipeline that first predicts the region-of-interest (ROI) via segmentation and then only focuses on the ROI parts to estimate gaze directions, greatly reducing redundant computations and data movements. On the hardware level, we further develop a dedicated accelerator that (1) integrates a novel workload orchestration between the aforementioned segmentation and gaze estimation models, (2) leverages intra-channel reuse opportunities for depth-wise layers, (3) utilizes input feature-wise partition to save activation memory size, and (4) develops a sequential-write-parallel-read input buffer to alleviate the bandwidth requirement for the activation global buffer. On-silicon measurement and extensive experiments validate that our EyeCoD consistently reduces both the communication and computation costs, leading to an overall system speedup of 10.95×, 3.21×, and 12.85× over general computing platforms including CPUs and GPUs, and a prior-art eye tracking processor called CIS-GEP, respectively, while maintaining the tracking accuracy. Codes are available at https://github.com/RICE-EIC/EyeCoD.
Haoran You, Cheng Wan 0005, Yang Zhao 0013, Zhongzhi Yu, Yonggan Fu, Jiayi Yuan 0001, Shang Wu 0003, Yongan Zhang, Chaojian Li, Vivek Boominathan, Ashok Veeraraghavan, Ziyun Li 0001, Yingyan (Celine) Lin
ISCA11
2022 FlatNet: Towards Photorealistic Scene Reconstruction From Lensless Measurements
abstract
Lensless imaging has emerged as a potential solution towards realizing ultra-miniature cameras by eschewing the bulky lens in a traditional camera. Without a focusing lens, the lensless cameras rely on computational algorithms to recover the scenes from multiplexed measurements. However, the current iterative-optimization-based reconstruction algorithms produce noisier and perceptually poorer images. In this work, we propose a non-iterative deep learning-based reconstruction approach that results in orders of magnitude improvement in image quality for lensless reconstructions. Our approach, called FlatNet, lays down a framework for reconstructing high-quality photorealistic images from mask-based lensless cameras, where the camera's forward model formulation is known. FlatNet consists of two stages: (1) an inversion stage that maps the measurement into a space of intermediate reconstruction by learning parameters within the forward model formulation, and (2) a perceptual enhancement stage that improves the perceptual quality of this intermediate reconstruction. These stages are trained together in an end-to-end manner. We show high-quality reconstructions by performing extensive experiments on real and challenging scenes using two different types of lensless prototypes: one which uses a separable forward model and another, which uses a more general non-separable cropped-convolution model. Our end-to-end approach is fast, produces photorealistic reconstructions, and is easy to adopt for other mask-based lensless cameras.
Salman Siddique Khan, Varun Sundar, Vivek Boominathan, Ashok Veeraraghavan, Kaushik Mitra
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 SACoD: Sensor Algorithm Co-Design Towards Efficient CNN-powered Intelligent PhlatCam
abstract
There has been a booming demand for integrating Convolutional Neural Networks (CNNs) powered functionalities into Internet-of-Thing (IoT) devices to enable ubiquitous intelligent "IoT cameras". However, more extensive applications of such IoT systems are still limited by two challenges. First, some applications, especially medicine-and wearable-related ones, impose stringent requirements on the camera form factor. Second, powerful CNNs often require considerable storage and energy cost, whereas IoT devices often suffer from limited resources. PhlatCam, with its form factor potentially reduced by orders of magnitude, has emerged as a promising solution to the first aforementioned challenge, while the second one remains a bottleneck. Existing compression techniques, which can potentially tackle the second challenge, are far from realizing the full potential in storage and energy reduction, because they mostly focus on the CNN algorithm itself. To this end, this work proposes SACoD, a Sensor Algorithm Co-Design framework to develop more efficient CNN-powered PhlatCam. In particular, the mask coded in the Phlat-Cam sensor and the backend CNN model are jointly optimized in terms of both model parameters and architectures via differential neural architecture search. Extensive experiments including both simulation and physical measurement on manufactured masks show that the proposed SACoD framework achieves aggressive model compression and energy savings while maintaining or even boosting the task accuracy, when benchmarking over two state-of-the-art (SOTA) designs with six datasets across four different vision tasks including classification, segmentation, image translation, and face recognition. Our codes are available at: https://github.com/RICE-EIC/SACoD.
Yonggan Fu, Yang Zhang 0001, Yue Wang 0036, Zhihan Lyu, Vivek Boominathan, Ashok Veeraraghavan, Yingyan (Celine) Lin
ICCV5
2020 FreeCam3D: Snapshot Structured Light 3D with Freely-Moving Cameras
Vivek Boominathan, Jacob T. Robinson, Hiroshi Kawasaki, Aswin C. Sankaranarayanan, Ashok Veeraraghavan
ECCV (27)2
2020 CANOPIC: Pre-Digital Privacy-Enhancing Encodings for Computer Vision
abstract
The standard pipeline for many vision tasks uses a conventional camera to capture an image that is then passed to a digital processor for information extraction. In some deployments, such as private locations, the captured digital imagery contains sensitive information exposed to digital vulnerabilities such as spyware, Trojans, etc. However, in many applications, the full imagery is unnecessary for the vision task at hand. In this paper we propose an optical and analog system that preprocesses the light from the scene before it reaches the digital imager to destroy sensitive information. We explore analog and optical encodings consisting of easily implementable operations such as convolution, pooling, and quantization. We perform a case study to evaluate how such encodings can destroy face identity information while preserving enough information for face detection. The encoding parameters are learned via an alternating optimization scheme based on adversarial learning with deep neural networks. We name our system CAnOPIC (Camera with Analog and Optical Privacy-Integrating Computations) and show that it has better performance in terms of both privacy and utility than conventional optical privacy-enhancing methods such as blurring and pixelation.
Jasper Tan, Salman Siddique Khan, Vivek Boominathan, Jeffrey Byrne, Richard G. Baraniuk, Kaushik Mitra, Ashok Veeraraghavan
ICME3
2020 PhlatCam: Designed Phase-Mask Based Thin Lensless Camera
abstract
We demonstrate a versatile thin lensless camera with a designed phase-mask placed at sub-2 mm from an imaging CMOS sensor. Using wave optics and phase retrieval methods, we present a general-purpose framework to create phase-masks that achieve desired sharp point-spread-functions (PSFs) for desired camera thicknesses. From a single 2D encoded measurement, we show the reconstruction of high-resolution 2D images, computational refocusing, and 3D imaging. This ability is made possible by our proposed high-performance contour-based PSF. The heuristic contour-based PSF is designed using concepts in signal processing to achieve maximal information transfer to a bit-depth limited sensor. Due to the efficient coding, we can use fast linear methods for high-quality image reconstructions and switch to iterative nonlinear methods for higher fidelity reconstructions and 3D imaging.
Vivek Boominathan, Jesse K. Adams, Jacob T. Robinson, Ashok Veeraraghavan
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 PhaseCam3D - Learning Phase Masks for Passive Single View Depth Estimation
abstract
There is an increasing need for passive 3D scanning in many applications that have stringent energy constraints. In this paper, we present an approach for single frame, single viewpoint, passive 3D imaging using a phase mask at the aperture plane of a camera. Our approach relies on an end-to-end optimization framework to jointly learn the optimal phase mask and the reconstruction algorithm that allows an accurate estimation of range image from captured data. Using our optimization framework, we design a new phase mask that performs significantly better than existing approaches. We build a prototype by inserting a phase mask fabricated using photolithography into the aperture plane of a conventional camera and show compelling performance in 3D imaging.
Vivek Boominathan, Huaijin G. Chen, Aswin C. Sankaranarayanan, Ashok Veeraraghavan
ICCP2
2019 Towards Photorealistic Reconstruction of Highly Multiplexed Lensless Images
abstract
Recent advancements in fields like Internet of Things (IoT), augmented reality, etc. have led to an unprecedented demand for miniature cameras with low cost that can be integrated anywhere and can be used for distributed monitoring. Mask-based lensless imaging systems make such inexpensive and compact models realizable. However, reduction in the size and cost of these imagers comes at the expense of their image quality due to the high degree of multiplexing inherent in their design. In this paper, we present a method to obtain image reconstructions from mask-based lensless measurements that are more photorealistic than those currently available in the literature. We particularly focus on FlatCam, a lensless imager consisting of a coded mask placed over a bare CMOS sensor. Existing techniques for reconstructing FlatCam measurements suffer from several drawbacks including lower resolution and dynamic range than lens-based cameras. Our approach overcomes these drawbacks using a fully trainable non-iterative deep learning based model. Our approach is based on two stages: an inversion stage that maps the measurement into the space of intermediate reconstruction and a perceptual enhancement stage that improves this intermediate reconstruction based on perceptual and signal distortion metrics. Our proposed method is fast and produces photo-realistic reconstruction as demonstrated on many real and challenging scenes.
Salman Siddique Khan, Adarsh V. R, Vivek Boominathan, Jasper Tan, Ashok Veeraraghavan, Kaushik Mitra
ICCV3
2017 Flat focus: depth of field analysis for the FlatCam lensless imaging system
abstract
Lensless imaging systems, such as the recently proposed FlatCam, offer numerous advantages over lens-based systems such as a thin form-factor, low cost, and higher light throughput. However, little work has been done in analyzing these systems' depth of field characteristics. A depth-dependent calibration step is necessary to obtain the image from the FlatCam measurements, and this calibration determines the system's depth of field. In this paper, we characterize the FlatCam's depth of field properties and show that (a) for scene depths on the order of tens of centimeters, it is possible to perform depth-selective refocusing from a single captured image and (b) for sufficiently large scene depths, calibratin.g for one depth can provide a very large depth of field.
Jasper Tan, Vivek Boominathan, Ashok Veeraraghavan, Richard G. Baraniuk
ICASSP2
2014 Improving resolution and depth-of-field of light field cameras using a hybrid imaging system
abstract
Current light field (LF) cameras provide low spatial resolution and limited depth-of-field (DOF) control when compared to traditional digital SLR (DSLR) cameras. We show that a hybrid imaging system consisting of a standard LF camera and a high-resolution standard camera enables (a) achieve high-resolution digital refocusing, (b) better DOF control than LF cameras, and (c) render graceful high-resolution viewpoint variations, all of which were previously unachievable. We propose a simple patch-based algorithm to super-resolve the low-resolution views of the light field using the high-resolution patches captured using a high-resolution SLR camera. The algorithm does not require the LF camera and the DSLR to be co-located or for any calibration information regarding the two imaging systems. We build an example prototype using a Lytro camera (380×380 pixel spatial resolution) and a 18 megapixel (MP) Canon DSLR camera to generate a light field with 11 MP resolution (9× super-resolution) and about 1 over 9thof the DOF of the Lytro camera. We show several experimental results on challenging scenes containing occlusions, specularities and complex non-lambertian materials, demonstrating the effectiveness of our approach.
Vivek Boominathan, Kaushik Mitra, Ashok Veeraraghavan
ICCP1
2012 Speaker recognition via sparse representations using orthogonal matching pursuit
abstract
The objective of this paper is to demonstrate the effectiveness of sparse representation techniques for speaker recognition. In this approach, each feature vector from unknown utterance is expressed as linear weighted sum of a dictionary of feature vectors belonging to many speakers. The weights associated with feature vectors in the dictionary are evaluated using orthogonal matching pursuit algorithm, which is a greedy approximation to l0 optimization. The weights thus obtained exhibit high level of sparsity, and only a few of them will have nonzero values. The feature vectors which belong to the correct speaker carry significant weights. The proposed method gives an equal error rate (EER) of 10.84% on NIST-2003 database, whereas the existing GMM-UBM system gives an EER of 9.67%. By combining evidence from both the systems an EER of 8.15% is achieved, indicating that both the systems carry complimentary information.
Vivek Boominathan, K. Sri Rama Murty
ICASSP1