EDBT 2026 Demo / reviewers in the wild / expert
Kaushik Mitra
dblp:26/3767
· DBLP profile ↗
54ranked-venue papers
5as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 3 first-author · 26 since 2021Artificial intelligence and machine learning · 21 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-Guided 3D Gaussian Splatting for Sparse View Reconstruction and Segmentation
S. Meena Padnekar, Kaushik Mitra, Sukhendu Das |
ICPR (13) | 2 |
| 2026 | Dark Stereo Multi-Exposure Dataset (DSMD) for extreme low-light image restoration and depth estimation
Rongali Simhachala Venkata Girish, M. V. A. Suhas Kumar, Mohit Lamba, Devaganthan S. S., Adithya Lenka, Kaushik Mitra |
Comput. Vis. Image Underst. | 6 |
| 2026 | Distilling auxiliary RGB-T features for unsupervised semantic segmentation
S. Meena Padnekar, Kaushik Mitra, Sukhendu Das |
Image Vis. Comput. | 2 |
| 2025 | PhotonSplat: 3D Scene Reconstruction and Colorization from SPAD SensorsabstractAdvances in 3D reconstruction using neural rendering have enabled high-quality 3D capture. However, they often fail when the input imagery is corrupted by motion blur, due to fast motion of the camera or the objects in the scene. This work advances neural rendering techniques in such scenarios by using single-photon avalanche diode (SPAD) arrays, an emerging sensing technology capable of sensing images at extremely high speeds. However, the use of SPADs presents its own set of unique challenges in the form of binary images, that are driven by stochastic photon arrivals. To address this, we introduce PhotonSplat, a framework designed to reconstruct 3D scenes directly from SPAD binary images, effectively navigating the noise vs. blur trade-off. Our approach incorporates a novel 3D spatial filtering technique to reduce noise in the renderings. The framework also supports both no-reference using generative priors and reference-based colorization from a single blurry image, enabling downstream applications such as segmentation, object detection and appearance editing tasks. Additionally, we extend our method to incorporate dynamic scene representations, making it suitable for scenes with moving objects. We further contribute PhotonScenes, a real-world multi-view dataset captured with the SPAD sensors. Code, data and video results are available at vinayak-vg.github.io/PhotonSplat/. Kuppa Sai Sri Teja, Sreevidya Chintalapati, Mukund Varma T., Haejoon Lee, Aswin C. Sankaranarayanan, Kaushik Mitra |
ICCP | 7 |
| 2025 | RT-X Net: RGB-Thermal Cross Attention Network for Low-Light Image EnhancementabstractIn nighttime conditions, high noise levels and bright illumination sources degrade image quality, making low-light image enhancement challenging. Thermal images provide complementary information, offering richer textures and structural details. We propose RT-X Net, a cross-attention network that fuses RGB and thermal images for nighttime image enhancement. We leverage self-attention networks for feature extraction and a cross-attention mechanism for fusion to effectively integrate information from both modalities. To support research in this domain, we introduce the Visible-Thermal Image Enhancement Evaluation (V-TIEE) dataset, comprising 50 co-located visible and thermal images captured under diverse nighttime conditions. Extensive evaluations on the publicly available LLVIP dataset and our V-TIEE dataset demonstrate that RT-X Net outperforms state-of-the-art methods in low-light image enhancement. The code and the V-TIEE can be found here https://github.com/jhakrraman/rt-xnet. Raman Jha, Adithya Lenka, Manikandasriram Srinivasan Ramanagopal, Aswin C. Sankaranarayanan, Kaushik Mitra |
ICIP | 5 |
| 2025 | SPC TO 3d: Novel View Synthesis from Binary SPC VIA I2I TranslationabstractSingle Photon Avalanche Diodes (SPADs) represent a cutting-edge imaging technology, capable of detecting individual photons with remarkable timing precision. Building on this sensitivity, Single Photon Cameras (SPCs) enable image capture at exceptionally high speeds under both low and high illumination. Enabling 3D reconstruction and radiance field recovery from such SPC data holds significant promise. However, the binary nature of SPC images leads to severe information loss, particularly in texture and color, making traditional 3D synthesis techniques ineffective. To address this challenge, we propose a modular two-stage framework that converts binary SPC images into high-quality colorized novel views. The first stage performs image-to-image (I2I) translation using generative models such as Pix2PixHD, converting binary SPC inputs into plausible RGB representations. The second stage employs 3D scene reconstruction techniques like Neural Radiance Fields (NeRF) or Gaussian Splatting (3DGS) to generate novel views. We validate our two-stage pipeline (Pix2PixHD + Nerf/3DGS) through extensive qualitative and quantitative experiments, demonstrating significant improvements in perceptual quality and geometric consistency over the alternative baseline. Sumit Sharma 0014, Gopi Matta, Kaushik Mitra |
ICIP | 3 |
| 2025 | GANESH: Generalizable NeRF for Lensless ImagingabstractLensless imaging offers a significant opportunity to develop ultra-compact cameras by removing the conventional bulky lens system. However, without a focusing element, the sensor's output is no longer a direct image but a complex multiplexed scene representation. Traditional methods have attempted to address this challenge by employing learnable inversions and refinement models, but these methods are primarily designed for 2D reconstruction and do not generalize well to 3D reconstruction. We introduce GANESH, a novel framework designed to enable simultaneous refinement and novel view synthesis from multi-view lensless images. Unlike existing methods that require scene-specific training, our approach supports on-the-fiy inference without retraining on each scene. Moreover, our framework allows us to tune our model to specific scenes, enhancing the rendering and refinement quality. To facilitate research in this area, we also present the first multi-view lensless dataset, LenslessScenes. Extensive experiments demonstrate that our method outperforms current approaches in reconstruction accuracy and refinement quality. Code and video results are available here. Rakesh Raj Madavan, Akshat Kaimal, Badhrinarayanan K. V, Rohit Choudhary, Chandrakala Shanmuganathan, Kaushik Mitra |
WACV | 7 |
| 2024 | GN-FR: Generalizable Neural Radinace Fields for Flare Removal
Gopi Matta, Rahul Siddartha, Rongali Simhachala Venkata Girish, Sumit Sharma 0014, Kaushik Mitra |
BMVC | 5 |
| 2024 | HDRSplat: Gaussian Splatting for High Dynmaic Range 3D Scene Reconstruction from Raw Images
Shreyas Singh, Aryan Garg, Kaushik Mitra |
BMVC | 3 |
| 2024 | Passive Snapshot Coded Aperture Dual-Pixel RGB-D ImagingabstractPassive, compact, single-shot 3D sensing is useful in many application areas such as microscopy, medical imaging, surgical navigation, and autonomous driving where form factor, time, and power constraints can exist. Ob-taining RGB-D scene information over a short imaging distance, in an ultra-compact form factor, and in a passive, snapshot manner is challenging. Dual-pixel (DP) sensors are a potential solution to achieve the same. DP sensors collect light rays from two different halves of the lens in two interleaved pixel arrays, thus capturing two slightly different views of the scene, like a stereo camera system. However, imaging with a DP sensor implies that the defocus blur size is directly proportional to the disparity seen between the views. This creates a tradeoff between disparity estimation vs. deblurring accuracy. To improve this tradeoff effect, we propose CADS (Coded Aperture Dual-Pixel Sensing), in which we use a coded aperture in the imaging lens along with a DP sensor. In our approach, we jointly learn an optimal coded pattern and the reconstruction algorithm in an end-to-end optimization setting. Our resulting CADS imaging system demonstrates improvement of> 1.5 dB PSNR in all-in-focus (AIF) estimates and 5-6% in depth estimation quality over naive DP sensing for a wide range of aperture settings. Furthermore, we build the proposed CADS prototypes for DSLR photography settings and in an endoscope and a dermoscope form factor. Our novel coded dual-pixel sensing approach demonstrates accurate RGB-D reconstruction results in simulations and real-world experiments in a passive, snapshot, and compact manner. Bhargav Ghanekar, Salman Siddique Khan, Pranav Sharma, Shreyas Singh, Vivek Boominathan, Kaushik Mitra, Ashok Veeraraghavan |
CVPR | 6 |
| 2024 | GAURA: Generalizable Approach for Unified Restoration and Rendering of Arbitrary Views
Rongali Simhachala Venkata Girish, Mukund Varma T., Ayush Tewari, Kaushik Mitra |
ECCV (14) | 5 |
| 2024 | Near-Field Neural Rendering Guided by Single-Shot Photometric StereoabstractWe present a novel near-field neural rendering approach that combines single-shot RGB photometric stereo and SDF-based Neural Rendering. Photometric stereo-based cues guide the neural rendering of 3D meshes. Recent studies have shown that SDF-based neural implicit surface reconstruction methods can produce smoother, more comprehensive 3D reconstructions. However, their performance tends to decline when capturing fine details in complex near-field images due to the ambiguity in RGB reconstruction loss. For instance, 3D endoscopy, characterized by near-field sparse views and complex surfaces, poses considerable challenges for existing volumetric SDF methods. Motivated by advancements in near-field photometric stereo, our work explores the utilization of these cues to enhance neural implicit surface reconstruction from diverse perspectives. To simplify the acquisition of photometric stereo images in endoscopy setups, we employ a strategy that involves single-shot RGB photometric stereo capture for each view. Utilizing a learning-based near-field photometric stereo network, we extract depth and normals for each view. These cues contribute to improved performance when dealing with very near-field natural objects as well as endoscopic scenes. We have extensively tested this approach on various rendered and real near-field datasets, and it consistently outperforms existing photometric NeRF techniques, especially when dealing with near-field multi-view visuals. Joshna Manoj Reddy, Tony Fredrick, Salman Siddique Khan, Kaushik Mitra |
ICASSP | 4 |
| 2024 | Stereo-Knowledge Distillation from dpMV to Dual Pixels for Light Field Video ReconstructionabstractDual pixels contain disparity cues arising from the defocus blur. This disparity information is useful for many vision tasks ranging from autonomous driving to 3D creative realism. However, directly estimating disparity from dual pixels is less accurate. This work hypothesizes that distilling high-precision dark stereo knowledge, implicitly or explicitly, to efficient dual-pixel student networks enables faithful reconstructions. This dark knowledge distillation should also alleviate stereo-synchronization setup and calibration costs while dramatically increasing parameter and inference time efficiency. We collect the first and largest 3-view dual-pixel video dataset, dpMV, to validate our explicit dark knowledge distillation hypothesis. We show that these methods outperform purely monocular solutions, especially in challenging foreground-background separation regions using faithful guidance from dual pixels. Finally, we demonstrate an unconventional use case unlocked by dpMV and implicit dark knowledge distillation from an ensemble of teachers for Light Field (LF) video reconstruction. Our LF video reconstruction method is the fastest and most temporally consistent to date. It remains competitive in reconstruction fidelity while offering many other essential properties like high parameter efficiency, implicit disocclusion handling, zero-shot cross-dataset transfer, geometrically consistent inference on higher spatial-angular resolutions, and adaptive baseline control. All source code is available at the repository https://github.com/Aryan-Garg. Aryan Garg, Raghav Mallampali, Akshat Joshi, Shrisudhan Govindarajan, Kaushik Mitra |
ICCP | 5 |
| 2024 | R2SFD: Improving Single Image Reflection Removal using Semantic Feature DictionaryabstractSingle image reflection removal is a severely ill-posed problem and it is very hard to separate the desirable transmission and undesirable reflection layers. Most of the existing single image reflection removal methods try to recover the transmission layer by exploiting cues that are extracted only from the given input image. However, there is abundant unutilized information in the form of millions of reflection free images available publicly. Even though this information is easily available, utilizing the same for effectively removing reflections is non-trivial. In this paper, we propose a novel method, termed R^2SFD, for improving single image reflection removal using a Semantic Feature Dictionary (SFD) constructed from a database of reflection-free images. The SFD is constructed using a novel Reflection Aware Feature Extractor (RAFENet) that extracts features invariant to the presence of reflections. The SFD and the input image are then passed to another novel network termed SFDNet. This network first extracts RAFENet features from the reflection-corrupted input image, searches for similar features in the SFD, and transfers the semantic content to generate the final output. To further improve reflection removal, we also introduce a Large Scale Reflection Removal (LSRR) dataset consisting of 2650 image pairs comprising of a variety of real world reflection scenarios. The proposed method achieves superior results both qualitatively and quantitatively compared to the state of the art single image reflection removal methods on real public datasets as well as our LSRR dataset. We will release the dataset at https://github.com/ee19d005/r2sfd. Green Rosh K. S, B. H. Pawan Prasad, Lokesh R. Boregowda, Kaushik Mitra |
ACM Multimedia | 4 |
| 2023 | Deep Unsupervised Reflection Removal Using Diffusion ModelsabstractReflections caused due to surfaces such as glass affect the aesthetics of an image, and are hence undesirable. Most of the recent works on reflection removal use supervised learning based approaches using deep neural networks. However, most of these methods require large amount of paired data for training, which is difficult to obtain. Moreover, it is difficult to deploy existing deep learning based algorithms on multiple devices with different computational power, since it is very hard to control the trade-off between the strength of reflection removal and computational complexity during inference. To address these challenges, we propose a novel deep learning based approach for reflection removal, that is both unsupervised and controllable. We use Denoising Diffusion Probability Models to learn a distribution of reflection-free images. The learnt model is then used to generate reflection-free images using an input conditioned forward diffusion process during inference. We also perform qualitative and quantitative comparison and show that our method is at par or better than existing methods for deep supervised reflection removal, while outperforming unsupervised method by ~ 6.5 dB. Green Rosh K. S, B. H. Pawan Prasad, Lokesh R. Boregowda, Kaushik Mitra |
ICIP | 4 |
| 2023 | Designing Optics and Algorithm for Ultra-Thin, High-Speed Lensless CamerasabstractThere is a growing demand for small, light-weight and low-latency cameras in the robotics and AR/VR community. Mask-based lensless cameras, by design, provide a combined advantage of form-factor, weight and speed. They do so by replacing the classical lens with a thin optical mask and computation. Recent works have explored deep learning based post-processing operations on lensless captures that allow high quality scene reconstruction. However, the ability of deep learning to find the optimal optics for thin lensless cameras has not been explored. In this work, we propose a learning based framework for designing the optics of thin lensless cameras. To highlight the effectiveness of our framework, we learn the optical phase mask for multiple tasks using physics-based neural networks. Specifically, we learn the optimal mask using a weighted loss defined for the following tasks-2D scene reconstructions, optical flow estimation and face detection. We show that mask learned through this framework is better than heuristically designed masks especially for small sensors sizes that allow lower bandwidth and faster readout. Finally, we verify the performance of our learned phase-mask on real data. Salman Siddique Khan, Vivek Boominathan, Ashok Veeraraghavan, Kaushik Mitra |
ICME | 4 |
| 2023 | Real-Time Restoration of Dark Stereo ImagesabstractLow-light image enhancement has been an actively researched area for decades and has produced excellent night-time single-image, video, and Light Field restoration methods. Despite these advances, the problem of extreme low-light stereo image restoration has been mostly ignored and addressing it can enable night-time capabilities to several applications such as smartphones and self-driving cars. We propose an especially light-weight and fast hybrid U-net architecture for extreme low-light stereo image restoration. In the initial few scale spaces, we process the left and right features individually, because the two features do not align well due to large disparity. At coarser scale-spaces, the disparity between left and right features decreases and the network’s receptive field increases. We use this fact to reduce computations by simultaneously processing the left and right features, which also benefits epipole preservation. As our architecture does not use any 3D convolution for fast inference, we use a Depth-Aware loss module to train our network. This module computes quick and coarse depth estimates to better enforce the stereo epipolar constraints. Extensive benchmarking in terms of visual enhancement and downstream depth estimation shows that our architecture not only restores dark stereo images faithfully but also offers 4−60× speed-up with 15−100× lower floating point operations, necessary for real-world applications. Mohit Lamba, M. V. A. Suhas Kumar, Kaushik Mitra |
WACV | 3 |
| 2023 | Burst Reflection Removal using Reflection Motion Aggregation CuesabstractSingle image reflection removal has attracted lot of interest in the recent past with data driven approaches demonstrating significant improvements. However deep learning based approaches for multi-image reflection removal remains relatively less explored. The existing multi-image methods require input images to be captured at sufficiently different view points with wide baselines. This makes it cumbersome for the user who is required to capture the scene by moving the camera in multiple directions. A more convenient way is to capture a burst of images in a short time duration without providing any specific instructions to the user. A burst of images captured on a hand-held device provide crucial cues that rely on the subtle handshakes created during the capture process to separate the reflection and the transmission layers. In this paper, we propose a multi-stage deep learning based approach for burst reflection removal. In the first stage, we perform reflection suppression on the individual images. In the second stage, a novel reflection motion aggregation (RMA) cue is extracted that emphasizes the transmission layer more than the reflection layer to aid better layer separation. In our final stage we use this RMA cue as a guide to remove reflections from the input. We provide the first real world burst images dataset along with ground truth for reflection removal that can enable future benchmarking. We evaluate both qualitatively and quantitatively to demonstrate the superiority of the proposed approach. Our method achieves ~ 2dB improvement in PSNR over single image based methods and ~ 1dB over multi-image based methods. B. H. Pawan Prasad, Green Rosh K. S, R. B. Lokesh, Kaushik Mitra |
WACV | 4 |
| 2022 | Synthesizing Light Field Video from Monocular Video
Shrisudhan Govindarajan, Prasan A. Shedligeri, Sarah, Kaushik Mitra |
ECCV (7) | 4 |
| 2022 | LWGNet - Learned Wirtinger Gradients for Fourier Ptychographic Phase Retrieval
Atreyee Saha, Salman Siddique Khan, Sagar Sehrawat, Sanjana S. Prabhu, Shanti Bhattacharya, Kaushik Mitra |
ECCV (7) | 6 |
| 2022 | Reference Guided Reflection Removal Using Deep Visual Attribute CuesabstractReflections in images are typically caused due to presence of glass like reflective objects or surfaces that affect the overall visual appeal and hence undesirable. There has been extensive interest in the past to use data driven approaches for both single image as well as multi image reflection removal. However, recently there has been only minor incremental improvements in single image reflection removal given the challenging ill-posed nature of the problem. Reference based methods has yielded state of the art performance in areas such as super resolution, however has been unexplored for reflection removal. In this paper, we propose a novel multi-stage deep learning based method for reference based reflection removal. We also propose a novel visual attribute cue that represents the reflection free semantic content of the input image. This cue is generated using the reference image while maintaining the geometric structure of the input image. We formulate the reference based reflection removal problem as extraction of visual attribute cues followed by a guided image restoration. We perform qualitative and quantitative evaluation to demonstrate the superiority of the proposed approach over the existing state of the art single image reflection removal methods. B. H. Pawan Prasad, Green Rosh K. S, R. B. Lokesh, Kaushik Mitra |
ICIP | 4 |
| 2022 | Fast and Efficient Restoration of Extremely Dark Light FieldsabstractThe ability of Light Field (LF) cameras to capture the 3D geometry of a scene in a single photographic exposure has become central to several applications ranging from passive depth estimation to post-capture refocusing and view synthesis. But these LF applications break down in extreme low-light conditions due to excessive noise and poor image photometry. Existing low-light restoration techniques are inappropriate because they either do not leverage LF’s multi-view perspective or have enormous time and memory complexity. We propose a three-stage network that is simultaneously fast and accurate for real world applications. Our accuracy comes from the fact that our three stage architecture utilizes global, local and view-specific information present in low-light LFs and fuse them using an RNN inspired feedforward network. We are fast because we restore multiple views simultaneously and so require less number of forward passes. Besides these advantages, our network is flexible enough to restore a m × m LF during inference even if trained for a smaller n × n (n < m) LF without any finetuning. Extensive experiments on real low-light LF demonstrate that compared to the current state-of-the-art, our model can achieve up to 1 dB higher restoration PSNR, with 9× speedup, 23% smaller model size and about 5× lower floating-point operations. Mohit Lamba, Kaushik Mitra |
WACV | 2 |
| 2022 | FlatNet: Towards Photorealistic Scene Reconstruction From Lensless MeasurementsabstractLensless imaging has emerged as a potential solution towards realizing ultra-miniature cameras by eschewing the bulky lens in a traditional camera. Without a focusing lens, the lensless cameras rely on computational algorithms to recover the scenes from multiplexed measurements. However, the current iterative-optimization-based reconstruction algorithms produce noisier and perceptually poorer images. In this work, we propose a non-iterative deep learning-based reconstruction approach that results in orders of magnitude improvement in image quality for lensless reconstructions. Our approach, called FlatNet, lays down a framework for reconstructing high-quality photorealistic images from mask-based lensless cameras, where the camera's forward model formulation is known. FlatNet consists of two stages: (1) an inversion stage that maps the measurement into a space of intermediate reconstruction by learning parameters within the forward model formulation, and (2) a perceptual enhancement stage that improves the perceptual quality of this intermediate reconstruction. These stages are trained together in an end-to-end manner. We show high-quality reconstructions by performing extensive experiments on real and challenging scenes using two different types of lensless prototypes: one which uses a separable forward model and another, which uses a more general non-separable cropped-convolution model. Our end-to-end approach is fast, produces photorealistic reconstructions, and is easy to adopt for other mask-based lensless cameras. Salman Siddique Khan, Varun Sundar, Vivek Boominathan, Ashok Veeraraghavan, Kaushik Mitra |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Toward Unaligned Guided Thermal Super-ResolutionabstractThermography is a useful imaging technique as it works well in poor visibility conditions. High-resolution thermal imaging sensors are usually expensive and this limits the general applicability of such imaging systems. Many thermal cameras are accompanied by a high-resolution visible-range camera, which can be used as a guide to super-resolve the low-resolution thermal images. However, the thermal and visible images form a stereo pair and the difference in their spectral range makes it very challenging to pixel-wise align the two images. The existing guided super-resolution (GSR) methods are based on aligned image pairs and hence are not appropriate for this task. In this paper, we attempt to remove the necessity of pixel-to-pixel alignment for GSR by proposing two models: the first one employs a correlation-based feature-alignment loss to reduce the misalignment in the feature-space itself and the second model includes a misalignment-map estimation block as a part of an end-to-end framework that adequately aligns the input images for performing guided super-resolution. We conduct multiple experiments to compare our methods with existing state-of-the-art single and guided super-resolution techniques and show that our models are better suited for the task of unaligned guided super-resolution from very low-resolution thermal images. Honey Gupta, Kaushik Mitra |
IEEE Trans. Image Process. | 2 |
| 2021 | Restoring Extremely Dark Images in Real TimeabstractA practical low-light enhancement solution must be computationally fast, memory-efficient, and achieve a visually appealing restoration. Most of the existing methods target restoration quality and thus compromise on speed and memory requirements, raising concerns about their real-world deployability. We propose a new deep learning architecture for extreme low-light single image restoration, which despite its fast & lightweight inference, produces a restoration that is perceptually at par with state-of-the-art computationally intense models. To achieve this, we do most of the processing in the higher scale-spaces, skipping the intermediate-scales wherever possible. Also unique to our model is the potential to process all the scale-spaces concurrently, offering an additional 30% speedup without compromising the restoration quality. Pre-amplification of the dark raw-image is an important step in extreme lowlight image enhancement. Most of the existing state of the art methods need GT exposure value to estimate the pre-amplification factor, which is not practically feasible. Thus, we propose an amplifier module that estimates the amplification factor using only the input raw image and can be used "off-the-shelf" with pre-trained models without any fine-tuning. We show that our model can restore an ultrahigh-definition 4K resolution image in just 1 sec. on a CPU and at 32 fps on a GPU and yet maintain a competitive restoration quality. We also show that our proposed model, without any fine-tuning, generalizes well to cameras not seen during training and to subsequent tasks such as object detection. Mohit Lamba, Kaushik Mitra |
CVPR | 2 |
| 2021 | V-DESIRR: Very Fast Deep Embedded Single Image Reflection RemovalabstractReal world images often gets corrupted due to unwanted reflections and their removal is highly desirable. A major share of such images originate from smart phone cameras capable of very high resolution captures. Most of the existing methods either focus on restoration quality by compromising on processing speed and memory requirements or, focus on removing reflections at very low resolutions, there by limiting their practical deploy-ability. We propose a light weight deep learning model for reflection removal using a novel scale space architecture. Our method processes the corrupted image in two stages, a Low Scale Sub-network (LSSNet) to process the lowest scale and a Progressive Inference (PI) stage to process all the higher scales. In order to reduce the computational complexity, the sub-networks in PI stage are designed to be much shallower than LSSNet. Moreover, we employ weight sharing between various scales within the PI stage to limit the model size. This also allows our method to generalize to very high resolutions without explicit retraining. Our method is superior both qualitatively and quantitatively compared to the state of the art methods and at the same time 20× faster with 50× less number of parameters compared to the most recent state-of-the-art algorithm RAGNet. We implemented our method on an android smart phone, where a high resolution 12 MP image is restored in under 5 seconds. B. H. Pawan Prasad, Green Rosh K. S, R. B. Lokesh, Kaushik Mitra, Sanjoy Chowdhury |
ICCV | 4 |
| 2021 | SeLFVi: Self-supervised Light-Field Video Reconstruction from Stereo VideoabstractLight-field imaging is appealing to the mobile devices market because of its capability for intuitive post-capture processing. Acquiring light field (LF) data with high angular, spatial and temporal resolution poses significant challenges, especially with space constraints preventing bulky optics. At the same time, stereo video capture, now available on many consumer devices, can be interpreted as a sparse LF-capture. We explore the application of small baseline stereo videos for reconstructing high fidelity LF videos.We propose a self-supervised learning-based algorithm for LF video reconstruction from stereo video. The self- supervised LF video reconstruction is guided via the geometric information from the individual stereo pairs and the temporal information from the video sequence. LF estimation is further regularized by a low-rank constraint based on layered LF displays. The proposed self-supervised algorithm facilitates advantages such as post-training finetuning on test sequences and variable angular view interpolation and extrapolation. Quantitatively the reconstructed LF videos show higher fidelity than previously proposed unsupervised approaches. We demonstrate our results via LF videos generated from publicly available stereo videos acquired from commercially available stereoscopic cameras. Finally, we demonstrate that our reconstructed LF videos allow applications such as post-capture focus control and region-of-interest (RoI) based focus tracking for videos. Prasan A. Shedligeri, Florian Schiffers, Sushobhan Ghosh, Oliver Cossairt, Kaushik Mitra |
ICCV | 5 |
| 2021 | A Unified Framework for Compressive Video Recovery from Coded Exposure TechniquesabstractSeveral coded exposure techniques have been proposed for acquiring high frame rate videos at low bandwidth. Most recently, a Coded-2-Bucket camera has been proposed that can acquire two compressed measurements in a single exposure, unlike previously proposed coded exposure techniques, which can acquire only a single measurement. Although two measurements are better than one for an effective video recovery, we are yet unaware of the clear advantage of two measurements, either quantitatively or qualitatively. Here, we propose a unified learning-based framework to make such a qualitative and quantitative comparison between those which capture only a single coded image (Flutter Shutter, Pixel-wise coded exposure) and those that capture two measurements per exposure (C2B). Our learning-based framework consists of a shift-variant convolutional layer followed by a fully convolutional deep neural network. Our proposed unified framework achieves the state of the art reconstructions in all three sensing techniques. Further analysis shows that when most scene points are static, the C2B sensor has a significant advantage over acquiring a single pixel-wise coded measurement. However, when most scene points undergo motion, the C2B sensor has only a marginal benefit over the single pixel-wise coded exposure measurement. Prasan A. Shedligeri, Anupama S, Kaushik Mitra |
WACV | 3 |
| 2021 | Whole genome analysis of more than 10 000 SARS-CoV-2 virus unveils global genetic diversity and target region of NSP6abstractWhole genome analysis of SARS-CoV-2 is important to identify its genetic diversity. Moreover, accurate detection of SARS-CoV-2 is required for its correct diagnosis. To address these, first we have analysed publicly available 10 664 complete or near-complete SARS-CoV-2 genomes of 73 countries globally to find mutation points in the coding regions as substitution, deletion, insertion and single nucleotide polymorphism (SNP) globally and country wise. In this regard, multiple sequence alignment is performed in the presence of reference sequence from NCBI. Once the alignment is done, a consensus sequence is build to analyse each genomic sequence to identify the unique mutation points as substitutions, deletions, insertions and SNPs globally, thereby resulting in 7209, 11700, 119 and 53 such mutation points respectively. Second, in such categories, unique mutations for individual countries are determined with respect to other 72 countries. In case of India, unique 385, 867, 1 and 11 substitutions, deletions, insertions and SNPs are present in 566 SARS-CoV-2 genomes while 458, 1343, 8 and 52 mutation points in such categories are common with other countries. In majority (above 10%) of virus population, the most frequent and common mutation points between global excluding India and India are L37F, P323L, F506L, S507G, D614G and Q57H in NSP6, RdRp, Exon, Spike and ORF3a respectively. While for India, the other most frequent mutation points are T1198K, A97V, T315N and P13L in NSP3, RdRp, Spike and ORF8 respectively. These mutations are further visualised in protein structures and phylogenetic analysis has been done to show the diversity in virus genomes. Third, a web application is provided for searching mutation points globally and country wise. Finally, we have identified the potential conserved region as target that belongs to the coding region of ORF1ab, specifically to the NSP6 gene. Subsequently, we have provided the primers and probes using that conserved region so that it can be used for detecting SARS-CoV-2. Contact:[email protected] information: Supplementary data are available at http://www.nitttrkol.ac.in/indrajit/projects/COVID-Mutation-10K. Indrajit Saha, Nimisha Ghosh, Ayan Pradhan, Debasree Maity, Kaushik Mitra |
Briefings Bioinform. | 6 |
| 2021 | High frame rate optical flow estimation from event sensors via intensity estimation
Prasan A. Shedligeri, Kaushik Mitra |
Comput. Vis. Image Underst. | 2 |
| 2021 | Harnessing Multi-View Perspective of Light Fields for Low-Light ImagingabstractLight Field (LF) offers unique advantages such as post-capture refocusing and depth estimation, but low-light conditions severely limit these capabilities. To restore low-light LFs we should harness the geometric cues present in different LF views, which is not possible using single-frame low-light enhancement techniques. We propose a deep neural network L3Fnet for Low-Light Light Field (L3F) restoration, which not only performs visual enhancement of each LF view but also preserves the epipolar geometry across views. We achieve this by adopting a two-stage architecture for L3Fnet. Stage-I looks at all the LF views to encode the LF geometry. This encoded information is then used in Stage-II to reconstruct each LF view. To facilitate learning-based techniques for low-light LF imaging, we collected a comprehensive LF dataset of various scenes. For each scene, we captured four LFs, one with near-optimal exposure and ISO settings and the others at different levels of low-light conditions varying from low to extreme low-light settings. The effectiveness of the proposed L3Fnet is supported by both visual and numerical comparisons on this dataset. To further analyze the performance of low-light restoration methods, we also propose the L3F-wild dataset that contains LF captured late at night with almost zero lux values. No ground truth is available in this dataset. To perform well on the L3F-wild dataset, any method must adapt to the light level of the captured scene. To do this we use a pre-processing block that makes L3Fnet robust to various degrees of low-light conditions. Lastly, we show that L3Fnet can also be used for low-light enhancement of single-frame images, despite it being engineered for LF data. We do so by converting the single-frame DSLR image into a form suitable to L3Fnet, which we call as pseudo-LF. Our code and dataset is available for download at https://mohitlamba94.github.io/L3Fnet/. Mohit Lamba, Kranthi Kumar Rachavarapu, Kaushik Mitra |
IEEE Trans. Image Process. | 3 |
| 2020 | Towards Fast and Light-Weight Restoration of Dark Images
Mohit Lamba, Atul Balaji, Kaushik Mitra |
BMVC | 3 |
| 2020 | Multi-Patch Aggregation Models for Resampling DetectionabstractImages captured nowadays are of varying dimensions with smartphones and DSLR's allowing users to choose from a list of available image resolutions. It is therefore imperative for forensic algorithms such as resampling detection to scale well for images of varying dimensions. However, in our experiments we observed that many state-of-the-art forensic algorithms are sensitive to image size and their performance quickly degenerates when operated on images of diverse dimensions despite re-training them using multiple image sizes. To handle this issue, we propose two novel deep neural networks - Iterative Pooling Network (IPN), which does not assume any prior information about the original image size, and Branched Network (BN), which uses this prior knowledge to produce better results. IPN adopts a novel iterative pooling strategy that converts tensors of multiple sizes to tensors of a fixed size, as required by deep learning models with fully connected layers. BN alternatively adopts a branched architecture with dedicated pathways for images of different sizes. The effectiveness of the proposed solution is demonstrated on two problems, resampling detection and photorealism detection, which are generally solved as independent problems with different deep learning models. The code is available at https://github.com/MohitLamba94/Iterative-Pooling. Mohit Lamba, Kaushik Mitra |
ICASSP | 2 |
| 2020 | CANOPIC: Pre-Digital Privacy-Enhancing Encodings for Computer VisionabstractThe standard pipeline for many vision tasks uses a conventional camera to capture an image that is then passed to a digital processor for information extraction. In some deployments, such as private locations, the captured digital imagery contains sensitive information exposed to digital vulnerabilities such as spyware, Trojans, etc. However, in many applications, the full imagery is unnecessary for the vision task at hand. In this paper we propose an optical and analog system that preprocesses the light from the scene before it reaches the digital imager to destroy sensitive information. We explore analog and optical encodings consisting of easily implementable operations such as convolution, pooling, and quantization. We perform a case study to evaluate how such encodings can destroy face identity information while preserving enough information for face detection. The encoding parameters are learned via an alternating optimization scheme based on adversarial learning with deep neural networks. We name our system CAnOPIC (Camera with Analog and Optical Privacy-Integrating Computations) and show that it has better performance in terms of both privacy and utility than conventional optical privacy-enhancing methods such as blurring and pixelation. Jasper Tan, Salman Siddique Khan, Vivek Boominathan, Jeffrey Byrne, Richard G. Baraniuk, Kaushik Mitra, Ashok Veeraraghavan |
ICME | 6 |
| 2020 | Video reconstruction by spatio-temporal fusion of blurred-coded image pairabstractLearning-based methods have enabled the recovery of a video sequence from a single motion-blurred image or a single coded exposure image. Recovering video from a single motion-blurred image is a very ill-posed problem and the recovered video usually has many artifacts. In addition to this, the direction of motion is lost and it results in motion ambiguity. However, it has the advantage of fully preserving the information in the static parts of the scene. The traditional coded exposure framework is better-posed but it only samples a fraction of the space-time volume, which is at best 50% of the space-time volume. Here, we propose to use the complementary information present in the fully-exposed (blurred) image along with the coded exposure image to recover a high fidelity video without any motion ambiguity. Our framework consists of a shared encoder followed by an attention module to selectively combine the spatial information from the fully-exposed image with the temporal information from the coded image, which is then super-resolved to recover a non-ambiguous high-quality video. The input to our algorithm is a fully-exposed and coded image pair. Such an acquisition system already exists in the form of a Coded-two-bucket (C2B) camera. We demonstrate that our proposed deep learning approach using blurred-coded image pair produces much better results than those from just a blurred image or just a coded image. Anupama S, Prasan A. Shedligeri, Abhishek Pal, Kaushik Mitra |
ICPR | 4 |
| 2019 | Towards Photorealistic Reconstruction of Highly Multiplexed Lensless ImagesabstractRecent advancements in fields like Internet of Things (IoT), augmented reality, etc. have led to an unprecedented demand for miniature cameras with low cost that can be integrated anywhere and can be used for distributed monitoring. Mask-based lensless imaging systems make such inexpensive and compact models realizable. However, reduction in the size and cost of these imagers comes at the expense of their image quality due to the high degree of multiplexing inherent in their design. In this paper, we present a method to obtain image reconstructions from mask-based lensless measurements that are more photorealistic than those currently available in the literature. We particularly focus on FlatCam, a lensless imager consisting of a coded mask placed over a bare CMOS sensor. Existing techniques for reconstructing FlatCam measurements suffer from several drawbacks including lower resolution and dynamic range than lens-based cameras. Our approach overcomes these drawbacks using a fully trainable non-iterative deep learning based model. Our approach is based on two stages: an inversion stage that maps the measurement into the space of intermediate reconstruction and a perceptual enhancement stage that improves this intermediate reconstruction based on perceptual and signal distortion metrics. Our proposed method is fast and produces photo-realistic reconstruction as demonstrated on many real and challenging scenes. Salman Siddique Khan, Adarsh V. R, Vivek Boominathan, Jasper Tan, Ashok Veeraraghavan, Kaushik Mitra |
ICCV | 6 |
| 2019 | Unsupervised Single Image Underwater Depth EstimationabstractDepth estimation from a single underwater image is one of the most challenging problems and is highly ill-posed. Due to the absence of large generalized underwater depth datasets and the difficulty in obtaining ground truth depth-maps, supervised learning techniques such as direct depth regression cannot be used. In this paper, we propose an unsupervised method for depth estimation from a single underwater image taken "in the wild" by using haze as a cue for depth. Our approach is based on indirect depth-map estimation where we learn the mapping functions between unpaired RGB-D terrestrial images and arbitrary underwater images to estimate the required depth-map. We propose a method which is based on the principles of cycle-consistent learning and uses dense-block based auto-encoders as generator networks. We evaluate and compare our method both quantitatively and qualitatively on various underwater images with diverse attenuation and scattering conditions and show that our method produces state-of-the-art results for unsupervised depth estimation from a single underwater image. Honey Gupta, Kaushik Mitra |
ICIP | 2 |
| 2019 | Neural Decoder for Topological Codes using Pseudo-Inverse of Parity Check MatrixabstractRecent developments in the field of deep learning have motivated many researchers to apply these methods to problems in quantum information. Torlai and Melko first proposed a decoder for surface codes based on neural networks. Since then, many other researchers have applied neural networks to study a variety of problems in the context of decoding. An important development in this regard was due to Varsamopoulos et at. who proposed a two-step decoder using neural networks. Subsequent work of Maskara et at. used the same concept for decoding for various noise models. We propose a similar two-step neural decoder using inverse parity-check matrix for topological color codes. We show that it outperforms the state-of-the-art performance of non-neural decoders for independent Pauli errors noise model on a 2D hexagonal color code. Our final decoder achieves a threshold of 10%. Our result is comparable to the recent work on neural decoder for quantum error correction by Maskara et at. It appears that our decoder has advantages with respect to training cost and complexity of the network for higher distances when compared to that of Maskara et at. Chaitanya Chinni, Abhishek Kulkarni, Dheeraj M. Pai, Kaushik Mitra, Pradeep Kiran Sarvepalli |
ITW | 4 |
| 2019 | Fully Convolutional Networks for Monocular Retinal Depth Estimation and Optic Disc-Cup SegmentationabstractGlaucoma is a serious ocular disorder for which the screening and diagnosis are carried out by the examination of the optic nerve head (ONH). The color fundus image (CFI) is the most common modality used for ocular screening. In CFI, the central region which is the optic disc and the optic cup region within the disc are examined to determine one of the important cues for glaucoma diagnosis called the optic cup-to-disc ratio (CDR). CDR calculation requires accurate segmentation of optic disc and cup. Another important cue for glaucoma progression is the variation of depth in ONH region. In this paper, we first propose a deep learning framework to estimate depth from a single fundus image. For the case of monocular retinal depth estimation, we are also plagued by the labeled data insufficiency. To overcome this problem we adopt the technique of pretraining the deep network where, instead of using a denoising autoencoder, we propose a new pretraining scheme called pseudo-depth reconstruction, which serves as a proxy task for retinal depth estimation. Empirically, we show pseudo-depth reconstruction to be a better proxy task than denoising. Our results outperform the existing techniques for depth estimation on the INSPIRE dataset. To extend the use of depth map for optic disc and cup segmentation, we propose a novel fully convolutional guided network, where, along with the color fundus image the network uses the depth map as a guide. We propose a convolutional block called multimodal feature extraction block to extract and fuse the features of the color image and the guide image. We extensively evaluate the proposed segmentation scheme on three datasets- ORIGA, RIMONEr3, and DRISHTI-GS. The performance of the method is comparable and in many cases, outperforms the most recent state of the art. Sharath M. Shankaranarayana, Keerthi Ram, Kaushik Mitra, Mohanasankar Sivaprakasam |
IEEE J. Biomed. Health Informatics | 3 |
| 2018 | Phase retrieval for Fourier Ptychography under varying amount of measurements
Lokesh Boominathan, Mayug Maniparambil, Honey Gupta, Rahul Baburajan, Kaushik Mitra |
BMVC | 5 |
| 2017 | Compressive image recovery using recurrent generative modelabstractReconstruction of signals from compressively sensed measurements is an ill-posed problem. In this paper, we leverage the recurrent generative model, RIDE, as an image prior for compressive image reconstruction. Recurrent networks can model long-range dependencies in images and hence can handle global multiplexing in compressive imaging. We perform MAP inference with RIDE using back-propagation to the inputs and projected gradient method. We propose an entropy thresholding based approach for preserving texture in images well. Our approach shows superior reconstructions compared to recent global reconstruction approaches like D-AMP and TVAL3 on both simulated and real data. Akshat Dave, Anil Kumar Vadathya, Kaushik Mitra |
ICIP | 3 |
| 2017 | Data driven coded aperture design for depth recoveryabstractInserting a patterned occluder at the aperture of a camera lens has been shown to improve the recovery of depth map and all-focus image compared to a fully open aperture. However, design of the aperture pattern plays a very critical role. Previous approaches for designing aperture codes make simple assumptions on image distributions to obtain metrics for evaluating aperture codes. However, real images may not follow those assumptions and hence the designed code may not be optimal for them. To address this drawback we propose a data driven approach for learning the optimal aperture pattern to recover depth map from a single coded image. We propose a two stage architecture where, in the first stage we simulate coded aperture images from a training dataset of all-focus images and depth maps and in the second stage we recover the depth map using a deep neural network. We demonstrate that our learned aperture code performs better than previously designed codes even on code design metrics proposed by previous approaches. Prasan A. Shedligeri, Sreyas Mohan, Kaushik Mitra |
ICIP | 3 |
| 2016 | Focal-sweep for large aperture time-of-flight camerasabstractTime-of-flight (ToF) imaging is an active method that utilizes a temporally modulated light source and a correlation-based (or lock-in) imager that computes the round-trip travel time from source to scene and back. Much like conventional imaging ToF cameras suffer from the trade-off between depth of field (DOF) and light throughput-larger apertures allow for more light collection but results in lower DoF. This trade-off is especially crucial in ToF systems since they require active illumination and have limited power, which limits performance in long-range imaging or imaging in strong ambient illumination (such as outdoors). Motivated by recent work in extended depth of field imaging for photography, we propose a focal sweep-based image acquisition methodology to increase depth-of-field and eliminate defocus blur. Our approach allows for a simple inversion algorithm to recover all-in-focus images. We validate our technique through simulation and experimental results. We demonstrate a proof-of-concept focal sweep time-of-flight acquisition system and show results for a real scene. Sagar Honnungar, Jason Holloway, Adithya Kumar Pediredla, Ashok Veeraraghavan, Kaushik Mitra |
ICIP | 5 |
| 2016 | Spatial Phase-Sweep: Increasing temporal resolution of transient imaging using a light source arrayabstractTransient imaging techniques capture the propagation of an ultra-short pulse of light through a scene, which in effect captures the optical impulse response of the scene. Recently, it has been shown that we can capture transient images using commercial, correlation imager based Time-of-Flight (ToF) systems. But the temporal resolution of these transient images are currently limited by high-speed electronics. In this paper, we propose `Spatial Phase-Sweep' (SPS), a technique that exploits the speed of light to increase the temporal resolution of transient imaging beyond the limit imposed by electronic circuits in these commercial ToF sensors. SPS uses a linear array of light sources with a controlled spatial separation between these sources. The differential positioning of these sources introduce sub nano-second time shifts in the light wavefront, improving the time resolution of captured transients. As a proof of concept, we demonstrate a prototype which improves the temporal resolution of transient imaging by a factor of 10x, without any modification to the underlying electronics. Ryuichi Tadano, Adithya Kumar Pediredla, Kaushik Mitra, Ashok Veeraraghavan |
ICIP | 3 |
| 2015 | Generalized Assorted Camera Arrays: Robust Cross-Channel Registration and ApplicationsabstractOne popular technique for multimodal imaging is generalized assorted pixels (GAP), where an assorted pixel array on the image sensor allows for multimodal capture. Unfortunately, GAP is limited in its applicability because of the need for multimodal filters that are amenable with semiconductor fabrication processes and results in a fixed multimodal imaging configuration. In this paper, we advocate for generalized assorted camera (GAC) arrays for multimodal imaging--i.e., a camera array with filters of different characteristics placed in front of each camera aperture. The GAC provides us with three distinct advantages over GAP: ease of implementation, flexible application-dependent imaging since filters are external and can be changed and depth information that can be used for enabling novel applications (e.g., postcapture refocusing). The primary challenge in GAC arrays is that since the different modalities are obtained from different viewpoints, there is a need for accurate and efficient cross-channel registration. Traditional approaches such as sum-of-squared differences, sum-of-absolute differences, and mutual information all result in multimodal registration errors. Here, we propose a robust cross-channel matching cost function, based on aligning normalized gradients, which allows us to compute cross-channel subpixel correspondences for scenes exhibiting nontrivial geometry. We highlight the promise of GAC arrays with our cross-channel normalized gradient cost for several applications such as low-light imaging, postcapture refocusing, skin perfusion imaging using color + near infrared, and hyperspectral imaging. Jason Holloway, Kaushik Mitra, Sanjeev J. Koppal, Ashok Veeraraghavan |
IEEE Trans. Image Process. | 2 |
| 2014 | Improving resolution and depth-of-field of light field cameras using a hybrid imaging systemabstractCurrent light field (LF) cameras provide low spatial resolution and limited depth-of-field (DOF) control when compared to traditional digital SLR (DSLR) cameras. We show that a hybrid imaging system consisting of a standard LF camera and a high-resolution standard camera enables (a) achieve high-resolution digital refocusing, (b) better DOF control than LF cameras, and (c) render graceful high-resolution viewpoint variations, all of which were previously unachievable. We propose a simple patch-based algorithm to super-resolve the low-resolution views of the light field using the high-resolution patches captured using a high-resolution SLR camera. The algorithm does not require the LF camera and the DSLR to be co-located or for any calibration information regarding the two imaging systems. We build an example prototype using a Lytro camera (380×380 pixel spatial resolution) and a 18 megapixel (MP) Canon DSLR camera to generate a light field with 11 MP resolution (9× super-resolution) and about 1 over 9thof the DOF of the Lytro camera. We show several experimental results on challenging scenes containing occlusions, specularities and complex non-lambertian materials, demonstrating the effectiveness of our approach. Vivek Boominathan, Kaushik Mitra, Ashok Veeraraghavan |
ICCP | 2 |
| 2014 | Can we beat Hadamard multiplexing? Data driven design and analysis for computational imaging systemsabstractComputational Imaging (CI) systems that exploit optical multiplexing and algorithmic demultiplexing have been shown to improve imaging performance in tasks such as motion deblurring, extended depth of field, light field and hyper-spectral imaging. Design and performance analysis of many of these approaches tend to ignore the role of image priors. It is well known that utilizing statistical image priors significantly improves demultiplexing performance. In this paper, we extend the Gaussian Mixture Model as a data-driven image prior (proposed by Mitra et. al [21]) to under-determined linear systems and study compressive CI methods such as light-field and hyper-spectral imaging. Further, we derive a novel algorithm for optimizing multiplexing matrices that simultaneously accounts for (a) sensor noise (b) image priors and (c) CI design constraints. We use our algorithm to design data-optimal multiplexing matrices for a variety of existing CI designs, and we use these matrices to analyze the performance of CI systems as a function of noise level. Our analysis gives new insight into the optimal performance of CI systems, and how this relates to the performance of classical multiplexing designs such as Hadamard matrices. Kaushik Mitra, Oliver Cossairt, Ashok Veeraraghavan |
ICCP | 1 |
| 2014 | A Framework for Analysis of Computational Imaging Systems: Role of Signal Prior, Sensor Noise and MultiplexingabstractOver the last decade, a number of computational imaging (CI) systems have been proposed for tasks such as motion deblurring, defocus deblurring and multispectral imaging. These techniques increase the amount of light reaching the sensor via multiplexing and then undo the deleterious effects of multiplexing by appropriate reconstruction algorithms. Given the widespread appeal and the considerable enthusiasm generated by these techniques, a detailed performance analysis of the benefits conferred by this approach is important. Unfortunately, a detailed analysis of CI has proven to be a challenging problem because performance depends equally on three components: (1) the optical multiplexing, (2) the noise characteristics of the sensor, and (3) the reconstruction algorithm which typically uses signal priors. A few recent papers [12], [30], [49] have performed analysis taking multiplexing and noise characteristics into account. However, analysis of CI systems under state-of-the-art reconstruction algorithms, most of which exploit signal prior models, has proven to be unwieldy. In this paper, we present a comprehensive analysis framework incorporating all three components. In order to perform this analysis, we model the signal priors using a Gaussian Mixture Model (GMM). A GMM prior confers two unique characteristics. First, GMM satisfies the universal approximation property which says that any prior density function can be approximated to any fidelity using a GMM with appropriate number of mixtures. Second, a GMM prior lends itself to analytical tractability allowing us to derive simple expressions for the `minimum mean square error' (MMSE) which we use as a metric to characterize the performance of CI systems. We use our framework to analyze several previously proposed CI techniques (focal sweep, flutter shutter, parabolic exposure, etc.), giving conclusive answer to the question: `How much performance gain is due to use of a signal prior and how much is due to multiplexing? Our analysis also clearly shows that multiplexing provides significant performance gains above and beyond the gains obtained due to use of signal priors. Kaushik Mitra, Oliver Cossairt, Ashok Veeraraghavan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Compressive epsilon photography for post-capture control in digital imagingabstractA traditional camera requires the photographer to select the many parameters at capture time. While advances in light field photography have enabled post-capture control of focus and perspective, they suffer from several limitations including lower spatial resolution, need for hardware modifications, and restrictive choice of aperture and focus setting. In this paper, we propose "compressive epsilon photography," a technique for achieving complete post-capture control of focus and aperture in a traditional camera by acquiring a carefully selected set of 8 to 16 images and computationally reconstructing images corresponding to all other focus-aperture settings. We make the following contributions: first, we learn the statistical redundancies in focal-aperture stacks using a Gaussian Mixture Model; second, we derive a greedy sampling strategy for selecting the best focus-aperture settings; and third, we develop an algorithm for reconstructing the entire focal-aperture stack from a few captured images. As a consequence, only a burst of images with carefully selected camera settings are acquired. Post-capture, the user can then select any focal-aperture setting of choice and the corresponding image can be rendered using our algorithm. We show extensive results on several real data sets. Atsushi Ito, Salil Tambe, Kaushik Mitra, Aswin C. Sankaranarayanan, Ashok Veeraraghavan |
ACM Trans. Graph. | 3 |
| 2013 | Blur and Illumination Robust Face Recognition via Set-Theoretic CharacterizationabstractWe address the problem of unconstrained face recognition from remotely acquired images. The main factors that make this problem challenging are image degradation due to blur, and appearance variations due to illumination and pose. In this paper, we address the problems of blur and illumination. We show that the set of all images obtained by blurring a given image forms a convex set. Based on this set-theoretic characterization, we propose a blur-robust algorithm whose main step involves solving simple convex optimization problems. We do not assume any parametric form for the blur kernels, however, if this information is available it can be easily incorporated into our algorithm. Furthermore, using the low-dimensional model for illumination variations, we show that the set of all images obtained from a face image by blurring it and by changing the illumination conditions forms a bi-convex set. Based on this characterization, we propose a blur and illumination-robust algorithm. Our experiments on a challenging real dataset obtained in uncontrolled settings illustrate the importance of jointly modeling blur and illumination. Priyanka Vageeswaran, Kaushik Mitra, Rama Chellappa |
IEEE Trans. Image Process. | 2 |
| 2012 | A hierarchical approach for human age estimationabstractWe consider the problem of automatic age estimation from face images. Age estimation is usually formulated as a regression problem relating the facial features and the age variable, and a single regression model is learnt for all ages. We propose a hierarchical approach, where we first divide the face images into various age groups and then learn a separate regression model for each group. Given a test image, we first classify the image into one of the age groups and then use the regression model for that particular group. To improve our classification result, we use many different classifiers and fuse them using the majority rule. Experiments show that our approach outperforms many state of the art regression methods for age estimation. Pavleen Thukral, Kaushik Mitra, Rama Chellappa |
ICASSP | 2 |
| 2010 | Robust RVM regression using sparse outlier modelabstractKernel regression techniques such as Relevance Vector Machine (RVM) regression, Support Vector Regression and Gaussian processes are widely used for solving many computer vision problems such as age, head pose, 3D human pose and lighting estimation. However, the presence of outliers in the training dataset makes the estimates from these regression techniques unreliable. In this paper, we propose robust versions of the RVM regression that can handle outliers in the training dataset. We decompose the noise term in the RVM formulation into a (sparse) outlier noise term and a Gaussian noise term. We then estimate the outlier noise along with the model parameters. We present two approaches for solving this estimation problem: (1) a Bayesian approach, which essentially follows the RVM framework and (2) an optimization approach based on Basis Pursuit Denoising. In the Bayesian approach, the robust RVM problem essentially becomes a bigger RVM problem with the advantage that it can be solved efficiently by a fast algorithm. Empirical evaluations, and real experiments on image de-noising and age estimation demonstrate the better performance of the robust RVM algorithms over that of the RVM reg ression. Kaushik Mitra, Ashok Veeraraghavan, Rama Chellappa |
CVPR | 1 |
| 2010 | Robust regression using sparse learning for high dimensional parameter estimation problemsabstractAlgorithms such as Least Median of Squares (LMedS) and Random Sample Consensus (RANSAC) have been very successful for low-dimensional robust regression problems. However, the combinatorial nature of these algorithms makes them practically unusable for high-dimensional applications. In this paper, we introduce algorithms that have cubic time complexity in the dimension of the problem, which make them computationally efficient for high-dimensional problems. We formulate the robust regression problem by projecting the dependent variable onto the null space of the independent variables which receives significant contributions only from the outliers. We then identify the outliers using sparse representation/learning based algorithms. Under certain conditions, that follow from the theory of sparse representation, these polynomial algorithms can accurately solve the robust regression problem which is, in general, a combinatorial problem. We present experimental results that demonstrate the efficacy of the proposed algorithms. We also analyze the intrinsic parameter space of robust regression and identify an efficient and accurate class of algorithms for different operating conditions. An application to facial age estimation is presented. Kaushik Mitra, Ashok Veeraraghavan, Rama Chellappa |
ICASSP | 1 |
| 2010 | Large-Scale Matrix Factorization with Missing Data under Additional ConstraintsabstractMatrix factorization in the presence of missing data is at the core of many computer vision problems such as structure from motion (SfM), non-rigid SfM and photometric stereo. We formulate the problem of matrix factorization with missing data as a low-rank semidefinite program (LRSDP) with the advantage that: $1)$ an efficient quasi-Newton implementation of the LRSDP enables us to solve large-scale factorization problems, and $2)$ additional constraints such as ortho-normality, required in orthographic SfM, can be directly incorporated in the new formulation. Our empirical evaluations suggest that, under the conditions of matrix completion theory, the proposed algorithm finds the optimal solution, and also requires fewer observations compared to the current state-of-the-art algorithms. We further demonstrate the effectiveness of the proposed algorithm in solving the affine SfM problem, non-rigid SfM and photometric stereo problems. Kaushik Mitra, Sameer Sheorey, Rama Chellappa |
NIPS | 1 |