VLDB 2026 Research / reviewers in the wild / expert
Suhas Lohit
dblp:169/9097
· DBLP profile ↗
24ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0002-0392-3818ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Recovering Pulse Waves From Video Using Deep Unrolling and Deep Equilibrium ModelsabstractCamera-based contactless monitoring of vital signs, also known as imaging photoplethysmography (iPPG), has seen applications in driver-monitoring, perfusion assessment, affective computing, and more. iPPG involves sensing the underlying cardiac pulse from video of the skin and estimating vital signs such as the pulse rate or a full pulse waveform. Some previous iPPG methods impose model-based sparse priors on the pulse signals and use iterative optimization for pulse wave recovery, while others use end-to-end black-box deep learning methods. In contrast, we introduce methods that combine signal processing and deep learning methods in an inverse problem framework. Our methods estimate the underlying pulse signal, pulse rate, and pulse rate variability from facial video by learning deep-network-based denoising operators that leverage deep algorithm unfolding and deep equilibrium models. Experiments show that our methods can denoise an acquired signal from the face and infer the correct underlying pulse rate and pulse rate variability, achieving pulse rate estimation performance consistent with the state-of-the-art on well-known benchmarks, all with less than one-fifth the number of learnable parameters as the closest competing method. Vineet R. Shenoy, Suhas Lohit, Hassan Mansour, Rama Chellappa, Tim K. Marks |
IEEE Trans. Image Process. | 2 |
| 2024 | Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-Aware Spatio-Temporal SamplingabstractExtensions of Neural Radiance Fields (NeRFs) to model dynamic scenes have enabled their near photo-realistic, free-viewpoint rendering. Although these methods have shown some potential in creating immersive experiences, two drawbacks limit their ubiquity: ( i) a significant reduction in reconstruction quality when the computing budget is limited, and (ii) a lack of semantic understanding of the underlying scenes. To address these issues, we introduce Gear-NeRF, which leverages semantic information from powerful image segmentation models. Our approach presents a principled way for learning a spatio-temporal (4D) semantic embedding, based on which we introduce the concept of gears to allow for stratified modeling of dynamic regions of the scene based on the extent of their motion. Such differentiation allows us to adjust the spatio- temporal sampling resolution for each region in proportion to its motion scale, achieving more photo-realistic dynamic novel view synthesis. At the same time, almost for free, our approach enables free-viewpoint tracking of objects of interest - a functionality not yet achieved by existing NeRF-based methods. Empirical studies validate the effectiveness of our method, where we achieve state-of-the-art rendering and tracking performance on multiple challenging datasets. The project page is available at: https://merl.com/research/highlights/gear-nerf Xinhang Liu, Yu-Wing Tai, Chi-Keung Tang, Pedro Miraldo, Suhas Lohit, Moitreya Chatterjee |
CVPR | 5 |
| 2024 | TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion ModelsabstractText-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g., a woman's photo) and a text description (e.g., “a woman is drinking water.”). Existing TI2V frameworks often require costly training on video-text datasets and spe-cific model designs for text and image conditioning. In this paper, we propose TI2V-Zero, a zero-shot, tuning-free method that empowers a pretrained text-to-video (T2V) diffusion model to be conditioned on a provided image, enabling TI2V generation without any optimization, fine-tuning, or introducing external modules. Our approach leverages a pretrained T2V diffusion foundation model as the generative prior. To guide video generation with the additional image input, we propose a “repeat-and-slide” strategy that modulates the reverse denoising process, al-lowing the frozen diffusion model to synthesize a video frame-by-frame starting from the provided image. To ensure temporal continuity, we employ a DDPM inversion strategy to initialize Gaussian noise for each newly synthesized frame and a resampling technique to help preserve visual details. We conduct comprehensive experiments on both domain-specific and open-domain datasets, where TI2V-Zero consistently outperforms a recent open-domain TI2V model. Furthermore, we show that TI2V-Zero can seam-lessly extend to other tasks such as video infilling and pre-diction when provided with more images. Its autoregressive design also supports long video generation. Haomiao Ni, Bernhard Egger 0001, Suhas Lohit, Anoop Cherian, Ye Wang 0001, Toshiaki Koike-Akino, Sharon X. Huang, Tim K. Marks |
CVPR | 3 |
| 2024 | Equivariant Spatio-temporal Self-supervision for LiDAR Object Detection
Deepti Hegde, Suhas Lohit, Kuan-Chuan Peng, Michael J. Jones 0001, Vishal M. Patel |
ECCV (26) | 2 |
| 2024 | Evaluating Large Vision-and-Language Models on Children's Mathematical OlympiadsabstractRecent years have seen a significant progress in the general-purpose problem solving abilities of large vision and language models (LVLMs), such as ChatGPT, Gemini, etc.; some of these breakthroughs even seem to enable AI models to outperform human abilities in varied tasks that demand higher-order cognitive skills. Are the current large AI models indeed capable of generalized problem solving as humans do? A systematic analysis of AI capabilities for joint vision and text reasoning, however, is missing in the current scientific literature. In this paper, we make an effort towards filling this gap, by evaluating state-of-the-art LVLMs on their mathematical and algorithmic reasoning abilities using visuo-linguistic problems from children's Olympiads. Specifically, we consider problems from the Mathematical Kangaroo (MK) Olympiad, which is a popular international competition targeted at children from grades 1-12, that tests children's deeper mathematical abilities using puzzles that are appropriately gauged to their age and skills. Using the puzzles from MK, we created a dataset, dubbed SMART-840, consisting of 840 problems from years 2020-2024. With our dataset, we analyze LVLMs power on mathematical reasoning; their responses on our puzzles offer a direct way to compare against that of children. Our results show that modern LVLMs do demonstrate increasingly powerful reasoning skills in solving problems for higher grades, but lack the foundations to correctly answer problems designed for younger children. Further analysis shows that there is no significant correlation between the reasoning capabilities of AI models and that of young children, and their capabilities appear to be based on a different type of reasoning than the cumulative knowledge that underlies children's mathematical skills. Anoop Cherian, Kuan-Chuan Peng, Suhas Lohit, Joanna Matthiesen, Kevin A. Smith 0001, Josh Tenenbaum |
NeurIPS | 3 |
| 2024 | Pixel-Grounded Prototypical Part NetworksabstractPrototypical part neural networks (ProtoPartNNs), namely ProtoPNet and its derivatives, are an intrinsically interpretable approach to machine learning. Their prototype learning scheme enables intuitive explanations of the form, this (prototype) looks like that (testing image patch). But, does this actually look like that? In this work, we delve into why object part localization and associated heat maps in past work are misleading. Rather than localizing to object parts, existing ProtoPartNNs localize to the entire image, contrary to generated explanatory visualizations. We argue that detraction from these underlying issues is due to the alluring nature of visualizations and an over-reliance on intuition. To alleviate these issues, we devise new receptive field-based architectural constraints for meaningful localization and a principled pixel space mapping for ProtoPartNNs. To improve interpretability, we propose additional architectural improvements, including a simplified classification head. We also make additional corrections to ProtoPNet and its derivatives, such as the use of a validation set, rather than a test set, to evaluate generalization during training. Our approach, PixPNet (Pixel-grounded Prototypical part Network), is the only ProtoPartNN that truly learns and localizes to prototypical object parts. We demonstrate that PixPNet achieves quantifiably improved interpretability without sacrificing accuracy1. Zachariah Carmichael, Suhas Lohit, Anoop Cherian, Michael J. Jones 0001, Walter J. Scheirer |
WACV | 2 |
| 2023 | Are Deep Neural Networks SMARTer Than Second Graders?abstractRecent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating art, question answering (e.g., ChatGPT), etc. Such a dramatic progress raises the question: how generalizable are neural networks in solving problems that demand broad skills? To answer this question, we propose SMART: a Simple Multimodal Algorithmic Reasoning Task and the associated SMART-101 dataset11The SMART-101 dataset is publicly available at: https://doi.org/10.5281/zenodo.7761800, for evaluating the abstraction, deduction, and generalization abilities of neural networks in solving visuo-linguistic puzzles designed specifically for children in the 6–8 age group. Our dataset consists of 101 unique puzzles; each puzzle comprises a picture and a question, and their solution needs a mix of several elementary skills, including arithmetic, algebra, and spatial reasoning, among others. To scale our dataset towards training deep neural networks, we programmatically generate entirely new instances for each puzzle while retaining their solution algorithm. To benchmark the performance on the SMART-101 dataset, we propose a vision-and-language meta-learning model that can incorporate varied state-of-the-art neural backbones. Our experiments reveal that while powerful deep models offer reasonable performances on puzzles in a supervised setting, they are not better than random accuracy when analyzed for generalization –filling this gap may demand new multimodal learning approaches. Anoop Cherian, Kuan-Chuan Peng, Suhas Lohit, Kevin A. Smith 0001, Josh Tenenbaum |
CVPR | 3 |
| 2023 | Robust Time Series Recovery and Classification Using Test-Time Noise Simulator NetworksabstractTime-series are commonly susceptible to various types of corruption due to sensor-level changes and defects which can result in missing samples, sensor and quantization noise, unknown calibration, unknown phase shifts etc. These corruptions cannot be easily corrected as the noise model may be unknown at the time of deployment. This also results in the inability to employ pre-trained classifiers, trained on (clean) source data. In this paper, we present a general framework and models for time-series that can make use of (unlabeled) test samples to estimate the noise model-entirely at test time. To this end, we use a coupled decoder model and an additional neural network which acts as a learned noise model simulator. We show that the framework is able to "clean" the data so as to match the source training data statistics and the cleaned data can be directly used with a pre-trained classifier for robust predictions. We perform empirical studies on diverse application domains with different types of sensors, clearly demonstrating the effectiveness and generality of this method. Eun Som Jeon, Suhas Lohit, Rushil Anirudh, Pavan Turaga |
ICASSP | 2 |
| 2023 | Steered Diffusion: A Generalized Framework for Plug-and-Play Conditional Image SynthesisabstractConditional generative models typically demand large annotated training sets to achieve high-quality synthesis. As a result, there has been significant interest in designing models that perform plug-and-play generation, i.e., to use a predefined or pretrained model, which is not explicitly trained on the generative task, to guide the generative process (e.g., using language). However, such guidance is typically useful only towards synthesizing high-level semantics rather than editing fine-grained details as in image-to-image translation tasks. To this end, and capitalizing on the powerful fine-grained generative control offered by the recent diffusion-based generative models, we introduce Steered Diffusion, a generalized framework for photorealistic zero-shot conditional image generation using a diffusion model trained for unconditional generation. The key idea is to steer the image generation of the diffusion model at inference time via designing a loss using a pre-trained inverse model that characterizes the conditional task. This loss modulates the sampling trajectory of the diffusion process. Our framework allows for easy incorporation of multiple conditions during inference. We present experiments using steered diffusion on several tasks including inpainting, colorization, text-guided semantic editing, and image super-resolution. Our results demonstrate clear qualitative and quantitative improvements over state-of-the-art diffusion-based plug-and-play models while adding negligible additional computational cost. Nithin Gopalakrishnan Nair, Anoop Cherian, Suhas Lohit, Ye Wang 0001, Toshiaki Koike-Akino, Vishal M. Patel, Tim K. Marks |
ICCV | 3 |
| 2023 | Unrolled iPPG: Video Heart Rate Estimation via Unrolling Proximal Gradient DescentabstractImaging photoplethysmography (iPPG) is the process of estimating a person’s heart rate from video. In this work, we propose Unrolled iPPG, in which we integrate iterative optimization updates with deep learning-based signal priors to estimate the pulse waveform and heart rate from facial videos. We model the signal extracted from video as the sum of an underlying pulse signal and noise, but instead of explicitly imposing a handcrafted prior (e.g., sparsity in the frequency domain) on the signal, we learn priors on the signal and noise using neural networks. We solve for the underlying pulse signal by unrolling proximal gradient descent; the algorithm alternates between gradient descent steps and application of learned denoisers, which replace handcrafted priors and their proximal operators. Using this method, we achieve state-of-the-art heart rate estimation on the challenging MMSE-HR dataset. Vineet R. Shenoy, Tim K. Marks, Hassan Mansour, Suhas Lohit |
ICIP | 4 |
| 2022 | Cross-Modal Knowledge Transfer Without Task-Relevant Source Data
Sk Miraj Ahmed, Suhas Lohit, Kuan-Chuan Peng, Michael J. Jones 0001, Amit K. Roy-Chowdhury |
ECCV (34) | 2 |
| 2022 | Distributed Radar Autofocus Imaging Using Deep PriorsabstractAntenna position ambiguity is a common problem that affects radar imaging systems that are mounted on mobile platforms. Existing approaches that aim to recover a sharp radar image despite this ambiguity aim to estimate the shift in the antenna position by modeling the radar scene as a sparse image with a small number of targets using explicit analytical models for the statistical distribution of the targets in a radar image. The radar imaging problem is then solved by alternating between estimating the radar image, followed by estimating the shift in the antenna positions, until convergence is reached. While such approaches have shown tremendous success, they still struggle to recover the true target positions and may arrive at incorrect local optima when the measurement noise level is high. In this work, we develop a data-driven learning-based strategy for modeling the image of the radar scene instead of relying on explicit analytical models. We adopt a residual Unet architecture of a neural network to act as a denoising operator which takes a backprojected radar image as input and outputs a true target image. While deep denoisers may generally result in unstable iterative algorithms, we introduce a simple filtering step that suppresses noise belonging to the null space of the radar operator from the iterates to stabilize the iterative procedure. We evaluate the effectiveness of our solution using simulated numerical experiments and demonstrate its superiority over the analytic signal prior. Hassan Mansour, Suhas Lohit, Petros Boufounos |
ICIP | 2 |
| 2022 | Learning Partial Equivariances From DataabstractGroup Convolutional Neural Networks (G-CNNs) constrain learned features to respect the symmetries in the selected group, and lead to better generalization when these symmetries appear in the data. If this is not the case, however, equivariance leads to overly constrained models and worse performance. Frequently, transformations occurring in data can be better represented by a subset of a group than by a group as a whole, e.g., rotations in $[-90^{\circ}, 90^{\circ}]$. In such cases, a model that respects equivariance partially is better suited to represent the data. In addition, relevant transformations may differ for low and high-level features. For instance, full rotation equivariance is useful to describe edge orientations in a face, but partial rotation equivariance is better suited to describe face poses relative to the camera. In other words, the optimal level of equivariance may differ per layer. In this work, we introduce Partial G-CNNs: G-CNNs able to learn layer-wise levels of partial and full equivariance to discrete, continuous groups and combinations thereof as part of training. Partial G-CNNs retain full equivariance when beneficial, e.g., for rotated MNIST, but adjust it whenever it becomes harmful, e.g., for classification of 6/9 digits or natural images. We empirically show that partial G-CNNs pair G-CNNs when full equivariance is advantageous, and outperform them otherwise. Our code is publicly available at www.github.com/merlresearch/partial_gcnn . David W. Romero, Suhas Lohit |
NeurIPS | 2 |
| 2022 | What Makes a "Good" Data Augmentation in Knowledge Distillation - A Statistical PerspectiveabstractKnowledge distillation (KD) is a general neural network training approach that uses a teacher model to guide the student model. Existing works mainly study KD from the network output side (e.g., trying to design a better KD loss function), while few have attempted to understand it from the input side. Especially, its interplay with data augmentation (DA) has not been well understood. In this paper, we ask: Why do some DA schemes (e.g., CutMix) inherently perform much better than others in KD? What makes a "good" DA in KD? Our investigation from a statistical perspective suggests that a good DA scheme should reduce the covariance of the teacher-student cross-entropy. A practical metric, the stddev of teacher’s mean probability (T. stddev), is further presented and well justified empirically. Besides the theoretical understanding, we also introduce a new entropy-based data-mixing DA scheme, CutMixPick, to further enhance CutMix. Extensive empirical studies support our claims and demonstrate how we can harvest considerable performance gains simply by using a better DA scheme in knowledge distillation. Code: https://github.com/MingSun-Tse/Good-DA-in-KD. Huan Wang 0014, Suhas Lohit, Michael J. Jones 0001, Yun Fu 0001 |
NeurIPS | 2 |
| 2022 | Model Compression Using Optimal TransportabstractModel compression methods are important to allow for easier deployment of deep learning models in compute, memory and energy-constrained environments such as mobile phones. Knowledge distillation is a class of model compression algorithms where knowledge from a large teacher network is transferred to a smaller student network thereby improving the student’s performance. In this paper, we show how optimal transport-based loss functions can be used for training a student network which encourages learning student network parameters that help bring the distribution of student features closer to that of the teacher features. We present image classification results on CIFAR-100, SVHN and ImageNet and show that the proposed optimal transport loss functions perform comparably to or better than other loss functions. Suhas Lohit, Michael J. Jones 0001 |
WACV | 1 |
| 2021 | Turnip: Time-Series U-Net With Recurrence For NIR Imaging PPGabstractImaging photoplethysmography (iPPG) is the process of estimating the waveform of a person’s pulse by processing a video of their face to detect minute color or intensity changes in the skin. Typically, iPPG methods use three-channel RGB video to address challenges due to motion. In situations such as driving, however, illumination in the visible spectrum is often quickly varying (e.g., daytime driving through shadows of trees and buildings) or insufficient (e.g., night driving). In such cases, a practical alternative is to use active illumination and bandpass-filtering from a monochromatic near-infrared (NIR) light source and camera. Contrary to learning-based iPPG solutions designed for multi-channel RGB, previous work in single-channel NIR iPPG has been based on hand-crafted models (with only a few manually tuned parameters), exploiting the sparsity of the PPG signal in the frequency domain. In contrast, we propose a modular framework for iPPG estimation of the heartbeat signal, in which the first module extracts a time-series signal from monochromatic NIR face video. The second module consists of a novel time-series U-net architecture in which a GRU (gated recurrent unit) network has been added to the passthrough layers. We test our approach on the challenging MR-NIRP Car Dataset, which consists of monochromatic NIR videos taken in both stationary and driving conditions. Our model’s iPPG estimation performance on NIR video outperforms both the state-of-the-art model-based method and a recent end-to-end deep learning method that we adapted to monochromatic video. Armand Comas, Tim K. Marks, Hassan Mansour, Suhas Lohit, Yechi Ma, Xiaoming Liu 0002 |
ICIP | 4 |
| 2021 | Generative Patch Priors for Practical Compressive Image RecoveryabstractIn this paper, we propose the generative patch prior (GPP) that defines a generative prior for compressive image recovery, based on patch-manifold models. Unlike learned, image-level priors that are restricted to the range space of a pre-trained generator, GPP can recover a wide variety of natural images using a pre-trained patch generator. Additionally, GPP retains the benefits of generative priors like high reconstruction quality at extremely low sensing rates, while also being much more generally applicable. We show that GPP outperforms several unsupervised and supervised techniques on three different sensing models - linear compressive sensing with known, and unknown calibration settings, and the non-linear phase retrieval problem. Finally, we propose an alternating optimization strategy using GPP for joint calibration-and-reconstruction which performs favorably against several baselines on a real world, uncalibrated compressive sensing dataset. The code and models for GPP are available on github.1. Rushil Anirudh, Suhas Lohit, Pavan Turaga |
WACV | 2 |
| 2021 | Recovering Trajectories of Unmarked Joints in 3D Human Actions Using Latent Space OptimizationabstractMotion capture (mocap) and time-of-flight based sensing of human actions are becoming increasingly popular modalities to perform robust activity analysis. Applications range from action recognition to quantifying movement quality for health applications. While marker-less motion capture has made great progress, in critical applications such as healthcare, marker-based systems, especially active markers, are still considered gold-standard. However, there are several practical challenges in both modalities such as visibility, tracking errors, and simply the need to keep marker setup convenient wherein movements are recorded with a reduced marker-set. This implies that certain joint locations will not even be marked-up, making downstream analysis of full body movement challenging. To address this gap, we first pose the problem of reconstructing the unmarked joint data as an ill-posed linear inverse problem. We recover missing joints for a given action by projecting it onto the manifold of human actions, this is achieved by optimizing the latent space representation of a deep autoencoder. Experiments on both mocap and Kinect datasets clearly demonstrate that the proposed method performs very well in recovering semantics of the actions and dynamics of missing joints. We will release all the code and models publicly. Suhas Lohit, Rushil Anirudh, Pavan Turaga |
WACV | 1 |
| 2020 | Rate-Invariant Autoencoding of Time-SeriesabstractFor time-series classification and retrieval applications, an important requirement is to develop representations/metrics that are robust to re-parametrization of the time-axis. Temporal re-parametrization as a model can account for variability in the underlying generative process, sampling rate variations, or plain temporal mis-alignment. In this paper, we extend prior work in disentangling latent spaces of autoencoding models, to design a novel architecture to learn rate-invariant latent codes in a completely unsupervised fashion. Unlike conventional neural network architectures, this method allows to explicitly disentangle temporal parameters in the form of order-preserving diffeomorphisms with respect to a learnable template. This makes the latent space more easily interpretable. We show the efficacy of our approach on a synthetic dataset and a real dataset for hand action-recognition. Kaushik Koneripalli, Suhas Lohit, Rushil Anirudh, Pavan Turaga |
ICASSP | 2 |
| 2019 | Temporal Transformer Networks: Joint Learning of Invariant and Discriminative Time WarpingabstractMany time-series classification problems involve developing metrics that are invariant to temporal misalignment. In human activity analysis, temporal misalignment arises due to various reasons including differing initial phase, sensor sampling rates, and elastic time-warps due to subject-specific biomechanics. Past work in this area has only looked at reducing intra-class variability by elastic temporal alignment. In this paper, we propose a hybrid model-based and data-driven approach to learn warping functions that not just reduce intra-class variability, but also increase inter-class separation. We call this a temporal transformer network (TTN). TTN is an interpretable differentiable module, which can be easily integrated at the front end of a classification network. The module is capable of reducing intra-class variance by generating input-dependent warping functions which lead to rate-robust representations. At the same time, it increases inter-class variance by learning warping functions that are more discriminative. We show improvements over strong baselines in 3D action recognition on challenging datasets using the proposed framework. The improvements are especially pronounced when training sets are smaller. Suhas Lohit, Pavan Turaga |
CVPR | 1 |
| 2019 | Unrolled Projected Gradient Descent for Multi-spectral Image FusionabstractIn this paper, we consider the problem of fusing low spatial resolution multi-spectral (MS) aerial images with their associated high spatial resolution panchromatic image. To solve this problem, various methods have been proposed, using either model-based or model-agnostic algorithms such as deep learning techniques. In this paper, we aim to utilize more interpretable architectures to solve the MS fusion problem by integrating existing ideas from image processing with deep learning. In particular, we develop a signal processing-inspired learning solution, where we unroll the iterations of the projected gradient descent (PGD) algorithm, and each iteration contains a projection operation carried out by a deep convolutional neural network. We observe that our proposed method provides a new perspective on existing deep-learning solutions, and under certain circumstance it reduces to current black-box deep learning methods. Our extensive experimental results show significant improvements of the proposed approach over several baselines. Suhas Lohit, Dehong Liu, Hassan Mansour, Petros Boufounos |
ICASSP | 1 |
| 2018 | CS-VQA: Visual Question Answering with Compressively Sensed ImagesabstractVisual Question Answering (VQA) is a complex semantic task requiring both natural language processing and visual recognition. In this paper, we explore whether VQA is solvable when images are captured in a sub-Nyquist compressive paradigm. We develop a series of deep-network architectures that exploit available compressive data to increasing degrees of accuracy, and show that VQA is indeed solvable in the compressed domain. Our results show that there is nominal degradation in VQA performance when using compressive measurements, but that accuracy can be recovered when VQA pipelines are used in conjunction with state-of-the-art deep neural networks for CS reconstruction. The results presented yield important implications for resource-constrained VQA applications. Li-Chi Huang, Kuldeep Kulkarni, Anik Jha, Suhas Lohit, Suren Jayasuriya, Pavan Turaga |
ICIP | 4 |
| 2016 | ReconNet: Non-Iterative Reconstruction of Images from Compressively Sensed MeasurementsabstractThe goal of this paper is to present a non-iterative and more importantly an extremely fast algorithm to reconstruct images from compressively sensed (CS) random measurements. To this end, we propose a novel convolutional neural network (CNN) architecture which takes in CS measurements of an image as input and outputs an intermediate reconstruction. We call this network, ReconNet. The intermediate reconstruction is fed into an off-the-shelf denoiser to obtain the final reconstructed image. On a standard dataset of images we show significant improvements in reconstruction results (both in terms of PSNR and time complexity) over state-of-the-art iterative CS reconstruction algorithms at various measurement rates. Further, through qualitative experiments on real data collected using our block single pixel camera (SPC), we show that our network is highly robust to sensor noise and can recover visually better quality images than competitive algorithms at extremely low sensing rates of 0.1 and 0.04. To demonstrate that our algorithm can recover semantically informative images even at a low measurement rate of 0.01, we present a very robust proof of concept real-time visual tracking application. Kuldeep Kulkarni, Suhas Lohit, Pavan Turaga, Ronan Kerviche, Amit Ashok |
CVPR | 2 |
| 2016 | Direct inference on compressive measurements using convolutional neural networksabstractCompressive imagers, e.g. the single-pixel camera (SPC), acquire measurements in the form of random projections of the scene instead of pixel intensities. Compressive Sensing (CS) theory allows accurate reconstruction of the image even from a small number of such projections. However, in practice, most reconstruction algorithms perform poorly at low measurement rates and are computationally very expensive. But perfect reconstruction is not the goal of high-level computer vision applications. Instead, we are interested in only determining certain properties of the image. Recent work has shown that effective inference is possible directly from the compressive measurements, without reconstruction, using correlational features. In this paper, we show that convolutional neural networks (CNNs) can be employed to extract discriminative non-linear features directly from CS measurements. Using these features, we demonstrate that effective high-level inference can be performed. Experimentally, using hand written digit recognition (MNIST dataset) and image recognition (ImageNet) as examples, we show that recognition is possible even at low measurement rates of about 0.1. Suhas Lohit, Kuldeep Kulkarni, Pavan Turaga |
ICIP | 1 |