Vishwanath Saragadam

dblp:172/1229 · DBLP profile ↗
← Back
20ranked-venue papers
12as first author
13since 2021 · last 2025
0000-0001-8028-7520ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 10 first-author · 8 since 2021Artificial intelligence and machine learning · 11 · 5 first-author · 10 since 2021
YearPublicationVenuePosition
2025 Bias for Action: Video Implicit Neural Representations with Bias Modulation
abstract
We propose a new continuous video modeling framework based on implicit neural representations (INRs) called ActINR. At the core of our approach is the observation that INRs can be considered as a learnable dictionary, with the shapes of the basis functions governed by the weights of the INR, and their locations governed by the biases. Given compact non-linear activation functions, we hypothesize that an INR’s biases are suitable to capture motion across images, and facilitate compact representations for video sequences. Using these observations, we design ActINR to share INR weights across frames of a video sequence, while using unique biases for each frame. We further model the biases as the output of a separate INR conditioned on time index to promote smoothness. By training the video INR and this bias INR together, we demonstrate unique capabilities, including 10× video slow motion, 4× spatial super resolution along with 2× slow motion, denoising, and video inpainting. ActINR performs remarkably well across numerous video processing tasks (often achieving more than 6dB improvement), setting a new standard for continuous modeling of videos.
Alper Kayabasi, Anil Kumar Vadathya, Guha Balakrishnan, Vishwanath Saragadam
CVPR4
2025 PS$^{2}$2 F: Polarized Spiral Point Spread Function for Single-Shot 3D Sensing
abstract
We propose a compact snapshot monocular depth estimation technique that relies on an engineered point spread function (PSF). Traditional approaches used in microscopic super-resolution imaging such as the Double-Helix PSF (DHPSF) are ill-suited for scenes that are more complex than a sparse set of point light sources. We show, using the Cramér-Rao lower bound, that separating the two lobes of the DHPSF and thereby capturing two separate images leads to a dramatic increase in depth accuracy. A special property of the phase mask used for generating the DHPSF is that a separation of the phase mask into two halves leads to a spatial separation of the two lobes. We leverage this property to build a compact polarization-based optical setup, where we place two orthogonal linear polarizers on each half of the DHPSF phase mask and then capture the resulting image with a polarization-sensitive camera. Results from simulations and a lab prototype demonstrate that our technique achieves up to $50\%$50% lower depth error compared to state-of-the-art designs including the DHPSF and the Tetrapod PSF, with little to no loss in spatial resolution.
Bhargav Ghanekar, Vishwanath Saragadam, Dushyant Mehra, Anna-Karin Gustavsson, Aswin C. Sankaranarayanan, Ashok Veeraraghavan
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Snapshot Lidar: Fourier Embedding of Amplitude and Phase for Single-Image Depth Reconstruction
abstract
Amplitude modulated continuous-wave time-of-flight (AMCW-ToF) cameras are finding applications as flash Lidars in autonomous navigation, robotics, and AR/VR applications. A conventional CW-ToF camera requires illuminating the scene with a temporally varying light source and demodulating a set of quadrature measurements to recover the scene's depth and intensity. Capturing the four measurements in sequence renders the system slow, invariably causing inaccuracies in depth estimates due to motion in the scene or the camera. To mitigate this problem, we propose a snapshot Lidar that captures amplitude and phase simultaneously as a single time-of-flight hologram. Uniquely, our approach requires minimal changes to existing CW- ToF imaging hardware. To demonstrate the efficacy of the proposed system, we design and build a lab prototype, and evaluate it under varying scene geometries, illumination conditions, and compare the reconstructed depth measurements against conventional techniques. We rigorously evaluate the robustness of our system on diverse real-world scenes to show that our technique results in a significant reduction in data bandwidth with minimal loss in reconstruction accuracy. As high-resolution CW-ToF cameras are becoming ubiquitous, increasing their temporal resolution by four times enables robust real-time capture of geometries of dynamic scenes.
Sarah Friday, Yunzi Shi, Yaswanth Cherivirala, Vishwanath Saragadam, Adithya Kumar Pediredla
CVPR4
2024 Titan: Bringing the Deep Image Prior to Implicit Representations
abstract
We study the interpolation capabilities of implicit neural representations (INRs) of images. In principle, INRs promise a number of advantages, such as continuous derivatives and arbitrary sampling, being freed from the restrictions of a raster grid. However, empirically, INRs have been observed to poorly interpolate between the pixels of the fit image; in other words, they do not inherently possess a suitable prior for natural images. In this paper, we propose to address and improve INRs’ interpolation capabilities by explicitly integrating image prior information into the INR architecture via deep decoder, a specific implementation of the deep image prior (DIP). Our method, which we call TITAN, leverages a residual connection from the input which enables integrating the principles of the grid-based DIP into the grid-free INR. Through super-resolution and computed tomography experiments, we demonstrate that our method significantly improves upon classic INRs, thanks to the induced natural image bias. We also find that by constraining the weights to be sparse, image quality and sharpness are enhanced, increasing the Lipschitz constant.
Lorenzo Luzi, Daniel LeJeune, Ali Siahkoohi, Sina Alemohammad, Vishwanath Saragadam, Hossein Babaei, Naiming Liu, Zichao Wang 0001, Richard G. Baraniuk
ICASSP5
2024 Implicit Neural Representations and the Algebra of Complex Wavelets
abstract
Implicit neural representations (INRs) have arisen as useful methods for representing signals on Euclidean domains. By parameterizing an image as a multilayer perceptron (MLP) on Euclidean space, INRs effectively couple spatial and spectral features of the represented signal in a way that is not obvious in the usual discrete representation. Although INRs using sinusoidal activation functions have been studied in terms of Fourier theory, recent works have shown the advantage of using wavelets instead of sinusoids as activation functions, due to their ability to simultaneously localize in both frequency and space. In this work, we approach such INRs and demonstrate how they resolve high-frequency features of signals from coarse approximations performed in the first layer of the MLP. This leads to multiple prescriptions for the design of INR architectures, including the use of progressive wavelets, decoupling of low and high-pass approximations, and initialization schemes based on the singularities of the target signal.
T. Mitchell Roddenberry, Vishwanath Saragadam, Maarten V. de Hoop, Richard G. Baraniuk
ICLR2
2024 DeepTensor: Low-Rank Tensor Decomposition With Deep Network Priors
abstract
DeepTensor is a computationally efficient framework for low-rank decomposition of matrices and tensors using deep generative networks. We decompose a tensor as the product of low-rank tensor factors (e.g., a matrix as the outer product of two vectors), where each low-rank tensor is generated by a deep network (DN) that is trained in a self-supervised manner to minimize the mean-square approximation error. Our key observation is that the implicit regularization inherent in DNs enables them to capture nonlinear signal structures (e.g., manifolds) that are out of the reach of classical linear methods like the singular value decomposition (SVD) and principal components analysis (PCA). Furthermore, in contrast to the SVD and PCA, whose performance deteriorates when the tensor's entries deviate from additive white Gaussian noise, we demonstrate that the performance of DeepTensor is robust to a wide range of distributions. We validate that DeepTensor is a robust and computationally efficient drop-in replacement for the SVD, PCA, nonnegative matrix factorization (NMF), and similar decompositions by exploring a range of real-world applications, including hyperspectral image denoising, 3D MRI tomography, and image classification. In particular, DeepTensor offers a 6 dB signal-to-noise ratio improvement over standard denoising methods for signal corrupted by Poisson noise and learns to decompose 3D tensors 60 times faster than a single DN equipped with 3D convolutions.
Vishwanath Saragadam, Randall Balestriero, Ashok Veeraraghavan, Richard G. Baraniuk
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Thermal Spread Functions (TSF): Physics-Guided Material Classification
abstract
Robust and non-destructive material classification is a challenging but crucial first-step in numerous vision applications. We propose a physics-guided material classification framework that relies on thermal properties of the object. Our key observation is that the rate of heating and cooling of an object depends on the unique intrinsic properties of the material, namely the emissivity and diffusivity. We leverage this observation by gently heating the objects in the scene with a low-power laser for a fixed duration and then turning it off, while a thermal camera captures measurements during the heating and cooling process. We then take this spatial and temporal “thermal spread function” (TSF) to solve an inverse heat equation using the finite-differences approach, resulting in a spatially varying estimate of diffusivity and emissivity. These tuples are then used to train a classifier that produces a fine-grained material label at each spatial pixel. Our approach is extremely simple requiring only a small light source (low power laser) and a thermal camera, and produces robust classification results with 86% accuracy over 16 classes11Code: https://github.com/aniketdashpute/TSF.
Aniket Dashpute, Vishwanath Saragadam, Emma Alexander, Florian Willomitzer, Aggelos K. Katsaggelos, Ashok Veeraraghavan, Oliver Cossairt
CVPR2
2023 WIRE: Wavelet Implicit Neural Representations
abstract
Implicit neural representations (INRs) have recently advanced numerous vision-related areas. INR performance depends strongly on the choice of activation function employed in its MLP network. A wide range of nonlinearities have been explored, but, unfortunately, current INRs designed to have high accuracy also suffer from poor robustness (to signal noise, parameter variation, etc.). Inspired by harmonic analysis, we develop a new, highly accurate and robust INR that does not exhibit this trade off. Our Wavelet Implicit neural REpresentation (WIRE) uses as its activation function the complex Gabor wavelet that is well-known to be optimally concentrated in space-frequency and to have excellent biases for representing images. A wide range of experiments (image denoising, image inpainting, super-resolution, computed tomography reconstruction, image over fitting, and novel view synthesis with neural radiance fields) demonstrate that WIRE defines the new state of the art in INR accuracy, training time, and robustness.
Vishwanath Saragadam, Daniel LeJeune, Jasper Tan, Guha Balakrishnan, Ashok Veeraraghavan, Richard G. Baraniuk
CVPR1
2023 Programmable Spectral Filter Arrays using Phase Spatial Light Modulators
abstract
Computational imaging has always benefited from tools that modulate light along the many dimensions of its plenoptic function. This paper provides a practical architecture for achieving spatially varying spectral modulation using a liquid crystal phase spatial light modulator (SLM). The use of a phase SLM, however, results in strong optical aberrations due to the unintended phase modulation, thereby precluding spectral modulation at high spatial resolutions. To mitigate this, we provide a careful and systematic analysis of the aberrations arising out of phase SLMs for the purpose of spatially varying spectral modulation; this analysis results in a dual strategy of “good patterns” that minimize the optical aberrations and a deep restoration network that overcomes any residual aberrations. We show a number of unique operating points with our prototype including single- and multi-image hyperspectral imaging, material classification (fewer than two images), and dynamic spectral filtering at video rates.
Vishwanath Saragadam, Vijay Rengarajan, Ryuichi Tadano, Tuo Zhuang, Hideki Oyaizu, Jun Murayama, Aswin C. Sankaranarayanan
ICCP1
2022 MINER: Multiscale Implicit Neural Representation
Vishwanath Saragadam, Jasper Tan, Guha Balakrishnan, Richard G. Baraniuk, Ashok Veeraraghavan
ECCV (23)1
2022 Improving Transformer with an Admixture of Attention Heads
abstract
Transformers with multi-head self-attention have achieved remarkable success in sequence modeling and beyond. However, they suffer from high computational and memory complexities for computing the attention matrix at each head. Recently, it has been shown that those attention matrices lie on a low-dimensional manifold and, thus, are redundant. We propose the Transformer with a Finite Admixture of Shared Heads (FiSHformers), a novel class of efficient and flexible transformers that allow the sharing of attention matrices between attention heads. At the core of FiSHformer is a novel finite admixture model of shared heads (FiSH) that samples attention matrices from a set of global attention matrices. The number of global attention matrices is much smaller than the number of local attention matrices generated. FiSHformers directly learn these global attention matrices rather than the local ones as in other transformers, thus significantly improving the computational and memory efficiency of the model. We empirically verify the advantages of the FiSHformer over the baseline transformers in a wide range of practical applications including language modeling, machine translation, and image classification. On the WikiText-103, IWSLT'14 De-En and WMT'14 En-De, FiSHformers use much fewer floating-point operations per second (FLOPs), memory, and parameters compared to the baseline transformers.
Hai Do, Vishwanath Saragadam, Minh Pham 0003, Duy Khuong Nguyen, Nhat Ho, Stanley J. Osher
NeurIPS5
2021 SliceNets - A Scalable Approach for Object Detection in 3D CT Scans
abstract
One of the most promising approaches for automated detection of guns and other prohibited items in aviation baggage screening is the use of 3D computed tomography (CT) scans. However, automated detection, especially with deep neural networks, faces two key challenges: the high dimensionality of individual 3D scans, and the lack of labelled training data. We address these challenges using a novel image-based detection and segmentation technique that we call the slice-and-fuse framework. Our approach relies on slicing the input 3D volumes, generating 2D predictions on each slice using 2D Convolutional Neural Networks (CNNs), and fusing them to obtain a 3D prediction. We develop two distinct detectors based on this slice-and-fuse strategy: the Retinal-SliceNet that uses a unified, single network with end-to-end training, and the U-SliceNet that uses a two-stage paradigm, first generating proposals using a voxel labeling network and, subsequently, refining the proposals by a 3D classification network. The networks are trained using a data augmentation approach that creates a very large training dataset by inserting weapons into 3D CT scans of threat-free bags. We demonstrate that the two SliceNets outperform state-of-the-art methods on a large-scale 3D baggage CT dataset for baggage classification, 3D object detection, and 3D semantic segmentation.
Anqi Yang, Vishwanath Saragadam, Duy Dao, Zhuo Hui, Jen-Hao Rick Chang, Aswin C. Sankaranarayanan
WACV3
2021 SASSI - Super-Pixelated Adaptive Spatio-Spectral Imaging
abstract
We introduce a novel video-rate hyperspectral imager with high spatial, temporal and spectral resolutions. Our key hypothesis is that spectral profiles of pixels within each super-pixel tend to be similar. Hence, a scene-adaptive spatial sampling of a hyperspectral scene, guided by its super-pixel segmented image, is capable of obtaining high-quality reconstructions. To achieve this, we acquire an RGB image of the scene, compute its super-pixels, from which we generate a spatial mask of locations where we measure high-resolution spectrum. The hyperspectral image is subsequently estimated by fusing the RGB image and the spectral measurements using a learnable guided filtering approach. Due to low computational complexity of the superpixel estimation step, our setup can capture hyperspectral images of the scenes with little overhead over traditional snapshot hyperspectral cameras, but with significantly higher spatial and spectral resolutions. We validate the proposed technique with extensive simulations as well as a lab prototype that measures hyperspectral video at a spatial resolution of 600 ×900 pixels, at a spectral resolution of 10 nm over visible wavebands, and achieving a frame rate at 18fps.
Vishwanath Saragadam, Michael DeZeeuw, Richard G. Baraniuk, Ashok Veeraraghavan, Aswin C. Sankaranarayanan
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Programmable Spectrometry: Per-pixel Material Classification using Learned Spectral Filters
abstract
Many materials have distinct spectral profiles, which facilitates estimation of the material composition of a scene by processing its hyperspectral image (HSI). However, this process is inherently wasteful since high-dimensional HSIs are expensive to acquire and only a set of linear projections of the HSI contribute to the classification task. This paper proposes the concept of programmable spectrometry for per-pixel material classification, where instead of sensing the HSI of the scene and then processing it, we optically compute the spectrally-filtered images. This is achieved using a computational camera with a programmable spectral response. Our approach provides gains both in terms of acquisition speed - since only the relevant measurements are acquired - and in signal-to-noise ratio - since we invariably avoid narrowband filters that are light inefficient. Given ample training data, we use learning techniques to identify the bank of spectral profiles that facilitate material classification. We verify the method in simulations, as well as validate our findings using a lab prototype of the camera.
Vishwanath Saragadam, Aswin C. Sankaranarayanan
ICCP1
2019 Wavelet Tree Parsing with Freeform Lensing
abstract
We propose an architecture for adaptive sensing of images by progressively measuring its wavelet coefficients. Our approach, commonly referred to as wavelet tree parsing, adaptively selects the specific wavelet coefficients to be sensed by modeling the children of dominant coefficients to be dominant themselves. A key challenge for practical implementation of this technique is that the wavelet patterns, especially at finer scales, occupy a tiny portion of the field of view and, hence, the resulting measurements have very poor light levels and signal-to-noise ratios (SNR). To address this, we propose a novel imaging architecture that uses a phase-only spatial light modulator as a freeform lens to concentrate a light source and create the wavelet patterns. This ensures that the SNR of measurements remain constant across different spatial scales. Using a lab prototype, we demonstrate successful reconstruction on a wide range of real scenes and show that concentrating illumination enables us to outperform non-adaptive techniques as well as adaptive techniques based on traditional projectors.
Vishwanath Saragadam, Aswin C. Sankaranarayanan
ICCP1
2019 Micro-Baseline Structured Light
abstract
We propose Micro-baseline Structured Light (MSL), a novel 3D imaging approach designed for small form-factor devices such as cell-phones and miniature robots. MSL operates with small projector-camera baseline and low-cost projection hardware, and can recover scene depths with computationally lightweight algorithms. The main observation is that a small baseline leads to small disparities, enabling a first-order approximation of the non-linear SL image formation model. This leads to the key theoretical result of the paper: the MSL equation, a linearized version of SL image formation. MSL equation is under-constrained due to two unknowns (depth and albedo) at each pixel, but can be efficiently solved using a local least squares approach. We analyze the performance of MSL in terms of various system parameters such as projected pattern and baseline, and provide guidelines for optimizing performance. Armed with these insights, we build a prototype to experimentally examine the theory and its practicality.
Vishwanath Saragadam, Raja Venkata, Jian Wang 0100, Shree K. Nayar, Mohit Gupta 0001
ICCV1
2019 Cross-Scale Predictive Dictionaries
abstract
Sparse representations using data dictionaries provide an efficient model particularly for signals that do not enjoy alternate analytic sparsifying transformations. However, solving inverse problems with sparsifying dictionaries can be computationally expensive, especially when the dictionary under consideration has a large number of atoms. In this paper, we incorporate additional structure on to dictionary-based sparse representations for visual signals to enable speedups when solving sparse approximation problems. The specific structure that we endow onto sparse models is that of a multi-scale modeling where the sparse representation at each scale is constrained by the sparse representation at coarser scales. We show that this cross-scale predictive model delivers significant speedups, often in the range of , with little loss in accuracy for linear inverse problems associated with images, videos, and light fields.
Vishwanath Saragadam, Xin Li 0001, Aswin C. Sankaranarayanan
IEEE Trans. Image Process.1
2019 KRISM - Krylov Subspace-based Optical Computing of Hyperspectral Images
abstract
We present an adaptive imaging technique that optically computes a low-rank approximation of a scene’s hyperspectral image, conceptualized as a matrix. Central to the proposed technique is the optical implementation of two measurement operators: a spectrally coded imager and a spatially coded spectrometer. By iterating between the two operators, we show that the top singular vectors and singular values of a hyperspectral image can be adaptively and optically computed with only a few iterations. We present an optical design that uses pupil plane coding for implementing the two operations and show several compelling results using a lab prototype to demonstrate the effectiveness of the proposed hyperspectral imager.
Vishwanath Saragadam, Aswin C. Sankaranarayanan
ACM Trans. Graph.1
2017 Compressive spectral anomaly detection
abstract
We propose a novel compressive imager for detecting anomalous spectral profiles in a scene. We model the background spectrum as a low-dimensional subspace while assuming the anomalies to form a spatially-sparse set of spectral profiles different from the background. Our core contributions are in the form of a two-stage sensing mechanism. In the first stage, we estimate the subspace for the background spectrum by acquiring spectral measurements at a few randomly-selected pixels. In the second stage, we acquire spatially-multiplexed spectral measurements of the scene. We remove the contributions of the background spectrum from the spatially-multiplexed measurements by projecting onto the complementary subspace of the background spectrum; the resulting measurements are of a sparse matrix that encodes the presence and spectra of anomalies, which can be recovered using a Multiple Measurement Vector formulation. Theoretical analysis and simulations show significant speed up in acquisition time over other anomaly detection techniques. A lab prototype based on a DMD and a visible spectrometer validates our proposed imager.
Vishwanath Saragadam, Jian Wang 0100, Xin Li 0001, Aswin C. Sankaranarayanan
ICCP1
2016 Cross-scale predictive dictionaries for image and video restoration
abstract
We propose a novel signal model, based on sparse representations, that captures cross-scale features for visual signals. We show that cross-scale predictive model enables faster solutions to sparse approximation problems. This is achieved by first solving the sparse approximation problem for the downsampled signal and using the support of the solution to constrain the support at the original resolution. The speedups obtained are especially compelling for high-dimensional signals that require large dictionaries to provide precise sparse approximations. We demonstrate speedups in the order of 10-20x for denoising and up to 9x speed-ups for compressive sensing of images and videos.
Vishwanath Saragadam, Aswin C. Sankaranarayanan, Xin Li 0001
ICIP1