Saiprasad Ravishankar

dblp:46/1532 · DBLP profile ↗
← Back
43ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0002-5792-5827ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Robust Physics-Based Deep MRI Reconstruction via Diffusion Purification
abstract
Deep learning (DL) supervised techniques have been extensively employed in magnetic resonance imaging (MRI) reconstruction, delivering notable performance enhancements over traditional non-DL methods. Nonetheless, these models have vulnerabilities during testing such as their susceptibility to worst-case or noise-based measurement perturbations, variations in training/testing settings like acceleration factors, contrast, $k$ -space sampling locations, and distribution shifts stemming from unseen lesions and different anatomies. This article addresses these robustness challenges by leveraging diffusion models (DMs). In particular, we present a robustification strategy that improves the resilience of DL-based MRI reconstruction methods by utilizing pretrained DMs as purifiers. We dub our method as robust DL-based MRI with diffusion purification (RODIO). In contrast to conventional robustification methods for DL-based MRI reconstruction, such as adversarial training (AT), our proposed approach eliminates the need to tackle a minimax optimization problem. It only necessitates efficient fine-tuning on purified examples. Our experimental results underscore the effectiveness of our approach in addressing the mentioned instabilities, outperforming standalone diffusion-based MRI reconstructors and leading robustification methods for deep supervised MRI reconstruction, including AT and randomized smoothing (RS). Our experiments demonstrate: 1) the adaptability of our approach across multiple DL-based supervised MRI reconstruction models; 2) compatibility with accelerated diffusion-based samplers; 3) robustness to data with unseen lesions; and 4) effectiveness when applied to unsupervised single-shot generative reconstructors.
Ismail Alkhouri, Shijun Liang 0001, Qing Qu 0001, Saiprasad Ravishankar
IEEE Trans. Neural Networks Learn. Syst.5
2025 Sequential Diffusion-Guided Deep Image Prior for Medical Image Reconstruction
abstract
Deep learning (DL) methods have been extensively applied to various image recovery problems, including magnetic resonance imaging (MRI) and computed tomography (CT) reconstruction. Beyond supervised models, other approaches have been recently explored including two key recent schemes: deep image prior (DIP) that is an unsupervised scan-adaptive method that leverages the network architecture as implicit regularization but can suffer from noise over-fitting, and diffusion models (DMs), where the sampling procedure of a pre-trained generative model is modified to allow sampling from the measurement-conditioned distribution through approximations. In this paper, we propose combining DIP and DMs for MRI and CT reconstruction, motivated by (i) the impact of the DIP network input and (ii) the use of DMs as diffusion purifiers (DPs). Specifically, we propose a sequential procedure that iteratively optimizes the DIP network with a DM-refined adaptive input using a loss with data consistency and autoencoding terms. We term the approach Sequential Diffusion-Guided DIP (uDiG-DIP). Our experimental results demonstrate that uDiG-DIP achieves superior reconstruction results compared to leading DM-based baselines and the original DIP for MRI and CT tasks.
Shijun Liang 0001, Ismail Alkhouri, Qing Qu 0001, Saiprasad Ravishankar
ICASSP5
2025 Learning Dynamics of Deep Matrix Factorization Beyond the Edge of Stability
abstract
Deep neural networks trained using gradient descent with a fixed learning rate $\eta$ often operate in the regime of ``edge of stability'' (EOS), where the largest eigenvalue of the Hessian equilibrates about the stability threshold $2/\eta$. In this work, we present a fine-grained analysis of the learning dynamics of (deep) linear networks (DLNs) within the deep matrix factorization loss beyond EOS. For DLNs, loss oscillations beyond EOS follow a period-doubling route to chaos. We theoretically analyze the regime of the 2-period orbit and show that the loss oscillations occur within a small subspace, with the dimension of the subspace precisely characterized by the learning rate. The crux of our analysis lies in showing that the symmetry-induced conservation law for gradient flow, defined as the balancing gap among the singular values across layers, breaks at EOS and decays monotonically to zero. Overall, our results contribute to explaining two key phenomena in deep networks: (i) shallow models and simple tasks do not always exhibit EOS; and (ii) oscillations occur within top features}. We present experiments to support our theory, along with examples demonstrating how these phenomena occur in nonlinear networks and how they differ from those which have benign landscape such as in DLNs.
Avrajit Ghosh, Soo Min Kwon, Saiprasad Ravishankar, Qing Qu 0001
ICLR4
2025 SITCOM: Step-wise Triple-Consistent Diffusion Sampling For Inverse Problems
abstract
Diffusion models (DMs) are a class of generative models that allow sampling from a distribution learned over a training set. When applied to solving inverse problems, the reverse sampling steps are modified to approximately sample from a measurement-conditioned distribution. However, these modifications may be unsuitable for certain settings (e.g., presence of measurement noise) and non-linear tasks, as they often struggle to correct errors from earlier steps and generally require a large number of optimization and/or sampling steps. To address these challenges, we state three conditions for achieving measurement-consistent diffusion trajectories. Building on these conditions, we propose a new optimization-based sampling method that not only enforces standard data manifold measurement consistency and forward diffusion consistency, as seen in previous studies, but also incorporates our proposed step-wise and network-regularized backward diffusion consistency that maintains a diffusion trajectory by optimizing over the input of the pre-trained model at every sampling step. By enforcing these conditions (implicitly or explicitly), our sampler requires significantly fewer reverse steps. Therefore, we refer to our method as **S**tep-w**i**se **T**riple-**Co**nsistent Sa**m**pling (**SITCOM**). Compared to SOTA baselines, our experiments across several linear and non-linear tasks (with natural and medical images) demonstrate that SITCOM achieves competitive or superior results in terms of standard similarity metrics and run-time.
Ismail Alkhouri, Shijun Liang 0001, Cheng-Han Huang, Jimmy Dai, Qing Qu 0001, Saiprasad Ravishankar
ICML6
2025 Variational Learning Finds Flatter Solutions at the Edge of Stability
abstract
Variational Learning (VL) has recently gained popularity for training deep neural networks. Part of its empirical success can be explained by theories such as PAC-Bayes bounds, minimum description length and marginal likelihood, but little has been done to unravel the implicit regularization in play. Here, we analyze the implicit regularization of VL through the Edge of Stability (EoS) framework. EoS has previously been used to show that gradient descent can find flat solutions and we extend this result to show that VL can find even flatter solutions. This result is obtained by controlling the shape of the variational posterior as well as the number of posterior samples used during training. The derivation follows in a similar fashion as in the standard EoS literature for deep learning, by first deriving a result for a quadratic problem and then extending it to deep neural networks. We empirically validate these findings on a wide variety of large networks, such as ResNet and ViT, to find that the theoretical results closely match the empirical ones. Ours is the first work to analyze the EoS dynamics of~VL.
Avrajit Ghosh, Bai Cong, Rio Yokota, Saiprasad Ravishankar, Molei Tao, Mohammad Emtiyaz Khan, Thomas Möllenhoff
NeurIPS4
2025 UGoDIT: Unsupervised Group Deep Image Prior Via Transferable Weights
abstract
Recent advances in data-centric deep generative models have led to significant progress in solving inverse imaging problems. However, these models (e.g., diffusion models (DMs)) typically require large amounts of fully sampled (clean) training data, which is often impractical in medical and scientific settings such as dynamic imaging. On the other hand, training-data-free approaches like the Deep Image Prior (DIP) do not require clean ground-truth images but suffer from noise overfitting and can be computationally expensive as the network parameters need to be optimized for each measurement vector independently. Moreover, DIP-based methods often overlook the potential of learning a prior using a small number of sub-sampled measurements (or degraded images) available during training. In this paper, we propose **UGoDIT**—an **U**nsupervised **G**r**o**up **DI**P with **T**ransferable weights—designed for the low-data regime where only a very small number, $M$, of sub-sampled measurement vectors are available during training. Our method learns a set of transferable weights by optimizing a shared encoder and $M$ disentangled decoders. At test time, we reconstruct the unseen degraded image using a DIP network, where part of the parameters are fixed to the learned weights, while the remaining are optimized to enforce measurement consistency. We evaluate \our on both medical (multi-coil MRI) and natural (super resolution and non-linear deblurring) image recovery tasks under various settings. Compared to recent standalone DIP methods, \our provides accelerated convergence and notable improvement in reconstruction quality. Furthermore, our method achieves performance competitive with SOTA DM-based and supervised approaches, despite not requiring large amounts of clean training data. Our code is available at: https://github.com/sjames40/UGoDIT.
Shijun Liang 0001, Ismail Alkhouri, Siddhant Gautam, Qing Qu 0001, Saiprasad Ravishankar
NeurIPS5
2024 Improving Training Efficiency of Diffusion Models via Multi-Stage Framework and Tailored Multi-Decoder Architecture
abstract
Diffusion models, emerging as powerful deep generative tools, excel in various applications. They operate through a two-steps process: introducing noise into training samples and then employing a model to convert random noise into new samples (e.g., images). However, their remarkable generative performance is hindered by slow training and sampling. This is due to the necessity of tracking extensive forward and reverse diffusion trajectories, and employing a large model with numerous parameters across multiple timesteps (i.e., noise levels). To tackle these challenges, we present a multi-stage framework inspired by our empirical findings. These observations indicate the advantages of employing distinct parameters tailored to each timestep while retaining universal parameters shared across all time steps. Our approach involves segmenting the time interval into multiple stages where we employ custom multi-decoder U-net architecture that blends time-dependent models with a universally shared encoder. Our framework enables the efficient distribution of computational resources and mitigates inter-stage interference, which substantially improves training efficiency. Extensive numerical experiments affirm the effectiveness of our framework, showcasing significant training and sampling efficiency enhancements on three state-of-the-art diffusion models, including large-scale latent diffusion models. Furthermore, our ablation studies illustrate the impact of two important components in our framework: (i) a novel timestep clustering algorithm for stage division, and (ii) an innovative multi-decoder U-net architecture, seamlessly integrating universal and customized hyperparameters.
Yifu Lu, Ismail Alkhouri, Saiprasad Ravishankar, Dogyoon Song, Qing Qu 0001
CVPR4
2024 Diffusion-Based Adversarial Purification for Robust Deep Mri Reconstruction
abstract
Deep learning (DL) methods have been extensively employed in magnetic resonance imaging (MRI) reconstruction, demonstrating remarkable performance improvements compared to traditional non-DL methods. However, recent studies have uncovered the susceptibility of these models to carefully engineered adversarial perturbations. In this paper, we tackle this issue by leveraging diffusion models. Specifically, we introduce a defense strategy that enhances the robustness of DL-based MRI reconstruction methods through the utilization of pre-trained diffusion models as adversarial purifiers. Unlike conventional state-of-the-art adversarial defense methods (e.g., adversarial training), our proposed approach eliminates the need to solve a minimax optimization problem to train the image reconstruction model from scratch, and only requires fine-tuning on purified adversarial examples. Our experimental findings underscore the effectiveness of our proposed technique when benchmarked against leading defense methodologies for MRI reconstruction such as adversarial training and randomized smoothing.
Ismail Alkhouri, Shijun Liang 0001, Qing Qu 0001, Saiprasad Ravishankar
ICASSP5
2024 Patient-Adaptive and Learned Mri Data Undersampling Using Neighborhood Clustering
abstract
There has been much recent interest in adapting undersampled trajectories in MRI based on training data. In this work, we propose a novel patient-adaptive MRI sampling algorithm based on grouping scans within a training set. Scan-adaptive sampling patterns are optimized together with an image reconstruction network for the training scans. The training optimization alternates between determining the best sampling pattern for each scan (based on a greedy search or iterative coordinate descent (ICD)) and training a reconstructor across the dataset. The eventual scan-adaptive sampling patterns on the training set are used as labels to predict sampling design using nearest neighbor search at test time. The proposed algorithm is applied to the fastMRI knee multicoil dataset and demonstrates improved performance over several baselines.
Siddhant Gautam, Angqi Li, Saiprasad Ravishankar
ICASSP3
2024 Optimal Eye Surgeon: Finding image priors through sparse generators at initialization
abstract
We introduce Optimal Eye Surgeon (OES), a framework for pruning and training deep image generator networks. Typically, untrained deep convolutional networks, which include image sampling operations, serve as effective image priors. However, they tend to overfit to noise in image restoration tasks due to being overparameterized. OES addresses this by adaptively pruning networks at random initialization to a level of underparameterization. This process effectively captures low-frequency image components even without training, by just masking. When trained to fit noisy image, these pruned subnetworks, which we term Sparse-DIP, resist overfitting to noise. This benefit arises from underparameterization and the regularization effect of masking, constraining them in the manifold of image priors. We demonstrate that subnetworks pruned through OES surpass other leading pruning methods, such as the Lottery Ticket Hypothesis, which is known to be suboptimal for image recovery tasks. Our extensive experiments demonstrate the transferability of OES-masks and the characteristics of sparse-subnetworks for image generation. Code is available at https://github.com/Avra98/Optimal-Eye-Surgeon.
Avrajit Ghosh, Xitong Zhang, Kenneth K. Sun, Qing Qu 0001, Saiprasad Ravishankar
ICML5
2024 Image Reconstruction Via Autoencoding Sequential Deep Image Prior
abstract
Recently, Deep Image Prior (DIP) has emerged as an effective unsupervised one-shot learner, delivering competitive results across various image recovery problems. This method only requires the noisy measurements and a forward operator, relying solely on deep networks initialized with random noise to learn and restore the structure of the data. However, DIP is notorious for its vulnerability to overfitting due to the overparameterization of the network. Building upon insights into the impact of the DIP input and drawing inspiration from the gradual denoising process in cutting-edge diffusion models, we introduce Autoencoding Sequential DIP (aSeqDIP) for image reconstruction. This method progressively denoises and reconstructs the image through a sequential optimization of network weights. This is achieved using an input-adaptive DIP objective, combined with an autoencoding regularization term. Compared to diffusion models, our method does not require training data and outperforms other DIP-based methods in mitigating noise overfitting while maintaining a similar number of parameter updates as Vanilla DIP. Through extensive experiments, we validate the effectiveness of our method in various image reconstruction tasks, such as MRI and CT reconstruction, as well as in image restoration tasks like image denoising, inpainting, and non-linear deblurring.
Ismail Alkhouri, Shijun Liang 0001, Evan Bell, Qing Qu 0001, Saiprasad Ravishankar
NeurIPS6
2024 Learning Sparsity-Promoting Regularizers Using Bilevel Optimization
abstract
Abstract. We present a gradient-based heuristic method for supervised learning of sparsity-promoting regularizers for denoising signals and images. Sparsity-promoting regularization is a key ingredient in solving modern signal reconstruction problems; however, the operators underlying these regularizers are usually either designed by hand or learned from data in an unsupervised way. The recent success of supervised learning (e.g., with convolutional neural networks) in solving image reconstruction problems suggests that it could be a fruitful approach to designing regularizers. Towards this end, we propose to denoise signals using a variational formulation with a parametric, sparsity-promoting regularizer, where the parameters of the regularizer are learned to minimize the mean squared error of reconstructions on a training set of ground truth image and measurement pairs. Training involves solving a challenging bilevel optimization problem; we derive an expression for the gradient of the training loss using the closed-form solution of the denoising problem and provide an accompanying gradient descent algorithm to minimize it. Our experiments with structured 1D signals and natural images indicate that the proposed method can learn an operator that outperforms well-known regularizers (total variation, DCT-sparsity, and unsupervised dictionary learning) and collaborative filtering for denoising.
Avrajit Ghosh, Michael T. McCann, Madeline Mitchell, Saiprasad Ravishankar
SIAM J. Imaging Sci.4
2023 Robust Self-Guided Deep Image Prior
abstract
In this work, we study the deep image prior (DIP) for reconstruction problems in magnetic resonance imaging (MRI). DIP has become a popular approach for image reconstruction, where it recovers the clear image by fitting an overparameterized convolutional neural network (CNN) to the corrupted/undersampled measurements. To improve the performance of DIP, recent work shows that using a reference image as an input often leads to improved reconstruction results compared to vanilla DIP with random input. However, obtaining the reference input image often requires supervision and hence is difficult in practice. In this work, we propose a self-guided reconstruction scheme that uses no training data other than the set of undersampled measurements to simultaneously estimate the network weights and input (reference). We introduce a new regularization that aids the joint estimation by requiring the CNN to act as a powerful denoiser. The proposed self-guided method gives significantly improved image reconstructions for MRI with limited measurements compared to the conventional DIP and the reference-guided method while eliminating the need for any additional data.
Evan Bell, Shijun Liang 0001, Qing Qu 0001, Saiprasad Ravishankar
ICASSP4
2023 SMUG: Towards Robust Mri Reconstruction by Smoothed Unrolling
abstract
Although deep learning (DL) has gained much popularity for accelerated magnetic resonance imaging (MRI), recent studies have shown that DL-based MRI reconstruction models could be over-sensitive to tiny input perturbations (that are called ‘adversarial perturbations’), which cause unstable, low-quality reconstructed images. This raises the question of how to design robust DL methods for MRI reconstruction. To address this problem, we propose a novel image reconstruction framework, termed SMOOTHED UNROLLING (SMUG), which advances a deep unrolling-based MRI reconstruction model using a randomized smoothing (RS)-based robust learning operation. RS, which improves the tolerance of a model against input noises, has been widely used in the design of adversarial defense for image classification. Yet, we find that the conventional design that applies RS to the entire DL process is ineffective for MRI reconstruction. We show that SMUG addresses the above issue by customizing the RS operation based on the unrolling architecture of the DL-based MRI reconstruction model. Compared to the vanilla RS approach and several variants of SMUG, we show that SMUG improves the robustness of MRI reconstruction with respect to a diverse set of perturbation sources, including perturbations to input measurements, different measurement sampling rates, and different unrolling steps. Code for SMUG will be available at https://github.com/LGM70/SMUG.
Jinghan Jia, Shijun Liang 0001, Yuguang Yao, Saiprasad Ravishankar, Sijia Liu 0001
ICASSP5
2022 Bilevel Learning of ℓ1 Regularizers with Closed-Form Gradients (BLORC)
abstract
We present a method for supervised learning of sparsity-promoting regularizers, which are a key ingredient in many modern signal reconstruction approaches. The parameters of the regularizer are learned to minimize the mean squared error of reconstruction on a training set of ground truth signal and measurement pairs. Training involves solving a challenging bilevel optimization problem with a nonsmooth lower-level objective. We derive an expression for the gradient of the training loss using the implicit closed-form solution of the lower-level variational problem given by its dual problem, and provide an accompanying gradient descent algorithm (dubbed BLORC) to minimize the loss. Our experiments on simple natural images and for denoising 1D signals show that the proposed method can learn meaningful operators and the analytical gradients are calculated faster than standard automatic differentiation methods. While the approach we present is applied to denoising, we believe that it could be adapted to a wide variety of inverse problems with linear measurement models, thus giving it applicability in a wide range of scenarios.
Avrajit Ghosh, Michael T. McCann, Saiprasad Ravishankar
ICASSP3
2022 REPNP: Plug-and-Play with Deep Reinforcement Learning Prior for Robust Image Restoration
abstract
Image restoration schemes based on the pre-trained deep models have received great attention due to their unique flexibility for solving various inverse problems. In particular, the Plug-and-Play (PnP) framework is a popular and powerful tool that can integrate an off-the-shelf deep denoiser for different image restoration tasks with known observation models. However, obtaining the observation model that exactly matches the actual one can be challenging in practice. Thus, the PnP schemes with conventional deep denoisers may fail to generate satisfying results in some real-world image restoration tasks. We argue that the robustness of the PnP framework is largely limited by using the off-the-shelf deep denoisers that are trained by deterministic optimization. To this end, we propose a novel deep reinforcement learning (DRL) based PnP framework, dubbed RePNP, by leveraging a light-weight DRL-based denoiser for robust image restoration tasks. Experimental results demonstrate that the proposed RePNP is robust to the observation model used in the PnP scheme deviating from the actual one. Thus, RePNP can generate more reliable restoration results for image deblurring and super resolution tasks. Compared with several state-of-the-art deep image restoration baselines, RePNP achieves better results subjective to model deviation with fewer model parameters.
Chong Wang 0011, Rongkai Zhang 0001, Saiprasad Ravishankar, Bihan Wen
ICIP3
2022 Exploiting Non-Local Priors via Self-Convolution for Highly-Efficient Image Restoration
abstract
Constructing effective priors is critical to solving ill-posed inverse problems in image processing and computational imaging. Recent works focused on exploiting non-local similarity by grouping similar patches for image modeling, and demonstrated state-of-the-art results in many image restoration applications. However, compared to classic methods based on filtering or sparsity, non-local algorithms are more time-consuming, mainly due to the highly inefficient block matching step, i.e., distance between every pair of overlapping patches needs to be computed. In this work, we propose a novel Self-Convolution operator to exploit image non-local properties in a unified framework. We prove that the proposed Self-Convolution based formulation can generalize the commonly-used non-local modeling methods, as well as produce results equivalent to standard methods, but with much cheaper computation. Furthermore, by applying Self-Convolution, we propose an effective multi-modality image restoration scheme, which is much more efficient than conventional block matching for non-local modeling. Experimental results demonstrate that (1) Self-Convolution with fast Fourier transform implementation can significantly speed up most of the popular non-local image restoration algorithms, with two-fold to nine-fold faster block matching, and (2) the proposed online multi-modality image restoration scheme achieves superior denoising results than competing methods in both efficiency and effectiveness on RGB-NIR images. The code for this work is publicly available at https://github.com/GuoLanqing/Self-Convolution.
Lanqing Guo, Zhiyuan Zha, Saiprasad Ravishankar, Bihan Wen
IEEE Trans. Image Process.3
2021 Self-Convolution: A Highly-Efficient Operator for Non-Local Image Restoration
abstract
Constructing effective image priors is critical to solving ill-posed inverse problems, such as image restoration. Recent works proposed to exploit image non-local similarity for inverse problems by grouping similar patches, and demonstrated state-of-the-art results in many applications. However, comparing to classic local methods based on filtering or sparsity, most of the non-local algorithms are time-consuming, mainly due to the highly inefficient and redundant block matching step, where the distance between each pair of overlapping patches needs to be computed. In this work, we propose a novel Self-Convolution operator to exploit image non-local similarity in a self-supervised way. The proposed Self-Convolution can generalize the commonly-used block matching step, and produce the equivalent results with much cheaper computation. Based on Self-Convolution, we propose an effective multi-modality image restoration scheme, which is much more efficient than conventional block matching for non-local modeling. Experimental results also demonstrate that Self-Convolution can significantly speed up most of the popular non-local image restoration algorithms, with two-fold to nine-fold faster block matching. The codes will be released on GitHub.
Lanqing Guo, Zhiyuan Zha, Saiprasad Ravishankar, Bihan Wen
ICASSP3
2021 Learning Sparsifying Transforms for Image Reconstruction in Electrical Impedance Tomography
abstract
Electrical Impedance Tomography (EIT) is a fast and non-invasive imaging technology that reconstructs the internal electrical properties of a subject. However, its functionality is limited by low spatial resolution arising from an ill-posed and ill-conditioned inverse problem. Several sparsity-promoting regularization methods have been applied to improve the quality of EIT image reconstruction, including various ℓ0and ℓ1-based analytical models (TV, TwIST, etc.), and a patch-based sparse representation via a learned dictionary (using the K-SVD algorithm), dubbed CS-EIT. To further exploit the potential of compressed sensing in Electrical Impedance Tomography, this paper incorporates the recent novel method of transform learning for EIT image reconstruction. We propose a blind compressed sensing algorithm, dubbed TL-EIT, which simultaneously optimizes the sparsifying transform and updates the reconstructed image. We demonstrate using both synthetic and in vivo data that the proposed TL-EIT is more effective than other sparsity-based algorithms for reconstructing high-quality EIT images. In addition, TL-EIT also accelerates the reconstruction process in comparison to other learning-based algorithms like CS-EIT.
Kaiyi Yang, Narong Borijindargoon, Boon Poh Ng, Saiprasad Ravishankar, Bihan Wen
ICASSP4
2021 Labmat: Learned Feature-Domain Block Matching For Image Restoration
abstract
Grouping of similar patches, called block matching, has been widely used in image restoration applications. Popular block matching algorithms exploit image non-local similarities in spatial or a fixed transform domain, e.g., wavelets and DCT. However, applying these methods on corrupted patches usually leads to degraded matching accuracy, thus limiting the image restoration performance. In this work, we develop a novel methodology for performing block matching in a supervised way by learning multi-layer sparsifying transforms. The proposed learned transform-domain block matching method for image restoration, dubbed LABMAT, is shown to have better accuracy in terms of clustering similar blocks in the presence of noise, and it also achieves an improved denoising performance when it is incorporated into popular non-local denoising schemes.
Shijun Liang 0001, Berk Iskender, Bihan Wen, Saiprasad Ravishankar
ICIP4
2021 Blind Primed Supervised (BLIPS) Learning for MR Image Reconstruction
abstract
This paper examines a combined supervised-unsupervised framework involving dictionary-based blind learning and deep supervised learning for MR image reconstruction from under-sampled k-space data. A major focus of the work is to investigate the possible synergy of learned features in traditional shallow reconstruction using adaptive sparsity-based priors and deep prior-based reconstruction. Specifically, we propose a framework that uses an unrolled network to refine a blind dictionary learning-based reconstruction. We compare the proposed method with strictly supervised deep learning-based reconstruction approaches on several datasets of varying sizes and anatomies. We also compare the proposed method to alternative approaches for combining dictionary-based methods with supervised learning in MR image reconstruction. The improvements yielded by the proposed framework suggest that the blind dictionary-based approach preserves fine image details that the supervised approach can iteratively refine, suggesting that the features learned using the two methods are complementary.
Anish Lahiri, Saiprasad Ravishankar, Jeffrey A. Fessler
IEEE Trans. Medical Imaging3
2021 Unified Supervised-Unsupervised (SUPER) Learning for X-Ray CT Image Reconstruction
abstract
Traditional model-based image reconstruction (MBIR) methods combine forward and noise models with simple object priors. Recent machine learning methods for image reconstruction typically involve supervised learning or unsupervised learning, both of which have their advantages and disadvantages. In this work, we propose a unified supervised-unsupervised (SUPER) learning framework for X-ray computed tomography (CT) image reconstruction. The proposed learning formulation combines both unsupervised learning-based priors (or even simple analytical priors) together with (supervised) deep network-based priors in a unified MBIR framework based on a fixed point iteration analysis. The proposed training algorithm is also an approximate scheme for a bilevel supervised training optimization problem, wherein the network-based regularizer in the lower-level MBIR problem is optimized using an upper-level reconstruction loss. The training problem is optimized by alternating between updating the network weights and iteratively updating the reconstructions based on those weights. We demonstrate the learned SUPER models' efficacy for low-dose CT image reconstruction, for which we use the NIH AAPM Mayo Clinic Low Dose CT Grand Challenge dataset for training and testing. In our experiments, we studied different combinations of supervised deep network priors and unsupervised learning-based or analytical priors. Both numerical and visual results show the superiority of the proposed unified SUPER methods over standalone supervised learning-based methods, iterative MBIR methods, and variations of SUPER obtained via ablation studies. We also show that the proposed algorithm converges rapidly in practice.
Siqi Ye, Zhipeng Li 0003, Michael T. McCann, Yong Long, Saiprasad Ravishankar
IEEE Trans. Medical Imaging5
2020 Image Reconstruction: From Sparsity to Data-Adaptive Methods and Machine Learning
abstract
The field of medical image reconstruction has seen roughly four types of methods. The first type tended to be analytical methods, such as filtered backprojection (FBP) for X-ray computed tomography (CT) and the inverse Fourier transform for magnetic resonance imaging (MRI), based on simple mathematical models for the imaging systems. These methods are typically fast, but have suboptimal properties such as poor resolution-noise tradeoff for CT. A second type is iterative reconstruction methods based on more complete models for the imaging system physics and, where appropriate, models for the sensor statistics. These iterative methods improved image quality by reducing noise and artifacts. The U.S. Food and Drug Administration (FDA)-approved methods among these have been based on relatively simple regularization models. A third type of methods has been designed to accommodate modified data acquisition methods, such as reduced sampling in MRI and CT to reduce scan time or radiation dose. These methods typically involve mathematical image models involving assumptions such as sparsity or low rank. A fourth type of methods replaces mathematically designed models of signals and systems with data-driven or adaptive models inspired by the field of machine learning. This article focuses on the two most recent trends in medical image reconstruction: methods based on sparsity or low-rank models and data-driven methods based on machine learning techniques.
Saiprasad Ravishankar, Jong Chul Ye, Jeffrey A. Fessler
Proc. IEEE1
2020 DECT-MULTRA: Dual-Energy CT Image Decomposition With Learned Mixed Material Models and Efficient Clustering
abstract
Dual-energy computed tomography (DECT) imaging plays an important role in advanced imaging applications due to its material decomposition capability. Image-domain decomposition operates directly on CT images using linear matrix inversion, but the decomposed material images can be severely degraded by noise and artifacts. This paper proposes a new method dubbed DECT-MULTRA for image-domain DECT material decomposition that combines conventional penalized weighted-least squares (PWLS) estimation with regularization based on a mixed union of learned transforms (MULTRA) model. Our proposed approach pre-learns a union of common-material sparsifying transforms from patches extracted from all the basis materials, and a union of cross-material sparsifying transforms from multi-material patches. The common-material transforms capture the common properties among different material images, while the cross-material transforms capture the cross-dependencies. The proposed PWLS formulation is optimized efficiently by alternating between an image update step and a sparse coding and clustering step, with both of these steps having closed-form solutions. The effectiveness of our method is validated with both XCAT phantom and clinical head data. The results demonstrate that our proposed method provides superior material image quality and decomposition accuracy compared to other competing methods.
Zhipeng Li 0003, Saiprasad Ravishankar, Yong Long, Jeffrey A. Fessler
IEEE Trans. Medical Imaging2
2020 SPULTRA: Low-Dose CT Image Reconstruction With Joint Statistical and Learned Image Models
abstract
Low-dose CT image reconstruction has been a popular research topic in recent years. A typical reconstruction method based on post-log measurements is called penalized weighted-least squares (PWLS). Due to the underlying limitations of the post-log statistical model, the PWLS reconstruction quality is often degraded in low-dose scans. This paper investigates a shifted-Poisson (SP) model based likelihood function that uses the pre-log raw measurements that better represents the measurement statistics, together with a data-driven regularizer exploiting a Union of Learned TRAnsforms (SPULTRA). Both the SP induced data-fidelity term and the regularizer in the proposed framework are nonconvex. The proposed SPULTRA algorithm uses quadratic surrogate functions for the SP induced data-fidelity term. Each iteration involves a quadratic subproblem for updating the image, and a sparse coding and clustering subproblem that has a closed-form solution. The SPULTRA algorithm has a similar computational cost per iteration as its recent counterpart PWLS-ULTRA that uses post-log measurements, and it provides better image reconstruction quality than PWLS-ULTRA, especially in low-dose scans.
Siqi Ye, Saiprasad Ravishankar, Yong Long, Jeffrey A. Fessler
IEEE Trans. Medical Imaging2
2019 VIDOSAT: High-Dimensional Sparsifying Transform Learning for Online Video Denoising
abstract
Techniques exploiting the sparsity of images in a transform domain are effective for various applications in image and video processing. In particular, transform learning methods involve cheap computations and have been demonstrated to perform well in applications, such as image denoising and medical image reconstruction. Recently, we proposed methods for online learning of sparsifying transforms from streaming signals, which enjoy good convergence guarantees and involve lower computational costs than online synthesis dictionary learning. In this paper, we apply online transform learning to video denoising. We present a novel framework for online video denoising based on high-dimensional sparsifying transform learning for spatio-temporal patches. The patches are constructed either from corresponding 2D patches in successive frames or using an online block matching technique. The proposed online video denoising requires little memory and offers efficient processing. Numerical experiments evaluate the performance of the proposed video denoising algorithms on multiple video data sets. The proposed methods outperform several related and recent techniques, including denoising with 3D DCT, prior schemes based on dictionary learning, non-local means, background separation, and deep learning, as well as the popular VBM3D and VBM4D.
Bihan Wen, Saiprasad Ravishankar, Yoram Bresler
IEEE Trans. Image Process.2
2018 PWLS-ULTRA: An Efficient Clustering and Learning-Based Approach for Low-Dose 3D CT Image Reconstruction
abstract
The development of computed tomography (CT) image reconstruction methods that significantly reduce patient radiation exposure, while maintaining high image quality is an important area of research in low-dose CT imaging. We propose a new penalized weighted least squares (PWLS) reconstruction method that exploits regularization based on an efficient Union of Learned TRAnsforms (PWLS-ULTRA). The union of square transforms is pre-learned from numerous image patches extracted from a dataset of CT images or volumes. The proposed PWLS-based cost function is optimized by alternating between a CT image reconstruction step, and a sparse coding and clustering step. The CT image reconstruction step is accelerated by a relaxed linearized augmented Lagrangian method with ordered-subsets that reduces the number of forward and back projections. Simulations with 2-D and 3-D axial CT scans of the extended cardiac-torso phantom and 3-D helical chest and abdomen scans show that for both normal-dose and low-dose levels, the proposed method significantly improves the quality of reconstructed images compared to PWLS reconstruction with a nonadaptive edge-preserving regularizer. PWLS with regularization based on a union of learned transforms leads to better image reconstructions than using a single learned square transform. We also incorporate patch-based weights in PWLS-ULTRA that enhance image quality and help improve image resolution uniformity. The proposed approach achieves comparable or better image quality compared to learned overcomplete synthesis dictionaries, but importantly, is much faster (computationally more efficient).
Xuehang Zheng, Saiprasad Ravishankar, Yong Long, Jeffrey A. Fessler
IEEE Trans. Medical Imaging2
2017 Online data-driven dynamic image restoration using DINO-KAT models
abstract
Sparsity-based techniques have been popular for reconstructing images and videos from limited or corrupted measurements. Methods such as dictionary or transform learning have been demonstrated to be useful in applications such as denoising, inpainting, and medical image reconstruction. In this work, we propose a new framework for online or sequential adaptive reconstruction of dynamic image sequences from linear (typically undersampled) measurements. In particular, the spatiotemporal patches of the underlying dynamic image sequence are assumed to be sparse in a DIctioNary with lOw-ranK AToms (DINO-KAT), and the dictionary model and images are simultaneously and sequentially estimated from streaming measurements. The proposed online algorithm involves efficient memory usage and simple and efficient updates of the low-rank atoms, sparse coefficients, and images. Our numerical experiments show the usefulness of the proposed scheme in inverse problem settings such as video reconstruction or inpainting from limited and noisy pixels.
Brian E. Moore, Saiprasad Ravishankar
ICIP2
2017 Low-Rank and Adaptive Sparse Signal (LASSI) Models for Highly Accelerated Dynamic Imaging
abstract
Sparsity-based approaches have been popular in many applications in image processing and imaging. Compressed sensing exploits the sparsity of images in a transform domain or dictionary to improve image recovery fromundersampledmeasurements. In the context of inverse problems in dynamic imaging, recent research has demonstrated the promise of sparsity and low-rank techniques. For example, the patches of the underlying data are modeled as sparse in an adaptive dictionary domain, and the resulting image and dictionary estimation from undersampled measurements is called dictionary-blind compressed sensing, or the dynamic image sequence is modeled as a sum of low-rank and sparse (in some transform domain) components (L+S model) that are estimated from limited measurements. In this work, we investigate a data-adaptive extension of the L+S model, dubbed LASSI, where the temporal image sequence is decomposed into a low-rank component and a component whose spatiotemporal (3D) patches are sparse in some adaptive dictionary domain. We investigate various formulations and efficient methods for jointly estimating the underlying dynamic signal components and the spatiotemporal dictionary from limited measurements. We also obtain efficient sparsity penalized dictionary-blind compressed sensing methods as special cases of our LASSI approaches. Our numerical experiments demonstrate the promising performance of LASSI schemes for dynamicmagnetic resonance image reconstruction from limited k-t space data compared to recent methods such as k-t SLR and L+S, and compared to the proposed dictionary-blind compressed sensing method.
Saiprasad Ravishankar, Brian E. Moore, Raj Rao Nadakuditi, Jeffrey A. Fessler
IEEE Trans. Medical Imaging1
2016 Learning flipping and rotation invariant sparsifying transforms
abstract
Adaptive sparse representation has been heavily exploited in signal processing and computer vision. Recently, sparsifying transform learning received interest for its cheap computation and optimal updates in the alternating algorithms. In this work, we develop a methodology for learning a Flipping and Rotation Invariant Sparsifying Transform, dubbed FRIST, to better represent natural images that contain textures with various geometrical directions. The proposed alternating learning algorithm involves efficient optimal updates. We demonstrate empirical convergence behavior of the proposed learning algorithm. Preliminary experiments show the usefulness of FRIST for image sparse representation, segmentation, robust inpainting, and MRI reconstruction with promising performances.
Bihan Wen, Saiprasad Ravishankar, Yoram Bresler
ICIP2
2015 Video denoising by online 3D sparsifying transform learning
abstract
Exploiting the sparsity of signals in an adaptive dictionary or transform domain benefits various applications in image/video processing. As opposed to synthesis dictionary learning, transform learning allows for cheap computations, and has been demonstrated to perform well in applications such as image denoising. Very recently, we proposed methods for online sparsifying transform learning, which are particularly useful for processing large-scale or streaming data. Online transform learning has good convergence guarantees and enjoys a much lower computational cost than online synthesis dictionary learning. In this work, we present a video denoising framework based on online 3D spatio-temporal sparsifying transform learning. The proposed scheme has low computational and memory costs, and can potentially handle streaming video. Our numerical experiments show promising performance for the proposed video denoising method compared to popular prior or state-of-the-art methods.
Bihan Wen, Saiprasad Ravishankar, Yoram Bresler
ICIP2
2015 Structured Overcomplete Sparsifying Transform Learning with Convergence Guarantees and Applications
Bihan Wen, Saiprasad Ravishankar, Yoram Bresler
Int. J. Comput. Vis.2
2015 Efficient Blind Compressed Sensing Using Sparsifying Transforms with Convergence Guarantees and Application to Magnetic Resonance Imaging
abstract
Natural signals and images are well known to be approximately sparse in transform domains such as wavelets and discrete cosine transform. This property has been heavily exploited in various applications in image processing and medical imaging. Compressed sensing exploits the sparsity of images or image patches in a transform domain or synthesis dictionary to reconstruct images from undersampled measurements. In this work, we focus on blind compressed sensing, where the underlying sparsifying transform is a priori unknown, and propose a framework to simultaneously reconstruct the underlying image as well as the sparsifying transform from highly undersampled measurements. The proposed block coordinate descent-type algorithms involve highly efficient optimal updates. Importantly, we prove that although the proposed blind compressed sensing formulations are highly nonconvex, our algorithms are globally convergent (i.e., they converge from any initialization) to the set of critical points of the objectives defining the formulations. These critical points are guaranteed to be at least partial global and partial local minimizers. The exact point(s) of convergence may depend on initialization. We illustrate the usefulness of the proposed framework for magnetic resonance image reconstruction from highly undersampled k-space measurements. As compared to previous methods involving the synthesis dictionary model, our approach is much faster, while also providing promising reconstruction quality.
Saiprasad Ravishankar, Yoram Bresler
SIAM J. Imaging Sci.1
2014 Doubly sparse transform learning with convergence guarantees
abstract
The sparsity of natural signals in transform domains such as the DCT has been heavily exploited in various applications. Recently, we introduced the idea of learning sparsifying transforms from data, and demonstrated the usefulness of learnt transforms in image representation, and denoising. However, the learning formulations therein were non-convex, and the algorithms lacked strong convergence properties. In this work, we propose a novel convex formulation for square sparsifying transform learning. We also enforce a doubly sparse structure on the transform, which makes its learning, storage, and implementation efficient. Our algorithm is guaranteed to converge to a global optimum, and moreover converges quickly. We also introduce a non-convex variant of the convex formulation, for which the algorithm is locally convergent. We show the superior promise of our learnt transforms as compared to analytical sparsifying transforms such as the DCT for image representation.
Saiprasad Ravishankar, Yoram Bresler
ICASSP1
2014 Learning overcomplete sparsifying transforms with block cosparsity
abstract
The sparsity of images in a transform domain or dictionary has been widely exploited in image processing. Compared to the synthesis dictionary model, sparse coding in the (single) transform model is cheap. However, natural images typically contain diverse textures that cannot be sparsified well by a single transform. Hence, we propose a union of sparsifying transforms model, which is equivalent to an overcomplete transform model with block cosparsity (OC-TOBOS). Our alternating algorithm for transform learning involves simple closed-form updates. When applied to images, our algorithm learns a collection of well-conditioned transforms, and a good clustering of the patches or textures. Our learnt transforms provide better image representations than learned square transforms. We also show the promising denoising performance and speedups provided by the proposed method compared to synthesis dictionary-based denoising.
Bihan Wen, Saiprasad Ravishankar, Yoram Bresler
ICIP2
2013 Learning overcomplete sparsifying transforms for signal processing
abstract
Adaptive sparse representations have been very popular in numerous applications in recent years. The learning of synthesis sparsifying dictionaries has particularly received much attention, and such adaptive dictionaries have been shown to be useful in applications such as image denoising, and magnetic resonance image reconstruction. In this work, we focus on the alternative sparsifying transform model, for which sparse coding is cheap and exact, and study the learning of tall or overcomplete sparsifying transforms from data. We propose various penalties that control the sparsifying ability, condition number, and incoherence of the learnt transforms. Our alternating algorithm for transform learning converges empirically, and significantly improves the quality of the learnt transform over the iterations. We present examples demonstrating the promising performance of adaptive overcomplete transforms over adaptive overcomplete synthesis dictionaries learnt using K-SVD, in the application of image denoising.
Saiprasad Ravishankar, Yoram Bresler
ICASSP1
2013 Closed-form solutions within sparsifying transform learning
abstract
Many applications in signal processing benefit from the sparsity of signals in a certain transform domain or dictionary. Synthesis sparsifying dictionaries that are directly adapted to data have been popular in applications such as image denoising, and medical image reconstruction. In this work, we focus specifically on the learning of orthonormal as well as well-conditioned square sparsifying transforms. The proposed algorithms alternate between a sparse coding step, and a transform update step. We derive the exact analytical solution for each of these steps. Adaptive well-conditioned transforms are shown to perform better in applications compared to adapted orthonormal ones. Moreover, the closed form solution for the transform update step achieves the global minimum in that step, and also provides speedups over iterative solutions involving conjugate gradients. We also present examples illustrating the promising performance and significant speed-ups of transform learning over synthesis K-SVD in image denoising.
Saiprasad Ravishankar, Yoram Bresler
ICASSP1
2013 Learning Doubly Sparse Transforms for Images
abstract
The sparsity of images in a transform domain or dictionary has been exploited in many applications in image processing. For example, analytical sparsifying transforms, such as wavelets and discrete cosine transform (DCT), have been extensively used in compression standards. Recently, synthesis sparsifying dictionaries that are directly adapted to the data have become popular especially in applications such as image denoising. Following up on our recent research, where we introduced the idea of learning square sparsifying transforms, we propose here novel problem formulations for learning doubly sparse transforms for signals or image patches. These transforms are a product of a fixed, fast analytic transform such as the DCT, and an adaptive matrix constrained to be sparse. Such transforms can be learnt, stored, and implemented efficiently. We show the superior promise of our learnt transforms as compared with analytical sparsifying transforms such as the DCT for image representation. We also show promising performance in image denoising that compares favorably with approaches involving learnt synthesis dictionaries such as the K-SVD algorithm. The proposed approach is also much faster than K-SVD denoising.
Saiprasad Ravishankar, Yoram Bresler
IEEE Trans. Image Process.1
2012 Learning sparsifying transforms for image processing
abstract
The sparsity of signals and images in a certain analytically defined transform domain or dictionary such as discrete cosine transform or wavelets has been exploited in many applications in signal and image processing. Recently, the idea of learning a dictionary for sparse representation of data has become popular. However, while there has been extensive research on learning synthesis dictionaries, the idea of learning analysis sparsifying transforms has received only little attention. We propose a novel problem formulation and an alternating algorithm for learning well-conditioned square sparsifying transforms from data. We show the superiority of our approach for image representation over analytical sparsifying transforms such as the DCT. We also show promise in image denoising. Denoising using the learnt analysis transforms is not only better than by synthesis dictionaries learnt using the K-SVD algorithm but also faster.
Saiprasad Ravishankar, Yoram Bresler
ICIP1
2012 Learning doubly sparse transforms for image representation
abstract
The sparsity of images in a fixed analytic transform domain or dictionary such as DCT or Wavelets has been exploited in many applications in image processing including image compression. Recently, synthesis sparsifying dictionaries that are directly adapted to the data have become popular in image processing. However, the idea of learning sparsifying transforms has received only little attention. We propose a novel problem formulation for learning doubly sparse transforms for signals or image patches. These transforms are a product of a fixed, fast analytic transform such as the DCT, and an adaptive matrix constrained to be sparse. Such transforms can be learnt, stored, and implemented efficiently. We show the superior promise of our approach as compared to analytical sparsifying transforms such as DCT for image representation.
Saiprasad Ravishankar, Yoram Bresler
ICIP1
2011 MR Image Reconstruction From Highly Undersampled k-Space Data by Dictionary Learning
abstract
Compressed sensing (CS) utilizes the sparsity of magnetic resonance (MR) images to enable accurate reconstruction from undersampled k-space data. Recent CS methods have employed analytical sparsifying transforms such as wavelets, curvelets, and finite differences. In this paper, we propose a novel framework for adaptively learning the sparsifying transform (dictionary), and reconstructing the image simultaneously from highly undersampled k-space data. The sparsity in this framework is enforced on overlapping image patches emphasizing local structure. Moreover, the dictionary is adapted to the particular image instance thereby favoring better sparsities and consequently much higher undersampling rates. The proposed alternating reconstruction algorithm learns the sparsifying dictionary, and uses it to remove aliasing and noise in one step, and subsequently restores and fills-in the k-space data in the other step. Numerical experiments are conducted on MR images and on real MR data of several anatomies with a variety of sampling schemes. The results demonstrate dramatic improvements on the order of 4-18 dB in reconstruction error and doubling of the acceptable undersampling factor using the proposed adaptive dictionary as compared to previous CS methods. These improvements persist over a wide range of practical data signal-to-noise ratios, without any parameter tuning.
Saiprasad Ravishankar, Yoram Bresler
IEEE Trans. Medical Imaging1
2009 Automated feature extraction for early detection of diabetic retinopathy in fundus images
abstract
Automated detection of lesions in retinal images can assist in early diagnosis and screening of a common disease: Diabetic Retinopathy. A robust and computationally efficient approach for the localization of the different features and lesions in a fundus retinal image is presented in this paper. Since many features have common intensity properties, geometric features and correlations are used to distinguish between them. We propose a new constraint for optic disk detection where we first detect the major blood vessels and use the intersection of these to find the approximate location of the optic disk. This is further localized using color properties. We also show that many of the features such as the blood vessels, exudates and microaneurysms and hemorrhages can be detected quite accurately using different morphological operations applied appropriately. Extensive evaluation of the algorithm on a database of 516 images with varied contrast, illumination and disease stages yields 97.1% success rate for optic disk localization, a sensitivity and specificity of 95.7%and 94.2%respectively for exudate detection and 95.1% and 90.5% for microaneurysm/hemorrhage detection. These compare very favorably with existing systems and promise real deployment of these systems.
Saiprasad Ravishankar, Anurag Mittal
CVPR1
2008 Multi-stage Contour Based Detection of Deformable Objects
Saiprasad Ravishankar, Anurag Mittal
ECCV (1)1