Sina Farsiu

dblp:34/2309 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-4872-2902ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
8 papers
Image and video processing · 77% Geometric modeling and processing · 20% Virtual and augmented reality · 3%
Artificial intelligence
4 papers
Segmentation and scene understanding · 90% 3D vision · 10%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing
image restoration
2.052025
RUN: Reversible Unfolding Network for Concealed Object Segmentation · ICML 2025
Reti-Diff: Illumination Degradation Image Restoration with Retinex-based Latent Diffusion Model · ICLR 2025
Efficient Fourier-Wavelet Super-Resolution · IEEE Trans. Image Process. 2010
Computer vision › Segmentation and scene understanding › semantic segmentation
weakly supervised semantic segmentation
0.912025
Segment Concealed Objects With Incomplete Supervision · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Image and video processing › compressive sensing
deep unfolding network
0.912025
RUN: Reversible Unfolding Network for Concealed Object Segmentation · ICML 2025
Computer vision › Segmentation and scene understanding
medical image segmentation
0.712023
Directional Connectivity-based Segmentation of Medical Images · CVPR 2023
Geometric modeling and processing
3d reconstruction
0.512021
Mesoscopic Photogrammetry With an Unstabilized Phone Camera · CVPR 2021
Geometric modeling and processing › 3d reconstruction
photogrammetry
0.512021
Mesoscopic Photogrammetry With an Unstabilized Phone Camera · CVPR 2021
Image and video processing › super-resolution
multi-frame super-resolution
0.232010
Efficient Fourier-Wavelet Super-Resolution · IEEE Trans. Image Process. 2010
Multiframe demosaicing and super-resolution of color images · IEEE Trans. Image Process. 2006
Fast and robust multiframe super resolution · IEEE Trans. Image Process. 2004
Computer vision › 3D vision
3d reconstruction
0.212015
Tree Topology Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Computer vision › 3D vision
structure from motion
0.212015
Tree Topology Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Image and video processing › image restoration
image denoising
0.222010
Efficient Fourier-Wavelet Super-Resolution · IEEE Trans. Image Process. 2010
Kernel Regression for Image Processing and Reconstruction · IEEE Trans. Image Process. 2007
Virtual and augmented reality › tracking
camera pose estimation
0.112021
Mesoscopic Photogrammetry With an Unstabilized Phone Camera · CVPR 2021
Image and video processing › image restoration
image deblurring
0.122010
Deblurring Using Regularized Locally Adaptive Kernel Regression · IEEE Trans. Image Process. 2008
Efficient Fourier-Wavelet Super-Resolution · IEEE Trans. Image Process. 2010
Image and video processing
super-resolution
0.112010
Efficient Fourier-Wavelet Super-Resolution · IEEE Trans. Image Process. 2010
Image and video processing › image restoration › image denoising
wavelet-based denoising
0.112010
Efficient Fourier-Wavelet Super-Resolution · IEEE Trans. Image Process. 2010
Image and video processing › image restoration › multi-task image restoration
denoising and deblurring
0.112008
Deblurring Using Regularized Locally Adaptive Kernel Regression · IEEE Trans. Image Process. 2008
Image and video processing
image fusion
0.112007
Kernel Regression for Image Processing and Reconstruction · IEEE Trans. Image Process. 2007
Image and video processing › image restoration › multichannel image restoration
color image restoration
0.112006
Multiframe demosaicing and super-resolution of color images · IEEE Trans. Image Process. 2006
Image and video processing › image restoration
demosaicing
0.112006
Multiframe demosaicing and super-resolution of color images · IEEE Trans. Image Process. 2006

Methods — techniques the papers use, named apart from their topics

state space model · 1.7reversible network · 1.7transformer · 0.9segment anything model · 0.9pseudo-labeling · 0.9mean-teacher framework · 0.9latent diffusion · 0.9feature grouping · 0.9diffusion model · 0.9latent space disentanglement · 0.7deep network · 0.7untrained encoder-decoder · 0.5nonparametric distortion model · 0.5convolutional neural network · 0.5heuristic search · 0.2generative tree-growth model · 0.2kernel regression · 0.2l1-norm minimization · 0.1
YearPublicationVenuePosition
2026 Spatial coherence loss: All objects matter in salient and camouflaged object detection Spatial coherence loss: All objects matter in salient and camouflaged object detection
Ziyun Yang, Kevin Choy, Sina Farsiu
Pattern Recognit.3
2025 Reti-Diff: Illumination Degradation Image Restoration with Retinex-based Latent Diffusion Model
abstract
Illumination degradation image restoration (IDIR) techniques aim to improve the visibility of degraded images and mitigate the adverse effects of deteriorated illumination. Among these algorithms, diffusion-based models (DM) have shown promising performance but are often burdened by heavy computational demands and pixel misalignment issues when predicting the image-level distribution. To tackle these problems, we propose to leverage DM within a compact latent space to generate concise guidance priors and introduce a novel solution called Reti-Diff for the IDIR task. Specifically, Reti-Diff comprises two significant components: the Retinex-based latent DM (RLDM) and the Retinex-guided transformer (RGformer). RLDM is designed to acquire Retinex knowledge, extracting reflectance and illumination priors to facilitate detailed reconstruction and illumination correction. RGformer subsequently utilizes these compact priors to guide the decomposition of image features into their respective reflectance and illumination components. Following this, RGformer further enhances and consolidates these decomposed features, resulting in the production of refined images with consistent content and robustness to handle complex degradation scenarios. Extensive experiments demonstrate that Reti-Diff outperforms existing methods on three IDIR tasks, as well as downstream applications.
Chunming He, Chengyu Fang 0001, Yulun Zhang 0001, Longxiang Tang, Jinfa Huang, Kai Li 0012, Zhenhua Guo 0001, Xiu Li 0001, Sina Farsiu
ICLR9
2025 RUN: Reversible Unfolding Network for Concealed Object Segmentation
abstract
Concealed object segmentation (COS) is a challenging problem that focuses on identifying objects that are visually blended into their background. Existing methods often employ reversible strategies to concentrate on uncertain regions but only focus on the mask level, overlooking the valuable of the RGB domain. To address this, we propose a Reversible Unfolding Network (RUN) in this paper. RUN formulates the COS task as a foreground-background separation process and incorporates an extra residual sparsity constraint to minimize segmentation uncertainties. The optimization solution of the proposed model is unfolded into a multistage network, allowing the original fixed parameters to become learnable. Each stage of RUN consists of two reversible modules: the Segmentation-Oriented Foreground Separation (SOFS) module and the Reconstruction-Oriented Background Extraction (ROBE) module. SOFS applies the reversible strategy at the mask level and introduces Reversible State Space to capture non-local information. ROBE extends this to the RGB domain, employing a reconstruction network to address conflicting foreground and background regions identified as distortion-prone areas, which arise from their separate estimation by independent modules. As the stages progress, RUN gradually facilitates reversible modeling of foreground and background in both the mask and RGB domains, reducing false-positive and false-negative regions. Extensive experiments demonstrate the superior performance of RUN and underscore the promise of unfolding-based frameworks for COS and other high-level vision tasks. Code is available at https://github.com/ChunmingHe/RUN.
Chunming He, Rihan Zhang, Fengyang Xiao, Chengyu Fang 0001, Longxiang Tang, Yulun Zhang 0001, Linghe Kong, Deng-Ping Fan, Kai Li 0012, Sina Farsiu
ICML10
2025 Segment Concealed Objects With Incomplete Supervision
abstract
Incompletely-Supervised Concealed Object Segmentation (ISCOS) involves segmenting objects that seamlessly blend into their surrounding environments, utilizing incompletely annotated data, such as weak and semi-annotations, for model training. This task remains highly challenging due to (1) the limited supervision provided by the incompletely annotated training data, and (2) the difficulty of distinguishing concealed objects from the background, which arises from the intrinsic similarities in concealed scenarios. In this paper, we introduce the first unified method for ISCOS to address these challenges. To tackle the issue of incomplete supervision, we propose a unified mean-teacher framework, SEE, that leverages the vision foundation model, "Segment Anything Model (SAM)", to generate pseudo-labels using coarse masks produced by the teacher model as prompts. To mitigate the effect of low-quality segmentation masks, we introduce a series of strategies for pseudo-label generation, storage, and supervision. These strategies aim to produce informative pseudo-labels, store the best pseudo-labels generated, and select the most reliable components to guide the student model, thereby ensuring robust network training. Additionally, to tackle the issue of intrinsic similarity, we design a hybrid-granularity feature grouping module that groups features at different granularities and aggregates these results. By clustering similar features, this module promotes segmentation coherence, facilitating more complete segmentation for both single-object and multiple-object images. We validate the effectiveness of our approach across multiple ISCOS tasks, and experimental results demonstrate that our method achieves state-of-the-art performance. Furthermore, SEE can serve as a plug-and-play solution, enhancing the performance of existing models.
Chunming He, Kai Li 0012, Yachao Zhang 0001, Ziyun Yang, Youwei Pang, Longxiang Tang, Chengyu Fang 0001, Yulun Zhang 0001, Linghe Kong, Xiu Li 0001, Sina Farsiu
IEEE Trans. Pattern Anal. Mach. Intell.11
2023 Directional Connectivity-based Segmentation of Medical Images
abstract
Anatomical consistency in biomarker segmentation is crucial for many medical image analysis tasks. A promising paradigm for achieving anatomically consistent segmentation via deep networks is incorporating pixel connectivity, a basic concept in digital topology, to model inter-pixel relationships. However, previous works on connectivity modeling have ignored the rich channel-wise directional information in the latent space. In this work, we demonstrate that effective disentanglement of directional sub-space from the shared latent space can significantly enhance the feature representation in the connectivity-based network. To this end, we propose a directional connectivity modeling scheme for segmentation that decouples, tracks, and utilizes the directional information across the network. Experiments on various public medical image segmentation benchmarks show the effectiveness of our model as compared to the state-of-the-art methods. Code is available at https://github.com/Zyun-Y/DconnNet.
Ziyun Yang, Sina Farsiu
CVPR2
2023 RetiFluidNet: A Self-Adaptive and Multi-Attention Deep Convolutional Network for Retinal OCT Fluid Segmentation
abstract
Optical coherence tomography (OCT) helps ophthalmologists assess macular edema, accumulation of fluids, and lesions at microscopic resolution. Quantification of retinal fluids is necessary for OCT-guided treatment management, which relies on a precise image segmentation step. As manual analysis of retinal fluids is a time-consuming, subjective, and error-prone task, there is increasing demand for fast and robust automatic solutions. In this study, a new convolutional neural architecture named RetiFluidNet is proposed for multi-class retinal fluid segmentation. The model benefits from hierarchical representation learning of textural, contextual, and edge features using a new self-adaptive dual-attention (SDA) module, multiple self-adaptive attention-based skip connections (SASC), and a novel multi-scale deep self-supervision learning (DSL) scheme. The attention mechanism in the proposed SDA module enables the model to automatically extract deformation-aware representations at different levels, and the introduced SASC paths further consider spatial-channel interdependencies for concatenation of counterpart encoder and decoder units, which improve representational capability. RetiFluidNet is also optimized using a joint loss function comprising a weighted version of dice overlap and edge-preserved connectivity-based losses, where several hierarchical stages of multi-scale local losses are integrated into the optimization process. The model is validated based on three publicly available datasets: RETOUCH, OPTIMA, and DUKE, with comparisons against several baselines. Experimental results on the datasets prove the effectiveness of the proposed model in retinal OCT fluid segmentation and reveal that the suggested method is more effective than existing state-of-the-art fluid segmentation algorithms in adapting to retinal OCT scans recorded by various image scanning instruments.
Reza Rasti, Armin Biglari, Mohammad Rezapourian, Ziyun Yang, Sina Farsiu
IEEE Trans. Medical Imaging5
2022 Modeling extremes with d-max-decreasing neural networks
abstract
We propose a neural network architecture that enables non-parametric calibration and generation of multivariate extreme value distributions (MEVs). MEVs arise from Extreme Value Theory (EVT) as the necessary class of models when extrapolating a distributional fit over large spatial and temporal scales based on data observed in intermediate scales. In turn, EVT dictates that $d$-max-decreasing, a stronger form of convexity, is an essential shape constraint in the characterization of MEVs. As far as we know, our proposed architecture provides the first class of non-parametric estimators for MEVs that preserve these essential shape constraints. We show that the architecture approximates the dependence structure encoded by MEVs at parametric rate. Moreover, we present a new method for sampling high-dimensional MEVs using a generative model. We demonstrate our methodology on a wide range of experimental settings, ranging from environmental sciences to financial mathematics and verify that the structural properties of MEVs are retained compared to existing methods.
Ali Hasan, Khalil Elkhalil, Yuting Ng, João M. Pereira 0002, Sina Farsiu, Jose H. Blanchet, Vahid Tarokh
UAI5
2022 BiconNet: An edge-preserved connectivity-based approach for salient object detection
Ziyun Yang, Somayyeh Soltanian-Zadeh, Sina Farsiu
Pattern Recognit.3
2021 Fisher Auto-Encoders
abstract
It has been conjectured that the Fisher divergence is more robust to model uncertainty than the conventional Kullback-Leibler (KL) divergence. This motivates the design of a new class of robust generative auto-encoders (AE) referred to as Fisher auto-encoders. Our approach is to design Fisher AEs by minimizing the Fisher divergence between the intractable joint distribution of observed data and latent variables, with that of the postulated/modeled joint distribution. In contrast to KL-based variational AEs (VAEs), the Fisher AE can exactly quantify the distance between the true and the model-based posterior distributions. Qualitative and quantitative results are provided on both MNIST and celebA datasets demonstrating the competitive performance of Fisher AEs in terms of robustness compared to other AEs such as VAEs and Wasserstein AEs.
Khalil Elkhalil, Ali Hasan, Jie Ding 0002, Sina Farsiu, Vahid Tarokh
AISTATS4
2021 Mesoscopic Photogrammetry With an Unstabilized Phone Camera
abstract
We present a feature-free photogrammetric technique that enables quantitative 3D mesoscopic (mm-scale height variation) imaging with tens-of-micron accuracy from sequences of images acquired by a smartphone at close range (several cm) under freehand motion without additional hardware. Our end-to-end, pixel-intensity-based approach jointly registers and stitches all the images by estimating a coaligned height map, which acts as a pixel-wise radial deformation field that orthorectifies each camera image to allow plane-plus-parallax registration. The height maps themselves are reparameterized as the output of an untrained encoder-decoder convolutional neural network (CNN) with the raw camera images as the input, which effectively removes many reconstruction artifacts. Our method also jointly estimates both the camera’s dynamic 6D pose and its distortion using a nonparametric model, the latter of which is especially important in mesoscopic applications when using cameras not designed for imaging at short working distances, such as smartphone cameras. We also propose strategies for reducing computation time and memory, applicable to other multi-frame registration problems. Finally, we demonstrate our method using sequences of multi-megapixel images captured by an un-stabilized smartphone on a variety of samples (e.g., painting brushstrokes, circuit board, seeds).
Kevin C. Zhou, Colin L. V. Cooke, Jaehee Park, Ruobing Qian, Roarke Horstmeyer, Joseph A. Izatt, Sina Farsiu
CVPR7
2021 Open-Source Automatic Segmentation of Ocular Structures and Biomarkers of Microbial Keratitis on Slit-Lamp Photography Images Using Deep Learning
abstract
We propose a fully-automatic deep learning-based algorithm for segmentation of ocular structures and microbial keratitis (MK) biomarkers on slit-lamp photography (SLP) images. The dataset consisted of SLP images from 133 eyes with manual annotations by a physician, P1. A modified region-based convolutional neural network, SLIT-Net, was developed and trained using P1's annotations to identify and segment four pathological regions of interest (ROIs) on diffuse white light images (stromal infiltrate (SI), hypopyon, white blood cell (WBC) border, corneal edema border), one pathological ROI on diffuse blue light images (epithelial defect (ED)), and two non-pathological ROIs on all images (corneal limbus, light reflexes). To assess inter-reader variability, 75 eyes were manually annotated for pathological ROIs by a second physician, P2. Performance was evaluated using the Dice similarity coefficient (DSC) and Hausdorff distance (HD). Using seven-fold cross-validation, the DSC of the algorithm (as compared to P1) for all ROIs was good (range: 0.62-0.95) on all 133 eyes. For the subset of 75 eyes with manual annotations by P2, the DSC for pathological ROIs ranged from 0.69-0.85 (SLIT-Net) vs. 0.37-0.92 (P2). DSCs for SLIT-Net were not significantly different than P2 for segmenting hypopyons (p > 0.05) and higher than P2 for WBCs (p < 0.001) and edema (p < 0.001). DSCs were higher for P2 for segmenting SIs (p < 0.001) and EDs (p < 0.001). HDs were lower for P2 for segmenting SIs (p = 0.005) and EDs (p < 0.001) and not significantly different for hypopyons (p > 0.05), WBCs (p > 0.05), and edema (p > 0.05). This prototype fully-automatic algorithm to segment MK biomarkers on SLP images performed to expectations on an exploratory dataset and holds promise for quantification of corneal physiology and pathology.
Jessica Loo, Matthias F. Kriegel, Megan M. Tuohy, Kyeong Hwan Kim, Venkatesh Prajna, Maria A. Woodward, Sina Farsiu
IEEE J. Biomed. Health Informatics7
2020 Learning Partial Differential Equations From Data Using Neural Networks
abstract
We develop a framework for estimating unknown partial differential equations (PDEs) from noisy data, using a deep learning approach. Given noisy samples of a solution to an unknown PDE, our method interpolates the samples using a neural network, and extracts the PDE by equating derivatives of the neural network approximation. Our method applies to PDEs which are linear combinations of user-defined dictionary functions, and generalizes previous methods that only consider parabolic PDEs. We introduce a regularization scheme that prevents the function approximation from overfitting the data and forces it to be a solution of the underlying PDE. We validate the model on simulated data generated by the known PDEs and added Gaussian noise, and we study our method under different levels of noise. We also compare the error of our method with a Cramer-Rao lower bound for an ordinary differential equation (ODE). Our results indicate that our method outperforms other methods in estimating PDEs, especially in the low signal-to-noise (SNR) regime.
Ali Hasan, João M. Pereira 0002, Robert J. Ravier, Sina Farsiu, Vahid Tarokh
ICASSP4
2020 MimickNet, Mimicking Clinical Image Post- Processing Under Black-Box Constraints
abstract
Image post-processing is used in clinical-grade ultrasound scanners to improve image quality (e.g., reduce speckle noise and enhance contrast). These post-processing techniques vary across manufacturers and are generally kept proprietary, which presents a challenge for researchers looking to match current clinical-grade workflows. We introduce a deep learning framework, MimickNet, that transforms conventional delay-and-summed (DAS) beams into the approximate Dynamic Tissue Contrast Enhanced (DTCE™) post-processed images found on Siemens clinical-grade scanners. Training MimickNet only requires post-processed image samples from a scanner of interest without the need for explicit pairing to DAS data. This flexibility allows MimickNet to hypothetically approximate any manufacturer's post-processing without access to the pre-processed data. MimickNet post-processing achieves a 0.940 ± 0.018 structural similarity index measurement (SSIM) compared to clinical-grade post-processing on a 400 cine-loop test set, 0.937 ± 0.025 SSIM on a prospectively acquired dataset, and 0.928 ± 0.003 SSIM on an out-of-distribution cardiac cine-loop after gain adjustment. To our knowledge, this is the first work to establish deep learning models that closely approximate ultrasound post-processing found in current medical practice. MimickNet serves as a clinical post-processing baseline for future works in ultrasound image formation to compare against. Additionally, it can be used as a pretrained model for fine-tuning towards different post-processing techniques. To this end, we have made the MimickNet software, phantom data, and permitted in vivo data open-source at https://github.com/ouwen/MimickNet.
Ouwen Huang, Will Long, Nick Bottenus, Marcelo Lerendegui, Gregg E. Trahey, Sina Farsiu, Mark Palmeri
IEEE Trans. Medical Imaging6
2018 Statistical Models of Signal and Noise and Fundamental Limits of Segmentation Accuracy in Retinal Optical Coherence Tomography
abstract
Optical coherence tomography (OCT) has revolutionized diagnosis and prognosis of ophthalmic diseases by visualization and measurement of retinal layers. To speed up the quantitative analysis of disease biomarkers, an increasing number of automatic segmentation algorithms have been proposed to estimate the boundary locations of retinal layers. While the performance of these algorithms has significantly improved in recent years, a critical question to ask is how far we are from a theoretical limit to OCT segmentation performance. In this paper, we present the Cramèr-Rao lower bounds (CRLBs) for the problem of OCT layer segmentation. In deriving the CRLBs, we address the important problem of defining statistical models that best represent the intensity distribution in each layer of the retina. Additionally, we calculate the bounds under an optimal affine bias, reflecting the use of prior knowledge in many segmentation algorithms. Experiments using in vivo images of human retina from a commercial spectral domain OCT system are presented, showing potential for improvement of automated segmentation accuracy. Our general mathematical model can be easily adapted for virtually any OCT system. Furthermore, the statistical models of signal and noise developed in this paper can be utilized for the future improvements of OCT image denoising, reconstruction, and many other applications.
Theodore B. Dubose, David Cunefare, Elijah Cole, Peyman Milanfar, Joseph A. Izatt, Sina Farsiu
IEEE Trans. Medical Imaging6
2017 Segmentation Based Sparse Reconstruction of Optical Coherence Tomography Images
abstract
We demonstrate the usefulness of utilizing a segmentation step for improving the performance of sparsity based image reconstruction algorithms. In specific, we will focus on retinal optical coherence tomography (OCT) reconstruction and propose a novel segmentation based reconstruction framework with sparse representation, termed segmentation based sparse reconstruction (SSR). The SSR method uses automatically segmented retinal layer information to construct layer-specific structural dictionaries. In addition, the SSR method efficiently exploits patch similarities within each segmented layer to enhance the reconstruction performance. Our experimental results on clinical-grade retinal OCT images demonstrate the effectiveness and efficiency of the proposed SSR method for both denoising and interpolation of OCT images.
Leyuan Fang, Shutao Li 0001, David Cunefare, Sina Farsiu
IEEE Trans. Medical Imaging4
2015 Tree Topology Estimation
abstract
Tree-like structures are fundamental in nature, and it is often useful to reconstruct the topology of a tree - what connects to what - from a two-dimensional image of it. However, the projected branches often cross in the image: the tree projects to a planar graph, and the inverse problem of reconstructing the topology of the tree from that of the graph is ill-posed. We regularize this problem with a generative, parametric tree-growth model. Under this model, reconstruction is possible in linear time if one knows the direction of each edge in the graph - which edge endpoint is closer to the root of the tree - but becomes NP-hard if the directions are not known. For the latter case, we present a heuristic search algorithm to estimate the most likely topology of a rooted, three-dimensional tree from a single two-dimensional image. Experimental results on retinal vessel, plant root, and synthetic tree data sets show that our methodology is both accurate and efficient.
Rolando Estrada, Carlo Tomasi, Scott C. Schmidler, Sina Farsiu
IEEE Trans. Pattern Anal. Mach. Intell.4
2015 Retinal Artery-Vein Classification via Topology Estimation
abstract
We propose a novel, graph-theoretic framework for distinguishing arteries from veins in a fundus image. We make use of the underlying vessel topology to better classify small and midsized vessels. We extend our previously proposed tree topology estimation framework by incorporating expert, domain-specific features to construct a simple, yet powerful global likelihood model. We efficiently maximize this model by iteratively exploring the space of possible solutions consistent with the projected vessels. We tested our method on four retinal datasets and achieved classification accuracies of 91.0%, 93.5%, 91.7%, and 90.9%, outperforming existing methods. Our results show the effectiveness of our approach, which is capable of analyzing the entire vasculature, including peripheral vessels, in wide field-of-view fundus photographs. This topology-based method is a potentially important tool for diagnosing diseases with retinal vascular manifestation.
Rolando Estrada, Michael J. Allingham, Priyatham S. Mettu, Scott W. Cousins, Carlo Tomasi, Sina Farsiu
IEEE Trans. Medical Imaging6
2015 3-D Adaptive Sparsity Based Image Compression With Applications to Optical Coherence Tomography
abstract
We present a novel general-purpose compression method for tomographic images, termed 3D adaptive sparse representation based compression (3D-ASRC). In this paper, we focus on applications of 3D-ASRC for the compression of ophthalmic 3D optical coherence tomography (OCT) images. The 3D-ASRC algorithm exploits correlations among adjacent OCT images to improve compression performance, yet is sensitive to preserving their differences. Due to the inherent denoising mechanism of the sparsity based 3D-ASRC, the quality of the compressed images are often better than the raw images they are based on. Experiments on clinical-grade retinal OCT images demonstrate the superiority of the proposed 3D-ASRC over other well-known compression methods.
Leyuan Fang, Shutao Li 0001, Xudong Kang, Joseph A. Izatt, Sina Farsiu
IEEE Trans. Medical Imaging5
2013 Fast Acquisition and Reconstruction of Optical Coherence Tomography Images via Sparse Representation
abstract
In this paper, we present a novel technique, based on compressive sensing principles, for reconstruction and enhancement of multi-dimensional image data. Our method is a major improvement and generalization of the multi-scale sparsity based tomographic denoising (MSBTD) algorithm we recently introduced for reducing speckle noise. Our new technique exhibits several advantages over MSBTD, including its capability to simultaneously reduce noise and interpolate missing data. Unlike MSBTD, our new method does not require an a priori high-quality image from the target imaging subject and thus offers the potential to shorten clinical imaging sessions. This novel image restoration method, which we termed sparsity based simultaneous denoising and interpolation (SBSDI), utilizes sparse representation dictionaries constructed from previously collected datasets. We tested the SBSDI algorithm on retinal spectral domain optical coherence tomography images captured in the clinic. Experiments showed that the SBSDI algorithm qualitatively and quantitatively outperforms other state-of-the-art methods.
Leyuan Fang, Shutao Li 0001, Ryan P. McNabb, Qing Nie, Anthony N. Kuo, Cynthia A. Toth, Joseph A. Izatt, Sina Farsiu
IEEE Trans. Medical Imaging8
2010 Efficient Fourier-Wavelet Super-Resolution
abstract
Super-resolution (SR) is the process of combining multiple aliased low-quality images to produce a high-resolution high-quality image. Aside from registration and fusion of low-resolution images, a key process in SR is the restoration and denoising of the fused images. We present a novel extension of the combined Fourier-wavelet deconvolution and denoising algorithm ForWarD to the multiframe SR application. Our method first uses a fast Fourier-base multiframe image restoration to produce a sharp, yet noisy estimate of the high-resolution image. Our method then applies a space-variant nonlinear wavelet thresholding that addresses the nonstationarity inherent in resolution-enhanced fused images. We describe a computationally efficient method for implementing this space-variant processing that leverages the efficiency of the fast Fourier transform (FFT) to minimize complexity. Finally, we demonstrate the effectiveness of this algorithm for regular imagery as well as in digital mammography.
M. Dirk Robinson, Cynthia A. Toth, Joseph Y. Lo, Sina Farsiu
IEEE Trans. Image Process.4
2009 Optimal Registration Of Aliased Images Using Variable Projection With Applications To Super-Resolution
abstract
Accurate registration of images is the most important and challenging aspect of multiframe image restoration problems such as super-resolution. The accuracy of super-resolution algorithms is quite often limited by the ability to register a set of low-resolution images. The main challenge in registering such images is the presence of aliasing. In this paper, we analyse the problem of jointly registering a set of aliased images and its relationship to super-resolution. We describe a statistically optimal approach to multiframe registration which exploits the concept of variable projections to achieve very efficient algorithms. Finally, we demonstrate how the proposed algorithm offers accurate estimation under various conditions when standard approaches fail to provide sufficient accuracy for super-resolution.
M. Dirk Robinson, Sina Farsiu, Peyman Milanfar
Comput. J.2
2008 Efficient restoration and enhancement of super-resolved X-ray images
abstract
Our previous work demonstrates the ability to reconstruct a single higher resolution image from fusing a collection of multiple extremely low-dosage aliased X-ray images. While this computationally efficient method eliminates aliasing artifacts associated with undersampling, it does not address the problem of deblurring the reconstructed image. In this paper, we present a fast nonlinear deblurring algorithm, specifically designed to address the nonstationary noise associated with multiframe reconstructed images. The algorithm uses a combination of Fourier sharpening and wavelet denoising similar to the ForWarD algorithm. Experimental results on enhancing digital mammogram images attest to the effectiveness of the presented method.
M. Dirk Robinson, Sina Farsiu, Joseph Y. Lo, Cynthia A. Toth
ICIP2
2008 Deblurring Using Regularized Locally Adaptive Kernel Regression
abstract
Kernel regression is an effective tool for a variety of image processing tasks such as denoising and interpolation [1]. In this paper, we extend the use of kernel regression for deblurring applications. In some earlier examples in the literature, such nonparametric deblurring was suboptimally performed in two sequential steps, namely denoising followed by deblurring. In contrast, our optimal solution jointly denoises and deblurs images. The proposed algorithm takes advantage of an effective and novel image prior that generalizes some of the most popular regularization techniques in the literature. Experimental results demonstrate the effectiveness of our method.
Hiroyuki Takeda, Sina Farsiu, Peyman Milanfar
IEEE Trans. Image Process.2
2007 Multi-Scale Statistical Detection and Ballistic Imaging Through Turbid Media
abstract
We exploit recent advances in the physical design of fast optical systems which enable active imaging with "ballistic" light. In this modality, fast bursts of optical energy are propagated into a medium, and the ballistic component of light (which travels with minimal diffusive distortion) is detected after transmission through the target and the medium. To improve the detection rate of the common single pixel optimal detectors, we exploit sampling at a diversity of locations in space, and develop a multi-scale algorithm based upon the generalized likelihood ratio test (GLRT) framework, which takes advantage of the spatial correlation of nearby samples. Experimental results show that objects of different size and shape that are completely unrecognizable using the common single pixel detection techniques, are detectable with very high accuracy using the said multi-scale GLRT technique.
Sina Farsiu, Peyman Milanfar
ICIP (3)1
2007 Kernel Regression for Image Processing and Reconstruction
abstract
In this paper, we make contact with the field of nonparametric statistics and present a development and generalization of tools and results for use in image processing and reconstruction. In particular, we adapt and expand kernel regression ideas for use in image denoising, upscaling, interpolation, fusion, and more. Furthermore, we establish key relationships with some popular existing methods and show how several of these algorithms, including the recently popularized bilateral filter, are special cases of the proposed framework. The resulting algorithms and analyses are amply illustrated with practical examples.
Hiroyuki Takeda, Sina Farsiu, Peyman Milanfar
IEEE Trans. Image Process.2
2006 Robust Kernel Regression for Restoration and Reconstruction of Images from Sparse Noisy Data
abstract
We introduce a class of robust non-parametric estimation methods which are ideally suited for the reconstruction of signals and images from noise-corrupted or sparsely collected samples. The filters derived from this class are locally adapted kernels which take into account both the local density of the available samples, and the actual values of these samples. As such, they are automatically steered and adapted to both the given sampling "geometry", and the samples' "radiometry". As the framework we proposed does not rely upon specific assumptions about noise or sampling distributions, it is applicable to a wide class of problems including efficient image upscaling, high quality reconstruction of an image from as little as 15% of its (irregularly sampled) pixels, super-resolution from noisy and under-determined data sets, state of the art denoising of images corrupted by Gaussian and other noise, effective removal of compression artifacts; and more.
Hiroyuki Takeda, Sina Farsiu, Peyman Milanfar
ICIP2
2006 Multiframe demosaicing and super-resolution of color images
abstract
In the last two decades, two related categories of problems have been studied independently in image restoration literature: super-resolution and demosaicing. A closer look at these problems reveals the relation between them, and, as conventional color digital cameras suffer from both low-spatial resolution and color-filtering, it is reasonable to address them in a unified context. In this paper, we propose a fast and robust hybrid method of super-resolution and demosaicing, based on a maximum a posteron estimation technique by minimizing a multiterm cost function. The L1 norm is used for measuring the difference between the projected estimate of the high-resolution image and each low-resolution image, removing outliers in the data and errors due to possibly inaccurate motion estimation. Bilateral regularization is used for spatially regularizing the luminance component, resulting in sharp edges and forcing interpolation along the edges and not across them. Simultaneously, Tikhonov regularization is used to smooth the chrominance components. Finally, an additional regularization term is used to force similar edge location and orientation in different color channels. We show that the minimization of the total cost function is relatively easy and fast. Experimental results on synthetic and real data sets confirm the effectiveness of our method.
Sina Farsiu, Michael Elad, Peyman Milanfar
IEEE Trans. Image Process.1
2004 Fast and robust multiframe super resolution
abstract
Super-resolution reconstruction produces one or a set of high-resolution images from a set of low-resolution images. In the last two decades, a variety of super-resolution methods have been proposed. These methods are usually very sensitive to their assumed model of data and noise, which limits their utility. This paper reviews some of these methods and addresses their short-comings. We propose an alternate approach using L1 norm minimization and robust regularization based on a bilateral prior to deal with different data and noise models. This computationally inexpensive method is robust to errors in motion and blur estimation and results in images with sharp edges. Simulation results confirm the effectiveness of our method and demonstrate its superiority to other super-resolution methods.
Sina Farsiu, M. Dirk Robinson, Michael Elad, Peyman Milanfar
IEEE Trans. Image Process.1
2003 Fast and robust super-resolution
abstract
In the last two decades, many papers have been published, proposing a variety methods of multiframe resolution enhancement. These methods are usually very sensitive to their assumed model of data and noise, which limits their utility. This paper reviews some of these methods and addresses their shortcomings. We propose a different implementation using L/sub 1/ norm minimization and robust regularization to deal with different data and noise models. This computationally inexpensive method is robust to errors in motion and blur estimation, and results in sharp edges. Simulation results confirm the effectiveness of our method and demonstrate its superiority to other robust super-resolution methods.
Sina Farsiu, M. Dirk Robinson, Michael Elad, Peyman Milanfar
ICIP (2)1