VLDB 2026 Research / reviewers in the wild / expert
Sina Farsiu
dblp:34/2309
· DBLP profile ↗
29ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-4872-2902ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
8 papers |
Image and video processing · 77% Geometric modeling and processing · 20% Virtual and augmented reality · 3% | |
| Artificial intelligence
4 papers |
Segmentation and scene understanding · 90% 3D vision · 10% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing
image restoration |
2.0 | 5 | 2025 | RUN: Reversible Unfolding Network for Concealed Object Segmentation · ICML 2025 Reti-Diff: Illumination Degradation Image Restoration with Retinex-based Latent Diffusion Model · ICLR 2025 Efficient Fourier-Wavelet Super-Resolution · IEEE Trans. Image Process. 2010 |
Computer vision › Segmentation and scene understanding › semantic segmentation
weakly supervised semantic segmentation |
0.9 | 1 | 2025 | Segment Concealed Objects With Incomplete Supervision · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Image and video processing › compressive sensing
deep unfolding network |
0.9 | 1 | 2025 | RUN: Reversible Unfolding Network for Concealed Object Segmentation · ICML 2025 |
Computer vision › Segmentation and scene understanding
medical image segmentation |
0.7 | 1 | 2023 | Directional Connectivity-based Segmentation of Medical Images · CVPR 2023 |
Geometric modeling and processing
3d reconstruction |
0.5 | 1 | 2021 | Mesoscopic Photogrammetry With an Unstabilized Phone Camera · CVPR 2021 |
Geometric modeling and processing › 3d reconstruction
photogrammetry |
0.5 | 1 | 2021 | Mesoscopic Photogrammetry With an Unstabilized Phone Camera · CVPR 2021 |
Image and video processing › super-resolution
multi-frame super-resolution |
0.2 | 3 | 2010 | Efficient Fourier-Wavelet Super-Resolution · IEEE Trans. Image Process. 2010 Multiframe demosaicing and super-resolution of color images · IEEE Trans. Image Process. 2006 Fast and robust multiframe super resolution · IEEE Trans. Image Process. 2004 |
Computer vision › 3D vision
3d reconstruction |
0.2 | 1 | 2015 | Tree Topology Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2015 |
Computer vision › 3D vision
structure from motion |
0.2 | 1 | 2015 | Tree Topology Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2015 |
Image and video processing › image restoration
image denoising |
0.2 | 2 | 2010 | Efficient Fourier-Wavelet Super-Resolution · IEEE Trans. Image Process. 2010 Kernel Regression for Image Processing and Reconstruction · IEEE Trans. Image Process. 2007 |
Virtual and augmented reality › tracking
camera pose estimation |
0.1 | 1 | 2021 | Mesoscopic Photogrammetry With an Unstabilized Phone Camera · CVPR 2021 |
Image and video processing › image restoration
image deblurring |
0.1 | 2 | 2010 | Deblurring Using Regularized Locally Adaptive Kernel Regression · IEEE Trans. Image Process. 2008 Efficient Fourier-Wavelet Super-Resolution · IEEE Trans. Image Process. 2010 |
Image and video processing
super-resolution |
0.1 | 1 | 2010 | Efficient Fourier-Wavelet Super-Resolution · IEEE Trans. Image Process. 2010 |
Image and video processing › image restoration › image denoising
wavelet-based denoising |
0.1 | 1 | 2010 | Efficient Fourier-Wavelet Super-Resolution · IEEE Trans. Image Process. 2010 |
Image and video processing › image restoration › multi-task image restoration
denoising and deblurring |
0.1 | 1 | 2008 | Deblurring Using Regularized Locally Adaptive Kernel Regression · IEEE Trans. Image Process. 2008 |
Image and video processing
image fusion |
0.1 | 1 | 2007 | Kernel Regression for Image Processing and Reconstruction · IEEE Trans. Image Process. 2007 |
Image and video processing › image restoration › multichannel image restoration
color image restoration |
0.1 | 1 | 2006 | Multiframe demosaicing and super-resolution of color images · IEEE Trans. Image Process. 2006 |
Image and video processing › image restoration
demosaicing |
0.1 | 1 | 2006 | Multiframe demosaicing and super-resolution of color images · IEEE Trans. Image Process. 2006 |
Methods — techniques the papers use, named apart from their topics
state space model · 1.7reversible network · 1.7transformer · 0.9segment anything model · 0.9pseudo-labeling · 0.9mean-teacher framework · 0.9latent diffusion · 0.9feature grouping · 0.9diffusion model · 0.9latent space disentanglement · 0.7deep network · 0.7untrained encoder-decoder · 0.5nonparametric distortion model · 0.5convolutional neural network · 0.5heuristic search · 0.2generative tree-growth model · 0.2kernel regression · 0.2l1-norm minimization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spatial coherence loss: All objects matter in salient and camouflaged object detection Spatial coherence loss: All objects matter in salient and camouflaged object detection
Ziyun Yang, Kevin Choy, Sina Farsiu |
Pattern Recognit. | 3 |
| 2025 | Reti-Diff: Illumination Degradation Image Restoration with Retinex-based Latent Diffusion ModelabstractIllumination degradation image restoration (IDIR) techniques aim to improve the visibility of degraded images and mitigate the adverse effects of deteriorated illumination. Among these algorithms, diffusion-based models (DM) have shown promising performance but are often burdened by heavy computational demands and pixel misalignment issues when predicting the image-level distribution. To tackle these problems, we propose to leverage DM within a compact latent space to generate concise guidance priors and introduce a novel solution called Reti-Diff for the IDIR task. Specifically, Reti-Diff comprises two significant components: the Retinex-based latent DM (RLDM) and the Retinex-guided transformer (RGformer). RLDM is designed to acquire Retinex knowledge, extracting reflectance and illumination priors to facilitate detailed reconstruction and illumination correction. RGformer subsequently utilizes these compact priors to guide the decomposition of image features into their respective reflectance and illumination components. Following this, RGformer further enhances and consolidates these decomposed features, resulting in the production of refined images with consistent content and robustness to handle complex degradation scenarios. Extensive experiments demonstrate that Reti-Diff outperforms existing methods on three IDIR tasks, as well as downstream applications. Chunming He, Chengyu Fang 0001, Yulun Zhang 0001, Longxiang Tang, Jinfa Huang, Kai Li 0012, Zhenhua Guo 0001, Xiu Li 0001, Sina Farsiu |
ICLR | 9 |
| 2025 | RUN: Reversible Unfolding Network for Concealed Object SegmentationabstractConcealed object segmentation (COS) is a challenging problem that focuses on identifying objects that are visually blended into their background. Existing methods often employ reversible strategies to concentrate on uncertain regions but only focus on the mask level, overlooking the valuable of the RGB domain. To address this, we propose a Reversible Unfolding Network (RUN) in this paper. RUN formulates the COS task as a foreground-background separation process and incorporates an extra residual sparsity constraint to minimize segmentation uncertainties. The optimization solution of the proposed model is unfolded into a multistage network, allowing the original fixed parameters to become learnable. Each stage of RUN consists of two reversible modules: the Segmentation-Oriented Foreground Separation (SOFS) module and the Reconstruction-Oriented Background Extraction (ROBE) module. SOFS applies the reversible strategy at the mask level and introduces Reversible State Space to capture non-local information. ROBE extends this to the RGB domain, employing a reconstruction network to address conflicting foreground and background regions identified as distortion-prone areas, which arise from their separate estimation by independent modules. As the stages progress, RUN gradually facilitates reversible modeling of foreground and background in both the mask and RGB domains, reducing false-positive and false-negative regions. Extensive experiments demonstrate the superior performance of RUN and underscore the promise of unfolding-based frameworks for COS and other high-level vision tasks. Code is available at https://github.com/ChunmingHe/RUN. Chunming He, Rihan Zhang, Fengyang Xiao, Chengyu Fang 0001, Longxiang Tang, Yulun Zhang 0001, Linghe Kong, Deng-Ping Fan, Kai Li 0012, Sina Farsiu |
ICML | 10 |
| 2025 | Segment Concealed Objects With Incomplete SupervisionabstractIncompletely-Supervised Concealed Object Segmentation (ISCOS) involves segmenting objects that seamlessly blend into their surrounding environments, utilizing incompletely annotated data, such as weak and semi-annotations, for model training. This task remains highly challenging due to (1) the limited supervision provided by the incompletely annotated training data, and (2) the difficulty of distinguishing concealed objects from the background, which arises from the intrinsic similarities in concealed scenarios. In this paper, we introduce the first unified method for ISCOS to address these challenges. To tackle the issue of incomplete supervision, we propose a unified mean-teacher framework, SEE, that leverages the vision foundation model, "Segment Anything Model (SAM)", to generate pseudo-labels using coarse masks produced by the teacher model as prompts. To mitigate the effect of low-quality segmentation masks, we introduce a series of strategies for pseudo-label generation, storage, and supervision. These strategies aim to produce informative pseudo-labels, store the best pseudo-labels generated, and select the most reliable components to guide the student model, thereby ensuring robust network training. Additionally, to tackle the issue of intrinsic similarity, we design a hybrid-granularity feature grouping module that groups features at different granularities and aggregates these results. By clustering similar features, this module promotes segmentation coherence, facilitating more complete segmentation for both single-object and multiple-object images. We validate the effectiveness of our approach across multiple ISCOS tasks, and experimental results demonstrate that our method achieves state-of-the-art performance. Furthermore, SEE can serve as a plug-and-play solution, enhancing the performance of existing models. Chunming He, Kai Li 0012, Yachao Zhang 0001, Ziyun Yang, Youwei Pang, Longxiang Tang, Chengyu Fang 0001, Yulun Zhang 0001, Linghe Kong, Xiu Li 0001, Sina Farsiu |
IEEE Trans. Pattern Anal. Mach. Intell. | 11 |
| 2023 | Directional Connectivity-based Segmentation of Medical ImagesabstractAnatomical consistency in biomarker segmentation is crucial for many medical image analysis tasks. A promising paradigm for achieving anatomically consistent segmentation via deep networks is incorporating pixel connectivity, a basic concept in digital topology, to model inter-pixel relationships. However, previous works on connectivity modeling have ignored the rich channel-wise directional information in the latent space. In this work, we demonstrate that effective disentanglement of directional sub-space from the shared latent space can significantly enhance the feature representation in the connectivity-based network. To this end, we propose a directional connectivity modeling scheme for segmentation that decouples, tracks, and utilizes the directional information across the network. Experiments on various public medical image segmentation benchmarks show the effectiveness of our model as compared to the state-of-the-art methods. Code is available at https://github.com/Zyun-Y/DconnNet. Ziyun Yang, Sina Farsiu |
CVPR | 2 |
| 2023 | RetiFluidNet: A Self-Adaptive and Multi-Attention Deep Convolutional Network for Retinal OCT Fluid SegmentationabstractOptical coherence tomography (OCT) helps ophthalmologists assess macular edema, accumulation of fluids, and lesions at microscopic resolution. Quantification of retinal fluids is necessary for OCT-guided treatment management, which relies on a precise image segmentation step. As manual analysis of retinal fluids is a time-consuming, subjective, and error-prone task, there is increasing demand for fast and robust automatic solutions. In this study, a new convolutional neural architecture named RetiFluidNet is proposed for multi-class retinal fluid segmentation. The model benefits from hierarchical representation learning of textural, contextual, and edge features using a new self-adaptive dual-attention (SDA) module, multiple self-adaptive attention-based skip connections (SASC), and a novel multi-scale deep self-supervision learning (DSL) scheme. The attention mechanism in the proposed SDA module enables the model to automatically extract deformation-aware representations at different levels, and the introduced SASC paths further consider spatial-channel interdependencies for concatenation of counterpart encoder and decoder units, which improve representational capability. RetiFluidNet is also optimized using a joint loss function comprising a weighted version of dice overlap and edge-preserved connectivity-based losses, where several hierarchical stages of multi-scale local losses are integrated into the optimization process. The model is validated based on three publicly available datasets: RETOUCH, OPTIMA, and DUKE, with comparisons against several baselines. Experimental results on the datasets prove the effectiveness of the proposed model in retinal OCT fluid segmentation and reveal that the suggested method is more effective than existing state-of-the-art fluid segmentation algorithms in adapting to retinal OCT scans recorded by various image scanning instruments. Reza Rasti, Armin Biglari, Mohammad Rezapourian, Ziyun Yang, Sina Farsiu |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Modeling extremes with d-max-decreasing neural networksabstractWe propose a neural network architecture that enables non-parametric calibration and generation of multivariate extreme value distributions (MEVs). MEVs arise from Extreme Value Theory (EVT) as the necessary class of models when extrapolating a distributional fit over large spatial and temporal scales based on data observed in intermediate scales. In turn, EVT dictates that $d$-max-decreasing, a stronger form of convexity, is an essential shape constraint in the characterization of MEVs. As far as we know, our proposed architecture provides the first class of non-parametric estimators for MEVs that preserve these essential shape constraints. We show that the architecture approximates the dependence structure encoded by MEVs at parametric rate. Moreover, we present a new method for sampling high-dimensional MEVs using a generative model. We demonstrate our methodology on a wide range of experimental settings, ranging from environmental sciences to financial mathematics and verify that the structural properties of MEVs are retained compared to existing methods. Ali Hasan, Khalil Elkhalil, Yuting Ng, João M. Pereira 0002, Sina Farsiu, Jose H. Blanchet, Vahid Tarokh |
UAI | 5 |
| 2022 | BiconNet: An edge-preserved connectivity-based approach for salient object detection
Ziyun Yang, Somayyeh Soltanian-Zadeh, Sina Farsiu |
Pattern Recognit. | 3 |
| 2021 | Fisher Auto-EncodersabstractIt has been conjectured that the Fisher divergence is more robust to model uncertainty than the conventional Kullback-Leibler (KL) divergence. This motivates the design of a new class of robust generative auto-encoders (AE) referred to as Fisher auto-encoders. Our approach is to design Fisher AEs by minimizing the Fisher divergence between the intractable joint distribution of observed data and latent variables, with that of the postulated/modeled joint distribution. In contrast to KL-based variational AEs (VAEs), the Fisher AE can exactly quantify the distance between the true and the model-based posterior distributions. Qualitative and quantitative results are provided on both MNIST and celebA datasets demonstrating the competitive performance of Fisher AEs in terms of robustness compared to other AEs such as VAEs and Wasserstein AEs. Khalil Elkhalil, Ali Hasan, Jie Ding 0002, Sina Farsiu, Vahid Tarokh |
AISTATS | 4 |
| 2021 | Mesoscopic Photogrammetry With an Unstabilized Phone CameraabstractWe present a feature-free photogrammetric technique that enables quantitative 3D mesoscopic (mm-scale height variation) imaging with tens-of-micron accuracy from sequences of images acquired by a smartphone at close range (several cm) under freehand motion without additional hardware. Our end-to-end, pixel-intensity-based approach jointly registers and stitches all the images by estimating a coaligned height map, which acts as a pixel-wise radial deformation field that orthorectifies each camera image to allow plane-plus-parallax registration. The height maps themselves are reparameterized as the output of an untrained encoder-decoder convolutional neural network (CNN) with the raw camera images as the input, which effectively removes many reconstruction artifacts. Our method also jointly estimates both the camera’s dynamic 6D pose and its distortion using a nonparametric model, the latter of which is especially important in mesoscopic applications when using cameras not designed for imaging at short working distances, such as smartphone cameras. We also propose strategies for reducing computation time and memory, applicable to other multi-frame registration problems. Finally, we demonstrate our method using sequences of multi-megapixel images captured by an un-stabilized smartphone on a variety of samples (e.g., painting brushstrokes, circuit board, seeds). Kevin C. Zhou, Colin L. V. Cooke, Jaehee Park, Ruobing Qian, Roarke Horstmeyer, Joseph A. Izatt, Sina Farsiu |
CVPR | 7 |
| 2021 | Open-Source Automatic Segmentation of Ocular Structures and Biomarkers of Microbial Keratitis on Slit-Lamp Photography Images Using Deep LearningabstractWe propose a fully-automatic deep learning-based algorithm for segmentation of ocular structures and microbial keratitis (MK) biomarkers on slit-lamp photography (SLP) images. The dataset consisted of SLP images from 133 eyes with manual annotations by a physician, P1. A modified region-based convolutional neural network, SLIT-Net, was developed and trained using P1's annotations to identify and segment four pathological regions of interest (ROIs) on diffuse white light images (stromal infiltrate (SI), hypopyon, white blood cell (WBC) border, corneal edema border), one pathological ROI on diffuse blue light images (epithelial defect (ED)), and two non-pathological ROIs on all images (corneal limbus, light reflexes). To assess inter-reader variability, 75 eyes were manually annotated for pathological ROIs by a second physician, P2. Performance was evaluated using the Dice similarity coefficient (DSC) and Hausdorff distance (HD). Using seven-fold cross-validation, the DSC of the algorithm (as compared to P1) for all ROIs was good (range: 0.62-0.95) on all 133 eyes. For the subset of 75 eyes with manual annotations by P2, the DSC for pathological ROIs ranged from 0.69-0.85 (SLIT-Net) vs. 0.37-0.92 (P2). DSCs for SLIT-Net were not significantly different than P2 for segmenting hypopyons (p > 0.05) and higher than P2 for WBCs (p < 0.001) and edema (p < 0.001). DSCs were higher for P2 for segmenting SIs (p < 0.001) and EDs (p < 0.001). HDs were lower for P2 for segmenting SIs (p = 0.005) and EDs (p < 0.001) and not significantly different for hypopyons (p > 0.05), WBCs (p > 0.05), and edema (p > 0.05). This prototype fully-automatic algorithm to segment MK biomarkers on SLP images performed to expectations on an exploratory dataset and holds promise for quantification of corneal physiology and pathology. Jessica Loo, Matthias F. Kriegel, Megan M. Tuohy, Kyeong Hwan Kim, Venkatesh Prajna, Maria A. Woodward, Sina Farsiu |
IEEE J. Biomed. Health Informatics | 7 |
| 2020 | Learning Partial Differential Equations From Data Using Neural NetworksabstractWe develop a framework for estimating unknown partial differential equations (PDEs) from noisy data, using a deep learning approach. Given noisy samples of a solution to an unknown PDE, our method interpolates the samples using a neural network, and extracts the PDE by equating derivatives of the neural network approximation. Our method applies to PDEs which are linear combinations of user-defined dictionary functions, and generalizes previous methods that only consider parabolic PDEs. We introduce a regularization scheme that prevents the function approximation from overfitting the data and forces it to be a solution of the underlying PDE. We validate the model on simulated data generated by the known PDEs and added Gaussian noise, and we study our method under different levels of noise. We also compare the error of our method with a Cramer-Rao lower bound for an ordinary differential equation (ODE). Our results indicate that our method outperforms other methods in estimating PDEs, especially in the low signal-to-noise (SNR) regime. Ali Hasan, João M. Pereira 0002, Robert J. Ravier, Sina Farsiu, Vahid Tarokh |
ICASSP | 4 |
| 2020 | MimickNet, Mimicking Clinical Image Post- Processing Under Black-Box ConstraintsabstractImage post-processing is used in clinical-grade ultrasound scanners to improve image quality (e.g., reduce speckle noise and enhance contrast). These post-processing techniques vary across manufacturers and are generally kept proprietary, which presents a challenge for researchers looking to match current clinical-grade workflows. We introduce a deep learning framework, MimickNet, that transforms conventional delay-and-summed (DAS) beams into the approximate Dynamic Tissue Contrast Enhanced (DTCE™) post-processed images found on Siemens clinical-grade scanners. Training MimickNet only requires post-processed image samples from a scanner of interest without the need for explicit pairing to DAS data. This flexibility allows MimickNet to hypothetically approximate any manufacturer's post-processing without access to the pre-processed data. MimickNet post-processing achieves a 0.940 ± 0.018 structural similarity index measurement (SSIM) compared to clinical-grade post-processing on a 400 cine-loop test set, 0.937 ± 0.025 SSIM on a prospectively acquired dataset, and 0.928 ± 0.003 SSIM on an out-of-distribution cardiac cine-loop after gain adjustment. To our knowledge, this is the first work to establish deep learning models that closely approximate ultrasound post-processing found in current medical practice. MimickNet serves as a clinical post-processing baseline for future works in ultrasound image formation to compare against. Additionally, it can be used as a pretrained model for fine-tuning towards different post-processing techniques. To this end, we have made the MimickNet software, phantom data, and permitted in vivo data open-source at https://github.com/ouwen/MimickNet. Ouwen Huang, Will Long, Nick Bottenus, Marcelo Lerendegui, Gregg E. Trahey, Sina Farsiu, Mark Palmeri |
IEEE Trans. Medical Imaging | 6 |
| 2018 | Statistical Models of Signal and Noise and Fundamental Limits of Segmentation Accuracy in Retinal Optical Coherence TomographyabstractOptical coherence tomography (OCT) has revolutionized diagnosis and prognosis of ophthalmic diseases by visualization and measurement of retinal layers. To speed up the quantitative analysis of disease biomarkers, an increasing number of automatic segmentation algorithms have been proposed to estimate the boundary locations of retinal layers. While the performance of these algorithms has significantly improved in recent years, a critical question to ask is how far we are from a theoretical limit to OCT segmentation performance. In this paper, we present the Cramèr-Rao lower bounds (CRLBs) for the problem of OCT layer segmentation. In deriving the CRLBs, we address the important problem of defining statistical models that best represent the intensity distribution in each layer of the retina. Additionally, we calculate the bounds under an optimal affine bias, reflecting the use of prior knowledge in many segmentation algorithms. Experiments using in vivo images of human retina from a commercial spectral domain OCT system are presented, showing potential for improvement of automated segmentation accuracy. Our general mathematical model can be easily adapted for virtually any OCT system. Furthermore, the statistical models of signal and noise developed in this paper can be utilized for the future improvements of OCT image denoising, reconstruction, and many other applications. Theodore B. Dubose, David Cunefare, Elijah Cole, Peyman Milanfar, Joseph A. Izatt, Sina Farsiu |
IEEE Trans. Medical Imaging | 6 |
| 2017 | Segmentation Based Sparse Reconstruction of Optical Coherence Tomography ImagesabstractWe demonstrate the usefulness of utilizing a segmentation step for improving the performance of sparsity based image reconstruction algorithms. In specific, we will focus on retinal optical coherence tomography (OCT) reconstruction and propose a novel segmentation based reconstruction framework with sparse representation, termed segmentation based sparse reconstruction (SSR). The SSR method uses automatically segmented retinal layer information to construct layer-specific structural dictionaries. In addition, the SSR method efficiently exploits patch similarities within each segmented layer to enhance the reconstruction performance. Our experimental results on clinical-grade retinal OCT images demonstrate the effectiveness and efficiency of the proposed SSR method for both denoising and interpolation of OCT images. Leyuan Fang, Shutao Li 0001, David Cunefare, Sina Farsiu |
IEEE Trans. Medical Imaging | 4 |
| 2015 | Tree Topology EstimationabstractTree-like structures are fundamental in nature, and it is often useful to reconstruct the topology of a tree - what connects to what - from a two-dimensional image of it. However, the projected branches often cross in the image: the tree projects to a planar graph, and the inverse problem of reconstructing the topology of the tree from that of the graph is ill-posed. We regularize this problem with a generative, parametric tree-growth model. Under this model, reconstruction is possible in linear time if one knows the direction of each edge in the graph - which edge endpoint is closer to the root of the tree - but becomes NP-hard if the directions are not known. For the latter case, we present a heuristic search algorithm to estimate the most likely topology of a rooted, three-dimensional tree from a single two-dimensional image. Experimental results on retinal vessel, plant root, and synthetic tree data sets show that our methodology is both accurate and efficient. Rolando Estrada, Carlo Tomasi, Scott C. Schmidler, Sina Farsiu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Retinal Artery-Vein Classification via Topology EstimationabstractWe propose a novel, graph-theoretic framework for distinguishing arteries from veins in a fundus image. We make use of the underlying vessel topology to better classify small and midsized vessels. We extend our previously proposed tree topology estimation framework by incorporating expert, domain-specific features to construct a simple, yet powerful global likelihood model. We efficiently maximize this model by iteratively exploring the space of possible solutions consistent with the projected vessels. We tested our method on four retinal datasets and achieved classification accuracies of 91.0%, 93.5%, 91.7%, and 90.9%, outperforming existing methods. Our results show the effectiveness of our approach, which is capable of analyzing the entire vasculature, including peripheral vessels, in wide field-of-view fundus photographs. This topology-based method is a potentially important tool for diagnosing diseases with retinal vascular manifestation. Rolando Estrada, Michael J. Allingham, Priyatham S. Mettu, Scott W. Cousins, Carlo Tomasi, Sina Farsiu |
IEEE Trans. Medical Imaging | 6 |
| 2015 | 3-D Adaptive Sparsity Based Image Compression With Applications to Optical Coherence TomographyabstractWe present a novel general-purpose compression method for tomographic images, termed 3D adaptive sparse representation based compression (3D-ASRC). In this paper, we focus on applications of 3D-ASRC for the compression of ophthalmic 3D optical coherence tomography (OCT) images. The 3D-ASRC algorithm exploits correlations among adjacent OCT images to improve compression performance, yet is sensitive to preserving their differences. Due to the inherent denoising mechanism of the sparsity based 3D-ASRC, the quality of the compressed images are often better than the raw images they are based on. Experiments on clinical-grade retinal OCT images demonstrate the superiority of the proposed 3D-ASRC over other well-known compression methods. Leyuan Fang, Shutao Li 0001, Xudong Kang, Joseph A. Izatt, Sina Farsiu |
IEEE Trans. Medical Imaging | 5 |
| 2013 | Fast Acquisition and Reconstruction of Optical Coherence Tomography Images via Sparse RepresentationabstractIn this paper, we present a novel technique, based on compressive sensing principles, for reconstruction and enhancement of multi-dimensional image data. Our method is a major improvement and generalization of the multi-scale sparsity based tomographic denoising (MSBTD) algorithm we recently introduced for reducing speckle noise. Our new technique exhibits several advantages over MSBTD, including its capability to simultaneously reduce noise and interpolate missing data. Unlike MSBTD, our new method does not require an a priori high-quality image from the target imaging subject and thus offers the potential to shorten clinical imaging sessions. This novel image restoration method, which we termed sparsity based simultaneous denoising and interpolation (SBSDI), utilizes sparse representation dictionaries constructed from previously collected datasets. We tested the SBSDI algorithm on retinal spectral domain optical coherence tomography images captured in the clinic. Experiments showed that the SBSDI algorithm qualitatively and quantitatively outperforms other state-of-the-art methods. Leyuan Fang, Shutao Li 0001, Ryan P. McNabb, Qing Nie, Anthony N. Kuo, Cynthia A. Toth, Joseph A. Izatt, Sina Farsiu |
IEEE Trans. Medical Imaging | 8 |
| 2010 | Efficient Fourier-Wavelet Super-ResolutionabstractSuper-resolution (SR) is the process of combining multiple aliased low-quality images to produce a high-resolution high-quality image. Aside from registration and fusion of low-resolution images, a key process in SR is the restoration and denoising of the fused images. We present a novel extension of the combined Fourier-wavelet deconvolution and denoising algorithm ForWarD to the multiframe SR application. Our method first uses a fast Fourier-base multiframe image restoration to produce a sharp, yet noisy estimate of the high-resolution image. Our method then applies a space-variant nonlinear wavelet thresholding that addresses the nonstationarity inherent in resolution-enhanced fused images. We describe a computationally efficient method for implementing this space-variant processing that leverages the efficiency of the fast Fourier transform (FFT) to minimize complexity. Finally, we demonstrate the effectiveness of this algorithm for regular imagery as well as in digital mammography. M. Dirk Robinson, Cynthia A. Toth, Joseph Y. Lo, Sina Farsiu |
IEEE Trans. Image Process. | 4 |
| 2009 | Optimal Registration Of Aliased Images Using Variable Projection With Applications To Super-ResolutionabstractAccurate registration of images is the most important and challenging aspect of multiframe image restoration problems such as super-resolution. The accuracy of super-resolution algorithms is quite often limited by the ability to register a set of low-resolution images. The main challenge in registering such images is the presence of aliasing. In this paper, we analyse the problem of jointly registering a set of aliased images and its relationship to super-resolution. We describe a statistically optimal approach to multiframe registration which exploits the concept of variable projections to achieve very efficient algorithms. Finally, we demonstrate how the proposed algorithm offers accurate estimation under various conditions when standard approaches fail to provide sufficient accuracy for super-resolution. M. Dirk Robinson, Sina Farsiu, Peyman Milanfar |
Comput. J. | 2 |
| 2008 | Efficient restoration and enhancement of super-resolved X-ray imagesabstractOur previous work demonstrates the ability to reconstruct a single higher resolution image from fusing a collection of multiple extremely low-dosage aliased X-ray images. While this computationally efficient method eliminates aliasing artifacts associated with undersampling, it does not address the problem of deblurring the reconstructed image. In this paper, we present a fast nonlinear deblurring algorithm, specifically designed to address the nonstationary noise associated with multiframe reconstructed images. The algorithm uses a combination of Fourier sharpening and wavelet denoising similar to the ForWarD algorithm. Experimental results on enhancing digital mammogram images attest to the effectiveness of the presented method. M. Dirk Robinson, Sina Farsiu, Joseph Y. Lo, Cynthia A. Toth |
ICIP | 2 |
| 2008 | Deblurring Using Regularized Locally Adaptive Kernel RegressionabstractKernel regression is an effective tool for a variety of image processing tasks such as denoising and interpolation [1]. In this paper, we extend the use of kernel regression for deblurring applications. In some earlier examples in the literature, such nonparametric deblurring was suboptimally performed in two sequential steps, namely denoising followed by deblurring. In contrast, our optimal solution jointly denoises and deblurs images. The proposed algorithm takes advantage of an effective and novel image prior that generalizes some of the most popular regularization techniques in the literature. Experimental results demonstrate the effectiveness of our method. Hiroyuki Takeda, Sina Farsiu, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2007 | Multi-Scale Statistical Detection and Ballistic Imaging Through Turbid MediaabstractWe exploit recent advances in the physical design of fast optical systems which enable active imaging with "ballistic" light. In this modality, fast bursts of optical energy are propagated into a medium, and the ballistic component of light (which travels with minimal diffusive distortion) is detected after transmission through the target and the medium. To improve the detection rate of the common single pixel optimal detectors, we exploit sampling at a diversity of locations in space, and develop a multi-scale algorithm based upon the generalized likelihood ratio test (GLRT) framework, which takes advantage of the spatial correlation of nearby samples. Experimental results show that objects of different size and shape that are completely unrecognizable using the common single pixel detection techniques, are detectable with very high accuracy using the said multi-scale GLRT technique. Sina Farsiu, Peyman Milanfar |
ICIP (3) | 1 |
| 2007 | Kernel Regression for Image Processing and ReconstructionabstractIn this paper, we make contact with the field of nonparametric statistics and present a development and generalization of tools and results for use in image processing and reconstruction. In particular, we adapt and expand kernel regression ideas for use in image denoising, upscaling, interpolation, fusion, and more. Furthermore, we establish key relationships with some popular existing methods and show how several of these algorithms, including the recently popularized bilateral filter, are special cases of the proposed framework. The resulting algorithms and analyses are amply illustrated with practical examples. Hiroyuki Takeda, Sina Farsiu, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2006 | Robust Kernel Regression for Restoration and Reconstruction of Images from Sparse Noisy DataabstractWe introduce a class of robust non-parametric estimation methods which are ideally suited for the reconstruction of signals and images from noise-corrupted or sparsely collected samples. The filters derived from this class are locally adapted kernels which take into account both the local density of the available samples, and the actual values of these samples. As such, they are automatically steered and adapted to both the given sampling "geometry", and the samples' "radiometry". As the framework we proposed does not rely upon specific assumptions about noise or sampling distributions, it is applicable to a wide class of problems including efficient image upscaling, high quality reconstruction of an image from as little as 15% of its (irregularly sampled) pixels, super-resolution from noisy and under-determined data sets, state of the art denoising of images corrupted by Gaussian and other noise, effective removal of compression artifacts; and more. Hiroyuki Takeda, Sina Farsiu, Peyman Milanfar |
ICIP | 2 |
| 2006 | Multiframe demosaicing and super-resolution of color imagesabstractIn the last two decades, two related categories of problems have been studied independently in image restoration literature: super-resolution and demosaicing. A closer look at these problems reveals the relation between them, and, as conventional color digital cameras suffer from both low-spatial resolution and color-filtering, it is reasonable to address them in a unified context. In this paper, we propose a fast and robust hybrid method of super-resolution and demosaicing, based on a maximum a posteron estimation technique by minimizing a multiterm cost function. The L1 norm is used for measuring the difference between the projected estimate of the high-resolution image and each low-resolution image, removing outliers in the data and errors due to possibly inaccurate motion estimation. Bilateral regularization is used for spatially regularizing the luminance component, resulting in sharp edges and forcing interpolation along the edges and not across them. Simultaneously, Tikhonov regularization is used to smooth the chrominance components. Finally, an additional regularization term is used to force similar edge location and orientation in different color channels. We show that the minimization of the total cost function is relatively easy and fast. Experimental results on synthetic and real data sets confirm the effectiveness of our method. Sina Farsiu, Michael Elad, Peyman Milanfar |
IEEE Trans. Image Process. | 1 |
| 2004 | Fast and robust multiframe super resolutionabstractSuper-resolution reconstruction produces one or a set of high-resolution images from a set of low-resolution images. In the last two decades, a variety of super-resolution methods have been proposed. These methods are usually very sensitive to their assumed model of data and noise, which limits their utility. This paper reviews some of these methods and addresses their short-comings. We propose an alternate approach using L1 norm minimization and robust regularization based on a bilateral prior to deal with different data and noise models. This computationally inexpensive method is robust to errors in motion and blur estimation and results in images with sharp edges. Simulation results confirm the effectiveness of our method and demonstrate its superiority to other super-resolution methods. Sina Farsiu, M. Dirk Robinson, Michael Elad, Peyman Milanfar |
IEEE Trans. Image Process. | 1 |
| 2003 | Fast and robust super-resolutionabstractIn the last two decades, many papers have been published, proposing a variety methods of multiframe resolution enhancement. These methods are usually very sensitive to their assumed model of data and noise, which limits their utility. This paper reviews some of these methods and addresses their shortcomings. We propose a different implementation using L/sub 1/ norm minimization and robust regularization to deal with different data and noise models. This computationally inexpensive method is robust to errors in motion and blur estimation, and results in sharp edges. Simulation results confirm the effectiveness of our method and demonstrate its superiority to other robust super-resolution methods. Sina Farsiu, M. Dirk Robinson, Michael Elad, Peyman Milanfar |
ICIP (2) | 1 |