VLDB 2026 Research / reviewers in the wild / expert
Kyong Hwan Jin
dblp:161/9868
· DBLP profile ↗
35ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0001-7885-4792ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 19 since 2021Artificial intelligence and machine learning · 16 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Theory of computation · 2Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GRAPE (Gaussian Rendering for Accelerated Pixel Enhancement) Brings Fast and Lightweight Arbitrary Super-ResolutionabstractWe present GRAPE—Gaussian Rendering for Accelerated Pixel Enhancement, a fast, lightweight method for arbitrary-scale super-resolution (ASSR) based on 2D Gaussian splatting. Lookup-table (LUT) schemes are limited to preset scale factors and struggle with varied textures, while implicit neural representations (INRs) slow down because they require per-coordinate queries; moreover, prior Gaussian-splatting approaches rely on heavy networks or complex processing. GRAPE overcomes these limitations with a compact design in which a single point-wise layer predicts anisotropic Gaussian parameters—RGB value, rotation, scale, and offset—and a differentiable rasterizer then renders the high-resolution image in one pass. The entire model, including both encoder and decoder, contains just 1.56 M parameters and requires only 1.10 GB of GPU memory, yet achieves 69.33 FPS on Urban100 at ×4 whose average image size is 985 × 798. This is more than 315 × faster than GSASR, a 20.45 M-parameter model that runs at 0.22 FPS. Although GRAPE does not further improve perceptual fidelity over heavier networks, it remains competitively close, providing an attractive quality-efficiency trade-off across Set5, Set14, BSD100, DIV2K and Urban100. Consequently, GRAPE is ideal for resource-limited deployments or interactive applications that require rapid screen updates. The source code will be made publicly available at github.com/mulkkog/GRAPE. Jung In Jang, Kyong Hwan Jin |
WACV | 2 |
| 2025 | Towards Lossless Implicit Neural Representation via Bit Plane DecompositionabstractWe quantify the upper bound on the size of the implicit neural representation (INR) model from a digital perspective. The upper bound of the model size increases exponentially as the required bit-precision increases. To this end, we present a bit-plane decomposition method that makes INR predict bit-planes, producing the same effect as reducing the upper bound of the model size. We validate our hypothesis that reducing the upper bound leads to faster convergence with constant model size. Our method achieves lossless representation in 2D image and audio fitting, even for high bit-depth signals, such as 16-bit, which was previously unachievable. We pioneered the presence of bit bias, which INR prioritizes as the most significant bit (MSB). We expand the application of the INR task to bit depth expansion, lossless image compression, and extreme network quantization. Our source code is available at https://github.com/WooKyoungHan/LosslessINR. Woo Kyoung Han, Byeonghun Lee, Hyunmin Cho, Sunghoon Im 0001, Kyong Hwan Jin |
CVPR | 5 |
| 2025 | BF-STVSR: B-Splines and Fourier - Best Friends for High Fidelity Spatial-Temporal Video Super-ResolutionabstractWhile prior methods in Continuous Spatial-Temporal Video Super-Resolution (C-STVSR) employ Implicit Neural Representation (INR) for continuous encoding, they often struggle to capture the complexity of video data, relying on simple coordinate concatenation and pre-trained optical flow networks for motion representation. Interestingly, we find that adding position encoding, contrary to common observations, does not improve—and even degrades—performance. This issue becomes particularly pronounced when combined with pre-trained optical flow networks, which can limit the model’s flexibility. To address these issues, we propose BF-STVSR, a C-STVSR framework with two key modules tailored to better represent spatial and temporal characteristics of video: 1) B-spline Mapper for smooth temporal interpolation, and 2) Fourier Mapper for capturing dominant spatial frequencies. Our approach achieves state-of-the-art in various metrics, including PSNR and SSIM, showing enhanced spatial details and natural temporal consistency. Our code is available ${\color{Cyan}\text{here}}$. Kyong Hwan Jin, Jaejun Yoo 0001 |
CVPR | 3 |
| 2025 | Identity-preserving Distillation Sampling by Fixed-Point IteratorabstractScore distillation sampling (SDS) demonstrates a powerful capability for text-conditioned 2D image and 3D object generation by distilling the knowledge from learned score functions. However, SDS often suffers from blurriness caused by noisy gradients. When SDS meets the image editing, such degradations can be reduced by adjusting bias shifts using reference pairs, but the de-biasing techniques are still corrupted by erroneous gradients. To this end, we introduce Identity-preserving Distillation Sampling (IDS), which compensates for the gradient leading to undesired changes in the results. Based on the analysis that these errors come from the text-conditioned scores, a new regularization technique, called fixed-point iterative regularization (FPR), is proposed to modify the score itself, driving the preservation of the identity even including poses and structures. Thanks to a self-correction by FPR, the proposed method provides clear and unambiguous representations corresponding to the given prompts in image-to-image editing and editable neural radiance field (NeRF). The structural consistency between the source and the edited data is obviously maintained compared to other state-of-the-art methods. Our code is https://github.com/shhh0620/IDS Seonhwa Kim, Soobin Park, Donghoon Ahn, Seungryong Kim, Kyong Hwan Jin, Eun Ju Cha |
CVPR | 7 |
| 2025 | JPEG Processing Neural Operator for Backward-Compatible CodingabstractDespite significant advances in learning-based lossy compression algorithms, standardizing codecs remains a critical challenge. In this paper, we present the JPEG Processing Neural Operator (JPNeO), a next-generation JPEG algorithm that maintains full backward compatibility with the current JPEG format. Our JPNeO improves chroma component preservation and enhances reconstruction fidelity compared to existing artifact removal methods by incorporating neural operators in both the encoding and decoding stages. JPNeO achieves practical benefits in terms of reduced memory usage and parameter count. We further validate our hypothesis about the existence of a space with high mutual information through empirical evidence. In summary, the JPNeO functions as a high-performance out-of-the-box image compression pipeline without changing source coding's protocol. Our source code is available at https://github.com/WooKyoungHan/JPNeO. Woo Kyoung Han, Yongjun Lee, Byeonghun Lee, Sanghyun Park 0004, Sunghoon Im 0001, Kyong Hwan Jin |
ICCV | 6 |
| 2025 | Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image Diffusion
Pu-Reum Kim, SeonHwa Kim, Soobin Park, Eun Ju Cha, Kyong Hwan Jin |
ICCV | 6 |
| 2025 | Reference-Based Super-Resolution via Image-Based Retrieval-Augmented Generation Diffusion
Byeonghun Lee, Hyunmin Cho, Hong Gyu Choi, Soo Min Kang, Iljun Ahn, Kyong Hwan Jin |
ICCV | 6 |
| 2025 | IM-LUT: Interpolation Mixing Look-Up Tables for Image Super-Resolution
Sejin Park 0002, Kyong Hwan Jin, Seung-Won Jung |
ICCV | 3 |
| 2025 | Efficient one-shot federated learning on medical data using knowledge distillation with image synthesis and client model adaptation
Myeongkyun Kang, Philip Chikontwe, Soopil Kim, Kyong Hwan Jin, Ehsan Adeli-Mosabbeb, Kilian M. Pohl, Sanghyun Park 0004 |
Medical Image Anal. | 4 |
| 2025 | Communication Efficient Federated Learning for Multi-Organ Segmentation via Knowledge Distillation With Image SynthesisabstractFederated learning (FL) methods for multi-organ segmentation in CT scans are gaining popularity, but generally require numerous rounds of parameter exchange between a central server and clients. This repetitive sharing of parameters between server and clients may not be practical due to the varying network infrastructures of clients and the large transmission of data. Further increasing repetitive sharing results from data heterogeneity among clients, i.e., clients may differ with respect to the type of data they share. For example, they might provide label maps of different organs (i.e. partial labels) as segmentations of all organs shown in the CT are not part of their clinical protocol. To this end, we propose an efficient communication approach for FL with partial labels. Specifically, parameters of local models are transmitted once to a central server and the global model is trained via knowledge distillation (KD) of the local models. While one can make use of unlabeled public data as inputs for KD, the model accuracy is often limited due to distribution shifts between local and public datasets. Herein, we propose to generate synthetic images from clients' models as additional inputs to mitigate data shifts between public and local data. In addition, our proposed method offers flexibility for additional finetuning through several rounds of communication using existing FL algorithms, leading to enhanced performance. Extensive evaluation on public datasets in few communication FL scenario reveals that our approach substantially improves over state-of-the-art methods. Soopil Kim, Heejung Park, Philip Chikontwe, Myeongkyun Kang, Kyong Hwan Jin, Ehsan Adeli-Mosabbeb, Kilian M. Pohl, Sanghyun Park 0004 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | JDEC: JPEG Decoding via Enhanced Continuous Cosine CoefficientsabstractWe propose a practical approach to JPEG image de-coding, utilizing a local implicit neural representation with continuous cosine formulation. The JPEG algorithm sig-nificantly quantizes discrete cosine transform (DCT) spec-tra to achieve a high compression rate, inevitably resulting in quality degradation while encoding an image. We have designed a continuous cosine spectrum estimator to address the quality degradation issue that restores the distorted spectrum. By leveraging local DCT formulations, our network has the privilege to exploit dequantization and upsampling simultaneously. Our proposed model enables decoding compressed images directly across different quality factors using a single pre-trained model without relying on a conventional JPEG decoder. As a result, our proposed network achieves state-of-the-art performance in flexible color image JPEG artifact removal tasks. Our source code is available at https://github.com/WooKyoungHan/Jdec. Woo Kyoung Han, Sunghoon Im 0001, Jaedeok Kim, Kyong Hwan Jin |
CVPR | 4 |
| 2024 | Self-rectifying Diffusion Sampling with Perturbed-Attention Guidance
Donghoon Ahn, Hyoungwon Cho, Jaewon Min, Woo-seok Jang, Seonhwa Kim, Hyun Hee Park, Kyong Hwan Jin, Seungryong Kim |
ECCV (73) | 8 |
| 2024 | BurstM: Deep Burst Multi-scale SR Using Fourier Space with Optical Flow
EungGu Kang, Byeonghun Lee, Sunghoon Im 0001, Kyong Hwan Jin |
ECCV (42) | 4 |
| 2024 | BroadBEV: Collaborative LiDAR-camera Fusion for Broad-sighted Bird's Eye View Map ConstructionabstractA recent sensor fusion in a Bird’s Eye View (BEV) space has shown its utility in various tasks such as 3D detection, map segmentation, etc. However, the approach struggles with inaccurate camera BEV estimation, and a perception of distant areas due to the sparsity of LiDAR points. In this paper, we propose a BEV fusion (BroadBEV) that aims to enhance camera BEV estimation for broad perception in the pre-defined BEV range, while simultaneously improving the completion of LiDAR’s sparsity in the entire BEV space. Toward that end, we devise Point-scattering that scatters LiDAR BEV distribution to camera depth distribution. The method boosts the learning of depth estimation of the camera branch and induces accurate location of dense camera features in BEV space. For an effective BEV fusion between the spatially synchronized features, we suggest ColFusion that applies self-attention weights of LiDAR and camera BEV features to each other. Our extensive experiments demonstrate that the suggested methods enable a broad BEV perception with remarkable performance gains. Giseop Kim, Kyong Hwan Jin, Sunwook Choi |
ICRA | 3 |
| 2024 | Learning Residual Elastic Warps for Image Stitching under Dirichlet Boundary ConditionabstractTrendy suggestions for learning-based elastic warps enable the deep image stitchings to align images exposed to large parallax errors. Despite the remarkable alignments, the methods struggle with occasional holes or discontinuity between overlapping and non-overlapping regions of a target image as the applied training strategy mostly focuses on overlap region alignment. As a result, they require additional modules such as seam finder and image inpainting for hiding discontinuity and filling holes, respectively. In this work, we suggest Recurrent Elastic Warps (REwarp) that address the problem with Dirichlet boundary condition and boost performances by residual learning for recurrent misalign correction. Specifically, REwarp predicts a homography and a Thin-plate Spline (TPS) under the boundary constraint for discontinuity and hole-free image stitching. Our experiments show the favorable aligns and the competitive computational costs of REwarp compared to the existing stitching methods. Our source code is available at https://github.com/minshu-kim/REwarp. Yongjun Lee, Woo Kyoung Han, Kyong Hwan Jin |
WACV | 4 |
| 2024 | Implicit Neural Image Stitching With Enhanced and Blended Feature ReconstructionabstractExisting frameworks for image stitching often provide visually reasonable stitchings. However, they suffer from blurry artifacts and disparities in illumination, depth level, etc. Although the recent learning-based stitchings relax such disparities, the required methods impose sacrifice of image qualities failing to capture high-frequency details for stitched images. To address the problem, we propose a novel approach, implicit Neural Image Stitching (NIS) that extends arbitrary-scale super-resolution. Our method estimates Fourier coefficients of images for quality-enhancing warps. Then, the suggested model blends color mismatches and misalignment in the latent space and decodes the features into RGB values of stitched images. Our experiments show that our approach achieves improvement in resolving the low-definition imaging of the previous deep image stitching with favorable accelerated image-enhancing methods. Our source code is available at https://github.com/minshu-kim/NIS. Byeonghun Lee, Sunghoon Im 0001, Kyong Hwan Jin |
WACV | 5 |
| 2024 | Federated learning with knowledge distillation for multi-organ segmentation with partially labeled datasets
Soopil Kim, Heejung Park, Myeongkyun Kang, Kyong Hwan Jin, Ehsan Adeli-Mosabbeb, Kilian M. Pohl, Sanghyun Park 0004 |
Medical Image Anal. | 4 |
| 2024 | FedNN: Federated learning on concept drift data using weight and adaptive group normalizations
Myeongkyun Kang, Soopil Kim, Kyong Hwan Jin, Ehsan Adeli-Mosabbeb, Kilian M. Pohl, Sanghyun Park 0004 |
Pattern Recognit. | 3 |
| 2023 | ABCD : Arbitrary Bitwise Coefficient for De-QuantizationabstractModern displays and contents support more than 8bits image and video. However, bit-starving situations such as compression codecs make low bit-depth (LBD) images (<8bits), occurring banding and blurry artifacts. Previous bit depth expansion (BDE) methods still produce unsatisfactory high bit-depth (HBD) images. To this end, we propose an implicit neural function with a bit query to recover de-quantized images from arbitrarily quantized inputs. We develop a phasor estimator to exploit the information of the nearest pixels. Our method shows superior performance against prior BDE methods on natural and animation images. We also demonstrate our model on YouTube UGC datasets for de-banding. Our source code is available at https://github.com/WooKyoungHan/ABCD Woo Kyoung Han, Byeonghun Lee, Sanghyun Park 0004, Kyong Hwan Jin |
CVPR | 4 |
| 2023 | B-Spline Texture Coefficients Estimator for Screen Content Image Super-ResolutionabstractScreen content images (SCIs) include many informative components, e.g., texts and graphics. Such content creates sharp edges or homogeneous areas, making a pixel distribution of SCI different from the natural image. Therefore, we need to properly handle the edges and textures to minimize information distortion of the contents when a display device's resolution differs from SCIs. To achieve this goal, we propose an implicit neural representation using B-splines for screen content image super-resolution (SCI SR) with arbitrary scales. Our method extracts scaling, translating, and smoothing parameters of B-splines. The followed multilayer perceptron (MLP) uses the estimated B-splines to recover high-resolution SCI. Our network outperforms both a transformer-based reconstruction and an implicit Fourier representation method in almost upscaling factor, thanks to the positive constraint and compact support of the B-spline basis. Moreover, our SR results are recognized as correct text letters with the highest confidence by a pre-trained scene text recognition network. Source code is available at https://github.com/ByeongHyunPak/btc. Byeonghyun Pak, Kyong Hwan Jin |
CVPR | 3 |
| 2023 | One-Shot Federated Learning on Medical Data Using Knowledge Distillation with Image Synthesis and Client Model Adaptation
Myeongkyun Kang, Philip Chikontwe, Soopil Kim, Kyong Hwan Jin, Ehsan Adeli-Mosabbeb, Kilian M. Pohl, Sanghyun Park 0004 |
MICCAI (2) | 4 |
| 2023 | DeepFold: enhancing protein structure prediction through optimized loss functions, improved template features, and re-optimized energy functionabstractMOTIVATION: Predicting protein structures with high accuracy is a critical challenge for the broad community of life sciences and industry. Despite progress made by deep neural networks like AlphaFold2, there is a need for further improvements in the quality of detailed structures, such as side-chains, along with protein backbone structures. RESULTS: Building upon the successes of AlphaFold2, the modifications we made include changing the losses of side-chain torsion angles and frame aligned point error, adding loss functions for side chain confidence and secondary structure prediction, and replacing template feature generation with a new alignment method based on conditional random fields. We also performed re-optimization by conformational space annealing using a molecular mechanics energy function which integrates the potential energies obtained from distogram and side-chain prediction. In the CASP15 blind test for single protein and domain modeling (109 domains), DeepFold ranked fourth among 132 groups with improvements in the details of the structure in terms of backbone, side-chain, and Molprobity. In terms of protein backbone accuracy, DeepFold achieved a median GDT-TS score of 88.64 compared with 85.88 of AlphaFold2. For TBM-easy/hard targets, DeepFold ranked at the top based on Z-scores for GDT-TS. This shows its practical value to the structural biology community, which demands highly accurate structures. In addition, a thorough analysis of 55 domains from 39 targets with publicly available structures indicates that DeepFold shows superior side-chain accuracy and Molprobity scores among the top-performing groups. AVAILABILITY AND IMPLEMENTATION: DeepFold tools are open-source software available at https://github.com/newtonjoo/deepfold. Jae-Won Lee, Jong-Hyun Won, Seonggwang Jeon, Yujin Choo, Yubin Yeon, Jin-Seon Oh, Seonhwa Kim, InSuk Joung, Cheongjae Jang, Sung Jong Lee, Kyong Hwan Jin, Giltae Song, Eun-Sol Kim, Jejoong Yoo, Eunok Paek, Yung-Kyun Noh, Keehyoung Joo |
Bioinform. | 13 |
| 2022 | Local Texture Estimator for Implicit Representation FunctionabstractRecent works with an implicit neural function shed light on representing images in arbitrary resolution. However, a standalone multi-layer perceptron shows limited performance in learning high-frequency components. In this paper, we propose a Local Texture Estimator (LTE), a dominant-frequency estimator for natural images, enabling an implicit function to capture fine details while reconstructing images in a continuous manner. When jointly trained with a deep super-resolution (SR) architecture, LTE is capable of characterizing image textures in 2D Fourier space. We show that an LTE-based neuralfunction achieves favorable performance against existing deep SR methods within an arbitrary-scale factor. Furthermore, we demonstrate that our implementation takes the shortest running time compared to previous works. Kyong Hwan Jin |
CVPR | 2 |
| 2022 | Learning Local Implicit Fourier Representation for Image Warping
Kwangpyo Choi, Kyong Hwan Jin |
ECCV (18) | 3 |
| 2021 | Deep Block Transform for AutoencodersabstractWe discover that a trainable convolution layer with a stride over 1 and kernel ≥ stride is identical to a trainable block transform. A block transform is performed when we use a convolution layer with a stride ≥ 2 and a kernel ≥ the stride. For instance, if we use the same widths, such as a 2×2 convolution kernel and stride-2, there are no overlaps between sliding windows, so this layer operates a block transform on the partitioned 2×2 blocks. A block transform reduces the computational complexity due to a stride ≥ 2. To keep the original size, we apply a transposed convolution (stride $=$ kernel ≥ 2), an adjoint operator of a forward block transform. Based on this relationship, we propose a trainable multi-scale block transform for autoencoders. The proposed method has an encoder consisting of two sequential convolutions with stride-2, a 2×2 kernel, and a decoder consisting of the encoder's two adjoint operators (transposed convolution). Clipping is used for nonlinear activations. Inspired by the zero-frequency element in the dictionary learning method, the proposed method uses DC values for residual learning. The proposed method shows high-resolution representations, whereas the stride-1 convolutional autoencoder with 3×3 kernels generates blurry images. Kyong Hwan Jin |
IEEE Signal Process. Lett. | 1 |
| 2021 | Time-Dependent Deep Image Prior for Dynamic MRIabstractWe propose a novel unsupervised deep-learning-based algorithm for dynamic magnetic resonance imaging (MRI) reconstruction. Dynamic MRI requires rapid data acquisition for the study of moving organs such as the heart. We introduce a generalized version of the deep-image-prior approach, which optimizes the weights of a reconstruction network to fit a sequence of sparsely acquired dynamic MRI measurements. Our method needs neither prior training nor additional data. In particular, for cardiac images, it does not require the marking of heartbeats or the reordering of spokes. The key ingredients of our method are threefold: 1) a fixed low-dimensional manifold that encodes the temporal variations of images; 2) a network that maps the manifold into a more expressive latent space; and 3) a convolutional neural network that generates a dynamic series of MRI images from the latent variables and that favors their consistency with the measurements in k -space. Our method outperforms the state-of-the-art methods quantitatively and qualitatively in both retrospective and real fetal cardiac datasets. To the best of our knowledge, this is the first unsupervised deep-learning-based method that can reconstruct the continuous variation of dynamic MRI sequences with high spatial resolution. Jaejun Yoo 0001, Kyong Hwan Jin, Jérôme Yerly, Matthias Stuber, Michael Unser |
IEEE Trans. Medical Imaging | 2 |
| 2018 | Sparse and Low-Rank Decomposition of a Hankel Structured Matrix for Impulse Noise RemovalabstractRecently, the annihilating filter-based low-rank Hankel matrix (ALOHA) approach was proposed as a powerful image inpainting method. Based on the observation that smoothness or textures within an image patch correspond to sparse spectral components in the frequency domain, ALOHA exploits the existence of annihilating filters and the associated rank-deficient Hankel matrices in an image domain to estimate any missing pixels. By extending this idea, we propose a novel impulse-noise removal algorithm that uses the sparse and low-rank decomposition of a Hankel structured matrix. This method, referred to as the robust ALOHA, is based on the observation that an image corrupted with the impulse noise has intact pixels; consequently, the impulse noise can be modeled as sparse components, whereas the underlying image can still be modeled using a low-rank Hankel structured matrix. To solve the sparse and low-rank matrix decomposition problem, we propose an alternating direction method of multiplier approach, with initial factorized matrices coming from a low-rank matrix-fitting algorithm. To adapt local image statistics that have distinct spectral distributions, the robust ALOHA is applied in a patch-by-patch manner. Experimental results from impulse noise for both single-channel and multichannel color images demonstrate that the robust ALOHA is superior to existing approaches, especially during the reconstruction of complex texture patterns. Kyong Hwan Jin, Jong Chul Ye |
IEEE Trans. Image Process. | 1 |
| 2018 | Grid-Free Localization Algorithm Using Low-Rank Hankel Matrix for Super-Resolution MicroscopyabstractLocalization microscopy, such as STORM / PALM, can reconstruct super-resolution images with a nanometer resolution through the iterative localization of fluorescence molecules. Recent studies in this area have focused mainly on the localization of densely activated molecules to improve temporal resolutions. However, higher density imaging requires an advanced algorithm that can resolve closely spaced molecules. Accordingly, sparsitydriven methods have been studied extensively. One of the major limitations of existing sparsity-driven approaches is the need for a fine sampling grid or for Taylor series approximation which may result in some degree of localization bias toward the grid. In addition, prior knowledge of the point-spread function (PSF) is required. To address these drawbacks, here we propose a true grid-free localization algorithm with adaptive PSF estimation. Specifically, based on the observation that sparsity in the spatial domain implies a low rank in the Fourier domain, the proposed method converts source localization problems into Fourier-domain signal processing problems so that a truly gridfree localization is possible. We verify the performance of the newly proposed method with several numerical simulations and a live-cell imaging experiment. Junhong Min, Kyong Hwan Jin, Michael Unser, Jong Chul Ye |
IEEE Trans. Image Process. | 2 |
| 2018 | Unified Theory for Recovery of Sparse Signals in a General Transform DomainabstractCompressed sensing is provided a data-acquisition paradigm for sparse signals. Remarkably, it has been shown that the practical algorithms provide robust recovery from noisy linear measurements acquired at a near optimal sampling rate. In many real-world applications, a signal of interest is typically sparse not in the canonical basis but in a certain transform domain, such as wavelets or the finite difference. The theory of compressed sensing was extended to the analysis sparsity model, but known extensions are limited to the specific choices of sensing matrix and sparsifying transform. In this paper, we propose a unified theory for robust recovery of sparse signals in a general transform domain by convex programming. In particular, our results apply to the general acquisition and sparsity models and show how the number of measurements for recovery depends on properties of measurement and sparsifying transforms. Moreover, we also provide extensions of our results to the scenarios where the atoms in the transform have varying incoherence parameters and the unknown signal exhibits a structured sparsity pattern. In particular, for the partial Fourier recovery of sparse signals over a circulant transform, our main results suggest a uniformly random sampling. Numerical results demonstrate that the variable density random sampling by our main results provides a superior recovery performance over the known sampling strategies. Kiryung Lee, Yanjun Li 0001, Kyong Hwan Jin, Jong Chul Ye |
IEEE Trans. Inf. Theory | 3 |
| 2018 | CNN-Based Projected Gradient Descent for Consistent CT Image ReconstructionabstractWe present a new image reconstruction method that replaces the projector in a projected gradient descent (PGD) with a convolutional neural network (CNN). Recently, CNNs trained as image-to-image regressors have been successfully used to solve inverse problems in imaging. However, unlike existing iterative image reconstruction algorithms, these CNN-based approaches usually lack a feedback mechanism to enforce that the reconstructed image is consistent with the measurements. We propose a relaxed version of PGD wherein gradient descent enforces measurement consistency, while a CNN recursively projects the solution closer to the space of desired reconstruction images. We show that this algorithm is guaranteed to converge and, under certain conditions, converges to a local minimum of a non-convex inverse problem. Finally, we propose a simple scheme to train the CNN to act like a projector. Our experiments on sparse-view computed-tomography reconstruction show an improvement over total variation-based regularization, dictionary learning, and a state-of-the-art deep learning-based direct reconstruction technique. Kyong Hwan Jin, Ha Q. Nguyen 0001, Michael T. McCann, Michael Unser |
IEEE Trans. Medical Imaging | 2 |
| 2017 | Deep Convolutional Neural Network for Inverse Problems in ImagingabstractIn this paper, we propose a novel deep convolutional neural network (CNN)-based algorithm for solving ill-posed inverse problems. Regularized iterative algorithms have emerged as the standard approach to ill-posed inverse problems in the past few decades. These methods produce excellent results, but can be challenging to deploy in practice due to factors including the high computational cost of the forward and adjoint operators and the difficulty of hyperparameter selection. The starting point of this paper is the observation that unrolled iterative methods have the form of a CNN (filtering followed by pointwise nonlinearity) when the normal operator (H*H, where H* is the adjoint of the forward imaging operator, H) of the forward model is a convolution. Based on this observation, we propose using direct inversion followed by a CNN to solve normal-convolutional inverse problems. The direct inversion encapsulates the physical model of the system, but leads to artifacts when the problem is ill posed; the CNN combines multiresolution decomposition and residual learning in order to learn to remove these artifacts while preserving image structure. We demonstrate the performance of the proposed network in sparse-view reconstruction (down to 50 views) on parallel beam X-ray computed tomography in synthetic phantoms as well as in real experimental sinograms. The proposed network outperforms total variation-regularized iterative reconstruction for the more realistic phantoms and requires less than a second to reconstruct a 512 × 512 image on the GPU. Kyong Hwan Jin, Michael T. McCann, Emmanuel Froustey, Michael Unser |
IEEE Trans. Image Process. | 1 |
| 2017 | Compressive Sampling Using Annihilating Filter-Based Low-Rank InterpolationabstractWhile the recent theory of compressed sensing provides an opportunity to overcome the Nyquist limit in recovering sparse signals, a solution approach usually takes the form of an inverse problem of an unknown signal, which is crucially dependent on specific signal representation. In this paper, we propose a drastically different two-step Fourier compressive sampling framework in a continuous domain that can be implemented via measurement domain interpolation, after which signal reconstruction can be done using classical analytic reconstruction methods. The main idea originates from the fundamental duality between the sparsity in the primary space and the low-rankness of a structured matrix in the spectral domain, showing that a low-rank interpolator in the spectral domain can enjoy all of the benefits of sparse recovery with performance guarantees. Most notably, the proposed low-rank interpolation approach can be regarded as a generalization of recent spectral compressed sensing to recover large classes of finite rate of innovations (FRI) signals at a near-optimal sampling rate. Moreover, for the case of cardinal representation, we can show that the proposed low-rank interpolation scheme will benefit from inherent regularization and an optimal incoherence parameter. Using a powerful dual certificate and the golfing scheme, we show that the new framework still achieves a near-optimal sampling rate for a general class of FRI signal recovery, while the sampling rate can be further reduced for a class of cardinal splines. Numerical results using various types of FRI signals confirm that the proposed low-rank interpolation approach offers significantly better phase transitions than conventional compressive sampling approaches. Jong Chul Ye, Jong Min Kim 0002, Kyong Hwan Jin, Kiryung Lee |
IEEE Trans. Inf. Theory | 3 |
| 2016 | Recent progresses of accelerated MRI using annihilating filter-based low-rank interpolationabstractRecently, an annihilating filter based low-rank Hankel matrix approach (ALOHA) was proposed as a general framework for sparsity-driven k-space interpolation method for compressed sensing MRI (CS-MRI). The principle of ALOHA framework is based on the fundamental duality between the transform domain sparsity in the primary space and the low-rankness of weighted Hankel matrix in Fourier domain, which converts CS-MRI to a k-space interpolation problem using structured matrix completion. In this review, we explain the theory behind ALOHA. Experimental results with in vivo data for multi-coil dynamic imaging, parametric mapping as well as Nyquist ghost correction confirmed that the proposed method has potential to be a general solution of various MR imaging problems. Kyong Hwan Jin, Dongwook Lee 0005, Jong Chul Ye |
ICIP | 1 |
| 2016 | Random impulse noise removal using sparse and low rank decomposition of annihilating filter-based Hankel matrixabstractAnnihilating filer-based low rank Hankel matrix (ALOHA) approach was recently proposed as an intrinsic image model for image inpainting estimation. Based on the observation that smoothness or textures within an image patch are represented as sparse spectral components in the frequency domain, ALOHA exploits the existence of annihilating filters and the associated rank-deficient Hankel matrices in the image domain to estimate the missing pixels. As a extension, here we propose a novel impulse noise removal algorithm using sparse + low rank decomposition of an annihilating filter-based Hankel matrix. This novel approach, what we call robust ALOHA, is inspired by the observation that an image corrupted with impulse noises has intact pixels; so the impulse noises can be modeled as sparse outliers, whereas the underlying image can be still modeled using a low-rank Hankel structured matrix. Numerical results confirm that robust ALOHA has significant performance improvements compared to the state-of-the-art impulse removal algorithms. Kyong Hwan Jin, Jong Chul Ye |
ICIP | 1 |
| 2015 | Annihilating Filter-Based Low-Rank Hankel Matrix Approach for Image InpaintingabstractIn this paper, we propose a patch-based image inpainting method using a low-rank Hankel structured matrix completion approach. The proposed method exploits the annihilation property between a shift-invariant filter and image data observed in many existing inpainting algorithms. In particular, by exploiting the commutative property of the convolution, the annihilation property results in a low-rank block Hankel structure data matrix, and the image inpainting problem becomes a low-rank structured matrix completion problem. The block Hankel structured matrices are obtained patch-by-patch to adapt to the local changes in the image statistics. To solve the structured low-rank matrix completion problem, we employ an alternating direction method of multipliers with factorization matrix initialization using the low-rank matrix fitting algorithm. As a side product of the matrix factorization, locally adaptive dictionaries can be also easily constructed. Despite the simplicity of the algorithm, the experimental results using irregularly subsampled images as well as various images with globally missing patterns showed that the proposed method outperforms existing state-of-the-art image inpainting methods. Kyong Hwan Jin, Jong Chul Ye |
IEEE Trans. Image Process. | 1 |