Il Yong Chun

dblp:174/1982 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0002-4226-3760ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Theory of computation · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Image and video processing · 77% Multimedia analysis and retrieval · 19% Computational photography and imaging · 4%
Artificial intelligence
3 papers
Representation and self-supervised learning · 40% Vision and language · 20% Image recognition and object detection · 20%
Theoretical computer science
1 paper
Information theory · 100%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing
image reconstruction
1.122023
Momentum-Net: Fast and Convergent Iterative Neural Network for Inverse Problems · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Convolutional Analysis Operator Learning: Acceleration and Convergence · IEEE Trans. Image Process. 2020
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
LaB-CL: Localized and Balanced Contrastive Learning for Improving Parking Slot Detection · ICRA 2025
Computer vision › Image recognition and object detection › object detection › category-specific object detection
parking slot detection
0.912025
LaB-CL: Localized and Balanced Contrastive Learning for Improving Parking Slot Detection · ICRA 2025
Machine learning › Representation and self-supervised learning › contrastive learning
supervised contrastive learning
0.912025
LaB-CL: Localized and Balanced Contrastive Learning for Improving Parking Slot Detection · ICRA 2025
Computer vision › Vision and language
video captioning
0.912025
MAMS: Model-Agnostic Module Selection Framework for Video Captioning · AAAI 2025
Multimedia analysis and retrieval
video captioning
0.912025
MAMS: Model-Agnostic Module Selection Framework for Video Captioning · AAAI 2025
Image and video processing › image restoration
inverse problem
0.712023
Momentum-Net: Fast and Convergent Iterative Neural Network for Inverse Problems · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Image and video processing › image reconstruction › tomographic reconstruction
sparse-view CT reconstruction
0.412020
Convolutional Analysis Operator Learning: Acceleration and Convergence · IEEE Trans. Image Process. 2020
Image and video processing › sparse representation
dictionary learning
0.312018
Convolutional Dictionary Learning: Acceleration and Convergence · IEEE Trans. Image Process. 2018
Image and video processing › image restoration › image denoising › sparse representation based denoising
dictionary learning-based denoising
0.312018
Convolutional Dictionary Learning: Acceleration and Convergence · IEEE Trans. Image Process. 2018
Image and video processing › image restoration
image denoising
0.312018
Convolutional Dictionary Learning: Acceleration and Convergence · IEEE Trans. Image Process. 2018
Information theory › signal processing
compressed sensing
0.312017
Compressed Sensing and Parallel Acquisition · IEEE Trans. Inf. Theory 2017
Information theory › signal processing › compressed sensing › sparse recovery
recovery guarantees
0.312017
Compressed Sensing and Parallel Acquisition · IEEE Trans. Inf. Theory 2017
Information theory › signal processing › compressed sensing
sparse recovery
0.312017
Compressed Sensing and Parallel Acquisition · IEEE Trans. Inf. Theory 2017
Robotics › Autonomous driving
perception
0.312025
LaB-CL: Localized and Balanced Contrastive Learning for Improving Parking Slot Detection · ICRA 2025
Medical and health informatics › medical imaging
medical image analysis
0.212023
Momentum-Net: Fast and Convergent Iterative Neural Network for Inverse Problems · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computational photography and imaging › light field imaging
light field photography
0.212023
Momentum-Net: Fast and Convergent Iterative Neural Network for Inverse Problems · IEEE Trans. Pattern Anal. Mach. Intell. 2023

Methods — techniques the papers use, named apart from their topics

majorization · 2.7momentum · 2.0block-wise optimization · 2.0multimodal transformer · 1.7module selection · 1.7attention masking · 1.7hard negative sampling · 0.9contrastive learning · 0.9class prototype · 0.9tight-frame constraint · 0.4block proximal extrapolated gradient · 0.4analysis sparsifying regularizer · 0.4ADMM · 0.3phase transition analysis · 0.3nonuniform recovery · 0.3
YearPublicationVenuePosition
2025 MAMS: Model-Agnostic Module Selection Framework for Video Captioning
abstract
Multi-modal transformers are rapidly gaining attention in video captioning tasks. Existing multi-modal video captioning methods extract a fixed number of frames, but this has critical challenges. If a limited number of frames are extracted, important frames with essential information for caption generation may be missed. Conversely, extracting an excessive number of frames includes consecutive frames, potentially causing redundancy in visual tokens extracted from consecutive video frames. To extract an appropriate number of frames for each video, this paper proposes the first model-agnostic module selection framework in video captioning that has two main functions: (1) selecting a caption generation module with an appropriate size based on visual tokens extracted from video frames, and (2) constructing subsets of visual tokens for the selected caption generation module. Furthermore, we propose a new adaptive attention masking scheme that enhances attention on important visual tokens. Our numerical experiments with three different benchmark datasets demonstrate that the proposed framework significantly improves the performances of three recent video captioning models.
Il Yong Chun, Hogun Park
AAAI2
2025 DX2CT: Diffusion Model for 3D CT Reconstruction from Bi or Mono-planar 2D X-ray(s)
abstract
Computational tomography (CT) provides high-resolution medical imaging, but it can expose patients to high radiation. X-ray scanners have low radiation exposure, but their resolutions are low. This paper proposes a new conditional diffusion model, DX2CT, that reconstructs three-dimensional (3D) CT volumes from bi or mono-planar X-ray image(s). Proposed DX2CT consists of two key components: 1) modulating feature maps extracted from two-dimensional (2D) X-ray(s) with 3D positions of CT volume using a new transformer and 2) effectively using the modulated 3D position-aware feature maps as conditions of DX2CT. In particular, the proposed transformer can provide conditions with rich information of a target CT slice to the conditional diffusion model, enabling high-quality CT reconstruction. Our experiments with the bi or mono-planar X-ray(s) benchmark datasets show that proposed DX2CT outperforms several state-of-the-art methods. Our codes and model will be available at: https://www.github.com/intyeger/DX2CT.
Yun Su Jeong, Hye Bin Yoo, Il Yong Chun
ICASSP3
2025 Autoregression-Free Video Prediction Using Diffusion Model for Mitigating Error Propagation
abstract
Existing long-term video prediction methods often rely on an autoregressive video prediction mechanism. However, this approach suffers from error propagation, particularly in distant future frames. To address this limitation, this paper proposes the first AutoRegression-Free (ARFree) video prediction framework using diffusion models. Different from an autoregressive video prediction mechanism, ARFree directly predicts any future frame tuples from the context frame tuple. The proposed ARFree consists of two key components: 1) a motion prediction module that predicts a future motion using motion feature extracted from the context frame tuple; 2) a training method that improves motion continuity and contextual consistency between adjacent future frame tuples. Our experiments with two benchmark datasets show that the proposed ARFree video prediction framework outperforms several state-of-the-art video prediction methods.
Woonho Ko, Jin Bok Park, Il Yong Chun
ICIP3
2025 LaB-CL: Localized and Balanced Contrastive Learning for Improving Parking Slot Detection
abstract
Parking slot detection is an essential technology in autonomous parking systems. In general, the classification problem of parking slot detection consists of two tasks, a task determining whether localized candidates are junctions of parking slots or not, and the other that identifies a shape of detected junctions. Both classification tasks can easily face biased learning toward the majority class, degrading classification performances. Yet, the data imbalance issue has been overlooked in parking slot detection. We propose the first supervised contrastive learning framework for parking slot detection, Localized and Balanced Contrastive Learning for improving parking slot detection (LaB-CL). The proposed LaBCL framework uses two main approaches. First, we propose to include class prototypes to consider representations from all classes in every mini batch, from the local perspective. Second, we propose a new hard negative sampling scheme that selects local representations with high prediction error. Experiments with the benchmark dataset demonstrate that the proposed LaB-CL framework can outperform existing parking slot detection methods.
U. Jin Jeong, Sumin Roh, Il Yong Chun
ICRA3
2024 Improving Neural Radiance Fields Using Near-Surface Sampling with Point Cloud Generation
abstract
Abstract Neural radiance field (NeRF) is an emerging view synthesis method that samples points in a three-dimensional (3D) space and estimates their existence and color probabilities. The disadvantage of NeRF is that it requires a long training time since it samples many 3D points. In addition, if one samples points from occluded regions or in the space where an object is unlikely to exist, the rendering quality of NeRF can be degraded. These issues can be solved by estimating the geometry of 3D scene. This paper proposes a near-surface sampling framework to improve the rendering quality of NeRF. To this end, the proposed method estimates the surface of a 3D object using depth images of the training set and performs sampling only near the estimated surface. To obtain depth information on a novel view, the paper proposes a 3D point cloud generation method and a simple refining method for projected depth from a point cloud. Experimental results show that the proposed near-surface sampling NeRF framework can significantly improve the rendering quality, compared to the original NeRF and three different state-of-the-art NeRF methods. In addition, one can significantly accelerate the training time of a NeRF model with the proposed near-surface sampling framework.
Hye Bin Yoo, Hyunmin Han, Sung Soo Hwang, Il Yong Chun
Neural Process. Lett.4
2023 Momentum-Net: Fast and Convergent Iterative Neural Network for Inverse Problems
abstract
Iterative neural networks (INN) are rapidly gaining attention for solving inverse problems in imaging, image processing, and computer vision. INNs combine regression NNs and an iterative model-based image reconstruction (MBIR) algorithm, often leading to both good generalization capability and outperforming reconstruction quality over existing MBIR optimization models. This paper proposes the first fast and convergent INN architecture, Momentum-Net, by generalizing a block-wise MBIR algorithm that uses momentum and majorizers with regression NNs. For fast MBIR, Momentum-Net uses momentum terms in extrapolation modules, and noniterative MBIR modules at each iteration by using majorizers, where each iteration of Momentum-Net consists of three core modules: image refining, extrapolation, and MBIR. Momentum-Net guarantees convergence to a fixed-point for general differentiable (non)convex MBIR functions (or data-fit terms) and convex feasible sets, under two asymptomatic conditions. To consider data-fit variations across training and testing samples, we also propose a regularization parameter selection scheme based on the "spectral spread" of majorization matrices. Numerical experiments for light-field photography using a focal stack and sparse-view computational tomography demonstrate that, given identical regression NN architectures, Momentum-Net significantly improves MBIR speed and accuracy over several existing INNs; it significantly improves reconstruction quality compared to a state-of-the-art MBIR method in each application.
Il Yong Chun, Hongki Lim, Jeffrey A. Fessler
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Improved Real-Time Monocular SLAM Using Semantic Segmentation on Selective Frames
abstract
Monocular simultaneous localization and mapping (SLAM) is emerging in advanced driver assistance systems and autonomous driving, because a single camera is cheap and easy to install. Conventional monocular SLAM has two major challenges leading inaccurate localization and mapping. First, it is challenging to estimate scales in localization and mapping. Second, conventional monocular SLAM uses inappropriate mapping factors such as dynamic objects and low-parallax areas in mapping. This paper proposes an improved real-time monocular SLAM that resolves the aforementioned challenges by efficiently using deep learning-based semantic segmentation. To achieve the real-time execution of the proposed method, we apply semantic segmentation only to downsampled keyframes in parallel with mapping processes. In addition, the proposed method corrects scales of camera poses and three-dimensional (3D) points, using estimated ground plane from road-labeled 3D points and the real camera height. The proposed method also removes inappropriate corner features labeled as moving objects and low parallax areas. Experiments with eight video sequences demonstrate that the proposed monocular SLAM system achieves significantly improved and comparable trajectory tracking accuracy, compared to existing state-of-the-art monocular and stereo SLAM systems, respectively. The proposed system can achieve real-time tracking on a standard CPU potentially with a standard GPU support, whereas existing segmentation-aided monocular SLAM does not.
Jinkyu Lee 0006, Muhyun Back, Sung Soo Hwang, Il Yong Chun
IEEE Trans. Intell. Transp. Syst.4
2020 Light-Field Reconstruction and Depth Estimation from Focal Stack Images Using Convolutional Neural Networks
abstract
Light-field (LF) reconstruction from focal stack images has diverse applications including face recognition, autonomous driving, and 3D reconstruction in virtual reality. It is a large-scale ill-conditioned inverse problem and typically requires regularized iterative algorithms to solve, which can be slow. This paper proposes a non-iterative LF reconstruction and depth estimation method based on three sequential convolutional neural networks (CNNs). The first CNN estimates an all-in-focus image from focal stack images. The second CNN estimates 4D ray depth from the estimated all-in-focus image via the first CNN, and focal stack images. The third CNN refines a Lambertian LF that is rendered using the all-in-focus image and ray depth estimated by the first and second CNNs, respectively. Numerical experiments show that the proposed CNN-based method achieves significantly more accurate and/or faster LF reconstruction, compared to a state-of-the-art sequential CNN using a single image, conventional model-based image reconstruction from a focal stack, and direct regression CNN from a focal stack.
Jeffrey A. Fessler, Theodore B. Norris, Il Yong Chun
ICASSP4
2020 Convolutional Analysis Operator Learning: Acceleration and Convergence
abstract
Convolutional operator learning is gaining attention in many signal processing and computer vision applications. Learning kernels has mostly relied on so-called patch-domain approaches that extract and store many overlapping patches across training signals. Due to memory demands, patch-domain methods have limitations when learning kernels from large datasets - particularly with multi-layered structures, e.g., convolutional neural networks - or when applying the learned kernels to high-dimensional signal recovery problems. The so-called convolution approach does not store many overlapping patches, and thus overcomes the memory problems particularly with careful algorithmic designs; it has been studied within the "synthesis" signal model, e.g., convolutional dictionary learning. This paper proposes a new convolutional analysis operator learning (CAOL) framework that learns an analysis sparsifying regularizer with the convolution perspective, and develops a new convergent Block Proximal Extrapolated Gradient method using a Majorizer (BPEG-M) to solve the corresponding block multi-nonconvex problems. To learn diverse filters within the CAOL framework, this paper introduces an orthogonality constraint that enforces a tight-frame filter condition, and a regularizer that promotes diversity between filters. Numerical experiments show that, with sharp majorizers, BPEG-M significantly accelerates the CAOL convergence rate compared to the state-of-the-art block proximal gradient (BPG) method. Numerical experiments for sparse-view computational tomography show that a convolutional sparsifying regularizer learned via CAOL significantly improves reconstruction quality compared to a conventional edge-preserving regularizer. Using more and wider kernels in a learned regularizer better preserves edges in reconstructed images.
Il Yong Chun, Jeffrey A. Fessler
IEEE Trans. Image Process.1
2020 Improved Low-Count Quantitative PET Reconstruction With an Iterative Neural Network
abstract
Image reconstruction in low-count PET is particularly challenging because gammas from natural radioactivity in Lu-based crystals cause high random fractions that lower the measurement signal-to-noise-ratio (SNR). In model-based image reconstruction (MBIR), using more iterations of an unregularized method may increase the noise, so incorporating regularization into the image reconstruction is desirable to control the noise. New regularization methods based on learned convolutional operators are emerging in MBIR. We modify the architecture of an iterative neural network, BCD-Net, for PET MBIR, and demonstrate the efficacy of the trained BCD-Net using XCAT phantom data that simulates the low true coincidence count-rates with high random fractions typical for Y-90 PET patient imaging after Y-90 microsphere radioembolization. Numerical results show that the proposed BCD-Net significantly improves CNR and RMSE of the reconstructed images compared to MBIR methods using non-trained regularizers, total variation (TV) and non-local means (NLM). Moreover, BCD-Net successfully generalizes to test data that differs from the training data. Improvements were also demonstrated for the clinically relevant phantom measurement data where we used training and testing datasets having very different activity distributions and count-levels.
Hongki Lim, Il Yong Chun, Yuni K. Dewaraja, Jeffrey A. Fessler
IEEE Trans. Medical Imaging2
2019 BCD-Net for Low-Dose CT Reconstruction: Acceleration, Convergence, and Generalization
Il Yong Chun, Xuehang Zheng, Yong Long, Jeffrey A. Fessler
MICCAI (6)1
2019 Convolutional Analysis Operator Learning: Dependence on Training Data
abstract
Convolutional analysis operator learning (CAOL) enables the unsupervised training of (hierarchical) convolutional sparsifying operators or autoencoders from large datasets. One can use many training images for CAOL, but a precise understanding of the impact of doing so has remained an open question. This letter presents a series of results that lend insight into the impact of dataset size on the filter update in CAOL. The first result is a general deterministic bound on errors in the estimated filters, and is followed by a bound on the expected errors as the number of training samples increases. The second result provides a high probability analogue. The bounds depend on properties of the training data, and we investigate their empirical values with real data. Taken together, these results provide evidence for the potential benefit of using more training data in CAOL.
Il Yong Chun, David Hong, Ben Adcock, Jeffrey A. Fessler
IEEE Signal Process. Lett.1
2018 Convolutional Dictionary Learning: Acceleration and Convergence
abstract
Convolutional dictionary learning (CDL or sparsifying CDL) has many applications in image processing and computer vision. There has been growing interest in developing efficient algorithms for CDL, mostly relying on the augmented Lagrangian (AL) method or the variant alternating direction method of multipliers (ADMM). When their parameters are properly tuned, AL methods have shown fast convergence in CDL. However, the parameter tuning process is not trivial due to its data dependence and, in practice, the convergence of AL methods depends on the AL parameters for nonconvex CDL problems. To moderate these problems, this paper proposes a new practically feasible and convergent Block Proximal Gradient method using a Majorizer (BPG-M) for CDL. The BPG-M-based CDL is investigated with different block updating schemes and majorization matrix designs, and further accelerated by incorporating some momentum coefficient formulas and restarting techniques. All of the methods investigated incorporate a boundary artifacts removal (or, more generally, sampling) operator in the learning model. Numerical experiments show that, without needing any parameter tuning process, the proposed BPG-M approach converges more stably to desirable solutions of lower objective values than the existing state-of-the-art ADMM algorithm and its memory-efficient variant do. Compared with the ADMM approaches, the BPG-M method using a multi-block updating scheme is particularly useful in single-threaded CDL algorithm handling large data sets, due to its lower memory requirement and no polynomial computational complexity. Image denoising experiments show that, for relatively strong additive white Gaussian noise, the filters learned by BPG-M-based CDL outperform those trained by the ADMM approach.
Il Yong Chun, Jeffrey A. Fessler
IEEE Trans. Image Process.1
2017 Compressed Sensing and Parallel Acquisition
abstract
Parallel acquisition systems arise in various applications to moderate problems caused by insufficient measurements in single-sensor systems. These systems allow simultaneous data acquisition in multiple sensors, thus alleviating such problems by providing more overall measurements. In this paper, we consider the combination of compressed sensing with parallel acquisition. We establish the theoretical improvements of such systems by providing nonuniform recovery guarantees for which, subject to appropriate conditions, the number of measurements required per sensor decreases linearly with the total number of sensors. Throughout, we consider two different sampling scenarios-distinct (i.e., independent sampling in each sensor) and identical (i.e., dependent sampling between sensors)-and a general mathematical framework that allows for a wide range of sensing matrices. We also consider not just the standard sparse signal model, but also the so-called sparse in levels signal model. As our results show, optimal recovery guarantees for both distinct and identical sampling are possible under much broader conditions on the so-called sensor profile matrices (which characterize environmental conditions between a source and the sensors) for the sparse in levels model than for the sparse model. To verify our recovery guarantees, we provide numerical results showing phase transitions for different multi-sensor environments.
Il Yong Chun, Ben Adcock
IEEE Trans. Inf. Theory1
2016 Optimal sparse recovery for multi-sensor measurements
abstract
Many practical sensing applications involve multiple sensors simultaneously acquiring measurements of a single object. Conversely, most existing sparse recovery guarantees in compressed sensing concern only single-sensor acquisition scenarios. In this paper, we address the optimal recovery of compressible signals from multi-sensor measurements using compressed sensing techniques. This confirms the benefits of multi-over single-sensor environments in the sense of reducing the number of measurements required per sensor, and therefore, depending on the application, the total time, power or cost. Throughout the paper we consider a broad class of sensing matrices, and two fundamentally different sampling scenarios (distinct and identical respectively), both of which are relevant to applications. For the case of diagonal sensor profile matrices (which characterize environmental conditions between a source and the sensors), this paper presents two key improvements over existing results. First, a simpler optimal recovery guarantee for distinct sampling, and second, an improved recovery guarantee for identical sampling, based on the so-called sparsity in levels signal model.
Il Yong Chun, Ben Adcock
ITW1
2016 Efficient Compressed Sensing SENSE pMRI Reconstruction With Joint Sparsity Promotion
abstract
The theory and techniques of compressed sensing (CS) have shown their potential as a breakthrough in accelerating k-space data acquisition for parallel magnetic resonance imaging (pMRI). However, the performance of CS reconstruction models in pMRI has not been fully maximized, and CS recovery guarantees for pMRI are largely absent. To improve reconstruction accuracy from parsimonious amounts of k-space data while maintaining flexibility, a new CS SENSitivity Encoding (SENSE) pMRI reconstruction framework promoting joint sparsity (JS) across channels (JS CS SENSE) is proposed in this paper. The recovery guarantee derived for the proposed JS CS SENSE model is demonstrated to be better than that of the conventional CS SENSE model and similar to that of the coil-by-coil CS model. The flexibility of the new model is better than the coil-by-coil CS model and the same as that of CS SENSE. For fast image reconstruction and fair comparisons, all the introduced CS-based constrained optimization problems are solved with split Bregman, variable splitting, and combined-variable splitting techniques. For the JS CS SENSE model in particular, these techniques lead to an efficient algorithm. Numerical experiments show that the reconstruction accuracy is significantly improved by JS CS SENSE compared with the conventional CS SENSE. In addition, an accurate residual-JS regularized sensitivity estimation model is also proposed and extended to calibration-less (CaL) JS CS SENSE. Numerical results show that CaL JS CS SENSE outperforms other state-of-the-art CS-based calibration-less methods in particular for reconstructing non-piecewise constant images.
Il Yong Chun, Ben Adcock, Thomas M. Talavage
IEEE Trans. Medical Imaging1