Hyung Il Koo

dblp:35/834 · DBLP profile ↗
← Back
43ranked-venue papers
19as first author
6since 2021 · last 2025
0000-0002-6955-8083ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 19 first-author · 3 since 2021Artificial intelligence and machine learning · 15 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 28% Reinforcement learning · 14% Language models and text generation · 14%
Computer graphics and multimedia
10 papers
Image and video processing · 71% Visual content generation and editing · 25% Geometric modeling and processing · 3%

Topics — the 23 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing
document image analysis
0.962016
Text-Line Detection in Camera-Captured Document Images Using the State Estimation of Connected Components · IEEE Trans. Image Process. 2016
Segmentation and Rectification of Pictures in the Camera-Captured Images of Printed Documents · IEEE Trans. Multim. 2013
Text-Line Extraction in Handwritten Chinese Documents Based on an Energy Minimization Framework · IEEE Trans. Image Process. 2012
Natural language and speech › Language models and text generation › large language model inference
inference-time computation
0.912025
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data · ICML 2025
Machine learning › Deep learning architectures and training › state space model
mamba
0.912025
Parameter-Efficient Fine-Tuning of State Space Models · ICML 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
Parameter-Efficient Fine-Tuning of State Space Models · ICML 2025
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model
0.912025
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data · ICML 2025
Machine learning › Deep learning architectures and training
state space model
0.912025
Parameter-Efficient Fine-Tuning of State Space Models · ICML 2025
Machine learning › Learning theory
weighted majority vote
0.912025
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data · ICML 2025
Machine learning › Generative modeling
diffusion model
0.812024
Eta Inversion: Designing an Optimal Eta Function for Diffusion-Based Real Image Editing · ECCV (14) 2024
Visual content generation and editing
image editing
0.812024
Eta Inversion: Designing an Optimal Eta Function for Diffusion-Based Real Image Editing · ECCV (14) 2024
Image and video processing › image preprocessing › image correction
perspective distortion correction
0.322013
Segmentation and Rectification of Pictures in the Camera-Captured Images of Printed Documents · IEEE Trans. Multim. 2013
Rectification of figures and photos in document images using bounding box interface · CVPR 2010
Machine learning › Transfer learning and domain adaptation › domain generalization
multi-source domain generalization
0.312025
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data · ICML 2025
Image and video processing › document image analysis › document layout analysis
text line detection
0.212016
Text-Line Detection in Camera-Captured Document Images Using the State Estimation of Connected Components · IEEE Trans. Image Process. 2016
Image and video processing › pattern detection
scene text detection
0.212013
Scene Text Detection via Connected Component Clustering and Nontext Filtering · IEEE Trans. Image Process. 2013
Image and video processing
image fusion
0.112011
Design of Interchannel MRF Model for Probabilistic Multichannel Image Processing · IEEE Trans. Image Process. 2011
Image and video processing › binary image processing
connected component labeling
0.122016
Text-Line Detection in Camera-Captured Document Images Using the State Estimation of Connected Components · IEEE Trans. Image Process. 2016
Scene Text Detection via Connected Component Clustering and Nontext Filtering · IEEE Trans. Image Process. 2013
Image and video processing › image restoration › document image restoration
document dewarping
0.112009
Composition of a Dewarped and Enhanced Document Image From Two View Images · IEEE Trans. Image Process. 2009
Image and video processing › image segmentation › graph-based segmentation
graph cut segmentation
0.112009
Graph cuts using a Riemannian metric induced by tensor voting · ICCV 2009
Image and video processing
image segmentation
0.112009
Graph cuts using a Riemannian metric induced by tensor voting · ICCV 2009
Geometric modeling and processing
tensor voting
0.112009
Graph cuts using a Riemannian metric induced by tensor voting · ICCV 2009
Mathematical optimization › discrete optimization
energy minimization
0.122012
Text-Line Extraction in Handwritten Chinese Documents Based on an Energy Minimization Framework · IEEE Trans. Image Process. 2012
Graph cuts using a Riemannian metric induced by tensor voting · ICCV 2009
Image and video processing › image restoration
image denoising
0.012011
Design of Interchannel MRF Model for Probabilistic Multichannel Image Processing · IEEE Trans. Image Process. 2011
Image and video processing
document image processing
0.012010
State Estimation in a Document Image and Its Application in Text Block Identification and Text Line Extraction · ECCV (2) 2010
Computational photography and imaging
image stitching
0.012009
Composition of a Dewarped and Enhanced Document Image From Two View Images · IEEE Trans. Image Process. 2009

Methods — techniques the papers use, named apart from their topics

eta inversion · 1.5diffusion model · 1.5synthetic reasoning data generation · 0.9sparse dimension tuning · 0.9process reward modeling · 0.9LoRA · 0.9maximally stable extremal region · 0.4adaboost · 0.4energy minimization · 0.4state estimation · 0.4graph cuts · 0.3alternating optimization · 0.3multilayer perceptron · 0.2boundary extraction · 0.2tensor voting · 0.1riemannian metric · 0.1
YearPublicationVenuePosition
2025 VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
abstract
Process Reward Models (PRMs) have proven effective at enhancing mathematical reasoning for Large Language Models (LLMs) by leveraging increased inference-time computation. However, they are predominantly trained on mathematical data and their generalizability to non-mathematical domains has not been rigorously studied. In response, this work first shows that current PRMs have poor performance in other domains. To address this limitation, we introduce ***VersaPRM***, a multi-domain PRM trained on synthetic reasoning data generated using our novel data generation and annotation method. VersaPRM achieves consistent performance gains across diverse domains. For instance, in the MMLU-Pro category of Law, VersaPRM via weighted majority voting, achieves a 7.9% performance gain over the majority voting baseline–surpassing Qwen2.5-Math-PRM's gain of 1.3%. We further contribute to the community by open-sourcing all data, code and models for VersaPRM.
Thomas Zeng 0003, Shuibai Zhang, Shutong Wu, Christian Classen, Daewon Chae, Ethan Ewer, Heeju Kim, Wonjun Kang, Jackson Kunde, Jungtaek Kim 0001, Hyung Il Koo, Kannan Ramchandran, Dimitris S. Papailiopoulos, Kangwook Lee 0001
ICML13
2025 Parameter-Efficient Fine-Tuning of State Space Models
abstract
Deep State Space Models (SSMs), such as Mamba (Gu & Dao, 2024), have become powerful tools for language modeling, offering high performance and linear scalability with sequence length. However, the application of parameter-efficient fine-tuning (PEFT) methods to SSM-based models remains largely underexplored. We start by investigating two fundamental questions on existing PEFT methods: (i) How do they perform on SSM-based models? (ii) Which parameters should they target for optimal results? Our analysis shows that LoRA and its variants consistently outperform all other PEFT methods. While LoRA is effective for linear projection matrices, it fails on SSM modules—yet still outperforms other methods applicable to SSMs, indicating their limitations. This underscores the need for a specialized SSM tuning approach. To address this, we propose Sparse Dimension Tuning (SDT), a PEFT method tailored for SSM modules. Combining SDT for SSMs with LoRA for linear projection matrices, we achieve state-of-the-art performance across extensive experiments.
Kevin Galim, Wonjun Kang, Hyung Il Koo, Kangwook Lee 0001
ICML4
2025 Counting Guidance for High Fidelity Text-to-Image Synthesis
abstract
Recently, there have been significant improvements in the quality and performance of text-to-image generation, largely due to the impressive results attained by diffusion models. However, text-to-image diffusion models sometimes struggle to create high-fidelity content for the given input prompt. One specific issue is their difficulty in generating the precise number of objects specified in the text prompt. For example, when provided with the prompt “five apples and ten lemons on a table,” images generated by diffusion models often contain an incorrect number of objects. In this paper, we present a method to improve diffusion models so that they accurately produce the correct object count based on the input prompt. We adopt a counting network that performs reference-less class-agnostic counting for any given image. We calculate the gradients of the counting network and refine the predicted noise for each step. To address the presence of multiple types of objects in the prompt, we utilize novel attention map guidance to obtain high-quality masks for each object. Finally, we guide the denoising process using the calculated gradients for each object. Through extensive experiments and evaluation, we demonstrate that the proposed method significantly enhances the fidelity of diffusion models with respect to object count.
Wonjun Kang, Kevin Galim, Hyung Il Koo, Nam Ik Cho
WACV3
2024 Eta Inversion: Designing an Optimal Eta Function for Diffusion-Based Real Image Editing
Wonjun Kang, Kevin Galim, Hyung Il Koo
ECCV (14)3
2022 Deep-learning and graph-based approach to table structure recognition
Jaewoo Park 0005, Hyung Il Koo, Nam Ik Cho
Multim. Tools Appl.3
2022 Inverse-Based Approach to Explaining and Visualizing Convolutional Neural Networks
abstract
This article presents a new method for understanding and visualizing convolutional neural networks (CNNs). Most existing approaches to this problem focus on a global score and evaluate the pixelwise contribution of inputs to the score. The analysis of CNNs for multilabeled outputs or regression has not yet been considered in the literature, despite their success on image classification tasks with well-defined global scores. To address this problem, we propose a new inverse-based approach that computes the inverse of a feedforward pass to identify activations of interest in lower layers. We developed a layerwise inverse procedure based on two observations: 1) inverse results should have consistent internal activations to the original forward pass and 2) a small amount of activation in inverse results is desirable for human interpretability. Experimental results show that the proposed method allows us to analyze CNNs for classification and regression in the same framework. We demonstrated that our method successfully finds attributions in the inputs for image classification with comparable performance to state-of-the-art methods. To visualize the tradeoff between various methods, we developed a novel plot that shows the tradeoff between the amount of activations and the rate of class reidentification. In the case of regression, our method showed that conventional CNNs for single image super-resolution overlook a portion of frequency bands that may result in performance degradation.
Hyuk Jin Kwon, Hyung Il Koo, Jae Woong Soh, Nam Ik Cho
IEEE Trans. Neural Networks Learn. Syst.2
2020 Improving Explainability of Integrated Gradients with Guided Non-Linearity
abstract
Along with the performance improvements of neural network models, developing methods that enable the explanation of their behavior is a significant research topic. For convolutional neural networks, the explainability is usually achieved with attribution (heatmap) that visualizes pixel-level importance or contribution of input to its corresponding result. This attribution should reflect the relation (dependency) between inputs and outputs, which has been studied with a variety of methods, e.g., derivative of an output with respect to an input pixel value, a weighted sum of gradients, amount of output changes to input perturbations, and so on. In this paper, we present a new method that improves the measure of attribution, and incorporates it into the integrated gradients method. To be precise, rather than using the conventional chain-rule, we propose a method called guided non-linearity that propagates gradients more effectively through non-linear units (e.g., ReLU and max-pool) so that only positive gradients backpropagate through nonlinear units. Our method is inspired by the mechanism of action potential generation in postsynaptic neurons, where the firing of action potentials depends on the sum of excitatory (EPSP) and inhibitory postsynaptic potentials (IPSP). We believe that paths consisting of EPSP-giving-neurons faithfully reflect the contribution of inputs to the output, and we make gradients flow only along those paths (i.e., paths of positive chain reactions). Experiments with 5 deep neural networks have shown that the proposed method outperforms others in terms of the deletion metrics, and yields fine-grained and more human-interpretable attribution.
Hyuk Jin Kwon, Hyung Il Koo, Nam Ik Cho
ICPR2
2020 Handwritten Text Segmentation via End-to-End Learning of Convolutional Neural Networks
Junho Jo, Hyung Il Koo, Jae Woong Soh, Nam Ik Cho
Multim. Tools Appl.2
2019 Age Estimation Using Trainable Gabor Wavelet Layers In A Convolutional Neural Network
abstract
In this paper, we propose a trainable Gabor wavelet (TGW) layer and cascade it with a convolutional neural network (CNN) for the age estimation. Unlike an existing method that uses fixed (hand-tuned) Gabor filters at the head of a CNN, we use Gabor wavelets that can be adapted for the given input as well as for the targeting task. This is enabled by (a) estimating hyperparameters of Gabor wavelets from the input and (b) using a 1 × 1 convolution layer for the selection of orientation parameter. The proposed TGW layers are trained with the standard gradient-descent method and can be easily incorporated with conventional CNNs in an end-to-end training manner. We conduct experiments on the Adience dataset and show that the proposed network outperforms the baseline CNN without TGW layers and efficiently used trainable parameters than ordinary CNN based methods.
Hyuk Jin Kwon, Hyung Il Koo, Jae Woong Soh, Nam Ik Cho
ICIP2
2018 Scene text rectification using glyph and character alignment properties
abstract
Scene text images usually suffer from perspective distortions, and hence their rectification has been an essential pre-processing step for many applications. Existing methods for scene text rectification mainly exploited the glyph property, which means that the characters in many languages have horizontal/vertical strokes and also have some symmetries in their shapes. In this paper, we propose to use an additional property that the characters need to be well aligned when rectified. For this, character alignment, as well as glyph properties, are encoded in the proposed cost function, and its minimization generates the transformation parameters. For encoding the alignment constraints, we perform the character segmentation using a projection profile method before optimizing the cost function. Since better segmentation needs better rectification and vice versa, the overall algorithm is designed to perform character segmentation and rectification iteratively. We evaluate our method on real and synthetic scene text images, and the experimental results show that our method achieves higher optical character recognition (OCR) rate than the previous approaches and also yields visually pleasing results.
Tae Ho Kil, Hyung Il Koo, Nam Ik Cho
ICPR2
2017 Robust Document Image Dewarping Method Using Text-Lines and Line Segments
abstract
Conventional text-line based document dewarping methods have problems when handling complex layout and/or very few text-lines. When there are few aligned text-lines in the image, this usually means that photos, graphics and/or tables take large portion of the input instead. Hence, for the robust document dewarping, we propose to use line segments in the image in addition to the aligned text-lines. Based on the assumption and observation that many of the line segments in the image are horizontally or vertically aligned in the well-rectified images, we encode this property into the cost function in addition to the text-line alignment cost. By minimizing the function, we can obtain transformation parameters for camera pose, page curve, etc., which are used for document rectification. Considering that there are many outliers in line segment directions and missed text-lines in some cases, the overall algorithm is designed in an iterative manner. At each step, we remove text components and line segments that are not well aligned, and then minimize the cost function with the updated information. Experimental results show that the proposed method is robust to the variety of page layouts.
Tae Ho Kil, Wonkyo Seo, Hyung Il Koo, Nam Ik Cho
ICDAR3
2017 Rectification of planar targets using line segments
Jaehyun An, Hyung Il Koo, Nam Ik Cho
Mach. Vis. Appl.2
2017 Open-Contour Tracking Using a New State-Space Model and Nonrigid Motion Training
abstract
Object tracking in a video sequence is usually achieved by tracking the bounding box over the object or the object’s boundary, each of which has somewhat different applications. In this paper, we present a new open-contour tracking algorithm based on a Bayesian framework in which the contour is a part of the object’s boundary. We first propose a new state-space model for the representation of contours, which can handle the rigid and nonrigid motions of contours independently. This model enables us to focus on the nonrigid motions during the training, and the model works for challenging rigid motion scenarios. In addition, for the robust tracking of contours, we propose a measurement function that considers the contrast on object boundaries, target appearance, and temporal coherence. We applied the proposed method to two kinds of open-contours targets, and the experimental results show that the proposed method achieves superior performance to the conventional contour tracking methods. The proposed method is also compared with recent bounding box tracking methods for the object tracking purposes, and the comparison shows that the proposed method works robustly to fast motions and yields a more accurate estimate of an object’s location than the conventional bounding box tracking methods.
Seon Heo, Hyung Il Koo, Nam Ik Cho
IEEE Trans. Circuits Syst. Video Technol.2
2016 Fast and simple text replacement algorithm for text-based augmented reality
abstract
In this paper, we present a novel text-based augmented reality system that performs optical character recognition on natural images and replaces the recognized texts with other informative contents. For the goal, we implement text detection and recognition functions, and develop an image augmentation algorithm for the realistic contents replacement. To be precise, we reconstruct background with a linear interpolation method and insert new contents to the reconstructed backgrounds. Finally, we get natural results by applying proper geometric distortions to them. In order to reduce visual artifacts caused by noisy boundaries, we also develop an optimal path selection method based on dynamic programming. Experimental results show that our method provides very natural results and runs in real-time even in mobile devices.
Hyung Il Koo, Beom Su Kim, Young Ki Baik, Nam Ik Cho
VCIP1
2016 Text-Line Detection in Camera-Captured Document Images Using the State Estimation of Connected Components
abstract
Camera-based text processing has attracted considerable attention and numerous methods have been proposed. However, most of these methods have focused on the scene text detection problem and relatively little work has been performed on camera-captured document images. In this paper, we present a text-line detection algorithm for camera-captured document images, which is an essential step toward document understanding. In particular, our method is developed by incorporating state estimation (an extension of scale selection) into a connected component (CC)-based framework. To be precise, we extract CCs with the maximally stable extremal region algorithm and estimate the scales and orientations of CCs from their projection profiles. Since this state estimation facilitates a merging process (bottom-up clustering) and provides a stopping criterion, our method is able to handle arbitrarily oriented text-lines and works robustly for a range of scales. Finally, a text-line/non-text-line classifier is trained and non-text candidates (e.g., background clutters) are filtered out with the classifier. Experimental results show that the proposed method outperforms conventional methods on a standard dataset and works well for a new challenging dataset.
Hyung Il Koo
IEEE Trans. Image Process.1
2015 Junction-based table detection in camera-captured document images
Wonkyo Seo, Hyung Il Koo, Nam Ik Cho
Int. J. Document Anal. Recognit.2
2015 Document dewarping via text-line based optimization
Beom Su Kim, Hyung Il Koo, Nam Ik Cho
Pattern Recognit.2
2015 Efficient Unwrap Representation of Faces for Video Editing
abstract
Unwrap mosaic is a method for decomposing a video into a 2D texture and a dense mapping that enable the reconstruction of the video from the texture. This representation is useful in some frameworks because we can edit videos by simply retouching 2D textures. However, the complexity of conventional approaches is too high to be adopted in time-critical applications (it takes up to several hours). In this letter, we focus on face-related applications such as face editing and replacement, and propose a face unwrap approach for these applications. To be precise, we adopt the view-based active appearance model (AAM) trackers and estimate dense mappings from the AAM results. The AAM also provides pose information which is also exploited in building the texture map. Experimental results show that our method is very efficient compared with the conventional unwrap mosaic approach. Moreover, based on the proposed system, we develop face-related applications.
Byeongyong Ahn, Hyung Il Koo, Hong Il Kim, Jichull Jeong, Nam Ik Cho
IEEE Signal Process. Lett.2
2015 Word Segmentation Method for Handwritten Documents based on Structured Learning
abstract
Segmentation of handwritten document images into text-lines and words is an essential task for optical character recognition. However, since the features of handwritten document are irregular and diverse depending on the person, it is considered a challenging problem. In order to address the problem, we formulate the word segmentation problem as a binary quadratic assignment problem that considers pairwise correlations between the gaps as well as the likelihoods of individual gaps. Even though many parameters are involved in our formulation, we estimate all parameters based on the Structured SVM framework so that the proposed method works well regardless of writing styles and written languages without user-defined parameters. Experimental results on ICDAR 2009/2013 handwriting segmentation databases show that proposed method achieves the state-of-the-art performance on Latin-based and Indian languages.
Jewoong Ryu, Hyung Il Koo, Nam Ik Cho
IEEE Signal Process. Lett.2
2014 Language-Independent Text-Line Extraction Algorithm for Handwritten Documents
abstract
Text-line extraction in handwritten documents is an important step for document image understanding, and a number of algorithms have been proposed to address this problem. However, most of them exploit features of specific languages and work only for a given language. In order to overcome this limitation, we develop a language-independent text-line extraction algorithm. Our method is based on connected components (CCs), however, unlike conventional methods, we analyze strokes and partition under-segmented CCs into normalized ones. Due to this normalization, the proposed method is able to estimate the states of CCs for a range of different languages and writing styles. From the estimated states, we build a cost function whose minimization yields text-lines. Experimental results show that the proposed method yields the state-of-the-art performance on Latin-based and Chinese script databases. Further, we submitted the proposed algorithm to the ICDAR 2013 handwriting segmentation competition and our method showed the best text-line extraction performance among 10 participant methods.
Jewoong Ryu, Hyung Il Koo, Nam Ik Cho
IEEE Signal Process. Lett.2
2013 Scene Text Detection via Connected Component Clustering and Nontext Filtering
abstract
In this paper, we present a new scene text detection algorithm based on two machine learning classifiers: one allows us to generate candidate word regions and the other filters out nontext ones. To be precise, we extract connected components (CCs) in images by using the maximally stable extremal region algorithm. These extracted CCs are partitioned into clusters so that we can generate candidate regions. Unlike conventional methods relying on heuristic rules in clustering, we train an AdaBoost classifier that determines the adjacency relationship and cluster CCs by using their pairwise relations. Then we normalize candidate word regions and determine whether each region contains text or not. Since the scale, skew, and color of each candidate can be estimated from CCs, we develop a text/nontext classifier for normalized images. This classifier is based on multilayer perceptrons and we can control recall and precision rates with a single free parameter. Finally, we extend our approach to exploit multichannel information. Experimental results on ICDAR 2005 and 2011 robust reading competition datasets show that our method yields the state-of-the-art performance both in speed and accuracy.
Hyung Il Koo, Duck Hoon Kim
IEEE Trans. Image Process.1
2013 Segmentation and Rectification of Pictures in the Camera-Captured Images of Printed Documents
abstract
This paper presents an algorithm that segments and rectifies pictures in camera-captured document images. Most of the conventional methods for this purpose require the 3-D shape of document surface, which are usually measured or inferred by a depth-measuring device, structured light, or stereo system. Unlike these methods, our method requires only a single-view image and a user-provided rough bounding box on the picture. Hence, the main features of the proposed algorithm are simple user interaction and short processing time: a mega-pixel size image can be segmented and rectified within 1-2 s, on receiving the user's bounding box. To achieve this goal, we develop a novel boundary extraction algorithm that exploits the specific properties of printed material. In the method, a set of boundary candidates is generated, and the optimal boundary is found by using an alternating optimization scheme. In addition to the segmentation method, we also propose a new rectification method, which can largely remove perspective distortions. Experimental results on a variety of images show that our method is efficient, robust, and easy to use.
Hyung Il Koo
IEEE Trans. Multim.1
2012 Text-Line Extraction in Handwritten Chinese Documents Based on an Energy Minimization Framework
abstract
Text-line extraction in unconstrained handwritten documents remains a challenging problem due to nonuniform character scale, spatially varying text orientation, and the interference between text lines. In order to address these problems, we propose a new cost function that considers the interactions between text lines and the curvilinearity of each text line. Precisely, we achieve this goal by introducing normalized measures for them, which are based on an estimated line spacing. We also present an optimization method that exploits the properties of our cost function. Experimental results on a database consisting of 853 handwritten Chinese document images have shown that our method achieves a detection rate of 99.52% and an error rate of 0.32%, which outperforms conventional methods.
Hyung Il Koo, Nam Ik Cho
IEEE Trans. Image Process.1
2011 Design of Interchannel MRF Model for Probabilistic Multichannel Image Processing
abstract
In this paper, we present a novel framework that exploits an informative reference channel in the processing of another channel. We formulate the problem as a maximum a posteriori estimation problem considering a reference channel and develop a probabilistic model encoding the interchannel correlations based on Markov random fields. Interestingly, the proposed formulation results in an image-specific and region-specific linear filter for each site. The strength of filter response can also be controlled in order to transfer the structural information of a channel to the others. Experimental results on satellite image fusion and chrominance image interpolation with denoising show that our method provides improved subjective and objective performance compared with conventional approaches.
Hyung Il Koo, Nam Ik Cho
IEEE Trans. Image Process.1
2010 Rectification of figures and photos in document images using bounding box interface
abstract
This paper proposes an algorithm for the segmentation and rectification of figures and photos in document images. The algorithm requires just a rough user-provided bounding box for the objects in a single-view image. On receiving the user's bounding box, it takes about 1-2 seconds to segment and rectify mega-pixel sized figures. The main feature of the algorithm is a novel segmentation method that exploits the properties of printed figures. Specifically, a set of boundary candidates is generated using the properties, and the optimal boundary in the set is found by using an alternating optimization scheme. This segmentation result is further refined so that it is well localized to the true boundary. In addition to our segmentation method, we also propose a new boundary interpolation method for the rectification of segmented figures. The method improves the quality of output by largely removing perspective distortions compared to conventional boundary interpolation methods. Experimental results on a variety of images show that the method is efficient, robust, and easy to use.
Hyung Il Koo, Nam Ik Cho
CVPR1
2010 State Estimation in a Document Image and Its Application in Text Block Identification and Text Line Extraction
Hyung Il Koo, Nam Ik Cho
ECCV (2)1
2010 A video object segmentation algorithm based on the feature learning and shape tracking
abstract
This paper proposes a video object segmentation algorithm based on the conditional random field (CRF) framework. A foreground object in the first frame is segmented by training the CRF on user interaction, i.e., by using user scribbles corresponding to foreground and background respectively for CRF training. The data term of the energy function in this CRF framework is designed as a function of the score of texture-color classifier trained by AdaBoost. From the second frame, a weighted data term that encodes the shape of the object is added to this energy function. The boundary pixels of the current frame are predicted by the optical flow, and a smaller cost is given to a pixel closer to the boundary and vice versa. Also, a confidence of optical flow is defined, and a larger weight is given to the data term when the confident is high. As a result, the data term related with the shape becomes important when the motion estimation is reliable, and conversely the color-texture term becomes important otherwise. Experimental results show that the proposed data term keeps the boundary correctly in most cases and provides comparable result when compared to a state-of-the-art method.
Sang-Hak Lee, Hyung Il Koo, Nam Ik Cho
ICIP2
2010 A new image projection method for panoramic image stitching
abstract
We propose a new image projection method in an attempt to reduce the perceptual distortion in panoramic image mosaics. Specifically, we reduce the stretching distortion of some image patches and bending of straight lines. Since the stretching distortion usually occurs when projecting a viewing sphere to the cylindrical image surface in an oblique direction, we propose to use an adjustable cylindrical surface to match the viewing direction with the equator of the cylindrical surface. Also, in order to find the trade-off between the stretching distortion and bending of straight lines, we also adjust the curvature of cylindrical surface according to the object of interest in the image. The warping function from the viewing sphere to the adjustable image surface is derived and the amount of distortion caused by this warping function is also defined. From the measure of distortion, the optimal pose of the cylindrical image plane and its curvature are determined, and the image on the viewing sphere is projected on the optimal plane. The experimental results show that the proposed method produces the panoramic image with less distortion than the existing methods.
Beom Su Kim, Hyung Il Koo, Nam Ik Cho
MMSP2
2010 Image segmentation algorithms based on the machine learning of features
Sang-Hak Lee, Hyung Il Koo, Nam Ik Cho
Pattern Recognit. Lett.2
2009 A new method to find an optimal warping function in image stitching
abstract
In image stitching applications, it is very important to find a suitable warping function for the visual quality of a composite (stitching result). In this paper, we present a new mathematical criterion to select an optimal warping function among a set of possible candidates (e.g., parametric family). The proposed criterion can be considered as a direct view condition for image stitching , i.e., it is desirable that each part of the composite image looks like its corresponding input image. More specifically, we do not use an explicit modeling of a compositing surface, but, we focus on the differential properties of a warping function. That is we design a cost function so that the Jacobian matrix of a warping function is close to a shape preserving matrix such as rotation and reflection matrices. The proposed cost function can be effectively minimized by using Levenberg-Marquardt algorithm. The experimental results show that the proposed method results in visually pleasing stitched results because the original shape of each image is preserved in the composite.
Hyung Il Koo, Beom Su Kim, Nam Ik Cho
ICASSP1
2009 Graph cuts using a Riemannian metric induced by tensor voting
abstract
In this paper, we present a new algorithm that combines the advantages of tensor voting into graph cuts. Tensor voting has been a popular tool for a number of early vision problems since it can use principles of perceptual grouping, which are not well considered in graph cuts. We attempt to encode the power of tensor voting into an energy minimization framework. For this, we assume that the tensor map obtained by tensor voting induces a Riemannian metric in image domain, and the metric is constructed according to the conventional ways of tensor interpretation. Finally, by embedding the induced Riemannian metric into the graph via edge weights, the graph cuts algorithm can have priors considering principles of perceptual grouping. The proposed method can be used in the labeling of occluded regions, object segmentation using only edge information, and boundary regularization.
Hyung Il Koo, Nam Ik Cho
ICCV1
2009 Eliminating structure misalignments using robust matching and image editing based on seam carving
abstract
In this paper, we propose an algorithm that generates a natural composite from the misaligned images. The image stitching for this purpose is usually performed in two steps: correspondence matching of salient features followed by appropriated warping or editing. The proposed correspondence matching problem is formulated as the one-dimensional registration along the stitching boundary, where an appropriate energy function is proposed. The designed energy function consists of three complementary terms that encode appearance, smoothness and ordering of points, whereas the existing method considers the edge strength and ordering. Then we develop an algorithm that makes several salient points move to desired positions by using seam carving/inserting, which produces visually pleasing results compared to the conventional warping methods with less computations. The experimental results show that the proposed method efficiently and robustly generates natural composite images.
Hyung Il Koo, Jung Gap Kuk, Nam Ik Cho
ICIP1
2009 An unsupervised image segmentation algorithm based on the machine learning of appropriate features
abstract
This paper proposes a new approach to the feature based unsupervised image segmentation. The difficulty with the conventional unsupervised segmentation lies in finding appropriate features that discriminate a meaningful region from the others. In this paper, the appropriate features are automatically learnt by machine learning with boosting scheme. At the initial step, the image is split into many small regions (blocks at first) and strong classifiers for every region, which discriminate the region from the others, are found by AdaBoosting. Each strong classifier so obtained is the weighted sum of several popular weak classifiers (features), which best describes the coherence of the region and thus well discriminates the region from the others. The output of this classifier is used in designing the energy function for the labeling, in the form of conditional random fields (CRFs). Minimization of the energy function produces the labeling result which reflects the property learnt by the classifier. For the labeling result, the machine learning is again performed and the process iterates until some conditions are met. Experimental results show that the proposed method provides competitive result compared to the conventional feature based methods.
Sang-Hak Lee, Hyung Il Koo, Nam Ik Cho
ICIP2
2009 Composition of a Dewarped and Enhanced Document Image From Two View Images
abstract
In this paper, we propose an algorithm to compose a geometrically dewarped and visually enhanced image from two document images taken by a digital camera at different angles. Unlike the conventional works that require special equipment or assumptions on the contents of books or complicated image acquisition steps, we estimate the unfolded book or document surface from the corresponding points between two images. For this purpose, the surface and camera matrices are estimated using structure reconstruction, 3-D projection analysis, and random sample consensus-based curve fitting with the cylindrical surface model. Because we do not need any assumption on the contents of books, the proposed method can be applied not only to optical character recognition (OCR), but also to the high-quality digitization of pictures in documents. In addition to the dewarping for a structurally better image, image mosaic is also performed for further improving the visual quality. By finding better parts of images (with less out of focus blur and/or without specular reflections) from either of views, we compose a better image by stitching and blending them. These processes are formulated as energy minimization problems that can be solved using a graph cut method. Experiments on many kinds of book or document images show that the proposed algorithm robustly works and yields visually pleasing results. Also, the OCR rate of the resulting image is comparable to that of document images from a flatbed scanner.
Hyung Il Koo, Nam Ik Cho
IEEE Trans. Image Process.1
2008 Image denoising based on a statistical model for wavelet coefficients
abstract
In this paper, we propose a new statistical model for the relationship of wavelet coefficients and its application to image denoising. The magnitude of a wavelet coefficient usually shows high correlations with the nearby ones. This property has been exploited in many wavelet-based image processing techniques. However, conventional works consider only the local neighborhood of a coefficient when inferring its hidden state. Consequently, the image context is not faithfully reflected and thus there are sometimes visually annoying artifacts. We attempt to alleviate this problem by developing a new statistical model for the random field that is consisted of hidden variables of the overall band and thus includes global relationship of wavelet coefficients. In this model, the image context is encoded by the relations of hidden states, and the state plane is efficiently inferred by the sum-product algorithm. In the experiment, the proposed model is incorporated with the state-of-the-art denoising algorithm, namely BLS GSM (Bayes Least Square - Gaussian Scale Mixture). The results show that the proposed algorithm suppresses many annoying artifacts that exist in the conventional denoising methods, and thus improves the subjective quality.
Hyung Il Koo, Nam Ik Cho
ICASSP1
2008 Camera-based document digitization using multiple images
abstract
Recently, there have been some attempts to use a potable digital camera for the document digitization. But unlike the conventional flatbed scanners, the document images taken by digital cameras suffer from the perspective distortion and the geometric distortion. These problems deteriorate not only the subjective quality but also the character recognition rate. In this paper, a new document dewarping algorithm that utilizes two input images is presented. The proposed algorithm requires neither auxiliary hardwares such that measure the 3D shape nor the impractical assumptions on the pose of books and cameras. Instead, the 3D shape of the book (document) surface is reconstructed from the geometric correspondence of two images and some post-processings. Contrast to the conventional works, the proposed algorithm do not need to find textlines. So it can be applied to not only text regions but also pictures, tables and mathematical equations. That is, we can rectify distorted images irrespective of languages or contents without auxiliary hardwares.
Hyung Il Koo, Nam Ik Cho
ICIP2
2008 Boosting image segmentation
abstract
This paper presents a new approach to image segmentation, based on the conditional random fields (CRF) modeling and AdaBoost. In the proposed segmentation algorithm, the discriminating characteristics are first learned online using a training machine, and then the learnt characteristics are used to improve the region segmentation. The proposed algorithm is devised to include any kind of features even if they have different semantics, and to learn the difference of regions by selecting and combining only a few discriminating features among them. These novel properties are accomplished by a new Gibbs energy derived from CRF, AdaBoost, and probabilistic interpretation of its strong classifier. Experimental results on various images show the effectiveness of the proposed method.
Hyung Il Koo, Nam Ik Cho
ICIP1
2008 Non-rigid image registration based on the globally optimized correspondences
abstract
In this paper, we propose a new approach to the non-rigid image registration. This problem can be easily attacked if we can find regularly distributed correspondence points over the whole image or over the objects of interest. Dense and stable image registration can be achieved by using some natural mapping (e.g., thin plate spline) of these correspondences. However, the problems with conventional correspondence matching methods are that the features can rarely be found at the textureless regions and the matching accuracy is degraded at the parts with non-rigid motions. In order to find the regularly spaced correspondences and their accurate matching even under the non-rigid motion, we place mesh nodes over the image and develop a new cost function that considers three complementary terms: similarity, smoothness and some topological constraint that prevents unlikely mappings. Experimental results demonstrate that the proposed method can find correct correspondences in the presence of non-rigid motions, multi-layers (motion discontinuity) and even in the textureless regions. Experimental results also show that the proposed method can be applied to old film restoration as well as image registration.
Hyung Il Koo, Jung Gap Kuk, Nam Ik Cho
ICPR1
2007 Prior Model for the MRF Modeling of Multi-Channel Images
abstract
In multi-channel images (e.g. color images with R, G, B channel, and multi-spectral images), there exist higher-order correlations among the channels. We develop a new MRP-MAP (Markov random field - maximum a posteriori) framework that can be used for various multi-channel image processing. Main features of the proposed framework is that the higher-order correlation between the channels is considered, whereas it is not well addressed in the conventional works. Given a channel image, the prior probability of another channel is computed based on the MRF modeling that the channel correlation is described as piecewise linear relationship. An optimization algorithm for the MAP estimation is also developed. The effectiveness of the proposed priors is demonstrated with a simple application, i.e., image denoising.
Hyung Il Koo, Nam Ik Cho
ICASSP (1)1
2007 Bayesian object extraction from uncalibrated image pairs
Hyung Il Koo, Sang Hwa Lee, Nam Ik Cho
Signal Process. Image Commun.1
2006 MAP-Based Object Extraction from Uncalibrated Image Pair
abstract
In this paper, we propose a new MAP algorithm for foreground objects extraction from uncalibrated image pairs. The segmentation is performed in the MAP framework with MRF (Markov random field) modeling of images. The proposed algorithm estimates several spatial transformations between two images by corresponding SIFT (scale invariant feature transform) points and sequential RANSAC (random sample consensus) algorithm. The area-ratio criterion is applied to the each transformation so that we select the transformation of foreground object. Using these transformations, we compute the likelihood of color-segmented subregions. We model the prior information which is based on smoothness condition. Finally, the object extraction is performed by the Bayesian belief propagation. Experiments on various image pairs and video sequences show promising results in extracting the foreground objects from the backgrounds.
Hyung Il Koo, Sang Hwa Lee, Nam Ik Cho, Seong Keun Kim, Dong Hahk Lee, Sunghoon Lee
ICIP1
2006 Stochastic Approach to Separate Diffuse and Specular Reflections
abstract
This paper presents separation of specular and diffuse reflection components from an image pair. The proposed approach is based on the dichromatic reflectance model and Markov random field models. The proposed method estimates specular and diffuse components by minimizing observation color noise and prior potential of specular reflectance. The specular reflection component is modelled as an MRF and estimated in maximum a posteriori framework. This paper proposes likelihood term of specular component and the prior model of specular reflectance based on Phong's shading. Some experiments show that the proposed approach separates the specular and diffuse reflection components effectively. The separated specular components can be utilized in image-based lighting which renders a scene with virtual lighting sources.
Sang Hwa Lee, Hyung Il Koo, Nam Ik Cho, Jong-Il Park
ICIP2
2003 A relevance feedback algorithm based on the clustering and Parzen window
abstract
A relevance feedback algorithm based on the nonparametric approach is proposed. In the feature space, the algorithm generates multiple hyper-spheres around the regions where the images relevant with the query are densely populated, whereas the conventional algorithm searches the images in a single hyper-ellipsoid region. Then the Parzen window approach is applied to estimate the probability of relevance of each image in these multiple clusters (hyper-spheres). As a result, the relevance region in the feature space expands rapidly and covers arbitrarily shaped spaces with a small number of parameters. Also, since the user needs to determine only the positive images not the ambiguous negative ones, it is more convenient to use compared to some of the existing algorithms requiring negative feedback.
Hyung Il Koo, Nam Ik Cho
ICIP (2)1