ByoungChul Ko

dblp:30/6174 · also Byoung Chul Ko · DBLP profile ↗
← Back
38ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-7284-0768ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 PINet: Improving the Stability of Prototype Networks via Phantasia-Inspired Uncertain Representations
abstract
Self-interpretable models are increasingly valued for their inherent explainability. Among them, part-prototype networks stand out by mimicking human reasoning through the use of learned prototypes. However, their explanations often lack stability, becoming sensitive to subtle input perturbations. In this work, we propose Prototype in Imagery Network (PINet), a framework that improves the stability of prototype-based explanations. Rather than training on all possible input variations, which is computationally infeasible, PINet draws inspiration from visual mental imagery. Specifically, we incorporate empty inputs and apply coarse location guidance to simulate the human ability to imagine rough object features (a process akin to Phantasia). PINet mimics this process by incorporating empty inputs and applying coarse location guidance. These imagined, or uncertain, representations are contrasted with those derived from actual inputs (certain representations). We model the differences between the two by computing similarity at both the feature and prototype levels, allowing uncertainty to be explicitly encoded during prototype learning. Comprehensive evaluations on CUB-200-2011 and Stanford Cars demonstrate that PINet consistently achieves robust accuracy and localization, even under noisy conditions. These results represent the ability of PINet to produce stable and interpretable explanations under uncertainty.
Ho Kyung Shin, Soeun Bae, Sang Min Kim, ByoungChul Ko, Woo-Jeoung Nam
AAAI4
2026 ZeRA: Zero-Reindex Multimodal RAG via Heterogeneous Embedding Alignment for Lightweight Query Encoding
Dasom Ahn, Hye Rim Kim, Sangwon Kim 0004, Kwang-Ju Kim, ByoungChul Ko
ICPR (11)6
2026 Energy-ensemble concept bottleneck models for enhancing interpretability and accuracy in concept-based learning
Dasom Ahn, Sangwon Kim 0004, ByoungChul Ko
Knowl. Based Syst.3
2025 AuGQ: Augmented quantization granularity to overcome accuracy degradation for sub-byte quantized deep neural networks
Ahmed Mujtaba, Wai-Kong Lee, ByoungChul Ko, Hyung Jin Chang, Seong Oun Hwang
Appl. Intell.3
2025 LAttE: A label-free and multimodal framework for context-aware person re-identification
Dasom Ahn, Sangwon Kim 0004, Kwang-Ju Kim, ByoungChul Ko
Neurocomputing4
2025 Semantic scene graph generation based on an edge dual scene graph and message passing neural network
ByoungChul Ko
Image Vis. Comput.2
2025 Balanced clustering contrastive learning for long-tailed visual recognition
Byeong-Il Kim, ByoungChul Ko
Pattern Anal. Appl.2
2024 EQ-CBM: A Probabilistic Concept Bottleneck with Energy-Based Models and Quantized Vectors
Sangwon Kim 0004, Dasom Ahn, ByoungChul Ko, In-Su Jang, Kwang-Ju Kim
ACCV (7)3
2024 Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency
abstract
Scene graph generation (SGG) is an important task in image understanding because it represents the relationships between objects in an image as a graph structure, making it possible to understand the semantic relationships between objects intuitively. Previous SGG studies used a message-passing neural networks (MPNN) to update features, which can effectively reflect information about surrounding objects. However, these studies have failed to reflect the co-occurrence of objects during SGG generation. In addition, they only addressed the long-tail problem of the training dataset from the perspectives of sampling and learning methods. To address these two problems, we propose CooK, which reflects the Co-occurrence Knowledge between objects, and the learnable term frequency-inverse document frequency (TF-$l$-IDF) to solve the long-tail problem. We applied the proposed model to the SGG benchmark dataset, and the results showed a performance improvement of up to 3.8% compared with existing state-of-the-art models in SGGen subtask. The proposed method exhibits generalization ability from the results obtained, showing uniform performance improvement for all MPNN models.
Sangwon Kim 0004, Dasom Ahn, Jong Taek Lee, ByoungChul Ko
ICML5
2024 BTD-RF: 3D scene reconstruction using block-term tensor decomposition
Seon Bin Kim, Sangwon Kim 0004, Dasom Ahn, ByoungChul Ko
Appl. Intell.4
2024 Domain-free fire detection using the spatial-temporal attention transform of the YOLO backbone
Sangwon Kim 0004, In-Su Jang, ByoungChul Ko
Pattern Anal. Appl.3
2023 Cross-Modal Learning with 3D Deformable Attention for Action Recognition
abstract
An important challenge in vision-based action recognition is the embedding of spatiotemporal features with two or more heterogeneous modalities into a single feature. In this study, we propose a new 3D deformable transformer for action recognition with adaptive spatiotemporal receptive fields and a cross-modal learning scheme. The 3D deformable transformer consists of three attention modules: 3D deformability, local joint stride, and temporal stride attention. The two cross-modal tokens are input into the 3D deformable attention module to create a cross-attention token with a reflected spatiotemporal correlation. Local joint stride attention is applied to spatially combine attention and pose tokens. Temporal stride attention temporally reduces the number of input tokens in the attention module and supports temporal expression learning without the simultaneous use of all tokens. The deformable transformer iterates L-times and combines the last cross-modal token for classification. The proposed 3D deformable transformer was tested on the NTU60, NTU120, FineGYM, and PennAction datasets, and showed results better than or similar to pre-trained state-of-the-art methods even without a pre-training process. In addition, by visualizing important joints and correlations during action recognition through spatial joint and temporal stride attention, the possibility of achieving an explainable potential for action recognition is presented.
Sangwon Kim 0004, Dasom Ahn, ByoungChul Ko
ICCV3
2023 STAR-Transformer: A Spatio-temporal Cross Attention Transformer for Human Action Recognition
abstract
In action recognition, although the combination of spatiotemporal videos and skeleton features can improve the recognition performance, a separate model and balancing feature representation for cross-modal data are required. To solve these problems, we propose Spatio-TemporAl cRoss (STAR)-transformer, which can effectively represent two cross-modal features as a recognizable vector. First, from the input video and skeleton sequence, video frames are output as global grid tokens and skeletons are output as joint map tokens, respectively. These tokens are then aggregated into multi-class tokens and input into STAR-transformer. The STAR-transformer encoder consists of a full spatio-temporal attention (FAttn) module and a proposed zigzag spatio-temporal attention (ZAttn) module. Similarly, the continuous decoder consists of a FAttn module and a proposed binary spatio-temporal attention (BAttn) module. STAR-transformer learns an efficient multi-feature representation of the spatio-temporal features by properly arranging pairings of the FAttn, ZAttn, and BAttn modules. Experimental results on the Penn-Action, NTU-RGB+D 60, and 120 datasets show that the proposed method achieves a promising improvement in performance in comparison to previous state-of-the-art methods.
Dasom Ahn, Sangwon Kim 0004, Hyunsu Hong, ByoungChul Ko
WACV4
2023 STAR++: Rethinking spatio-temporal cross attention transformer for video action recognition
Dasom Ahn, Sangwon Kim 0004, ByoungChul Ko
Appl. Intell.3
2023 SSL-MOT: self-supervised learning based multi-object tracking
Sangwon Kim 0004, Jimi Lee, ByoungChul Ko
Appl. Intell.3
2022 ViT-NeT: Interpretable Vision Transformers with Neural Tree Decoder
abstract
Vision transformers (ViTs), which have demonstrated a state-of-the-art performance in image classification, can also visualize global interpretations through attention-based contributions. However, the complexity of the model makes it difficult to interpret the decision-making process, and the ambiguity of the attention maps can cause incorrect correlations between image patches. In this study, we propose a new ViT neural tree decoder (ViT-NeT). A ViT acts as a backbone, and to solve its limitations, the output contextual image patches are applied to the proposed NeT. The NeT aims to accurately classify fine-grained objects with similar inter-class correlations and different intra-class correlations. In addition, it describes the decision-making process through a tree structure and prototype and enables a visual interpretation of the results. The proposed ViT-NeT is designed to not only improve the classification performance but also provide a human-friendly interpretation, which is effective in resolving the trade-off between performance and interpretability. We compared the performance of ViT-NeT with other state-of-art methods using widely used fine-grained visual categorization benchmark datasets and experimentally proved that the proposed method is superior in terms of the classification performance and interpretability. The code and models are publicly available at https://github.com/jumpsnack/ViT-NeT.
Sangwon Kim 0004, Jae-Yeal Nam, ByoungChul Ko
ICML3
2022 Lightweight surrogate random forest support for model simplification and feature relevance
Sangwon Kim 0004, Mira Jeong, ByoungChul Ko
Appl. Intell.3
2019 Intelligent Driver Emotion Monitoring Based on Lightweight Multilayer Random Forests
abstract
In this paper, we propose a lightweight multi-layer random forest (LMRF) model. The LMRF model is a non-neural network-style deep model composed of arbitrary forests rather than layers. DNN is a powerful algorithm for facial recognition (FER), but there are too many parameters, careful parameter tuning, large amounts of training data, black box models, and pretrained architecture required for a current DNN. To overcome the burden of real-time processing DNN, we use the proposed LMRF with two tree structures per layer and a small number of trees for high-speed FER. We conducted experiments using an actual driving database captured using a near-infrared (NIR) camera to monitor the driver's emotions. The proposed LMRF provides similar FER accuracy to DNN with a small number of hyperparameters, and the faster processing time using the CPU.
Mira Jeong, Minji Park, ByoungChul Ko
INDIN3
2018 Driver Facial Landmark Detection in Real Driving Situations
abstract
This paper proposes a novel facial landmark detection (FLD) algorithm for use in real driving situations. The proposed algorithm is based on an ensemble of local weighted random forest regressor (WRFR) with random sampling consensus (RANSAC) and explicit global shape models, and considers the dynamic and irregular characteristics of driving. In this paper, to estimate the offset distance from a landmark and a reference point, we first detect the nose region as a reference point. Next, we propose a local WRFR to maintain the generality with a small number of regression trees. With the WRFR, we adopt RANSAC instead of the averaging or median of offsets to handle the problem of sensitivity to outlier offsets, and we estimate the accurate 2D offset vector. To identify the erroneous positions of local landmarks and rearrange the overall landmark layout, we adopt the global face models based on the spatial relation between landmarks. Using the unified framework of the proposed FLD, our proposed algorithm is robust to large head poses and partial occlusions caused by a driver's hair or sunglasses. For a benchmark data set considering real driving situations, we construct a data set called a face alignment data set used in driving (FADID) using a near-infrared camera for FLD under real driving situations. We apply the proposed algorithm to various driving sequences in FADID, and the results show that its FLD performance is better than that of other state-of-the-art methods, while the computational speed is high for real-time applications such as driver-state monitoring systems.
Mira Jeong, ByoungChul Ko, Soo Yeong Kwak, Jae-Yeal Nam
IEEE Trans. Circuits Syst. Video Technol.2
2017 Early Detection of Sudden Pedestrian Crossing for Safe Driving During Summer Nights
abstract
Sudden pedestrian crossing (SPC) is the major reason for pedestrian-vehicle crashes. In this paper, we focus on detecting SPCs at night for supporting an advanced driver assistance system using a far-infrared (FIR) camera mounted on the front roof of a vehicle. Although the thermal temperature of the road is similar to or higher than that of the pedestrians during summer nights, many previous researches have focused on pedestrian detection during the winter, spring, or autumn seasons. However, our research concentrates on SPC during the hot summer season because the number of collisions between pedestrians and vehicles in Korea is higher at that time than during the other seasons. For real-time processing, we first decide the optimal levels of image scaling and search area. We then use our proposed method for detecting virtual reference lines that are associated with road segmentation without using color information and change these lines according to the turning direction of the vehicle. Pedestrian detection is conducted using a cascade random forest with low-dimensional Haar-like features and oriented center-symmetric local binary patterns. The SPC is predicted based on the likelihood and the spatiotemporal features of the pedestrians, such as their overlapping ratio with virtual reference lines, as well as the direction and magnitude of each pedestrian's movement. The proposed algorithm was successfully applied to various pedestrian data sets captured by an FIR camera, and the results show that its SPC detection performance is better than those of other methods.
Mira Jeong, ByoungChul Ko, Jae-Yeal Nam
IEEE Trans. Circuits Syst. Video Technol.2
2017 Pedestrian Tracking Using Online Boosted Random Ferns Learning in Far-Infrared Imagery for Safe Driving at Night
abstract
Pedestrian-vehicle accidents that occur at night are a major social problem worldwide. Advanced driver assistance systems that are equipped with cameras have been designed to automatically prevent such accidents. Among the various types of cameras used in such systems, far-infrared (FIR) cameras are favorable because they are invariant to illumination changes. Therefore, this paper focuses on a pedestrian nighttime tracking system with an FIR camera that is able to discern thermal energy and is mounted on the forward roof part of a vehicle. Since the temperature difference between the pedestrian and background depends on the season and the weather, we therefore propose two models to detect pedestrians according to the season and the weather, which are determined using Weber-Fechner's law. For tracking pedestrians, we perform real-time online learning to track pedestrians using boosted random ferns and update the trackers at each frame. In particular, we link detection responses to trajectories based on similarities in position, size, and appearance. There is no standard data set for evaluating the tracking performance using an FIR camera; thus, we created the Keimyung University tracking data set (KMUTD) by combining the KMU sudden pedestrian crossing (SPC) data set [21] for summer nights with additional tracking data for winter nights. The KMUTD contains video sequences involving a moving camera, moving pedestrians, sudden shape deformations, unexpected motion changes, and partial or full occlusions between pedestrians at night. The proposed algorithm is successfully applied to various pedestrian video sequences of the KMUTD; specifically, the proposed algorithm yields more accurate tracking performance than other existing methods.
Joon Young Kwak, ByoungChul Ko, Jae-Yeal Nam
IEEE Trans. Intell. Transp. Syst.2
2016 Facial landmark detection based on an ensemble of local weighted regressors during real driving situation
abstract
In this study, we propose a novel method for facial landmark detection (FLD) based on an ensemble of local weighted regressors and a global face shape model under real driving situations. Unlike other FLD approaches, the method proposed in this study first detects the nose region instead of a face-bounding box as a reference point for estimating the offset from a landmark and a reference point. Next, a weighted random forest regressor (WRFR) is used for designing a regressor that maintains the generality while utilizing a small number of decision trees. During the training period, some of the trees having low accuracy are removed and the remaining trees of the WRFR have different weights according to their regression accuracy. As a global face shape model, we use the spatial relationship between three landmarks to identify erroneous estimates of the local regressors and provide valid alternatives. Using the unified framework of the proposed FLD, our algorithm is robust to facial expressions and partial occlusions caused by a subject's hair or sunglasses. For our experiment, using a near-infrared camera, we constructed a benchmark dataset for FLD under real driving situations, which we call the Face Alignment Dataset used In Driving (FADID). The proposed algorithm was successfully applied to various driving sequences in FADID, and the results show that its FLD detection performance is better than that of other state-of-the-art methods.
Mira Jeong, Joon Young Kwak, ByoungChul Ko, Jae-Yeal Nam
ICPR3
2016 Online learning based multiple pedestrians tracking in thermal imagery for safe driving at night
abstract
According to a report, night time and poor illumination driving is overall 2-3 times more dangerous then day time. For example, young people aged 18-24 were killed between 21:00 and 05:59 (the night-time and early morning) on week-days in the EU-23 countries because of road accident in 2010. As the similar pattern with EU, most pedestrian-vehicle accidents occur between 6 p.m. and 8 a.m., and the rate of pedestrian fatalities is highest between 4 a.m. and 6 a.m. in South Korea. Among several factors such as inebriated drivers and pedestrian, drowsiness, decreased visibility is the major cause of pedestrian-vehicle accident at night. To reduce the accident owing to driver's inattention at night, recent advanced driver assistance system (ADAS) has been researching on automatic pedestrian detection and tracking using night vision camera. Therefore, this tutorial focuses on introducing a multiple pedestrians tracking system using a thermal camera that is able to discern thermal energy at night-time. In a pedestrian tracking-by-detection system, multi-pedestrian detection accuracy is essential for post tracking process. Since the temperature difference between the pedestrian and background depends on the season and weather, we therefore first introduce two models for detecting pedestrians according to the season and weather, which are determined using Weber-Fechner's law. Two detection models use the optimal levels of the image scaling and search area instead of image pyramid to reduce the computational cost of image scaling for detecting multiple pedestrians of various sizes. Online learning is appropriate in the case that image frames is obtained sequentially. Theoretically offline learning could obtain global optimal solution while it is not as practical as online learning. Therefore, we introduce some state-of-the-art real-time online learning algorithms with our online learning based on boosted random ferns (BRFs) in detail based on the references as the second topic. Third, for association checking of multiple pedestrians, we explain the advantages and disadvantages of feed-forward system and global association system. Feed-forward system uses only current and past observations, which is called tracklet to estimate the current tracker's state. On contrary, global association system uses future and global information to estimate the current tracker's state in an offline step, for example, bipartite graph matching, Hungarian algorithm, dynamic programming, and min-cost max-flow network flow. Because maintaining the tracker's ID in successive frames is a challenging task owing to overlapping pedestrians in the multiple pedestrians tracking system, we introduce a few association checking algorithms to maintain the tracker's ID. As the feature for association checking, we also explain popular features in computer vision, such as the spatial proximity, velocity orientation and context (shape, size, color). Fourth, we introduce a few evaluation video sequences for pedestrian tracking such as OSli thermal pedestrian database, CVC-09 sequences, KAIST benchmark dataset, and KMliTD dataset. In particular, we are focusing on the KMliTD which contains video sequences involving a moving camera, moving multiple pedestrians, sudden shape deformations, unexpected motion changes, and partial or full occlusions between pedestrians at summer and winter night. Finally, we introduce the evaluation methods to measure the performance of the pedestrian tracking system such as Multiple Object Tracking Precision (MOTP) and Multiple Object Tracking Accuracy (MOTA), and Tracking Distance Error (TDE). The performance comparison among different tracking approaches is also presented when the proposed online tracking method is applied to benchmark data. As the further research in multiple pedestrians tracking, we will guide the fusion of sensors such as Radio Detection and Ranging (RADAR) sensor or Light Detection and Ranging (LIDAR) sensor with a camera for overcome the limitations occurred in a standalone sensor. In addition, with the increasing sensor resolutions, we mention the plan how to develop computationally feasible multi-extended object tracking algorithm.
ByoungChul Ko, Joon Young Kwak, Jae-Yeal Nam
Intelligent Vehicles Symposium1
2015 Multi-person Tracking Based on Body Parts and Online Random Ferns Learning of Thermal Images
abstract
This paper presents a novel algorithm for tracking multiple persons with thermal imaging. The algorithm uses online random ferns (RF) learning to update the model of the person and particle filters to approximate the person's location. To estimate the observational likelihood for particle weighting, we perform online training for the initial ferns using boosted random ferns (BRF) in the first frame in regions where persons are detected. Then, RF for the tracker model is re-trained based on the observed distribution of selected ferns in consecutive frames. To design a robust tracking model impervious to occlusion, we divide person regions into 4 x 4 sub-blocks and then train the RF using concatenated feature vectors from 16 sub-blocks. In addition, we propose an occlusion-check algorithm to distinguish normal object-tracking from long and short-term occlusion. The proposed algorithm is compared with similar existing algorithms to show that its tracking performance is superior to those of other classifiers and tracking methods.
Joon Young Kwak, ByoungChul Ko, Jae-Yeal Nam
WACV2
2013 Wildfire smoke detection using spatiotemporal bag-of-features of smoke
abstract
This paper presents a wildfire smoke detection method based on a spatiotemporal bag-of-features (BoF) and a random forest classifier. First, candidate blocks are detected using key-frame differences and non-parametric color models to reduce the computation time. Subsequently, spatiotemporal three-dimensional (3D) volumes are built by combining the candidate blocks in the current key-frame and the corresponding blocks in previous frames. A histogram of gradient (HOG) is extracted as a spatial feature, and a histogram of optical flow (HOF) is extracted as a temporal feature based on the fact that the diffusion direction of smoke is upward owing to thermal convection. Using these spatiotemporal features, a codebook and a BoF histogram are generated from training data. For smoke verification, a random forest classifier is built during the training phase by using the BoF histogram. The random forest with BoF histogram can increase the detection accuracy and allow smoke detection to be carried out in near real-time.
JunOh Park, ByoungChul Ko, Jae-Yeal Nam, Soo Yeong Kwak
WACV2
2013 Spatiotemporal bag-of-features for early wildfire smoke detection
ByoungChul Ko, JunOh Park, Jae-Yeal Nam
Image Vis. Comput.1
2012 Human Detection Using Wavelet-Based CS-LBP and a Cascade of Random Forests
abstract
In this paper, we propose a novel human detection approach combining wavelet-based center symmetric LBP (WCS-LBP) with a cascade of random forests. To detect human regions, we first extract three types of WCS-LBP features from a scanning window of wavelet transformed sub-images to reduce the feature dimension. Then, the extracted WCS-LBP descriptors are applied to a cascade of random forests, which are ensembles of random decision trees. Using a cascade of random forests with WCS-LBP, human detection is performed in near real-time, and the detection accuracy is also increased, as compared to combinations of other features and classifiers. The proposed algorithm is successfully applied to various human and non-human images from the INRIA dataset, and it performs better than other related algorithms.
Deok-Yeon Kim, Joon Young Kwak, ByoungChul Ko, Jae-Yeal Nam
ICME3
2011 Modeling and Formalization of Fuzzy Finite Automata for Detection of Irregular Fire Flames
abstract
Fire-flame detection using a video camera is difficult because a flame has irregular characteristics, i.e., vague shapes and color patterns. Therefore, in this paper, we propose a novel fire-flame detection method using fuzzy finite automata (FFA) with probability density functions based on visual features, thereby providing a systemic approach to handling irregularity in computational systems and the ability to handle continuous spaces by combining the capabilities of automata with fuzzy logic. First, moving regions are detected via background subtraction, and the candidate flame regions are then identified by applying flame color models. In general, flame regions have a continuous irregular pattern; therefore, probability density functions are generated for the variation in intensity, wavelet energy, and motion orientation and applied to the FFA. The proposed algorithm is successfully applied to various fire/non-fire videos, and its detection performance is better than that of other methods.
ByoungChul Ko, SunJae Ham, Jae-Yeal Nam
IEEE Trans. Circuits Syst. Video Technol.1
2010 Fire-Flame Detection Based on Fuzzy Finite Automation
abstract
This paper proposes a new fire-flame detection method using probabilistic membership function of visual features and Fuzzy Finite Automata (FFA). First, moving regions are detected by analyzing the background subtraction and candidate flame regions then identified by applying flame color models. Since flame regions generally have an irregular pattern continuously, membership functions of variance of intensity, wavelet energy and motion orientation are generate and applied to FFA. Since FFA combines the capabilities of automata with fuzzy logic, it not only provides a systemic approach to handle uncertainty in computational systems, but also can handle continuous spaces. The proposed algorithm is successfully applied to various fire videos and shows a better detection performance when compared with other methods.
SunJae Ham, ByoungChul Ko, Jae-Yeal Nam
ICPR2
2009 X-Ray Image Classification and Retrieval Using Ensemble Combination of Visual Descriptors
JeongHee Shim, KiHee Park, ByoungChul Ko, Jae-Yeal Nam
PSIVT3
2007 Salient human detection for robot vision
Soo Yeong Kwak, ByoungChul Ko, Hyeran Byun
Pattern Anal. Appl.2
2005 FRIP: a region-based image retrieval tool using automatic image segmentation and stepwise Boolean AND matching
abstract
We present our region-based image retrieval tool, finding region in the picture (FRIP), that is able to accommodate, to the extent possible, region scaling, rotation, and translation. Our goal is to develop an effective retrieval system to overcome a few limitations associated with existing systems. To do this, we propose adaptive circular filters used for semantic image segmentation, which are based on both Bayes' theorem and texture distribution of image. In addition, to decrease the computational complexity without losing the accuracy of the search results, we extract optimal feature vectors from segmented regions and apply them to our stepwise Boolean AND matching scheme. The experimental results using real world images show that our system can indeed improve retrieval performance compared to other global property-based or region-of-interest-based image retrieval methods.
ByoungChul Ko, Hyeran Byun
IEEE Trans. Multim.1
2003 Robust Face Detection and Tracking for Real-Life Applications
abstract
In this paper, we propose a new face detection and tracking algorithm for real-life telecommunication applications, such as video conferencing, cellular phone and PDA. We combine template-based face detection and tracking method with color information to track a face regardless of various lighting conditions and complex backgrounds as well as the race. Based on our experiments, we generate robust face templates from wavelet-transformed lowpass and two highpass subimages at the second level low-resolution. However, since template matching is generally sensitive to the change of illumination conditions, we propose a new type of preprocessing method. Tracking method is applied to reduce the computation time and predict precise face candidate region even though the movement is not uniform. Facial components are also detected using k-means clustering and their geometrical properties. Finally, from the relative distance of two eyes, we verify the real face and estimate the size of facial ellipse. To validate face detection and tracking performance of our algorithm, we test our method using six different video categories of QCIF size which are recorded in dynamic environments.
Hyeran Byun, ByoungChul Ko
Int. J. Pattern Recognit. Artif. Intell.2
2003 Extracting Salient Regions And Learning Importance Scores In Region-Based Image Retrieval
abstract
In this paper, we propose a new method for extracting salient regions and learning their importance scores in region-based image retrieval. In Region-Based Image Retrieval (RBIR), not all the regions are important for retrieving similar images and rather, in retrieval, the user is often interested in performing a query on only one or a few regions rather than the whole image. Therefore, for a successful retrieval system, it is an important issue to specify which regions are important for retrieving an image. To extract salient regions from images automatically, we make three assumptions and determine salient regions with their importance scores. In this paper, we apply the relevance feedback algorithm to the matching process as two different purposes: one is for updating importance scores of salient regions and the other is for updating weights of feature vectors. By using our relevance feedback method, the matching process can improve retrieval performance interactively and allow progressive refinement of query results according to the user's feedback action. Through experiments and comparison with other methods, our proposed method shows good performance as well as easy and semantic interface for region-based image retrieval. The efficacy of our method is validated using a set of 3000 images from Corel-photo CD.
ByoungChul Ko, Hyeran Byun
Int. J. Pattern Recognit. Artif. Intell.1
2002 Automatic text extraction in news images using morphology
InYoung Jang, ByoungChul Ko, Hyeran Byun, Yeongwoo Choi
VCIP2
2001 A New Content-Based Image Retrieval System Using Hang Gesture And Relevange Feedback
ByoungChul Ko, Hyeran Byun
ICME1
2001 Region-based Image Retrieval Using Probabilistic Feature Relevance Learning
ByoungChul Ko, Hyeran Byun
Pattern Anal. Appl.1
2000 Region-Based Image Retrieval System Using Efficient Feature Description
abstract
In this paper we introduce a region-based image retrieval system, FRIP. This system includes a robust image segmentation scheme using scaled and shifted color and shape description scheme using modified radius-based signature. For image segmentation, by using our proposed circular filter, we can keep the boundary of object naturally and merge small senseless regions of object into a whole body. For efficient shape description, we extract 5 features from each region: color, texture, scale, location, and shape. From these features, we calculate the similarity distance between the query and database regions and it returns the top K-nearest neighbor regions.
ByoungChul Ko, Hae-Sung Lee, Hyeran Byun
ICPR1