Daijin Kim 0001

dblp:k/DaijinKim · DBLP profile ↗
← Back
136ranked-venue papers
12as first author
11since 2021 · last 2023
0000-0002-8046-8521ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 86 · 11 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 57 · 1 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 12Applied, interdisciplinary, general and emerging computing · 12Systems, architecture and hardware · 3Computer networks · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2023 Weight-Based Mask For Domain Adaptation
abstract
In computer vision, unsupervised domain adaptation (UDA) is an approach to transferring knowledge from a label-rich source domain to a fully-unlabeled target domain. Conventional UDA approaches have two problems. The first problem is that a class classifier can be biased to the source domain because it is trained using only source samples. The second is that previous approaches align image-level features regardless of foreground and background, although the classifier requires foreground features. To solve these problems, we introduce Weight-based Mask Network (WEMNet) composed of Domain Ignore Module (DIM) and Semantic Enhancement Module (SEM). DIM obtains domain-agnostic feature representations via the weight of the domain discriminator and predicts categories. In addition, SEM obtains class-related feature representations using the classifier weight and focuses on the foreground features for domain adaptation. Extensive experimental results reveal that the proposed WEMNet outperforms the competitive accuracy on representative UDA datasets.
Eunseop Lee, Inhan Kim, Daijin Kim 0001
ICASSP3
2022 Learning Mixture of Domain-Specific Experts via Disentangled Factors for Autonomous Driving
Inhan Kim, Joonyeong Lee, Daijin Kim 0001
AAAI3
2022 Revisiting Image Pyramid Structure for High Resolution Salient Object Detection
Kunhee Kim, Joonyeong Lee, Dongmin Cha, Daijin Kim 0001
ACCV (7)6
2022 A Style-aware Discriminator for Controllable Image Translation
abstract
Current image-to-image translations do not control the output domain beyond the classes used during training, nor do they interpolate between different domains well, leading to implausible results. This limitation largely arises because labels do not consider the semantic distance. To mitigate such problems, we propose a style-aware discriminator that acts as a critic as well as a style encoder to provide conditions. The style-aware discriminator learns a controllable style space using prototype-based self-supervised learning and simultaneously guides the generator. Experiments on multiple datasets verify that the proposed model outperforms current state-of-the-art image-to-image translation methods. In contrast with current methods, the proposed approach supports various applications, including style interpolation, content transplantation, and local image translation. The code is available at github.com/kunheek/style-aware-discriminator.
Kunhee Kim, Sanghun Park, Eunyeong Jeon, Daijin Kim 0001
CVPR5
2022 Object Discovery via Contrastive Learning for Weakly Supervised Object Detection
Jinhwan Seo, Wonho Bae, Danica J. Sutherland, Junhyug Noh, Daijin Kim 0001
ECCV (31)5
2022 Systematized event-aware learning for multi-object tracking
abstract
We propose an end-to-end online multi-object tracking (MOT) framework with a systematized event-aware loss, which is designed to control possible occurrences in an online MOT situation and compel the tracker to take appropriate actions when such events occur. Training samples from real candidates using a simulation tracker are generated, and a systematized event-aware association matrix is constructed for every frame to enable the tracker to learn the ideal action in a running environment. Several experiments, including ablation studies on various public MOT benchmark datasets, are conducted. The experimental results verify that each event affecting the tracking measure can be controlled, and the proposed method presents optimal results compared with recent state-of-the-art MOT methods.
Hyemin Lee, Daijin Kim 0001
UAI2
2021 FA-GAN: Feature-Aware GAN for Text to Image Synthesis
abstract
Text-to-image synthesis aims to generate a photo-realistic image from a given natural language description. Previous works have made significant progress with Generative Adversarial Networks (GANs). Nonetheless, it is still hard to generate intact objects or clear textures (Fig 1). To address this issue, we propose Feature-Aware Generative Adversarial Network (FA-GAN) to synthesize a high-quality image by integrating two techniques: a self-supervised discriminator and a feature-aware loss. First, we design a self-supervised discriminator with an auxiliary decoder so that the discriminator can extract better representation. Secondly, we introduce a feature-aware loss to provide the generator more direct supervision by employing the feature representation from the self-supervised discriminator. Experiments on the MSCOCO dataset show that our proposed method significantly advances the state-of-the-art FID score from 28.92 to 24.58.
Eunyeong Jeon, Kunhee Kim, Daijin Kim 0001
ICIP3
2021 SpaceMeshLab: Spatial Context Memoization And Meshgrid Atrous Convolution Consensus For Semantic Segmentation
abstract
Semantic segmentation networks adopt transfer learning from image classification networks causing a shortage of spatial context information. For this reason, we propose Spatial Context Memoization (SpaM), a bypassing branch for spatial context by retaining the input dimension and constantly communicating its spatial context and rich semantic information mutually with the backbone network. Multi-scale context information for semantic segmentation is crucial for dealing with diverse sizes and shapes of target objects in the given scene. Conventional multi-scale context scheme adopts multiple effective receptive fields by multiple dilation rates or pooling operations, but often suffer from misalignment problem with respect to the target pixel. To this end, we propose Meshgrid Atrous Convolution Consensus (MetroCon2) which brings multi-scale scheme into fine-grained multi-scale object context using convolutions with meshgrid-like scattered dilation rates. SpaceMeshLab (ResNet-101 + SpaM + MetroCon2) achieves 82.0% mIoU in Cityscapes test and 53.5% mIoU on Pascal-Context val set.
Jinseong Kim, Daijin Kim 0001
ICIP3
2021 Localization Uncertainty-Based Attention For Object Detection
abstract
Object detection has been applied in a wide variety of real world scenarios, so detection algorithms must provide confidence in the results to ensure that appropriate decisions can be made based on their results. Accordingly, several studies have investigated the probabilistic confidence of bounding box regression. However, such approaches have been restricted to anchor-based detectors, which use box confidence values as additional screening scores during non-maximum suppression (NMS) procedures. In this paper, we propose a more efficient uncertainty-aware dense detector (UADET) that predicts four-directional localization uncertainties via Gaussian modeling. Furthermore, a simple uncertainty attention module (UAM) that exploits box confidence maps is proposed to improve performance through feature refinement. Experiments using the MS COCO benchmark show that our UADET consistently surpasses baseline FCOS, and that our best model, ResNext-64x4d-101-DCN, obtains a single model, single-scale AP of 48.3% on COCO test-dev, thus achieving the state-of-the-art among various object detectors.
Sanghun Park, Kunhee Kim, Eunseop Lee, Daijin Kim 0001
ICIP4
2021 UACANet: Uncertainty Augmented Context Attention for Polyp Segmentation
abstract
We propose Uncertainty Augmented Context Attention network (UACANet) for polyp segmentation which considers an uncertain area of the saliency map. We construct a modified version of U-Net shape network with additional encoder and decoder and compute a saliency map in each bottom-up stream prediction module and propagate to the next prediction module. In each prediction module, previously predicted saliency map is utilized to compute foreground, background and uncertain area map and we aggregate the feature map with three area maps for each representation. Then we compute the relation between each representation and each pixel in the feature map. We conduct experiments on five popular polyp segmentation benchmarks, Kvasir, CVC-ClinicDB, ETIS, CVC-ColonDB and CVC-300, and our method achieves state-of-the-art performance. Especially, we achieve 76.6% mean Dice on ETIS dataset which is 13.8% improvement compared to the previous state-of-the-art method. Source code is publicly available at https://github.com/plemeri/UACANet
Hyemin Lee, Daijin Kim 0001
ACM Multimedia3
2021 ACN: Occlusion-tolerant face alignment by attentional combination of heterogeneous regression networks
Hyunsung Park, Daijin Kim 0001
Pattern Recognit.2
2020 Multi-task Learning with Future States for Vision-Based Autonomous Driving
Inhan Kim, Hyemin Lee, Joonyeong Lee, Eunseop Lee, Daijin Kim 0001
ACCV (3)5
2020 VAN: Versatile Affinity Network for End-to-End Online Multi-object Tracking
Hyemin Lee, Inhan Kim, Daijin Kim 0001
ACCV (2)3
2020 Spatio-Temporal Slowfast Self-Attention Network For Action Recognition
abstract
We propose Spatio-Temporal SlowFast Self-Attention network for action recognition. Conventional Convolutional Neural Networks have the advantage of capturing the local area of the data. However, to understand a human action, it is appropriate to consider both human and the overall context of given scene. Therefore, we repurpose a self-attention mechanism from Self-Attention GAN (SAGAN) to our model for retrieving global semantic context when making action recognition. Using the self-attention mechanism, we propose a module that can extract four features in video information: spatial information, temporal information, slow action information, and fast action information. We train and test our network on the Atomic Visual Actions (AVA) dataset and show significant frame-AP improvements on 28 categories.
Myeongjun Kim, Daijin Kim 0001
ICIP3
2020 Fusion of Saliency Map and Deep Feature-Based Correlation Filter for Enhancing Tracking Performances
abstract
This paper proposes the fusion of a saliency map and a deep feature-based correlation filter to enhance tracking accuracy by reflecting spatial attention and foreground information in the tracking process. The saliency map enables the tracker to focus on a salient object region. The target foreground region is roughly segmented from the background region using a pixel-wise likelihood map derived from the color model and the shape model. Given that only visual information in the foreground region is used, the model is prevented from learning background, and the tracker is made robust to background change. The proposed saliency map can be easily combined with any tracking methods by providing the map as the weight value to the target response map. The saliency map is combined with a correlation filter-based tracker, and we prove that the saliency map successfully improves tracking performance. Experiments are conducted to validate the proposed method on public benchmark datasets. The proposed method achieves remarkable results compared with existing state-of-the-art trackers and successfully improves the tracking performance combined with two version of correlation filter-based methods.
Hyemin Lee, Daijin Kim 0001
ICIP2
2020 A complementary regression network for accurate face alignment
Hyunsung Park, Daijin Kim 0001
Image Vis. Comput.2
2020 A CNN-based 3D human pose estimation based on projection of depth and ridge data
Yeonho Kim, Daijin Kim 0001
Pattern Recognit.2
2019 Attentional Feature-Pair Relation Networks for Accurate Face Recognition
abstract
Human face recognition is one of the most important research areas in biometrics. However, the robust face recognition under a drastic change of the facial pose, expression, and illumination is a big challenging problem for its practical application. Such variations make face recognition more difficult. In this paper, we propose a novel face recognition method, called Attentional Feature-pair Relation Network (AFRN), which represents the face by the relevant pairs of local appearance block features with their attention scores. The AFRN represents the face by all possible pairs of the 9x9 local appearance block features, the importance of each pair is considered by the attention map that is obtained from the low-rank bilinear pooling, and each pair is weighted by its corresponding attention score. To increase the accuracy, we select top-K pairs of local appearance block features as relevant facial information and drop the remaining irrelevant. The weighted top-K pairs are propagated to extract the joint feature-pair relation by using bilinear attention network. In experiments, we show the effectiveness of the proposed AFRN and achieve the outstanding performance in the 1:1 face verification and 1:N face identification tasks compared to existing state-of-the-art methods on the challenging LFW, YTF, CALFW, CPLFW, CFP, AgeDB, IJB-A, IJB-B, and IJB-C datasets.
Bong-Nam Kang, Bongjin Jun, Daijin Kim 0001
ICCV4
2019 Accurate traffic light detection using deep neural network with focal regression loss
Eunseop Lee, Daijin Kim 0001
Image Vis. Comput.2
2018 BAN: Focusing on Boundary Context for Object Detection
Taewook Kim 0003, Bong-Nam Kang, Daijin Kim 0001
ACCV (6)5
2018 Pairwise Relational Networks for Face Recognition
Bong-Nam Kang, Daijin Kim 0001
ECCV (2)3
2018 SAN: Learning Relationship Between Convolutional Features for Multi-scale Object Detection
Bong-Nam Kang, Daijin Kim 0001
ECCV (5)3
2018 Salient Region-Based Online Object Tracking
abstract
In this paper, we propose a salient region-based tracking method that discriminates the exact target region from background by using a probabilistic color model. The color model is updated using image pixels included in salient region. From the extracted salient region, we derive shape model which can be combined with color model that enable the tracker to be robust when the color distribution of target object is similar with other objects. Additionally, we adopt template matching weighted by the shape model to discriminate the target when the background has very similar color distribution with target object. The weight between color matching and template matching is automatically determined based on the confidence of the response map. The proposed method is robust to scale change, object transformation, and rotation. In experiments on public datasets, the proposed method achieved a higher result compared with existing state-of-the-art methods in terms of Expected Overlap Ratio (EAO) only using color model and template matching. The internal analysis proves that the combination of salient region and shape model can increase the tracking performance.
Hyemin Lee, Daijin Kim 0001
WACV2
2018 Real-time dance evaluation by markerless human pose estimation
Yeonho Kim, Daijin Kim 0001
Multim. Tools Appl.2
2017 Detector with focus: Normalizing gradient in image pyramid
abstract
An image pyramid can extend many object detection algorithms to solve detection on multiple scales. However, interpolation during the resampling process of an image pyramid causes gradient variation, which is the difference of the gradients between the original image and the scaled images. Our key insight is that the increased variance of gradients makes the classifiers have difficulty in correctly assigning categories. We prove the existence of the gradient variation by formulating the ratio of gradient expectations between an original image and scaled images, then propose a simple and novel gradient normalization method to eliminate the effect of this variation. The proposed normalization method reduce the variance in an image pyramid and allow the classifier to focus on a smaller coverage. We show the improvement in three different visual recognition problems: pedestrian detection, pose estimation, and object detection. The method is generally applicable to many vision algorithms based on an image pyramid with gradients.
Bong-Nam Kang, Daijin Kim 0001
ICIP3
2017 Robust pedestrian detection under deformation using simple boosted features
Hak Kyoung Kim, Daijin Kim 0001
Image Vis. Comput.2
2017 A variety of local structure patterns and their hybridization for accurate eye detection
Inho Choi, Daijin Kim 0001
Pattern Recognit.2
2017 Robust human activity recognition from depth video using spatiotemporal multi-fused features
Ahmad Jalal, Yeonho Kim, Yong-Joong Kim, Shaharyar Kamal, Daijin Kim 0001
Pattern Recognit.5
2016 A two-stage foreground propagation for moving object detection in a non-stationary
abstract
In this paper, we propose a two-stage foreground propagation that uses clues to adapt to the environment and detect moving objects in a non-stationary camera. The first stage creates a weight matrix to instantaneously regulate the background model by responding to clues from frame differencing and background subtraction. The regulated background model is less affected by inaccurate motion compensation. In the second stage, an iterative approach is taken to refine the threshold for each pixel location by initially using pixels with high foreground probability as clues. Foreground regions detected from the refined threshold are less likely to be false detections and capture true object regions with completeness. Experimental results showed that the two-stage foreground propagation had significantly higher recall with comparable precision and outperformed other methods.
WonTaek Chung, Yong-Joong Kim, Daijin Kim 0001
AVSS4
2016 Patch-based visual microphone for improving quality of sound
abstract
Visual microphone is a technique introduced to recover the sound from a silence video. And traditional method of sound recovery involves extracting and combining subtle motion signals from the entire image. However, there are two possible drawbacks of recovering the sound using the entire image. First, motion signals extracted from plain and edge regions may contain noise due to an aperture problem. Although plain regions are penalised by their squared amplitude values, summation of all the pixels present in the image could introduce notable amount of noise. Second, it is unclear which part of the surface of an object is hit by the sound wave. Utilizing only the region hit by the sound wave is expected to lead to a better sound recovery. The proposed patch-based visual microphone framework addresses these two problems by recovering the sound from a sub-region (patch) in the image centered at a key point (corner). Since we are unable to know which sub-region in the image is good for sound recovery, speeches are recovered from patches centered at each key point (corner), and then the best speech with the least noise is selected as a recovered speech. Extensive experiment results show that utilizing motion signals from a small region in the image near a key point can improve quality of the recovered speech.
Juhyun Ahn, Yong-Joong Kim, Daijin Kim 0001
ICPR3
2016 Integrating hidden Markov models based on Mixture-of-Templates and k-NN2 ensemble for activity recognition
abstract
This paper considers the activity recognition problem using inertial sensor data. It is a challenging temporal pattern recognition problem as the sensor data can be easily mixed with noise and also has large intra-class variation, resulting from different characteristic of people doing same activity, and interclass similarity among several similar activities. To handle these problems concentrating on the classification method, this paper proposes a novel ensemble scheme of hidden Markov models for activity recognition. To improve the performance of activity recognition, our method models the outputs of multiple hidden Markov models by using enhanced template-based classifier fusion method, in which multiple local templates are generated as Mixture-of-Templates, and k-NN2ensemble method is proposed to elaborately recognize activities. To show the effectiveness of the proposed method, we carried out several experiments on UCI Human Activity Recognition dataset and compared our method with several alternative methods. As a result, our method outperforms other methods.
Yong-Joong Kim, Juhyun Ahn, Daijin Kim 0001
ICPR4
2016 Deep convolution neural network with stacks of multi-scale convolutional layer block using triplet of faces for face recognition in the wild
abstract
Recently, deep convolutional neural networks have set a new trend in fields of face recognition by improving the state-of-the-art performance. By using deep neural networks, much more sophisticated and high level abstracted features can be learned automatically. In this paper, we propose a method for face recognition using multi-scale convolution layer blocks and triplets of faces in unconstrained environments. We use the ensemble of deep convolution neural networks trained on differently scaled and aligned face images. This extracts low dimensional but high-level abstraction and discriminative features for face recognition. With these features, we employ the jointly Bayesian model and transfer learning which adapts the knowledge trained from the source domain to target domain. Experiment shows that our proposed method achieves 98.33% pair-wise verification accuracy on the LFW dataset.
Bong-Nam Kang, Daijin Kim 0001
SMC3
2016 Face spoofing detection with highlight removal effect and distortions
abstract
With rapid development of face recognition and detection techniques, the face has been frequently used as a biometric to find illegitimate access. It relates to a security issues of system directly, and hence, the face spoofing detection is an important issue. However, correctly classifying spoofing or genuine faces is challenging due to diverse environment conditions such as brightness and color of a face skin. Therefore we propose a novel approach to robustly find the spoofing faces using the highlight removal effect, which is based on the reflection information. Because spoofing face image is recaptured by a camera, it has additional light information. It means that spoofing image could have much more highlighted areas and abnormal reflection information. By extracting these differences, we are able to generate features for robust face spoofing detection. In addition, the spoofing face image and genuine face image have distinct textures because of surface material of medium. The skin and spoofing medium are expected to have different texture, and some genuine image characteristics are distorted such as color distribution. We achieve state-of-the-art performance by concatenating these features. It significantly outperforms especially for the error rate.
Inhan Kim, Juhyun Ahn, Daijin Kim 0001
SMC3
2016 Detecting Korean characters in natural scenes by alphabet detection and agglomerative character construction
abstract
This paper considers the Korean character detection problem. Unlike English where an alphabet constitutes a character, the Korean character is composed of more than two Korean alphabets, where they could be either connected or separated, relying on the Korean character font. Also, the Korean has two character structures which constitute a nested structure. These properties make the Korean character detection problem difficult. In this paper, we divide the Korean character detection problem into two subproblems, Korean alphabet detection and Korean character construction, and redefine the Korean character structures to efficiently detect Korean characters. Based on the new structures, we train four independent Korean alphabet detectors, and perform a sequential alphabet detection process with a specific detection order, to eliminate false alarms caused during detection procedure. Finally, the detected alphabets are grouped into the Korean characters by an agglomerative character construction algorithm. To evaluate our method, we carried out some experiments on a public dataset with several alternatives, and showed that our proposed Korean character detection method has outperformed other methods.
Jangho Kim, Yong-Joong Kim, Daijin Kim 0001
SMC4
2016 A deep neural network based hashing for efficient image retrieval
abstract
Learning valid similarities is a vital problem in hashing methods, especially in large-scale image search. Similarity pertained hashing method is widely used in image retrieval for its high quality compact binary code mapping. The hashing scheme of most existing hashing methods is that the input data is encoded as a vector of visual features and hashed into binary hash codes via projection functions or quantization methods afterward. However, this separated pipeline may prone to lose accurate similarities of images, since the limited compatible domain between visual feature vectors generation and binary codes mapping process. Encouraged by the extraordinary image representation learning ability of deep neural networks in classification, we propose a structure that merges binary code generation process within deep neural networks for efficient image retrieval. The proposed architecture contains two fundamental blocks. The stacked convolution layers of Network In Network with global average pooling compute effective image representation and the embedded latent layer with binary activation functions learn binary hash codes simultaneously. Experiments show that the proposed method gains improvement over several state-of-the-art hashing methods.
Bong-Nam Kang, Daijin Kim 0001
SMC3
2016 Robust face alignment and tracking by combining local search and global fitting
Jongju Shin, Daijin Kim 0001
Image Vis. Comput.2
2015 Scene text detection with robust character candidate extraction method
abstract
The maximally stable extremal region (MSER) method has been widely used to extract character candidates, but because of its requirement for maximum stability, high text detection performance is difficult to obtain. To overcome this problem, we propose a robust character candidate extraction method that performs ER tree construction, sub-path partitioning, sub-path pruning, and character candidate selection sequentially. Then, we use the AdaBoost trained character classifier to verify the extracted character candidates. Then, we use heuristics to refine the classified character candidates and group the refined character candidates into text regions according to their geometric adjacency and color similarity. We also apply the proposed text detection method to two different color channels Crand Cband obtain the final detection result by combining the detection results on the three different channels. The proposed text detection method on ICDAR 2013 dataset achieved 8%, 1%, and 4% improvements in recall rate, precision rate and f-score, respectively, compared to the state-of-the-art methods.
Myung-Chul Sung, Bongjin Jun, Hojin Cho, Daijin Kim 0001
ICDAR4
2015 Combined Document/Business Card Detector for Proactive Document-Based Services on the Smartphone
Yong-Joong Kim, Bong-Nam Kang, Daijin Kim 0001
ICONIP (4)4
2015 Business card region segmentation by block-based line fitting and largest quadrilateral search with constraints
abstract
In this paper, we propose a novel segmentation method to extract business card region from the image. In our method, an input image is partitioned into four blocks and the probabilistic Hough transform is applied to each block to detect the line segment candidates of business card boundary. Then, our method searches the largest quadrilateral, which is formed by using the detected line segments, under some constraints through RANSAC-like method. To evaluate the proposed method, we test our method on the collected business card images having various kinds of backgrounds. As a result of experiments, we show that our method has achieved about segmentation rate of 90%.
Yong-Joong Kim, Insu Kim, Daijin Kim 0001
ISPA3
2015 Hidden Markov Model Ensemble for Activity Recognition Using Tri-Axis Accelerometer
abstract
Recently, thanks to a variety of sensors equipped on smartphone, a lot of research about mobile activity recognition using accelerometer have been studied for context inference of mobile user and healthcare applications. Previous works, however, have a limitation in classifying some activities because of intra-class variations and inter-class similarities. To handle this problem, in this paper we propose a novel method to recognize activity of smart phone user based on hidden Markov model, where an ensemble method of hidden Markov models is proposed and used to recognize activity. To evaluate our method, we have carried out some experiments by using UCI Human Activity Recognition dataset, and as a result we have achieved about 83.51% accuracy when using two simple features, mean and standard deviation. It is a comparable result to other powerful discriminative methods such as support vector machine and multilayer perceptron.
Yong-Joong Kim, Bong-Nam Kang, Daijin Kim 0001
SMC3
2015 Adaptive Deformation Handling for Pedestrian Detection
abstract
Despite the abundance of successful models for pedestrian detection, many are limited in their ability to handle deformations, such as large appearance variations. In view of insufficient number of models with the ability to handle deformations, we propose a simple strategy, which incorporates deformation handling with a spatial pyramid method in basic classifier learning. By using the max pooling method, this approach aggregates a set of randomly selected basic features from a local region. The spatial pyramid method has been integrated to our method to construct a richer feature in a local region. We show how to train the model with this deformation handling method using a boosting process. Our best detector outperforms the state-of-the-art of pedestrian detection on the INRIA and the Caltech-USA datasets. It achieves a log average miss rate of 12.21% on the INRIA and a log average miss rate of 24.03% on the Caltech-USA datasets.
Hak Kyoung Kim, Daijin Kim 0001
WACV3
2015 Accurate abandoned and removed object classification using hierarchical finite state machine
Jiman Kim, Daijin Kim 0001
Image Vis. Comput.2
2015 Accurate Human Pose Estimation by Aggregating Multiple Pose Hypotheses Using Modified Kernel Density Approximation
abstract
This letter proposes an accurate human pose estimation method that uses a modified kernel density approximation (m-KDA) to multiple pose hypotheses. Existing methods show poor human pose estimation because of cluttered background or self-occlusion by the human. To improve the pose estimation accuracy, we propose to use m-KDA to aggregate multiple pose estimation results. First, we use the flexible mixture-of-parts model (FMM) to estimate the human poses then use the top-M scores to choose the good pose hypotheses. Second, we aggregate the top-M pose hypotheses with the m-KDA, in which each kernel density function is modified by each pose's score value and each pose's compatibility function that represents how far each pose hypothesis is departed from the nominal value of top-M pose hypotheses. Third, we determine the optimal pose configuration by repeating the above m-KDA computation, starting from the root part (head) to the leaf parts (hands and feet), sequentially. In pose estimation experiments on two benchmark datasets (PARSE and LSP), the proposed method achieved 1.5-4.0% improvement in the percentage of correct localized parts (PCP) over the state-of-the-art methods.
Eunji Cho, Daijin Kim 0001
IEEE Signal Process. Lett.2
2014 Static region classification using hierarchical finite state machine
abstract
The ability of most existing approaches to classify static regions such as abandoned and removed objects in images is affected by illumination and traffic volume because of several predefined threshold values. To reduce these effects, we propose an accurate static region classification method using a hierarchical finite state machine that consists of three layers. Each FSM is defined by a Mealy state machine, where a support vector machine (SVM) determines the state transition based on the current state and input features. Because the proposed method uses optimally trained by SVM classifiers, it does not require threshold values and guarantees better classification accuracy under severe environmental changes. In experiments, the proposed method provided much higher classification accuracy and lower false alarm rate than the state-of-the-art methods.
Jiman Kim, Daijin Kim 0001
ICIP2
2014 Accurate Static Region Classification Using Multiple Cues for ARO Detection
abstract
This letter proposes an accurate static region classification for detecting abandoned or removed objects (ARO) using multiple cues. Most existing ARO detection approaches show many falsely detected static regions and low ARO detection performance in real situations because they use single cue and a number of pre-defined threshold values. The proposed method presents multiple cues as intensity, motion, and shape to characterize the true static regions and classifies their candidates into true/false static regions using a SVM classifier, which avoids any dependency on pre-defined threshold values. Experimental results show that the proposed method achieved better ARO detection accuracy and lower false detection rate than the existing methods. In addition, the proposed method can be utilized to several practical applications such as illegal parking detection, garbage throwing detection, thief detection, forest fire detection, and camouflaged solider detection.
Jiman Kim, Daijin Kim 0001
IEEE Signal Process. Lett.2
2014 Hybrid Approach for Facial Feature Detection and Tracking under Occlusion
abstract
When a face is partially occluded in an image, the existing discriminative or generative methods often do not find facial features. This is due to the limitations of local facial feature detectors and appearance modeling in discriminative and generative methods, respectively. To solve this problem, we propose a new facial feature detection method that hybridizes the discriminative and generative methods. The proposed method consists of an initialization stage and optimization stage. The initialization stage detects the face, estimates the facial pose, and obtains the initial parameter set by locating the pose-specific mean shape on the detected face. The optimization stage obtains the facial features by updating the parameter set using the combined Hessian matrix and gradient vector of shape and appearance errors obtained from two methods. Further, we extend the proposed facial feature detection to face tracking by adding a template face obtained from the previous image frame. In experiments, the proposed method yields more accurate facial feature detection or tracking under heavy occlusions and pose variations than the existing methods.
Jongju Shin, Daijin Kim 0001
IEEE Signal Process. Lett.2
2013 Accurate eye detection using generalized binary pattern
abstract
This paper proposes the eye detection using generalized binary pattern (GBP). The GBP can generate all possible patterns of ordered comparisons within a 3 × 3 neighborhood. Since existing local structure patterns consider all neighboring pixels around a given pixel, the number of possible patterns is fixed and limited to 2n. However, since the GBP takes the ordered comparisons of some partial neighboring pixels around a given pixel, a total of 502 different types can be generated in 3 × 3 block. So, our proposed GBP generates 19,162 binary patterns at the given pixel. Among the possible binary patterns, we take an effective set of pattern and position by the AdaBoost feature selection algorithm. Experimental results shows that the GBP provides higher eye detection accuracy on the BioID and FERET databases than other existing local structure patterns such as LBP and MCT.
Inho Choi, Hyunsung Park, Daijin Kim 0001
RO-MAN3
2013 Fast moving object detection with non-stationary background
Jiman Kim, Xiaofei Wang 0001, Chunsheng Zhu, Daijin Kim 0001
Multim. Tools Appl.5
2013 Local Transform Features and Hybridization for Accurate Face and Human Detection
abstract
We propose two novel local transform features: local gradient patterns (LGP) and binary histograms of oriented gradients (BHOG). LGP assigns one if the neighboring gradient of a given pixel is greater than its average of eight neighboring gradients and zero otherwise, which makes the local intensity variations along the edge components robust. BHOG assigns one if the histogram bin has a higher value than the average value of the total histogram bins, and zero otherwise, which makes the computation time fast due to no further postprocessing and SVM classification. We also propose a hybrid feature that combines several local transform features by means of the AdaBoost method, where the best feature having the lowest classification error is sequentially selected until we obtain the required classification performance. This hybridization makes face and human detection robust to global illumination changes by LBP, local intensity changes by LGP, and local pose changes by BHOG, which considerably improves detection performance. We apply the proposed features to face detection using the MIT+CMU and FDDB databases and human detection using the INRIA and Caltech databases. Our experimental results indicate that the proposed LGP and BHOG feature attain accurate detection performance and fast computation time, respectively, and the hybrid feature improves face and human detection performance considerably.
Bongjin Jun, Inho Choi, Daijin Kim 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Generalized Binary Pattern for Eye Detection
abstract
This letter proposes a novel local structure pattern, the generalized binary pattern (GBP), which can represent all possible binary patterns of ordered comparisons within a 3 × 3 neighborhood. Since most existing local structure patterns consider ordered comparisons of all neighboring pixels (8 or 9 pixels) around a given pixel, the number of possible binary patterns is fixed and limited to 256 (or 511). In contrast, our proposed GBP takes the ordered comparisons of some partial neighboring pixels (2 to 9 pixels) around a given pixel, which generates a total of 502 different structure types and thus extends the number of possible binary patterns to 19,162. Among these possible binary patterns, we choose an effective set of binary patterns for a given problem by means of the AdaBoost feature selection method. Experimental results show that our proposed GBP provides higher eye detection accuracy on the BioID and LFW face databases than other existing local structure patterns.
Inho Choi, Daijin Kim 0001
IEEE Signal Process. Lett.2
2012 Abnormal Object Detection Using Feedforward Model and Sequential Filters
abstract
Abnormal object detection and discrimisnation is a critical research area for vision-based surveillance systems. This paper proposes a novel algorithm for the detection and discrimination of abnormal objects, such as abandoned and stolen objects. The proposed algorithm consists of three stages and three different filters. The three stages cooperate with each other using the feedforward model to enhance detection and discrimination performance, while the sequential filters efficiently reject falsely detected regions using three categories of information. The results of experiments conducted using public datasets indicate that the proposed algorithm is more accurate and has a lower false alarm ratio than the existing system.
Jiman Kim, Bong-Nam Kang, Daijin Kim 0001
AVSS4
2012 Robust face detection using local gradient patterns and evidence accumulation
Bongjin Jun, Daijin Kim 0001
Pattern Recognit.2
2012 Design and Implementation of a Pipelined Datapath for High-Speed Face Detection Using FPGA
abstract
This paper presents design and implementation of a pipelined datapath for real-time face detection using cascades of boosted classifiers. We propose following methods: symmetric image downscaling, classifier sharing, and cascade merging, to achieve the desired processing speed and area efficiency. First, an image pyramid with 16 levels is generated from the input image to simultaneously detect faces with different scales. The downscaled images are then transferred to the first stage of the cascade that is shared between the corresponding image pairs based on the pixel validity of the symmetric image pyramid. The last method exploits the different hit ratios of the cascade stages. We use a tree-structured cascade of classifiers since most of the nonface elements are eliminated during the early stages of the classifier. The use of a synthesis tool confirms that the proposed design reduces resource utilization by one-eighth without accuracy loss, compared to the fully parallelized implementation of the same algorithm. We implemented the proposed hardware architecture on a Xilinx Virtex-5 LX330 FPGA. The indicative throughput is 307 frames/s irrespective of the number of faces in the scene for standard VGA (640 × 480) images with an operating frequency of 125.59 MHz. We may ensure that face detection results are generated at each clock cycle after the initial pipeline delay, using this fully pipelined datapath for tree-structured cascade classifiers.
Seunghun Jin, Dongkyun Kim, Thuy Tuong Nguyen, Daijin Kim 0001, Jaewook Jeon
IEEE Trans. Ind. Informatics4
2011 Separating Occluded Humans by Bayesian Pixel Classifier with Re-weighted Posterior Probability
Yeonho Kim, Daijin Kim 0001
ACIVS3
2011 Eye Detection and Eye Blink Detection Using AdaBoost Learning and Grouping
abstract
This paper proposes a precise eye detection and eye blink detection algorithm. Eye detection combines and separates scanning results based on an MCT-based AdaBoost detector. The algorithm detects eyes by applying the eye detector to eye candidate regions of a face. To eliminate outliers, we select an eye candidate group by grouping eye candidates. A refinement process using the average position of eye candidates in the selected eye candidate group obtains reliable detection results. Eye blink detection uses an MCT-based AdaBoost classifier, which discriminates between opened and closed eyes. The eye detection rate is 99.34% at the 0.1 normalized error on the BioID database. The eye blink detection accuracy is 96% at the 0.03 FAR on our blink database, which contains 400 images. The average processing time is 1 ms and 30 ms in a PC (Core2Duo 3.2GHz) and smart phone (PXA312), respectively.
Inho Choi, Seungchul Han, Daijin Kim 0001
ICCCN3
2011 A compact local binary pattern using maximization of mutual information for face analysis
Bongjin Jun, Daijin Kim 0001
Pattern Recognit.3
2011 Real-time lip reading system for isolated Korean word recognition
Jongju Shin, Daijin Kim 0001
Pattern Recognit.3
2011 A novel illumination-robust face recognition using statistical and non-statistical method
Bongjin Jun, Daijin Kim 0001
Pattern Recognit. Lett.3
2010 Frontal face classifier using AdaBoost with MCT features
abstract
In this paper, we describe how to classify frontal face from the results of face detection which include non-frontal faces. To do this, we use AdaBoost learning method with Modified Census Transform (MCT) to construct a two-class classifier. As a result of that, our frontal face classifier achieves high classification rate above 96% and fast performance about 10 frames/sec in mobile device.
Jongmin Yoon, Daijin Kim 0001
ICARCV2
2010 Moving object detection under free-moving camera
abstract
The detection of moving objects in complex environments with various types of motion is a difficult problem because the camera motion and the object motion are mixed. This paper proposes a moving object detection algorithm that uses motion clustering and classification from only two consecutive image frames captured by a free-moving camera. The proposed moving object detection has no assumption about the camera motion and the environmental conditions. The experimental results show that the proposed moving object detection is accurate within an accuracy of 7 pixels on the average.
Jiman Kim, Guensu Ye, Daijin Kim 0001
ICIP3
2010 Sound source localization in reverberant environment using visual information
abstract
Recently, many researchers have carried out works on audio-video integration. It is worth exploring because service robots are supposed to interact with human beings using both visual and auditory sensors. In this paper, we propose an audio-video method for sound source localization in reverberant environment. Using visual information from a vision camera, we could train our audio localizer to distinguish a real source from fake sources and improved the performance of audio localizer in reverberant environment.
Byoung-gi Lee, Jongsuk Choi, Daijin Kim 0001
IROS3
2010 Pose Robust Human Detection in Depth Images Using Multiply-Oriented 2D Elliptical Filters
abstract
This paper proposes a pose robust human detection and identification method for sequences of stereo images using multiply-oriented 2D elliptical filters (MO2DEFs), which can detect and identify humans regardless of scale and pose. Four 2D elliptical filters with specific orientations are applied to a 2D spatial-depth histogram, and threshold values are used to detect humans. The human pose is then determined by finding the filter whose convolution result was maximal. Candidates are verified by either detecting the face or matching head-shoulder shapes. Human identification employs the human detection method for a sequence of input stereo images and identifies them as a registered human or a new human using the Bhattacharyya distance of the color histogram. Experimental results show that (1) the accuracy of pose angle estimation is about 88%, (2) human detection using the proposed method outperforms that of using the existing Object Oriented Scale Adaptive Filter (OOSAF) by 15–20%, especially in the case of posed humans, and (3) the human identification method has a nearly perfect accuracy.
Sang-Ho Cho, Daijin Kim 0001
Int. J. Pattern Recognit. Artif. Intell.3
2010 A Fast ICP Algorithm for 3-D Human Body Motion Tracking
abstract
Iterative closest point (ICP) algorithm has been widely used for registering the geometry, shape and color of the 3-D meshes. However, ICP requires a long computation time to find the corresponding closest points between the model points and the data points. To overcome this problem, we propose a fast ICP algorithm that consists of two acceleration techniques: hierarchical model point selection (HMPS) and logarithmic data point search (LDPS). HMPS accelerates the search by reducing the search region of the data points corresponding to a model point effectively: it selects the model points in a coarse-to-fine manner and employs the four neighboring closest data points in the upper layer to make the search region for finding the closest data point corresponding to a model point in the lower layer. LDPS accelerates the search by visiting the data points within the search region using 2-D logarithm search. The HMPS method and the LDPS method can be operating separately or together. To evaluate the speed of the proposed ICP, we apply it to the 3-D human body motion tracking. The proposed fast ICP is about 3.17 times faster than the existing ICP such as the K-D tree.
Daijin Kim 0001
IEEE Signal Process. Lett.2
2009 An FPGA-based Parallel Hardware Architecture for Real-Time Face Detection Using a Face Certainty Map
abstract
This paper presents an FPGA-based parallel hardware architecture for real-time face detection. An image pyramid with twenty depth levels is generated using the input image. For these scaled-down images, a local binary pattern transform and feature evaluation are performed in parallel by using the proposed block RAM-based window processing architecture. By sharing the feature look-up tables between two corresponding scaled-down images, we can reduce the use of routing resources by half. For prototyping and evaluation purposes, the hardware architecture was integrated into a Virtex-5 FPGA. The experimental result shows around 300 frames per second speed performance for processing standard VGA (640times480times8) images. In addition, the throughput of the implementation can be adjusted in proportion to the frame rate of the camera, by synchronizing each individual module with the pixel sampling clock.
Seunghun Jin, Dongkyun Kim, Thuy Tuong Nguyen, Bongjin Jun, Daijin Kim 0001, Jaewook Jeon
ASAP5
2009 Pose Robust Human Detection in Depth Image Using Four Directional 2D Elliptical Filters
abstract
This paper proposes a pose robust human detection method for sequences of stereo images using four directional 2D elliptical filters (4D2DEFs), which can detect humans regardless of scale and pose. Four 2D elliptical filters with specific orientations are applied to a 2D spatial depth histogram, and threshold values are used to detect human candidates. These candidates are verified by either detecting the face or matching head-shoulder shapes. Experimental results show that human detection using the proposed method outperforms that of using the existing Object Oriented Scale Adaptive Filter (OOSAF) by 15~20%, especially in the case of posed humans.
Sang-Ho Cho, Jongmin Yoon, Daijin Kim 0001
ISM4
2009 Fast Car/Human Classification Using Triple Directional Edge Property and Local Relations
abstract
For fast and robust car/human classification, novel features, ELR and DELR, using triple directional edge property of objects and an efficient method based on two relations are proposed. The proposed feature is considerably less sensitive to distance, occlusion, and the existence of groups than other existing features. The proposed method using the temporal and spatial relation complements the temporary classification error by historic accumulation and multi-resolution block-region referencing. Because the proposed method has high average classification accuracy at various conditions of distance, occlusion, and group and fast processing time, it is useful to the real-time surveillance embedding system.
Jiman Kim, Daijin Kim 0001
ISM2
2009 MMI-Based Optimal LBP Code Selection for Face Recognition
abstract
Many variants of local binary patterns (LBPs) are widely used for face analysis due to their inherent simplicity and robustness. However, it has not yet been proven that LBPsare optimal for this task in regards to achieving the best balance between minimizing code numbers and reducing classification error. We propose an effective code selection method for selecting optimal LBP (OLBP) based on the maximization of mutual information (MMI) between features and class labels. We demonstrate the effectiveness of the proposed OLBP through several face recognition experiments. Experimental results show that the OLBP outperforms other features such as LBP, ULBP, and MCT in terms of minimizing the number of codes and reducing the classification error.
Jongmin Yoon, Daijin Kim 0001
ISM3
2009 Real-time facial expression recognition using STAAM and layered GDA classifier
Jaewon Sung, Daijin Kim 0001
Image Vis. Comput.2
2009 Tensor-Based AAM with Continuous Variation Estimation: Application to Variation-Robust Face Recognition
abstract
The Active appearance model (AAM) is a well-known model that can represent a non-rigid object effectively. However, the fitting result is often unsatisfactory when an input image deviates from the training images due to its fixed shape and appearance model. To obtain more robust AAM fitting, we propose a tensor-based AAM that can handle a variety of subjects, poses, expressions, and illuminations in the tensor algebra framework, which consists of an image tensor and a model tensor. The image tensor estimates image variations such as pose, expression, and illumination of the input image using two different variation estimation techniques: discrete and continuous variation estimation. The model tensor generates variation-specific AAM basis vectors from the estimated image variations, which leads to more accurate fitting results. To validate the usefulness of the tensor-based AAM, we performed variation-robust face recognition using the tensor-based AAM fitting results. To do, we propose indirect AAM feature transformation. Experimental results show that tensor-based AAM with continuous variation estimation outperforms that with discrete variation estimation and conventional AAM in terms of the average fitting error and the face recognition rate.
Hyung-Soo Lee, Daijin Kim 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Natural facial expression recognition using differential-AAM and manifold learning
Yeongjae Cheon, Daijin Kim 0001
Pattern Recognit.2
2009 Subtle facial expression recognition using motion magnification
Sungsoo Park, Daijin Kim 0001
Pattern Recognit. Lett.2
2009 Adaptive active appearance model with incremental learning
Jaewon Sung, Daijin Kim 0001
Pattern Recognit. Lett.2
2008 An efficient and accurate hierarchical ICIA fitting method for 3D Morphable Models
abstract
We propose the efficient and accurate hierarchical ICIA fitting method for 3D Morphable Models (3DMMs). The conventional ICIA fitting method for 3DMMs requires a long computation time because the 3D face model contains a large number of vertices and it also requires to compute the Hessian matrix using the visible vertices every iteration. For the efficient fitting, we use the hierarchical fitting that use a set of multi-resolution 3D face model and the Gaussian image pyramid. For more accurate fitting, we use a two-stage parameter update that only update the rigid and the texture parameters and then update all parameters after the initial convergence. We present several experiment results to prove that our proposed method shows better performance than previous works.
Bong-Nam Kang, Daijin Kim 0001, Hyeran Byun
FG2
2008 Illumination-robust face recognition using tensor-based active appearance model
abstract
In this paper, we propose an illumination-robust face recognition using tensor-based active appearance model (AAM). First, we introduce the tensor-based AAM which generates variation specific AAM basis vectors for improving the fitting performance of AAM. Then, we propose a novel illumination transformation method which adopts ratio image concept into tensor model framework. Experimental results show that the proposed tensor-based AAM improves the performance of face recognition significantly.
Hyung-Soo Lee, Daijin Kim 0001
FG2
2008 The POSTECH face database (PF07) and performance evaluation
abstract
We constructed a face database POSTECH face database (PF07). PF07 contains the true-color face images of 200 people, 100 men and 100 women, representing 320 various images (5 pose variations times 4 expression variations times 16 illumination variations) per person. All of the people in the database are Korean. We also present the results of face recognition experiments under various conditions using three baseline face recognition algorithms in order to provide an example evaluation protocol on the database. The database is expected to be used to evaluate the algorithm of face recognition for Korean people or for people with systematic variations.
Hyoung-Soo Lee, Sungsoo Park, Bong-Nam Kang, Jongju Shin, Hong-Mo Je, Bongjin Jun, Daijin Kim 0001
FG8
2008 Spontaneous facial expression classification with facial motion vectors
abstract
This paper proposes a novel spontaneous facial expression classification method using the facial motion magnification which transforms the subtle facial expressions into the corresponding exaggerated facial expressions. Facial motion magnification consists of four steps: First, we perform the active appearance model (AAM) fitting to extract 70 facial feature points in the face image sequence. Second, we align the face image sequence using the static three feature points. Third, we estimate the motion vectors of 27 feature points using the feature point tracking method. Finally, we obtain the exaggerated facial expressions by magnifying the motion vectors of the 27 feature points. After facial motion magnification, we recognize the exaggerated facial expressions using the support vector machines (SVM) to classify the facial expression features. Experimental results of the subtle facial expression recognition show promising results of the proposed method.
Sungsoo Park, Daijin Kim 0001
FG2
2008 SLAM by Combining Multidimensional Scaling and Particle Filtering
Hong-Mo Je, Daijin Kim 0001
ICIC (1)2
2008 Multi-resolution 3D morphable models and its matching method
abstract
The inverse compositional image alignment (ICIA) is known as an efficient matching method for 3D morphable models (3DMMs). However, it requires a long computation time since the 3D face models consist of a large number of vertices. Also, it requires to recompute the Hessian matrix using the visible vertices every iteration. For a fast and an efficient matching, we propose the efficient and accurate hierarchical ICIA (HICIA) matching method for 3DMMs. The proposed matching method requires multi-resolution 3D face models and the Gaussian image pyramid. The multi-resolution 3D face models are built by sub-sampling at the 2:1 sampling rate to construct the lower-resolution 3D face models. For more accurate matching, we use a two-stage model parameter update that only updates the rigid and the texture parameters and then updates all parameters after the initial convergence. We present several experimental results to prove that the proposed method shows better performance than that of the conventional ICIA matching method.
Bong-Nam Kang, Hyeran Byun, Daijin Kim 0001
ICPR3
2008 Facial expression analysis with facial expression deformation
abstract
In this paper, we proposes an effective and novel approach to recognize subtle facial expression method which is facial expression deformation. The proposed method deforms subtle facial expressions into corresponding extreme facial expressions. Facial expression deformation processes by extracting subtle motion vector of the predefined feature points and amplifying them. By adding amplified motion vector to Active Appearance Models (AAMs) fitted feature points, the extreme facial expression images is recovered (obtained) by the piece-wise affine warping. After facial expression deformation, we extract the shape and appearance features by projecting deformed facial expression image to the AAM shape and appearance model. We use the multi-class Support Vector Machines (SVMs) to classify the shape and appearance features. The facial expression recognition performance shows promising results of the proposed method.
Sungsoo Park, Jongju Shin, Daijin Kim 0001
ICPR3
2008 Enhanced ResolutionAaware Fitting algorithm using interpolation operator
abstract
This paper proposes a new method to speed up the Resolution-Aware Fitting (RAF) algorithm. An interpolation operator is used instead of a blur operator in the RAF algorithm. The RAF algorithm with the interpolation operator is twice as fast as the RAF algorithm with the blur operator without losing fitting accuracy. It concluded that the RAF algorithm with the interpolation operator is superior to the RAF algorithm with the blur operator.
Jongju Shin, Daijin Kim 0001
ICPR2
2008 A Natural Facial Expression Recognition Using Differential-AAM and k-NNS
abstract
This paper proposes a novel natural facial expression recognition method that recognizes a sequence of dynamic facial expression images using the differential active appearance model (AAM) and k-NNS as follows. First, we use the differential-AAM features (DAFs) that are computed from the difference of the AAM parameters between an input face image and a reference face image. Second, we perform the manifold learning. Third, we recognize the facial expression of the input face image in the embedded feature space using sequence based k-NN, k-NNS. Since we use DAFs, we also propose an effective way of finding the neutral facial expression as kernel density approximation. Experimental results show that (1) the DAFs improves the facial expression recognition performance than the conventional AAM features by 20% and (2) the sequence-based k-nearest neighbors classifier provides a 95% of facial expression recognition performance on the facial expression database (FED06).
Yeongjae Cheon, Daijin Kim 0001
ISM2
2008 3D Face Fitting Using Multi-stage Parameter Updating in the 3D Morphable Face Model
abstract
This paper proposes a 3D face fitting algorithm based on the 3D Morphable Face Model. This algorithm updates parameters at each stage in the fitting process. To simplify measurement of parameters, this algorithm updates the parameters in cylindrical coordinates. The parameters are updated using a cost function which is the difference between the input 3D face data and the fitted 3D face model. This proposed algorithm shows good results when shape, texture and extrinsic variations occur in the 3D domain. The average correlation of the fitted shape and texture parameters to the test values are 0.72 and 0.99, respectively. This 3D face fitting algorithm can be widely used for 3D face analysis and 3D face recognition.
Inho Choi, Daijin Kim 0001
ISM2
2008 Pose Robust Face Tracking by Combining Active Appearance Models and Cylinder Head Models
Jaewon Sung, Takeo Kanade, Daijin Kim 0001
Int. J. Comput. Vis.3
2008 Robust head tracking using 3D ellipsoidal head model in particle filter
Sukwon Choi, Daijin Kim 0001
Pattern Recognit.2
2008 Real-time object recognition using relational dependency based on graphical model
Woo-han Yun, Sung Yang Bang, Daijin Kim 0001
Pattern Recognit.3
2008 Expression-invariant face recognition by facial expression transformations
Hyung-Soo Lee, Daijin Kim 0001
Pattern Recognit. Lett.2
2008 Illumination-robust face recognition using ridge regressive bilinear models
Dongsoo Shin, Hyung-Soo Lee, Daijin Kim 0001
Pattern Recognit. Lett.3
2008 Tensor-Based Active Appearance Model
abstract
The active appearance model (AAM) is a well-known model that can represent a nonrigid object effectively. However, the fitting result is often unsatisfactory when the input image has pose, expression, and illumination variations. To overcome this problem, we propose a tensor-based AAM which consists of two kinds of tensors: image tensor and model tensor. The image tensor is used to estimate the image variation such as the pose, the expression, and the illumination by finding the basis subtensor with minimal reconstruction error. The model tensor is used to generate the specific AAM basis vectors by indexing the model tensor in terms of the estimated image variations. Experimental results show that the proposed tensor-based AAM reduces the fitting error of the conventional AAM by about four pixels.
Hyung-Soo Lee, Daijin Kim 0001
IEEE Signal Process. Lett.2
2008 Pose-Robust Facial Expression Recognition Using View-Based 2D + 3D AAM
abstract
This paper proposes a pose-robust face tracking and facial expression recognition method using a view-based 2D$+$3D active appearance model (AAM) that extends the 2D$+$3D AAM to the view-based approach, where one independent face model is used for a specific view and an appropriate face model is selected for the input face image. Our extension has been conducted in many aspects. First, we use principal component analysis with missing data to construct the 2D$+$3D AAM due to the missing data in the posed face images. Second, we develop an effective model selection method that directly uses the estimated pose angle from the 2D$+$3D AAM, which makes face tracking pose-robust and feature extraction for facial expression recognition accurate. Third, we propose a double-layered generalized discriminant analysis (GDA) for facial expression recognition. Experimental results show the following: 1) The face tracking by the view-based 2D$+$3D AAM, which uses multiple face models with one face model per each view, is more robust to pose change than that by an integrated 2D$+$3D AAM, which uses an integrated face model for all three views; 2) the double-layered GDA extracts good features for facial expression recognition; and 3) the view-based 2D$+$3D AAM outperforms other existing models at pose-varying facial expression recognition.
Jaewon Sung, Daijin Kim 0001
IEEE Trans. Syst. Man Cybern. Part A2
2007 Eye Correction Using Correlation Information
Inho Choi, Daijin Kim 0001
ACCV (1)2
2007 Hand Gesture Recognition To Understand Musical Conducting Action
abstract
This paper deals with the understanding of four musical time patterns and three tempos that are generated by a human conductor of robot orchestra or an operator of computer-based music play system using the hand gesture recognition. We use only a stereo vision camera with no extra special devices such as sensor glove, 3D motion capture system, infra-red camera, electronic baton and so on. We propose a simple and reliable vision-based hand gesture recognition using the conducting feature point (CFP), the motion-direction code, and the motion history matching. The proposed hand gesture recognition system operates as follows: First, it extracts the human hand region by segmenting the depth information generated by stereo matching of image sequences. Next, it follows the motion of the center of the gravity(COG) of the extracted hand region and generates the gesture features such as CFP and the direction-code. Finally, we obtain the current timing pattern of the music's beat and tempo by the proposed hand gesture recognition using either CFP tracking or motion histogram matching. The experimental results show that the musical time pattern and tempo recognition rate are over 86% on the test data set when the motion histogram matching is used.
Hong-Mo Je, Jiman Kim, Daijin Kim 0001
RO-MAN3
2007 Real-time 3D Head Tracking and Head Gesture Recognition
abstract
This paper proposes a fast 3D head tracking method that is working robustly under a variety of difficult conditions such as the rapidly changing pose, head movement and illumination. First, we obtain the pose robustness by using the 3D cylindrical head model (CHM) and dynamic template. Second, we also obtain the robustness about the fast head movement by using the dynamic template. Third, we obtain the illumination robustness by modeling the illumination basis vectors and by adding them to the previous input image to adapt the current input image. Experimental results show that the proposed head tracking method outperforms the other tracking method and it tracks the head successfully even if the head moves fast under the rapidly changing poses and illuminations. The proposed head tracking method has a versatile applications such as the head gesture recognition for the human robot interaction, the head gesture TV remote controller for the handicapped people, and the drawing tool by the head movement for the entertainment.
Wooju Ryu, Daijin Kim 0001
RO-MAN2
2007 Combining Local and Global Motion Estimators for Robust Face Tracking
abstract
Faces show both global and local motions, where the former represents rigid head movements due to 3D translation and rotation and the local motion represents non-rigid deformation due to speech, or facial expressions. Although non- rigid face models can represent both types of the facial motions, they are not enough to track the facial motions correctly. The non-rigid face models have large number of model parameters to explain various deformation of the face and the high dimensionality of their model parameter space make them sensitive to initial model parameters, apt to be stuck to local minimum, and difficult to be recovered (re-initialized) from failure when iterative gradient descent optimization techniques are used. To alleviate these problems, we propose to use two types of face trackers that are suitable for estimating the global and local motions, respectively. In the proposed algorithm, the global motion estimator is applied at first and the estimated global motion is used to compute proper initial model parameters of the local motion estimator to make it converge correctly. In this paper, we used active appearance model (MM) and cylinder head model (CHM) as the representative examples of the non- rigid and rigid face models. Experimental results showed that face tracking combining AAMs and CHMs improved the face tracking performance than that of AAMs in terms of 170% higher tracking rate and the 115% wider pose coverage.
Jaewon Sung, Daijin Kim 0001
RO-MAN2
2007 A Unified Gradient-Based Approach for Combining ASM into AAM
Jaewon Sung, Takeo Kanade, Daijin Kim 0001
Int. J. Comput. Vis.3
2007 Robust location tracking using a dual layer particle filter
Keunho Yun, Daijin Kim 0001
Pervasive Mob. Comput.2
2007 Simultaneous gesture segmentation and recognition based on forward spotting accumulative HMMs
Jinyoung Song, Daijin Kim 0001
Pattern Recognit.3
2007 Robust face tracking by integration of two separate trackers: Skin color and facial shape
Hyung-Soo Lee, Daijin Kim 0001
Pattern Recognit.2
2007 A background robust active appearance model using active contour technique
Jaewon Sung, Daijin Kim 0001
Pattern Recognit.2
2006 A Real-Time Face Tracking using the Stereo Active Appearance Model
abstract
This paper proposes a real-time 3D face tracking system. This system uses a stereo active appearance model fitting (STAAM) algorithm that uses multiple calibrated perspective cameras to compute the 3D shape and rigid motion parameters. The use of calibration information reduces the number of model parameters, restricts the degree of freedom in the model parameters, and increases the accuracy and speed of fitting. The real-time face tracking system works alternating two different modes: the detection mode detects the face and eyes in an image and the tracking mode tracks the detected face using the STAAM. In addition, it utilizes a histogram matching to make the system robust to lighting conditions, and uses the motion information to compensate the temporal fluctuation of the captured images. The experimental results show that the proposed system operates robustly under varying lighting conditions and fast in the real-time manner at above 8 frames/sec.
Daijin Kim 0001, Jaewon Sung
ICIP1
2006 Staam: Fitting a 2D+3D AAM to Stereo Images
abstract
This paper proposes a stereo active appearance model fitting algorithm (STAAM), that fits a 2D+3D active appearance model to stereo images acquired from calibrated perspective cameras. The STAAM uses geometrical relationship between cameras to compute the 3D shape and rigid motion parameters. The use of calibration information reduces the number of model parameters, restricts the degree of freedom in the model parameters, and increases the accuracy and speed of fitting. Also, the STAAM uses a modified simultaneous update fitting method that reduces the fitting computation greatly. Experimental results show that (2) the STAAM shows a better fitting stability than the existing multi-view AAM, (2) the modified simultaneous update algorithm accelerates the AAM fitting speed.
Jaewon Sung, Daijin Kim 0001
ICIP2
2006 Large Motion Object Tracking using Active Contour Combined Active Appearance Model
abstract
Because the Active Appearance Model (AAM) is sensitive to the initial parameters, it is difficult to track an object that shows a large motion. This paper proposes an active contour combined Active Appearance Model that can track an object whose motion is large. The proposed AAM fitting algorithm consists of two alternating procedures: active contour fitting to find the contour sample that best fits the face image and then the active appearance model fitting that begins from the estimated motion parameters. Experimental results show that the proposed active contour combined AAM provides better accuracy and convergence characteristics in terms of RMS error and convergence rate than the existing AAM. The combination of the existing robust AAM and the proposed active contour based AAM (AC-R-AAM) had the best accuracy and convergence performances.
Jaewon Sung, Daijin Kim 0001
ICVS2
2006 Gender Classification with Bayesian Kernel Methods
abstract
We consider the gender classification task of discriminating between images of faces of men and women from face images. In appearance-based approaches, the initial images are preprocessed (e.g. normalized) and input into classifiers. Recently, SVMs which are popular kernel classifiers have been applied to gender classification and have shown excellent performance. We propose to use one of Bayesian kernel methods which is Gaussian Process Classifiers (GPCs) for gender classification. The main advantage of Bayesian kernel methods such as GPCs over SVMs is that they determine the hyperparameters of the kernel based on Bayesian model selection criterion. Our results show that GPCs outperformed SVMs with cross validation.
Daijin Kim 0001, Zoubin Ghahramani, Sung Yang Bang
IJCNN2
2006 A Unified Approach for Combining ASM into AAM
Jaewon Sung, Daijin Kim 0001
PSIVT2
2006 A Robust Location Tracking Using Ubiquitous RFID Wireless Network
Keunho Yun, Seokwon Choi, Daijin Kim 0001
UIC3
2006 Appearance-based gender classification with Gaussian processes
Daijin Kim 0001, Zoubin Ghahramani, Sung Yang Bang
Pattern Recognit. Lett.2
2006 Generating frontal view face image for pose invariant face recognition
Hyung-Soo Lee, Daijin Kim 0001
Pattern Recognit. Lett.2
2005 Real-time multiple people tracking using competitive condensation
Heegu Kang, Daijin Kim 0001
Pattern Recognit.2
2005 Face membership authentication using SVM classification tree generated by membership-based LLE data partition
abstract
This paper presents a new membership authentication method by face classification using a support vector machine (SVM) classification tree, in which the size of membership group and the members in the membership group can be changed dynamically. Unlike our previous SVM ensemble-based method, which performed only one face classification in the whole feature space, the proposed method employed a divide and conquer strategy that first performs a recursive data partition by membership-based locally linear embedding (LLE) data clustering, then does the SVM classification in each partitioned feature subset. Our experimental results show that the proposed SVM tree not only keeps the good properties that the SVM ensemble method has, such as a good authentication accuracy and the robustness to the change of members, but also has a considerable improvement on the stability under the change of membership group size.
Shaoning Pang 0001, Daijin Kim 0001, Sung Yang Bang
IEEE Trans. Neural Networks2
2004 Extension of aam with 3d shape model for facial shape tracking
abstract
This paper represents a 3D model-based object tracking algorithm to extract rigid and non-rigid motion of an object simultaneously. Previous AAM is extended by replacing 2D shape model with 3D shape model, and appropriate model fitting algorithm is derived. Proposed algorithm is applied to face tracking in video sequence to extract nonrigidly deforming facial shape.
Jaewon Sung, Daijin Kim 0001
ICIP2
2004 Real-Time Facial Pose Identification With Hierarchically Structured Ml Pose Classifier
abstract
Since pose-varying face images form nonlinear convex manifold in high dimensional image space, it is difficult to model their pose distribution in terms of a simple probabilistic density function. To solve this difficulty, we divide the pose space into many constituent pose classes and treat the continuous pose estimation problem as a discrete pose-class identification problem. We propose to use a hierarchically structured ML (Maximum Likelihood) pose classifiers in the reduced feature space to decrease the computation time for pose identification, where pose space is divided into several pose groups and each group consists of a number of similar neighboring poses. We use the CONDENSATION algorithm to find a newly appearing face and track the face with a variety of poses in real-time. Simulation results show that our proposed pose identification using the hierarchically structured ML pose classifiers can perform a faster pose identification than conventional pose identification using the flat structured ML pose classifiers. A real-time facial pose tracking system is built with high speed hierarchically structured ML pose classifiers.
Jaewon Sung, Daijin Kim 0001
Int. J. Pattern Recognit. Artif. Intell.2
2004 Prediction of the suitability for image-matching based on self-similarity of vision contents
Shaoning Pang 0001, Daijin Kim 0001, Sung Yang Bang
Image Vis. Comput.3
2004 Face recognition using the second-order mixture-of-eigenfaces method
Daijin Kim 0001, Sung Yang Bang, Sang-Youn Lee
Pattern Recognit.2
2003 Face Retrieval Using 1st- and 2nd-order PCA Mixture Model
Sang-Youn Lee, Daijin Kim 0001, YoungSik Choi
ICCSA (2)3
2003 Real-time automatic vehicle management system using vehicle tracking and car plate number identification
abstract
This paper proposes a real-time vehicle management system using a vehicle tracking and a car plate number identification technique. The system uses two cameras: one for tracking vehicles and another for capturing LP (license plate). We track the vehicles by applying the CONDENSATION algorithm over the vehicle's movement image captured from the first camera. To render the CONDENSATION algorithm more effective, we build a discrete vehicle shape model by training vehicle patterns with a SOM (self organizing map), which makes the system suitable for real-time application. Next, we take the probabilistic dynamic model such as HMM (hidden Markov model) to reflect the temporal change in shape of various vehicles. As a vehicle reaches the designated target line, a signal is sent to the second camera for capturing the vehicle's front side. The captured image is transferred to an LPR (vehicle LP recognition system) which recognizes the vehicle's category and LP. LPR system detects the vehicle LP using the only the vertical edge of the captured vehicle image, and effectively accomplishes the character segmentation of the LP region using the geometric transformation without respect to the position and angle of the CCD camera. The segmented characters are recognized using the SVM (support vector machine). By combining these two techniques, we construct a real-time automatic vehicle management system that can be used to control vehicle parking and searching for specific vehicles.
Hwajeong Lee, Daijin Kim 0001, Sung Yang Bang
ICME3
2003 Human Face Detection in Digital Video Using SVMEnsemble
Hong-Mo Je, Daijin Kim 0001, Sung Yang Bang
Neural Process. Lett.2
2003 Extensions of LDA by PCA mixture model and class-wise features
Daijin Kim 0001, Sung Yang Bang
Pattern Recognit.2
2003 Face recognition using the embedded HMM with second-order block-specific observations
Min-Sub Kim, Daijin Kim 0001, Sang-Youn Lee
Pattern Recognit.2
2003 Constructing support vector machine ensemble
Shaoning Pang 0001, Hong-Mo Je, Daijin Kim 0001, Sung Yang Bang
Pattern Recognit.4
2003 An efficient model order selection for PCA mixture model
Daijin Kim 0001, Sung Yang Bang
Pattern Recognit. Lett.2
2003 Face recognition using LDA mixture model
Daijin Kim 0001, Sung Yang Bang
Pattern Recognit. Lett.2
2003 Membership authentication in the dynamic group by face classification using SVM ensemble
Shaoning Pang 0001, Daijin Kim 0001, Sung Yang Bang
Pattern Recognit. Lett.2
2002 Real-time multiple people tracking using competitive condensation
abstract
The CONDENSATION algorithm is attractive as it has robust tracking performance and potential for real-time implementation. However the CONDENSATION tracker has difficulty with real-time implementation for multiple people tracking since it requires complicated shape model and large number of samples for precise tracking performance. This paper presents two improvements for real-time multiple object tracking: the discrete shape model with a small search space and the competition rule which requires a small number of samples to track multiple people. We show that they achieve robust and real-time tracking for image sequences of a crowd of people.
Heegu Kang, Daijin Kim 0001, Sung Yang Bang
ICIP (3)2
2002 Face retrieval using 1st- and 2nd-order PCA mixture model
abstract
This paper deals with face retrieval using the 1st- and 2nd-order PCA mixture model. The well-known eigenface method uses one set of holistic facial features obtained by PCA. However, the single set of eigenfaces is not enough to represent face images with large variations. To overcome this weakness, we propose the method that uses more than one set of eigenfaces obtained from the EM learning in PCA mixture model. 2nd-order eigenface method can be extended to the method using 2nd-order PCA mixture model, also. Simulation results show that the method using 2nd-order PCA mixture model is the best for the face images with illumination variations and the method using 1st-order PCA mixture model is the best for the face images with pose variations in, terms of ANMRR (average of the normalized modified retrieval rank).
Daijin Kim 0001, Sung Yang Bang
ICIP (2)2
2002 A design of CMAC-based fuzzy logic controller with fast learning and accurate approximation
Daijin Kim 0001
Fuzzy Sets Syst.1
2002 An accurate COG defuzzifier design using Lamarckian co-adaptation of learning and evolution
Daijin Kim 0001, YoungSik Choi, Sang-Youn Lee
Fuzzy Sets Syst.1
2002 A numeral character recognition using the PCA mixture model
Daijin Kim 0001, Sung Yang Bang
Pattern Recognit. Lett.2
2002 Face recognition using the mixture-of-eigenfaces method
Daijin Kim 0001, Sung Yang Bang
Pattern Recognit. Lett.2
2001 An optimal design of neuro-FLC by Lamarckian co-adaptation of learning and evolution
Daijin Kim 0001, Han-Pyul Lee
Fuzzy Sets Syst.1
2001 Data classification based on tolerant rough set
Daijin Kim 0001
Pattern Recognit.1
2000 Co-adaptation of self-organizing maps by evolution and learning
Daijin Kim 0001, Sunha Ahn, Dae-Seong Kang
Neurocomputing1
2000 A Handwritten Numeral Character Classification Using Tolerant Rough Set
abstract
Proposes a data classification method based on the tolerant rough set that extends the existing equivalent rough set. A similarity measure between two data is described by a distance function of all constituent attributes and they are defined to be tolerant when their similarity measure exceeds a similarity threshold value. The determination of optimal similarity threshold value is very important for accurate classification. So, we determine it optimally by using the genetic algorithm (GA), where the goal of evolution is to balance two requirements such that: 1) some tolerant objects are required to be included in the same class as many as possible; and 2) some objects in the same class are required to be tolerant as much as possible. After finding the optimal similarity threshold value, a tolerant set of each object is obtained and the data set is grouped into the lower and upper approximation set depending on the coincidence of their classes. We propose a two-stage classification method such that all data are classified by using the lower approximation at the first stage and then the nonclassified data at the first stage are classified again by using the rough membership functions obtained from the upper approximation set. We apply the proposed classification method to the handwritten numeral character classification problem and compare its classification performance and learning time with those of the feedforward neural network's backpropagation algorithm.
Daijin Kim 0001, Sung Yang Bang
IEEE Trans. Pattern Anal. Mach. Intell.1
1999 An accurate and cost-effective COG defuzzifier without the multiplier and the divider
Daijin Kim 0001, In-Hyun Cho
Fuzzy Sets Syst.1
1999 A MS-GS VQ codebook design for wireless image communication using genetic algorithms
abstract
An image compression technique is proposed that attempts to achieve both robustness to transmission bit errors common to wireless image communication, as well as sufficient visual quality of the reconstructed images. Error robustness is achieved by using biorthogonal wavelet subband image coding with multistage gain-shape vector quantization (MS-GS VQ) which uses three stages of signal decomposition in an attempt to reduce the effect of transmission bit errors by distributing image information among many blocks. Good visual quality of the reconstructed images is obtained by applying genetic algorithms (GAs) to codebook generation to produce reconstruction capabilities that are superior to the conventional techniques. The proposed decomposition scheme also supports the use of GAs because decomposition reduces the problem size. Some simulations for evaluating the performance of the proposed coding scheme on both transmission bit errors and distortions of the reconstructed images are performed. Simulation results show that the proposed MS-GS VQ with good codebooks designed by GAs provides not only better robustness to transmission bit errors but also higher peak signal-to-noise ratio even under high bit error rate conditions.
Daijin Kim 0001, Sunha Ahn
IEEE Trans. Evol. Comput.1
1998 Two Co-adaptation Schemes of Evolution and Learning for an Optimal VQ Codebook
Daijin Kim 0001, Sunha Ahn
ICONIP1
1998 Improving the fuzzy system performance by fuzzy system ensemble
Daijin Kim 0001
Fuzzy Sets Syst.1
1997 Forecasting time series with genetic fuzzy predictor ensemble
abstract
This paper proposes a genetic fuzzy predictor ensemble (GFPE) for the accurate prediction of the future in the chaotic or nonstationary time series. Each fuzzy predictor in the GFPE is built from two design stages, where each stage is performed by different genetic algorithms (GA). The first stage generates a fuzzy rule base that covers as many of training examples as possible. The second stage builds fine-tuned membership functions that make the prediction error as small as possible. These two design stages are repeated independently upon the different partition combinations of input-output variables. The prediction error will be reduced further by invoking the GFPE that combines multiple fuzzy predictors by an equal prediction error weighting method. Applications to both the Mackey-Glass chaotic time series and the nonstationary foreign currency exchange rate prediction problem are presented. The prediction accuracy of the proposed method is compared with that of other fuzzy and neural network predictors in terms of the root mean squared error (RMSE).
Daijin Kim 0001, Chulhyun Kim
IEEE Trans. Fuzzy Syst.1