Jae-Joon Han

dblp:74/7985 · DBLP profile ↗
← Back
31ranked-venue papers
0as first author
11since 2021 · last 2023
0000-0002-6505-6529ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 9 since 2021Artificial intelligence and machine learning · 16 · 10 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2023 Generative Multi-Label Correlation Learning
abstract
In real-world applications, a single instance could have more than one label. To solve this task, multi-label learning methods emerged in recent years. It is a more challenging problem for many reasons, such as complex label correlation, long-tail label distribution, and data shortage. In general, overcoming these challenges and bettering learning performance could be achieved by utilizing more training samples and including label correlations. However, these solutions are expensive and inflexible. Large-scale, well-labeled datasets are difficult to obtain, and building label correlation maps requires task-specific semantic information as prior knowledge. To address these limitations, we propose a general and compact Multi-Label Correlation Learning (MUCO) framework. MUCO explicitly and effectively learns the latent label correlations by updating a label correlation tensor, which provides highly accurate and interpretable prediction results. In addition, a multi-label generative strategy is deployed to handle the long-tail label distribution challenge. It borrows the visual clues from limited samples and synthesizes more diverse samples. All networks in our model are optimized simultaneously. Extensive experiments illustrate the effectiveness and efficiency of MUCO. Ablation studies further prove the effectiveness of all the modules.
Lichen Wang, Zhengming Ding, Kasey Lee, Seungju Han 0001, Jae-Joon Han, Changkyu Choi, Yun Fu 0001
ACM Trans. Knowl. Discov. Data5
2022 Self-Supervised Dense Consistency Regularization for Image-to-Image Translation
abstract
Unsupervised image-to-image translation has gained considerable attention due to recent impressive advances in generative adversarial networks (GANs). This paper presents a simple but effective regularization technique for improving GAN-based image-to-image translation. To generate images with realistic local semantics and structures, we propose an auxiliary self-supervision loss that enforces point-wise consistency of the overlapping region between a pair of patches cropped from a single real image during training the discriminator of a GAN. Our experiment shows that the proposed dense consistency regularization improves performance substantially on various image-to-image translation scenarios. It also leads to extra performance gains through the combination with instance-level regularization methods. Furthermore, we verify that the proposed model captures domain-specific characteristics more effectively with only a small fraction of training data.
Minsu Ko, Eun Ju Cha, Sungjoo Suh, Huijin Lee, Jae-Joon Han, Jinwoo Shin, Bohyung Han
CVPR5
2022 Towards Accurate Facial Landmark Detection via Cascaded Transformers
abstract
Accurate facial landmarks are essential prerequisites for many tasks related to human faces. In this paper, an accurate facial landmark detector is proposed based on cascaded transformers. We formulate facial landmark detection as a coordinate regression task such that the model can be trained end-to-end. With self-attention in transformers, our model can inherently exploit the structured relationships between landmarks, which would benefit landmark detection under challenging conditions such as large pose and occlusion. During cascaded refinement, our model is able to extract the most relevant image features around the target landmark for coordinate prediction, based on deformable attention mechanism, thus bringing more accurate alignment. In addition, we propose a novel decoder that refines image features and landmark positions simultaneously. With few parameter increasing, the detection performance improves further. Our model achieves new state-of-the-art performance on several standard facial landmark detection benchmarks, and shows good generalization ability in cross-dataset evaluation.
Hui Li 0031, Zidong Guo, Seon-Min Rhee, Seungju Han 0001, Jae-Joon Han
CVPR5
2022 Pushing the Performance Limit of Scene Text Recognizer without Human Annotation
abstract
Scene text recognition (STR) attracts much attention over the years because of its wide application. Most methods train STR model in a fully supervised manner which requires large amounts of labeled data. Although synthetic data contributes a lot to STR, it suffers from the real-to-synthetic domain gap the restricts model performance. In this work, we aim to boost STR models by leveraging both synthetic data and the numerous real unlabeled images, exempting human annotation cost thoroughly. A robust con-sistency regularization based semi-supervised framework is proposed for STR, which can effectively solve the instability issue due to domain inconsistency between synthetic and real images. A character-level consistency regularization is designed to mitigate the misalignment between characters in sequence recognition. Extensive experiments on standard text recognition benchmarks demonstrate the effectiveness of the proposed method. It can steadily improve existing STR models, and boost an STR model to achieve new state-of-the-art results. To our best knowledge, this is the first consistency regularization based framework that applies successfully to STR.
Caiyuan Zheng, Hui Li 0031, Seon-Min Rhee, Seungju Han 0001, Jae-Joon Han, Peng Wang 0015
CVPR5
2022 Slot-VPS: Object-centric Representation Learning for Video Panoptic Segmentation
abstract
Video Panoptic Segmentation (VPS) aims at assigning a class label to each pixel, uniquely segmenting and identifying all object instances consistently across all frames. Classic solutions usually decompose the VPS task into several subtasks and utilize multiple surrogates (e.g. boxes and masks, centers and offsets) to represent objects. However, this divide-and-conquer strategy requires complex post-processing in both spatial and temporal domains and is vulnerable to failures from surrogate tasks. In this paper, inspired by object-centric learning which learns compact and robust object representations, we present Slot- VPS, the first end-to-end framework for this task. We encode all panoptic entities in a video, including both foreground instances and background semantics, with a unified representation called panoptic slots. The coherent spatio-temporal object's information is retrieved and encoded into the panoptic slots by the proposed Video Panoptic Retriever, enabling to localize, segment, differentiate, and associate objects in a unified manner. Finally, the output panoptic slots can be directly converted into the class, mask, and object ID of panoptic objects in the video. We conduct extensive ablation studies and demonstrate the effectiveness of our approach on two benchmark datasets, Cityscapes- VP S (val and test sets) and VIPER (val set), achieving new state-of-the-art performance of 63.7, 63.3 and 56.2 VPQ, respectively.
Yi Zhou 0020, Hui Zhang 0093, Hana Lee, Shuyang Sun, Pingjun Li, Yangguang Zhu, ByungIn Yoo, Xiaojuan Qi 0001, Jae-Joon Han
CVPR9
2021 Order Regularization on Ordinal Loss for Head Pose, Age and Gaze Estimation
abstract
Ordinal loss is widely used in solving regression problems with deep learning technologies. Its basic idea is to convert regression to classification while preserving the natural order. However, the order constraint is enforced only by ordinal label implicitly, leading to the real output values not strictly in order. It causes the network to learn separable feature rather than discriminative feature, and possibly overfit on training set. In this paper, we propose order regularization on ordinal loss, which makes the outputs in order by explicitly constraining the ordinal classifiers in order. The proposed method contains two parts, i.e. similar-weights constraint, which reduces the ineffective space between classifiers, and differential-bias constraint, which enforces the decision planes in order and enhances the discrimination power of the classifiers. Experimental results show that our proposed method boosts the performance of original ordinal loss on various regression problems such as head pose, age, and gaze estimation, with significant error reduction of around 5%. Furthermore, our method outperforms the state of the art on all these tasks, with the performance gain of 14.4%, 2.2% and 6.5% on head pose, age and gaze estimation respectively.
Tianchu Guo, ByungIn Yoo, Youngjun Kwak, Jae-Joon Han
AAAI6
2021 Quality-Agnostic Image Recognition via Invertible Decoder
abstract
Despite the remarkable performance of deep models on image recognition tasks, they are known to be susceptible to common corruptions such as blur, noise, and low-resolution. Data augmentation is a conventional way to build a robust model by considering these common corruptions during the training. However, a naive data augmentation scheme may result in a non-specialized model for particular corruptions, as the model tends to learn the averaged distribution among corruptions. To mitigate the issue, we propose a new paradigm of training deep image recognition networks that produce clean-like features from any quality image via an invertible neural architecture. The proposed method consists of two stages. In the first stage, we train an invertible network with only clean images under the recognition objective. In the second stage, its inversion, i.e., the invertible decoder, is attached to a new recognition network and we train this encoder-decoder network using both clean and corrupted images by considering recognition and reconstruction objectives. Our two-stage scheme allows the network to produce clean-like and robust features from any quality images, by reconstructing their clean images via the invertible decoder. We demonstrate the effectiveness of our method on image classification and face recognition tasks.
Seungju Han 0001, Jiwon Baek, Seong-Jin Park, Jae-Joon Han, Jinwoo Shin
CVPR5
2021 Controllable Image Restoration for Under-Display Camera in Smartphones
abstract
Under-display camera (UDC) technology is essential for full-screen display in smartphones and is achieved by removing the concept of drilling holes on display. However, this causes inevitable image degradation in the form of spatially variant blur and noise because of the opaque display in front of the camera. To address spatially variant blur and noise in UDC images, we propose a novel controllable image restoration algorithm utilizing pixel-wise UDC-specific kernel representation and a noise estimator. The kernel representation is derived from an elaborate optical model that reflects the effect of both normal and oblique light incidence. Also, noise-adaptive learning is introduced to control noise levels, which can be utilized to provide optimal results depending on the user preferences. The experiments showed that the proposed method achieved superior quantitative performance as well as higher perceptual quality on both a real-world dataset and a monitor-based aligned dataset compared to conventional image restoration algorithms.
Kinam Kwon, Eunhee Kang, Su-Jin Lee, Hyong-Euk Lee, ByungIn Yoo, Jae-Joon Han
CVPR7
2021 RaScaNet: Learning Tiny Models by Raster-Scanning Images
abstract
Deploying deep convolutional neural networks on ultra-low power systems is challenging due to the extremely limited resources. Especially, the memory becomes a bottleneck as the systems put a hard limit on the size of on-chip memory. Because peak memory explosion in the lower layers is critical even in tiny models, the size of an input image should be reduced with sacrifice in accuracy. To overcome this drawback, we propose a novel Raster-Scanning Network, named RaScaNet, inspired by raster-scanning in image sensors. RaScaNet reads only a few rows of pixels at a time using a convolutional neural network and then sequentially learns the representation of the whole image using a recurrent neural network. The proposed method operates on an ultra-low power system without input size reduction; it requires 15.9–24.3× smaller peak memory and 5.3–12.9× smaller weight memory than the state-of-the-art tiny models. Moreover, RaScaNet fully exploits on-chip SRAM and cache memory of the system as the sum of the peak memory and the weight memory does not exceed 60 KB, improving the power efficiency of the system. In our experiments, we demonstrate the binary classification performance of RaScaNet on Visual Wake Words and Pascal VOC datasets.
Jaehyoung Yoo, Changyong Son, Sangil Jung, ByungIn Yoo, Changkyu Choi, Jae-Joon Han, Bohyung Han
CVPR7
2021 Large Scale Multi-Illuminant (LSMI) Dataset for Developing White Balance Algorithm under Mixed Illumination
abstract
We introduce a Large Scale Multi-Illuminant (LSMI) Dataset that contains 7,486 images, captured with three different cameras on more than 2,700 scenes with two or three illuminants. For each image in the dataset, the new dataset provides not only the pixel-wise ground truth illumination but also the chromaticity of each illuminant in the scene and the mixture ratio of illuminants per pixel. Images in our dataset are mostly captured with illuminants existing in the scene, and the ground truth illumination is computed by taking the difference between the images with different illumination combination. Therefore, our dataset captures natural composition in the real-world setting with wide field-of-view, providing more extensive dataset compared to existing datasets for multi-illumination white balance. As conventional single illuminant white balance algorithms cannot be directly applied, we also apply per-pixel DNN-based white balance algorithm and show its effectiveness against using patch-wise white balancing. We validate the benefits of our dataset through extensive analysis including a user-study, and expect the dataset to make meaningful contribution for future work in white balancing.
Dongyoung Kim, Jinwoo Kim 0007, Seonghyeon Nam, Yeonkyung Lee, Nahyup Kang, Hyong-Euk Lee, ByungIn Yoo, Jae-Joon Han, Seon Joo Kim
ICCV9
2021 Learning Generalized Intersection Over Union for Dense Pixelwise Prediction
abstract
Intersection over union (IoU) score, also named Jaccard Index, is one of the most fundamental evaluation methods in machine learning. The original IoU computation cannot provide non-zero gradients and thus cannot be directly optimized by nowadays deep learning methods. Several recent works generalized IoU for bounding box regression, but they are not straightforward to adapt for pixelwise prediction. In particular, the original IoU fails to provide effective gradients for the non-overlapping and location-deviation cases, which results in performance plateau. In this paper, we propose PixIoU, a generalized IoU for pixelwise prediction that is sensitive to the distance for non-overlapping cases and the locations in prediction. We provide proofs that PixIoU holds many nice properties as the original IoU. To optimize the PixIoU, we also propose a loss function that is proved to be submodular, hence we can apply the Lovász functions, the efficient surrogates for submodular functions for learning this loss. Experimental results show consistent performance improvements by learning PixIoU over the original IoU for several different pixelwise prediction tasks on Pascal VOC, VOT-2020 and Cityscapes.
Jiaqian Yu, Jingtao Xu, Qiang Wang 0023, ByungIn Yoo, Jae-Joon Han
ICML7
2020 DiscFace: Minimum Discrepancy Learning for Deep Face Recognition
Seungju Han 0001, Seong-Jin Park, Jiwon Baek, Jinwoo Shin, Jae-Joon Han, Changkyu Choi
ACCV (5)6
2020 Meta Variance Transfer: Learning to Augment from the Others
abstract
Humans have the ability to robustly recognize objects with various factors of variations such as nonrigid transformations, background noises, and changes in lighting conditions. However, training deep learning models generally require huge amount of data instances under diverse variations, to ensure its robustness. To alleviate the need of collecting large amount of data and better learn to generalize with scarce data instances, we propose a novel meta-learning method which learns to transfer factors of variations from one class to another, such that it can improve the classification performance on unseen examples. Transferred variations generate virtual samples that augment the feature space of the target class during training, simulating upcoming query samples with similar variations. By sharing the factors of variations across different classes, the model becomes more robust to variations in the unseen examples and tasks using small number of examples per class. We validate our model on multiple benchmark datasets for few-shot classification and face recognition, on which our model significantly improves the performance of the base model, outperforming relevant baselines.
Seong-Jin Park, Seungju Han 0001, Jiwon Baek, Juhwan Song, Haebeom Lee, Jae-Joon Han, Sung Ju Hwang
ICML7
2019 Learning to Quantize Deep Networks by Optimizing Quantization Intervals With Task Loss
abstract
Reducing bit-widths of activations and weights of deep networks makes it efficient to compute and store them in memory, which is crucial in their deployments to resource-limited devices, such as mobile phones. However, decreasing bit-widths with quantization generally yields drastically degraded accuracy. To tackle this problem, we propose to learn to quantize activations and weights via a trainable quantizer that transforms and discretizes them. Specifically, we parameterize the quantization intervals and obtain their optimal values by directly minimizing the task loss of the network. This quantization-interval-learning (QIL) allows the quantized networks to maintain the accuracy of the full-precision (32-bit) networks with bit-width as low as 4-bit and minimize the accuracy degeneration with further bit-width reduction (i.e., 3 and 2-bit). Moreover, our quantizer can be trained on a heterogeneous dataset, and thus can be used to quantize pretrained networks without access to their training data. We demonstrate the effectiveness of our trainable quantizer on ImageNet dataset with various network architectures such as ResNet-18, -34 and AlexNet, on which it outperforms existing methods to achieve the state-of-the-art accuracy.
Sangil Jung, Changyong Son, Seohyung Lee, JinWoo Son, Jae-Joon Han, Youngjun Kwak, Sung Ju Hwang, Changkyu Choi
CVPR5
2019 Generative Correlation Discovery Network for Multi-label Learning
abstract
The goal of Multi-label learning is to predict multiple labels of each single instance. This is a challenging problem since the training data is limited, long-tail label distribution, and complicated label correlations. Generally, more training samples and label correlation knowledge would benefit the learning performance. However, it is difficult to obtain large-scale well-labeled datasets, and building such a label correlation map requires sophisticated semantic knowledge. To this end, we propose an end-to-end Generative Correlation Discovery Network (GCDN) method for multi-label learning in this paper. GCDN captures the existing data distribution, and synthesizes diverse data to enlarge the diversity of the training features; meanwhile, it also learns the label correlations based on a specifically-designed, simple but effective correlation discovery network to automatically discover the label correlations and considerately improve the label prediction accuracy. Extensive experiments on several benchmarks are provided to demonstrate the effectiveness, efficiency, and high accuracy of our approach.
Lichen Wang, Zhengming Ding, Seungju Han 0001, Jae-Joon Han, Changkyu Choi, Yun Fu 0001
ICDM4
2019 Robust Discriminative Metric Learning for Image Representation
abstract
Metric learning has attracted significant attention in the past decades, because of its appealing advances in various real-world tasks, e.g., person re-identification and face recognition. Traditional supervised metric learning attempts to seek a discriminative metric, which could minimize the pairwise distance of within-class data samples, while maximizing the pairwise distance of data samples from various classes. However, it is still a challenge to build a robust and discriminative metric, especially for corrupted data in the real-world application. In this paper, we propose a Robust Discriminative Metric Learning algorithm through fast low-rank representation and denoising strategy. To be specific, the metric learning problem is guided by a discriminative regularization by incorporating the pair-wise or class-wise information. Moreover, the low-rank basis learning is jointly optimized with the metric to better uncover the global data structure and remove noise. Furthermore, the fast low-rank representation is implemented to mitigate the computational burden and ensure the scalability on large-scale datasets. Finally, we evaluate our learned metric on several challenging tasks, e.g., face recognition/verification, object recognition, image clustering, and person re-identification. The experimental results verify the effectiveness of our proposed algorithm in comparison to many metric learning algorithms, even deep learning ones.
Zhengming Ding, Ming Shao, Wonjun Hwang, Sungjoo Suh, Jae-Joon Han, Changkyu Choi, Yun Fu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2018 Residual Encoder Decoder Network and Adaptive Prior for Face Parsing
abstract
Face Parsing assigns every pixel in a facial image with a semantic label, which could be applied in various applications including face recognition, facial beautification, affective computing and animation. While lots of progress have been made in this field, current state-of-the-art methods still fail to extract real effective feature and restore accurate score map, especially for those facial parts which have large variations of deformation and fairly similar appearance, e.g. mouth, eyes and thin eyebrows. In this paper, we propose a novel pixel-wise face parsing method called Residual Encoder Decoder Network (RED-Net), which combines a feature-rich encoder-decoder framework with adaptive prior mechanism. Our encoder-decoder framework extracts feature with ResNet and decodes the feature by elaborately fusing the residual architectures in to deconvolution. This framework learns more effective feature comparing to that learnt by decoding with interpolation or classic deconvolution operations. To overcome the appearance ambiguity between facial parts, an adaptive prior mechanism is proposed in term of the decoder prediction confidence, allowing refining the final result. The experimental results on two public datasets demonstrate that our method outperforms the state-of-the-arts significantly, achieving improvements of F-measure from 0.854 to 0.905 on Helen dataset, and pixel accuracy from 95.12% to 97.59% on the LFW dataset. In particular, convincing qualitative examples show that our method parses eye, eyebrow, and lip regins more accurately.
Tianchu Guo, Youngsung Kim, Deheng Qian, ByungIn Yoo, Jingtao Xu, Dongqing Zou, Jae-Joon Han, Changkyu Choi
AAAI8
2017 Directional coherence-based spatiotemporal descriptor for object detection in static and dynamic scenes
Wonjun Kim 0001, Jae-Joon Han
Mach. Vis. Appl.2
2016 A fast multi-view face detector for mobile phone
abstract
A new face detector with very high accuracy and real-time speed for mobile phone is introduced. The method achieves the fastest speed and the highest accuracy compared with other similar methods. A series of ideas are proposed in order to accelerate detection speed of the traditional Adaboost detector. First of all, the threshold for weak classifier is learned based on a new multiple instance pruning method regarding not only positive samples but also negative samples, by which, weak classifier is able to reject background more efficiently. Then, a coarse-to-fine scan is applied. Coarse scan is used to find possible face location, and the fine scan refines the face location and rejects false alarms. We further improve the speed of multi-scale face detection by introducing two different template sizes for detector training. By which, the smaller faces can be rapidly detected and the high performance is kept for larger faces. The proposed method is evaluated on public dataset FDDB, the result shows competitive performance against all Adaboost based methods. The method has been implemented on mobile phone and the speed is superior to all competitors.
Wonjun Hwang, Jae-Joon Han, Changkyu Choi, Haitao Wang 0006
ICIP6
2015 Robust pose normalization for face recognition under varying views
abstract
Unconstrained face recognition under varying views is one of the most challenging tasks, since the difference in appearances caused by poses may be even larger than that due to identity. In this paper, we exploit and analyze a novel pose normalization scheme for facial images under varying views via robust 3D shape reconstruction from single, unconstrained photos in the wild. Specifically, to address the problem of ambiguous 2D-to-3D landmark correspondence and imperfect landmark detector, for each input 2D face, the 3D shape is suggested to be learned by iteratively refining the 3D landmarks and the weighting coefficients of each landmark. Experimental results on both LFW and a large-scale self-collected face databases demonstrate that the proposed approach performs better than the existing representative technologies.
Xuetao Feng, Lujin Gong, Wonjun Hwang, Jae-Joon Han
ICIP6
2015 Combining nonuniform sampling, hybrid super vector, and random forest with discriminative decision trees for action recognition
abstract
Trajectory-based features have become popular for action recognition and achieve the state-of-the-art results on a variety of datasets. In this paper, we propose a novel framework to improve the performance of action recognition. Specifically, we first apply the nonuniform sampling method to efficiently select features for given actions. The proposed hybrid super vector, namely fisher vector (FV) combined with vector of locally aggregated descriptors (VLAD), is then employed to encode sampled trajectories. A random forest with discriminative decision trees, where every tree node is a discriminative classifier, is finally applied to predict action labels. We have achieved 88.2% in average accuracy on the UCF101 dataset, which outperforms the best results that have been reported in the literature.
Kuanhong Xu, Ya Lu, Xuetao Feng, Jae-Joon Han
ICIP6
2015 Face Liveness Detection From a Single Image via Diffusion Speed Model
abstract
Spoofing using photographs or videos is one of the most common methods of attacking face recognition and verification systems. In this paper, we propose a real-time and nonintrusive method based on the diffusion speed of a single image to address this problem. In particular, inspired by the observation that the difference in surface properties between a live face and a fake one is efficiently revealed in the diffusion speed, we exploit antispoofing features by utilizing the total variation flow scheme. More specifically, we propose defining the local patterns of the diffusion speed, the so-called local speed patterns, as our features, which are input into the linear SVM classifier to determine whether the given face is fake or not. One important advantage of the proposed method is that, in contrast to previous approaches, it accurately identifies diverse malicious attacks regardless of the medium of the image, e.g., paper or screen. Moreover, the proposed method does not require any specific user action. Experimental results on various data sets show that the proposed method is effective for face liveness detection as compared with previous approaches proposed in studies in the literature.
Sungjoo Suh, Jae-Joon Han
IEEE Trans. Image Process.3
2014 HDO: A novel local image descriptor
abstract
This paper presents a simple, yet powerful local image descriptor, called the histograms of dominant orientations (HDO). The HDO consists of two components, namely the dominant orientation and its coherence, which represents how intensively gradients in the local region are distributed along the dominant orientation. For a given image patch, we incorporate these two components into a 1-D histogram and define it as our HDO descriptor. Compared to previous approaches suffering from the presence of clutters and significant distortions, our HDO descriptor has a great ability to preserve the underlying image structure, and it can thus be successfully applied to various applications (e.g., object detection). The proposed method has been extensively tested on several challenging data sets and results show that our HDO descriptor is effective for object detection in images.
ByungIn Yoo, Jae-Joon Han
ICIP3
2014 Hierarchical gaze estimation based on adaptive feature learning
abstract
Existing appearance-based gaze estimation methods suffer from tedious calibration and appearance variation caused by head movement. In this paper, to handle this problem, we propose a novel appearance-based gaze estimation method by introducing supervised adaptive feature extraction and hierarchical mapping model. Firstly, an adaptive feature learning method is proposed to extract topology-preserving (TOP) feature individually. Then hierarchical mapping method is proposed to localize gaze position based on coarse-to-fine strategy. Appearance synthesis approach is used to increase the refer sample density. Experiments show that under the condition of sparse calibration, proposed method has better performance in accuracy than existing methods under fixed head pose without chinrest. Moreover, our method can be easily extended for head pose-varying gaze estimation.
Kang Xue, Dongkyung Nam, Jae-Joon Han, Haitao Wang 0006
ICIP4
2014 Randomized decision bush: Combining global shape parameters and local scalable descriptors for human body parts recognition
abstract
This paper presents a novel method which combines global shape parameters and scalable local descriptors for accurate body parts recognition from a single depth image in real-time. Human poses are of extremely large variation in aspects of visual shapes, because human can take poses from daily activities to gymnastic actions. In order to cover wide-range of the human poses, the proposed algorithm employs a unified structure which combines pose clustering and body parts classification. We name the proposed method Randomized Decision Bush (RDB). Specifically, global shape parameters which can discriminate coarse level shapes are utilized for pose clustering while scalable local shape descriptors are employed for accurate classification. RDB splits the various human poses into multiple clusters which contain similar shapes of the poses. As a result, it provides robust clustering which enables fine level classification within the cluster. The experimental results show improvements on recognizing body parts due to the pose clustering and classification with scalable local descriptors. Additionally, we significantly reduce the complexity of training a large number of human shapes.
ByungIn Yoo, Jae-Joon Han, Changkyu Choi, Du-Sik Park, Junmo Kim 0002
ICIP3
2014 Video Saliency Detection Using Contrast of Spatiotemporal Directional Coherence
abstract
Saliency detection in video sequences has attracted great attention in recent years due to its promising contributions for various computer vision applications. However, most existing methods often fail to correctly find salient regions in complex scenes due to ambiguities between salient motions and irrelevant ones generated from the background. In this letter, we present a simple and powerful framework for video saliency detection based on the contrast of the spatiotemporal directional coherence. Our approach is designed in a general way so that it can be applied to videos taken under various environments including dynamic illuminations and textures in the background, which still bother previous saliency detection methods. Based on various challenging datasets, we compare ours with several competitive approaches proposed in literature and results demonstrate that the proposed method is effective and robust for detecting salient regions in diverse video sequences.
Jae-Joon Han
IEEE Signal Process. Lett.2
2014 SVD Face: Illumination-Invariant Face Representation
abstract
In this letter, we propose a novel method to extract illumination-invariant features for face recognition and verification under varying illuminations. Inspired by the fact that normalized coefficients of the singular value decomposition (SVD) are insensitive to different illumination conditions, we exploit a simple, yet powerful scheme for describing underlying structures of faces, so-called SVD face. In contrast to previous approaches still suffering from the loss of details, our SVD face greatly preserves textures of the original image based on the relaxation of SVD coefficients. Theoretical analysis shows that our SVD face is an illumination-invariant measure and has an ability to discover meaningful components (e.g., eyes, mouth, etc.) of face images while suppressing the effect of various illuminations. Experimental results on both Yale B and our illuminated face (IF) datasets demonstrate that the SVD face is effective for face recognition and verification compared to previous approaches proposed in literature.
Sungjoo Suh, Wonjun Hwang, Jae-Joon Han
IEEE Signal Process. Lett.4
2013 Interactive manipulation and visualization of a deformable 3D organ model for medical diagnostic support
abstract
In this paper, an interactive medical image visualization system to support medical therapy has been introduced, where 3D organ model with the corresponding medical image is visualized interactively for diagnosis and surgical planning. To show effectiveness of the proposed system, 3D liver model generated from CT data has been utilized in consideration of its deformable characteristics by respiration as well as appearance. In addition, a hand gesture interface is applied on the graphical user interface for providing more natural and intuitive interactivity.
Hyong-Euk Lee, Nahyup Kang, Jae-Joon Han, James D. K. Kim, Chang-Yeong Kim
CCNC3
2013 Connecting users to virtual worlds within MPEG-V standardization
Seungju Han 0001, Jae-Joon Han, James D. K. Kim, Chang-Yeong Kim
Signal Process. Image Commun.2
2013 Virtual world control system using sensed information and adaptation engine
Sang-Kyun Kim, Yong Soo Joo, Minho Shin, Seungju Han 0001, Jae-Joon Han
Signal Process. Image Commun.5
2010 Controlling virtual world by the real world devices with an MPEG-V framework
abstract
The recent online networked virtual worlds such as SecondLife, World of Warcraft and Lineage have been increasingly popular. A life-scale virtual world presentation and the intuitive interaction between the users and the virtual worlds would provide more natural and immersive experience for users. The emergence of novel interaction technologies such as sensing the facial expression and the motion of the users and the real world environments could be used to provide a strong connection between them. For the wide acceptance and use of the virtual world, a various type of novel interaction devices should have a unified interaction formats between the real world and the virtual world and interoperability among virtual worlds. Thus, MPEG-V Media Context and Control (ISO/IEC 23005) standardizes such connecting information. The paper provides an overview and its usage example of MPEG-V from the real world to the virtual world (R2V) on interfaces for controlling avatars and virtual objects in the virtual world by the real world devices. In particular, we investigate how the MPEG-V framework can be applied for the facial animation of an avatar in various types of virtual worlds.
Seungju Han 0001, Jae-Joon Han, Youngkyoo Hwang, Jung-Bae Kim, Won-Chul Bang, James D. K. Kim, Chang-Yeong Kim
MMSP2