VLDB 2026 Research / reviewers in the wild / expert
Hyeran Byun
dblp:77/5360
· DBLP profile ↗
95ranked-venue papers
4as first author
32since 2021 · last 2025
0000-0002-3082-3214ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 67 · 2 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 52 · 19 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploiting Domain Properties in Language-Driven Domain Generalization for Semantic SegmentationabstractRecent domain generalized semantic segmentation (DGSS) studies have achieved notable improvements by distilling semantic knowledge from Vision-Language Models (VLMs). However, they overlook the semantic misalignment between visual and textual contexts, which arises due to the rigidity of a fixed context prompt learned on a single source domain. To this end, we present a novel domain generalization framework for semantic segmentation, namely Domain-aware Prompt-driven Masked Transformer (DPMFormer). Firstly, we introduce domain-aware prompt learning to facilitate semantic alignment between visual and textual cues. To capture various domain-specific properties with a single source dataset, we propose domain-aware contrastive learning along with the texture perturbation that diversifies the observable domains. Lastly, to establish a framework resilient against diverse environmental changes, we have proposed the domain-robust consistency learning which guides the model to minimize discrepancies of prediction from original and the augmented images. Through experiments and analyses, we demonstrate the superiority of the proposed framework, which establishes a new state-of-the-art on various DGSS benchmarks. The code is available at https://github.com/jone1222/DPMFormer. Seogkyu Jeon, Kibeom Hong, Hyeran Byun |
ICCV | 3 |
| 2024 | Fair-VPT: Fair Visual Prompt Tuning for Image ClassificationabstractDespite the remarkable success of Vision Transformers (ViT) across diverse fields in computer vision, they have a clear drawback of expensive adaption cost for downstream tasks due to the increased scale. To address this, Visual Prompt Tuning (VPT) incorporates learnable parameters in the input space of ViT. While freezing the ViT backbone and tuning only the prompts, it exhibits superior performances to full fine-tuning. However, despite the outstanding advantage, we point out that VPT may lead to serious unfairness in downstream classification. Initially, we investigate the causes of unfairness in VPT, identifying the biasedly pre-trained ViT as a principal factor. Motivated by this observation, we propose a Fair Visual Prompt Tuning (Fair-VPT) which removes biased information in the pre-trained ViT while adapting it to downstream classification tasks. To this end, we categorize prompts into “cleaner prompts” and “target prompts”. Based on this, we encode the class token in two different ways by either masking or not masking the target prompts in the self-attention process. These encoded tokens are trained with distinct objective functions, resulting in the inclusion of different information in the target and cleaner prompts. Moreover, we introduce a disentanglement loss based on contrastive learning to further decorrelate them. In experiments across diverse benchmarks, the proposed method demonstrates the most superior performance in terms of balanced classification accuracy and fairness. Hyeran Byun |
CVPR | 2 |
| 2024 | BAM-DETR: Boundary-Aligned Moment Detection Transformer for Temporal Sentence Grounding in Videos
Pilhyeon Lee, Hyeran Byun |
ECCV (2) | 2 |
| 2024 | Small Objects Matters in Weakly-supervised Semantic SegmentationabstractWeakly-supervised semantic segmentation (WSSS) performs pixel-wise classification given only image-level labels for training. Despite the difficulty of this task, the research community has achieved promising results over the last five years. Still, current WSSS literature misses the detailed sense of how well the methods perform on different sizes of objects. Thus we propose a novel evaluation metric to provide a comprehensive assessment across different object sizes and collect a size-balanced evaluation set to complement PASCAL VOC. With these two gadgets, we reveal that the existing WSSS methods struggle in capturing small objects. Furthermore, we propose a size-balanced cross-entropy loss coupled with a proper training strategy. It generally improves existing WSSS methods as validated upon ten baselines on three different datasets. Cheolhyun Mun, Sanghuk Lee, Youngjung Uh, Junsuk Choe, Hyeran Byun |
WACV | 5 |
| 2024 | Discovering an inference recipe for weakly-supervised object localization
Sanghuk Lee, Cheolhyun Mun, Youngjung Uh, Junsuk Choe, Hyeran Byun |
Pattern Recognit. | 5 |
| 2023 | Decomposed Cross-Modal Distillation for RGB-based Temporal Action DetectionabstractTemporal action detection aims to predict the time intervals and the classes of action instances in the video. Despite the promising performance, existing two-stream models exhibit slow inference speed due to their reliance on computationally expensive optical flow. In this paper, we introduce a decomposed cross-modal distillation framework to build a strong RGB-based detector by transferring knowledge of the motion modality. Specifically, instead of direct distillation, we propose to separately learn RGB and motion representations, which are in turn combined to perform action localization. The dual-branch design and the asymmetric training objectives enable effective motion knowledge transfer while preserving RGB information intact. In addition, we introduce a local attentive fusion to better exploit the multimodal complementarity. It is designed to preserve the local discriminability of the features that is important for action localization. Extensive experiments on the benchmarks verify the effectiveness of the proposed method in enhancing RGB-based action detectors. Notably, our framework is agnostic to backbones and detection heads, bringing consistent gains across different model combinations. Pilhyeon Lee, Taeoh Kim, Minho Shim, Dongyoon Wee, Hyeran Byun |
CVPR | 5 |
| 2023 | AesPA-Net: Aesthetic Pattern-Aware Style Transfer NetworksabstractTo deliver the artistic expression of the target style, recent studies exploit the attention mechanism owing to its ability to map the local patches of the style image to the corresponding patches of the content image. However, because of the low semantic correspondence between arbitrary content and artworks, the attention module repeatedly abuses specific local patches from the style image, resulting in disharmonious and evident repetitive artifacts. To overcome this limitation and accomplish impeccable artistic style transfer, we focus on enhancing the attention mechanism and capturing the rhythm of patterns that organize the style. In this paper, we introduce a novel metric, namely pattern repeatability, that quantifies the repetition of patterns in the style image. Based on the pattern repeatability, we propose Aesthetic Pattern-Aware style transfer Networks (AesPA-Net) that discover the sweet spot of local and global style expressions. In addition, we propose a novel self-supervisory task to encourage the attention mechanism to learn precise and meaningful semantic correspondence. Lastly, we introduce the patch-wise style loss to transfer the elaborate rhythm of local patterns. Through qualitative and quantitative evaluations, we verify the reliability of the proposed pattern repeatability that aligns with human perception, and demonstrate the superiority of the proposed framework. All codes and pre-trained weights are available at Kibeom-Hong/AesPA-Net. Kibeom Hong, Seogkyu Jeon, Junsoo Lee 0002, Namhyuk Ahn, Kunhee Kim, Pilhyeon Lee, Youngjung Uh, Hyeran Byun |
ICCV | 9 |
| 2023 | Improving Diversity in Zero-Shot GAN Adaptation with Semantic VariationsabstractTraining deep generative models usually requires a large amount of data. To alleviate the data collection cost, the task of zero-shot GAN adaptation aims to reuse well-trained generators to synthesize images of an unseen target domain without any further training samples. Due to the data absence, the textual description of the target domain and the vision-language models, e.g., CLIP, are utilized to effectively guide the generator. However, with only a single representative text feature instead of real images, the synthesized images gradually lose diversity as the model is optimized, which is also known as mode collapse. To tackle the problem, we propose a novel method to find semantic variations of the target text in the CLIP space. Specifically, we explore diverse semantic variations based on the informative text feature of the target domain while regularizing the uncontrolled deviation of the semantic information. With the obtained variations, we design a novel directional moment loss that matches the first and second moments of image and text direction distributions. Moreover, we introduce elastic weight consolidation and a relation consistency loss to effectively preserve valuable content information from the source domain, e.g., appearances. Through extensive experiments, we demonstrate the efficacy of the proposed methods in ensuring sample diversity in various scenarios of zero-shot GAN adaptation. We also conduct ablation studies to validate the effect of each proposed component. Notably, our model achieves a new state-of-the-art on zero-shot GAN adaptation in terms of both diversity and quality. Seogkyu Jeon, Bei Liu 0001, Pilhyeon Lee, Kibeom Hong, Jianlong Fu, Hyeran Byun |
ICCV | 6 |
| 2023 | BallGAN: 3D-aware Image Synthesis with a Spherical Backgroundabstract3D-aware GANs aim to synthesize realistic 3D scenes that can be rendered in arbitrary camera viewpoints, generating high-quality images with well-defined geometry. As 3D content creation becomes more popular, the ability to generate foreground objects separately from the background has become a crucial property. Existing methods have been developed regarding overall image quality, but they can not generate foreground objects only and often show degraded 3D geometry. In this work, we propose to represent the background as a spherical surface for multiple reasons inspired by computer graphics. Our method naturally provides foreground-only 3D synthesis facilitating easier 3D content creation. Furthermore, it improves the foreground geometry of 3D-aware GANs and the training stability on datasets with complex backgrounds. Project page: https://minjung-s.github.io/ballgan/ Minjung Shin, Yunji Seo, Jeongmin Bae 0001, Young Sun Choi, Hyunsu Kim, Hyeran Byun, Youngjung Uh |
ICCV | 6 |
| 2023 | Task-aware network: Mitigation of task-aware and task-free performance gap in online continual learning
Yongwon Hong, Hyeran Byun |
Neurocomputing | 3 |
| 2023 | Fair classification by loss balancing via fairness-aware batch sampling
Do-Hyung Kim 0004, Sunhee Hwang, Hyeran Byun |
Neurocomputing | 4 |
| 2023 | Expert-guided contrastive learning for video-text retrieval
Jewook Lee, Pilhyeon Lee, Hyeran Byun |
Neurocomputing | 4 |
| 2022 | Fair Contrastive Learning for Facial Attribute ClassificationabstractLearning visual representation of high quality is essential for image classification. Recently, a series of contrastive representation learning methods have achieved preeminent success. Particularly, SupCon [18] outperformed the dominant methods based on cross-entropy loss in representation learning. However, we notice that there could be potential ethical risks in supervised contrastive learning. In this paper, we for the first time analyze unfairness caused by supervised contrastive learning and propose a new Fair Supervised Contrastive Loss (FSCL) for fair visual representation learning. Inheriting the philosophy of supervised contrastive learning, it encourages representation of the same class to be closer to each other than that of different classes, while ensuring fairness by penalizing the inclusion of sensitive attribute information in representation. In addition, we introduce a group-wise normalization to diminish the disparities of intra-group compactness and inter-class separability between demographic groups that arouse unfair classification. Through extensive experiments on CelebA and UTK Face, we validate that the proposed method significantly outperforms SupCon and existing state-of-the-art methods in terms of the trade-off between top-l accuracy and fairness. Moreover, our method is robust to the intensity of data bias and effectively works in incomplete supervised settings. Our code is available at https://github.com/sungho-CoolG/FSCL. Jewook Lee, Pilhyeon Lee, Sunhee Hwang, Do-Hyung Kim 0004, Hyeran Byun |
CVPR | 6 |
| 2022 | Exploiting domain transferability for collaborative inter-level domain adaptive object detection
Mirae Do, Seogkyu Jeon, Pilhyeon Lee, Kibeom Hong, Yu-Seung Ma, Hyeran Byun |
Expert Syst. Appl. | 6 |
| 2022 | WeatherGAN: Unsupervised multi-weather image-to-image translation via single content-preserving UResNet generator
Sunhee Hwang, Seogkyu Jeon, Yu-Seung Ma, Hyeran Byun |
Multim. Tools Appl. | 4 |
| 2022 | Return of the normal distribution: Flexible deep continual learning with variational auto-encoders
Yongwon Hong, Martin Mundt, Yungjung Uh, Hyeran Byun |
Neural Networks | 5 |
| 2022 | Exploiting shape cues for weakly supervised semantic segmentation
Sungpil Kho, Pilhyeon Lee, Wonyoung Lee 0004, Minsong Ki, Hyeran Byun |
Pattern Recognit. | 5 |
| 2022 | Discriminative deep attributes for generalized zero-shot learning
Hoseong Kim, Jewook Lee, Hyeran Byun |
Pattern Recognit. | 3 |
| 2021 | Weakly-supervised Temporal Action Localization by Uncertainty ModelingabstractWeakly-supervised temporal action localization aims to learn detecting temporal intervals of action classes with only video-level labels. To this end, it is crucial to separate frames of action classes from the background frames (i.e., frames not belonging to any action classes). In this paper, we present a new perspective on background frames where they are modeled as out-of-distribution samples regarding their inconsistency. Then, background frames can be detected by estimating the probability of each frame being out-of-distribution, known as uncertainty, but it is infeasible to directly learn uncertainty without frame-level labels. To realize the uncertainty learning in the weakly-supervised setting, we leverage the multiple instance learning formulation. Moreover, we further introduce a background entropy loss to better discriminate background frames by encouraging their in-distribution (action) probabilities to be uniformly distributed over all action classes. Experimental results show that our uncertainty modeling is effective at alleviating the interference of background frames and brings a large performance gain without bells and whistles. We demonstrate that our model significantly outperforms state-of-the-art methods on the benchmarks, THUMOS'14 and ActivityNet (1.2 & 1.3). Our code is available at https://github.com/Pilhyeon/WTAL-Uncertainty-Modeling. Pilhyeon Lee, Jinglu Wang, Yan Lu 0001, Hyeran Byun |
AAAI | 4 |
| 2021 | Learning Disentangled Representation for Fair Facial Attribute Classification via Fairness-aware Information AlignmentabstractAlthough AI systems archive a great success in various societal fields, there still exists a challengeable issue of outputting discriminatory results with respect to protected attributes (e.g., gender and age). The popular approach to solving the issue is to remove protected attribute information in the decision process. However, this approach has a limitation that beneficial information for target tasks may also be eliminated. To overcome the limitation, we propose Fairness-aware Disentangling Variational Auto-Encoder (FD-VAE) that disentangles data representation into three subspaces: 1) Target Attribute Latent (TAL), 2) Protected Attribute Latent (PAL), 3) Mutual Attribute Latent (MAL). On top of that, we propose a decorrelation loss that aligns the overall information into each subspace, instead of removing the protected attribute information. After learning the representation, we re-encode MAL to include only target information and combine it with TAL to perform downstream tasks. In our experiments on CelebA and UTK Face datasets, we show that the proposed method mitigates unfairness in facial attribute classification tasks with respect to gender and age. Ours outperforms previous methods by large margins on two standard fairness metrics, equal opportunity and equalized odds. Sunhee Hwang, Do-Hyung Kim 0004, Hyeran Byun |
AAAI | 4 |
| 2021 | Foreground Mining via Contrastive Guidance for Weakly Supervised Object Localization
Wonyoung Lee 0004, Minsong Ki, Cheolhyun Mun, Sungpil Kho, Hyeran Byun |
BMVC | 5 |
| 2021 | Mitigating Inter-Subject Brain Signal Variability FOR EEG-Based Driver Fatigue State ClassificationabstractWith great research advances on Brain-Computer-Interface (BCI) systems, Electroencephalography (EEG) based driver fatigue state classification models have shown its effectiveness. However, EEG signals contain large differences between individuals, making it hard to build a unified model among individuals. In this paper, we propose a subject- independent EEG-based driver fatigue state (i.e., awake, tired, and drowsy) classification model that mitigates a performance gap between subjects. To this end, we exploit an adversarial training strategy to make our classification model misclassify the subject labels. Besides, we propose an Intersubject Feature Distance Minimization (IFDM) method that minimizes the Wasserstein distance between two different subject groups of the same class to reduce the individual performance discrepancy. Our method is also designed to enable training even if the subject labels are not sufficiently included in the EEG dataset. To demonstrate the ability of the proposed method, we conduct a drowsiness classification task on a publicly available SEED-VIG dataset. The experimental results show our model achieves the highest accuracy and the lowest individual performance variability. Sunhee Hwang, Do-Hyung Kim 0004, Jewook Lee, Hyeran Byun |
ICASSP | 5 |
| 2021 | Continuous Face Aging Generative Adversarial NetworksabstractFace aging is the task aiming to translate the faces in input images to designated ages. To simplify the problem, previous methods have limited themselves only able to produce discrete age groups, each of which consists of ten years. Consequently, the exact ages of the translated results are unknown and it is unable to obtain the faces of different ages within groups. To this end, we propose the continuous face aging generative adversarial networks (CFA-GAN). Specifically, to make the continuous aging feasible, we propose to decompose image features into two orthogonal features: the identity and the age basis features. Moreover, we introduce the novel loss function for identity preservation which maximizes the cosine similarity between the original and the generated identity basis features. With the qualitative and quantitative evaluations on MORPH, we demonstrate the realistic and continuous aging ability of our model, validating its superiority against existing models. To the best of our knowledge, this work is the first attempt to handle continuous target ages. Seogkyu Jeon, Pilhyeon Lee, Kibeom Hong, Hyeran Byun |
ICASSP | 4 |
| 2021 | Domain-Aware Universal Style TransferabstractStyle transfer aims to reproduce content images with the styles from reference images. Existing universal style transfer methods successfully deliver arbitrary styles to original images either in an artistic or a photo-realistic way. However, the range of "arbitrary style" defined by existing works is bounded in the particular domain due to their structural limitation. Specifically, the degrees of content preservation and stylization are established according to a predefined target domain. As a result, both photo-realistic and artistic models have difficulty in performing the desired style transfer for the other domain. To overcome this limitation, we propose a unified architecture, Domain-aware Style Transfer Networks (DSTN) that transfer not only the style but also the property of domain (i.e., domainness) from a given reference image. To this end, we design a novel domainness indicator that captures the domainness value from the texture and structural features of reference images. Moreover, we introduce a unified framework with domainaware skip connection to adaptively transfer the stroke and palette to the input contents guided by the domainness indicator. Our extensive experiments validate that our model produces better qualitative results and outperforms previous methods in terms of proxy metrics on both artistic and photo-realistic stylizations. All codes and pre-trained weights are available at Kibeom-Hong/Domain-Aware-Style-Transfer. Kibeom Hong, Seogkyu Jeon, Huan Yang 0005, Jianlong Fu, Hyeran Byun |
ICCV | 5 |
| 2021 | Contrastive Attention Maps for Self-supervised Co-localizationabstractThe goal of unsupervised co-localization is to locate the object in a scene under the assumptions that 1) the dataset consists of only one superclass, e.g., birds, and 2) there are no human-annotated labels in the dataset. The most recent method achieves impressive co-localization performance by employing self-supervised representation learning approaches such as predicting rotation. In this paper, we introduce a new contrastive objective directly on the attention maps to enhance co-localization performance. Our contrastive loss function exploits rich information of location, which induces the model to activate the extent of the object effectively. In addition, we propose a pixel-wise attention pooling that selectively aggregates the feature map regarding their magnitudes across channels. Our methods are simple and shown effective by extensive qualitative and quantitative evaluation, achieving state-of-the-art co-localization performances by large margins on four datasets: CUB-200-2011, Stanford Cars, FGVC-Aircraft, and Stanford Dogs. Our code will be publicly available online for the research community. Minsong Ki, Youngjung Uh, Junsuk Choe, Hyeran Byun |
ICCV | 4 |
| 2021 | Learning Action Completeness from Points for Weakly-supervised Temporal Action LocalizationabstractWe tackle the problem of localizing temporal intervals of actions with only a single frame label for each action instance for training. Owing to label sparsity, existing work fails to learn action completeness, resulting in fragmentary action predictions. In this paper, we propose a novel framework, where dense pseudo-labels are generated to provide completeness guidance for the model. Concretely, we first select pseudo background points to supplement point-level action labels. Then, by taking the points as seeds, we search for the optimal sequence that is likely to contain complete action instances while agreeing with the seeds. To learn completeness from the obtained sequence, we introduce two novel losses that contrast action instances with background ones in terms of action score and feature similarity, respectively. Experimental results demonstrate that our completeness guidance indeed helps the model to locate complete action instances, leading to large performance gains especially under high IoU thresholds. Moreover, we demonstrate the superiority of our method over existing state-of-the-art methods on four benchmarks: THUMOS’14, GTEA, BEOID, and ActivityNet. Notably, our method even performs comparably to recent fully-supervised methods, at the 6× cheaper annotation cost. Our code is available at https://github.com/Pilhyeon. Pilhyeon Lee, Hyeran Byun |
ICCV | 2 |
| 2021 | Feature Stylization and Domain-aware Contrastive Learning for Domain GeneralizationabstractDomain generalization aims to enhance the model robustness against domain shift without accessing the target domain. Since the available source domains for training are limited, recent approaches focus on generating samples of novel domains. Nevertheless, they either struggle with the optimization problem when synthesizing abundant domains or cause the distortion of class semantics. To these ends, we propose a novel domain generalization framework where feature statistics are utilized for stylizing original features to ones with novel domain properties. To preserve class information during stylization, we first decompose features into high and low frequency components. Afterward, we stylize the low frequency components with the novel domain styles sampled from the manipulated statistics, while preserving the shape cues in high frequency ones. As the final step, we re-merge both the components to synthesize novel domain features. To enhance domain robustness, we utilize the stylized features to maintain the model consistency in terms of features as well as outputs. We achieve the feature consistency with the proposed domain-aware supervised contrastive loss, which ensures domain invariance while increasing class discriminability. Experimental results demonstrate the effectiveness of the proposed feature stylization and the domain-aware contrastive loss. Through quantitative comparisons, we verify the lead of our method upon existing state-of-the-art methods on two benchmarks, PACS and Office-Home. Seogkyu Jeon, Kibeom Hong, Pilhyeon Lee, Jewook Lee, Hyeran Byun |
ACM Multimedia | 5 |
| 2021 | ArrowGAN : Learning to generate videos by learning Arrow of Time
Kibeom Hong, Youngjung Uh, Hyeran Byun |
Neurocomputing | 3 |
| 2021 | Contrastive and consistent feature learning for weakly supervised object localization and semantic segmentation
Minsong Ki, Youngjung Uh, Wonyoung Lee 0004, Hyeran Byun |
Neurocomputing | 4 |
| 2021 | The feature generator of hard negative samples for fine-grained image recognition
Taehung Kim, Kibeom Hong, Hyeran Byun |
Neurocomputing | 3 |
| 2021 | Zero-shot learning with self-supervision by shuffling semantic embeddings
Hoseong Kim, Jewook Lee, Hyeran Byun |
Neurocomputing | 3 |
| 2021 | Sequence feature generation with temporal unrolling network for zero-shot action recognition
Jewook Lee, Hoseong Kim, Hyeran Byun |
Neurocomputing | 3 |
| 2020 | Background Suppression Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization is a very challenging problem because frame-wise labels are not given in the training stage while the only hint is video-level labels: whether each video contains action frames of interest. Previous methods aggregate frame-level class scores to produce video-level prediction and learn from video-level action labels. This formulation does not fully model the problem in that background frames are forced to be misclassified as action classes to predict video-level labels accurately. In this paper, we design Background Suppression Network (BaS-Net) which introduces an auxiliary class for background and has a two-branch weight-sharing architecture with an asymmetrical training strategy. This enables BaS-Net to suppress activations from background frames to improve localization performance. Extensive experiments demonstrate the effectiveness of BaS-Net and its superiority over the state-of-the-art methods on the most popular benchmarks – THUMOS'14 and ActivityNet. Our code and the trained model are available at https://github.com/Pilhyeon/BaSNet-pytorch. Pilhyeon Lee, Youngjung Uh, Hyeran Byun |
AAAI | 3 |
| 2020 | Exploiting Transferable Knowledge for Fairness-Aware Image Classification
Sunhee Hwang, Pilhyeon Lee, Seogkyu Jeon, Do-Hyung Kim 0004, Hyeran Byun |
ACCV (4) | 6 |
| 2020 | In-sample Contrastive Learning and Consistent Attention for Weakly Supervised Object Localization
Minsong Ki, Youngjung Uh, Wonyoung Lee 0004, Hyeran Byun |
ACCV (4) | 4 |
| 2020 | FairFaceGAN: Fairness-aware Facial Image-to-Image Translation
Sunhee Hwang, Do-Hyung Kim 0004, Mirae Do, Hyeran Byun |
BMVC | 5 |
| 2020 | Learning Texture Invariant Representation for Domain Adaptation of Semantic SegmentationabstractSince annotating pixel-level labels for semantic segmentation is laborious, leveraging synthetic data is an attractive solution. However, due to the domain gap between synthetic domain and real domain, it is challenging for a model trained with synthetic data to generalize to real data. In this paper, considering the fundamental difference between the two domains as the texture, we propose a method to adapt to the target domain's texture. First, we diversity the texture of synthetic images using a style transfer algorithm. The various textures of generated images prevent a segmentation model from overfitting to one specific (synthetic) texture. Then, we fine-tune the model with self-training to get direct supervision of the target texture. Our results achieve state-of-the-art performance and we analyze the properties of the model trained on the stylized dataset with extensive experiments. Myeongjin Kim, Hyeran Byun |
CVPR | 2 |
| 2020 | Unsupervised Image-to-Image Translation Via Fair Representation of Gender BiasabstractFairness becomes a critical issue of computer vision to reduce discriminative factors in various systems. Among computer vision tasks, Image-to-Image translation for facial attributes editing can yield discriminative results. The unexpected gender changed results can be generated instead of editing target attributes due to the dataset imbalance problem. In this work, we propose a framework of unsupervised Image-to-Image translation that learns a fair representation by separating the latent space of our model into two purposes: 1) Target Attribute Editing, 2) Gender Preserving. We evaluate the proposed framework on CelebA dataset. Both quantitive and qualitative results demonstrate that our method improves image quality and fairness than the prior Image-to-Image translation method. Sunhee Hwang, Hyeran Byun |
ICASSP | 2 |
| 2020 | Coping with Pandemics: Opportunities and Challenges for AI Multimedia in the "New Normal"abstractTheworld iswelcoming the newnormal - the coronavirus pandemic has significantly changed the way people live, work, communicate and learn. Almost everyone now is wearing a face mask when they go in public. People are working from home, some taking care of children at the same time. Bars and restaurants are limited to carry-out and delivery only. Meetings and conferences go online. Schools are closed and educators are instead holding video conference classes regularly. All these become the new normal as our ways of life. The panel thus provides a valuable opportunity for people from a variety of backgrounds to exchange views on opportunities and challenges for AI multimedia in the current and post pandemics era. Jiaying Liu 0001, Wen-Huang Cheng, Klara Nahrstedt, Ramesh Jain 0001, Elisa Ricci 0001, Hyeran Byun |
ACM Multimedia | 6 |
| 2020 | Selective residual learning for Visual Question Answering
Jongkwang Hong, Hyeran Byun |
Neurocomputing | 3 |
| 2020 | Unseen image generating domain-free networks for generalized zero-shot learning
Hoseong Kim, Jewook Lee, Hyeran Byun |
Neurocomputing | 3 |
| 2020 | Learning CNN features from DE features for EEG-based emotion recognition
Sunhee Hwang, Kibeom Hong, Guiyoung Son, Hyeran Byun |
Pattern Anal. Appl. | 4 |
| 2019 | Exploiting hierarchical visual features for visual question answering
Jongkwang Hong, Jianlong Fu, Youngjung Uh, Tao Mei 0001, Hyeran Byun |
Neurocomputing | 5 |
| 2018 | Face Identification for an in-vehicle Surveillance System Using Near Infrared CameraabstractFace identification is an essential topic in surveillance system research. Surveillance systems have many unconstrained conditions, e.g., brightness, occlusion, and user state variations. In this paper, we propose a multi-SVM based face recognition method using a near-infrared camera. Our method has a face identification scenario optimized for an in-vehicle surveillance system, which comprises two steps: (i) registering a driver and (ii) recognizing whether the driver is a registered. We perform feature extraction and recognition for each facial landmark. In the case of extreme exposure to light, we convert normal face images into simulated light overexposed images for learning. Thus, face classifiers for normal and extreme illumination conditions are simultaneously generated. We also create a new face dataset and evaluate our method with both our new and PolyU NIR datasets. Experimental results show that we achieve significantly higher recognition accuracy than existing methods. Minsong Ki, Bora Cho, Taejun Jeon, Yeongwoo Choi, Hyeran Byun |
AVSS | 5 |
| 2018 | Vision-Based Recognition of Road Regulation for Intelligent VehicleabstractIn this paper, we present a new framework to detect and recognize entire lanes and symbolic marks on high resolution road images. The first part of the framework utilizes local threshold to overcome the limitations of fixed threshold determination in road marking segmentation. The second part of the framework handles false detections caused by nearby objects on the roads such as vehicles and buildings by re-moving the areas that are not related to road surface using semantic segmentation. It also boosts recognition performance with a cascaded classifier structure that combines CNN for symbolic mark recognition and SVM for lane verification. The proposed lane detection achieves average Fl-score of 0.96 and symbol recognition achieves average Fl-score of 0.91. The proposed method is expected to advance the vehicle industry; with a GPU device, the proposed method can easily be embedded in smart vehicles. Kwangyong Lim, Yongwon Hong, Minsong Ki, Yeongwoo Choi, Hyeran Byun |
Intelligent Vehicles Symposium | 5 |
| 2018 | CoVieW'18: The 1st Workshop and Challenge on Comprehensive Video Understanding in the WildabstractThe 1st Workshop and Challenge on Comprehensive Video Understanding in the Wild, dubbed CoVieW'18, is held in Seoul, Korea on October 22, 2018, in conjuction with ACM Multimedia 2018. The workshop aims to solve the joint and comprehensive understanding problem in untrimmed videos with a particular emphasis on joint action and scene recognition. The workshop encourages researchers to participate in joint action and scene recognition challenge in untrimmed videos and to report their results. The workshop program includes 1 keynote speech, 2 invited speakers, 6 regular and challenge papers. The developments made in the workshop will deliver a step change in a variety of video applications. Kwanghoon Sohn, Ming-Hsuan Yang 0001, Hyeran Byun, Jongwoo Lim, Gee-Sern Hsu, Stephen Lin 0001, Euntai Kim, Seungryong Kim |
ACM Multimedia | 3 |
| 2018 | Exploiting Web Images for Video Highlight Detection With Triplet Deep RankingabstractHighlight detection from videos has been widely studied due to the fast growth of video contents. However, most existing approaches to highlight detection, either handcraft feature based or deep learning based, heavily rely on human-curated training data, which is very expensive to obtain and, thus, hinders the scalability to large datasets and unlabeled video categories. We observe that the largely available Web images can be applied as a weak supervision for highlight detection. For example, the top-ranked images in reference to the query “skiing” returned by a search engine may contain considerable positive samples of “skiing” highlights. Motivated by this observation, we propose a novel triplet deep ranking approach to video highlight detection using Web images as a weak supervision. The approach handles the relative preference of highlight scores between highlighting frames, nonhighlighting frames, and Web images by the triplet ranking constraints. Our approach can iteratively train two interdependent deep models (i.e., a triplet highlight model and a pairwise noise model) to deal with the noisy Web images in a single framework. We train the two models with relative preferences to generalize the capability regardless of the categories of training data. Therefore, our approach is fully category independent and exploits weakly supervised Web images. We evaluate our approach on two challenging datasets and achieve impressive results compared with the state-of-the-art pairwise ranking support vector machines, a robust recurrent autoencoder, and spatial deep convolution neural networks. We also empirically verify through cross-dataset evaluation that our category-independent model is fairly generalizable even if two different datasets do not share exactly the same categories. Hoseong Kim, Tao Mei 0001, Hyeran Byun, Ting Yao 0003 |
IEEE Trans. Multim. | 3 |
| 2017 | Discovering overlooked objects: Context-based boosting of object detection in indoor scenes
Jongkwang Hong, Yongwon Hong, Youngjung Uh, Hyeran Byun |
Pattern Recognit. Lett. | 4 |
| 2016 | Multi-view 3D reconstruction by random-search and propagation with view-dependent patch maps
Youngjung Uh, Hyeran Byun |
Multim. Tools Appl. | 2 |
| 2016 | A Space-Time Graph Optimization Approach Based on Maximum Cliques for Action DetectionabstractWe present an efficient action detection method that takes a space-time (ST) graph optimization approach for real-world videos. Given an ST graph representing the entire action video, our method identifies a maximum-weight connected subgraph (MWCS) indicating an action region by applying an optimization approach based on clique information. We define an energy function based on maximum weight cliques for subregions of the graph and formulate it using an optimization problem that can be represented as a linear system. Our energy function includes the maximum and connectivity properties for finding the MWCS, and its optimization solution indicates the probability of belonging to the maximum subgraph for each node. Our graph optimization method efficiently solves the detection problem by applying the clique-based approach and simple linear system solver. We demonstrate that our detection method results in a more accurate localization compared with conventional methods through our experimental results with real-world data sets, such as the Hollywood and MSR action data sets. We also show that our method outperforms the state-of-the-art methods of action detection. Sunyoung Cho, Hyeran Byun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | An Efficient Eye Tracking Using POMDP for Robust Human Computer Interaction
Ji Hye Rhee, Won Jun Sung, Mi Young Nam, Hyeran Byun, Phill-Kyu Rhee |
ICVS | 4 |
| 2014 | Efficient Colorization of Large-Scale Point Cloud Using Multi-pass Z-OrderingabstractWe present an efficient colorization method for a large scale point cloud using multi-view images. To address the practical issues of noisy camera parameters and color inconsistencies across multi-view images, our method takes an optimization approach for achieving visually pleasing point cloud colorization. We introduce a multi-pass Z-ordering technique that efficiently defines a graph structure to a large-scale and un-ordered set of 3D points, and use the graph structure for optimizing the point colors to be assigned. Our technique is useful for defining minimal but sufficient connectivities among 3D points so that the optimization can exploit the sparsity for efficiently solving the problem. We demonstrate the effectiveness of our method using synthetic datasets and a large-scale real-world data in comparison with other graph construction techniques. Sunyoung Cho, Jizhou Yan, Yasuyuki Matsushita, Hyeran Byun |
3DV | 4 |
| 2014 | Efficient Multiview Stereo by Random-Search and PropagationabstractWe present an efficient multi-view 3D reconstruction method based on randomization and propagation scheme. Our method progressively refines 3D point estimates by randomly perturbing the initial guess of 3D points and propagates photo-consistent ones to their neighbors. In contrast to previous refinement methods that perform local optimization for a better photo-consistency, our randomization approach takes lucky matchings for reducing the computational complexity. Experiments show favorable efficiency of the proposed method with the accuracy that is close to the state-of-the-art methods. Youngjung Uh, Yasuyuki Matsushita, Hyeran Byun |
3DV | 3 |
| 2013 | A unified approach to background adaptation and initialization in public scenes
Daeyong Park, Hyeran Byun |
Pattern Recognit. | 2 |
| 2013 | Recognizing human-human interaction activities using visual and textual information
Sunyoung Cho, Soo Yeong Kwak, Hyeran Byun |
Pattern Recognit. Lett. | 3 |
| 2012 | Generating panorama image by synthesizing multiple homographyabstractThis paper presents a method to generate image mosaics of a panoramic scene. In general, the relation between images which is required for mosaicing cannot be expressed by a single homography due to geometrical condition of the scene, even if the images are taken at the same position. Many existing methods are using only one homography to make panorama image while ignoring the geometrical variations. Therefore, they experience a lot of distortions and misalignments from input images which contain several planes which cannot be handled by one homography. In this paper, we present a novel method that utilizes synthesis of multiple homography to warp the images. Moreover, our method determines the number of homography automatically, without user's input. By our method, various distortions of shapes and mismatches can be reduced. Youngjung Uh, Hyeran Byun |
ICIP | 3 |
| 2012 | Color and shape feature-based detection of speed sign in real-timeabstractThis paper presents a method for detecting speed sign based on color and shape features in real-time under real-life environment. In our method, Region Of Interest(ROI) is extracted and verified based on shape feature. In the first step, ROI is roughly extracted by segmentation of a red rim and the segments are optimized by the boundary using guided image filtering. Next step, the shape-based detection verifies the extracted red rim. We compare three different shape-based detection methods, RSD, BCT, and STVUE, and the RSD shows the best speed sign detection rate of 93% on the experimental data of 62 images containing 85 speed sign. Seunggyu Kim, Youngjung Uh, Hyeran Byun |
SMC | 4 |
| 2012 | Incremental face recognition for large-scale social network services
Kwontaeg Choi, Kar-Ann Toh, Hyeran Byun |
Pattern Recognit. | 3 |
| 2012 | Dynamic curve color model for image matting
Sunyoung Cho, Hyeran Byun |
Pattern Recognit. Lett. | 2 |
| 2012 | Service-oriented architecture based on biometric using random features and incremental neural networks
Kwontaeg Choi, Kar-Ann Toh, Youngjung Uh, Hyeran Byun |
Soft Comput. | 4 |
| 2011 | Plot preservation approach for video summarizationabstractThis paper presents an effective method for summarizing content of the video with plot preserved. The proposed method provides static video summary and consists of three major procedures (1) extracting keyframes regarding temporal information, (2) estimating Region of Interest (ROI) from the extracted frames, and (3) assembling the ROI into one image by arranging them according to the temporal order and their size, while they blend each other smoothly. In each process, we make the method concise for computational complexity. Experiment on various types of video (e.g. movie, animation, home video) demonstrates that the proposed method generates expressive video summarization and conserves the plot information. The main contribution relies on reflecting temporal information in video summarization. Yeosun Lim, Youngjung Uh, Hyeran Byun |
SMC | 3 |
| 2011 | Realtime training on mobile devices for face recognition applications
Kwontaeg Choi, Kar-Ann Toh, Hyeran Byun |
Pattern Recognit. | 3 |
| 2010 | Adaptive Color Curve Models for Image MattingabstractImage matting is the process of extracting a foreground element from a single image with limited user input. To solve the inherently ill-posed problem, there exist various methods which use specific color model. One representative method assumes that the colors of the foreground and background elements satisfy the linear color model. The other recent method considers line-point color model and point-point color model. In this paper we present a new adaptive color curve model for image matting. We assume that the colors of local region form curve. Based on these pixels in the local region, we adaptively construct a curve model using quadratic Bézier curve model. This curve model enables us to derive a matting equation for estimating alphas of pixels forming a curve using quadratic formula. We show that our model estimates alpha mattes comparable or more accurately than recent existing methods. Sunyoung Cho, Hyeran Byun |
ICPR | 2 |
| 2010 | Coarse-to-Fine Particle Filter by Implicit Motion Estimation for 3D Head Tracking on Mobile DevicesabstractDue to the widely spread mobile devices over the years, a low cost implementation of an efficient head tracking system is becoming more useful for a wide range of applications. In this paper, we make an attempt to solving real-time 3D head tracking problem on mobile devices by enhancing the fitness of the dynamics. In our method, the particles are generated by implicit motion estimation between two particles rather than the explicit motion estimation using corresponding point matching between consecutive two frames. This generation is applied iteratively using coarse-to fine strategy in order to handle a large motion using a small number of particle. This reduces the computational cost while preserving the performance. We evaluate the efficiency and effectiveness of the proposed algorithm by empirical experiments. Finally, we demonstrate our method on a recent mobile phone. Hachoen Sung, Kwontaeg Choi, Sunyoung Cho, Hyeran Byun |
ICPR | 4 |
| 2009 | Object-Wise Multilayer Background Ordering for Public Area SurveillanceabstractPublic area is one of the most significant places which need video surveillance. However, pixel-wise adaptive background subtraction methods are disturbed by incessantly passing or temporally staying foreground due to its adaptability. In such an environment, even the initialization of background is not free from the influence of foregrounds. If the adaptability is modified carelessly for selective learning, the stability of the background model will be damaged. Adjusting or fusing the learning rate slows down the false learning rate but cannot solve the problems. In this paper, we present a multilayer background modeling algorithm for public area surveillance. We efficiently cluster regions in object-wise using spatiotemporal cohesion together with spectral similarity by comparing inputs with background layer. And we classify the clustered regions and update the multi-layer model according to the results. Using the PETS data, we show that the proposed method not only maintain the background robustly but also initialize background with stationary object detection in crowded public area. Daeyong Park, Hyeran Byun |
AVSS | 2 |
| 2009 | Easy real object insertion tool for composite photographsabstractIn this paper, we present a tool for easily creating natural composite photographs with just a few mouse clicks. The proposed tool works with three simple steps: 1) draw a bounding box around reference object of background photograph for object height estimation; 2) draw a trimap on the object photograph with mouse drag; 3) move a mouse in location to place an object. We present automatic algorithm for object height estimation in single image and matting technique for extracting object in detail. Experimental results show that our results are created easily and naturally using our tool. Sunyoung Cho, Hyeran Byun |
MoMM | 2 |
| 2009 | Human tracking and silhouette extraction for human-robot interaction systems
Jung-Ho Ahn 0002, Cheolmin Choi, Soo Yeong Kwak, Kil-Cheon Kim, Hyeran Byun |
Pattern Anal. Appl. | 5 |
| 2008 | A collaborative face recognition framework on a social network platformabstractFace recognition has many useful applications spanning surveillance, law enforcement, information security, smart card and entertainment technologies. Very recently, a learning based face recognition system is also seen to be applied to web platform combining face recognition and web service. However, many existing methods which focused on recognition accuracy cannot cope with the new social network platform because the adopted static learning approach is not adaptive to daily updated photographs among the massive number of users. In this paper, we discuss the difference between a stand-alone based system and a social network based system and propose a new collaborative face recognition framework where a redundant tagging can be avoided via sharing the identification information for efficient update under the social network platform. Our Experiments (including a web stress test) using a public database show that the proposed method records a better accuracy than that of the state-of-the-art classifier SVM adopting a polynomial kernel and has fast execution time for both training and testing. Kwontaeg Choi, Hyeran Byun, Kar-Ann Toh |
FG | 2 |
| 2008 | An efficient and accurate hierarchical ICIA fitting method for 3D Morphable ModelsabstractWe propose the efficient and accurate hierarchical ICIA fitting method for 3D Morphable Models (3DMMs). The conventional ICIA fitting method for 3DMMs requires a long computation time because the 3D face model contains a large number of vertices and it also requires to compute the Hessian matrix using the visible vertices every iteration. For the efficient fitting, we use the hierarchical fitting that use a set of multi-resolution 3D face model and the Gaussian image pyramid. For more accurate fitting, we use a two-stage parameter update that only update the rigid and the texture parameters and then update all parameters after the initial convergence. We present several experiment results to prove that our proposed method shows better performance than previous works. Bong-Nam Kang, Daijin Kim 0001, Hyeran Byun |
FG | 3 |
| 2008 | Multi Subspaces Active Appearance ModelsabstractThe original Active Appearance Model(AAM) uses the mean matrix of gradient matrixes instead of a gradient matrix which should be recomputed with respect to a varying parameter at a fitting phase. By this property, the original AAM can guarantee a fast fitting speed because it avoids computation of a gradient matrix of which a computation complexity is high. However, the fixed gradient matrix is not a good choice when the distribution of a training database is nonlinear because the mean can not represent the variation of a training database. To overcome this problem, this paper proposes multi subspaces AAM. First, we divide a training database into multi subspaces along the illumination direction, and build the independent AAM for each subspace. At a fitting phase, we adaptively choose a subspace well fit to a target image. However, the parameter update problem is occurred because a subspace can be changed during a fitting phase. To solve this problem, we propose a linear transform matrix on an eigenspace. In experiments, we apply the proposed method to Yale Face Database B and demonstrate that the method is robust for facial images under various illuminations. Junyeong Yang, Hyeran Byun |
FG | 2 |
| 2008 | Multi-resolution 3D morphable models and its matching methodabstractThe inverse compositional image alignment (ICIA) is known as an efficient matching method for 3D morphable models (3DMMs). However, it requires a long computation time since the 3D face models consist of a large number of vertices. Also, it requires to recompute the Hessian matrix using the visible vertices every iteration. For a fast and an efficient matching, we propose the efficient and accurate hierarchical ICIA (HICIA) matching method for 3DMMs. The proposed matching method requires multi-resolution 3D face models and the Gaussian image pyramid. The multi-resolution 3D face models are built by sub-sampling at the 2:1 sampling rate to construct the lower-resolution 3D face models. For more accurate matching, we use a two-stage model parameter update that only updates the rigid and the texture parameters and then updates all parameters after the initial convergence. We present several experimental results to prove that the proposed method shows better performance than that of the conventional ICIA matching method. Bong-Nam Kang, Hyeran Byun, Daijin Kim 0001 |
ICPR | 2 |
| 2008 | Curve fitting algorithm using iterative error minimization for sketch beautificationabstractIn previous sketch recognition systems, curve has been fitted by a bit heuristic method. In this paper, we solved the problem by finding the optimal parameter of quadratic Bezier curve and utilize the error minimization between an input curve and a fitting curve by using iterative error minimization. First, we interpolated the input curve to compute the distance because the input curve consists of a set of sparse points. Then, we define the objective function. To find the optimal parameter, we assume that the initial parameter is known. Then, we derive the gradient vector with respect to the current parameter, and the parameter is updated by the gradient vector. This two steps are repeated until the error is not reduced. From the experiment, the average approximation error of the proposed algorithm was 0.946433 about 1400 synthesized curves, and this result demonstrates that the given curve can be fitted very closely by using the proposed fitting algorithm. Junyeong Yang, Hyeran Byun |
ICPR | 2 |
| 2008 | Feature extraction method based on cascade noise elimination for sketch recognitionabstractFreehand sketching is a very efficient means for us to communicate each other. As Table PC is widely popularized, the research about sketch recognition became one of important research issue. To recognize sketch, the feature point should be extracted and then each feature point is analyzed as line or curve. However, most of feature extraction algorithms suffers from noise which is occurred from the bad drawing sketch. In this paper, we propose the feature extraction algorithm robust to noise. The proposed algorithm consists of three cascade steps: candidate feature point extraction, noise reduction, and hook elimination. At the candidate feature point extraction step, the feature points is selected among input points. Then, in second step, we reduce the noise which is occurred from the previous step by using noise reduction rule based on inner product between two neighbor vectors. Finally, the hook, which can not be eliminated from two previous steps, is eliminated by the proposed hook elimination method. The experimental result shows that the average approximation error is less than 1 about 1004 line-curve hybrid shapes, and the proposed algorithm is the good feature methods. Junyeong Yang, Hyeran Byun |
ICPR | 2 |
| 2008 | A fast image retrieval system using index lookup table on mobile deviceabstractThe development of mobile-based image retrieval system is required for the efficient management about various image data with fast growth of mobile devices. The resource of mobile phone is a bit limited due to low memory capacity and low CPU performance. Hence, the system which is applicable on mobile phone should overcome the problem of limited resource. In this paper, we propose the fast feature extraction method which is applicable on mobile environment. The proposed algorithm is based on the lookup table which converts RGB color space to one of 36 indices on HSV color space. By the lookup table-based feature extraction, we extract three histograms which is about color distribution and the location of the distribution. Then, we load the proposed algorithm to mobile phone. Our system can query by three different quries: query-by-image, query-by-color, and query-by-blob. The experimental result shows that the proposed method is a fast and stable engine enough to utilize on mobile platform. Junyeong Yang, Sanghyuk Park, Hacheon Seong, Hyeran Byun, Yeong-Kyu Lim |
ICPR | 4 |
| 2007 | Salient human detection for robot vision
Soo Yeong Kwak, ByoungChul Ko, Hyeran Byun |
Pattern Anal. Appl. | 3 |
| 2006 | Accurate Foreground Extraction Using Graph Cut with Trimap Estimation
Jung-Ho Ahn 0002, Hyeran Byun |
PSIVT | 2 |
| 2005 | FRIP: a region-based image retrieval tool using automatic image segmentation and stepwise Boolean AND matchingabstractWe present our region-based image retrieval tool, finding region in the picture (FRIP), that is able to accommodate, to the extent possible, region scaling, rotation, and translation. Our goal is to develop an effective retrieval system to overcome a few limitations associated with existing systems. To do this, we propose adaptive circular filters used for semantic image segmentation, which are based on both Bayes' theorem and texture distribution of image. In addition, to decrease the computational complexity without losing the accuracy of the search results, we extract optimal feature vectors from segmented regions and apply them to our stepwise Boolean AND matching scheme. The experimental results using real world images show that our system can indeed improve retrieval performance compared to other global property-based or region-of-interest-based image retrieval methods. ByoungChul Ko, Hyeran Byun |
IEEE Trans. Multim. | 2 |
| 2003 | Multi-class Support Vector Machines with Case-Based Combination for Face Recognition
Jaepil Ko, Hyeran Byun |
CAIP | 2 |
| 2003 | Various decomposition methods applied to face recognitionabstractFace recognition has mainly focused on face representation, so a simple classifier is frequently used. For a robust system, it is common to construct a multiclass classifier by combining outputs of several binary ones. In this paper, we overviews basic decomposition and decoding schemes and propose new methods then give empirical results of recognition performance on the ORL face dataset. Jaepil Ko, Eunju Kim, Hyeran Byun |
IJCNN | 3 |
| 2003 | Robust Face Detection and Tracking for Real-Life ApplicationsabstractIn this paper, we propose a new face detection and tracking algorithm for real-life telecommunication applications, such as video conferencing, cellular phone and PDA. We combine template-based face detection and tracking method with color information to track a face regardless of various lighting conditions and complex backgrounds as well as the race. Based on our experiments, we generate robust face templates from wavelet-transformed lowpass and two highpass subimages at the second level low-resolution. However, since template matching is generally sensitive to the change of illumination conditions, we propose a new type of preprocessing method. Tracking method is applied to reduce the computation time and predict precise face candidate region even though the movement is not uniform. Facial components are also detected using k-means clustering and their geometrical properties. Finally, from the relative distance of two eyes, we verify the real face and estimate the size of facial ellipse. To validate face detection and tracking performance of our algorithm, we test our method using six different video categories of QCIF size which are recorded in dynamic environments. Hyeran Byun, ByoungChul Ko |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2003 | Real-Time Pedestrian Detection Using Support Vector MachinesabstractIn this paper, we present a real-time pedestrian detection method in outdoor environments. It is necessary for pedestrian detection to implement obstacle and face detection which are major parts of a walking guidance system for the visually impaired. It detects foreground objects on the ground, discriminates pedestrians from other noninterest objects, and extracts candidate regions for face detection and recognition. For effective real-time pedestrian detection, we have developed a method using stereo-based segmentation and the SVM (Support Vector Machines), which works well particularly in binary classification problem (e.g. object detection). We used vertical edge features extracted from arms, legs and torso. In our experiments, test results on a large number of outdoor scenes demonstrated the effectiveness of the proposed pedestrian detection method. Seonghoon Kang, Hyeran Byun, Seong-Whan Lee |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2003 | Extracting Salient Regions And Learning Importance Scores In Region-Based Image RetrievalabstractIn this paper, we propose a new method for extracting salient regions and learning their importance scores in region-based image retrieval. In Region-Based Image Retrieval (RBIR), not all the regions are important for retrieving similar images and rather, in retrieval, the user is often interested in performing a query on only one or a few regions rather than the whole image. Therefore, for a successful retrieval system, it is an important issue to specify which regions are important for retrieving an image. To extract salient regions from images automatically, we make three assumptions and determine salient regions with their importance scores. In this paper, we apply the relevance feedback algorithm to the matching process as two different purposes: one is for updating importance scores of salient regions and the other is for updating weights of feature vectors. By using our relevance feedback method, the matching process can improve retrieval performance interactively and allow progressive refinement of query results according to the user's feedback action. Through experiments and comparison with other methods, our proposed method shows good performance as well as easy and semantic interface for region-based image retrieval. The efficacy of our method is validated using a set of 3000 images from Corel-photo CD. ByoungChul Ko, Hyeran Byun |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2003 | A Survey on Pattern Recognition Applications of Support Vector MachinesabstractIn this paper, we present a survey on pattern recognition applications of Support Vector Machines (SVMs). Since SVMs show good generalization performance on many real-life data and the approach is properly motivated theoretically, it has been applied to wide range of applications. This paper describes a brief introduction of SVMs and summarizes its various pattern recognition applications. Hyeran Byun, Seong-Whan Lee |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2003 | N-division output coding method applied to face recognition
Jaepil Ko, Hyeran Byun |
Pattern Recognit. Lett. | 2 |
| 2002 | Text Extraction in Digital News Video Using Morphology
Hyeran Byun, InYoung Jang, Yeongwoo Choi |
Document Analysis Systems | 1 |
| 2002 | Scene Text Extraction in Complex Images
Hyeran Byun, Myung-Cheol Roh, Kil-Cheon Kim, Yeongwoo Choi, Seong-Whan Lee |
Document Analysis Systems | 1 |
| 2002 | A Simple Illumination Normalization Algorithm for Face Recognition
Jaepil Ko, Eunju Kim, Hyeran Byun |
PRICAI | 3 |
| 2002 | Automatic text extraction in news images using morphology
InYoung Jang, ByoungChul Ko, Hyeran Byun, Yeongwoo Choi |
VCIP | 3 |
| 2002 | Automatic generation of structured hyperdocuments from document images
Jeong-Seon Park, Hyeran Byun, Jongsub Moon, Seong-Whan Lee |
Pattern Recognit. | 3 |
| 2001 | A New Content-Based Image Retrieval System Using Hang Gesture And Relevange Feedback
ByoungChul Ko, Hyeran Byun |
ICME | 3 |
| 2001 | Region-based Image Retrieval Using Probabilistic Feature Relevance Learning
ByoungChul Ko, Hyeran Byun |
Pattern Anal. Appl. | 3 |
| 2001 | Blending video with synthetic images created from matched viewing environments
Moonho Park, Heedong Ko, Hyeran Byun |
Pattern Recognit. | 3 |
| 2000 | Region-Based Image Retrieval System Using Efficient Feature DescriptionabstractIn this paper we introduce a region-based image retrieval system, FRIP. This system includes a robust image segmentation scheme using scaled and shifted color and shape description scheme using modified radius-based signature. For image segmentation, by using our proposed circular filter, we can keep the boundary of object naturally and merge small senseless regions of object into a whole body. For efficient shape description, we extract 5 features from each region: color, texture, scale, location, and shape. From these features, we calculate the similarity distance between the query and database regions and it returns the top K-nearest neighbor regions. ByoungChul Ko, Hae-Sung Lee, Hyeran Byun |
ICPR | 3 |
| 1999 | A mixed environment for tele-meeting between real and virtual humanabstractWe propose a mixed environment for tele-meeting between a real human and computer-generated virtual humans. For the natural interactions among the participants, the mixed environment is augmented by the visual and aural cues that are crucial for tele-meeting. As the real and virtual worlds are seamlessly coupled, the proposed system can give the participants the sense of tele-presence. The system may be applied to immersive tele-meeting, tele-education, and interactive TV programs. Moonho Park, Laehyun Kim, Heedong Ko, Hyeran Byun |
IROS | 4 |
| 1999 | Interactive virtual studio and immersive televiewer environmentabstractIn this paper, we propose a novel virtual studio system in which an anchor in the virtual set interacts with televiewers as if they were sharing the same environment. A televiewer participates in the virtual studio environment by sensing and controlling a dummy head equipped with camera, speaker and microphones. The dummy head acts as a surrogate televiewer, providing the viewpoint experienced by the televiewer via a video camera and the sound experienced by the televiewer via microphone in its head. The anchor can not only interact with the virtual set elements but also share the physical studio with the surrogate televiewers. A televiewer with a head-mounted display (HMD) may feel immersed in the virtual studio environment seamlessly combining the virtual set elements with the real studio elements and interact with the anchor and vice versa. The proposed system consists of Interactive Virtual Studio (IVS) environment and Immersive Televiewer Environment (ITE) in which all the physical elements are collected and managed through IVS and the seamlessly mixed virtual and real elements are experienced via ITE. The essential idea is to have a dual universe where what makes a natural interface only physically should remain as physical and what makes easier to represent virtually should remain virtual and these two parallel universe should be coordinated seamlessly to provide the proper mix of the virtual and real mixed reality experience. In practice, this new interactive virtual studio for the immersive tele-meeting environment may be applied to the production of interactive TV program, tele-conferencing, tele-education and others. Laehyun Kim, Heedong Ko, Moonho Park, Hyeran Byun |
VRST | 4 |