Hyunjung Shim

dblp:72/4620 · DBLP profile ↗
← Back
67ranked-venue papers
11as first author
43since 2021 · last 2026
0000-0001-6796-1058ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 10 first-author · 20 since 2021Databases, data management, data science and information retrieval · 10 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Rethinking Direct Preference Optimization in Diffusion Models
abstract
Aligning text-to-image (T2I) diffusion models with human preferences has emerged as a critical research challenge. While Direct Preference Optimization (DPO) has established a foundation for preference learning in large language models (LLMs), its extension to diffusion models remains limited in alignment performance. In this work, we propose an enhanced version of Diffusion-DPO by introducing a stable reference model update strategy. This strategy facilitates the exploration of better alignment solutions while maintaining training stability. Moreover, we design a timestep-aware optimization strategy that further boosts performance by addressing preference learning imbalance across timesteps. Through the synergistic combination of our exploration and timestep-aware optimization, our method significantly improves the alignment performance of Diffusion-DPO on human preference evaluation benchmarks, achieving state-of-the-art results.
Junyong Kang, Seohyun Lim, Kyungjune Baek, Hyunjung Shim
AAAI4
2026 PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation
abstract
Automating scientific poster generation requires hierarchical document understanding and coherent content-layout planning. Existing methods often rely on flat summarization or optimize content and layout separately. As a result, they often suffer from information loss, weak logical flow, and poor visual balance. We present PosterForest, a training-free framework for scientific poster generation. Our method introduces the Poster Tree, a structured intermediate representation that captures document hierarchy and visual-textual semantics across multiple levels. Building on this representation, content and layout agents perform hierarchical reasoning and recursive refinement, progressively optimizing the poster from global organization to local composition. This joint optimization improves semantic coherence, logical flow, and visual harmony. Experiments show that PosterForest outperforms prior methods in both automatic and human evaluations, without additional training or domain-specific supervision.
Jiho Choi, Seojeong Park, Seongjong Song, Hyunjung Shim
ACL (1)4
2026 MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
abstract
Video Moment Retrieval (MR) aims to localize moments within a video based on a given natural language query. Given the prevalent use of platforms like YouTube for information retrieval, the demand for MR techniques is significantly growing. Recent DETR-based models have made notable advances in performance but still struggle with accurately localizing short moments. Through data analysis, we identified limited feature diversity in short moments, which motivated the development of MomentMix. MomentMix generates new short-moment samples by employing two augmentation strategies: ForegroundMix and BackgroundMix, each enhancing the ability to understand the query-relevant and irrelevant frames, respectively. Additionally, our analysis of prediction bias revealed that short moments particularly struggle with accurately predicting their center positions and length of moments. To address this, we propose a Length-Aware Decoder, which conditions length through a novel bipartite matching process. Our extensive studies demonstrate the efficacy of our length-aware approach, especially in localizing short moments, leading to improved overall performance. Our method surpasses state-of-the-art DETR-based methods on benchmark datasets, achieving the highest R1 and mAP on QVHighlights and the highest [email protected] on TACoS and Charades-STA (such as a 9.62% gain in [email protected] and an 16.9% gain in mAP average for QVHighlights). The code is available at https://github.com/sjpark5800/LA-DETR.
Seojeong Park, Jiho Choi, Kyungjune Baek, Hyunjung Shim
WACV4
2025 Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
abstract
Despite the impressive success of text-to-image (TTI) models, existing studies overlook the issue of whether these models accurately convey factual information. In this paper, we focus on the problem of image hallucination, where images created by TTI models fail to faithfully depict factual content. To address this, we introduce I-HallA (Image Hallucination evaluation with Question Answering), a novel automated evaluation metric that measures the factuality of generated images through visual question answering (VQA). We also introduce I-HallA v1.0, a curated benchmark dataset for this purpose. As part of this process, we develop a pipeline that generates high-quality question-answer pairs using multiple GPT-4 Omni-based agents, with human judgments to ensure accuracy. Our evaluation protocols measure image hallucination by testing if images from existing TTI models can correctly respond to these questions. The I-HallA v1.0 dataset comprises 1.2K diverse image-text pairs across nine categories with 1,000 rigorously curated questions covering various compositional challenges. We evaluate five TTI models using I-HallA and reveal that these state-of-the-art models often fail to accurately convey factual information. Moreover, we validate the reliability of our metric by demonstrating a strong Spearman correlation (ρ=0.95) with human judgments. Our benchmark dataset and metric can serve as a foundation for developing factually accurate TTI models.
Youngsun Lim, Hojun Choi, Hyunjung Shim
AAAI3
2025 I0T: Embedding Standardization Method Towards Zero Modality Gap
abstract
Contrastive Language-Image Pretraining (CLIP) enables zero-shot inference in downstream tasks such as image-text retrieval and classification.However, recent works extending CLIP suffer from the issue of modality gap, which arises when the image and text embeddings are projected to disparate manifolds, deviating from the intended objective of image-text contrastive learning.We discover that this phenomenon is linked to the modality-specific characteristic that each image or text encoder independently possesses.Herein, we propose two methods to address the modality gap: (1) a post-hoc embedding standardization method, I0T post that reduces the modality gap approximately to zero and (2) a trainable method, I0T async , to alleviate the modality gap problem by adding two normalization layers for each encoder.Our I0T framework can significantly reduce the modality gap while preserving the original embedding representations of trained models with their locked parameters.In practice, I0T post can serve as an alternative explainable automatic evaluation metric of widely used CLIPScore (CLIP-S).The code is available in https://github.com/xfactlab/I0T.
Na Min An, Eunki Kim, James Thorne, Hyunjung Shim
ACL (1)4
2025 Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation
abstract
Open-Vocabulary Part Segmentation (OVPS) is an emerging field for recognizing fine-grained parts in unseen categories. We identify two primary challenges in OVPS: (1) the difficulty in aligning part-level image-text correspondence, and (2) the lack of structural understanding in segmenting object parts. To address these issues, we propose Part-CATSeg, a novel framework that integrates object-aware part-level cost aggregation, compositional loss, and structural guidance from DINO. Our approach employs a disentangled cost aggregation strategy that handles object and part-level costs separately, enhancing the precision of part-level segmentation. We also introduce a compositional loss to better capture part-object relationships, compensating for the limited part annotations. Additionally, structural guidance from DINO features improves boundary delineation and inter-part understanding. Extensive experiments on Pascal-Part-116, ADE20K-Part-234, and PartImageNet datasets demonstrate that our method significantly outperforms state-of-the-art approaches, setting a new baseline for robust generalization to unseen part categories.
Jiho Choi, Seonho Lee, Minhyun Lee, Hyunjung Shim
CVPR5
2025 Classifier-guided CLIP Distillation for Unsupervised Multi-label Classification
abstract
Multi-label classification is crucial for comprehensive image understanding, yet acquiring accurate annotations is challenging and costly. To address this, a recent study suggests exploiting unsupervised multi-label classification leveraging CLIP, a powerful vision-language model. Despite CLIP’s proficiency, it suffers from view-dependent predictions and inherent bias, limiting its effectiveness. We propose a novel method that addresses these issues by leveraging multiple views near target objects, guided by Class Activation Mapping (CAM) of the classifier, and debiasing pseudo-labels derived from CLIP predictions. Our Classifier-guided CLIP Distillation (CCD) enables selecting multiple local views without extra labels and debiasing predictions to enhance classification performance. Experimental results validate our method’s superiority over existing techniques across diverse datasets. The code is available at https://github.com/k0u-id/CCD.
Dongseob Kim, Hyunjung Shim
CVPR2
2025 No Thing, Nothing: Highlighting Safety-Critical Classes for Robust LiDAR Semantic Segmentation in Adverse Weather
abstract
Existing domain generalization methods for LiDAR semantic segmentation under adverse weather struggle to accurately predict "things" categories compared to "stuff" categories. In typical driving scenes, "things" categories can be dynamic and associated with higher collision risks, making them crucial for safe navigation and planning. Recognizing the importance of "things" categories, we identify their performance drop as a serious bottleneck in existing approaches. We observed that adverse weather induces degradation of semantic-level features and both corruption of local features, leading to a misprediction of "things" as "stuff". To mitigate these corruptions, we suggests our method, NTN - segmeNt Things for No-accident. To address semantic-level feature corruption, we bind each point feature to its superclass, preventing the misprediction of things classes into visually dissimilar categories. Additionally, to enhance robustness against local corruption caused by adverse weather, we define each LiDAR beam as a local region and propose a regularization term that aligns the clean data with its corrupted counterpart in feature space. NTN achieves state-of-the-art performance with a +2.6 mIoU gain on the SemanticKITTI-to-SemanticSTF benchmark and +7.9 mIoU on the SemanticPOSS-to-SemanticSTF benchmark. Notably, NTN achieves a +4.8 and +7.9 mIoU improvement on "things" classes, respectively, highlighting its effectiveness.
Hwijeong Lee, Inha Kang, Hyunjung Shim
CVPR4
2025 Scribble-Guided Diffusion for Training-Free Text-to-Image Generation
abstract
Recent advancements in text-to-image diffusion models have shown impressive results but often fail to fully capture user’s intent. Existing methods combining textual inputs with bounding boxes or region masks lack precise spatial guidance, leading to misaligned or unintended object orientations. To address these issues, we propose Scribble-Guided Diffusion (ScribbleDiff), a training-free approach that employs user-provided scribbles as visual prompts for image generation. However, incorporating scribbles poses challenges due to their sparse and thin nature, which complicates accurate alignment. To resolve this, we introduce moment alignment and scribble propagation, enabling effective and flexible alignment between generated images and scribble inputs. Experiments on the PASCAL-Scribble dataset demonstrate notable improvements in spatial control and consistency, validating the effectiveness of our method in scribble-guidance.
Seonho Lee, Jiho Choi, Seohyun Lim, Jiwook Kim, Hyunjung Shim
ICIP5
2025 DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity Preservation
abstract
Score distillation sampling (SDS) has emerged as an effective framework in text-driven 3D editing tasks, leveraging diffusion models for 3D-consistent editing. However, existing SDS-based 3D editing methods suffer from long training times and produce low-quality results. We identify that the root cause of this performance degradation is their conflict with the sampling dynamics of diffusion models. Addressing this conflict allows us to treat SDS as a diffusion reverse process for 3D editing via sampling from data space. In contrast, existing methods naively distill the score function using diffusion models. From these insights, we propose DreamCatalyst, a novel framework that considers these sampling dynamics in the SDS framework. Specifically, we devise the optimization process of our DreamCatalyst to approximate the diffusion reverse process in editing tasks, thereby aligning with diffusion sampling dynamics. As a result, DreamCatalyst successfully reduces training time and improves editing quality. Our method offers two modes: (1) a fast mode that edits Neural Radiance Fields (NeRF) scenes approximately 23 times faster than current state-of-the-art NeRF editing methods, and (2) a high-quality mode that produces superior results about 8 times faster than these methods. Notably, our high-quality mode outperforms current state-of-the-art NeRF editing methods in terms of both speed and quality. DreamCatalyst also surpasses the state-of-the-art 3D Gaussian Splatting (3DGS) editing methods, establishing itself as an effective and model-agnostic 3D editing solution.
Jiwook Kim, Seonho Lee, Jaeyo Shin, Jiho Choi, Hyunjung Shim
ICLR5
2025 DGQ: Distribution-Aware Group Quantization for Text-to-Image Diffusion Models
abstract
Despite the widespread use of text-to-image diffusion models across various tasks, their computational and memory demands limit practical applications. To mitigate this issue, quantization of diffusion models has been explored. It reduces memory usage and computational costs by compressing weights and activations into lower-bit formats. However, existing methods often struggle to preserve both image quality and text-image alignment, particularly in lower-bit($<$ 8bits) quantization. In this paper, we analyze the challenges associated with quantizing text-to-image diffusion models from a distributional perspective. Our analysis reveals that activation outliers play a crucial role in determining image quality. Additionally, we identify distinctive patterns in cross-attention scores, which significantly affects text-image alignment. To address these challenges, we propose Distribution-aware Group Quantization (DGQ), a method that identifies and adaptively handles pixel-wise and channel-wise outliers to preserve image quality. Furthermore, DGQ applies prompt-specific logarithmic quantization scales to maintain text-image alignment. Our method demonstrates remarkable performance on datasets such as MS-COCO and PartiPrompts. We are the first to successfully achieve low-bit quantization of text-to-image diffusion models without requiring additional fine-tuning of weight quantization parameters. Code is available at \link{https://github.com/ugonfor/DGQ}.
Hyogon Ryu, NaHyeon Park, Hyunjung Shim
ICLR3
2025 Label-Augmented Dataset Distillation
abstract
Traditional dataset distillation primarily focuses on image representation while often overlooking the important role of labels. In this study, we introduce Label-Augmented Dataset Distillation (LADD), a new dataset distillation framework enhancing dataset distillation with label augmentations. LADD sub-samples each synthetic image, generating additional dense labels to capture rich semantics. These dense labels require only a 2.5% increase in storage (ImageNet subsets) with significant performance benefits, providing strong learning signals. Our label-generation strategy can complement existing dataset distillation methods and significantly enhance their training efficiency and performance. Experimental results demonstrate that LADD outperforms existing methods in terms of computational overhead and accuracy. With three high-performance dataset distillation algorithms, LADD achieves remarkable gains by an average of 14.9% in accuracy. Furthermore, the effectiveness of our method is proven across various datasets, distillation hyperparameters, and algorithms. Finally, our method improves the cross-architecture robustness of the distilled dataset, which is important in the application scenario.
Seoungyoon Kang, Youngsun Lim, Hyunjung Shim
WACV3
2025 Text optimization with latent inversion for non-rigid image editing
Yunji Jung, Seokju Lee, Tair Djanibekov, Jong Chul Ye, Hyunjung Shim
Pattern Recognit. Lett.5
2024 Weakly Supervised Semantic Segmentation for Driving Scenes
abstract
State-of-the-art techniques in weakly-supervised semantic segmentation (WSSS) using image-level labels exhibit severe performance degradation on driving scene datasets such as Cityscapes. To address this challenge, we develop a new WSSS framework tailored to driving scene datasets. Based on extensive analysis of dataset characteristics, we employ Contrastive Language-Image Pre-training (CLIP) as our baseline to obtain pseudo-masks. However, CLIP introduces two key challenges: (1) pseudo-masks from CLIP lack in representing small object classes, and (2) these masks contain notable noise. We propose solutions for each issue as follows. (1) We devise Global-Local View Training that seamlessly incorporates small-scale patches during model training, thereby enhancing the model's capability to handle small-sized yet critical objects in driving scenes (e.g., traffic light). (2) We introduce Consistency-Aware Region Balancing (CARB), a novel technique that discerns reliable and noisy regions through evaluating the consistency between CLIP masks and segmentation predictions. It prioritizes reliable pixels over noisy pixels via adaptive loss weighting. Notably, the proposed method achieves 51.8\% mIoU on the Cityscapes test dataset, showcasing its potential as a strong WSSS baseline on driving scene datasets. Experimental results on CamVid and WildDash2 demonstrate the effectiveness of our method across diverse datasets, even with small-scale datasets or visually challenging conditions. The code is available at https://github.com/k0u-id/CARB.
Dongseob Kim, Junsuk Choe, Hyunjung Shim
AAAI4
2024 SeiT++: Masked Token Modeling Improves Storage-Efficient Training
Minhyun Lee, Song Park, Byeongho Heo, Dongyoon Han, Hyunjung Shim
ECCV (28)5
2024 Rethinking Data Augmentation for Robust LiDAR Semantic Segmentation in Adverse Weather
Hyunjung Shim
ECCV (16)3
2024 Memory-Efficient Fine-Tuning for Quantized Diffusion Model
Hyogon Ryu, Seohyun Lim, Hyunjung Shim
ECCV (16)3
2024 Learning from Spatio-temporal Correlation for Semi-Supervised LiDAR Semantic Segmentation
abstract
We address the challenges of the semi-supervised LiDAR segmentation (SSLS) problem, particularly in low-budget scenarios. The two main issues in low-budget SSLS are the poor-quality pseudo-labels for unlabeled data, and the performance drops due to the significant imbalance between ground-truth and pseudo-labels. This imbalance leads to a vicious training cycle. To overcome these challenges, we leverage the spatio-temporal prior by recognizing the substantial overlap between temporally adjacent LiDAR scans. We propose a proximity-based label estimation, which generates highly accurate pseudo-labels for unlabeled data by utilizing semantic consistency with adjacent labeled data. Additionally, we enhance this method by progressively expanding the pseudo-labels from the nearest unlabeled scans, which helps significantly reduce errors linked to dynamic classes. Additionally, we employ a dual-branch structure to mitigate performance degradation caused by data imbalance. Experimental results demonstrate remarkable performance in low-budget settings (i.e., ≤ 5%) and meaningful improvements in normal budget settings (i.e., 5 – 50%). Finally, our method has achieved new state-of-the-art results on SemanticKITTI and nuScenes in semi-supervised LiDAR segmentation. With only 5% labeled data, it offers competitive results against fully-supervised counterparts. Moreover, it surpasses the performance of the previous state-of-the-art at 100% labeled data (75.2%) using only 20% of labeled data (76.0%) on nuScenes. The code is available on https://github.com/halbielee/PLE.
Hwijeong Lee, Hyunjung Shim
IROS3
2024 Understanding Multi-Granularity for Open-Vocabulary Part Segmentation
abstract
Open-vocabulary part segmentation (OVPS) is an emerging research area focused on segmenting fine-grained entities using diverse and previously unseen vocabularies. Our study highlights the inherent complexities of part segmentation due to intricate boundaries and diverse granularity, reflecting the knowledge-based nature of part identification. To address these challenges, we propose PartCLIPSeg, a novel framework utilizing generalized parts and object-level contexts to mitigate the lack of generalization in fine-grained parts. PartCLIPSeg integrates competitive part relationships and attention control, alleviating ambiguous boundaries and underrepresented parts. Experimental results demonstrate that PartCLIPSeg outperforms existing state-of-the-art OVPS methods, offering refined segmentation and an advanced understanding of part relationships within images. Through extensive experiments, our model demonstrated a significant improvement over the state-of-the-art models on the Pascal-Part-116, ADE20K-Part-234, and PartImageNet datasets.
Jiho Choi, Seonho Lee, Minhyun Lee, Hyunjung Shim
NeurIPS5
2024 Few-Shot Font Generation With Weakly Supervised Localized Representations
abstract
Automatic few-shot font generation aims to solve a well-defined, real-world problem because manual font designs are expensive and sensitive to the expertise of designers. Existing methods learn to disentangle style and content elements by developing a universal style representation for each font style. However, this approach limits the model in representing diverse local styles because it is unsuitable for complicated letter systems. For example, Chinese characters consist of a varying number of components (often called "radical") with a highly complex structure. In this paper, we propose a novel font generation method that learns localized styles, namely component-wise style representations, instead of universal styles. The proposed style representations enable synthesizing complex local details in text designs. However, learning component-wise styles solely from a few reference glyphs is infeasible when a target script has a large number of components, for example, over 200 for Chinese. To reduce the number of required reference glyphs, we represent component-wise styles by a product of component and style factors inspired by low-rank matrix factorization. Owing to the combination of strong representation and a compact factorization strategy, our method shows remarkably better few-shot font generation results (with only eight reference glyphs) than other state-of-the-art methods. Moreover, strong locality supervision was not utilized, such as the location of each component, skeleton, or strokes. The source code is available at https://github.com/clovaai/lffont.
Song Park, Sanghyuk Chun, Junbum Cha, Bado Lee, Hyunjung Shim
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Your lottery ticket is damaged: Towards all-alive pruning for extremely sparse networks
Min-Soo Kim 0002, Hyunjung Shim, Jongwuk Lee
Inf. Sci.3
2023 CoMix: Collaborative filtering with mixup for implicit datasets
Jaewan Moon, Yoonki Jeong, Dong-Kyu Chae, Hyunjung Shim, Jongwuk Lee
Inf. Sci.5
2023 Evaluation for Weakly Supervised Object Localization: Protocol, Metrics, and Datasets
abstract
Weakly-supervised object localization (WSOL) has gained popularity over the last years for its promise to train localization models with only image-level labels. Since the seminal WSOL work of class activation mapping (CAM), the field has focused on how to expand the attention regions to cover objects more broadly and localize them better. However, these strategies rely on full localization supervision for validating hyperparameters and model selection, which is in principle prohibited under the WSOL setup. In this paper, we argue that WSOL task is ill-posed with only image-level labels, and propose a new evaluation protocol where full supervision is limited to only a small held-out set not overlapping with the test set. We observe that, under our protocol, the five most recent WSOL methods have not made a major improvement over the CAM baseline. Moreover, we report that existing WSOL methods have not reached the few-shot learning baseline, where the full-supervision at validation time is used for model training instead. Based on our findings, we discuss some future directions for WSOL. Source code and dataset are available at https://github.com/clovaai/wsolevaluation https://github.com/clovaai/wsolevaluation.
Junsuk Choe, Seong Joon Oh, Sanghyuk Chun, Zeynep Akata, Hyunjung Shim
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Saliency as Pseudo-Pixel Supervision for Weakly and Semi-Supervised Semantic Segmentation
abstract
Existing studies on semantic segmentation using image-level weak supervision have several limitations, including sparse object coverage, inaccurate object boundaries, and co-occurring pixels from non-target objects. To overcome these challenges, we propose a novel framework, an improved version of Explicit Pseudo-pixel Supervision (EPS++), which learns from pixel-level feedback by combining two types of weak supervision. Specifically, the image-level label provides the object identity via the localization map, and the saliency map from an off-the-shelf saliency detection model offers rich object boundaries. We devise a joint training strategy to fully utilize the complementary relationship between disparate information. Notably, we suggest an Inconsistent Region Drop (IRD) strategy, which effectively handles errors in saliency maps using fewer hyper-parameters than EPS. Our method can obtain accurate object boundaries and discard co-occurring pixels, significantly improving the quality of pseudo-masks. Experimental results show that EPS++ effectively resolves the key challenges of semantic segmentation using weak supervision, resulting in new state-of-the-art performances on three benchmark datasets in a weakly supervised semantic segmentation setting. Furthermore, we show that the proposed method can be extended to solve the semi-supervised semantic segmentation problem using image-level weak supervision. Surprisingly, the proposed model also achieves new state-of-the-art performances on two popular benchmark datasets.
Minhyun Lee, Jongwuk Lee, Hyunjung Shim
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Entropy regularization for weakly supervised object localization
Dongjun Hwang, Jung-Woo Ha 0001, Hyunjung Shim, Junsuk Choe
Pattern Recognit. Lett.3
2022 Long-tail Mixup for Extreme Multi-label Classification
abstract
Extreme multi-label classification (XMC) aims at finding multiple relevant labels for a given sample from a huge label set at the industrial scale. The XMC problem inherently poses two challenges: scalability and label sparsity - the number of labels is too large, and labels follow the long-tail distribution. To resolve these problems, we propose a novel Mixup-based augmentation method for long-tail labels, called TailMix. Building upon the partition-based model, TailMix utilizes the context vectors generated from the label attention layer. It first selectively chooses two context vectors using the inverse propensity score of labels and the label proximity graph representing the co-occurrence of labels. Using two context vectors, it augments new samples with the long-tail label to improve the accuracy of long-tail labels. Despite its simplicity, experimental results show that TailMix consistently outperforms other augmentation methods on three benchmark datasets, especially for long-tail labels in terms of two metrics, [email protected] and [email protected]
Sangwoo Han, Eunseong Choi, Chan Lim, Hyunjung Shim, Jongwuk Lee
CIKM4
2022 Commonality in Natural Images Rescues GANs: Pretraining GANs with Generic and Privacy-free Synthetic Data
abstract
Transfer learning for GANs successfully improves generation performance under low-shot regimes. However, existing studies show that the pretrained model using a single benchmark dataset is not generalized to various target datasets. More importantly, the pretrained model can be vulnerable to copyright or privacy risks as membership inference attack advances. To resolve both issues, we propose an effective and unbiased data synthesizer, namely Primitives - PS, inspired by the generic characteristics of natural images. Specifically, we utilize 1) the generic statistics on the frequency magnitude spectrum, 2) the elementary shape (i.e., image composition via elementary shapes) for representing the structure information, and 3) the existence of saliency as prior. Since our synthesizer only considers the generic properties of natural images, the single model pretrained on our dataset can be consistently transferred to various target datasets, and even outperforms the previous methods pretrained with the natural images in terms of Fréchet inception distance. Extensive analysis, ablation study, and evaluations demonstrate that each component of our data synthesizer is effective, and provide insights on the desirable nature of the pretrained model for the transferability of GANs.
Kyungjune Baek, Hyunjung Shim
CVPR2
2022 Threshold Matters in WSSS: Manipulating the Activation for the Robust and Accurate Segmentation Model Against Thresholds
abstract
Weakly-supervised semantic segmentation (WSSS) has recently gained much attention for its promise to train segmentation models only with image-level labels. Existing WSSS methods commonly argue that the sparse coverage of CAM incurs the performance bottleneck of WSSS. This paper provides analytical and empirical evidence that the actual bottleneck may not be sparse coverage but a global thresholding scheme applied after CAM. Then, we show that this issue can be mitigated by satisfying two conditions; 1) reducing the imbalance in the foreground activation and 2) increasing the gap between the foreground and the background activation. Based on these findings, we propose a novel activation manipulation network with a per-pixel classification loss and a label conditioning module. Per-pixel classification naturally induces two-level activation in activation maps, which can penalize the most discriminative parts, promote the less discriminative parts, and deactivate the background regions. Label conditioning imposes that the output label of pseudo-masks should be any of true image-level labels; it penalizes the wrong activation assigned to non-target classes. Based on extensive analysis and evaluations, we demonstrate that each component helps produce accurate pseudo-masks, achieving the robustness against the choice of the global threshold. Finally, our model achieves state-of-the-art records on both PAS-CAL VOC 2012 and MS COCO 2014 datasets. The code is available at https://github.com/gaviotas/AMN.
Minhyun Lee, Dongseob Kim, Hyunjung Shim
CVPR3
2022 Learning from Better Supervision: Self-distillation for Learning with Noisy Labels
abstract
The remarkable performance of deep neural networks heavily rely on large-scale datasets with high-quality annotations. Since the data collection process such as web crawling naturally involves unreliable supervision (i.e., noisy label), handling samples with noisy labels has been actively studied. Existing methods in learning with noisy labels (LNL) 1) develop the sampling strategy for filtering out the noisy labels or 2) devise the robust loss function against noisy labels. As a result of these efforts, existing LNL models achieve impressive performance, recording a higher accuracy than the ratio of the clean samples in the dataset. Based on this observation, we propose a self-distillation framework to utilize the prediction of existing LNL models and further improve the performance via rectified distillation; hard pseudo label and feature distillation. Our rectified distillation can be easily applied to existing LNL models, thus we can enjoy their state-of-the-art performances. From extensive evaluations, we confirm that our model is effective on both synthetic and real noisy datasets with state-of-the-art performances on four benchmark datasets.
Kyungjune Baek, Hyunjung Shim
ICPR3
2022 Logit Mixing Training for More Reliable and Accurate Prediction
abstract
When a person solves the multi-choice problem, she considers not only what is the answer but also what is not the answer. Knowing what choice is not the answer and utilizing the relationships between choices, she can improve the prediction accuracy. Inspired by this human reasoning process, we propose a new training strategy to fully utilize inter-class relationships, namely LogitMix. Our strategy is combined with recent data augmentation techniques, e.g., Mixup, Manifold Mixup, CutMix, and PuzzleMix. Then, we suggest using a mixed logit, i.e., a mixture of two logits, as an auxiliary training objective. Since the logit can preserve both positive and negative inter-class relationships, it can impose a network to learn the probability of wrong answers correctly. Our extensive experimental results on the image- and language-based tasks demonstrate that LogitMix achieves state-of-the-art performance among recent data augmentation techniques regarding calibration error and prediction accuracy.
Duhyeon Bang, Kyungjune Baek, Yunho Jeon, Jin-Hwa Kim, Jongwuk Lee, Hyunjung Shim
IJCAI8
2022 S-Walk: Accurate and Scalable Session-based Recommendation with Random Walks
abstract
Session-based recommendation (SR) predicts the next items from a sequence of previous items consumed by an anonymous user. Most existing SR models focus only on modeling intra-session characteristics but pay less attention to inter-session relationships of items, which has the potential to improve accuracy. Another critical aspect of recommender systems is computational efficiency and scalability, considering practical feasibility in commercial applications. To account for both accuracy and scalability, we propose a novel session-based recommendation with a random walk, namely S-Walk. Precisely, S-Walk effectively captures intra- and inter-session correlations by handling high-order relationships among items using random walks with restart (RWR). By adopting linear models with closed-form solutions for transition and teleportation matrices that constitute RWR, S-Walk is highly efficient and scalable. Extensive experiments demonstrate that S-Walk achieves comparable or state-of-the-art performance in various metrics on four benchmark datasets. Moreover, the model learned by S-Walk can be highly compressed without sacrificing accuracy, conducting two or more orders of magnitude faster inference than existing DNN-based models, making it suitable for large-scale commercial systems.
Minjin Choi 0001, Jinhong Kim, Joonseok Lee, Hyunjung Shim, Jongwuk Lee
WSDM4
2022 Knowledge distillation meets recommendation: collaborative distillation for top-N recommendation
Jae-woong Lee, Minjin Choi 0001, Lee Sael, Hyunjung Shim, Jongwuk Lee
Knowl. Inf. Syst.4
2021 Few-shot Font Generation with Localized Style Representations and Factorization
abstract
Automatic few-shot font generation is a practical and widely studied problem because manual designs are expensive and sensitive to the expertise of designers. Existing few-shot font generation methods aim to learn to disentangle the style and content element from a few reference glyphs, and mainly focus on a universal style representation for each font style. However, such approach limits the model in representing diverse local styles, and thus makes it unsuitable to the most complicated letter system, e.g., Chinese, whose characters consist of a varying number of components (often called ``radical'') with a highly complex structure. In this paper, we propose a novel font generation method by learning localized styles, namely component-wise style representations, instead of universal styles. The proposed style representations enable us to synthesize complex local details in text designs. However, learning component-wise styles solely from reference glyphs is infeasible in the few-shot font generation scenario, when a target script has a large number of components, e.g., over 200 for Chinese. To reduce the number of reference glyphs, we simplify component-wise styles by a product of component factor and style factor, inspired by low-rank matrix factorization. Thanks to the combination of strong representation and a compact factorization strategy, our method shows remarkably better few-shot font generation results (with only 8 reference glyph images) than other state-of-the-arts, without utilizing strong locality supervision, e.g., location of each component, skeleton, or strokes. The source code is available at https://github.com/clovaai/lffont.
Song Park, Sanghyuk Chun, Junbum Cha, Bado Lee, Hyunjung Shim
AAAI5
2021 Railroad Is Not a Train: Saliency As Pseudo-Pixel Supervision for Weakly Supervised Semantic Segmentation
abstract
Existing studies in weakly-supervised semantic segmentation (WSSS) using image-level weak supervision have several limitations: sparse object coverage, inaccurate object boundaries, and co-occurring pixels from non-target objects. To overcome these challenges, we propose a novel framework, namely Explicit Pseudo-pixel Supervision (EPS), which learns from pixel-level feedback by combining two weak supervisions; the image-level label provides the object identity via the localization map and the saliency map from the off-the-shelf saliency detection model offers rich boundaries. We devise a joint training strategy to fully utilize the complementary relationship between both information. Our method can obtain accurate object boundaries and discard co-occurring pixels, thereby significantly improving the quality of pseudo-masks. Experimental results show that the proposed method remarkably outperforms existing methods by resolving key challenges of WSSS and achieves the new state-of-the-art performance on both PASCAL VOC 2012 and MS COCO 2014 datasets. The code is available at https://github.com/halbielee/EPS.
Minhyun Lee, Jongwuk Lee, Hyunjung Shim
CVPR4
2021 Rethinking the Truly Unsupervised Image-to-Image Translation
abstract
Every recent image-to-image translation model inherently requires either image-level (i.e. input-output pairs) or set-level (i.e. domain labels) supervision. However, even set-level supervision can be a severe bottleneck for data collection in practice. In this paper, we tackle image-to-image translation in a fully unsupervised setting, i.e., neither paired images nor domain labels. To this end, we propose a truly unsupervised image-to-image translation model (TUNIT) that simultaneously learns to separate image domains and translates input images into the estimated domains. Experimental results show that our model achieves comparable or even better performance than the set-level supervised model trained with full labels, generalizes well on various datasets, and is robust against the choice of hyperparameters (e.g. the preset number of pseudo domains). Furthermore, TUNIT can be easily extended to semi-supervised learning with a few labeled data.
Kyungjune Baek, Yunjey Choi, Youngjung Uh, Jaejun Yoo 0001, Hyunjung Shim
ICCV5
2021 Multiple Heads are Better than One: Few-shot Font Generation with Multiple Localized Experts
abstract
A few-shot font generation (FFG) method has to satisfy two objectives: the generated images should preserve the underlying global structure of the target character and present the diverse local reference style. Existing FFG methods aim to disentangle content and style either by extracting a universal representation style or extracting multiple component-wise style representations. However, previous methods either fail to capture diverse local styles or cannot be generalized to a character with unseen components, e.g., unseen language systems. To mitigate the issues, we propose a novel FFG method, named Multiple Localized Experts Few-shot Font Generation Network (MX-Font). MX-Font extracts multiple style features not explicitly conditioned on component labels, but automatically by multiple experts to represent different local concepts, e.g., left-side sub-glyph. Owing to the multiple experts, MX-Font can capture diverse local concepts and show the generalizability to unseen languages. During training, we utilize component labels as weak supervision to guide each expert to be specialized for different local concepts. We formulate the component assign problem to each expert as the graph matching problem, and solve it by the Hungarian algorithm. We also employ the independence loss and the content-style adversarial loss to impose the content-style disentanglement. In our experiments, MX-Font outperforms previous state-of-the-art FFG methods in the Chinese generation and cross-lingual, e.g., Chinese to Korean, generation. Source code is available at https://github.com/clovaai/mxfont.
Song Park, Sanghyuk Chun, Junbum Cha, Bado Lee, Hyunjung Shim
ICCV5
2021 Session-aware Linear Item-Item Models for Session-based Recommendation
abstract
Session-based recommendation aims at predicting the next item given a sequence of previous items consumed in the session, e.g., on e-commerce or multimedia streaming services. Specifically, session data exhibits some unique characteristics, i.e., session consistency and sequential dependency over items within the session, repeated item consumption, and session timeliness. In this paper, we propose simple-yet-effective linear models for considering the holistic aspects of the sessions. The comprehensive nature of our models helps improve the quality of session-based recommendation. More importantly, it provides a generalized framework for reflecting different perspectives of session data. Furthermore, since our models can be solved by closed-form solutions, they are highly scalable. Experimental results demonstrate that the proposed linear models show competitive or state-of-the-art performance in various metrics on several real-world datasets.
Minjin Choi 0001, Jinhong Kim, Joonseok Lee, Hyunjung Shim, Jongwuk Lee
WWW4
2021 Distilling from professors: Enhancing the knowledge distillation of teachers
Duhyeon Bang, Jongwuk Lee, Hyunjung Shim
Inf. Sci.3
2021 Weakly-supervised progressive denoising with unpaired CT images
Byeongjoon Kim, Hyunjung Shim, Jongduk Baek
Medical Image Anal.2
2021 Rigid and non-rigid motion artifact reduction in X-ray CT using attention module
Youngjun Ko, Seunghyuk Moon, Jongduk Baek, Hyunjung Shim
Medical Image Anal.4
2021 Attention-Based Dropout Layer for Weakly Supervised Single Object Localization and Semantic Segmentation
abstract
Both weakly supervised single object localization and semantic segmentation techniques learn an object's location using only image-level labels. However, these techniques are limited to cover only the most discriminative part of the object and not the entire object. To address this problem, we propose an attention-based dropout layer, which utilizes the attention mechanism to locate the entire object efficiently. To achieve this, we devise two key components, 1) hiding the most discriminative part from the model to capture the entire object, and 2) highlighting the informative region to improve the classification power of the model. These allow the classifier to be maintained with a reasonable accuracy while the entire object is covered. Through extensive experiments, we demonstrate that the proposed method effectively improves the weakly supervised single object localization accuracy, thereby achieving a new state-of-the-art localization accuracy on the CUB-200-2011 and a comparable accuracy existing state-of-the-arts on the ImageNet-1k. The proposed method is also effective in improving the weakly supervised semantic segmentation performance on the Pascal VOC and MS COCO. Furthermore, the proposed method is more efficient than existing techniques in terms of parameter and computation overheads. Additionally, the proposed method can be easily applied in various backbone networks.
Junsuk Choe, Hyunjung Shim
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 GridMix: Strong regularization through local context mapping
Kyungjune Baek, Duhyeon Bang, Hyunjung Shim
Pattern Recognit.3
2021 Region-based dropout with attention prior for weakly supervised object localization
Junsuk Choe, Dongyoon Han, Sangdoo Yun, Jung-Woo Ha 0001, Seong Joon Oh, Hyunjung Shim
Pattern Recognit.6
2020 PsyNet: Self-Supervised Approach to Object Localization Using Point Symmetric Transformation
abstract
Existing co-localization techniques significantly lose performance over weakly or fully supervised methods in accuracy and inference time. In this paper, we overcome common drawbacks of co-localization techniques by utilizing self-supervised learning approach. The major technical contributions of the proposed method are two-fold. 1) We devise a new geometric transformation, namely point symmetric transformation and utilize its parameters as an artificial label for self-supervised learning. This new transformation can also play the role of region-drop based regularization. 2) We suggest a heat map extraction method for computing the heat map from the network trained by self-supervision, namely class-agnostic activation mapping. It is done by computing the spatial attention map. Based on extensive evaluations, we observe that the proposed method records new state-of-the-art performance in three fine-grained datasets for unsupervised object localization. Moreover, we show that the idea of the proposed method can be adopted in a modified manner to solve the weakly supervised object localization task. As a result, we outperform the current state-of-the-art technique in weakly supervised object localization by a significant gap.
Kyungjune Baek, Minhyun Lee, Hyunjung Shim
AAAI3
2020 Evaluating Weakly Supervised Object Localization Methods Right
abstract
Weakly-supervised object localization (WSOL) has gained popularity over the last years for its promise to train localization models with only image-level labels. Since the seminal WSOL work of class activation mapping (CAM), the field has focused on how to expand the attention regions to cover objects more broadly and localize them better. However, these strategies rely on full localization supervision to validate hyperparameters and for model selection, which is in principle prohibited under the WSOL setup. In this paper, we argue that WSOL task is ill-posed with only image-level labels, and propose a new evaluation protocol where full supervision is limited to only a small held-out set not overlapping with the test set. We observe that, under our protocol, the five most recent WSOL methods have not made a major improvement over the CAM baseline. Moreover, we report that existing WSOL methods have not reached the few-shot learning baseline, where the full-supervision at validation time is used for model training instead. Based on our findings, we discuss some future directions for WSOL.
Junsuk Choe, Seong Joon Oh, Sanghyuk Chun, Zeynep Akata, Hyunjung Shim
CVPR6
2020 Discriminator Feature-Based Inference by Recycling the Discriminator of GANs
Duhyeon Bang, Seoungyoon Kang, Hyunjung Shim
Int. J. Comput. Vis.3
2019 Attention-Based Dropout Layer for Weakly Supervised Object Localization
abstract
Weakly Supervised Object Localization (WSOL) techniques learn the object location only using image-level labels, without location annotations. A common limitation for these techniques is that they cover only the most discriminative part of the object, not the entire object. To address this problem, we propose an Attention-based Dropout Layer (ADL), which utilizes the self-attention mechanism to process the feature maps of the model. The proposed method is composed of two key components: 1) hiding the most discriminative part from the model for capturing the integral extent of object, and 2) highlighting the informative region for improving the recognition power of the model. Based on extensive experiments, we demonstrate that the proposed method is effective to improve the accuracy of WSOL, achieving a new state-of-the-art localization accuracy in CUB-200-2011 dataset. We also show that the proposed method is much more efficient in terms of both parameter and computation overheads than existing techniques.
Junsuk Choe, Hyunjung Shim
CVPR2
2019 Collaborative Distillation for Top-N Recommendation
abstract
Knowledge distillation (KD) is a well-known method to reduce inference latency by compressing a cumbersome teacher model to a small student model. Despite the success of KD in the classification task, applying KD to recommender models is challenging due to the sparsity of positive feedback, the ambiguity of missing feedback, and the ranking problem associated with the top-N recommendation. To address the issues, we propose a new KD model for the collaborative filtering approach, namely collaborative distillation (CD). Specifically, (1) we reformulate a loss function to deal with the ambiguity of missing feedback. (2) We exploit probabilistic rank-aware sampling for the top-N recommendation. (3) To train the proposed model effectively, we develop two training strategies for the student model, called the teacher-and the student-guided training methods, selecting the most useful feedback from the teacher model. Via experimental results, we demonstrate that the proposed model outperforms the state-of-the-art method by 5.5-29.7% and 4.8-27.8% in hit rate (HR) and normalized discounted cumulative gain (NDCG), respectively. Moreover, the proposed model achieves the performance comparable to the teacher model.
Jae-woong Lee, Minjin Choi 0001, Jongwuk Lee, Hyunjung Shim
ICDM4
2019 Dual Neural Personalized Ranking
abstract
Implicit user feedback is a fundamental dataset for personalized recommendation models. Because of its inherent characteristics of sparse one-class values, it is challenging to uncover meaningful user/item representations. In this paper, we propose dual neural personalized ranking (DualNPR), which fully exploits both user- and item-side pairwise rankings in a unified manner. The key novelties of the proposed model are three-fold: (1) DualNPR discovers mutual correlation among users and items by utilizing both user- and item-side pairwise rankings, alleviating the data sparsity problem. We stress that, unlike existing models that require extra information, DualNPR naturally augments both user- and item-side pairwise rankings from a user-item interaction matrix. (2) DualNPR is built upon deep matrix factorization to capture the variability of user/item representations. In particular, it chooses raw user/item vectors as an input and learns latent user/item representations effectively. (3) DualNPR employs a dynamic negative sampling method using an exponential function, further improving the accuracy of top-N recommendation. In experimental results over three benchmark datasets, DualNPR outperforms baseline models by 21.9-86.7% in hit rate, 14.5-105.8% in normalized discounted cumulative gain, and 5.1-23.3% in the area under the ROC curve.
Seunghyeon Kim, Jongwuk Lee, Hyunjung Shim
WWW3
2019 Semantic-aware neural style transfer
Joo-Hyun Park, Song Park, Hyunjung Shim
Image Vis. Comput.3
2018 Editable Generative Adversarial Networks: Generating and Editing Faces Simultaneously
Kyungjune Baek, Duhyeon Bang, Hyunjung Shim
ACCV (1)3
2018 Depth Reconstruction of Translucent Objects from a Single Time-of-Flight Camera Using Deep Residual Networks
Seongjong Song, Hyunjung Shim
ACCV (5)2
2018 Resembled Generative Adversarial Networks: Two Domains with Similar Attributes
Duhyeon Bang, Hyunjung Shim
BMVC2
2018 Improved Training of Generative Adversarial Networks using Representative Features
abstract
Despite the success of generative adversarial networks (GANs) for image generation, the trade-off between visual quality and image diversity remains a significant issue. This paper achieves both aims simultaneously by improving the stability of training GANs. The key idea of the proposed approach is to implicitly regularize the discriminator using representative features. Focusing on the fact that standard GAN minimizes reverse Kullback-Leibler (KL) divergence, we transfer the representative feature, which is extracted from the data distribution using a pre-trained autoencoder (AE), to the discriminator of standard GANs. Because the AE learns to minimize forward KL divergence, our GAN training with representative features is influenced by both reverse and forward KL divergence. Consequently, the proposed approach is verified to improve visual quality and diversity of state of the art GANs using extensive evaluations.
Duhyeon Bang, Hyunjung Shim
ICML2
2018 Session details: Interactive Art
Hyunjung Shim
ACM Multimedia1
2018 Robust approach to inverse lighting using RGB-D images
Junsuk Choe, Hyunjung Shim
Inf. Sci.2
2016 Recovering Translucent Objects Using a Single Time-of-Flight Depth Camera
abstract
Translucency introduces great challenges to 3-D acquisition because of complicated light behaviors such as refraction and transmittance. In this paper, we describe the development of a unified 3-D data acquisition framework that reconstructs translucent objects using a single commercial time-of-flight (ToF) camera. In our capture scenario, we record a depth map and intensity image of the scene twice using a static ToF camera; first, we capture the depth map and intensity image of an arbitrary background, and then we position the translucent foreground object and record a second depth map and intensity image with both the foreground and the background. As a result of material characteristics, the translucent object yields systematic distortions in the depth map. We developed a new distance representation that interprets the depth distortion induced as a result of translucency. By analyzing ToF depth sensing principles, we constructed a distance model governed by the level of translucency, foreground depth, and background depth. Using an analysis-by-synthesis approach, we can recover the 3-D geometry of a translucent object from a pair of depth maps and their intensity images. Extensive evaluation and case studies demonstrate that our method is effective for modeling the nonlinear depth distortion due to translucency and for reconstruction of a 3-D translucent object.
Hyunjung Shim, Seungkyu Lee 0001
IEEE Trans. Circuits Syst. Video Technol.1
2015 Skewed stereo time-of-flight camera for translucent object imaging
Seungkyu Lee 0001, Hyunjung Shim
Image Vis. Comput.2
2012 Light weight multiview capturing
abstract
The framework of capturing multiview images has received a great attention from industry as well as academy due to its usability in various applications; 3D scene reconstruction, super-resolution, scene understanding, surveillance etc. This paper presents a multiview capturing system that uses a single camera with multiple mirrors. We place multiple mirrors pointing to the target scene and record an image of those mirrors. From this image of mirrors, we can extract the multiview images of the target scene, images reflected by mirrors. Based on the proposed method, we achieve some success in geometry reconstruction and view synthesis using multiview images, which shows the usability of the proposed work. The main contribution of this work is to introduce a light weight implementation for multiview capturing with a number of useful properties; the automatic synchronization across multiple views, the consistent color/white balance setting across multiple views, the capability of recording a wide range of view and the cost-effective implementation.
Hyunjung Shim, Tsuhan Chen, Seungkyu Lee 0001, James Dokyoon Kim, Chang-Yeong Kim
CCNC1
2012 Estimating all frequency lighting using a color/depth image
abstract
This paper presents a novel approach to estimating the lighting using a pair of one color and one depth image. To effectively model all frequency lighting, we introduce a hybrid representation for lighting; the combination of spherical harmonic basis functions and point lights. Upon the existing framework of spherical harmonics based diffuse reflection, we divide the color image into diffuse and non-diffuse reflections. Then, we use the diffuse reflection for estimating the low frequency lighting. For high frequency lighting, we obtain the specular reflections by analyzing the non-diffuse reflection. Knowing specular reflections and scene geometry, we are capable of computing the direction of point lights, inverting the reflected ray with respect to the surface normal. Then, we optimize the intensity of point lights by analysis by synthesis paradigm. By superimposing the low and high frequency lighting, we recover the lighting present in the scene. While existing methods use the low frequency lighting to infer the high frequency lighting, we propose to use the nondiffuse reflection for directly estimating the high frequency lighting. In this way, we make good use of the non-diffuse reflections in scene analysis and understanding. Experimental results show that the proposed approach is an effective solution for the lighting estimation of real world environment.
Hyunjung Shim
ICIP1
2012 Automatic color realism enhancement for virtual reality
abstract
Photorealism has been one of essential goals for virtual reality. The state-of-the-art techniques employ various rendering algorithms to simulate physically accurate light transport for generating the photorealistic appearance of scene. However, they still require a labor-intensive tone mapping and color tunes by an experienced artist. In this paper, we propose an automatic photorealism enhancement algorithm by manipulating the color distribution of graphics to match with that of real photographs. Based on the hypothesis that photorealism is highly correlated with the frequency of color characteristics appearing in real photographs, we find principal color components. Then, we transfer the statistical characteristics of photographs onto graphics so to enhance their photorealism. Experiments and a user study have confirmed the effectiveness of proposed method.
Hyunjung Shim, Seungkyu Lee 0001
VR1
2012 Automatic color realism enhancement for computer generated images
Hyunjung Shim, Seungkyu Lee 0001
Comput. Graph.1
2012 Probabilistic Approach to Realistic Face Synthesis With a Single Uncalibrated Image
abstract
This paper presents a novel approach to automatic face modeling for realistic synthesis from an unknown face image, using a probabilistic face diffuse model and a generic face specular map. We construct a probabilistic face diffuse model for estimating the albedo and normals of the input face. Then, we develop a generic face specular map for estimating the specularity of face. Using the estimated albedo, normal and specular information, we can synthesize the face under arbitrary lighting and viewing directions realistically. Unlike many existing techniques, our approach can extract both the diffuse and specular information of face without involving an intensive 3D matching procedure. We conduct three different experiments to show our improvement over the prior art. First, we compare the proposed algorithm with previous techniques, including the state of the art, to demonstrate our achievement in realistic face synthesis. Moreover, we evaluate the proposed algorithm over non-automatic face modeling techniques through a subjective study. This evaluation is meaningful in that it tells us how far the proposed algorithm as well as others are from the real photograph in terms of the perceptual quality. Finally, we apply our face model for improving the face recognition performance under varying illumination conditions and show that the proposed algorithm is effective to enhance the face recognition rate. Thanks to the compact representation and the effective inference scheme, our technique is applicable for many practical applications, such as avatar creation, digital face cloning, face normalization, deidentification and many others.
Hyunjung Shim
IEEE Trans. Image Process.1
2012 Time-of-flight sensor and color camera calibration for multi-view acquisition
Hyunjung Shim, Rolf Adelsberger, James Dokyoon Kim, Seon-Min Rhee, Taehyun Rhee, Jae-Young Sim, Markus Gross 0001, Chang-Yeong Kim
Vis. Comput.1
2010 A probabilistic approach to realistic face synthesis
abstract
This paper presents a novel approach to face modeling for realistic synthesis, powered by a probabilistic face diffuse model and a generic face specular map. We first construct a probabilistic face diffuse model for estimating the albedo and the normals of a face from an unknown input image. Then, we introduce a generic face specular map for estimating the specularity of the face. Using the estimated albedo, normal and specular information, we can synthesize the face under arbitrary lighting and viewing directions realistically. Unlike many existing face modeling techniques, our approach can retain both the diffuse and specular properties of the face without involving an elaborating 3D matching procedure. Thanks to the compact representation and the effective inference scheme, our technique can be applied to many practical applications, such as face normalization, avatar creation and de-identification.
Hyunjung Shim, Inwoo Ha, Taehyun Rhee, James Dokyoon Kim, Chang-Yeong Kim
ICIP1
2008 A Subspace Model-Based Approach to Face Relighting Under Unknown Lighting and Poses
abstract
We present a new approach to face relighting by jointly estimating the pose, reflectance functions, and lighting from as few as one image of a face. Upon such estimation, we can synthesize the face image under any prescribed new lighting condition. In contrast to commonly used face shape models or shape-dependent models, we neither recover nor assume the 3-D face shape during the estimation process. Instead, we train a pose- and pixel-dependent subspace model of the reflectance function using a face database that contains samples of pose and illumination for a large number of individuals (e.g., the CMU PIE database and the Yale database). Using this subspace model, we can estimate the pose, the reflectance functions, and the lighting condition of any given face image. Our approach lends itself to practical applications thanks to many desirable properties, including the preservation of the non-Lambertian skin reflectance properties and facial hair, as well as reproduction of various shadows on the face. Extensive experiments show that, compared to recent representative face relighting techniques, our method successfully produces better results, in terms of subjective and objective quality, without reconstructing a 3-D shape.
Hyunjung Shim, Jiebo Luo 0001, Tsuhan Chen
IEEE Trans. Image Process.1
2005 A statistical framework for image-based relighting
abstract
With image-based relighting (IBL), one can render realistic relit images of a scene without prior knowledge of object geometry in the scene. However, traditional IBL methods require a large number of basis images, each corresponding to a lighting pattern, to estimate the surface reflectance function (SRF) of the scene. We present a statistical approach to estimating the SRF which requires fewer basis images. We formulate the SRF estimation problem in a signal reconstruction framework. We use principal component analysis (PCA) (Duda, R.O. et al., "Pattern Classification, 2nd edition", p.115-17, Wiley Interscience, 2000) to show that the most effective lighting patterns for the data acquisition process are the SRF covariance matrix eigenvectors corresponding to the largest eigenvalues. In addition, we show that, for typical SRFs, especially when the objects have Lambertian surfaces, DCT-based lighting patterns perform as well as the optimal PCA-based lighting patterns. We compare SRF estimation performance of the statistical approach with traditional IBL techniques. Experimental results show that the statistical approach can achieve better performance with fewer basis images.
Hyunjung Shim, Tsuhan Chen
ICASSP (2)1