EDBT 2026 Demo / reviewers in the wild / expert
Youngjung Uh
dblp:57/10511
· DBLP profile ↗
42ranked-venue papers
2as first author
29since 2021 · last 2026
0000-0001-8173-3334ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 19 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 4D Scaffold Gaussian Splatting with Dynamic-Aware Anchor Growing for Efficient and High-Fidelity Dynamic Scene ReconstructionabstractModeling dynamic scenes through 4D Gaussians offers high visual fidelity and fast rendering speeds, but comes with significant storage overhead. Recent approaches mitigate this cost by aggressively reducing the number of Gaussians. However, this inevitably removes Gaussians essential for high-quality rendering, leading to severe degradation in dynamic regions. In this paper, we introduce a novel 4D anchor-based framework that tackles the storage cost in different perspective. Rather than reducing the number of Gaussians, our method retains a sufficient quantity to accurately model dynamic contents, while compressing them into compact, grid-aligned 4D anchor features. Each anchor is processed by an MLP to spawn a set of neural 4D Gaussians, which represent a local spatiotemporal region. We design these neural 4D Gaussians to capture temporal changes with minimal parameters, making them well-suited for the MLP-based spawning. Moreover, we introduce a dynamic-aware anchor growing strategy to effectively assign additional anchors to under-reconstructed dynamic regions. Our method adjusts the accumulated gradients with Gaussians' temporal coverage, significantly improving reconstruction quality in dynamic regions. Experimental results highlight that our method achieves state-of-the-art visual quality in dynamic regions, outperforming all baselines by a large margin with practical storage costs. Woong Oh Cho, In Cho, Seoha Kim, Jeongmin Bae 0001, Youngjung Uh, Seon Joo Kim |
AAAI | 5 |
| 2026 | Eye-for-an-eye: Appearance Transfer with Dense Semantic Correspondence in Diffusion ModelsabstractAs pre-trained text-to-image diffusion models have become a useful tool for image synthesis, people want to specify the results in various ways. This paper tackles training-free appearance transfer, which produces an image with the structure of a target image from the appearance of a reference image. Existing methods usually do not reflect semantic correspondence, as they rely on query-key similarity within the self-attention layer to establish correspondences between images. To this end, we propose explicitly rearranging the features according to the dense semantic correspondences. Extensive experiments show the superiority of our method in various aspects: preserving the structure of the target and reflecting the correct color from the reference, even when the two images are not aligned. Sooyeon Go, Kyungmook Choi, Minjung Shin, Youngjung Uh |
WACV | 4 |
| 2025 | Rethinking Open-Vocabulary Segmentation of Radiance Fields in 3D SpaceabstractUnderstanding the 3D semantics of a scene is a fundamental problem for various scenarios such as embodied agents. While NeRFs and 3DGS excel at novel-view synthesis, previous methods for understanding their semantics have been limited to incomplete 3D understanding: their segmentation results are rendered as 2D masks that do not represent the entire 3D space. To address this limitation, we redefine the problem to segment the 3D volume and propose the following methods for better 3D understanding. We directly supervise the 3D points to train the language embedding field, unlike previous methods that anchor supervision at 2D pixels. We transfer the learned language field to 3DGS, achieving the first real-time rendering speed without sacrificing training time or accuracy. Lastly, we introduce a 3D querying and evaluation protocol for assessing the reconstructed geometry and semantics together. Code, checkpoints, and annotations are available at the project page. Hyunjee Lee, Youngsik Yun, Jeongmin Bae 0001, Seoha Kim, Youngjung Uh |
AAAI | 5 |
| 2025 | TCFG: Tangential Damping Classifier-free GuidanceabstractDiffusion models have achieved remarkable success in text-to-image synthesis, largely attributed to the use of classifier-free guidance (CFG), which enables high-quality, condition-aligned image generation. CFG combines the conditional score (e.g., text-conditioned) with the unconditional score to control the output. However, the unconditional score is in charge of estimating the transition between manifolds of adjacent timesteps from xtto xt−1, which may inadvertently interfere with the trajectory toward the specific condition. In this work, we introduce a novel approach that leverages a geometric perspective on the unconditional score to enhance CFG performance when conditional scores are available. Specifically, we propose a method that filters the singular vectors of both conditional and unconditional scores using singular value decomposition. This filtering process aligns the unconditional score with the conditional score, thereby refining the sampling trajectory to stay closer to the manifold. Our approach improves image quality with negligible additional computation. We provide deeper insights into the score function behavior in diffusion models and present a practical technique for achieving more accurate and contextually coherent image synthesis. project page Mingi Kwon, Shin seong Kim, Jaeseok Jeong 0001, Yi Ting Hsiao, Youngjung Uh |
CVPR | 5 |
| 2025 | Stylekeeper: Prevent Content Leakage using Negative Visual Query GuidanceabstractIn the domain of text-to-image generation, diffusion models have emerged as powerful tools. Recently, studies on visual prompting, where images are used as prompts, have enabled more precise control over style and content. However, existing methods often suffer from content leakage, where undesired elements of the visual style prompt are transferred along with the intended style. To address this issue, we 1) extend classifier-free guidance (CFG) to utilize swapping self-attention and propose 2) negative visual query guidance (NVQG) to reduce the transfer of unwanted contents. NVQG employs negative score by intentionally simulating content leakage scenarios that swap queries instead of key and values of self-attention layers from visual style prompts. This simple yet effective method significantly reduces content leakage. Furthermore, we provide careful solutions for using a real image as visual style prompts. Through extensive evaluation across various styles and text prompts, our method demonstrates superiority over existing approaches, reflecting the style of the references, and ensuring that resulting images match the text prompts. Our code is available \href{https://github.com/naver-ai/StyleKeeper}{here}. Jaeseok Jeong 0002, Gayoung Lee, Yunjey Choi, Youngjung Uh |
ICCV | 5 |
| 2025 | Balanced Conic Rectified FlowabstractRectified flow is a generative model that learns smooth transport mappings between two distributions through an ordinary differential equation (ODE). The model learns a straight ODE by reflow steps which iteratively update the supervisory flow. It allows for a relatively simple and efficient generation of high-quality images. However, rectified flow still faces several challenges. 1) The reflow process is slow because it requires a large number of generated pairs to model the target distribution. 2) It is well known that the use of suboptimal fake samples in reflow can lead to performance degradation of the learned flow model. This issue is further exacerbated by error accumulation across reflow steps and model collapse in denoising autoencoder models caused by self-consuming training.
In this work, we go one step further and empirically demonstrate that the reflow process causes the learned model to drift away from the target distribution, which in turn leads to a growing discrepancy in reconstruction error between fake and real images. We reveal the drift problem and design a new reflow step, namely the conic reflow. It supervises the model by the inversions of real data points through the previously learned model and its interpolation with random initial points. Our conic reflow leads to multiple advantages. 1) It keeps the ODE paths toward real samples, evaluated by reconstruction. 2) We use only a small number of generated samples instead of large generated samples, 600K and 4M, respectively. 3) The learned model generates images with higher quality evaluated by FID, IS, and Recall. 4) The learned flow is more straight than others, evaluated by curvature. We achieve much lower FID in both one-step and full-step generation in CIFAR-10. The conic reflow generalizes to various datasets such as LSUN Bedroom and ImageNet. Shin seong Kim, Mingi Kwon, Youngjung Uh |
NeurIPS | 4 |
| 2025 | FLoD: Integrating Flexible Level of Detail into 3D Gaussian Splatting for Customizable Renderingabstract3D Gaussian Splatting (3DGS) has significantly advanced computer graphics by enabling high-quality 3D reconstruction and fast rendering speeds, inspiring numerous follow-up studies. However, 3DGS and its subsequent works are restricted to specific hardware setups, either on only low-cost or on only high-end configurations. Approaches aimed at reducing 3DGS memory usage enable rendering on low-cost GPU but compromise rendering quality, which fails to leverage the hardware capabilities in the case of higher-end GPU. Conversely, methods that enhance rendering quality require high-end GPU with large VRAM, making such methods impractical for lower-end devices with limited memory capacity. Consequently, 3DGS-based works generally assume a single hardware setup and lack the flexibility to adapt to varying hardware constraints. To overcome this limitation, we propose Flexible Level of Detail (FLoD) for 3DGS. FLoD constructs a multi-level 3DGS representation through level-specific 3D scale constraints, where each level independently reconstructs the entire scene with varying detail and GPU memory usage. A level-by-level training strategy is introduced to ensure structural consistency across levels. Furthermore, the multi-level structure of FLoD allows selective rendering of image regions at different detail levels, providing additional memory-efficient rendering options. To our knowledge, among prior works which incorporate the concept of Level of Detail (LoD) with 3DGS, FLoD is the first to follow the core principle of LoD by offering adjustable options for a broad range of GPU settings. Experiments demonstrate that FLoD provides various rendering options with trade-offs between quality and memory usage, enabling real-time rendering under diverse memory constraints. Furthermore, we show that FLoD generalizes to different 3DGS frameworks, indicating its potential for integration into future state-of-the-art developments. Yunji Seo, Young Sun Choi, Hyun Seung Son, Youngjung Uh |
ACM Trans. Graph. | 4 |
| 2024 | Sync-NeRF: Generalizing Dynamic NeRFs to Unsynchronized VideosabstractRecent advancements in 4D scene reconstruction using neural radiance fields (NeRF) have demonstrated the ability to represent dynamic scenes from multi-view videos. However, they fail to reconstruct the dynamic scenes and struggle to fit even the training views in unsynchronized settings. It happens because they employ a single latent embedding for a frame while the multi-view images at the same frame were actually captured at different moments. To address this limitation, we introduce time offsets for individual unsynchronized videos and jointly optimize the offsets with NeRF. By design, our method is applicable for various baselines and improves them with large margins. Furthermore, finding the offsets always works as synchronizing the videos without manual effort. Experiments are conducted on the common Plenoptic Video Dataset and a newly built Unsynchronized Dynamic Blender Dataset to verify the performance of our method. Project page: https://seoha-kim.github.io/sync-nerf Seoha Kim, Jeongmin Bae 0001, Youngsik Yun, Hahyun Lee, Gun Bang, Youngjung Uh |
AAAI | 6 |
| 2024 | Characterizing and Quantifying Expert Input Behavior in League of LegendsabstractTo achieve high performance in esports, players must be able to effectively and efficiently control input devices such as a computer mouse and keyboard (i.e., input skills). Characterizing and quantifying a player’s input skills can provide useful insights, but collecting and analyzing sufficient amounts of data in ecologically valid settings remains a challenge. Targeting the popular esports game, League of Legends, we go beyond the limitations of previous studies and demonstrate a holistic pipeline of input behavior analysis: from quantifying the quality of players’ input behavior (i.e., input skill) to training players based on the analysis. Based on interviews with five top-tier professionals and analysis of input behavior logs from 4,835 matches played freely at home collected from 193 players (including 18 professionals), we confirmed that players with higher ranks in the game implement eight different input skills with higher quality. In a three-week follow-up study using a training aid that visualizes a player’s input skill levels, we found that the analysis provided players with actionable lessons, potentially leading to meaningful changes in their input behavior. Hanbyeol Lee, Seyeon Lee, Rohan Nallapati, Youngjung Uh, Byungjoo Lee |
CHI | 4 |
| 2024 | Per-Gaussian Embedding-Based Deformation for Deformable 3D Gaussian Splatting
Jeongmin Bae 0001, Seoha Kim, Youngsik Yun, Hahyun Lee, Gun Bang, Youngjung Uh |
ECCV (15) | 6 |
| 2024 | HARIVO: Harnessing Text-to-Image Models for Video Generation
Mingi Kwon, Seoung Wug Oh, Yang Zhou 0009, Difan Liu, Joon-Young Lee, Haoran Cai, Baqiao Liu, Feng Liu 0015, Youngjung Uh |
ECCV (53) | 9 |
| 2024 | Attribute Based Interpretable Evaluation Metrics for Generative ModelsabstractWhen the training dataset comprises a 1:1 proportion of dogs to cats, a generative model that produces 1:1 dogs and cats better resembles the training species distribution than another model with 3:1 dogs and cats. Can we capture this phenomenon using existing metrics? Unfortunately, we cannot, because these metrics do not provide any interpretability beyond “diversity". In this context, we propose a new evaluation protocol that measures the divergence of a set of generated images from the training set regarding the distribution of attribute strengths as follows. Singleattribute Divergence (SaD) reveals the attributes that are generated excessively or insufficiently by measuring the divergence of PDFs of individual attributes. Paired-attribute Divergence (PaD) reveals such pairs of attributes by measuring the divergence of joint PDFs of pairs of attributes. For measuring the attribute strengths of an image, we propose Heterogeneous CLIPScore (HCS) which measures the cosine similarity between image and text vectors with heterogeneous initial points. With SaD and PaD, we reveal the following about existing generative models. ProjectedGAN generates implausible attribute relationships such as baby with beard even though it has competitive scores of existing metrics. Diffusion models struggle to capture diverse colors in the datasets. The larger sampling timesteps of the latent diffusion model generate the more minor objects including earrings and necklace. Stable Diffusion v1.5 better captures the attributes than v2.1. Our metrics lay a foundation for explainable evaluations of generative models. Dongkyun Kim, Mingi Kwon, Youngjung Uh |
ICML | 3 |
| 2024 | Training-free Content Injection using h-space in Diffusion ModelsabstractDiffusion models (DMs) synthesize high-quality images in various domains. However, controlling their generative process is still hazy because the intermediate variables in the process are not rigorously studied. Recently, the bottleneck feature of the U-Net, namely h-space, is found to convey the semantics of the resulting image. It enables StyleCLIP-like latent editing within DMs. In this paper, we explore further usage of h-space beyond attribute editing, and introduce a method to inject the content of one image into another image by combining their features in the generative processes. Briefly, given the original generative process of the other image, 1) we gradually blend the bottleneck feature of the content with proper normalization, and 2) we calibrate the skip connections to match the injected content. Unlike custom-diffusion approaches, our method does not require time-consuming optimization or fine-tuning. Instead, our method manipulates intermediate features within a feed-forward generative process. Furthermore, our method does not require supervision from external networks. Project page: https://curryjung.github.io/DiffStyle/ Jaeseok Jeong 0002, Mingi Kwon, Youngjung Uh |
WACV | 3 |
| 2024 | Small Objects Matters in Weakly-supervised Semantic SegmentationabstractWeakly-supervised semantic segmentation (WSSS) performs pixel-wise classification given only image-level labels for training. Despite the difficulty of this task, the research community has achieved promising results over the last five years. Still, current WSSS literature misses the detailed sense of how well the methods perform on different sizes of objects. Thus we propose a novel evaluation metric to provide a comprehensive assessment across different object sizes and collect a size-balanced evaluation set to complement PASCAL VOC. With these two gadgets, we reveal that the existing WSSS methods struggle in capturing small objects. Furthermore, we propose a size-balanced cross-entropy loss coupled with a proper training strategy. It generally improves existing WSSS methods as validated upon ten baselines on three different datasets. Cheolhyun Mun, Sanghuk Lee, Youngjung Uh, Junsuk Choe, Hyeran Byun |
WACV | 3 |
| 2024 | Discovering an inference recipe for weakly-supervised object localization
Sanghuk Lee, Cheolhyun Mun, Youngjung Uh, Junsuk Choe, Hyeran Byun |
Pattern Recognit. | 3 |
| 2023 | LANIT: Language-Driven Image-to-Image Translation for Unlabeled DataabstractExisting techniques for image-to-image translation commonly have suffered from two critical problems: heavy reliance on per-sample domain annotation and/or inability to handle multiple attributes per image. Recent truly-unsupervised methods adopt clustering approaches to easily provide per-sample one-hot domain labels. However, they cannot account for the real-world setting: one sample may have multiple attributes. In addition, the semantics of the clusters are not easily coupled to human understanding. To overcome these, we present LANguage-driven Image-to-image Translation model, dubbed LANIT. We leverage easy-to-obtain candidate attributes given in texts for a dataset: the similarity between images and attributes indicates per-sample domain labels. This formulation naturally enables multi-hot labels so that users can specify the target domain with a set of attributes in language. To account for the case that the initial prompts are inaccurate, we also present prompt learning. We further present domain regularization loss that enforces translated images to be mapped to the corresponding domain. Experiments on several standard benchmarks demonstrate that LANIT achieves comparable or superior performance to existing models. The code is available at github.com/KU-CVLAB/LANIT. Seokju Cho, Jaejun Yoo 0001, Youngjung Uh, Seungryong Kim |
CVPR | 6 |
| 2023 | AesPA-Net: Aesthetic Pattern-Aware Style Transfer NetworksabstractTo deliver the artistic expression of the target style, recent studies exploit the attention mechanism owing to its ability to map the local patches of the style image to the corresponding patches of the content image. However, because of the low semantic correspondence between arbitrary content and artworks, the attention module repeatedly abuses specific local patches from the style image, resulting in disharmonious and evident repetitive artifacts. To overcome this limitation and accomplish impeccable artistic style transfer, we focus on enhancing the attention mechanism and capturing the rhythm of patterns that organize the style. In this paper, we introduce a novel metric, namely pattern repeatability, that quantifies the repetition of patterns in the style image. Based on the pattern repeatability, we propose Aesthetic Pattern-Aware style transfer Networks (AesPA-Net) that discover the sweet spot of local and global style expressions. In addition, we propose a novel self-supervisory task to encourage the attention mechanism to learn precise and meaningful semantic correspondence. Lastly, we introduce the patch-wise style loss to transfer the elaborate rhythm of local patterns. Through qualitative and quantitative evaluations, we verify the reliability of the proposed pattern repeatability that aligns with human perception, and demonstrate the superiority of the proposed framework. All codes and pre-trained weights are available at Kibeom-Hong/AesPA-Net. Kibeom Hong, Seogkyu Jeon, Junsoo Lee 0002, Namhyuk Ahn, Kunhee Kim, Pilhyeon Lee, Youngjung Uh, Hyeran Byun |
ICCV | 8 |
| 2023 | BallGAN: 3D-aware Image Synthesis with a Spherical Backgroundabstract3D-aware GANs aim to synthesize realistic 3D scenes that can be rendered in arbitrary camera viewpoints, generating high-quality images with well-defined geometry. As 3D content creation becomes more popular, the ability to generate foreground objects separately from the background has become a crucial property. Existing methods have been developed regarding overall image quality, but they can not generate foreground objects only and often show degraded 3D geometry. In this work, we propose to represent the background as a spherical surface for multiple reasons inspired by computer graphics. Our method naturally provides foreground-only 3D synthesis facilitating easier 3D content creation. Furthermore, it improves the foreground geometry of 3D-aware GANs and the training stability on datasets with complex backgrounds. Project page: https://minjung-s.github.io/ballgan/ Minjung Shin, Yunji Seo, Jeongmin Bae 0001, Young Sun Choi, Hyunsu Kim, Hyeran Byun, Youngjung Uh |
ICCV | 7 |
| 2023 | Diffusion Models Already Have A Semantic Latent Space
Mingi Kwon, Jaeseok Jeong 0002, Youngjung Uh |
ICLR | 3 |
| 2023 | Semantic Image Synthesis with Unconditional GeneratorabstractSemantic image synthesis (SIS) aims to generate realistic images according to semantic masks given by a user. Although recent methods produce high quality results with fine spatial control, SIS requires expensive pixel-level annotation of the training images. On the other hand, manipulating intermediate feature maps in a pretrained unconditional generator such as StyleGAN supports coarse spatial control without heavy annotation. In this paper, we introduce a new approach, for reflecting user's detailed guiding masks on a pretrained unconditional generator. Our method converts a user's guiding mask to a proxy mask through a semantic mapper. Then the proxy mask conditions the resulting image through a rearranging network based on cross-attention mechanism. The proxy mask is simple clustering of intermediate feature maps in the generator. The semantic mapper and the rearranging network are easy to train (less than half an hour). Our method is useful for many tasks: semantic image synthesis, spatially editing real images, and unaligned local transplantation. Last but not least, it is generally applicable to various datasets such as human faces, animal faces, and churches. Jungwoo Chae, Hyunin Cho, Sooyeon Go, Kyungmook Choi, Youngjung Uh |
NeurIPS | 5 |
| 2023 | Understanding the Latent Space of Diffusion Models through the Lens of Riemannian GeometryabstractDespite the success of diffusion models (DMs), we still lack a thorough understanding of their latent space. To understand the latent space $\mathbf{x}_t \in \mathcal{X}$, we analyze them from a geometrical perspective. Our approach involves deriving the local latent basis within $\mathcal{X}$ by leveraging the pullback metric associated with their encoding feature maps. Remarkably, our discovered local latent basis enables image editing capabilities by moving $\mathbf{x}_t$, the latent space of DMs, along the basis vector at specific timesteps. We further analyze how the geometric structure of DMs evolves over diffusion timesteps and differs across different text conditions. This confirms the known phenomenon of coarse-to-fine generation, as well as reveals novel insights such as the discrepancy between $\mathbf{x}_t$ across timesteps, the effect of dataset complexity, and the time-varying influence of text prompts. To the best of our knowledge, this paper is the first to present image editing through $\mathbf{x}$-space traversal, editing only once at specific timestep $t$ without any additional training, and providing thorough analyses of the latent structure of DMs.
The code to reproduce our experiments can be found at the [link](https://github.com/enkeejunior1/Diffusion-Pullback). Yong-Hyun Park, Mingi Kwon, Jaewoong Choi, Junghyo Jo, Youngjung Uh |
NeurIPS | 5 |
| 2022 | Feature Statistics Mixing Regularization for Generative Adversarial NetworksabstractIn generative adversarial networks, improving discriminators is one of the key components for generation performance. As image classifiers are biased toward texture and debiasing improves accuracy, we investigate 1) if the discriminators are biased, and 2) if debiasing the discriminators will improve generation performance. Indeed, we find empirical evidence that the discriminators are sensitive to the style (e.g., texture and color) of images. As a remedy, we propose feature statistics mixing regularization (FSMR) that encourages the discriminator's prediction to be invariant to the styles of input images. Specifically, we generate a mixed feature of an original and a reference image in the discriminator's feature space and we apply regularization so that the prediction for the mixed feature is consistent with the prediction for the original image. We conduct extensive experiments to demonstrate that our regularization leads to reduced sensitivity to style and consistently improves the performance of various GAN architectures on nine datasets. In addition, adding FSMR to recently-proposed augmentation-based GAN methods further improves image quality. Our code is available at https://github.com/naver-ai/FSMR. Yunjey Choi, Youngjung Uh |
CVPR | 3 |
| 2022 | FurryGAN: High Quality Foreground-Aware Image Synthesis
Jeongmin Bae 0001, Mingi Kwon, Youngjung Uh |
ECCV (14) | 3 |
| 2021 | Exploiting Spatial Dimensions of Latent in GAN for Real-Time Image EditingabstractGenerative adversarial networks (GANs) synthesize realistic images from random latent vectors. Although manipulating the latent vectors controls the synthesized outputs, editing real images with GANs suffers from i) time-consuming optimization for projecting real images to the latent vectors, ii) or inaccurate embedding through an encoder. We propose StyleMapGAN: the intermediate latent space has spatial dimensions, and a spatially variant modulation replaces AdaIN. It makes the embedding through an encoder more accurate than existing optimization-based methods while maintaining the properties of GANs. Experimental results demonstrate that our method significantly outperforms state-of-the-art models in various image manipulation tasks such as local editing and image interpolation. Last but not least, conventional editing methods on GANs are still valid on our StyleMapGAN. Source code is available at https://github.com/naver-ai/StyleMapGAN. Hyunsu Kim, Yunjey Choi, Sungjoo Yoo, Youngjung Uh |
CVPR | 5 |
| 2021 | Rethinking the Truly Unsupervised Image-to-Image TranslationabstractEvery recent image-to-image translation model inherently requires either image-level (i.e. input-output pairs) or set-level (i.e. domain labels) supervision. However, even set-level supervision can be a severe bottleneck for data collection in practice. In this paper, we tackle image-to-image translation in a fully unsupervised setting, i.e., neither paired images nor domain labels. To this end, we propose a truly unsupervised image-to-image translation model (TUNIT) that simultaneously learns to separate image domains and translates input images into the estimated domains. Experimental results show that our model achieves comparable or even better performance than the set-level supervised model trained with full labels, generalizes well on various datasets, and is robust against the choice of hyperparameters (e.g. the preset number of pseudo domains). Furthermore, TUNIT can be easily extended to semi-supervised learning with a few labeled data. Kyungjune Baek, Yunjey Choi, Youngjung Uh, Jaejun Yoo 0001, Hyunjung Shim |
ICCV | 3 |
| 2021 | Contrastive Attention Maps for Self-supervised Co-localizationabstractThe goal of unsupervised co-localization is to locate the object in a scene under the assumptions that 1) the dataset consists of only one superclass, e.g., birds, and 2) there are no human-annotated labels in the dataset. The most recent method achieves impressive co-localization performance by employing self-supervised representation learning approaches such as predicting rotation. In this paper, we introduce a new contrastive objective directly on the attention maps to enhance co-localization performance. Our contrastive loss function exploits rich information of location, which induces the model to activate the extent of the object effectively. In addition, we propose a pixel-wise attention pooling that selectively aggregates the feature map regarding their magnitudes across channels. Our methods are simple and shown effective by extensive qualitative and quantitative evaluation, achieving state-of-the-art co-localization performances by large margins on four datasets: CUB-200-2011, Stanford Cars, FGVC-Aircraft, and Stanford Dogs. Our code will be publicly available online for the research community. Minsong Ki, Youngjung Uh, Junsuk Choe, Hyeran Byun |
ICCV | 2 |
| 2021 | AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant Weights
Byeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han, Sangdoo Yun, Gyuwan Kim, Youngjung Uh, Jung-Woo Ha 0001 |
ICLR | 7 |
| 2021 | ArrowGAN : Learning to generate videos by learning Arrow of Time
Kibeom Hong, Youngjung Uh, Hyeran Byun |
Neurocomputing | 2 |
| 2021 | Contrastive and consistent feature learning for weakly supervised object localization and semantic segmentation
Minsong Ki, Youngjung Uh, Wonyoung Lee 0004, Hyeran Byun |
Neurocomputing | 2 |
| 2020 | Background Suppression Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization is a very challenging problem because frame-wise labels are not given in the training stage while the only hint is video-level labels: whether each video contains action frames of interest. Previous methods aggregate frame-level class scores to produce video-level prediction and learn from video-level action labels. This formulation does not fully model the problem in that background frames are forced to be misclassified as action classes to predict video-level labels accurately. In this paper, we design Background Suppression Network (BaS-Net) which introduces an auxiliary class for background and has a two-branch weight-sharing architecture with an asymmetrical training strategy. This enables BaS-Net to suppress activations from background frames to improve localization performance. Extensive experiments demonstrate the effectiveness of BaS-Net and its superiority over the state-of-the-art methods on the most popular benchmarks – THUMOS'14 and ActivityNet. Our code and the trained model are available at https://github.com/Pilhyeon/BaSNet-pytorch. Pilhyeon Lee, Youngjung Uh, Hyeran Byun |
AAAI | 2 |
| 2020 | In-sample Contrastive Learning and Consistent Attention for Weakly Supervised Object Localization
Minsong Ki, Youngjung Uh, Wonyoung Lee 0004, Hyeran Byun |
ACCV (4) | 2 |
| 2020 | StarGAN v2: Diverse Image Synthesis for Multiple DomainsabstractA good image-to-image translation model should learn a mapping between different visual domains while satisfying the following properties: 1) diversity of generated images and 2) scalability over multiple domains. Existing methods address either of the issues, having limited diversity or multiple models for all domains. We propose StarGAN v2, a single framework that tackles both and shows significantly improved results over the baselines. Experiments on CelebA-HQ and a new animal faces dataset (AFHQ) validate our superiority in terms of visual quality, diversity, and scalability. To better assess image-to-image translation models, we release AFHQ, high-quality animal faces with large inter- and intra-domain differences. The code, pretrained models, and dataset are available at https://github.com/clovaai/stargan-v2. Yunjey Choi, Youngjung Uh, Jaejun Yoo 0001, Jung-Woo Ha 0001 |
CVPR | 2 |
| 2020 | Reliable Fidelity and Diversity Metrics for Generative ModelsabstractDevising indicative evaluation metrics for the image generation task remains an open problem. The most widely used metric for measuring the similarity between real and generated images has been the Frechet Inception Distance (FID) score. Since it does not differentiate the fidelity and diversity aspects of the generated images, recent papers have introduced variants of precision and recall metrics to diagnose those properties separately. In this paper, we show that even the latest version of the precision and recall metrics are not reliable yet. For example, they fail to detect the match between two identical distributions, they are not robust against outliers, and the evaluation hyperparameters are selected arbitrarily. We propose density and coverage metrics that solve the above issues. We analytically and experimentally show that density and coverage provide more interpretable and reliable signals for practitioners than the existing metrics. Muhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi, Jaejun Yoo 0001 |
ICML | 3 |
| 2019 | Photorealistic Style Transfer via Wavelet TransformsabstractRecent style transfer models have provided promising artistic results. However, given a photograph as a reference style, existing methods are limited by spatial distortions or unrealistic artifacts, which should not happen in real photographs. We introduce a theoretically sound correction to the network architecture that remarkably enhances photorealism and faithfully transfers the style. The key ingredient of our method is wavelet transforms that naturally fits in deep networks. We propose a wavelet corrected transfer based on whitening and coloring transforms (WCT2) that allows features to preserve their structural information and statistical properties of VGG feature space during stylization. This is the first and the only end-to-end model that can stylize a 1024x1024 resolution image in 4.7 seconds, giving a pleasing and photorealistic quality without any post-processing. Last but not least, our model provides a stable video stylization without temporal constraints. Our code, generated images, pre-trained models and supplementary documents are all available at https://github.com/ClovaAI/WCT2. Jaejun Yoo 0001, Youngjung Uh, Sanghyuk Chun, Byeongkyu Kang, Jung-Woo Ha 0001 |
ICCV | 2 |
| 2019 | Exploiting hierarchical visual features for visual question answering
Jongkwang Hong, Jianlong Fu, Youngjung Uh, Tao Mei 0001, Hyeran Byun |
Neurocomputing | 3 |
| 2017 | Discovering overlooked objects: Context-based boosting of object detection in indoor scenes
Jongkwang Hong, Yongwon Hong, Youngjung Uh, Hyeran Byun |
Pattern Recognit. Lett. | 3 |
| 2016 | Multi-view 3D reconstruction by random-search and propagation with view-dependent patch maps
Youngjung Uh, Hyeran Byun |
Multim. Tools Appl. | 1 |
| 2014 | Efficient Multiview Stereo by Random-Search and PropagationabstractWe present an efficient multi-view 3D reconstruction method based on randomization and propagation scheme. Our method progressively refines 3D point estimates by randomly perturbing the initial guess of 3D points and propagates photo-consistent ones to their neighbors. In contrast to previous refinement methods that perform local optimization for a better photo-consistency, our randomization approach takes lucky matchings for reducing the computational complexity. Experiments show favorable efficiency of the proposed method with the accuracy that is close to the state-of-the-art methods. Youngjung Uh, Yasuyuki Matsushita, Hyeran Byun |
3DV | 1 |
| 2012 | Generating panorama image by synthesizing multiple homographyabstractThis paper presents a method to generate image mosaics of a panoramic scene. In general, the relation between images which is required for mosaicing cannot be expressed by a single homography due to geometrical condition of the scene, even if the images are taken at the same position. Many existing methods are using only one homography to make panorama image while ignoring the geometrical variations. Therefore, they experience a lot of distortions and misalignments from input images which contain several planes which cannot be handled by one homography. In this paper, we present a novel method that utilizes synthesis of multiple homography to warp the images. Moreover, our method determines the number of homography automatically, without user's input. By our method, various distortions of shapes and mismatches can be reduced. Youngjung Uh, Hyeran Byun |
ICIP | 2 |
| 2012 | Color and shape feature-based detection of speed sign in real-timeabstractThis paper presents a method for detecting speed sign based on color and shape features in real-time under real-life environment. In our method, Region Of Interest(ROI) is extracted and verified based on shape feature. In the first step, ROI is roughly extracted by segmentation of a red rim and the segments are optimized by the boundary using guided image filtering. Next step, the shape-based detection verifies the extracted red rim. We compare three different shape-based detection methods, RSD, BCT, and STVUE, and the RSD shows the best speed sign detection rate of 93% on the experimental data of 62 images containing 85 speed sign. Seunggyu Kim, Youngjung Uh, Hyeran Byun |
SMC | 3 |
| 2012 | Service-oriented architecture based on biometric using random features and incremental neural networks
Kwontaeg Choi, Kar-Ann Toh, Youngjung Uh, Hyeran Byun |
Soft Comput. | 3 |
| 2011 | Plot preservation approach for video summarizationabstractThis paper presents an effective method for summarizing content of the video with plot preserved. The proposed method provides static video summary and consists of three major procedures (1) extracting keyframes regarding temporal information, (2) estimating Region of Interest (ROI) from the extracted frames, and (3) assembling the ROI into one image by arranging them according to the temporal order and their size, while they blend each other smoothly. In each process, we make the method concise for computational complexity. Experiment on various types of video (e.g. movie, animation, home video) demonstrates that the proposed method generates expressive video summarization and conserves the plot information. The main contribution relies on reflecting temporal information in video summarization. Yeosun Lim, Youngjung Uh, Hyeran Byun |
SMC | 2 |