Namhyuk Ahn

dblp:217/1998 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0003-1990-9516ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 6 since 2021
YearPublicationVenuePosition
2026 T2LF: LLM-Guided Multimodal Diffusion for Text-to-Light Field Synthesis
abstract
We present a novel text-driven approach for light field (LF) synthesis. Existing methods typically generate LFs from given images, requiring users to find reference images, which makes it difficult to construct the desired scene directly and limits scene diversity. Moreover, existing methods are mainly designed for limited baselines from training datasets, making it difficult to implement various viewpoint changes and consequently limiting the flexibility of motion. In contrast, our method directly synthesizes LFs from user-provided text descriptions by leveraging the scene understanding capabilities of a multi-modal large language model (LLM) and the generative power of a diffusion model. Given a text prompt describing the desired LF, the multimodal LLM extracts relevant information for LF synthesis, which then guides a diffusion model to produce diverse scenes and motions. This approach enables LF synthesis even with a pre-trained model not initially designed for this purpose, requiring only minimal fine-tuning. The proposed framework enables visually diverse LF synthesis with only text input. Experimental results demonstrate that the synthesized LFs exhibit geometric consistency and achieve advanced synthesis quality compared to existing methods.
Soyoung Yoon, Namhyuk Ahn, In Kyu Park
WACV2
2026 DiffBlender: Composable and versatile multimodal text-to-image diffusion models
Sungnyun Kim, Junsoo Lee 0002, Kibeom Hong, Namhyuk Ahn
Expert Syst. Appl.5
2026 Imperceptible Protection Against Style Imitation From Diffusion Models
abstract
Recent progress in diffusion models has profoundly enhanced the fidelity of image generation, but it has raised concerns about copyright infringements. While prior methods have introduced adversarial perturbations to prevent style imitation, most are accompanied by the degradation of artworks' visual quality. Recognizing the importance of maintaining this, we intro duce a visually improved protection method while preserving its protection capability. To this end, we devise a perceptual map to highlight areas sensitive to human eyes, guided by instance-aware refinement, which refines the protection intensity accordingly. We also introduce a difficulty-aware protection by predicting how difficult the artwork is to protect and dynamically adjusting the intensity based on this. Lastly, we integrate a perceptual constraints bank to further improve the imperceptibility. Results show that our method substantially elevates the quality of the protected image without compromising on protection efficacy.
Namhyuk Ahn, Wonhyuk Ahn, KiYoon Yoo, Seung-Hun Nam
IEEE Trans. Multim.1
2025 Nearly Zero-Cost Protection Against Mimicry by Personalized Diffusion Models
abstract
Recent advancements in diffusion models revolutionize image generation but pose risks of misuse, such as replicating artworks or generating deepfakes. Existing image protection methods, though effective, struggle to balance protection efficacy, invisibility, and latency, thus limiting practical use. We introduce perturbation pre-training to reduce latency and propose a mixture-of-perturbations approach that dynamically adapts to input images to minimize performance degradation. Our novel training strategy computes protection loss across multiple VAE feature spaces, while adaptive targeted protection at inference enhances robustness and invisibility. Experiments show comparable protection performance with improved invisibility and drastically reduced inference time. The code and demo are available at https://webtoon.github.io/impasto
Namhyuk Ahn, KiYoon Yoo, Wonhyuk Ahn, Seung-Hun Nam
CVPR1
2025 Magnitude Attention-based Dynamic Pruning
abstract
Existing pruning methods often rely on weight importance to identify sparse structures but typically apply this information statically, without leveraging it adaptively during training. In this work, we propose a novel approach - M agnitude A ttention-based Dynamic P runing (MAP) method, which applies the importance of weights throughout both the forward and backward paths to explore sparse model structures dynamically. Magnitude attention is defined based on the magnitude of weights as continuous real-valued numbers enabling a seamless transition from a redundant to an effective sparse network by promoting efficient exploration. Additionally, the attention mechanism ensures more effective updates for important layers within the sparse network. Later, our approach shifts from exploration to exploitation, exclusively updating the sparse model composed of crucial weights based on the explored structure, resulting in pruned models that not only achieve performance comparable to dense models but also outperform previous pruning methods on CIFAR-10/100 and ImageNet.
Jihye Back, Namhyuk Ahn, Jangho Kim
Expert Syst. Appl.2
2024 DreamStyler: Paint by Style Inversion with Text-to-Image Diffusion Models
abstract
Recent progresses in large-scale text-to-image models have yielded remarkable accomplishments, finding various applications in art domain. However, expressing unique characteristics of an artwork (e.g. brushwork, colortone, or composition) with text prompts alone may encounter limitations due to the inherent constraints of verbal description. To this end, we introduce DreamStyle, a novel framework designed for artistic image synthesis, proficient in both text-to-image synthesis and style transfer. DreamStyle optimizes a multi-stage textual embedding with a context-aware text prompt, resulting in prominent image quality. In addition, with content and style guidance, DreamStyle exhibits flexibility to accommodate a range of style references. Experimental results demonstrate its superior performance across multiple scenarios, suggesting its promising potential in artistic product creation. Project page: https://nmhkahn.github.io/dreamstyler/
Namhyuk Ahn, Junsoo Lee 0002, Chunggi Lee, Kunhee Kim, Seung-Hun Nam, Kibeom Hong
AAAI1
2024 Decomposing texture and semantic for out-of-distribution detection
abstract
The out-of-distribution (OOD) detection task assumes samples that follow the distribution of training data as in-distribution (ID), while samples from other data distributions are considered OOD. In recent years, the OOD detection tasks have made significant progress since many studies observed that the distribution mismatch between training and real datasets can severely deteriorate the reliability of AI systems. Nevertheless, the lack of precise interpretation for the in-distribution (ID) limits the application of the OOD detection methods to real-world systems. To tackle this, we decompose the definition of the ID into texture and semantics, motivated by the demands of real-world scenarios. We also design new benchmarks to measure the robustness that OOD detection methods should have. Our proposed benchmark verifies not only the precision but also the robustness of the detection models. It is crucial to measure both factors in OOD detection as they indicate different traits of the model. For instance, precision is relevant to scenarios that detect minor cracks in the conveyor belt of a smart factory, whereas robustness pertains to maintaining performance under diverse weather conditions, as required by autonomous driving. To achieve a good balance between the OOD detection performance and robustness, our method takes a divide-and-conquer approach. Specifically, the proposed model first handles each component of the texture and semantics separately and then fuses these later. This philosophy is empirically proven by a series of benchmarks including both the proposed and the conventional counterpart. By decomposing the prior “unclear” definition of the ID into texture and semantic components, our novel approach better suits the demands of a reliable machine learning system, which requires robustness and consistent performance across varied scenarios. Unlike prior works, our approach does not rely on any extra datasets or labels. This prevents our proposed framework from being dependent on a particular dataset distribution.
Jeong-Hyeon Moon, Namhyuk Ahn, Kyung-Ah Sohn 0001
Expert Syst. Appl.2
2024 Data Augmentation for Low-Level Vision: CutBlur and Mixture-of-Augmentation
Namhyuk Ahn, Jaejun Yoo 0001, Kyung-Ah Sohn 0001
Int. J. Comput. Vis.1
2023 Interactive Cartoonization with Controllable Perceptual Factors
abstract
Cartoonization is a task that renders natural photos into cartoon styles. Previous deep cartoonization methods only have focused on end-to-end translation, which may hinder editability. Instead, we propose a novel solution with editing features of texture and color based on the cartoon creation process. To do that, we design a model architecture to have separate decoders, texture and color, to decouple these attributes. In the texture decoder, we propose a texture controller, which enables a user to control stroke style and abstraction to generate diverse cartoon textures. We also introduce an HSV color augmentation to induce the networks to generate diverse and controllable color translation. To the best of our knowledge, our work is the first deep approach to control the cartoonization at inference while showing profound quality improvement over to baselines.
Namhyuk Ahn, Patrick Kwon, Jihye Back, Kibeom Hong, Seungkwon Kim
CVPR1
2023 AesPA-Net: Aesthetic Pattern-Aware Style Transfer Networks
abstract
To deliver the artistic expression of the target style, recent studies exploit the attention mechanism owing to its ability to map the local patches of the style image to the corresponding patches of the content image. However, because of the low semantic correspondence between arbitrary content and artworks, the attention module repeatedly abuses specific local patches from the style image, resulting in disharmonious and evident repetitive artifacts. To overcome this limitation and accomplish impeccable artistic style transfer, we focus on enhancing the attention mechanism and capturing the rhythm of patterns that organize the style. In this paper, we introduce a novel metric, namely pattern repeatability, that quantifies the repetition of patterns in the style image. Based on the pattern repeatability, we propose Aesthetic Pattern-Aware style transfer Networks (AesPA-Net) that discover the sweet spot of local and global style expressions. In addition, we propose a novel self-supervisory task to encourage the attention mechanism to learn precise and meaningful semantic correspondence. Lastly, we introduce the patch-wise style loss to transfer the elaborate rhythm of local patterns. Through qualitative and quantitative evaluations, we verify the reliability of the proposed pattern repeatability that aligns with human perception, and demonstrate the superiority of the proposed framework. All codes and pre-trained weights are available at Kibeom-Hong/AesPA-Net.
Kibeom Hong, Seogkyu Jeon, Junsoo Lee 0002, Namhyuk Ahn, Kunhee Kim, Pilhyeon Lee, Youngjung Uh, Hyeran Byun
ICCV4
2022 Efficient deep neural network for photo-realistic image super-resolution
abstract
Recent progress in deep learning-based models has improved photo-realistic (or perceptual) single-image super-resolution significantly. However, despite their powerful performance, many methods are difficult to apply to real-world applications because of the heavy computational requirements. To facilitate the use of a deep model under such demands, we focus on keeping the network efficient while maintaining its performance. In detail, we design an architecture that implements a cascading mechanism on a residual network to boost the performance with limited resources via multi-level feature fusion. In addition, our proposed model adopts group convolution and recursive schemes in order to achieve extreme efficiency. We further improve the perceptual quality of the output by employing the adversarial learning paradigm and a multi-scale discriminator approach. The performance of our method is investigated through extensive internal experiments and benchmarks using various datasets. Our results show that our models outperform the recent methods with similar complexity, for both traditional pixel-based and perception-based tasks.
Namhyuk Ahn, Byungkon Kang, Kyung-Ah Sohn 0001
Pattern Recognit.1
2020 Restoring Spatially-Heterogeneous Distortions Using Mixture of Experts Network
Sijin Kim, Namhyuk Ahn, Kyung-Ah Sohn 0001
ACCV (2)2
2020 Rethinking Data Augmentation for Image Super-resolution: A Comprehensive Analysis and a New Strategy
abstract
Data augmentation is an effective way to improve the performance of deep networks. Unfortunately, current methods are mostly developed for high-level vision tasks (e.g., classification) and few are studied for low-level vision tasks (e.g., image restoration). In this paper, we provide a comprehensive analysis of the existing augmentation methods applied to the super-resolution task. We find that the methods discarding or manipulating the pixels or features too much hamper the image restoration, where the spatial relationship is very important. Based on our analyses, we propose CutBlur that cuts a low-resolution patch and pastes it to the corresponding high-resolution image region and vice versa. The key intuition of CutBlur is to enable a model to learn not only "how" but also "where" to super-resolve an image. By doing so, the model can understand "how much", instead of blindly learning to apply super-resolution to every given pixel. Our method consistently and significantly improves the performance across various scenarios, especially when the model size is big and the data is collected under real-world environments. We also show that our method improves other low-level vision tasks, such as denoising and compression artifact removal.
Jaejun Yoo 0001, Namhyuk Ahn, Kyung-Ah Sohn 0001
CVPR2
2018 Fast, Accurate, and Lightweight Super-Resolution with Cascading Residual Network
Namhyuk Ahn, Byungkon Kang, Kyung-Ah Sohn 0001
ECCV (10)1