Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Daneul Kim

dblp:354/8351 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0003-2223-8063ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 51% Efficient and distributed learning · 24% Video understanding and tracking · 14%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing
image editing
1.722025
Exploring Multimodal Diffusion Transformers for Enhanced Prompt-Based Image Editing · ICCV 2025
Improving Editability in Image Generation with Layer-wise Memory · CVPR 2025
Machine learning › Generative modeling
diffusion model
0.912025
Exploring Multimodal Diffusion Transformers for Enhanced Prompt-Based Image Editing · ICCV 2025
Machine learning › Generative modeling › diffusion model
diffusion transformer
0.912025
Exploring Multimodal Diffusion Transformers for Enhanced Prompt-Based Image Editing · ICCV 2025
Computer vision › Video understanding and tracking › video event understanding
event boundary detection
0.912025
Online Generic Event Boundary Detection · ICCV 2025
Visual content generation and editing › image editing › interactive image editing
iterative image editing
0.912025
Improving Editability in Image Generation with Layer-wise Memory · CVPR 2025
Visual content generation and editing › image editing › local image editing
mask-based image editing
0.912025
Improving Editability in Image Generation with Layer-wise Memory · CVPR 2025
Visual content generation and editing › image editing
text-guided image editing
0.912025
Exploring Multimodal Diffusion Transformers for Enhanced Prompt-Based Image Editing · ICCV 2025
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
channel pruning
0.812024
Pick-or-Mix: Dynamic Channel Sampling for ConvNets · CVPR 2024
Machine learning › Deep learning architectures and training › convolutional neural network
convolutional neural network architecture
0.812024
Pick-or-Mix: Dynamic Channel Sampling for ConvNets · CVPR 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
Pick-or-Mix: Dynamic Channel Sampling for ConvNets · CVPR 2024
Machine learning › Generative modeling › cross-modal generation
story visualization
0.712023
Story Visualization by Online Text Augmentation with Context Memory · ICCV 2023
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.712023
Story Visualization by Online Text Augmentation with Context Memory · ICCV 2023
Machine learning › Generative modeling › synthetic data generation
text data augmentation
0.212023
Story Visualization by Online Text Augmentation with Context Memory · ICCV 2023

Methods — techniques the papers use, named apart from their topics

multimodal attention analysis · 1.7statistical test · 0.9layer-wise memory · 0.9event segmentation theory · 0.9cross-attention · 0.9background consistency guidance · 0.9dynamic channel sampling · 0.8channel squeezing · 0.8online text augmentation · 0.7memory architecture · 0.7bidirectional transformer · 0.7
YearPublicationVenuePosition
2025 Improving Editability in Image Generation with Layer-wise Memory
abstract
Most real-world image editing tasks require multiple sequential edits to achieve desired results. Current editing approaches, primarily designed for single-object modifications, struggle with sequential editing: especially with maintaining previous edits along with adapting new objects naturally into the existing content. These limitations significantly hinder complex editing scenarios where multiple objects need to be modified while preserving their contextual relationships. We address this fundamental challenge through two key proposals: enabling rough mask inputs that preserve existing content while naturally integrating new elements and supporting consistent editing across multiple modifications. Our framework achieves this through layer-wise memory, which stores latent representations and prompt embeddings from previous edits. We propose Background Consistency Guidance that leverages memorized latents to maintain scene coherence and Multi-Query Disentanglement in cross-attention that ensures natural adaptation to existing content. To evaluate our method, we present a new benchmark dataset incorporating semantic alignment metrics and interactive editing scenarios. Through comprehensive experiments, we demonstrate superior performance in iterative image editing tasks with minimal user effort, requiring only rough masks while maintaining high-quality results throughout multiple editing steps.
Daneul Kim, Jaeah Lee, Jaesik Park
CVPR1
2025 Online Generic Event Boundary Detection
abstract
Generic Event Boundary Detection (GEBD) aims to interpret long-form videos through the lens of human perception. However, current GEBD methods require processing complete video frames to make predictions, unlike humans processing data online and in real-time. To bridge this gap, we introduce a new task, Online Generic Event Boundary Detection (On-GEBD), aiming to detect boundaries of generic events immediately in streaming videos. This task faces unique challenges of identifying subtle, taxonomy-free event changes in real-time, without the access to future frames. To tackle these challenges, we propose a novel On-GEBD framework, Estimator, inspired by Event Segmentation Theory (EST) which explains how humans segment ongoing activity into events by leveraging the discrepancies between predicted and actual information. Our framework consists of two key components: the Consistent Event Anticipator (CEA), and the Online Boundary Discriminator (OBD). Specifically, the CEA generates a prediction of the future frame reflecting current event dynamics based solely on prior frames. Then, the OBD measures the prediction error and adaptively adjusts the threshold using statistical tests on past errors to capture diverse, subtle event transitions. Experimental results demonstrate that Estimator outperforms all baselines adapted from recent online video understanding models and achieves performance comparable to prior offline-GEBD methods on the Kinetics-GEBD and TAPOS datasets.
Hyungrok Jung, Daneul Kim, Seunggyun Lim, Jeany Son
ICCV2
2025 Exploring Multimodal Diffusion Transformers for Enhanced Prompt-Based Image Editing
abstract
Transformer-based diffusion models have recently superseded traditional U-Net architectures, with multimodal diffusion transformers (MM-DiT) emerging as the dominant approach in state-of-the-art models like Stable Diffusion 3 and Flux.1. Previous approaches have relied on unidirectional cross-attention mechanisms, with information flowing from text embeddings to image latents. In contrast, MMDiT introduces a unified attention mechanism that concatenates input projections from both modalities and performs a single full attention operation, allowing bidirectional information flow between text and image branches. This architectural shift presents significant challenges for existing editing techniques. In this paper, we systematically analyze MM-DiT's attention mechanism by decomposing attention matrices into four distinct blocks, revealing their inherent characteristics. Through these analyses, we propose a robust, prompt-based image editing method for MM-DiT that supports global to local edits across various MM-DiT variants, including few-step models. We believe our findings bridge the gap between existing U-Net-based methods and emerging architectures, offering deeper insights into MMDiT's behavioral patterns.
Joonghyuk Shin, Alchan Hwang, Daneul Kim, Jaesik Park
ICCV4
2024 Pick-or-Mix: Dynamic Channel Sampling for ConvNets
abstract
Channel pruning approaches for convolutional neural networks (ConvNets) deactivate the channels, statically or dynamically, and require special implementation. In addition, channel squeezing in representative ConvNets is carried out via$1\times 1$convolutions which dominates a large portion of computations and network parameters. Given these chal-lenges, we propose an effective multi-purpose module for dynamic channel sampling, namely Pick-or-Mix (PiX), which does not require special implementation. PiX divides a set of channels into subsets and then picks from them, where the picking decision is dynamically made per each pixel based on the input activations. We plug PiX into prominent ConvNet architectures and verify its multi-purpose utilities. After replacing$1\times 1$channel squeezing layers in ResNet with PiX, the network becomes 25% faster without losing accuracy. We show that PiX allows ConvNets to learn bet-ter data representation than widely adopted approaches to enhance networks' representation power (e.g., SE, CBAM, AFF, SKNet, and DWP). We also show that PiX achieves state-of-the-art performance on network downscaling and dynamic channel pruning applications. Code: https://github.com/ashishkumar822/PiX
Ashish Kumar 0006, Daneul Kim, Jaesik Park, Laxmidhar Behera
CVPR2
2023 Story Visualization by Online Text Augmentation with Context Memory
abstract
Story visualization (SV) is a challenging text-to-image generation task for the difficulty of not only rendering visual details from the text descriptions but also encoding a long-term context across multiple sentences. While prior efforts mostly focus on generating a semantically relevant image for each sentence, encoding a context spread across the given paragraph to generate contextually convincing images (e.g., with a correct character or with a proper background of the scene) remains a challenge. To this end, we propose a novel memory architecture for the Bi-directional Transformer framework with an online text augmentation that generates multiple pseudo-descriptions as supplementary supervision during training for better generalization to the language variation at inference. In extensive experiments on the two popular SV benchmarks, i.e., the Pororo-SV and Flintstones-SV, the proposed method significantly outperforms the state of the arts in various metrics including FID, character F1, frame accuracy, BLEU-2/3, and R-precision with similar or less computational complexity.
Daechul Ahn, Daneul Kim, Gwangmo Song, Honglak Lee, Dongyeop Kang
ICCV2