Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Dongnam Byun

dblp:394/6600 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Generative modeling · 65% Trustworthy machine learning · 35%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing › image generation
text-to-image generation
1.012026
DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation · AAAI 2026
Machine learning › Trustworthy machine learning › interpretability
concept intervention
0.912025
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models · ICLR 2025
Machine learning › Generative modeling
diffusion model
0.912025
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models · ICLR 2025
Machine learning › Generative modeling › diffusion model
image editing
0.912025
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models · ICLR 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models · ICLR 2025
Machine learning › Trustworthy machine learning
interpretability
0.312025
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models · ICLR 2025
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.312025
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models · ICLR 2025

Methods — techniques the papers use, named apart from their topics

CLIP embedding manipulation · 1.0ordered weakening analysis · 0.9head relevance vector · 0.9concept strengthening · 0.9concept adjusting · 0.9
YearPublicationVenuePosition
2026 DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
abstract
Recent progress in text-to-image (T2I) generative models has led to significant improvements in generating high-quality images aligned with text prompts. However, these models still struggle with prompts involving multiple objects, often resulting in object neglect or object mixing. Through extensive studies, we identify four problematic scenarios, Similar Shapes, Similar Textures, Dissimilar Background Biases, and Many Objects, where inter-object relationships frequently lead to such failures. Motivated by two key observations about CLIP embeddings, we propose DOS (Directional Object Separation), a method that modifies three types of CLIP text embeddings before passing them into text-to-image models. Experimental results show that DOS consistently improves the success rate of multi-object image generation and reduces object mixing. In human evaluations, DOS significantly outperforms four competing methods, receiving 26.24%-43.04% more votes across four benchmarks. These results highlight DOS as a practical and effective solution for improving multi-object image generation.
Dongnam Byun, Jungwon Park, Jungmin Ko, Changin Choi, Wonjong Rhee
AAAI1
2025 Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models
abstract
Recent text-to-image diffusion models leverage cross-attention layers, which have been effectively utilized to enhance a range of visual generative tasks. However, our understanding of cross-attention layers remains somewhat limited. In this study, we introduce a mechanistic interpretability approach for diffusion models by constructing Head Relevance Vectors (HRVs) that align with human-specified visual concepts. An HRV for a given visual concept has a length equal to the total number of cross-attention heads, with each element representing the importance of the corresponding head for the given visual concept. To validate HRVs as interpretable features, we develop an ordered weakening analysis that demonstrates their effectiveness. Furthermore, we propose concept strengthening and concept adjusting methods and apply them to enhance three visual generative tasks. Our results show that HRVs can reduce misinterpretations of polysemous words in image generation, successfully modify five challenging attributes in image editing, and mitigate catastrophic neglect in multi-concept generation. Overall, our work provides an advancement in understanding cross-attention layers and introduces new approaches for fine-controlling these layers at the head level.
Jungwon Park, Jungmin Ko, Dongnam Byun, Jangwon Suh, Wonjong Rhee
ICLR3