VLDB 2026 Research / reviewers in the wild / expert
Dongnam Byun
dblp:394/6600
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Generative modeling · 65% Trustworthy machine learning · 35% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Visual content generation and editing › image generation
text-to-image generation |
1.0 | 1 | 2026 | DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation · AAAI 2026 |
Machine learning › Trustworthy machine learning › interpretability
concept intervention |
0.9 | 1 | 2025 | Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models · ICLR 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
image editing |
0.9 | 1 | 2025 | Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models · ICLR 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2025 | Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models · ICLR 2025 |
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
0.3 | 1 | 2025 | Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
CLIP embedding manipulation · 1.0ordered weakening analysis · 0.9head relevance vector · 0.9concept strengthening · 0.9concept adjusting · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DOS: Directional Object Separation in Text Embeddings for Multi-Object Image GenerationabstractRecent progress in text-to-image (T2I) generative models has led to significant improvements in generating high-quality images aligned with text prompts. However, these models still struggle with prompts involving multiple objects, often resulting in object neglect or object mixing. Through extensive studies, we identify four problematic scenarios, Similar Shapes, Similar Textures, Dissimilar Background Biases, and Many Objects, where inter-object relationships frequently lead to such failures. Motivated by two key observations about CLIP embeddings, we propose DOS (Directional Object Separation), a method that modifies three types of CLIP text embeddings before passing them into text-to-image models. Experimental results show that DOS consistently improves the success rate of multi-object image generation and reduces object mixing. In human evaluations, DOS significantly outperforms four competing methods, receiving 26.24%-43.04% more votes across four benchmarks. These results highlight DOS as a practical and effective solution for improving multi-object image generation. Dongnam Byun, Jungwon Park, Jungmin Ko, Changin Choi, Wonjong Rhee |
AAAI | 1 |
| 2025 | Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative ModelsabstractRecent text-to-image diffusion models leverage cross-attention layers, which have been effectively utilized to enhance a range of visual generative tasks. However, our understanding of cross-attention layers remains somewhat limited. In this study, we introduce a mechanistic interpretability approach for diffusion models by constructing Head Relevance Vectors (HRVs) that align with human-specified visual concepts. An HRV for a given visual concept has a length equal to the total number of cross-attention heads, with each element representing the importance of the corresponding head for the given visual concept. To validate HRVs as interpretable features, we develop an ordered weakening analysis that demonstrates their effectiveness. Furthermore, we propose concept strengthening and concept adjusting methods and apply them to enhance three visual generative tasks. Our results show that HRVs can reduce misinterpretations of polysemous words in image generation, successfully modify five challenging attributes in image editing, and mitigate catastrophic neglect in multi-concept generation. Overall, our work provides an advancement in understanding cross-attention layers and introduces new approaches for fine-controlling these layers at the head level. Jungwon Park, Jungmin Ko, Dongnam Byun, Jangwon Suh, Wonjong Rhee |
ICLR | 3 |