EDBT 2026 Demo / reviewers in the wild / expert
Jungwon Park
dblp:248/8135
· DBLP profile ↗
14ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 9 since 2021Systems, architecture and hardware · 5 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DOS: Directional Object Separation in Text Embeddings for Multi-Object Image GenerationabstractRecent progress in text-to-image (T2I) generative models has led to significant improvements in generating high-quality images aligned with text prompts. However, these models still struggle with prompts involving multiple objects, often resulting in object neglect or object mixing. Through extensive studies, we identify four problematic scenarios, Similar Shapes, Similar Textures, Dissimilar Background Biases, and Many Objects, where inter-object relationships frequently lead to such failures. Motivated by two key observations about CLIP embeddings, we propose DOS (Directional Object Separation), a method that modifies three types of CLIP text embeddings before passing them into text-to-image models. Experimental results show that DOS consistently improves the success rate of multi-object image generation and reduces object mixing. In human evaluations, DOS significantly outperforms four competing methods, receiving 26.24%-43.04% more votes across four benchmarks. These results highlight DOS as a practical and effective solution for improving multi-object image generation. Dongnam Byun, Jungwon Park, Jungmin Ko, Changin Choi, Wonjong Rhee |
AAAI | 2 |
| 2025 | Task-Specific Preconditioner for Cross-Domain Few-Shot LearningabstractCross-Domain Few-Shot Learning (CDFSL) methods typically parameterize models with task-agnostic and task-specific parameters. To adapt task-specific parameters, recent approaches have utilized fixed optimization strategies, despite their potential sub-optimality across varying domains or target tasks. To address this issue, we propose a novel adaptation mechanism called Task-Specific Preconditioned gradient descent (TSP). Our method first meta-learns Domain-Specific Preconditioners (DSPs) that capture the characteristics of each meta-training domain, which are then linearly combined using task-coefficients to form the Task-Specific Preconditioner. The preconditioner is applied to gradient descent, making the optimization adaptive to the target task. We constrain our preconditioners to be positive definite, guiding the preconditioned gradient toward the direction of steepest descent. Empirical evaluations on the Meta-Dataset show that TSP achieves state-of-the-art performance across diverse experimental scenarios. Suhyun Kang, Jungwon Park, Wonseok Lee 0002, Wonjong Rhee |
AAAI | 2 |
| 2025 | Hearable Image: On-Device Image-Driven Sound Effect Generation for Hearing What You SeeabstractThere have been various studies in audio generation from image, text, or video. However, the existing approaches have not consider on-device environment because audio generation models are computationally expensive and require heavy storage capacity to save large number of weights. In addition, it is difficult to get stable generation outputs because unexpected results may occur depending on various model inputs. In image-to-audio generation, there are diverse images in smartphones, and too many visual contexts are contained in image features. Therefore, it is sometimes unpredictable which audio categories are generated from images. In this paper, we propose a robust on-device sound effect generation framework that is image-to-audio generation based on latent diffusion. First, to avoid unstable and unpredictable audio generation results, we propose a stable sound generation framework with Audio Feature Dictionary and Audio-Image Matching Pipeline to generate sound effects from predefined sound effect categories. If an image matches to sound effect categories, proposed framework directly generate sound effects from audio features corresponding to the matched categories. Second, we propose Multi-Category Generation and Generation Flow Map to generate robust and diverse sound effects depending on audio categories. Using global and local features of an image, we can select multiple categories of sound effects. Third, the framework can be implemented in smartphone devices because we train the proposed model with low computational cost and small number of model weights under 4-step latent diffusion inference. Various experiments show the proposed framework solves on-device sound generation problem with maintaining generation quality and audio-image matching performances compared to large scale models. Our demo is available at: https://youtu.be/Y5HTr8wwqOA. Deokjun Eom, Nahyun Kim, Woo Hyun Nam, Kyung-Rae Kim, Chaebin Im, Jungwon Park |
CIKM | 6 |
| 2025 | ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention AdaptationabstractRectified Flow text-to-image models surpass diffusion models in image quality and text alignment, but adapting ReFlow for real-image editing remains challenging. We propose a new real-image editing method for ReFlow by analyzing the intermediate representations of multimodal transformer blocks and identifying three key features. To extract these features from real images with sufficient structural preservation, we leverage mid-step latent, which is inverted only up to the mid-step. We then adapt attention during injection to improve editability and enhance alignment to the target text. Our method is training-free, requires no user-provided mask, and can be applied even without a source prompt. Extensive experiments on two benchmarks with nine baselines demonstrate its superior performance over prior methods, further validated by human evaluations confirming a strong user preference for our approach. Jimyeong Kim, Jungwon Park, Yeji Song, Nojun Kwak, Wonjong Rhee |
ICCV | 2 |
| 2025 | Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative ModelsabstractRecent text-to-image diffusion models leverage cross-attention layers, which have been effectively utilized to enhance a range of visual generative tasks. However, our understanding of cross-attention layers remains somewhat limited. In this study, we introduce a mechanistic interpretability approach for diffusion models by constructing Head Relevance Vectors (HRVs) that align with human-specified visual concepts. An HRV for a given visual concept has a length equal to the total number of cross-attention heads, with each element representing the importance of the corresponding head for the given visual concept. To validate HRVs as interpretable features, we develop an ordered weakening analysis that demonstrates their effectiveness. Furthermore, we propose concept strengthening and concept adjusting methods and apply them to enhance three visual generative tasks. Our results show that HRVs can reduce misinterpretations of polysemous words in image generation, successfully modify five challenging attributes in image editing, and mitigate catastrophic neglect in multi-concept generation. Overall, our work provides an advancement in understanding cross-attention layers and introduces new approaches for fine-controlling these layers at the head level. Jungwon Park, Jungmin Ko, Dongnam Byun, Jangwon Suh, Wonjong Rhee |
ICLR | 1 |
| 2025 | QP Chaser: Polynomial Trajectory Generation for Autonomous Aerial TrackingabstractMaintaining the visibility of the target is one of the major objectives of aerial tracking missions. This paper proposes a target-visible trajectory planning pipeline using quadratic programming (QP). Our approach can handle various tracking settings, including 1) single- and dual-target following and 2) both static and dynamic environments, unlike other works that focus on a single specific setup. In contrast to other studies that fully trust the predicted trajectory of the target and consider only the visibility of the target’s center, our pipeline considers error in target path prediction and the entire body of the target to maintain the target visibility robustly. First, a prediction module uses a sample-check strategy to quickly calculate the reachable areas of moving objects, which represent the areas their bodies can reach, considering obstacles. Subsequently, the planning module formulates a single QP problem, considering path homotopy, to generate a tracking trajectory that maximizes the visibility of the target’s reachable area among obstacles. The performance of the planner is validated in multiple scenarios, through high-fidelity simulations and real-world experiments. Yunwoo Lee, Jungwon Park, Seungwoo Jung, Boseong Jeon, Dahyun Oh, H. Jin Kim |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image PersonalizationabstractIn text-to-image personalization, a timely and crucial challenge is the tendency of generated images overfitting to the biases present in the reference images. We initiate our study with a comprehensive categorization of the biases into background, nearby-object, tied-object, substance (in style re-contextualization), and pose biases. These biases manifest in the generated images due to their entanglement into the subject embedding. This undesired embedding entanglement not only results in the reflection of biases from the reference images into the generated images but also notably diminishes the alignment of the generated images with the given generation prompt. To address this challenge, we propose SID (Selectively Informative Description), a text description strategy that deviates from the prevalent approach of only characterizing the subject's class identification. SID is generated utilizing multimodal GPT-4 and can be seamlessly integrated into optimization-based models. We present comprehensive experimental results along with analyses of cross-attention maps, subject-alignment, non-subject-disentanglement, and text-alignment. Jimyeong Kim, Jungwon Park, Wonjong Rhee |
CVPR | 2 |
| 2023 | Safe and Distributed Multi-Agent Motion Planning under Minimum Speed ConstraintsabstractThe motion planning problem for multiple unstop-pable agents is of interest in many robotics applications, for example, autonomous traffic management for multiple fixed-wing aircraft. Unfortunately, many of the existing algorithms cannot provide safety for such agents, because they require the agents to be able to brake to a complete stop for safety and feasibility insurance. In this paper, we present a distributed multi-agent motion planner that guarantees collision avoidance and persistent feasibility, which can be applied to a team of homogeneous mobile vehicles that cannot stop. The planner is built on top of the idea that a collision-free trajectory in form of a loop can safely accommodate multiple unstoppable agents, while avoiding collisions among them and static obstacles. At every time step, in a distributed manner, the agents generate trajectory-manipulating actions that preserve the loop structure. Then, a deconfliction process selects a conflict-free subset of the generated actions, which are applied at the next time step. Through simulation using an unstoppable Dubins car model, we show that the proposed motion planner is able to provide persistent safety guarantees for such agents in obstacle-cluttered space in real-time. Inkyu Jang, Jungwon Park, H. Jin Kim |
ICRA | 2 |
| 2023 | Decentralized Deadlock-free Trajectory Planning for Quadrotor Swarm in Obstacle-rich EnvironmentsabstractThis paper presents a decentralized multi-agent trajectory planning (MATP) algorithm that guarantees to generate a safe, deadlock-free trajectory in an obstacle-rich environment under a limited communication range. The proposed algorithm utilizes a grid-based multi-agent path planning (MAPP) algorithm for deadlock resolution, and we introduce the subgoal optimization method to make the agent converge to the waypoint generated from the MAPP without deadlock. In addition, the proposed algorithm ensures the feasibility of the optimization problem and collision avoidance by adopting a linear safe corridor (LSC). We verify that the proposed algorithm does not cause a deadlock in both random forests and dense mazes regardless of communication range, and it outperforms our previous work in flight time and distance. We validate the proposed algorithm through a hardware demonstration with ten quadrotors. Jungwon Park, Inkyu Jang, H. Jin Kim |
ICRA | 1 |
| 2023 | Isotropic Representation Can Improve Dense Retrieval
Euna Jung, Jungwon Park, Jaekeol Choi, Sungyoon Kim, Wonjong Rhee |
PAKDD (3) | 2 |
| 2023 | DLSC: Distributed Multi-Agent Trajectory Planning in Maze-Like Dynamic Environments Using Linear Safe CorridorabstractThis article presents an online distributed trajectory planning algorithm for a quadrotor swarm in a maze-like dynamic environment. We utilize a dynamic linear safe corridor to construct the feasible collision constraints that can ensure interagent collision avoidance and consider the uncertainty of moving obstacles. We introduce mode-based subgoal planning to resolve deadlock faster in a complex environment using only previously shared information. For dynamic obstacle avoidance, we adopt heuristic methods such as collision alert propagation and escape point planning to deal with the situation where dynamic obstacles approach the agents clustered in a narrow corridor. We prove that the proposed algorithm guarantees the feasibility of the optimization problem for every replanning step. In an obstacle-free space, the proposed method can compute the trajectories for 60 agents on average 7.66 ms per agent with an Intel i7 laptop and shows the perfect success rate. Also, our method shows 64.5$\%$shorter flight time than buffered Voronoi cell and 34.6$\%$shorter than with our previous work. We conduct the simulation in a random forest and maze with four dynamic obstacles, and the proposed algorithm shows the highest success rate and shortest flight time compared to state-of-the-art baseline algorithms. In particular, the proposed algorithm shows over 97$\%$success rate when the velocity of moving obstacles is below the agent's maximum speed. We validate the safety and robustness of the proposed algorithm through a hardware demonstration with ten quadrotors and two pedestrians in a maze-like environment. Jungwon Park, Yunwoo Lee, Inkyu Jang, H. Jin Kim |
IEEE Trans. Robotics | 1 |
| 2021 | Target-visible Polynomial Trajectory Generation within an MAV TeamabstractAutonomous aerial videography is a challenging task, which involves collision avoidance against obstacles and visibility guaranteed target tracking in unstructured environments. In this paper, we organize a two micro aerial vehicle (MAV) team, which consists of a target agent responsible for a specific mission and a camera agent for filming the target agent. Especially, this paper focuses on trajectory planning of the camera agent to chase without occlusion of target agent. Our trajectory planner module includes two phases of guaranteeing target visibility. In the first phase, we generate homotopic safe flight corridor (SFC) to attain target-visible regions. In the subsequent phase, we generate a safe and smooth trajectory with the continuous visibility constraint based on the SFC, using quadratic programming (QP). Regardless of complexity of map, our planner converts an overall problem to a single QP and generates a steady flight trajectory without undesirable fluctuating motion, while guaranteeing all-time visibility. We validate our approach in Gazebo simulations and a real-world experiment. Yunwoo Lee, Jungwon Park, Boseong Jeon, H. Jin Kim |
IROS | 2 |
| 2020 | Efficient Multi-Agent Trajectory Planning with Feasibility Guarantee using Relative Bernstein PolynomialabstractThis paper presents a new efficient algorithm which guarantees a solution for a class of multi-agent trajectory planning problems in obstacle-dense environments. Our algorithm combines the advantages of both grid-based and optimization-based approaches, and generates safe, dynamically feasible trajectories without suffering from an erroneous optimization setup such as imposing infeasible collision constraints. We adopt a sequential optimization method with dummy agents to improve the scalability of the algorithm, and utilize the convex hull property of Bernstein and relative Bernstein polynomial to replace non-convex collision avoidance constraints to convex ones. The proposed method can compute the trajectory for 64 agents on average 6.36 seconds with Intel Core i7-7700 @ 3.60GHz CPU and 16G RAM, and it reduces more than 50% of the objective cost compared to our previous work. We validate the proposed algorithm through simulation and flight tests. Jungwon Park, Junha Kim, Inkyu Jang, H. Jin Kim |
ICRA | 1 |
| 2019 | Fast Trajectory Planning for Multiple Quadrotors using Relative Safe Flight CorridorabstractThis paper presents a new trajectory planning method for multiple quadrotors in obstacle-dense environments. We suggest a relative safe flight corridor (RSFC) to model safe region between a pair of agents, and it is used to generate linear constraints for inter-collision avoidance by utilizing the convex hull property of relative Bernstein polynomial. Our approach employs a graph-based multi-agent pathfinding algorithm to generate an initial trajectory, which is used to construct a safe flight corridor (SFC) and RSFC. We express the trajectory as a piecewise Bernstein polynomial and formulate the trajectory planning problem into one quadratic programming problem using linear constraints from SFC and RSFC. The proposed method can compute collision-free trajectory for 16 agents within a second and for 64 agents less than a minute, and it is validated both through simulation and indoor flight test. Jungwon Park, H. Jin Kim |
IROS | 1 |