EDBT 2026 Demo / reviewers in the wild / expert
Sangmin Woo
dblp:295/8548
· DBLP profile ↗
19ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0003-4451-9675ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What and when to look? Temporal span proposal network for video relation detection
Sangmin Woo, Junhyug Noh, Kangil Kim |
Expert Syst. Appl. | 1 |
| 2025 | Diffusion Model Patching via Mixture-of-PromptsabstractWe present Diffusion Model Patching (DMP), a simple method to boost the performance of pre-trained diffusion models that have already reached convergence, with a negligible increase in parameters. DMP inserts a small, learnable set of prompts into the model's input space while keeping the original model frozen. The effectiveness of DMP is not merely due to the addition of parameters but stems from its dynamic gating mechanism, which selects and combines a subset of learnable prompts at every step of the generative process (i.e., reverse denoising steps). This strategy, which we term "mixture-of-prompts'', enables the model to draw on the distinct expertise of each prompt, essentially "patching'' the model's functionality at every step with minimal yet specialized parameters. Uniquely, DMP enhances the model by further training on the original dataset already used for training, even in a scenario where significant improvements are typically not expected due to model convergence. Experiments show that DMP significantly enhances the converged FID of DiT-L/2 on FFHQ by 10.38%, achieved with only a 1.43% parameter increase and 50K additional training iterations. Seokil Ham, Sangmin Woo, Hyojun Go, Byeongjun Park, Changick Kim |
AAAI | 2 |
| 2025 | Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear TransformationabstractDespite the growing interest in Mamba architecture as a potential replacement for Transformer architecture, parameter-efficient fine-tuning (PEFT) approaches for Mamba remain largely unexplored. In our study, we introduce two key insights-driven strategies for PEFT in Mamba architecture: (1) While state-space models (SSMs) have been regarded as the cornerstone of Mamba architecture, then expected to play a primary role in transfer learning, our findings reveal that Projectors—not SSMs—are the predominant contributors to transfer learning. (2) Based on our observation, we propose a novel PEFT method specialized to Mamba architecture: Projector-targeted Diagonalcentric Linear Transformation (ProDiaL). ProDiaL focuses on optimizing only the pretrained Projectors for new tasks through diagonal-centric linear transformation matrices, without directly fine-tuning the Projector weights. This targeted approach allows efficient task adaptation, utilizing less than 1% of the total parameters, and exhibits strong performance across both vision and language Mamba models, highlighting its versatility and effectiveness. Seokil Ham, Hee-Seon Kim, Sangmin Woo, Changick Kim |
CVPR | 3 |
| 2025 | A Systematic Survey of Automatic Prompt Optimization TechniquesabstractKiran Ramnath, Kang Zhou, Sheng Guan, Soumya Smruti Mishra, Xuan Qi, Zhengyuan Shen, Shuai Wang, Sangmin Woo, Sullam Jeoung, Yawei Wang, Haozhu Wang, Han Ding, Yuzhe Lu, Zhichao Xu, Yun Zhou, Balasubramaniam Srinivasan, Qiaojing Yan, Yueyan Chen, Haibo Ding, Panpan Xu, Lin Lee Cheong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Kiran Ramnath, Sheng Guan, Soumya Smruti Mishra, Xuan Qi, Zhengyuan Shen, Sangmin Woo, Sullam Jeoung, Haozhu Wang, Han Ding 0004, Yuzhe Lu, Zhichao Xu 0001, Qiaojing Yan, Yueyan Chen, Haibo Ding, Lin Lee Cheong |
EMNLP | 8 |
| 2025 | Modality mixer exploiting complementary information for multi-modal action recognition
Sangmin Woo, Muhammad Adi Nugroho, Changick Kim |
Comput. Vis. Image Underst. | 2 |
| 2024 | HarmonyView: Harmonizing Consistency and Diversity in One-Image-to-3DabstractRecent progress in single-image 3D generation highlights the importance of multi-view coherency, leveraging 3D priors from large-scale diffusion models pretrained on Internet-scale images. However, the aspect of novel-view diversity remains underexplored within the research landscape due to the ambiguity in converting a 2D image into 3D content, where numerous potential shapes can emerge. Here, we aim to address this research gap by simultaneously addressing both consistency and diversity. Yet, striking a balance between these two aspects poses a considerable challenge due to their inherent trade-offs. This work introduces HarmonyView, a simple yet effective diffusion sampling technique adept at decomposing two intricate aspects in single-image 3D generation: consistency and diversity. This approach paves the way for a more nuanced exploration of the two critical dimensions within the sampling process. Moreover, we propose a new evaluation metric based on CLIP image and text encoders to comprehensively assess the diversity of the generated views, which closely aligns with human evaluators' judgments. In experiments, HarmonyView achieves a harmonious balance, demonstrating a win-win scenario in both consistency and diversity. Sangmin Woo, Byeongjun Park, Hyojun Go, Changick Kim |
CVPR | 1 |
| 2024 | Spatio-Temporal Proximity-Aware Dual-Path Model for Panoramic Activity Recognition
Yooseung Wang, Sangmin Woo, Changick Kim |
ECCV (8) | 3 |
| 2024 | Flow-Assisted Motion Learning Network for Weakly-Supervised Group Activity Recognition
Muhammad Adi Nugroho, Sangmin Woo, Jinyoung Park 0001, Yooseung Wang, Changick Kim |
ECCV (48) | 2 |
| 2024 | Switch Diffusion Transformer: Synergizing Denoising Tasks with Sparse Mixture-of-Experts
Byeongjun Park, Hyojun Go, Sangmin Woo, Seokil Ham, Changick Kim |
ECCV (53) | 4 |
| 2024 | Denoising Task Routing for Diffusion ModelsabstractDiffusion models generate highly realistic images by learning a multi-step denoising process, naturally embodying the principles of multi-task learning (MTL). Despite the inherent connection between diffusion models and MTL, there remains an unexplored area in designing neural architectures that explicitly incorporate MTL into the framework of diffusion models. In this paper, we present Denoising Task Routing (DTR), a simple add-on strategy for existing diffusion model architectures to establish distinct information pathways for individual tasks within a single architecture by selectively activating subsets of channels in the model. What makes DTR particularly compelling is its seamless integration of prior knowledge of denoising tasks into the framework: (1) Task Affinity: DTR activates similar channels for tasks at adjacent timesteps and shifts activated channels as sliding windows through timesteps, capitalizing on the inherent strong affinity between tasks at adjacent timesteps. (2) Task Weights: During the early stages (higher timesteps) of the denoising process, DTR assigns a greater number of task-specific channels, leveraging the insight that diffusion models prioritize reconstructing global structure and perceptually rich contents in earlier stages, and focus on simple noise removal in later stages. Our experiments reveal that DTR not only consistently boosts diffusion models' performance across different evaluation protocols without adding extra parameters but also accelerates training convergence. Finally, we show the complementarity between our architectural approach and existing MTL optimization techniques, providing a more complete view of MTL in the context of diffusion training. Significantly, by leveraging this complementarity, we attain matched performance of DiT-XL using the smaller DiT-L with a reduction in training iterations from 7M to 2M. Our project page is available at https://byeongjun-park.github.io/DTR/ Byeongjun Park, Sangmin Woo, Hyojun Go, Changick Kim |
ICLR | 2 |
| 2024 | Sketch-based Video Object LocalizationabstractWe introduce Sketch-based Video Object Localization (SVOL), a new task aimed at localizing spatio-temporal object boxes in video queried by the input sketch. We first outline the challenges in the SVOL task and build the Sketch-Video Attention Network (SVANet) with the following design principles: (i) to consider temporal information of video and bridge the domain gap between sketch and video; (ii) to accurately identify and localize multiple objects simultaneously; (iii) to handle various styles of sketches; (iv) to be classification-free. In particular, SVANet is equipped with a Cross-modal Transformer that models the interaction between learnable object tokens, query sketch, and video through attention operations, and learns upon a per-frame set matching strategy that enables frame-wise prediction while utilizing global video context. We evaluate SVANet on a newly curated SVOL dataset. By design, SVANet successfully learns the mapping between the query sketches and video objects, achieving state-of-the-art results on the SVOL benchmark. We further confirm the effectiveness of SVANet via extensive ablation studies and visualizations. Lastly, we demonstrate its transfer capability on unseen datasets and novel categories, suggesting its high scalability in real-world applications. Codes are available at https://github.com/sangminwoo/SVOL. Sangmin Woo, So-Yeong Jeon, Jinyoung Park 0001, Minji Son, Changick Kim |
WACV | 1 |
| 2023 | Towards Good Practices for Missing Modality Robust Action RecognitionabstractStandard multi-modal models assume the use of the same modalities in training and inference stages. However, in practice, the environment in which multi-modal models operate may not satisfy such assumption. As such, their performances degrade drastically if any modality is missing in the inference stage. We ask: how can we train a model that is robust to missing modalities? This paper seeks a set of good practices for multi-modal action recognition, with a particular interest in circumstances where some modalities are not available at an inference time. First, we show how to effectively regularize the model during training (e.g., data augmentation). Second, we investigate on fusion methods for robustness to missing modalities: we find that transformer-based fusion shows better robustness for missing modality than summation or concatenation. Third, we propose a simple modular network, ActionMAE, which learns missing modality predictive coding by randomly dropping modality features and tries to reconstruct them with the remaining modality features. Coupling these good practices, we build a model that is not only effective in multi-modal action recognition but also robust to modality missing. Our model achieves the state-of-the-arts on multiple benchmarks and maintains competitive performances even in missing modality scenarios. Sangmin Woo, Yeonju Park, Muhammad Adi Nugroho, Changick Kim |
AAAI | 1 |
| 2023 | Audio-Visual Glance Network for Efficient Video RecognitionabstractDeep learning has made significant strides in video understanding tasks, but the computation required to classify lengthy and massive videos using clip-level video classifiers remains impractical and prohibitively expensive. To address this issue, we propose Audio-Visual Glance Network (AVGN), which leverages the commonly available audio and visual modalities to efficiently process the spatio-temporally important parts of a video. AVGN firstly divides the video into snippets of image-audio clip pair and employs lightweight unimodal encoders to extract global visual features and audio features. To identify the important temporal segments, we use an Audio-Visual Temporal Saliency Transformer (AV-TeST) that estimates the saliency scores of each frame. To further increase efficiency in the spatial dimension, AVGN processes only the important patches instead of the whole images. We use an Audio-Enhanced Spatial Patch Attention (AESPA) module to produce a set of enhanced coarse visual features, which are fed to a policy network that produces the coordinates of the important patches. This approach enables us to focus only on the most important spatio-temporally parts of the video, leading to more efficient video recognition. Moreover, we incorporate various training techniques and multi-modal feature fusion to enhance the robustness and effectiveness of our AVGN. By combining these strategies, our AVGN sets new state-of-the-art performance in multiple video recognition benchmarks while achieving faster processing speed. Muhammad Adi Nugroho, Sangmin Woo, Changick Kim |
ICCV | 2 |
| 2023 | Multi-modal Social Group Activity Recognition in Panoramic SceneabstractGroup Activity Recognition (GAR) is a challenging problem in computer vision due to the intricate dynamics and interactions among individuals. The existing methods utilize RGB videos face challenges in panoramic environments with numerous individuals and social groups. In this paper, we propose Multimodal Group Activity Recognition network (MGAR-net), that leverages the combined power of RGB and LiDAR modalities. Our approach effectively utilizes information from both modalities thus robustly and accurately captures individual relationships and detects social groups in face of optical challenges. By harnessing the capability of LiDAR with our new fusion module, called Distance Aware Fusion Module (DAFM), MGAR-net acquires valuable 3D structure information. We conduct experiments on the JRDB-Act dataset, which contains challenging scenarios with numerous people. The results demonstrate that LiDAR data provide valuable information for social grouping and recognizing individual action and group activities, particularly in crowded group settings. For social grouping, our MGAR-net improve performance by about 12% compared to the existing state-of-the-art models in terms of the AP metric. Sangmin Woo, Jinyoung Park 0001, Muhammad Adi Nugroho, Changick Kim |
VCIP | 3 |
| 2023 | AHFu-Net: Align, Hallucinate, and Fuse Network for Missing Multimodal Action RecognitionabstractIn this work, we explore the multimodal action recognition problem, specifically in the context of RGB-Depth modalities scenario, where a subset of the learning modalities is missing at inference time. To address this issue, we construct a hallucination network to generate missing modality information from the available modality at inference time. We propose key components of an effective spatio-temporal encoder for strong unimodal performance with Local Patch Temporal Transformer (LPTT) and Spatial Encoder Transformer (SET), alignment of multi-modal features, and fusion strategy with our Multimodal Bottleneck Transformer Fusion Module (MMBTF). We incorporate these ideas into a novel framework named AHFu-Net (Align, Hallucinate, and Fuse network) for RGB-Depth action recognition. Our experiments demonstrate that AHFu Net achieves state-of-the-art performance while maintaining high accuracy in the case of missing modality on multimodal datasets of NTU-RGB+D and NWUCLA. Muhammad Adi Nugroho, Sangmin Woo, Changick Kim |
VCIP | 2 |
| 2023 | Modality Mixer for Multi-modal Action RecognitionabstractIn multi-modal action recognition, it is important to consider not only the complementary nature of different modalities but also global action content. In this paper, we propose a novel network, named Modality Mixer (M-Mixer) network, to leverage complementary information across modalities and temporal context of an action for multi-modal action recognition. We also introduce a simple yet effective recurrent unit, called Multi-modal Contextualization Unit (MCU), which is a core component of M-Mixer. Our MCU temporally encodes a sequence of one modality (e.g., RGB) with action content features of other modalities (e.g., depth, IR). This process encourages M-Mixer to exploit global action content and also to supplement complementary information of other modalities. As a result, our proposed method outperforms state-of-the-art methods on NTU RGB+D 60, NTU RGB+D 120, and NW-UCLA datasets. Moreover, we demonstrate the effectiveness of M-Mixer by conducting comprehensive ablation studies. Sangmin Woo, Yeonju Park, Muhammad Adi Nugroho, Changick Kim |
WACV | 2 |
| 2023 | Cross-modal alignment and translation for missing modality action recognition
Yeonju Park, Sangmin Woo, Muhammad Adi Nugroho, Changick Kim |
Comput. Vis. Image Underst. | 2 |
| 2023 | Tackling the Challenges in Scene Graph Generation With Local-to-Global InteractionsabstractIn this work, we seek new insights into the underlying challenges of the scene graph generation (SGG) task. Quantitative and qualitative analysis of the visual genome (VG) dataset implies: 1) ambiguity: even if interobject relationship contains the same object (or predicate), they may not be visually or semantically similar; 2) asymmetry: despite the nature of the relationship that embodied the direction, it was not well addressed in previous studies; and 3) higher-order contexts: leveraging the identities of certain graph elements can help generate accurate scene graphs. Motivated by the analysis, we design a novel SGG framework, Local-to-global interaction networks (LOGINs). Locally, interactions extract the essence between three instances of subject, object, and background, while baking direction awareness into the network by explicitly constraining the input order of subject and object. Globally, interactions encode the contexts between every graph component (i.e., nodes and edges). Finally, Attract and Repel loss is utilized to fine-tune the distribution of predicate embeddings. By design, our framework enables predicting the scene graph in a bottom-up manner, leveraging the possible complementariness. To quantify how much LOGIN is aware of relational direction, a new diagnostic task called Bidirectional Relationship Classification (BRC) is also proposed. Experimental results demonstrate that LOGIN can successfully distinguish relational direction than existing methods (in BRC task), while showing state-of-the-art results on the VG benchmark (in SGG task). Sangmin Woo, Junhyug Noh, Kangil Kim |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Temporal Flow Mask Attention for Open-Set Long-Tailed Recognition of Wild Animals in Camera-Trap ImagesabstractCamera traps, unmanned observation devices, and deep learning-based image recognition systems have greatly reduced human effort in collecting and analyzing wildlife images. However, data collected via above apparatus exhibits 1) long-tailed and 2) open-ended distribution problems. To tackle the open-set long-tailed recognition problem, we propose the Temporal Flow Mask Attention Network that comprises three key building blocks: 1) an optical flow module, 2) an attention residual module, and 3) a meta-embedding classifier. We extract temporal features of sequential frames using the optical flow module and learn informative representation using attention residual blocks. Moreover, we show that applying the meta-embedding technique boosts the performance of the method in open-set long-tailed recognition. We apply this method on a Korean De-militarized Zone (DMZ) dataset. We conduct extensive experiments, and quantitative and qualitative analyses to prove that our method effectively tackles the open-set long-tailed recognition problem while being robust to unknown classes. Jeongsoo Kim, Sangmin Woo, Byeongjun Park, Changick Kim |
ICIP | 2 |