Sangwon Kim 0004

dblp:37/320-4 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-7452-3897ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 ZeRA: Zero-Reindex Multimodal RAG via Heterogeneous Embedding Alignment for Lightweight Query Encoding
Dasom Ahn, Hye Rim Kim, Sangwon Kim 0004, Kwang-Ju Kim, ByoungChul Ko
ICPR (11)4
2026 Energy-ensemble concept bottleneck models for enhancing interpretability and accuracy in concept-based learning
Dasom Ahn, Sangwon Kim 0004, ByoungChul Ko
Knowl. Based Syst.2
2025 LAttE: A label-free and multimodal framework for context-aware person re-identification
Dasom Ahn, Sangwon Kim 0004, Kwang-Ju Kim, ByoungChul Ko
Neurocomputing2
2024 EQ-CBM: A Probabilistic Concept Bottleneck with Energy-Based Models and Quantized Vectors
Sangwon Kim 0004, Dasom Ahn, ByoungChul Ko, In-Su Jang, Kwang-Ju Kim
ACCV (7)1
2024 Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency
abstract
Scene graph generation (SGG) is an important task in image understanding because it represents the relationships between objects in an image as a graph structure, making it possible to understand the semantic relationships between objects intuitively. Previous SGG studies used a message-passing neural networks (MPNN) to update features, which can effectively reflect information about surrounding objects. However, these studies have failed to reflect the co-occurrence of objects during SGG generation. In addition, they only addressed the long-tail problem of the training dataset from the perspectives of sampling and learning methods. To address these two problems, we propose CooK, which reflects the Co-occurrence Knowledge between objects, and the learnable term frequency-inverse document frequency (TF-$l$-IDF) to solve the long-tail problem. We applied the proposed model to the SGG benchmark dataset, and the results showed a performance improvement of up to 3.8% compared with existing state-of-the-art models in SGGen subtask. The proposed method exhibits generalization ability from the results obtained, showing uniform performance improvement for all MPNN models.
Sangwon Kim 0004, Dasom Ahn, Jong Taek Lee, ByoungChul Ko
ICML2
2024 BTD-RF: 3D scene reconstruction using block-term tensor decomposition
Seon Bin Kim, Sangwon Kim 0004, Dasom Ahn, ByoungChul Ko
Appl. Intell.2
2024 Domain-free fire detection using the spatial-temporal attention transform of the YOLO backbone
Sangwon Kim 0004, In-Su Jang, ByoungChul Ko
Pattern Anal. Appl.1
2023 Cross-Modal Learning with 3D Deformable Attention for Action Recognition
abstract
An important challenge in vision-based action recognition is the embedding of spatiotemporal features with two or more heterogeneous modalities into a single feature. In this study, we propose a new 3D deformable transformer for action recognition with adaptive spatiotemporal receptive fields and a cross-modal learning scheme. The 3D deformable transformer consists of three attention modules: 3D deformability, local joint stride, and temporal stride attention. The two cross-modal tokens are input into the 3D deformable attention module to create a cross-attention token with a reflected spatiotemporal correlation. Local joint stride attention is applied to spatially combine attention and pose tokens. Temporal stride attention temporally reduces the number of input tokens in the attention module and supports temporal expression learning without the simultaneous use of all tokens. The deformable transformer iterates L-times and combines the last cross-modal token for classification. The proposed 3D deformable transformer was tested on the NTU60, NTU120, FineGYM, and PennAction datasets, and showed results better than or similar to pre-trained state-of-the-art methods even without a pre-training process. In addition, by visualizing important joints and correlations during action recognition through spatial joint and temporal stride attention, the possibility of achieving an explainable potential for action recognition is presented.
Sangwon Kim 0004, Dasom Ahn, ByoungChul Ko
ICCV1
2023 STAR-Transformer: A Spatio-temporal Cross Attention Transformer for Human Action Recognition
abstract
In action recognition, although the combination of spatiotemporal videos and skeleton features can improve the recognition performance, a separate model and balancing feature representation for cross-modal data are required. To solve these problems, we propose Spatio-TemporAl cRoss (STAR)-transformer, which can effectively represent two cross-modal features as a recognizable vector. First, from the input video and skeleton sequence, video frames are output as global grid tokens and skeletons are output as joint map tokens, respectively. These tokens are then aggregated into multi-class tokens and input into STAR-transformer. The STAR-transformer encoder consists of a full spatio-temporal attention (FAttn) module and a proposed zigzag spatio-temporal attention (ZAttn) module. Similarly, the continuous decoder consists of a FAttn module and a proposed binary spatio-temporal attention (BAttn) module. STAR-transformer learns an efficient multi-feature representation of the spatio-temporal features by properly arranging pairings of the FAttn, ZAttn, and BAttn modules. Experimental results on the Penn-Action, NTU-RGB+D 60, and 120 datasets show that the proposed method achieves a promising improvement in performance in comparison to previous state-of-the-art methods.
Dasom Ahn, Sangwon Kim 0004, Hyunsu Hong, ByoungChul Ko
WACV2
2023 STAR++: Rethinking spatio-temporal cross attention transformer for video action recognition
Dasom Ahn, Sangwon Kim 0004, ByoungChul Ko
Appl. Intell.2
2023 SSL-MOT: self-supervised learning based multi-object tracking
Sangwon Kim 0004, Jimi Lee, ByoungChul Ko
Appl. Intell.1
2022 ViT-NeT: Interpretable Vision Transformers with Neural Tree Decoder
abstract
Vision transformers (ViTs), which have demonstrated a state-of-the-art performance in image classification, can also visualize global interpretations through attention-based contributions. However, the complexity of the model makes it difficult to interpret the decision-making process, and the ambiguity of the attention maps can cause incorrect correlations between image patches. In this study, we propose a new ViT neural tree decoder (ViT-NeT). A ViT acts as a backbone, and to solve its limitations, the output contextual image patches are applied to the proposed NeT. The NeT aims to accurately classify fine-grained objects with similar inter-class correlations and different intra-class correlations. In addition, it describes the decision-making process through a tree structure and prototype and enables a visual interpretation of the results. The proposed ViT-NeT is designed to not only improve the classification performance but also provide a human-friendly interpretation, which is effective in resolving the trade-off between performance and interpretability. We compared the performance of ViT-NeT with other state-of-art methods using widely used fine-grained visual categorization benchmark datasets and experimentally proved that the proposed method is superior in terms of the classification performance and interpretability. The code and models are publicly available at https://github.com/jumpsnack/ViT-NeT.
Sangwon Kim 0004, Jae-Yeal Nam, ByoungChul Ko
ICML1
2022 Lightweight surrogate random forest support for model simplification and feature relevance
Sangwon Kim 0004, Mira Jeong, ByoungChul Ko
Appl. Intell.1