Chull Hwan Song

dblp:77/1371 · DBLP profile ↗
← Back
8ranked-venue papers
7as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 6 since 2021
YearPublicationVenuePosition
2024 SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
abstract
Vision-language models (VLMs) have made significant strides in cross-modal understanding through large-scale paired datasets. However, in fashion domain, datasets often exhibit a disparity between the information conveyed in image and text. This issue stems from datasets containing multiple images of a single fashion item all paired with one text, leading to cases where some textual details are not visible in individual images. This mismatch, particularly when non-co-occurring elements are masked, undermines the training of conventional VLM objectives like Masked Language Modeling and Masked Image Modeling, thereby hindering the model's ability to accurately align fine-grained visual and textual features. Addressing this problem, we propose Synchronized attentional Masking (SyncMask), which generate masks that pinpoint the image patches and word tokens where the information co-occur in both image and text. This synchronization is accomplished by harnessing cross-attentional features obtained from a momentum model, ensuring a precise alignment between the two modalities. Additionally, we enhance grouped batch sampling with semi-hard negatives, effectively mitigating false negative issues in Image-Text Matching and Image-Text Contrastive learning objectives within fashion datasets. Our experiments demonstrate the effectiveness of the proposed approach, outper-forming existing methods in three downstream tasks.
Chull Hwan Song, Taebaek Hwang, Jooyoung Yoon, Shunghyun Choi, Yeong Hyeon Gu
CVPR1
2024 On Train-Test Class Overlap and Detection for Image Retrieval
abstract
How important is it for training and evaluation sets to not have class overlap in image retrieval? We revisit Google Landmarks v2 clean [56], the most popular training set, by identifying and removing class overlap with Revisited Oxford and Paris [34], the most popular evaluation set. By comparing the original and the new RGLDv2-clean on a benchmark of reproduced state-of-the-art methods, our findings are striking. Not only is there a dramatic drop in performance, but it is inconsistent across methods, changing the ranking. What does it take to focus on objects or interest and ignore background clutter when indexing? Do we need to train an object detector and the representation separately? Do we need location supervision? We introduce Single-stage Detect-to-Retrieve (CiDeR), an end-to-end, single-stage pipeline to detect objects of interest and extract a global image representation. We outperform previous state-of-the-art on both existing training sets and the new RGLDv2-clean. Our dataset is available at https://github.com/dealicious-inc/RGLDv2-clean.
Chull Hwan Song, Jooyoung Yoon, Taebaek Hwang, Shunghyun Choi, Yeong Hyeon Gu, Yannis Avrithis
CVPR1
2023 Conditional Cross Attention Network for Multi-Space Embedding without Entanglement in Only a SINGLE Network
abstract
Many studies in vision tasks have aimed to create effective embedding spaces for single-label object prediction within an image. However, in reality, most objects possess multiple specific attributes, such as shape, color, and length, with each attribute composed of various classes. To apply models in real-world scenarios, it is essential to be able to distinguish between the granular components of an object. Conventional approaches to embedding multiple specific attributes into a single network often result in entanglement, where fine-grained features of each attribute cannot be identified separately. To address this problem, we propose a Conditional Cross-Attention Network that induces disentangled multi-space embeddings for various specific attributes with only a single backbone. Firstly, we employ a cross-attention mechanism to fuse and switch the information of conditions (specific attributes), and we demonstrate its effectiveness through a diverse visualization example. Secondly, we leverage the vision transformer for the first time to a fine-grained image retrieval task and present a simple yet effective framework compared to existing methods. Unlike previous studies where performance varied depending on the benchmark dataset, our proposed method achieved consistent state-of-the-art performance on the FashionAI, DARN, DeepFashion, and Zappos50K benchmark datasets.
Chull Hwan Song, Taebaek Hwang, Jooyoung Yoon, Shunghyun Choi, Yeong Hyeon Gu
ICCV1
2023 Boosting vision transformers for image retrieval
abstract
Vision transformers have achieved remarkable progress in vision tasks such as image classification and detection. However, in instance-level image retrieval, transformers have not yet shown good performance compared to convolutional networks. We propose a number of improvements that make transformers outperform the state of the art for the first time. (1) We show that a hybrid architecture is more effective than plain transformers, by a large margin. (2) We introduce two branches collecting global (classification token) and local (patch tokens) information, from which we form a global image representation. (3) In each branch, we collect multilayer features from the transformer encoder, corresponding to skip connections across distant layers. (4) We enhance locality of interactions at the deeper layers of the encoder, which is the relative weakness of vision transformers. We train our model on all commonly used training sets and, for the first time, we make fair comparisons separately per training set. In all cases, we outperform previous models based on global representation. Public code is available at https://github.com/dealicious-inc/DToP.
Chull Hwan Song, Jooyoung Yoon, Shunghyun Choi, Yannis Avrithis
WACV1
2022 Convolutional Attribute Mask with Two-step Attention for Fashion Image Retrieval
abstract
We propose a method to learn multiple latent spaces for attribute-specific fashion image retrieval. Our network learns multiple deep image features for a given set of fashion attributes through convolutional attribute masks and two-step attention. The masks promote our network to learn image features for a specific attribute. The two-step attention helps our network collect important spatial and channel information for the fashion attribute using spatial or channel attention. We visually show that our network correctly attend to essential regions for a given fashion attribute and learns well-distanced embeddings in latent spaces. We achieve state-of-the-art performance for attribute-specific fashion image retrieval.
Chull Hwan Song, Hye Joo Han
ICPR1
2022 All the attention you need: Global-local, spatial-channel attention for image retrieval
abstract
We address representation learning for large-scale instance-level image retrieval. Apart from backbone, training pipelines and loss functions, popular approaches have focused on different spatial pooling and attention mechanisms, which are at the core of learning a powerful global image representation. There are different forms of attention according to the interaction of elements of the feature tensor (local and global) and the dimensions where it is applied (spatial and channel). Unfortunately, each study addresses only one or two forms of attention and applies it to different problems like classification, detection or retrieval.We present global-local attention module (GLAM), which is attached at the end of a backbone network and incorporates all four forms of attention: local and global, spatial and channel. We obtain a new feature tensor and, by spatial pooling, we learn a powerful embedding for image retrieval. Focusing on global descriptors, we provide empirical evidence of the interaction of all forms of attention and improve the state of the art on standard benchmarks.
Chull Hwan Song, Hye Joo Han, Yannis Avrithis
WACV1
2005 A Hybrid Approach to Combine HMM and SVM Methods for the Prediction of the Transmembrane Spanning Region
Min Kyung Kim 0002, Chull Hwan Song, Seong Joon Yoo, Hyun Seok Park
KES (3)2
2005 An Ontology for Integrating Multimedia Databases
Chull Hwan Song, Young Hyun Koo, Seong Joon Yoo, ByeongHo Choi
KES (3)1