EDBT 2026 Demo / reviewers in the wild / expert
Young Kyun Jang
dblp:241/5546
· DBLP profile ↗
15ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 11 since 2021Artificial intelligence and machine learning · 12 · 7 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Delving into Pre-training for Domain Transfer: A Broad Study of Pre-training for Domain Generalization and Domain Adaptation
Jungmyung Wi, Young Kyun Jang, Dujin Lee, Myeongseok Nam, Donghyun Kim 0006 |
Int. J. Comput. Vis. | 2 |
| 2026 | FAST-GOAL: Fast and Efficient Global-Local Object Alignment LearningabstractVision-language models such as CLIP have shown impressive capabilities in aligning images and text, but they often struggle with lengthy and detailed text descriptions due to pre-training on short and concise captions. We present FAST-GOAL (Fast and Efficient Global-local Object Alignment Learning), an efficient fine-tuning method that enhances ability of CLIP to handle lengthy text through global-local semantic alignment. Our method consists of two key components. First, Fast Local Image-Sentence Matching (FLISM) efficiently extracts local image regions through object detection and spatial division, then matches them with corresponding sentences. Second, Token Similarity-based Learning (TSL) maximizes the similarity between patch tokens from specific regions in the image and their corresponding region embeddings, applying the same principle to text, which enhances the ability of the model to capture detailed correspondences. Additionally, we introduce GLIT100k, a dataset that provides both global image-lengthy caption pairs and context-derived local pairs, where local descriptions are extracted from global captions to maintain semantic coherence. Through extensive experiments on long caption datasets (DOCCI, DCI) and short caption datasets (MSCOCO, Flickr30k), we demonstrate that FAST-GOAL achieves significant improvements over baselines, enabling effective adaptation of CLIP to detailed textual descriptions while maintaining computational efficiency. Hyungyu Choi, Young Kyun Jang, Chanho Eom |
IEEE Trans. Image Process. | 2 |
| 2025 | GOAL: Global-local Object Alignment LearningabstractVision-language models like CLIP have shown impressive capabilities in aligning images and text, but they often struggle with lengthy and detailed text descriptions because of their training focus on short and concise captions. We present GOAL (Global-local Object Alignment Learning), a novel fine-tuning method that enhances CLIP’s ability to handle lengthy text by leveraging both global and local semantic alignments between image and lengthy text. Our approach consists of two key components: Local Image-Sentence Matching (LISM), which identifies corresponding pairs between image segments and descriptive sentences, and Token Similarity-based Learning (TSL), which efficiently propagates local element attention through these matched pairs. Evaluating GOAL on three new benchmarks for image-lengthy text retrieval, we demonstrate significant improvements over baseline CLIP fine-tuning, establishing a simple yet effective approach for adapting CLIP to detailed textual descriptions. Through extensive experiments, we show that our method’s focus on local semantic alignment alongside global context leads to more nuanced and representative embeddings, particularly beneficial for tasks requiring fine-grained understanding of lengthy text descriptions. Hyungyu Choi, Young Kyun Jang, Chanho Eom |
CVPR | 2 |
| 2025 | MA-CIR: A Multimodal Arithmetic Benchmark for Composed Image Retrieval
Jaeseok Byun, Young Kyun Jang, Seokhyeon Jeong, Taesup Moon |
ICCV | 2 |
| 2025 | Towards Cross-Modal Backward-Compatible Representation Learning for Vision-Language ModelsabstractModern retrieval systems often struggle with upgrading to new and more powerful models due to the incompatibility of embeddings between the old and new models. This necessitates a costly process known as backfilling, which involves re-computing the embeddings for a large number of data samples. In vision, Backward-compatible Training (BT) has been proposed to ensure that the new model aligns with the old model's embeddings. This paper extends the concept of vision-only BT to the field of cross-modal retrieval, marking the first attempt to address Cross-modal BT (XBT). Our goal is to achieve backward-compatibility between Vision-Language Pretraining (VLP) models, such as CLIP, for the cross-modal retrieval task. To address XBT challenges, we propose an efficient solution: a projection module that maps the new model's embeddings to those of the old model. This module, pretrained solely with text data, significantly reduces the number of image-text pairs required for XBT learning, and, once it is pretrained, it avoids using the old model during training. Furthermore, we utilize parameter-efficient training strategies that improve efficiency and preserve the off-the-shelf new model's knowledge by avoiding any modifications. Experimental results on cross-modal retrieval datasets demonstrate the effectiveness of XBT and its potential to enable backfill-free upgrades when a new VLP model emerges. Young Kyun Jang, Ser-Nam Lim |
ICCV | 1 |
| 2024 | MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video UnderstandingabstractWith the success of large language models (LLMs), integrating the vision model into LLMs to build vision-language foundation models has gained much more interest recently. However, existing LLM-based large multimodal models (e.g., Video-LLaMA, VideoChat) can only take in a limited number of frames for short video understanding. In this study, we mainly focus on designing an efficient and effective model for long-term video understanding. Instead of trying to process more frames simultaneously like most existing work, we propose to process videos in an online manner and store past video information in a memory bank. This allows our model to reference historical video content for long-term analysis without exceeding LLMs' context length constraints or GPU memory limits. Our memory bank can be seamlessly integrated into current multimodal LLMs in an off-the-shelf manner. We conduct extensive experiments on various video understanding tasks, such as long-video understanding, video question answering, and video captioning, and our model can achieve state-of-the-art performances across multiple datasets. Bo He 0004, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, Ser-Nam Lim |
CVPR | 3 |
| 2024 | On the Robustness of Large Multimodal Models Against Image Adversarial AttacksabstractRecent advances in instruction tuning have led to the development of State-of-the-Art Large Multimodal Models (LMMs). Given the novelty of these models, the impact of visual adversarial attacks on LMMs has not been thoroughly examined. We conduct a comprehensive study of the robustness of various LMMs against different adversarial attacks, evaluated across tasks including image classification, image captioning, and Visual Question Answer (VQA). We find that in general LMMs are not robust to visual adversarial inputs. However, our findings suggest that context provided to the model via prompts—such as questions in a QA pair—helps to mitigate the effects of visual adversarial inputs. Notably, the LMMs evaluated demonstrated remarkable resilience to such attacks on the ScienceQA task with only an 8.10% drop in performance compared to their visual counterparts which dropped 99.73%. We also propose a new approach to real-world image classification which we term query decomposition. By incorporating existence queries into our input prompt we observe diminished attack effectiveness and improvements in image classification accuracy. This research highlights a previously under explored facet of LMM robustness and sets the stage for future work aimed at strengthening the resilience of multimodal systems in adversarial environments. Xuanming Cui, Alejandro Aparcedo, Young Kyun Jang, Ser-Nam Lim |
CVPR | 3 |
| 2024 | Visual Delta Generator with Large Multi-Modal Models for Semi-Supervised Composed Image RetrievalabstractComposed Image Retrieval (CIR) is a task that retrieves images similar to a query, based on a provided textual modification. Current techniques rely on supervised learning for CIR models using labeled triplets of the. These specific triplets are not as commonly available as simple image-text pairs, limiting the widespread use of CIR and its scalability. On the other hand, zero-shot CIR can be relatively easily trained with image-caption pairs without considering the image-to-image relation, but this approach tends to yield lower accuracy. We propose a new semi-supervised CIR approach where we search for a reference and its related target images in auxiliary data and learn our large language model-based Visual Delta Generator (VDG) to generate text describing the visual difference (i.e., visual delta) between the two. VDG, equipped with fluent language knowledge and being model agnostic, can generate pseudo triplets to boost the performance of CIR models. Our approach significantly improves the existing supervised learning approaches and achieves state-of-the-art results on the CIR benchmarks. Young Kyun Jang, Zihang Meng, Dat Huynh, Ser-Nam Lim |
CVPR | 1 |
| 2024 | Spherical Linear Interpolation and Text-Anchoring for Zero-Shot Composed Image Retrieval
Young Kyun Jang, Dat Huynh, Ashish Shah, Wen-Kai Chen, Ser-Nam Lim |
ECCV (19) | 1 |
| 2022 | Deep Hash Distillation for Image Retrieval
Young Kyun Jang, Geonmo Gu, Byungsoo Ko, Isaac Kang, Nam Ik Cho |
ECCV (14) | 1 |
| 2022 | Self-Supervised Pretraining for Deep Hash-Based Image RetrievalabstractDeep hashing aims to produce discriminative binary hash codes for fast image retrieval through a deep baseline network and additional trainable hash function. In a supervised deep hashing network, the baseline network is generally initialized with classification-based pretrained models, and the overall hashing network is trained in a supervised fashion. However, since classification and retrieval are two different tasks, it is necessary to reconsider the initial model for the baseline network. In this paper, we propose to use a self-supervised pretrained model as the baseline for the first time. We investigate the impact of pretrained model types by comparing deep hashing networks that use the baseline network with 1) randomly initialized weights, 2) conventional supervised pretrained weights, and 3) proposed self-supervised pretrained weights. As a result, we confirm that the performance of deep hashing differs depending on the initial baseline setting, and the proposed self-supervised baseline model shows comparable or better performance over the supervised one. Our code is released at https://github.com/HaeyoonYang/SSPH. Haeyoon Yang, Young Kyun Jang, Isaac Kang, Nam Ik Cho |
ICIP | 2 |
| 2021 | Self-supervised Product Quantization for Deep Unsupervised Image RetrievalabstractSupervised deep learning-based hash and vector quantization are enabling fast and large-scale image retrieval systems. By fully exploiting label annotations, they are achieving outstanding retrieval performances compared to the conventional methods. However, it is painstaking to assign labels precisely for a vast amount of training data, and also, the annotation process is error-prone. To tackle these issues, we propose the first deep unsupervised image retrieval method dubbed Self-supervised Product Quantization (SPQ) network, which is label-free and trained in a self-supervised manner. We design a Cross Quantized Contrastive learning strategy that jointly learns codewords and deep visual descriptors by comparing individually transformed images (views). Our method analyzes the image contents to extract descriptive features, allowing us to understand image representations for accurate retrieval. By conducting extensive experiments on benchmarks, we demonstrate that the proposed method yields state-of-the-art results even without supervised pretraining. Young Kyun Jang, Nam Ik Cho |
ICCV | 1 |
| 2020 | Generalized Product Quantization Network for Semi-Supervised Image RetrievalabstractImage retrieval methods that employ hashing or vector quantization have achieved great success by taking advantage of deep learning. However, these approaches do not meet expectations unless expensive label information is sufficient. To resolve this issue, we propose the first quantization-based semi-supervised image retrieval scheme: Generalized Product Quantization (GPQ) network. We design a novel metric learning strategy that preserves semantic similarity between labeled data, and employ entropy regularization term to fully exploit inherent potentials of unlabeled data. Our solution increases the generalization capacity of the quantization network, which allows overcoming previous limitations in the retrieval community. Extensive experimental results demonstrate that GPQ yields state-of-the-art performance on large-scale real image benchmark datasets. Young Kyun Jang, Nam Ik Cho |
CVPR | 1 |
| 2019 | Deep Face Image Retrieval for Cancelable Biometric AuthenticationabstractThis paper presents a cancelable biometric system for face authentication by exploiting the convolutional neural network (CNN)-based face image retrieval system. For the cancelable biometrics we must build a template that achieves good performance while maintaining some essential conditions. First the same template should not be used in different applications. Second if the compromise event occurs original biometric data should not be retrieved from the template. Last the template should be easily discarded and recreated. Hence we propose a Deep Table-based Hashing (DTH) framework that encodes CNN-based features into a binary code by utilizing the index of the hashing table. We employ noise embedding and intra-normalization that distorts biometric data which enhances the non-invertibility. For training we propose a new segment-clustering loss and pairwise Hamming loss with two classification losses. The final authentication results are obtained by voting on the outcome of the retrieval system. Experiments conducted on two large scale face image datasets demonstrate that the proposed method works as a proper cancelable biometric system. Young Kyun Jang, Nam Ik Cho |
AVSS | 1 |
| 2018 | Deep Clustering and Block Hashing Network for Face Image Retrieval
Young Kyun Jang, Dong-ju Jeong, Seok Hee Lee, Nam Ik Cho |
ACCV (6) | 1 |