Dong-Jin Kim 0003

dblp:16/9611-3 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0001-7231-7494ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 13 since 2021
YearPublicationVenuePosition
2026 ADuLTS: Appearance Descriptions under Long-Tailed Scenarios with diverse synthesized images
SeungJu Cha 0001, Seunghee Choi, Kwanyoung Lee, Dong-Jin Kim 0003
Comput. Vis. Image Underst.4
2025 ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
abstract
Recent lightweight image captioning models using retrieved data mainly focus on text prompts. However, previous works only utilize the retrieved text as text prompts, and the visual information relies only on the CLIP visual embedding. Because of this issue, there is a limitation that the image descriptions inherent in the prompt are not sufficiently reflected in the visual embedding space. To tackle this issue, we propose ViPCap, a novel retrieval text-based visual prompt for lightweight image captioning. ViPCap leverages the retrieved text with image information as visual prompts to enhance the ability of the model to capture relevant visual information. By mapping text prompts into the CLIP space and generating multiple randomized Gaussian distributions, our method leverages sampling to explore randomly augmented distributions and effectively retrieves the semantic features that contain image information. These retrieved features are integrated into the image and designated as the visual prompt, leading to performance improvements on the datasets such as COCO, Flickr30k, and NoCaps. Experimental results demonstrate that ViPCap significantly outperforms prior lightweight captioning models in efficiency and effectiveness, demonstrating the potential for a plug-and-play solution.
Taewhan Kim 0002, Soeun Lee, Si-Woo Kim, Dong-Jin Kim 0003
AAAI4
2025 VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
abstract
Recent large-scale text-to-image diffusion models generate photorealistic images but often struggle to accurately depict interactions between humans and objects due to their limited ability to differentiate various interaction words. In this work, we propose VerbDiff to address the challenge of capturing nuanced interactions within text-to-image diffusion models. VerbDiff is a novel text-to-image generation model that weakens the bias between interaction words and objects, enhancing the understanding of interactions. Specifically, we disentangle various interaction words from frequency-based anchor words and leverage localized interaction regions from generated images to help the model better capture semantics in distinctive words without extra conditions. Our approach enables the model to accurately understand the intended interaction between humans and objects, producing high-quality images with accurate interactions aligned with specified verbs. Extensive experiments on the HICO-DET dataset demonstrate the effectiveness of our method compared to previous approaches.
SeungJu Cha 0001, Kwanyoung Lee, Ye-Chan Kim, Hyunwoo Oh, Dong-Jin Kim 0003
CVPR5
2025 SIDA: Synthetic Image Driven Zero-shot Domain Adaptation
Ye-Chan Kim, SeungJu Cha 0001, Si-Woo Kim, Taewhan Kim 0002, Dong-Jin Kim 0003
ACM Multimedia5
2025 SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
abstract
Zero-shot Image Captioning (ZIC) increasingly utilizes synthetic datasets generated by text-to-image (T2I) models to mitigate the need for costly manual annotation. However, these T2I models often produce images that exhibit semantic misalignments with their corresponding input captions (e.g., missing objects, incorrect attributes), resulting in noisy synthetic image-caption pairs that can hinder model training. Existing dataset pruning techniques are largely designed for removing noisy text in web-crawled data. However, these methods are ill-suited for the distinct challenges of synthetic data, where captions are typically well-formed, but images may be inaccurate representations. To address this gap, we introduce SynC, a novel framework specifically designed to refine synthetic image-caption datasets for ZIC. Instead of conventional filtering or regeneration, SynC focuses on reassigning captions to the most semantically aligned images already present within the synthetic image pool. Our approach employs a one-to-many mapping strategy by initially retrieving multiple relevant candidate images for each caption. We then apply a cycle-consistency-inspired alignment scorer that selects the best image by verifying its ability to retrieve the original caption via image-to-text retrieval. Extensive evaluations demonstrate that SynC consistently and significantly improves performance across various ZIC models on standard benchmarks (MS-COCO, Flickr30k, NoCaps), achieving state-of-the-art results in several scenarios. SynC offers an effective strategy for curating refined synthetic data to enhance ZIC.
Si-Woo Kim, MinJu Jeon, Ye-Chan Kim, Soeun Lee, Taewhan Kim 0002, Dong-Jin Kim 0003
ACM Multimedia6
2025 CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
abstract
We propose CatchPhrase, a novel audio-to-image generation framework designed to mitigate semantic misalignment between audio inputs and generated images. While recent advances in multi-modal encoders have enabled progress in cross-modal generation, ambiguity stemming from homographs and auditory illusions continues to hinder accurate alignment. To address this issue, CatchPhrase generates enriched cross-modal semantic prompts (EXPrompt Mining ) from weak class labels by leveraging large language models (LLMs) and audio captioning models (ACMs). To address both class-level and instance-level misalignment, we apply multi-modal filtering and retrieval to select the most semantically aligned prompt for each audio sample (EXPrompt Selector ). A lightweight mapping network is then trained to adapt pre-trained text-to-image generation models to audio input. Extensive experiments on multiple audio classification datasets demonstrate that CatchPhrase improves audio-to-image alignment and consistently enhances generation quality by mitigating semantic misalignment.
Hyunwoo Oh, SeungJu Cha 0001, Kwanyoung Lee, Si-Woo Kim, Dong-Jin Kim 0003
ACM Multimedia5
2025 ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic Diffusion
abstract
Text-to-image diffusion models often exhibit degraded performance when generating images beyond their training resolution. Recent training-free methods can mitigate this limitation, but they often require substantial computation or are incompatible with recent Diffusion Transformer models. In this paper, we propose ScaleDiff, a model-agnostic and highly efficient framework for extending the resolution of pretrained diffusion models without any additional training. A core component of our framework is Neighborhood Patch Attention (NPA), an efficient mechanism that reduces computational redundancy in the self-attention layer with non-overlapping patches. We integrate NPA into an SDEdit pipeline and introduce Latent Frequency Mixing (LFM) to better generate fine details. Furthermore, we apply Structure Guidance to enhance global structure during the denoising process. Experimental results demonstrate that ScaleDiff achieves state-of-the-art performance among training-free methods in terms of both image quality and inference speed on both U-Net and Diffusion Transformer architectures.
Sungho Koh, SeungJu Cha 0001, Hyunwoo Oh, Kwanyoung Lee, Dong-Jin Kim 0003
NeurIPS5
2024 Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
abstract
In this paper, we propose a new method to enhance compositional understanding in pretrained vision and language models (VLMs) without sacrificing performance in zero-shot multi-modal tasks.Traditional fine-tuning approaches often improve compositional reasoning at the cost of degrading multi-modal capabilities, primarily due to the use of global hard negative (HN) loss, which contrasts global representations of images and texts.This global HN loss pushes HN texts that are highly similar to the original ones, damaging the model's multi-modal representations.To overcome this limitation, we propose Fine-grained Selective Calibrated CLIP (FSC-CLIP), which integrates local hard negative loss and selective calibrated regularization.These innovations provide fine-grained negative supervision while preserving the model's representational integrity.Our extensive evaluations across diverse benchmarks for both compositionality and multi-modal tasks show that FSC-CLIP not only achieves compositionality on par with state-of-the-art models but also retains strong multi-modal capabilities.
Youngtaek Oh, Jae-Won Cho, Dong-Jin Kim 0003, In-So Kweon, Junmo Kim 0002
EMNLP3
2023 Generative Bias for Robust Visual Question Answering
abstract
The task of Visual Question Answering (VQA) is known to be plagued by the issue of VQA models exploiting biases within the dataset to make its final prediction. Various previous ensemble based debiasing methods have been proposed where an additional model is purposefully trained to be biased in order to train a robust target model. However, these methods compute the bias for a model simply from the label statistics of the training data or from single modal branches. In this work, in order to better learn the bias a target VQA model suffers from, we propose a generative method to train the bias model directly from the target model, called GenB. In particular, GenB employs a generative network to learn the bias in the target model through a combination of the adversarial objective and knowledge distillation. We then debias our target model with GenB as a bias model, and show through extensive experiments the effects of our method on various VQA bias datasets including VQA-CP2, VQA-CP1, GQA-OOD, and VQA-CE, and show state-of-the-art results with the LXMERT architecture on VQA-CP2.
Jae-Won Cho, Dong-Jin Kim 0003, Hyeonggon Ryu, In-So Kweon
CVPR2
2023 Self-Sufficient Framework for Continuous Sign Language Recognition
abstract
The goal of this work is to develop self-sufficient framework for Continuous Sign Language Recognition (CSLR) that addresses key issues of sign language recognition. These include the need for complex multi-scale features such as hands, face, and mouth for understanding, and absence of frame-level annotations. To this end, we propose (1) Divide and Focus Convolution (DFConv) which extracts both manual and non-manual features without the need for additional networks or annotations, and (2) Dense Pseudo-Label Refinement (DPLR) which propagates non-spiky frame-level pseudo-labels by combining the ground truth gloss sequence labels with the predicted sequence. We demonstrate that our model achieves state-of-the-art performance among RGB-based methods on large-scale CSLR benchmarks, PHOENIX-2014 and PHOENIX-2014-T, while showing comparable results with better efficiency when compared to other approaches that use multi-modality or extra annotations.
Youngjoon Jang 0001, Youngtaek Oh, Jae-Won Cho, Myungchul Kim 0002, Dong-Jin Kim 0003, In-So Kweon, Joon Son Chung
ICASSP5
2023 Empirical study on using adapters for debiased Visual Question Answering
Jae-Won Cho, Dawit Mureja Argaw, Youngtaek Oh, Dong-Jin Kim 0003, In-So Kweon
Comput. Vis. Image Underst.4
2023 Local pseudo-attributes for long-tailed recognition
Dong-Jin Kim 0003, Tsung-Wei Ke, Stella X. Yu
Pattern Recognit. Lett.1
2023 Modeling Semantic Correlation and Hierarchy for Real-World Wildlife Recognition
abstract
We explore the challenges of human-in-the-loop frameworks to label wildlife recognition datasets with a neural network. In wildlife imagery, the main challenges for a model to assist human annotation are two-fold: (1) the training dataset is usually imbalanced, which makes the model's suggestion biased, and (2) there are complex taxonomies in the classes. We establish a simple and efficient baseline, including the debiasing loss function and the hyperbolic network architecture, to address these issues. Moreover, we propose leveraging the semantic correlation to train the model more effectively by adding a co-occurrence layer to our model during training. We demonstrate the efficacy of our method in both a real-world wildlife areal survey recognition dataset and the public image classification dataset, CIFAR100-LT, CIFAR10-LT, and iNaturalist.
Dong-Jin Kim 0003, Zhongqi Miao, Yunhui Guo, Stella X. Yu
IEEE Signal Process. Lett.1
2023 MCDAL: Maximum Classifier Discrepancy for Active Learning
abstract
Recent state-of-the-art active learning methods have mostly leveraged generative adversarial networks (GANs) for sample acquisition; however, GAN is usually known to suffer from instability and sensitivity to hyperparameters. In contrast to these methods, in this article, we propose a novel active learning framework that we call Maximum Classifier Discrepancy for Active Learning (MCDAL) that takes the prediction discrepancies between multiple classifiers. In particular, we utilize two auxiliary classification layers that learn tighter decision boundaries by maximizing the discrepancies among them. Intuitively, the discrepancies in the auxiliary classification layers' predictions indicate the uncertainty in the prediction. In this regard, we propose a novel method to leverage the classifier discrepancies for the acquisition function for active learning. We also provide an interpretation of our idea in relation to existing GAN-based active learning methods and domain adaptation frameworks. Moreover, we empirically demonstrate the utility of our approach where the performance of our approach exceeds the state-of-the-art methods on several image classification and semantic segmentation datasets in active learning setups.
Jae-Won Cho, Dong-Jin Kim 0003, Yunjae Jung, In-So Kweon
IEEE Trans. Neural Networks Learn. Syst.2
2022 Signing Outside the Studio: Benchmarking Background Robustness for Continuous Sign Language Recognition
Youngjoon Jang 0001, Youngtaek Oh, Jae-Won Cho, Dong-Jin Kim 0003, Joon Son Chung, In-So Kweon
BMVC4
2022 DASO: Distribution-Aware Semantics-Oriented Pseudo-label for Imbalanced Semi-Supervised Learning
abstract
The capability of the traditional semi-supervised learning (SSL) methods is far from real-world application due to severely biased pseudo-labels caused by (1) class imbalance and (2) class distribution mismatch between labeled and unlabeled data. This paper addresses such a relatively under-explored problem. First, we propose a general pseudo-labeling framework that class-adaptively blends the semantic pseudo-label from a similarity-based classifier to the linear one from the linear classifier, after making the observation that both types of pseudo-labels have complementary properties in terms of bias. We further introduce a novel semantic alignment loss to establish balanced feature representation to reduce the biased predictions from the classifier. We term the whole framework as Distribution-Aware Semantics-Oriented (DASO) Pseudo-label. We conduct extensive experiments in a wide range of imbalanced benchmarks: CIFAR10/100-LT, STL10-LT, and large-scale long-tailed Semi-Aves with open-set class, and demonstrate that, the proposed DASO framework reliably improves SSL learners with unlabeled data especially when both (1) class imbalance and (2) distribution mismatch dominate.
Youngtaek Oh, Dong-Jin Kim 0003, In-So Kweon
CVPR2
2022 Dense Relational Image Captioning via Multi-Task Triple-Stream Networks
abstract
We introduce dense relational captioning, a novel image captioning task which aims to generate multiple captions with respect to relational information between objects in a visual scene. Relational captioning provides explicit descriptions for each relationship between object combinations. This framework is advantageous in both diversity and amount of information, leading to a comprehensive image understanding based on relationships, e.g., relational proposal generation. For relational understanding between objects, the part-of-speech (POS; i.e., subject-object-predicate categories) can be a valuable prior information to guide the causal sequence of words in a caption. We enforce our framework to learn not only to generate captions but also to understand the POS of each word. To this end, we propose the multi-task triple-stream network (MTTSNet) which consists of three recurrent units responsible for each POS which is trained by jointly predicting the correct captions and POS for each word. In addition, we found that the performance of MTTSNet can be improved by modulating the object embeddings with an explicit relational module. We demonstrate that our proposed model can generate more diverse and richer captions, via extensive experimental analysis on large scale datasets and several metrics. Then, we present applications of our framework to holistic image captioning, scene graph generation, and retrieval tasks.
Dong-Jin Kim 0003, Tae-Hyun Oh, Jinsoo Choi, In-So Kweon
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Single-Modal Entropy based Active Learning for Visual Question Answering
Dong-Jin Kim 0003, Jae-Won Cho, Jinsoo Choi, Yunjae Jung, In-So Kweon
BMVC1
2021 LabOR: Labeling Only if Required for Domain Adaptive Semantic Segmentation
abstract
Unsupervised Domain Adaptation (UDA) for semantic segmentation has been actively studied to mitigate the domain gap between label-rich source data and unlabeled target data. Despite these efforts, UDA still has a long way to go to reach the fully supervised performance. To this end, we propose a Labeling Only if Required strategy, LabOR, where we introduce a human-in-the-loop approach to adaptively give scarce labels to points that a UDA model is uncertain about. In order to find the uncertain points, we generate an inconsistency mask using the proposed adaptive pixel selector and we label these segment-based regions to achieve near supervised performance with only a small fraction (about 2.2%) ground truth points, which we call "Segment based Pixel-Labeling (SPL)." To further reduce the efforts of the human annotator, we also propose "Point based Pixel-Labeling (PPL)," which finds the most representative points for labeling within the generated inconsistency mask. This reduces efforts from 2.2% segment label → 40 points label while minimizing performance degradation. Through extensive experimentation, we show the advantages of this new framework for domain adaptive semantic segmentation while minimizing human labor costs.
Inkyu Shin, Dong-Jin Kim 0003, Jae-Won Cho, Sanghyun Woo, Kwanyong Park, In-So Kweon
ICCV2
2021 ACP++: Action Co-Occurrence Priors for Human-Object Interaction Detection
abstract
A common problem in the task of human-object interaction (HOI) detection is that numerous HOI classes have only a small number of labeled examples, resulting in training sets with a long-tailed distribution. The lack of positive labels can lead to low classification accuracy for these classes. Towards addressing this issue, we observe that there exist natural correlations and anti-correlations among human-object interactions. In this paper, we model the correlations as action co-occurrence matrices and present techniques to learn these priors and leverage them for more effective training, especially on rare classes. The efficacy of our approach is demonstrated experimentally, where the performance of our approach consistently improves over the state-of-the-art methods on both of the two leading HOI detection benchmark datasets, HICO-Det and V-COCO.
Dong-Jin Kim 0003, Xiao Sun 0001, Jinsoo Choi, Stephen Lin 0001, In-So Kweon
IEEE Trans. Image Process.1
2020 Detecting Human-Object Interactions with Action Co-occurrence Priors
Dong-Jin Kim 0003, Xiao Sun 0001, Jinsoo Choi, Stephen Lin 0001, In-So Kweon
ECCV (21)1
2019 Dense Relational Captioning: Triple-Stream Networks for Relationship-Based Captioning
abstract
Our goal in this work is to train an image captioning model that generates more dense and informative captions. We introduce "relational captioning," a novel image captioning task which aims to generate multiple captions with respect to relational information between objects in an image. Relational captioning is a framework that is advantageous in both diversity and amount of information, leading to image understanding based on relationships. Part-of-speech (POS, i.e. subject-object-predicate categories) tags can be assigned to every English word. We leverage the POS as a prior to guide the correct sequence of words in a caption. To this end, we propose a multi-task triple-stream network (MTTSNet) which consists of three recurrent units for the respective POS and jointly performs POS prediction and captioning. We demonstrate more diverse and richer representations generated by the proposed model against several baselines and competing methods.
Dong-Jin Kim 0003, Jinsoo Choi, Tae-Hyun Oh, In-So Kweon
CVPR1
2019 Image Captioning with Very Scarce Supervised Data: Adversarial Semi-Supervised Learning Approach
abstract
Dong-Jin Kim, Jinsoo Choi, Tae-Hyun Oh, In So Kweon. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Dong-Jin Kim 0003, Jinsoo Choi, Tae-Hyun Oh, In-So Kweon
EMNLP/IJCNLP (1)1
2018 Disjoint Multi-task Learning Between Heterogeneous Human-Centric Tasks
abstract
Human behavior understanding is arguably one of the most important mid-level components in artificial intelligence. In order to efficiently make use of data, multi-task learning has been studied in diverse computer vision tasks including human behavior understanding. However, multitask learning relies on task specific datasets and constructing such datasets can be cumbersome. It requires huge amounts of data, labeling efforts, statistical consideration etc. In this paper, we leverage existing single-task datasets for human action classification and captioning data for efficient human behavior learning. Since the data in each dataset has respective heterogeneous annotations, traditional multi-task learning is not effective in this scenario. To this end, we propose a novel alternating directional optimization method to efficiently learn from the heterogeneous data. We demonstrate the effectiveness of our model and show performance improvements on both classification and sentence retrieval tasks in comparison to the models trained on each of the single-task datasets.
Dong-Jin Kim 0003, Jinsoo Choi, Tae-Hyun Oh, Youngjin Yoon, In-So Kweon
WACV1