Si-Woo Kim

dblp:96/4573 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Vision and language · 77% Generative modeling · 13% Transfer learning and domain adaptation · 8%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
image captioning
2.532025
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning · ACM Multimedia 2025
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning · AAAI 2025
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning · EMNLP 2024
Computer vision › Vision and language
video captioning
1.622025
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning · EMNLP 2025
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning · EMNLP 2024
Computer vision › Vision and language › image captioning › low-shot image captioning
zero-shot image captioning
1.622025
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning · ACM Multimedia 2025
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning · EMNLP 2024
Machine learning › Generative modeling › cross-modal generation
audio-to-image generation
0.912025
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation · ACM Multimedia 2025
Computer vision › Vision and language › cross-modal alignment
cross-modal semantic alignment
0.912025
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation · ACM Multimedia 2025
Computer vision › Vision and language › video captioning
dense video captioning
0.912025
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning · EMNLP 2025
Machine learning › Transfer learning and domain adaptation › domain adaptation › low-resource domain adaptation
zero-shot domain adaptation
0.912025
SIDA: Synthetic Image Driven Zero-shot Domain Adaptation · ACM Multimedia 2025
Information retrieval
retrieval models
0.912025
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning · AAAI 2025
Machine learning › Generative modeling
cross-modal generation
0.312025
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation · ACM Multimedia 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.312025
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning · ACM Multimedia 2025
Image and video processing
saliency detection
0.312025
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning · EMNLP 2025
Information retrieval
cross-modal retrieval
0.212024
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

visual prompt · 1.7saliency-aware reweighting · 1.7adaptive segmentation · 1.7CLIP · 1.7synthetic image generation · 0.9mapping network · 0.9large language model · 0.9image-to-text retrieval · 0.9cycle-consistency scoring · 0.9audio captioning model · 0.9text-only training · 0.8fusion module · 0.8frequency-based entity filtering · 0.8
YearPublicationVenuePosition
2025 ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
abstract
Recent lightweight image captioning models using retrieved data mainly focus on text prompts. However, previous works only utilize the retrieved text as text prompts, and the visual information relies only on the CLIP visual embedding. Because of this issue, there is a limitation that the image descriptions inherent in the prompt are not sufficiently reflected in the visual embedding space. To tackle this issue, we propose ViPCap, a novel retrieval text-based visual prompt for lightweight image captioning. ViPCap leverages the retrieved text with image information as visual prompts to enhance the ability of the model to capture relevant visual information. By mapping text prompts into the CLIP space and generating multiple randomized Gaussian distributions, our method leverages sampling to explore randomly augmented distributions and effectively retrieves the semantic features that contain image information. These retrieved features are integrated into the image and designated as the visual prompt, leading to performance improvements on the datasets such as COCO, Flickr30k, and NoCaps. Experimental results demonstrate that ViPCap significantly outperforms prior lightweight captioning models in efficiency and effectiveness, demonstrating the potential for a plug-and-play solution.
Taewhan Kim 0002, Soeun Lee, Si-Woo Kim, Dong-Jin Kim 0003
AAAI3
2025 Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
abstract
Dense video captioning aims to temporally localize events in video and generate captions for each event.While recent works propose end-to-end models, they suffer from two limitations: (1) applying timestamp supervision only to text while treating all video frames equally, and (2) retrieving captions from fixedsize video chunks, overlooking scene transitions.To address these, we propose Sali4Vid, a simple yet effective saliency-aware framework.We introduce Saliency-aware Video Reweighting, which converts timestamp annotations into sigmoid-based frame importance weights, and Semantic-based Adaptive Caption Retrieval, which segments videos by frame similarity to capture scene transitions and improve caption retrieval.Sali4Vid achieves state-of-the-art results on YouCook2 and ViTT, demonstrating the benefit of jointly improving video weighting and retrieval for dense video captioning. 1
MinJu Jeon, Si-Woo Kim, Ye-Chan Kim, HyunGee Kim
EMNLP2
2025 SIDA: Synthetic Image Driven Zero-shot Domain Adaptation
Ye-Chan Kim, SeungJu Cha 0001, Si-Woo Kim, Taewhan Kim 0002, Dong-Jin Kim 0003
ACM Multimedia3
2025 SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
abstract
Zero-shot Image Captioning (ZIC) increasingly utilizes synthetic datasets generated by text-to-image (T2I) models to mitigate the need for costly manual annotation. However, these T2I models often produce images that exhibit semantic misalignments with their corresponding input captions (e.g., missing objects, incorrect attributes), resulting in noisy synthetic image-caption pairs that can hinder model training. Existing dataset pruning techniques are largely designed for removing noisy text in web-crawled data. However, these methods are ill-suited for the distinct challenges of synthetic data, where captions are typically well-formed, but images may be inaccurate representations. To address this gap, we introduce SynC, a novel framework specifically designed to refine synthetic image-caption datasets for ZIC. Instead of conventional filtering or regeneration, SynC focuses on reassigning captions to the most semantically aligned images already present within the synthetic image pool. Our approach employs a one-to-many mapping strategy by initially retrieving multiple relevant candidate images for each caption. We then apply a cycle-consistency-inspired alignment scorer that selects the best image by verifying its ability to retrieve the original caption via image-to-text retrieval. Extensive evaluations demonstrate that SynC consistently and significantly improves performance across various ZIC models on standard benchmarks (MS-COCO, Flickr30k, NoCaps), achieving state-of-the-art results in several scenarios. SynC offers an effective strategy for curating refined synthetic data to enhance ZIC.
Si-Woo Kim, MinJu Jeon, Ye-Chan Kim, Soeun Lee, Taewhan Kim 0002, Dong-Jin Kim 0003
ACM Multimedia1
2025 CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
abstract
We propose CatchPhrase, a novel audio-to-image generation framework designed to mitigate semantic misalignment between audio inputs and generated images. While recent advances in multi-modal encoders have enabled progress in cross-modal generation, ambiguity stemming from homographs and auditory illusions continues to hinder accurate alignment. To address this issue, CatchPhrase generates enriched cross-modal semantic prompts (EXPrompt Mining ) from weak class labels by leveraging large language models (LLMs) and audio captioning models (ACMs). To address both class-level and instance-level misalignment, we apply multi-modal filtering and retrieval to select the most semantically aligned prompt for each audio sample (EXPrompt Selector ). A lightweight mapping network is then trained to adapt pre-trained text-to-image generation models to audio input. Extensive experiments on multiple audio classification datasets demonstrate that CatchPhrase improves audio-to-image alignment and consistently enhances generation quality by mitigating semantic misalignment.
Hyunwoo Oh, SeungJu Cha 0001, Kwanyoung Lee, Si-Woo Kim, Dong-Jin Kim 0003
ACM Multimedia4
2024 IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
abstract
Recent advancements in image captioning have explored text-only training methods to overcome the limitations of paired image-text data.However, existing text-only training methods often overlook the modality gap between using text data during training and employing images during inference.To address this issue, we propose a novel approach called Image-like Retrieval, which aligns text features with visually relevant features to mitigate the modality gap.Our method further enhances the accuracy of generated captions by designing a Fusion Module that integrates retrieved captions with input features.Additionally, we introduce a Frequency-based Entity Filtering technique that significantly improves caption quality.We integrate these methods into a unified framework, which we refer to as IFCap (Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning).Through extensive experimentation, our straightforward yet powerful approach has demonstrated its efficacy, outperforming the state-of-the-art methods by a significant margin in both image captioning and video captioning compared to zero-shot captioning based on text-only training.1
Soeun Lee, Si-Woo Kim
EMNLP2