Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Taehoon Song

dblp:397/3932 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 44% Image recognition and object detection · 28% Transfer learning and domain adaptation · 15%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-language model
1.722025
Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection · NeurIPS 2025
Super-Class Guided Transformer for Zero-Shot Attribute Classification · AAAI 2025
Machine learning › Transfer learning and domain adaptation › foundation model adaptation
model-agnostic adaptation
1.012026
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization · AAAI 2026
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
1.012026
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization · AAAI 2026
Computer vision › Vision and language › vision-language model
vision-language model adaptation
1.012026
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization · AAAI 2026
Computer vision › Image recognition and object detection
attribute recognition
0.912025
Super-Class Guided Transformer for Zero-Shot Attribute Classification · AAAI 2025
Computer vision › Vision and language › vision-language model
prompt learning
0.912025
Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection · NeurIPS 2025
Computer vision › Image recognition and object detection › human-object interaction detection
zero-shot human-object interaction detection
0.912025
Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection · NeurIPS 2025
Computer vision › Image recognition and object detection
visual recognition
0.312026
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization · AAAI 2026
Machine learning › Transfer learning and domain adaptation › cross-domain transfer
cross-dataset transfer
0.312025
Super-Class Guided Transformer for Zero-Shot Attribute Classification · AAAI 2025

Methods — techniques the papers use, named apart from their topics

weak-to-strong generalization · 1.0unsupervised knowledge transfer · 1.0adapter · 1.0transformer · 0.9super-class query initialization · 0.9prompt learning · 0.9gaussian perturbation · 0.9consistency regularization · 0.9CLIP · 0.9
YearPublicationVenuePosition
2026 Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
abstract
Vision-Language Models (VLMs) have been widely used in various visual recognition tasks due to their remarkable generalization capabilities. As these models grow in size and complexity, fine-tuning becomes costly, emphasizing the need to reuse adaptation knowledge from 'weaker' models to efficiently enhance 'stronger' ones. However, existing adaptation transfer methods exhibit limited transferability across models due to their model-specific design and high computational demands. To tackle this, we propose Transferable Model-agnostic adapter (TransMiter), a light-weight adapter that improves vision-language models 'without backpropagation'. TransMiter captures the knowledge gap between pre-trained and fine-tuned VLMs, in an 'unsupervised' manner. Once trained, this knowledge can be seamlessly transferred across different models without the need for backpropagation. Moreover, TransMiter consists of only a few layers, inducing a negligible additional inference cost. Notably, supplementing the process with a few labeled data further yields additional performance gain, often surpassing a fine-tuned stronger model, with a marginal training cost. Experimental results and analyses demonstrate that TransMiter effectively and efficiently transfers adaptation knowledge while preserving generalization abilities across VLMs of different sizes and architectures in visual recognition tasks.
Taehoon Song, Sanghyeok Lee, Miso Choi, Hyunwoo J. Kim
AAAI2
2025 Super-Class Guided Transformer for Zero-Shot Attribute Classification
abstract
Attribute classification is crucial for identifying specific characteristics within image regions. Vision-Language Models (VLMs) have been effective in zero-shot tasks by leveraging their general knowledge from large-scale datasets. Recent studies demonstrate that transformer-based models with class-wise queries can effectively address zero-shot multi-label classification. However, poor utilization of the relationship between seen and unseen attributes makes the model lack generalizability. Additionally, attribute classification generally involves many attributes, making maintaining the model’s scalability difficult. To address these issues, we propose Super-class guided transFormer (SugaFormer), a novel framework that leverages super-classes to enhance scalability and generalizability for zero-shot attribute classification. SugaFormer employs Super-class Query Initialization (SQI) to reduce the number of queries, utilizing common semantic information from super-classes, and incorporates Multi-context Decoding (MD) to handle diverse visual cues. To strengthen generalizability, we introduce two knowledge transfer strategies that utilize VLMs. During training, Super-class guided Consistency Regularization (SCR) aligns model’s features with VLMs using super-class guided prompts, and during inference, Zero-shot Retrieval-based Score Enhancement (ZRSE) refines predictions for unseen attributes. Extensive experiments demonstrate that SugaFormer achieves state-of-the-art performance across three widely-used attribute classification benchmarks under zero-shot, and cross-dataset transfer settings.
Sehyung Kim, Chanhyeong Yang, Taehoon Song, Hyunwoo J. Kim
AAAI4
2025 Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection
abstract
Zero-shot Human-Object Interaction detection aims to localize humans and objects in an image and recognize their interaction, even when specific verb-object pairs are unseen during training. Recent works have shown promising results using prompt learning with pretrained vision-language models such as CLIP, which align natural language prompts with visual features in a shared embedding space. However, existing approaches still fail to handle the *visual complexity of interaction*—including (1) *intra-class visual diversity*, where instances of the same verb appear in diverse poses and contexts, and (2) *inter-class visual entanglement*, where distinct verbs yield visually similar patterns. To address these challenges, we propose **VDRP**, a framework for *Visual Diversity and Region-aware Prompt learning*. First, we introduce a visual diversity-aware prompt learning strategy that injects group-wise visual variance into the context embedding. We further apply Gaussian perturbation to encourage the prompt to capture diverse visual variations of a verb. Second, we retrieve region-specific concepts from the human, object, and union regions. These are used to augment the diversity-aware prompt embeddings, yielding region-aware prompts that improve verb-level discrimination. Experiments on the HICO-DET benchmark demonstrate that our method achieves state-of-the-art performance under four zero-shot evaluation settings, effectively addressing both intra-class diversity and inter-class visual entanglement. Code is available at https://github.com/mlvlab/VDRP.
Chanhyeong Yang, Taehoon Song, Hyunwoo J. Kim
NeurIPS2