Sehyung Kim

dblp:372/6337 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0003-6381-9291ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Image recognition and object detection · 60% Vision and language · 35% Transfer learning and domain adaptation · 5%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-language model
1.122025
Super-Class Guided Transformer for Zero-Shot Attribute Classification · AAAI 2025
Retrieval-Augmented Open-Vocabulary Object Detection · CVPR 2024
Computer vision › Image recognition and object detection
attribute recognition
0.912025
Super-Class Guided Transformer for Zero-Shot Attribute Classification · AAAI 2025
Computer vision › Image recognition and object detection › object detection › detector training
label assignment
0.812024
Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-Based Visual Relationship Detection · CVPR 2024
Computer vision › Image recognition and object detection
object detection
0.812024
Retrieval-Augmented Open-Vocabulary Object Detection · CVPR 2024
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection
0.812024
Retrieval-Augmented Open-Vocabulary Object Detection · CVPR 2024
Computer vision › Vision and language
visual relationship detection
0.812024
Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-Based Visual Relationship Detection · CVPR 2024
Machine learning › Transfer learning and domain adaptation › cross-domain transfer
cross-dataset transfer
0.312025
Super-Class Guided Transformer for Zero-Shot Attribute Classification · AAAI 2025

Methods — techniques the papers use, named apart from their topics

transformer · 1.6super-class query initialization · 0.9consistency regularization · 0.9retrieval augmentation · 0.8pseudo-labeling · 0.8multi-assignment · 0.8large language model · 0.8
YearPublicationVenuePosition
2026 Improved query specialization for transformer-based visual relationship detection
abstract
Visual Relationship Detection (VRD) has significantly advanced with Transformer-based architectures. However, we identify two fundamental drawbacks in conventional label assignment methods used for training Transformer-based VRD models, where ground-truth (GT) annotations are matched to model predictions. In conventional assignment, queries are trained to detect all relations rather than specializing in specific ones, resulting in ‘unspecialized’ queries. Also, each ground-truth (GT) annotation is assigned to only one prediction under conventional assignment, suppressing other near-correct predictions by labeling them as ‘no relation’. To address these issues, we introduce a novel method called Groupwise Query Spe ci a lization and Q uality-Aware Multi-Assignment (SpeaQ). Groupwise Query Specialization clusters queries and relations into exclusive groups, promoting specialization by assigning a set of relations only to a corresponding query group. Quality-Aware Multi-Assignment enhances training signals by allowing multiple predictions closely matching the GT to be positively assigned. Additionally, we introduce dynamic query reallocation, which transfers queries from high- to low-performing groups for balanced training. Experimental results demonstrate that SpeaQ+, combining SpeaQ with dynamic query reallocation, consistently improves performance across seven baseline models on five benchmarks without additional inference cost.
Jongha Kim, Jinyoung Park 0005, Jinyoung Kim 0007, Sehyung Kim, Hyunwoo J. Kim
Inf. Sci.5
2025 Super-Class Guided Transformer for Zero-Shot Attribute Classification
abstract
Attribute classification is crucial for identifying specific characteristics within image regions. Vision-Language Models (VLMs) have been effective in zero-shot tasks by leveraging their general knowledge from large-scale datasets. Recent studies demonstrate that transformer-based models with class-wise queries can effectively address zero-shot multi-label classification. However, poor utilization of the relationship between seen and unseen attributes makes the model lack generalizability. Additionally, attribute classification generally involves many attributes, making maintaining the model’s scalability difficult. To address these issues, we propose Super-class guided transFormer (SugaFormer), a novel framework that leverages super-classes to enhance scalability and generalizability for zero-shot attribute classification. SugaFormer employs Super-class Query Initialization (SQI) to reduce the number of queries, utilizing common semantic information from super-classes, and incorporates Multi-context Decoding (MD) to handle diverse visual cues. To strengthen generalizability, we introduce two knowledge transfer strategies that utilize VLMs. During training, Super-class guided Consistency Regularization (SCR) aligns model’s features with VLMs using super-class guided prompts, and during inference, Zero-shot Retrieval-based Score Enhancement (ZRSE) refines predictions for unseen attributes. Extensive experiments demonstrate that SugaFormer achieves state-of-the-art performance across three widely-used attribute classification benchmarks under zero-shot, and cross-dataset transfer settings.
Sehyung Kim, Chanhyeong Yang, Taehoon Song, Hyunwoo J. Kim
AAAI1
2024 Retrieval-Augmented Open-Vocabulary Object Detection
abstract
Open-vocabulary object detection (OVD) has been stud-ied with Vision-Language Models (VLMs) to detect novel objects beyond the pre-trained categories. Previous ap-proaches improve the generalization ability to expand the knowledge of the detector, using ‘positive’ pseudo-labels with additional ‘class' names, e.g., sock, iPod, and alli-gator. To extend the previous methods in two aspects, we propose Retrieval-Augmented Losses and visual Features (RALF). Our method retrieves related ‘negative’ classes and augments loss functions. Also, visual features are aug-mented with ‘verbalized concepts' of classes, e.g., worn on the feet, handheld music player, and sharp teeth. Specif-ically, RALF consists of two modules: Retrieval Aug-mented Losses (RAL) and Retrieval-Augmented visual Fea-tures (RAF). RAL constitutes two losses reflecting the se-mantic similarity with negative vocabularies. In addition, RAF augments visual features with the verbalized con-cepts from a large language model (LLM). Our experiments demonstrate the effectiveness of RALF on COCO and LVIS benchmark datasets. We achieve improvement up to 3.4 box APN50on novel categories of the COCO dataset and 3.6 mask APr gains on the LVIS dataset. Code is available at https://github.com/mlvlab/RALF.
Eulrang Cho, Sehyung Kim, Hyunwoo J. Kim
CVPR3
2024 Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-Based Visual Relationship Detection
abstract
Visual Relationship Detection (VRD) has seen significant advancements with Transformer-based architectures recently. However, we identify two key limitations in a conventional label assignment for training Transformer-based VRD models, which is a process of mapping a ground-truth (GT) to a prediction. Under the conventional assignment, an ‘unspecialized’ query is trained since a query is expected to detect every relation, which makes it difficult for a query to specialize in specific relations. Furthermore, a query is also insufficiently trained since a GT is assigned only to a single prediction, therefore near-correct or even correct predictions are suppressed by being assigned ‘no relation (⊘)‘ as a GT. To address these issues, we propose Groupwise Query Specialization and Quality-Aware Multi-Assignment (SpeaQ). Groupwise Query Specialization trains a ‘specialized’ query by dividing queries and relations into disjoint groups and directing a query in a specific query group solely toward relations in the corresponding relation group. Quality-Aware Multi-Assignment further facilitates the training by assigning a GT to multiple predictions that are significantly close to a GT in terms of a subject, an object, and the relation in between. Experimental results and analyses show that SpeaQ effectively trains ‘specialized’ queries, which better utilize the capacity of a model, resulting in consistent performance gains with ‘zero’ additional inference cost across multiple VRD models and benchmarks. Code is available at https://github.com/m1vlab/SpeaQ.
Jongha Kim, Jinyoung Park 0005, Jinyoung Kim 0007, Sehyung Kim, Hyunwoo J. Kim
CVPR5