Seongheon Park

dblp:331/8077 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 35% Vision and language · 19% Trustworthy machine learning · 15%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
hallucination detection
1.722025
GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity · NeurIPS 2025
Steer LLM Latents for Hallucination Detection · ICML 2025
Machine learning › Trustworthy machine learning
uncertainty estimation
1.012026
Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model
large language model representation
0.912025
Steer LLM Latents for Hallucination Detection · ICML 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity · NeurIPS 2025
Computer vision › Vision and language › multimodal hallucination
object hallucination
0.912025
GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity · NeurIPS 2025
Natural language and speech › Language models and text generation › model steering
representation steering
0.912025
Steer LLM Latents for Hallucination Detection · ICML 2025
Machine learning › Transfer learning and domain adaptation › zero-shot learning › compositional zero-shot learning
attribute-object composition
0.712023
Hierarchical Visual Primitive Experts for Compositional Zero-Shot Learning · ICCV 2023
Machine learning › Transfer learning and domain adaptation › zero-shot learning
compositional zero-shot learning
0.712023
Hierarchical Visual Primitive Experts for Compositional Zero-Shot Learning · ICCV 2023
Machine learning › Deep learning architectures and training
data augmentation
0.712023
PartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-Identification · CVPR 2023
Machine learning › Trustworthy machine learning
fairness and bias
0.712023
Hierarchical Visual Primitive Experts for Compositional Zero-Shot Learning · ICCV 2023
Computer vision › Face, body and person analysis
person re-identification
0.712023
PartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-Identification · CVPR 2023
Computer vision › Face, body and person analysis › person re-identification › multi-modal person re-identification
visible-infrared person re-identification
0.712023
PartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-Identification · CVPR 2023
Natural language and speech › Language models and text generation
LLM agents
0.312026
Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

pseudo-labeling · 0.9optimal transport · 0.9global-local similarity · 0.9embedding similarity · 0.9confidence filtering · 0.9object-guided attention · 0.7mixup · 0.7minority attribute augmentation · 0.7entropy-based mining · 0.7contrastive learning · 0.7
YearPublicationVenuePosition
2026 Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities
abstract
Changdae Oh, Seongheon Park, To Eun Kim, Jiatong Li, Wendi Li, Samuel Yeh, Sean Du, Hamed Hassani, Paul Bogdan, Dawn Song, Sharon Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Changdae Oh, Seongheon Park, To Eun Kim, Wendi Li, Samuel Yeh 0001, Sean Du, Seyed Hamed Hassani, Paul Bogdan, Dawn Song, Yixuan Li 0001
ACL (1)2
2025 Steer LLM Latents for Hallucination Detection
abstract
Hallucinations in LLMs pose a significant concern to their safe deployment in real-world applications. Recent approaches have leveraged the latent space of LLMs for hallucination detection, but their embeddings, optimized for linguistic coherence rather than factual accuracy, often fail to clearly separate truthful and hallucinated content. To this end, we propose the **T**ruthfulness **S**eparator **V**ector (**TSV**), a lightweight and flexible steering vector that reshapes the LLM’s representation space during inference to enhance the separation between truthful and hallucinated outputs, without altering model parameters. Our two-stage framework first trains TSV on a small set of labeled exemplars to form compact and well-separated clusters. It then augments the exemplar set with unlabeled LLM generations, employing an optimal transport-based algorithm for pseudo-labeling combined with a confidence-based filtering process. Extensive experiments demonstrate that TSV achieves state-of-the-art performance with minimal labeled data, exhibiting strong generalization across datasets and providing a practical solution for real-world LLM applications.
Seongheon Park, Xuefeng Du, Min-Hsuan Yeh, Haobo Wang 0001, Yixuan Li 0001
ICML1
2025 GeoRanker: Distance-Aware Ranking for Worldwide Image Geolocalization
abstract
Worldwide image geolocalization—the task of predicting GPS coordinates from images taken anywhere on Earth—poses a fundamental challenge due to the vast diversity in visual content across regions. While recent approaches adopt a two-stage pipeline of retrieving candidates and selecting the best match, they typically rely on simplistic similarity heuristics and point-wise supervision, failing to model spatial relationships among candidates. In this paper, we propose **GeoRanker**, a distance-aware ranking framework that leverages large vision-language models to jointly encode query–candidate interactions and predict geographic proximity. In addition, we introduce a *multi-order distance loss* that ranks both absolute and relative distances, enabling the model to reason over structured spatial relationships. To support this, we curate GeoRanking, the first dataset explicitly designed for geographic ranking tasks with multimodal candidate information. GeoRanker achieves state-of-the-art results on two well-established benchmarks (IM2GPS3K and YFCC4K), significantly outperforming current best methods. We also release our code, checkpoint, and dataset online for ease of reproduction.
Pengyue Jia, Seongheon Park, Xiangyu Zhao 0001, Yixuan Li 0001
NeurIPS2
2025 GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
abstract
Object hallucination in large vision-language models presents a significant challenge to their safe deployment in real-world applications. Recent works have proposed object-level hallucination scores to estimate the likelihood of object hallucination; however, these methods typically adopt either a global or local perspective in isolation, which may limit detection reliability. In this paper, we introduce GLSim, a novel training-free object hallucination detection framework that leverages complementary global and local embedding similarity signals between image and text modalities, enabling more accurate and reliable hallucination detection in diverse scenarios. We comprehensively benchmark existing object hallucination detection methods and demonstrate that GLSim achieves superior detection performance, outperforming competitive baselines by a significant margin.
Seongheon Park, Yixuan Li 0001
NeurIPS1
2023 PartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-Identification
abstract
Modern data augmentation using a mixture-based technique can regularize the models from overfitting to the training data in various computer vision applications, but a proper data augmentation technique tailored for the part-based Visible-Infrared person Re-IDentification (VI-ReID) models remains unexplored. In this paper, we present a novel data augmentation technique, dubbed PartMix, that synthesizes the augmented samples by mixing the part descriptors across the modalities to improve the performance of part-based VI-ReID models. Especially, we synthesize the positive and negative samples within the same and across different identities and regularize the backbone model through contrastive learning. In addition, we also present an entropy-based mining strategy to weaken the adverse impact of unreliable positive and negative samples. When incorporated into existing part-based VI-ReID model, PartMix consistently boosts the performance. We conduct experiments to demonstrate the effectiveness of our PartMix over the existing VI-ReID methods and provide ablation studies.
Seungryong Kim, Jungin Park, Seongheon Park, Kwanghoon Sohn
CVPR4
2023 Hierarchical Visual Primitive Experts for Compositional Zero-Shot Learning
abstract
Compositional zero-shot learning (CZSL) aims to recognize unseen compositions with prior knowledge of known primitives (attribute and object). Previous works for CZSL often suffer from grasping the contextuality between attribute and object, as well as the discriminability of visual features, and the long-tailed distribution of real-world compositional data. We propose a simple and scalable framework called Composition Transformer (CoT) to address these issues. CoT employs object and attribute experts in distinctive manners to generate representative embeddings, using the visual network hierarchically. The object expert extracts representative object embeddings from the final layer in a bottom-up manner, while the attribute expert makes attribute embeddings in a top-down manner with a proposed object-guided attention module that models contextuality explicitly. To remedy biased prediction caused by imbalanced data distribution, we develop a simple minority attribute augmentation (MAA) that synthesizes virtual samples by mixing two images and oversampling minority attribute classes. Our method achieves SoTA performance on several benchmarks, including MIT-States, C-GQA, and VAW-CZSL. We also demonstrate the effectiveness of CoT in improving visual discrimination and addressing the model bias from the imbalanced data distribution. The code is available at https://github.com/HanjaeKim98/CoT.
Hanjae Kim, Jiyoung Lee 0005, Seongheon Park, Kwanghoon Sohn
ICCV3
2023 Language-free Training for Zero-shot Video Grounding
abstract
Given an untrimmed video and a language query depicting a specific temporal moment in the video, video grounding aims to localize the time interval by understanding the text and video simultaneously. One of the most challenging issues is an extremely time- and cost-consuming annotation collection, including video captions in a natural language form and their corresponding temporal regions. In this paper, we present a simple yet novel training framework for video grounding in the zero-shot setting, which learns a network with only video data without any annotation. Inspired by the recent language-free paradigm, i.e. training without language data, we train the network without compelling the generation of fake (pseudo) text queries into a natural language form. Specifically, we propose a method for learning a video grounding model by selecting a temporal interval as a hypothetical correct answer and considering the visual feature selected by our method in the interval as a language feature, with the help of the well-aligned visual-language space of CLIP. Extensive experiments demonstrate the prominence of our language-free training framework, outperforming the existing zero-shot video grounding method and even several weakly-supervised approaches with large margins on two standard datasets.
Dahye Kim 0004, Jungin Park, Jiyoung Lee 0005, Seongheon Park, Kwanghoon Sohn
WACV4
2023 Normality Guided Multiple Instance Learning for Weakly Supervised Video Anomaly Detection
abstract
Weakly supervised Video Anomaly Detection (wVAD) aims to distinguish anomalies from normal events based on video-level supervision. Most existing works utilize Multiple Instance Learning (MIL) with ranking loss to tackle this task. These methods, however, rely on noisy predictions from a MIL-based classifier for target instance selection in ranking loss, degrading model performance. To overcome this problem, we propose Normality Guided Multiple Instance Learning (NG-MIL) framework, which encodes diverse normal patterns from noise-free normal videos into prototypes for constructing a similarity-based classifier. By ensembling predictions of two classifiers, our method could refine the anomaly scores, reducing training instability from weak labels. Moreover, we introduce normality clustering and normality guided triplet loss constraining inner bag instances to boost the effect of NG-MIL and increase the discriminability of classifiers. Extensive experiments on three public datasets (ShanghaiTech, UCF-Crime, XD-Violence) demonstrate that our method is comparable to or better than existing weakly supervised methods, achieving state-of-the-art results.
Seongheon Park, Hanjae Kim, Dahye Kim 0004, Kwanghoon Sohn
WACV1