VLDB 2026 Research / reviewers in the wild / expert
Seongheon Park
dblp:331/8077
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 35% Vision and language · 19% Trustworthy machine learning · 15% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
hallucination detection |
1.7 | 2 | 2025 | GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity · NeurIPS 2025 Steer LLM Latents for Hallucination Detection · ICML 2025 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
1.0 | 1 | 2026 | Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities · ACL (1) 2026 |
Natural language and speech › Language models and text generation › large language model
large language model representation |
0.9 | 1 | 2025 | Steer LLM Latents for Hallucination Detection · ICML 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity · NeurIPS 2025 |
Computer vision › Vision and language › multimodal hallucination
object hallucination |
0.9 | 1 | 2025 | GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity · NeurIPS 2025 |
Natural language and speech › Language models and text generation › model steering
representation steering |
0.9 | 1 | 2025 | Steer LLM Latents for Hallucination Detection · ICML 2025 |
Machine learning › Transfer learning and domain adaptation › zero-shot learning › compositional zero-shot learning
attribute-object composition |
0.7 | 1 | 2023 | Hierarchical Visual Primitive Experts for Compositional Zero-Shot Learning · ICCV 2023 |
Machine learning › Transfer learning and domain adaptation › zero-shot learning
compositional zero-shot learning |
0.7 | 1 | 2023 | Hierarchical Visual Primitive Experts for Compositional Zero-Shot Learning · ICCV 2023 |
Machine learning › Deep learning architectures and training
data augmentation |
0.7 | 1 | 2023 | PartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-Identification · CVPR 2023 |
Machine learning › Trustworthy machine learning
fairness and bias |
0.7 | 1 | 2023 | Hierarchical Visual Primitive Experts for Compositional Zero-Shot Learning · ICCV 2023 |
Computer vision › Face, body and person analysis
person re-identification |
0.7 | 1 | 2023 | PartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-Identification · CVPR 2023 |
Computer vision › Face, body and person analysis › person re-identification › multi-modal person re-identification
visible-infrared person re-identification |
0.7 | 1 | 2023 | PartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-Identification · CVPR 2023 |
Natural language and speech › Language models and text generation
LLM agents |
0.3 | 1 | 2026 | Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
pseudo-labeling · 0.9optimal transport · 0.9global-local similarity · 0.9embedding similarity · 0.9confidence filtering · 0.9object-guided attention · 0.7mixup · 0.7minority attribute augmentation · 0.7entropy-based mining · 0.7contrastive learning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and OpportunitiesabstractChangdae Oh, Seongheon Park, To Eun Kim, Jiatong Li, Wendi Li, Samuel Yeh, Sean Du, Hamed Hassani, Paul Bogdan, Dawn Song, Sharon Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Changdae Oh, Seongheon Park, To Eun Kim, Wendi Li, Samuel Yeh 0001, Sean Du, Seyed Hamed Hassani, Paul Bogdan, Dawn Song, Yixuan Li 0001 |
ACL (1) | 2 |
| 2025 | Steer LLM Latents for Hallucination DetectionabstractHallucinations in LLMs pose a significant concern to their safe deployment in real-world applications. Recent approaches have leveraged the latent space of LLMs for hallucination detection, but their embeddings, optimized for linguistic coherence rather than factual accuracy, often fail to clearly separate truthful and hallucinated content.
To this end, we propose the **T**ruthfulness **S**eparator **V**ector (**TSV**), a lightweight and flexible steering vector that reshapes the LLM’s representation space during inference to enhance the separation between truthful and hallucinated outputs, without altering model parameters.
Our two-stage framework first trains TSV on a small set of labeled exemplars to form compact and well-separated clusters.
It then augments the exemplar set with unlabeled LLM generations, employing an optimal transport-based algorithm for pseudo-labeling combined with a confidence-based filtering process.
Extensive experiments demonstrate that TSV achieves state-of-the-art performance with minimal labeled data, exhibiting strong generalization across datasets and providing a practical solution for real-world LLM applications. Seongheon Park, Xuefeng Du, Min-Hsuan Yeh, Haobo Wang 0001, Yixuan Li 0001 |
ICML | 1 |
| 2025 | GeoRanker: Distance-Aware Ranking for Worldwide Image GeolocalizationabstractWorldwide image geolocalization—the task of predicting GPS coordinates from images taken anywhere on Earth—poses a fundamental challenge due to the vast diversity in visual content across regions. While recent approaches adopt a two-stage pipeline of retrieving candidates and selecting the best match, they typically rely on simplistic similarity heuristics and point-wise supervision, failing to model spatial relationships among candidates. In this paper, we propose **GeoRanker**, a distance-aware ranking framework that leverages large vision-language models to jointly encode query–candidate interactions and predict geographic proximity. In addition, we introduce a *multi-order distance loss* that ranks both absolute and relative distances, enabling the model to reason over structured spatial relationships. To support this, we curate GeoRanking, the first dataset explicitly designed for geographic ranking tasks with multimodal candidate information. GeoRanker achieves state-of-the-art results on two well-established benchmarks (IM2GPS3K and YFCC4K), significantly outperforming current best methods. We also release our code, checkpoint, and dataset online for ease of reproduction. Pengyue Jia, Seongheon Park, Xiangyu Zhao 0001, Yixuan Li 0001 |
NeurIPS | 2 |
| 2025 | GLSim: Detecting Object Hallucinations in LVLMs via Global-Local SimilarityabstractObject hallucination in large vision-language models presents a significant challenge to their safe deployment in real-world applications. Recent works have proposed object-level hallucination scores to estimate the likelihood of object hallucination; however, these methods typically adopt either a global or local perspective in isolation, which may limit detection reliability. In this paper, we introduce GLSim, a novel training-free object hallucination detection framework that leverages complementary global and local embedding similarity signals between image and text modalities, enabling more accurate and reliable hallucination detection in diverse scenarios. We comprehensively benchmark existing object hallucination detection methods and demonstrate that GLSim achieves superior detection performance, outperforming competitive baselines by a significant margin. Seongheon Park, Yixuan Li 0001 |
NeurIPS | 1 |
| 2023 | PartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-IdentificationabstractModern data augmentation using a mixture-based technique can regularize the models from overfitting to the training data in various computer vision applications, but a proper data augmentation technique tailored for the part-based Visible-Infrared person Re-IDentification (VI-ReID) models remains unexplored. In this paper, we present a novel data augmentation technique, dubbed PartMix, that synthesizes the augmented samples by mixing the part descriptors across the modalities to improve the performance of part-based VI-ReID models. Especially, we synthesize the positive and negative samples within the same and across different identities and regularize the backbone model through contrastive learning. In addition, we also present an entropy-based mining strategy to weaken the adverse impact of unreliable positive and negative samples. When incorporated into existing part-based VI-ReID model, PartMix consistently boosts the performance. We conduct experiments to demonstrate the effectiveness of our PartMix over the existing VI-ReID methods and provide ablation studies. Seungryong Kim, Jungin Park, Seongheon Park, Kwanghoon Sohn |
CVPR | 4 |
| 2023 | Hierarchical Visual Primitive Experts for Compositional Zero-Shot LearningabstractCompositional zero-shot learning (CZSL) aims to recognize unseen compositions with prior knowledge of known primitives (attribute and object). Previous works for CZSL often suffer from grasping the contextuality between attribute and object, as well as the discriminability of visual features, and the long-tailed distribution of real-world compositional data. We propose a simple and scalable framework called Composition Transformer (CoT) to address these issues. CoT employs object and attribute experts in distinctive manners to generate representative embeddings, using the visual network hierarchically. The object expert extracts representative object embeddings from the final layer in a bottom-up manner, while the attribute expert makes attribute embeddings in a top-down manner with a proposed object-guided attention module that models contextuality explicitly. To remedy biased prediction caused by imbalanced data distribution, we develop a simple minority attribute augmentation (MAA) that synthesizes virtual samples by mixing two images and oversampling minority attribute classes. Our method achieves SoTA performance on several benchmarks, including MIT-States, C-GQA, and VAW-CZSL. We also demonstrate the effectiveness of CoT in improving visual discrimination and addressing the model bias from the imbalanced data distribution. The code is available at https://github.com/HanjaeKim98/CoT. Hanjae Kim, Jiyoung Lee 0005, Seongheon Park, Kwanghoon Sohn |
ICCV | 3 |
| 2023 | Language-free Training for Zero-shot Video GroundingabstractGiven an untrimmed video and a language query depicting a specific temporal moment in the video, video grounding aims to localize the time interval by understanding the text and video simultaneously. One of the most challenging issues is an extremely time- and cost-consuming annotation collection, including video captions in a natural language form and their corresponding temporal regions. In this paper, we present a simple yet novel training framework for video grounding in the zero-shot setting, which learns a network with only video data without any annotation. Inspired by the recent language-free paradigm, i.e. training without language data, we train the network without compelling the generation of fake (pseudo) text queries into a natural language form. Specifically, we propose a method for learning a video grounding model by selecting a temporal interval as a hypothetical correct answer and considering the visual feature selected by our method in the interval as a language feature, with the help of the well-aligned visual-language space of CLIP. Extensive experiments demonstrate the prominence of our language-free training framework, outperforming the existing zero-shot video grounding method and even several weakly-supervised approaches with large margins on two standard datasets. Dahye Kim 0004, Jungin Park, Jiyoung Lee 0005, Seongheon Park, Kwanghoon Sohn |
WACV | 4 |
| 2023 | Normality Guided Multiple Instance Learning for Weakly Supervised Video Anomaly DetectionabstractWeakly supervised Video Anomaly Detection (wVAD) aims to distinguish anomalies from normal events based on video-level supervision. Most existing works utilize Multiple Instance Learning (MIL) with ranking loss to tackle this task. These methods, however, rely on noisy predictions from a MIL-based classifier for target instance selection in ranking loss, degrading model performance. To overcome this problem, we propose Normality Guided Multiple Instance Learning (NG-MIL) framework, which encodes diverse normal patterns from noise-free normal videos into prototypes for constructing a similarity-based classifier. By ensembling predictions of two classifiers, our method could refine the anomaly scores, reducing training instability from weak labels. Moreover, we introduce normality clustering and normality guided triplet loss constraining inner bag instances to boost the effect of NG-MIL and increase the discriminability of classifiers. Extensive experiments on three public datasets (ShanghaiTech, UCF-Crime, XD-Violence) demonstrate that our method is comparable to or better than existing weakly supervised methods, achieving state-of-the-art results. Seongheon Park, Hanjae Kim, Dahye Kim 0004, Kwanghoon Sohn |
WACV | 1 |