EDBT 2026 Demo / reviewers in the wild / expert
Eunseop Yoon
dblp:331/3764
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0002-5580-5354ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Trustworthy machine learning · 26% Language models and text generation · 25% Vision and language · 12% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 26 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
calibration |
1.4 | 2 | 2024 | C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion · ICLR 2024 ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure · ICLR 2023 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
1.4 | 2 | 2024 | C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion · ICLR 2024 ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure · ICLR 2023 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | A Gradient Guidance Perspective on Stepwise Preference Optimization for Diffusion Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation › alignment › preference alignment
direct alignment algorithms |
0.9 | 1 | 2025 | ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization · ICML 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation › alignment
preference alignment |
0.9 | 1 | 2025 | A Gradient Guidance Perspective on Stepwise Preference Optimization for Diffusion Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization · ICML 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | A Gradient Guidance Perspective on Stepwise Preference Optimization for Diffusion Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model training
token-level credit assignment |
0.9 | 1 | 2025 | ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization · ICML 2025 |
Computer vision › Vision and language › video-language model
video large language model |
0.9 | 1 | 2025 | Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models · ICLR 2025 |
Machine learning › Deep learning architectures and training
data augmentation |
0.8 | 1 | 2024 | SimPSI: A Simple Strategy to Preserve Spectral Information in Time Series Data Augmentation · AAAI 2024 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation |
0.8 | 1 | 2024 | BI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation · ECCV (31) 2024 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
multimodal response generation |
0.8 | 1 | 2024 | BI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation · ECCV (31) 2024 |
Natural language and speech › Language models and text generation
prompt tuning |
0.8 | 1 | 2024 | C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion · ICLR 2024 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
spectral representation |
0.8 | 1 | 2024 | SimPSI: A Simple Strategy to Preserve Spectral Information in Time Series Data Augmentation · AAAI 2024 |
Machine learning › Transfer learning and domain adaptation
test-time adaptation |
0.8 | 1 | 2024 | C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion · ICLR 2024 |
Machine learning › Trustworthy machine learning › calibration
test-time calibration |
0.8 | 1 | 2024 | C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion · ICLR 2024 |
Machine learning › Deep learning architectures and training › data augmentation
time series data augmentation |
0.8 | 1 | 2024 | SimPSI: A Simple Strategy to Preserve Spectral Information in Time Series Data Augmentation · AAAI 2024 |
Machine learning › Trustworthy machine learning › calibration
calibration measures |
0.7 | 1 | 2023 | ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure · ICLR 2023 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.6 | 1 | 2022 | Information-Theoretic Text Hallucination Reduction for Video-grounded Dialogue · EMNLP 2022 |
Machine learning › Trustworthy machine learning
robustness |
0.6 | 1 | 2022 | Information-Theoretic Text Hallucination Reduction for Video-grounded Dialogue · EMNLP 2022 |
Natural language and speech › Question answering and dialogue systems › multimodal dialogue system
video-grounded dialogue |
0.6 | 1 | 2022 | Information-Theoretic Text Hallucination Reduction for Video-grounded Dialogue · EMNLP 2022 |
Multimedia analysis and retrieval › video retrieval › video moment retrieval
video corpus moment retrieval |
0.6 | 1 | 2022 | Selective Query-Guided Debiasing for Video Corpus Moment Retrieval · ECCV (36) 2022 |
Computer vision › Video understanding and tracking
video question answering |
0.3 | 1 | 2025 | Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models · ICLR 2025 |
Computer vision › Vision and language › vision-language model
CLIP |
0.2 | 1 | 2024 | C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion · ICLR 2024 |
Computer vision › Vision and language
vision-language model |
0.2 | 1 | 2024 | C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
spectral mixing · 1.5preservation map · 1.5variance-informed noise reduction · 0.9gradient guidance · 0.9direct preference optimization · 0.9dataset construction · 0.9confidence-based token selection · 0.9alignment training · 0.9KL divergence · 0.9image history bridging · 0.8debiasing · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language ModelsabstractIn the broader context of deep learning, Multimodal Large Language Models have achieved significant breakthroughs by leveraging powerful Large Language Models as a backbone to align different modalities into the language space. A prime exemplification is the development of Video Large Language Models (Video-LLMs). While numerous advancements have been proposed to enhance the video understanding capabilities of these models, they are predominantly trained on questions generated directly from video content. However, in real-world scenarios, users often pose questions that extend beyond the informational scope of the video, highlighting the need for Video-LLMs to assess the relevance of the question. We demonstrate that even the best-performing Video-LLMs fail to reject unfit questions-not necessarily due to a lack of video understanding, but because they have not been trained to identify and refuse such questions. To address this limitation, we propose alignment for answerability, a framework that equips Video-LLMs with the ability to evaluate the relevance of a question based on the input video and appropriately decline to answer when the question exceeds the scope of the video, as well as an evaluation framework with a comprehensive set of metrics designed to measure model behavior before and after alignment. Furthermore, we present a pipeline for creating a dataset specifically tailored for alignment for answerability, leveraging existing video-description paired datasets. Eunseop Yoon, Hee Suk Yoon, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICLR | 1 |
| 2025 | ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference OptimizationabstractWe introduce ConfPO, a method for preference learning in Large Language Models (LLMs) that identifies and optimizes preference-critical tokens based solely on the training policy’s confidence, without requiring any auxiliary models or compute. Unlike prior Direct Alignment Algorithms (DAAs) such as Direct Preference Optimization (DPO), which uniformly adjust all token probabilities regardless of their relevance to preference, ConfPO focuses optimization on the most impactful tokens. This targeted approach improves alignment quality while mitigating overoptimization (i.e., reward hacking) by using the KL divergence budget more efficiently. In contrast to recent token-level methods that rely on credit-assignment models or AI annotators, raising concerns about scalability and reliability, ConfPO is simple, lightweight, and model-free. Experimental results on challenging alignment benchmarks, including AlpacaEval 2 and Arena-Hard, demonstrate that ConfPO consistently outperforms uniform DAAs across various LLMs, delivering better alignment with zero additional computational overhead. Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson, Sungwoong Kim, Chang Dong Yoo |
ICML | 2 |
| 2025 | A Gradient Guidance Perspective on Stepwise Preference Optimization for Diffusion ModelsabstractDirect Preference Optimization (DPO) is a key framework for aligning text-to-image models with human preferences, extended by Stepwise Preference Optimization (SPO) to leverage intermediate steps for preference learning, generating more aesthetically pleasing images with significantly less computational cost. While effective, SPO's underlying mechanisms remain underexplored. In light of this, we critically re-examine SPO by formalizing its mechanism as gradient guidance. This new lens shows that SPO uses biased temporal weighting, giving too little weight to later generative steps, and unlike likelihood centric views it reveals substantial noise in the gradient estimates. Leveraging these insights, our GradSPO algorithm introduces a simplified loss and a targeted, variance-informed noise reduction strategy, enhancing training stability. Evaluations on SD 1.5 and SDXL show GradSPO substantially outperforms leading baselines in human preference, yielding images with markedly improved aesthetics and semantic faithfulness, leading to more robust alignment. Code and models are available at https://github.com/JoshuaTTJ/GradSPO. Joshua Tian Jin Tee, Hee Suk Yoon, Abu Hanif Muhammad Syarubany, Eunseop Yoon, Chang Dong Yoo |
NeurIPS | 4 |
| 2024 | SimPSI: A Simple Strategy to Preserve Spectral Information in Time Series Data AugmentationabstractData augmentation is a crucial component in training neural networks to overcome the limitation imposed by data size, and several techniques have been studied for time series. Although these techniques are effective in certain tasks, they have yet to be generalized to time series benchmarks. We find that current data augmentation techniques ruin the core information contained within the frequency domain. To address this issue, we propose a simple strategy to preserve spectral information (SimPSI) in time series data augmentation. SimPSI preserves the spectral information by mixing the original and augmented input spectrum weighted by a preservation map, which indicates the importance score of each frequency. Specifically, our experimental contributions are to build three distinct preservation maps: magnitude spectrum, saliency map, and spectrum-preservative map. We apply SimPSI to various time series data augmentations and evaluate its effectiveness across a wide range of time series benchmarks. Our experimental results support that SimPSI considerably enhances the performance of time series data augmentations by preserving core spectral information. The source code used in the paper is available at https://github.com/Hyun-Ryu/simpsi. Hyun Ryu, Sunjae Yoon, Hee Suk Yoon, Eunseop Yoon, Chang Dong Yoo |
AAAI | 4 |
| 2024 | BI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation
Hee Suk Yoon, Eunseop Yoon, Joshua Tian Jin Tee, Kang Zhang 0008, Yu-Jung Heo, Du-Seong Chang, Chang Dong Yoo |
ECCV (31) | 2 |
| 2024 | AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech RecognitionabstractIn Automatic Speech Recognition (ASR) systems, a recurring obstacle is the generation of narrowly focused output distributions. This phenomenon emerges as a side effect of Connectionist Temporal Classification (CTC), a robust sequence learning tool that utilizes dynamic programming for sequence mapping. While earlier efforts have tried to combine the CTC loss with an entropy maximization regularization term to mitigate this issue, they employed a constant weighting term on the regularization during the training, which we find may not be optimal. In this work, we introduce Adaptive Maximum Entropy Regularization (AdaMER), a technique that can modulate the impact of entropy regularization throughout the training process. This approach not only refines ASR model training but ensures that as training proceeds, predictions display the desired model confidence. SooHwan Eom, Eunseop Yoon, Hee Suk Yoon, Chanwoo Kim 0001, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICASSP | 2 |
| 2024 | C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature DispersionabstractIn deep learning, test-time adaptation has gained attention as a method for model fine-tuning without the need for labeled data. A prime exemplification is the recently proposed test-time prompt tuning for large-scale vision-language models such as CLIP. Unfortunately, these prompts have been mainly developed to improve accuracy, overlooking the importance of calibration, which is a crucial aspect for quantifying prediction uncertainty. However, traditional calibration methods rely on substantial amounts of labeled data, making them impractical for test-time scenarios. To this end, this paper explores calibration during test-time prompt tuning by leveraging the inherent properties of CLIP. Through a series of observations, we find that the prompt choice significantly affects the calibration in CLIP, where the prompts leading to higher text feature dispersion result in better-calibrated predictions. Introducing the Average Text Feature Dispersion (ATFD), we establish its relationship with calibration error and present a novel method, Calibrated Test-time Prompt Tuning (C-TPT), for optimizing prompts during test-time with enhanced calibration. Through extensive experiments on different CLIP architectures and datasets, we show that C-TPT can effectively improve the calibration of test-time prompt tuning without needing labeled data. The code is publicly accessible at https://github.com/hee-suk-yoon/C-TPT. Hee Suk Yoon, Eunseop Yoon, Joshua Tian Jin Tee, Mark Hasegawa-Johnson, Yingzhen Li, Chang Dong Yoo |
ICLR | 2 |
| 2024 | LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition
Eunseop Yoon, Hee Suk Yoon, John B. Harvill, Mark Hasegawa-Johnson, Chang Dong Yoo |
INTERSPEECH | 1 |
| 2023 | Counterfactual Two-Stage Debiasing For Video Corpus Moment RetrievalabstractVideo Corpus Moment Retrieval aims to select a temporal video moment pertinent to a given language query from a large video corpus. Existing systems are prone to rely on a retrieval bias as a shortcut, which hinders the systems from accurately learning vision-language association. The retrieval bias is spurious correlations between query and scene. For a given query, systems tend to retrieve incorrectly correlated scenes due to biased annotations that have predominant binding in a dataset. To this end, we present a Counterfactual Two-stage Debiasing Learning (CTDL), which incorporates a counterfactual bias network that intentionally learns the retrieval bias by providing a shortcut to learn the spurious correlation between keyword and scene, and performs two-stage debiasing learning that mitigates the bias via contrasting factual retrievals with counterfactually biased retrievals. Extensive experiments show the effectiveness of CTDL paradigm. Sunjae Yoon, Ji Woo Hong, SooHwan Eom, Hee Suk Yoon, Eunseop Yoon, Daehyeok Kim, Junyeong Kim, Chanwoo Kim 0001, Chang Dong Yoo |
ICASSP | 5 |
| 2023 | ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure
Hee Suk Yoon, Joshua Tian Jin Tee, Eunseop Yoon, Sunjae Yoon, Gwangsu Kim, Yingzhen Li, Chang Dong Yoo |
ICLR | 3 |
| 2023 | Mitigating the Exposure Bias in Sentence-Level Grapheme-to-Phoneme (G2P) Transduction
Eunseop Yoon, Hee Suk Yoon, Dhananjaya Gowda, SooHwan Eom, Daehyeok Kim, John B. Harvill, Heting Gao, Mark Hasegawa-Johnson, Chanwoo Kim 0001, Chang Dong Yoo |
INTERSPEECH | 1 |
| 2022 | Selective Query-Guided Debiasing for Video Corpus Moment Retrieval
Sunjae Yoon, Ji Woo Hong, Eunseop Yoon, Dahyun Kim 0002, Junyeong Kim, Hee Suk Yoon, Chang Dong Yoo |
ECCV (36) | 3 |
| 2022 | Information-Theoretic Text Hallucination Reduction for Video-grounded DialogueabstractVideo-grounded Dialogue (VGD) aims to decode an answer sentence to a question regarding a given video and dialogue context. Despite the recent success of multi-modal reasoning to generate answer sentences, existing dialogue systems still suffer from a text hallucination problem, which denotes indiscriminate text-copying from input texts without an understanding of the question. This is due to learning spurious correlations from the fact that answer sentences in the dataset usually include the words of input texts, thus the VGD system excessively relies on copying words from input texts by hoping those words to overlap with ground-truth texts. Hence, we design Text Hallucination Mitigating (THAM) framework, which incorporates Text Hallucination Regularization (THR) loss derived from the proposed information-theoretic text hallucination measurement approach. Applying THAM with current dialogue systems validates the effectiveness on VGD benchmarks (i.e., AVSD@DSTC7 and AVSD@DSTC8) and shows enhanced interpretability. Sunjae Yoon, Eunseop Yoon, Hee Suk Yoon, Junyeong Kim, Chang Dong Yoo |
EMNLP | 2 |