VLDB 2026 Research / reviewers in the wild / expert
Jihwan Oh
dblp:207/7481
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0003-3274-4533ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 46% Multi-agent systems · 26% Generative modeling · 20% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
1.0 | 1 | 2026 | MERIT Feedback Elicits Better Bargaining in LLM Negotiators · ACL (1) 2026 |
Knowledge, reasoning and agents › Multi-agent systems
automated negotiation |
1.0 | 1 | 2026 | MERIT Feedback Elicits Better Bargaining in LLM Negotiators · ACL (1) 2026 |
Machine learning › Generative modeling
flow matching |
0.8 | 1 | 2024 | Preference Alignment with Flow Matching · NeurIPS 2024 |
Natural language and speech › Language models and text generation › alignment
preference alignment |
0.8 | 1 | 2024 | Preference Alignment with Flow Matching · NeurIPS 2024 |
Machine learning › Reinforcement learning
preference learning |
0.3 | 1 | 2026 | MERIT Feedback Elicits Better Bargaining in LLM Negotiators · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
fine-tuning · 1.8utility theory · 1.0prompting · 1.0flow matching · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MERIT Feedback Elicits Better Bargaining in LLM NegotiatorsabstractBargaining is often regarded as a logical arena rather than an art or a matter of intuition, yet Large Language Models (LLMs) still struggle to navigate it due to limited strategic depth and difficulty adapting to complex human factors.Current benchmarks rarely capture this limitation.To bridge this gap, we present a utility feedback centric framework.Our contributions are: (i) AGORABENCH, a new benchmark spanning nine challenging settings (e.g., deception, monopoly) that supports diverse strategy modeling; (ii) human-aligned, economically grounded metrics derived from utility theory.This is operationalized via agent utility, negotiation power, and acquisition ratio that implicitly measure how well the negotiation aligns with human preference and (iii) a human preference grounded dataset with learning pipeline that strengthens LLMs' bargaining ability through both prompting and finetuning.Empirical results indicate that baseline LLM strategies often diverge from human preferences, while our mechanism substantially improves negotiation performance, yielding deeper strategic behavior and stronger opponent awareness. Jihwan Oh, Murad Aghazada, Yooju Shin, Se-Young Yun, Taehyeon Kim 0001 |
ACL (1) | 1 |
| 2026 | From Belief Entrenchment to Robust Reasoning in LLM AgentsabstractAbstract Multi-Agent Debate (MAD) has emerged as a promising inference scaling method for Large Language Model (LLM) reasoning. However, it frequently suffers from belief entrenchment, where agents reinforce shared errors rather than correcting them. Going beyond merely identifying this failure, we decompose it into two distinct root causes: (1) the model’s biased static initial belief and (2) homogenized debate dynamics that amplify the majority view regardless of correctness. To address these sequentially, we propose DReaMAD (Diverse Reasoning via MultiAgent Debate). Our framework first rectifies the static belief via strategic prior knowledge elicitation, then reshapes the debate dynamics by enforcing perspective diversity. Validated on our new MetaNIM Arena benchmark, DReaMAD significantly mitigates entrenchment, achieving a +9.5% accuracy gain over ReAct prompting and a +19.0% higher win rate than standard MAD. Jihwan Oh, Minchan Jeong, Jongwoo Ko, Se-Young Yun |
Trans. Assoc. Comput. Linguistics | 1 |
| 2025 | Learning to Verify Summary Facts with Fine-Grained LLM FeedbackabstractTraining automatic summary fact verifiers often faces the challenge of a lack of human-labeled data. In this paper, we explore alternative way of leveraging Large Language Model (LLM) generated feedback to address the inherent limitation of using human-labeled data. We introduce FineSumFact, a large-scale dataset containing fine-grained factual feedback on summaries. We employ 10 distinct LLMs for diverse summary generation and Llama-3-70B-Instruct for feedback. We utilize this dataset to fine-tune the lightweight open-source model Llama-3-8B-Instruct, optimizing resource efficiency while maintaining high performance. Our experimental results reveal that the model trained on extensive LLM-generated datasets surpasses that trained on smaller human-annotated datasets when evaluated using human-generated test sets. Fine-tuning fact verification models with LLM feedback can be more effective and cost-efficient than using human feedback. The dataset is available at https://github.com/DISL-Lab/FineSumFact. Jihwan Oh, Jeonghwan Choi, Nicole Hee-Yeon Kim, Taewon Yun, Hwanjun Song |
COLING | 1 |
| 2025 | Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power ImplicationsabstractThis paper characterizes the interplay between overlapping compute and communication in GPU-accelerated distributed training, a critical but often overlooked aspect impacting performance and power. We systematically evaluate state-of-the-art NVIDIA (H100/A100) and AMD (MI250/MI210) GPUs, investigating the effects of hardware features including numeric precision, specialized cores, and power capping. Experiments across GPT and LLaMA models show that overlap causes significant compute slowdown (an average of 18.9 % and upto 40.0 %) compared to an ideal scenario where compute executes without communication interference. However, overlapping execution is still faster than sequential execution by an average 10.2 % and maximum of 26.6 %. We find overlap increases peak power, specialized datapaths have mixed effects, and power capping severely exacerbates slowdowns. These results emphasize the need for balanced strategies in optimizing distributed training throughput and power efficiency. Seonho Lee, Jihwan Oh, Seokjin Go, Divya Mahajan 0001 |
ISPASS | 2 |
| 2025 | Learning to Summarize from LLM-generated FeedbackabstractHwanjun Song, Taewon Yun, Yuho Lee, Jihwan Oh, Gihun Lee, Jason Cai, Hang Su. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hwanjun Song, Taewon Yun, Yuho Lee, Jihwan Oh, Gihun Lee, Jason Cai |
NAACL (Long Papers) | 4 |
| 2024 | Preference Alignment with Flow MatchingabstractWe present Preference Flow Matching (PFM), a new framework for preference alignment that streamlines the integration of preferences into an arbitrary class of pre-trained models. Existing alignment methods require fine-tuning pre-trained models, which presents challenges such as scalability, inefficiency, and the need for model modifications, especially with black-box APIs like GPT-4. In contrast, PFM utilizes flow matching techniques to directly learn from preference data, thereby reducing the dependency on extensive fine-tuning of pre-trained models. By leveraging flow-based models, PFM transforms less preferred data into preferred outcomes, and effectively aligns model outputs with human preferences without relying on explicit or implicit reward function estimation, thus avoiding common issues like overfitting in reward models. We provide theoretical insights that support our method’s alignment with standard preference alignment objectives. Experimental results indicate the practical effectiveness of our method, offering a new direction in aligning a pre-trained model to preference. Our code is available at https://github.com/jadehaus/preference-flow-matching. Minu Kim 0001, Yongsik Lee, Sehyeok Kang, Jihwan Oh, Song Chong, Se-Young Yun |
NeurIPS | 4 |
| 2020 | GRIA: Graphical Regularization for Integrative AnalysisabstractIntegrative analysis jointly analyzes multiple data sets to overcome curse of dimensionality. It can detect important but weak signals by jointly selecting features for all data sets, but unfortunately the sets of important features are not always the same for all data sets. Variations which allows heterogeneous sparsity structure-a subset of data sets can have a zero coefficient for a selected feature-have been proposed, but it compromises the effect of integrative analysis recalling the problem of losing weak important signals. We propose a new integrative analysis approach which not only aggregates weak important signals well in homogeneity setting but also substantially alleviates the problem of losing weak important signals in heterogeneity setting. Our approach exploits a priori known graphical structure of features by forcing joint selection of adjacent features, and integrating such information over multiple data sets can increase the power while taking into account the heterogeneity across data sets. We confirm the problem of existing approaches and demonstrate the superiority of our method through a simulation study and an application to gene expression data from ADNI. Changgee Chang, Jihwan Oh, Qi Long |
SDM | 2 |