Jaehun Jung

dblp:192/7707 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Low-light image enhancement via distribution of latent transitions
Jaehun Jung
J. Vis. Commun. Image Represent.1
2026 Data Adaptive Stochastic Ensemble Net: Optimizing Infection Predictions for COVID-19 Cluster Analysis
abstract
Machine learning has garnered significant interest and is extensively utilized in the medical field due to its direct impact on human life. Two components are necessary to develop an AI-based infection prediction assistance system: a training dataset and machine learning prediction model. For AI-based infection prediction model, we first gathered a real-world COVID-19 cluster dataset, consisting of 8,844 confirmed cases across 519 clusters, which includes individual properties and contact relationships between confirmed cases. Second, we introduce the Data Adaptive Stochastic Ensemble Network (DASEN) to enhance prediction robustness. DASEN dynamically adjusts the weight of each component by optimizing the Dirichlet distribution concentration parameter based on the data distribution. We demonstrate the validity of DASEN, showing that different models focus on distinct features and perform well on data with varying characteristics, thus preventing overfitting to majority labels. Notably, DASEN provides superior robustness across all settings with minimal overhead for parameter optimization.
Sungjun Lim 0002, YongTaek Lim, Hojun Park, Junggu Lee, Jaehun Jung, Kyungwoo Song
IEEE J. Biomed. Health Informatics5
2025 Causal Effect Variational Transformer for Public Health Measures and COVID-19 Infection Cluster Analysis
abstract
Recent research increasingly integrates causal inference into deep learning models to enhance the explainability and robustness of medical applications. However, data scarcity remains a fundamental challenge due to privacy constraints and the high cost of data collection. This issue, compounded by complex variable dependencies and unobserved latent confounders, hinders the reliable estimation of causal effects. To address these challenges, we collect two real-world COVID-19 infection cluster datasets, including public health measures, from distinct distributions in collaboration with local governments, a medical university, and a hospital. We also propose a cut-off augmentation method that generates diverse feature-label pairs by slicing time-series sequences at different observation windows, effectively simulating partial observations common in real-world settings. We further introduce the Causal Effect Variational Transformer (CEVT), a Transformer-based model that captures temporal structure and addresses the difficulty of causal estimation under scarce data, complex dependencies, and latent confounding by modeling multiple treatments through an iterative conditioning mechanism. We validate the causal modeling capability of CEVT on synthetic datasets and demonstrate that, on two distinct COVID-19 datasets, it consistently outperforms baselines in infection prediction. Notably, the causal effects estimated by CEVT converge with findings from medical studies on infection control, reinforcing its reliability and underscoring its potential to inform public health decision-making.
Jinho Kang, Sungjun Lim 0002, Hojun Park, Jiyoung Jung, Jaehun Jung, Kyungwoo Song
CIKM5
2025 Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions
abstract
Recent research in vision-language models (VLMs) has centered around the possibility of equipping them with implicit long-form chainof-thought reasoning-akin to the success observed in language models-via distillation and reinforcement learning.But what about the non-reasoning models already trained and deployed across the internet?Should we simply abandon them, or is there hope for a search mechanism that can elicit hidden knowledge and induce long reasoning traces-without any additional training or supervision?In this paper, we explore this possibility using a Monte Carlo Tree Search (MCTS)-inspired algorithm, which injects subquestion-subanswer pairs into the model's output stream.We show that framing reasoning as a search process-where subquestions act as latent decisions within a broader inference trajectory-helps the model "connect the dots" between fragmented knowledge and produce extended reasoning traces in nonreasoning models.We evaluate our method across three benchmarks and observe consistent improvements.Notably, our approach yields a 2% overall improvement on MMMU-PRO, including a significant 9% gain in Liberal Arts.Problem (P) What is the position of the blue car in the image?A: parked on the sidewalk B: driving on the road C: parked on the grass D: parked on the road CoT: To determine the position of the blue car in the image, let's analyze the scene: * The image shows a busy urban street with several vehicles.* The blue car is located on the right side of the image.* It is positioned on what appears to be a paved surface, which is consistent with a road.* The car is not on the grass or the sidewalk, as those areas are clearly distinguishable in the image.* The car is stationary, suggesting it is parked.Given these observations, the blue car is parked on the road.
David Acuna, Ximing Lu, Jaehun Jung, Hyunwoo Kim 0002, Amlan Kar, Sanja Fidler, Yejin Choi 0001
EMNLP3
2025 Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
abstract
We present a principled approach to provide LLM-based evaluation with a rigorous guarantee of human agreement. We first propose that a reliable evaluation method should not uncritically rely on model preferences for pairwise evaluation, but rather assess the confidence of judge models and selectively decide when to trust its judgement. We then show that under this *selective evaluation* framework, human agreement can be provably guaranteed---such that the model evaluation aligns with that of humans to a user-specified agreement level. As part of our framework, we also introduce *Simulated Annotators*, a novel confidence estimation method that significantly improves judge calibration and thus enables high coverage of evaluated instances. Finally, we propose *Cascaded Selective Evaluation*, where we use cheaper models as initial judges and escalate to stronger models only when necessary---again, while still providing a provable guarantee of human agreement. Experimental results show that Cascaded Selective Evaluation guarantees strong alignment with humans, far beyond what LLM judges could achieve without selective evaluation. For example, on a subset of Chatbot Arena where GPT-4 almost never achieves 80% human agreement, our method, even while employing substantially cost-effective models such as Mistral-7B, *guarantees* over 80% human agreement with almost 80% test coverage.
Jaehun Jung, Faeze Brahman, Yejin Choi 0001
ICLR1
2025 Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning
abstract
Data diversity is crucial for training a strong language model. Yet metrics of diversity often diverge from this goal, measuring variations in heuristic features—like n-grams or embeddings—that are detached from how the model actually performs on a target task. This motivates us to ask: *Can we redefine data diversity—beyond measuring variations in heuristic features—in a way that better predicts model generalization?* Through large-scale empirical analyses spanning over 300 training runs, carefully controlled for data scale and quality, we show that data diversity can be a strong predictor of generalization in LLM reasoning—as measured by average model performance on unseen out-of-distribution benchmarks. We introduce **G-Vendi**, a metric that quantifies diversity via the entropy of model-induced loss gradients. G-Vendi scales to million-sample datasets and yet consistently outperforms heuristic alternatives, achieving strong correlation ($\text{Spearman's } \rho \approx 0.9$) with out-of-distribution (OOD) performance across both natural language inference (NLI) and math reasoning tasks. Building on this insight, we present **Prismatic Synthesis**, a framework for generating diverse synthetic data by targeting underrepresented regions in gradient space. Experimental results show that Prismatic Synthesis consistently improves model performance as we scale synthetic data—not just on in-distribution test but across unseen, out-of-distribution benchmarks—significantly outperforming state-of-the-art models in both domains. For example, PrismMath-7B, our model distilled from a 32B LLM without human verification, outperforms R1-Distill-Qwen-7B—trained on proprietary data generated by 671B R1—on 6 out of 7 challenging math benchmarks.
Jaehun Jung, Seungju Han 0002, Ximing Lu, Skyler Hallinan, David Acuna, Shrimai Prabhumoye, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Yejin Choi 0001
NeurIPS1
2025 COVID-19 prediction with doubly multi-task Gaussian Process
Sooyon Kim, Yongtaek Lim, Sungjun Lim 0002, Gyeongdeok Seo, Hojun Park, Jaehun Jung, Kyungwoo Song
J. Biomed. Informatics7
2024 JAMDEC: Unsupervised Authorship Obfuscation using Constrained Decoding over Small Language Models
abstract
Jillian Fisher, Ximing Lu, Jaehun Jung, Liwei Jiang, Zaid Harchaoui, Yejin Choi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jillian Fisher, Ximing Lu, Jaehun Jung, Zaïd Harchaoui, Yejin Choi 0001
NAACL-HLT3
2024 Impossible Distillation for Paraphrasing and Summarization: How to Make High-quality Lemonade out of Small, Low-quality Model
abstract
Jaehun Jung, Peter West, Liwei Jiang, Faeze Brahman, Ximing Lu, Jillian Fisher, Taylor Sorensen, Yejin Choi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jaehun Jung, Peter West, Faeze Brahman, Ximing Lu, Jillian Fisher, Taylor Sorensen, Yejin Choi 0001
NAACL-HLT1
2023 DataHalo: A Customizable Notification Visualization System for Personalized and Longitudinal Interactions
abstract
People struggle with the overflow of smartphone notifications but often face two challenges: (1) prioritizing the informative notifications as they wish and (2) retaining the delivered information as long as they want to utilize it. In this paper, we present DataHalo, a customizable notification visualization system that represents notifications as prolonged ambient visualizations on the home screen. DataHalo supports keyword-based filtering and categorization, and draws graphical marks based on time-varying importance model to enable longitudinal interaction with the notifications. We evaluated DataHalo through a usability study (N = 17), from which we improved the interface. We then conducted a three-week deployment study (N = 12) to assess how people use DataHalo in their domestic contexts. Our study revealed that people generated various visualization settings for different kinds of apps. Drawing on both quantitative and qualitative findings, we discussed implications for supporting effective notification management through customizable ambient visualizations.
GuHyun Han, Jaehun Jung, Young-Ho Kim, Jinwook Seo
CHI2
2023 Inference-Time Policy Adapters (IPA): Tailoring Extreme-Scale LMs without Fine-tuning
abstract
Ximing Lu, Faeze Brahman, Peter West, Jaehun Jung, Khyathi Chandu, Abhilasha Ravichander, Prithviraj Ammanabrolu, Liwei Jiang, Sahana Ramnath, Nouha Dziri, Jillian Fisher, Bill Lin, Skyler Hallinan, Lianhui Qin, Xiang Ren, Sean Welleck, Yejin Choi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Ximing Lu, Faeze Brahman, Peter West, Jaehun Jung, Khyathi Raghavi Chandu, Abhilasha Ravichander, Prithviraj Ammanabrolu, Sahana Ramnath, Nouha Dziri, Jillian Fisher, Bill Y. Lin, Skyler Hallinan, Lianhui Qin, Xiang Ren 0001, Sean Welleck, Yejin Choi 0001
EMNLP4
2022 Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations
abstract
Pre-trained language models (LMs) struggle with consistent reasoning; recently, prompting LMs to generate explanations that self-guide the inference has emerged as a promising direction to amend this.However, these approaches are fundamentally bounded by the correctness of explanations, which themselves are often noisy and inconsistent.In this work, we develop MAIEUTIC PROMPTING, which aims to infer a correct answer to a question even from the unreliable generations of LM.MAIEUTIC PROMPTING induces a tree of explanations abductively (e.g.X is true, because . . . ) and recursively, then frames the inference as a satisfiability problem over these explanations and their logical relations.We test MAIEUTIC PROMPTING for true/false QA on three challenging benchmarks that require complex commonsense reasoning.MAIEU-TIC PROMPTING achieves up to 20% better accuracy than state-of-the-art prompting methods, and as a fully unsupervised approach, performs competitively with supervised models.We also show that MAIEUTIC PROMPTING improves robustness in inference while providing interpretable rationales. 1
Jaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman, Chandra Bhagavatula, Ronan Le Bras 0001, Yejin Choi 0001
EMNLP1
2021 Learning to Walk across Time for Interpretable Temporal Knowledge Graph Completion
abstract
Static knowledge graphs (KGs), despite their wide usage in relational reasoning and downstream tasks, fall short of realistic modeling of knowledge and facts that are only temporarily valid. Compared to static knowledge graphs, temporal knowledge graphs (TKGs) inherently reflect the transient nature of real-world knowledge. Naturally, automatic TKG completion has drawn much research interests for a more realistic modeling of relational reasoning. However, most of the existing models for TKG completion extend static KG embeddings that do not fully exploit TKG structure, thus lacking in 1) accounting for temporally relevant events already residing in the local neighborhood of a query, and 2) path-based inference that facilitates multi-hop reasoning and better interpretability. In this paper, we propose T-GAP, a novel model for TKG completion that maximally utilizes both temporal information and graph structure in its encoder and decoder. T-GAP encodes query-specific substructure of TKG by focusing on the temporal displacement between each event and the query timestamp, and performs path-based inference by propagating attention through the graph. Our empirical experiments demonstrate that T-GAP not only achieves superior performance against state-of-the-art baselines, but also competently generalizes to queries with unseen timestamps. Through extensive qualitative analyses, we also show that T-GAP enjoys transparent interpretability, and follows human intuition in its reasoning process.
Jaehun Jung, Jinhong Jung, U Kang
KDD1
2020 AttnIO: Knowledge Graph Exploration with In-and-Out Attention Flow for Knowledge-Grounded Dialogue
abstract
Retrieving the proper knowledge relevant to conversational context is an important challenge in dialogue systems, to engage users with more informative response.Several recent works propose to formulate this knowledge selection problem as a path traversal over an external knowledge graph (KG), but show only a limited utilization of KG structure, leaving rooms of improvement in performance.To this effect, we present AttnIO, a new dialog-conditioned path traversal model that makes a full use of rich structural information in KG based on two directions of attention flows.Through the attention flows, At-tnIO is not only capable of exploring a broad range of multi-hop knowledge paths, but also learns to flexibly adjust the varying range of plausible nodes and edges to attend depending on the dialog context.Empirical evaluations present a marked performance improvement of AttnIO compared to all baselines in OpenDi-alKG dataset.Also, we find that our model can be trained to generate an adequate knowledge path even when the paths are not available and only the destination nodes are given as label, making it more applicable to real-world dialogue systems.
Jaehun Jung, Bokyung Son, Sungwon Lyu
EMNLP (1)1