VLDB 2026 Research / reviewers in the wild / expert
Sohee Yang
dblp:236/5847
· DBLP profile ↗
15ranked-venue papers
3as first author
12since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reasoning Models Better Express Their ConfidenceabstractDespite their strengths, large language models (LLMs) often fail to communicate their confidence accurately, making it difficult to assess when they might be wrong and limiting their reliability. In this work, we demonstrate that reasoning models that engage in extended chain-of-thought (CoT) reasoning exhibit superior performance not only in problem-solving but also in accurately expressing their confidence.
Specifically, we benchmark six reasoning models across six datasets and find that they achieve strictly better confidence calibration than their non-reasoning counterparts in 33 out of the 36 settings. Our detailed analysis reveals that these gains in calibration stem from the slow thinking behaviors of reasoning models (e.g., exploring alternative approaches and backtracking) which enable them to adjust their confidence dynamically throughout their CoT, making it progressively more accurate. In particular, we find that reasoning models become increasingly better calibrated as their CoT unfolds, a trend not observed in non-reasoning models. Moreover, removing slow thinking behaviors from the CoT leads to a significant drop in calibration. Lastly, we show that non-reasoning models also demonstrate enhanced calibration when simply guided to slow think via in-context learning, fully isolating slow thinking as the source of the calibration gains. Dongkeun Yoon, Seungone Kim, Sohee Yang, Sunkyoung Kim 0002, Yongil Kim, Eunbi Choi, Yireun Kim, Minjoon Seo |
NeurIPS | 3 |
| 2024 | Investigating the Effectiveness of Task-Agnostic Prefix Prompt for Instruction FollowingabstractIn this paper, we present our finding that prepending a Task-Agnostic Prefix Prompt (TAPP) to the input improves the instruction-following ability of various Large Language Models (LLMs) during inference. TAPP is different from canonical prompts for LLMs in that it is a fixed prompt prepended to the beginning of every input regardless of the target task for zero-shot generalization. We observe that both base LLMs (i.e. not fine-tuned to follow instructions) and instruction-tuned models benefit from TAPP, resulting in 34.58% and 12.26% improvement on average, respectively. This implies that the instruction-following ability of LLMs can be improved during inference time with a fixed prompt constructed with simple heuristics. We hypothesize that TAPP assists language models to better estimate the output distribution by focusing more on the instruction of the target task during inference. In other words, such ability does not seem to be sufficiently activated in not only base LLMs but also many instruction-fine-tuned LLMs. Seonghyeon Ye, Hyeonbin Hwang, Sohee Yang, Hyeongu Yun, Yireun Kim, Minjoon Seo |
AAAI | 3 |
| 2024 | Do Large Language Models Latently Perform Multi-Hop Reasoning?abstractWe study whether Large Language Models (LLMs) latently perform multi-hop reasoning with complex prompts such as "The mother of the singer of 'Superstition' is".We look for evidence of a latent reasoning pathway where an LLM (1) latently identifies "the singer of 'Superstition"' as Stevie Wonder, the bridge entity, and (2) uses its knowledge of Stevie Wonder's mother to complete the prompt.We analyze these two hops individually and consider their co-occurrence as indicative of latent multi-hop reasoning.For the first hop, we test if changing the prompt to indirectly mention the bridge entity instead of any other entity increases the LLM's internal recall of the bridge entity.For the second hop, we test if increasing this recall causes the LLM to better utilize what it knows about the bridge entity.We find strong evidence of latent multi-hop reasoning for the prompts of certain relation types, with the reasoning pathway used in more than 80% of the prompts.However, the utilization is highly contextual, varying across different types of prompts.Also, on average, the evidence for the second hop and the full multi-hop traversal is rather moderate and only substantial for the first hop.Moreover, we find a clear scaling trend with increasing model size for the first hop of reasoning but not for the second hop.Our experimental findings suggest potential challenges and opportunities for future development and applications of LLMs. Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, Sebastian Riedel 0001 |
ACL (1) | 1 |
| 2024 | Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop QueriesabstractLarge language models (LLMs) can solve complex multi-step problems, but little is known about how these computations are implemented internally.Motivated by this, we study how LLMs answer multi-hop queries such as "The spouse of the performer of Imagine is".These queries require two information extraction steps: a latent one for resolving the first hop ("the performer of Imagine") into the bridge entity (John Lennon), and another for resolving the second hop ("the spouse of John Lennon") into the target entity (Yoko Ono).Understanding how the latent step is computed internally is key to understanding the overall computation.By carefully analyzing the internal computations of transformer-based LLMs, we discover that the bridge entity is resolved in the early layers of the model.Then, only after this resolution, the two-hop query is solved in the later layers.Because the second hop commences in later layers, there could be cases where these layers no longer encode the necessary knowledge for correctly predicting the answer.Motivated by this, we propose a novel "back-patching" analysis method whereby a hidden representation from a later layer is patched back to an earlier layer.We find that in up to 66% of previously incorrect cases there exists a back-patch that results in the correct generation of the answer, showing that the later layers indeed sometimes lack the needed functionality.Overall, our methods and findings open further opportunities for understanding and improving latent reasoning in transformer-based LLMs. Eden Biran, Daniela Gottesman, Sohee Yang, Mor Geva, Amir Globerson |
EMNLP | 3 |
| 2024 | Exploring the Practicality of Generative Retrieval on Dynamic CorporaabstractBenchmarking the performance of information retrieval (IR) is mostly conducted with a fixed set of documents (static corpora).However, in realistic scenarios, this is rarely the case and the documents to be retrieved are constantly updated and added.In this paper, we focus on Generative Retrievals (GR), which apply autoregressive language models to IR problems, and explore their adaptability and robustness in dynamic scenarios.We also conduct an extensive evaluation of computational and memory efficiency, crucial factors for real-world deployment of IR systems handling vast and ever-changing document collections.Our results on the StreamingQA benchmark demonstrate that GR is more adaptable to evolving knowledge (4 -11%), robust in learning knowledge with temporal information, and efficient in terms of inference FLOPs (ˆ2), indexing time (ˆ6), and storage footprint (ˆ4) compared to Dual Encoders (DE), which are commonly used in retrieval systems.Our paper highlights the potential of GR for future use in practical IR systems within dynamic environments. Chaeeun Kim, Soyoung Yoon, Hyunji Lee, Joel Jang, Sohee Yang, Minjoon Seo |
EMNLP | 5 |
| 2024 | How Do Large Language Models Acquire Factual Knowledge During Pretraining?abstractDespite the recent observation that large language models (LLMs) can store substantial factual knowledge, there is a limited understanding of the mechanisms of how they acquire factual knowledge through pretraining. This work addresses this gap by studying how LLMs acquire factual knowledge during pretraining. The findings reveal several important insights into the dynamics of factual knowledge acquisition during pretraining. First, counterintuitively, we observe that pretraining on more data shows no significant improvement in the model's capability to acquire and maintain factual knowledge. Next, LLMs undergo forgetting of memorization and generalization of factual knowledge, and LLMs trained with duplicated training data exhibit faster forgetting. Third, training LLMs with larger batch sizes can enhance the models' robustness to forgetting. Overall, our observations suggest that factual knowledge acquisition in LLM pretraining occurs by progressively increasing the probability of factual knowledge presented in the pretraining data at each step. However, this increase is diluted by subsequent forgetting. Based on this interpretation, we demonstrate that we can provide plausible explanations on recently observed behaviors of LLMs, such as the poor performance of LLMs on long-tail knowledge and the benefits of deduplicating the pretraining corpus. Hoyeon Chang, Jinho Park 0005, Seonghyeon Ye, Sohee Yang, Youngkyung Seo, Du-Seong Chang, Minjoon Seo |
NeurIPS | 4 |
| 2024 | Improving Probability-based Prompt Selection Through Unified Evaluation and AnalysisabstractAbstract Previous work in prompt engineering for large language models has introduced different gradient-free probability-based prompt selection methods that aim to choose the optimal prompt among the candidates for a given task but have failed to provide a comprehensive and fair comparison between each other. In this paper, we propose a unified framework to interpret and evaluate the existing probability-based prompt selection methods by performing extensive experiments on 13 common and diverse NLP tasks. We find that each of the existing methods can be interpreted as some variant of the method that maximizes mutual information between the input and the predicted output (MI). Utilizing this finding, we develop several other combinatorial variants of MI and increase the effectiveness of the oracle prompt selection method from 87.79% to 94.98%, measured as the ratio of the performance of the selected prompt to that of the optimal oracle prompt. Furthermore, considering that all the methods rely on the output probability distribution of the model that might be biased, we propose a novel calibration method called Calibration by Marginalization (CBM) that is orthogonal to the existing methods and helps increase the prompt selection effectiveness of the best method to 96.85%, achieving 99.44% of the oracle prompt F1 without calibration.1 Sohee Yang, Jonghyeon Kim, Joel Jang, Seonghyeon Ye, Hyunji Lee, Minjoon Seo |
Trans. Assoc. Comput. Linguistics | 1 |
| 2023 | Knowledge Unlearning for Mitigating Privacy Risks in Language ModelsabstractJoel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, Minjoon Seo. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, Minjoon Seo |
ACL (1) | 3 |
| 2022 | TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language ModelsabstractLanguage Models (LMs) become outdated as the world changes; they often fail to perform tasks requiring recent factual information which was absent or different during training, a phenomenon called temporal misalignment.This is especially a challenging problem because the research community still lacks a coherent dataset for assessing the adaptability of LMs to frequently-updated knowledge corpus such as Wikipedia.To this end, we introduce TEMPORALWIKI, a lifelong benchmark for ever-evolving LMs that utilizes the difference between consecutive snapshots of English Wikipedia and English Wikidata for training and evaluation, respectively.The benchmark hence allows researchers to periodically track an LM's ability to retain previous knowledge and acquire updated/new knowledge at each point in time.We also find that training an LM on the diff data through continual learning methods achieves similar or better perplexity than on the entire snapshot in our benchmark with 12 times less computational cost, which verifies that factual knowledge in LMs can be safely updated with minimal training data via continual learning.The dataset and the code is made available at this link. What is the most dominant COVID-19 variant?What is the most dominant COVID-19 variant?Difference Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Minjoon Seo |
EMNLP | 4 |
| 2022 | Generative Multi-hop RetrievalabstractA common practice for text retrieval is to use an encoder to map the documents and the query to a common vector space and perform a nearest neighbor search (NNS); multi-hop retrieval also often adopts the same paradigm, usually with a modification of iteratively reformulating the query vector so that it can retrieve different documents at each hop.However, such a biencoder approach has limitations in multi-hop settings; (1) the reformulated query gets longer as the number of hops increases, which further tightens the embedding bottleneck of the query vector, and (2) it is prone to error propagation.In this paper, we focus on alleviating these limitations in multi-hop settings by formulating the problem in a fully generative way.We propose an encoder-decoder model that performs multi-hop retrieval by simply generating the entire text sequences of the retrieval targets, which means the query and the documents interact in the language model's parametric space rather than L2 or inner product space as in the bi-encoder approach.Our approach, Generative Multi-hop Retrieval (GMR), consistently achieves comparable or higher performance than bi-encoder models in five datasets while demonstrating superior GPU memory and storage footprint.1 Hyunji Lee, Sohee Yang, Hanseok Oh, Minjoon Seo |
EMNLP | 2 |
| 2022 | Towards Continual Knowledge Learning of Language Models
Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, Minjoon Seo |
ICLR | 3 |
| 2021 | Designing a Minimal Retrieve-and-Read System for Open-Domain Question AnsweringabstractIn open-domain question answering (QA), retrieve-and-read mechanism has the inherent benefit of interpretability and the easiness of adding, removing, or editing knowledge compared to the parametric approaches of closedbook QA models.However, it is also known to suffer from its large storage footprint due to its document corpus and index.Here, we discuss several orthogonal strategies to drastically reduce the footprint of a retrieve-andread open-domain QA system by up to 160x.Our results indicate that retrieve-and-read can be a viable option even in a highly constrained serving environment such as edge devices, as we show that it can achieve better accuracy than a purely parametric model with comparable docker-level system size. 1 Sohee Yang, Minjoon Seo |
NAACL-HLT | 1 |
| 2020 | Efficient Dialogue State Tracking by Selectively Overwriting MemoryabstractRecent works in dialogue state tracking (DST) focus on an open vocabulary-based setting to resolve scalability and generalization issues of the predefined ontology-based approaches.However, they are inefficient in that they predict the dialogue state at every turn from scratch.Here, we consider dialogue state as an explicit fixed-sized memory and propose a selectively overwriting mechanism for more efficient DST.This mechanism consists of two steps: (1) predicting state operation on each of the memory slots, and (2) overwriting the memory with new values, of which only a few are generated according to the predicted state operations.Our method decomposes DST into two sub-tasks and guides the decoder to focus only on one of the tasks, thus reducing the burden of the decoder.This enhances the effectiveness of training and DST performance.Our SOM-DST (Selectively Overwriting Memory for Dialogue State Tracking) model achieves state-of-theart joint goal accuracy with 51.72% in Mul-tiWOZ 2.0 and 53.01% in MultiWOZ 2.1 in an open vocabulary-based DST setting.In addition, we analyze the accuracy gaps between the current and the ground truth-given situations and suggest that it is a promising direction to improve state operation prediction to boost the DST performance. 1 Sungdong Kim, Sohee Yang, Gyuwan Kim, Sang-Woo Lee 0001 |
ACL | 2 |
| 2020 | ClovaCall: Korean Goal-Oriented Dialog Speech Corpus for Automatic Speech Recognition of Contact CentersabstractAutomatic speech recognition (ASR) via call is essential for various applications, including AI for contact center (AICC) services. Despite the advancement of ASR, however, most publicly available call-based speech corpora such as Switchboard are old-fashioned. Also, most existing call corpora are in English and mainly focus on open domain dialog or general scenarios such as audiobooks. Here we introduce a new large-scale Korean call-based speech corpus under a goal-oriented dialog scenario from more than 11,000 people, i.e., ClovaCall corpus. ClovaCall includes approximately 60,000 pairs of a short sentence and its corresponding spoken utterance in a restaurant reservation domain. We validate the effectiveness of our dataset with intensive experiments using two standard ASR models. Furthermore, we release our ClovaCall dataset and baseline source codes to be available via https://github.com/ClovaAI/ClovaCall. Copyright © 2020 ISCA Jung-Woo Ha 0001, Kihyun Nam, Sang-Woo Lee 0001, Sohee Yang, Hyunhoon Jung, Hyeji Kim, Eunmi Kim, Soojin Kim, Hyun Ah Kim, Kyoungtae Doh, Chan Kyu Lee, Nako Sung, Sunghun Kim 0001 |
INTERSPEECH | 5 |
| 2019 | Large-Scale Answerer in Questioner's Mind for Visual Dialog Question Generation
Sang-Woo Lee 0001, Sohee Yang, Jaejun Yoo 0001, Jung-Woo Ha 0001 |
ICLR (Poster) | 3 |