Eunseong Choi

dblp:291/4794 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0003-1400-5227ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 GRAM: Generative Recommendation via Semantic-aware Multi-granular Late Fusion
abstract
Generative recommendation is an emerging paradigm that leverages the extensive knowledge of large language models by formulating recommendations into a text-to-text generation task.However, existing studies face two key limitations in (i) incorporating implicit item relationships and (ii) utilizing rich yet lengthy item information.To address these challenges, we propose a Generative Recommender via semantic-Aware Multi-granular late fusion (GRAM), introducing two synergistic innovations.First, we design semantic-to-lexical translation to encode implicit hierarchical and collaborative item relationships into the vocabulary space of LLMs.Second, we present multi-granular late fusion to integrate rich semantics efficiently with minimal information loss.It employs separate encoders for multigranular prompts, delaying the fusion until the decoding stage.Experiments on four benchmark datasets show that GRAM outperforms eight state-of-the-art generative recommendation models, achieving significant improvements of 11.5-16.0% in Recall@5 and 5.3-13.6% in NDCG@5.
Sunkyung Lee 0001, Minjin Choi 0001, Eunseong Choi, Hye-young Kim, Jongwuk Lee
ACL (1)3
2025 Conflict-Aware Soft Prompting for Retrieval-Augmented Generation
abstract
Retrieval-augmented generation (RAG) enhances the capabilities of large language models (LLMs) by incorporating external knowledge into their input prompts.However, when the retrieved context contradicts the LLM's parametric knowledge, it often fails to resolve the conflict between incorrect external context and correct parametric knowledge, known as context-memory conflict.To tackle this problem, we introduce Conflict-Aware REtrieval-Augmented Generation (CARE), consisting of a context assessor and a base LLM.The context assessor encodes external context into compact memory embeddings.Through grounded/adversarial soft prompting, the context assessor is trained to discern unreliable context and capture a guidance signal that directs reasoning toward the more reliable knowledge source.Extensive experiments show that CARE effectively mitigates context-memory conflicts, leading to an average performance gain of 5.0% on QA and fact-checking benchmarks, establishing a promising direction for trustworthy and adaptive RAG systems 1 .
Eunseong Choi, June Park, Hyeri Lee, Jongwuk Lee
EMNLP1
2025 Multi-view-guided Passage Reranking with Large Language Models
abstract
Recent advances in large language models (LLMs) have shown impressive performance in passage reranking tasks.Despite their success, LLM-based methods still face challenges in efficiency and sensitivity to external biases.(i) Existing models rely mostly on autoregressive generation and sliding window strategies to rank passages, which incurs heavy computational overhead as the number of passages increases.(ii) External biases, such as position or selection bias, hinder the model's ability to accurately represent passages and the inputorder sensitivity.To address these limitations, we introduce a novel passage reranking model, called Multi-View-guided Passage Reranking (MVP).MVP is a non-generative LLM-based reranking method that encodes query-passage information into diverse view embeddings without being influenced by external biases.For each view, it combines query-aware passage embeddings to produce a distinct anchor vector, used to directly compute relevance scores in a single decoding step.Besides, it employs an orthogonal loss to make the views more distinctive.Extensive experiments demonstrate that MVP, with just 220M parameters, matches the performance of much larger 7B-scale finetuned models while achieving a 100× reduction in inference latency.Notably, the 3B-parameter variant of MVP achieves state-of-the-art performance on both in-domain and out-of-domain benchmarks.The source code is available at https://github.com/bulbna/MVP.
Jeongwoo Na, Jun Kwon, Eunseong Choi, Jongwuk Lee
EMNLP3
2023 Forgetting-aware Linear Bias for Attentive Knowledge Tracing
abstract
Knowledge Tracing (KT) aims to track proficiency based on a question-solving history, allowing us to offer a streamlined curriculum. Recent studies actively utilize attention-based mechanisms to capture the correlation between questions and combine it with the learner's characteristics for responses. However, our empirical study shows that existing attention-based KT models neglect the learner's forgetting behavior, especially as the interaction history becomes longer. This problem arises from the bias that overprioritizes the correlation of questions while inadvertently ignoring the impact of forgetting behavior. This paper proposes a simple-yet-effective solution, namely Forgetting-aware Linear Bias (FoLiBi), to reflect forgetting behavior as a linear bias. Despite its simplicity, FoLiBi is readily equipped with existing attentive KT models by effectively decomposing question correlations with forgetting behavior. FoLiBi plugged with several KT models yields a consistent improvement of up to 2.58% in AUC over state-of-the-art KT models on four benchmark datasets.
Yoonjin Im, Eunseong Choi, Heejin Kook, Jongwuk Lee
CIKM2
2023 ConQueR: Contextualized Query Reduction using Search Logs
abstract
Query reformulation is a key mechanism to alleviate the linguistic chasm of query in ad-hoc retrieval. Among various solutions, query reduction effectively removes extraneous terms and specifies concise user intent from long queries. However, it is challenging to capture hidden and diverse user intent. This paper proposes Contextualized Query Reduction (ConQueR) using a pre-trained language model (PLM). Specifically, it reduces verbose queries with two different views: core term extraction and sub-query selection. One extracts core terms from an original query at the term level, and the other determines whether a sub-query is a suitable reduction for the original query at the sequence level. Since they operate at different levels of granularity and complement each other, they are finally aggregated in an ensemble manner. We evaluate the reduction quality of ConQueR on real-world search logs collected from a commercial web search engine. It achieves up to 8.45% gains in exact match scores over the best competing model.
Hye-young Kim, Minjin Choi 0001, Sunkyung Lee 0001, Eunseong Choi, Young-In Song, Jongwuk Lee
SIGIR4
2022 SpaDE: Improving Sparse Representations using a Dual Document Encoder for First-stage Retrieval
abstract
Sparse document representations have been widely used to retrieve relevant documents via exact lexical matching. Owing to the pre-computed inverted index, it supports fast ad-hoc search but incurs the vocabulary mismatch problem. Although recent neural ranking models using pre-trained language models can address this problem, they usually require expensive query inference costs, implying the trade-off between effectiveness and efficiency. Tackling the trade-off, we propose a novel uni-encoder ranking model, Sparse retriever using a Dual document Encoder (SpaDE), learning document representation via the dual encoder. Each encoder plays a central role in (i) adjusting the importance of terms to improve lexical matching and (ii) expanding additional terms to support semantic matching. Furthermore, our co-training strategy trains the dual encoder effectively and avoids unnecessary intervention in training each other. Experimental results on several benchmarks show that SpaDE outperforms existing uni-encoder ranking models.
Eunseong Choi, Sunkyung Lee 0001, Minjin Choi 0001, Hyeseon Ko, Young-In Song, Jongwuk Lee
CIKM1
2022 Long-tail Mixup for Extreme Multi-label Classification
abstract
Extreme multi-label classification (XMC) aims at finding multiple relevant labels for a given sample from a huge label set at the industrial scale. The XMC problem inherently poses two challenges: scalability and label sparsity - the number of labels is too large, and labels follow the long-tail distribution. To resolve these problems, we propose a novel Mixup-based augmentation method for long-tail labels, called TailMix. Building upon the partition-based model, TailMix utilizes the context vectors generated from the label attention layer. It first selectively chooses two context vectors using the inverse propensity score of labels and the label proximity graph representing the co-occurrence of labels. Using two context vectors, it augments new samples with the long-tail label to improve the accuracy of long-tail labels. Despite its simplicity, experimental results show that TailMix consistently outperforms other augmentation methods on three benchmark datasets, especially for long-tail labels in terms of two metrics, [email protected] and [email protected]
Sangwoo Han, Eunseong Choi, Chan Lim, Hyunjung Shim, Jongwuk Lee
CIKM2
2021 MelBERT: Metaphor Detection via Contextualized Late Interaction using Metaphorical Identification Theories
abstract
Minjin Choi, Sunkyung Lee, Eunseong Choi, Heesoo Park, Junhyuk Lee, Dongwon Lee, Jongwuk Lee. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Minjin Choi 0001, Sunkyung Lee 0001, Eunseong Choi, Heesoo Park, Junhyuk Lee, Dongwon Lee 0001, Jongwuk Lee
NAACL-HLT3