Kyungjae Lee 0002

dblp:13/7265-2 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0003-0586-3748ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
abstract
Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Choi 0001, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Y. Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee 0002, Minjoon Seo
NAACL (Long Papers)31
2024 Co-Creating Question-and-Answer Style Articles with Large Language Models for Research Promotion
abstract
Research promotion enables researchers to share advanced knowledge with pertinent academic communities. The question-and-answer (QA) style articles are effective for researchers to promote their research by enabling readers to understand research on complex subjects. Recent advances in large language models (LLMs) have opened avenues for supporting researchers in creating QA-style articles for research promotion. However, without the authors’ involvement, these models may only partially capture the researcher’s intention and voice. We developed AQUA, a research probe that enables researchers to co-create QA-style articles with LLMs to promote their research papers. A user study (n=12) reveals that LLMs reduced authors’ burden and helped them understand the readers’ perspectives. Nevertheless, LLMs failed to capture the unique intent of the authors, and their automated generation discouraged authors from carefully revising their answers. Based on our findings, we discuss human-LLM interaction design to enable authors to create QA-style articles that reflect their intention.
Hyunseung Lim, Ji Yong Cho, Taewan Kim 0004, Jeongeon Park, Hyungyu Shin, Seulgi Choi, Sunghyun Park 0005, Kyungjae Lee 0002, Juho Kim 0001, Moontae Lee, Hwajung Hong
Conference on Designing Interactive Systems8
2024 Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
abstract
Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Y. Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee 0002, Minjoon Seo
EMNLP9
2023 On Complementarity Objectives for Hybrid Retrieval
abstract
Dense retrieval has shown promising results in various information retrieval tasks, and hybrid retrieval, combined with the strength of sparse retrieval, has also been actively studied.A key challenge in hybrid retrieval is to make sparse and dense complementary to each other.Existing models have focused on dense models to capture "residual" features neglected in the sparse models.Our key distinction is to show how this notion of residual complementarity is limited, and propose a new objective, denoted as RoC (Ratio of Complementarity), which captures a fuller notion of complementarity.We propose a two-level orthogonality designed to improve RoC, then show that the improved RoC of our model, in turn, improves the performance of hybrid retrieval.Our method outperforms all state-of-the-art methods on three representative IR benchmarks: MSMARCO-Passage, Natural Questions, and TREC Ro-bust04, with statistical significance.Our finding is also consistent in various adversarial settings.
Dohyeon Lee, Seung-won Hwang, Kyungjae Lee 0002, Seungtaek Choi, Sunghyun Park 0005
ACL (1)3
2023 PreWoMe: Exploiting Presuppositions as Working Memory for Long Form Question Answering
abstract
Information-seeking questions in long-form question answering (LFQA) often prove misleading due to ambiguity or false presupposition in the question.While many existing approaches handle misleading questions, they are tailored to limited questions, which are insufficient in a real-world setting with unpredictable input characteristics.In this work, we propose PreWoMe, a unified approach capable of handling any type of information-seeking question.The key idea of PreWoMe involves extracting presuppositions in the question and exploiting them as working memory to generate feedback and action about the question.Our experiment shows that PreWoMe is effective not only in tackling misleading questions but also in handling normal ones, thereby demonstrating the effectiveness of leveraging presuppositions, feedback, and action for real-world QA settings.
Wookje Han, Jinsol Park, Kyungjae Lee 0002
EMNLP3
2023 Exploring the Benefits of Training Expert Language Models over Instruction Tuning
abstract
Recently, Language Models (LMs) instruction-tuned on multiple tasks, also known as multitask-prompted fine-tuning (MT), have shown capabilities to generalize to unseen tasks. Previous work has shown that scaling the number of finetuning datasets and instructions is the key component in making stronger MT LMs. In this work, we report surprising findings that show an expert LM trained on just a single task can outperform an MT LM trained with 300+ different tasks on 11 different unseen datasets and on 13 datasets of the BIG-bench benchmark by an average of 3.20% and 1.29%, respectively. This finding casts doubt on the previously held belief that simply scaling the number of tasks makes stronger MT LMs. Leveraging this finding, we further show that this distributed approach of training multiple expert LMs instead of a single MT LM for zero-shot inference possesses many benefits including (1) avoiding negative task transfer that often occurs during instruction tuning, (2) being able to continually learn new tasks without having to re-train on previous tasks to avoid catastrophic forgetting, and (3) showing compositional capabilities when merging individual experts together.
Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim 0001, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee 0002, Minjoon Seo
ICML7
2023 QASA: Advanced Question Answering on Scientific Articles
abstract
Reasoning is the crux of intellectual thinking. While question answering (QA) tasks are prolific with various computational models and benchmark datasets, they mostly tackle factoid or shallow QA without asking deeper understanding. Dual process theory asserts that human reasoning consists of associative thinking to collect relevant pieces of knowledge and logical reasoning to consciously conclude grounding on evidential rationale. Based on our intensive think-aloud study that revealed the three types of questions: surface, testing, and deep questions, we first propose the QASA benchmark that consists of 1798 novel question answering pairs that require full-stack reasoning on scientific articles in AI and ML fields. Then we propose the QASA approach that tackles the full-stack reasoning with large language models via associative selection, evidential rationale-generation, and systematic composition. Our experimental results show that QASA's full-stack inference outperforms the state-of-the-art InstructGPT by a big margin. We also find that rationale-generation is critical for the performance gain, claiming how we should rethink advanced question answering. The dataset is available at https://github.com/lgresearch/QASA.
Yoonjoo Lee, Kyungjae Lee 0002, Sunghyun Park 0005, Dasol Hwang, Jaehyeon Kim, Hong-In Lee, Moontae Lee
ICML2
2023 On Monotonic Aggregation for Open-domain QA
Sang-eun Han, Yeonseok Jeong, Seung-won Hwang, Kyungjae Lee 0002
INTERSPEECH4
2021 Robustifying Multi-hop QA through Pseudo-Evidentiality Training
abstract
Kyungjae Lee, Seung-won Hwang, Sang-eun Han, Dohyeon Lee. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Kyungjae Lee 0002, Seung-won Hwang, Sang-eun Han, Dohyeon Lee
ACL/IJCNLP (1)1
2021 Query Generation for Multimodal Documents
abstract
Kyungho Kim, Kyungjae Lee, Seung-won Hwang, Young-In Song, Seungwook Lee. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Kyungho Kim, Kyungjae Lee 0002, Seung-won Hwang, Young-In Song, Seungwook Lee
EACL2
2020 Segment-Then-Rank: Non-Factoid Question Answering on Instructional Videos
Kyungjae Lee 0002, Nan Duan 0001, Lei Ji 0001, Seung-won Hwang
AAAI1
2019 Learning with Limited Data for Multilingual Reading Comprehension
abstract
Kyungjae Lee, Sunghyun Park, Hojae Han, Jinyoung Yeo, Seung-won Hwang, Juho Lee. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Kyungjae Lee 0002, Sunghyun Park 0005, Hojae Han, Jinyoung Yeo, Seung-won Hwang
EMNLP/IJCNLP (1)1
2019 Categorical Metadata Representation for Customized Text Classification
abstract
The performance of text classification has improved tremendously using intelligently engineered neural-based models, especially those injecting categorical metadata as additional information, e.g., using user/product information for sentiment classification. This information has been used to modify parts of the model (e.g., word embeddings, attention mechanisms) such that results can be customized according to the metadata. We observe that current representation methods for categorical metadata, which are devised for human consumption, are not as effective as claimed in popular classification methods, outperformed even by simple concatenation of categorical features in the final layer of the sentence encoder. We conjecture that categorical features are harder to represent for machine use, as available context only indirectly describes the category, and even such context is often scarce (for tail category). To this end, we propose using basis vectors to effectively incorporate categorical metadata on various parts of a neural-based model. This additionally decreases the number of parameters dramatically, especially when the number of categorical features is large. Extensive experiments on various data sets with different properties are performed and show that through our method, we can represent categorical metadata more effectively to customize parts of the model, including unexplored ones, and increase the performance of the model greatly.
Jihyeok Kim, Reinald Kim Amplayo, Kyungjae Lee 0002, Sua Sung, Minji Seo, Seung-won Hwang
Trans. Assoc. Comput. Linguistics3
2018 Translations as Additional Contexts for Sentence Classification
abstract
In sentence classification tasks, additional contexts, such as the neighboring sentences, may improve the accuracy of the classifier. However, such contexts are domain-dependent and thus cannot be used for another classification task with an inappropriate domain. In contrast, we propose the use of translated sentences as domain-free context that is always available regardless of the domain. We find that naive feature expansion of translations gains only marginal improvements and may decrease the performance of the classifier, due to possible inaccurate translations thus producing noisy sentence vectors. To this end, we present multiple context fixing attachment (MCFA), a series of modules attached to multiple sentence vectors to fix the noise in the vectors using the other sentence vectors as context. We show that our method performs competitively compared to previous models, achieving best classification performance on multiple data sets. We are the first to use translations as domain-free contexts for sentence classification.
Reinald Kim Amplayo, Kyungjae Lee 0002, Jinyeong Yeo, Seung-won Hwang
IJCAI2
2018 Semi-supervised Training Data Generation for Multilingual Question Answering
Kyungjae Lee 0002, Kyoungho Yoon, Sunghyun Park 0005, Seung-won Hwang
LREC1
2017 Gradable Adjective Embedding for Commonsense Knowledge
Kyungjae Lee 0002, Hyunsouk Cho, Seung-won Hwang
PAKDD (2)1