VLDB 2026 Research / reviewers in the wild / expert
Yoonjoo Lee
dblp:31/10855
· DBLP profile ↗
15ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0001-7491-986XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evalet: Evaluating Large Language Models through Functional Fragmentation
Tae Soo Kim 0002, Heechan Lee, Yoonjoo Lee, Joseph Seering, Juho Kim 0001 |
CHI | 3 |
| 2025 | Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM ReviewsabstractHyungyu Shin, Jingyu Tang, Yoonjoo Lee, Nayoung Kim, Hyunseung Lim, Ji Yong Cho, Hwajung Hong, Moontae Lee, Juho Kim. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Hyungyu Shin, Yoonjoo Lee, Hyunseung Lim, Ji Yong Cho, Hwajung Hong, Moontae Lee, Juho Kim 0001 |
EMNLP | 3 |
| 2025 | The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language ModelsabstractSeungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Choi 0001, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Y. Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee 0002, Minjoon Seo |
NAACL (Long Papers) | 22 |
| 2025 | PANORAMA: A Dataset and Benchmarks Capturing Decision Trails and Rationales in Patent ExaminationabstractPatent examination remains an ongoing challenge in the NLP literature even after the advent of large language models (LLMs), as it requires an extensive yet nuanced human judgment on whether a submitted $\textit{claim}$ meets the statutory standards of $\textit{novelty}$ and $\textit{non-obviousness}$ against previously granted claims—$\textit{prior art}$—in expert domains. Previous NLP studies have approached this challenge as a prediction task (e.g., forecasting grant outcomes) with high-level proxies such as similarity metrics or classifiers trained on historical labels. However, this approach often overlooks the step-by-step evaluations that examiners must make with profound information, including rationales for the decisions provided in $\textit{office actions}$ documents, which also makes it harder to measure the current state of techniques in patent review processes. To fill this gap, we construct PANORAMA, a dataset of 8,143 U.S. patent examination records that preserves the full decision trails, including original applications, all cited references, $\textit{Non-Final Rejections}$, and $\textit{Notices of Allowance}$. Also, PANORAMA decomposes the trails into sequential benchmarks that emulate patent professionals' patent review processes and allow researchers to examine large language models' capabilities at each step of them. Our findings indicate that, although LLMs are relatively effective at retrieving relevant prior art and pinpointing the pertinent paragraphs, they struggle to assess the novelty and non-obviousness of patent claims. We discuss these results and argue that advancing NLP, including LLMs, in the patent domain requires a deeper understanding of real-world patent examination. Our dataset is openly available at https://huggingface.co/datasets/LG-AI-Research/PANORAMA. Hyunseung Lim, Sooyohn Nam, Sungmin Na, Ji Yong Cho, June Yong Yang, Hyungyu Shin, Yoonjoo Lee, Juho Kim 0001, Moontae Lee, Hwajung Hong |
NeurIPS | 7 |
| 2024 | EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaabstractBy simply composing prompts, developers can prototype novel generative applications with Large Language Models (LLMs). To refine prototypes into products, however, developers must iteratively revise prompts by evaluating outputs to diagnose weaknesses. Formative interviews (N=8) revealed that developers invest significant effort in manually evaluating outputs as they assess context-specific and subjective criteria. We present EvalLM, an interactive system for iteratively refining prompts by evaluating multiple outputs on user-defined criteria. By describing criteria in natural language, users can employ the system’s LLM-based evaluator to get an overview of where prompts excel or fail, and improve these based on the evaluator’s feedback. A comparative study (N=12) showed that EvalLM, when compared to manual evaluation, helped participants compose more diverse criteria, examine twice as many outputs, and reach satisfactory prompts with 59% fewer revisions. Beyond prompts, our work can be extended to augment model evaluation and alignment in specific application contexts. Tae Soo Kim 0002, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, Juho Kim 0001 |
CHI | 2 |
| 2024 | VIVID: Human-AI Collaborative Authoring of Vicarious Dialogues from Lecture VideosabstractThe lengthy monologue-style online lectures cause learners to lose engagement easily. Designing lectures in a “vicarious dialogue” format can foster learners’ cognitive activities more than monologue-style. However, designing online lectures in a dialogue style catered to the diverse needs of learners is laborious for instructors. We conducted a design workshop with eight educational experts and seven instructors to present key guidelines and the potential use of large language models (LLM) to transform a monologue lecture script into pedagogically meaningful dialogue. Applying these design guidelines, we created VIVID which allows instructors to collaborate with LLMs to design, evaluate, and modify pedagogical dialogues. In a within-subjects study with instructors (N=12), we show that VIVID helped instructors select and revise dialogues efficiently, thereby supporting the authoring of quality dialogues. Our findings demonstrate the potential of LLMs to assist instructors with creating high-quality educational dialogues across various learning stages. Seulgi Choi, Yoonjoo Lee, Juho Kim 0001 |
CHI | 3 |
| 2024 | PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected PapersabstractWith the rapid growth of scholarly archives, researchers subscribe to “paper alert’’ systems that periodically provide them with recommendations of recently published papers that are similar to previously collected papers. However, researchers sometimes struggle to make sense of nuanced connections between recommended papers and their own research context, as existing systems only present paper titles and abstracts. To help researchers spot these connections, we present PaperWeaver, an enriched paper alerts system that provides contextualized text descriptions of recommended papers based on user-collected papers. PaperWeaver employs a computational method based on Large Language Models (LLMs) to infer users’ research interests from their collected papers, extract context-specific aspects of papers, and compare recommended and collected papers on these aspects. Our user study (N=15) showed that participants using PaperWeaver were able to better understand the relevance of recommended papers and triage them more confidently when compared to a baseline that presented the related work sections from recommended papers. Yoonjoo Lee, Hyeonsu B. Kang, Matt Latzke, Juho Kim 0001, Jonathan Bragg, Joseph Chee Chang, Pao Siangliulue |
CHI | 1 |
| 2024 | ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language ModelsabstractBenjamin Newman, Yoonjoo Lee, Aakanksha Naik, Pao Siangliulue, Raymond Fok, Juho Kim, Daniel S Weld, Joseph Chee Chang, Kyle Lo. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Benjamin Newman, Yoonjoo Lee, Aakanksha Naik, Pao Siangliulue, Raymond Fok, Juho Kim 0001, Daniel S. Weld, Joseph Chee Chang, Kyle Lo |
EMNLP | 2 |
| 2023 | DAPIE: Interactive Step-by-Step Explanatory Dialogues to Answer Children's Why and How QuestionsabstractChildren acquire an understanding of the world by asking “why” and “how” questions. Conversational agents (CAs) like smart speakers or voice assistants can be promising respondents to children’s questions as they are more readily available than parents or teachers. However, CAs’ answers to “why” and “how” questions are not designed for children, as they can be difficult to understand and provide little interactivity to engage the child. In this work, we propose design guidelines for creating interactive dialogues that promote children’s engagement and help them understand explanations. Applying these guidelines, we propose DAPIE, a system that answers children’s questions through interactive dialogue by employing an AI-based pipeline that automatically transforms existing long-form answers from online sources into such dialogues. A user study (N=16) showed that, with DAPIE, children performed better in an immediate understanding assessment while also reporting higher enjoyment than when explanations were presented sentence-by-sentence. Yoonjoo Lee, Tae Soo Kim 0002, Sungdong Kim, Yohan Yun, Juho Kim 0001 |
CHI | 1 |
| 2023 | QASA: Advanced Question Answering on Scientific ArticlesabstractReasoning is the crux of intellectual thinking. While question answering (QA) tasks are prolific with various computational models and benchmark datasets, they mostly tackle factoid or shallow QA without asking deeper understanding. Dual process theory asserts that human reasoning consists of associative thinking to collect relevant pieces of knowledge and logical reasoning to consciously conclude grounding on evidential rationale. Based on our intensive think-aloud study that revealed the three types of questions: surface, testing, and deep questions, we first propose the QASA benchmark that consists of 1798 novel question answering pairs that require full-stack reasoning on scientific articles in AI and ML fields. Then we propose the QASA approach that tackles the full-stack reasoning with large language models via associative selection, evidential rationale-generation, and systematic composition. Our experimental results show that QASA's full-stack inference outperforms the state-of-the-art InstructGPT by a big margin. We also find that rationale-generation is critical for the performance gain, claiming how we should rethink advanced question answering. The dataset is available at https://github.com/lgresearch/QASA. Yoonjoo Lee, Kyungjae Lee 0002, Sunghyun Park 0005, Dasol Hwang, Jaehyeon Kim, Hong-In Lee, Moontae Lee |
ICML | 1 |
| 2023 | Cells, Generators, and Lenses: Design Framework for Object-Oriented Interaction with Large Language ModelsabstractLarge Language Models (LLMs) have become the backbone of numerous writing interfaces with the goal of supporting end-users across diverse writing tasks. While LLMs reduce the effort of manual writing, end-users may need to experiment and iterate with various generation configurations (e.g., inputs and model parameters) until results meet their goals. However, these interfaces are not designed for experimentation and iteration, and can restrict how end-users track, compare, and combine configurations. In this work, we present “cells, generators, and lenses”, a framework to designing interfaces that support interactive objects that embody configuration components (i.e., input, model, output). Interface designers can apply our framework to produce interfaces that enable end-users to create variations of these objects, combine and recombine them into new configurations, and compare them in parallel to efficiently iterate and experiment with LLMs. To showcase how our framework generalizes to diverse writing tasks, we redesigned three different interfaces—story writing, copywriting, and email composing—and, to demonstrate its effectiveness in supporting end-users, we conducted a comparative study (N=18) where participants used our interactive objects to generate and experiment more. Finally, we investigate the usability of the framework through a workshop with designers (N=3) where we observed that our framework served as both bootstrapping and inspiration in the design process. Tae Soo Kim 0002, Yoonjoo Lee, Minsuk Chang, Juho Kim 0001 |
UIST | 2 |
| 2022 | Promptiverse: Scalable Generation of Scaffolding Prompts Through Human-AI Hybrid Knowledge Graph AnnotationabstractOnline learners are hugely diverse with varying prior knowledge, but most instructional videos online are created to be one-size-fits-all. Thus, learners may struggle to understand the content by only watching the videos. Providing scaffolding prompts can help learners overcome these struggles through questions and hints that relate different concepts in the videos and elicit meaningful learning. However, serving diverse learners would require a spectrum of scaffolding prompts, which incurs high authoring effort. In this work, we introduce Promptiverse, an approach for generating diverse, multi-turn scaffolding prompts at scale, powered by numerous traversal paths over knowledge graphs. To facilitate the construction of the knowledge graphs, we propose a hybrid human-AI annotation tool, Grannotate. In our study (N=24), participants produced 40 times more on-par quality prompts with higher diversity, through Promptiverse and Grannotate, compared to hand-designed prompts. Promptiverse presents a model for creating diverse and adaptive learning experiences online. Yoonjoo Lee, John Joon Young Chung, Tae Soo Kim 0002, Jean Y. Song, Juho Kim 0001 |
CHI | 1 |
| 2022 | XDesign: Integrating Interface Design into Explainable AI EducationabstractWe introduce XDesign, a web-based interactive platform that guides learners through a multi-stage design process for creating user-centered explanations of AI models. Results from a course deployment show that students were able to identify concrete user needs in interacting with explanations, highlight user tasks to support the needs, and design a user interface that aids the tasks. Hyungyu Shin, Nabila Sindi, Yoonjoo Lee, Jaeryoung Ka, Jean Y. Song, Juho Kim 0001 |
SIGCSE (2) | 3 |
| 2021 | Personalizing Ambience and Illusionary Presence: How People Use "Study with me" Videos to Create Effective Studying Environmentsabstract“Study with me” videos contain footage of people studying for hours, in which social components like conversations or informational content like instructions are absent. Recently, they became increasingly popular on video-sharing platforms. This paper provides the first broad look into what “study with me” videos are and how people use them. We analyzed 30 “study with me” videos and conducted 12 interviews with their viewers to understand their motivation and viewing practices. We identified a three-factor model that explains the mechanism for shaping a satisfactory studying experience in general. One of the factors, a well-suited ambience, was difficult to achieve because of two common challenges: external conditions that prevent studying in study-friendly places and extra cost needed to create a personally desired ambience. We found that the viewers used “study with me” videos to create a personalized ambience at a lower cost, to find controllable peer pressure, and to get emotional support. These findings suggest that the viewers self-regulate their learning through watching “study with me” videos to improve efficiency even when studying alone at home. Yoonjoo Lee, John Joon Young Chung, Jean Y. Song, Minsuk Chang, Juho Kim 0001 |
CHI | 1 |
| 2011 | GMM-Based KLT-Domain Switched-Split Vector Quantization for LSF CodingabstractFor quantization of line spectral frequency (LSF), Gaussian mixture model (GMM) based switched split vector quantization (SSVQ) has been reported as the best performing intra-frame coding method. However, GMM-SSVQ partly recovers correlations between the subvectors of split vector quantization (SVQ). In the proposed GMM-SSVQ with the Karhunen-Loève Transform (KLT), KLT-domain quantization for each mixture with a novel region-clustering algorithm is applied to GMM-SSVQ. Compared with SVQ and GMM-SSVQ, it provides 4 and 1 bit higher performance in terms of average spectral distortion and outliers, respectively. Computational complexity and memory requirements are similar to GMM-SSVQ. Yoonjoo Lee, Wonjin Jung, Moo Young Kim 0001 |
IEEE Signal Process. Lett. | 1 |