Daehyun Kim 0005

dblp:85/5149-5 · also Dae Hyun Kim 0005 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-8657-9986ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 13 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 DiaryPlay: AI-Assisted Creation of Interactive Story Vignettes for Everyday Storytelling
abstract
An interactive vignette is a popular and immersive visual storytelling approach that invites viewers to role-play a character and influences the narrative in an interactive environment. However, it has not been widely used by everyday storytellers yet due to authoring complexity, which conflicts with the immediacy of everyday storytelling. We introduce DiaryPlay, an AI-assisted authoring system for interactive vignette creation in everyday storytelling. It takes a natural language story as input and extracts the three core elements of an interactive vignette (environment, characters, and events), enabling authors to focus on refining these elements instead of constructing them from scratch. Then, it automatically transforms the single-branch story input into a branch-and-bottleneck structure using an LLM-powered narrative planner, which enables flexible viewer interactions while freeing the author from multi-branching. A technical evaluation (N=16) shows that DiaryPlay-generated character activities are on par with human-authored ones regarding believability. A user study (N=16) shows that DiaryPlay effectively supports authors in creating interactive vignette elements, maintains authorial intent while reacting to viewer interactions, and provides engaging viewing experiences.
Jiangnan Xu, Haeseul Cha, Gosu Choi, Gyu-cheol Lee, Yeo-Jin Yoon, Zucheul Lee, Konstantinos Papangelis, Daehyun Kim 0005, Juho Kim 0001
CHI8
2026 Understanding Spatiotemporal-Aware Multimodal Conversational Search in the Outdoor Urban Space
abstract
Emerging multimodal conversational search (MCS) tools (e.g., Gemini Live) allow users to search for spatiotemporal information through natural language dialogues as they move through urban space. Despite the growing popularity of these tools, there is limited understanding of how people engage with this technology. To address this gap, we developed UrbanSearch, an MCS technology probe designed to capture the user’s current geolocation, time, and visual surroundings. A contextual inquiry (N=23) revealed that MCS tools provide two core values: requiring low effort in forming queries while offering highly relevant responses, and functioning as a central information gateway. As a promising technology, MCS supports environmental learning, in-situ decision making, and personalized navigation. Participants also revealed unmet needs for spatial reasoning and transparent integration of multi-source information, along with concerns related to peripheral awareness, social context, and personal space. Drawing from the findings, we discuss design implications for future MCS tools in urban spaces.
Jiangnan Xu, Suyeon Seo, Joni Salminen, Michael Saker, Joon Gi Shin, Alan Chamberlain, Konstantinos Papangelis, Daehyun Kim 0005
CHI8
2026 Cerebra: Aligning Implicit Knowledge in Interactive SQL Authoring
abstract
LLM-driven tools have significantly lowered barriers to writing SQL queries. However, user instructions are often underspecified, assuming the model understands implicit knowledge, such as dataset schemas, domain conventions, and task-specific requirements, that isn’t explicitly provided. This results in frequently erroneous scripts that require users to repeatedly clarify their intent. Additionally, users struggle to validate generated scripts because they cannot verify whether the model correctly applied implicit knowledge. We present Cerebra, an interactive NL-to-SQL tool that aligns implicit knowledge between users and LLMs during SQL authoring. Cerebra automatically retrieves implicit knowledge from historical SQL scripts based on user instructions, presents this knowledge in an interactive tree view for code review, and supports iterative refinement to improve generated scripts. To evaluate the effectiveness and usability of Cerebra, we conducted a user study with 16 participants, demonstrating its improved support for customized SQL authoring. The source code of Cerebra is available at https://github.com/zjuidg/CHI26-Cerebra.
Yunfan Zhou, Qiming Shi, Zhongsu Luo, Xiwen Cai, Yanwei Huang, Daehyun Kim 0005, Di Weng, Yingcai Wu
CHI6
2026 Dataset-Adaptive Dimensionality Reduction
abstract
Selecting the appropriate dimensionality reduction (DR) technique and determining its optimal hyperparameter settings that maximize the accuracy of the output projections typically involves extensive trial and error, often resulting in unnecessary computational overhead. To address this challenge, we propose a dataset-adaptive approach to DR optimization guided by structural complexity metrics. These metrics quantify the intrinsic complexity of a dataset, predicting whether higher-dimensional spaces are necessary to represent it accurately. Since complex datasets are often inaccurately represented in two-dimensional projections, leveraging these metrics enables us to predict the maximum achievable accuracy of DR techniques for a given dataset, eliminating redundant trials in optimizing DR. We introduce the design and theoretical foundations of these structural complexity metrics. We quantitatively verify that our metrics effectively approximate the ground truth complexity of datasets and confirm their suitability for guiding dataset-adaptive DR workflow. Finally, we empirically show that our dataset-adaptive workflow significantly enhances the efficiency of DR optimization without compromising accuracy.
Hyeon Jeon, Jeongin Park, Daehyun Kim 0005, Sungbok Shin, Jinwook Seo
IEEE Trans. Vis. Comput. Graph.4
2025 PlanTogether: Facilitating AI Application Planning Using Information Graphs and Large Language Models
Daehyun Kim 0005, Daeheon Jeong, Shakhnozakhon Yadgarova, Hyungyu Shin, Jinho Son, Hariharan Subramonyam, Juho Kim 0001
CHI1
2025 Generative AI personas considered harmful? Putting forth twenty challenges of algorithmic user representation in human-computer interaction
abstract
• Shows how GenAI fundamentally transforms existing persona development issues through evolutionary amplification rather than creating entirely new problems, with traditional biases becoming algorithmic discrimination and manual inconsistencies becoming convincing AI hallucinations. • Reveals how traditional limitations manifest differently in GenAI contexts across transparency, fairness, reliability, and control domains, with expert validation showing 60% of challenges are more problematic for GenAIPs than conventional approaches. • Documents how GenAI transforms not just technical challenges but harm distribution, with persona developers facing operational complexity while target user groups bear severe consequences through systematic misrepresentation and exclusion. • Provides evidence that while GenAIPs appear to solve traditional limitations, they transform existing challenges into more complex forms requiring novel validation approaches and human-AI collaboration frameworks for responsible implementation. Generative AI personas (GenAIPs) promise user-centred design efficiency, but their impact on different persona challenges remains unexplored. Inspired by Dijkstra’s classic essay on harmful programming constructs, we analyze twenty challenges in persona development using Human-Centered AI principles. Through literature review and expert survey (n=17), we find that GenAIPs transform rather than eliminate traditional persona challenges. Experts rated all challenges as problematic for GenAIPs (M > 4.0), with the highest concerns for hallucinations (M=5.94), over-sanitization (M=5.82), and lack of standardization (M=5.59). 12 out of 20 challenges are considered more problematic for GenAIPs than conventional personas, particularly bias amplification, validation challenges, and accessibility without expertise. We provide HCAI-grounded guidelines demonstrating that effective GenAIP implementation requires human-AI collaboration rather than automation and prioritizing user welfare over technical efficiency.
Danial Amin 0001, Joni Salminen, Jim Jansen, Joon Gi Shin, Daehyun Kim 0005
Int. J. Hum. Comput. Stud.5
2024 A Context-Aware Onboarding Agent for Metaverse Powered by Large Language Models
abstract
One common asset of metaverse is that users can freely explore places and actions without linear procedures. Thus, it is hard yet important to understand the divergent challenges each user faces when onboarding metaverse. Our formative study (N = 16) shows that first-time users ask questions about metaverse that concern 1) a short-term spatiotemporal context, regarding the user’s current location, recent conversation, and actions, and 2) a long-term exploration context regarding the user’s experience history. Based on the findings, we present PICAN, a Large Language Model-based pipeline that generates context-aware answers to users when onboarding metaverse. An ablation study (N = 20) reveals that PICAN’s usage of context made responses more useful and immersive than those generated without contexts. Furthermore, a user study (N = 21) shows that the use of long-term exploration context promotes users’ learning about the locations and activities within the virtual environment.
Jihyeong Hong, Yokyung Lee, Daehyun Kim 0005, Daeun Choi, Yeo-Jin Yoon, Gyu-cheol Lee, Zucheul Lee, Juho Kim 0001
Conference on Designing Interactive Systems3
2024 AINeedsPlanner: A Workbook to Support Effective Collaboration Between AI Experts and Clients
abstract
Clients often partner with AI experts to develop AI applications tailored to their needs. In these partnerships, careful planning and clear communication are critical, as inaccurate or incomplete specifications can result in misaligned model characteristics, expensive reworks, and potential friction between collaborators. Unfortunately, given the complexity of requirements ranging from functionality, data, and governance, effective guidelines for collaborative specification of requirements in client-AI expert collaborations are missing. In this work, we introduce AINeedsPlanner, a workbook that AI experts and clients can use to facilitate effective interchange of clear specifications. The workbook is based on (1) an interview of 10 completed AI application project teams, which identifies and characterizes steps in AI application planning and (2) a study with 12 AI experts, which defines a taxonomy of AI experts’ information needs and dimensions that affect the information needs. Finally, we demonstrate the workbook’s utility with two case studies in real-world settings.
Daehyun Kim 0005, Hyungyu Shin, Shakhnozakhon Yadgarova, Jinho Son, Hariharan Subramonyam, Juho Kim 0001
Conference on Designing Interactive Systems1
2024 Natural Language Dataset Generation Framework for Visualizations Powered by Large Language Models
abstract
We introduce VL2NL, a Large Language Model (LLM) framework that generates rich and diverse NL datasets using Vega-Lite specifications as input, thereby streamlining the development of Natural Language Interfaces (NLIs) for data visualization. To synthesize relevant chart semantics accurately and enhance syntactic diversity in each NL dataset, we leverage 1) a guided discovery incorporated into prompting so that LLMs can steer themselves to create faithful NL datasets in a self-directed manner; 2) a score-based paraphrasing to augment NL syntax along with four language axes. We also present a new collection of 1,981 real-world Vega-Lite specifications that have increased diversity and complexity than existing chart collections. When tested on our chart collection, VL2NL extracted chart semantics and generated L1/L2 captions with 89.4% and 76.0% accuracy, respectively. It also demonstrated generating and paraphrasing utterances and questions with greater diversity compared to the benchmarks. Last, we discuss how our NL datasets and framework can be utilized in real-world scenarios. The codes and chart collection are available at https://github.com/hyungkwonko/chart-llm.
Hyung-Kwon Ko, Hyeon Jeon, Gwanmo Park, Daehyun Kim 0005, Juho Kim 0001, Jinwook Seo
CHI4
2024 DataDive: Supporting Readers' Contextualization of Statistical Statements with Data Exploration
abstract
Statistical statements that refer to data to support narratives or claims are commonly used to inform readers about the magnitude of social issues. While contextualizing statistical statements with relevant data supports readers in building their own interpretation of statements, the complexity of finding contextual information on the web and linking statistical statements with it impedes readers’ efforts to do so. We present DataDive, an interactive tool for contextualizing statistical statements for the readers of online texts. Based on users’ selections of statistical statements, our tool uses an LLM-powered pipeline to generate candidates of relevant contexts and poses them as guiding questions to the user as potential contexts for exploration. When the user selects a question, DataDive employs visualizations to further help the user compare and explore contextually relevant data. A technical evaluation shows that DataDive generates important and diverse questions that facilitate exploration around statistical statements and retrieves relevant data for comparison. Moreover, a user study with 21 participants suggests that DataDive facilitates users to explore diverse contexts and to be more aware of how statistical data could relate to the text.
Khanh-Duy Le, Gionnieve Lim, Daehyun Kim 0005, Yoo Jin Hong, Juho Kim 0001
IUI4
2024 EC: A Tool for Guiding Chart and Caption Emphasis
abstract
Recent work has shown that when both the chart and caption emphasize the same aspects of the data, readers tend to remember the doubly-emphasized features as takeaways; when there is a mismatch, readers rely on the chart to form takeaways and can miss information in the caption text. Through a survey of 280 chart-caption pairs in real-world sources (e.g., news media, poll reports, government reports, academic articles, and Tableau Public), we find that captions often do not emphasize the same information in practice, which could limit how effectively readers take away the authors' intended messages. Motivated by the survey findings, we present EMPHASISCHECKER, an interactive tool that highlights visually prominent chart features as well as the features emphasized by the caption text along with any mismatches in the emphasis. The tool implements a time-series prominent feature detector based on the Ramer-Douglas-Peucker algorithm and a text reference extractor that identifies time references and data descriptions in the caption and matches them with chart data. This information enables authors to compare features emphasized by these two modalities, quickly see mismatches, and make necessary revisions. A user study confirms that our tool is both useful and easy to use when authoring charts and captions.
Daehyun Kim 0005, Seulgi Choi, Juho Kim 0001, Vidya Setlur, Maneesh Agrawala
IEEE Trans. Vis. Comput. Graph.1
2021 Towards Understanding How Readers Integrate Charts and Captions: A Case Study with Line Charts
abstract
Charts often contain visually prominent features that draw attention to aspects of the data and include text captions that emphasize aspects of the data. Through a crowdsourced study, we explore how readers gather takeaways when considering charts and captions together. We first ask participants to mark visually prominent regions in a set of line charts. We then generate text captions based on the prominent features and ask participants to report their takeaways after observing chart-caption pairs. We find that when both the chart and caption describe a high-prominence feature, readers treat the doubly emphasized high-prominence feature as the takeaway; when the caption describes a low-prominence chart feature, readers rely on the chart and report a higher-prominence feature as the takeaway. We also find that external information that provides context, helps further convey the caption’s message to the reader. We use these findings to provide guidelines for authoring effective chart-caption pairs.
Daehyun Kim 0005, Vidya Setlur, Maneesh Agrawala
CHI1
2020 Answering Questions about Charts and Generating Visual Explanations
abstract
People often use charts to analyze data, answer questions and explain their answers to others. In a formative study, we find that such human-generated questions and explanations commonly refer to visual features of charts. Based on this study, we developed an automatic chart question answering pipeline that generates visual explanations describing how the answer was obtained. Our pipeline first extracts the data and visual encodings from an input Vega-Lite chart. Then, given a natural language question about the chart, it transforms references to visual attributes into references to the data. It next applies a state-of-the-art machine learning algorithm to answer the transformed question. Finally, it uses a template-based approach to explain in natural language how the answer is determined from the chart's visual features. A user study finds that our pipeline-generated visual explanations significantly outperform in transparency and are comparable in usefulness and trust to human-generated explanations.
Daehyun Kim 0005, Enamul Hoque Prince, Maneesh Agrawala
CHI1
2020 Sneak Pique: Exploring Autocompletion as a Data Discovery Scaffold for Supporting Visual Analysis
abstract
Natural language interaction has evolved as a useful modality to help users explore and interact with their data during visual analysis. Little work has been done to explore how autocompletion can help with data discovery while helping users formulate analytical questions. We developed a system called \system as a design probe to better understand the usefulness of autocompletion for visual analysis. We ran three Mechanical Turk studies to evaluate user preferences for various text- and visualization widget-based autocompletion design variants for helping with partial search queries. Our findings indicate that users found data previews to be useful in the suggestions. Widgets were preferred for previewing temporal, geospatial, and numerical data while text autocompletion was preferred for categorical and hierarchical data. We conducted an exploratory analysis of our system implementing this specific subset of preferred autocompletion variants. Our insights regarding the efficacy of these autocompletion suggestions can inform the future design of natural language interfaces supporting visual analysis.
Vidya Setlur, Enamul Hoque Prince, Daehyun Kim 0005, Angel X. Chang
UIST3
2018 Facilitating Document Reading by Linking Text and Tables
abstract
Document authors commonly use tables to support arguments presented in the text. But, because tables are usually separate from the main body text, readers must split their attention between different parts of the document. We present an interactive document reader that automatically links document text with corresponding table cells. Readers can select a sentence (or tables cells) and our reader highlights the relevant table cells (or sentences). We provide an automatic pipeline for extracting such references between sentence text and table cells for existing PDF documents that combines structural analysis of tables with natural language processing and rule-based matching. On a test corpus of 330 (sentence, table) pairs, our pipeline correctly extracts 48.8% of the references. An additional 30.5% contain only false negatives (FN) errors -- the reference is missing table cells. The remaining 20.7% contain false positives (FP) errors -- the reference includes extraneous table cells and could therefore mislead readers. A user study finds that despite such errors, our interactive document reader helps readers match sentences with corresponding table cells more accurately and quickly than a baseline document reader.
Daehyun Kim 0005, Enamul Hoque Prince, Juho Kim 0001, Maneesh Agrawala
UIST1