VLDB 2026 Research / reviewers in the wild / expert
Lujie Karen Chen
dblp:289/6421
· DBLP profile ↗
24ranked-venue papers
10as first author
24since 2021 · last 2026
0000-0002-7185-8405ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 18 · 9 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Show and Tell: Explainable AI-Supported Feedback for Developing Data Visualization Sense-Making Skills
Supakit Boonsongprasert, Sachin Pathak, Chris Song, Lujie Karen Chen |
AIED (1) | 4 |
| 2026 | Exploring Student Feedback Needs and Design Opportunities in Data Storytelling EducationabstractData storytelling workflows ask learners to integrate analytical, design, and narrative skills, but instructors rarely have the capacity to provide detailed feedback at each step. Computational and AI-assisted storytelling offers opportunities to support student learning, but how feedback should be structured effectively remains unclear. To address this gap, we conducted a two-phase participatory design study. Through participant observations (N=8) and interviews (N=6), the first phase explored learners and educators’ feedback needs and challenges in a data storytelling course. The second phase conducted two design workshops (N=8/10) to design and evaluate feedback strategies (frequency, seamlessness, accountability) for Story Studio: an AI-assisted narrative storytelling application. Our findings show that participants perceived on-demand and process feedback modes as effective, but automatic and outcome feedback as slightly more persuasive. We discuss implications for designing AI-augmented storytelling systems that adapt their feedback modes to the diverse needs and expectations of students. Jennifer Posada, Taha Hassan, Lujie Karen Chen, Louise Yarnall, Jiaqi Gong |
CHI | 3 |
| 2026 | An LLM-Based Agentic AI System for Automated Construction of Knowledge Taxonomies in Data Science Problem SolvingabstractThe widespread integration of artificial intelligence into data science practice is fundamentally reconfiguring the roles of data science practitioners. As procedural and execution-oriented tasks such as coding, data preprocessing, and routine analysis become increasingly automated, human contribution is shifting toward complex cognitive task of problem solving including higher-order reasoning and decision making. This shift challenges prevailing models of data science education, which has been focusing on tools and techniques while offering limited support to develop those critical problem solving competencies. A prerequisite for addressing this gap is formalizing the Data Science Problem Solving (DSPS) competency as structured knowledge that can be systematically represented and evaluated, for example, by adopting principled approach of constructing knowledge taxonomies for DSPS. However, constructing a coherent and verifiable taxonomy of such knowledge remains a nontrivial knowledge engineering task. In this paper, we present a multi-agent LLM framework for the automated generation and evaluation of a DSPS knowledge taxonomy, grounded in established taxonomy construction and evaluation methods. The system comprises an Author agent that proposes and revises taxonomies and a Critique agent that evaluates them and provides feedback based on explicit assessment criteria. Additionally, an Orchestrator agent manages iterative refinement by routing feedback and enforcing stopping conditions. Our results demonstrate the potential of organizing LLMs into collaborative-adversarial architectures that support reflective, iterative knowledge construction and refinement rather than single-pass content generation. When further developed, the proposed multi-agent LLM systems can function as a scalable framework for principled knowledge engineering, enabling the systematic design of DSPS curriculum, assessments, instructional scaffolds, and AI-assisted learning environments at scale. Md Sakib Ul Rahman Sourove, Shimei Pan, Lujie Karen Chen |
L@S | 3 |
| 2026 | Developing Problem-Solving Competency in Data Science: Exploring A Case-Based Approach
Lujie Karen Chen, Maryam Alomair, Muhammad Ali Yousuf, Shimei Pan |
SIGCSE (1) | 1 |
| 2026 | Fostering Cross-disciplinary Competency in Undergraduate Data Science Education: Exploring a Collaborative Teaching ApproachabstractWe report on a collaborative teaching experiment conducted in Fall 2024 between an introductory data science course and a health communication course at a four-year college on the U.S. East Coast. Four peer-learning sessions were designed to align with key problem-solving stages in the health communication curriculum, with students working in mixed teams under the guidance of instructors from both disciplines. To evaluate outcomes, we administered pre- and post-tests on disciplinary knowledge and collected student reflections on the value of interdisciplinary collaboration. Results show discernible cross-disciplinary learning gains, with data science students demonstrating stronger improvements in health communication knowledge, while health communication students reported moderate gains in data science knowledge. Survey and open-ended responses further reveal that data science students valued communication, teamwork, and public health perspectives, while health communication students highlighted learning about data use and collection in real-world contexts. These findings suggest that structured cross-disciplinary teaching collaboration can effectively foster mutual learning and broaden students' appreciation of complementary skills across fields. Lujie Karen Chen, Katie Birger |
SIGCSE (2) | 1 |
| 2026 | Show or Tell? Piloting an AI Feedback Tool for Data Story Reading in Introductory Data ScienceabstractThe ability to make sense of data visualizations is a core component of data literacy and an essential skill in data science and analytics education. Yet, it remains unclear how this skill can be deliberately taught and practiced. To address this challenge, we designed a classroom activity called data story reading, in which students generate a title and narrative in response to a given visualization. Narratives are distinguished as either ''Show'' (surface-level descriptions of what is displayed) or ''Tell'' (deeper interpretations that uncover trends, patterns, or anomalies). Building on this distinction, we developed an AI-supported feedback tool that automatically classifies student narratives at the sentence level as Show or Tell, presents the results to students, and invites them to critique and reflect on the AI's feedback. We piloted the tool in an introductory data science course in Spring 2025, enrolling students from diverse backgrounds. Data for analysis were drawn from system logs capturing sentence-level classifications, student critiques, and reflective comments. Findings suggest that AI feedback helped students recognize the difference between descriptive and interpretive sense-making, while also surfacing tensions in providing nuanced feedback. We discuss implications for AI-supported learning, limitations of the current tool, and directions for future refinement. Lujie Karen Chen, Supakit Boonsongprasert, Sachin Pathak |
SIGCSE (2) | 1 |
| 2026 | Story Studio Plus: Coaching Data Storytelling Competency in the era of AIabstractAs data storytelling—communicating insights through data—becomes increasingly essential across disciplines, the need for scalable instructional support in this area has never been greater. While the fields of information visualization and narrative design offer rich foundations, their integration into data science education remains limited. Educators are often left without adequate tools or pedagogical frameworks to teach data storytelling at scale. To address this gap, we introduce Story Studio Plus, an AI-empowered coaching tool designed to support collaborative learning among students, educators, and intelligent agents. Developed through iterative co-design with teachers, students, and domain experts, Story Studio blends principles from learning sciences, human-computer interaction, and narrative visualization. The tool provides formative feedback, scaffolded prompts, and interactive examples tailored to students' developmental stages. Uniquely, it positions AI not as a replacement for human instruction but as a collaborative partner—enhancing teacher facilitation, supporting student agency, and fostering a co-creative classroom culture. In this tutorial, participants will engage hands-on with Story Studio's latest features and explore how human educators, learners, and AI systems can co-construct knowledge through data storytelling. Participants will also contribute feedback to shape future iterations of the tool. This session offers an applied lens on how AI can be meaningfully integrated into data science education. This project is partially supported by National Science Foundation grants 2302794 and 2302795. We would like to thank software development work by Emily Jackson and other members from the UA SAIL lab and the human-centered design work led by Jennifer Posada from University of Maryland, Baltimore County. Lujie Karen Chen, Taha Hassan, Louise Yarnall, Jiaqi Gong |
SIGCSE (2) | 1 |
| 2026 | Personal Informatics in Undergraduate Data Science: Learning by Analyzing the SelfabstractPersonal informatics refers to the data-driven, iterative process of collecting personal data, reflecting on it, and using the resulting insights to inform positive behavioral change. This practice, which grows in popularity with the widespread availability of wearable devices, can also be viewed as a form of data-driven scientific inquiry into personal behavior. When embedded in educational contexts, personal informatics offers college students an authentic and active learning experience that cultivates both data science skills and self-regulation, which are critical competencies for learners whose executive functions are still developing. In this work, we describe a personal informatics project piloted in an introductory data science course. The project began with scaffolded supports to help students identify areas of their lives they wished to improve and to formulate testable hypotheses informed by existing literature. Students then collected personal data continuously over a five-week period, and met with their accountability partner weekly to discuss and reflect on the data based on initial analysis. In the second half of the semester, students engaged in hands-on labs designed to help them generate deeper insights into their own behaviors while building key data analytics skills. These included data wrangling and visualization, making sense of data artifacts, and applying statistical inference techniques. Students wraped up the projects by crafting creative products (e.g. Infographics or Memes) that can be shared with peer students. We present an overview of analysis of students' final reflection and creative artifacts. Our findings offer insights into how this multi-purpose learning activity can be improved to support students' development as data science learners, while empowering them to engage in meaningful, evidence-based inquiry grounded in their lived experiences. Lujie Karen Chen, Jennifer Posada, Sydnee Angus |
SIGCSE (2) | 1 |
| 2026 | 'Why Do I Feel Like a Fraud?': Understanding Imposter Phenomenon in Computing Students through Ecological Momentary AssessmentabstractImposter Phenomenon (IP), described as the persistent feeling of intellectual fraudulence despite evidence of competence, affects many computing students and has been linked to diminished academic engagement, and social withdrawal. While prior research has established IP's prevalence and contributing factors in computing education, little is known about how it manifests in computing students' daily academic lives. This paper presents findings from a seven-day Ecological Momentary Assessment (EMA) study in which 34 computing students submitted three daily reports capturing their experiences related to IP. Participants reported on stress levels, context (people, location, thoughts, tasks, and feedback), and six IP characteristics(feelings of over-preparation, feelings of procrastination, feeling a need to be the best, fear of being shamed/humiliated because of performance, feeling incapable/incompetent and a fear of success leading to higher expectations) in context. Our analysis reveals patterns between feedback, daily stress, and the emergence of IP characteristics, offering insight into how everyday interactions may reinforce or relieve IP feelings. By linking momentary experiences with psychological outcomes, we contribute empirical evidence for the design of intervention(s) aimed at improving student well-being in computing education. This is the first known EMA study to examine IP in computing students' daily context. Yetunde Esther Okueso, Sydnee Angus, Sri Kavya Penta, Lujie Karen Chen |
SIGCSE (1) | 4 |
| 2026 | Rethinking the Future of Data Science Education: A Case for Thoughtful Design to Integrate AI into the College Classroom
Louise Yarnall, Hui Yang 0025, Sophia Ouyang, Lujie Karen Chen |
SIGCSE (1) | 4 |
| 2025 | Causal Explanation of Quality of Parent-Child Interactions with Multimodal Behavioral Features (Student Abstract)abstractThe quality of interactions between parents and children is a critical factor in child development. Recent years have seen programs to improve parenting behaviors through evidence-based approaches, such as attachment-based interventions. A vital element of these programs is to assess the quality of parenting behaviors via video recordings of parent-child interactions, which is often time-intensive. In our previous work, we explored machine learning models to predict expert ratings of parenting behaviors from video recordings of semi-structured parent-child play. However, the large set of low-level multimodal features struggled to provide explainable insights, which created barriers to communicating with domain experts and improving the models further. In this work, we developed a machine learning pipeline that combines sparse multiple canonical correlation analysis with causal discovery techniques to uncover explainable causal relationships between nine categories of behavioral features and the quality ratings of parent-child interactions. This approach offers valuable insights into the otherwise black-box models and contributes to the growing body of work on transparent and trustworthy machine learning models of parenting behaviors. Katherine Guerrerio, Lujie Karen Chen, Lisa Berlin, Brenda Jones Harden |
AAAI | 2 |
| 2025 | Causal Explanation of the Quality of Parent-Child Interactions with Multimodal Behavioral Features
Katherine Guerrerio, Lujie Karen Chen, Lisa Berlin, Brenda Jones Harden |
ICMI | 2 |
| 2025 | Systematically Identifying, Defining and Organizing Knowledge Components for Data Science Problem Solving through Human-LLM CollaborationabstractAs demand grows for job-ready data science professionals, there is increasing recognition that traditional training often falls short in cultivating the higher-order reasoning and real-world problem-solving skills essential to the field. A foundational step toward addressing this gap is the identification and organization of knowledge components (KCs) that underlie data science problem solving (DSPS). KCs represent conditional knowledge-knowing about appropriate actions given particular contexts or conditions-and correspond to the critical decisions data scientists must make throughout the problem-solving process. While existing taxonomies in data science education support curriculum development, they often lack the granularity and focus needed to support the assessment and development of DSPS skills. In this paper, we present a novel framework that combines the strengths of large language models (LLMs) and human expertise to identify, define, and organize KCs specific to DSPS. We treat LLMs as ''knowledge engineering assistants'' capable of generating candidate KCs by drawing on their extensive training data, which includes a vast amount of domain knowledge and diverse sets of real-world DSPS cases. Our process involves prompting multiple LLMs to generate decision points, synthesizing and refining KC definitions across models, and using sentence-embedding models to infer the underlying structure of the resulting taxonomy. Human experts then review and iteratively refine the taxonomy to ensure validity. This human-AI collaborative workflow offers a scalable and efficient proof-of-concept for LLM-assisted knowledge engineering. The resulting KC taxonomy lays the groundwork for developing fine-grained assessment tools and adaptive learning systems that support deliberate practice in DSPS. Furthermore, the framework illustrates the potential of LLMs not just as content generators but as partners in structuring domain knowledge to inform instructional design. Future work will involve extending the framework by generating a directed graph of KCs based on their input-output dependencies and validating the taxonomy through expert consensus and learner studies. This approach contributes to both the practical advancement of DSPS coaching in data science education and the broader methodological toolkit for AI-supported knowledge engineering. Priyanka Rani, Maryam Alomair, Shimei Pan, Lujie Karen Chen |
L@S | 4 |
| 2025 | Laying Foundations for Scalable Coaching for Data Storytelling: A Evidence-Centered Design ApproachabstractAs demand for data scientists has increased to inform decision-making across multiple fields of societal importance, postsecondary institutions have expanded data science course offerings. Despite such growth, educators struggle to teach students all the skills central to data science. They focus on programming and statistical tools and lack time for mentoring students in data storytelling. This working paper reviewed literature and interviewed experts to model the domain knowledge of data storytelling to inform the design of intelligent technology to support data storytelling instruction at scale. The paper closes with a recommendation of two ways that artificial intelligence tools can support the development of students' data storytelling knowledge and skills: ''direct'' feedback to students on routine data science tasks and ''facilitated'' summaries of students' data story progress to inform instructors' feedback. We intend to apply these insights to the design of intelligent coaching in an online platform to support the development of storytelling competency at scale. Louise Yarnall, Hui Yang 0025, Sophia Ouyang, Lujie Karen Chen |
L@S | 4 |
| 2025 | Story Studio: A Coaching Tool to Support the Development of Data Storytelling Competency at ScaleabstractThe demand for data storytelling, or communication with data, has surged significantly in recent years. Despite its growing recognition, large-scale coaching support tools for students and instructors remain lacking. The field of information visualization has explored data visualization and storytelling. Still, this knowledge has yet to be integrated with learning science in classroom settings to enhance student learning. Post-secondary data science educators require specialized tools and guidance to teach data storytelling on a larger scale. This tutorial invites those interested in the pedagogy of data storytelling to explore the initial version of a new coaching tool, Story Studio, developed by our team. The tool's design is based on our understanding of the data storytelling process and the application of relevant learning science principles, incorporating insights from experts, educators, and students. During the tutorial, participants will engage in hands-on experiences with the tool and interact with the research and development team, providing valuable feedback to support the tool's iterative improvement. This session aims to bridge the gap between data storytelling research and educational practice, equipping educators with the necessary resources to effectively teach this critical skill set. This material is based upon work supported by the National Science Foundation under Grant No. 2302794 and 2302795. Lujie Karen Chen, Louise Yarnall, Jiaqi Gong |
SIGCSE (2) | 1 |
| 2025 | Engaging K-12 Learners in Data Annotation for AI Climate ModelsabstractDue to the climate crisis, summers in Greenland have been rapidly getting warmer, causing increasing rates of ice melt on the Greenland ice sheet and speeding up sea-level rise. Evidence of this change can be measured by the number and location (elevation) of water pools and lakes that form on the surface of the ice sheet. In addition, crevasses can cause lakes to drain extremely rapidly causing the ice to flow faster, contributing to sea-level rise. However, the lack of annotated data makes it difficult to automatically detect and track these behavioral changes in the polar ice sheet lakes. This study describes how a team of polar and data scientists actively engaged middle and high school students in their classrooms in a data annotation process through an engaging curriculum unit to identify multiple ice sheet phenomena observed in satellite imagery. The findings describe the learning outcomes from both student and teacher perspectives. It also projects learners' understanding and sentiments about climate change and the role of artificial intelligence (AI) models coupled as an extension of citizen science in addressing climate change. Michael MacFerrin, Edward Boyda, Kimberly Young, Josephine M. Namayanja, Aneesh Subramanian, Mohamed F. Mokbel, Lujie Karen Chen, Vandana Pursnani Janeja |
SIGCSE (2) | 7 |
| 2025 | Predicting Students' Interest from Small Group Conversational Characteristics: Insights from an AI Literacy Education with High School StudentsabstractRecent years have seen developments in AI instructional practices for K-12 students. In literature, students' interest in AI is shown to correlate with gaining AI knowledge; however, little is known about how AI interest manifests in classroom discourses during AI literacy lessons. This study examined students' participation in an integrated AI curriculum delivered to a cognitive science class in a high school in the southern US. Students worked in small groups and built a supervised machine learning model to recognize kids' drawings at different stages of artistic development. Our analysis showed that semantic features extracted from students' small group conversations significantly predicted their interest in learning AI. However, we found no significant relationship between students' social construction of knowledge and their interests. This study sheds light on the relationship between the learning process and interest; when further developed, this analysis may be developed into a classroom activity analytics tool that may provide real-time feedback to teachers engaged in AI literacy education to enhance teaching effectiveness in this nascent content area. Shenghua Zha, Lujie Karen Chen, Woei Hung, Na Gong, Pamela Moore, Bethany Klemetsrud |
SIGCSE (2) | 2 |
| 2024 | Towards Automated Multiple Choice Question Generation and Evaluation: Aligning with Bloom's Taxonomy
Kevin Hwang, Kenneth Wang 0004, Maryam Alomair, Fow-Sen Choa, Lujie Karen Chen |
AIED (2) | 5 |
| 2024 | Foundational Tools for Coaching Data StorytellingabstractData storytelling is the skill to communicate data effectively and efficiently. Effective data storytelling goes beyond data visualization and focuses on explanation with clear rhetorical functions. It starts with a set of data insights collected from the data science workflow and involves iterative and interactive processes of filtering those insights into story slices, from which data stories can be created through ordering, organizing and narration. Data storytelling is an integral component of a well-rounded data science education, which complements foundational skills like quantitative reasoning and programming. Despite its significance, solid understanding of the theory and practice of developing data storytelling competency is lacking. Data storytelling is often perceived as a mythical process where quantitative information magically transforms into compelling narratives. Designing scalable coaching tools for data storytelling requires leveraging multidisciplinary expertise from learning science, computer science, data science, communication science, and human-centered design. In this workshop, we will share some initial findings and reflections from our interdisciplinary team searching for effective coaching methods and tools to support coaching data storytelling at scale. We will present results from literature reviews and expert interviews which will be packaged into a set of foundational tools such as mental model, cognitive processes and schema for story construction, assessment strategy, as well as preliminary ideas of tools to support data storytelling coaching. We hope to use this workshop to build a community of researchers and practitioners in coaching data storytelling in postsecondary formal and informal learning context. Lujie Karen Chen, Jiaqi Gong, Louise Yarnall |
SIGCSE (2) | 1 |
| 2024 | Assessment-via-Teaching: Exploring an Alternative Assessment Strategy in Undergraduate Introductory Data Science CourseabstractLearning-by-teaching is an active learning method that has the promise of engaging students and enhancing their learning. These benefits include the improvement of the metacognitive process, increased motivation and self-efficacy, and the opportunity to refine communication skills. These benefits are particularly valuable in data science education. Though not as widely studied, teaching assignments, where students demonstrate competency through teaching others can be used as a formative assessment tool. This paper describes an "alternative midterm" experiment conducted in an undergraduate introductory data science class. In this pilot, students were asked to demonstrate their newly learned data science skills by teaching them to individuals without data science backgrounds. The evidence of learning will be illustrated through an in-depth analysis of the students' teaching products. These products consist of two parts: (1) "Teaching Preparation" materials, such as Python Notebooks containing problems and solutions crafted by the students, and (2) "Teaching Implementation", video recordings of real-time interactions between the students and their "tutees." The paper will also present qualitative feedback from the students, highlighting promising signs of engagement and acceptance. We will conclude by discussing the challenges and opportunities of implementing this assessment strategy on a larger scale. Lujie Karen Chen, Justin Thai |
SIGCSE (2) | 1 |
| 2024 | Adopting Foundational Data Science Curriculum with Diverse Institutional ContextsabstractThe prevalence of data across all disciplines and the large workforce demand from industry has led to the rise in interest of data science courses. Educators are increasingly recognizing the value of building communities of practice and adapting and translating courses and programs that have been shown to be successful and sharing lessons learned in increasing diversity in data science education. We describe and analyze our experiences translating a lower-division data science curriculum from one university, University of California, Berkeley, to another setting with very different student populations and institutional context, University of Maryland, Baltimore County (UMBC). We present our findings from student interviews across two semesters of the course offering at UMBC specifically focusing on the challenges and positive experiences that the students had in the UMBC course. We highlight lessons learned to reflect on the existing large scale program at UC Berkeley, its adaptation and opportunities for increasing diversity in new settings. Our findings emphasize the importance of adapting courses and programs to existing curricula, student populations, cyberinfrastructure, and faculty and staff resources. Smaller class sizes open up the possibility of more individualized assignments, tailored to the majors, career interests, and social change motivations of diverse students. While students across institutional contexts may need varying degrees of support, we found that often students from diverse backgrounds, if engaged deeply, show significant enthusiasm for data science and its applications. Vandana Pursnani Janeja, Maria Sanchez, Yi Xuan Khoo, Claudia von Vacano, Lujie Karen Chen |
SIGCSE (1) | 5 |
| 2023 | Modeling Metacognitive and Cognitive Processes in Data Science Problem Solving (Student Abstract)abstractData Science (DS) is an interdisciplinary topic that is applicable to many domains. In this preliminary investigation, we use caselet, a mini-version of a case study, as a learning tool to allow students to practice data science problem solving (DSPS). Using a dataset collected from a real-world classroom, we performed correlation analysis to reveal the structure of cognition and metacognition processes. We also explored the similarity of different DS knowledge components based on students’ performance. In addition, we built a predictive model to characterize the relationship between metacognition, cognition, and learning gain. Maryam Alomair, Shimei Pan, Lujie Karen Chen |
AAAI | 3 |
| 2021 | Affect, Support and Personal Factors: Multimodal Causal Models of One-on-one Coaching
Lujie Karen Chen, Joseph D. Ramsey, Artur Dubrawski |
EDM | 1 |
| 2021 | Timing of Support in One-on-one Math Problem Solving Coaching: A Survival Analysis Approach with Multimodal DataabstractIn this paper, we explore a kind of teaching-oriented temporal analytics on the timing of support in the context of one-on-one math problem-solving coaching. We build the analytical framework upon the human-human multimodal interaction data collected from the naturalist environments. We demonstrated the potential utility of leveraging survival analysis, a class of statistical methods to model time-to-event data, to gain insights into the timing decisions. We shed light on the heterogeneity of coaching decisions as to when to render support in connection to the problem-solving stages, coaching dyads, as well as the pre-intervention event characteristics. This work opens future avenues into a different type of human tutoring study supported by multimodal data, computational models, and statistical frameworks. This model framework may yield useful reflective teaching analytics to tutors, coaches, or teachers when further developed. We also envision that those analyses may ultimately inform the design of AI-supported autonomous agents that could learn the tutorial interaction logic from empirical data. Lujie Karen Chen |
LAK | 1 |