VLDB 2026 Research / reviewers in the wild / expert
Benjamin Xie
dblp:170/7526
· DBLP profile ↗
16ranked-venue papers
10as first author
9since 2021 · last 2024
0000-0003-3275-992XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 12 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | From Consumers to Critical Users: Prompty, an AI Literacy Tool for High School StudentsabstractIn an age where Large Language Models (LLMs) expedite the generation of text, the skills for critically evaluating and creating meaningful text using these models are often lacking. To help classroom teachers address this, we introduce Prompty, a specialized teaching tool co-designed to facilitate both critical and effective use of LLMs. Prompty serves multiple learning goals: it allows students to critically evaluate text generated by LLMs, aids in their writing practice, and provides a deeper understanding of how LLMs function—all within a student-friendly environment secured by essential guardrails. Prompty was co-designed in collaboration with high school teachers as part of CRAFT, an initiative by Stanford University to promote AI literacy. It was pilot-tested in a high school English class to serve as an AI writing assistant, focusing on the critical evaluation of machine-generated text. This trial yielded preliminary evidence that attests to the tool's effectiveness in fulfilling its educational goals. The findings from the pilot study indicate that easy-to-use tools like Prompty have great potential. These tools can be adapted to fit the goals of individual teachers. They can help in achieving subject-specific learning goals while serving as an effective way to teach AI concepts in high school. Deepak Varuvel Dennison, Raycelle C. C. Garcia, Parth Sarin, Jacob Wolf, Christine Bywater, Benjamin Xie, Victor R. Lee |
AAAI | 6 |
| 2024 | Co-designing AI Education Curriculum with Cross-Disciplinary High School TeachersabstractHigh school teachers from many disciplines have growing interests in teaching about artificial intelligence (AI). This cross-disciplinary interest reflects the prevalence of AI tools across society, such as Generative AI tools built upon Large Language Models (LLM). However, high school classes are unique and complex environments, led by teachers with limited time and resources with priorities that vary by class and the students they serve. Therefore, developing curricula about AI for classes that span many disciplines (e.g. history, art, math) must involve centering the expertise of cross-disciplinary teachers. In this study, we conducted five collaborative curricular co-design sessions with eight teachers who taught high school humanities and STEM classes. We sought to understand how teachers considered AI when it was taught in art, math, and social studies contexts, as well as opportunities and challenges they identified with incorporating AI tools into their instruction. We found that teachers considered technical skills and ethical debates around AI, opportunities for "dual exploration" between AI and disciplinary learning, and limitations of AI tools as supporting engagement and reflection but also potentially distracting. We interpreted our findings relative to co-designing adaptable AI curricula to support teaching about and with AI across high school disciplines. Benjamin Xie, Parth Sarin, Jacob Wolf, Raycelle C. C. Garcia, Victoria Delaney, Isabel Sieh, Anika Fuloria, Deepak Varuvel Dennison, Christine Bywater, Victor R. Lee |
AAAI | 1 |
| 2024 | Using Benchmarking Infrastructure to Evaluate LLM Performance on CS Concept Inventories: Challenges, Opportunities, and CritiquesabstractBACKGROUND AND CONTEXT. The pace of advancement of large language models (LLMs) motivates the use of existing infrastructure to automate the evaluation of LLM performance on computing education tasks. Concept inventories are well suited for evaluation because of their careful design and prior validity evidence. Prerna Rao, Yifan Mai 0001, Benjamin Xie |
ICER (1) | 4 |
| 2023 | Developing Novice Programmers' Self-Regulation Skills with Code ReplaysabstractLearning programming benefits from self-regulation, but novices lack support for developing these skills of cognitive control. To support their development, we designed Code Replayer, an online tool that enables novice programmers to practice programming and then replay their coding process to reflect and identify process improvements. To evaluate the impact of replaying code on self-regulation, we conducted a formative qualitative evaluation with 21 novice programmers who used Code Replayer to practice writing code. We found that after watching code replays, participants more frequently interpreted problem prompts and planned their solutions, two crucial self-regulation behaviors that novices often overlook. We interpret our results by focusing on two focal points in the design of code replays as a programming self-regulation intervention: interpreting pauses in replays and ensuring replays of struggle are more informative and less detrimental. Benjamin Xie, Jared Ordona Lim, Paul K. D. Pham, Min Li 0090, Amy J. Ko |
ICER (1) | 1 |
| 2023 | Centering Environmental Justice in Computing EducationabstractIn this Birds of a Feather, we will discuss the roles of computing education in preparing students to understand and address the disparate impacts of climate change in local and global contexts. We intend to have open discussions on the challenges and opportunities related to connecting computing education with climate change and the injustices that climate change exacerbates. We will focus discussions around three questions: (1) How can we center justice-based perspectives on understanding and addressing climate change in computing education? (2) What are the relationships between computing, climate change, and overall environmental impacts? and (3) How should we reimagine traditional notions of "development" and "progress" in computing in ways that challenge how current framings within computing misunderstand or misteach computing's environmental impacts? We invite all computing educators, researchers, administrators, practitioners and anyone else with any level of curiosity about climate change to join this discussion (because it affects all of us!). Expected outcomes for this Birds of a Feather include sharing resources and experiences to build a community for knowledge sharing and collaborations. Benjamin Xie, Greg L. Nelson, Francisco Enrique Vicente Castro, Nicholas Lytle, Briana Bettin |
SIGCSE (2) | 1 |
| 2022 | A Decade of Demographics in Computing Education Research: A Critical Review of Trends in Collection, Reporting, and UseabstractComputing education research (CER) has used demographic data to understand learners’ identities, backgrounds, and contexts for efforts such as culturally-responsive computing. Prior work indicates that failing to elucidate and critically engage with the implicit assumptions of a field can unintentionally reinforce power structures that further marginalize people from non-dominant groups. The goal of this paper is two-fold: to understand what populations CER researchers have studied, and to surface implicit assumptions about how researchers have collected, reported, and used demographic data on these populations. We conducted a content analysis of 510 peer-reviewed papers published in 12 CER venues from 2012 to 2021. We found that (1) 60% of papers studied older learners in formal contexts (i.e. post-secondary education); (2) 68% of papers left unclear how researchers collected demographic data; and (3) while 94% of papers were single-site studies, only 14% addressed the limitations of their contexts. We also identified hegemonic norms through ambiguous aggregate term usage (e.g. underrepresented, diverse) in 23% of papers, and through incomplete reporting of demographics (i.e. leaving out demographics for some participants in their sample) in 35% of papers. We discuss the implications of these findings for the CER field, raising considerations for CER researchers to keep in mind when collecting, reporting, and using demographic data. Alannah Oleson, Benjamin Xie, Jean Salac, Jayne Everson, F. Megumi Kivuva, Amy J. Ko |
ICER (1) | 2 |
| 2022 | Surfacing Equity Issues in Large Computing Courses with Peer-Ranked, Demographically-Labeled Student FeedbackabstractAs computing courses become larger, students of minoritized groups continue to disproportionately face challenges that hinder their academic and professional success (e.g. implicit bias, microaggressions, lack of resources, assumptions of preparatory privilege). This can impact career aspirations and sense of belonging in computing communities. Instructors have the power to make immediate changes to support more equitable learning, but they are often unaware of students' challenges. To help both instructors and students understand the inequities in their classes, we developed StudentAmp, an interactive system that uses student feedback and self-reported demographic information (e.g. gender, ethnicity, disability, educational background) to show challenges and how they affect students differently. To help instructors make sense of feedback, StudentAmp ranks challenges by student-perceived disruptiveness. We conducted formative evaluations with five large college computing courses (150 - 750 students) being taught remotely during the COVID-19 pandemic. We found that students shared challenges beyond the scope of the course, perceived sharing information about who they were as useful but potentially dangerous, and that teaching teams were able to use this information to consider the positionality of students sharing challenges. Our findings relate to a central design tension of supporting equity by sharing contextualized information about students while also ensuring their privacy and well-being. Benjamin Xie, Alannah Oleson, Jayne Everson, Amy J. Ko |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2021 | Interpretations and Uses of Data for Equity in Computing EducationabstractComputing education’s booming enrollment exacerbates inclusion challenges ranging from tools that do not support diverse learners to instructors not being aware of unique challenges that students of minoritized groups face. While data often perpetuates inequities in many contexts, it could also serve to support equity-related goals if properly contextualized. To understand how data could support equitable learning, I explore how affording information and agency supports students’ self-directed learning of Python programming, how contextualizing psychometric data on test bias with curriculum designers’ domain expertise could support equitable curriculum improvements, and how contextualizing student feedback with demographic information and peer perspectives could help instructors become aware of challenges that students from minoritized groups face while preserving student privacy and well-being. By studying how students, curriculum designers, and teachers interpreted and used data relating to experiences learning computing, I contribute techniques that contextualize equity-oriented interpretations and uses of data with stakeholders’ domain expertise. Benjamin Xie |
ICER | 1 |
| 2021 | Domain Experts' Interpretations of Assessment Bias in a Scaled, Online Computer Science CurriculumabstractUnderstanding inequity at scale is necessary for designing equitable online learning experiences, but also difficult. Statistical techniques like differential item functioning (DIF) can help identify whether items/questions in an assessment exhibit potential bias by disadvantaging certain groups (e.g. whether item disadvantages woman vs man of equivalent knowledge). While testing companies typically use DIF to identify items to remove, we explored how domain-experts such as curriculum designers could use DIF to better understand how to design instructional materials to better serve students from diverse groups. Using Code.org's online Computer Science Discoveries (CSD) curriculum, we analyzed 139,097 responses from 19,617 students to identify DIF by gender and race in assessment items (e.g. multiple choice questions). Of the 17 items, we identified six that disadvantaged students who reported as female when compared to students who reported as non-binary or male. We also identified that most (13) items disadvantaged AHNP (African/Black, Hispanic/Latinx, Native American/Alaskan Native, Pacific Islander) students compared to WA (white, Asian) students. We then conducted a workshop and interviews with seven curriculum designers and found that they interpreted item bias relative to an intersection of item features and student identity, the broader curriculum, and differing uses for assessments. We interpreted these findings in the broader context of using data on assessment bias to inform domain-experts' efforts to design more equitable learning experiences. Benjamin Xie, Matthew J. Davidson, Baker Franke, Emily McLeod, Min Li 0090, Amy J. Ko |
L@S | 1 |
| 2020 | The Effect of Informing Agency in Self-Directed Online Learning EnvironmentsabstractChoices learners make when navigating a self-directed online learning tool can impact the effectiveness of the experience. But these tools often do not afford learners the agency or the information to make decisions beneficial to their learning. We evaluated the effect of varying levels of information and agency in a self-directed environment designed to teach programming. We investigated three design alternatives: informed high-agency, informed low-agency, and less informed high-agency. To investigate the effect of these alternatives on learning, we conducted a study with 79 novice programmers. Our results indicated that increased agency and information may have translated to more motivation, but not improved learning. Qualitative results suggest this was due to the burden that agency and information placed on decision-making. We interpret our results in relation to informing the design of self-directed online tools for learner agency. Benjamin Xie, Greg L. Nelson, Harshitha Akkaraju, William Kwok, Amy J. Ko |
L@S | 1 |
| 2020 | Investigating Novices' In Situ Reflections on Their Programming ProcessabstractPrior work on novice programmers' self-regulation have shown it to be inconsistent and shallow, but trainable through direct instruction. However, prior work has primarily studied self-regulation retrospectively, which relies on students to remember how they regulated their process, or in laboratory settings, limiting the ecological validity of findings. To address these limitations, we investigated 31 novice programmers' self-regulation in situ over 10 weeks. We had them to keep journals about their work and later had them to reflect on their journaling. Through a series of qualitative analyses of journals and survey responses, we found that all participants monitored their process and evaluated their work, that few interpreted the problems they were solving or adapted prior solutions. We also found that some students self-regulated their programming in many ways, while others in almost none. Students reported many difficulties integrating reflection into their work; some were completely unaware of their process, some struggled to integrate reflection into their process, and others found reflection conflicted with their work. These results suggest that self-regulation during programming is highly variable in practice, and that teaching self-regulation skills to improve programming outcomes may require differentiated instruction based on students self-awareness and existing programming practices. Dastyni Loksa, Benjamin Xie, Harrison Kwik, Amy J. Ko |
SIGCSE | 2 |
| 2019 | An Item Response Theory Evaluation of a Language-Independent CS1 Knowledge AssessmentabstractTests serve an important role in computing education, measuring achievement and differentiating between learners with varying knowledge. But tests may have flaws that confuse learners or may be too difficult or easy, making test scores less valid and reliable. We analyzed the Second Computer Science 1 (SCS1) concept inventory, a widely used assessment of introductory computer science (CS1) knowledge, for such flaws. The prior validation study of the SCS1 used Classical Test Theory and was unable to determine whether differences in scores were a result of question properties or learner knowledge. We extended this validation by modeling question difficulty and learner knowledge separately with Item Response Theory (IRT) and performing expert review on problematic questions. We found that three questions measured knowledge that was unrelated to the rest of the SCS1, and four questions were too difficult for our sample of 489 undergrads from two universities. Benjamin Xie, Matthew J. Davidson, Min Li 0090, Amy J. Ko |
SIGCSE | 1 |
| 2018 | Experiences of Computer Science Transfer StudentsabstractAbout half of recent computer and information science graduates attended community college at some point. Prior work on transfer students in general suggests that the transfer process can engage people from underrepresented communities, but can also be academically and socially "shocking". However, we know little about the experiences of transfer students in computer science in particular. We used the Laanan-Transfer Student Questionnaire (L-TSQ) to survey 25 transfer students and 135 native (non-transfer) students and conducted follow-up interviews with 8 transfer students attending a large public 4-year university in a city with significant technology industry presence. We found that while transfer students were more diverse demographically, the support of the university for transfer student orientation tended to mitigate social shocks of transferring. This did not, however, eliminate gaps in academic performance. These findings suggest that there are other non-social factors that influence academic performance that CS programs must support to equitably engage students who transfer. Harrison Kwik, Benjamin Xie, Amy J. Ko |
ICER | 2 |
| 2018 | An Explicit Strategy to Scaffold Novice Program TracingabstractWe propose and evaluate a lightweight strategy for tracing code that can be efficiently taught to novice programmers, building off of recent findings on "sketching" when tracing. This strategy helps novices apply the syntactic and semantic knowledge they are learning by encouraging line-by-line tracing and providing an external representation of memory for them to update. To evaluate the effect of teaching this strategy, we conducted a block-randomized experiment with 24 novices enrolled in a university-level CS1 course. We spent only 5-10 minutes introducing the strategy to the experimental condition. We then asked both conditions to think-aloud as they predicted the output of short programs. Students using this strategy scored on average 15% higher than students in the control group for the tracing problems used the study (p<0.05). Qualitative analysis of think-aloud and interview data showed that tracing systematically (line-by-line and "sketching" intermediate values) led to better performance and that the strategy scaffolded and encouraged systematic tracing. Students who learned the strategy also scored on average 7% higher on the course midterm. These findings suggest that in <1 hour and without computer-based tools, we can improve CS1 students' tracing abilities by explicitly teaching a strategy. Benjamin Xie, Greg L. Nelson, Amy J. Ko |
SIGCSE | 1 |
| 2017 | Comprehension First: Evaluating a Novel Pedagogy and Tutoring System for Program Tracing in CS1abstractWhat knowledge does learning programming require? Prior work has focused on theorizing program writing and problem solving skills. We examine program comprehension and propose a formal theory of program tracing knowledge based on control flow paths through an interpreter program's source code. Because novices cannot understand the interpreter's programming language notation, we transform it into causal relationships from code tokens to instructions to machine state changes. To teach this knowledge, we propose a comprehension-first pedagogy based on causal inference, by showing, explaining, and assessing each path by stepping through concrete examples within many example programs. To assess this pedagogy, we built PLTutor, a tutorial system with a fixed curriculum of example programs. We evaluate learning gains among self-selected CS1 students using a block randomized lab study comparing PLTutor with Codecademy, a writing tutorial. In our small study, we find some evidence of improved learning gains on the SCS1, with average learning gains of PLTutor 60% higher than Codecademy (gain of 3.89 vs. 2.42 out of 27 questions). These gains strongly predicted midterms (R2=.64) only for PLTutor participants, whose grades showed less variation and no failures. Greg L. Nelson, Benjamin Xie, Amy J. Ko |
ICER | 2 |
| 2016 | Skill progression in MIT app inventorabstractThis paper contributes to the growing body of research that attempts to measure online, informal learning. We analyze skill progression in MIT App Inventor, an informal online learning environment with over 5 million users and 15.9 million projects/apps created. Our objective is to understand how people learn computational thinking concepts while creating mobile applications with App Inventor. In particular, we are interested in the relationship between the progression of skill in using App Inventor functionality and in using computational thinking concepts as learners create more apps. We model skill progression along two dimensions: breadth and depth of capability. Given a sample of 10,571 random users who have each created at least 20 apps, we analyze the relationship between demonstrating domain-specific skills by using App Inventor functionality and generalizable skills by using computational thinking concepts. Our findings indicate that domain-specific and generalizable skills progress similarly; there is a common pattern of expanding breadth of capability by using new skills over the first 10 projects, then developing depth of capability by using previously introduced skills to build more sophisticated apps. Benjamin Xie, Harold Abelson |
VL/HCC | 1 |