Ruiwei Xiao

dblp:314/8067 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-6461-7611ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning to Use AI for Learning: Teaching Responsible Use of AI Chatbot to K-12 Students Through an AI Literacy Module
abstract
As Artificial Intelligence (AI) becomes increasingly integrated into daily life, there is a growing need to equip the next generation with the ability to apply, interact with, evaluate, and collaborate with AI systems responsibly. Prior research highlights the urgent demand from K-12 educators to teach students the ethical and effective use of AI for learning. To address this need, we designed a Large-Language Model (LLM)-based module to teach prompting literacy. This includes scenario-based deliberate practice activities with direct interaction with intelligent LLM agents, aiming to foster secondary school students' responsible engagement with AI chatbots. We conducted two iterations of classroom deployment in 11 authentic secondary education classrooms, and evaluated 1) AI-based auto-grader's capability; 2) students' prompting performance and confidence changes towards using AI for learning; and 3) the quality of learning and assessment materials. Results indicated that the AI-based auto-grader could grade student-written prompts with satisfactory quality. In addition, the instructional materials supported students in improving their prompting skills through practice and led to positive shifts in their perceptions of using AI for learning. Furthermore, data from Study 1 informed assessment revisions in Study 2. Analyses of item difficulty and discrimination in Study 2 showed that True/False and open-ended questions could measure prompting literacy more effectively than multiple-choice questions for our target learners. These promising outcomes highlight the potential for broader deployment and highlight the need for broader studies to assess learning effectiveness and assessment design.
Ruiwei Xiao, Xinying Hou, Ying-Jui Tseng, Hsuan Nieu, Guanze Liao, John C. Stamper, Kenneth R. Koedinger
AAAI1
2026 Enabling Multi-agent Systems as Learning Designers: Applying Learning Sciences to AI Instructional Design
Ruiwei Xiao, Xinying Hou, John C. Stamper
AIED (3)2
2026 Do Teachers Dream of GenAI Widening Educational (In)equality? Envisioning the Future of K-12 GenAI Education from Global Teachers' Perspectives
abstract
Generative artificial intelligence (GenAI) is rapidly entering K-12 classrooms worldwide, initiating urgent debates about its potential to either reduce or exacerbate educational inequalities. Drawing on interviews with 30 K-12 teachers across the United States, South Africa, and Taiwan, this study examines how teachers navigate this GenAI tension around educational equalities. We found teachers actively framed GenAI education as an equality-oriented practice: they used it to alleviate pre-existing inequalities while simultaneously working to prevent new inequalities from emerging. Despite these efforts, teachers confronted persistent systemic barriers, i.e., unequal infrastructure, insufficient professional training, and restrictive social norms, that individual initiative alone could not overcome. Teachers thus articulated normative visions for more inclusive GenAI education. By centering teachers’ practices, constraints, and future envisions, this study contributes a global account of how GenAI education is being integrated into K-12 contexts and highlights what is required to make its adoption genuinely equal.
Ruiwei Xiao, Qing Xiao 0002, Xinying Hou, Phenyo Phemelo Moletsane, Hanqi Jane Li, Hong Shen 0004, John C. Stamper
CHI1
2026 How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures
Shan Zhang 0003, Ruiwei Xiao, Anthony Botelho, Guanze Liao, Thomas K. F. Chiu, John C. Stamper, Kenneth R. Koedinger
LAK2
2026 Exploring Student Choice and the Use of Multimodal Generative AI in Programming Learning
abstract
The broad adoption of Generative AI (GenAI) is impacting Computer Science education, and recent studies found its benefits and potential concerns when students use it for programming learning. However, most existing explorations focus on GenAI tools that primarily support text-to-text interaction. With recent developments, GenAI applications have begun supporting multiple modes of communication, known as multimodality. In this work, we explored how undergraduate programming novices choose and work with multimodal GenAI tools, and their criteria for choices. We selected a commercially available multimodal GenAI platform for interaction, as it supports multiple input and output modalities, including text, audio, image upload, and real-time screen-sharing. Through 16 think-aloud sessions that combined participant observation with follow-up semi-structured interviews, we investigated student modality choices for GenAI tools when completing programming problems and the underlying criteria for modality selections. With multimodal communication emerging as the future of AI in education, this work aims to spark continued exploration on understanding student interaction with multimodal GenAI in the context of CS education.
Xinying Hou, Ruiwei Xiao, Runlong Ye 0002, Michael Liut, John C. Stamper
SIGCSE (1)2
2025 "From Unseen Needs to Classroom Solutions": Exploring AI Literacy Challenges & Opportunities with Project-Based Learning Toolkit in K-12 Education
abstract
As artificial intelligence (AI) becomes increasingly central to various fields, there is a growing need to equip K-12 students with AI literacy skills that extend beyond computer science. This paper explores the integration of a Project-Based Learning (PBL) AI toolkit into diverse subject areas, aimed at helping educators teach AI concepts more effectively. Through interviews and co-design sessions with K-12 teachers, we examined current AI literacy levels and how teachers adapt AI tools like the AI Art Lab, AI Music Studio, and AI Chatbot into their course designs. While teachers appreciated the potential of AI tools to foster creativity and critical thinking, they also expressed concerns about the accuracy, trustworthiness, and ethical implications of AI-generated content. Our findings reveal the challenges teachers face, including limited resources, varying student and instructor skill levels, and the need for scalable, adaptable AI tools. This research contributes insights that can inform the development of AI curricula tailored to diverse educational contexts.
Ruiwei Xiao, Hsuan Nieu, Ying-Jui Tseng, Guanze Liao
AAAI2
2025 Generating AI Literacy MCQs: A Multi-Agent LLM Approach
abstract
Artificial intelligence (AI) is transforming society, making it crucial to prepare the next generation through AI literacy in K-12 education. However, scalable and reliable AI literacy materials and assessment resources are lacking. To address this gap, our study presents a novel approach to generating multiple-choice questions (MCQs) for AI literacy assessments. Our method utilizes large language models (LLMs) to automatically generate scalable, high-quality assessment questions. These questions align with user-provided learning objectives, grade levels, and Bloom's Taxonomy levels. We introduce an iterative workflow incorporating LLM-powered critique agents to ensure the generated questions meet pedagogical standards. In the preliminary evaluation, experts expressed strong interest in using the LLM-generated MCQs, indicating that this system could enrich existing AI literacy materials and provide a valuable addition to the toolkit of K-12 educators.
Ruiwei Xiao, Ying-Jui Tseng
SIGCSE (2)2
2024 Supporting Self-Reflection at Scale with Large Language Models: Insights from Randomized Field Experiments in Classrooms
abstract
Self-reflection on learning experiences constitutes a fundamental cognitive process, essential for consolidating knowledge and enhancing learning efficacy. However, traditional methods to facilitate reflection often face challenges in personalization, immediacy of feedback, engagement, and scalability. Integration of Large Language Models (LLMs) into the reflection process could mitigate these limitations. In this paper, we conducted two randomized field experiments in undergraduate computer science courses to investigate the potential of LLMs to help students engage in post-lesson reflection. In the first experiment (N=145), students completed a take-home assignment with the support of an LLM assistant; half of these students were then provided access to an LLM designed to facilitate self-reflection. The results indicated that the students assigned to LLM-guided reflection reported somewhat increased self-confidence compared to peers in a no-reflection control and a non-significant trend towards higher scores on a later assessment. Thematic analysis of students' interactions with the LLM showed that the LLM often affirmed the student's understanding, expanded on the student's reflection, and prompted additional reflection; these behaviors suggest ways LLM-interaction might facilitate reflection. In the second experiment (N=112), we evaluated the impact of LLM-guided self-reflection against other scalable reflection methods, such as questionnaire-based activities and review of key lecture slides, after assignment. Our findings suggest that the students in the questionnaire and LLM-based reflection groups performed equally well and better than those who were only exposed to lecture slides, according to their scores on a proctored exam two weeks later on the same subject matter. These results underscore the utility of LLM-guided reflection and questionnaire-based activities in improving learning outcomes. Our work highlights that focusing solely on the accuracy of LLMs can overlook their potential to enhance metacognitive skills through practices such as self-reflection. We discuss the implications of our research for the learning-at-scale community, highlighting the potential of LLMs to enhance learning experiences through personalized, engaging, and scalable reflection practices.
Ruiwei Xiao, Benjamin Lawson, Ilya Musabirov, Jiakai Shi, Huayin Luo, Joseph Jay Williams, Anna N. Rafferty, John C. Stamper, Michael Liut
L@S2
2024 Assessing the Efficacy of Goal-Based Scenarios in Scaling AI Literacy for Non-Technical Learners
abstract
AI's pervasive role in various fields highlights the imperative for the workforce to adeptly leverage its potential. While numerous courses cater to developers, there exists a discernible void for the wider community of AI users. To address this, our study introduces 'AI User'-a suite of interactive modules hosted on the Sail() platform, designed specifically for non-technical individuals utilizing Goal-Based Scenario (GBS) learning. We conducted a controlled experiment to ascertain whether GBS offers superior learning gains in AI literacy compared to traditional deliberate practice using multiple choice questions.
Ying-Jui Tseng, Ruiwei Xiao, Christopher Bogart, Jaromír Savelka, Majd F. Sakr
SIGCSE (2)2
2023 Detecting Cheating in Online Take-Home Exams with Randomized Questions
abstract
The last three years were a significant challenge for educational institutions, due to the loss of face-to-face instruction and exam proctoring. Many instructors turned to asynchronous, online exams as a replacement for standard pen-and-paper exams. It is no surprise that many tools aimed at delivering computer-based assessments have become popular and are centers of research and development. This poster discusses our attempt to build a post-exam cheating detection system for the PrairieLearn open-source platform that supports randomized question generators, to uncover irregularities in submissions. Our system compares all pairs of students using four rules: Times (did students take the exam synchronously?), Answers (did they have similar wrong answers?), Orders (did they answer the questions in the same order?), and Scores (did they achieve the same scores?). It adds one final individual rule, the Score-Time-Ratio, that measures how many "points per minute" a student has earned, to flag students who open the exam, copy in a perfect answer, and submit. We deliver a detailed report to the instructor, allowing them to sort their students based on these measures, providing a data-driven way for them to investigate.
Ruiwei Xiao, Eduardo Huerta-Mercado, Dan Garcia 0001
SIGCSE (2)1
2022 Improved Testing of PrairieLearn Question Generators
abstract
With many institutions forced online due to the pandemic, assessments became a challenge for many educators. Take-home exams provided the flexibility required for varied student needs (and time zones), but they were vulnerable to cheating. In response, many turned to tools that could present a different exam for every student. PrairieLearn is a feature-rich open-source package that allows educators to author randomized Question Generators; we have been using the tool extensively for the last two years, and it has a fast-growing educator user base. One of the first issues we noticed with the system was that the only way to quality assure (QA) a question was to click the new variant button, which would spin whatever internal random number generators were used again to produce a new question. Sometimes it was the same one you had just seen, and other times it would never seem to "hit'' on the variant you were looking to debug. This poster describes our team's work to solve this problem through the design of an API that would allow a question to declare how many total variants it had, and be asked to render variant i. The user interface could then be extended to list what variant the QA team was viewing out of the total (e.g., 7/50), and a next, previous and go to a particular variant buttons would allow for the team to easily QA all variants.
Aayush Shah, Alan Lee, Chris Chi, Ruiwei Xiao, Pranav Sukumar, Jesus Villalobos, Dan Garcia 0001
SIGCSE (2)4