Min Li 0090

dblp:82/0-90 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0008-4523-0391ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 6 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Development of a Versatile Assessment for CS1
abstract
In this poster, we present an ongoing project to develop and validate a subscale-based introductory computer science (CS1) assessment for broad use in computer science education research. This project addresses the need for a validated, adaptable assessment of CS1 knowledge across diverse institutional contexts and programming languages. We build upon the existing Second CS1 (SCS1) assessment, a pseudocode-based assessment for undergraduate introductory computer science course. While SCS1 has been widely used, it lacks flexibility and subscale validation, limiting its suitability for varied course content and research needs. Our poster will present our initial work towards creating an enhanced version of SCS1, termed SCS1++. We report on our progress developing and piloting a broader range of questions with varying difficulty levels. We will also use the poster as an opportunity to recruit field test sites for testing and validating SCS1++ with a broad student population.
Miranda C. Parker, Lolo Aboufoul, Braxton Haight, Min Li 0090
SIGCSE (2)5
2026 Investigating Answer Choice Bias within a College-Level Introductory Computing Assessment
abstract
Assessments of student learning can take many forms, but for assessing learning at scale, a multiple-choice exam is often used. A multiple-choice question is comprised of a stem, followed by a number of distractors and a correct option. All parts of a question need to be carefully designed to support assessment validity, but each part can also exhibit bias and negatively impact students in an inequitable way. In this research, we extend psychometric methods to explore the bias within the distractors of multiple choice questions on an introductory computing assessment for undergraduate students. We use Differential Distractor Functioning (DDF) on 259 student responses to identify problematic distractors. We discuss the distractors that were flagged in our analysis, which did vary in regards to biasing for male or female students. This work contributes a deeper understanding of the issues that can lie within an assessment, furthering our efforts to create fairer measures of learning for all students.
Miranda C. Parker, Sin Yu Ciou, Yale Quan, Min Li 0090
SIGCSE (1)6
2024 Intersectional Biases Within an Introductory Computing Assessment
abstract
Assessments that can measure student understanding of concepts in a reliable and valid way are incredibly valuable in research. Unfortunately, assessments can be a source of bias, differentially impacting students along various demographic lines. Differential Item Functioning (DIF) is a method to explore assessment bias. However, DIF is primarily limited to a single binary demographic variable (e.g. white and non-white; male and female). In this paper, we describe a novel expansion of DIF methods to explore intersections of student identities. We demonstrate the use of classic DIF on a data set of 255 complete responses to a CS1 assessment using binary race and gender variables in our analyses. Then, we present the importance of intersectional DIF by running a similar analysis on intersectional data. Using these methods, we identify problematic items on the assessment that are biased against certain groups of test-takers. Our work contributes an innovative method to help interpret assessment results and inform changes to assessments.
Miranda C. Parker, Min Li 0090
SIGCSE (1)3
2023 Developing Novice Programmers' Self-Regulation Skills with Code Replays
abstract
Learning programming benefits from self-regulation, but novices lack support for developing these skills of cognitive control. To support their development, we designed Code Replayer, an online tool that enables novice programmers to practice programming and then replay their coding process to reflect and identify process improvements. To evaluate the impact of replaying code on self-regulation, we conducted a formative qualitative evaluation with 21 novice programmers who used Code Replayer to practice writing code. We found that after watching code replays, participants more frequently interpreted problem prompts and planned their solutions, two crucial self-regulation behaviors that novices often overlook. We interpret our results by focusing on two focal points in the design of code replays as a programming self-regulation intervention: interpreting pauses in replays and ensuring replays of struggle are more informative and less detrimental.
Benjamin Xie, Jared Ordona Lim, Paul K. D. Pham, Min Li 0090, Amy J. Ko
ICER (1)4
2021 Domain Experts' Interpretations of Assessment Bias in a Scaled, Online Computer Science Curriculum
abstract
Understanding inequity at scale is necessary for designing equitable online learning experiences, but also difficult. Statistical techniques like differential item functioning (DIF) can help identify whether items/questions in an assessment exhibit potential bias by disadvantaging certain groups (e.g. whether item disadvantages woman vs man of equivalent knowledge). While testing companies typically use DIF to identify items to remove, we explored how domain-experts such as curriculum designers could use DIF to better understand how to design instructional materials to better serve students from diverse groups. Using Code.org's online Computer Science Discoveries (CSD) curriculum, we analyzed 139,097 responses from 19,617 students to identify DIF by gender and race in assessment items (e.g. multiple choice questions). Of the 17 items, we identified six that disadvantaged students who reported as female when compared to students who reported as non-binary or male. We also identified that most (13) items disadvantaged AHNP (African/Black, Hispanic/Latinx, Native American/Alaskan Native, Pacific Islander) students compared to WA (white, Asian) students. We then conducted a workshop and interviews with seven curriculum designers and found that they interpreted item bias relative to an intersection of item features and student identity, the broader curriculum, and differing uses for assessments. We interpreted these findings in the broader context of using data on assessment bias to inform domain-experts' efforts to design more equitable learning experiences.
Benjamin Xie, Matthew J. Davidson, Baker Franke, Emily McLeod, Min Li 0090, Amy J. Ko
L@S5
2021 Investigating Item Bias in a CS1 Exam with Differential Item Functioning
abstract
Reliable and valid exams are a crucial part of both sound research design and trustworthy assessment of student knowledge. Assessing and addressing item bias is a crucial step in building a validity argument for any assessment instrument. Despite calls for valid assessment tools in CS, item bias is rarely investigated. What kinds of item bias might appear in conventional CS1 exams? To investigate this, we examined responses to a final exam in a large CS1 course. We used differential item functioning (DIF) methods and specifically investigated bias related to binary gender and year of study. Although not a published assessment instrument, the exam had a similar format to many exams in higher education and research: students are asked to trace code and write programs, using paper and pencil. One item with significant DIF was detected on the exam, though the magnitude was negligible. This case study shows how to detect DIF items so that future researchers and practitioners can do these analyses.
Matthew J. Davidson, Brett Wortzman, Amy J. Ko, Min Li 0090
SIGCSE4
2019 An Item Response Theory Evaluation of a Language-Independent CS1 Knowledge Assessment
abstract
Tests serve an important role in computing education, measuring achievement and differentiating between learners with varying knowledge. But tests may have flaws that confuse learners or may be too difficult or easy, making test scores less valid and reliable. We analyzed the Second Computer Science 1 (SCS1) concept inventory, a widely used assessment of introductory computer science (CS1) knowledge, for such flaws. The prior validation study of the SCS1 used Classical Test Theory and was unable to determine whether differences in scores were a result of question properties or learner knowledge. We extended this validation by modeling question difficulty and learner knowledge separately with Item Response Theory (IRT) and performing expert review on problematic questions. We found that three questions measured knowledge that was unrelated to the rest of the SCS1, and four questions were too difficult for our sample of 489 undergrads from two universities.
Benjamin Xie, Matthew J. Davidson, Min Li 0090, Amy J. Ko
SIGCSE3