Matthew J. Davidson

dblp:236/5411 · also Matt J. Davidson · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0001-6848-2876ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 ContextQ: Generated Questions to Support Meaningful Parent-Child Dialogue While Co-Reading
abstract
Much of early literacy education happens at home with caretakers reading books to young children. Prior research demonstrates how having dialogue with children during co-reading can develop critical reading readiness skills, but most adult readers are unsure if and how to lead effective conversations. We present ContextQ, a tablet-based reading application to unobtrusively present auto-generated dialogic questions to caretakers to support this dialogic reading practice. An ablation study demonstrates how our method of encoding educator expertise into the question generation pipeline can produce high-quality output; and through a user study with 12 parent-child dyads (child age: 4–6), we demonstrate that this system can serve as a guide for parents in leading contextually meaningful dialogue, leading to significantly more conversational turns from both the parent and the child and deeper conversations with connections to the child’s everyday life.
Griffin Dietz, Siddhartha Prasad, Matthew J. Davidson, Leah Findlater, R. Benjamin Shapiro
IDC3
2021 Towards an Understanding of Program Writing as a Cognitive Process: Analysis of Keystroke Logs
abstract
Program writing is difficult to teach, learn, and assess. One challenge is a lack of theory or understanding about what program writing is. My dissertation will address this challenge by applying theories from natural language (NL) writing to try and understand program writing as a cognitive process. By analyzing keystroke logs collected during program writing, I plan to identify similarities with NL writing, potential diagnostic information, and how program writing changes as students become more proficient.
Matthew J. Davidson
ICER1
2021 Domain Experts' Interpretations of Assessment Bias in a Scaled, Online Computer Science Curriculum
abstract
Understanding inequity at scale is necessary for designing equitable online learning experiences, but also difficult. Statistical techniques like differential item functioning (DIF) can help identify whether items/questions in an assessment exhibit potential bias by disadvantaging certain groups (e.g. whether item disadvantages woman vs man of equivalent knowledge). While testing companies typically use DIF to identify items to remove, we explored how domain-experts such as curriculum designers could use DIF to better understand how to design instructional materials to better serve students from diverse groups. Using Code.org's online Computer Science Discoveries (CSD) curriculum, we analyzed 139,097 responses from 19,617 students to identify DIF by gender and race in assessment items (e.g. multiple choice questions). Of the 17 items, we identified six that disadvantaged students who reported as female when compared to students who reported as non-binary or male. We also identified that most (13) items disadvantaged AHNP (African/Black, Hispanic/Latinx, Native American/Alaskan Native, Pacific Islander) students compared to WA (white, Asian) students. We then conducted a workshop and interviews with seven curriculum designers and found that they interpreted item bias relative to an intersection of item features and student identity, the broader curriculum, and differing uses for assessments. We interpreted these findings in the broader context of using data on assessment bias to inform domain-experts' efforts to design more equitable learning experiences.
Benjamin Xie, Matthew J. Davidson, Baker Franke, Emily McLeod, Min Li 0090, Amy J. Ko
L@S2
2021 Investigating Item Bias in a CS1 Exam with Differential Item Functioning
abstract
Reliable and valid exams are a crucial part of both sound research design and trustworthy assessment of student knowledge. Assessing and addressing item bias is a crucial step in building a validity argument for any assessment instrument. Despite calls for valid assessment tools in CS, item bias is rarely investigated. What kinds of item bias might appear in conventional CS1 exams? To investigate this, we examined responses to a final exam in a large CS1 course. We used differential item functioning (DIF) methods and specifically investigated bias related to binary gender and year of study. Although not a published assessment instrument, the exam had a similar format to many exams in higher education and research: students are asked to trace code and write programs, using paper and pencil. One item with significant DIF was detected on the exam, though the magnitude was negligible. This case study shows how to detect DIF items so that future researchers and practitioners can do these analyses.
Matthew J. Davidson, Brett Wortzman, Amy J. Ko, Min Li 0090
SIGCSE1
2019 An Item Response Theory Evaluation of a Language-Independent CS1 Knowledge Assessment
abstract
Tests serve an important role in computing education, measuring achievement and differentiating between learners with varying knowledge. But tests may have flaws that confuse learners or may be too difficult or easy, making test scores less valid and reliable. We analyzed the Second Computer Science 1 (SCS1) concept inventory, a widely used assessment of introductory computer science (CS1) knowledge, for such flaws. The prior validation study of the SCS1 used Classical Test Theory and was unable to determine whether differences in scores were a result of question properties or learner knowledge. We extended this validation by modeling question difficulty and learner knowledge separately with Item Response Theory (IRT) and performing expert review on problematic questions. We found that three questions measured knowledge that was unrelated to the rest of the SCS1, and four questions were too difficult for our sample of 489 undergrads from two universities.
Benjamin Xie, Matthew J. Davidson, Min Li 0090, Amy J. Ko
SIGCSE2