VLDB 2026 Research / reviewers in the wild / expert
Erin Gatz
dblp:365/4574
· DBLP profile ↗
10ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0002-6880-5740ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Open-Response Assessment with LearnLM
Danielle R. Thomas, Conrad Borchers, Shambhavi Bhushan, Sanjit Kakarla, Alex Houk, Ralph Abboud, Shivang Gupta, Erin Gatz, Kenneth R. Koedinger |
AIED (5) | 8 |
| 2025 | Leveraging LLMs to Assess Tutor Moves in Real-Life Dialogues: A Feasibility Study
Danielle R. Thomas, Conrad Borchers, Jionghao Lin, Sanjit Kakarla, Shambhavi Bhushan, Erin Gatz, Shivang Gupta, Ralph Abboud, Kenneth R. Koedinger |
EC-TEL (2) | 6 |
| 2025 | Does Multiple Choice Have a Future in the Age of Generative AI? A Posttest-only RCTabstractThe role of multiple-choice questions (MCQs) as effective learning tools has been debated in past research. While MCQs are widely used due to their ease in grading, open response questions are increasingly used for instruction, given advances in large language models (LLMs) for automated grading. This study evaluates MCQs effectiveness relative to open-response questions, both individually and in combination, on learning. These activities are embedded within six tutor lessons on advocacy. Using a posttest-only randomized control design, we compare the performance of 234 tutors (790 lesson completions) across three conditions: MCQ only, open response only, and a combination of both. We find no significant learning differences across conditions at posttest, but tutors in the MCQ condition took significantly less time to complete instruction. These findings suggest that MCQs are as effective, and more efficient, than open response tasks for learning when practice time is limited. To further enhance efficiency, we autograded open responses using GPT-4o and GPT-4-turbo. GPT models demonstrate proficiency for purposes of low-stakes assessment, though further research is needed for broader use. This study contributes a dataset of lesson log data, human annotation rubrics, and LLM prompts to promote transparency and reproducibility. Danielle R. Thomas, Conrad Borchers, Sanjit Kakarla, Jionghao Lin, Shambhavi Bhushan, Boyuan Guo, Erin Gatz, Kenneth R. Koedinger |
LAK | 7 |
| 2025 | Do Tutors Learn from Equity Training and Can Generative AI Assess It?abstractEquity is a core concern of learning analytics. However, applications that teach and assess equity skills, particularly at scale are lacking, often due to barriers in evaluating language. Advances in generative AI via large language models (LLMs) are being used in a wide range of applications, with this present work assessing its use in the equity domain. We evaluate tutor performance within an online lesson on enhancing tutors' skills when responding to students in potentially inequitable situations. We apply a mixed-method approach to analyze the performance of 81 undergraduate remote tutors. We find marginally significant learning gains with increases in tutors' self-reported confidence in their knowledge in responding to middle school students experiencing possible inequities from pretest to posttest. Both GPT-4o and GPT-4-turbo demonstrate proficiency in assessing tutors ability to predict and explain the best approach. Balancing performance, efficiency, and cost, we determine that few-shot learning using GPT-4o is the preferred model. This work makes available a dataset of lesson log data, tutor responses, rubrics for human annotation, and generative AI prompts. Future work involves leveling the difficulty among scenarios and enhancing LLM prompts for large-scale grading and assessment. Danielle R. Thomas, Conrad Borchers, Sanjit Kakarla, Jionghao Lin, Shambhavi Bhushan, Boyuan Guo, Erin Gatz, Kenneth R. Koedinger |
LAK | 7 |
| 2025 | Towards Equitable Community-Industry Collaborations: Understanding the Experiences of Nonprofits' Collaborations with Tech CompaniesabstractCommunity-based partnerships are essential to creating inclusive and equitable technologies and design practices. Though recent scholarship in HCI focuses on equitable design practices, there is less focus on understanding the experiences of community-based nonprofit organizations (CBOs) when partnering with technology companies. In this paper, we focus on understanding the perspectives of CBOs by answering the following research question: What are the experiences of CBOs that have collaborated with technology companies? Through a series of design workshops with 18 participants who work at community-based nonprofits that have collaborated with technology firms, we identified four elements of community-industry collaborations that collectively shape the overall experience: divergences in cultural and organizational norms, ''setting the table,'' project relationship dynamics, and affective qualities. We conclude by discussing the power structures that impact community-industry collaboration and suggest reflective practices to guide equitable collaborations between CBOs and tech companies. Sheena Lewis Erete, Eric Corbett, Natasha Smith-Walker, Jay L. Cunningham, Erin Gatz, Tina M. Park, Tam Perry, Lauren Wilcox, Remi Denton |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2025 | A Node on the Constellation: The Role of Feminist Makerspaces in Building and Sustaining Alternative Cultures of Technology ProductionabstractFeminist makerspaces offer community-led alternatives to dominant tech cultures by centering care, mutual aid, and collective knowledge production. While prior CSCW research has explored their inclusive practices, less is known about how these spaces sustain themselves over time. Drawing on interviews with 18 founders and members across 8 U.S. feminist makerspaces as well as autoethnographic reflection, we examine the organizational and relational practices that support long-term endurance. We find that sustainability is not achieved through growth or institutionalization, but through care-driven stewardship, solidarity with local justice movements, and shared governance. These social practices position feminist makerspaces as prefigurative counterspaces-sites that enact, rather than defer, feminist values in everyday practice. This paper offers empirical insight into how feminist makerspaces persist amid structural precarity, and highlights the forms of labor and coalition-building that underpin alternative sociotechnical infrastructures. Erin Gatz, Yasmine Kotturi, Andrea Afua Kwamya, Sarah E. Fox |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2024 | The Neglected 15%: Positive Effects of Hybrid Human-AI Tutoring Among Students with Disabilities
Danielle R. Thomas, Erin Gatz, Shivang Gupta, Vincent Aleven, Kenneth R. Koedinger |
AIED (1) | 2 |
| 2024 | Improving Student Learning with Hybrid Human-AI Tutoring: A Three-Study Quasi-Experimental InvestigationabstractArtificial intelligence (AI) applications to support human tutoring have potential to significantly improve learning outcomes, but engagement issues persist, especially among students from low-income backgrounds. We introduce an AI-assisted tutoring model that combines human and AI tutoring and hypothesize this synergy will have positive impacts on learning processes. To investigate this hypothesis, we conduct a three-study quasi-experiment across three urban and low-income middle schools: 1) 125 students in a Pennsylvania school; 2) 385 students (50% Latinx) in a California school, and 3) 75 students (100% Black) in a Pennsylvania charter school, all implementing analogous tutoring models. We compare learning analytics of students engaged in human-AI tutoring compared to students using math software only. We find human-AI tutoring has positive effects, particularly in student’s proficiency and usage, with evidence suggesting lower achieving students may benefit more compared to higher achieving students. We illustrate the use of quasi-experimental methods adapted to the particulars of different schools and data-availability contexts so as to achieve the rapid data-driven iteration needed to guide an inspired creation into effective innovation. Future work focuses on improving the tutor dashboard and optimizing tutor-student ratios, while maintaining annual costs per student of approximately $700 annually. Danielle R. Thomas, Jionghao Lin, Erin Gatz, Ashish Gurung, Shivang Gupta, Kole Norberg, Stephen Fancsali, Vincent Aleven, Lee G. Branstetter, Emma Brunskill, Kenneth R. Koedinger |
LAK | 3 |
| 2024 | Learning and AI Evaluation of Tutors Responding to Students Engaging in Negative Self-TalkabstractAddressing negative self-talk by students, such as responding to a student when saying, "I am dumb"or "I can't do this"can be difficult for even the most experienced tutor. Despite potential tutor learning from scenario-based lessons on this topic, human-graded assessment remains time-consuming. Leveraging generative AI for evaluating textual responses in online training presents a scalable solution. Research suggests a tutor validates student's feelings when they speak negatively of themselves, e.g., by a tutor responding, "I understand how you feel"or "I recognize this is difficult."This ongoing work assesses the performance of 60 undergraduate tutors within an online lesson on enhancing tutors' abilities to respond to students engaging in negative self-talk. We find statistically significant tutor learning gains from pretest to posttest. Additionally, we describe a method of using generative AI for assessing tutors' responses to predict the best approach and subsequently explain the rationale behind it. Using the large language model GPT-4, we find high absolute performance when evaluating tutor responses involving predicting (F1 = 0.85) and explaining (F1 = 0.83) the best approach. Minor improvements are needed to the lesson itself. A future goal of this work is to fully develop automated systems of assessing tutor learning attending to barriers to students' motivation and doing so at scale. Danielle R. Thomas, Jionghao Lin, Shambhavi Bhushan, Ralph Abboud, Erin Gatz, Shivang Gupta, Kenneth R. Koedinger |
L@S | 5 |
| 2024 | Peerdea: Co-Designing a Peer Support Platform with Creative EntrepreneursabstractCreative entrepreneurs rely on online platforms to build community and overcome isolated work conditions. However, because of frequent attempts by larger brands to use their work without permission, creative entrepreneurs constrain their use of social platforms to safeguard their intellectual property. In this paper, we describe a multi-year partnership with a feminist makerspace to build a social platform, called Peerdea, that centered creative entrepreneurs' needs such that online feedback, information exchange, goal setting, and accountability were more readily available to them. Through an iterative, community-collaborative approach with 46 creative entrepreneurs, we report on the kinds of peer support entrepreneurs sought on Peerdea such as feedback on in-progress and unpolished work. We argue that by aligning Peerdea's design with the makerspace's community of practice, Peerdea leveraged the relationship and trust building that occurs more readily in person for entrepreneurs. In addition, we highlight the role of a community leader who actively managed the relationships between researchers and entrepreneurs, surfaced research failures and championed successes, and provided critical mediation for co-design when participants' livelihoods were implicated. Yasmine Kotturi, Jenny Yu, Pranav Khadpe, Erin Gatz, Harvey Zheng, Sarah E. Fox, Chinmay Kulkarni 0001 |
Proc. ACM Hum. Comput. Interact. | 4 |