VLDB 2026 Research / reviewers in the wild / expert
Shambhavi Bhushan
dblp:380/7063
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
0009-0004-3695-2334ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Open-Response Assessment with LearnLM
Danielle R. Thomas, Conrad Borchers, Shambhavi Bhushan, Sanjit Kakarla, Alex Houk, Ralph Abboud, Shivang Gupta, Erin Gatz, Kenneth R. Koedinger |
AIED (5) | 3 |
| 2025 | Leveraging LLMs to Assess Tutor Moves in Real-Life Dialogues: A Feasibility Study
Danielle R. Thomas, Conrad Borchers, Jionghao Lin, Sanjit Kakarla, Shambhavi Bhushan, Erin Gatz, Shivang Gupta, Ralph Abboud, Kenneth R. Koedinger |
EC-TEL (2) | 5 |
| 2025 | Does Multiple Choice Have a Future in the Age of Generative AI? A Posttest-only RCTabstractThe role of multiple-choice questions (MCQs) as effective learning tools has been debated in past research. While MCQs are widely used due to their ease in grading, open response questions are increasingly used for instruction, given advances in large language models (LLMs) for automated grading. This study evaluates MCQs effectiveness relative to open-response questions, both individually and in combination, on learning. These activities are embedded within six tutor lessons on advocacy. Using a posttest-only randomized control design, we compare the performance of 234 tutors (790 lesson completions) across three conditions: MCQ only, open response only, and a combination of both. We find no significant learning differences across conditions at posttest, but tutors in the MCQ condition took significantly less time to complete instruction. These findings suggest that MCQs are as effective, and more efficient, than open response tasks for learning when practice time is limited. To further enhance efficiency, we autograded open responses using GPT-4o and GPT-4-turbo. GPT models demonstrate proficiency for purposes of low-stakes assessment, though further research is needed for broader use. This study contributes a dataset of lesson log data, human annotation rubrics, and LLM prompts to promote transparency and reproducibility. Danielle R. Thomas, Conrad Borchers, Sanjit Kakarla, Jionghao Lin, Shambhavi Bhushan, Boyuan Guo, Erin Gatz, Kenneth R. Koedinger |
LAK | 5 |
| 2025 | Do Tutors Learn from Equity Training and Can Generative AI Assess It?abstractEquity is a core concern of learning analytics. However, applications that teach and assess equity skills, particularly at scale are lacking, often due to barriers in evaluating language. Advances in generative AI via large language models (LLMs) are being used in a wide range of applications, with this present work assessing its use in the equity domain. We evaluate tutor performance within an online lesson on enhancing tutors' skills when responding to students in potentially inequitable situations. We apply a mixed-method approach to analyze the performance of 81 undergraduate remote tutors. We find marginally significant learning gains with increases in tutors' self-reported confidence in their knowledge in responding to middle school students experiencing possible inequities from pretest to posttest. Both GPT-4o and GPT-4-turbo demonstrate proficiency in assessing tutors ability to predict and explain the best approach. Balancing performance, efficiency, and cost, we determine that few-shot learning using GPT-4o is the preferred model. This work makes available a dataset of lesson log data, tutor responses, rubrics for human annotation, and generative AI prompts. Future work involves leveling the difficulty among scenarios and enhancing LLM prompts for large-scale grading and assessment. Danielle R. Thomas, Conrad Borchers, Sanjit Kakarla, Jionghao Lin, Shambhavi Bhushan, Boyuan Guo, Erin Gatz, Kenneth R. Koedinger |
LAK | 5 |
| 2024 | Learning and AI Evaluation of Tutors Responding to Students Engaging in Negative Self-TalkabstractAddressing negative self-talk by students, such as responding to a student when saying, "I am dumb"or "I can't do this"can be difficult for even the most experienced tutor. Despite potential tutor learning from scenario-based lessons on this topic, human-graded assessment remains time-consuming. Leveraging generative AI for evaluating textual responses in online training presents a scalable solution. Research suggests a tutor validates student's feelings when they speak negatively of themselves, e.g., by a tutor responding, "I understand how you feel"or "I recognize this is difficult."This ongoing work assesses the performance of 60 undergraduate tutors within an online lesson on enhancing tutors' abilities to respond to students engaging in negative self-talk. We find statistically significant tutor learning gains from pretest to posttest. Additionally, we describe a method of using generative AI for assessing tutors' responses to predict the best approach and subsequently explain the rationale behind it. Using the large language model GPT-4, we find high absolute performance when evaluating tutor responses involving predicting (F1 = 0.85) and explaining (F1 = 0.83) the best approach. Minor improvements are needed to the lesson itself. A future goal of this work is to fully develop automated systems of assessing tutor learning attending to barriers to students' motivation and doing so at scale. Danielle R. Thomas, Jionghao Lin, Shambhavi Bhushan, Ralph Abboud, Erin Gatz, Shivang Gupta, Kenneth R. Koedinger |
L@S | 3 |