VLDB 2026 Research / reviewers in the wild / expert
Dorottya Demszky
dblp:227/2614 · also Dora Demszky
· DBLP profile ↗
20ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0002-6759-9367ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 10 since 2021Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Marked Pedagogies: Examining Linguistic Biases in Personalized Automated Writing FeedbackabstractEffective personalized feedback is critical to students’ literacy development. Though LLM-powered tools now promise to automate such feedback at scale, LLMs are not language-neutral: they privilege standard academic English and reproduce social stereotypes, raising concerns about how “personalization” shapes the feedback students receive. We examine how four widely used LLMs (GPT-4o, GPT-3.5-turbo, Llama-3.3 70B, Llama-3.1 8B) adapt written feedback in response to student attributes. Using 600 eighth-grade persuasive essays from the PERSUADE dataset, we generated feedback under prompt conditions embedding gender, race/ethnicity, learning needs, achievement, and motivation. We analyze lexical shifts across model outputs by adapting the Marked Words framework. Our results reveal systematic, stereotype-aligned shifts in feedback conditioned on presumed student attributes—even when essay content was identical. Feedback for students marked by race, language, or disability often exhibited positive feedback bias and feedback withholding bias—overuse of praise, less substantive critique, and assumptions of limited ability. Across attributes, models tailored not only what content was emphasized but also how writing was judged and how students were addressed. We term these instructional orientations Marked Pedagogies and highlight the need for transparency and accountability in automated feedback tools. Mei Tan, Lena Phalen, Dorottya Demszky |
LAK | 3 |
| 2025 | Multi-Stage Speaker Diarization for Noisy Classrooms
Ali Sartaz Khan, Tolúlopé Ògúnrèmí, Ahmed Adel Attia, Dorottya Demszky |
EDM | 4 |
| 2025 | CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom EnvironmentsabstractCreating Automatic Speech Recognition (ASR) systems that are robust and resilient to classroom conditions is paramount to the development of AI tools to aid teachers and students. In this work, we study the efficacy of continued pretraining (CPT) in adapting Wav2vec2.0 to the classroom domain. We show that CPT is a powerful tool in that regard and reduces the Word Error Rate (WER) of Wav2vec2.0-based models by upwards of 10%. More specifically, CPT improves the model’s robustness to different noises, microphones and classroom conditions. Ahmed Adel Attia, Dorottya Demszky, Tolúlopé Ògúnrèmí, Jing Liu 0064, Carol Y. Espy-Wilson |
ICASSP | 2 |
| 2025 | From Weak Labels to Strong Results: Utilizing 5, 000 Hours of Noisy Classroom Transcripts with Minimal Accurate DataabstractRecent progress in speech recognition has relied on models trained on vast amounts of labeled data. However, classroom Automatic Speech Recognition (ASR) faces the real-world challenge of abundant weak transcripts paired with only a small amount of accurate, gold-standard data. In such low-resource settings, high transcription costs make re-transcription impractical. To address this, we ask: what is the best approach when abundant inexpensive weak transcripts coexist with limited gold-standard data, as is the case for classroom speech data? We propose Weakly Supervised Pretraining (WSP), a two-step process where models are first pretrained on weak transcripts in a supervised manner, and then fine-tuned on accurate data. Our results, based on both synthetic and real weak transcripts, show that WSP outperforms alternative methods, establishing it as an effective training methodology for low-resource ASR in real-world scenarios. Ahmed Adel Attia, Dorottya Demszky, Jing Liu 0064, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2025 | Exploring the Benefit of Customizing Feedback Interventions For Educators and Students With Offline Contextual Multi-Armed Bandits
Joy Yun, Allen Nie, Emma Brunskill, Dorottya Demszky |
LAK | 4 |
| 2025 | MathemaTikZ: A Dataset and Benchmark for Mathematical Diagram GenerationabstractDiagrams play a fundamental role in mathematics education, serving both as essential components of mathematical problems and as powerful scaffolding tools to support student comprehension. While AI tools have shown promise in supporting teachers with lesson preparation, especially with text-based mathematical content, they still struggle with reliably generating visual diagrams. Our work makes two main contributions: (1) We introduce MathemaTikZ, a dataset derived from the Illustrative Mathematics curriculum, comprising 3,793 mathematical diagrams paired with their natural language descriptions, problem contexts, and TikZ implementations. These span the full range of diagrams utilized in the K12 math curriculum. (2) We conduct comprehensive baseline evaluations using state-of-the-art language models (GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash) to assess current capabilities in mathematical diagram generation. Our findings reveal that even the best-performing models achieve a 73.9% success rate in accurately generating mathematical diagrams, with performance varying significantly across different types of visualizations. Through detailed error analysis, we identify four key challenge areas that future work should address: spatial reasoning and element placement, adherence to geometric constraints, pedagogical knowledge of mathematical diagrams, and preservation of mathematical relationships. Our results establish baselines for mathematical diagram generation and highlight critical areas for improvement in making AI tools more effective for mathematics education. Rizwaan Malik, Rebecca Li Hao, Ritika Kacholia, Dorottya Demszky |
L@S | 4 |
| 2025 | Advancing the Science of Teaching with Tutoring Data: A Collaborative Workshop with the National Tutoring ObservatoryabstractL@S ’25, Palermo, Italy Danielle R. Thomas, Dorottya Demszky, Kenneth R. Koedinger, Josh Marland, Doug Pietrzak, Justin Reich, Rachel Slama, Amalia Christina Toutziaridi, René F. Kizilcec |
L@S | 2 |
| 2024 | Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. AdultsabstractRecent advancements in Automatic Speech Recognition (ASR) systems, exemplified by Whisper, have demonstrated the potential of these systems to approach human-level performance given sufficient data. However, this progress doesn’t readily extend to ASR for children due to the lim- ited availability of suitable child-specific databases and the distinct characteristics of children’s speech. A recent study investigated leveraging the My Science Tutor (MyST) chil- dren’s speech corpus to enhance Whisper’s performance in recognizing children’s speech. They were able to demon- strate some improvement on a limited testset. This paper builds on these findings by enhancing the utility of the MyST dataset through more efficient data preprocessing. We reduce the Word Error Rate (WER) on the MyST testset 13.93% to 9.11% with Whisper-Small and from 13.23% to 8.61% with Whisper-Medium and show that this improvement can be generalized to unseen datasets. We also highlight important challenges towards improving children’s ASR performance and the effect of fine-tuning in improving the transcription of disfluent speech. Ahmed Adel Attia, Jing Liu 0064, Wei Ai 0002, Dorottya Demszky, Carol Y. Espy-Wilson |
AIES (1) | 4 |
| 2024 | Does Feedback on Talk Time Increase Student Engagement? Evidence from a Randomized Controlled Trial on a Math Tutoring PlatformabstractProviding ample opportunities for students to express their thinking is pivotal to their learning of mathematical concepts. We introduce the Talk Meter, which provides in-the-moment automated feedback on student-teacher talk ratios. We conduct a randomized controlled trial on a virtual math tutoring platform (n=742 tutors) to evaluate the effectiveness of the Talk Meter at increasing student talk. In one treatment arm, we show the Talk Meter only to the tutor, while in the other arm we show it to both the student and the tutor. We find that the Talk Meter increases student talk ratios in both treatment conditions by 13-14%; this trend is driven by the tutor talking less in the tutor-facing condition, whereas in the student-facing condition it is driven by the student expressing significantly more mathematical thinking. Through interviews with tutors, we find the student-facing Talk Meter was more motivating to students, especially those with introverted personalities, and was effective at encouraging joint effort towards balanced talk time. These results demonstrate the promise of in-the-moment joint talk time feedback to both teachers and students as a low cost, engaging, and scalable way to increase students’ mathematical reasoning. Dorottya Demszky, Rose E. Wang, Sean Geraghty, Carol Yu |
LAK | 1 |
| 2024 | Scaling High-Leverage Curriculum Scaffolding in Middle-School MathematicsabstractDespite well-designed curriculum materials, teachers often face significant challenges in their implementation due to the diverse learning needs present in classrooms. This paper examines whether and how Large Language Models (LLMs) can be leveraged to enhance K-12 math education by facilitating the creation of high-quality curriculum scaffolds that reflect expert teachers' strategies. Through an in-depth qualitative analysis with experienced middle-school math teachers, we identified crucial instructional supplements such as warm-up tasks and example-problem pairs that are essential for engaging students and supporting diverse learner needs. Building on these insights, we developed ScaffGen, an LLM-powered tool designed to generate curriculum-aligned educational materials. While LLMs alone may fall short in educational contexts, when enhanced with expert teacher insights, they can effectively emulate the cognitive processes required for pedagogically robust material creation. We plan to assess the effectiveness of these AI-generated materials through rigorous evaluations involving comparisons with expert-written benchmarks and field tests in classroom settings. This research highlights the potential of LLMs to mimic expert decision-making in educational material creation, offering significant implications for scalable instructional support. Rizwaan Malik, Dorna Abdi, Rose E. Wang, Dorottya Demszky |
L@S | 4 |
| 2024 | Reframing Authority: A Computational Measure of Power-Affirming Feedback on Student WritingabstractFeedback has the power to support students' development as agentive writers or to reinforce the authority of the teacher. The degree to which feedback is power-affirming---legitimizing students' ideas and positioning students as authors in the writing process---is a critical yet challenging construct to implement and measure. As teachers increasingly adopt AI-powered feedback solutions, it is important to understand how such feedback positions students. In this work, we collect a dataset of 1,012 in-line feedback comments on 7-10th grade student essays by 20 experienced English Language Arts (ELA) teachers, along with 400 in-line comments by GPT3. We adapt a framework of power-affirming feedback to annotate and characterize this data, and we train RoBERTa-based models to automatically classify feedback comments along key dimensions of the framework. We find that teacher-written comments are significantly more power-affirming than AI comments and that power-affirming feedback occurs infrequently in both teacher-written and AI-generated feedback. Our results indicate that RoBERTa-based models can effectively measure dimensions of power-affirming feedback, indicating that computational measures can serve as a scalable method for identifying and facilitating power-affirming feedback. Mei Tan, Christopher Mah, Dorottya Demszky |
L@S | 3 |
| 2024 | Enhancing Tutoring Effectiveness Through Automated Feedback: Preliminary Findings from a Pilot Randomized Controlled Trial on SAT TutoringabstractTo address educational inequities, high-quality SAT tutoring is crucial for students who need support. Many tutors, however, are novices and require coaching to be effective. To examine this issue, we conducted a pilot randomized controlled trial (RCT) to evaluate the effectiveness of personalized, automated feedback for novice tutors. This feedback aimed to enhance their tutoring skills and, consequently, improve student outcomes. In our RCT, we not only assessed the effectiveness of the feedback provided to novice tutors but also examined the impact of extending this feedback to both tutors and their students, compared to just the tutors alone. Furthermore, we explored how the use of social versus personal goal-oriented language in the feedback influences educational outcomes. Our preliminary findings from this pilot indicate that providing feedback to both tutors and learners led to a statistically significant improvement in average SAT practice test scores. Additionally, this approach significantly increased the tutor talk time ratio, an outcome that was somewhat unexpected and requires further investigation. These initial results form the foundation for a more comprehensive RCT currently underway. This ongoing study aims to delve deeper into these initial findings and refine our understanding of the effectiveness of feedback in tutoring settings. Joy Yun, Yann Hicke, Mariah Olson, Dorottya Demszky |
L@S | 4 |
| 2024 | Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math MistakesabstractRose Wang, Qingyang Zhang, Carly Robinson, Susanna Loeb, Dorottya Demszky. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Rose E. Wang, Carly Robinson, Susanna Loeb, Dorottya Demszky |
NAACL-HLT | 5 |
| 2023 | "Mistakes Help Us Grow": Facilitating and Evaluating Growth Mindset Supportive Language in ClassroomsabstractTeachers' growth mindset supportive language (GMSL)-rhetoric emphasizing that one's skills can be improved over time-has been shown to significantly reduce disparities in academic achievement and enhance students' learning outcomes.Although teachers espouse growth mindset principles, most find it difficult to adopt GMSL in their practice due the lack of effective coaching in this area.We explore whether large language models (LLMs) can provide automated, personalized coaching to support teachers' use of GMSL.We establish an effective coaching tool to reframe unsupportive utterances to GMSL by developing (i) a parallel dataset containing GMSL-trained teacher reframings of unsupportive statements with an accompanying annotation guide, (ii) a GMSL prompt framework to revise teachers' unsupportive language, and (iii) an evaluation framework grounded in psychological theory for evaluating GMSL with the help of students and teachers.1 We conduct a large-scale evaluation involving 174 teachers and 1,006 students, finding that both teachers and students perceive GMSL-trained teacher and model reframings as more effective in fostering a growth mindset and promoting challenge-seeking behavior, among other benefits.We also find that model-generated reframings outperform those from the GMSL-trained teachers.These results show promise for harnessing LLMs to provide automated GMSL feedback for teachers and, more broadly, LLMs' potentiality for supporting students' learning in the classroom.Our findings also demonstrate the benefit of largescale human evaluations when applying LLMs in educational domains. Kunal Handa, Margaret Clapper, Jessica Boyle, Rose E. Wang, Diyi Yang, David S. Yeager, Dorottya Demszky |
EMNLP | 7 |
| 2023 | MD3: The Multi-Dialect Dataset of Dialogues
Jacob Eisenstein, Vinodkumar Prabhakaran, Clara Rivera, Dorottya Demszky, Devyani Sharma |
INTERSPEECH | 4 |
| 2023 | M-Powering Teachers: Natural Language Processing Powered Feedback Improves 1: 1 Instruction and Student OutcomesabstractAlthough learners are being connected 1:1 with instructors at an increasing scale, most of these instructors do not receive effective, consistent feedback to help them improve. We deployed M-Powering Teachers, an automated tool based on natural language processing to give instructors feedback on dialogic instructional practices ---including their uptake of student contributions, talk time and questioning practices --- in a 1:1 online learning context. We conducted a randomized controlled trial on Polygence, a research mentorship platform for high schoolers (n=414 mentors) to evaluate the effectiveness of the feedback tool. We find that the intervention improved mentors' uptake of student contributions by 10%, reduced their talk time by 5% and improved student's experience with the program as well as their relative optimism about their academic future. These results corroborate existing evidence that scalable and low-cost automated feedback can improve instruction and learning in online educational contexts. Dorottya Demszky, Jing Liu 0064 |
L@S | 1 |
| 2021 | Measuring Conversational Uptake: A Case Study on Student-Teacher InteractionsabstractDorottya Demszky, Jing Liu, Zid Mancenido, Julie Cohen, Heather Hill, Dan Jurafsky, Tatsunori Hashimoto. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Dorottya Demszky, Jing Liu 0064, Zid Mancenido, Julie Cohen, Heather Hill, Daniel Jurafsky, Tatsunori B. Hashimoto |
ACL/IJCNLP (1) | 1 |
| 2021 | Learning to Recognize Dialect FeaturesabstractDorottya Demszky, Devyani Sharma, Jonathan Clark, Vinodkumar Prabhakaran, Jacob Eisenstein. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Dorottya Demszky, Devyani Sharma, Jonathan H. Clark, Vinodkumar Prabhakaran, Jacob Eisenstein |
NAACL-HLT | 1 |
| 2020 | GoEmotions: A Dataset of Fine-Grained EmotionsabstractUnderstanding emotion expressed in language has a wide range of applications, from building empathetic chatbots to detecting harmful online behavior. Advancement in this area can be improved using large-scale datasets with a fine-grained typology, adaptable to multiple downstream tasks. We introduce GoEmotions, the largest manually annotated dataset of 58k English Reddit comments, labeled for 27 emotion categories or Neutral. We demonstrate the high quality of the annotations via Principal Preserved Component Analysis. We conduct transfer learning experiments with existing emotion benchmarks to show that our dataset generalizes well to other domains and different emotion taxonomies. Our BERT-based model achieves an average F1-score of .46 across our proposed taxonomy, leaving much room for improvement. Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, Sujith Ravi |
ACL | 1 |
| 2020 | Pártélet: A Hungarian Corpus of Propaganda Texts from the Hungarian Socialist EraabstractIn this paper, we present Pártélet, a digitized Hungarian corpus of Communist propaganda texts. Pártélet was the official journal of the governing party during the Hungarian socialism from 1956 to 1989, hence it represents the direct political agitation and propaganda of the dictatorial system in question. The paper has a dual purpose: first, to present a general review of the corpus compilation process and the basic statistical data of the corpus, and second, to demonstrate through two case studies what the dataset can be used for. We show that our corpus provides a unique opportunity for conducting research on Hungarian propaganda discourse, as well as analyzing changes of this discourse over a 35-year period of time with computer-assisted methods. Zoltán Kmetty, Veronika Vincze, Dorottya Demszky, Orsolya Ring, Balázs Nagy 0004, Martina Katalin Szabó |
LREC | 3 |