VLDB 2026 Research / reviewers in the wild / expert
Rosy Southwell
dblp:278/3892
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0003-4141-523XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | "It feels like we're not meeting the criteria": Examining and Mitigating the Cascading Effects of Bias in Automatic Speech Recognition in Spoken Language InterfacesabstractResearchers have demonstrated that Automatic Speech Recognition (ASR) systems perform differently across demographic groups (i.e. show bias), yet their downstream impact on spoken language interfaces remains unexplored. We examined this question in the context of a real-world AI-powered interface that provides tutors with feedback on the quality of their discourse. We found that the Whisper ASR had lower accuracy for Black vs. white tutors, likely due to differences in acoustic patterns of speech. The downstream automated discourse classifiers of tutor talk were correspondingly less accurate for Black tutors when presented with ASR input. As a result, although Black tutors demonstrated higher-quality discourse on human transcripts, this trend was not evident on ASR transcripts. We experimented with methods to reduce ASR bias, finding that fine-tuning the ASR on Black speech reduced, but did not eliminate, ASR bias and its downstream effects. We discuss implications for AI-based spoken language interfaces aimed at providing unbiased assessments to improve performance outcomes. Kelechi Ezema, Chelsea Chandler, Rosy Southwell, Niranjan Cholendiran, Sidney K. D'Mello |
CHI | 3 |
| 2024 | Speaker Diarization in the Classroom: How Much Does Each Student Speak in Group Discussions?
Shiran Dudy, Xinlu He, Rosy Southwell, Jacob Whitehill |
EDM | 5 |
| 2024 | Automatic Speech Recognition Tuned for Child Speech in the ClassroomabstractK-12 school classrooms have proven to be a challenging environment for Automatic Speech Recognition (ASR) systems, both due to background noise and conversation, and differences in linguistic and acoustic properties from adult speech, on which the majority of ASR systems are trained and evaluated. We report on experiments to improve ASR for child speech in the classroom by training and fine-tuning transformer models on public corpora of adult and child speech augmented with classroom background noise. By tuning OpenAI’s Whisper model we achieve a 38% relative reduction in word error rate (WER) to 9.2% on the public MyST dataset of child speech – the lowest yet reported – and a 7% relative reduction to reach 54% WER on a more challenging classroom speech dataset (ISAT). We also introduce a novel beam hypothesis rescoring method that incorporates a speed-aware term to capture prior knowledge of human speaking rates, as well as a Large Language Model, to select among hypotheses. We demonstrate the effectiveness of this technique on both publicly-available datasets and a classroom speech dataset. Rosy Southwell, Wayne H. Ward, Viet Anh Trinh, Charis Clevenger, Clay Clevenger, Emily Watts, Jason G. Reitman, Sidney K. D'Mello, Jacob Whitehill |
ICASSP | 1 |
| 2024 | Putting the "Brain" Back in the Eye-Mind Link: Aligning Eye Movements and Brain Activations During Naturalistic ReadingabstractEye movements have long been used to reflect ongoing cognitive processing to develop explanatory and predictive models of mental states and processes. This relationship, deemed the eye-mind link, contains underlying assumptions of the mental processes occurring in the brain, which have rarely been explicitly investigated. We propose a multimodal approach to investigate alignment of eye movements and brain activations (eye-brain alignment) and how it might be predicted by unfolding cognitive processes. We applied this method to a dataset of 76 participants who read long, connected texts while their eye movements and the hemodynamic responses in their brains were tracked using functional near-infrared spectroscopy (fNIRS). We found that reliable eye-brain alignment signals varied based on the participants’ cognitive state during reading. Implications for multimodal modeling of cognitive processes are discussed. Megan Caruso, Rosy Southwell, Leanne M. Hirshfield, Sidney K. D'Mello |
ICMI | 2 |
| 2023 | Getting the Wiggles Out: Movement Between Tasks Predicts Future Mind Wandering During Learning Activities
Rosy Southwell, Candace E. Peacock, Sidney K. D'Mello |
AIED | 1 |
| 2023 | A Comparative Analysis of Automatic Speech Recognition Errors in Small Group Classroom DiscourseabstractIn collaborative learning environments, effective intelligent learning systems need to accurately analyze and understand the collaborative discourse between learners (i.e., group modeling) to provide adaptive support. We investigate how automatic speech recognition (ASR) errors influence discourse models of small group collaboration in noisy real-world classrooms. Our dataset consisted of 30 students recorded by consumer off-the-shelf microphones (Yeti Blue) while engaging in dyadic- and triadic- collaborative learning in a multi-day STEM curriculum unit. We found that two state-of-the-art ASR systems (Google Speech and OpenAI Whisper) yielded very high word error rates (0.822, 0.847) but very different profiles of error with Google being more conservative, rejecting 38% of utterances instead of 12% for Whisper. Next, we examined how these ASR errors influenced down-stream small group modeling based on pre-trained large language models for three tasks: Abstract Meaning Representation parsing (AMRParsing), on-task/off-task detection (OnTask), and Accountable Productive Talk prediction (TalkMove). As expected, models trained on clean human transcripts yielded degraded performance on all three tasks, measured by the transfer ratio (TR). However, the TR of the specific sentence-level AMRParsing task (.39 - .62) was much lower than that of the abstract discourse-level OnTask (.63- .94) and TalkMove tasks (.64-.72). Furthermore, different training strategies that incorporated ASR transcripts alone or as augmentations of human transcripts increased accuracy for the discourse-level tasks (OnTask and TalkMove) but not AMRParsing. Simulation experiments suggested that the models were tolerant of missing utterances in the dialog context, and that jointly improving ASR accuracy on important word classes (e.g., verbs and nouns) can improve performance across all tasks. Overall, our results provide insights into how different types of NLP-based tasks might be tolerant of ASR errors under extremely noisy conditions and provide suggestions for how to improve accuracy in small group modeling settings for a more equitable, engaging, and adaptive collaborative learning environment. Jie Cao 0010, Ananya Ganesh, Jon Z. Cai, Rosy Southwell, Margaret Perkoff, Michael Regan, Katharina Kann, James H. Martin, Martha Palmer, Sidney K. D'Mello |
UMAP | 4 |
| 2023 | Gaze-based predictive models of deep reading comprehension
Rosy Southwell, Caitlin Mills 0001, Megan Caruso, Sidney K. D'Mello |
User Model. User Adapt. Interact. | 1 |
| 2022 | Going Deep and Far: Gaze-based Models Predict Multiple Depths of Comprehension During and One Week Following Reading
Megan Caruso, Candace E. Peacock, Rosy Southwell, Guojing Zhou, Sidney K. D'Mello |
EDM | 3 |
| 2022 | Challenges and Feasibility of Automatic Speech Recognition for Modeling Student Collaborative Discourse in Classrooms
Rosy Southwell, Samuel L. Pugh, Margaret Perkoff, Charis Clevenger, Jeffrey Bush 0001, Rachel Lieber, Wayne H. Ward, Peter W. Foltz, Sidney K. D'Mello |
EDM | 1 |