Babette Bühler

dblp:287/4523 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-1679-4979ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Democratizing Writing Support with AI: Insights from One Year of Real-World Interactions with an Open-Access Writing Feedback Tool
abstract
Writing is a foundational skill for educational, professional, and civic participation, yet access to frequent and timely writing feedback remains deeply unequal. Teachers face significant workload constraints, particularly in large classes, and many learners lack alternative sources of individualized feedback. While large language models (LLMs) offer the opportunity for scalable, adaptive support, little is known about how students engage with such feedback tools in real-world, self-directed settings. We present a large-scale, year-long analysis of 23,650 voluntary interactions with an open-access AI writing feedback system used by students across diverse educational contexts and age groups, conducted in accordance with strict data protection standards. Using a clustering approach, we identify 2,800 iterative revision chains and apply a validated LLM-based multidimensional scoring framework to assess text quality over time. Our findings reveal that students who revised their texts after receiving AI feedback demonstrated statistically significant, albeit modest, improvements across both content and language-related dimensions (overall writing quality: ∆ = 0.067, p < .001, r = .17), with the greatest gains observed among initially low-performing writers. Revision frequency was positively associated with improvement, particularly in higher-order writing skills. However, engagement was uneven, with higher usage among students in academically oriented schools. These results demonstrate both the technical feasibility and social potential of deploying generative AI for educational support at scale, while highlighting the need for inclusive infrastructure, accessible design, and targeted outreach to truly democratize educational benefits.
Babette Bühler, Ivo Bueno, Enkelejda Kasneci
AAAI1
2026 Should AI Ask First? Investigating the Effects of Proactive vs Reactive AI Mentoring in Self-directed Learning
Khaoula Otmani, Anna Bodonhelyi, Babette Bühler, Enkelejda Kasneci
AIED (3)3
2026 Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
abstract
Observation of classroom interactions can provide concrete feedback to teachers, but current methods rely on manual annotation, which is resource-intensive and hard to scale. This work explores AI-driven analysis of classroom recordings, focusing on multimodal instructional activity and discourse recognition as a foundation for actionable feedback. Using a densely annotated dataset of 164 hours of video and 68 lesson transcripts, we design parallel, modality-specific pipelines. For video, we evaluate zero-shot multimodal LLMs, fine-tuned vision–language models, and self-supervised video transformers on 24 activity labels. For transcripts, we fine-tune a transformer-based classifier with contextualized inputs and compare it against prompting-based LLMs on 19 discourse labels. To handle class imbalance and multi-label complexity, we apply per-label thresholding, context windows, and imbalance-aware loss functions. The results show that fine-tuned models consistently outperform prompting-based approaches, achieving macro-F1 scores of 0.577 for video and 0.460 for transcripts. These results demonstrate the feasibility of automated classroom analysis and establish a foundation for scalable teacher feedback systems.
Ivo Bueno, Ruikun Hou, Babette Bühler, Tim Fütterer, James Drimalla, Jonathan K. Foster, Peter A. Youngs, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci
WACV3
2025 Multimodal Assessment of Classroom Discourse Quality: A Text-Centered Attention-Based Multi-Task Learning Approach
Ruikun Hou, Babette Bühler, Tim Fütterer, Efe Bozkir, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci
EDM2
2025 Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring
abstract
The manual assessment and grading of student writing is a time-consuming yet critical task for teachers. Recent developments in generative AI offer potential solutions to facilitate essay-scoring tasks for teachers. In our study, we evaluate the performance (e.g. alignment and reliability) of both open-source and closed-source LLMs in assessing German student essays, comparing their evaluations to those of 37 teachers across 10 pre-defined criteria (i.e., plot logic, expression). A corpus of 20 real-world essays from Year 7 and 8 students was analyzed using five LLMs: GPT-3.5, GPT-4, o1-preview, LLaMA 3-70B, and Mixtral 8x7B, aiming to provide in-depth insights into LLMs’ scoring capabilities. Closed-source GPT models outperform open-source models in both internal consistency and alignment with human ratings, particularly excelling in language-related criteria. The o1 model outperforms all other LLMs, achieving Spearman’s = .74 with human assessments in the Overall score, and an internal consistency of = .80, though biased towards higher scores. These findings indicate that LLM-based assessment can be a useful tool to reduce teacher workload by supporting the evaluation of essays, especially with regard to language-related criteria. However, due to their tendency to overrate and their remaining issues to capture the content quality, the models require further refinement.
Kathrin Seßler, Maurice Fürstenberg, Babette Bühler, Enkelejda Kasneci
LAK3
2024 Automated Assessment of Encouragement and Warmth in Classrooms Leveraging Multimodal Emotional Features and ChatGPT
Ruikun Hou, Tim Fütterer, Babette Bühler, Efe Bozkir, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci
AIED (1)3
2024 Detecting Aware and Unaware Mind Wandering During Lecture Viewing: A Multimodal Machine Learning Approach Using Eye Tracking, Facial Videos and Physiological Data
abstract
Learners often experience aware and unaware mind wandering during educational tasks, both negatively impacting learning outcomes. Differentiating these types of task-unrelated thoughts is crucial, as they stem from different cognitive processes and warrant tailored support that addresses the specific nature of mind wandering. Automated detection of these episodes could help mitigate their adverse effects, for example, by developing adaptive, attention-aware learning environments. In this study (N = 87), we explored a novel multimodal approach, combining eye tracking, facial videos, and physiological wristbands (i.e., electrodermal activity and heart rate), to predict aware and unaware mind wandering during lecture video watching. In addition, to allow comparison to previous research, we also predicted an integrated mind-wandering category. Mind wandering was assessed using 15 two-stage thought probes to determine task-unrelated thoughts and the participants’ awareness of their mind wandering. Our findings indicate that a multimodal approach outperforms unimodal methods, utilizing the top 100 features from the fused data. Specifically, aware mind wandering was detected at 20% above chance (AUC-PR = 0.396), unaware mind wandering at 14% above chance (AUC-PR = 0.267), and the combined category at 40% above chance (AUC-PR = 0.637). Eye tracking and video features proved more predictive than physiological measures when used as standalone modalities. SHAP analysis, employed to explain the results, highlighted the significance of integrating features from all three modalities for effective detection, particularly emphasizing the role of video-based facial expressions in identifying unaware mind wandering. Going beyond the current state of the art, this study demonstrates the potential of leveraging multimodal data to enhance the precision of aware and unaware mind-wandering detection and differentiation, setting a foundation for advancing educational technologies that respond dynamically to learners’ cognitive states.
Babette Bühler, Efe Bozkir, Hannah Deininger, Patricia Goldberg, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci
ICMI1
2024 On Task and in Sync: Examining the Relationship between Gaze Synchrony and Self-reported Attention During Video Lecture Learning
abstract
Successful learning depends on learners' ability to sustain attention, which is particularly challenging in online education due to limited teacher interaction. A potential indicator for attention is gaze synchrony, demonstrating predictive power for learning achievements in video-based learning in controlled experiments focusing on manipulating attention. This study (N=84) examines the relationship between gaze synchronization and self-reported attention of learners, using experience sampling, during realistic online video learning. Gaze synchrony was assessed through Kullback-Leibler Divergence of gaze density maps and MultiMatch algorithm scanpath comparisons. Results indicated significantly higher gaze synchronization in attentive participants for both measures and self-reported attention significantly predicted post-test scores. In contrast, synchrony measures did not correlate with learning outcomes. While supporting the hypothesis that attentive learners exhibit similar eye movements, the direct use of synchrony as an attention indicator poses challenges, requiring further research on the interplay of attention, gaze synchrony, and video content type.
Babette Bühler, Efe Bozkir, Hannah Deininger, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci
Proc. ACM Hum. Comput. Interact.1
2023 Automated Hand-Raising Detection in Classroom Videos: A View-Invariant and Occlusion-Robust Machine Learning Approach
Babette Bühler, Ruikun Hou, Efe Bozkir, Patricia Goldberg, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci
AIED1
2023 Watch out for those bananas! Gaze Based Mario Kart Performance Classification
abstract
This paper is about a small eye tracking study for scan path classification. Seven participants played Mario Kart while wearing a head mounted eye tracker. In total, we had 64 recordings, but one had to be removed (Only 79 gaze samples were recorded). We compared different scan path classification features to estimate the performance of the participants based on the ranking they achieved. The best performing feature was ENCODJI which incooperates saccades and the heatmap in one feature. HOV, which uses saccade angles, performed well for all tasks but was outperformed by the heatmap (HEAT) for the last two groups.
Wolfgang Fuhl, Björn Severitt, Nora Castner, Babette Bühler, Johannes Meyer 0001, Daniel Weber 0003, Regine Lendway, Ruikun Hou, Enkelejda Kasneci
ETRA4
2021 Web Table Classification Based on Visual Features
Babette Bühler, Heiko Paulheim
ICWE1