EDBT 2026 Demo / reviewers in the wild / expert
Tanja Käser
dblp:95/11458
· DBLP profile ↗
63ranked-venue papers
8as first author
51since 2021 · last 2026
0000-0003-0672-0415ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 46 · 7 first-author · 35 since 2021Human-computer interaction and ubiquitous computing · 30 · 5 first-author · 24 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 12 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student AttacksabstractLarge Language Models (LLMs) are increasingly used in education, yet their default helpfulness often conflicts with pedagogical principles.Prior work evaluates pedagogical quality via answer leakage-the disclosure of complete solutions instead of scaffolding-but typically assumes well-intentioned learners, leaving tutor robustness under student misuse largely unexplored.In this paper, we study scenarios where students behave adversarially and aim to obtain the correct answer from the tutor.We evaluate a broad set of LLM-based tutor models, including different model families, pedagogically aligned models, and a multi-agent design, under a range of adversarial student attacks.We adapt six groups of adversarial and persuasive techniques to the educational setting and use them to probe how likely a tutor is to reveal the final answer.We evaluate answer leakage robustness using different types of in-context adversarial student agents, finding that they often fail to carry out effective attacks.We therefore introduce an adversarial student agent that we fine-tune to jailbreak LLM-based tutors, which we propose as the core of a standardized benchmark for evaluating tutor robustness.Finally, we present simple but effective defense strategies that reduce answer leakage and strengthen the robustness of LLM-based tutors in adversarial scenarios.Code available at this link. Marta Knezevic, Tanja Käser |
ACL (1) | 3 |
| 2026 | REFINE: Real-World Exploration of Interactive Feedback and Student Behaviour
Fares Fawzi, Seyed Parsa Neshaei, Marta Knezevic, Tanya Nazaretsky, Tanja Käser |
AIED (1) | 5 |
| 2026 | Scalable and Explainable Learner-Video Interaction Prediction Using Multimodal Large Language Models
Dominik Glandorf, Fares Fawzi, Tanja Käser |
AIED | 3 |
| 2026 | Moving Beyond Review: Applying Language Models to Planning and Translation in Reflection
Seyed Parsa Neshaei, Richard Lee Davis, Tanja Käser |
AIED | 3 |
| 2026 | Towards Real-Time Personalized Feedback in Open-Ended Learning Environments
Ethan Prihar, Hugues Saltini, Peter Bühlmann, Tanja Käser |
AIED | 4 |
| 2026 | Teaching Multivariational Reasoning Through AI-Guided Inquiry in Interactive Simulations
Ekaterina Shved, Engin Bumbacher, Seyed Parsa Neshaei, Tanja Käser |
AIED | 4 |
| 2026 | Structuring versus Problematizing: How LLM-based Agents Scaffold Learning in Diagnostic ReasoningabstractSupporting students in developing diagnostic reasoning is a key challenge across educational domains. Novices often face cognitive biases such as premature closure and over-reliance on heuristics, and they struggle to transfer diagnostic strategies to new cases. Scenario-based learning (SBL) enhanced by Learning Analytics (LA) and large language models (LLM) offers a promising approach by combining realistic case experiences with personalized scaffolding. Yet, how different scaffolding approaches shape reasoning processes remains insufficiently explored. This study introduces PharmaSim Switch, an SBL environment for pharmacy technician training, extended with an LA- and LLM-powered pharmacist agent that implements pedagogical conversations rooted in two theory-driven scaffolding approaches: structuring and problematizing, as well as a student learning trajectory. In a between-groups experiment, 63 vocational students completed a learning scenario, a near-transfer scenario, and a far-transfer scenario under one of the two scaffolding conditions. Results indicate that both scaffolding approaches were effective in supporting the use of diagnostic strategies. Performance outcomes were primarily influenced by scenario complexity rather than students’ prior knowledge or the scaffolding approach used. The structuring approach was associated with more accurate Active and Interactive participation, whereas problematizing elicited more Constructive engagement. These findings underscore the value of combining scaffolding approaches when designing LA- and LLM-based systems to effectively foster diagnostic reasoning. Fatma Betül Güres, Tanya Nazaretsky, Seyed Parsa Neshaei, Tanja Käser |
LAK | 4 |
| 2026 | Turning 500+ Students into Teachers: A Semester-Long Study of an AI Teachable Agent in an Undergraduate Algorithms Course
Christopher Petrie, Miltiadis Stouras, Nicolas Ettlin, Amaury George, Paola Mejia-Domenzain, Vinitra Swamy, Tanja Käser, Ola Svensson |
L@S | 8 |
| 2025 | iLLuMinaTE: An LLM-XAI Framework Leveraging Social Science Explanation Theories Towards Actionable Student Performance FeedbackabstractRecent advances in eXplainable AI (XAI) for education have highlighted a critical challenge: ensuring that explanations for state-of-the-art models are understandable for non-technical users such as educators and students. In response, we introduce iLLuMinaTE, a zero-shot, chain-of-prompts LLM-XAI pipeline inspired by Miller (2019)'s cognitive model of explanation. iLLuMinaTE is designed to deliver theory-driven, actionable feedback to students in online courses. iLLuMinaTE navigates three main stages — causal connection, explanation selection, and explanation presentation — with variations drawing from eight social science theories (e.g. Abnormal Conditions, Pearl's Model of Explanation, Necessity and Robustness Selection, Contrastive Explanation). We extensively evaluate 21,915 natural language explanations of iLLuMinaTE extracted from three LLMs (GPT-4o, Gemma2-9B, Llama3-70B), with three different underlying XAI methods (LIME, Counterfactuals, MC-LIME), across students from three diverse online courses. Our evaluation involves analyses of explanation alignment to the social science theory, understandability of the explanation, and a real-world user preference study with 114 university students containing a novel actionability simulation. We find that students prefer iLLuMinaTE explanations over traditional explainers 89.52% of the time. Our work provides a robust, ready-to-use framework for effectively communicating hybrid XAI-driven insights in education, with significant generalization potential for other human-centric fields. Vinitra Swamy, Davide Romano, Bhargav Srinivasa Desikan, Oana-Maria Camburu, Tanja Käser |
AAAI | 5 |
| 2025 | One Code to Predict Them All: Universal Encoding for Inquiry Modeling
Jade Cock, Valentine Delevaux, Ido Roll, Richard Lee Davis, Tanja Käser |
AIED (5) | 5 |
| 2025 | Towards A Student-Facing Dashboard to Support Learning of Inquiry Competencies in Interactive Simulations
Eman Ganaiem, Tanja Käser, Ido Roll |
AIED (5) | 2 |
| 2025 | How Instructional Sequence and Personalized Support Impact Diagnostic Strategy Learning
Fatma Betül Güres, Tanya Nazaretsky, Bahar Radmehr, Martina A. Rau, Tanja Käser |
AIED (6) | 5 |
| 2025 | "Piecing Data Connections Together Like a Puzzle": Effects of Increasing Task Complexity on the Effectiveness of Data Storytelling Enhanced VisualisationsabstractThe emerging concept of data storytelling (DS) suggests that enhancing visualisations with annotations and narratives can make complex data more insightful than conventional visualisations. Previous works found that DS-enhanced visualisations are more effective than conventional visualisations for simple tasks like identifying key data points or the main message. However, no previous work has explored the extent to which DS enhancements influence task completion across different levels of cognitive complexity. We address this gap by presenting the results of a study where 128 participants completed tasks based on four visualisations (two line charts and two choropleth maps, either with or without DS elements) spanning a range of complexity based on Bloom's taxonomy, which has been applied in data visualisation to categorise tasks hierarchically from lower to higher-order thinking. Results suggest that while DS-enhanced visualisations effectively support lower-order tasks (finding data points and understanding insights), they don't necessarily aid the correct completion of higher-order tasks (application, analysis, evaluation and creation). However, DS enhancements improve how efficiently participants complete complex tasks. Mikaela Elizabeth Milesi, Paola Mejia-Domenzain, Laura Brandl, Vanessa Echeverría, Yueqiao Jin, Dragan Gasevic, Yi-Shan Tsai, Tanja Käser, Roberto Martínez-Maldonado |
CHI | 8 |
| 2025 | Applying DebiasEd: A Package for Mitigating Unfairness in Educational Data
Jade Cock, Frank Stinar, René F. Kizilcec, Tanja Käser |
EDM | 4 |
| 2025 | Bridging the Data Gap: Using LLMs to Augment Datasets for Text Classification
Seyed Parsa Neshaei, Richard Lee Davis, Paola Mejia-Domenzain, Tanya Nazaretsky, Tanja Käser |
EDM | 5 |
| 2025 | The 2nd Human-Centric eXplainable AI in Education (HEXED) Workshop
Vinitra Swamy, Jakub Kuzilek, Juan D. Pinto, Luc Paquette, Tanja Käser, Qianhui Liu, Lea Cohausz |
EDM | 5 |
| 2025 | SCRIBE: Structured Chain Reasoning for Interactive Behaviour Explanations using Tool CallingabstractLanguage models can be used to provide interactive, personalized student feedback in educational settings.However, real-world deployment faces three key challenges: privacy concerns, limited computational resources, and the need for pedagogically valid responses.These constraints require small, open-source models that can run locally and reliably ground their outputs in correct information.We introduce SCRIBE, a framework for multi-hop, tool-augmented reasoning designed to generate valid responses to student questions about feedback reports.SCRIBE combines domainspecific tools with a self-reflective inference pipeline that supports iterative reasoning, tool use, and error recovery.We distil these capabilities into 3B and 8B models via two-stage LoRA fine-tuning on synthetic GPT-4o-generated data.Evaluation with a human-aligned GPT-Judge and a user study with 108 students shows that 8B-SCRIBE models achieve comparable or superior quality to much larger models in key dimensions such as relevance and actionability, while being perceived on par with GPT-4o and Llama-3.3 70B by students.These findings demonstrate the viability of SCRIBE for low-resource, privacy-sensitive educational applications.How can I improve my performance to pass the course? UsefulRelevant Actionable Correct Tools sort_student_features_with_importance Improve your performance by watching course videos more regularly and following a steady study routine -these habits strongly influence learning outcomes.I need to understand which specific behaviors were most affecting your performance in the course.This helped identify the most influential factors-mainly video load frequency and the number of study sessions. Fares Fawzi, Vinitra Swamy, Dominik Glandorf, Tanya Nazaretsky, Tanja Käser |
EMNLP | 5 |
| 2025 | Intrinsic User-Centric Interpretability through Global Mixture of ExpertsabstractIn human-centric settings like education or healthcare, model accuracy and model explainability are key factors for user adoption. Towards these two goals, intrinsically interpretable deep learning models have gained popularity, focusing on accurate predictions alongside faithful explanations. However, there exists a gap in the human-centeredness of these approaches, which often produce nuanced and complex explanations that are not easily actionable for downstream users. We present InterpretCC (interpretable conditional computation), a family of intrinsically interpretable neural networks at a unique point in the design space that optimizes for ease of human understanding and explanation faithfulness, while maintaining comparable performance to state-of-the-art models. InterpretCC achieves this through adaptive sparse activation of features before prediction, allowing the model to use a different, minimal set of features for each instance. We extend this idea into an interpretable, global mixture-of-experts (MoE) model that allows users to specify topics of interest, discretely separates the feature space for each data point into topical subnetworks, and adaptively and sparsely activates these topical subnetworks for prediction. We apply InterpretCC for text, time series and tabular data across several real-world datasets, demonstrating comparable performance with non-interpretable baselines and outperforming intrinsically interpretable baselines. Through a user study involving 56 teachers, InterpretCC explanations are found to have higher actionability and usefulness over other intrinsically interpretable approaches. Vinitra Swamy, Syrielle Montariol, Julian Blackwell, Jibril Frej, Martin Jaggi, Tanja Käser |
ICLR | 6 |
| 2025 | TeamTeachingViz: Benefits, Challenges, and Ethical Considerations of Using a Multimodal Analytics Dashboard to Support Team Teaching ReflectionabstractTeam teaching in higher education can be challenging, especially for educators managing large classes with limited pedagogical training and few opportunities to reflect on their practices. Emerging sensing technologies and analytics can capture and analyse patterns of collaboration, communication, and movement of team teaching. Yet, few studies have presented these data to educators for reflection. To address this gap, we examine the benefits, challenges, and concerns of presenting multimodal teaching data (positional, audio, and spatial pedagogy observations) to educators via the TeamTeachingViz dashboard. We evaluated TeamTeachingViz in an authentic classroom context where educators explored their own data and team teaching strategies. Multimodal data was collected from 36 in-the-wild classroom sessions involving 12 educators grouped in various combinations over 4 weeks, followed by semi-structured interviews to reflect on their practices. Findings suggest that educators improved their self-awareness by using data-driven insights to understand their movements and interactions, enabling continuous improvement in team teaching. However, they noted the need for additional data, such as student behaviours and speech content, to better contextualise these insights. Riordan Alfredo, Paola Mejia-Domenzain, Vanessa Echeverría, Dwi Rahayu, Linxuan Zhao, Haya Alajlan, Zach Swiecki, Tanja Käser, Dragan Gasevic, Roberto Martínez-Maldonado |
LAK | 8 |
| 2025 | The Effect of Different Support Strategies on Student AffectabstractWithin many online learning platforms, struggling students are provided with support to guide them through challenging material. Support comes in many forms, and is typically evaluated based on its ability to improve students' performance on future tasks. However, there is little experimentation to evaluate how these supports impact students' emotional states. Student's emotional state, or affect, significantly impacts their motivation to engage with learning material and persist through challenges. Positive emotions can foster intrinsic engagement and deeper commitment, whereas negative emotions may lead to disengagement and avoidance of challenging tasks. In this work, we use publicly available data from online experiments and affect modeling to causally evaluate the impact that different support strategies have on students' affect. Through analysis of 25 experiments with 6,463 total participants, we find multiple significant positive and negative changes in students' affect when receiving hints, examples, or scaffolding questions, despite all three having a positive impact on performance, revealing the need for more nuanced evaluations of support strategies to uncover their impact beyond just performance. The code for this project is available at https://osf.io/74dgx. Julie Le Tallec, Ethan Prihar, Tanja Käser |
LAK | 3 |
| 2025 | Viewpoint: The Future of Human-Centric Explainable Artificial Intelligence (XAI) is not Post-Hoc ExplanationsabstractExplainable Artificial Intelligence (XAI) plays a crucial role in enabling human understanding and trust in deep learning systems. As models get larger, more ubiquitous, and pervasive in aspects of daily life, explainability is necessary to minimize adverse effects of model mistakes. Unfortunately, current approaches in human-centric XAI (e.g. predictive tasks in healthcare, education, or personalized ads) tend to rely on a single post-hoc explainer, whereas recent work has identified systematic disagreement between post-hoc explainers when applied to the same instances of underlying black-box models. In this viewpoint paper, we therefore present a call for action to address the limitations of current state-of-the-art explainers. We propose a shift from post-hoc explainability to designing interpretable neural network architectures. We identify five needs of human-centric XAI (real-time, accurate, actionable, human-interpretable, and consistent) and propose two possible routes forward for interpretable-by-design neural network workflows (adaptive routing and temporal diagnostics). We postulate that the future of human-centric XAI is neither in explaining black-boxes nor in reverting to traditional, interpretable models, but in neural networks that are intrinsically interpretable. Vinitra Swamy, Jibril Frej, Tanja Käser |
J. Artif. Intell. Res. | 3 |
| 2024 | Navigating Self-regulated Learning Dimensions: Exploring Interactions Across Modalities
Paola Mejia-Domenzain, Tanya Nazaretsky, Simon Schultze, Jan Hochweber, Tanja Käser |
AIED (2) | 5 |
| 2024 | Teaching and Measuring Multidimensional Inquiry Skills Using Interactive Simulations
Ekaterina Shved, Engin Bumbacher, Paola Mejia-Domenzain, Manu Kapur, Tanja Käser |
AIED (1) | 5 |
| 2024 | Fashioning Creative Expertise with Generative AI: Graphical Interfaces for Design Space Exploration Better Support Ideation Than Text PromptsabstractThis paper investigates the potential impact of deep generative models on the work of creative professionals. We argue that current generative modeling tools lack critical features that would make them useful creativity support tools, and introduce our own tool, generative.fashion1, which was designed with theoretical principles of design space exploration in mind. Through qualitative studies with fashion design apprentices, we demonstrate how generative.fashion supported both divergent and convergent thinking, and compare it with a state-of-the-art text-based interface using Stable Diffusion. In general, the apprentices preferred generative.fashion, citing the features explicitly designed to support ideation. In two follow-up studies, we provide quantitative results that support and expand on these insights. We conclude that text-only prompts in existing models restrict creative exploration, especially for novices. Our work demonstrates that interfaces which are theoretically aligned with principles of design space exploration are essential for unlocking the full creative potential of generative AI. Richard Lee Davis, Thiemo Wambsganss, Kevin Gonyop Kim, Tanja Käser, Pierre Dillenbourg |
CHI | 5 |
| 2024 | AI or Human? Evaluating Student Feedback Perceptions in Higher Education
Tanya Nazaretsky, Paola Mejia-Domenzain, Vinitra Swamy, Jibril Frej, Tanja Käser |
EC-TEL (1) | 5 |
| 2024 | Investigation of behavioral Differences: Uncovering Behavioral Sources of Demographic Bias in Educational Algorithms
Jade Cock, Hugues Saltini, Haoyu Sheng, Riya Ranjan, Richard Lee Davis, Tanja Käser |
EDM | 6 |
| 2024 | Student Answer Forecasting: Transformer-Driven Answer Choice Prediction for Language Learning
Elena Grazia Gado, Tommaso Martorella, Luca Zunino, Paola Mejia-Domenzain, Vinitra Swamy, Jibril Frej, Tanja Käser |
EDM | 7 |
| 2024 | Leveraging Large Language Models for Next-Generation Educational Technologies
Neil T. Heffernan, Rose E. Wang, Christopher J. MacLellan, Arto Hellas, Chenglu Li, Candace A. Walkington, Joshua Littenberg-Tobias, David Joyner, Steven Moore, Adish Singla, Zachary A. Pardos, Maciej Pankiewicz, Juho Kim 0001, Shashank Sonkar, Clayton Cohn, Anthony Botelho, Andrew S. Lan, Mingyu Feng, Tanja Käser, Eamon Worden |
EDM | 20 |
| 2024 | Towards Modeling Learner Performance with Large Language Models
Seyed Parsa Neshaei, Richard Lee Davis, Adam Hazimeh, Bojan Lazarevski, Pierre Dillenbourg, Tanja Käser |
EDM | 6 |
| 2024 | Human-Centric eXplainable AI in Education (HEXED) Workshop
Juan D. Pinto, Luc Paquette, Vinitra Swamy, Tanja Käser, Qianhui Liu, Lea Cohausz |
EDM | 4 |
| 2024 | Towards Generalizable Agents in Text-Based Educational Environments: A Study of Integrating RL with LLMs
Bahar Radmehr, Adish Singla, Tanja Käser |
EDM | 3 |
| 2024 | Let Me Teach You: Pedagogical Foundations of Feedback for Language ModelsabstractNatural Language Feedback (NLF) is an increasingly popular mechanism for aligning Large Language Models (LLMs) to human preferences.Despite the diversity of the information it can convey, NLF methods are often handdesigned and arbitrary, with little systematic grounding.At the same time, research in learning sciences has long established several effective feedback models.In this opinion piece, we compile ideas from pedagogy to introduce FELT, a feedback framework for LLMs that outlines various characteristics of the feedback space, and a feedback content taxonomy based on these variables, providing a general mapping of the feedback space.In addition to streamlining NLF designs, FELT also brings out new, unexplored directions for research in NLF.We make our taxonomy available to the community, providing guides and examples for mapping our categorizations to future research. Beatriz Borges, Niket Tandon, Tanja Käser, Antoine Bosselut |
EMNLP | 3 |
| 2024 | Finding Paths for Explainable MOOC Recommendation: A Learner PerspectiveabstractThe increasing availability of Massive Open Online Courses (MOOCs) has created a necessity for personalized course recommendation systems. These systems often combine neural networks with Knowledge Graphs (KGs) to achieve richer representations of learners and courses. While these enriched representations allow more accurate and personalized recommendations, explainability remains a significant challenge which is especially problematic for certain domains with significant impact such as education and online learning. Recently, a novel class of recommender systems that uses reinforcement learning and graph reasoning over KGs has been proposed to generate explainable recommendations in the form of paths over a KG. Despite their accuracy and interpretability on e-commerce datasets, these approaches have scarcely been applied to the educational domain and their use in practice has not been studied. In this work, we propose an explainable recommendation system for MOOCs that uses graph reasoning. To validate the practical implications of our approach, we conducted a user study examining user perceptions of our new explainable recommendations. We demonstrate the generalizability of our approach by conducting experiments on two educational datasets: COCO and Xuetang. Jibril Frej, Marta Knezevic, Tanya Nazaretsky, Tanja Käser |
LAK | 5 |
| 2024 | Course Recommender Systems Need to Consider the Job MarketabstractCurrent course recommender systems primarily leverage learner-course interactions, course content, learner preferences, and supplementary course details like instructor, institution, ratings, and reviews, to make their recommendation. However, these systems often overlook a critical aspect: the evolving skill demand of the job market. This paper focuses on the perspective of academic researchers, working in collaboration with the industry, aiming to develop a course recommender system that incorporates job market skill demands. In light of the job market's rapid changes and the current state of research in course recommender systems, we outline essential properties for course recommender systems to address these demands effectively, including explainable, sequential, unsupervised, and aligned with the job market and user's goals. Our discussion extends to the challenges and research questions this objective entails, including unsupervised skill extraction from job listings, course descriptions, and resumes, as well as predicting recommendations that align with learner objectives and the job market and designing metrics to evaluate this alignment. Furthermore, we introduce an initial system that addresses some existing limitations of course recommender systems using large Language Models (LLMs) for skill extraction and Reinforcement Learning (RL) for alignment with the job market. We provide empirical results using open-source data to demonstrate its effectiveness. Jibril Frej, Anna Dai, Syrielle Montariol, Antoine Bosselut, Tanja Käser |
SIGIR | 5 |
| 2023 | Ripple: Concept-Based Interpretation for Raw Time Series Models in EducationabstractTime series is the most prevalent form of input data for educational prediction tasks. The vast majority of research using time series data focuses on hand-crafted features, designed by experts for predictive performance and interpretability. However, extracting these features is labor-intensive for humans and computers. In this paper, we propose an approach that utilizes irregular multivariate time series modeling with graph neural networks to achieve comparable or better accuracy with raw time series clickstreams in comparison to hand-crafted features. Furthermore, we extend concept activation vectors for interpretability in raw time series models. We analyze these advances in the education domain, addressing the task of early student performance prediction for downstream targeted interventions and instructional support. Our experimental analysis on 23 MOOCs with millions of combined interactions over six behavioral dimensions show that models designed with our approach can (i) beat state-of-the-art educational time series baselines with no feature extraction and (ii) provide interpretable insights for personalized interventions. Source code: https://github.com/epfl-ml4ed/ripple/. Mohammad Asadi, Vinitra Swamy, Jibril Frej, Julien Tuan Tu Vignoud, Mirko Marras, Tanja Käser |
AAAI | 6 |
| 2023 | Understanding Revision Behavior in Adaptive Writing Support Systems for Education
Luca Mouchel, Thiemo Wambsganss, Paola Mejia-Domenzain, Tanja Käser |
EDM | 4 |
| 2023 | Protected Attributes Tell Us Who, Behavior Tells Us How: A Comparison of Demographic and Behavioral Oversampling for Fair Student Success ModelingabstractAlgorithms deployed in education can shape the learning experience and success of a student. It is therefore important to understand whether and how such algorithms might create inequalities or amplify existing biases. In this paper, we analyze the fairness of models which use behavioral data to identify at-risk students and suggest two novel pre-processing approaches for bias mitigation. Based on the concept of intersectionality, the first approach involves intelligent oversampling on combinations of demographic attributes. The second approach does not require any knowledge of demographic attributes and is based on the assumption that such attributes are a (noisy) proxy for student behavior. We hence propose to directly oversample different types of behaviors identified in a cluster analysis. We evaluate our approaches on data from (i) an open-ended learning environment and (ii) a flipped classroom course. Our results show that both approaches can mitigate model bias. Directly oversampling on behavior is a valuable alternative, when demographic metadata is not available. Source code and extended results are provided in https://github.com/epfl-ml4ed/behavioral-oversampling. Jade Cock, Richard Lee Davis, Mirko Marras, Tanja Käser |
LAK | 5 |
| 2023 | Do Not Trust a Model Because It is Confident: Uncovering and Characterizing Unknown Unknowns to Student Success Predictors in Online-Based LearningabstractStudent success models might be prone to develop weak spots, i.e., examples hard to accurately classify due to insufficient representation during model creation. This weakness is one of the main factors undermining users’ trust, since model predictions could for instance lead an instructor to not intervene on a student in need. In this paper, we unveil the need of detecting and characterizing unknown unknowns in student success prediction in order to better understand when models may fail. Unknown unknowns include the students for which the model is highly confident in its predictions, but is actually wrong. Therefore, we cannot solely rely on the model’s confidence when evaluating the predictions quality. We first introduce a framework for the identification and characterization of unknown unknowns. We then assess its informativeness on log data collected from flipped courses and online courses using quantitative analyses and interviews with instructors. Our results show that unknown unknowns are a critical issue in this domain and that our framework can be applied to support their detection. The source code is available at https://github.com/epfl-ml4ed/unknown-unknowns. Roberta Galici, Tanja Käser, Gianni Fenu, Mirko Marras |
LAK | 2 |
| 2023 | Trusting the Explainers: Teacher Validation of Explainable Artificial Intelligence for Course DesignabstractDeep learning models for learning analytics have become increasingly popular over the last few years; however, these approaches are still not widely adopted in real-world settings, likely due to a lack of trust and transparency. In this paper, we tackle this issue by implementing explainable AI methods for black-box neural networks. This work focuses on the context of online and blended learning and the use case of student success prediction models. We use a pairwise study design, enabling us to investigate controlled differences between pairs of courses. Our analyses cover five course pairs that differ in one educationally relevant aspect and two popular instance-based explainable AI methods (LIME and SHAP). We quantitatively compare the distances between the explanations across courses and methods. We then validate the explanations of LIME and SHAP with 26 semi-structured interviews of university-level educators regarding which features they believe contribute most to student success, which explanations they trust most, and how they could transform these insights into actionable course design decisions. Our results show that quantitatively, explainers significantly disagree with each other about what is important, and qualitatively, experts themselves do not agree on which explanations are most trustworthy. All code, extended results, and the interview protocol are provided at https://github.com/epfl-ml4ed/trusting-explainers. Vinitra Swamy, Sijia Du, Mirko Marras, Tanja Käser |
LAK | 4 |
| 2023 | MultiMoDN - Multimodal, Multi-Task, Interpretable Modular NetworksabstractPredicting multiple real-world tasks in a single model often requires a particularly diverse feature space. Multimodal (MM) models aim to extract the synergistic predictive potential of multiple data types to create a shared feature space with aligned semantic meaning across inputs of drastically varying sizes (i.e. images, text, sound). Most current MM architectures fuse these representations in parallel, which not only limits their interpretability but also creates a dependency on modality availability. We present MultiModN, a multimodal, modular network that fuses latent representations in a sequence of any number, combination, or type of modality while providing granular real-time predictive feedback on any number or combination of predictive tasks. MultiModN's composable pipeline is interpretable-by-design, as well as innately multi-task and robust to the fundamental issue of biased missingness. We perform four experiments on several benchmark MM datasets across 10 real-world tasks (predicting medical diagnoses, academic performance, and weather), and show that MultiModN's sequential MM fusion does not compromise performance compared with a baseline of parallel fusion. By simulating the challenging bias of missing not-at-random (MNAR), this work shows that, contrary to MultiModN, parallel fusion baselines erroneously learn MNAR and suffer catastrophic failure when faced with different patterns of MNAR at inference. To the best of our knowledge, this is the first inherently MNAR-resistant approach to MM modeling. In conclusion, MultiModN provides granular insights, robustness, and flexibility without compromising performance. Vinitra Swamy, Malika Satayeva, Jibril Frej, Thierry Bossy, Thijs Vogels, Martin Jaggi, Tanja Käser, Mary-Anne Hartley |
NeurIPS | 7 |
| 2023 | How Close are Predictive Models to Teachers in Detecting Learners at Risk?abstractDetecting learners in need of support is a complex process for both teachers and machines. Most prior work has devised visualization tools that allow teachers to do so by analyzing educational indicators. Other recent efforts have been devoted to models that predict whether learners might be at risk. However, the question on how teacher-like is the model behaving under this detection task still remains unanswered. In this paper, we investigate the (dis)agreement between teachers and model decisions, using a real-world flipped course as a case study. From the model perspective, we considered a well-known neural network, trained on educational indicators extracted from online pre-class logs. To gather teachers’ understanding, we employed a crowd sourcing approach including over 360 human intelligence tasks from 60 university teachers. We asked each recruited teacher to analyze visualizations pertaining to four relevant educational indicators of a given learner, and reason about their probability of failing the course (and so requiring support). Learners presented to teachers were selected to address different aspects of model confidence and (in)accuracy. Our results show that teacher and model predictions diverged for students who passed the course, while predictions were similar for students who failed the course. Moreover, confidence and correctness were more aligned in teachers than the model, reducing the unknown risks originally present in models. The source code is available at https://github.com/epfl-ml4ed/unknown-unknowns. Roberta Galici, Tanja Käser, Gianni Fenu, Mirko Marras |
UMAP | 2 |
| 2022 | Identifying and Comparing Multi-dimensional Student Profiles Across Flipped Classrooms
Paola Mejia-Domenzain, Mirko Marras, Christian Giang, Tanja Käser |
AIED (1) | 4 |
| 2022 | Bias at a Second Glance: A Deep Dive into Bias for German Educational Peer-Review Data ModelingabstractNatural Language Processing (NLP) has become increasingly utilized to provide adaptivity in educational applications. However, recent research has highlighted a variety of biases in pre-trained language models. While existing studies investigate bias in different domains, they are limited in addressing fine-grained analysis on educational corpora and text that is not English. In this work, we analyze bias across text and through multiple architectures on a corpus of 9,165 German peer-reviews collected from university students over five years. Notably, our corpus includes labels such as helpfulness, quality, and critical aspect ratings from the peer-review recipient as well as demographic attributes. We conduct a Word Embedding Association Test (WEAT) analysis on (1) our collected corpus in connection with the clustered labels, (2) the most common pre-trained German language models (T5, BERT, and GPT-2) and GloVe embeddings, and (3) the language models after fine-tuning on our collected data-set. In contrast to our initial expectations, we found that our collected corpus does not reveal many biases in the co-occurrence analysis or in the GloVe embeddings. However, the pre-trained German language models find substantial conceptual, racial, and gender bias and have significant changes in bias across conceptual and racial axes during fine-tuning on the peer-review data. With our research, we aim to contribute to the fourth UN sustainability goal (quality education) with a novel dataset, an understanding of biases in natural language education data, and the potential harms of not counteracting biases in language models for educational tasks. Thiemo Wambsganss, Vinitra Swamy, Roman Rietsche, Tanja Käser |
COLING | 4 |
| 2022 | Generalisable Methods for Early Prediction in Interactive Simulations for Education
Jade Cock, Mirko Marras, Christian Giang, Tanja Käser |
EDM | 4 |
| 2022 | Evaluating the Explainers: Black-Box Explainable Machine Learning for Student Success Prediction in MOOCs
Vinitra Swamy, Bahar Radmehr, Natasa Krco, Mirko Marras, Tanja Käser |
EDM | 5 |
| 2022 | Meta Transfer Learning for Early Success Prediction in MOOCsabstractDespite the increasing popularity of massive open online courses (MOOCs), many suffer from high dropout and low success rates. Early prediction of student success for targeted intervention is therefore essential to ensure no student is left behind in a course. There exists a large body of research in success prediction for MOOCs, focusing mainly on training models from scratch for individual courses. This setting is impractical in early success prediction as the performance of a student is only known at the end of the course. In this paper, we aim to create early success prediction models that can be transferred between MOOCs from different domains and topics. To do so, we present three novel strategies for transfer: 1) pre-training a model on a large set of diverse courses, 2) leveraging the pre-trained model by including meta information about courses, and 3) fine-tuning the model on previous course iterations. Our experiments on 26 MOOCs with over 145,000 combined enrollments and millions of interactions show that models combining interaction data and course information have comparable or better performance than models which have access to previous iterations of the course. With these models, we aim to effectively enable educators to warm-start their predictions for new and ongoing courses. Vinitra Swamy, Mirko Marras, Tanja Käser |
L@S | 3 |
| 2022 | Improving Students Argumentation Learning with Adaptive Self-Evaluation NudgingabstractRecent advantages from computational linguists can be leveraged to nudge students with adaptive self-evaluation based on their argumentation skill level. To investigate how individual argumentation self-evaluation will help students write more convincing texts, we designed an intelligent argumentation writing support system called ArgumentFeedback based on nudging theory and evaluated it in a series of three qualitative and quantitative studies with a total of 83 students. We found that students who received a self-evaluation nudge wrote more convincing texts with a better quality of formal and perceived argumentation compared to the control group. The measured self-efficacy and the technology acceptance provide promising results for embedding adaptive argumentation writing support tools in combination with digital nudging in traditional learning settings to foster self-regulated learning. Our results indicate that the design of nudging-based learning applications for self-regulated learning combined with computational methods for argumentation self-evaluation has a beneficial use to foster better writing skills of students. Thiemo Wambsganss, Andreas Janson, Tanja Käser, Jan Marco Leimeister |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2022 | Designing Conversational Evaluation Tools: A Comparison of Text and Voice Modalities to Improve Response Quality in Course EvaluationsabstractConversational agents (CAs) provide opportunities for improving the interaction in evaluation surveys. To investigate if and how a user-centered conversational evaluation tool impacts users' response quality and their experience, we build EVA - a novel conversational course evaluation tool for educational scenarios. In a field experiment with 128 students, we compared EVA against a static web survey. Our results confirm prior findings from literature about the positive effect of conversational evaluation tools in the domain of education. Second, we then investigate the differences between a voice-based and text-based conversational human-computer interaction of EVA in the same experimental set-up. Against our prior expectation, the students of the voice-based interaction answered with higher information quality but with lower quantity of information compared to the text-based modality. Our findings indicate that using a conversational CA (voice and text-based) results in a higher response quality and user experience compared to a static web survey interface. Thiemo Wambsganss, Naim Zierau, Matthias Söllner 0001, Tanja Käser, Kenneth R. Koedinger, Jan Marco Leimeister |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2021 | Early Prediction of Conceptual Understanding in Interactive Simulations
Jade Cock, Mirko Marras, Christian Giang, Tanja Käser |
EDM | 4 |
| 2021 | Can Feature Predictive Power Generalize? Benchmarking Early Predictors of Student Success across Flipped and Online Courses
Mirko Marras, Julien Tuan Tu Vignoud, Tanja Käser |
EDM | 3 |
| 2021 | L2D 2021: First International Workshop on Enabling Data-Driven Decisions from Learning on the WebabstractBy offering courses and resources, learning platforms on the Web have been attracting lots of participants, and the interactions with these systems have generated a vast amount of learning-related data. Their collection, processing and analysis have promoted a significant growth of learning analytics and have opened up new opportunities for supporting and assessing educational experiences. To provide all the stakeholders involved in the educational process with a timely guidance, being able to understand student's behavior and enable models which provide data-driven decisions pertaining to the learning domain is a primary property of online platforms, aiming at maximizing learning outcomes. In this workshop, we focus on collecting new contributions in this emerging area and on providing a common ground for researchers and practitioners (Website: https://mirkomarras.github.io/l2d-wsdm2021). Danilo Dessì, Tanja Käser, Mirko Marras, Elvira Popescu, Harald Sack |
WSDM | 2 |
| 2019 | Exploring Neural Network Models for the Classification of Students in Highly Interactive Environments
Tanja Käser, Daniel L. Schwartz 0001 |
EDM | 1 |
| 2017 | Efficient Feature Embeddings for Student Classification with Variational Auto-encoders
Severin Klingler, Rafael Wampfler, Tanja Käser, Barbara Solenthaler, Markus Gross 0001 |
EDM | 3 |
| 2017 | Modeling exploration strategies to predict student performance within a learning environment and beyondabstractModeling and predicting student learning is an important task in computer-based education. A large body of work has focused on representing and predicting student knowledge accurately. Existing techniques are mostly based on students' performance and on timing features. However, research in education, psychology and educational data mining has demonstrated that students' choices and strategies substantially influence learning. In this paper, we investigate the impact of students' exploration strategies on learning and propose the use of a probabilistic model jointly representing student knowledge and strategies. Our analyses are based on data collected from an interactive computer-based game. Our results show that exploration strategies are a significant predictor of the learning outcome. Furthermore, the joint models of performance and knowledge significantly improve the prediction accuracy within the game as well as on external post-test data, indicating that this combined representation provides a better proxy for learning. Tanja Käser, Nicole R. Hallinen, Daniel L. Schwartz 0001 |
LAK | 1 |
| 2016 | Temporally Coherent Clustering of Student Data
Severin Klingler, Tanja Käser, Barbara Solenthaler, Markus Gross 0001 |
EDM | 2 |
| 2016 | Stealth Assessment in ITS - A Study for Developmental Dyscalculia
Severin Klingler, Tanja Käser, Alberto Giovanni Busetto, Barbara Solenthaler, Juliane Kohn, Michael von Aster, Markus Gross 0001 |
ITS | 2 |
| 2016 | When to stop?: towards universal instructional policiesabstractThe adaptivity of intelligent tutoring systems relies on the accuracy of the student model and the design of the instructional policy. Recently an instructional policy has been presented that is compatible with all common student models. In this work we present the next step towards a universal instructional policy. We introduce a new policy that is applicable to an even wider range of student models including DBNs modeling skill topologies and forgetting. We theoretically and empirically compare our policy to previous policies. Using synthetic and real world data sets we show that our policy can effectively handle wheel-spinning students as well as forgetting across a wide range of student models. Tanja Käser, Severin Klingler, Markus Gross 0001 |
LAK | 1 |
| 2015 | On the Performance Characteristics of Latent-Factor and Knowledge Tracing Models
Severin Klingler, Tanja Käser, Barbara Solenthaler, Markus Gross 0001 |
EDM | 2 |
| 2014 | Computational Education using Latent Structured PredictionabstractComputational education offers an important add-on to conventional teaching. To provide optimal learning conditions, accurate representation of students’ current skills and adaptation to newly acquired knowledge are essential. To obtain sufficient representational power we investigate suitability of general graphical models and discuss adaptation by learning parameters of a log-linear distribution. For interpretability we propose to constrain the parameter space a-priori by leveraging domain knowledge. We show the benefits of general graphical models and of regularizing the parameter space by evaluation of our models on data collected from a computational education software for children having difficulties in learning mathematics. Tanja Käser, Alexander G. Schwing, Tamir Hazan, Markus Gross 0001 |
AISTATS | 1 |
| 2014 | Different parameters - same prediction: An analysis of learning curves
Tanja Käser, Kenneth R. Koedinger, Markus Gross 0001 |
EDM | 1 |
| 2014 | Beyond Knowledge Tracing: Modeling Skill Topologies with Bayesian Networks
Tanja Käser, Severin Klingler, Alexander G. Schwing, Markus Gross 0001 |
Intelligent Tutoring Systems | 1 |
| 2013 | Cluster-Based Prediction of Mathematical Learning Patterns
Tanja Käser, Alberto Giovanni Busetto, Barbara Solenthaler, Juliane Kohn, Michael von Aster, Markus Gross 0001 |
AIED | 1 |
| 2012 | Modelling and Optimizing the Process of Learning Mathematics
Tanja Käser, Alberto Giovanni Busetto, Gian-Marco Baschera, Juliane Kohn, Karin Kucian, Michael von Aster, Markus Gross 0001 |
ITS | 1 |