Manu Kapur

dblp:45/4612 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0002-2232-6111ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2025 MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors
abstract
Evaluating the pedagogical capabilities of AIbased tutoring models is critical for making guided progress in the field.Yet, we lack a reliable, easy-to-use, and simple-to-run evaluation that reflects the pedagogical abilities of models.To fill this gap, we present MATH-TUTORBENCH, an open-source benchmark for holistic tutoring model evaluation.MATHTU-TORBENCH contains a collection of datasets and metrics that broadly cover tutor abilities as defined by learning sciences research in dialogbased teaching.To score the pedagogical quality of open-ended teacher responses, we train a reward model and show it can discriminate expert from novice teacher responses with high accuracy.We evaluate a wide set of closed-and open-weight models on MATHTUTORBENCH and find that subject expertise, indicated by solving ability, does not immediately translate to good teaching.Rather, pedagogy and subject expertise appear to form a trade-off that is navigated by the degree of tutoring specialization of the model.Furthermore, tutoring appears to become more challenging in longer dialogs, where simpler questioning strategies begin to fail.We release the benchmark, code, and leaderboard openly to enable rapid benchmarking of future models. 1 github.com/eth-lre/mathtutorbench
Jakub Macina, Nico Daheim, Ido Hakimi, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan
EMNLP4
2024 Designing for Embodied Sense-making of Mathematics: Perspectives on Directed and Spontaneous Bodily Actions
abstract
While mathematics is conventionally viewed as an abstract discipline, contemporary perspectives on embodied cognition underscore the significance of integrating students’ bodily experiences into the learning process. However, the efficacy of embodied learning activities, as compared to traditional methods, remains under scrutiny. We argue that both directed and spontaneous bodily actions should be considered when designing embodied learning activities, and explore such bodily actions through two studies. A quantitative user study involving directed bodily actions in Virtual Reality and on tablet reveals vr’s support for math-anxious and body-aware learners, and distinct movement patterns related to varying mathematical abilities. A subsequent qualitative analysis identifies key characteristics of spontaneous bodily actions, namely coarseness, muscle tension, repetitions, anchors, perspective, and metaphors. Derived from both studies, we propose design recommendations, advocating for expanded embodied interaction design, consideration of embodied metaphors, coarse gesturing for deep features identification, supporting of sense-making anchors, and in-vr learning assessments.
Julia Chatain, Venera Gashaj, Bibin Muttappillil, Robert W. Sumner, Manu Kapur
Conference on Designing Interactive Systems5
2024 Teaching and Measuring Multidimensional Inquiry Skills Using Interactive Simulations
Ekaterina Shved, Engin Bumbacher, Paola Mejia-Domenzain, Manu Kapur, Tanja Käser
AIED (1)4
2024 Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
abstract
Large language models (LLMs) present an opportunity to scale high-quality personalized education to all.A promising approach towards this means is to build dialog tutoring models that scaffold students' problem-solving.However, even though existing LLMs perform well in solving reasoning questions, they struggle to precisely detect student's errors and tailor their feedback to these errors.Inspired by realworld teaching practice where teachers identify student errors and customize their response based on them, we focus on verifying student solutions and show how grounding to such verification improves the overall quality of tutor response generation.We collect a dataset of 1K stepwise math reasoning chains with the first error step annotated by teachers.We show empirically that finding the mistake in a student solution is challenging for current models.We propose and evaluate several verifiers for detecting these errors.Using both automatic and human evaluation we show that the student solution verifiers steer the generation model towards highly targeted responses to student errors which are more often correct with less hallucinations compared to existing baselines.https://github.com/eth-lre/ verify-then-generate Teacher If the height is 6, what is the length of the box?Volume of a box is height * width * length.Student Multi-turn dialog tutoring task Goal: Generate next teacher utterance.Not quite.Is the length you computed 2-times more than height?targeted and correct A. Error reason (baseline): Student made a careless mistake.B. Correctness verification: incorrect C. Stepwise verification: Step 2 -We set an equation 2 * length = 6 ... D. Error Description: length is used as a label instead of a variable representing the number.E. Alignment: Missing student steps: We know height is 6,... Matching steps are: Next we know length...<=>We set an equation... Verification-based Conditional Generation ModelEquation is height * width * length.Volume of a box is height * width * length.We set an equation 2 * length = 6, so length is 3.Next we know length = 2 * height, so length is 12. Student Reasoning Chain-of-Thought (CoT) SolutionThe volume is 6 * 4 * 3 = 72.We know height is 6, width is 4, and we found the length is 12.So the volume is 6 * 4 * 12 = 288.I think the answer is 72.Great work, this is correct!factually incorrect Conditional Generation Model (baseline) Targeted Correct Actionable 1. Stepwise verification + -
Nico Daheim, Jakub Macina, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan
EMNLP3
2023 The Development of Multivariable Causality Strategy: Instruction or Simulation First?
Janan Saba, Manu Kapur, Ido Roll
AIED2
2023 Opportunities and Challenges in Neural Dialog Tutoring
abstract
Jakub Macina, Nico Daheim, Lingzhi Wang, Tanmay Sinha, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Jakub Macina, Nico Daheim, Lingzhi Wang 0001, Tanmay Sinha, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan
EACL5
2023 Grounding Graph Theory in Embodied Concreteness with Virtual Reality
abstract
Abstract mathematics can be difficult to grasp, in part because it relies on symbols and formalisms that are powerful yet meaningless to novices unless grounded in concreteness. Although a wide corpus of research focuses on concreteness in mathematics education, the notion of concreteness can be apprehended in various ways and it is not yet clear which specific aspects of concreteness help the learners. In this paper, we explore embodiment as a form of concreteness to ground abstract mathematics. First, we designed and evaluated an embodied learning activity on graph theory. Through a user study with 89 participants, we then compared three approaches: abstraction, manipulated concreteness, and embodied concreteness. Our results show that, compared to abstraction, both forms of concreteness increase learners’ perceived attention, confidence, and satisfaction. However, only embodied concreteness increases perceived relevance and grounding. Moreover, unlike manipulated concreteness, embodied concreteness does not impair learning outcomes nor transfer abilities.
Julia Chatain, Rudolf Varga, Violaine Fayolle, Manu Kapur, Robert W. Sumner
TEI4
2022 Grasping Derivatives: Teaching Mathematics through Embodied Interactions using Tablets and Virtual Reality
abstract
Grasping mathematics can be difficult. Often, students struggle to connect mathematical concepts with their own experiences and even believe that math has nothing to do with the real world. To create more concreteness in mathematics education, we focus on the role of the body in learning, and more specifically, embodied interactions for learning derivatives. In this project, we designed an embodied game to teach derivatives, and validated our design with a panel of experts. We then used this prototype to explore different embodied interactions in terms of usability, sense of embodiment, and learning outcomes. In particular, we evaluated different degrees of embodied interactions, and different types of embodied interactions in Virtual Reality. We conclude with insights and recommendations for mathematics education with embodied interactions.
Julia Chatain, Virginia Ramp, Venera Gashaj, Violaine Fayolle, Manu Kapur, Robert W. Sumner, Stéphane Magnenat
IDC5
2022 Math Self-Concept, Stereotypes, Achievement and Anxiety - A Cross-Sectional Study
Alexander von Bergen, Dario Cvencek, Esther Ziegler, Venera Gashaj, Andrew N. Meltzoff, Manu Kapur
CogSci6
2022 The Impact of Prior Knowledge in Narrative-Based Learning on Understanding Biological Concepts in Higher Education
Samuel Tobler, Tanmay Sinha, Katja Köhler, Ernst Hafen, Manu Kapur
CogSci5
2022 Automatic Generation of Socratic Subquestions for Teaching Math Word Problems
abstract
Socratic questioning is an educational method that allows students to discover answers to complex problems by asking them a series of thoughtful questions.Generation of didactically sound questions is challenging, requiring understanding of the reasoning process involved in the problem.We hypothesize that such questioning strategy can not only enhance the human performance, but also assist the math word problem (MWP) solvers.In this work, we explore the ability of large language models (LMs) in generating sequential questions for guiding math word problem-solving.We propose various guided question generation schemes based on input conditioning and reinforcement learning.On both automatic and human quality evaluations, we find that LMs constrained with desirable question properties generate superior questions and improve the overall performance of a math word problem solver.We conduct a preliminary user study to examine the potential value of such question generation models in the education domain.Results suggest that the difficulty level of problems plays an important role in determining whether questioning improves or hinders human performance.We discuss the future of using such questioning strategies in education.https://github.com/eth-nlped/ scaffolding-generation
Kumar Shridhar, Jakub Macina, Mennatallah El-Assady, Tanmay Sinha, Manu Kapur, Mrinmaya Sachan
EMNLP5
2022 Productive Failure
abstract
There is now a substantive body of research that demonstrates the effectiveness of Productive Failure (PF) for learning novel concepts across STEM domains. I will start by describing what PF is, and situate the case for PF both theoretically and empirically. I will then delineate boundary conditions for how, when and why PF succeeds, ending with implications for teaching and learning.
Manu Kapur
ICER (1)1
2019 Resource-Rich versus Resource-Poor Assessment in Introductory Computer Science and its Implications on Models of Cognition: An in-Class Experimental Study
Tobias Halbherr, Hermann Lehner, Manu Kapur
CogSci3
2019 Impact of Explicit Failure and Success-driven Preparatory Activities on Learning
Tanmay Sinha, Manu Kapur, Robert West 0001, Michele Catasta, Matthias Hauswirth, Dragan Trninic
CogSci2
2019 When Productive Failure Fails
Tanmay Sinha, Manu Kapur
CogSci2
2019 The Disappearing "Advantages of Abstract Examples in Learning Math"
Dragan Trninic, Manu Kapur, Tanmay Sinha
CogSci2
2017 Preparatory Effects of Problem Posing on Learning from Instruction
Manu Kapur
CogSci1
2011 Classroom-based Experiments in Productive Failure
Manu Kapur, Katerine Bielaczyc
CogSci1
2007 Assessment of Group and Individual Learning through Intelligent Visualization (AGILeViz)
Eric Hamilton 0001, Amy L. Baylor, Mark J. Chavez, Ronald A. Cole, Chris DiGiano, Andrew Hurford, Michael J. Jacobson, Manu Kapur, Leo Lee, Richard Lesh, Manuel Lima, Philip Vahey
AIED8
2007 Diffusion of Pedagogical Innovations as a Complex Adaptive Process - Agent-Based Modeling as Research Method
Jun-Song Huang, Manu Kapur
ICCE2