Nico Daheim

dblp:285/5587 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021
YearPublicationVenuePosition
2025 From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning
abstract
Large language models (LLMs) can transform education, but their optimization for direct question-answering often undermines effective pedagogy which requires strategically withholding answers.To mitigate this, we propose an online reinforcement learning (RL)-based alignment framework that can quickly adapt LLMs into effective tutors using simulated student-tutor interactions by emphasizing pedagogical quality and guided problem-solving over simply giving away answers.We use our method to train a 7B parameter tutor model without human annotations which reaches similar performance to larger proprietary models like LearnLM.We introduce a controllable reward weighting to balance pedagogical support and student solving accuracy, allowing us to trace the Pareto frontier between these two objectives.Our models better preserve reasoning capabilities than single-turn SFT baselines and can optionally enhance interpretability through thinking tags that expose the model's instructional planning.
David Dinucu-Jianu, Jakub Macina, Nico Daheim, Ido Hakimi, Iryna Gurevych, Mrinmaya Sachan
EMNLP3
2025 MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors
abstract
Evaluating the pedagogical capabilities of AIbased tutoring models is critical for making guided progress in the field.Yet, we lack a reliable, easy-to-use, and simple-to-run evaluation that reflects the pedagogical abilities of models.To fill this gap, we present MATH-TUTORBENCH, an open-source benchmark for holistic tutoring model evaluation.MATHTU-TORBENCH contains a collection of datasets and metrics that broadly cover tutor abilities as defined by learning sciences research in dialogbased teaching.To score the pedagogical quality of open-ended teacher responses, we train a reward model and show it can discriminate expert from novice teacher responses with high accuracy.We evaluate a wide set of closed-and open-weight models on MATHTUTORBENCH and find that subject expertise, indicated by solving ability, does not immediately translate to good teaching.Rather, pedagogy and subject expertise appear to form a trade-off that is navigated by the degree of tutoring specialization of the model.Furthermore, tutoring appears to become more challenging in longer dialogs, where simpler questioning strategies begin to fail.We release the benchmark, code, and leaderboard openly to enable rapid benchmarking of future models. 1 github.com/eth-lre/mathtutorbench
Jakub Macina, Nico Daheim, Ido Hakimi, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan
EMNLP2
2025 A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputs
abstract
Artem Shelmanov, Ekaterina Fadeeva, Akim Tsvigun, Ivan Tsvigun, Zhuohan Xie, Igor Kiselev, Nico Daheim, Caiqi Zhang, Artem Vazhentsev, Mrinmaya Sachan, Preslav Nakov, Timothy Baldwin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Artem Shelmanov, Ekaterina Fadeeva, Akim Tsvigun, Ivan Tsvigun, Zhuohan Xie, Igor Kiselev, Nico Daheim, Caiqi Zhang, Artem Vazhentsev, Mrinmaya Sachan, Preslav Nakov, Timothy Baldwin
EMNLP7
2025 Uncertainty-Aware Decoding with Minimum Bayes Risk
abstract
Despite their outstanding performance in the majority of scenarios, contemporary language models still occasionally generate undesirable outputs, for example, hallucinated text. While such behaviors have previously been linked to uncertainty, there is a notable lack of methods that actively consider uncertainty during text generation. In this work, we show how Minimum Bayes Risk (MBR) decoding, which selects model generations according to an expected risk, can be generalized into a principled uncertainty-aware decoding method. In short, we account for model uncertainty during decoding by incorporating a posterior over model parameters into MBR’s computation of expected risk. We show that this modified expected risk is useful for both choosing outputs and deciding when to abstain from generation and can provide improvements without incurring overhead. We benchmark different methods for learning posteriors and show that performance improves with prediction diversity. We release our code publicly.
Nico Daheim, Clara Meister, Thomas Möllenhoff, Iryna Gurevych
ICLR1
2024 Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
abstract
Large language models (LLMs) present an opportunity to scale high-quality personalized education to all.A promising approach towards this means is to build dialog tutoring models that scaffold students' problem-solving.However, even though existing LLMs perform well in solving reasoning questions, they struggle to precisely detect student's errors and tailor their feedback to these errors.Inspired by realworld teaching practice where teachers identify student errors and customize their response based on them, we focus on verifying student solutions and show how grounding to such verification improves the overall quality of tutor response generation.We collect a dataset of 1K stepwise math reasoning chains with the first error step annotated by teachers.We show empirically that finding the mistake in a student solution is challenging for current models.We propose and evaluate several verifiers for detecting these errors.Using both automatic and human evaluation we show that the student solution verifiers steer the generation model towards highly targeted responses to student errors which are more often correct with less hallucinations compared to existing baselines.https://github.com/eth-lre/ verify-then-generate Teacher If the height is 6, what is the length of the box?Volume of a box is height * width * length.Student Multi-turn dialog tutoring task Goal: Generate next teacher utterance.Not quite.Is the length you computed 2-times more than height?targeted and correct A. Error reason (baseline): Student made a careless mistake.B. Correctness verification: incorrect C. Stepwise verification: Step 2 -We set an equation 2 * length = 6 ... D. Error Description: length is used as a label instead of a variable representing the number.E. Alignment: Missing student steps: We know height is 6,... Matching steps are: Next we know length...<=>We set an equation... Verification-based Conditional Generation ModelEquation is height * width * length.Volume of a box is height * width * length.We set an equation 2 * length = 6, so length is 3.Next we know length = 2 * height, so length is 12. Student Reasoning Chain-of-Thought (CoT) SolutionThe volume is 6 * 4 * 3 = 72.We know height is 6, width is 4, and we found the length is 12.So the volume is 6 * 4 * 12 = 288.I think the answer is 72.Great work, this is correct!factually incorrect Conditional Generation Model (baseline) Targeted Correct Actionable 1. Stepwise verification + -
Nico Daheim, Jakub Macina, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan
EMNLP1
2024 Model Merging by Uncertainty-Based Gradient Matching
abstract
Models trained on different datasets can be merged by a weighted-averaging of their parameters, but why does it work and when can it fail? Here, we connect the inaccuracy of weighted-averaging to mismatches in the gradients and propose a new uncertainty-based scheme to improve the performance by reducing the mismatch. The connection also reveals implicit assumptions in other schemes such as averaging, task arithmetic, and Fisher-weighted averaging. Our new method gives consistent improvements for large language models and vision transformers, both in terms of performance and robustness to hyperparameters.
Nico Daheim, Thomas Möllenhoff, Edoardo Maria Ponti, Iryna Gurevych, Mohammad Emtiyaz Khan
ICLR1
2024 Variational Learning is Effective for Large Deep Networks
abstract
We give extensive empirical evidence against the common belief that variational learning is ineffective for large neural networks. We show that an optimizer called Improved Variational Online Newton (IVON) consistently matches or outperforms Adam for training large networks such as GPT-2 and ResNets from scratch. IVON's computational costs are nearly identical to Adam but its predictive uncertainty is better. We show several new use cases of IVON where we improve finetuning and model merging in Large Language Models, accurately predict generalization error, and faithfully estimate sensitivity to data. We find overwhelming evidence that variational learning is effective. Code is available at https://github.com/team-approx-bayes/ivon.
Yuesong Shen, Nico Daheim, Bai Cong, Peter Nickl, Gian Maria Marconi, Clement Bazan, Rio Yokota, Iryna Gurevych, Daniel Cremers, Mohammad Emtiyaz Khan, Thomas Möllenhoff
ICML2
2024 Elastic Weight Removal for Faithful and Abstractive Dialogue Generation
abstract
Nico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, Edoardo Ponti. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Nico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, Edoardo Maria Ponti
NAACL-HLT1
2024 Task-Oriented Document-Grounded Dialog Systems by HLTPR@RWTH for DSTC9 and DSTC10
abstract
This paper summarizes our contributions to the document-grounded dialog tasks at the 9th and 10th Dialog System Technology Challenges (DSTC9 and DSTC10). In both iterations the task consists of three subtasks: first detect whether the current turn is knowledge seeking, second select a relevant knowledge document, and third generate a response grounded on the selected document. For DSTC9 we proposed different approaches to make the selection task more efficient. The best method, Hierarchical Selection, actually improves the results compared to the original baseline and gives a speedup of 24x. In the DSTC10 iteration of the task, the challenge was to adapt systems trained on written dialogs to perform well on noisy automatic speech recognition transcripts. Therefore, we proposed data augmentation techniques to increase the robustness of the models as well as methods to adapt the style of generated responses to fit well into the proceeding dialog. Additionally, we proposed a noisy channel model that allows for increasing the factuality of the generated responses. In addition to summarizing our previous contributions, in this work, we also report on a few small improvements and reconsider the automatic evaluation metrics for the generation task which have shown a low correlation to human judgments.
David Thulke, Nico Daheim, Christian Dugast, Hermann Ney
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Opportunities and Challenges in Neural Dialog Tutoring
abstract
Jakub Macina, Nico Daheim, Lingzhi Wang, Tanmay Sinha, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Jakub Macina, Nico Daheim, Lingzhi Wang 0001, Tanmay Sinha, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan
EACL2
2023 Poor Man's Quality Estimation: Predicting Reference-Based MT Metrics Without the Reference
abstract
Vilém Zouhar, Shehzaad Dhuliawala, Wangchunshu Zhou, Nico Daheim, Tom Kocmi, Yuchen Eleanor Jiang, Mrinmaya Sachan. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Vilém Zouhar, Shehzaad Dhuliawala, Wangchunshu Zhou, Nico Daheim, Tom Kocmi, Yuchen Eleanor Jiang, Mrinmaya Sachan
EACL4