VLDB 2026 Research / reviewers in the wild / expert
Nico Daheim
dblp:285/5587
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement LearningabstractLarge language models (LLMs) can transform education, but their optimization for direct question-answering often undermines effective pedagogy which requires strategically withholding answers.To mitigate this, we propose an online reinforcement learning (RL)-based alignment framework that can quickly adapt LLMs into effective tutors using simulated student-tutor interactions by emphasizing pedagogical quality and guided problem-solving over simply giving away answers.We use our method to train a 7B parameter tutor model without human annotations which reaches similar performance to larger proprietary models like LearnLM.We introduce a controllable reward weighting to balance pedagogical support and student solving accuracy, allowing us to trace the Pareto frontier between these two objectives.Our models better preserve reasoning capabilities than single-turn SFT baselines and can optionally enhance interpretability through thinking tags that expose the model's instructional planning. David Dinucu-Jianu, Jakub Macina, Nico Daheim, Ido Hakimi, Iryna Gurevych, Mrinmaya Sachan |
EMNLP | 3 |
| 2025 | MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM TutorsabstractEvaluating the pedagogical capabilities of AIbased tutoring models is critical for making guided progress in the field.Yet, we lack a reliable, easy-to-use, and simple-to-run evaluation that reflects the pedagogical abilities of models.To fill this gap, we present MATH-TUTORBENCH, an open-source benchmark for holistic tutoring model evaluation.MATHTU-TORBENCH contains a collection of datasets and metrics that broadly cover tutor abilities as defined by learning sciences research in dialogbased teaching.To score the pedagogical quality of open-ended teacher responses, we train a reward model and show it can discriminate expert from novice teacher responses with high accuracy.We evaluate a wide set of closed-and open-weight models on MATHTUTORBENCH and find that subject expertise, indicated by solving ability, does not immediately translate to good teaching.Rather, pedagogy and subject expertise appear to form a trade-off that is navigated by the degree of tutoring specialization of the model.Furthermore, tutoring appears to become more challenging in longer dialogs, where simpler questioning strategies begin to fail.We release the benchmark, code, and leaderboard openly to enable rapid benchmarking of future models. 1 github.com/eth-lre/mathtutorbench Jakub Macina, Nico Daheim, Ido Hakimi, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan |
EMNLP | 2 |
| 2025 | A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM OutputsabstractArtem Shelmanov, Ekaterina Fadeeva, Akim Tsvigun, Ivan Tsvigun, Zhuohan Xie, Igor Kiselev, Nico Daheim, Caiqi Zhang, Artem Vazhentsev, Mrinmaya Sachan, Preslav Nakov, Timothy Baldwin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Artem Shelmanov, Ekaterina Fadeeva, Akim Tsvigun, Ivan Tsvigun, Zhuohan Xie, Igor Kiselev, Nico Daheim, Caiqi Zhang, Artem Vazhentsev, Mrinmaya Sachan, Preslav Nakov, Timothy Baldwin |
EMNLP | 7 |
| 2025 | Uncertainty-Aware Decoding with Minimum Bayes RiskabstractDespite their outstanding performance in the majority of scenarios, contemporary language models still occasionally generate undesirable outputs, for example, hallucinated text. While such behaviors have previously been linked to uncertainty, there is a notable lack of methods that actively consider uncertainty during text generation. In this work, we show how Minimum Bayes Risk (MBR) decoding, which selects model generations according to an expected risk, can be generalized into a principled uncertainty-aware decoding method. In short, we account for model uncertainty during decoding by incorporating a posterior over model parameters into MBR’s computation of expected risk. We show that this modified expected risk is useful for both choosing outputs and deciding when to abstain from generation and can provide improvements without incurring overhead. We benchmark different methods for learning posteriors and show that performance improves with prediction diversity. We release our code publicly. Nico Daheim, Clara Meister, Thomas Möllenhoff, Iryna Gurevych |
ICLR | 1 |
| 2024 | Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model TutorsabstractLarge language models (LLMs) present an opportunity to scale high-quality personalized education to all.A promising approach towards this means is to build dialog tutoring models that scaffold students' problem-solving.However, even though existing LLMs perform well in solving reasoning questions, they struggle to precisely detect student's errors and tailor their feedback to these errors.Inspired by realworld teaching practice where teachers identify student errors and customize their response based on them, we focus on verifying student solutions and show how grounding to such verification improves the overall quality of tutor response generation.We collect a dataset of 1K stepwise math reasoning chains with the first error step annotated by teachers.We show empirically that finding the mistake in a student solution is challenging for current models.We propose and evaluate several verifiers for detecting these errors.Using both automatic and human evaluation we show that the student solution verifiers steer the generation model towards highly targeted responses to student errors which are more often correct with less hallucinations compared to existing baselines.https://github.com/eth-lre/ verify-then-generate Teacher If the height is 6, what is the length of the box?Volume of a box is height * width * length.Student Multi-turn dialog tutoring task Goal: Generate next teacher utterance.Not quite.Is the length you computed 2-times more than height?targeted and correct A. Error reason (baseline): Student made a careless mistake.B. Correctness verification: incorrect C. Stepwise verification: Step 2 -We set an equation 2 * length = 6 ... D. Error Description: length is used as a label instead of a variable representing the number.E. Alignment: Missing student steps: We know height is 6,... Matching steps are: Next we know length...<=>We set an equation... Verification-based Conditional Generation ModelEquation is height * width * length.Volume of a box is height * width * length.We set an equation 2 * length = 6, so length is 3.Next we know length = 2 * height, so length is 12. Student Reasoning Chain-of-Thought (CoT) SolutionThe volume is 6 * 4 * 3 = 72.We know height is 6, width is 4, and we found the length is 12.So the volume is 6 * 4 * 12 = 288.I think the answer is 72.Great work, this is correct!factually incorrect Conditional Generation Model (baseline) Targeted Correct Actionable 1. Stepwise verification + - Nico Daheim, Jakub Macina, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan |
EMNLP | 1 |
| 2024 | Model Merging by Uncertainty-Based Gradient MatchingabstractModels trained on different datasets can be merged by a weighted-averaging of their parameters, but why does it work and when can it fail? Here, we connect the inaccuracy of weighted-averaging to mismatches in the gradients and propose a new uncertainty-based scheme to improve the performance by reducing the mismatch. The connection also reveals implicit assumptions in other schemes such as averaging, task arithmetic, and Fisher-weighted averaging. Our new method gives consistent improvements for large language models and vision transformers, both in terms of performance and robustness to hyperparameters. Nico Daheim, Thomas Möllenhoff, Edoardo Maria Ponti, Iryna Gurevych, Mohammad Emtiyaz Khan |
ICLR | 1 |
| 2024 | Variational Learning is Effective for Large Deep NetworksabstractWe give extensive empirical evidence against the common belief that variational learning is ineffective for large neural networks. We show that an optimizer called Improved Variational Online Newton (IVON) consistently matches or outperforms Adam for training large networks such as GPT-2 and ResNets from scratch. IVON's computational costs are nearly identical to Adam but its predictive uncertainty is better. We show several new use cases of IVON where we improve finetuning and model merging in Large Language Models, accurately predict generalization error, and faithfully estimate sensitivity to data. We find overwhelming evidence that variational learning is effective. Code is available at https://github.com/team-approx-bayes/ivon. Yuesong Shen, Nico Daheim, Bai Cong, Peter Nickl, Gian Maria Marconi, Clement Bazan, Rio Yokota, Iryna Gurevych, Daniel Cremers, Mohammad Emtiyaz Khan, Thomas Möllenhoff |
ICML | 2 |
| 2024 | Elastic Weight Removal for Faithful and Abstractive Dialogue GenerationabstractNico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, Edoardo Ponti. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Nico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, Edoardo Maria Ponti |
NAACL-HLT | 1 |
| 2024 | Task-Oriented Document-Grounded Dialog Systems by HLTPR@RWTH for DSTC9 and DSTC10abstractThis paper summarizes our contributions to the document-grounded dialog tasks at the 9th and 10th Dialog System Technology Challenges (DSTC9 and DSTC10). In both iterations the task consists of three subtasks: first detect whether the current turn is knowledge seeking, second select a relevant knowledge document, and third generate a response grounded on the selected document. For DSTC9 we proposed different approaches to make the selection task more efficient. The best method, Hierarchical Selection, actually improves the results compared to the original baseline and gives a speedup of 24x. In the DSTC10 iteration of the task, the challenge was to adapt systems trained on written dialogs to perform well on noisy automatic speech recognition transcripts. Therefore, we proposed data augmentation techniques to increase the robustness of the models as well as methods to adapt the style of generated responses to fit well into the proceeding dialog. Additionally, we proposed a noisy channel model that allows for increasing the factuality of the generated responses. In addition to summarizing our previous contributions, in this work, we also report on a few small improvements and reconsider the automatic evaluation metrics for the generation task which have shown a low correlation to human judgments. David Thulke, Nico Daheim, Christian Dugast, Hermann Ney |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Opportunities and Challenges in Neural Dialog TutoringabstractJakub Macina, Nico Daheim, Lingzhi Wang, Tanmay Sinha, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Jakub Macina, Nico Daheim, Lingzhi Wang 0001, Tanmay Sinha, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan |
EACL | 2 |
| 2023 | Poor Man's Quality Estimation: Predicting Reference-Based MT Metrics Without the ReferenceabstractVilém Zouhar, Shehzaad Dhuliawala, Wangchunshu Zhou, Nico Daheim, Tom Kocmi, Yuchen Eleanor Jiang, Mrinmaya Sachan. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Vilém Zouhar, Shehzaad Dhuliawala, Wangchunshu Zhou, Nico Daheim, Tom Kocmi, Yuchen Eleanor Jiang, Mrinmaya Sachan |
EACL | 4 |