VLDB 2026 Research / reviewers in the wild / expert
Philipp Mondorf
dblp:371/2609
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 46% Trustworthy machine learning · 34% Knowledge representation and reasoning · 13% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
2.5 | 3 | 2025 | Reason to Rote: Rethinking Memorization in Reasoning · EMNLP 2025 Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models · ACL (1) 2025 Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning · ACL (1) 2024 |
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit analysis |
1.7 | 2 | 2025 | The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It · EMNLP 2025 Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models · ACL (1) 2025 |
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
1.7 | 2 | 2025 | The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It · EMNLP 2025 Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models · ACL (1) 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
logic-based reasoning |
1.5 | 2 | 2024 | Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models · EMNLP 2024 Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning · ACL (1) 2024 |
Natural language and speech › Language models and text generation
large language model reasoning |
1.1 | 2 | 2025 | Reason to Rote: Rethinking Memorization in Reasoning · EMNLP 2025 Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning · ACL (1) 2024 |
Natural language and speech › Language models and text generation › mathematical reasoning › numerical reasoning
number representation |
1.0 | 1 | 2026 | Language Models Learn Universal Representations of Numbers and Here's Why You Should Care · ACL (1) 2026 |
Natural language and speech › Language models and text generation › mathematical reasoning
numerical reasoning |
1.0 | 1 | 2026 | Language Models Learn Universal Representations of Numbers and Here's Why You Should Care · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis
error detection |
0.9 | 1 | 2025 | The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model › knowledge in language models › memorization
memorization in language models |
0.9 | 1 | 2025 | Reason to Rote: Rethinking Memorization in Reasoning · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model › knowledge in language models › memorization
memorization mechanisms |
0.9 | 1 | 2025 | Reason to Rote: Rethinking Memorization in Reasoning · EMNLP 2025 |
Natural language and speech › Language models and text generation › language modeling › language model architecture
transformer language model |
0.9 | 1 | 2025 | Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models · ACL (1) 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
deductive reasoning |
0.8 | 1 | 2024 | Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning · ACL (1) 2024 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.8 | 1 | 2024 | Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models · EMNLP 2024 |
Natural language and speech › Language models and text generation › large language model evaluation
reasoning benchmark |
0.8 | 1 | 2024 | Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models · EMNLP 2024 |
Natural language and speech › Language models and text generation › evaluation of language models
reasoning evaluation |
0.8 | 1 | 2024 | Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning · ACL (1) 2024 |
Machine learning › Representation and self-supervised learning › representation learning
emergent representations |
0.3 | 1 | 2026 | Language Models Learn Universal Representations of Numbers and Here's Why You Should Care · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
probing · 1.0synthetic reasoning datasets · 0.9probabilistic context-free grammar · 0.9circuit composition · 0.9circuit analysis · 0.9propositional logic · 0.8cognitive psychology analysis · 0.8benchmark construction · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Language Models Learn Universal Representations of Numbers and Here's Why You Should CareabstractMichal Štefánik, Timothee Mickus, Marek Kadlčík, Bertram Højer, Michal Spiegel, Raúl Vázquez, Aman Sinha, Josef Kuchař, Philipp Mondorf, Pontus Stenetorp. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Michal Stefánik, Timothee Mickus, Marek Kadlcík, Bertram Højer, Michal Spiegel, Raúl Vázquez, Aman Sinha 0002, Josef Kuchar, Philipp Mondorf, Pontus Stenetorp |
ACL (1) | 9 |
| 2025 | Circuit Compositions: Exploring Modular Structures in Transformer-Based Language ModelsabstractA fundamental question in interpretability research is to what extent neural networks, particularly language models, implement reusable functions through subnetworks that can be composed to perform more complex tasks.Recent advances in mechanistic interpretability have made progress in identifying circuits, which represent the minimal computational subgraphs responsible for a model's behavior on specific tasks.However, most studies focus on identifying circuits for individual tasks without investigating how functionally similar circuits relate to each other.To address this gap, we study the modularity of neural networks by analyzing circuits for highly compositional subtasks within a transformer-based language model.Specifically, given a probabilistic context-free grammar, we identify and compare circuits responsible for ten modular string-edit operations.Our results indicate that functionally similar circuits exhibit both notable node overlap and crosstask faithfulness.Moreover, we demonstrate that the circuits identified can be reused and combined through set operations to represent more complex functional model capabilities. Philipp Mondorf, Sondre Wold, Barbara Plank |
ACL (1) | 1 |
| 2025 | The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate ItabstractThe ability of large language models (LLMs) to validate their output and identify potential errors is crucial for ensuring robustness and reliability. However, current research indicates that LLMs struggle with self-correction, encountering significant challenges in detecting errors. While studies have explored methods to enhance self-correction in LLMs, relatively little attention has been given to understanding the models’ internal mechanisms underlying error detection. In this paper, we present a mechanistic analysis of error detection in LLMs, focusing on simple arithmetic problems. Through circuit analysis, we identify the computational subgraphs responsible for detecting arithmetic errors across four smaller-sized LLMs. Our findings reveal that all models heavily rely on \textit{consistency heads}\textemdash{}attention heads that assess surface-level alignment of numerical values in arithmetic solutions. Moreover, we observe that the models’ internal arithmetic computation primarily occurs in higher layers, whereas validation takes place in middle layers, before the final arithmetic results are fully encoded. This structural dissociation between arithmetic computation and validation seems to explain why smaller-sized LLMs struggle to detect even simple arithmetic errors. Leonardo Bertolazzi, Philipp Mondorf, Barbara Plank, Raffaella Bernardi |
EMNLP | 2 |
| 2025 | Reason to Rote: Rethinking Memorization in ReasoningabstractLarge language models readily memorize arbitrary training instances, such as label noise, yet they perform strikingly well on reasoning tasks.In this work, we investigate how language models memorize label noise, and why such memorization in many cases does not heavily affect generalizable reasoning capabilities.Using two controllable synthetic reasoning datasets with noisy labels, four-digit addition (FDA) and two-hop relational reasoning (THR), we discover a reliance of memorization on generalizable reasoning mechanisms: models continue to compute intermediate reasoning outputs even when retrieving memorized noisy labels, and intervening reasoning adversely affects memorization.We further show that memorization operates through distributed encoding, i.e., aggregating various inputs and intermediate results, rather than building a look-up mechanism from inputs to noisy labels.Moreover, our FDA case study reveals memorization occurs via outlier heuristics, where existing neuron activation patterns are slightly shifted to fit noisy labels.Together, our findings suggest that memorization of label noise in language models builds on, rather than overrides, the underlying reasoning mechanisms, shedding lights on the intriguing phenomenon of benign memorization.1 Yupei Du, Philipp Mondorf, Silvia Casola, Yuekun Yao, Robert Litschko, Barbara Plank |
EMNLP | 2 |
| 2024 | Comparing Inferential Strategies of Humans and Large Language Models in Deductive ReasoningabstractDeductive reasoning plays a pivotal role in the formulation of sound and cohesive arguments.It allows individuals to draw conclusions that logically follow, given the truth value of the information provided.Recent progress in the domain of large language models (LLMs) has showcased their capability in executing deductive reasoning tasks.Nonetheless, a significant portion of research primarily assesses the accuracy of LLMs in solving such tasks, often overlooking a deeper analysis of their reasoning behavior.In this study, we draw upon principles from cognitive psychology to examine inferential strategies employed by LLMs, through a detailed evaluation of their responses to propositional logic problems.Our findings indicate that LLMs display reasoning patterns akin to those observed in humans, including strategies like supposition following or chain construction.Moreover, our research demonstrates that the architecture and scale of the model significantly affect its preferred method of reasoning, with more advanced models tending to adopt strategies more frequently than less sophisticated ones.Importantly, we assert that a model's accuracy, that is the correctness of its final conclusion, does not necessarily reflect the validity of its reasoning process.This distinction underscores the necessity for more nuanced evaluation procedures in the field. Problem 1:Statements: 1.Either there is a white marble in the box or else there is a black marble in the box, but not both.2. Either there is a black marble in the box or else there is a pink marble in the box, but not both.3.There is a pink marble in the box if and only if there is a gray marble in the box.Conclusion: If there is a white marble in the box then there is a gray marble in the box.Response: Sure, I'd be happy to help!Here's my reasoning process for this problem:1. First, I'll start by analyzing the given statements: From Statement 1, we know that the box contains either a white marble or a black marble, but not both.From Statement 2, we know that the box contains either a black marble or a pink marble, but not both.From Statement 3, we know that if there is a pink marble in the box, then there must be a gray marble in the box.2. Now, let's consider the conclusion: If there is a white marble in the box then there is a gray marble in the box.Can we logically deduce this conclusion from the given statements?Well, if there is a white marble in the box, then by Statement 1, there cannot be a black marble in the box.If there is no black marble in the box, then by Statement 2, there must be a pink marble in the box.But if there is a pink marble in the box, then by Statement 3, there must be a gray marble in the box.So, yes, we can logically deduce the conclusion that if there is a white marble in the box, then there is a gray marble in the box. Philipp Mondorf, Barbara Plank |
ACL (1) | 1 |
| 2024 | Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language ModelsabstractKnights and knaves problems represent a classic genre of logical puzzles where characters either tell the truth or lie.The objective is to logically deduce each character's identity based on their statements.The challenge arises from the truth-telling or lying behavior, which influences the logical implications of each statement.Solving these puzzles requires not only direct deductions from individual statements, but the ability to assess the truthfulness of statements by reasoning through various hypothetical scenarios.As such, knights and knaves puzzles serve as compelling examples of suppositional reasoning.In this paper, we introduce TruthQuest, a benchmark for suppositional reasoning based on the principles of knights and knaves puzzles.Our benchmark presents problems of varying complexity, considering both the number of characters and the types of logical statements involved.Evaluations on TruthQuest show that large language models like Llama 3 and Mixtral-8x7B exhibit significant difficulties solving these tasks.A detailed error analysis of the models' output reveals that lower-performing models exhibit a diverse range of reasoning errors, frequently failing to grasp the concept of truth and lies.In comparison, more proficient models primarily struggle with accurately inferring the logical implications of potentially false statements. Philipp Mondorf, Barbara Plank |
EMNLP | 1 |