EDBT 2026 Demo / reviewers in the wild / expert
Martin Tutek
dblp:186/7079
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-5227-5397ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Trustworthy machine learning · 46% Language models and text generation · 33% Reinforcement learning · 8% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › machine unlearning
concept unlearning |
1.0 | 1 | 2026 | CRISP: Persistent Concept Unlearning via Sparse Autoencoders · ACL (1) 2026 |
Machine learning › Trustworthy machine learning
machine unlearning |
1.0 | 1 | 2026 | CRISP: Persistent Concept Unlearning via Sparse Autoencoders · ACL (1) 2026 |
Machine learning › Trustworthy machine learning
robustness |
1.0 | 1 | 2026 | PRAGWORLD: A Benchmark Evaluating LLMs' Local World Model Under Minimal Linguistic Alterations and Conversational Dynamics · AAAI 2026 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
1.0 | 1 | 2026 | PRAGWORLD: A Benchmark Evaluating LLMs' Local World Model Under Minimal Linguistic Alterations and Conversational Dynamics · AAAI 2026 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps · EMNLP 2025 |
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
faithfulness |
0.9 | 1 | 2025 | Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps · EMNLP 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | MIB: A Mechanistic Interpretability Benchmark · ICML 2025 |
Natural language and speech › Language models and text generation
knowledge editing |
0.9 | 1 | 2025 | Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps · EMNLP 2025 |
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
0.9 | 1 | 2025 | MIB: A Mechanistic Interpretability Benchmark · ICML 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
conditional reasoning |
0.8 | 1 | 2024 | Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs · EMNLP 2024 |
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection |
0.8 | 1 | 2024 | Out-of-Distribution Detection by Leveraging Between-Layer Transformation Smoothness · ICLR 2024 |
Natural language and speech › Language models and text generation
prompting |
0.8 | 1 | 2024 | Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs · EMNLP 2024 |
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering |
0.3 | 1 | 2026 | PRAGWORLD: A Benchmark Evaluating LLMs' Local World Model Under Minimal Linguistic Alterations and Conversational Dynamics · AAAI 2026 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2026 | CRISP: Persistent Concept Unlearning via Sparse Autoencoders · ACL (1) 2026 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering |
0.3 | 1 | 2025 | Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps · EMNLP 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.2 | 1 | 2024 | Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs · EMNLP 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.2 | 1 | 2024 | Out-of-Distribution Detection by Leveraging Between-Layer Transformation Smoothness · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
sparse autoencoder · 1.9layer regularization · 1.0interpretability · 1.0fine-tuning · 1.0activation suppression · 1.0unlearning · 0.9parametric faithfulness measurement · 0.9distributed alignment search · 0.9attribution patching · 0.9chain-of-prompts · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PRAGWORLD: A Benchmark Evaluating LLMs' Local World Model Under Minimal Linguistic Alterations and Conversational DynamicsabstractReal-world conversations are rich with pragmatic elements, such as entity mentions, references, and implicatures. Understanding such nuances is a requirement for successful natural communication, and often requires building a local _world model_ which encodes such elements and captures the dynamics of their evolving states. However, it is not well-understood whether language models (LMs) construct or maintain a robust implicit representation of conversations. In this work, we evaluate the ability of LMs to encode and update their internal world model in dyadic conversations and test their _malleability_ under linguistic alterations. To facilitate this, we apply seven minimal linguistic alterations to conversations sourced from popular conversational QA datasets and construct a benchmark with two variants (i.e., Manual and Synthetic) comprising yes-no questions. We evaluate nine open and one closed source LMs and observe that they struggle to maintain robust accuracy. Our analysis unveils that LMs struggle to memorize crucial details, such as tracking entities under linguistic alterations to conversations. We then propose a dual-perspective interpretability framework which identifies transformer layers that are _useful_ or _harmful_ and highlights linguistic alterations most influenced by harmful layers, typically due to encoding spurious signals or relying on shortcuts. Inspired by these insights, we propose two layer-regularization based fine-tuning strategies that suppress the effect of the harmful layers. Sachin Vashistha, Aryan Bibhuti, Atharva Naik, Martin Tutek, Somak Aditya |
AAAI | 4 |
| 2026 | CRISP: Persistent Concept Unlearning via Sparse AutoencodersabstractAs large language models (LLMs) are increasingly deployed in real-world applications, the need to selectively remove unwanted knowledge while preserving model utility has become paramount. Recent work has explored sparse autoencoders (SAEs) to perform precise interventions on monosemantic features. However, most SAE-based methods operate at inference time, which does not create persistent changes in the model's parameters. Such interventions can be bypassed or reversed by malicious actors with parameter access. We introduce CRISP, a parameter-efficient method for persistent concept unlearning using SAEs. CRISP automatically identifies salient SAE features across multiple layers and suppresses their activations. We experiment with two LLMs and show that our method outperforms prior approaches on safety-critical unlearning tasks from the WMDP benchmark, successfully removing harmful knowledge while preserving general and in-domain capabilities. Feature-level analysis reveals that CRISP achieves semantically coherent separation between target and benign concepts, allowing precise suppression of the target features. Tomer Ashuach, Dana Arad, Aaron Mueller, Martin Tutek, Yonatan Belinkov |
ACL (1) | 4 |
| 2025 | Measuring Chain of Thought Faithfulness by Unlearning Reasoning StepsabstractWhen prompted to think step-by-step, language models (LMs) produce a chain of thought (CoT), a sequence of reasoning steps that the model supposedly used to produce its prediction.Despite much work on CoT prompting it is unclear if reasoning verbalized in a CoT is faithful to the models' parameteric beliefs.We introduce a framework for measuring parametric faithfulness of generated reasoning, and propose Faithfulness by Unlearning Reasoning steps (FUR), an instance of this framework.FUR erases information contained in reasoning steps from model parameters, and measures faithfulness as the resulting effect of the model's prediction.Our experiments with four LMs and five multi-hop multi-choice question answering (MCQA) datasets show that FUR is frequently able to precisely change the underlying models' prediction for a given instance by unlearning key steps, indicating when a CoT is parametrically faithful.Further analysis shows that CoTs generated by models post-unlearning support different answers, hinting at a deeper effect of unlearning. 1 Martin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan Belinkov |
EMNLP | 1 |
| 2025 | MIB: A Mechanistic Interpretability BenchmarkabstractHow can we know whether new mechanistic interpretability methods achieve real improvements?
In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretability Benchmark, with two tracks spanning four tasks and five models. MIB favors methods that precisely and concisely recover relevant causal pathways or causal variables in neural language models. The circuit localization track compares methods that locate the model components---and connections between them---most important for performing a task (e.g., attribution patching or information flow routes). The causal variable track compares methods that featurize a hidden vector, e.g., sparse autoencoders (SAE) or distributed alignment search (DAS), and align those features to a task-relevant causal variable. Using MIB, we find that attribution and mask optimization methods perform best on circuit localization. For causal variable localization, we find that the supervised DAS method performs best, while SAEs features are not better than neurons, i.e., non-featurized hidden vectors. These findings illustrate that MIB enables meaningful comparisons, and increases our confidence that there has been real progress in the field. Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Iván Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fiotto-Kaufman, Tal Haklay, Michael Hanna 0001, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, Yonatan Belinkov |
ICML | 20 |
| 2024 | CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and CalibrationabstractIn recent years, large language models (LLMs) have shown remarkable capabilities at scale, particularly at generating text conditioned on a prompt.In our work, we investigate the use of LLMs to augment training data of smaller language models (SLMs) with automatically generated counterfactual (CF) instances -i.e.minimally altered inputs -in order to improve out-of-domain (OOD) performance of SLMs in the extractive question answering (QA) setup.We show that, across various LLM generators, such data augmentation consistently enhances OOD performance and improves model calibration for both confidence-based and rationaleaugmented calibrator models.Furthermore, these performance improvements correlate with higher diversity of CF instances in terms of their surface form and semantic content.Finally, we show that CF augmented models which are easier to calibrate also exhibit much lower entropy when assigning importance, indicating that rationale-augmented calibrators prefer concise explanations.1 Rachneet Sachdeva, Martin Tutek, Iryna Gurevych |
EACL (1) | 2 |
| 2024 | Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMsabstractReasoning is a fundamental component of language understanding.Recent prompting techniques, such as chain of thought, have consistently improved LLMs' performance on various reasoning tasks.Nevertheless, there is still little understanding of what triggers reasoning abilities in LLMs in the inference stage.In this paper, we investigate the effect of the input representation on the reasoning abilities of LLMs.We hypothesize that representing natural language tasks as code can enhance specific reasoning abilities such as entity tracking or logical reasoning.To study this, we propose code prompting, a methodology we operationalize as a chain of prompts that transforms a natural language problem into code and directly prompts the LLM using the generated code without resorting to external code execution.We find that code prompting exhibits a high-performance boost for multiple LLMs (up to 22.52 percentage points on GPT 3.5, 7.75 on Mixtral, and 16.78 on Mistral) across multiple conditional reasoning datasets.We then conduct comprehensive experiments to understand how the code representation triggers reasoning abilities and which capabilities are elicited in the underlying models.Our analysis on GPT 3.5 reveals that the code formatting of the input problem is essential for performance improvement.Furthermore, the code representation improves sample efficiency of in-context learning and facilitates state tracking of entities.1 Haritz Puerto, Martin Tutek, Somak Aditya, Xiaodan Zhu 0001, Iryna Gurevych |
EMNLP | 2 |
| 2024 | Out-of-Distribution Detection by Leveraging Between-Layer Transformation SmoothnessabstractEffective out-of-distribution (OOD) detection is crucial for reliable machine learning models, yet most current methods are limited in practical use due to requirements like access to training data or intervention in training. We present a novel method for detecting OOD data in Transformers based on transformation smoothness between intermediate layers of a network (BLOOD), which is applicable to pre-trained models without access to training data. BLOOD utilizes the tendency of between-layer representation transformations of in-distribution (ID) data to be smoother than the corresponding transformations of OOD data, a property that we also demonstrate empirically. We evaluate BLOOD on several text classification tasks with Transformer networks and demonstrate that it outperforms methods with comparable resource requirements. Our analysis also suggests that when learning simpler tasks, OOD data transformations maintain their original sharpness, whereas sharpness increases with more complex tasks. Fran Jelenic, Josip Jukic, Martin Tutek, Mate Puljiz, Jan Snajder |
ICLR | 3 |
| 2016 | Detecting and Ranking Conceptual Links between Texts Using a Knowledge BaseabstractRecent research has explored the use of Knowledge Bases (KBs) to represent documents as subgraphs of a KB concept graph and define metrics to characterize semantic relatedness of documents in terms of properties of the document concept graphs. However, none of the studies so far have examined to what degree such metrics capture a user-perceived relatedness of documents. Considering the users' explanations of how pairs of documents are related, the aim is to identify concepts in a KB graph that express the same notion of document relatedness. Our algorithm generates paths through the KB graph that originate from the terms in two documents. KB concepts where these paths intersect capture the semantic relatedness of the two starting terms and therefore the two documents. We consider how such intersecting concepts relate to the concepts in the users' explanations. The higher the users' concepts appear in the ranked list of intersecting concepts, the better the method in capturing the users' notion of document relatedness. Our experiments show that our approach outperforms a simpler graph method that uses properties of the concept nodes alone. Martin Tutek, Goran Glavas, Jan Snajder, Natasa Milic-Frayling, Bojana Dalbelo Basic |
CIKM | 1 |