VLDB 2026 Research / reviewers in the wild / expert
Catherine Chen 0001
dblp:05/5358-1
· DBLP profile ↗
9ranked-venue papers in the field
5as first author
9since 2021 · last 2026
0009-0009-8734-436XORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 9 (5 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tutorial on Mechanistic Interpretability
Catherine Chen 0001, Maria Heuss, Carsten Eickhoff |
ECIR (4) | 1 |
| 2026 | How Role-Play Shapes Relevance Judgment in Zero-Shot LLM Rankers
Yumeng Wang 0001, Jirui Qi, Catherine Chen 0001, Panagiotis Eustratiadis, Suzan Verberne |
ECIR (1) | 3 |
| 2026 | Second Workshop on Explainability in Information RetrievalabstractAs models grow more complex and societal demands for transparency increase with emerging regulations, explainability has become an increasingly important research area. However, despite its recognized relevance, progress in explainability research in information retrieval (IR) has been slower than in related fields. This full day workshop aims to advance research in explainable IR by providing a more in-depth platform to reflect on recent developments and facilitate discussions across both new and persistent challenges. Building upon the first edition of the workshop, which was a great success in bringing together multiple perspectives on explainability in IR, this second edition will focus on synthesizing a common agenda for the research community. Through a set of interactive activities, the workshop will bring together a diverse group of researchers to build a shared understanding of key tasks and challenges, and to help shape future directions for explainable IR research. The workshop will have as concrete outcomes a roadmap document and a special issue proposal for a journal issue on explainability in IR. Catherine Chen 0001, Maria Heuss, Tanya Chowdhury, James Allan 0001, Avishek Anand, Carsten Eickhoff, Suzan Verberne |
SIGIR | 1 |
| 2025 | MechIR: A Mechanistic Interpretability Framework for Information Retrieval
Andrew Parry, Catherine Chen 0001, Carsten Eickhoff, Sean MacAvaney |
ECIR (5) | 2 |
| 2025 | Workshop on Explainability in Information RetrievalabstractAs models grow more complex and societal demands for transparency increase with emerging regulations, explainability has become an even more important research area. However, despite its recognized relevance, explainability research in IR has seen slower progress than in related fields. This full day workshop aims to advance research in explainable information retrieval by providing a more in-depth platform to reflect on recent developments and facilitate discussions to address new and persistent challenges. Our goal is to bring together a diverse group of researchers to build a shared understanding of key tasks and challenges that will lay the foundation for the future of explainable IR research. Maria Heuss, Catherine Chen 0001, Avishek Anand, Carsten Eickhoff, Suzan Verberne |
SIGIR | 2 |
| 2025 | Towards Best Practices of Axiomatic Activation Patching in Information RetrievalabstractMechanistic interpretability research, which aims to uncover the internal processes of machine learning models, has gained significant attention. One state-of-the-art technique, activation patching, has been applied to analyzing neural ranker behavior in relation to information retrieval (IR) axioms. To date, however, this remains a rapidly evolving topic in IR, with no established methodology for measuring results or constructing datasets to ensure pronounced, robust, and consistent patching effects. In this study, based on experimental results, we provide recommendations on measuring patching effects and designing diagnostic datasets for investigating term frequency. We identify the rareness and informativeness of injected terms as a key factor influencing the magnitude of patching effects. Additionally, we find that low score differences between baseline and perturbed documents introduce significant noise, which can be mitigated by filtering or applying penalty scores to the metric. More generally, we provide practical recommendations for the reliable application of activation patching in IR, advancing future interpretability research of neural ranking models. Our code is available at https://github.com/polgrisha/best-practices-ir-patching. Gregory Polyakov, Catherine Chen 0001, Carsten Eickhoff |
SIGIR | 2 |
| 2024 | Evaluating Search System Explainability with Psychometrics and CrowdsourcingabstractAs information retrieval (IR) systems, such as search engines and conversational agents, become ubiquitous in various domains, the need for transparent and explainable systems grows to ensure accountability, fairness, and unbiased results. Despite recent advances in explainable AI and IR techniques, there is no consensus on the definition of explainability. Existing approaches often treat it as a singular notion, disregarding the multidimensional definition postulated in the literature. In this paper, we use psychometrics and crowdsourcing to identify human-centered factors of explainability in Web search systems and introduce SSE (Search System Explainability), an evaluation metric for explainable IR (XIR) search systems. In a crowdsourced user study, we demonstrate SSE's ability to distinguish between explainable and non-explainable systems, showing that systems with higher scores indeed indicate greater interpretability. We hope that aside from these concrete contributions to XIR, this line of work will serve as a blueprint for similar explainability evaluation efforts in other domains of machine learning and natural language processing. Catherine Chen 0001, Carsten Eickhoff |
SIGIR | 1 |
| 2024 | Axiomatic Causal Interventions for Reverse Engineering Relevance Computation in Neural Retrieval ModelsabstractNeural models have demonstrated remarkable performance across diverse ranking tasks. However, the processes and internal mechanisms along which they determine relevance are still largely unknown. Existing approaches for analyzing neural ranker behavior with respect to IR properties rely either on assessing overall model behavior or employing probing methods that may offer an incomplete understanding of causal mechanisms. To provide a more granular understanding of internal model decision-making processes, we propose the use of causal interventions to reverse engineer neural rankers, and demonstrate how mechanistic interpretability methods can be used to isolate components satisfying term-frequency axioms within a ranking model. We identify a group of attention heads that detect duplicate tokens in earlier layers of the model, then communicate with downstream heads to compute overall document relevance. More generally, we propose that this style of mechanistic analysis opens up avenues for reverse engineering the processes neural retrieval models use to compute relevance. This work aims to initiate granular interpretability efforts that will not only benefit retrieval model development and training, but ultimately ensure safer deployment of these models. Catherine Chen 0001, Jack Merullo, Carsten Eickhoff |
SIGIR | 1 |
| 2023 | Quantifying and Advancing Information Retrieval System ExplainabilityabstractAs information retrieval (IR) systems, such as search engines and conversational agents, become ubiquitous in various domains, the need for transparent and explainable systems grows to ensure accountability, fairness, and unbiased results. Despite many recent advances toward explainable AI and IR techniques, there is no consensus on what it means for a system to be explainable. Although a growing body of literature suggests that explainability is comprised of multiple subfactors [2, 5, 6], virtually all existing approaches treat it as a singular notion. Additionally, while neural retrieval models (NRMs) have become popular for their ability to achieve high performance[3, 4, 7, 8], research on the explainability of NRMs has been largely unexplored until recent years. Numerous questions remain unanswered regarding the most effective means of comprehending how these intricate models arrive at their decisions and the extent to which these methods will function efficiently for both developers and end-users. Catherine Chen 0001 |
SIGIR | 1 |