EDBT 2026 Demo / reviewers in the wild / expert
Tim Hagen
dblp:389/9527
· DBLP profile ↗
10ranked-venue papers in the field
1as first author
10since 2021 · last 2026
0009-0000-4854-7249ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 10 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating the Efficiency and Effectiveness of Learned Sparse Retrieval with the lsr_benchmark
Maik Fröbe, Ferdinand Schlatt, Cosimo Rulli, Tim Hagen, Jan Heinrich Merker, Gijs Hendriksen, Carlos Eduardo Rosar Kós Lassance, Franco Maria Nardini, Rossano Venturini, Martin Potthast |
ECIR (4) | 4 |
| 2026 | Overview of Touché 2026: Argumentation Systems - Extended Abstract
Johannes Kiesel, Marc Feger, Tim Hagen, Sebastian Heineking, Maximilian Heinrich, Maik Fröbe, Katarina Boland, Wilhelm Pertsch, Julia Romberg, Ines Zelch, Stefan Dietze, Matthias Hagen, Martin Potthast, Benno Stein 0001 |
ECIR (4) | 3 |
| 2026 | Auto-Judge: A Cross-Task Benchmark for Comparing LLM Judges for Citation-Grounded RAG SystemsabstractWe present the Auto-Judge resource for the meta-evaluation of automated LLM judges, especially judges that evaluate Retrieval-Augmented Generation (RAG) systems that ground their response with citations. The resource couples (i) a data release of topics, pooled RAG responses, and human judgments, with (ii) a standardized protocol and software infrastructure for implementing "LLM-as-a-judge" methods in a reproducible and extensible way, including support for parameter sweeps and variant tracking. Naghmeh Farzi, Tim Hagen, Eugene Yang 0001, Maik Fröbe, Ronak Pradeep, Hossein A. Rahmani, Xi Wang 0012, Oleg Zendel, Martin Potthast, Laura Dietz |
SIGIR | 2 |
| 2026 | ReNeuIR at SIGIR 2026: The Fifth Workshop on Reaching Efficiency in Neural Information RetrievalabstractThe lack of efficiency in neural information retrieval remains one of the primary obstacles to deploying neural retrieval models as a first-stage retriever at scale. While recent tools have improved the standardized measurement of model efficiency, substantial progress is still needed to enable systematic comparative evaluation, for example, in terms of standards for systems and hardware configurations, cloud-based evaluation, benchmarks, and reproducibility. Beyond measurement, the IR community also needs stronger incentives to move in this direction, such as cost-efficiency as a review criterion or as efficiency and/or effectiveness measures in shared tasks, related teaching materials, efficiency-oriented user studies, and specialized awards for efficiency achievements. In particular, developing more efficient variants of highly effective retrieval algorithms should become an admissible research goal for PhD students if cost-efficiency is to become a first-class design objective in~IR. With ReNeuIR, we have established a recurring forum where these questions and new ideas are discussed and where the community comes together to collaboratively evaluate and improve efficiency benchmarking frameworks---most notably through the organization of a shared task focused on efficiency and reproducibility. Maik Fröbe, Tim Hagen, Franco Maria Nardini, Martin Potthast |
SIGIR | 2 |
| 2025 | Overview of Touché 2025: Argumentation Systems - Extended Abstract
Johannes Kiesel, Çagri Çöltekin, Marcel Gohsen, Sebastian Heineking, Maximilian Heinrich, Maik Fröbe, Tim Hagen, Mohammad Aliannejadi, Tomaz Erjavec, Matthias Hagen, Matyás Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Harrisen Scells, Ines Zelch, Martin Potthast, Benno Stein 0001 |
ECIR (5) | 7 |
| 2025 | Web-Scale Retrieval Experimentation with chatnoir-pyterrier
Jan Heinrich Merker, Janek Bevendorff, Maik Fröbe, Tim Hagen, Harrisen Scells, Matti Wiegmann, Benno Stein 0001, Matthias Hagen, Martin Potthast |
ECIR (5) | 4 |
| 2025 | ReNeuIR at SIGIR 2025: The Fourth Workshop on Reaching Efficiency in Neural Information RetrievalabstractMeasuring effectiveness and efficiency in information retrieval has a strong empirical background. While modern retrieval systems substantially improve effectiveness, the community has not yet agreed on how to measure efficiency, making it difficult to contrast effectiveness and efficiency fairly. Efficiency-oriented system comparisons are difficult due to factors such as hardware configurations, software versioning, and experimental settings. Efficiency affects users, researchers, and the environment and can be measured in many dimensions beyond time and space, such as resource consumption, water usage, and sample efficiency. Analyzing the efficiency of algorithms and their trade-off with effectiveness requires revisiting and establishing new standards and principles, from defining relevant concepts to designing new measures and guidelines to assess the findings' significance. ReNeuIR's fourth iteration aims to bring the community together to debate these questions and collaboratively test and improve benchmarking frameworks for efficiency based on discussions and collaborations of its previous iterations, including a shared task focused on efficiency and reproducibility. Sebastian Bruch 0001, Maik Fröbe, Tim Hagen, Franco Maria Nardini, Martin Potthast |
SIGIR | 3 |
| 2025 | The Viability of Crowdsourcing for RAG EvaluationabstractHow good are humans at writing and judging responses in retrieval-augmented generation (RAG) scenarios? To answer this question, we investigate the efficacy of crowdsourcing for RAG through two complementary studies: response writing and response utility judgment. Our new Webis Crowd RAG Corpus 2025 (Webis-CrowdRAG-25) consists of 903 human-written and 903 LLM-generated responses for the 301 topics of the TREC 2024 RAG~track, with each response composed according to one of the three discourse styles 'bullet list', 'essay', or 'news'. For a selection of 65 topics, the corpus further contains 47,320 pairwise human judgments and 10,556 pairwise LLM judgments across seven utility dimensions (e.g., coverage and coherence). Our analyses give insights into human writing behavior for RAG and the viability of crowdsourcing for RAG evaluation. We find that human pairwise judgments provide reliable and cost-effective results. This is much less the case for LLM-based pairwise and human/LLM-based pointwise judgments, nor for automated comparisons with human-written reference responses. All our data and tools are freely available. Lukas Gienapp, Tim Hagen, Maik Fröbe, Matthias Hagen, Benno Stein 0001, Martin Potthast, Harrisen Scells |
SIGIR | 2 |
| 2025 | TIREx Tracker: The Information Retrieval Experiment TrackerabstractThe reproducibility and transparency of retrieval experiments depends on the availability of information about the experimental setup. However, the manual collection of experiment metadata can be tedious, error-prone, and inconsistent, which calls for an automated systematic collection. Expanding ir_metadata, we present the TIREx tracker, a tool that records hardware configurations, power/CPU/RAM/GPU usage, and experiment/system versions. Implemented as a lightweight platform-independent C binary, the TIREx tracker integrates seamlessly into Python, Java, or C/C++ workflows and can be easily integrated into shard task submissions, as we demonstrate for the TIRA/TIREx platform. Code, binaries, and documentation of the TIREx tracker are publicly available at https://github.com/tira-io/tirex-tracker. Tim Hagen, Maik Fröbe, Jan Heinrich Merker, Harrisen Scells, Matthias Hagen, Martin Potthast |
SIGIR | 1 |
| 2025 | TITE: Token-Independent Text Encoder for Information RetrievalabstractTransformer-based retrieval approaches typically use the contextualized embedding of the first input token as a dense vector representation for queries and documents. The embeddings of all other tokens are also computed but then discarded, wasting resources. In this paper, we propose the Token-Independent Text Encoder (TITE) as a more efficient modification of the backbone encoder model. Using an attention-based pooling technique, TITE iteratively reduces the sequence length of hidden states layer by layer so that the final output is already a single sequence representation vector. Our empirical analyses on the TREC 2019 and 2020 Deep Learning tracks and the BEIR benchmark show that TITE is on par in terms of effectiveness compared to standard bi-encoder retrieval models while being up to 3.3 times faster at encoding queries and documents. Our code is available at: https://github.com/webis-de/SIGIR-25. Ferdinand Schlatt, Tim Hagen, Martin Potthast, Matthias Hagen |
SIGIR | 2 |