Wojciech Kusa

dblp:203/9078 · DBLP profile ↗
← Back
10ranked-venue papers in the field
4as first author
10since 2021 · last 2026
0000-0003-4420-4147ORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 10 (4 first)
YearPublicationVenuePosition
2026 Talmud-IR: A Talmud-Inspired Interface for Discussing RAG Response Quality
Wojciech Kusa, Niklas Deckers, Maik Fröbe, Laura Dietz, Birte Platow, Mark Sanderson
ECIR (4)1
2026 CTCL: A Cross-Language Benchmark for Matching Patients to Clinical Trials
abstract
Publisher Copyright: © 2026 Owner/Author.
Maciej Rybinski, Wojciech Kusa, Necva Bölücü, Georgios Peikos, Aditya Joshi 0001, Sarvnaz Karimi, Aitziber Atutxa, Javier Del Ser, Ahmet Bölücü, Monica Chierichetti, Pritam Dasgupta, Nicolás Jiménez García, Borja Pedruzo, Angelika Romanska, Ioulia Symeonidou
SIGIR2
2026 The LLM Effect on IR Benchmarks: A Meta-Analysis of Effectiveness, Baselines, and Contamination
abstract
Benchmark collections have long enabled controlled comparison and cumulative progress in Information Retrieval (IR). However, prior meta-analyses show that reported effectiveness gains often fail to accumulate, in part due to weak or outdated baselines. Large language models (LLMs) are increasingly used in retrieval pipelines, yet their impact on established IR benchmarks has not been systematically analyzed. We analyze 179 publications reporting on the TREC Robust04 collection and the TREC Deep Learning 2020 (DL20) Passage Retrieval benchmark, using ACM Digital Library keyword search supplemented by citation-graph backtracking for Robust04. We observe what we term an LLM effect: recent systems incorporating LLM components achieve 8.8% higher nDCG@10 on DL20 than the best TREC 2020 result and 11.9% higher on Robust04 than the strongest pre-2024 result. However, evaluation practice has shifted from MAP to nDCG@10 over the same window, and our adaptation of the Data Contamination Quiz reveals 12-41% contamination across two widely-used LLM rerankers. Filtering contaminated topics shows no statistically significant effectiveness difference, but small samples and the uncertainty of adapting contamination detection to reranking prevent us from ruling out memorization as a contributing factor. We read the LLM effect as real but unverified: visible in the aggregate numbers, but not cleanly separable from metric drift or pretraining overlap.
Moritz Staudinger, Wojciech Kusa, Allan Hanbury
SIGIR2
2025 Compare: A Framework for Scientific Comparisons
abstract
Navigating the vast and rapidly increasing sea of academic publications to identify institutional synergies, benchmark research contributions and pinpoint key research contributions has become an increasingly daunting task, especially with the current exponential increase in new publications. Existing tools provide useful overviews or single-document insights, but none supports structured, qualitative comparisons across institutions or publications. To address this, we demonstrate Compare, a novel framework that tackles this challenge by enabling sophisticated long-context comparisons of scientific contributions. Compare empowers users to explore and analyze research overlaps and differences at both the institutional and publication granularity, all driven by user-defined questions and automatic retrieval over online resources. For this we leverage on Retrieval-Augmented Generation over evolving data sources to foster long context knowledge synthesis. Unlike traditional scientometric tools, Compare goes beyond quantitative indicators by providing qualitative, citation-supported comparisons.
Moritz Staudinger, Wojciech Kusa, Matteo Cancellieri, David Pride, Petr Knoth, Allan Hanbury
CIKM2
2025 ASPIRE: Assistive System for Performance Evaluation in IR
Georgios Peikos, Wojciech Kusa, Symeon Symeonidis
ECIR (5)2
2025 TimIR: Time-Traveling Through IR History
Moritz Staudinger, Wojciech Kusa, Florina Piroi, Andreas Rauber, Allan Hanbury
ECIR (4)2
2023 CRUISE-Screening: Living Literature Reviews Toolbox
abstract
Keeping up with research and finding related work is still a time-consuming task for academics. Researchers sift through thousands of studies to identify a few relevant ones. Automation techniques can help by increasing the efficiency and effectiveness of this task. To this end, we developed CRUISE-Screening, a web-based application for conducting living literature reviews -- a type of literature review that is continuously updated to reflect the latest research in a particular field. CRUISE-Screening is connected to several search engines via an API, which allows for updating the search results periodically. Moreover, it can facilitate the process of screening for relevant publications by using text classification and question answering models. CRUISE-Screening can be used both by researchers conducting literature reviews and by those working on automating the citation screening process to validate their algorithms. The application is open-source, and a demo is available under this URL: https://citation-screening.ec.tuwien.ac.at.
Wojciech Kusa, Petr Knoth, Allan Hanbury
CIKM1
2023 VoMBaT: A Tool for Visualising Evaluation Measure Behaviour in High-Recall Search Tasks
abstract
The objective of High-Recall Information Retrieval (HRIR) is to retrieve as many relevant documents as possible for a given search topic. One approach to HRIR is Technology-Assisted Review (TAR), which uses information retrieval and machine learning techniques to aid the review of large document collections. TAR systems are commonly used in legal eDiscovery and systematic literature reviews. Successful TAR systems are able to find the majority of relevant documents using the least number of assessments. Commonly used retrospective evaluation assumes that the system achieves a specific, fixed recall level first, and then measures the precision or work saved (e.g., precision at r% recall). This approach can cause problems related to understanding the behaviour of evaluation measures in a fixed recall setting. It is also problematic when estimating time and money savings during technology-assisted reviews.
Wojciech Kusa, Aldo Lipani, Petr Knoth, Allan Hanbury
SIGIR1
2022 Automation of Citation Screening for Systematic Literature Reviews Using Neural Networks: A Replicability Study
Wojciech Kusa, Allan Hanbury, Petr Knoth
ECIR (1)1
2022 ORCAS-I: Queries Annotated with Intent using Weak Supervision
abstract
User intent classification is an important task in information retrieval. In this work, we introduce a revised taxonomy of user intent. We take the widely used differentiation between navigational, transactional and informational queries as a starting point, and identify three different sub-classes for the informational queries: instrumental, factual and abstain. The resulting classification of user queries is more fine-grained, reaches a high level of consistency between annotators, and can serve as the basis for an effective automatic classification process. The newly introduced categories help distinguish between types of queries that a retrieval system could act upon, for example by prioritizing different types of results in the ranking.
Daria Alexander, Wojciech Kusa, Arjen P. de Vries
SIGIR2