VLDB 2026 Research / reviewers in the wild / expert
Maciej Rybinski
dblp:157/9497 · also Maciek Rybinski
· DBLP profile ↗
10ranked-venue papers in the field
8as first author
9since 2021 · last 2026
0000-0002-0174-0567ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 9 (7 first)Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CTCL: A Cross-Language Benchmark for Matching Patients to Clinical TrialsabstractPublisher Copyright: © 2026 Owner/Author. Maciej Rybinski, Wojciech Kusa, Necva Bölücü, Georgios Peikos, Aditya Joshi 0001, Sarvnaz Karimi, Aitziber Atutxa, Javier Del Ser, Ahmet Bölücü, Monica Chierichetti, Pritam Dasgupta, Nicolás Jiménez García, Borja Pedruzo, Angelika Romanska, Ioulia Symeonidou |
SIGIR | 1 |
| 2025 | PEQQS: a Dataset for Probing Extractive Quantity-focused Question Answering from Scientific LiteratureabstractQuestion Answering (QA) and Information Retrieval (IR) play a crucial role in information-seeking pipelines implemented in many emerging AI research assistant applications. Large Language Models (LLMs) have demonstrated exceptional effectiveness on QA tasks, with Retrieval Augmented Generation (RAG) techniques often boosting the results. However, in many of those emerging applications, the onus of conducting the actual literature search falls on the user, i.e. the user searches for the relevant literature and the LLM-based assistant extracts the solicited answers from each of the user-supplied documents. The interplay between the quality of the user-conducted search and the quality of the final results remains understudied. Maciej Rybinski, Necva Bölücü, Huichen Yang, Stephen Wan 0001 |
CIKM | 1 |
| 2024 | An adaptive approach to noisy annotations in scientific information extraction
Necva Bölücü, Maciej Rybinski, Xiang Dai 0001, Stephen Wan 0001 |
Inf. Process. Manag. | 2 |
| 2023 | A Self-Learning Resource-Efficient Re-Ranking Method for Clinical Trials SearchabstractComplex search scenarios, such as those in biomedical settings, can be challenging. One such scenario is matching a patient's profile to relevant clinical trials. There are multiple criteria that should match for a document (clinical trial) to be considered relevant to a query (patient's profile represented with an admission note). While different neural ranking methods have been proposed for searching clinical trials, resource-efficient approaches to ranker training are less studied. A resource-efficient method uses training data in moderation. We propose a self-learning reranking method that achieves results comparable to those of more complicated, fully supervised, systems. Our experiments demonstrate our method's robustness and competitiveness compared to the state-of-the-art approaches in clinical trial search. Maciej Rybinski, Sarvnaz Karimi |
CIKM | 1 |
| 2023 | SciHarvester: Searching Scientific Documents for Numerical ValuesabstractA challenge for search technologies is to support scientific literature surveys that present overviews of the reported numerical values documented for specific physical properties. We present SciHarvester, a system tailored to address this problem for agronomic science. It provides an interface to search PubAg documents, allowing complex queries involving restrictions on numerical values. SciHarvester identifies relevant documents and generates overview of reported parameter values. The system allows interrogation of the results to explain the system's performance. Our evaluations demonstrate the promise of incorporating information extraction techniques with the use of neural scoring mechanisms. Maciej Rybinski, Stephen Wan 0001, Sarvnaz Karimi, Cécile Paris, Brian Jin, Neil I. Huth, Peter J. Thorburn, Dean P. Holzworth |
SIGIR | 1 |
| 2022 | A2A-API: A Prototype for Biomedical Information Retrieval Research and BenchmarkingabstractFinding relevant literature is crucial for biomedical research and in the practice of evidence-based medicine, making biomedical search an important application area within the field of information retrieval. This is recognised by the broader IR community, and in particular by the organisers of Text Retrieval Conference (TREC) as early as 2003. While TREC provides crucial evaluation resources, to get started in biomedical IR one needs to tackle an important software engineering hurdle of parsing, indexing, and deploying several large document collections. Moreover, many newcomers to the field often face a steep learning curve, where theoretical concepts are tangled up with technical aspects. Finally, many of the existing baselines and systems are difficult to reproduce. Maciej Rybinski, Liam Watts, Sarvnaz Karimi |
SIGIR | 1 |
| 2021 | SearchEHR: A Family History Search System for Clinical Decision SupportabstractFinding patients with specific clinical conditions, such as having a familial disease history of diabetes, is an important task for clinical decision support. Clinical notes in Electronic Health Records (EHR), which document the patient medical history and familial disease history, are valuable resources for patient cohort selection. However, such information is difficult to discover in clinical text, and full-text search techniques often fail due to the unique characteristics of clinical language. We describe a system---SearchEHR---that combines Natural Language Processing (NLP) and Information Retrieval (IR) techniques to facilitate utilising clinical notes to find cohorts of patients, with a special focus on family disease history. Xiang Dai 0001, Maciej Rybinski, Sarvnaz Karimi |
CIKM | 2 |
| 2021 | Will Sorafenib Help?: Treatment-aware Reranking in Precision Medicine SearchabstractHigh-quality evidence from the biomedical literature is crucial for decision making of oncologists who treat cancer patients. Search for evidence on a specific treatment for a patient is the challenge set by the precision medicine track of TREC in 2020. To address this challenge, we propose a two-step method to incorporate treatment into the query formulation and ranking. Training of such ranking function uses a zero-shot setup to incorporate the novel focus on treatments which did not exist in any of the previous TREC tracks. Our treatment-aware neural reranking approach, FAT, achieves state-of-the-art effectiveness for TREC Precision Medicine 2020. Our analysis indicates that the BERT-based rerankers automatically learn to score documents through identifying concepts relevant to precision medicine, similar to hand-crafted heuristics successful in the earlier studies. Maciej Rybinski, Sarvnaz Karimi |
CIKM | 1 |
| 2021 | Science2Cure: A Clinical Trial Search PrototypeabstractWith the advances in precision medicine, identifying clinical trials relevant to a specific patient profile becomes more challenging. Often very specific molecular-level patient features need to be matched for the trial to be deemed relevant. Clinical trials contain strict inclusion and exclusion criteria, often written in free-text. Patients profiles are also semi-structured, with some important information hidden in clinical notes. We present a search system that given a patient profile searches over clinical trials for potential matches. It enables the users to leverage the powerful querying language that comes with Apache Lucene query syntax in combination with state-of-the-art Divergence From Randomness retrieval coupled with a BERT-based neural ranking component. This system aims to assist in clinical decision making. Maciej Rybinski, Sarvnaz Karimi, Aleney Khoo |
SIGIR | 1 |
| 2017 | DomESA: a novel approach for extending domain-oriented lexical relatedness calculations with domain-specific semantics
Maciej Rybinski, José Francisco Aldana-Montes |
J. Intell. Inf. Syst. | 1 |