EDBT 2026 Demo / reviewers in the wild / expert
Sarvnaz Karimi
dblp:85/1259
· DBLP profile ↗
19ranked-venue papers in the field
4as first author
9since 2021 · last 2026
0000-0002-4927-3937ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 17 (4 first)Database Systems & Data Management · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CTCL: A Cross-Language Benchmark for Matching Patients to Clinical TrialsabstractPublisher Copyright: © 2026 Owner/Author. Maciej Rybinski, Wojciech Kusa, Necva Bölücü, Georgios Peikos, Aditya Joshi 0001, Sarvnaz Karimi, Aitziber Atutxa, Javier Del Ser, Ahmet Bölücü, Monica Chierichetti, Pritam Dasgupta, Nicolás Jiménez García, Borja Pedruzo, Angelika Romanska, Ioulia Symeonidou |
SIGIR | 6 |
| 2026 | RegionSLM: Region-aware Question Answering on Document ScreenshotsabstractReal-world document question-answering that relies on screenshots, such as bills and forms, requires evidence that is often spatially localised and visually cluttered. However, most Screenshot Language Models (SLMs) encode the entire page holistically and rely on implicit attention to ''find'' relevant content, which limits both accuracy and efficiency. We present RegionSLM, a region-aware SLM designed to explicitly connect the question to its supporting regions. RegionSLM has two key components: (1) a patch-relevance router that learns a query–region relevance distribution, enabling the model to produce a box-free relevance prior at inference; and (2) Relevance-Guided Region Pooling (RGRP), a query-conditional attention–pooling module that aggregates dense features into a small set of region tokens, which preserves grounding signals while reducing computational overhead. To support training and evaluation, we further curate ReDoc, a region-supervised corpus with 105k documents and 350k question-answer pairs, obtained via a question-guided two-step filtering procedure. Extensive experiments on 12 datasets demonstrate that explicitly learning query–region relevance and pooling it into compact region tokens is an effective and practical recipe for document retrieval and understanding. Chao Wang 0102, Hehe Fan, Huichen Yang, Sarvnaz Karimi, Lina Yao 0001, Yi Yang 0001 |
SIGIR | 4 |
| 2026 | The Second Workshop on Evaluation of Multimodal GenerationabstractMultimodal generation and retrieval systems are increasingly central to modern information retrieval, powering retrieval-augmented generation (RAG), multimodal search, recommendation, and knowledge intensive applications. Despite rapid progress in multimodal large language models (MLLMs), robust and principled evaluation of multimodal generation and retrieval remains a major open challenge for the IR community. This workshop aims to foster discussions and research efforts by bringing together researchers and practitioners in information retrieval, natural language processing, computer vision, and multimodal AI. Our goal is to establish evaluation methods for multimodal research and advance research efforts in this direction. Wei Zhang 0098, Xiang Dai 0001, Sarvnaz Karimi, Desmond Elliott, Biaoyan Fang, Mong Yuan Sim |
SIGIR | 3 |
| 2023 | A Self-Learning Resource-Efficient Re-Ranking Method for Clinical Trials SearchabstractComplex search scenarios, such as those in biomedical settings, can be challenging. One such scenario is matching a patient's profile to relevant clinical trials. There are multiple criteria that should match for a document (clinical trial) to be considered relevant to a query (patient's profile represented with an admission note). While different neural ranking methods have been proposed for searching clinical trials, resource-efficient approaches to ranker training are less studied. A resource-efficient method uses training data in moderation. We propose a self-learning reranking method that achieves results comparable to those of more complicated, fully supervised, systems. Our experiments demonstrate our method's robustness and competitiveness compared to the state-of-the-art approaches in clinical trial search. Maciej Rybinski, Sarvnaz Karimi |
CIKM | 3 |
| 2023 | SciHarvester: Searching Scientific Documents for Numerical ValuesabstractA challenge for search technologies is to support scientific literature surveys that present overviews of the reported numerical values documented for specific physical properties. We present SciHarvester, a system tailored to address this problem for agronomic science. It provides an interface to search PubAg documents, allowing complex queries involving restrictions on numerical values. SciHarvester identifies relevant documents and generates overview of reported parameter values. The system allows interrogation of the results to explain the system's performance. Our evaluations demonstrate the promise of incorporating information extraction techniques with the use of neural scoring mechanisms. Maciej Rybinski, Stephen Wan 0001, Sarvnaz Karimi, Cécile Paris, Brian Jin, Neil I. Huth, Peter J. Thorburn, Dean P. Holzworth |
SIGIR | 3 |
| 2022 | A2A-API: A Prototype for Biomedical Information Retrieval Research and BenchmarkingabstractFinding relevant literature is crucial for biomedical research and in the practice of evidence-based medicine, making biomedical search an important application area within the field of information retrieval. This is recognised by the broader IR community, and in particular by the organisers of Text Retrieval Conference (TREC) as early as 2003. While TREC provides crucial evaluation resources, to get started in biomedical IR one needs to tackle an important software engineering hurdle of parsing, indexing, and deploying several large document collections. Moreover, many newcomers to the field often face a steep learning curve, where theoretical concepts are tangled up with technical aspects. Finally, many of the existing baselines and systems are difficult to reproduce. Maciej Rybinski, Liam Watts, Sarvnaz Karimi |
SIGIR | 3 |
| 2021 | SearchEHR: A Family History Search System for Clinical Decision SupportabstractFinding patients with specific clinical conditions, such as having a familial disease history of diabetes, is an important task for clinical decision support. Clinical notes in Electronic Health Records (EHR), which document the patient medical history and familial disease history, are valuable resources for patient cohort selection. However, such information is difficult to discover in clinical text, and full-text search techniques often fail due to the unique characteristics of clinical language. We describe a system---SearchEHR---that combines Natural Language Processing (NLP) and Information Retrieval (IR) techniques to facilitate utilising clinical notes to find cohorts of patients, with a special focus on family disease history. Xiang Dai 0001, Maciej Rybinski, Sarvnaz Karimi |
CIKM | 3 |
| 2021 | Will Sorafenib Help?: Treatment-aware Reranking in Precision Medicine SearchabstractHigh-quality evidence from the biomedical literature is crucial for decision making of oncologists who treat cancer patients. Search for evidence on a specific treatment for a patient is the challenge set by the precision medicine track of TREC in 2020. To address this challenge, we propose a two-step method to incorporate treatment into the query formulation and ranking. Training of such ranking function uses a zero-shot setup to incorporate the novel focus on treatments which did not exist in any of the previous TREC tracks. Our treatment-aware neural reranking approach, FAT, achieves state-of-the-art effectiveness for TREC Precision Medicine 2020. Our analysis indicates that the BERT-based rerankers automatically learn to score documents through identifying concepts relevant to precision medicine, similar to hand-crafted heuristics successful in the earlier studies. Maciej Rybinski, Sarvnaz Karimi |
CIKM | 2 |
| 2021 | Science2Cure: A Clinical Trial Search PrototypeabstractWith the advances in precision medicine, identifying clinical trials relevant to a specific patient profile becomes more challenging. Often very specific molecular-level patient features need to be matched for the trial to be deemed relevant. Clinical trials contain strict inclusion and exclusion criteria, often written in free-text. Patients profiles are also semi-structured, with some important information hidden in clinical notes. We present a search system that given a patient profile searches over clinical trials for potential matches. It enables the users to leverage the powerful querying language that comes with Apache Lucene query syntax in combination with state-of-the-art Divergence From Randomness retrieval coupled with a BERT-based neural ranking component. This system aims to assist in clinical decision making. Maciej Rybinski, Sarvnaz Karimi, Aleney Khoo |
SIGIR | 2 |
| 2020 | Beyond mean rating: Probabilistic aggregation of star ratings based on helpfulnessabstractAbstract The star‐rating mechanism of customer reviews is used universally by the online population to compare and select merchants, movies, products, and services. The consensus opinion from aggregation of star ratings is used as a proxy for item quality. Online reviews are noisy and effective aggregation of star ratings to accurately reflect the “true quality” of products and services is challenging. The mean‐rating aggregation model is widely used and other aggregation models are also proposed. These existing aggregation models rely on a large number of reviews to tolerate noise. However, many products rarely have reviews. We propose probabilistic aggregation models for review ratings based on the Dirichlet distribution to combat data sparsity in reviews. We further propose to exploit the “helpfulness” social information and time to filter noisy reviews and effectively aggregate ratings to compute the consensus opinion. Our experiments on an Amazon data set show that our probabilistic aggregation models based on “helpfulness” achieve better performance than the statistical and heuristic baseline approaches. Wenyi Tay, Xiuzhen Zhang 0001, Sarvnaz Karimi |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2019 | An Experimentation Platform for Precision MedicineabstractPrecision medicine - where data from patients, their genes, their lifestyles and the available treatments and their combination are taken into account for finding a suitable treatment - requires searching the biomedical literature and other resources such as clinical trials with the patients' information. The retrieved information could then be used in curating data for clinicians for decision-making. We present information retrieval researchers with an on-line system which enables experimentation in search for precision medicine within the framework provided by the TREC Precision Medicine (PM) track. A number of query and document processing and ranking approaches are provided. These include some ofthe most promising gene mention expansion methods, as well as learning-to-rank using neural networks. Sarvnaz Karimi, Brian Jin |
SIGIR | 2 |
| 2018 | A2A: Benchmark Your Clinical Decision Support SearchabstractClinical Decision Support (CDS) systems aim to assist clinicians in their daily decision-making related to diagnosis, tests, and treatments of patients by providing relevant evidence from the scientific literature. This promise however is yet to be fulfilled, with search for relevant literature for a given patient condition still being an active research topic. The TREC CDS track was designed to address this research gap. We developed a platform to facilitate experimentation and hypothesis testing for information retrieval researchers working on this topic. It provides a large range of query and document processing techniques that are explored in the biomedical search domain. Sarvnaz Karimi, Falk Scholer, Brian Jin, Sara Falamaki |
SIGIR | 1 |
| 2016 | Parallel Duplicate Detection in Adverse Drug Reaction Databases with SparkabstractThe World Health Organization (WHO) and drug regulators in many countries maintain databases for adverse drug reaction reports. Data duplication is a significant problem in such databases as reports often come from a variety of sources. Most duplicate detection techniques either have limitations on handling large amount of data or lack effective means to deal with data with imbalanced label distribution. In this paper, we propose a scalable duplicate detection method built on top of Spark to address these problems. Our method uses the kNN (k nearest neighbors) classifier to identify labelled report pairs that are most useful for classifying new report pairs. To deal with the high computational cost of kNN, we partition the labelled data into clusters for parallel computing. We give a method to minimize the crosscluster kNN search. Our experimental results show that the proposed method is able to produce robust duplicate detection results and scalable performance. Chen Wang 0008, Sarvnaz Karimi |
EDBT | 2 |
| 2012 | ESA: emergency situation awareness via microbloggersabstractDuring a disastrous event, such as an earthquake or river flooding, information on what happened, who was affected and how, where help is needed, and how to aid people who were affected, is crucial. While communication is important in such times of crisis, damage to infrastructure such as telephone lines makes it difficult for authorities and victims to communicate. Microblogging has played a critical role as an important communication platform during crises when other media has failed. We demonstrate our ESA (Emergency Situation Awareness) system that mines microblogs in real-time to extract and visualise useful information about incidents and their impact on the community in order to equip the right authorities and the general public with situational awareness. Jie Yin 0001, Sarvnaz Karimi, Bella Robinson, Mark A. Cameron |
CIKM | 2 |
| 2012 | Quantifying the impact of concept recognition on biomedical information retrieval
Sarvnaz Karimi, Justin Zobel, Falk Scholer |
Inf. Process. Manag. | 1 |
| 2011 | Diverse retrieval via greedy optimization of expected 1-call@k in a latent subtopic relevance modelabstractIt has been previously observed that optimization of the [email protected] relevance objective (i.e., a set-based objective that is 1 if at least one document is relevant, otherwise 0) empirically correlates with diverse retrieval. In this paper, we proceed one step further and show theoretically that greedily optimizing expected [email protected] w.r.t. a latent subtopic model of binary relevance leads to a diverse retrieval algorithm sharing many features of existing diversification approaches. This new result is complementary to a variety of diverse retrieval algorithms derived from alternate rank-based relevance criteria such as average precision and reciprocal rank. As such, the derivation presented here for expected [email protected] provides a novel theoretical perspective on the emergence of diversity via a latent subtopic model of relevance --- an idea underlying both ambiguous and faceted subtopic retrieval that have been used to motivate diverse retrieval. Scott Sanner, Shengbo Guo, Thore Graepel, Sadegh Kharazmi, Sarvnaz Karimi |
CIKM | 5 |
| 2011 | Domain expert topic familiarity and search behaviorabstractUsers of information retrieval systems employ a variety of strategies when searching for information. One factor that can directly influence how searchers go about their information finding task is the level of familiaritywith a search topic. We investigate how the search behavior of domain experts changes based on their previous level of familiarity with a search topic, reporting on a user study of biomedical experts searching for a range of domain-specific material. The results of our study show that topic familiarity can influence the number of queries that are employed to complete a task, the types of queries that are entered, and the overall number of query terms. Our findings suggest that biomedical search systems should enable searching through a variety of querying modes, to support the different search strategies that users were found to employ depending on their familiarity with the information that they are searching for. Sarvnaz Karimi, Falk Scholer, Adam Clark, Sadegh Kharazmi |
SIGIR | 1 |
| 2010 | Visualizing search results and document collections using topic maps
David Newman 0001, Timothy Baldwin, Lawrence Cavedon, Sarvnaz Karimi, David Martínez 0001, Falk Scholer, Justin Zobel |
J. Web Semant. | 5 |
| 2006 | English to Persian Transliteration
Sarvnaz Karimi, Andrew Turpin, Falk Scholer |
SPIRE | 1 |