EDBT 2026 Demo / reviewers in the wild / expert
Allan Hanbury
dblp:55/6683
· DBLP profile ↗
59ranked-venue papers in the field
2as first author
23since 2021 · last 2026
0000-0002-7149-5843ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 54 (2 first)Big Data, Cloud & Distributed Data Systems · 2Other / Interdisciplinary · 2Database Systems & Data Management · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The LLM Effect on IR Benchmarks: A Meta-Analysis of Effectiveness, Baselines, and ContaminationabstractBenchmark collections have long enabled controlled comparison and cumulative progress in Information Retrieval (IR). However, prior meta-analyses show that reported effectiveness gains often fail to accumulate, in part due to weak or outdated baselines. Large language models (LLMs) are increasingly used in retrieval pipelines, yet their impact on established IR benchmarks has not been systematically analyzed. We analyze 179 publications reporting on the TREC Robust04 collection and the TREC Deep Learning 2020 (DL20) Passage Retrieval benchmark, using ACM Digital Library keyword search supplemented by citation-graph backtracking for Robust04. We observe what we term an LLM effect: recent systems incorporating LLM components achieve 8.8% higher nDCG@10 on DL20 than the best TREC 2020 result and 11.9% higher on Robust04 than the strongest pre-2024 result. However, evaluation practice has shifted from MAP to nDCG@10 over the same window, and our adaptation of the Data Contamination Quiz reveals 12-41% contamination across two widely-used LLM rerankers. Filtering contaminated topics shows no statistically significant effectiveness difference, but small samples and the uncertainty of adapting contamination detection to reranking prevent us from ruling out memorization as a contributing factor. We read the LLM effect as real but unverified: visible in the aggregate numbers, but not cleanly separable from metric drift or pretraining overlap. Moritz Staudinger, Wojciech Kusa, Allan Hanbury |
SIGIR | 3 |
| 2025 | Compare: A Framework for Scientific ComparisonsabstractNavigating the vast and rapidly increasing sea of academic publications to identify institutional synergies, benchmark research contributions and pinpoint key research contributions has become an increasingly daunting task, especially with the current exponential increase in new publications. Existing tools provide useful overviews or single-document insights, but none supports structured, qualitative comparisons across institutions or publications. To address this, we demonstrate Compare, a novel framework that tackles this challenge by enabling sophisticated long-context comparisons of scientific contributions. Compare empowers users to explore and analyze research overlaps and differences at both the institutional and publication granularity, all driven by user-defined questions and automatic retrieval over online resources. For this we leverage on Retrieval-Augmented Generation over evolving data sources to foster long context knowledge synthesis. Unlike traditional scientometric tools, Compare goes beyond quantitative indicators by providing qualitative, citation-supported comparisons. Moritz Staudinger, Wojciech Kusa, Matteo Cancellieri, David Pride, Petr Knoth, Allan Hanbury |
CIKM | 6 |
| 2025 | TimIR: Time-Traveling Through IR History
Moritz Staudinger, Wojciech Kusa, Florina Piroi, Andreas Rauber, Allan Hanbury |
ECIR (4) | 5 |
| 2025 | 6th Workshop on Patent Text Mining and Semantic Technologies (PatentSemTech2025)abstractInformation retrieval systems for the patent domain have a long and evolving history, serving as effective tools to support patent experts in a variety of daily tasks.They facilitate patent landscape analysis, help in the drafting and evaluation tasks in the patenting process, and enable efficient information extraction to gain practical insights into new technologies and innovations.Moreover, they assist in identifying existing solutions, knowledge gaps, trends, and persistent challenges within specific technological fields, thereby informing strategic decision-making and innovation management.Advances in machine learning and natural language processing allow to further automate such tasks, e.g.paragraph retrieval, question answering (QA) or patent text generation.The exploration of semantic technologies for the intellectual property (IP) industry is still in its early stages, with significant potential yet to be unlocked.Investigating the use of artificial intelligence (AI) methods for the patent domain is therefore not only of academic interest, but also highly relevant for practitioners.Compared to other domains, high quality, semi-structured, annotated data is available in large volumes (a requirement for supervised machine learning models), making training large models easier.On the other hand, domain-specific challenges arise, such as very technical language or legal requirements for patent documents, and data from various disciplines and technological areas.With the 6th edition of this workshop we will provide a platform for researchers and industry to discuss recent developments for semantic patent retrieval and analysis employing sophisticated methods ranging from patent text mining, domain-specific information retrieval to large language models (LLMs) targeting next generation applications and use cases for the IP and related domains. Ralf Krestel, Hidir Aras, Linda Andersson, Florina Piroi, Allan Hanbury, Dean Alderucci |
SIGIR | 5 |
| 2024 | 5th Workshop on Patent Text Mining and Semantic Technologies (PatentSemTech2024)abstractInformation retrieval systems for the patent domain have a long history.They can support patent experts in a variety of daily tasks: from analyzing the patent landscape to support experts in the patenting process and large-scale information extraction.Advances in machine learning and natural language processing allow to further automate tasks, such as paragraph retrieval, question answering (QA) or even patent text generation.Uncovering the potential of semantic technologies for the intellectual property (IP) industry is just getting started.Investigating the use of artificial intelligence methods for the patent domain is therefore not only of academic interest, but also highly relevant for practitioners.Compared to other domains, high quality, semi-structured, annotated data is available in large volumes (a requirement for supervised machine learning models), making training large models easier.On the other hand, domain-specific challenges arise, such as very technical language or legal requirements for patent documents.With the 5th edition of this workshop we will provide a platform for researchers and industry to learn about novel and emerging technologies for semantic patent retrieval and big analytics employing sophisticated methods ranging from patent text mining, domain-specific information retrieval to large language models targeting next generation applications and use cases for the IP and related domains. Ralf Krestel, Hidir Aras, Linda Andersson, Florina Piroi, Allan Hanbury, Dean Alderucci |
SIGIR | 5 |
| 2023 | CRUISE-Screening: Living Literature Reviews ToolboxabstractKeeping up with research and finding related work is still a time-consuming task for academics. Researchers sift through thousands of studies to identify a few relevant ones. Automation techniques can help by increasing the efficiency and effectiveness of this task. To this end, we developed CRUISE-Screening, a web-based application for conducting living literature reviews -- a type of literature review that is continuously updated to reflect the latest research in a particular field. CRUISE-Screening is connected to several search engines via an API, which allows for updating the search results periodically. Moreover, it can facilitate the process of screening for relevant publications by using text classification and question answering models. CRUISE-Screening can be used both by researchers conducting literature reviews and by those working on automating the citation screening process to validate their algorithms. The application is open-source, and a demo is available under this URL: https://citation-screening.ec.tuwien.ac.at. Wojciech Kusa, Petr Knoth, Allan Hanbury |
CIKM | 3 |
| 2023 | Readability Measures as Predictors of Understandability and Engagement in Searching to Learn
Yasin Ghafourian, Allan Hanbury, Petr Knoth |
TPDL | 2 |
| 2023 | Ranking for Learning: Studying Users' Perceptions of Relevance, Understandability, and Engagement
Yasin Ghafourian, Allan Hanbury, Petr Knoth |
TPDL | 2 |
| 2023 | SCI-3000: A Dataset for Figure, Table and Caption Extraction from Scientific PDFs
Filip Darmanovic, Allan Hanbury, Markus Zlabinger |
ICDAR (1) | 2 |
| 2023 | 4th Workshop on Patent Text Mining and Semantic Technologies (PatentSemTech2023)abstractInformation retrieval systems for the patent domain have a long history. They can support patent experts in a variety of daily tasks: from analyzing the patent landscape to support experts in the patenting process and large-scale information extraction. Advances in machine learning and natural language processing allow to further automate tasks, such as paragraph retrieval or even patent text generation. Uncovering the potential of semantic technologies for the intellectual property (IP) industry is just getting started. Investigating the use of artificial intelligence methods for the patent domain is therefore not only of academic interest, but also highly relevant for practitioners. Compared to other domains, high quality, semi-structured, annotated data is available in large volumes (a requirement for supervised machine learning models), making training large models easier. On the other hand, domain-specific challenges arise, such as very technical language or legal requirements for patent documents. The focus of the 4th edition of this workshop will be on two-way communication between industry and academia from all areas of information retrieval in particular with the Asian community. We want to bring together novel research results and the latest systems and methods employed by practitioners in the field. Ralf Krestel, Hidir Aras, Linda Andersson, Florina Piroi, Allan Hanbury, Dean Alderucci |
SIGIR | 5 |
| 2023 | VoMBaT: A Tool for Visualising Evaluation Measure Behaviour in High-Recall Search TasksabstractThe objective of High-Recall Information Retrieval (HRIR) is to retrieve as many relevant documents as possible for a given search topic. One approach to HRIR is Technology-Assisted Review (TAR), which uses information retrieval and machine learning techniques to aid the review of large document collections. TAR systems are commonly used in legal eDiscovery and systematic literature reviews. Successful TAR systems are able to find the majority of relevant documents using the least number of assessments. Commonly used retrospective evaluation assumes that the system achieves a specific, fixed recall level first, and then measures the precision or work saved (e.g., precision at r% recall). This approach can cause problems related to understanding the behaviour of evaluation measures in a fixed recall setting. It is also problematic when estimating time and money savings during technology-assisted reviews. Wojciech Kusa, Aldo Lipani, Petr Knoth, Allan Hanbury |
SIGIR | 4 |
| 2022 | TripJudge: A Relevance Judgement Test Collection for TripClick Health RetrievalabstractRobust test collections are crucial for Information Retrieval research. Recently there is a growing interest in evaluating retrieval systems for domain-specific retrieval tasks, however these tasks often lack a reliable test collection with human-annotated relevance assessments following the Cranfield paradigm. In the medical domain, the TripClick collection was recently proposed, which contains click log data from the Trip search engine and includes two click-based test sets. However the clicks are biased to the retrieval model used, which remains unknown, and a previous study shows that the test sets have a low judgement coverage for the Top-10 results of lexical and neural retrieval models. In this paper we present the novel, relevance judgement test collection TripJudge for TripClick health retrieval. We collect relevance judgements in an annotation campaign and ensure the quality and reusability of TripJudge by a variety of ranking methods for pool creation, by multiple judgements per query-document pair and by an at least moderate inter-annotator agreement. We compare system evaluation with TripJudge and TripClick and find that that click and judgement-based evaluation can lead to substantially different system rankings. Sophia Althammer, Sebastian Hofstätter, Suzan Verberne, Allan Hanbury |
CIKM | 4 |
| 2022 | Introducing Neural Bag of Whole-Words with ColBERTer: Contextualized Late Interactions using Enhanced ReductionabstractRecent progress in neural information retrieval has demonstrated large gains in quality, while often sacrificing efficiency and interpretability compared to classical approaches. We propose ColBERTer, a neural retrieval model using contextualized late interaction (ColBERT) with enhanced reduction. Along the effectiveness Pareto frontier, ColBERTer dramatically lowers ColBERT's storage requirements while simultaneously improving the interpretability of its token-matching scores. To this end, ColBERTer fuses single-vector retrieval, multi-vector refinement, and optional lexical matching components into one model. For its multi-vector component, ColBERTer reduces the number of stored vectors by learning unique whole-word representations and learning to identify and remove word representations that are not essential to effective scoring. We employ an explicit multi-task, multi-stage training to facilitate using very small vector dimensions. Results on the MS MARCO and TREC-DL collection show that ColBERTer reduces the storage footprint by up to 2.5x, while maintaining effectiveness. With just one dimension per token in its smallest setting, ColBERTer achieves index storage parity with the plaintext size, with very strong effectiveness results. Finally, we demonstrate ColBERTer's robustness on seven high-quality out-of-domain collections, yielding statistically significant gains over traditional retrieval baselines. Sebastian Hofstätter, Omar Khattab, Sophia Althammer, Mete Sertkan, Allan Hanbury |
CIKM | 5 |
| 2022 | PARM: A Paragraph Aggregation Retrieval Model for Dense Document-to-Document Retrieval
Sophia Althammer, Sebastian Hofstätter, Mete Sertkan, Suzan Verberne, Allan Hanbury |
ECIR (1) | 5 |
| 2022 | Establishing Strong Baselines For TripClick Health Retrieval
Sebastian Hofstätter, Sophia Althammer, Mete Sertkan, Allan Hanbury |
ECIR (2) | 4 |
| 2022 | Automation of Citation Screening for Systematic Literature Reviews Using Neural Networks: A Replicability Study
Wojciech Kusa, Allan Hanbury, Petr Knoth |
ECIR (1) | 2 |
| 2022 | 3rd Workshop on Patent Text Mining and Semantic Technologies (PatentSemTech2022)abstractSteadily increasing numbers of patent applications per year and large amounts of available patent data necessitate highly efficient and interactive next-generation information retrieval systems in the patent domain. AI and Machine Learning (ML) methods such as Deep Learning (DL) are successfully adopted in many domains, so patent researchers and practitioners start to employ AI-based approaches as well, to support experts in the patenting process or to automate patent analysis and retrieval processes. AI-enhanced Information Retrieval systems can improve patent search and analysis but also require millions of annotated sample data for training the ML models. When working with patent data, particular challenges arise that call for adaption of existing IR and AI methods as well as development of novel approaches suited for the patent domain. The focus of the 3rd edition of this workshop will be on two-way communication between industry and academia from all areas of Information Retrieval, such as Natural Language Processing (NLP), Text and Data Mining (TDM), and Semantic Technologies (ST). We want to bring together novel research results and the latest systems and methods employed by the Intellectual Property (IP) industry. Ralf Krestel, Hidir Aras, Linda Andersson, Florina Piroi, Allan Hanbury, Dean Alderucci |
SIGIR | 5 |
| 2021 | Cross-Domain Retrieval in the Legal and Patent Domains: A Reproducibility Study
Sophia Althammer, Sebastian Hofstätter, Allan Hanbury |
ECIR (2) | 3 |
| 2021 | Mitigating the Position Bias of Transformer Models in Passage Re-ranking
Sebastian Hofstätter, Aldo Lipani, Sophia Althammer, Markus Zlabinger, Allan Hanbury |
ECIR (1) | 5 |
| 2021 | Measuring Societal Biases from Text Corpora with Smoothed First-Order Co-occurrence
Navid Rekabsaz, Robert West 0001, James Henderson 0001, Allan Hanbury |
ICWSM | 4 |
| 2021 | Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingabstractA vital step towards the widespread adoption of neural retrieval models is their resource efficiency throughout the training, indexing and query workflows. The neural IR community made great advancements in training effective dual-encoder dense retrieval (DR) models recently. A dense text retrieval model uses a single vector representation per query and passage to score a match, which enables low-latency first-stage retrieval with a nearest neighbor search. Increasingly common, training approaches require enormous compute power, as they either conduct negative passage sampling out of a continuously updating refreshing index or require very large batch sizes. Instead of relying on more compute capability, we introduce an efficient topic-aware query and balanced margin sampling technique, called TAS-Balanced. We cluster queries once before training and sample queries out of a cluster per batch. We train our lightweight 6-layer DR model with a novel dual-teacher supervision that combines pairwise and in-batch negative teachers. Our method is trainable on a single consumer-grade GPU in under 48 hours. We show that our TAS-Balanced training method achieves state-of-the-art low-latency (64ms per query) results on two TREC Deep Learning Track query sets. Evaluated on [email protected], we outperform BM25 by 44%, a plainly trained DR by 19%, docT5query by 11%, and the previous best DR model by 5%. Additionally, TAS-Balanced produces the first dense retriever that outperforms every other method on recall at any cutoff on TREC-DL and allows more resource intensive re-ranking models to operate on fewer passages to improve results further. Sebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, Allan Hanbury |
SIGIR | 5 |
| 2021 | Intra-Document Cascading: Learning to Select Passages for Neural Document RankingabstractAn emerging recipe for achieving state-of-the-art effectiveness in neural document re-ranking involves utilizing large pre-trained language models - e.g., BERT - to evaluate all individual passages in the document and then aggregating the outputs by pooling or additional Transformer layers. A major drawback of this approach is high query latency due to the cost of evaluating every passage in the document with BERT. To make matters worse, this high inference cost and latency varies based on the length of the document, with longer documents requiring more time and computation. To address this challenge, we adopt an intra-document cascading strategy, which prunes passages of a candidate document using a less expensive model, called ESM, before running a scoring model that is more expensive and effective, called ETM. We found it best to train ESM (short for Efficient Student Model) via knowledge distillation from the ETM (short for Effective Teacher Model) e.g., BERT. This pruning allows us to only run the ETM model on a smaller set of passages whose size does not vary by document length. Our experiments on the MS MARCO and TREC Deep Learning Track benchmarks suggest that the proposed Intra-Document Cascaded Ranking Model (IDCM) leads to over 400% lower query latency by providing essentially the same effectiveness as the state-of-the-art BERT-based document ranking models. Sebastian Hofstätter, Bhaskar Mitra 0001, Hamed Zamani, Nick Craswell, Allan Hanbury |
SIGIR | 5 |
| 2021 | 2nd Workshop on Patent Text Mining and Semantic Technologies (PatentSemTech2021)abstractInformation retrieval plays a crucial role in the patent domain. With the success of deep learning (DL) in other domains, patent practitioners and researchers are increasingly developing DL-based approaches to support experts in the patenting process or to automate processes for patent analysis. AI-enhanced information retrieval systems can improve patent search but also require lots of annotated data. When working with patent data, particular challenges arise that call for adaption and novel approaches of general IR and AI methods. with this workshop series we want to establish a two-way communication channel between industry and academia from relevant fields in information retrieval, such as natural language processing (NLP), text and data mining (TDM), and semantic technologies (ST), in order to explore and transfer new knowledge, methods and technologies for the benefit of industrial applications as well as support interdisciplinary research in applied sciences forthe intellectual property (IP) and neighbouring domains. Ralf Krestel, Hidir Aras, Linda Andersson, Florina Piroi, Allan Hanbury, Dean Alderucci |
SIGIR | 5 |
| 2020 | Country-wide Mobility Changes Observed Using Mobile Phone Data During COVID-19 PandemicabstractIn March 2020, the Austrian government introduced a widespread lock-down in response to the COVID-19 pandemic. Based on subjective impressions and anecdotal evidence, Austrian public and private life came to a sudden halt. Here we assess the effect of the lock-down quantitatively for all regions in Austria and present an analysis of daily changes of human mobility throughout Austria using near-real-time anonymized mobile phone data. We describe an efficient data aggregation pipeline and analyze the mobility by quantifying mobile-phone traffic at specific point of interests (POIs), analyzing individual trajectories and investigating the cluster structure of the origin-destination graph. We found a reduction of commuters at Viennese metro stations of over 80% and the number of devices with a radius of gyration of less than 500 m almost doubled. The results of studying crowd-movement behavior highlight considerable changes in the structure of mobility networks, revealed by a higher modularity and an increase from 12 to 20 detected communities. We demonstrate the relevance of mobility data for epidemiological studies by showing a significant correlation of the outflow from the town of Ischgl (an early COVID-19 hotspot) and the reported COVID-19 cases with an 8-day time lag. This research indicates that mobile phone usage data permits the moment-by-moment quantification of mobility behavior for a whole country. We emphasize the need to improve the availability of such data in anonymized form to empower rapid response to combat COVID-19 and future pandemics. Georg Heiler, Tobias Reisch, Jan Hurt, Mohammad Forghani, Aida Omani, Allan Hanbury, Farid Karimipour |
IEEE BigData | 6 |
| 2020 | Learning to Re-Rank with Contextualized StopwordsabstractThe use of stopwords has been thoroughly studied in traditional Information Retrieval systems, but remains unexplored in the context of neural models. Neural re-ranking models take the full text of both the query and document into account. Naturally, removing tokens that do not carry relevance information provides us with an opportunity to improve the effectiveness by reducing noise and lower document representation caching-storage requirements. In this work we propose a novel contextualized stopword detection mechanism for neural re-ranking models. This mechanism consists of training a sparse vector in order to filter out document tokens from the ranking decision. This vector is learned end-to-end based on the contextualized document representations, allowing the model to filter terms on a per occurrence basis. This leads to a more explainable model, as it reduces noise. We integrate our component into the state-of-the-art interaction-based TK neural re-ranking model. Our experiments on the MS MARCO passage collection and queries from the TREC 2019 Deep Learning Track show that filtering out traditional stopwords prior to the neural model reduces its effectiveness, while learning to filter out contextualized representations improves it. Sebastian Hofstätter, Aldo Lipani, Markus Zlabinger, Allan Hanbury |
CIKM | 4 |
| 2020 | Fine-Grained Relevance Annotations for Multi-Task Document Ranking and Question AnsweringabstractThere are many existing retrieval and question answering datasets. However, most of them either focus on ranked list evaluation or single-candidate question answering. This divide makes it challenging to properly evaluate approaches concerned with ranking documents and providing snippets or answers for a given query. In this work, we present FiRA: a novel dataset of Fine-Grained Relevance Annotations. We extend the ranked retrieval annotations of the Deep Learning track of TREC 2019 with passage and word level graded relevance annotations for all relevant documents. We use our newly created data to study the distribution of relevance in long documents, as well as the attention of annotators to specific positions of the text. As an example, we evaluate the recently introduced TKL document ranking model. We find that although TKL exhibits state-of-the-art retrieval results for long documents, it misses many relevant passages. Sebastian Hofstätter, Markus Zlabinger, Mete Sertkan, Michael Schröder 0005, Allan Hanbury |
CIKM | 5 |
| 2020 | Neural-IR-Explorer: A Content-Focused Tool to Explore Neural Re-ranking Results
Sebastian Hofstätter, Markus Zlabinger, Allan Hanbury |
ECIR (2) | 3 |
| 2020 | DSR: A Collection for the Evaluation of Graded Disease-Symptom RelationsabstractThe effective extraction of ranked disease-symptom relationships is a critical component in various medical tasks, including computer-assisted medical diagnosis or the discovery of unexpected associations between diseases. While existing disease-symptom relationship extraction methods are used as the foundation in the various medical tasks, no collection is available to systematically evaluate the performance of such methods. In this paper, we introduce the D isease- S ymptom R elation Collection ( dsr -collection), created by five physicians as expert annotators. We provide graded symptom judgments for diseases by differentiating between relevant symptoms and primary symptoms . Further, we provide several strong baselines, based on the methods used in previous studies. The first method is based on word embeddings, and the second on co-occurrences of MeSH-keywords of medical articles. For the co-occurrence method, we propose an adaption in which not only keywords are considered, but also the full text of medical articles. The evaluation on the dsr -collection shows the effectiveness of the proposed adaption in terms of nDCG, precision, and recall. Markus Zlabinger, Sebastian Hofstätter, Navid Rekabsaz, Allan Hanbury |
ECIR (2) | 4 |
| 2020 | Local Self-Attention over Long Text for Efficient Document RetrievalabstractNeural networks, particularly Transformer-based architectures, have achieved significant performance improvements on several retrieval benchmarks. When the items being retrieved are documents, the time and memory cost of employing Transformers over a full sequence of document terms can be prohibitive. A popular strategy involves considering only the first n terms of the document. This can, however, result in a biased system that under retrieves longer documents. In this work, we propose a local self-attention which considers a moving window over the document terms and for each term attends only to other terms in the same window. This local attention incurs a fraction of the compute and memory cost of attention over the whole document. The windowed approach also leads to more compact packing of padded documents in minibatches resulting in additional savings. We also employ a learned saturation function and a two-staged pooling strategy to identify relevant regions of the document. The Transformer-Kernel pooling model with these changes can efficiently elicit relevance information from documents with thousands of tokens. We benchmark our proposed modifications on the document ranking task from the TREC 2019 Deep Learning track and observe significant improvements in retrieval quality as well as increased retrieval of longer documents at moderate increase in compute and memory costs. Sebastian Hofstätter, Hamed Zamani, Bhaskar Mitra 0001, Nick Craswell, Allan Hanbury |
SIGIR | 5 |
| 2020 | DEXA: Supporting Non-Expert Annotators with Dynamic Examples from ExpertsabstractThe success of crowdsourcing based annotation of text corpora depends on ensuring that crowdworkers are sufficiently well-trained to perform the annotation task accurately. To that end, a frequent approach to train annotators is to provide instructions and a few example cases that demonstrate how the task should be performed (referred to as the CONTROL approach). These globally defined "task-level examples", however, (i) often only cover the common cases that are encountered during an annotation task; and (ii) require effort from crowdworkers during the annotation process to find the most relevant example for the currently annotated sample. To overcome these limitations, we propose to support workers in addition to task-level examples, also with "task-instance level" examples that are semantically similar to the currently annotated data sample (referred to as Dynamic Examples for Annotation, DEXA). Such dynamic examples can be retrieved from collections previously labeled by experts, which are usually available as gold standard dataset. We evaluate DEXA on a complex task of annotating participants, interventions, and outcomes (known as PIO) in sentences of medical studies. The dynamic examples are retrieved using BioSent2Vec, an unsupervised semantic sentence similarity method specific to the biomedical domain. Results show that (i) workers of the DEXA approach reach on average much higher agreements (Cohen's Kappa) to experts than workers of the the CONTROL approach (avg. of 0.68 to experts in DEXA vs. 0.40 in CONTROL); (ii) already three per majority voting aggregated annotations of the DEXA approach reach substantial agreements to experts of 0.78/0.75/0.69 for P/I/O (in CONTROL 0.73/0.58/0.46). Finally, (iii) we acquire explicit feedback from workers and show that in the majority of cases (avg. 72%) workers find the dynamic examples useful. Markus Zlabinger, Marta Sabou, Sebastian Hofstätter, Mete Sertkan, Allan Hanbury |
SIGIR | 5 |
| 2019 | Comparing Implementation Variants Of Distributed Spatial Join on SparkabstractAs an increasing number of sensor devices (Internet of Things) is used, more and more spatio-temporal data becomes available. Being able to process and analyze large quantities of such datasets is therefore critical. Spatial joins in classical geo-information systems do not scale well. Nevertheless, distributed implementations are promising to solve this. Various implementation variants for distributed spatial joins are documented in literature, with some being only suitable for specific use cases. We compared broadcast and multiple variants of a distributed spatially partitioned join. We anticipate that this comparison will give guidance to when to use which implementation strategy. Georg Heiler, Allan Hanbury |
IEEE BigData | 2 |
| 2019 | Enriching Word Embeddings for Patent Retrieval with Global Context
Sebastian Hofstätter, Navid Rekabsaz, Mihai Lupu, Carsten Eickhoff, Allan Hanbury |
ECIR (1) | 5 |
| 2019 | On the Effect of Low-Frequency Terms on Neural-IR ModelsabstractLow-frequency terms are a recurring challenge for information retrieval models, especially neural IR frameworks struggle with adequately capturing infrequently observed words. While these terms are often removed from neural models - mainly as a concession to efficiency demands - they traditionally play an important role in the performance of IR models. In this paper, we analyze the effects of low-frequency terms on the performance and robustness of neural IR models. We conduct controlled experiments on three recent neural IR models, trained on a large-scale passage retrieval collection. We evaluate the neural IR models with various vocabulary sizes for their respective word embeddings, considering different levels of constraints on the available GPU memory. We observe that despite the significant benefits of using larger vocabularies, the performance gap between the vocabularies can be, to a great extent, mitigated by extensive tuning of a related parameter: the number of documents to re-rank. We further investigate the use of subword-token embedding models, and in particular FastText, for neural IR models. Our experiments show that using FastText brings slight improvements to the overall performance of the neural IR models in comparison to models trained on the full vocabulary, while the improvement becomes much more pronounced for queries containing low-frequency terms. Sebastian Hofstätter, Navid Rekabsaz, Carsten Eickhoff, Allan Hanbury |
SIGIR | 4 |
| 2018 | MM: A new Framework for Multidimensional Evaluation of Search EnginesabstractIn this paper, we proposed a framework to evaluate information retrieval systems in presence of multidimensional relevance. This is an important problem in tasks such as consumer health search, where the understandability and trustworthiness of information greatly influence people's decisions based on the search engine results, but common topicality-only evaluation measures ignore these aspects. We used synthetic and real data to compare our proposed framework, named MM, to the understandability-biased information evaluation (UBIRE), an existing framework used in the context of consumer health search. We showed how the proposed approach diverges from the UBIRE framework, and how MM can be used to better understand the trade-offs between topical relevance and the other relevance dimensions. João R. M. Palotti, Guido Zuccon, Allan Hanbury |
CIKM | 3 |
| 2018 | A systematic approach to normalization in probabilistic modelsabstractEvery information retrieval (IR) model embeds in its scoring function a form of term frequency (TF) quantification. The contribution of the term frequency is determined by the properties of the function of the chosen TF quantification, and by its TF normalization. The first defines how independent the occurrences of multiple terms are, while the second acts on mitigating the a priori probability of having a high term frequency in a document (estimation usually based on the document length). New test collections, coming from different domains (e.g. medical, legal), give evidence that not only document length, but in addition, verboseness of documents should be explicitly considered. Therefore we propose and investigate a systematic combination of document verboseness and length. To theoretically justify the combination, we show the duality between document verboseness and length. In addition, we investigate the duality between verboseness and other components of IR models. We test these new TF normalizations on four suitable test collections. We do this on a well defined spectrum of TF quantifications. Finally, based on the theoretical and experimental observations, we show how the two components of this new normalization, document verboseness and length, interact with each other. Our experiments demonstrate that the new models never underperform existing models, while sometimes introducing statistically significantly better results, at no additional computational cost. Aldo Lipani, Thomas Roelleke, Mihai Lupu, Allan Hanbury |
Inf. Retr. J. | 4 |
| 2017 | Does Online Evaluation Correspond to Offline Evaluation in Query Auto Completion?
Alexandros Bampoulidis, João R. M. Palotti, Mihai Lupu, Jon Brassey, Allan Hanbury |
ECIR | 5 |
| 2017 | Fixed-Cost Pooling Strategies Based on IR Evaluation Measures
Aldo Lipani, João R. M. Palotti, Mihai Lupu, Florina Piroi, Guido Zuccon, Allan Hanbury |
ECIR | 6 |
| 2017 | Exploration of a Threshold for Similarity Based on Uncertainty in Word Embedding
Navid Rekabsaz, Mihai Lupu, Allan Hanbury |
ECIR | 3 |
| 2017 | Visual Pool: A Tool to Visualize and Interact with the Pooling MethodabstractEvery year more than 25 test collections are built among the main Information Retrieval (IR) evaluation campaigns. They are extremely important in IR because they become the evaluation praxis for the forthcoming years. Test collections are built mostly using the pooling method. The main advantage of this method is that it drastically reduces the number of documents to be judged. It does so at the cost of introducing biases, which are sometimes aggravated by non optimal configuration. In this paper we develop a novel visualization technique for the pooling method, and integrate it in a demo application named Visual Pool. This demo application enables the user to interact with the pooling method with ease, and develops visual hints in order to analyze existing test collections, and build better ones. Aldo Lipani, Mihai Lupu, Allan Hanbury |
SIGIR | 3 |
| 2017 | Word Embedding Causes Topic Shifting; Exploit Global Context!abstractExploitation of term relatedness provided by word embedding has gained considerable attention in recent IR literature. However, an emerging question is whether this sort of relatedness fits to the needs of IR with respect to retrieval effectiveness. While we observe a high potential of word embedding as a resource for related terms, the incidence of several cases of topic shifting deteriorates the final performance of the applied retrieval models. To address this issue, we revisit the use of global context (i.e. the term co-occurrence in documents) to measure the term relatedness. We hypothesize that in order to avoid topic shifting among the terms with high word embedding similarity, they should often share similar global contexts as well. We therefore study the effectiveness of post filtering of related terms by various global context relatedness measures. Experimental results show significant improvements in two out of three test collections, and support our initial hypothesis regarding the importance of considering global context in retrieval. Navid Rekabsaz, Mihai Lupu, Allan Hanbury, Hamed Zamani |
SIGIR | 3 |
| 2016 | When is the Time Ripe for Natural Language Processing for Patent Passage Retrieval?abstractPatent text is a mixture of legal terms and domain specific terms. In technical English text, a multi-word unit method is often deployed as a word formation strategy in order to expand the working vocabulary, i.e. introducing a new concept without the invention of an entirely new word. In this paper we explore query generation using natural language processing technologies in order to capture domain specific concepts represented as multi-word units. In this paper we examine a range of query generation methods using both linguistic and statistical information. We also propose a new method to identify domain specific terms from other more general phrases. We apply a machine learning approach using domain knowledge and corpus linguistic information in order to learn domain specific terms in relation to phrases' Termhood values. The experiments are conducted on the English part of the CLEF-IP 2013 test collection. The outcome of the experiments shows that the favoured method in terms of PRES and recall is when a language model is used and search terms are extracted with a part-of-speech tagger and a noun phrase chunker. With our proposed methods we improve each evaluation metric significantly compared to the existing state-of-the-art for the CLEP-IP 2013 test collection: for [email protected] by 26% (0.544 from 0.433), for [email protected] by 17% (0.631 from 0.540) and on document MAP by 57% (0.300 from 0.191). Linda Andersson, Mihai Lupu, João R. M. Palotti, Allan Hanbury, Andreas Rauber |
CIKM | 4 |
| 2016 | The Solitude of Relevant Documents in the PoolabstractPool bias is a well understood problem of test-collection based benchmarking in information retrieval. The pooling method itself is designed to identify all relevant documents. In practice, 'all' translates to `as many as possible given some budgetary constraints' and the problem persists, albeit mitigated. Recently, methods to address this pool bias for previously created test collections have been proposed, for the evaluation measure precision at cut-off ([email protected]). Analyzing previous methods, we make the empirical observation that the distribution of the probability of providing new relevant documents to the pool, over the runs, is log-normal (when the pooling strategy is fixed depth at cut-off). We use this observation to calculate a prior probability of providing new relevant documents, which we then use in a pool bias estimator that improves upon previous estimates of precision at cut-off. Through extensive experimental results, covering 15 test collections, we show that the proposed bias correction method is the new state of the art, providing the closest estimates yet when compared to the original pool. Aldo Lipani, Mihai Lupu, Evangelos Kanoulas, Allan Hanbury |
CIKM | 4 |
| 2016 | Generalizing Translation Models in the Probabilistic Relevance FrameworkabstractA recurring question in information retrieval is whether term associations can be properly integrated in traditional information retrieval models while preserving their robustness and effectiveness. In this paper, we revisit a wide spectrum of existing models (Pivoted Document Normalization, BM25, BM25 Verboseness Aware, Multi-Aspect TF, and Language Modelling) by introducing a generalisation of the idea of the translation model. This generalisation is a de facto transformation of the translation models from Language Modelling to the probabilistic models. In doing so, we observe a potential limitation of these generalised translation models: they only affect the term frequency based components of all the models, ignoring changes in document and collection statistics. We correct this limitation by extending the translation models with the 15 statistics of term associations and provide extensive experimental results to demonstrate the benefit of the newly proposed methods. Additionally, we compare the translation models with query expansion methods based on the same term association resources, as well as based on Pseudo-Relevance Feedback (PRF). We observe that translation models always outperform the first, but provide complementary information with the second, such that by using PRF and our translation models together we observe results better than the current state of the art. Navid Rekabsaz, Mihai Lupu, Allan Hanbury, Guido Zuccon |
CIKM | 3 |
| 2016 | Query Variations and their Effect on Comparing Information Retrieval SystemsabstractWe explore the implications of using query variations for evaluating information retrieval systems and how these variations should be exploited to compare system effectiveness. Current evaluation approaches consider the availability of a set of topics (information needs), and only one expression of each topic in the form of a query is used for evaluation and system comparison. While there is strong evidence that considering query variations better models the usage of retrieval systems and accounts for the important user aspect of user variability, it is unclear how to best exploit query variations for evaluating and comparing information retrieval systems. Guido Zuccon, João R. M. Palotti, Allan Hanbury |
CIKM | 3 |
| 2016 | The Curious Incidence of Bias Corrections in the Pool
Aldo Lipani, Mihai Lupu, Allan Hanbury |
ECIR | 3 |
| 2016 | Ranking Health Web Pages with Relevance and UnderstandabilityabstractWe propose a method that integrates relevance and understandability to rank health web documents. We use a learning to rank approach with standard retrieval features to determine topical relevance and additional features based on readability measures and medical lexical aspects to determine understandability. Our experiments measured the effectiveness of the learning to rank approach integrating understandability on a consumer health benchmark. The findings suggest that this approach promotes documents that are at the same time topically relevant and understandable. João R. M. Palotti, Lorraine Goeuriot, Guido Zuccon, Allan Hanbury |
SIGIR | 4 |
| 2016 | How users search and what they search for in the medical domain - Understanding laypeople and experts through query logs
João R. M. Palotti, Allan Hanbury, Henning Müller, Charles E. Kahn Jr. |
Inf. Retr. J. | 2 |
| 2015 | The Influence of Pre-processing on the Estimation of Readability of Web DocumentsabstractThis paper investigates the effect that text pre-processing approaches have on the estimation of the readability of web pages. Readability has been highlighted as an important aspect of web search result personalisation in previous work. The most widely used text readability measures rely on surface level characteristics of text, such as the length of words and sentences. We demonstrate that different tools for extracting text from web pages lead to very different estimations of readability. This has an important implication for search engines because search result personalisation strategies that consider users reading ability may fail if incorrect text readability estimations are computed. João R. M. Palotti, Guido Zuccon, Allan Hanbury |
CIKM | 3 |
| 2015 | Workshop Multimodal Retrieval in the Medical Domain (MRMD) 2015
Henning Müller, Oscar Alfonso Jiménez del Toro, Allan Hanbury, Georg Langs, Antonio Foncubierta-Rodríguez |
ECIR | 3 |
| 2015 | DASyR(IR) - document analysis system for systematic reviews (in Information Retrieval)abstractCreating systematic reviews is a painstaking task undertaken especially in domains where experimental results are the primary method to knowledge creation. For the review authors, analysing documents to extract relevant data is a demanding activity. To support the creation of systematic reviews, we have created DASyR-a semi-automatic document analysis system. DASyR is our solution to annotating published papers for the purpose of ontology population. For domains where dictionaries are not existing or inadequate, DASyR relies on a semi-automatic annotation bootstrapping method based on positional Random Indexing, followed by traditional Machine Learning algorithms to extend the annotation set. We provide an example of the method application to a subdomain of Computer Science, the Information Retrieval evaluation domain. The reliance of this domain on large scale experimental studies makes it a perfect domain to test on. We show the utility of DASyR through experimental results for different parameter values for the bootstrap procedure, evaluated in terms of annotator agreement, error rate, precision and recall. Florina Piroi, Aldo Lipani, Mihai Lupu, Allan Hanbury |
ICDAR | 4 |
| 2015 | Splitting Water: Precision and Anti-Precision to Reduce Pool BiasabstractFor many tasks in evaluation campaigns, especially those modeling narrow domain-specific challenges, lack of participation leads to a potential pooling bias due to the scarce number of pooled runs. It is well known that the reliability of a test collection is proportional to the number of topics and relevance assessments provided for each topic, but also to same extent to the diversity in participation in the challenges. Hence, in this paper we present a new perspective in reducing the pool bias by studying the effect of merging an unpooled run with the pooled runs. We also introduce an indicator used by the bias correction method to decide whether the correction needs to be applied or not. This indicator gives strong clues about the potential of a "good" run tested on an "unfriendly" test collection (i.e. a collection where the pool was contributed to by runs very different from the one at hand). We demonstrate the correctness of our method on a set of fifteen test collections from the Text REtrieval Conference (TREC). We observe a reduction in system ranking error and absolute score difference error. Aldo Lipani, Mihai Lupu, Allan Hanbury |
SIGIR | 3 |
| 2014 | A Glimpse into the State and Future of (Big) Data Analytics in Austria - Results from an Online SurveyabstractWe present results from questionnaire data that were collected from leading data analytics experts in Austria. The online survey addresses very current and pressing questions in the area of (big) data analysis. Our i??ndings provide valuable insights about what top Austrian data scientists think about data analytics, what they consider as important application areas that can benei??t from big data and data processing, the challenges of the future and how soon these challenges will become important, and the potential research topics of tomorrow. We
visualize results, summarize our i??ndings and suggest a possible roadmap for future decision making. Ralf Bierig, Allan Hanbury, Martina Haas, Florina Piroi, Helmut Berger, Mihai Lupu, Michael Dittenbach |
DATA | 2 |
| 2014 | Khresmoi Professional: Multilingual, Multimodal Professional Medical Search
Liadh Kelly, Sebastian Dungs, Sascha Kriewel, Allan Hanbury, Lorraine Goeuriot, Gareth J. F. Jones, Georg Langs, Henning Müller |
ECIR | 4 |
| 2014 | A System Framework for Concept- and Credibility-Based Multimedia RetrievalabstractWe present a multimedia retrieval system framework that incorporates components for processing multimedia content in different modes and languages. The framework provides concept-based information retrieval facilities that applies credibility information for result re-ranking. The architecture combines both a direct user interface and a batched evaluation interface for reproducible research in multimedia IR. The demo presents a preliminary version of the system framework and shows a use case based on the ImageCLEF 2011 Wikipedia test collection. Ralf Bierig, Cristina Serban, Alexandra Siriteanu, Mihai Lupu, Allan Hanbury |
ICMR | 5 |
| 2014 | Guest editorial: Special issue on information retrieval in the intellectual property domain
Allan Hanbury, Mihai Lupu, Noriko Kando, Barrou Diallo |
Inf. Retr. | 1 |
| 2013 | Exploring Patent Passage Retrieval Using Nouns Phrases
Linda Andersson, Parvaz Mahdabi, Allan Hanbury, Andreas Rauber |
ECIR | 3 |
| 2013 | Integrating IR Technologies for Professional Search - (Full-Day Workshop)
Michail Salampasis, Norbert Fuhr, Allan Hanbury, Mihai Lupu, Birger Larsen, Henrik Strindberg |
ECIR | 3 |
| 2012 | Medical information retrieval: an instance of domain-specific searchabstractDue to an explosion in the amount of medical information available, search techniques are gaining importance in the medical domain. This tutorial discusses recent results on search in the medical domain, including the outcome of surveys on end user requirements, research relevant to the field, and current medical and health search applications available. Finally, the extent to which available techniques meet user requirements are discussed, and open challenges in the field are identified. Allan Hanbury |
SIGIR | 1 |
| 2011 | 4th international workshop on patent information retrieval (PaIR'11)abstractThe 4th International Workshop on Patent Information Retrieval builds on the experiences of the first three workshops, to provide its participants an exciting, scientifically challenging and interactive event, where specific issues of patent retrieval may be put into the general context of Information Retrieval and Knowledge Management, in order to explore innovative solutions to new and old problems, but also to evaluate and adapt traditional or classic approaches to new problems. This year, we observe an increase in the use of standardized test collections in the contributions received, and, at the same time, new discussion points on how to make such standardized evaluation exercises more accessible to the larger IP community. Mihai Lupu, Allan Hanbury, Andreas Rauber |
CIKM | 2 |