Petr Knoth

dblp:01/7968 · DBLP profile ↗
← Back
21ranked-venue papers in the field
3as first author
12since 2021 · last 2026
0000-0003-1161-7359ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 19 (3 first)Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 Evaluating Information Retrieval Models Along Time: The LongEval Lab at CLEF 2026
Timo Breuer 0002, Matteo Cancellieri, Alaa El-Ebshihy, Maik Fröbe, Petra Galuscáková, Lorraine Goeuriot, Gabriel Iturra-Bocaz, Jüri Keller, Petr Knoth, Andreas Konstantin Kruff, Philippe Mulhem, Florina Piroi, David Pride, Philipp Schaer, Didier Schwab
ECIR (4)9
2025 Compare: A Framework for Scientific Comparisons
abstract
Navigating the vast and rapidly increasing sea of academic publications to identify institutional synergies, benchmark research contributions and pinpoint key research contributions has become an increasingly daunting task, especially with the current exponential increase in new publications. Existing tools provide useful overviews or single-document insights, but none supports structured, qualitative comparisons across institutions or publications. To address this, we demonstrate Compare, a novel framework that tackles this challenge by enabling sophisticated long-context comparisons of scientific contributions. Compare empowers users to explore and analyze research overlaps and differences at both the institutional and publication granularity, all driven by user-defined questions and automatic retrieval over online resources. For this we leverage on Retrieval-Augmented Generation over evolving data sources to foster long context knowledge synthesis. Unlike traditional scientometric tools, Compare goes beyond quantitative indicators by providing qualitative, citation-supported comparisons.
Moritz Staudinger, Wojciech Kusa, Matteo Cancellieri, David Pride, Petr Knoth, Allan Hanbury
CIKM5
2025 LongEval at CLEF 2025: Longitudinal Evaluation of IR Model Performance
Matteo Cancellieri, Alaa El-Ebshihy, Tobias Fink, Petra Galuscáková, Gabriela González Sáez, Lorraine Goeuriot, David Iommi, Jüri Keller, Petr Knoth, Philippe Mulhem, Florina Piroi, David Pride, Philipp Schaer
ECIR (5)9
2023 Prompting Strategies for Citation Classification
abstract
Citation classification aims to identify the purpose of the cited article in the citing article. Previous citation classification methods rely largely on supervised approaches. The models are trained on datasets with citing sentences or citation contexts annotated for a citation's purpose or function or intent. Recent advancements in Large Language Models (LLMs) have dramatically improved the ability of NLP systems to achieve state-of-the-art performances under zero or few-shot settings. This makes LLMs particularly suitable for tasks where sufficiently large labelled datasets are not yet available, which remains to be the case for citation classification. This paper systematically investigates the effectiveness of different prompting strategies for citation classification and compares them to promptless strategies as a baseline. Specifically, we evaluate the following four strategies, two of which we introduce for the first time, which involve updating Language Model (LM) parameters while training the model: (1) Promptless fine-tuning, (2) Fixed-prompt LM tuning, (3) Dynamic Context-prompt LM tuning (proposed), (4) Prompt + LM fine-tuning (proposed). Additionally, we test the zero-shot performance of LLMs, GPT3.5, a (5) Tuning-free prompting strategy that involves no parameter updating. Our results show that prompting methods based on LM parameter updating significantly improve citation classification performances on both domain-specific and multi-disciplinary citation classifications. Moreover, our Dynamic Context-prompting method achieves top scores both for the ACL-ARC and ACT2 citation classification datasets, surpassing the highest-performing system in the 3C shared task benchmark. Interestingly, we observe zero-shot GPT3.5 to perform well on ACT2 but poorly on the ACL-ARC dataset.
Suchetha N. Kunnath, David Pride, Petr Knoth
CIKM3
2023 CRUISE-Screening: Living Literature Reviews Toolbox
abstract
Keeping up with research and finding related work is still a time-consuming task for academics. Researchers sift through thousands of studies to identify a few relevant ones. Automation techniques can help by increasing the efficiency and effectiveness of this task. To this end, we developed CRUISE-Screening, a web-based application for conducting living literature reviews -- a type of literature review that is continuously updated to reflect the latest research in a particular field. CRUISE-Screening is connected to several search engines via an API, which allows for updating the search results periodically. Moreover, it can facilitate the process of screening for relevant publications by using text classification and question answering models. CRUISE-Screening can be used both by researchers conducting literature reviews and by those working on automating the citation screening process to validate their algorithms. The application is open-source, and a demo is available under this URL: https://citation-screening.ec.tuwien.ac.at.
Wojciech Kusa, Petr Knoth, Allan Hanbury
CIKM2
2023 Readability Measures as Predictors of Understandability and Engagement in Searching to Learn
Yasin Ghafourian, Allan Hanbury, Petr Knoth
TPDL3
2023 Ranking for Learning: Studying Users' Perceptions of Relevance, Understandability, and Engagement
Yasin Ghafourian, Allan Hanbury, Petr Knoth
TPDL3
2023 CORE-GPT: Combining Open Access Research and Large Language Models for Credible, Trustworthy Question Answering
David Pride, Matteo Cancellieri, Petr Knoth
TPDL3
2023 VoMBaT: A Tool for Visualising Evaluation Measure Behaviour in High-Recall Search Tasks
abstract
The objective of High-Recall Information Retrieval (HRIR) is to retrieve as many relevant documents as possible for a given search topic. One approach to HRIR is Technology-Assisted Review (TAR), which uses information retrieval and machine learning techniques to aid the review of large document collections. TAR systems are commonly used in legal eDiscovery and systematic literature reviews. Successful TAR systems are able to find the majority of relevant documents using the least number of assessments. Commonly used retrospective evaluation assumes that the system achieves a specific, fixed recall level first, and then measures the precision or work saved (e.g., precision at r% recall). This approach can cause problems related to understanding the behaviour of evaluation measures in a fixed recall setting. It is also problematic when estimating time and money savings during technology-assisted reviews.
Wojciech Kusa, Aldo Lipani, Petr Knoth, Allan Hanbury
SIGIR3
2022 Automation of Citation Screening for Systematic Literature Reviews Using Neural Networks: A Replicability Study
Wojciech Kusa, Allan Hanbury, Petr Knoth
ECIR (1)3
2022 Cui Bono? Cumulative Advantage in Open Access Publishing
David Pride, Matteo Cancellieri, Petr Knoth
TPDL3
2022 Formal Analysis and Estimation of Chance in Datasets Based on Their Properties
abstract
Machine learning research, particularly in genomics, is often based on wide shaped datasets, i.e. datasets having a large number of features, but a small number of samples. Such configurations raise the possibility of chance influence (the increase of measured accuracy due to chance correlations) on the learning process and the evaluation results. Prior research underlined the problem of generalization of models obtained based on such data. In this paper, we investigate the influence of chance on prediction and show its significant effects on wide shaped datasets. First, we empirically demonstrate how significant the influence of chance in such datasets is by showing that prediction models trained on thousands of randomly generated datasets can achieve high accuracy. This is the case even when using cross-validation. We then provide a formal analysis of chance influence and design formal chance influence estimators based on the dataset parameters, namely its sample size, the number of features, the number of classes and the class distribution. Finally, we provide an in-depth discussion of the formal analysis including applications of the findings and recommendations on chance influence mitigation.
Abdel Aziz Taha, Luca Papariello, Alexandros Bampoulidis, Petr Knoth, Mihai Lupu
IEEE Trans. Knowl. Data Eng.4
2019 Online Evaluations for Everyone: Mr. DLib's Living Lab for Scholarly Recommendations
Jöran Beel, Andrew Collins 0002, Oliver Kopp, Linus W. Dietz, Petr Knoth
ECIR (2)5
2018 Peer Review and Citation Data in Predicting University Rankings, a Large-Scale Analysis
David Pride, Petr Knoth
TPDL2
2018 Using citation-context to reduce topic drifting on pure citation-based recommendation
abstract
Recent works in the area of academic recommender systems have demonstrated the effectiveness of co-citation and citation closeness in related-document recommendations. However, documents recommended from such systems may drift away from the main theme of the query document. In this work, we investigate whether incorporating the textual information in close proximity to a citation as well as the citation position could reduce such drifting and further increase the performance of the recommender system. To investigate this, we run experiments with several recommendation methods on a newly created and now publicly available dataset containing 53 million unique citation-based records. We then conduct a user-based evaluation with domain-knowledgeable participants. Our results show that a new method based on the combination of Citation Proximity Analysis (CPA), topic modelling and word embeddings achieves more than 20% improvement in Normalised Discounted Cumulative Gain (nDCG) compared to CPA.
Anita Khadka, Petr Knoth
RecSys2
2017 Classifying Document Types to Enhance Search and Recommendations in Digital Libraries
Aristotelis Charalampous, Petr Knoth
TPDL2
2017 What Others Say About This Work? Scalable Extraction of Citation Contexts from Research Papers
Petr Knoth, Philip Gooch, Kris Jack
TPDL1
2017 Incidental or Influential? - Challenges in Automatically Detecting Citation Importance Using Publication Full Texts
David Pride, Petr Knoth
TPDL2
2017 Workshop on Scholarly Web Mining (SWM 2017)
abstract
No abstract available.
Robert M. Patton, Thomas E. Potok, Petr Knoth, Drahomira Herrmannova
WSDM3
2011 Connecting Repositories in the Open Access Domain Using Text Mining and Semantic Data
Petr Knoth, Vojtech Robotka, Zdenek Zdráhal
TPDL1
2010 EUROGENE: Multilingual Retrieval and Machine Translation Applied to Human Genetics
Petr Knoth, Trevor D. Collins, Elsa Sklavounou, Zdenek Zdráhal
ECIR1