Florina Piroi

dblp:52/1991 · DBLP profile ↗
← Back
18ranked-venue papers in the field
1as first author
15since 2021 · last 2026
0000-0001-7584-6439ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 16Database Systems & Data Management · 1Other / Interdisciplinary · 1 (1 first)
YearPublicationVenuePosition
2026 Evaluating Information Retrieval Models Along Time: The LongEval Lab at CLEF 2026
Timo Breuer 0002, Matteo Cancellieri, Alaa El-Ebshihy, Maik Fröbe, Petra Galuscáková, Lorraine Goeuriot, Gabriel Iturra-Bocaz, Jüri Keller, Petr Knoth, Andreas Konstantin Kruff, Philippe Mulhem, Florina Piroi, David Pride, Philipp Schaer, Didier Schwab
ECIR (4)12
2025 LongEval at CLEF 2025: Longitudinal Evaluation of IR Model Performance
Matteo Cancellieri, Alaa El-Ebshihy, Tobias Fink, Petra Galuscáková, Gabriela González Sáez, Lorraine Goeuriot, David Iommi, Jüri Keller, Petr Knoth, Philippe Mulhem, Florina Piroi, David Pride, Philipp Schaer
ECIR (5)11
2025 Benchmark Creation for Narrative Knowledge Delta Extraction Tasks: Can LLMs Help?
Alaa El-Ebshihy, Annisa Maulida Ningtyas, Florina Piroi, Andreas Rauber
ECIR (3)3
2025 TimIR: Time-Traveling Through IR History
Moritz Staudinger, Wojciech Kusa, Florina Piroi, Andreas Rauber, Allan Hanbury
ECIR (4)3
2025 6th Workshop on Patent Text Mining and Semantic Technologies (PatentSemTech2025)
abstract
Information retrieval systems for the patent domain have a long and evolving history, serving as effective tools to support patent experts in a variety of daily tasks.They facilitate patent landscape analysis, help in the drafting and evaluation tasks in the patenting process, and enable efficient information extraction to gain practical insights into new technologies and innovations.Moreover, they assist in identifying existing solutions, knowledge gaps, trends, and persistent challenges within specific technological fields, thereby informing strategic decision-making and innovation management.Advances in machine learning and natural language processing allow to further automate such tasks, e.g.paragraph retrieval, question answering (QA) or patent text generation.The exploration of semantic technologies for the intellectual property (IP) industry is still in its early stages, with significant potential yet to be unlocked.Investigating the use of artificial intelligence (AI) methods for the patent domain is therefore not only of academic interest, but also highly relevant for practitioners.Compared to other domains, high quality, semi-structured, annotated data is available in large volumes (a requirement for supervised machine learning models), making training large models easier.On the other hand, domain-specific challenges arise, such as very technical language or legal requirements for patent documents, and data from various disciplines and technological areas.With the 6th edition of this workshop we will provide a platform for researchers and industry to discuss recent developments for semantic patent retrieval and analysis employing sophisticated methods ranging from patent text mining, domain-specific information retrieval to large language models (LLMs) targeting next generation applications and use cases for the IP and related domains.
Ralf Krestel, Hidir Aras, Linda Andersson, Florina Piroi, Allan Hanbury, Dean Alderucci
SIGIR4
2024 LongEval: Longitudinal Evaluation of Model Performance at CLEF 2024
Rabab Alkhalifa, Hsuvas Borkakoty, Romain Deveaud, Alaa El-Ebshihy, Luis Espinosa Anke, Tobias Fink, Gabriela González Sáez, Petra Galuscáková, Lorraine Goeuriot, David Iommi, Maria Liakata, Harish Tayyar Madabushi, Pablo Medina-Alias, Philippe Mulhem, Florina Piroi, Martin Popel, Christophe Servan, Arkaitz Zubiaga
ECIR (6)15
2024 5th Workshop on Patent Text Mining and Semantic Technologies (PatentSemTech2024)
abstract
Information retrieval systems for the patent domain have a long history.They can support patent experts in a variety of daily tasks: from analyzing the patent landscape to support experts in the patenting process and large-scale information extraction.Advances in machine learning and natural language processing allow to further automate tasks, such as paragraph retrieval, question answering (QA) or even patent text generation.Uncovering the potential of semantic technologies for the intellectual property (IP) industry is just getting started.Investigating the use of artificial intelligence methods for the patent domain is therefore not only of academic interest, but also highly relevant for practitioners.Compared to other domains, high quality, semi-structured, annotated data is available in large volumes (a requirement for supervised machine learning models), making training large models easier.On the other hand, domain-specific challenges arise, such as very technical language or legal requirements for patent documents.With the 5th edition of this workshop we will provide a platform for researchers and industry to learn about novel and emerging technologies for semantic patent retrieval and big analytics employing sophisticated methods ranging from patent text mining, domain-specific information retrieval to large language models targeting next generation applications and use cases for the IP and related domains.
Ralf Krestel, Hidir Aras, Linda Andersson, Florina Piroi, Allan Hanbury, Dean Alderucci
SIGIR4
2023 LongEval: Longitudinal Evaluation of Model Performance at CLEF 2023
Rabab Alkhalifa, Iman Munire Bilal, Hsuvas Borkakoty, José Camacho-Collados, Romain Deveaud, Alaa El-Ebshihy, Luis Espinosa Anke, Gabriela González Sáez, Petra Galuscáková, Lorraine Goeuriot, Elena Kochkina, Maria Liakata, Daniel Loureiro, Harish Tayyar Madabushi, Philippe Mulhem, Florina Piroi, Martin Popel, Christophe Servan, Arkaitz Zubiaga
ECIR (3)16
2023 ECIR 2023 Workshop: Legal Information Retrieval
Suzan Verberne, Evangelos Kanoulas, Gineke Wiggers, Florina Piroi, Arjen P. de Vries
ECIR (3)4
2023 Using Semi-automatic Annotation Platform to Create Corpus for Argumentative Zoning
Alaa El-Ebshihy, Annisa Maulida Ningtyas, Florina Piroi, Andreas Rauber, Ade Romadhony, Said al Faraby, Mira Kania Sabariah
TPDL3
2023 LongEval-Retrieval: French-English Dynamic Test Collection for Continuous Web Search Evaluation
abstract
LongEval-Retrieval is a Web document retrieval benchmark that focuses on continuous retrieval evaluation. This test collection is intended to be used to study the temporal persistence of Information Retrieval systems and will be used as the test collection in the Longitudinal Evaluation of Model Performance Track (LongEval) at CLEF 2023. This benchmark simulates an evolving information system environment - such as the one a Web search engine operates in - where the document collection, the query distribution, and relevance all move continuously, while following the Cranfield paradigm for offline evaluation. To do that, we introduce the concept of a dynamic test collection that is composed of successive sub-collections each representing the state of an information system at a given time step. In LongEval-Retrieval, each sub-collection contains a set of queries, documents, and soft relevance assessments built from click models. The data comes from Qwant, a privacy-preserving Web search engine that primarily focuses on the French market. LongEval-Retrieval also provides a 'mirror' collection: it is initially constructed in the French language to benefit from the majority of Qwant's traffic, before being translated to English. This paper presents the creation process of LongEval-Retrieval and provides baseline runs and analysis.
Petra Galuscáková, Romain Deveaud, Gabriela González Sáez, Philippe Mulhem, Lorraine Goeuriot, Florina Piroi, Martin Popel
SIGIR6
2023 4th Workshop on Patent Text Mining and Semantic Technologies (PatentSemTech2023)
abstract
Information retrieval systems for the patent domain have a long history. They can support patent experts in a variety of daily tasks: from analyzing the patent landscape to support experts in the patenting process and large-scale information extraction. Advances in machine learning and natural language processing allow to further automate tasks, such as paragraph retrieval or even patent text generation. Uncovering the potential of semantic technologies for the intellectual property (IP) industry is just getting started. Investigating the use of artificial intelligence methods for the patent domain is therefore not only of academic interest, but also highly relevant for practitioners. Compared to other domains, high quality, semi-structured, annotated data is available in large volumes (a requirement for supervised machine learning models), making training large models easier. On the other hand, domain-specific challenges arise, such as very technical language or legal requirements for patent documents. The focus of the 4th edition of this workshop will be on two-way communication between industry and academia from all areas of information retrieval in particular with the Asian community. We want to bring together novel research results and the latest systems and methods employed by practitioners in the field.
Ralf Krestel, Hidir Aras, Linda Andersson, Florina Piroi, Allan Hanbury, Dean Alderucci
SIGIR4
2022 A Platform for Argumentative Zoning Annotation and Scientific Summarization
abstract
Argumentative Zoning (AZ) is a tool to obtain informative summaries of scientific articles. Using AZ assumes the definition of the main rhetorical structure in scientific articles, which are, then, used for the summary creation. The unavailability of large AZ annotated benchmark datasets is a bottleneck to training AZ-based summarization algorithms. In this work, we present an annotation platform for an AZ that defines four categories (zones), Claim, Method, Result and Conclusion, that are used to label sentences selected from scientific articles. The proposed tool can be used both for collecting benchmark datasets, and to help the researchers to create their own sub-corpora.
Alaa El-Ebshihy, Annisa Maulida Ningtyas, Linda Andersson, Florina Piroi, Andreas Rauber
CIKM4
2022 3rd Workshop on Patent Text Mining and Semantic Technologies (PatentSemTech2022)
abstract
Steadily increasing numbers of patent applications per year and large amounts of available patent data necessitate highly efficient and interactive next-generation information retrieval systems in the patent domain. AI and Machine Learning (ML) methods such as Deep Learning (DL) are successfully adopted in many domains, so patent researchers and practitioners start to employ AI-based approaches as well, to support experts in the patenting process or to automate patent analysis and retrieval processes. AI-enhanced Information Retrieval systems can improve patent search and analysis but also require millions of annotated sample data for training the ML models. When working with patent data, particular challenges arise that call for adaption of existing IR and AI methods as well as development of novel approaches suited for the patent domain. The focus of the 3rd edition of this workshop will be on two-way communication between industry and academia from all areas of Information Retrieval, such as Natural Language Processing (NLP), Text and Data Mining (TDM), and Semantic Technologies (ST). We want to bring together novel research results and the latest systems and methods employed by the Intellectual Property (IP) industry.
Ralf Krestel, Hidir Aras, Linda Andersson, Florina Piroi, Allan Hanbury, Dean Alderucci
SIGIR4
2021 2nd Workshop on Patent Text Mining and Semantic Technologies (PatentSemTech2021)
abstract
Information retrieval plays a crucial role in the patent domain. With the success of deep learning (DL) in other domains, patent practitioners and researchers are increasingly developing DL-based approaches to support experts in the patenting process or to automate processes for patent analysis. AI-enhanced information retrieval systems can improve patent search but also require lots of annotated data. When working with patent data, particular challenges arise that call for adaption and novel approaches of general IR and AI methods. with this workshop series we want to establish a two-way communication channel between industry and academia from relevant fields in information retrieval, such as natural language processing (NLP), text and data mining (TDM), and semantic technologies (ST), in order to explore and transfer new knowledge, methods and technologies for the benefit of industrial applications as well as support interdisciplinary research in applied sciences forthe intellectual property (IP) and neighbouring domains.
Ralf Krestel, Hidir Aras, Linda Andersson, Florina Piroi, Allan Hanbury, Dean Alderucci
SIGIR4
2017 Fixed-Cost Pooling Strategies Based on IR Evaluation Measures
Aldo Lipani, João R. M. Palotti, Mihai Lupu, Florina Piroi, Guido Zuccon, Allan Hanbury
ECIR4
2015 DASyR(IR) - document analysis system for systematic reviews (in Information Retrieval)
abstract
Creating systematic reviews is a painstaking task undertaken especially in domains where experimental results are the primary method to knowledge creation. For the review authors, analysing documents to extract relevant data is a demanding activity. To support the creation of systematic reviews, we have created DASyR-a semi-automatic document analysis system. DASyR is our solution to annotating published papers for the purpose of ontology population. For domains where dictionaries are not existing or inadequate, DASyR relies on a semi-automatic annotation bootstrapping method based on positional Random Indexing, followed by traditional Machine Learning algorithms to extend the annotation set. We provide an example of the method application to a subdomain of Computer Science, the Information Retrieval evaluation domain. The reliance of this domain on large scale experimental studies makes it a perfect domain to test on. We show the utility of DASyR through experimental results for different parameter values for the bootstrap procedure, evaluated in terms of annotator agreement, error rate, precision and recall.
Florina Piroi, Aldo Lipani, Mihai Lupu, Allan Hanbury
ICDAR1
2014 A Glimpse into the State and Future of (Big) Data Analytics in Austria - Results from an Online Survey
abstract
We present results from questionnaire data that were collected from leading data analytics experts in Austria. The online survey addresses very current and pressing questions in the area of (big) data analysis. Our i??ndings provide valuable insights about what top Austrian data scientists think about data analytics, what they consider as important application areas that can benei??t from big data and data processing, the challenges of the future and how soon these challenges will become important, and the potential research topics of tomorrow. We visualize results, summarize our i??ndings and suggest a possible roadmap for future decision making.
Ralf Bierig, Allan Hanbury, Martina Haas, Florina Piroi, Helmut Berger, Mihai Lupu, Michael Dittenbach
DATA4