Emanuela Boros

dblp:96/8781 · DBLP profile ↗
← Back
15ranked-venue papers in the field
4as first author
14since 2021 · last 2026
0000-0001-6299-9452ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 10 (3 first)Other / Interdisciplinary · 5 (1 first)
YearPublicationVenuePosition
2026 CLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts
Juri Opitz, Corina Julia Raclé, Emanuela Boros, Andrianos Michail, Matteo Romanello, Maud Ehrmann, Simon Clematide
ECIR (4)3
2026 One Model, Many Guidelines: Instruction Fine-Tuning for Historical Named Entity Recognition
Tien-Nam Nguyen, Emanuela Boros, Adam Jatowt, Mickaël Coustaty, Ahmed Hamdi, Antoine Doucet
ICDAR (3)2
2026 A Dataset of Cultural Heritage Manipulation on English Wikipedia in the Russo-Ukrainian Context
abstract
Digital knowledge platforms increasingly shape how cultural heritage and historical responsibility are represented and remembered, yet they are vulnerable to subtle forms of narrative manipulation that rarely appear as explicit misinformation. We present a new dataset of English Wikipedia revision pairs designed to support the systematic study of cultural heritage manipulation in the context of the Russo--Ukrainian conflict. The resource captures fine-grained textual changes through which cultural attribution, agency, and legitimacy are incrementally reframed within a collaborative encyclopedic infrastructure. By providing paired revisions, editorial context, and consistent annotations, the dataset enables analysis of discursive interventions that remain procedurally legitimate while producing cumulative cultural effects. The dataset is released to support research in information retrieval, natural language processing, and computational social science on detecting and understanding subtle narrative change in large-scale, collaborative knowledge systems. The resource is available at https://zenodo.org/records/19883068.
Maxime Garambois, Hamest Tamrazyan, Emanuela Boros
SIGIR3
2024 Developing a Standardised Vocabulary for Ukrainian Epigraphy and Expanding Digital Epigraphic Resources
Hamest Tamrazyan, Emanuela Boros, Frédéric Kaplan
TPDL (2)2
2023 Injecting Temporal-Aware Knowledge in Historical Named Entity Recognition
Carlos E. González-Gallardo, Emanuela Boros, Edward Giamphy, Ahmed Hamdi, José G. Moreno 0001, Antoine Doucet
ECIR (1)2
2023 Analyzing the Impact of Tokenization on Multilingual Epidemic Surveillance in Low-Resource Languages
Stephen Mutuvi, Emanuela Boros, Antoine Doucet, Gaël Lejeune, Adam Jatowt, Moses Odeo
ICDAR (3)2
2023 Detecting Forged Receipts with Domain-Specific Ontology-Based Entities & Relations
Beatriz Martínez Tornés, Emanuela Boros, Antoine Doucet, Petra Gomez-Krämer, Jean-Marc Ogier
ICDAR (3)2
2023 Receipt Dataset for Document Forgery Detection
Beatriz Martínez Tornés, Théo Taburet, Emanuela Boros, Kais Rouis, Antoine Doucet, Petra Gomez-Krämer, Nicolas Sidere, Vincent Poulain D'Andecy
ICDAR (3)3
2022 Exploring Entities in Event Detection as Question Answering
Emanuela Boros, José G. Moreno 0001, Antoine Doucet
ECIR (1)1
2022 Integrated interdisciplinary workflows for research on historical newspapers: Perspectives from humanities scholars, computer scientists, and librarians
abstract
This article considers the interdisciplinary opportunities and challenges of working with digital cultural heritage, such as digitized historical newspapers, and proposes an integrated digital hermeneutics workflow to combine purely disciplinary research approaches from computer science, humanities, and library work. Common interests and motivations of the above-mentioned disciplines have resulted in interdisciplinary projects and collaborations such as the NewsEye project, which is working on novel solutions on how digital heritage data is (re)searched, accessed, used, and analyzed. We argue that collaborations of different disciplines can benefit from a good understanding of the workflows and traditions of each of the disciplines involved but must find integrated approaches to successfully exploit the full potential of digitized sources. The paper is furthermore providing an insight into digital tools, methods, and hermeneutics in action, showing that integrated interdisciplinary research needs to build something in between the disciplines while respecting and understanding each other's expertise and expectations.
Sarah Oberbichler, Emanuela Boros, Antoine Doucet, Jani Marjanen, Eva Pfanzelter, Juha Rautiainen, Hannu Toivonen, Mikko Tolonen
J. Assoc. Inf. Sci. Technol.2
2021 Event Detection with Entity Markers
Emanuela Boros, José G. Moreno 0001, Antoine Doucet
ECIR (2)1
2021 Token-Level Multilingual Epidemic Dataset for Event Extraction
Stephen Mutuvi, Emanuela Boros, Antoine Doucet, Gaël Lejeune, Adam Jatowt, Moses Odeo
TPDL2
2021 The Importance of Character-Level Information in an Event Detection Model
Emanuela Boros, Romaric Besançon, Olivier Ferret, Brigitte Grau
NLDB1
2021 A Multilingual Dataset for Named Entity Recognition, Entity Linking and Stance Detection in Historical Newspapers
abstract
Named entity processing over historical texts is more and more being used due to the massive documents and archives being stored in digital libraries. However, due to the poor annotated resources of historical nature, information extraction performances fall behind those on contemporary texts. In this paper, we introduce the development of the NewsEye resource, a multilingual dataset for named entity recognition and linking enriched with stances towards named entities. The dataset is comprised of diachronic historical newspaper material published between 1850 and 1950 in French, German, Finnish, and Swedish. Such historical resource is essential in the context of developing and evaluating named entity processing systems. It evenly allows enhancing the performances of existing approaches on historical documents which enables adequate and efficient semantic indexing of historical documents on digital cultural heritage collections.
Ahmed Hamdi, Elvys Linhares Pontes, Emanuela Boros, Thi-Tuyet-Hai Nguyen, Günter Hackl, José G. Moreno 0001, Antoine Doucet
SIGIR3
2019 Automatic Page Classification in a Large Collection of Manuscripts Based on the International Image Interoperability Framework
abstract
In patrimonial institutions such as libraries and archives, the valorization of the vast amount of documents that have been recently digitized is still a challenge. Most of these documents are freely accessible as images but their textual content remains largely unreachable and unknown. Research projects dedicated to specific collection allow creating meta-data or even transcriptions obtained through volunteers or crowd-sourcing. But the vast majority of the documents cannot be manually transcribed or indexed: automatic large-scale processes for indexing are needed. The increasing adoption of the International Image Interoperability Framework (IIIF) by the patrimonial institutions is a technological enabler for the development of such services. Images are accessible with a unique protocol across institutions and both images and data can be presented with standard tools. In this paper, we describe an architecture for automatic processing of historical documents owned by different institutions but processed and presented thanks to the IIIF framework. We implemented this architecture and processed a large collection of books of hours with a page classifier trained on an annotated sample. The result is freely distributed and can be viewed with any IIIF compatible viewer.
Emanuela Boros, Alexis Toumi, Erwan Rouchet, Bastien Abadie, Dominique Stutzmann, Christopher Kermorvant
ICDAR1