VLDB 2026 Research / reviewers in the wild / expert
Antoine Doucet
dblp:86/1691
· DBLP profile ↗
45ranked-venue papers in the field
3as first author
26since 2021 · last 2026
0000-0001-6160-3356ORCID · verified
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 22 (3 first)Information Retrieval & Web Search · 21Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RAGXDoc: Structured Knowledge-Guided Retrieval and Explainable Re-ranking for Academic Documents
Dipendra Sharma Kafle, Esma Talhi, Mickaël Coustaty, Antoine Doucet |
ICDAR (3) | 4 |
| 2026 | One Model, Many Guidelines: Instruction Fine-Tuning for Historical Named Entity Recognition
Tien-Nam Nguyen, Emanuela Boros, Adam Jatowt, Mickaël Coustaty, Ahmed Hamdi, Antoine Doucet |
ICDAR (3) | 6 |
| 2025 | Evaluating Robustness of LLMs in Question Answering on Multilingual Noisy OCR DataabstractOptical Character Recognition (OCR) plays a crucial role in digitizing historical and multilingual documents, yet OCR errors - imperfect extraction of text, including character insertion, deletion, and substitution can significantly impact downstream tasks like question-answering (QA). In this work, we conduct a comprehensive analysis of how OCR-induced noise affects the performance of Multilingual QA Systems. To support this analysis, we introduce a multilingual QA dataset MultiOCR-QA, comprising 50K question-answer pairs across three languages, English, French, and German. The dataset is curated from OCR-ed historical documents, which include different levels and types of OCR noise. We then evaluate how different state-of-the-art Large Language Models (LLMs) perform under different error conditions, focusing on three major OCR error types. Our findings show that QA systems are highly prone to OCR-induced errors and perform poorly on noisy OCR text. By comparing model performance on clean versus noisy texts, we provide insights into the limitations of current approaches and emphasize the need for more noise-resilient QA systems in historical digitization contexts. Bhawna Piryani, Jamshid Mozafari, Abdelrahman Abdallah, Antoine Doucet, Adam Jatowt |
CIKM | 4 |
| 2025 | Multidisciplinary End-to-End Document-Level Relation Extraction from Scientific Literature
Julien Delaunay, Tran Thi Hong Hanh, Carlos E. González-Gallardo, Georgeta Bordea, Nicolas Sidere, Antoine Doucet, Olivier de Viron |
ICDAR (4) | 6 |
| 2025 | Expertise Finding: Domain Extraction from Documents Using Fuzzy Clustering
Dipendra Sharma Kafle, Esma Talhi, Mickaël Coustaty, Antoine Doucet |
ICDAR (1) | 4 |
| 2025 | Few-Shot Document Classification in Real Applications: Boosting Precision with Novelty Detection
Tri-Cong Pham, Mickaël Coustaty, Aurélie Joseph, Gaspar Deloin, Vincent Poulain D'Andecy, Antoine Doucet |
ICDAR (3) | 6 |
| 2025 | Ar-Q-Former: Historical Newspaper Article Separation Based on Multimodal Transformer Structure
Nancy Girdhar, Tran Thi Hong Hanh, Carlos E. González-Gallardo, Mickaël Coustaty, Antoine Doucet |
ICDAR (3) | 6 |
| 2024 | Leveraging Transfer Learning for Article Segmentation in Historical Newspapers
Nancy Girdhar, Deepak Sharma 0005, Mickaël Coustaty, Antoine Doucet |
TPDL (1) | 4 |
| 2024 | Leveraging Open Large Language Models for Historical Named Entity Recognition
Carlos E. González-Gallardo, Tran Thi Hong Hanh, Ahmed Hamdi, Antoine Doucet |
TPDL (1) | 4 |
| 2024 | LIT: Label-Informed Transformers on Token-Based Classification
Tran Thi Hong Hanh, Carlos E. González-Gallardo, Mickaël Coustaty, Antoine Doucet |
TPDL (1) | 5 |
| 2024 | LIAS: Layout Information-Based Article Separation in Historical Newspapers
Tran Thi Hong Hanh, Carlos E. González-Gallardo, Mickaël Coustaty, Antoine Doucet |
TPDL (1) | 5 |
| 2024 | Global-SEG: Text Semantic Segmentation Based on Global Semantic Pair Relations
Tran Thi Hong Hanh, Carlos E. González-Gallardo, Mickaël Coustaty, Antoine Doucet |
ICDAR (4) | 5 |
| 2023 | Injecting Temporal-Aware Knowledge in Historical Named Entity Recognition
Carlos E. González-Gallardo, Emanuela Boros, Edward Giamphy, Ahmed Hamdi, José G. Moreno 0001, Antoine Doucet |
ECIR (1) | 6 |
| 2023 | Analyzing the Impact of Tokenization on Multilingual Epidemic Surveillance in Low-Resource Languages
Stephen Mutuvi, Emanuela Boros, Antoine Doucet, Gaël Lejeune, Adam Jatowt, Moses Odeo |
ICDAR (3) | 3 |
| 2023 | DocILE Benchmark for Document Information Localization and Extraction
Stepán Simsa, Milan Sulc, Michal Uricár, Ahmed Hamdi, Matej Kocián, Matyás Skalický, Jiri Matas, Antoine Doucet, Mickaël Coustaty, Dimosthenis Karatzas |
ICDAR (2) | 9 |
| 2023 | Detecting Forged Receipts with Domain-Specific Ontology-Based Entities & Relations
Beatriz Martínez Tornés, Emanuela Boros, Antoine Doucet, Petra Gomez-Krämer, Jean-Marc Ogier |
ICDAR (3) | 3 |
| 2023 | Receipt Dataset for Document Forgery Detection
Beatriz Martínez Tornés, Théo Taburet, Emanuela Boros, Kais Rouis, Antoine Doucet, Petra Gomez-Krämer, Nicolas Sidere, Vincent Poulain D'Andecy |
ICDAR (3) | 5 |
| 2022 | ReadOCR: A Novel Dataset and Readability Assessment of OCRed Texts
Thi-Tuyet-Hai Nguyen, Adam Jatowt, Mickaël Coustaty, Antoine Doucet |
DAS | 4 |
| 2022 | Exploring Entities in Event Detection as Question Answering
Emanuela Boros, José G. Moreno 0001, Antoine Doucet |
ECIR (1) | 3 |
| 2022 | Introducing the HIPE 2022 Shared Task: Named Entity Recognition and Linking in Multilingual Historical Documents
Maud Ehrmann, Matteo Romanello, Antoine Doucet, Simon Clematide |
ECIR (2) | 3 |
| 2022 | Integrated interdisciplinary workflows for research on historical newspapers: Perspectives from humanities scholars, computer scientists, and librariansabstractThis article considers the interdisciplinary opportunities and challenges of working with digital cultural heritage, such as digitized historical newspapers, and proposes an integrated digital hermeneutics workflow to combine purely disciplinary research approaches from computer science, humanities, and library work. Common interests and motivations of the above-mentioned disciplines have resulted in interdisciplinary projects and collaborations such as the NewsEye project, which is working on novel solutions on how digital heritage data is (re)searched, accessed, used, and analyzed. We argue that collaborations of different disciplines can benefit from a good understanding of the workflows and traditions of each of the disciplines involved but must find integrated approaches to successfully exploit the full potential of digitized sources. The paper is furthermore providing an insight into digital tools, methods, and hermeneutics in action, showing that integrated interdisciplinary research needs to build something in between the disciplines while respecting and understanding each other's expertise and expectations. Sarah Oberbichler, Emanuela Boros, Antoine Doucet, Jani Marjanen, Eva Pfanzelter, Juha Rautiainen, Hannu Toivonen, Mikko Tolonen |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2021 | Event Detection with Entity Markers
Emanuela Boros, José G. Moreno 0001, Antoine Doucet |
ECIR (2) | 3 |
| 2021 | A Comprehensive Extraction of Relevant Real-World-Event Qualifiers for Semantic Search Engines
Cyrille Suire, Cyril Faucher, Antoine Doucet |
TPDL | 4 |
| 2021 | Token-Level Multilingual Epidemic Dataset for Event Extraction
Stephen Mutuvi, Emanuela Boros, Antoine Doucet, Gaël Lejeune, Adam Jatowt, Moses Odeo |
TPDL | 3 |
| 2021 | Information Extraction from Invoices
Ahmed Hamdi, Elodie Carel, Aurélie Joseph, Mickaël Coustaty, Antoine Doucet |
ICDAR (2) | 5 |
| 2021 | A Multilingual Dataset for Named Entity Recognition, Entity Linking and Stance Detection in Historical NewspapersabstractNamed entity processing over historical texts is more and more being used due to the massive documents and archives being stored in digital libraries. However, due to the poor annotated resources of historical nature, information extraction performances fall behind those on contemporary texts. In this paper, we introduce the development of the NewsEye resource, a multilingual dataset for named entity recognition and linking enriched with stances towards named entities. The dataset is comprised of diachronic historical newspaper material published between 1850 and 1950 in French, German, Finnish, and Swedish. Such historical resource is essential in the context of developing and evaluating named entity processing systems. It evenly allows enhancing the performances of existing approaches on historical documents which enables adequate and efficient semantic indexing of historical documents on digital cultural heritage collections. Ahmed Hamdi, Elvys Linhares Pontes, Emanuela Boros, Thi-Tuyet-Hai Nguyen, Günter Hackl, José G. Moreno 0001, Antoine Doucet |
SIGIR | 7 |
| 2020 | Assessing and Minimizing the Impact of OCR Quality on Named Entity Recognition
Ahmed Hamdi, Axel Jean-Caurant, Nicolas Sidere, Mickaël Coustaty, Antoine Doucet |
TPDL | 5 |
| 2020 | Determining image age with rank-consistent ordinal classification and object-centered ensembleabstractA significant number of old photographs including ones that are posted online do not contain the information of the date at which they were taken, or this information needs to be verified. Many of such pictures are either scanned analog photographs or photographs taken using a digital camera with incorrect settings. Estimating the date of such pictures is useful for enhancing data quality and its consistency, improving information retrieval and for other related applications. In this study, we propose a novel approach for automatic estimation of the shooting dates of photographs based on a rank-consistent ordinal classification method for neural networks. We also introduce an ensemble approach that involves object segmentation. We conclude that assuring the rank consistency in the ordinal classification as well as combining models trained on segmented objects improve the results of the age determination task. Shota Ashida, Adam Jatowt, Antoine Doucet, Masatoshi Yoshikawa |
MMAsia | 3 |
| 2019 | Document in Context of its Time (DICT): Providing Temporal Context to Support Analysis of Past DocumentsabstractOld documents tend to be difficult to be analyzed and understood, not only for average users but oftentimes for professionals as well. This is due to the context shift, vocabulary evolution and, in general, the lack of precise knowledge about the writing styles in the past. We propose a concept of positioning document in the context of its time, and develop an interactive system to support such an objective. Our system helps users to know whether the vocabulary used by an author in the past were frequent at the time of text creation, whether the author used anachronisms or neologisms, and so on. It also enables detecting terms in text that underwent considerable semantic change and provides more information on the nature of such change. Overall, the proposed tool offers additional knowledge on the writing style and vocabulary choice in documents by drawing from data collected at the time of their creation or at other user-specified time. Adam Jatowt, Ricardo Campos 0001, Sourav S. Bhowmick, Antoine Doucet |
CIKM | 4 |
| 2019 | Post-OCR Error Detection by Generating Plausible CandidatesabstractThe accuracy of Optical Character Recognition (OCR) technologies considerably impacts the way digital documents are indexed, accessed and exploited. Post-processing approaches detect and correct remaining errors to improve the quality of OCR texts. However, state-of-the-art approaches still need to be improved. Most of the existing post-OCR techniques use predefined error position lists or apply simple techniques to detect errors. In this paper, we describe a novel error detector using different features from character-level (including character noisy channel, index of peculiarity) to word-level (such as frequencies of n-grams, skip-grams, part-of-speech) Experimental results show that our approach outperforms the best performing techniques in the ICDAR 2017 Competition on Post-OCR text correction. Thi-Tuyet-Hai Nguyen, Adam Jatowt, Mickaël Coustaty, Vincent Nguyen 0001, Antoine Doucet |
ICDAR | 5 |
| 2019 | ICDAR 2019 Competition on Post-OCR Text CorrectionabstractThis paper describes the second round of the ICDAR 2019 competition on post-OCR text correction and presents the different methods submitted by the participants. OCR has been an active research field for over the past 30 years but results are still imperfect, especially for historical documents. The purpose of this competition is to compare and evaluate automatic approaches for correcting (denoising) OCR-ed texts. The present challenge consists of two tasks: 1) error detection and 2) error correction. An original dataset of 22M OCR-ed symbols along with an aligned ground truth was provided to the participants with 80% of the dataset dedicated to training and 20% to evaluation. Different sources were aggregated and contain newspapers, historical printed documents as well as manuscripts and shopping receipts, covering 10 European languages (Bulgarian, Czech, Dutch, English, Finish, French, German, Polish, Spanish and Slovak). Five teams submitted results, the error detection scores vary from 41 to 95% and the best error correction improvement is 44%. This competition, which counted 34 registrations, illustrates the strong interest of the community to improve OCR output, which is a key issue to any digitization process involving textual data. Christophe Rigaud, Antoine Doucet, Mickaël Coustaty, Jean-Philippe Moreux |
ICDAR | 2 |
| 2018 | Unsupervised Crisis Information Extraction from Twitter DataabstractWhile microblogging-based Online Social Networks have become an attractive data source in emergency situations, overcoming information overload is still not trivial. We propose a framework which integrates natural language processing and clustering techniques in order to produce a ranking of relevant tweets based on their informativeness. Experiments on four Twitter collections in two languages (English and French) proved the significance of our approach. Roberto Interdonato, Antoine Doucet, Jean-Loup Guillaume |
ASONAM | 2 |
| 2018 | Every Word has its History: Interactive Exploration and Visualization of Word Sense EvolutionabstractHuman language constantly evolves due to the changing world and the need for easier forms of expression and communication. Our knowledge of language evolution is however still fragmentary despite significant interest of both researchers as well as wider public in the evolution of language. In this paper, we present an interactive framework that permits users study the evolution of words and concepts. The system we propose offers a rich online interface allowing arbitrary queries and complex analytics over large scale historical textual data, letting users investigate changes in meaning, context and word relationships across time. Adam Jatowt, Ricardo Campos 0001, Sourav S. Bhowmick, Nina Tahmasebi, Antoine Doucet |
CIKM | 5 |
| 2018 | Feature Selection for Document Flow SegmentationabstractIn this paper, we describe a method to restore a flow of continuous documents. The flow is a collection of consecutive scanned pages without explicit separation marks between documents. Our method is based on contextual and layout descriptors meant to specify the relationship between each pair of consecutive pages. The relationships are represented using vectors of features with boolean values indicating the presence or the absence of descriptors on concerned pages. The segmentation task therefore consists in classifying such vectors into continuities or breaks. The continuity class indicates that pages belong to the same document while the break class ends the ongoing document and starts a new one. The experimental part is based on a large collection of real administrative documents. Ahmed Hamdi, Mickaël Coustaty, Aurélie Joseph, Vincent Poulain D'Andecy, Antoine Doucet, Jean-Marc Ogier |
DAS | 5 |
| 2018 | Detecting prominent microblog users over crisis events phases
Imen Bizid, Nibal Nayef, Patrice Boursier, Antoine Doucet |
Inf. Syst. | 4 |
| 2017 | ICDAR2017 Competition on Post-OCR Text CorrectionabstractThis paper describes the ICDAR2017 competition on post-OCR text correction and presents the different methods submitted by the participants. OCR has been an active research field for over the past 30 years but results are still imperfect, especially for historical documents. The purpose of this competition is to compare and evaluate automatic approaches for correcting (denoising) OCR-ed texts. The challenge consists of two independent tasks: 1) error detection and 2) error correction. An original dataset of 12M OCR-ed symbols along with an aligned ground truth was provided to the participants with 80% of the dataset dedicated to the training and 20% to the evaluation. Different sources were aggregated and namely contain newspapers and monographs covering 2 languages (English and French). 11 teams submitted results, while the difficulty of the task was underlined by the fact that only half of the submitted methods were able to denoise the evaluation dataset on average. In any case, this competition, which counted 35 registrations, illustrates the strong interest of the community in this essential problem, which is key to any digitization process involving textual data. Guillaume Chiron, Antoine Doucet, Mickaël Coustaty, Jean-Philippe Moreux |
ICDAR | 2 |
| 2017 | Enhancing Table of Contents Extraction by System AggregationabstractThe OCR-ed books usually lack logical structure information, such as chapters, sections. To enrich the navigation experience of users, several approaches have been proposed to extract table of contents (ToC) from digitised books. In this paper, we introduce an aggregation-based method to enhance ToC extraction using system submissions from the ICDAR Book structure extraction competitions (2009, 2011, and 2013). Our experimental results show that the union of two best approaches outperforms the existing approaches using both the title-based and link-based evaluation measures on a dataset of more than 2000 books. By efficiently combining the results of existing systems in an unsupervised way, we consistently beat the state-of-the-art in book structure extraction, with performance improvements that are statistically significant. Thi-Tuyet-Hai Nguyen, Antoine Doucet, Mickaël Coustaty |
ICDAR | 2 |
| 2015 | Eighth Workshop on Exploiting Semantic Annotations in Information Retrieval (ESAIR'15)abstractThe amount of structured content published on the Web has been growing rapidly, making it possible to address increasingly complex information access tasks. Recent years have witnessed the emergence of large scale human-curated knowledge bases as well as a growing array of techniques that identify or extract information automatically from unstructured and semi-structured sources. The ESAIR workshop series aims to advance the general research agenda on the problem of creating and exploiting semantic annotations. The eighth edition of ESAIR sets its focus on applications. We dedicate a special "annotations in action" track to demonstrations that showcase innovative prototype systems, in addition to the regular research and position paper contributions. The workshop also features invited talks from leaders in the field. The desired outcome of ESAIR'15 is a roadmap and research agenda that guides academic efforts and aligns them with industrial directions and developments. Krisztian Balog, Jeff Dalton 0001, Antoine Doucet, Yusra Ibrahim |
CIKM | 3 |
| 2015 | Identification of Microblogs Prominent Users during Events by Learning Temporal Sequences of FeaturesabstractDuring specific real-world events, some users of microblogging platforms could provide exclusive information about those events. The identification of such prominent users depends on several factors such as the freshness and the relevance of their shared information. This work proposes a probabilistic model for the identification of prominent users in microblogs during specific events. The model is based on learning and classifying user behavior over time using Mixture of Gaussians Hidden Markov Models. A user is characterized by a temporal sequence of feature vectors describing his activities. The features computed at each time-stamp are designed to reflect both the on- and off-topic activities of users, and they are computationally feasible in real-time. Imen Bizid, Nibal Nayef, Patrice Boursier, Sami Faïz, Antoine Doucet |
CIKM | 5 |
| 2014 | Dating Color Images with Ordinal ClassificationabstractThis paper proposes a new approach for automatically dating a photograph, based solely on its content. Building on recent advances in computer vision, the images are first described by a set of features. Then, the age group of every image is predicted by a classifier trained with annotated data. The key strength of our approach -- which makes it perform better than existing ones -- is the introduction of an ordinal classification framework, particularly adapted to the type of data to be predicted (age groups). The approach is validated on a recent challenging dataset for which it produces state-of-the-art results. Paul Martin 0001, Antoine Doucet, Frédéric Jurie |
ICMR | 2 |
| 2014 | Document summarization based on word associationsabstractIn the age of big data, automatic methods for creating summaries of documents become increasingly important. In this paper we propose a novel, unsupervised method for (multi-)document summarization. In an unsupervised and language-independent fashion, this approach relies on the strength of word associations in the set of documents to be summarized. The summaries are generated by picking sentences which cover the most specific word associations of the document(s). We measure the performance on the DUC 2007 dataset. Our experiments indicate that the proposed method is the best-performing unsupervised summarization method in the state-of-the-art that makes no use of human-curated knowledge bases. Oskar Gross, Antoine Doucet, Hannu Toivonen |
SIGIR | 2 |
| 2013 | ICDAR 2013 Competition on Book Structure ExtractionabstractThis paper summarizes the 3rd Book Structure Extraction competition that was run at the ICDAR 2013. Its goal is to evaluate and compare automatic techniques for deriving structure information from digitized books, which could then be used to aid navigation inside the books. More specifically, the task that participants are faced with is to construct hyper linked tables of contents for a collection of 1,000 digitized books. This paper reviews the setup of the competition, the book collection used in the task, and the measures used for the evaluation. The main novelty of the 2013 competition is that we were able to rely on an external provider for the ground truthing phase, hence granting the consistency of the evaluation. In addition, this permitted to nearly double the number of annotated books from the 1,040 books annotated in 2009 and 2011 to over 2,000 books. The paper further presents the result performance of the 6 participating research teams, and briefly summarizes their approaches. Antoine Doucet, Gabriella Kazai, Sebastian Colutto, Günter Mühlberger |
ICDAR | 1 |
| 2011 | ICDAR 2011 Book Structure Extraction CompetitionabstractIn this paper, we summarize the 2nd Book Structure Extraction competition run at ICDAR 2011. Its goal is to evaluate and compare automatic techniques for deriving structure information from digitized books, which could then be used to aid navigation inside the books. More specifically, the task that participants are faced with is to construct hyper linked tables of contents for a collection of 1,000 digitized books. This paper reviews the setup of the competition, the book collection used in the task, and the measures used for the evaluation. It further presents the outcome of the competition: an additional ground truth of 513 book tables of contents, contributed by 6 institutions, and the result performance of the 4 participating research teams. Antoine Doucet, Gabriella Kazai, Jean-Luc Meunier |
ICDAR | 1 |
| 2009 | ICDAR 2009 Book Structure Extraction CompetitionabstractThis paper introduces the Book Structure Extraction competition run at ICDAR 2009. The goal of the competition is to evaluate and compare automatic techniques for deriving structure information from digitized books, which could then be used to aid navigation inside the books. More specifically, the task that participants are faced with is to construct hyperlinked tables of contents for a collection of 1,000 digitized books. This paper describes the setup of the competition, the book collection used in the task, and the proposed measures for the evaluation. Results of the evaluation will be presented at the ICDAR 2009 conference and will be published in the INEX 2009 proceedings. Antoine Doucet, Gabriella Kazai, Bodin Dresevic, Aleksandar Uzelac, Bogdan Radakovic, Nikola Todic |
ICDAR | 1 |
| 2008 | XML-aided phrase indexing for hypertext documentsabstractWe combine techniques of XML Mining and Text Mining for the benefit of Information Retrieval. By manipulating the word sequence according to the XML structure of the marked-up text, we strengthen phrase boundaries so that they are more obvious to the algorithms that extract multiword sequences from text. Consequently, the quality of the indexed phrases improves, which has a positive effect on the average precision measured by the INEX 2007 standards. Miro Lehtonen, Antoine Doucet |
SIGIR | 2 |