Mélodie Boillet

dblp:251/2008 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-0618-7852ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 METATR: A Multilingual, Evolving Benchmark for Automatic Text Recognition
Mélodie Boillet, Solène Tarride, Christopher Kermorvant
ICDAR (3)1
2024 The Socface Project: Large-Scale Collection, Processing, and Analysis of a Century of French Censuses
Mélodie Boillet, Solène Tarride, Yoann Schneider, Bastien Abadie, Lionel Kesztenbaum, Christopher Kermorvant
ICDAR (3)1
2024 Improving Automatic Text Recognition with Language Models in the PyLaia Open-Source Library
Solène Tarride, Yoann Schneider, Marie Generali-Lince, Mélodie Boillet, Bastien Abadie, Christopher Kermorvant
ICDAR (5)4
2023 Key-Value Information Extraction from Full Handwritten Pages
Solène Tarride, Mélodie Boillet, Christopher Kermorvant
ICDAR (2)2
2023 SIMARA: A Database for Key-Value Information Extraction from Full-Page Handwritten Documents
Solène Tarride, Mélodie Boillet, Jean-François Moufflet, Christopher Kermorvant
ICDAR (3)2
2023 Large-scale genealogical information extraction from handwritten Quebec parish records
Solène Tarride, Martin Maarand, Mélodie Boillet, James McGrath, Eugénie Capel, Hélène Vézina, Christopher Kermorvant
Int. J. Document Anal. Recognit.3
2023 Confidence Estimation for Object Detection in Document Images
Mélodie Boillet, Christopher Kermorvant, Thierry Paquet
Pattern Recognit. Lett.1
2022 Robust text line detection in historical documents: learning and evaluation methods
Mélodie Boillet, Christopher Kermorvant, Thierry Paquet
Int. J. Document Anal. Recognit.1
2020 Multiple Document Datasets Pre-training Improves Text Line Detection With Deep Neural Networks
abstract
In this paper, we introduce a fully convolutional network for the document layout analysis task. While state-of-the-art methods are using models pre-trained on natural scene images, our method Doc-UFCN relies on a U-shaped model trained from scratch for detecting objects from historical documents. We consider the line segmentation task and more generally the layout analysis problem as a pixel-wise classification task then our model outputs a pixel-labeling of the input images. We show that Doc-UFCN outperforms state-of-the-art methods on various datasets and also demonstrate that the pre-trained parts on natural scene images are not required to reach good results. In addition, we show that pre-training on multiple document datasets can improve the performances. We evaluate the models using various metrics to have a fair and complete comparison between the methods.
Mélodie Boillet, Christopher Kermorvant, Thierry Paquet
ICPR1
2020 Books of Hours. the First Liturgical Data Set for Text Segmentation
abstract
The Book of Hours was the bestseller of the late Middle Ages and Renaissance. It is a historical invaluable treasure, documenting the devotional practices of Christians in the late Middle Ages. Up to now, its textual content has been scarcely studied because of its manuscript nature, its length and its complex content. At first glance, it looks too standardized. However, the study of book of hours raises important challenges: (i) in image analysis, its often lavish ornamentation (illegible painted initials, line-fillers, etc.), abbreviated words, multilingualism are difficult to address in Handwritten Text Recognition (HTR); (ii) its hierarchical entangled structure offers a new field of investigation for text segmentation; (iii) in digital humanities, its textual content gives opportunities for historical analysis. In this paper, we provide the first corpus of books of hours, which consists of Latin transcriptions of 300 books of hours generated by Handwritten Text Recognition (HTR) - that is like Optical Character Recognition (OCR) but for handwritten and not printed texts. We designed a structural scheme of the book of hours and annotated manually two books of hours according to this scheme. Lastly, we performed a systematic evaluation of the main state of the art text segmentation approaches.
Amir Hazem, Béatrice Daille, Christopher Kermorvant, Dominique Stutzmann, Marie-Laurence Bonhomme, Martin Maarand, Mélodie Boillet
LREC7