Véronique Eglin

dblp:e/VEglin · DBLP profile ↗
← Back
29ranked-venue papers in the field
5as first author
6since 2021 · last 2026
0000-0001-8738-2088ORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 27 (5 first)Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 CCASL: Counterexamples to comparative analysis of scientific literature - Application to polymers
Aymar Tchagoue, Véronique Eglin, Sébastien Pruvost, Jean-Marc Petit, Jannick Duchet-Rumeau, Jean-François Gérard
Data Knowl. Eng.2
2025 Rx-PAD: Recognition and eXtraction - A Dataset for Prescription Analysis and Clinical Data Structuring
Jonathan Pattin Cottet, Véronique Eglin, Alex Aussem
ICDAR (5)2
2025 A Multimodal Evaluation Pipeline for Mathematical Expression Recognition: Comparisons of Datasets, Metrics, and Models
François Wieckowiak, Véronique Eglin, Tony Bonnet, Stéphane Bres, Laëtitia Rousseau
ICDAR (4)2
2023 Ensuring an Error-Free Transcription on a Full Engineering Tags Dataset Through Unsupervised Post-OCR Methods
Mathieu Francois, Véronique Eglin
ICDAR (5)2
2022 Text Detection and Post-OCR Correction in Engineering Documents
Mathieu Francois, Véronique Eglin, Maxime Biou
DAS2
2022 Challenging Children Handwriting Recognition Study Exploiting Synthetic, Mixed and Real Data
Sofiane Medjram, Véronique Eglin, Stéphane Bres
DAS2
2020 Classification of Phonetic Characters by Space-Filling Curves
Valentin Owczarek, Jordan Drapeau, Jean-Christophe Burie, Patrick Franco, Mickaël Coustaty, Rémy Mullot, Véronique Eglin
DAS7
2019 Recurrent Neural Network Approach for Table Field Extraction in Business Documents
abstract
Efficiently extracting information from documents issued by their partners is crucial for companies that face huge daily document flows. Particularly, tables contain most valuable information of business documents. However, their contents are challenging to automatically parse as tables from industrial contexts may have complex and ambiguous physical structure. Bypassing their structure recognition, we propose a generic method for end-to-end table field extraction that starts with the sequence of document tokens segmented by an OCR engine and directly tags each token with one of the possible field types. Similar to the state-of-the-art methods for non-tabular field extraction, our approach resorts to a token level recurrent neural network combining spatial and textual features. We empirically assess the effectiveness of recurrent connections for our task by comparing our method with a baseline feedforward network having local context knowledge added to its inputs. We train and evaluate both approaches on a dataset of 28,570 purchase orders to retrieve the ID numbers and quantities of the ordered products. Our method outperforms the baseline with micro F1 score on unknown document layouts of 0.821 compared to 0.764.
Clément Sage, Alex Aussem, Haytham Elghazel, Véronique Eglin, Jérémy Espinas
ICDAR4
2019 KeyWord Spotting using Siamese Triplet Deep Neural Networks
abstract
Deep neural networks has shown great success in computer vision fields by achieving considerable state-of-the-art results and are beginning to arouse big interest in the document analysis community. In this paper, we present a novel siamese deep network of three inputs that allows retrieving the most similar words to a given query. The proposed system follows a query-by-example approach according to a segmentation-based technique and aims to learn suitable representations of handwritten word images, for which a simple Euclidean distance could perform the matching. The results obtained for the George Washington dataset show the potential and the effectiveness of the proposed keyword spotting system.
Yasmine Serdouk, Véronique Eglin, Stéphane Bres, Mylène Pardoen
ICDAR2
2017 ICDAR2017 Competition on the Classification of Medieval Handwritings in Latin Script
abstract
This paper presents the results of the ICDAR2017 Competition on the Classification of Medieval Handwritings in Latin Script (CLaMM), jointly organized by Computer Scientists and Humanists (paleographers). This work follows a competition at ICFHR2016 and aims at providing a rich annotated database of European medieval manuscripts to the community on Handwriting Analysis and Recognition. We proposed four independent classification tasks which attracted 10 registered teams, with 6 submitted classifiers from 4 participants. Those classifiers are trained on a set of 3540 images with their ground truths. In task 1 (Script classification) and task 3 (Date classification), the classifiers have been evaluated by a test set of 2000 greyscale, tiff, 300 dpi images. In task 2 (Script classification) and task 4 (Date classification), the test set consists of 1000 images in different formats, resolutions and color representation. The best scores are respectively 85.2% for task 1, 76.5% for task 2, 59% for task 3, and 49.9% for task 4. An analysis based on the matrix of confusion of each classifier is also given.
Florence Cloppet, Véronique Eglin, Marlene Helias-Baron, Van Cuong Kieu, Nicole Vincent, Dominique Stutzmann
ICDAR2
2017 Discovering Motifs with Variants in Music Databases
Riyadh Benammar, Christine Largeron, Véronique Eglin, Mylène Pardoen
IDA3
2014 A Novel Learning-Free Word Spotting Approach Based on Graph Representation
abstract
Effective information retrieval on handwritten document images has always been a challenging task. In this paper, we propose a novel handwritten word spotting approach based on graph representation. The presented model comprises both topological and morphological signatures of handwriting. Skeleton-based graphs with the Shape Context labelled vertexes are established for connected components. Each word image is represented as a sequence of graphs. In order to be robust to the handwriting variations, an exhaustive merging process based on DTW alignment result is introduced in the similarity measure between word images. With respect to the computation complexity, an approximate graph edit distance approach using bipartite matching is employed for graph matching. The experiments on the George Washington dataset and the marriage records from the Barcelona Cathedral dataset demonstrate that the proposed approach outperforms the state-of-the-art structural methods.
Peng Wang 0006, Véronique Eglin, Christophe Garcia, Christine Largeron, Josep Lladós 0001, Alicia Fornés
Document Analysis Systems2
2013 Document Classification in a Non-stationary Environment: A One-Class SVM Approach
abstract
In this paper, we investigate a specific area of document classification in which the documents come as a flow over the time. Moreover, the exact number of classes of document to deal with is not known from the beginning and could evolve over the time. To be able to perform classification task in such area, we need specific classifiers that are able to perform incremental learning and change their modeling over the time. More specifically, we are focusing our study on SVM approaches, known to perform well, and for which incremental (i-SVM) procedures exist. Nevertheless, most of them are only able to deal with a fixed number of classes. So we designed a new incremental learning procedure based on one-class SVMs. This one is able to improve its classification accuracy over the time, with the arrival of new labeled data, without performing any complete retraining. Moreover, when instances are coming with a previously unknown label (appearance of a new class), the training procedure is able to modify the classifier model to recognize this corresponding new kind of documents. To investigate this area, waiting for collecting documents images as a flow, we did first experiments on the Optical Recognition of Handwritten Digits Data Set. These experiments show that our incremental approach is able: to perform, at each time, as well as a static one-class classifier fully retrained using all previously seen data, to model very quickly and efficiently new incoming classes.
Anh Khoi Ngo Ho, Nicolas Ragot, Jean-Yves Ramel, Véronique Eglin, Nicolas Sidere
ICDAR4
2013 A Comprehensive Representation Model for Handwriting Dedicated to Word Spotting
abstract
In this paper, we propose an original representation model for handwriting document images. Most state-of-the-art handwriting representation models only use separately textural properties, selective dominant features (such as stroke orientation or gradient orientation) or structural properties. To avoid the drawbacks of using the properties from a single aspect, we design a comprehensive model that contains both morphological and topological information of handwriting. After interest points (the starting/ending points, branch points and high-curved points) are selected, an adapted version of Shape Context (SC) descriptor built on the interest points is employed to describe the contour of the text. In order to model the structural characteristics of the handwritten text, a graph is constructed based on the interest points and the skeleton of the text. With the graph, loops and specific strokes in the handwriting are detected and analyzed. Based on this model, a coarse-to-fine approach for word spotting application is introduced. Without segmenting texts into words, a group of regions of interest are selected by comparing textural features (orientation, projection profile, upper and lower border projection) using the DTW method. Afterwards, regions of interest and queries are represented by the proposed model. The final similarity measure is a weighted mixture of the SC cost, loop difference, stroke analysis and texture comparison with different weights. The validation of the model shows the significance of combining the various properties of the handwriting envisaged in its different aspects.
Peng Wang 0006, Véronique Eglin, Christophe Garcia, Christine Largeron, Antony McKenna
ICDAR2
2011 A Mixed Approach for Handwritten Documents Structural Analysis
abstract
In this paper we propose a new method for document pages segmentation. First dedicated to handwritten documents, our method is designed to extract the different text zones, paragraph and fragment in unconstrained documents. The proposed approach is a mixed one, using both the advantages of top-down and bottom-up approaches. In this paper we proposed and evaluation of our methods on a 183 documents database, taken from a 19th century handwritten corpus : the "dossiers de Bouvard et Pécuchet" from Flaubert. With this evaluation we demonstrate that the combination of the top-down and the bottom-up approach allow to improve the obtained results.
Vincent Malleron, Véronique Eglin
ICDAR2
2010 A new approach for centerline extraction in handwritten strokes: an application to the constitution of a code book
abstract
It is the pleasure of the organizing committee to welcome all participants to the 2010 IAPR Workshop on Document Analysis Systems (DAS). This year's workshop is being held June 9-11th in Boston, Massachusetts, a location situated in the heart of beautiful New England in the northeastern United States. Boston has a rich history that dates back to the 1600's and was a major focal point in the history of the American Revolution and America's fight for independence. Over the years, Boston has developed into a center for industrial, academic and cultural excellence and draws millions of visitors each year. DAS 2010 is the ninth workshop in a series. The first DAS was held in Kaiserslautern, Germany in 1994, and was followed by Malvern, PA (1996); Nagano, Japan (1998); Rio de Janeiro, Brazil (2000); Princeton, NJ (2002); Florence, Italy (2004); Nelson, New Zealand (2006) and Nara, Japan (2008). The DAS tradition is to bring together industry, academic and government researchers interested in many aspects of document analysis systems and to provide opportunities for fruitful interaction and collaboration. This year's workshop is organized as a three-day, single track event, with oral and poster presentations, as well as working group discussions on the second and third afternoons. Special sessions on contributed datasets and a keynote talk on this same topic provide a compelling theme, and we hope this focus will help push the field toward greater sharing of data and accepted standards for evaluation. This year, we received 91 submissions from 25 countries on six continents. The papers were reviewed by 45 members of our research community and an international program committee representing 16 different countries. The overall quality was excellent and we have chosen 28 full papers for oral presentation and 37 as poster papers, as well as 15 short papers that will be presented either as posters or in a special short oral format. In addition, six groups are scheduled to present live demos during the poster sessions. Full papers underwent the standard peer review process and will appear in the official workshop proceedings to be published in the ACM International Conference Proceedings Series, available online as part of the ACM Digital Library. The short papers are included in the unofficial hardcopy proceedings distributed at the event as well as on the DAS 2010 website.
Hani Daher, Véronique Eglin, Stéphane Bres, Nicole Vincent
Document Analysis Systems2
2009 Graph b-Coloring for Automatic Recognition of Documents
abstract
In order to reduce the rejection rate of our automatic reading system, we propose to pre-classify the business documents by introducing an automatic recognition of documents stage (ARD) as a pre-processing step. This important step will guide the other stages involved in the recognition process of the documents contents. Once the document class identified, the reading system will use correct information from the ARD stage to improve the segmentation of the layout, the recognition of the document structure, the parameterization of the OCR, and the final decision for the rejection. We propose in this paper an original method for the classification of business documents suited for complex layouts having great variability. We introduce the graph coloring approach for both layout analysis and document classification. The proposed method is reliable, robust to various constraints and guarantees a real-time answer to the sorting of business documents.
Djamel Gaceb, Véronique Eglin, Frank Lebourgeois, Hubert Emptoz
ICDAR2
2009 Text Lines and Snippets Extraction for 19th Century Handwriting Documents Layout Analysis
abstract
In this paper we propose a new approach to improve electronic editions of human science corpus, providing an efficient estimation of manuscripts pages structure. In any handwriting documents analysis process, the text line segmentation is an important stage. The presence of variable inter-line spaces, of inconstant base-line skews, overlapping and occlusions in unconstrained ancient 19th handwritten documents complexifies the text lines segmentation task. In this paper, we only use as prior knowledge of script the fact that text lines skews can be random and irregular.In that context, we model text line detection as an image segmentation problem by enhancing text line structure using Hough transform and a clustering of connected components so as to make text line boundaries appear. The proposed approach of snippets decomposition for page layout analysislies on a first step of content pages classification in five visual and genetic taxonomies, and a second step of text line extraction and snippets decomposition. Experiments show that the proposed method achieves high accuracy for detecting text lines in regular and semi-regular handwritten pages in the corpus of digitized Flaubert manuscripts (”Dossiers documentaires de Bouvard et Pécuchet”, 1872-1880).
Vincent Malleron, Véronique Eglin, Hubert Emptoz, Stéphanie Dord-Crouslé, Philippe Régnier
ICDAR2
2008 Physical Layout Segmentation of Mail Application Dedicated to Automatic Postal Sorting System
abstract
Every day, the postal sorting systems diffuse several tons of mails. It is noted that the principal origin of mail rejection is related to the failure of address-block localization task, particularly, of the physical layout segmentation stage. The bottom-up and top-down segmentation methods bring different knowledge that should not be ignored when we need to increase the robustness. Hybrid methods combine the two strategies in order to take advantages of one strategy to the detriment of other. Starting from these remarks, our proposal makes use of a hybrid segmentation strategy more adapted to the postal mails. The high level stages are based on the hierarchical graphs coloring, allowing managing through a pyramidal data organization, the complex rules leading the interpretation of the connected components decomposition of interest zones. Today, no other work in this context has make use of the powerfulness of this tool. The performance evaluation of our approach was tested on a corpus of 10000 envelope images. The processing times and the rejection rate were considerably reduced.
Djamel Gaceb, Véronique Eglin, Frank Lebourgeois, Hubert Emptoz
Document Analysis Systems2
2007 Writer Identification Using Steered Hermite Features and SVM
abstract
Writer recognition is considered as a difficult problem to solve due to variations found in the writing, even from the same writer. In this paper, Steered Hermite Features are used to identify writer from a written document. We will show that Steered Hermite Features are highly useful for text images because they extract lot of information, no- tably for data characterized by oriented features, curves and segments. The algorithm we propose here, first calcu- lates the Steered Hermite Features of the images which are then passed on to Support Vector Machine for training and testing. The base of tests consists of sample of some lines of writings (five at most) of primarily diversified writings of authors from IAM database. With the proposed algorithm based on Steered Hermite Features, we were able to achieve an accuracy of around 83% percent for a set of 30 authors with non overlapping images of written text.
A. Imdad, Stéphane Bres, Véronique Eglin, C. Rivero-Moreno, Hubert Emptoz
ICDAR3
2007 A Proposition of Retrieval Tools for Historical Document Images Libraries
abstract
In this article, we propose a method of characterization of pictures of old documents based on a texture approach. This characterization is carried out with the help of a multi- resolution study of the textures contained in the pictures of the document. So, by extracting five features linked to the frequencies and to the orientations in the different parts of a page, it is possible to extract and to compare elements of high semantic level without expressing any hypothesis about the physical or logical structure of the analysed documents. Experiments show the feasibility of the fulfillment of tools for the navigation or the indexation help. In these experimentations, we will lay the emphasis upon the pertinence of these texture features and the advances that they represent in terms of characterization of content of a deeply heterogeneous corpus.
Nicholas Journet, Jean-Yves Ramel, Rémy Mullot, Véronique Eglin
ICDAR4
2007 Curvelets Based Queries for CBIR Application in Handwriting Collections
abstract
This paper presents a new use of the curvelet transform as a multiscale method for indexing linear singularities and curved handwritten shapes in documents images. As it belongs to the wavelet family, this representation can be useful at several scales of details. The proposed scheme for handwritten shape characterization targets to detect oriented and curved fragments at different scales so as to compose an unique signature for each handwritten analyzed samples. In this way, curvelets coefficients are used as a representation tool for handwriting when searching in large manuscripts databases by finding similar handwritten samples. Current results of ancient manuscripts retrieval are very promising with very satisfying precisions and recalls.
Guillaume Joutel, Véronique Eglin, Stéphane Bres, Hubert Emptoz
ICDAR2
2005 Frequencies Decomposition and Partial Similarities Retrieval for Ancient Handwriting Documents Compression
abstract
This paper presents a new segmentation free approach of partial similarities retrieval in ancient handwritten documents. The method has been developed to improve usual handwritings compression approaches that are not adapted to patrimonial images specificities. We present here the similarities characterization that lies on oriented handwriting shapes decomposition. The frequencies page decomposition realizes a pavement of handwritten regions stored in directional maps where partial similarities are estimated. This decomposition is obtained by a frequencies analysis implying Gabor bank filters with an adequate parameter setting based on the most significant directions of the text. For each map, we compute a similarity graph that reveals redundant shapes and determines a resulting redundancy rate. The resulting graph is the first part of the compression system currently under development.
Abir El Abed, Véronique Eglin, Frank Lebourgeois, Hubert Emptoz
ICDAR2
2005 Biological inspired Tools for Patrimonial Handwriting Denoising and Categorization
abstract
In this paper, we propose a global segmentation free methodology for patrimonial documents denoising, handwriting characterization and categorization that are based on biological inspired approaches. We widely used here the spectral domain of handwritten images by frequency decompositions: Hermite transforms and Gabor bank filters. Handwritten pages are described by a multiscale signature that is based on orientation features and that is at the basis of a similarity measure. The current results of handwriting categorization and indexing are very promising and show that it is possible to analyze handwritten drawings without any a priori graphemes segmentation.
Véronique Eglin, Stéphane Bres, Carlos Rivero, Hubert Emptoz
ICDAR1
2005 Text/Graphic labelling of Ancient Printed Documents
abstract
This paper presents a text/graphic labelling for ancient printed documents. Our approach is based on the extraction and the quantification of the various orientations that are present in ancient printed document images. The documents are initially cut into normalized square windows in which we analyze significant orientations with a directional rose. Each kind of information (textual or graphical) is typically identified and marked by its orientation distribution. This choice of characterization allows us to separate textual regions from graphics by minimizing the a priori knowledge. The evaluation of our proposition lies on a page classification using layout extraction criteria. The system has been tested over several ancient printed books of the Renaissance.
Nicholas Journet, Véronique Eglin, Jean-Yves Ramel, Rémy Mullot
ICDAR2
2004 Multiscale Handwriting Characterization for Writers' Classification
Véronique Eglin, Stéphane Bres, Carlos Rivero
Document Analysis Systems1
2003 Document page similarity based on layout visual saliency: Application to query by example and document classification
abstract
In this paper we propose to define a measure of visualsimilarity to compare different pages in a corpus. Thismeasure is based on the analysis of the visual layoutsaliency of the page composition. This similarity iscomputed using both the document layout andcharacteristics of the text itself. The text characterizationuses statistical features derived from textural primitives.Our purpose is to establish perceptive links betweendocuments in order to facilitate their storage and theirretrieval. In this paper we present two possibleapplications of this measure of similarity: the query ofthe corpus by example and the documents classification.In the first application, we extract documents that are themost visually similar to a document, given as query. Inthe second application, the similarity measure is used toclassify the document under investigation using its visualsimilarity to a reference set of documents. Our test corpusis extracted from the Finland MTDB Oulu multi-genredatabase that provides a great diversity of page layoutsand contents.
Véronique Eglin, Stéphane Bres
ICDAR1
2001 Visual Exploration and Functional Document Labeling
abstract
This paper presents a new approach to textual data labeling based on texture analysis. Texture is used here to show the impact of document composition on visual exploration. We demonstrate how textural properties are well adapted to typography characterization by categorizing document regions into visual text classes (such as headings, head- and footnotes, paragraphs, abstracts, etc.). We reference and classify different types of text fonts according to their visual aspect and the visual impression that emerges from the textual data. Experiments on a set of various document images show a good accuracy and robustness for our method.
Véronique Eglin, Antoine Gagneux
ICDAR1
1997 Logarithmic Spiral Grid and Gaze Control for the Development of Strategies of Visual Segmentation on a Document
abstract
The paper presents a page segmentation method which is based on perception phenomena and displays the unequal importance of information in the visual field. The access of information is directly linked to the search of attractive areas. This search is based on the idea of freeing oneself from an unbending physical structure and from a uniform vertical and horizontal scanning of the document, so as to classify the data in order of importance and interest. Using a space variant geometry for block selection, the page image, instead of being represented by a bitmap format, can be abstractly represented by the block format. This space variant geometry lays a sound basis for elaborating the kinetics of the ocular shifting on a document, which provides not only a meaningless document representation in blocks, but shows a unified view corresponding to the integration of time variant representations of the same visual field.
Véronique Eglin, Hubert Emptoz
ICDAR1