Véronique Eglin

dblp:e/VEglin · DBLP profile ↗
← Back
52ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0001-8738-2088ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 29 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CCASL: Counterexamples to comparative analysis of scientific literature - Application to polymers
Aymar Tchagoue, Véronique Eglin, Sébastien Pruvost, Jean-Marc Petit, Jannick Duchet-Rumeau, Jean-François Gérard
Data Knowl. Eng.2
2025 Unsupervised Energy-Based Model for the Identification of Out-Of-Distribution Copy Detection Patterns
abstract
The recent advances in deep neural networks have led us to revisit the safety of current mechanisms for the authentification of people’s identities and validating the provenance of traded goods. This paper studies the problem of identifying counterfeits of Copy Detection Patterns (CDPs). CDPs are used on product packaging to identify counterfeiting by relying on the information loss principle. Current verification techniques rely on a supervised machine learning paradigm, which needs examples from real and fake CDPs to learn a model capable of differentiating both classes of CDPs. This paper proposes an unsupervised forensic approach that is capable of training an Energy-Based Model for detecting out-of-distribution (OOD) images (fake CDP examples in our case) using only original CDP images. This paper also introduces a novel thresholding technique that only requires original CDPs for threshold value selection. We validate our approach on the Indigo dataset, with results demonstrating comparable counterfeit detection capabilities to prior work but using a method trained on less data.
Marc Chapus, Carlos Fernando Crispim, Véronique Eglin, Atilla Baskurt
AVSS3
2025 Rx-PAD: Recognition and eXtraction - A Dataset for Prescription Analysis and Clinical Data Structuring
Jonathan Pattin Cottet, Véronique Eglin, Alex Aussem
ICDAR (5)2
2025 A Multimodal Evaluation Pipeline for Mathematical Expression Recognition: Comparisons of Datasets, Metrics, and Models
François Wieckowiak, Véronique Eglin, Tony Bonnet, Stéphane Bres, Laëtitia Rousseau
ICDAR (4)2
2024 Study's system: a three-stage system for children handwriting recognition and spelling error detection
Sofiane Medjram, Véronique Eglin, Stéphane Bres
Multim. Tools Appl.2
2023 Ensuring an Error-Free Transcription on a Full Engineering Tags Dataset Through Unsupervised Post-OCR Methods
Mathieu Francois, Véronique Eglin
ICDAR (5)2
2022 Text Detection and Post-OCR Correction in Engineering Documents
Mathieu Francois, Véronique Eglin, Maxime Biou
DAS2
2022 Challenging Children Handwriting Recognition Study Exploiting Synthetic, Mixed and Real Data
Sofiane Medjram, Véronique Eglin, Stéphane Bres
DAS2
2021 One-Class Detection and Classification of Defects on Concrete Surfaces
abstract
Today, railway infrastructure is subject to regular inspections carried out by image acquisition. Experts need automatic tools to process these images and extract defects. These defects can be classified in two categories : linear defects, like cracks, and surface defects, like spall, humidity, graffiti… Cracks detection has been presented in a previous work, and is based on Local Binary Patterns (LBP) extraction and our FLASH algorithm [1]. In this paper, we propose a new method, using the same LBP descriptors, already extracted, to detect surface defects using generative adversarial networks (GAN). The detection is followed by a classification with a potentially scalable system. Experiments show that our detector achieves better performance than the state-of-the art in application to defect detection.
Yannick Faula, Véronique Eglin, Stéphane Bres
SMC2
2020 Classification of Phonetic Characters by Space-Filling Curves
Valentin Owczarek, Jordan Drapeau, Jean-Christophe Burie, Patrick Franco, Mickaël Coustaty, Rémy Mullot, Véronique Eglin
DAS7
2019 Recurrent Neural Network Approach for Table Field Extraction in Business Documents
abstract
Efficiently extracting information from documents issued by their partners is crucial for companies that face huge daily document flows. Particularly, tables contain most valuable information of business documents. However, their contents are challenging to automatically parse as tables from industrial contexts may have complex and ambiguous physical structure. Bypassing their structure recognition, we propose a generic method for end-to-end table field extraction that starts with the sequence of document tokens segmented by an OCR engine and directly tags each token with one of the possible field types. Similar to the state-of-the-art methods for non-tabular field extraction, our approach resorts to a token level recurrent neural network combining spatial and textual features. We empirically assess the effectiveness of recurrent connections for our task by comparing our method with a baseline feedforward network having local context knowledge added to its inputs. We train and evaluate both approaches on a dataset of 28,570 purchase orders to retrieve the ID numbers and quantities of the ordered products. Our method outperforms the baseline with micro F1 score on unknown document layouts of 0.821 compared to 0.764.
Clément Sage, Alex Aussem, Haytham Elghazel, Véronique Eglin, Jérémy Espinas
ICDAR4
2019 KeyWord Spotting using Siamese Triplet Deep Neural Networks
abstract
Deep neural networks has shown great success in computer vision fields by achieving considerable state-of-the-art results and are beginning to arouse big interest in the document analysis community. In this paper, we present a novel siamese deep network of three inputs that allows retrieving the most similar words to a given query. The proposed system follows a query-by-example approach according to a segmentation-based technique and aims to learn suitable representations of handwritten word images, for which a simple Euclidean distance could perform the matching. The results obtained for the George Washington dataset show the potential and the effectiveness of the proposed keyword spotting system.
Yasmine Serdouk, Véronique Eglin, Stéphane Bres, Mylène Pardoen
ICDAR2
2018 A Fast Local Analysis by Thresholding applied to image matching
abstract
Key structures extraction and matching are key steps in computer vision. Many fields of application need large image acquisition and fast extraction of fine structures. In this study, we focus on situations where existing local feature extractors give not enough satisfying results concerning both accuracy and time processing. Among good illustrations, we can quote short-line extraction in local weakly-contrasted images. We propose a new Fast Local Analysis by threSHolding (FLASH) designed to process large images under hard time constraints. We use “micro-line” points as key feature. These are used for shape reconstruction (like lines) and local signature design. We apply FLASH on the field of concrete infrastructure monitoring where robots and UAVs are more and more used for automated defect detection (like cracks). For large concrete surfaces, there are several hard constraints such as the computational time and the reliability. Results show us that the computations are faster than several existing algorithms in image matching and FLASH has invariance to rotation, partial occlusion, and scale range from 0.7 to 1.4 without scale-space exploration.
Yannick Faula, Stéphane Bres, Véronique Eglin
ICPR3
2017 ICDAR2017 Competition on the Classification of Medieval Handwritings in Latin Script
abstract
This paper presents the results of the ICDAR2017 Competition on the Classification of Medieval Handwritings in Latin Script (CLaMM), jointly organized by Computer Scientists and Humanists (paleographers). This work follows a competition at ICFHR2016 and aims at providing a rich annotated database of European medieval manuscripts to the community on Handwriting Analysis and Recognition. We proposed four independent classification tasks which attracted 10 registered teams, with 6 submitted classifiers from 4 participants. Those classifiers are trained on a set of 3540 images with their ground truths. In task 1 (Script classification) and task 3 (Date classification), the classifiers have been evaluated by a test set of 2000 greyscale, tiff, 300 dpi images. In task 2 (Script classification) and task 4 (Date classification), the test set consists of 1000 images in different formats, resolutions and color representation. The best scores are respectively 85.2% for task 1, 76.5% for task 2, 59% for task 3, and 49.9% for task 4. An analysis based on the matrix of confusion of each classifier is also given.
Florence Cloppet, Véronique Eglin, Marlene Helias-Baron, Van Cuong Kieu, Nicole Vincent, Dominique Stutzmann
ICDAR2
2017 Discovering Motifs with Variants in Music Databases
Riyadh Benammar, Christine Largeron, Véronique Eglin, Mylène Pardoen
IDA3
2017 A multi-one-class dynamic classifier for adaptive digitization of document streams
Anh Khoi Ngo Ho, Véronique Eglin, Nicolas Ragot, Jean-Yves Ramel
Int. J. Document Anal. Recognit.2
2016 SmartATID: A Mobile Captured Arabic Text Images Dataset for Multi-purpose Recognition Tasks
abstract
Today's smartphones are able to capture documents with a good and simple way as any personal scanners. The captured document images need to be processed by specific and automated document processing systems. The systems are dedicated to textual content analysis, indexing and recognition. For instance, they may be used for font identification, writer identification and word or line segmentation. The state-of-the-art works lack comprehensive database for Arabic document images which are captured by mobile phones. This paper presents the first public offline images database for both printed and handwriting Arabic mobile captured documents, named "SmartATID". The document images of the database are acquired under varying capture conditions (blur, perspective angles and light). This causes photometric and geometric distortions that influence the performance of OCR process but also the page segmentation in lines and paragraphs. Each document image of our database is provided with a ground truth file that contains the exact text transcription and all numerical capture parameters used for each image capture. The database is freely and publicly usable by the research community at the following address http:// sites.google.com/site/smartatid.
Fatma Chabchoub, Yousri Kessentini, Slim Kanoun, Véronique Eglin, Frank Lebourgeois
ICFHR4
2016 ICFHR2016 Competition on the Classification of Medieval Handwritings in Latin Script
abstract
This paper presents the results of the ICFHR2016 Competition on the Classification of Medieval Handwritings in Latin Script (CLaMM), jointly organized by Computer Scientists and Humanists (paleographers). This work aims at providing a rich database of European medieval manuscripts to the community on Handwriting Analysis and Recognition. At this competition, we proposed two independent classification tasks which attracted five participants with seven submitted classifiers. Those classifiers are trained on a set of 2000 images with their ground truths. In the first task of script crisp classification, the classifiers have been evaluated on a test set of 1000 single-type manuscripts. In the second task of "Fuzzy Classification", the classifiers have been carried out on a set of 2000 multi-script-type manuscripts. The results of the participants provide the first baseline evaluation up to the accuracy score of 83.9% for the task 1 and to the fuzzy weighted score of 2.96/4 for the task 2. An analysis based on the intra-class distance and matrix of confusion of each classifier is also given.
Florence Cloppet, Véronique Eglin, Van Cuong Kieu, Dominique Stutzmann, Nicole Vincent
ICFHR2
2016 libcrn, an Open-Source Document Image Processing Library
abstract
In this paper we introduce libcrn, a multiplatform open-source document image processing library aimed at researchers and companies. It is written in C++11 and has a non-contaminating license that makes it available for use in any project without legal constraints. The features include low-level image processing (color format conversion, binarization, convolution, PDE…), document images specific tools (connected components extraction, recursive block description, PDF export…), maths (matrix arithmetics, linear algebra, GMMs, equation solvers…), classification and clustering (kNN, k-means, HMMs…). The API is comprehensively documented and libcrn's architecture follows modern C++ guidelines to facilitate the handling of the library and enforce its safe usage. A sample OCR, which is only 30 lines long, is described to illustrate libcrn's scope of possibilities.
Yann Leydier, Jean Duong, Stéphane Bres, Véronique Eglin, Frank Lebourgeois, Martial Tola
ICFHR4
2014 A Novel Learning-Free Word Spotting Approach Based on Graph Representation
abstract
Effective information retrieval on handwritten document images has always been a challenging task. In this paper, we propose a novel handwritten word spotting approach based on graph representation. The presented model comprises both topological and morphological signatures of handwriting. Skeleton-based graphs with the Shape Context labelled vertexes are established for connected components. Each word image is represented as a sequence of graphs. In order to be robust to the handwriting variations, an exhaustive merging process based on DTW alignment result is introduced in the similarity measure between word images. With respect to the computation complexity, an approximate graph edit distance approach using bipartite matching is employed for graph matching. The experiments on the George Washington dataset and the marriage records from the Barcelona Cathedral dataset demonstrate that the proposed approach outperforms the state-of-the-art structural methods.
Peng Wang 0006, Véronique Eglin, Christophe Garcia, Christine Largeron, Josep Lladós 0001, Alicia Fornés
Document Analysis Systems2
2014 Arabic Font Recognition Based on a Texture Analysis
abstract
Existing works on the font recognition and based on texture analysis often used Gray Level Cooccurence Matrix (GLCM), Gabor Filters (GF) and wavelet. In this paper, we use Steer able Pyramid (SP) for texture analysis of Arabic homogeneous and normalized text block in order to font recognition. In this frameworks, we use K Nearest Neighbors (KNN) and Back-propagation Artificial Neural Network (BpANN) for classification. The Obtained experimental results on the APTID/MF database (Arabic Printed Text Image/ Multi-Font) are encouragents.
Faten Kallel Jaiem, Slim Kanoun, Véronique Eglin
ICFHR3
2014 Learning-Free Text-Image Alignment for Medieval Manuscripts
abstract
In this paper, we describe a new approach for text-image alignment of middle-age documents. The method is dedicated to word-to-word alignment in a segmentation-free and learning-free way. The best word-to-word matching lies on an adapted editing cost between signatures extracted on Unicode characters and on the images. The results are evaluated on the "Queste del saint Graal" (13th c.) by palaeographers through an intuitive validation interface that also offers a very fast interactive correction process. The gain of time resulting from the absence of learning stage offers the opportunity to pay more attention to the integration of different specificities and variations in middle-age handwritten documents (presence of typical abbreviations, allographs).
Yann Leydier, Véronique Eglin, Stéphane Bres, Dominique Stutzmann
ICFHR2
2014 Handwritten word spotting based on a hybrid optimal distance
abstract
In this paper, we develop a comprehensive representation model for handwriting, which contains both morphological and topological information. An adapted Shape Context descriptor built on structural points is employed to describe the contour of the text. Graphs are first constructed by using the structural points as nodes and the skeleton of the strokes as edges. Based on graphs, Topological Node Features (TNFs) of n-neighbourhood are extracted. Bag-of-Words representation model based on the TNFs is employed to depict the topological characteristics of word images. Moreover, a novel approach for word spotting application by using the proposed model is presented. The final distance is a weighted mixture of the SC cost, and the TNF distribution comparison. Linear Discriminant Analysis (LDA) is used to learn the optimal weight for each part of the distance with the consideration of writing styles. The evaluation of the proposed approach shows the significance of combining the properties of the handwriting from different aspects.
Peng Wang 0006, Véronique Eglin, Christine Largeron, Christophe Garcia
ICIP2
2014 A Coarse-to-Fine Word Spotting Approach for Historical Handwritten Documents Based on Graph Embedding and Graph Edit Distance
abstract
Effective information retrieval on handwritten document images has always been a challenging task, especially historical ones. In the paper, we propose a coarse-to-fine handwritten word spotting approach based on graph representation. The presented model comprises both the topological and morphological signatures of the handwriting. Skeleton-based graphs with the Shape Context labelled vertexes are established for connected components. Each word image is represented as a sequence of graphs. Aiming at developing a practical and efficient word spotting approach for large-scale historical handwritten documents, a fast and coarse comparison is first applied to prune the regions that are not similar to the query based on the graph embedding methodology. Afterwards, the query and regions of interest are compared by graph edit distance based on the Dynamic Time Warping alignment. The proposed approach is evaluated on a public dataset containing 50 pages of historical marriage license records. The results show that the proposed approach achieves a compromise between efficiency and accuracy.
Peng Wang 0006, Véronique Eglin, Christophe Garcia, Christine Largeron, Josep Lladós 0001, Alicia Fornés
ICPR2
2013 Exploring Interest Points and Local Descriptors for Word Spotting Application on Historical Handwriting Images
Peng Wang 0006, Véronique Eglin, Christine Largeron, Antony McKenna, Christophe Garcia
CAIP (2)2
2013 Document Classification in a Non-stationary Environment: A One-Class SVM Approach
abstract
In this paper, we investigate a specific area of document classification in which the documents come as a flow over the time. Moreover, the exact number of classes of document to deal with is not known from the beginning and could evolve over the time. To be able to perform classification task in such area, we need specific classifiers that are able to perform incremental learning and change their modeling over the time. More specifically, we are focusing our study on SVM approaches, known to perform well, and for which incremental (i-SVM) procedures exist. Nevertheless, most of them are only able to deal with a fixed number of classes. So we designed a new incremental learning procedure based on one-class SVMs. This one is able to improve its classification accuracy over the time, with the arrival of new labeled data, without performing any complete retraining. Moreover, when instances are coming with a previously unknown label (appearance of a new class), the training procedure is able to modify the classifier model to recognize this corresponding new kind of documents. To investigate this area, waiting for collecting documents images as a flow, we did first experiments on the Optical Recognition of Handwritten Digits Data Set. These experiments show that our incremental approach is able: to perform, at each time, as well as a static one-class classifier fully retrained using all previously seen data, to model very quickly and efficiently new incoming classes.
Anh Khoi Ngo Ho, Nicolas Ragot, Jean-Yves Ramel, Véronique Eglin, Nicolas Sidere
ICDAR4
2013 A Comprehensive Representation Model for Handwriting Dedicated to Word Spotting
abstract
In this paper, we propose an original representation model for handwriting document images. Most state-of-the-art handwriting representation models only use separately textural properties, selective dominant features (such as stroke orientation or gradient orientation) or structural properties. To avoid the drawbacks of using the properties from a single aspect, we design a comprehensive model that contains both morphological and topological information of handwriting. After interest points (the starting/ending points, branch points and high-curved points) are selected, an adapted version of Shape Context (SC) descriptor built on the interest points is employed to describe the contour of the text. In order to model the structural characteristics of the handwritten text, a graph is constructed based on the interest points and the skeleton of the text. With the graph, loops and specific strokes in the handwriting are detected and analyzed. Based on this model, a coarse-to-fine approach for word spotting application is introduced. Without segmenting texts into words, a group of regions of interest are selected by comparing textural features (orientation, projection profile, upper and lower border projection) using the DTW method. Afterwards, regions of interest and queries are represented by the proposed model. The final similarity measure is a weighted mixture of the SC cost, loop difference, stroke analysis and texture comparison with different weights. The validation of the model shows the significance of combining the various properties of the handwriting envisaged in its different aspects.
Peng Wang 0006, Véronique Eglin, Christophe Garcia, Christine Largeron, Antony McKenna
ICDAR2
2011 A Mixed Approach for Handwritten Documents Structural Analysis
abstract
In this paper we propose a new method for document pages segmentation. First dedicated to handwritten documents, our method is designed to extract the different text zones, paragraph and fragment in unconstrained documents. The proposed approach is a mixed one, using both the advantages of top-down and bottom-up approaches. In this paper we proposed and evaluation of our methods on a 183 documents database, taken from a 19th century handwritten corpus : the "dossiers de Bouvard et Pécuchet" from Flaubert. With this evaluation we demonstrate that the combination of the top-down and the bottom-up approach allow to improve the obtained results.
Vincent Malleron, Véronique Eglin
ICDAR2
2010 A new approach for centerline extraction in handwritten strokes: an application to the constitution of a code book
abstract
It is the pleasure of the organizing committee to welcome all participants to the 2010 IAPR Workshop on Document Analysis Systems (DAS). This year's workshop is being held June 9-11th in Boston, Massachusetts, a location situated in the heart of beautiful New England in the northeastern United States. Boston has a rich history that dates back to the 1600's and was a major focal point in the history of the American Revolution and America's fight for independence. Over the years, Boston has developed into a center for industrial, academic and cultural excellence and draws millions of visitors each year. DAS 2010 is the ninth workshop in a series. The first DAS was held in Kaiserslautern, Germany in 1994, and was followed by Malvern, PA (1996); Nagano, Japan (1998); Rio de Janeiro, Brazil (2000); Princeton, NJ (2002); Florence, Italy (2004); Nelson, New Zealand (2006) and Nara, Japan (2008). The DAS tradition is to bring together industry, academic and government researchers interested in many aspects of document analysis systems and to provide opportunities for fruitful interaction and collaboration. This year's workshop is organized as a three-day, single track event, with oral and poster presentations, as well as working group discussions on the second and third afternoons. Special sessions on contributed datasets and a keynote talk on this same topic provide a compelling theme, and we hope this focus will help push the field toward greater sharing of data and accepted standards for evaluation. This year, we received 91 submissions from 25 countries on six continents. The papers were reviewed by 45 members of our research community and an international program committee representing 16 different countries. The overall quality was excellent and we have chosen 28 full papers for oral presentation and 37 as poster papers, as well as 15 short papers that will be presented either as posters or in a special short oral format. In addition, six groups are scheduled to present live demos during the poster sessions. Full papers underwent the standard peer review process and will appear in the official workshop proceedings to be published in the ACM International Conference Proceedings Series, available online as part of the ACM Digital Library. The short papers are included in the unofficial hardcopy proceedings distributed at the event as well as on the DAS 2010 website.
Hani Daher, Véronique Eglin, Stéphane Bres, Nicole Vincent
Document Analysis Systems2
2010 Ancient Handwritings Decomposition Into Graphemes and Codebook Generation Based on Graph Coloring
abstract
We present in this paper a new method of analysis and decomposition of handwritten documents into glyphs (graphemes) and their associated code book. The different techniques that are involved in this paper are inspired by image processing methods in a large sense and mathematical models implying graph coloring. Our approaches provide firstly a rapid and detailed characterization of handwritten shapes based on dynamic tracking of the handwriting (curvature, thickness, direction, etc.) and also a very efficient analysis method for the categorization of basic shapes (graphemes). The tools that we have produced enable paleographers to study quickly and more accurately a large volume of manuscripts and to extract a large number of characteristics that are specific to individual writer or specific era.
Hani Daher, Djamel Gaceb, Véronique Eglin, Stéphane Bres, Nicole Vincent
ICFHR3
2009 Hierarchical Decomposition of Handwritten Manuscripts Layouts
Vincent Malleron, Véronique Eglin, Hubert Emptoz, Stéphanie Dord-Crouslé, Philippe Régnier
CAIP2
2009 Graph b-Coloring for Automatic Recognition of Documents
abstract
In order to reduce the rejection rate of our automatic reading system, we propose to pre-classify the business documents by introducing an automatic recognition of documents stage (ARD) as a pre-processing step. This important step will guide the other stages involved in the recognition process of the documents contents. Once the document class identified, the reading system will use correct information from the ARD stage to improve the segmentation of the layout, the recognition of the document structure, the parameterization of the OCR, and the final decision for the rejection. We propose in this paper an original method for the classification of business documents suited for complex layouts having great variability. We introduce the graph coloring approach for both layout analysis and document classification. The proposed method is reliable, robust to various constraints and guarantees a real-time answer to the sorting of business documents.
Djamel Gaceb, Véronique Eglin, Frank Lebourgeois, Hubert Emptoz
ICDAR2
2009 Text Lines and Snippets Extraction for 19th Century Handwriting Documents Layout Analysis
abstract
In this paper we propose a new approach to improve electronic editions of human science corpus, providing an efficient estimation of manuscripts pages structure. In any handwriting documents analysis process, the text line segmentation is an important stage. The presence of variable inter-line spaces, of inconstant base-line skews, overlapping and occlusions in unconstrained ancient 19th handwritten documents complexifies the text lines segmentation task. In this paper, we only use as prior knowledge of script the fact that text lines skews can be random and irregular.In that context, we model text line detection as an image segmentation problem by enhancing text line structure using Hough transform and a clustering of connected components so as to make text line boundaries appear. The proposed approach of snippets decomposition for page layout analysislies on a first step of content pages classification in five visual and genetic taxonomies, and a second step of text line extraction and snippets decomposition. Experiments show that the proposed method achieves high accuracy for detecting text lines in regular and semi-regular handwritten pages in the corpus of digitized Flaubert manuscripts (”Dossiers documentaires de Bouvard et Pécuchet”, 1872-1880).
Vincent Malleron, Véronique Eglin, Hubert Emptoz, Stéphanie Dord-Crouslé, Philippe Régnier
ICDAR2
2008 Physical Layout Segmentation of Mail Application Dedicated to Automatic Postal Sorting System
abstract
Every day, the postal sorting systems diffuse several tons of mails. It is noted that the principal origin of mail rejection is related to the failure of address-block localization task, particularly, of the physical layout segmentation stage. The bottom-up and top-down segmentation methods bring different knowledge that should not be ignored when we need to increase the robustness. Hybrid methods combine the two strategies in order to take advantages of one strategy to the detriment of other. Starting from these remarks, our proposal makes use of a hybrid segmentation strategy more adapted to the postal mails. The high level stages are based on the hierarchical graphs coloring, allowing managing through a pyramidal data organization, the complex rules leading the interpretation of the connected components decomposition of interest zones. Today, no other work in this context has make use of the powerfulness of this tool. The performance evaluation of our approach was tested on a corpus of 10000 envelope images. The processing times and the rejection rate were considerably reduced.
Djamel Gaceb, Véronique Eglin, Frank Lebourgeois, Hubert Emptoz
Document Analysis Systems2
2008 A Complete Pyramidal Geometrical Scheme for Text Based Image Description and Retrieval
Guillaume Joutel, Véronique Eglin, Hubert Emptoz
ICISP2
2008 Application of graph coloring in physical layout segmentation
abstract
Every-day, the postal sorting systems diffuse several tons of mails. It is noted that the principal origin of mail rejection is related to the failure of address-block localization task, particularly, of the physical layout segmentation stage. The bottom-up and top-down segmentation methods bring different knowledge that should not be ignored when we need to increase the robustness. Hybrid methods combine the two strategies in order to take advantages of one strategy to the detriment of other. Starting from these remarks, our proposal makes use of a hybrid segmentation strategy more adapted to the postal mails. The high level stages are based on the hierarchical graphs coloring. Today, no other work in this context has make use of the powerfulness of this tool. The performance evaluation of our approach was tested on a corpus of 10000 envelope images. The processing times and the rejection rate were considerably reduced.
Djamel Gaceb, Véronique Eglin, Frank Lebourgeois, Hubert Emptoz
ICPR2
2008 Generic scale-space process for handwriting documents analysis
abstract
This paper presents a generic architecture for handwriting documents analysis. It covers all analysis steps from the content description of the document (layout analysis, handwriting shape characterization) to three dedicated Digital Libraries applications (CBIR in great ancient documents images database, Paleo-graphical images classification and word spotting). The generic scale space tool is based on the Curvelets decomposition of images for the indexation of linear singularities of handwritten shapes. The proposed scheme for handwritten shape characterization targets to detect oriented and curved fragments at different scales: it is used in a first step to extract visual textual interest regions and secondly to use the Curvelets coefficients in various ways to satisfy the three designed applications. The complete implementation scheme is validated with a specific application of word spotting based on the orientations analysis. The proposed method is language independent and only visual orientation and appearance based. In that context, no lexical information nor any other statistical language models are required. The first proposed tests for this application are proposed on medieval documents images and on European 18th century correspondences corpus from the CERPHI. Precision-recall analysis testifies the relevance of the contribution.
Guillaume Joutel, Véronique Eglin, Hubert Emptoz
ICPR2
2008 Improvement of postal mail sorting system
Djamel Gaceb, Véronique Eglin, Frank Lebourgeois, Hubert Emptoz
Int. J. Document Anal. Recognit.2
2008 Document image characterization using a multiresolution analysis of the texture: application to old documents
Nicholas Journet, Jean-Yves Ramel, Rémy Mullot, Véronique Eglin
Int. J. Document Anal. Recognit.4
2007 Writer Identification Using Steered Hermite Features and SVM
abstract
Writer recognition is considered as a difficult problem to solve due to variations found in the writing, even from the same writer. In this paper, Steered Hermite Features are used to identify writer from a written document. We will show that Steered Hermite Features are highly useful for text images because they extract lot of information, no- tably for data characterized by oriented features, curves and segments. The algorithm we propose here, first calcu- lates the Steered Hermite Features of the images which are then passed on to Support Vector Machine for training and testing. The base of tests consists of sample of some lines of writings (five at most) of primarily diversified writings of authors from IAM database. With the proposed algorithm based on Steered Hermite Features, we were able to achieve an accuracy of around 83% percent for a set of 30 authors with non overlapping images of written text.
A. Imdad, Stéphane Bres, Véronique Eglin, C. Rivero-Moreno, Hubert Emptoz
ICDAR3
2007 A Proposition of Retrieval Tools for Historical Document Images Libraries
abstract
In this article, we propose a method of characterization of pictures of old documents based on a texture approach. This characterization is carried out with the help of a multi- resolution study of the textures contained in the pictures of the document. So, by extracting five features linked to the frequencies and to the orientations in the different parts of a page, it is possible to extract and to compare elements of high semantic level without expressing any hypothesis about the physical or logical structure of the analysed documents. Experiments show the feasibility of the fulfillment of tools for the navigation or the indexation help. In these experimentations, we will lay the emphasis upon the pertinence of these texture features and the advances that they represent in terms of characterization of content of a deeply heterogeneous corpus.
Nicholas Journet, Jean-Yves Ramel, Rémy Mullot, Véronique Eglin
ICDAR4
2007 Curvelets Based Queries for CBIR Application in Handwriting Collections
abstract
This paper presents a new use of the curvelet transform as a multiscale method for indexing linear singularities and curved handwritten shapes in documents images. As it belongs to the wavelet family, this representation can be useful at several scales of details. The proposed scheme for handwritten shape characterization targets to detect oriented and curved fragments at different scales so as to compose an unique signature for each handwritten analyzed samples. In this way, curvelets coefficients are used as a representation tool for handwriting when searching in large manuscripts databases by finding similar handwritten samples. Current results of ancient manuscripts retrieval are very promising with very satisfying precisions and recalls.
Guillaume Joutel, Véronique Eglin, Stéphane Bres, Hubert Emptoz
ICDAR2
2007 Hermite and Gabor transforms for noise reduction and handwriting classification in ancient manuscripts
Véronique Eglin, Stéphane Bres, Carlos Rivero
Int. J. Document Anal. Recognit.1
2005 Frequencies Decomposition and Partial Similarities Retrieval for Ancient Handwriting Documents Compression
abstract
This paper presents a new segmentation free approach of partial similarities retrieval in ancient handwritten documents. The method has been developed to improve usual handwritings compression approaches that are not adapted to patrimonial images specificities. We present here the similarities characterization that lies on oriented handwriting shapes decomposition. The frequencies page decomposition realizes a pavement of handwritten regions stored in directional maps where partial similarities are estimated. This decomposition is obtained by a frequencies analysis implying Gabor bank filters with an adequate parameter setting based on the most significant directions of the text. For each map, we compute a similarity graph that reveals redundant shapes and determines a resulting redundancy rate. The resulting graph is the first part of the compression system currently under development.
Abir El Abed, Véronique Eglin, Frank Lebourgeois, Hubert Emptoz
ICDAR2
2005 Biological inspired Tools for Patrimonial Handwriting Denoising and Categorization
abstract
In this paper, we propose a global segmentation free methodology for patrimonial documents denoising, handwriting characterization and categorization that are based on biological inspired approaches. We widely used here the spectral domain of handwritten images by frequency decompositions: Hermite transforms and Gabor bank filters. Handwritten pages are described by a multiscale signature that is based on orientation features and that is at the basis of a similarity measure. The current results of handwriting categorization and indexing are very promising and show that it is possible to analyze handwritten drawings without any a priori graphemes segmentation.
Véronique Eglin, Stéphane Bres, Carlos Rivero, Hubert Emptoz
ICDAR1
2005 Text/Graphic labelling of Ancient Printed Documents
abstract
This paper presents a text/graphic labelling for ancient printed documents. Our approach is based on the extraction and the quantification of the various orientations that are present in ancient printed document images. The documents are initially cut into normalized square windows in which we analyze significant orientations with a directional rose. Each kind of information (textual or graphical) is typically identified and marked by its orientation distribution. This choice of characterization allows us to separate textual regions from graphics by minimizing the a priori knowledge. The evaluation of our proposition lies on a page classification using layout extraction criteria. The system has been tested over several ancient printed books of the Renaissance.
Nicholas Journet, Véronique Eglin, Jean-Yves Ramel, Rémy Mullot
ICDAR2
2004 Multiscale Handwriting Characterization for Writers' Classification
Véronique Eglin, Stéphane Bres, Carlos Rivero
Document Analysis Systems1
2004 Analysis and interpretation of visual saliency for document functional labeling
Véronique Eglin, Stéphane Bres
Int. J. Document Anal. Recognit.1
2003 Document page similarity based on layout visual saliency: Application to query by example and document classification
abstract
In this paper we propose to define a measure of visualsimilarity to compare different pages in a corpus. Thismeasure is based on the analysis of the visual layoutsaliency of the page composition. This similarity iscomputed using both the document layout andcharacteristics of the text itself. The text characterizationuses statistical features derived from textural primitives.Our purpose is to establish perceptive links betweendocuments in order to facilitate their storage and theirretrieval. In this paper we present two possibleapplications of this measure of similarity: the query ofthe corpus by example and the documents classification.In the first application, we extract documents that are themost visually similar to a document, given as query. Inthe second application, the similarity measure is used toclassify the document under investigation using its visualsimilarity to a reference set of documents. Our test corpusis extracted from the Finland MTDB Oulu multi-genredatabase that provides a great diversity of page layoutsand contents.
Véronique Eglin, Stéphane Bres
ICDAR1
2001 Visual Exploration and Functional Document Labeling
abstract
This paper presents a new approach to textual data labeling based on texture analysis. Texture is used here to show the impact of document composition on visual exploration. We demonstrate how textural properties are well adapted to typography characterization by categorizing document regions into visual text classes (such as headings, head- and footnotes, paragraphs, abstracts, etc.). We reference and classify different types of text fonts according to their visual aspect and the visual impression that emerges from the textual data. Experiments on a set of various document images show a good accuracy and robustness for our method.
Véronique Eglin, Antoine Gagneux
ICDAR1
1998 Printed text featuring using the visual criteria of legibility and complexity
abstract
We present an approach to printed text characterization with a statistical analysis of texture. This method is based on a visibility and complexity criterion of type font families. The text is analyzed through its typographic form. So, we propose to label different kinds of text according to their visual aspect and their textural contents (especially their size, the line and letter spacing but also their complexity and their density). The texture is characterized with a statistical analysis based on samples randomly taken. This analysis is on the basis of text labeling. It is dedicated to the classification of texts according to their eye-catching properties. The global aim of this work is to recover the logical structure of a document according to the relative importance of font-types. This discrimination has been realized by the use of a scale of legibility, complexity, and of structural relief of forms.
Véronique Eglin, Stéphane Bres, Hubert Emptoz
ICPR1
1997 Logarithmic Spiral Grid and Gaze Control for the Development of Strategies of Visual Segmentation on a Document
abstract
The paper presents a page segmentation method which is based on perception phenomena and displays the unequal importance of information in the visual field. The access of information is directly linked to the search of attractive areas. This search is based on the idea of freeing oneself from an unbending physical structure and from a uniform vertical and horizontal scanning of the document, so as to classify the data in order of importance and interest. Using a space variant geometry for block selection, the page image, instead of being represented by a bitmap format, can be abstractly represented by the block format. This space variant geometry lays a sound basis for elaborating the kinetics of the ocular shifting on a document, which provides not only a meaningless document representation in blocks, but shows a unified view corresponding to the integration of time variant representations of the same visual field.
Véronique Eglin, Hubert Emptoz
ICDAR1