VLDB 2026 Research / reviewers in the wild / expert
Véronique Eglin
dblp:e/VEglin
· DBLP profile ↗
52ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0001-8738-2088ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 29 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CCASL: Counterexamples to comparative analysis of scientific literature - Application to polymers
Aymar Tchagoue, Véronique Eglin, Sébastien Pruvost, Jean-Marc Petit, Jannick Duchet-Rumeau, Jean-François Gérard |
Data Knowl. Eng. | 2 |
| 2025 | Unsupervised Energy-Based Model for the Identification of Out-Of-Distribution Copy Detection PatternsabstractThe recent advances in deep neural networks have led us to revisit the safety of current mechanisms for the authentification of people’s identities and validating the provenance of traded goods. This paper studies the problem of identifying counterfeits of Copy Detection Patterns (CDPs). CDPs are used on product packaging to identify counterfeiting by relying on the information loss principle. Current verification techniques rely on a supervised machine learning paradigm, which needs examples from real and fake CDPs to learn a model capable of differentiating both classes of CDPs. This paper proposes an unsupervised forensic approach that is capable of training an Energy-Based Model for detecting out-of-distribution (OOD) images (fake CDP examples in our case) using only original CDP images. This paper also introduces a novel thresholding technique that only requires original CDPs for threshold value selection. We validate our approach on the Indigo dataset, with results demonstrating comparable counterfeit detection capabilities to prior work but using a method trained on less data. Marc Chapus, Carlos Fernando Crispim, Véronique Eglin, Atilla Baskurt |
AVSS | 3 |
| 2025 | Rx-PAD: Recognition and eXtraction - A Dataset for Prescription Analysis and Clinical Data Structuring
Jonathan Pattin Cottet, Véronique Eglin, Alex Aussem |
ICDAR (5) | 2 |
| 2025 | A Multimodal Evaluation Pipeline for Mathematical Expression Recognition: Comparisons of Datasets, Metrics, and Models
François Wieckowiak, Véronique Eglin, Tony Bonnet, Stéphane Bres, Laëtitia Rousseau |
ICDAR (4) | 2 |
| 2024 | Study's system: a three-stage system for children handwriting recognition and spelling error detection
Sofiane Medjram, Véronique Eglin, Stéphane Bres |
Multim. Tools Appl. | 2 |
| 2023 | Ensuring an Error-Free Transcription on a Full Engineering Tags Dataset Through Unsupervised Post-OCR Methods
Mathieu Francois, Véronique Eglin |
ICDAR (5) | 2 |
| 2022 | Text Detection and Post-OCR Correction in Engineering Documents
Mathieu Francois, Véronique Eglin, Maxime Biou |
DAS | 2 |
| 2022 | Challenging Children Handwriting Recognition Study Exploiting Synthetic, Mixed and Real Data
Sofiane Medjram, Véronique Eglin, Stéphane Bres |
DAS | 2 |
| 2021 | One-Class Detection and Classification of Defects on Concrete SurfacesabstractToday, railway infrastructure is subject to regular inspections carried out by image acquisition. Experts need automatic tools to process these images and extract defects. These defects can be classified in two categories : linear defects, like cracks, and surface defects, like spall, humidity, graffiti… Cracks detection has been presented in a previous work, and is based on Local Binary Patterns (LBP) extraction and our FLASH algorithm [1]. In this paper, we propose a new method, using the same LBP descriptors, already extracted, to detect surface defects using generative adversarial networks (GAN). The detection is followed by a classification with a potentially scalable system. Experiments show that our detector achieves better performance than the state-of-the art in application to defect detection. Yannick Faula, Véronique Eglin, Stéphane Bres |
SMC | 2 |
| 2020 | Classification of Phonetic Characters by Space-Filling Curves
Valentin Owczarek, Jordan Drapeau, Jean-Christophe Burie, Patrick Franco, Mickaël Coustaty, Rémy Mullot, Véronique Eglin |
DAS | 7 |
| 2019 | Recurrent Neural Network Approach for Table Field Extraction in Business DocumentsabstractEfficiently extracting information from documents issued by their partners is crucial for companies that face huge daily document flows. Particularly, tables contain most valuable information of business documents. However, their contents are challenging to automatically parse as tables from industrial contexts may have complex and ambiguous physical structure. Bypassing their structure recognition, we propose a generic method for end-to-end table field extraction that starts with the sequence of document tokens segmented by an OCR engine and directly tags each token with one of the possible field types. Similar to the state-of-the-art methods for non-tabular field extraction, our approach resorts to a token level recurrent neural network combining spatial and textual features. We empirically assess the effectiveness of recurrent connections for our task by comparing our method with a baseline feedforward network having local context knowledge added to its inputs. We train and evaluate both approaches on a dataset of 28,570 purchase orders to retrieve the ID numbers and quantities of the ordered products. Our method outperforms the baseline with micro F1 score on unknown document layouts of 0.821 compared to 0.764. Clément Sage, Alex Aussem, Haytham Elghazel, Véronique Eglin, Jérémy Espinas |
ICDAR | 4 |
| 2019 | KeyWord Spotting using Siamese Triplet Deep Neural NetworksabstractDeep neural networks has shown great success in computer vision fields by achieving considerable state-of-the-art results and are beginning to arouse big interest in the document analysis community. In this paper, we present a novel siamese deep network of three inputs that allows retrieving the most similar words to a given query. The proposed system follows a query-by-example approach according to a segmentation-based technique and aims to learn suitable representations of handwritten word images, for which a simple Euclidean distance could perform the matching. The results obtained for the George Washington dataset show the potential and the effectiveness of the proposed keyword spotting system. Yasmine Serdouk, Véronique Eglin, Stéphane Bres, Mylène Pardoen |
ICDAR | 2 |
| 2018 | A Fast Local Analysis by Thresholding applied to image matchingabstractKey structures extraction and matching are key steps in computer vision. Many fields of application need large image acquisition and fast extraction of fine structures. In this study, we focus on situations where existing local feature extractors give not enough satisfying results concerning both accuracy and time processing. Among good illustrations, we can quote short-line extraction in local weakly-contrasted images. We propose a new Fast Local Analysis by threSHolding (FLASH) designed to process large images under hard time constraints. We use “micro-line” points as key feature. These are used for shape reconstruction (like lines) and local signature design. We apply FLASH on the field of concrete infrastructure monitoring where robots and UAVs are more and more used for automated defect detection (like cracks). For large concrete surfaces, there are several hard constraints such as the computational time and the reliability. Results show us that the computations are faster than several existing algorithms in image matching and FLASH has invariance to rotation, partial occlusion, and scale range from 0.7 to 1.4 without scale-space exploration. Yannick Faula, Stéphane Bres, Véronique Eglin |
ICPR | 3 |
| 2017 | ICDAR2017 Competition on the Classification of Medieval Handwritings in Latin ScriptabstractThis paper presents the results of the ICDAR2017 Competition on the Classification of Medieval Handwritings in Latin Script (CLaMM), jointly organized by Computer Scientists and Humanists (paleographers). This work follows a competition at ICFHR2016 and aims at providing a rich annotated database of European medieval manuscripts to the community on Handwriting Analysis and Recognition. We proposed four independent classification tasks which attracted 10 registered teams, with 6 submitted classifiers from 4 participants. Those classifiers are trained on a set of 3540 images with their ground truths. In task 1 (Script classification) and task 3 (Date classification), the classifiers have been evaluated by a test set of 2000 greyscale, tiff, 300 dpi images. In task 2 (Script classification) and task 4 (Date classification), the test set consists of 1000 images in different formats, resolutions and color representation. The best scores are respectively 85.2% for task 1, 76.5% for task 2, 59% for task 3, and 49.9% for task 4. An analysis based on the matrix of confusion of each classifier is also given. Florence Cloppet, Véronique Eglin, Marlene Helias-Baron, Van Cuong Kieu, Nicole Vincent, Dominique Stutzmann |
ICDAR | 2 |
| 2017 | Discovering Motifs with Variants in Music Databases
Riyadh Benammar, Christine Largeron, Véronique Eglin, Mylène Pardoen |
IDA | 3 |
| 2017 | A multi-one-class dynamic classifier for adaptive digitization of document streams
Anh Khoi Ngo Ho, Véronique Eglin, Nicolas Ragot, Jean-Yves Ramel |
Int. J. Document Anal. Recognit. | 2 |
| 2016 | SmartATID: A Mobile Captured Arabic Text Images Dataset for Multi-purpose Recognition TasksabstractToday's smartphones are able to capture documents with a good and simple way as any personal scanners. The captured document images need to be processed by specific and automated document processing systems. The systems are dedicated to textual content analysis, indexing and recognition. For instance, they may be used for font identification, writer identification and word or line segmentation. The state-of-the-art works lack comprehensive database for Arabic document images which are captured by mobile phones. This paper presents the first public offline images database for both printed and handwriting Arabic mobile captured documents, named "SmartATID". The document images of the database are acquired under varying capture conditions (blur, perspective angles and light). This causes photometric and geometric distortions that influence the performance of OCR process but also the page segmentation in lines and paragraphs. Each document image of our database is provided with a ground truth file that contains the exact text transcription and all numerical capture parameters used for each image capture. The database is freely and publicly usable by the research community at the following address http:// sites.google.com/site/smartatid. Fatma Chabchoub, Yousri Kessentini, Slim Kanoun, Véronique Eglin, Frank Lebourgeois |
ICFHR | 4 |
| 2016 | ICFHR2016 Competition on the Classification of Medieval Handwritings in Latin ScriptabstractThis paper presents the results of the ICFHR2016 Competition on the Classification of Medieval Handwritings in Latin Script (CLaMM), jointly organized by Computer Scientists and Humanists (paleographers). This work aims at providing a rich database of European medieval manuscripts to the community on Handwriting Analysis and Recognition. At this competition, we proposed two independent classification tasks which attracted five participants with seven submitted classifiers. Those classifiers are trained on a set of 2000 images with their ground truths. In the first task of script crisp classification, the classifiers have been evaluated on a test set of 1000 single-type manuscripts. In the second task of "Fuzzy Classification", the classifiers have been carried out on a set of 2000 multi-script-type manuscripts. The results of the participants provide the first baseline evaluation up to the accuracy score of 83.9% for the task 1 and to the fuzzy weighted score of 2.96/4 for the task 2. An analysis based on the intra-class distance and matrix of confusion of each classifier is also given. Florence Cloppet, Véronique Eglin, Van Cuong Kieu, Dominique Stutzmann, Nicole Vincent |
ICFHR | 2 |
| 2016 | libcrn, an Open-Source Document Image Processing LibraryabstractIn this paper we introduce libcrn, a multiplatform open-source document image processing library aimed at researchers and companies. It is written in C++11 and has a non-contaminating license that makes it available for use in any project without legal constraints. The features include low-level image processing (color format conversion, binarization, convolution, PDE…), document images specific tools (connected components extraction, recursive block description, PDF export…), maths (matrix arithmetics, linear algebra, GMMs, equation solvers…), classification and clustering (kNN, k-means, HMMs…). The API is comprehensively documented and libcrn's architecture follows modern C++ guidelines to facilitate the handling of the library and enforce its safe usage. A sample OCR, which is only 30 lines long, is described to illustrate libcrn's scope of possibilities. Yann Leydier, Jean Duong, Stéphane Bres, Véronique Eglin, Frank Lebourgeois, Martial Tola |
ICFHR | 4 |
| 2014 | A Novel Learning-Free Word Spotting Approach Based on Graph RepresentationabstractEffective information retrieval on handwritten document images has always been a challenging task. In this paper, we propose a novel handwritten word spotting approach based on graph representation. The presented model comprises both topological and morphological signatures of handwriting. Skeleton-based graphs with the Shape Context labelled vertexes are established for connected components. Each word image is represented as a sequence of graphs. In order to be robust to the handwriting variations, an exhaustive merging process based on DTW alignment result is introduced in the similarity measure between word images. With respect to the computation complexity, an approximate graph edit distance approach using bipartite matching is employed for graph matching. The experiments on the George Washington dataset and the marriage records from the Barcelona Cathedral dataset demonstrate that the proposed approach outperforms the state-of-the-art structural methods. Peng Wang 0006, Véronique Eglin, Christophe Garcia, Christine Largeron, Josep Lladós 0001, Alicia Fornés |
Document Analysis Systems | 2 |
| 2014 | Arabic Font Recognition Based on a Texture AnalysisabstractExisting works on the font recognition and based on texture analysis often used Gray Level Cooccurence Matrix (GLCM), Gabor Filters (GF) and wavelet. In this paper, we use Steer able Pyramid (SP) for texture analysis of Arabic homogeneous and normalized text block in order to font recognition. In this frameworks, we use K Nearest Neighbors (KNN) and Back-propagation Artificial Neural Network (BpANN) for classification. The Obtained experimental results on the APTID/MF database (Arabic Printed Text Image/ Multi-Font) are encouragents. Faten Kallel Jaiem, Slim Kanoun, Véronique Eglin |
ICFHR | 3 |
| 2014 | Learning-Free Text-Image Alignment for Medieval ManuscriptsabstractIn this paper, we describe a new approach for text-image alignment of middle-age documents. The method is dedicated to word-to-word alignment in a segmentation-free and learning-free way. The best word-to-word matching lies on an adapted editing cost between signatures extracted on Unicode characters and on the images. The results are evaluated on the "Queste del saint Graal" (13th c.) by palaeographers through an intuitive validation interface that also offers a very fast interactive correction process. The gain of time resulting from the absence of learning stage offers the opportunity to pay more attention to the integration of different specificities and variations in middle-age handwritten documents (presence of typical abbreviations, allographs). Yann Leydier, Véronique Eglin, Stéphane Bres, Dominique Stutzmann |
ICFHR | 2 |
| 2014 | Handwritten word spotting based on a hybrid optimal distanceabstractIn this paper, we develop a comprehensive representation model for handwriting, which contains both morphological and topological information. An adapted Shape Context descriptor built on structural points is employed to describe the contour of the text. Graphs are first constructed by using the structural points as nodes and the skeleton of the strokes as edges. Based on graphs, Topological Node Features (TNFs) of n-neighbourhood are extracted. Bag-of-Words representation model based on the TNFs is employed to depict the topological characteristics of word images. Moreover, a novel approach for word spotting application by using the proposed model is presented. The final distance is a weighted mixture of the SC cost, and the TNF distribution comparison. Linear Discriminant Analysis (LDA) is used to learn the optimal weight for each part of the distance with the consideration of writing styles. The evaluation of the proposed approach shows the significance of combining the properties of the handwriting from different aspects. Peng Wang 0006, Véronique Eglin, Christine Largeron, Christophe Garcia |
ICIP | 2 |
| 2014 | A Coarse-to-Fine Word Spotting Approach for Historical Handwritten Documents Based on Graph Embedding and Graph Edit DistanceabstractEffective information retrieval on handwritten document images has always been a challenging task, especially historical ones. In the paper, we propose a coarse-to-fine handwritten word spotting approach based on graph representation. The presented model comprises both the topological and morphological signatures of the handwriting. Skeleton-based graphs with the Shape Context labelled vertexes are established for connected components. Each word image is represented as a sequence of graphs. Aiming at developing a practical and efficient word spotting approach for large-scale historical handwritten documents, a fast and coarse comparison is first applied to prune the regions that are not similar to the query based on the graph embedding methodology. Afterwards, the query and regions of interest are compared by graph edit distance based on the Dynamic Time Warping alignment. The proposed approach is evaluated on a public dataset containing 50 pages of historical marriage license records. The results show that the proposed approach achieves a compromise between efficiency and accuracy. Peng Wang 0006, Véronique Eglin, Christophe Garcia, Christine Largeron, Josep Lladós 0001, Alicia Fornés |
ICPR | 2 |
| 2013 | Exploring Interest Points and Local Descriptors for Word Spotting Application on Historical Handwriting Images
Peng Wang 0006, Véronique Eglin, Christine Largeron, Antony McKenna, Christophe Garcia |
CAIP (2) | 2 |
| 2013 | Document Classification in a Non-stationary Environment: A One-Class SVM ApproachabstractIn this paper, we investigate a specific area of document classification in which the documents come as a flow over the time. Moreover, the exact number of classes of document to deal with is not known from the beginning and could evolve over the time. To be able to perform classification task in such area, we need specific classifiers that are able to perform incremental learning and change their modeling over the time. More specifically, we are focusing our study on SVM approaches, known to perform well, and for which incremental (i-SVM) procedures exist. Nevertheless, most of them are only able to deal with a fixed number of classes. So we designed a new incremental learning procedure based on one-class SVMs. This one is able to improve its classification accuracy over the time, with the arrival of new labeled data, without performing any complete retraining. Moreover, when instances are coming with a previously unknown label (appearance of a new class), the training procedure is able to modify the classifier model to recognize this corresponding new kind of documents. To investigate this area, waiting for collecting documents images as a flow, we did first experiments on the Optical Recognition of Handwritten Digits Data Set. These experiments show that our incremental approach is able: to perform, at each time, as well as a static one-class classifier fully retrained using all previously seen data, to model very quickly and efficiently new incoming classes. Anh Khoi Ngo Ho, Nicolas Ragot, Jean-Yves Ramel, Véronique Eglin, Nicolas Sidere |
ICDAR | 4 |
| 2013 | A Comprehensive Representation Model for Handwriting Dedicated to Word SpottingabstractIn this paper, we propose an original representation model for handwriting document images. Most state-of-the-art handwriting representation models only use separately textural properties, selective dominant features (such as stroke orientation or gradient orientation) or structural properties. To avoid the drawbacks of using the properties from a single aspect, we design a comprehensive model that contains both morphological and topological information of handwriting. After interest points (the starting/ending points, branch points and high-curved points) are selected, an adapted version of Shape Context (SC) descriptor built on the interest points is employed to describe the contour of the text. In order to model the structural characteristics of the handwritten text, a graph is constructed based on the interest points and the skeleton of the text. With the graph, loops and specific strokes in the handwriting are detected and analyzed. Based on this model, a coarse-to-fine approach for word spotting application is introduced. Without segmenting texts into words, a group of regions of interest are selected by comparing textural features (orientation, projection profile, upper and lower border projection) using the DTW method. Afterwards, regions of interest and queries are represented by the proposed model. The final similarity measure is a weighted mixture of the SC cost, loop difference, stroke analysis and texture comparison with different weights. The validation of the model shows the significance of combining the various properties of the handwriting envisaged in its different aspects. Peng Wang 0006, Véronique Eglin, Christophe Garcia, Christine Largeron, Antony McKenna |
ICDAR | 2 |
| 2011 | A Mixed Approach for Handwritten Documents Structural AnalysisabstractIn this paper we propose a new method for document pages segmentation. First dedicated to handwritten documents, our method is designed to extract the different text zones, paragraph and fragment in unconstrained documents. The proposed approach is a mixed one, using both the advantages of top-down and bottom-up approaches. In this paper we proposed and evaluation of our methods on a 183 documents database, taken from a 19th century handwritten corpus : the "dossiers de Bouvard et Pécuchet" from Flaubert. With this evaluation we demonstrate that the combination of the top-down and the bottom-up approach allow to improve the obtained results. Vincent Malleron, Véronique Eglin |
ICDAR | 2 |
| 2010 | A new approach for centerline extraction in handwritten strokes: an application to the constitution of a code bookabstractIt is the pleasure of the organizing committee to welcome all participants to the 2010 IAPR Workshop on Document Analysis Systems (DAS). This year's workshop is being held June 9-11th in Boston, Massachusetts, a location situated in the heart of beautiful New England in the northeastern United States. Boston has a rich history that dates back to the 1600's and was a major focal point in the history of the American Revolution and America's fight for independence. Over the years, Boston has developed into a center for industrial, academic and cultural excellence and draws millions of visitors each year. DAS 2010 is the ninth workshop in a series. The first DAS was held in Kaiserslautern, Germany in 1994, and was followed by Malvern, PA (1996); Nagano, Japan (1998); Rio de Janeiro, Brazil (2000); Princeton, NJ (2002); Florence, Italy (2004); Nelson, New Zealand (2006) and Nara, Japan (2008). The DAS tradition is to bring together industry, academic and government researchers interested in many aspects of document analysis systems and to provide opportunities for fruitful interaction and collaboration. This year's workshop is organized as a three-day, single track event, with oral and poster presentations, as well as working group discussions on the second and third afternoons. Special sessions on contributed datasets and a keynote talk on this same topic provide a compelling theme, and we hope this focus will help push the field toward greater sharing of data and accepted standards for evaluation. This year, we received 91 submissions from 25 countries on six continents. The papers were reviewed by 45 members of our research community and an international program committee representing 16 different countries. The overall quality was excellent and we have chosen 28 full papers for oral presentation and 37 as poster papers, as well as 15 short papers that will be presented either as posters or in a special short oral format. In addition, six groups are scheduled to present live demos during the poster sessions. Full papers underwent the standard peer review process and will appear in the official workshop proceedings to be published in the ACM International Conference Proceedings Series, available online as part of the ACM Digital Library. The short papers are included in the unofficial hardcopy proceedings distributed at the event as well as on the DAS 2010 website. Hani Daher, Véronique Eglin, Stéphane Bres, Nicole Vincent |
Document Analysis Systems | 2 |
| 2010 | Ancient Handwritings Decomposition Into Graphemes and Codebook Generation Based on Graph ColoringabstractWe present in this paper a new method of analysis and decomposition of handwritten documents into glyphs (graphemes) and their associated code book. The different techniques that are involved in this paper are inspired by image processing methods in a large sense and mathematical models implying graph coloring. Our approaches provide firstly a rapid and detailed characterization of handwritten shapes based on dynamic tracking of the handwriting (curvature, thickness, direction, etc.) and also a very efficient analysis method for the categorization of basic shapes (graphemes). The tools that we have produced enable paleographers to study quickly and more accurately a large volume of manuscripts and to extract a large number of characteristics that are specific to individual writer or specific era. Hani Daher, Djamel Gaceb, Véronique Eglin, Stéphane Bres, Nicole Vincent |
ICFHR | 3 |
| 2009 | Hierarchical Decomposition of Handwritten Manuscripts Layouts
Vincent Malleron, Véronique Eglin, Hubert Emptoz, Stéphanie Dord-Crouslé, Philippe Régnier |
CAIP | 2 |
| 2009 | Graph b-Coloring for Automatic Recognition of DocumentsabstractIn order to reduce the rejection rate of our automatic reading system, we propose to pre-classify the business documents by introducing an automatic recognition of documents stage (ARD) as a pre-processing step. This important step will guide the other stages involved in the recognition process of the documents contents. Once the document class identified, the reading system will use correct information from the ARD stage to improve the segmentation of the layout, the recognition of the document structure, the parameterization of the OCR, and the final decision for the rejection. We propose in this paper an original method for the classification of business documents suited for complex layouts having great variability. We introduce the graph coloring approach for both layout analysis and document classification. The proposed method is reliable, robust to various constraints and guarantees a real-time answer to the sorting of business documents. Djamel Gaceb, Véronique Eglin, Frank Lebourgeois, Hubert Emptoz |
ICDAR | 2 |
| 2009 | Text Lines and Snippets Extraction for 19th Century Handwriting Documents Layout AnalysisabstractIn this paper we propose a new approach to improve electronic editions of human science corpus, providing an efficient estimation of manuscripts pages structure. In any handwriting documents analysis process, the text line segmentation is an important stage. The presence of variable inter-line spaces, of inconstant base-line skews, overlapping and occlusions in unconstrained ancient 19th handwritten documents complexifies the text lines segmentation task. In this paper, we only use as prior knowledge of script the fact that text lines skews can be random and irregular.In that context, we model text line detection as an image segmentation problem by enhancing text line structure using Hough transform and a clustering of connected components so as to make text line boundaries appear. The proposed approach of snippets decomposition for page layout analysislies on a first step of content pages classification in five visual and genetic taxonomies, and a second step of text line extraction and snippets decomposition. Experiments show that the proposed method achieves high accuracy for detecting text lines in regular and semi-regular handwritten pages in the corpus of digitized Flaubert manuscripts (”Dossiers documentaires de Bouvard et Pécuchet”, 1872-1880). Vincent Malleron, Véronique Eglin, Hubert Emptoz, Stéphanie Dord-Crouslé, Philippe Régnier |
ICDAR | 2 |
| 2008 | Physical Layout Segmentation of Mail Application Dedicated to Automatic Postal Sorting SystemabstractEvery day, the postal sorting systems diffuse several tons of mails. It is noted that the principal origin of mail rejection is related to the failure of address-block localization task, particularly, of the physical layout segmentation stage. The bottom-up and top-down segmentation methods bring different knowledge that should not be ignored when we need to increase the robustness. Hybrid methods combine the two strategies in order to take advantages of one strategy to the detriment of other. Starting from these remarks, our proposal makes use of a hybrid segmentation strategy more adapted to the postal mails. The high level stages are based on the hierarchical graphs coloring, allowing managing through a pyramidal data organization, the complex rules leading the interpretation of the connected components decomposition of interest zones. Today, no other work in this context has make use of the powerfulness of this tool. The performance evaluation of our approach was tested on a corpus of 10000 envelope images. The processing times and the rejection rate were considerably reduced. Djamel Gaceb, Véronique Eglin, Frank Lebourgeois, Hubert Emptoz |
Document Analysis Systems | 2 |
| 2008 | A Complete Pyramidal Geometrical Scheme for Text Based Image Description and Retrieval
Guillaume Joutel, Véronique Eglin, Hubert Emptoz |
ICISP | 2 |
| 2008 | Application of graph coloring in physical layout segmentationabstractEvery-day, the postal sorting systems diffuse several tons of mails. It is noted that the principal origin of mail rejection is related to the failure of address-block localization task, particularly, of the physical layout segmentation stage. The bottom-up and top-down segmentation methods bring different knowledge that should not be ignored when we need to increase the robustness. Hybrid methods combine the two strategies in order to take advantages of one strategy to the detriment of other. Starting from these remarks, our proposal makes use of a hybrid segmentation strategy more adapted to the postal mails. The high level stages are based on the hierarchical graphs coloring. Today, no other work in this context has make use of the powerfulness of this tool. The performance evaluation of our approach was tested on a corpus of 10000 envelope images. The processing times and the rejection rate were considerably reduced. Djamel Gaceb, Véronique Eglin, Frank Lebourgeois, Hubert Emptoz |
ICPR | 2 |
| 2008 | Generic scale-space process for handwriting documents analysisabstractThis paper presents a generic architecture for handwriting documents analysis. It covers all analysis steps from the content description of the document (layout analysis, handwriting shape characterization) to three dedicated Digital Libraries applications (CBIR in great ancient documents images database, Paleo-graphical images classification and word spotting). The generic scale space tool is based on the Curvelets decomposition of images for the indexation of linear singularities of handwritten shapes. The proposed scheme for handwritten shape characterization targets to detect oriented and curved fragments at different scales: it is used in a first step to extract visual textual interest regions and secondly to use the Curvelets coefficients in various ways to satisfy the three designed applications. The complete implementation scheme is validated with a specific application of word spotting based on the orientations analysis. The proposed method is language independent and only visual orientation and appearance based. In that context, no lexical information nor any other statistical language models are required. The first proposed tests for this application are proposed on medieval documents images and on European 18th century correspondences corpus from the CERPHI. Precision-recall analysis testifies the relevance of the contribution. Guillaume Joutel, Véronique Eglin, Hubert Emptoz |
ICPR | 2 |
| 2008 | Improvement of postal mail sorting system
Djamel Gaceb, Véronique Eglin, Frank Lebourgeois, Hubert Emptoz |
Int. J. Document Anal. Recognit. | 2 |
| 2008 | Document image characterization using a multiresolution analysis of the texture: application to old documents
Nicholas Journet, Jean-Yves Ramel, Rémy Mullot, Véronique Eglin |
Int. J. Document Anal. Recognit. | 4 |
| 2007 | Writer Identification Using Steered Hermite Features and SVMabstractWriter recognition is considered as a difficult problem to solve due to variations found in the writing, even from the same writer. In this paper, Steered Hermite Features are used to identify writer from a written document. We will show that Steered Hermite Features are highly useful for text images because they extract lot of information, no- tably for data characterized by oriented features, curves and segments. The algorithm we propose here, first calcu- lates the Steered Hermite Features of the images which are then passed on to Support Vector Machine for training and testing. The base of tests consists of sample of some lines of writings (five at most) of primarily diversified writings of authors from IAM database. With the proposed algorithm based on Steered Hermite Features, we were able to achieve an accuracy of around 83% percent for a set of 30 authors with non overlapping images of written text. A. Imdad, Stéphane Bres, Véronique Eglin, C. Rivero-Moreno, Hubert Emptoz |
ICDAR | 3 |
| 2007 | A Proposition of Retrieval Tools for Historical Document Images LibrariesabstractIn this article, we propose a method of characterization of pictures of old documents based on a texture approach. This characterization is carried out with the help of a multi- resolution study of the textures contained in the pictures of the document. So, by extracting five features linked to the frequencies and to the orientations in the different parts of a page, it is possible to extract and to compare elements of high semantic level without expressing any hypothesis about the physical or logical structure of the analysed documents. Experiments show the feasibility of the fulfillment of tools for the navigation or the indexation help. In these experimentations, we will lay the emphasis upon the pertinence of these texture features and the advances that they represent in terms of characterization of content of a deeply heterogeneous corpus. Nicholas Journet, Jean-Yves Ramel, Rémy Mullot, Véronique Eglin |
ICDAR | 4 |
| 2007 | Curvelets Based Queries for CBIR Application in Handwriting CollectionsabstractThis paper presents a new use of the curvelet transform as a multiscale method for indexing linear singularities and curved handwritten shapes in documents images. As it belongs to the wavelet family, this representation can be useful at several scales of details. The proposed scheme for handwritten shape characterization targets to detect oriented and curved fragments at different scales so as to compose an unique signature for each handwritten analyzed samples. In this way, curvelets coefficients are used as a representation tool for handwriting when searching in large manuscripts databases by finding similar handwritten samples. Current results of ancient manuscripts retrieval are very promising with very satisfying precisions and recalls. Guillaume Joutel, Véronique Eglin, Stéphane Bres, Hubert Emptoz |
ICDAR | 2 |
| 2007 | Hermite and Gabor transforms for noise reduction and handwriting classification in ancient manuscripts
Véronique Eglin, Stéphane Bres, Carlos Rivero |
Int. J. Document Anal. Recognit. | 1 |
| 2005 | Frequencies Decomposition and Partial Similarities Retrieval for Ancient Handwriting Documents CompressionabstractThis paper presents a new segmentation free approach of partial similarities retrieval in ancient handwritten documents. The method has been developed to improve usual handwritings compression approaches that are not adapted to patrimonial images specificities. We present here the similarities characterization that lies on oriented handwriting shapes decomposition. The frequencies page decomposition realizes a pavement of handwritten regions stored in directional maps where partial similarities are estimated. This decomposition is obtained by a frequencies analysis implying Gabor bank filters with an adequate parameter setting based on the most significant directions of the text. For each map, we compute a similarity graph that reveals redundant shapes and determines a resulting redundancy rate. The resulting graph is the first part of the compression system currently under development. Abir El Abed, Véronique Eglin, Frank Lebourgeois, Hubert Emptoz |
ICDAR | 2 |
| 2005 | Biological inspired Tools for Patrimonial Handwriting Denoising and CategorizationabstractIn this paper, we propose a global segmentation free methodology for patrimonial documents denoising, handwriting characterization and categorization that are based on biological inspired approaches. We widely used here the spectral domain of handwritten images by frequency decompositions: Hermite transforms and Gabor bank filters. Handwritten pages are described by a multiscale signature that is based on orientation features and that is at the basis of a similarity measure. The current results of handwriting categorization and indexing are very promising and show that it is possible to analyze handwritten drawings without any a priori graphemes segmentation. Véronique Eglin, Stéphane Bres, Carlos Rivero, Hubert Emptoz |
ICDAR | 1 |
| 2005 | Text/Graphic labelling of Ancient Printed DocumentsabstractThis paper presents a text/graphic labelling for ancient printed documents. Our approach is based on the extraction and the quantification of the various orientations that are present in ancient printed document images. The documents are initially cut into normalized square windows in which we analyze significant orientations with a directional rose. Each kind of information (textual or graphical) is typically identified and marked by its orientation distribution. This choice of characterization allows us to separate textual regions from graphics by minimizing the a priori knowledge. The evaluation of our proposition lies on a page classification using layout extraction criteria. The system has been tested over several ancient printed books of the Renaissance. Nicholas Journet, Véronique Eglin, Jean-Yves Ramel, Rémy Mullot |
ICDAR | 2 |
| 2004 | Multiscale Handwriting Characterization for Writers' Classification
Véronique Eglin, Stéphane Bres, Carlos Rivero |
Document Analysis Systems | 1 |
| 2004 | Analysis and interpretation of visual saliency for document functional labeling
Véronique Eglin, Stéphane Bres |
Int. J. Document Anal. Recognit. | 1 |
| 2003 | Document page similarity based on layout visual saliency: Application to query by example and document classificationabstractIn this paper we propose to define a measure of visualsimilarity to compare different pages in a corpus. Thismeasure is based on the analysis of the visual layoutsaliency of the page composition. This similarity iscomputed using both the document layout andcharacteristics of the text itself. The text characterizationuses statistical features derived from textural primitives.Our purpose is to establish perceptive links betweendocuments in order to facilitate their storage and theirretrieval. In this paper we present two possibleapplications of this measure of similarity: the query ofthe corpus by example and the documents classification.In the first application, we extract documents that are themost visually similar to a document, given as query. Inthe second application, the similarity measure is used toclassify the document under investigation using its visualsimilarity to a reference set of documents. Our test corpusis extracted from the Finland MTDB Oulu multi-genredatabase that provides a great diversity of page layoutsand contents. Véronique Eglin, Stéphane Bres |
ICDAR | 1 |
| 2001 | Visual Exploration and Functional Document LabelingabstractThis paper presents a new approach to textual data labeling based on texture analysis. Texture is used here to show the impact of document composition on visual exploration. We demonstrate how textural properties are well adapted to typography characterization by categorizing document regions into visual text classes (such as headings, head- and footnotes, paragraphs, abstracts, etc.). We reference and classify different types of text fonts according to their visual aspect and the visual impression that emerges from the textual data. Experiments on a set of various document images show a good accuracy and robustness for our method. Véronique Eglin, Antoine Gagneux |
ICDAR | 1 |
| 1998 | Printed text featuring using the visual criteria of legibility and complexityabstractWe present an approach to printed text characterization with a statistical analysis of texture. This method is based on a visibility and complexity criterion of type font families. The text is analyzed through its typographic form. So, we propose to label different kinds of text according to their visual aspect and their textural contents (especially their size, the line and letter spacing but also their complexity and their density). The texture is characterized with a statistical analysis based on samples randomly taken. This analysis is on the basis of text labeling. It is dedicated to the classification of texts according to their eye-catching properties. The global aim of this work is to recover the logical structure of a document according to the relative importance of font-types. This discrimination has been realized by the use of a scale of legibility, complexity, and of structural relief of forms. Véronique Eglin, Stéphane Bres, Hubert Emptoz |
ICPR | 1 |
| 1997 | Logarithmic Spiral Grid and Gaze Control for the Development of Strategies of Visual Segmentation on a DocumentabstractThe paper presents a page segmentation method which is based on perception phenomena and displays the unequal importance of information in the visual field. The access of information is directly linked to the search of attractive areas. This search is based on the idea of freeing oneself from an unbending physical structure and from a uniform vertical and horizontal scanning of the document, so as to classify the data in order of importance and interest. Using a space variant geometry for block selection, the page image, instead of being represented by a bitmap format, can be abstractly represented by the block format. This space variant geometry lays a sound basis for elaborating the kinetics of the ocular shifting on a document, which provides not only a meaningless document representation in blocks, but shows a unified view corresponding to the integration of time variant representations of the same visual field. Véronique Eglin, Hubert Emptoz |
ICDAR | 1 |