Jean-Marc Ogier

dblp:99/720 · DBLP profile ↗
← Back
77ranked-venue papers in the field
2as first author
7since 2021 · last 2025
0000-0002-5666-475XORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 73 (2 first)Information Retrieval & Web Search · 4
YearPublicationVenuePosition
2025 WildKhmerST: A Comprehensive Dataset and Benchmark for Khmer Scene Text Detection and Recognition in the Wild
abstract
This study presents a large-scale dataset of Khmer scene text images captured in real-world environments. Khmer, the official language of Cambodia, is spoken by approximately 17 million people. While Optical Character Recognition (OCR) systems have achieved remarkable success in Roman (Latin) script languages such as English, Khmer script poses unique challenges due to its intricate structure, absence of clear word boundaries, and highly diverse character shapes and sizes. A significant limitation in Khmer OCR research has been the scarcity of high-quality training data, particularly for deep learning-based models, which require extensive datasets to achieve robust performance. To address these challenges, we introduce a newly constructed dataset of Khmer scene text, comprising 29,601 annotated text lines from 10,000 unique images. This dataset is highly diverse and challenging, encompassing artistic text, blurred text, low-light conditions, curved text, text in complex backgrounds, and occluded text. Each text line is annotated with polygonal bounding box coordinates and line-level transcriptions, alongside attributes describing background complexity, character appearance, and text style. To establish a foundational benchmark for future research in Khmer OCR, we provide baseline results for Khmer text detection and recognition. Additionally, we propose a robust evaluation metric tailored for Khmer OCR, enabling precise assessment of CER and WER while accounting for the unique characteristics of the Khmer script.
Vannkinh Nom, Saly Keo, Souhail Bakkali, Muhammad Muzzamil Luqman, Mickaël Coustaty, Marçal Rossinyol, Jean-Marc Ogier
ICDAR (5)7
2025 QUEST: Quality-Aware Semi-supervised Table Extraction for Business Documents
Eliott Thomas, Mickaël Coustaty, Aurélie Joseph, Gaspar Deloin, Elodie Carel, Vincent Poulain D'Andecy, Jean-Marc Ogier
ICDAR (5)7
2024 Confidence-Aware Document OCR Error Detection
Arthur Hemmer, Mickaël Coustaty, Nicola Bartolo, Jean-Marc Ogier
DAS4
2024 CHIC: Corporate Document for Visual Question Answering
Ibrahim Souleiman Mahamoud, Mickaël Coustaty, Aurélie Joseph, Vincent Poulain D'Andecy, Jean-Marc Ogier
ICDAR (6)5
2023 Detecting Forged Receipts with Domain-Specific Ontology-Based Entities & Relations
Beatriz Martínez Tornés, Emanuela Boros, Antoine Doucet, Petra Gomez-Krämer, Jean-Marc Ogier
ICDAR (3)5
2022 QAlayout: Question Answering Layout Based on Multimodal Attention for Visual Question Answering on Corporate Document
Ibrahim Souleiman Mahamoud, Mickaël Coustaty, Aurélie Joseph, Vincent Poulain D'Andecy, Jean-Marc Ogier
DAS5
2021 Multimodal Attention-Based Learning for Imbalanced Corporate Documents Classification
Ibrahim Souleiman Mahamoud, Joris Voerman, Mickaël Coustaty, Aurélie Joseph, Vincent Poulain D'Andecy, Jean-Marc Ogier
ICDAR (3)6
2020 Background Removal of French University Diplomas
Tanmoy Mondal, Mickaël Coustaty, Petra Gomez-Krämer, Jean-Marc Ogier
DAS4
2020 Evaluation of Neural Network Classification Systems on Document Stream
Joris Voerman, Aurélie Joseph, Mickaël Coustaty, Vincent Poulain D'Andecy, Jean-Marc Ogier
DAS5
2019 Discourse Descriptor for Document Incremental Classification Comparison with Deep Learning
abstract
We propose both a new strategy to weight a text vector for document classification and a comparison of a deep learning approach versus an incremental classification approach, integrating our novel strategy. Bag-of-word vectors are classic approaches to describe a textual document in document classification objective. A weakness of the bag of words is to lose the organization of the discourse within the document. Inspired by some Deep Learning approaches and Natural Language Processing for text classification, we suggest a simple strategy, featuring the terms according to their relative positions within the discourse sequence. For experimentations, we apply this strategy to a recent document incremental classification approach from the state-of-the-art. And, we propose an original comparison between Incremental learning and Deep learning, by comparing the incremental system with a CCN-RNN-based approach. It demonstrates that both approaches are competitive in similar contest.
Vincent Poulain D'Andecy, Aurélie Joseph, Joaquín Cuenca, Jean-Marc Ogier
ICDAR4
2019 Instance Aware Document Image Segmentation using Label Pyramid Networks and Deep Watershed Transformation
abstract
Segmentation of complex document images remains a challenge due to the large variability of layout and image degradation. In this paper, we propose a method to segment complex document images based on Label Pyramid Network (LPN) and Deep Watershed Transform (DWT). The method can segment document images into instance aware regions including text lines, text regions, figures, tables, etc. The backbone of LPN can be any type of Fully Convolutional Networks (FCN), and in training, label map pyramids on training images are provided to exploit the hierarchical boundary information of regions efficiently through multi-task learning. The label map pyramid is transformed from region class label map by distance transformation and multi-level thresholding. In segmentation, the outputs of multiple tasks of LPN are summed into one single probability map, on which watershed transformation is carried out to segment the document image into instance aware regions. In experiments on four public databases, our method is demonstrated effective and superior, yielding state of the art performance for text line segmentation, baseline detection and region segmentation.
Xiao-Hui Li 0012, Jean-Marc Ogier, Cheng-Lin Liu 0001
ICDAR5
2019 A Robust Data Hiding Scheme Using Generated Content for Securing Genuine Documents
abstract
Data hiding is an effective technique, compared to pervasive black-and-white code patterns such as barcode and quick response code, which can be used to secure document images against forgery or unauthorized intervention. In this work, we propose a robust digital watermarking scheme for securing genuine documents by leveraging generative adversarial networks (GAN). To begin with, the input document is adjusted to its right form by geometric correction. Next, the generated document is obtained from the input document by using the mentioned networks, and it is regarded as a reference for data hiding and detection. We then introduce an algorithm that hides a secret information into the document and produces a watermarked document whose content is minimally distorted in terms of normal observation. Furthermore, we also present a method that detects the hidden data from the watermarked document by measuring the distance of pixel values between the generated and watermarked document. For improving the security feature, we encode the secret information prior to hiding it by using pseudo random numbers. Lastly, we demonstrate that our approach gives high precision of data detection, and competitive performance compared to state-of-the-art approaches.
Cu Vinh Loc, Jean-Christophe Burie, Jean-Marc Ogier, Cheng-Lin Liu 0001
ICDAR3
2019 Hiding Security Feature Into Text Content for Securing Documents Using Generated Font
abstract
Motivated by increasing possibility of the tampering of genuine documents during a transmission over digital channels, we focus on developing a watermarking framework for determining whether a given document is genuine or falsified. The proposed framework is performed by hiding a security feature or secret information within the document. In order to hide the security feature, we replace the appropriate characters of legal document by the equivalent characters coming from generated fonts, called hereafter the variations of characters. These variations are produced by training generative adversarial networks (GAN) with the features of character's skeleton and normal shape. Regarding the process of detecting hidden information, we make use of fully convolutional networks (FCN) to produce salient regions from the watermarked document. The salient regions mark positions of document where the characters are substituted by their variations, and these positions are used as a reference for extracting the hidden information. Lastly, we demonstrate that our approach gives high precision of data detection, and competitive performance compared to state-of-the-art approaches.
Cu Vinh Loc, Jean-Christophe Burie, Jean-Marc Ogier, Cheng-Lin Liu 0001
ICDAR3
2019 Learning Free Document Image Binarization Based on Fast Fuzzy C-Means Clustering
abstract
In this paper, a novel local threshold binarization method using fast Fuzzy C-Means clustering is proposed. Historical document images with non-uniform background, stains, faded ink are first processed by removing the background using inpainting based method. Then using Fuzzy C-Means clustering is used to cluster out the pixels into three main clusters : sure text pixels, sure background pixels and confused pixels which may or may not be labeled as text. Based on the structural symmetry of pixels (SSP), these confused pixels are then classified into text or background pixels. The SSP is defined as those pixels around strokes whose gradient magnitudes are big enough and whose directions are opposite. As the gradient map is our basis for computing the SSP, we further propose to estimate the background surface first and to extract potential SSP in the compensated image so as to deal with degradations of document images such as uneven illumination, low contrast and stain. To prove the effectiveness of our method, tests on eight public document image datasets are preformed and the experimental results show that our method outperforms other local threshold binarization approaches on both F-measure and PSNR.
Tanmoy Mondal, Mickaël Coustaty, Petra Gomez-Krämer, Jean-Marc Ogier
ICDAR4
2019 ICDAR2019 Robust Reading Challenge on Multi-lingual Scene Text Detection and Recognition - RRC-MLT-2019
abstract
With the growing cosmopolitan culture of modern cities, the need of robust Multi-Lingual scene Text (MLT) detection and recognition systems has never been more immense. With the goal to systematically benchmark and push the state-of-the-art forward, the proposed competition builds on top of the RRC-MLT-2017 with an additional end-to-end task, an additional language in the real images dataset, a large scale multi-lingual synthetic dataset to assist the training, and a baseline End-to-End recognition method. The real dataset consists of 20,000 images containing text from 10 languages. The challenge has 4 tasks covering various aspects of multi-lingual scene text: (a) text detection, (b) cropped word script classification, (c) joint text detection and script classification and (d) end-to-end detection and recognition. In total, the competition received 60 submissions from the research and industrial communities. This paper presents the dataset, the tasks and the findings of the presented RRC-MLT-2019 challenge.
Nibal Nayef, Cheng-Lin Liu 0001, Jean-Marc Ogier, Michal Busta, Pinaki Nath Chowdhury, Dimosthenis Karatzas, Wafa Khlif, Jiri Matas, Umapada Pal 0001, Jean-Christophe Burie
ICDAR3
2019 Fast Text/non-Text Image Classification with Knowledge Distillation
abstract
How to efficiently judge whether a natural image contains texts or not is an important problem. Since text detection and recognition algorithms are usually time-consuming, and it is unnecessary to run them on images that do not contain any texts. In this paper, we investigate this problem from two perspectives: the speed and the accuracy. First, to achieve high speed for efficient filtering large number of images especially on CPU, we propose using small and shallow convolutional neural network, where the features from different layers are adaptively pooled into certain sizes to overcome difficulties caused by multiple scales and various locations. Although this can achieve high speed but its accuracy is not satisfactory due to limited capacity of small network. Therefore, our second contribution is using the knowledge distillation to improve the accuracy of the small network, by constructing a larger and deeper neural network as teacher network to instruct the learning process of the small network. With the above two strategies, we can achieve both high speed and high accuracy for filtering scene text images. Experimental results on a benchmark dataset have shown the effectiveness of our method: the teacher network yields state-of-the-art performance, and the distilled small network achieves high performance while maintaining high speed which is 176 times faster on CPU and 3.8 times faster on GPU than a compared benchmark method.
Miao Zhao, Rui-Qi Wang, Xu-Yao Zhang, Linlin Huang 0001, Jean-Marc Ogier
ICDAR6
2018 InDUS: Incremental Document Understanding System Focus on Document Classification
abstract
Our objective is to propose a Document Understanding System for Digital Mailroom application which can cope with three challenges: (1) process a full workflow with high accuracy, with the constraint of a partial training; (2) minimal requirement for configuration work from expert users; (3) adapt incrementally the system in quasi real-time to continuously maximize the recall. We describe an end-to-end system based on existing incremental algorithms for both document classification and field extraction. But in this paper, we really focus on the document classification issue. The main contribution is to adapt the Incremental Growing Neural Gas (A2ING) with a dynamic incremental feature vector. Moreover, a generic Framework automatically selects textual descriptors relying on performance. The quality assessment converges the A2ING and controls the system accuracy.
Vincent Poulain D'Andecy, Aurélie Joseph, Jean-Marc Ogier
DAS3
2018 Feature Selection for Document Flow Segmentation
abstract
In this paper, we describe a method to restore a flow of continuous documents. The flow is a collection of consecutive scanned pages without explicit separation marks between documents. Our method is based on contextual and layout descriptors meant to specify the relationship between each pair of consecutive pages. The relationships are represented using vectors of features with boolean values indicating the presence or the absence of descriptors on concerned pages. The segmentation task therefore consists in classifying such vectors into continuities or breaks. The continuity class indicates that pages belong to the same document while the break class ends the ongoing document and starts a new one. The experimental part is based on a large collection of real administrative documents.
Ahmed Hamdi, Mickaël Coustaty, Aurélie Joseph, Vincent Poulain D'Andecy, Antoine Doucet, Jean-Marc Ogier
DAS6
2018 Offline Arabic Handwriting Recognition Using BLSTMs Combination
abstract
We propose in this paper, an Arabic handwriting recognition system based on multiple BLSTM-CTC combination architectures. Given several feature sets, the low-level fusion consisted in projecting them into a unique feature space. Mid-level combination methods were performed using two techniques: the first one consists in averaging the a-posteriori probabilities of each individual BLSTM, and injecting them in the CTC decoding. The second is based on the training of a new BLSTM-CTC system using the sum of the a-posteriori probabilities generated by the individual systems. The high-level fusion is based on the combination of the individual decoding outputs. Lattice combination and ROVER strategies were evaluated in this context. The experiments conducted on the KHATT database showed that the high-level combination method significantly improves the recognition rate compared to the other fusion strategies.
Sana Khamekhem Jemni, Yousri Kessentini, Slim Kanoun, Jean-Marc Ogier
DAS4
2018 Learning Text Component Features via Convolutional Neural Networks for Scene Text Detection
abstract
Reading the text embedded in natural scene images is essential to many applications. In this paper, we propose a method for detecting text in scene images based on multi-level connected component (CC) analysis and learning text component features via convolutional neural networks (CNN), followed by a graph-based grouping of overlapping text boxes. The multi-level CC analysis allows the extraction of redundant text and non-text components at multiple binarization levels to minimize the loss of any potential text candidates. The features of the resulting raw text/non-text components of different granularity levels are learned via a CNN. Those two modules eliminate the need for complex ad-hoc preprocessing steps for finding initial candidates, and the need for hand-designed features to classify such candidates into text or non-text. The components classified as text at different granularity levels, are grouped in a graph based on the overlap of their extended bounding boxes, then, the connected graph components are retained. This eliminates redundant text components and forms words or textlines. When evaluated on the "Robust Reading Competition" dataset for natural scene images, our method achieved better detection results compared to state-of-the-art methods. In addition to its efficacy, our method can be easily adapted to detect multi-oriented or multi-lingual text as it operates at low level initial components, and it does not require such components to be characters.
Wafa Khlif, Nibal Nayef, Jean-Christophe Burie, Jean-Marc Ogier, Adel M. Alimi
DAS4
2018 Stable Regions and Object Fill-Based Approach for Document Images Watermarking
abstract
In the literature, the document image watermarking schemes in spatial domain mainly focus on text content, so they need to be further improved to be applied on general document content. In this paper, we propose a blindly invisible watermarking approach for grayscale document images in spatial domain, which is based on stable regions and object fill. In order to detect stable regions, the document is transformed into an intermediate form by taking advantage of image processing operations prior to applying nonsubsampled contourlet transform (NSCT). Next, the separated objects in stable regions are obtained by object segmentation. The stroke and fill of obtained objects are detected, and only the locations of object fill are marked as referential ones for mapping to gray level values where data hiding and detection are conducted. Then, the watermarking algorithm is developed by using every group of gray level values corresponding to locations of each object fill for carrying one watermark bit. The experiments are performed with various document contents, and our approach shows high performance in terms of imperceptibility, capacity and robustness against distortions like JPEG compression, geometric transformation and print-and-scan process.
Cu Vinh Loc, Jean-Christophe Burie, Jean-Marc Ogier
DAS3
2017 A Document Straight Line Based Segmentation for Complex Layout Extraction
abstract
Document layout extraction is a difficult step in the image interpretation process due to the high complexity of documents. The main challenge relies on the huge gap between both the physical and the logical structures of document images. In order to loose as few as possible information, most existing methods are working at pixel level. In this paper, we present a new framework for complex layout extraction based on features of high levels obtained from a document straight line based segmentation. We propose to capture the straight line segments thanks to a new transform integrating the local spatial organization of the segments contained in the document content. Such transform can be applied either on the foreground (related to the document content) or the background pixels, in order to take advantage of the duality of information present in both document parts. Experimental results obtained on the PRImA Layout Analysis dataset illustrate the robustness of our framework for the extraction of specific components of the document including text areas, images and separators.
Héloïse Alhéritière, Florence Cloppet, Camille Kurtz, Jean-Marc Ogier, Nicole Vincent
ICDAR4
2017 Local Binary Patterns for Document Forgery Detection
abstract
Document forgery is an increasing problem for both the public administration and private companies. It represents substantial losses in time and economical resources. Classical solutions to this problem such as watermarks or other integrated security patterns can not be applied in general for any unknown incoming document due to the large variability on types of documents. In that scenario it is important to resort to forensic techniques to seek and analyze inconsistencies on the intrinsic features of the document image. In this paper we present a classification-based approach for forgery detection. We use uniform Local Binary Patterns (LBP) to capture discriminant texture features that are common on forged regions. Besides, we combine multiple descriptors from neighboring regions to model contextual information. Results using Support Vector Machines (SVM) for patch classification show that we are able to detect several types of forgeries in a wide range of types of documents.
Francisco Cruz 0003, Nicolas Sidere, Mickaël Coustaty, Vincent Poulain D'Andecy, Jean-Marc Ogier
ICDAR5
2017 A Perceptual Image Hashing Algorithm for Hybrid Document Security
abstract
In order to create an automatic document security system one needs to secure the textual content but also the graphical content of the document. This paper proposes a hashing algorithm capable of securing the graphical parts of paper and digital documents with unprecedented performance and a very small digest. The main challenge for such an algorithm is that of stability, in particular with respect to print and scan noise. We define the generic notion of stability and how to evaluate it. To achieve such performance we use both dense local information and global descriptors. We have tested our method on two datasets totaling nearly 45000 images.
Sébastien Eskenazi, Boris Bodin, Petra Gomez-Krämer, Jean-Marc Ogier
ICDAR4
2017 A Complete Scheme of Spatially Categorized Glyph Recognition for the Transliteration of Balinese Palm Leaf Manuscripts
abstract
To open a wider access to the precious content of historical Balinese palm leaf manuscripts, an appropriate system to transliterate the Balinese script to the Roman script is needed. To achieve this goal, a Balinese glyph recognition scheme is very important. This scheme needs to be developed by taking into account the degraded condition of palm leaf manuscripts and the complexity of Balinese script. In this paper, we present a complete scheme of spatially categorized glyph recognition for the transliteration of Balinese palm leaf manuscripts. For this scheme, five different categories of glyph recognizers based on the spatial positions on the manuscript are proposed. These recognizers will be used to verify and to validate the recognition result of the global glyph recognizer. Each glyph recognizer is built based on the combination of some feature extraction methods and it is trained on a single layer neural network. The trained network is initialized by an unsupervised feature learning. The output of the glyph recognition scheme will be sent as the input to the phonological transliteration system. The results are evaluated with the ground truth of transliterated text provided by philologists. Our scheme shows a very promising result for Balinese palm leaf manuscripts transliteration and can be adapted to other type of script.
Made Windu Antara Kesiman, Jean-Christophe Burie, Jean-Marc Ogier
ICDAR3
2017 Semantic Text Detection in Born-Digital Images via Fully Convolutional Networks
abstract
Traditional layout analysis methods cannot be easily adapted to born-digital images which carry properties from both regular document images and natural scene images. One layout approach for analyzing born-digital images is to separate the text layer from the graphics layer before further analyzing any of them. In this paper, we propose a method for detecting text regions in such images by casting the detection problem as a semantic object segmentation problem. The text classification is done in a holistic approach using fully convolutional networks where the full image is fed as input to the network and the output is a pixel heat map of the same input image size. This solves the problem of low resolution images, and the variability of text scale within one image. It also eliminates the need for finding interest points, candidate text locations or low level components. The experimental evaluation of our method on the ICDAR 2013 dataset shows that our method outperforms state-of-the-art methods. The detected text regions also allow flexibility to later apply methods for finding text components at character, word or textline levels in different orientations.
Nibal Nayef, Jean-Marc Ogier
ICDAR2
2017 ICDAR2017 Robust Reading Challenge on Multi-Lingual Scene Text Detection and Script Identification - RRC-MLT
abstract
Text detection and recognition in a natural environment are key components of many applications, ranging from business card digitization to shop indexation in a street. This competition aims at assessing the ability of state-of-the-art methods to detect Multi-Lingual Text (MLT) in scene images, such as in contents gathered from the Internet media and in modern cities where multiple cultures live and communicate together. This competition is an extension of the Robust Reading Competition (RRC) which has been held since 2003 both in ICDAR and in an online context. The proposed competition is presented as a new challenge of the RRC. The dataset built for this challenge largely extends the previous RRC editions in many aspects: the multi-lingual text, the size of the dataset, the multi-oriented text, the wide variety of scenes. The dataset is comprised of 18,000 images which contain text belonging to 9 languages. The challenge is comprised of three tasks related to text detection and script classification. We have received a total of 16 participations from the research and industrial communities. This paper presents the dataset, the tasks and the findings of this RRC-MLT challenge.
Nibal Nayef, Imen Bizid, Hyunsoo Choi, Dimosthenis Karatzas, Zhenbo Luo, Umapada Pal 0001, Christophe Rigaud, Joseph Chazalon, Wafa Khlif, Muhammad Muzzamil Luqman, Jean-Christophe Burie, Cheng-Lin Liu 0001, Jean-Marc Ogier
ICDAR15
2016 Delaunay Triangulation-Based Features for Camera-Based Document Image Retrieval System
abstract
In this paper, we propose a new feature vector, named DElaunay TRIangulation-based Features (DETRIF), for real-time camera-based document image retrieval. DETRIF is computed based on the geometrical constraints from each pair of adjacency triangles in delaunay triangulation which is constructed from centroids of connected components. Besides, we employ a hashing-based indexing system in order to evaluate the performance of DETRIF and to compare it with other systems such as LLAH and SRIF. The experimentation is carried out on two datasets comprising of 400 heterogeneous-content complex linguistic map images (huge size, 9800 X 11768 pixels resolution) and 700 textual document images.
Quoc Bao Dang, Marçal Rusiñol, Mickaël Coustaty, Muhammad Muzzamil Luqman, De Cao Tran, Jean-Marc Ogier
DAS6
2016 Evaluation of the Stability of Four Document Segmentation Algorithms
abstract
The importance of having stable information extraction algorithms for security related applications and more generally for industrial use cases has been recently highlighted. Stability is what makes an algorithm reliable as it gives a guarantee that the results will be reproducible on similar data. Without it, security criteria such as the probability of false positives cannot be quantified. As a consequence, no security application can be built from an unstable algorithm. In a document verification framework, the probability of false positives indicates the probability that two different results are given for two copies of the same document. This paper builds on our previous work about a stable layout descriptor to study the stability of four segmentation algorithms. We consider that a segmentation algorithm is stable if it produces the same layout for all copies of the same document. The algorithms studied are two versions of PAL, Voronoi, and JSEG. We compare the stability of the different algorithms and study the factors influencing their stability.
Sébastien Eskenazi, Petra Gomez-Krämer, Jean-Marc Ogier
DAS3
2016 Semi-automatic Text and Graphics Extraction of Manga Using Eye Tracking Information
abstract
The popularity of storing, distributing and reading comic books electronically has made the task of comics analysis an interesting research problem. Different work have been carried out aiming at understanding their layout structure and the graphic content. However the results are still far from universally applicable, largely due to the huge variety in expression styles and page arrangement, especially in manga (Japanese comics). In this paper, we propose a comic image analysis approach using eye-tracking data recorded during manga reading sessions. As humans are extremely capable of interpreting the structured drawing content, and show different reading behaviors based on the nature of the content, their eye movements follow distinguishable patterns over text or graphic regions. Therefore, eye gaze data can add rich information to the understanding of the manga content. Experimental results show that the fixations and saccades indeed form consistent patterns among readers, and can be used for manga textual and graphical analysis.
Christophe Rigaud, Nam Le Thanh 0001, Jean-Christophe Burie, Jean-Marc Ogier, Shoya Ishimaru, Motoi Iwata, Koichi Kise
DAS4
2015 The Delaunay Document Layout Descriptor
abstract
Security applications related to document authentication require an exact match between an authentic copy and the original of a document. This implies that the documents analysis algorithms that are used to compare two documents (original and copy) should provide the same output. This kind of algorithm includes the computation of layout descriptors from the segmentation result, as the layout of a document is a part of its semantic content. To this end, this paper presents a new layout descriptor that significantly improves the state of the art. The basic of this descriptor is the use of a Delaunay triangulation of the centroids of the document regions. This triangulation is seen as a graph and the adjacency matrix of the graph forms the descriptor. While most layout descriptors have a stability of 0% with regard to an exact match, our descriptor has a stability of 74% which can be brought up to 100% with the use of an appropriate matching algorithm. It also achieves 100% accuracy and retrieval in a document retrieval scheme on a database of 960 document images. Furthermore, this descriptor is extremely efficient as it performs a search in constant time with respect to the size of the document database and it reduces the size of the index of the database by a factor 400.
Sébastien Eskenazi, Petra Gomez-Krämer, Jean-Marc Ogier
DocEng3
2015 A Conditional Random Field model for font forgery detection
abstract
Nowadays, document forgery is becoming a real issue. A large amount of documents that contain critical information as payment slips, invoices or contracts, are constantly subject to fraudster manipulation because of the lack of security regarding this kind of document. Previously, a system to detect fraudulent documents based on its intrinsic features has been presented. It was especially designed to retrieve copy-move forgery and imperfection due to fraudster manipulation. However, when a set of characters is not present in the original document, copy-move forgery is not feasible. Hence, the fraudster will use a text toolbox to add or modify information in the document by imitating the font or he will cut and paste characters from another document where the font properties are similar. This often results in font type errors. Thus, a clue to detect document forgery consists of finding characters, words or sentences in a document with font properties different from their surroundings. To this end, we present in this paper an automatic forgery detection method based on document font features. Using the Conditional Random Field a measurement of probability that a character belongs to a specific font is made by comparing the character font features to a knowledge database. Then, the character is classified as a genuine or a fake one by comparing its probability to belong to a certain font type with those of the neighboring characters.
Romain Bertrand, Oriol Ramos Terrades, Petra Gomez-Krämer, Patrick Franco, Jean-Marc Ogier
ICDAR5
2015 ICDAR2015 competition on smartphone document capture and OCR (SmartDoc)
abstract
Smartphones are enabling new ways of capture, hence arises the need for seamless and reliable acquisition and digitization of documents, in order to convert them to editable, searchable and a more human-readable format. Current state-of-the-art works lack databases and baseline benchmarks for digitizing mobile captured documents. We have organized a competition for mobile document capture and OCR in order to address this issue. The competition is structured into two independent challenges: smartphone document capture, and smartphone OCR. This report describes the datasets for both challenges along with their ground truth, details the performance evaluation protocols which we used, and presents the final results of the participating methods. In total, we received 13 submissions: 8 for challenge-1, and 5 for challenge-2.
Jean-Christophe Burie, Joseph Chazalon, Mickaël Coustaty, Sébastien Eskenazi, Muhammad Muzzamil Luqman, Maroua Mehri, Nibal Nayef, Jean-Marc Ogier, Sophea Prum, Marçal Rusiñol
ICDAR8
2015 Multiresolution approach based on adaptive superpixels for administrative documents segmentation into color layers
abstract
Administrative document images are usually processed in black and white what generates many problems due to the errors related to the binarization. Besides all semantic information provided by the color is lost. Document images have a rich and highly variable content. The presence of false colors and artefacts introduced by the scanning and the compression alter the segmentation of the regions. Problems arise when there is no correspondence between the point clouds which are detected in a color space and the real regions of an image. In order to help the segmentation, we propose the extraction of the main colors of an image as a set of binary layers. Due to the industrial context, our approach has to run unsupervised on a generic dataset of color administrative documents. The originality of this approach is the use of a multiresolution analysis to detect the number of colors automatically. At a low resolution, a set of local regions is obtained thanks to a SLIC-based approach which takes into account the structure of documents and which combines both colorimetric information and spatial information. Then, a merging stage is applied on each resolution separately based on the colors which have been extracted at a lower resolution. This contribution can both feed the traditional process and exploit colorimetric information.
Elodie Carel, Jean-Christophe Burie, Vincent Courboulay, Jean-Marc Ogier, Vincent Poulain D'Andecy
ICDAR4
2015 Improving document matching performance by local descriptor filtering
abstract
In this paper we propose an effective method aimed at reducing the amount of local descriptors to be indexed in a document matching framework. In an off-line training stage, the matching between the model document and incoming images is computed retaining the local descriptors from the model that steadily produce good matches. We have evaluated this approach by using the ICDAR2015 SmartDOC dataset containing near 25 000 images from documents to be captured by a mobile device. We have tested the performance of this filtering step by using ORB and SIFT local detectors and descriptors. The results show an important gain both in quality of the final matching as well as in time and space requirements.
Joseph Chazalon, Marçal Rusiñol, Jean-Marc Ogier
ICDAR3
2015 A semi-automatic groundtruthing tool for mobile-captured document segmentation
abstract
This paper presents a novel way to generate ground-truth data for the evaluation of mobile document capture systems, focusing on the first stage of the image processing pipeline involved: document object detection and segmentation in low-quality preview frames. We introduce and describe a simple, robust and fast technique based on color markers which enables a semi-automated annotation of page corners. We also detail a technique for marker removal. Methods and tools presented in the paper were successfully used to annotate, in few hours, 24889 frames in 150 video files for the smartDOC competition at ICDAR 2015.
Joseph Chazalon, Marçal Rusiñol, Jean-Marc Ogier, Josep Lladós 0001
ICDAR3
2015 Graph matching versus bag of graph: a comparative study for lettrines recognition
abstract
This paper proposes a comparison of three classification methods of graphical historical images. Historical image datasets are becoming bigger and bigger, and the use of classical computer vision techniques is not sufficient to deal with these large repositories. In the context of this paper, we propose to compare three methods by applying graph matching techniques on a dataset already used in many papers. The first one is based on a statistical approach, the second one on a graph-based classification, and finally the third one is an hybrid approach relying on the specificities of the two previous one. For this last method, we propose here to adapt it to this specific dataset. Some results are proposed and commented, what shows the superiority of the hybrid approach.
Mickaël Coustaty, Jean-Marc Ogier
ICDAR2
2015 SRIF: Scale and Rotation Invariant Features for camera-based document image retrieval
abstract
In this paper, we propose a new feature vector, named Scale and Rotation Invariant Features (SRIF), for real-time camera-based document image retrieval. SRIF is based on Locally Likely Arrangement Hashing (LLAH), which has been widely used and accepted as an efficient real-time camera-based document image retrieval method based on text. SRIF is computed based on geometrical constraints between pairs of nearest points around a keypoint. It can deal with feature point extraction errors which are introduced as a result of the camera capturing of documents. The experimental results show that SRIF outperforms LLAH in terms of retrieval accuracy and processing time.
Quoc Bao Dang, Muhammad Muzzamil Luqman, Mickaël Coustaty, De Cao Tran, Jean-Marc Ogier
ICDAR5
2015 Camera-based document image retrieval system using local features - comparing SRIF with LLAH, SIFT, SURF and ORB
abstract
In this paper, we present camera-based document retrieval systems using various local features as well as various indexing methods. We employ our recently developed features, named Scale and Rotation Invariant Features (SRIF), which are computed based on geometrical constraints between pairs of nearest points around a keypoint. We compare SRIF with state-of-the-art local features. The experimental results show that SRIF outperforms the state-of-the-art in terms of retrieval time with 90.8% retrieval accuracy.
Quoc Bao Dang, Viet Phuong Le, Muhammad Muzzamil Luqman, Mickaël Coustaty, De Cao Tran, Jean-Marc Ogier
ICDAR6
2015 Let's be done with thresholds!
abstract
Current security applications rely on the performances of the algorithms that they use. For document authentication, document analysis algorithms should be precise enough to detect any modification. They should also be stable enough so that a document and its photocopy yield the same result. This requirement is an absolute stability. Having close values is not enough. They need to be exactly the same. This paper presents our preliminary work on the case of a stable layout descriptor. While everyone knows that thresholds are a source of instability, they are still common practice. We describe a promising layout descriptor which drastically reduces the number of thresholds compared to the state of the art. Unfortunately, it is not stable enough when tested on real data. There are still too many thresholds. This paper opens and justifies the path towards algorithms without any threshold.
Sébastien Eskenazi, Petra Gomez-Krämer, Jean-Marc Ogier
ICDAR3
2015 A segmentation free Word Spotting for handwritten documents
abstract
In this paper, a Word Spotting model is presented, that is motivated by some characteristics of the human visual system. The proposed bio-inspired model works at two different levels. First, a Global Filtering module enables to define several candidate zones. Then, a Refining Filtering module facilitates the selection of good retrieved results. These two modules are based on a process of accumulation of votes resulting from the application of generalized Haar-Like-features. The process does not need the segmentation of documents neither in lines nor in words. The proposed approach is evaluated using the George Washington Database and outperforming state-of-the-art performances.
Adam Ghorbel, Jean-Marc Ogier, Nicole Vincent
ICDAR2
2015 An initial study on the construction of ground truth binarized images of ancient palm leaf manuscripts
abstract
Ancient palm leaf manuscripts are one of the very valuable cultural heritages that store various forms of knowledge and historical records of social life in Southeast Asia. The automatic analysis of these documents, in order to extract relevant information, is a real challenge. However, to evaluate the developed extraction algorithms, a ground truth is absolutely necessary. In this paper, we present some of the challenges of the state of the art binarization methods as an initial study for the construction of ground truth binarized images of palm leaf manuscripts. We propose and analyze the need for a specific scheme for the construction of the ground truth of binarized images. The aim of this scheme is to achieve a better ground truth for low quality palm leaf manuscripts. We experimentally tested and evaluated our proposed specific scheme and got promising results. This scheme adapts and performs better in constructing the ground truth of binarized images for palm leaf manuscripts.
Made Windu Antara Kesiman, Sophea Prum, Jean-Christophe Burie, Jean-Marc Ogier
ICDAR4
2015 Content-based comic retrieval using multilayer graph representation and frequent graph mining
abstract
Comics has its large audience and market throughout the world, yet despite the huge research interest given to content-based image retrieval (CBIR) systems, the question of how to effectively retrieve comic images has been little studied. In this paper, we propose a scheme to represent and retrieve comic-page images using attributed Region Adjacency Graphs (RAGs) and their frequent subgraphs. We first extract the graphical structures and local features of each panel of the whole comic volume, then separate different categories of local features to different layers of attributed RAGs. After that, a list of frequent subgraphs for each layer is obtained by using frequent subgraph mining (FSM) technique. For indexing and CBIR purpose, the recognition and ranking are done by checking for isomorphism between the graphs representing the query versus the discovered frequent subgraphs. Our experimental results show that the proposed approach can achieve reliable retrieval results of comic images using query-by-example (QBE) model.
Nam Le Thanh 0001, Muhammad Muzzamil Luqman, Jean-Christophe Burie, Jean-Marc Ogier
ICDAR4
2015 Text and non-text segmentation based on connected component features
abstract
Document image segmentation is crucial to OCR and other digitization processes. In this paper, we present a learning-based approach for text and non-text separation in document images. The training features are extracted at the level of connected components, a mid-level between the slow noise-sensitive pixel level, and the segmentation-dependent zone level. Given all types, shapes and sizes of connected components, we extract a powerful set of features based on size, shape, stroke width and position of each connected component. Adaboosting with Decision trees is used for labeling connected components. Finally, the classification of connected components into text and non-text is corrected based on classification probabilities and size as well as stroke width analysis of the nearest neighbors of a connected component. The performance of our approach has been evaluated on the two standard datasets: UW-III and ICDAR-2009 competition for document layout analysis. Our results demonstrate that the proposed approach achieves competitive performance for segmenting text and non-text in document images of variable content and degradation.
Viet Phuong Le, Nibal Nayef, Muriel Visani, Jean-Marc Ogier, De Cao Tran
ICDAR4
2015 A character degradation model for color document images
abstract
Degradation models are widely used to create large datasets with ground-truth for designing optimal denoising algorithms, and to assess the performance of different document analysis and recognition methods (OCR, segmentation, and so on). Since the processing of document images is usually performed in grayscale or in black and white, some degradation models have been proposed in order to simulate noise on these specific images. However, there is a growing interest in the study of color document images. In this context, there is a real need of tools able to simulate the effects of noise on color for an evaluation purpose. In this paper, we propose to extend a model from the state-of-the-art to color document images. Our model includes three main steps. First, seed-points are selected depending on color information. Then, they are associated to different kinds of noise in order to simulate the degradation effect occurring on different content. Last, the values of pixels located inside an elliptic region around the seed points are modified in order to obtain degraded regions.
Do Thi Luyen, Elodie Carel, Jean-Marc Ogier, Jean-Christophe Burie
ICDAR3
2015 SmartDoc-QA: A dataset for quality assessment of smartphone captured document images - single and multiple distortions
abstract
Smartphones are enabling new ways of capture, hence arises the need for seamless and reliable acquisition and digitization of documents. The quality assessment step is an important part of both the acquisition and the digitization processes. Assessing document quality could aid users during the capture process or help improve image enhancement methods after a document has been captured. Current state-of-the-art works lack databases in the field of document image quality assessment. In order to provide a baseline benchmark for quality assessment methods for mobile captured documents, we present in this paper a dataset for quality assessment that contains both singly- and multiply-distorted document images. The proposed dataset could be used for benchmarking quality assessment methods by the objective measure of OCR accuracy, and could be also used to benchmark quality enhancement methods. There are three types of documents in the dataset: modern documents, old administrative letters and receipts. The document images of the dataset are captured under varying capture conditions (light, different types of blur and perspective angles). This causes geometric and photometric distortions that hinder the OCR process. The ground truth of the dataset images consists of the text transcriptions of the documents, the OCR results of the captured documents and the values of the different capture parameters used for each image. We also present how the dataset could be used for evaluation in the field of no-reference quality assessment. The dataset is freely and publicly available for use by the research community at http://navidomass.univ-lr.fr/SmartDoc-QA.
Nibal Nayef, Muhammad Muzzamil Luqman, Sophea Prum, Sébastien Eskenazi, Joseph Chazalon, Jean-Marc Ogier
ICDAR6
2015 Text zone classification using unsupervised feature learning
abstract
Text zone classification is a vital step in the digitization process, without which OCR systems perform poorly. Prior methods to document zone classification have relied on large sets of hand-crafted features for training zone classifiers. Such features are usually database-dependent, and their computation is time consuming. In this work we propose a novel method for text zone classification that relies on the approach of unsupervised feature learning. Within our method, feature vectors of document zones are automatically learned by patches extraction, encoding and pooling, where feature encoding is based on a codebook of visual words. The training phase of the text classifier takes into consideration the unbalance between text zones and non-text zones of all types. The proposed method has been tested on publicly available standard databases, and achieved competitive or better results compared to state-of-the-art methods. The results show that our approach matches well the task of text classification, and is robust to zone shapes, orientations and size.
Nibal Nayef, Jean-Marc Ogier
ICDAR2
2015 Speech balloon and speaker association for comics and manga understanding
abstract
Comics and manga are one of the most important forms of publication and play a major role in spreading culture all over the world. In this paper we focus on balloons and their association to comic characters or more generally text and graphic links retrieval. This information is not directly encoded in the image, whether scanned or digital-born, it has to be understood according to other information present in the image. Such high level information allows new browsing experience and story understanding (e.g. dialog analysis, situation retrieval). We propose a speech balloon and comic character association method able to retrieve which character is emitting which speech balloon. The proposed method is based on geometric graph analysis and anchor point selection. This work has been evaluated over various comic book styles from the eBDtheque dataset and also a volume of the Kingdom manga series.
Christophe Rigaud, Nam Le Thanh 0001, Jean-Christophe Burie, Jean-Marc Ogier, Motoi Iwata, Eiki Imazu, Koichi Kise
ICDAR4
2015 A comparative study of local detectors and descriptors for mobile document classification
abstract
In this paper we conduct a comparative study of local key-point detectors and local descriptors for the specific task of mobile document classification. A classification architecture based on direct matching of local descriptors is used as baseline for the comparative study. A set of four different key-point detectors and four different local descriptors are tested in all the possible combinations. The experiments are conducted in a database consisting of 30 model documents acquired on 6 different backgrounds, totaling more than 36.000 test images.
Marçal Rusiñol, Joseph Chazalon, Jean-Marc Ogier, Josep Lladós 0001
ICDAR3
2014 Efficient Example-Based Super-Resolution of Single Text Images Based on Selective Patch Processing
abstract
Example-based super-resolution (SR) methods learn the correspondences between low resolution (LR) and high-resolution (HR) image patches, where the patches are extracted from a training database. To reconstruct a single LR image into a HR one, each LR image patch is processed by the previously trained model to recover its corresponding HR patch. For this reason, they are computationally inefficient. We propose the use of a selective patch processing technique to carry out the super-resolution step more efficiently, while maintaining the output quality. In this technique, only patches of high variance are processed by the costly reconstruction steps, while the rest of the patches are processed by fast bicubic interpolation. We have applied the proposed improvement on representative example-based SR methods to super-resolve text images. The results show a significant speed up for text SR without a drop in theocrat accuracy. In order to carry out an extensive and solid performance evaluation, we also present a public database of text images for training and testing example-based SR methods.
Nibal Nayef, Joseph Chazalon, Petra Gomez-Krämer, Jean-Marc Ogier
Document Analysis Systems4
2014 Color Descriptor for Content-Based Drawing Retrieval
abstract
Human detection in computer vision field is an active field of research. Extending this to human-like drawings such as the main characters in comic book stories is not trivial. Comics analysis is a very recent field of research at the intersection of graphics, texts, objects and people recognition. The detection of the main comic characters is an essential step towards a fully automatic comic book understanding. This paper presents a color-based approach for comics character retrieval using content-based drawing retrieval and color palette.
Christophe Rigaud, Dimosthenis Karatzas, Jean-Christophe Burie, Jean-Marc Ogier
Document Analysis Systems4
2014 Combining Focus Measure Operators to Predict OCR Accuracy in Mobile-Captured Document Images
abstract
Mobile document image acquisition is a new trend raising serious issues in business document processing workflows. Such digitization procedure is unreliable, and integrates many distortions which must be detected as soon as possible, on the mobile, to avoid paying data transmission fees, and losing information due to the inability to re-capture later a document with temporary availability. In this context, out-of-focus blur is major issue: users have no direct control over it, and it seriously degrades OCR recognition. In this paper, we concentrate on the estimation of focus quality, to ensure a sufficient legibility of a document image for OCR processing. We propose two contributions to improve OCR accuracy prediction for mobile-captured document images. First, we present 24 focus measures, never tested on document images, which are fast to compute and require no training. Second, we show that a combination of those measures enables state-of-the art performance regarding the correlation with OCR accuracy. The resulting approach is fast, robust, and easy to implement in a mobile device. Experiments are performed on a public dataset, and precise details about image processing are given.
Marçal Rusiñol, Joseph Chazalon, Jean-Marc Ogier
Document Analysis Systems3
2013 Dominant color segmentation of administrative document images by hierarchical clustering
abstract
This paper addresses the problem of color documents images segmentation in an industrial context. Automated Document Recognition (ADR) systems highly reduce time and resource costs of companies by managing their huge amount of administrative documents, and by optimizing their workflow. Most of the time, a binarization is performed due to their historical industrial process. Therefore, colorimetric information can improve the process. In this paper, we propose a hierarchical clustering based approach to extract dominant color masks of documents. Indeed, our dataset comprises different kind of scanned administrative document images such as invoices, forms, letters, and so on. We do not know a priori the number of dominant colors on our documents. These masks will further feed the inputs to an OCR in order to bring extra-information about the colorimetric context. This approach requires neither user interaction nor setting steps. Experiments on several types of documents show the relevance of the proposed approach
Elodie Carel, Vincent Courboulay, Jean-Christophe Burie, Jean-Marc Ogier
ACM Symposium on Document Engineering4
2013 Visual saliency and terminology extraction for document annotation
abstract
The document digitization process becomes a crucial economical issue in our society. Then, it becomes necessary to be able to organize this huge amount of documents. The work proposed in this paper tends to propose a new method to automatically classify document using a saliency-based segmentation process on one hand, and a terminology extraction and annotation on the other hand. The saliency-based segmentation is used to extract salient regions and by the way logo, while the terminology approach is used to annotate them and to automatically classify the document. The approach does not require human expertise, and use Google Images as a knowledge database. The results obtained on a real database of 1766 documents show the relevance of the approach.
Benjamin Duthil, Mickaël Coustaty, Vincent Courboulay, Jean-Marc Ogier
ACM Symposium on Document Engineering4
2013 Bag of subjects: lecture videos multimodal indexing
abstract
In this paper, we address multimodal indexing and retrieval for videos of lectures or seminars. This paper proposes a combination of technologies respectively issuing from image document analysis and text mining. Based on visual information and textual information extracted from slide images, we investigate a Bag of mixed Words (visual words and textual words) model to represent lecture slide's contents. Lecture videos are indexed and retrieved by using extended Bag of Words model. In this model, it is assumed that a video may contain multiple subjects; and this model discovers the visual representation of these subjects automatically and indexes the video accordingly. We discuss the mixed text/image query and proposed indexing approach for retrieval lecture videos and report a quantitative evaluation on lecture videos of our Lab.
Vincent Nguyen 0001, Jean-Marc Ogier, Franck Charneau
ACM Symposium on Document Engineering2
2013 A System Based on Intrinsic Features for Fraudulent Document Detection
abstract
Paper documents still represent a large amount of information supports used nowadays and may contain critical data. Even though official documents are secured with techniques such as printed patterns or artwork, paper documents suffer from a lack of security. However, the high availability of cheap scanning and printing hardware allows non-experts to easily create fake documents. As the use of a watermarking system added during the document production step is hardly possible, solutions have to be proposed to distinguish a genuine document from a forged one. In this paper, we present an automatic forgery detection method based on document's intrinsic features at character level. This method is based on the one hand on outlier character detection in a discriminant feature space and on the other hand on the detection of strictly similar characters. Therefore, a feature set is computed for all characters. Then, based on a distance between characters of the same class, the character is classified as a genuine one or a fake one.
Romain Bertrand, Petra Gomez-Krämer, Oriol Ramos Terrades, Patrick Franco, Jean-Marc Ogier
ICDAR5
2013 eBDtheque: A Representative Database of Comics
abstract
We present eBDtheque, a database of various comic book images and their ground truth for panels, balloons and text lines plus semantic annotations. The database consists of a hundred pages of various comic book albums, Franco-Belgian, American comics and mangas. Additionally, we present the piece of software used to establish the ground truth and a tool to validate results against this ground truth. Everything is publicly available for scientific use on http://ebdtheque.univ-lr.fr.
Clément Guérin, Christophe Rigaud, Antoine Mercier 0003, Farid Ammar-Boudjelal, Karell Bertet, Alain Bouju, Jean-Christophe Burie, Georges Louis, Jean-Marc Ogier, Arnaud Revel
ICDAR9
2013 Text-Independent Writer Identification on Online Arabic Handwriting
abstract
Most of existing works on online text-independent writer identification follow analytical approach based on grapheme or character as primitive. However, segmentation is time-consuming and becomes an obstacle in real time application systems. To avoid this problem, we propose a novel framework which follows global approach based on word as primitive. Different sets of statistic and dynamic features are used in different levels in the word. Features are extracted from the point, the stroke, the space between strokes and the whole word. In decision phase, Dynamic Time Warping (DTW) and Support Vector Machine (SVM) are used. To evaluate our framework, we use set 1 from the ADAB database (Arabic DAtaBase). A limited amount of data is available for some writers. In this paper, we focus on the effect of writers' number and words' number per writer on the obtained results. We highlight the accuracy of studied features and used classifiers. Experimental results are promising and show that our proposed framework can improve the identification rates.
Mariem Gargouri Kchaou, Slim Kanoun, Jean-Marc Ogier
ICDAR3
2013 Improving Logo Spotting and Matching for Document Categorization by a Post-Filter Based on Homography
abstract
Digital document categorization based on logo spotting and recognition has raised a great interest in the research community because logos in documents are sources of information for categorizing documents with low costs. In this paper, we present an approach to improve the result of our method for logo spotting and recognition based on key point matching and presented in our previous paper [7]. First, the key points from both the query document images and a given set of logos (logo gallery) are extracted and described by SIFT, and are matched in the SIFT feature space. Secondly, logo segmentation is performed using spatial density-based clustering. The contribution of this paper is to add a third step where homography is used to filter the matched key points as a post-processing. And finally, in the decision stage, logo classification is performed by using an accumulating histogram. Our approach is tested using a well-known benchmark database of real world documents containing logos, and achieves good performances compared to state-of-the-art approaches.
Viet Phuong Le, Muriel Visani, De Cao Tran, Jean-Marc Ogier
ICDAR4
2013 Interactive Knowledge Learning for Ancient Images
abstract
This paper deals with cultural heritage preservation and ancient document indexing. In the management of historical documents, ancient images are described using semantic information, often manually annotated by historians. In this paper, we propose an approach to interactively propagate the historians' knowledge to a database of drop caps images manually populated by historians with drop caps image annotations. Based on a novel document indexing processing scheme which combines the use of the Zipf law and the use of bag of patterns, our approach extends the Bag of Words model to represent the knowledge by visual features through relevance feedback. Then annotation propagation is automatically performed to propagate knowledge to the drop caps image database. In this article, our approach is presented together with preliminary experimental results and an illustrative example.
Vincent Nguyen 0001, Mickaël Coustaty, Alain Boucher, Jean-Marc Ogier
ICDAR4
2013 A Discriminative Approach to On-Line Handwriting Recognition Using Bi-character Models
abstract
Unconstrained on-line handwriting recognition is typically approached within the framework of generative HMM-based classifiers. In this paper, we introduce a novel discriminative method that relies, in contrast, on explicit grapheme segmentation and SVM-based character recognition. In addition to single character recognition with rejection, bi-characters are recognized in order to refine the recognition hypotheses. In particular, bi-character recognition is able to cope with the problem of shared character parts. Whole word recognition is achieved with an efficient dynamic programming method similar to the Viterbi algorithm. In an experimental evaluation on the Unipen-ICROW-03 database, we demonstrate improvements in recognition accuracy of up to 8% for a lexicon of 20,000 words with the proposed method when compared with an HMM-based baseline system. The computational speed is on par with the baseline system.
Sophea Prum, Muriel Visani, Andreas Fischer 0002, Jean-Marc Ogier
ICDAR4
2013 An Active Contour Model for Speech Balloon Detection in Comics
abstract
Comic books constitute an important cultural heritage asset in many countries. Digitization combined with subsequent comic book understanding would enable a variety of new applications, including content-based retrieval and content retargeting. Document understanding in this domain is challenging as comics are semi-structured documents, combining semantically important graphical and textual parts. Few studies have been done in this direction. In this work we detail a novel approach for closed and non-closed speech balloon localization in scanned comic book pages, an essential step towards a fully automatic comic book understanding. The approach is compared with existing methods for closed balloon localization found in the literature and results are presented.
Christophe Rigaud, Jean-Christophe Burie, Jean-Marc Ogier, Dimosthenis Karatzas, Joost van de Weijer 0001
ICDAR3
2013 Specific Comic Character Detection Using Local Feature Matching
abstract
Comic books are a kind of storytelling graphic publications mainly expressed by abstract line drawings. As a clue of story lines, comic characters play an important role in the story, and their detection is an essential part of comic book analysis. For this purpose, the task includes (1) locating characters in comics pages and (2) identifying them, which is called specific character detection. Corresponding to different scenes of comic books, one specific character can be represented by various expressions coupled with rotations, occlusions, and other perspective drawing effects, which challenge the detection. In this paper, we focus on stable features regarding the possible transformations and proposed a framework to detect them. Specifically, some discriminative features are selected as detectors for characterizing characters, on the basis of a training dataset. Based on the detectors, the drawings of the same characters in different scenes can be detected. The methodology has been experimented and validated on 6 titles of comics. Despite the terrific changes for different scenes, the proposed method achieved detection of 70% comic characters.
Weihan Sun, Jean-Christophe Burie, Jean-Marc Ogier, Koichi Kise
ICDAR3
2012 Panel and Speech Balloon Extraction from Comic Books
abstract
Comic books represent an important cultural heritage in many countries. However, few researches have been done in order to analyse the content of comics such as panels, speech balloons or characters. At first glance, the structure of a comic page may appear easy to determine. In practice, the configuration of the page, the size and the shape of the panels can be different from one page to the next. Moreover, authors often draw extended contents (speech balloon or comic art) that overlap two panels or more. In some situations, the panel extraction can become a real challenge. Speech balloons are other important elements of comics. Full text indexing is only possible if the text can be extracted. However the text is usually embedded among graphic elements. Moreover, unlike newspapers, the text layout in speech balloons can be irregular. Classic text extraction method can fail. We propose, in this paper, a method based on region growing and mathematical morphology to extract automatically the panels of a comic page and a method to detect speech balloons. Our approach is compared with other methods find in the literature. Results are presented and discussed.
Anh Khoi Ngo Ho, Jean-Christophe Burie, Jean-Marc Ogier
Document Analysis Systems3
2011 Writer Identification Using TF-IDF for Cursive Handwritten Word Recognition
abstract
In this paper, we present two text-independent writer identification methods in a closed-world context. Both methods use on-line and off-line features jointly with a classifier inspired from information retrieval methods. These methods are local, respectively based on the character and grapheme levels. This writer identification engine may be used to personalize our cursive word recognition engine [1] to the handwriting style of the writer, resulting in an adaptive cursive word recognizer. Experiments assess the effectiveness of the proposed approaches in a context of writer identification as well as integrated to our cursive word recognizer to make it adaptive.
Quang Anh Bui, Muriel Visani, Sophea Prum, Jean-Marc Ogier
ICDAR4
2011 Discrimination of Old Document Images Using Their Style
abstract
Based on the principle described by Pareti et al. in [1], [2], and by Chouaib et al. in [3], this paper proposes to combine the use of the Zipf law and the use of bag of patterns for the implementation of a document indexing processing scheme. Contrarily to these two mentioned approaches, we retain the most important patterns based on the TF-IDF criteria, and the pattern selection is local. This paper presents the different stages of our indexing process, as well as their application to historical documents. Results on comlex images are given, illustrated and discussed.
Mickaël Coustaty, Jean-Marc Ogier
ICDAR2
2011 Bags of Strokes Based Approach for Classification and Indexing of Drop Caps
abstract
This paper proposes an approach to process drop cap images - images of decorated letter that begin chapters of old documents that are preserved in libraries, museums - in the domain of characterization, classification and indexing of old documents. The originality of our proposal is based on the fact that we do not try to extract the letter of drop caps but to classify the drop caps according to period, author and style. The drop caps are characterized by using relevant visual features such as length, thickness, orientation, complexity and change of direction on their primitive elements: strokes. The purpose of this approach is to efficiently extract information embedded in the drop caps for the classification and the indexing of old documents. These new visual features based on bags of strokes are more easily calculable and generally applicable than texture or shape features. Experiments based on characterization, classification and indexing phases demonstrate the performance of our propositions and the advances that they represent in terms of content-based drop caps retrieval.
Thi Thuong Huyen Nguyen, Mickaël Coustaty, Jean-Marc Ogier
ICDAR3
2010 Form recognition from ink strokes on tablet
abstract
This paper proposes a method for form recognition from handwritten input captured as digital ink on a tablet. Form recognition is an important step of form processing to read the data on a filled form. This type of recognition is different from traditional image matching (searching and retrieving) because the query image (data ink) has but few common characteristics with the retrieval images (i.e. form templates). The fact is that form structure is free in variation. It is not rare that two forms are very close in their structures and in semantic of fields. Filling in a form is also free in variation. The same form may be filled with different contents and in different ways of online context. The main idea for matching between ink strokes and form template in this paper is featureless, based on Bhattacharyya measure. The distance between the distribution of ink strokes and the distribution of form fields is the matching measure. These distributions are spatial information which is based on the crossing of coordinates of ink points and the crossing of fields to be filled. However, these coordinates are not taken on the same coordinate system. Ink point coordinates are based on the tablet coordinate system (differs from A4 format) while field coordinates are in paper size (for example, A4 format). In order to deal with this problem, affine transform is used to standardize the coordinate system. The coordinate system on the tablet is transformed into paper format system.
De Cao Tran, Patrick Franco, Jean-Marc Ogier
Document Analysis Systems3
2009 Drop Caps Decomposition for Indexing a New Letter Extraction Method
abstract
This paper present a new method to extract shapes in drop caps and particularly the most important shape: letter itself. This method relies on a combination of a Aujol and Chambolle algorithm first, and a segmentation using a Zipf law in a second step. This method can be enhanced as a three-step process: 1) decomposition in layers 2) segmentation using a Zipf law 3) selection of connected components to only emphasize the required information - letter itself.
Mickaël Coustaty, Jean-Marc Ogier, Rudolf Pareti, Nicole Vincent
ICDAR2
2008 Object Extraction from Colour Cadastral Maps
abstract
In this paper, an object extraction method from ancient colour maps is proposed. It consists on the localization of quarters inside a given cadastral map. The colour aspect is exploited thanks to a colour restoration algorithm and the selection of a relevant hybrid colour model. Objects composing the map are located using a multi-components gradient. To identify quarters, a peeling the onion method is adopted. This selective method starts by separated text and graphics. On the graphic layer, a connected component analysis is carried out through the use of a neighbourhood graph. This graph is smartly pruned to consider only significant areas. Consequently, the quarter boundaries are found using a snake which is a computer-generated curve that moves within an image to fit a given object. The performance of our method is measured up in two steps: Firstly, the colour space selection is assessed according to the colour distinction capacity while being robust to variations/noise then the automatic extraction approach is compared to the user ground truth. Results show the good behaviour of the whole system.
Romain Raveaux, Jean-Christophe Burie, Jean-Marc Ogier
Document Analysis Systems3
2007 A Colour Document Interpretation: Application to Ancient Cadastral Maps
abstract
In this paper, a colour graphic document analysis is proposed with an application to ancient cadastral maps. The approach relies on the idea that images of document are fairly different than usual images, such as natural scenes or paintings. From this statement, we present an architecture for colour document understanding. It is based on two paradigms. Firstly, a dedicated colour representation named adapted colour space which aims to learn the image colour specificity and secondly a document oriented segmentation using a region growing algorithm supervised by a hierarchical strategy. Experiments are performed to judge the whole process and the first results show a good behaviour in term of information retrieval.
Romain Raveaux, Jean-Christophe Burie, Jean-Marc Ogier
ICDAR3
2006 The Fuzzy-Spatial Descriptor for the Online Graphic Recognition: Overlapping Matrix Algorithm
Noorazrin Zakaria, Jean-Marc Ogier, Josep Lladós 0001
Document Analysis Systems2
2004 DocMining: A Document Analysis System Builder
Sébastien Adam, Maurizio Rigamonti, Eric Clavier, Éric Trupin, Jean-Marc Ogier, Karl Tombre, Joël Gardes
Document Analysis Systems5
2003 Symbols Recognition by Global-Local Structural Approaches, Based on the Scenarios Use, and with a XML Representation of Data
abstract
symbols on the documents. We have based our system on a combination of local and global structural approaches. The global approach groups the connected components together according to some closeness and connection constraints. The local approach splits up each connected component into a graph of geometrical objects (vectors, arcs, curves). The extracted graphs are matched thanks to a structural classifier, which permits graph-subgraph and exact-inexact matching. The system adaptability is obtained thanks to the scenarios use. A XML data representation is used, allowing the data manipulations and the graphic representations of results.
Mathieu Delalandre, Stéphane Nicolas, Éric Trupin, Jean-Marc Ogier
ICDAR4
1999 Multi-scaled and Multi-oriented Character Recognition: An Original Strategy
abstract
We propose an original methodology allowing detection and recognition of multi-oriented and multi-scaled shapes. The supports on which the method is applied are technical documents representing the network of the French telephonic operator (France Telecom) overlaid on urban maps. The adopted technique, based on the Mellin Fourier Transform is integrated in a global strategy that permits one to solve ambiguous situations, through the provision of contextual information. The strategy, which is applied to solve the character/symbol classification problem, can be divided into two stages. The first one consists of constructing a moment invariants vector from each shape which is extracted from a character layer issued from the system approach. The second consists of detecting and recognising connected shapes. The results of the application of this technique are very encouraging, since the classification rate reaches excellent scores if we consider that no contextual information has been integrated in the recognition process (orientation of the string, integration of data issued from dictionaries stored on alpha-numeric databases).
Sébastien Adam, Jean-Marc Ogier, Claude Cariou, Rémy Mullot, Joël Gardes, Yves Lecourtier
ICDAR2
1997 An Image Interpretation Device cannot be Reliable without any Semantic Coherency Analysis of the Interpretated Objects - Application to French Cadastral Maps
abstract
The methodology we have used for the interpretation of French cadastral documents focuses on a number of studied results of the visual perception and a hierarchical description of the document. The strategy used has been based on the "model" document, employing a mixed approach including various "points of view" about the image to be processed. The results of this mixed analysis reveal the appearance of noninterpretable objects on the cadaster, due to the presence of semantic uncoherence. Thanks to the return cycles between the high and low level processing, an analytical strategy is proposed to independently cure the incoherence, thus to attain the most reliable interpretation of the cadastral map.
Jean-Marc Ogier, Rémy Mullot, Jacques Labiche, Yves Lecourtier
ICDAR1
1993 Attributes extraction for French map interpretation
abstract
The authors deal with a cadastral map interpretation device. The approach to solving this problem is twofold: first, it consists in vectorizing the image by extracting the lowest information level; then, using the low level primitives and introducing the knowledge of the cadastral map, it consists in reconstructing real cadastral entities. The authors present the decomposition of the information into different levels of abstraction (from the low level information to the real cadastral object). They also present the interactions between these different levels, interactions which are managed by the consistency of the objects on each of the different levels. They then present the original tools allowing extraction of the low level primitives.>
Jean-Marc Ogier, Jacques Labiche, Rémy Mullot, Yves Lecourtier
ICDAR1