Thierry Paquet

dblp:06/6769 · DBLP profile ↗
← Back
87ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-2044-7542ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 69 · 3 first-author · 13 since 2021Databases, data management, data science and information retrieval · 44 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 since 2021
YearPublicationVenuePosition
2026 Few-Shot Writer Adaptation via Multimodal In-Context Learning
Tom Simon, Pierrick Tranouez, Stéphane Nicolas, Clément Chatelain 0001, Thierry Paquet
ICDAR (2)5
2026 DenTab: A Dataset for Table Recognition and Visual QA on Real-World Dental Estimates
Laziz Hamdi, Amine Tamasna, Thierry Paquet
ICPR (4)3
2026 End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music
abstract
Abstract Optical Music Recognition (OMR) has made significant progress since its inception, with various approaches now capable of accurately transcribing music scores into digital formats. Despite these advancements, most so-called end-to-end OMR approaches still rely on multi-stage processing pipelines for transcribing full-page score images, which entails challenges such as the need for dedicated layout analysis and specific annotated data, thereby limiting the general applicability of such methods. In this paper, we present the first truly end-to-end approach for page-level OMR in complex layouts. Our system, which combines convolutional layers with autoregressive Transformers, processes an entire music score page and outputs a complete transcription in a music encoding format. This is made possible by both the architecture and the training procedure, which utilizes curriculum learning through incremental synthetic data generation. We evaluate the proposed system using pianoform corpora, which is one of the most complex sources in the OMR literature. This evaluation is conducted first in a controlled scenario with synthetic data, and subsequently against two real-world corpora of varying conditions. Our approach is compared with leading commercial OMR software. The results demonstrate that our system not only successfully transcribes full-page music scores but also outperforms the commercial tool in both zero-shot settings and after fine-tuning with the target domain, representing a significant contribution to the field of OMR.
Antonio Ríos-Vila, Jorge Calvo-Zaragoza, David Rizo, Thierry Paquet
Int. J. Comput. Vis.4
2025 Classifying the Unknown: In-Context Learning for Open-Vocabulary Text and Symbol Recognition
Tom Simon, William Mocaër, Pierrick Tranouez, Clément Chatelain 0001, Thierry Paquet
ICDAR (4)5
2025 DANIEL: a fast document attention network for information extraction and labelling of handwritten documents
Thomas Constum, Pierrick Tranouez, Thierry Paquet
Int. J. Document Anal. Recognit.3
2024 End-to-End Information Extraction in Handwritten Documents: Understanding Paris Marriage Records from 1880 to 1940
Thomas Constum, Lucas Preel, Théo Larcher, Thierry Paquet, Pierrick Tranouez, Sandra Brée
ICDAR (3)4
2024 Sheet Music Transformer: End-To-End Optical Music Recognition Beyond Monophonic Transcription
Antonio Ríos-Vila, Jorge Calvo-Zaragoza, Thierry Paquet
ICDAR (6)3
2023 Faster DAN: Multi-target Queries with Document Positional Encoding for End-to-End Handwritten Document Recognition
Denis Coquenet, Clément Chatelain 0001, Thierry Paquet
ICDAR (4)3
2023 End-to-End Handwritten Paragraph Text Recognition Using a Vertical Attention Network
abstract
Unconstrained handwritten text recognition remains challenging for computer vision systems. Paragraph text recognition is traditionally achieved by two models: the first one for line segmentation and the second one for text line recognition. We propose a unified end-to-end model using hybrid attention to tackle this task. This model is designed to iteratively process a paragraph image line by line. It can be split into three modules. An encoder generates feature maps from the whole paragraph image. Then, an attention module recurrently generates a vertical weighted mask enabling to focus on the current text line features. This way, it performs a kind of implicit line segmentation. For each text line features, a decoder module recognizes the character sequence associated, leading to the recognition of a whole paragraph. We achieve state-of-the-art character error rate at paragraph level on three popular datasets: 1.91% for RIMES, 4.45% for IAM and 3.59% for READ 2016. Our code and trained model weights are available at https://github.com/FactoDeepLearning/VerticalAttentionOCR.
Denis Coquenet, Clément Chatelain 0001, Thierry Paquet
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 DAN: A Segmentation-Free Document Attention Network for Handwritten Document Recognition
abstract
Unconstrained handwritten text recognition is a challenging computer vision task. It is traditionally handled by a two-step approach, combining line segmentation followed by text line recognition. For the first time, we propose an end-to-end segmentation-free architecture for the task of handwritten document recognition: the Document Attention Network. In addition to text recognition, the model is trained to label text parts using begin and end tags in an XML-like fashion. This model is made up of an FCN encoder for feature extraction and a stack of transformer decoder layers for a recurrent token-by-token prediction process. It takes whole text documents as input and sequentially outputs characters, as well as logical layout tokens. Contrary to the existing segmentation-based approaches, the model is trained without using any segmentation label. We achieve competitive results on the READ 2016 dataset at page level, as well as double-page level with a CER of 3.43% and 3.70%, respectively. We also provide results for the RIMES 2009 dataset at page level, reaching 4.54% of CER. We provide all source code and pre-trained model weights at https://github.com/FactoDeepLearning/DAN.
Denis Coquenet, Clément Chatelain 0001, Thierry Paquet
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Confidence Estimation for Object Detection in Document Images
Mélodie Boillet, Christopher Kermorvant, Thierry Paquet
Pattern Recognit. Lett.3
2022 Recognition and Information Extraction in Historical Handwritten Tables: Toward Understanding Early 20th Century Paris Census
Thomas Constum, Nicolas Kempf, Thierry Paquet, Pierrick Tranouez, Clément Chatelain 0001, Sandra Brée, François Merveille
DAS3
2022 Robust text line detection in historical documents: learning and evaluation methods
Mélodie Boillet, Christopher Kermorvant, Thierry Paquet
Int. J. Document Anal. Recognit.3
2021 SPAN: A Simple Predict & Align Network for Handwritten Paragraph Recognition
abstract
International audience
Denis Coquenet, Clément Chatelain 0001, Thierry Paquet
ICDAR (3)3
2020 A Named Entity Extraction System for Historical Financial Data
Wassim Swaileh, Thierry Paquet, Sébastien Adam, Andres Rojas Camacho
DAS2
2020 Recurrence-free unconstrained handwritten text recognition using gated fully convolutional network
abstract
Unconstrained handwritten text recognition is a major step in most document analysis tasks. This is generally processed by deep recurrent neural networks and more specifically with the use of Long Short-Term Memory cells. The main drawbacks of these components are the large number of parameters involved and their sequential execution during training and prediction. One alternative solution to using LSTM cells is to compensate the long time memory loss with an heavy use of convolutional layers whose operations can be executed in parallel and which imply fewer parameters. In this paper we present a Gated Fully Convolutional Network architecture that is a recurrence-free alternative to the well-known CNN+LSTM architectures. Our model is trained with the CTC loss and shows competitive results on both the RIMES and IAM datasets. We release all code to enable reproduction of our experiments: https://github.com/FactoDeepLearning/LinePytorchOCR.
Denis Coquenet, Clément Chatelain 0001, Thierry Paquet
ICFHR3
2020 Multiple Document Datasets Pre-training Improves Text Line Detection With Deep Neural Networks
abstract
In this paper, we introduce a fully convolutional network for the document layout analysis task. While state-of-the-art methods are using models pre-trained on natural scene images, our method Doc-UFCN relies on a U-shaped model trained from scratch for detecting objects from historical documents. We consider the line segmentation task and more generally the layout analysis problem as a pixel-wise classification task then our model outputs a pixel-labeling of the input images. We show that Doc-UFCN outperforms state-of-the-art methods on various datasets and also demonstrate that the pre-trained parts on natural scene images are not required to reach good results. In addition, we show that pre-training on multiple document datasets can improve the performances. We evaluate the models using various metrics to have a fair and complete comparison between the methods.
Mélodie Boillet, Christopher Kermorvant, Thierry Paquet
ICPR3
2020 Handwriting recognition using cohort of LSTM and lexicon verification with extremely large lexicon
Bruno Stuner, Clément Chatelain 0001, Thierry Paquet
Multim. Tools Appl.3
2020 Multi-scale Gated Fully Convolutional DenseNets for semantic labeling of historical newspaper images
Yann Soullard, Pierrick Tranouez, Clément Chatelain 0001, Stéphane Nicolas, Thierry Paquet
Pattern Recognit. Lett.5
2019 Improving Text Recognition using Optical and Language Model Writer Adaptation
abstract
State-of-the-art methods for handwriting text recognition are based on deep learning approaches and language modeling that require large data sets during training. In practice, there are some applications where the system processes mono-writer documents, and would thus benefit from being trained on examples from that writer. However, this is not common to have numerous examples coming from just one writer. In this paper, we propose an approach to adapt both the optical model and the language model to a particular writer, from a generic system trained on large data sets with a variety of examples. We show the benefits of the optical and language model writer adaptation. Our approach reaches competitive results on the READ 2018 data set, which is dedicated to model adaptation to particular writers.
Yann Soullard, Wassim Swaileh, Pierrick Tranouez, Thierry Paquet, Clément Chatelain 0001
ICDAR4
2019 A unified multilingual handwriting recognition system using multigrams sub-lexical units
Wassim Swaileh, Yann Soullard, Thierry Paquet
Pattern Recognit. Lett.3
2018 Fully convolutional network with dilated convolutions for handwritten text line segmentation
Guillaume Renton, Yann Soullard, Clément Chatelain 0001, Sébastien Adam, Christopher Kermorvant, Thierry Paquet
Int. J. Document Anal. Recognit.6
2017 Self-Training of BLSTM with Lexicon Verification for Handwriting Recognition
abstract
Deep learning approaches now provide state-of-the-art performance in many computer vision tasks such as handwriting recognition. However, the huge number of parameters of these models require big annotated training datasets which are difficult to obtain. Training neural networks with unlabeled data is one of the key problems to achieve significant progress in deep learning. In this article, we explore a new semi-supervised training strategy to train long-short term memory (LSTM) recurrent neural networks for isolated handwritten words recognition. The idea of our self-training strategy relies on the iteration of training Bidirectional LSTM recurrent neural network (BLSTM) using both labeled and unlabeled data. At each iteration the current trained network labels the unlabeled data and submit them to a very efficient "lexicon verification" rule. Verified unlabeled data are added to the labeled dataset at the end of each iteration. This verification stage has very low sensitivity to the lexicon size, and a full word coverage of the dataset is not necessary to make the semi-supervised method efficient. The strategy enables self-training with a single BLSTM and show promising results on the Rimes dataset.
Bruno Stuner, Clément Chatelain 0001, Thierry Paquet
ICDAR3
2017 Handwriting Recognition with Multigrams
abstract
We introduce a novel handwriting recognition approach based on sub-lexical units known as multigrams of characters, that are variable lengths characters sequences. A Hidden Semi Markov model is used to model the multigrams occurrences within the target language corpus. Decoding the training language corpus with this model provides an optimized multigram lexicon of reduced size with high coverage rate of OOV compared to the traditional word modeling approach. The handwriting recognition system is composed of two components: the optical model and the statistical n-grams of multigrams language model. The two models are combined together during the recognition process using a decoding technique based on Weighted Finite State Transducers (WFST). We experiment the approach on two Latin language datasets (the French RIMES and English IAM datasets) and we show that it outperforms words and character models language models for high Out Of Vocabulary (OOV) words rates, and that it performs similarly to these traditional models for low OOV rates, with the advantage of a reduced complexity.
Wassim Swaileh, Thierry Paquet, Yann Soullard, Pierrick Tranouez
ICDAR2
2017 Gesture sequence recognition with one shot learned CRF/HMM hybrid model
Selma Belgacem, Clément Chatelain 0001, Thierry Paquet
Image Vis. Comput.3
2016 A Lexicon Verification Strategy in a BLSTM Cascade Framework
abstract
Handwriting recognition always has been a difficult problem, with image related problems on the one hand and language processing on the other hand. Significant improvements have been made in handwriting recognition thanks to new recurrent neural networks based on LSTM cells. The high character recognition performances of these networks are almost systematically combined with linguistic knowledge, that is to say lexicon driven decoding method, to correct character misrecognitions. However with such high performance, we wonder on the possibility to use them without lexical decoding for word recognition. In this article, we explore this idea by proposing a lexicon verification strategy that provides a very low error rate, while conceding a consequent amount of rejects. Therefore, this verification approach perfectly fits in a cascade framework, where the rejects of a classifier are processed by the next cascade's classifier. The resulting system is nearly insensitive to the lexicon size, while providing a much faster decoding process than a standard lexicon driven decoding. Furthermore, when processing the final rejects of the cascade by a basic lexical decoding, our approach reach state of the art performance for isolated word recognition.
Bruno Stuner, Clément Chatelain 0001, Thierry Paquet
ICFHR3
2016 A Unified French/English Syllabic Model for Handwriting Recognition
abstract
In this paper we introduce a new unified syllabic model for French and English handwriting recognition, based on hidden Markov models (HMM). The recognition system training and recognition components such as optical models, lexicons and language models are designed to be language independent. In this purpose a syllable based model is proposed for French and English. This model is evaluated and compared to n-gram character and words models. A promising performance is achieved by the syllabic model, which meets the words model performance, with the advantage of a reduced system complexity. Furthermore, the unification of likely similar scripts improves the system performance over all models considering the English and French languages. The French RIMES and the English IAM datasets are used for the evaluation.
Wassim Swaileh, Julien Lerouge, Thierry Paquet
ICFHR3
2016 Cascading BLSTM networks for handwritten word recognition
abstract
Handwritten word recognition is a tough task, mixing image and natural language processing. Recently new recurrent neural networks with LSTM cells allowed significant improvements in this field. These networks are generally coupled with lexical and linguistic knowledge in order to correct character misrecognitions, namely using a lexicon driven decoding. Yet the high performances of LSTM networks let us think that there is a room to use them without lexical decoding. In this article we propose a lexicon-free decoding, combined with a lexicon verification method. This lexicon control method presents some interesting properties and enables us to efficiently combine LSTM networks in a cascade framework. This cascade process is not driven but simply controlled by the lexicon, allowing it to speed up the decoding while being nearly insensitive to the lexicon size. Our approach presents promising results with low error rate by conceding rejects. Those rejects can finally be processed by a standard lexical decoding, enabling us to reach state of the art performance, while being much faster than existing methods for decoding.
Bruno Stuner, Clément Chatelain 0001, Thierry Paquet
ICPR3
2015 Benchmarking discriminative approaches for word spotting in handwritten documents
abstract
In this article, we propose to benchmark the most popular methods for word spotting in handwritten documents. The benchmark includes a pure HMM approach, as well as hybrid discriminative methods MLP-HMM, CRF-HMM, RNN-HMM and BLSTM-CTC-HMM. This study enables us to observe the increase ratio of performance provided by each discriminative stage compared with the pure generative HMM approach. Moreover, we put forward the different abilities of all these discriminative stages from the simplest MLP to the most complex and current state of the art BLSTM-CTC. We also propose a more specific and original study on BLSTM-CTC, showing that when used as a lexicon-free recognizer, it can reach very interesting word-spotting performance.
Gautier Bideault, Luc Mioulet, Clément Chatelain 0001, Thierry Paquet
ICDAR4
2015 Unconstrained Bengali handwriting recognition with recurrent models
abstract
This paper presents a pioneering attempt for developing a recurrent neural net based connectionist system for unconstrained Bengali offline handwriting recognition. The major challenge in configuring such a classification system for a complex script like Bengali is to effectively define the character classes. A novel way of defining character classes is introduced making the recognition problem suitable for using a recurrent model. Indeed, it has to deal with more than nine hundred character classes for which the occurrence probability is very skewed in the language. An off-the-shelf BLSTM-CTC recognizer is used. An open-source dataset is developed for unconstrained Bengali offline handwriting recognition. The dataset contains 2,338 handwritten text lines consisting of about 21,000 word. Experiment shows that with the new definition of character classes the BLSTM-CTC provides an impressive performance for unconstrained Bengali offline handwriting recognition. The character level recognition accuracy is 75.40% without doing any post-processing on the BLSTM-CTC output. Among the 24.60% character level errors, the substitution, deletion and insertion errors are 18.91%, 4.69% and 0.98%, respectively.
Utpal Garain, Luc Mioulet, Bidyut B. Chaudhuri, Clément Chatelain 0001, Thierry Paquet
ICDAR5
2015 Keyword spotting in handwritten documents based on a generic text line HMM and a SVM verification
abstract
In this paper, we propose a novel system for keyword spotting in handwritten documents. Our approach proceeds in two steps: first a generic text line HMM provides a simple and flexible tool to localize the keyword and its character boundaries. In the second step, a SVM based verification system estimates and combines the character probabilities to provide keyword confidence scores which are further combined with the HMM score. The system has been evaluated on a public handwritten document database used for the 2011 ICDAR handwriting recognition competitions and shows that the verification stage improves the performance and outperforms some other state-of-the-art approaches.
Yousri Kessentini, Thierry Paquet
ICDAR2
2015 Language identification from handwritten documents
abstract
This paper presents a novel approach for language identification in handwritten documents. The approach is based on script identification followed by character recognition. BLSTM-CTC based handwriting recognizers are used and the OCR output is fed to a statistical language identifier for detecting the language of the input handwritten document. Documents in two scripts (Latin and Bengali) and four languages (English, French, Bengali and Assamese) are considered for evaluation. Several alternative frameworks have been explored, effects of handwriting recognition and text length on language detection have been studied. It is observed that with some empirical restrictions it is very much possible to achieve more that 80% language detection accuracy and based on the current research practical systems can be designed.
Luc Mioulet, Utpal Garain, Clément Chatelain 0001, Philippine Barlas, Thierry Paquet
ICDAR5
2015 OCR performance prediction using cross-OCR alignment
abstract
Since 2006 the national library of France (BnF) has developed many mass digitization projects on its collections. The indexation of digital documents on Gallica (the digital library of the BnF) is done through their textual content obtained thanks to service providers that use Optical Character Recognition software (OCR). The modern technologies of OCR achieve good performances on modern documents produced with uniform layout and known fonts. However, for old documents, OCR results are of lower quality. The OCR quality assessment is a real challenge for the BnF. On the one hand, due to the sequential architecture of OCR treatments, the identification of OCR errors sources is intractable. On the other hand, besides the word confidence, no additional quality information is reported in OCR outputs. In this paper, we present a study on OCR performance estimation aiming to control the quality of word transcriptions achieved by OCR. This quality assessment process has to operate without any comparison with ground truthed data. In this respect, our methodology relies on cross alignment of the OCR results with those of a secondary OCR called reference OCR. This secondary OCR provides uncertain but useful information that will be used as uncertain groundtruth. OCR performance is estimated using support vector regression. This predictor uses some global features computed on the cross-alignment results. The experimentations reported show that our estimate describes more faithfully the quality of OCR outputs than average word confidence scores that are computed by OCR. The proposed methodology can be adapted easily to various corpora by tuning the system using a training dataset of documents that have similar properties to those to be treated.
Ahmed Ben Salah, Jean-Philippe Moreux, Nicolas Ragot, Thierry Paquet
ICDAR4
2015 Multi-script iterative steerable directional filtering for handwritten text line extraction
abstract
In this paper, we introduce an iterative method for handwritten text line extraction. The proposed method improves the steerable filter approach by introducing an iterative scheme that iterates lines detection for various configurations of the filters, thus making the method auto adaptable with different types of scripts. We tested the method on different handwritten text datasets, used during earlier competitions organized at ICDAR or ICFHR, and using the same evaluation protocol. The tests carried out on the Open-Hart data set for Arabic scripts, as well as three other Latin and Greek scripts, show that state of the art performance are obtained without tuning any parameters.
Wassim Swaileh, Kamel Ait-Mohand, Thierry Paquet
ICDAR3
2015 A Hybrid BLSTM-HMM for Spotting Regular Expressions
Gautier Bideault, Luc Mioulet, Clément Chatelain 0001, Thierry Paquet
ICPRAM (2)4
2015 BLSTM-CTC Combination Strategies for Off-line Handwriting Recognition
Luc Mioulet, Gautier Bideault, Clément Chatelain 0001, Thierry Paquet, Stephan Brunessaux
ICPRAM (1)4
2015 A deep HMM model for multiple keywords spotting in handwritten documents
Simon Thomas 0002, Clément Chatelain 0001, Laurent Heutte, Thierry Paquet, Yousri Kessentini
Pattern Anal. Appl.4
2015 A Dempster-Shafer Theory based combination of handwriting recognition systems with multiple rejection strategies
Yousri Kessentini, Thomas Burger, Thierry Paquet
Pattern Recognit.3
2014 A Typed and Handwritten Text Block Segmentation System for Heterogeneous and Complex Documents
abstract
This paper presents a Document Image Analysis (DIA) system able to extract homogeneous typed and handwritten text regions from complex layout documents of various types. The method is based on two connected component classification stages that successively discriminate text/non text and typed/handwritten shapes, followed by an original block segmentation method based on white rectangles detection. We present the results obtained by the system during the first competition round of the MAURDOR campaign.
Philippine Barlas, Sébastien Adam, Clément Chatelain 0001, Thierry Paquet
Document Analysis Systems4
2014 OCR Performance Prediction Using a Bag of Allographs and Support Vector Regression
abstract
In this paper, we describe a novel and simple technique for prediction of OCR results without using any OCR. The technique uses a bag of allographs to characterize textual components. Then a support vector regression (SVR) technique is used to build a predictor based on the bag of allographs. The performance of the system is evaluated on a corpus of historical documents. The proposed technique produces correct prediction of OCR results on training and test documents within the range of standard deviation of 4.18% and 6.54% respectively. The proposed system has been designed as a tool to assist selection of corpora in libraries and specify the typical performance that can be expected on the selection.
Tapan Kumar Bhowmik, Thierry Paquet, Nicolas Ragot
Document Analysis Systems2
2014 Writing Type and Language Identification in Heterogeneous and Complex Documents
abstract
This paper presents a system dedicated to automatic recognition of both the writing type and the language of text regions in heterogeneous and complex documents. This system is able to process documents with mixed printed and handwritten text, in various languages (French, English and Arabic). To handle such a problem, we divided it into two sub-tasks: The writing type identification and the language identification. The method for the writing type recognition is based on the analysis of the connected components while the language identification approach combines the analysis of connected components and the analysis of character distributions. We present the results obtained by the system during the second competition round of the MAURDOR campaign, and show that the performance of our system compares favorably with other participants.
David Hebert, Philippine Barlas, Clément Chatelain 0001, Sébastien Adam, Thierry Paquet
ICFHR5
2014 Combining Structure and Parameter Adaptation of HMMs for Printed Text Recognition
abstract
We present two algorithms that extend existing HMM parameter adaptation algorithms (MAP and MLLR) by adapting the HMM structure. This improvement relies on a smart combination of MAP and MLLR with a structure optimization procedure. Our algorithms are semi-supervised: to adapt a given HMM model on new data, they require little labeled data for parameter adaptation and a moderate amount of unlabeled data to estimate the criteria used for HMM structure optimization. Structure optimization is based on state splitting and state merging operations and proceeds so as to optimize either the likelihood or a heuristic criterion. Our algorithms are successfully applied to the recognition of printed characters by adapting the HMM character models of a polyfont printed text recognizer to new fonts. Our experiments involve a total of 1,120,000 real and 3,100,000 synthetic character images and concern a set of 89 HMM models. A comparison of our results with those of state-of-the-art adaptation algorithms (MAP and MLLR) shows a significant increase in the accuracy of character recognition.
Kamel Ait-Mohand, Thierry Paquet, Nicolas Ragot
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Discrete CRF Based Combination Framework for Document Image Binarization
abstract
Document image binarization is still an active research area as shows the number of binarization techniques proposed since many decades. The binarization of degraded document images is still difficult and encourages the development of new algorithms. For the last decade, discrete conditional random fields have been successfully used for many domains such as automatic language analysis. In this paper, we propose a CRF based framework to explore the combination capabilities of this model by combining discrete outputs from several well known binarization algorithms. The framework uses two 1D CRF models on the horizontal and the vertical directions that are coupled for each pixel by the product of the marginal probabilities computed from the both models. Experiments are made on two datasets from the Document Image Binarization Contest (DIBCO) 2009 and 2011 and show best performances than most of the methods presented at DIBCO 2011.
David Hebert, Stéphane Nicolas, Thierry Paquet
ICDAR3
2013 Learning to Detect Tables in Scanned Document Images Using Line Information
abstract
This paper presents a method to detect table regions in document images by identifying the column and row line-separators and their properties. The method employs a run-length approach to identify the horizontal and vertical lines present in the input image. From each group of intersecting horizontal and vertical lines, a set of 26 low-level features are extracted and an SVM classifier is used to test if it belongs to a table or not. The performance of the method is evaluated on a heterogeneous corpus of French, English and Arabic documents that contain various types of table structures and compared with that of the Tesseract OCR system.
Thotreingam Kasar, Philippine Barlas, Sébastien Adam, Clément Chatelain 0001, Thierry Paquet
ICDAR5
2013 Word Spotting and Regular Expression Detection in Handwritten Documents
abstract
In this paper, we propose a novel system for word spotting and regular expression detection in Handwritten documents. The proposed approach is lexicon-free, i.e., able to spot arbitrary keywords that are not required to be known at the training stage. Furthermore, the proposed system is segmentation-free, i.e., text lines are not required to be segmented into words. The originalities of our approach is twofold. First we propose a new filler model which allows to speed-up the decoding process. Second, we extend the methodology to search for regular expressions. The system has been evaluated on a public handwritten document database used for the 2011 ICDAR handwriting recognition competitions.
Yousri Kessentini, Clément Chatelain 0001, Thierry Paquet
ICDAR3
2012 Logical segmentation for article extraction in digitized old newspapers
abstract
Newspapers are documents made of news item and informative articles. They are not meant to be read iteratively: the reader can pick his items in any order he fancies. Ignoring this structural property, most digitized newspaper archives only offer access by issue or at best by page to their content. We have built a digitization workflow that automatically extracts newspaper articles from images, which allows indexing and retrieval of information at the article level. Our back-end system extracts the logical structure of the page to produce the informative units: the articles. Each image is labelled at the pixel level, through a machine learning based method, then the page logical structure is constructed up from there by the detection of structuring entities such as horizontal and vertical separators, titles and text lines. This logical structure is stored in a METS wrapper associated to the ALTO file produced by the system including the OCRed text. Our front-end system provides a web high definition visualisation of images, textual indexing and retrieval facilities, searching and reading at the article level. Articles transcriptions can be collaboratively corrected, which as a consequence allows for better indexing. We are currently testing our system on the archives of the Journal de Rouen, one of France eldest local newspaper. These 250 years of publication amount to 300 000 pages of very variable image quality and layout complexity. Test year 1808 can be consulted at plair.univ-rouen.fr.
Thomas Palfray, David Hebert, Stéphane Nicolas, Pierrick Tranouez, Thierry Paquet
ACM Symposium on Document Engineering5
2012 A categorization system for handwritten documents
Thierry Paquet, Laurent Heutte, Guillaume Koch, Clément Chatelain 0001
Int. J. Document Anal. Recognit.1
2011 Constructing Dynamic Frames of Discernment in Cases of Large Number of Classes
Yousri Kessentini, Thomas Burger, Thierry Paquet
ECSQARU3
2011 Dempster-Shafer Based Rejection Strategy for Handwritten Word Recognition
abstract
In this paper, a novel rejection strategy is proposed to optimize the reliability of an handwritten word recognition system. The proposed approach is based on several steps. First, we combine the outputs of several HMM classifiers using the Dempster-Shafer theory (DST). Then, we take advantage of the expressivity of mass functions (the counter part of probability distributions in DST) to characterize the quality/reliability of the classification. Finally, we use this characterization to decide whether a test word is rejected or not. Experiments carried out on RIMES and IFN/ENIT datasets show that the proposed approach outperforms other state-of-the-art rejection methods.
Thomas Burger, Yousri Kessentini, Thierry Paquet
ICDAR3
2011 Continuous CRF with Multi-scale Quantization Feature Functions Application to Structure Extraction in Old Newspaper
abstract
We introduce quantization feature functions to represent continuous or large range discrete data into the symbolic CRF data representation. We show that doing this convertion in a simple way allows the CRF to automaticaly select discriminative features to achieve best performance. This system is evaluated on a segmentation task of degraded newspapers archives. The results obtained show the ability of the CRF model to deal with numerical features similarly as for symbolic representation thanks to the use of quantization feature functions. The segmentation task is achieved by the definition of a horizontal CRF model dedicated to pixel labelling.
David Hebert, Thierry Paquet, Stéphane Nicolas
ICDAR2
2011 An Optimized Multi-stream Decoding Algorithm for Handwritten Word Recognition
abstract
This paper is focused on the optimization of the computational efficiency of a multi-stream word recognition system. The aim of this work is to optimize the multi-stream decoding step in order to reduce the recognition time and the complexity to allow combining a large number of streams. Two different multi-stream decoding strategies are compared based on two-level and HMM-recombination algorithms. Experiments carried out on public handwritten word databases show significant speed gains at decoding while keeping the same performances, in addition to new insights for combining a large number of streams.
Yousri Kessentini, Thierry Paquet, Ahmed Guermazi
ICDAR2
2010 Dealing with Precise and Imprecise Decisions with a Dempster-Shafer Theory Based Algorithm in the Context of Handwritten Word Recognition
abstract
The classification process in handwriting recognition is designed to provide lists of results rather than single results, so that context models can be used as post-processing. Most of the time, the length of the list is determined once and for all the items to classify. Here, we present a method based on Dempster-Shafer theory that allows a different length list for each item, depending on the precision of the information involved in the decision process. As it is difficult to compare the results of such an algorithm to classical accuracy rates, we also propose a generic evaluation methodology. Finally, this algorithm is evaluated on Latin and Arabic handwritten isolated word datasets.
Thomas Burger, Yousri Kessentini, Thierry Paquet
ICFHR3
2010 Alpha-Numerical Sequences Extraction in Handwritten Documents
abstract
In this paper, we introduce an alpha-numerical sequences extraction system (keywords, numerical fields or alpha-numerical sequences) in unconstrained handwritten documents. Contrary to most of the approaches presented in the literature, our system relies on a global handwriting line model describing two kinds of information : i) the relevant information and ii) the irrelevant information represented by a shallow parsing model. The shallow parsing of isolated text lines allows quick information extraction in any document while rejecting at the same time irrelevant information. Results on a public french incoming mails database show the efficiency of the approach.
Simon Thomas 0002, Clément Chatelain 0001, Laurent Heutte, Thierry Paquet
ICFHR4
2010 Structure Adaptation of HMM Applied to OCR
abstract
In this paper we present a new algorithm for the adaptation of Hidden Markov Models (HMM models). The principle of our iterative adaptive algorithm is to alternate an HMM structure adaptation stage with an HMM Gaussian MAP adaptation stage of the parameters. This algorithm is applied to the recognition of printed characters to adapt the character models of a poly font general purpose character recognizer to new fonts of characters, never seen during training. A comparison of the results with those of MAP classical adaptation scheme show a slight increase in the recognition performance.
Kamel Ait-Mohand, Thierry Paquet, Nicolas Ragot, Laurent Heutte
ICPR2
2010 An Information Extraction Model for Unconstrained Handwritten Documents
abstract
In this paper, a new information extraction system by statistical shallow parsing in unconstrained handwritten documents is introduced. Unlike classical approaches found in the literature as keyword spotting or full document recognition, our approach relies on a strong and powerful global handwriting model. A entire text line is considered as an indivisible entity and is modeled with Hidden Markov Models. In this way, text line shallow parsing allows fast extraction of the relevant information in any document while rejecting at the same time irrelevant information. First results are promising and show the interest of the approach.
Simon Thomas 0002, Clément Chatelain 0001, Laurent Heutte, Thierry Paquet
ICPR4
2010 Evidential Combination of Multiple HMM Classifiers for Multi-script Handwritting Recognition
Yousri Kessentini, Thomas Burger, Thierry Paquet
IPMU3
2010 A multi-model selection framework for unknown and/or evolutive misclassification cost problems
Clément Chatelain 0001, Sébastien Adam, Yves Lecourtier, Laurent Heutte, Thierry Paquet
Pattern Recognit.5
2010 Off-line handwritten word recognition using multi-stream hidden Markov models
Yousri Kessentini, Thierry Paquet, Abdelmajid Ben Hamadou
Pattern Recognit. Lett.2
2009 Off-Line Multi-Script Writer Identification Using AR Coefficients
abstract
The problem of writer identification in a multi-script environment is attempted using a two-dimensional (2D) autoregressive (AR) modeling technique. Each writer is represented by a set of 2D AR model coefficients. A method to estimate AR model coefficients is proposed. This method is applied to an image of text written by a specific writer so that AR coefficients are obtained to characterize the writer. For a given sample, AR coefficients are computed and its L2distance with each of the stored (writer) prototypes identifies the writer for the sample. The method has been tested on datasets of two different scripts, namely RIMES containing 382 French writers and ISI consisting of samples from 40 Bengali writers. Modeling of writing styles using different context patterns at different image resolution has been investigated. Experimental results show that the technique achieves results comparable with that of the previous approaches.
Utpal Garain, Thierry Paquet
ICDAR2
2009 A Multi-Lingual Recognition System for Arabic and Latin Handwriting
abstract
Generally, handwritten word recognition systems use script specific methodologies. In this paper, we present a unified approach for multi-lingual recognition of alphabetic scripts. The proposed system operates independently of the nature of the script using the multi-stream paradigm. The experiments have been carried out on a multi-script database composed of Arabic and Latin handwritten words from the IFN/ENIT and the IRONOFF public databases and show interesting recognition performances with only 1.5% of script confusion and an overall word recognition rate of 84.5% using a multi-script lexicon of 1142 words.
Yousri Kessentini, Thierry Paquet, Abdelmajid Ben Hamadou
ICDAR2
2008 Multi-script handwriting recognition with N-streams low level features
abstract
The multi-stream paradigm provides an interesting framework for the integration of multiple sources of information. In this paper, we present our multi-script recognition system using the multi-stream formalism to combine low level feature streams. We Analyze how the combination of n streams (n=2,....,4) can improve the recognition performance. Significant experiments have been carried out on two publicly available word databases: IFN/ENIT benchmark database (Arabic script) and IRONOFF database (Latin script). The proposed framework shows interesting results in both cases thanks to the use of low level features in a segmentation free approach which ensure its applicability to various scripts.
Yousri Kessentini, Thierry Paquet, Abdelmajid Ben Hamadou
ICPR2
2007 Multi-Objective Optimization for SVM Model Selection
abstract
In this paper, we propose a multi-objective optimization method for SVM model selection using the well known NSGA-II algorithm. FA and FR rates are the two criteria used to find the optimal hyperparameters of a set of SVM classifiers. The proposed strategy is applied to a digit/outlier discrimination task embedded in a more global information extraction system that aims at locating and recognizing numerical fields in handwritten incoming mail documents. Experiments conducted on a large database of digits and outliers show clearly that our method compares favorably with the results obtained by a state-of-the- art mono-objective optimization technique using the classical Area Under ROC Curve criterion (AUC).
Clément Chatelain 0001, Sébastien Adam, Yves Lecourtier, Laurent Heutte, Thierry Paquet
ICDAR5
2007 Machine Dating of Handwritten Manuscripts
abstract
This paper presents a pioneering study on automatic dating of handwritten manuscripts. Analysis of handwriting style forms the core of the dating method. Initially, it is hypothesized that a manuscript can be dated, to a certain level of accuracy, by looking at the way it is written. The hypothesis is then verified with real samples of known dates. A general framework is proposed for machine dating of handwritten manuscripts. Experiments on a database containing manuscripts of Gustave Flaubert (1821- 1880), the famous French novelist reports about 62% accuracy when manuscripts are dated within a range of five calendar years with respect to their exact year of writing.
Utpal Garain, Swapan K. Parui, Thierry Paquet, Laurent Heutte
ICDAR3
2007 A Multi-stream Approach to Off-Line Handwritten Word Recognition
abstract
We present in this paper a new approach based on multi-stream hidden Markov models (HMM) for the recognition of off-line handwriting. Every word is presented by two HMM models: the first one is learned with features extracted from upper contour, the second with features extracted from lower contour. The combination of these two sources of information is studied using the multi-stream framework. We present experiment results obtained on a database composed of isolated words extracted from incoming mail documents.
Yousri Kessentini, Thierry Paquet, Abdelmajid Ben Hamadou
ICDAR2
2007 Document Image Segmentation Using a 2D Conditional Random Field Model
abstract
This work relates to the implementation of a 2D conditional random field model in the context of document image analysis. Our model makes it possible to take variability into account and to integrate contextual knowledge, while taking benefit from machine learning techniques. Experiments on handwritten drafts of Flaubert show that these models provide interesting solutions.
Stéphane Nicolas, J. Dardenne, Thierry Paquet, Laurent Heutte
ICDAR3
2007 Detection and recognition of erasures in on-line captured paper forms
Alain Wiart, Thierry Paquet, Laurent Heutte
Pattern Recognit. Lett.2
2006 Segmentation-Driven Recognition Applied to Numerical Field Extraction from Handwritten Incoming Mail Documents
Clément Chatelain 0001, Laurent Heutte, Thierry Paquet
Document Analysis Systems3
2006 A Multi-Agent Model and Tabu Search Optimization to Manage Agricultural Territories
Wassim Jaziri, Thierry Paquet
GeoInformatica2
2006 On foreground - background separation in low quality document images
Utpal Garain, Thierry Paquet, Laurent Heutte
Int. J. Document Anal. Recognit.2
2005 On Foreground-Background Separation in Low Quality Color Document Images
abstract
This paper proposes an adaptive method for separation of foreground and background in low quality color document images. A connected component labelling is initially implemented to capture the spatially connected similar color pixels. Next, dominant background components are determined to divide the entire image into number of grids each representing local uniformity in illumination, background, etc. Finally foreground parts are located using local information around them. Several color images of old historical documents including manuscripts of high importance are used in the experiment. Apart from a qualitative evaluation, results are quantitatively compared with one popular foreground/background separation technique.
Utpal Garain, Thierry Paquet, Laurent Heutte
ICDAR2
2005 Handwritten Document Segmentation Using Hidden Markov Random Fields
abstract
In this paper we present a method based on hidden Markov random fields and 2D dynamic programming image decoding, for segmenting pages of complex handwritten manuscripts such as novelist drafts. After a formal description of the theoretical framework and the principles of the decoding method, we describe the implementation of the model and the decoding method. Then we discuss the results obtained with this approach on the drafts of the French novelist Gustave Flaubert.
Stéphane Nicolas, Yousri Kessentini, Thierry Paquet, Laurent Heutte
ICDAR3
2005 A writer identification and verification system
Ameur Bensefia, Thierry Paquet, Laurent Heutte
Pattern Recognit. Lett.2
2005 Automatic extraction of numerical sequences in handwritten incoming mail documents
Guillaume Koch, Laurent Heutte, Thierry Paquet
Pattern Recognit. Lett.3
2004 Enriching Historical Manuscripts: The Bovary Project
Stéphane Nicolas, Thierry Paquet, Laurent Heutte
Document Analysis Systems2
2004 A multiple agent architecture for handwritten text recognition
Laurent Heutte, Ali Nosary, Thierry Paquet
Pattern Recognit.3
2004 Unsupervised writer adaptation applied to handwritten text recognition
Ali Nosary, Laurent Heutte, Thierry Paquet
Pattern Recognit.3
2003 Digitizing cultural heritage manuscripts: the Bovary project
abstract
In this paper we describe the Bovary Project, a manuscripts digitization project of the famous French writer Gustave FLAUBERT first great work. This project has just begun at the end of 2002 and should end in 2006 by providing an online access to an hypertextual edition of "Madame Bovary" drafts set. We develop the global context of this project, the main objectives, the first studies and the considered outlooks for the project's carried out.
Stéphane Nicolas, Thierry Paquet, Laurent Heutte
ACM Symposium on Document Engineering2
2003 Information Retrieval Based Writer Identification
abstract
This communication deals with the Writer Identificationtask. Our previous work has shown the interest of usingthe graphemes as features for describing the individualproperties of Handwriting. We propose here to exploit thesame feature set but using an information retrievalparadigm to describe and compare the handwritten queryto each sample of handwriting in the database. Using thistechnique the image processing stage is performed onlyonce and before the retrieval process can take place, thusleading to a significant saving in the computation of eachquery response, compared to our initial proposition. Themethod has been tested on two handwritten databases.The first one has been collected from 88 different writersat PSI Lab. while the second one contains 39 writers fromthe original correspondence of Emile Zola, a famousFrench novelist of the last 19th century. We also analyzethe proposed method when using concatenation ofgraphemes (bi and tri-gramme) as features.
Ameur Bensefia, Thierry Paquet, Laurent Heutte
ICDAR2
2003 Numerical Sequence Extraction in Handwritten Incoming Mail Documents
abstract
In this communication, we propose a method for the automatic extraction of numerical fields in handwritten documents. The approach exploits the known syntactic structure of the numerical field to extract, combined with a set of contextual morphological features to find the best label to each connected component. Applying an HMM based syntactic analyzer on the overall document allows to localize/extract fields of interest. Reported results on the extraction of zip codes, phone numbers and customer codes from handwritten incoming mail documents demonstrate the interest of the proposed approach.
Guillaume Koch, Laurent Heutte, Thierry Paquet
ICDAR3
1999 Defining Writer's Invariants to Adapt the Recognition Task
abstract
Investigates the automatic reading of unconstrained omni-writer handwritten texts. This paper shows how to endow the reading system with adaptation faculties for each writer's handwriting. The adaptation principles are of major importance for making robust decisions when neither simple lexical nor syntactic rules can be used, e.g. for a free lexicon or for full text recognition. The first part of this paper defines the concept of writer's invariants. In the second part, we explain how the recognition system can be adapted to a particular handwriting by exploiting the graphical context defined by the writer's invariants. This adaptation is guaranteed, thanks to the writer's invariants, by activating interaction links over the whole text between the recognition procedures for word entities and those for letter entities.
Ali Nosary, Laurent Heutte, Thierry Paquet, Yves Lecourtier
ICDAR3
1998 A structural/statistical feature based vector for handwritten character recognition
Laurent Heutte, Thierry Paquet, Jean-Vincent Moreau, Yves Lecourtier, Christian Olivier
Pattern Recognit. Lett.2
1997 Optimal Order of Markov Models Applied to Bankchecks
abstract
The aim of this study is to show that the optimal order of Markov Model of cursive words can be rigorously stated in order to fit the structural properties of the observed data using Akaike information criterion. The method has been tested on French Postal check amounts up to order 4. An original structural representation of cursive words based on graphemes is used. The conditional probability to have a word model given an observed sequence of graphemes is computed independently of the length of the sequence. The recognition results obtained confirm the optimal order found using Akaike criterion.
Christian Olivier, Thierry Paquet, Manuel Avila, Yves Lecourtier
Int. J. Pattern Recognit. Artif. Intell.2
1996 Combining structural and statistical features for the recognition of handwritten characters
abstract
The authors present a feature vector for the recognition of handwritten characters which combines the strengths of both statistical and structural feature extractors. Thanks to a combination of seven complementary families of features (ranging from pure structural to pure statistical and including both local and global features), a complete description of the characters can be achieved thus providing a wide range of identification clues. The recognition system has been tested on three categories of handwritten characters: handwritten well-segmented digits extracted from the NIST Database, uppercase letters collected from US dead letter envelopes and graphemes generated by a handwritten cursive word segmentation performed on US address word images. We thus demonstrate in this paper that high recognition rates with very low substitution rates can he achieved by means of the same general-purpose structural/statistical feature based vector.
Laurent Heutte, Jean-Vincent Moreau, Thierry Paquet, Yves Lecourtier, Christian Olivier
ICPR3
1995 Evaluation of codes and primitives: recognition of unconstrained handwritten numerals
abstract
In this paper, we present a mathematical model for evaluating codes and primitives in optical character recognition. The model is based on code efficiency, calculated from its average length and its transmitted information. This efficiency is obtained from entropies and conditional entropies which are estimated from probabilities of character recognition depending on the issued codes. This method is used to evaluate two sets of primitives in numeral recognition. Then, we propose a method of constructing binary decision trees for the recognition of handwritten numerals. This method is based on the mathematical model previously stated, which is used to process the transmitted information about the primitives. We demonstrate the performance of our system with experiments using real data.
N. Feray, Denis de Brucq, Katerin Romeo-Pakker, Thierry Paquet
ICDAR4
1995 Recognition of handwritten words using stochastic models
abstract
The paper deals with the global recognition of a small lexicon of words, based on a pseudo segmentation stage introducing anchor points. We avoid the difficult problem of segmentating the word into letters and the complexity involved by such models to build possible letter graphs. We use two structural representations of the word, strokes and graphemes, each of them being analyzed using a Markov model. These simple models are individually optimized by a rigorous choice of the order for fitting the structural properties of the observed data using Akaike information criteria. The conditional probability to have a word model, given the observation sequence, is computed by taking into account the length of the sequence. Results of the study are presented on French cheque images.
Christian Olivier, Thierry Paquet, Manuel Avila, Yves Lecourtier
ICDAR2
1993 Automatic reading of the literal amount of bank checks
Thierry Paquet, Yves Lecourtier
Mach. Vis. Appl.1
1993 Recognition of handwritten sentences using a restricted lexicon
Thierry Paquet, Yves Lecourtier
Pattern Recognit.1