VLDB 2026 Research / reviewers in the wild / expert
Nicolas Ragot
dblp:31/2380
· DBLP profile ↗
32ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0003-2321-942XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 15 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating Robustness of 3D Gaussian Splatting-Based 6D Camera Pose Refinement Under Degraded Conditions for Lightly Textured Industrial Synthetic ObjectsabstractIn this paper, 6D camera pose refinement is explored using 3D Gaussian Splatting (3DGS) on lightly textured industrial object datasets. The study employs datasets generated with Unity 3D rendering software, featuring objects such as a bicycle, MiR robot, Tiago robot, and UR robotic arm, each captured with ground-truth intrinsic and extrinsic camera parameters. A 3DGS model is trained to represent each scene, and 6D pose refinement is evaluated using a recent pose optimization approach, iComMa (inverting 3DGS via Comparing and Matching), which aligns rendered and query images. Experiments utilize 20% of the images for testing and 80% for training the 3DGS models. Camera poses are initialized with varying degrees of perturbation (δ) in both rotation and translation to assess the refinement capabilities. Robustness is further evaluated under degraded conditions by applying various types of noise to the query images, including Gaussian noise, salt-and-pepper noise, dilation, and erosion. Results demonstrate reliable pose refinement under photometric noise; however, with structural noise, the method maintains good rotation accuracy but struggles with translation due to changes in geometric features. This approach shows promise for industrial applications, where 3DGS models trained on synthetic datasets can refine camera poses in real-world industrial environments with common noise characteristics. Sunil Choudhary, Nicolas Ragot, Vincent Havard, Yohan Dupuis |
IECON | 2 |
| 2025 | Synthetic Data-Driven Augmentation for Precise 6-DoF Pose Estimation of Building Components in Automated Facility InspectionsabstractThis paper tackles the challenge of automating facility inspections by detecting building components, estimating their six-degree-of-freedom (6-DoF) poses (position and orientation), and comparing these estimations to Building Information Modeling (BIM) ground truth data. Vision based Deep learning methods offer promising results in pose estimation. They rely heavily on large annotated image datasets for training, which are often lacking or too specific for building settings. Manually annotating large datasets is both labor-intensive and time-consuming, particularly when they contain hundreds of images. This challenge further underscores the need for efficient and automated methods for dataset annotation, particularly in building environments where dataset size and complexity are significant. We present a study examining the point where the performance of a 6-DoF object detection algorithm is efficient. In particular, we analyze the proportion of real data required for detection and pose estimation models to perform well on double-leaf doors. Synthetic data are generated using the Digital Twin (DT), which is intended to serve as the future reference for BIM in facility inspections. Adding synthetic data reduces rotation error by up to 1.7°, starting from 1,000 images. We analyze the effectiveness of different proportions of real and synthetic data to provide insight into optimizing dataset composition. The accuracy of the 3D center point and the accuracy at 25 cm/25 degrees are better than 90% accuracy when more than 50 real images combined with synthetic data. Aristide Laignel, Nicolas Ragot, Fabrice Duval |
IECON | 2 |
| 2025 | Improving Image-Based Tool Detection in Industrial Workstations using Data AugmentationabstractWithin the framework of Industry 5.0, affordances enable intuitive and adaptive interactions between operators and their industrial work environments. Accurately perceiving these affordances enhances overall production performance, safety, and operator effectiveness. This paper focuses on the initial step of a larger affordance characterization pipeline: detecting tools used by operators during manual assembly tasks. To address the challenges of data scarcity and annotation effort in industrial contexts, we train a custom YOLOv9-based deep learning model on a data-augmented dataset combining real-world and synthetic images, automatically generated from a digital model of an industrial workstation in Unity3D. Through extensive experiments, we varied dataset sizes (50–300 images) and real-world data proportions in the data-augmented train datasets (0%–50%), to assess their impact on tool detection. Results show that only 10% of real-world data is sufficient to achieve strong performance across all data-augmented dataset sizes. A tool specific analysis reveals that visual characteristics such as size and shape influence detection. These findings highlight the effectiveness of combining synthetic and real data to reduce annotation effort while supporting robust tool detection for affordance characterization. Sarah Ouarab, Nicolas Ragot, Yohan Dupuis |
IECON | 3 |
| 2025 | The Augmented Perception: An emerging approach towards resilient manufacturing systems involving robotic agents and digital twin
Yassine Feddoul, Nicolas Ragot, Fabrice Duval, Vincent Havard, David Baudry |
Adv. Eng. Informatics | 2 |
| 2024 | ConvNeXt based semi-supervised approach with consistency regularization for weeds classificationabstractWeed recognition is an essential step for automatic weed control systems. Identifying weeds enables targeted control measures to be implemented, minimizing the use of chemicals and reducing the impact on the environment. Deep learning-based approaches proved to be effective for addressing various complex classification problems. However, to benefit fully from their capabilities, large amounts of labeled data are required, which represents a limitation for agricultural applications, consequence of the tedious and time-consuming process of data labeling. Conversely, unlabeled data could be acquired in large quantities, with relative ease. Hence, our aim is to develop robust and precise deep learning models, to carry-out the recognition and identification of weed species, using both types of data. To this end, we propose a method, that adopts the semi-supervised learning paradigm, to optimally combine labeled and unlabeled data. The method is based on a new deep neural networks architecture, which consists of a modernized convolutional encoder belonging to the family ConvNeXt and a thoroughly designed deep decoder network. This architecture, enables a successful integration of consistency regularization. The conducted experiments on DeepWeeds and 4-Weeds, showed that the semi-supervised models trained through our proposed method provide a stable and high classification performance, compared to other state-of-the-art deep learning models, which were affected negatively by the amount of labeled data available, and the presence of noise during inference. Furthermore, the effectiveness of the proposed method was demonstrated in comparison to other semi-supervised learning methods. The results obtained demonstrate the benefits of adopting the semi-supervised learning paradigm, especially in scenarios with very limited labeled data. Farouq Benchallal, Adel Hafiane, Nicolas Ragot, Raphaël Canals |
Expert Syst. Appl. | 3 |
| 2022 | Efficient Dynamic Texture Classification with Probabilistic MotifsabstractWe propose to tackle dynamic texture video classification as a pattern mining problem. In a nutshell, videos are represented by frequent sequences of representative patches. Firstly, we use a Gaussian Mixture Model to make the clustering of patches from training videos. Secondly, a soft assignment is used as an encoding method to construct sequences of probability vectors (p-sequences) representing sequences of spatio-temporal patches. Thirdly, for each class, we mine meaningful motifs appearing inside the training p-sequences by means of an adapted data mining approach. Finally, feature vectors are constructed from the mined motifs, using the probabilistic support, which quantifies the match between the p-sequences, of the video to be classified, and the key-motifs of the classes. Experimental results and analysis for dynamic texture classification on benchmark datasets (i.e. UCLA, Traffic) show the interest of the proposed method. Luong Phat Nguyen, Julien Mille, Dominique Li, Donatello Conte, Nicolas Ragot |
ICPR | 5 |
| 2022 | View Selection for Industrial Object RecognitionabstractThe last industrial revolutions and the digital transformation have led to a rise of robotics and to the emergence of the concept of digital twin. A major challenge falls within the update of this virtual representation, so that the supervision operator and the system itself can take appropriate decisions. One way to achieve that is to take advantage of the multi-robot perception capabilities by merging their individual observations to collectively enhance object recognition and robot environmental understanding. Since object recognition strongly depends on the viewing angles, one challenge deals with identifying the most relevant camera poses containing the most relevant information about the nature of the object. In this paper we propose a smart view selection approach which aims at determining the poses of the cameras and the number of the most informative views while maximising the object recognition. Based on a synthetic view dataset of traditional industrial objects, we adopt a clustering-based approach for maximising the inter-class distance and minimising the intra-class one. To do so, we compute a score for each view based on the Fowlkes-Mallows Index. This leads us to order the dataset and select a subset of views maximising the score. Then, this subset is used as a training dataset for a knn-classifier. The results, presented in terms of F1-score metric, are promising and highlight the relevance of our work: i) our smart selection enables the collection of a limited number of the most informative camera poses for object recognition; ii) feature extraction from a pre-trained CNN combined with a clustering algorithm allows the separability of industrial object categories; iii) our approach is robust since it provides good performances while the camera poses are in the neighbourhood of the exact camera positions provided by our processing pipeline. Kewei Xu, Nicolas Ragot, Yohan Dupuis |
IECON | 2 |
| 2021 | Deep Learning for Document Layout Generation: A First Reproducible Quantitative Evaluation and a Baseline Model
Romain Carletto, Hubert Cardot, Nicolas Ragot |
ICDAR (3) | 3 |
| 2018 | Comparative study of conventional time series matching techniques for word spotting
Tanmoy Mondal, Nicolas Ragot, Jean-Yves Ramel, Umapada Pal 0001 |
Pattern Recognit. | 2 |
| 2017 | A multi-one-class dynamic classifier for adaptive digitization of document streams
Anh Khoi Ngo Ho, Véronique Eglin, Nicolas Ragot, Jean-Yves Ramel |
Int. J. Document Anal. Recognit. | 3 |
| 2016 | Interactive Definition and Tuning of One-Class Classifiers for Document Image ClassificationabstractWith mass of data, document image classification systems have to face new trends like being able to process heterogeneous data streams efficiently. Generally, when processing data streams, few knowledge is available about the content of the possible streams. Furthermore, as getting labelled data is costly, the classification model has to be learned from few available labelled examples. To handle such specific context, we think that combining one-class classifiers could be a very interesting alternative to quickly define and tune classification systems dedicated to different document streams. The main interest of one-class classifiers is that no interdependence occurs between each classifier model allowing easy removal, addition or modification of classes of documents. Such reconfiguration will not have any impact on the other classifiers. It is also noticeable that each classifier can use a different set of features compared to the other to handle the same class or even different classes. In return, as only one class is well-specified during the learning step, one-class classifiers have to be defined carefully to obtain good performances. It is more difficult to select the representative training examples and the discriminative features with only positive examples. To overcome these difficulties, we have defined a complete framework offering different methods that can help a system designer to define and tune one-class classifier models. The aims are to make easier the selection of good training examples and of suitable features depending on the class to recognize into the document stream. For that purpose, the proposed methods compute different measures to evaluate the relevance of the available features and training examples. Moreover, a visualization of the decision space according to selected examples and features is proposed to help such a choice and, an automatic tuning is proposed for the parameters of the models according to the class to recognize when a validation stream is available. The pertinence of the proposed framework is illustrated on two different use cases (a real data stream and a public data set). Nathalie Girard, Roger Trullo, Sabine Barrat, Nicolas Ragot, Jean-Yves Ramel |
DAS | 4 |
| 2016 | Text Extraction in Document Images: Highlight on Using Corner PointsabstractDuring past years, text extraction in document images has been widely studied in the general context of Document Image Analysis (DIA) and especially in the framework of layout analysis. Many existing techniques rely on complex processes based on preprocessing, image transforms or component/edges extraction and their analysis. At the same time, text extraction inside videos has received an increased interest and the use of corner or key points has been proven to be very effective. Because it is noteworthy to notice that very few studies were performed on the use of corner points for text extraction in document images, we propose in this paper to evaluate the possibilities associated with this kind of approach for DIA. To do that, we designed a very simple technique based on FAST key points. A first stage divide the image into blocks and the density of points inside each one is computed. The more dense ones are kept as text blocks. Then, connectivity of blocks is checked to group them and to obtain complete text blocks. This technique has been evaluated on different kind of images: different languages (Telugu, Arabic, French), handwritten as well as typewritten, skewed documents, images at different resolution and with different kind and amount of noises (deformations, ink dot, bleed through, acquisition (blur, resolution)), etc. Even with fixed parameters for all such kind of documents images, the precision and recall are close or higher to 90% which makes this basic method already effective. Consequently, even if the proposed approach does not propose a breakthrough from theoretical aspects, it highlights that accurate text extraction could be achieved without complex approach. Moreover, this approach could also be easily improved to be more precise, robust and useful for more complex layout analysis. Vikas Yadav, Nicolas Ragot |
DAS | 2 |
| 2016 | Flexible Sequence Matching technique: An effective learning-free approach for word spotting
Tanmoy Mondal, Nicolas Ragot, Jean-Yves Ramel, Umapada Pal 0001 |
Pattern Recognit. | 2 |
| 2015 | Performance evaluation of DTW and its variants for word spotting in degraded documentsabstractIn word spotting literature, classical DTW has been widely employed. However there exists several other improved versions of DTW along with other robust sequence matching techniques. Very few of them have been studied in the context of word spotting and this scarcity of research work is the motivation of the paper. This paper presents a comparative study of classical Dynamic Time Warping (DTW) technique and many of its improved modifications, as well as other sequence matching techniques in the context of word spotting. An experimental study on historical documents is performed to evaluate the behavior of DTW's variants and other sequence matching techniques. A detailed comparative analysis along with wide range of experimentation is performed, which shows that classical DTW remains a good choice when there are no segmentation problems and when features are very local. In case of word segmentation errors, Continuous Dynamic Programming (CDP) seems to be a better choice. This research work has introduced several other improved sequence matching algorithms in the context of word spotting, which show interesting and improved results. Tanmoy Mondal, Nicolas Ragot, Jean-Yves Ramel, Umapada Pal 0001 |
ICDAR | 2 |
| 2015 | Exemplary Sequence Cardinality: An effective application for word spottingabstractIn this paper, a new sequence matching algorithm called as Exemplary Sequence Cardinality (ESC) is proposed. ESC combines several abilities of other sequence matching algorithms e.g. DTW, SSDTW, CDP, FSM, MVM, OSB1. Depending on the application domain, ESC can be tuned to behave such as these different sequence matching algorithms. Its generality and robustness comes from its ability to find subsequences (as in CDP and SSDTW), to skip outliers inside the target sequences (as in MVM and FSM) and also in the query sequence (as in OSB ) and it has the ability to have many to one and one to many correspondences (as in DTW) between the elements of the query and the target sequences. It's special characteristic of skipping noisy elements from query sequence along with other afore mentioned properties gives it an edge over FSM. In case of word spotting application, the outliers skipping capability of ESC makes it less sensible to local variations in the spelling of words, and also to noise present in the query and/or in the target word images. Due to it's capability of sub-sequence matching, the ESC algorithm has the ability to retrieve a query inside a line or piece of line. Finally, its multiple matching facilities (many to one and one to many matching) is proven to be well advantageous in case of different length of target and query sequences due to the variability in scale, font, type/size factors. By experimenting on printed historical document images, we have demonstrated the interest of proposed ESC algorithm in specific cases when incorrect word segmentation and word level local variations occur regularly. Tanmoy Mondal, Nicolas Ragot, Jean-Yves Ramel, Umapada Pal 0001 |
ICDAR | 2 |
| 2015 | OCR performance prediction using cross-OCR alignmentabstractSince 2006 the national library of France (BnF) has developed many mass digitization projects on its collections. The indexation of digital documents on Gallica (the digital library of the BnF) is done through their textual content obtained thanks to service providers that use Optical Character Recognition software (OCR). The modern technologies of OCR achieve good performances on modern documents produced with uniform layout and known fonts. However, for old documents, OCR results are of lower quality. The OCR quality assessment is a real challenge for the BnF. On the one hand, due to the sequential architecture of OCR treatments, the identification of OCR errors sources is intractable. On the other hand, besides the word confidence, no additional quality information is reported in OCR outputs. In this paper, we present a study on OCR performance estimation aiming to control the quality of word transcriptions achieved by OCR. This quality assessment process has to operate without any comparison with ground truthed data. In this respect, our methodology relies on cross alignment of the OCR results with those of a secondary OCR called reference OCR. This secondary OCR provides uncertain but useful information that will be used as uncertain groundtruth. OCR performance is estimated using support vector regression. This predictor uses some global features computed on the cross-alignment results. The experimentations reported show that our estimate describes more faithfully the quality of OCR outputs than average word confidence scores that are computed by OCR. The proposed methodology can be adapted easily to various corpora by tuning the system using a training dataset of documents that have similar properties to those to be treated. Ahmed Ben Salah, Jean-Philippe Moreux, Nicolas Ragot, Thierry Paquet |
ICDAR | 3 |
| 2014 | OCR Performance Prediction Using a Bag of Allographs and Support Vector RegressionabstractIn this paper, we describe a novel and simple technique for prediction of OCR results without using any OCR. The technique uses a bag of allographs to characterize textual components. Then a support vector regression (SVR) technique is used to build a predictor based on the bag of allographs. The performance of the system is evaluated on a corpus of historical documents. The proposed technique produces correct prediction of OCR results on training and test documents within the range of standard deviation of 4.18% and 6.54% respectively. The proposed system has been designed as a tool to assist selection of corpora in libraries and specify the typical performance that can be expected on the selection. Tapan Kumar Bhowmik, Thierry Paquet, Nicolas Ragot |
Document Analysis Systems | 3 |
| 2014 | Flexible Sequence Matching Technique: Application to Word Spotting in Degraded DocumentsabstractIn this paper, a new sequence-matching algorithm, called as Flexible Sequence Matching (FSM) algorithm is proposed. FSM combines several abilities of other sequence matching algorithms (especially DTW, CDP and MVM) that could be configured depending on the application domain. Its generality and robustness comes from its ability to find sub sequences (as in CDP), to skip outliers inside the match sequences (as in MVM) and to match multiple elements with a single one (as in CDP and DTW). These properties make it extremely suitable for robust word spotting. More precisely, the FSM algorithm has the capability to retrieve a query inside a line or piece of line. This facility is useful as word segmentation process may not work accurately or when only line segmentation information is available. Furthermore, thanks to its skipping capability, that makes the proposed FSM algorithm less sensible to local variations in the spelling of words, and also to local degradation effects. Finally, its multiple matching facilities (many to one and one to many matching) are useful in case of different length of target and query sequences due to the variability in scale factor. We demonstrate the superiority of proposed FSM algorithm in specific cases such as incorrect word segmentation and word level local variations. When different experiments were performed using handwritten George Washington dataset and also on historical typewritten document images, quite promising results were obtained. Tanmoy Mondal, Nicolas Ragot, Jean-Yves Ramel, Umapada Pal 0001 |
ICFHR | 2 |
| 2014 | Enhanced omnidirectional image unwrapping for face detectionabstractThis paper introduces a new framework to improve the performance of Viola and Jones face detector on omnidirectional unwrapped images. First, an optimization scheme is used to improve the unwrapped image specifically for rectangular Haar-like features. Then, we compare our unwrapping approach to the performance obtained with spherical unwrapping. The impact of the decision boundary and candidate window density are also investigated. Our work suggests that our new unwrapping technique improves significantly the performance of Viola and Jones detector on omnidirectional unwrapped images. Yohan Dupuis, A. Mendoza Quispe, Pascal Vasseur, Benjamín Castañeda, Nicolas Ragot |
ICIP | 5 |
| 2014 | Fast omni-image unwarping using pano-mapping pointers arrayabstractOmni-cameras are becoming ubiquitous in several applications that require a wide field of view (such as 3D reconstruction, video surveillance, robot vision, etc.) and thus the need to reduce computational time during omni-image un-warping. In this paper a computationally efficient alternative referred as pano-mapping pointers array (PMPA) is proposed. First, a buffer is created to save the omni-image. Then, a PMPA is created, where each entry points to a specific place in the buffer depending on the interpolation desired (nearest neighbors or bilinear interpolation). Finally, using the PMPA we unwarp the omni-image. The PMPA is created once for an omni-camera, interpolation method and panoramic image resolution. Experiments on a standard computer demonstrate that the proposed method is about 5.8 (nearest neighbors) and 2.1 (bilinear interpolation) times faster than the classic pano-mapping table method and about 1.7 (nearest neighbors) times faster than the one-eighth pano-mapping table method. Jaime Reategui, Paul Rodríguez 0001, Nicolas Ragot |
ICIP | 3 |
| 2014 | Word Spotting in Bangla and English Graphical DocumentsabstractWord spotting in graphical documents is a very challenging task. With an increase usage of electronic media, we are in a need of searching objects in graphical documents by some labeled text. To address such scenarios we propose a word spotting system dedicated to graphical documents with Bangla and English scripts. In our proposed system, first text-graphics layers are separated using Gabor filter. In the text layer, character segmentation approach is applied using water reservoir based method to extract each character from the document. Then recognition of these isolated characters is done using rotation invariant feature, coupled with SVM classifier. Well recognized characters are then grouped based on their sizes. Initial spotting is started to find a query word among those groups of characters. In case if the system could spot a word partially due to any noise, SIFT is applied to identify missing portion of that partial spotting. Experimental results on English and Bangla script document images show that the method is feasible to spot a location in text labeled graphical documents. Arundhati Tarafdar, Umapada Pal 0001, Jean-Yves Ramel, Nicolas Ragot, Bidyut B. Chaudhuri |
ICPR | 4 |
| 2014 | Combining Structure and Parameter Adaptation of HMMs for Printed Text RecognitionabstractWe present two algorithms that extend existing HMM parameter adaptation algorithms (MAP and MLLR) by adapting the HMM structure. This improvement relies on a smart combination of MAP and MLLR with a structure optimization procedure. Our algorithms are semi-supervised: to adapt a given HMM model on new data, they require little labeled data for parameter adaptation and a moderate amount of unlabeled data to estimate the criteria used for HMM structure optimization. Structure optimization is based on state splitting and state merging operations and proceeds so as to optimize either the likelihood or a heuristic criterion. Our algorithms are successfully applied to the recognition of printed characters by adapting the HMM character models of a polyfont printed text recognizer to new fonts. Our experiments involve a total of 1,120,000 real and 3,100,000 synthetic character images and concern a set of 89 HMM models. A comparison of our results with those of state-of-the-art adaptation algorithms (MAP and MLLR) shows a significant increase in the accuracy of character recognition. Kamel Ait-Mohand, Thierry Paquet, Nicolas Ragot |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | Document Classification in a Non-stationary Environment: A One-Class SVM ApproachabstractIn this paper, we investigate a specific area of document classification in which the documents come as a flow over the time. Moreover, the exact number of classes of document to deal with is not known from the beginning and could evolve over the time. To be able to perform classification task in such area, we need specific classifiers that are able to perform incremental learning and change their modeling over the time. More specifically, we are focusing our study on SVM approaches, known to perform well, and for which incremental (i-SVM) procedures exist. Nevertheless, most of them are only able to deal with a fixed number of classes. So we designed a new incremental learning procedure based on one-class SVMs. This one is able to improve its classification accuracy over the time, with the arrival of new labeled data, without performing any complete retraining. Moreover, when instances are coming with a previously unknown label (appearance of a new class), the training procedure is able to modify the classifier model to recognize this corresponding new kind of documents. To investigate this area, waiting for collecting documents images as a flow, we did first experiments on the Optical Recognition of Handwritten Digits Data Set. These experiments show that our incremental approach is able: to perform, at each time, as well as a static one-class classifier fully retrained using all previously seen data, to model very quickly and efficiently new incoming classes. Anh Khoi Ngo Ho, Nicolas Ragot, Jean-Yves Ramel, Véronique Eglin, Nicolas Sidere |
ICDAR | 2 |
| 2013 | A Fast Word Retrieval Technique Based on Kernelized Locality Sensitive HashingabstractIn this paper, we have presented a new and faster word retrieval approach, which is able to deal with heterogeneous document image collections. A certain amount of image features (statistical and Gabor Wavelet) are extracted, which inherently represent word's images. These features are used for generating hash table for fast retrieval of similar image from a very large image dataset. The decomposition and embedding of high-dimensional features and complex distance functions into a low-dimensional Hamming space helps to efficiently search items. However, existing methods do not apply for high-dimensional kernelized data when the underlying features' embedding for the kernel is unknown. The generalization of locality sensitive hashing (LSH) for arbitrary kernel is presented in the paper. The proposed algorithm provides sub-linear time similarity search and works for a wide class of similarity functions. Tanmoy Mondal, Nicolas Ragot, Jean-Yves Ramel, Umapada Pal 0001 |
ICDAR | 2 |
| 2013 | A Two-Stage Approach for Word Spotting in Graphical DocumentsabstractPresence of multi-oriented characters, connected characters with graphical lines, intersection of text and symbols with graphical lines/curves etc. are very common in graphical documents. As a result word spotting in graphical documents is still a challenging task that we try to solve (partially) in this paper. The proposed approach proceeds in two stages. In the first stage, recognition of isolated components is done using rotation invariant features and an SVM classifier. The characters having good recognition score and match in the query string are first selected for initial spotting. Because of structural complexity of graphical documents as well as of touching components, we may miss some of the query characters during initial spotting in some documents. In that case, based on the position, size and orientation of the recognized characters in the input document image, regions where missing characters may be located (candidate regions) are defined. In the second stage, Scale Invariant Feature Transform (SIFT) is used to find those missing characters in the candidate regions for possible spotting. Finally, using the position, size, orientation as well as intercharacter gap information of the recognized components, spotting is validated. Experimental results demonstrate that the method is efficient to locate a query word in multi-oriented and/or touching graphical documents. Arundhati Tarafdar, Umapada Pal 0001, Partha Pratim Roy 0001, Nicolas Ragot, Jean-Yves Ramel |
ICDAR | 4 |
| 2011 | Word Retrieval in Historical Document Using Character-PrimitivesabstractWord searching and indexing in historical document collections is a challenging problem because, characters in these documents are often touching or broken due to degradation/ ageing effects. For efficient searching in such historical documents, this paper presents a novel approach towards word spotting using string matching of character primitives. We describe the text string as a sequence of primitives which consists of a single character or a part of a character. Primitive segmentation is performed analyzing text background information that is obtained by water reservoir technique. Next, the primitives are clustered using template matching and a codebook of representative primitives is built. Using this primitive codebook, the text information in the document images are encoded and stored. For a query word, we segment it into primitives and encode the word by a string of representative primitives from codebook. Finally, an approximate string matching is applied to find similar words. The matching similarity is used to rank the retrieved words. The proposed method is tested on historical books of French alphabets and we have obtained encouraging results from the experiment. Partha Pratim Roy 0001, Jean-Yves Ramel, Nicolas Ragot |
ICDAR | 3 |
| 2010 | Structure Adaptation of HMM Applied to OCRabstractIn this paper we present a new algorithm for the adaptation of Hidden Markov Models (HMM models). The principle of our iterative adaptive algorithm is to alternate an HMM structure adaptation stage with an HMM Gaussian MAP adaptation stage of the parameters. This algorithm is applied to the recognition of printed characters to adapt the character models of a poly font general purpose character recognizer to new fonts of characters, never seen during training. A comparison of the results with those of MAP classical adaptation scheme show a slight increase in the recognition performance. Kamel Ait-Mohand, Thierry Paquet, Nicolas Ragot, Laurent Heutte |
ICPR | 3 |
| 2007 | Writer Style Adaptation in Online Handwriting Recognizers by a Fuzzy Mechanism Approach: the Adapt MethodabstractThis study presents an automatic online adaptation mechanism to the handwriting style of a writer for the recognition of isolated handwritten characters. The classifier we use here is based on a Fuzzy Inference System (FIS) similar to those we have designed for handwriting recognition. In this FIS each premise rule is composed of a fuzzy prototype which represents intrinsic properties of a class. Furthermore, the conclusion part of rules associates a score to the prototype for each class. The adaptation mechanism affects both the conclusions of the rules and the fuzzy prototypes by recentering and reshaping them thanks to a new approach called ADAPT inspired by the Learning Vector Quantization. Thus the FIS is automatically fitted to the handwriting style of the writer that currently uses the system. Our adaptation mechanism is compared with well known adaptation techniques. The tests were based on eight different writers and the results illustrate the benefits of the method in terms of error rate reduction (86% in average). This allows such kind of simple classifiers to achieve up to 98.4% of recognition accuracy on the 26 Latin letters in a writer dependent context. Harold Mouchère, Éric Anquetil, Nicolas Ragot |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2005 | Handwritten Gesture Recognition Driven by the Spatial Context of StrokesabstractIn this paper, we present a new approach that explicitly exploits the spatial context of strokes to drive the shape recognition. We call this recognition method "context driven recognition" (CDR). The underlying idea is that only a sub-set of all possible symbols can be recognized in a specific spatial context. The main challenge is to detect and model automatically the context areas of interest so that the recognition method can be independent of any specific information on the targeted pen-based application. The paper details the learning scheme of the CDR method and how the obtained model is used during the recognition process. The results on a real-world pen-based recognition problem show that the method can reach better performances than a classical approach by decreasing the shape recognition complexity. François Bouteruche, Éric Anquetil, Nicolas Ragot |
ICDAR | 3 |
| 2005 | On-line Writer Adaptation for Handwriting Recognition using Fuzzy Inference SystemsabstractWe present an automatic on-line adaptation mechanism to the writer's handwriting style for the recognition of isolated handwritten characters. The classifier is based on a fuzzy inference system (FIS). This FIS is composed of fuzzy prototypes which represent the intrinsic properties of the classes and it uses numeric conclusions. The proposed adaptation mechanism affects both the conclusions of the rules and the fuzzy prototypes of the premises by re-centering and re-shaping them. Doing so, the FIS is automatically fitted to the handwriting style of the writer that is currently using the system. This adaptation mechanism has been tested with 8 different writers. The results show the adaptation mechanism is able to improve the recognition rate from 88% to 98.2% in average for the 26 Latin letters. Harold Mouchère, Éric Anquetil, Nicolas Ragot |
ICDAR | 3 |
| 2003 | A Generic Hybrid Classifier Based on Hierarchical Fuzzy Modeling: Experiments on On-Line Handwritten Character RecognitionabstractIn our previous works, a recognition system named ResifCar was designed specifically for on-line handwritten character recognition. This system is based on an explicit modeling by hierarchical fuzzy rules. Thus, it is understandable an optimizable after the learning stage. We present in this article a new classifier that is an extension of ResifCar. Indeed it tries to combine ResifCar's advantages with a generic aspect to handle different recognition problems. This new hybrid system combines two complementary levels. The first one uses a robust modeling by an intrinsic fuzzy clustering of each class and determines their confusing areas. The second level, based on fuzzy decision trees, operates a progressive discrimination inside these areas. Both levels are formalized by fuzzy inference systems organized hierarchically and fused for final decision. Experiments were conducted on the one hand on classical benchmarks and on the other hand on on-line handwritten digits and lower-case letters. For all of these cases, the classifier achieves good recognition rates without final optimization. 1. Nicolas Ragot, Éric Anquetil |
ICDAR | 1 |
| 2001 | A New Hybrid Learning Method for Fuzzy Decision TreesabstractThis paper presents a new hybrid learning method for the construction of fuzzy decision trees. The main principle of this approach is to automatically generates a hierarchical organization of the knowledge coupled with local choice of the best feature subspace. To improve the representation, a double level of modeling is used. Firstly a pre-classification level searches fuzzy decision regions to operate a natural discrimination between classes. The second level refines the previous one, doing an intrinsic fuzzy modeling of the classes represented in the fuzzy regions. Moreover, the best feature subspace is determined locally by a genetic algorithm for each partitioning. Finally, to have an understandable and "transparent" representation, the fuzzy decision tree is formalized as a fuzzy inference system which is easily modifiable and can be optimized a posteriori. First experimental results conducted on classical benchmarks and on a handwritten digits database show the capacity of the hybrid learning approach to provide reliable and compact classification system. Nicolas Ragot, Éric Anquetil |
FUZZ-IEEE | 1 |