EDBT 2026 Demo / reviewers in the wild / expert
Salvatore Tabbone
dblp:13/3772 · also Salvatore-Antoine Tabbone
· DBLP profile ↗
29ranked-venue papers in the field
3as first author
3since 2021 · last 2026
0000-0002-0024-1280ORCID · verified
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 28 (3 first)Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Vision-Language Models on Historical Postcards
Matthieu Pelingre, Salvatore Tabbone |
ICDAR (3) | 2 |
| 2023 | Historical Document Image Segmentation Combining Deep Learning and Gabor Features
Maroua Mehri, Akrem Sellami, Salvatore Tabbone |
ICDAR (4) | 3 |
| 2021 | EDNets: Deep Feature Learning for Document Image Classification Based on Multi-view Encoder-Decoder Neural Networks
Akrem Sellami, Salvatore Tabbone |
ICDAR (4) | 2 |
| 2019 | Automatic Synthetic Document Image Generation using Generative Adversarial Networks: Application in Mobile-Captured Document AnalysisabstractIn this paper, we propose a method using Generative Adversarial Networks for automatically synthesizing document images that are similar to real printed documents captured by mobile phone's camera in unconstrained environment. We focus on the simulation of image defects for unconstrained mobile image acquisition procedure (non-uniform illumination, defocusing, optical and mechanical deformations, vibrations, noise in electronic components,...). Our approach is proven to be low-cost as it only requires a collection of real document images without any annotation. Experimental results show the effectiveness of our approach to improve OCR (Optical Character Recognition) recognition rate in a mobile-captured document images framework. Although in this paper, we focus on modern printed document images, our proposed approach could be extended to another type of documents, including historical one. Quang Anh Bui, David Mollard, Salvatore Tabbone |
ICDAR | 3 |
| 2018 | Predicting Mobile-Captured Document Images Sharpness QualityabstractNowadays the number of mobile applications is fast growing. Among them, mobile applications based on Optical Character Recognition (OCR) play an important role. One of the main challenge of such applications to overcome is that the image acquisition procedure is in a manner unreliable and may contain many distortions. As a consequence, a suitable OCR output requires efforts to enhance the quality of the captured image. This leads to an increase of computation time and cost. In this perspective, we focus on the prediction of image's sharpness quality. We choose to concentrate on image's sharpness quality because blur distortions seriously alter readability for both human and computer. Our contribution consists of a method combining focus and sharpness measures with a Support Vector Machine to classify image's sharpness quality. This approach is fast, reliable, and can be easily implemented on mobile devices. Experimental results show that our method is efficient for OCR based mobile-captured document images. Quang Anh Bui, David Mollard, Salvatore Tabbone |
DAS | 3 |
| 2017 | Selecting Automatically Pre-Processing Methods to Improve OCR PerformancesabstractIn this paper, we propose an approach that automatically selects suitable document pre-processing algorithms to increase OCR performances. We first provide an experimental evaluations protocol to study effects of document pre-processing methods on different OCR engines for document images that have different type of distorsions. We remark that, when distortions on the document image is unknown, a pre-processing methods does not always improve but sometimes decreases the OCR performance. We conclude that the effectiveness of a pre-processing algorithm depends on the nature of the OCR and type of distorsions. In the context that distortions on the document and information about OCR system's mechanism are unknown, we propose an automatic pre-processing selection method based on a convolutional neural network with 15 layers and where the last layer contains neurons representing our different pre-processing algorithms. Experimental results show the effectiveness of our approach to improve OCR performances in a mobile-captured document images framework. Quang Anh Bui, David Mollard, Salvatore Tabbone |
ICDAR | 3 |
| 2015 | Automatic annotation extension and classification of documents using a probabilistic graphical modelabstractWith the fast growth of document images, document annotation has become a research area of great interest. Annotation allows to describe the semantic content of documents and facilitates their use and research. However, for a huge number of documents, the manual annotation of each document becomes a tedious task. A solution is to annotate a small part of the documents and to extend it automatically to the whole dataset. In this paper, we propose a model for annotation extension and document classification using a probabilistic graphical model. In this latter, we combine visual and textual characteristics and we show that the integration of the user feedback improves the annotation step. Abdessalem Bouzaieni, Sabine Barrat, Salvatore Tabbone |
ICDAR | 3 |
| 2014 | Spotting Symbol Using Sparsity over Learned Dictionary of Local DescriptorsabstractThis paper proposes a new approach to spot symbols into graphical documents using sparse representations. More specifically, a dictionary is learned from a training database of local descriptors defined over the documents. Following their sparse representations, interest points sharing similar properties are used to define interest regions. Using an original adaptation of information retrieval techniques, a vector model for interest regions and for a query symbol is built based on its sparsity in a visual vocabulary where the visual words are columns in the learned dictionary. The matching process is performed comparing the similarity between vector models. Evaluation on SESYD datasets demonstrates that our method is promising. Thanh-Ha Do, Salvatore Tabbone, Oriol Ramos Terrades |
Document Analysis Systems | 2 |
| 2014 | Multiscale Stroke-Based Page Segmentation ApproachabstractIn this paper we present a new hybrid page segmentation approach based on connected component and region analysis. We first describe our stroke descriptor that detects text and line component candidates using the skeleton of the binarized document image. Then, an active contour model is applied to segment the rest of the image into photo and background regions. This classification is verified by studying the variation of each detected region. Finally, we cluster the text candidates using mean-shift analysis technique according to their corresponding sizes and we present our adaptive projection profile approach to gather separately horizontal and vertical text regions. The method is applied for segmenting realistic scanned document images (newspapers and magazines) that contain text, lines and photo regions. We evaluate the performances of our approach by comparing it to the existing methods that participated in ICDAR page segmentation competition. Mehdi Felhi, Salvatore Tabbone, Maria V. Ortiz Segovia |
Document Analysis Systems | 2 |
| 2013 | Document noise removal using sparse representations over learned dictionaryabstractIn this paper, we propose an algorithm for denoising document images using sparse representations. Following a training set, this algorithm is able to learn the main document characteristics and also, the kind of noise included into the documents. In this perspective, we propose to model the noise energy based on the normalized cross-correlation between pairs of noisy and non-noisy documents. Experimental results on several datasets demonstrate the robustness of our method compared with the state-of-the-art. Thanh-Ha Do, Salvatore Tabbone, Oriol Ramos Terrades |
ACM Symposium on Document Engineering | 2 |
| 2013 | New Approach for Symbol Recognition Combining Shape Context of Interest Points with Sparse RepresentationabstractIn this paper, we propose a new approach for symbol description. Our method is built based on the combination of shape context of interest points descriptor and sparse representation. More specifically, we first learn a dictionary describing shape context of interest point descriptors. Then, based on information retrieval techniques, we build a vector model for each symbol based on its sparse representation in a visual vocabulary whose visual words are columns in the learned dictionary. The retrieval task is performed by ranking symbols based on similarity between vector models. The evaluation of our method, using benchmark datasets, demonstrates the validity of our approach and shows that it outperforms related state-of-the-art methods. Thanh-Ha Do, Salvatore Tabbone, Oriol Ramos Terrades |
ICDAR | 2 |
| 2012 | Symbol Recognition Using a Galois Lattice of Frequent Graphical PatternsabstractGraphics recognition is an important task in many real-life applications. In this article, we propose a new approach to recognize graphical symbols by the use of a frequent Galois lattice. We propose to build a concept lattice not in terms of graphical patterns but in terms of frequent graphical patterns. The purpose of this paper is twofold : first, we try to identify the best primitives from a given graphical symbol based on a descriptor invariant to rotation, translation and scaling. Each symbol is decribed using a feature vector computed on stable neighborhood for a set of points chosen randomly from the symbol. Secondly, we propose a new recognition approach based on a frequent Galois lattice. The obtained concept lattice based on frequent patterns is used as a classifier. The retrieval performance and behavior of the method have been tested for graphics recognition. We have compared our method with others based on different descriptors and classifiers. Our approach proves that the symbol description method and the algorithm used to extract frequent attributes to build the frequent Galois lattice are suitable to the recognition process. Ameni Boumaiza, Salvatore Tabbone |
Document Analysis Systems | 2 |
| 2011 | A Novel Approach for Graphics Recognition Based on Galois Lattice and Bag of Words RepresentationabstractThis paper presents a new approach for graphical symbols recognition by combining a concept lattice with a bag of words representation. Visual words define the properties of a graphical symbol that will be modeled in the Galois Lattice. The algorithm of classification is based on the Galois lattice where intentions of its concepts are visual words. The words as visual primitives allow to evaluate the classifier with a symbolic approach that no longer need a signature discretization step to build the Galois Lattice. Our approach is compared to classical approaches on different graphical symbols and we show the relevance and the robustness of our proposal for the classification task. Ameni Boumaiza, Salvatore Tabbone |
ICDAR | 2 |
| 2011 | A Shape Descriptor Combining Logarithmic-Scale Histogram of Radon Transform and Phase-Only Correlation FunctionabstractA shape descriptor combining the histogram of the Radon transform, the logarithmic-scale histogram, and the phase-only correlation function is proposed. Applying a logarithmic-scale to the Radon transform, the shape scaling and rotation become two-dimensional translation in our descriptor without any normalization. The geometric invariance to translation, when we match two shapes, are kept using the phase-only correlation function. In addition, we can determine with this function the rotation angle and the scale parameter between two shapes. Our descriptor is robust to shape occlusion and noise also. Makoto Hasegawa, Salvatore Tabbone |
ICDAR | 2 |
| 2010 | Text extraction from graphical document images using sparse representationabstractA novel text extraction method from graphical document images is presented in this paper. Graphical document images containing text and graphics components are considered as two-dimensional signals by which text and graphics have different morphological characteristics. The proposed algorithm relies upon a sparse representation framework with two appropriately chosen discriminative overcomplete dictionaries, each one gives sparse representation over one type of signal and non-sparse representation over the other. Separation of text and graphics components is obtained by promoting sparse representation of input images in these two dictionaries. Some heuristic rules are used for grouping text components into text strings in post-processing steps. The proposed method overcomes the problem of touching between text and graphics. Preliminary experiments show some promising results on different types of document. Thai V. Hoang, Salvatore Tabbone |
Document Analysis Systems | 2 |
| 2010 | A system to detect rooms in architectural floor plan imagesabstractIn this article, a system to detect rooms in architectural floor plan images is described. We first present a primitive extraction algorithm for line detection. It is based on an original coupling of classical Hough transform with image vectorization in order to perform robust and efficient line detection. We show how the lines that satisfy some graphical arrangements are combined into walls. We also present the way we detect some door hypothesis thanks to the extraction of arcs. Walls and door hypothesis are then used by our room segmentation strategy; it consists in recursively decomposing the image until getting nearly convex regions. The notion of convexity is difficult to quantify, and the selection of separation lines between regions can also be rough. We take advantage of knowledge associated to architectural floor plans in order to obtain mostly rectangular rooms. Qualitative and quantitative evaluations performed on a corpus of real documents show promising results. Sébastien Macé, Hervé Locteau, Ernest Valveny, Salvatore Tabbone |
Document Analysis Systems | 4 |
| 2009 | Modeling, Classifying and Annotating Weakly Annotated Images Using Bayesian NetworkabstractWe propose a probabilistic graphical model to represent weakly annotated images. This model is used to classify images and automatically extend existing annotations to new images by taking into account semantic relations between keywords. The proposed method has been evaluated in classification and automatic annotation of images. The experimental results, obtained from a database of more than 30000 images, by combining visual and textual information, show an improvement by 50.5% in terms of recognition rate against only visual information classification. Taking into account semantic relations between keywords improves the recognition rate by 10.5% and the mean rate of good annotations by 6.9%. The proposed method is experimentally competitive with the state-of-art classifiers. Sabine Barrat, Salvatore Tabbone |
ICDAR | 2 |
| 2009 | Generic Feature Selection and Document ProcessingabstractThis paper presents a generic features selection method and its applications on some document analysis problems.The method is based on a genetic algorithm (GA), whose fitness function is defined by combining Adaboot classifiers associated with each feature. Our method is not linked to a classifier achieving the final recognition task; we have used a combination of weak classifiers to evaluate a subset of features. So we select features that can further be used in the most appropriate classifiers.This method has been tested on three applications: dropcaps classification, handwritten digits recognition and text detection. The results show the efficiency and robustness of the proposed approach. Hassan Chouaib, Nicole Vincent, Florence Cloppet, Salvatore Tabbone |
ICDAR | 4 |
| 2009 | Extraction of Nom Text Regions from Stele Images Using Area Voronoi DiagramabstractAutomatic processing of images of steles is a challenging problem due to the variation in their structures and body text characteristics. In this paper, area Voronoi diagram is used to represent the neighborhood of connected components in stele images containing Nom characters. Body text region is then extracted from stele images by the selection of appropriate adjacent Voronoi regions based on the information about the thickness of neighboring connected components. Experimental results show that the proposed method is highly accurate and robust to various types of stele. Thai V. Hoang, Salvatore Tabbone, Ngoc-Yen Pham |
ICDAR | 2 |
| 2009 | A Symbol Spotting Approach Based on the Vector Model and a Visual VocabularyabstractThis paper addresses the difficult problem of symbol spotting for graphic documents. We propose an approach where each graphic document is indexed as a text document by using the vector model and an inverted file structure. The method relies on a visual vocabulary built from a shape descriptor adapted to the document level and invariant under classical geometric transforms (rotation, scaling and translation). Regions of interest selected with high degree of confidence using a voting strategy are considered as occurrences of a query symbol. Experimental results are promising and show the feasibility of our approach. Salvatore Tabbone, Alain Boucher |
ICDAR | 2 |
| 2008 | Symbol Descriptor Based on Shape Context and Vector Model of Information RetrievalabstractIn this paper we present an adaptive method for graphic symbol representation based on shape contexts. The proposed descriptor is invariant under classical geometric transforms (rotation, scale) and based on interest points. To reduce the complexity of matching a symbol to a largeset of candidates we use the popular vector model for information retrieval. In this way, on the set of shape descriptors we build a visual vocabulary where each symbol is retrieved on visual words. Experimental results on complex and occluded symbols show that the approach is very promising. Salvatore Tabbone, Oriol Ramos Terrades |
Document Analysis Systems | 2 |
| 2007 | An Indexing Method for Graphical DocumentsabstractIn this paper, a method to browse symbols into graphical documents is presented. More precisely, we propose a combined filtering and indexing mechanism that retrieves in an efficient way the most similar symbols to a given input query. For a database of 200000 symbols the retrieval time has been divided by a factor of 4, 5 compared to a linear search. Salvatore Tabbone, Daniel Zuwala |
ICDAR | 1 |
| 2007 | A Review of Shape Descriptors for Document AnalysisabstractShape descriptors play an important role in many document analysis application. In this paper we review some of the shape descriptors proposed in the last years from a new point of view. We propose the definitions of descriptor and primitive and introduce the notion of feature extraction method. With these definitions, we propose a new classification of shape descriptors that permits to classify according to their properties pointing out their strengths and weaknesses. Oriol Ramos Terrades, Salvatore Tabbone, Ernest Valveny |
ICDAR | 2 |
| 2006 | A Method for Symbol Spotting in Graphical Documents
Daniel Zuwala, Salvatore Tabbone |
Document Analysis Systems | 2 |
| 2005 | Automatical Definition of Measures from the Combination of Shape DescriptorsabstractThis paper presents a novel approach to combine shape descriptors. Each approach is applied on several clusters of objects. For each cluster and for any descriptor a map is associated directly from the confusion matrix. Such a method allows to determine automatically the better weight associated to the descriptor for the object under consideration. At last, we show that the additive combination of such measures allows to improve the classification. J.-P. Salmon, Laurent Wendling, Salvatore Tabbone |
ICDAR | 3 |
| 2004 | A Hybrid Approach to Detect Graphical Symbols in Documents
Salvatore Tabbone, Laurent Wendling, Daniel Zuwala |
Document Analysis Systems | 1 |
| 2003 | Recognition of Arrows in Line Drawings based on the Aggregation of Geometric Criteria using the Choquet IntegralabstractA new way to detect arrows in line drawings is proposed in this paper. Our approach is based on the definition of the structure of such a symbol. Signatures of angular areas are computed and axiomatic properties and geometric characteristics are checked using the Choquet integral. Finally an experimental application on line-drawing documents shows the interest of our approach. Laurent Wendling, Salvatore Tabbone |
ICDAR | 2 |
| 2002 | Text/Graphics Separation Revisited
Karl Tombre, Salvatore Tabbone, Loïc Pélissier, Bart Lamiroy, Philippe Dosch |
Document Analysis Systems | 2 |
| 2001 | Indexing of Technical Line Drawings Based on F-SignaturesabstractWe propose a method for indexing technical drawings. Our features are based on the notion of F-signature, which is a particular histogram of forces. The force histogram has low time complexity and describes a signature which is invariant to scaling, translation, symmetry and rotation. This article presents a new application of such signatures in the field of document analysis, and tests different characteristics of the F-signatures, like their sensitivity to the shape of the objects and to noise. Finally, experimental results show the effectiveness of such an approach. A brief overview of pattern recognition approaches dedicated to technical document indexing is given. The notions of histogram of forces and of its properties are presented. Salvatore Tabbone, Laurent Wendling, Karl Tombre |
ICDAR | 1 |