EDBT 2026 Demo / reviewers in the wild / expert
Josep Lladós 0001
dblp:54/769 · also Josep Lladós Canet
· DBLP profile ↗
71ranked-venue papers in the field
3as first author
17since 2021 · last 2026
0000-0002-4533-4739ORCID · conflict
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 70 (3 first)Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Chunks to Graphs: Training-Free Multimodal Late Interaction for Document Understanding
Ayush Lodh, Souparni Mazumder, Sanket Biswas, Josep Lladós 0001, Nisha Singh Chauhan |
ICDAR (3) | 4 |
| 2026 | Robust Interpretation of Historical Documents in Knowledge Graphs Through Query Inference and Execution
Sebastià Nicolau, Adrià Molina, Oriol Ramos Terrades, Josep Lladós 0001 |
ICDAR (3) | 4 |
| 2026 | Conversational Retrieval and On-the-Fly Knowledge Modeling of Historical Penitentiary Repression Records
Paula Font Solà, Adrià Molina Rodríguez, Josep Lladós 0001 |
ICDAR (3) | 3 |
| 2025 | Where Layout Meets Language: Lightweight Spatial Enhancement to Large Language Models for Document Understanding
Nil Biescas, Sanket Biswas, Josep Lladós 0001, Jordy Van Landeghem |
ICDAR (4) | 3 |
| 2025 | Doc2GraphFormer: Bridging Structured Graph Learning with Transformer Attention for Efficient Document Understanding
Souparni Mazumder, Sanket Biswas, Aniket Pal, Alloy Das, Umapada Pal 0001, Josep Lladós 0001 |
ICDAR (4) | 6 |
| 2025 | ICDAR 2025 Handwritten Notes Understanding Challenge
Aniket Pal, Sanket Biswas, Alloy Das, Ayush Lodh, Priyanka Banerjee, Soumitri Chattopadhyay, Ajoy Mondal, Dimosthenis Karatzas, Josep Lladós 0001, C. V. Jawahar |
ICDAR (5) | 9 |
| 2024 | Fetch-A-Set: A Large-Scale OCR-Free Benchmark for Historical Document Retrieval
Adrià Molina, Oriol Ramos Terrades, Josep Lladós 0001 |
DAS | 3 |
| 2024 | GraphKD: Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
Ayan Banerjee 0002, Sanket Biswas, Josep Lladós 0001, Umapada Pal 0001 |
ICDAR (3) | 3 |
| 2024 | GeoContrastNet: Contrastive Key-Value Edge Learning for Language-Agnostic Document Understanding
Nil Biescas, Carlos Boned, Josep Lladós 0001, Sanket Biswas |
ICDAR (1) | 3 |
| 2024 | DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
Jordy Van Landeghem, Subhajit Maity, Ayan Banerjee 0002, Matthew B. Blaschko, Marie-Francine Moens, Josep Lladós 0001, Sanket Biswas |
ICDAR (4) | 6 |
| 2024 | SketchGPT: Autoregressive Modeling for Sketch Generation and Recognition
Adarsh Tiwari, Sanket Biswas, Josep Lladós 0001 |
ICDAR (5) | 3 |
| 2023 | SwinDocSegmenter: An End-to-End Unified Domain Adaptive Transformer for Document Instance Segmentation
Ayan Banerjee 0002, Sanket Biswas, Josep Lladós 0001, Umapada Pal 0001 |
ICDAR (1) | 3 |
| 2023 | SelfDocSeg: A Self-supervised Vision-Based Approach Towards Document Segmentation
Subhajit Maity, Sanket Biswas, Siladittya Manna, Ayan Banerjee 0002, Josep Lladós 0001, Saumik Bhattacharya, Umapada Pal 0001 |
ICDAR (1) | 5 |
| 2022 | A Generic Image Retrieval Method for Date Estimation of Historical Document Collections
Adrià Molina, Lluís Gómez i Bigorda, Oriol Ramos Terrades, Josep Lladós 0001 |
DAS | 4 |
| 2021 | DocSynth: A Layout Guided Approach for Controllable Document Image Synthesis
Sanket Biswas, Pau Riba, Josep Lladós 0001, Umapada Pal 0001 |
ICDAR (3) | 3 |
| 2021 | Date Estimation in the Wild of Scanned Historical Photos: An Image Retrieval Approach
Adrià Molina, Pau Riba, Lluís Gómez i Bigorda, Oriol Ramos Terrades, Josep Lladós 0001 |
ICDAR (2) | 5 |
| 2021 | Learning to Rank Words: Optimizing Ranking Metrics for Word Spotting
Pau Riba, Adrià Molina, Lluís Gómez i Bigorda, Oriol Ramos Terrades, Josep Lladós 0001 |
ICDAR (2) | 5 |
| 2019 | Recurrent Comparator with Attention Models to Detect Counterfeit DocumentsabstractThis paper is focused on the detection of counterfeit documents via the recurrent comparison of the security textured background regions of two images. The main contributions are twofold: first we apply and adapt a recurrent comparator architecture with attention mechanism to the counterfeit detection task, which constructs a representation of the background regions by recurrently condition the next observation, learning the difference between genuine and counterfeit images through iterative glimpses. Second we propose a new counterfeit document dataset to ensure the generalization of the learned model towards the detection of the lack of resolution during the counterfeit manufacturing. The presented network, outperforms state-of-the-art classification approaches for counterfeit detection as demonstrated in the evaluation. Albert Berenguel, Oriol Ramos Terrades, Josep Lladós 0001, Cristina Cañero Morales |
ICDAR | 3 |
| 2019 | Table Detection in Invoice Documents by Graph Neural NetworksabstractTabular structures in documents offer a complementary dimension to the raw textual data, representing logical or quantitative relationships among pieces of information. In digital mail room applications, where a large amount of administrative documents must be processed with reasonable accuracy, the detection and interpretation of tables is crucial. Table recognition has gained interest in document image analysis, in particular in unconstrained formats (absence of rule lines, unknown information of rows and columns). In this work, we propose a graph-based approach for detecting tables in document images. Instead of using the raw content (recognized text), we make use of the location, context and content type, thus it is purely a structure perception approach, not dependent on the language and the quality of the text reading. Our framework makes use of Graph Neural Networks (GNNs) in order to describe the local repetitive structural information of tables in invoice documents. Our proposed model has been experimentally validated in two invoice datasets and achieved encouraging results. Additionally, due to the scarcity of benchmark datasets for this task, we have contributed to the community a novel dataset derived from the RVL-CDIP invoice data. It will be publicly released to facilitate future research. Pau Riba, Anjan Dutta 0001, Lutz Goldmann, Alicia Fornés, Oriol Ramos Terrades, Josep Lladós 0001 |
ICDAR | 6 |
| 2018 | Joint Recognition of Handwritten Text and Named Entities with a Neural End-to-End ModelabstractWhen extracting information from handwritten documents, text transcription and named entity recognition are usually faced as separate subsequent tasks. This has the disadvantage that errors in the first module affect heavily the performance of the second module. In this work we propose to do both tasks jointly, using a single neural network with a common architecture used for plain text recognition. Experimentally, the work has been tested on a collection of historical marriage records. Results of experiments are presented to show the effect on the performance for different configurations: different ways of encoding the information, doing or not transfer learning and processing at text line or multi-line region level. The results are comparable to state of the art reported in the ICDAR 2017 Information Extraction competition, even though the proposed technique does not use any dictionaries, language modeling or post processing. Manuel Carbonell, Mauricio Villegas, Alicia Fornés, Josep Lladós 0001 |
DAS | 4 |
| 2017 | Pyramidal Stochastic Graphlet Embedding for Document Pattern ClassificationabstractDocument pattern classification methods using graphs have received a lot of attention because of its robust representation paradigm and rich theoretical background. However, the way of preserving and the process for delineating documents with graphs introduce noise in the rendition of underlying data, which creates instability in the graph representation. To deal with such unreliability in representation, in this paper, we propose Pyramidal Stochastic Graphlet Embedding (PSGE). Given a graph representing a document pattern, our method first computes a graph pyramid by successively reducing the base graph. Once the graph pyramid is computed, we apply Stochastic Graphlet Embedding (SGE) for each level of the pyramid and combine their embedded representation to obtain a global delineation of the original graph. The consideration of pyramid of graphs rather than just a base graph extends the representational power of the graph embedding, which reduces the instability caused due to noise and distortion. When plugged with support vector machine, our proposed PSGE has outperformed the state-of-the-art results in recognition of handwritten words as well as graphical symbols. Anjan Dutta 0001, Pau Riba, Josep Lladós 0001, Alicia Fornés |
ICDAR | 3 |
| 2017 | Evaluation of Texture Descriptors for Validation of Counterfeit DocumentsabstractThis paper describes an exhaustive comparative analysis and evaluation of different existing texture descriptor algorithms to differentiate between genuine and counterfeit documents. We include in our experiments different categories of algorithms and compare them in different scenarios with several counterfeit datasets, comprising banknotes and identity documents. Computational time in the extraction of each descriptor is important because the final objective is to use it in a real industrial scenario. HoG and CNN based descriptors stands out statistically over the rest in terms of the F1-score/time ratio performance. Albert Berenguel, Oriol Ramos Terrades, Josep Lladós 0001, Cristina Cañero Morales |
ICDAR | 3 |
| 2017 | ICDAR2017 Competition on Information Extraction in Historical Handwritten RecordsabstractThe extraction of relevant information from historical handwritten document collections is one of the key steps in order to make these manuscripts available for access and searches. In this competition, the goal is to detect the named entities and assign each of them a semantic category, and therefore, to simulate the filling in of a knowledge database. This paper describes the dataset, the tasks, the evaluation metrics, the participants methods and the results. Alicia Fornés, Verónica Romero 0001, Arnau Baró, Juan Ignacio Toledo, Joan-Andreu Sánchez, Enrique Vidal 0001, Josep Lladós 0001 |
ICDAR | 7 |
| 2017 | Improving Information Retrieval in Multiwriter Scenario by Exploiting the Similarity Graph of Document TermsabstractInformation Retrieval (IR) is the activity of obtaining information resources relevant to a questioned information. It usually retrieves a set of objects ranked according to the relevancy to the needed fact. In document analysis, information retrieval receives a lot of attention in terms of symbol and word spotting. However, through decades the community mostly focused either on printed or on single writer scenario, where the state-of-the-art results have achieved reasonable performance on the available datasets. Nevertheless, the existing algorithms do not perform accordingly on multiwriter scenario. A graph representing relations between a set of objects is a structure where each node delineates an individual element and the similarity between them is represented as a weight on the connecting edge. In this paper, we explore different analytics of graphs constructed from words or graphical symbols, such as diffusion, shortest path, etc. to improve the performance of information retrieval methods in multiwriter scenario. Pau Riba, Anjan Dutta 0001, Sounak Dey, Josep Lladós 0001, Alicia Fornés |
ICDAR | 4 |
| 2017 | Handwriting Recognition by Attribute Embedding and Recurrent Neural NetworksabstractHandwriting recognition consists in obtaining the transcription of a text image. Recent word spotting methods based on attribute embedding have shown good performance when recognizing words. However, they are holistic methods in the sense that they recognize the word as a whole (i.e. they find the closest word in the lexicon to the word image). Consequently, these kinds of approaches are not able to deal with out of vocabulary words, which are common in historical manuscripts. Also, they cannot be extended to recognize text lines. In order to address these issues, in this paper we propose a handwriting recognition method that adapts the attribute embedding to sequence learning. Concretely, the method learns the attribute embedding of patches of word images with a convolutional neural network. Then, these embeddings are presented as a sequence to a recurrent neural network that produces the transcription. We obtain promising results even without the use of any kind of dictionary or language model. Juan Ignacio Toledo, Sounak Dey, Alicia Fornés, Josep Lladós 0001 |
ICDAR | 4 |
| 2016 | Banknote Counterfeit Detection through Background Texture Printing AnalysisabstractThis paper is focused on the detection of counterfeit photocopy banknotes. The main difficulty is to work on a real industrial scenario without any constraint about the acquisition device and with a single image. The main contributions of this paper are twofold: first the adaptation and performance evaluation of existing approaches to classify the genuine and photocopy banknotes using background texture printing analysis, which have not been applied into this context before. Second, a new dataset of Euro banknotes images acquired with several cameras under different luminance conditions to evaluate these methods. Experiments on the proposed algorithms show that mixing SIFT features and sparse coding dictionaries achieves quasi perfect classification using a linear SVM with the created dataset. Approaches using dictionaries to cover all possible texture variations have demonstrated to be robust and outperform the state-of-the-art methods using the proposed benchmark. Albert Berenguel, Oriol Ramos Terrades, Josep Lladós 0001, Cristina Cañero Morales |
DAS | 3 |
| 2016 | An Interactive Transcription System of Census Records Using Word-Spotting Based Information TransferabstractThis paper presents a system to assist in the transcription of historical handwritten census records in a crowdsourcing platform. Census records have a tabular structured layout. They consist in a sequence of rows with information of homes ordered by street address. For each household snippet in the page, the list of family members is reported. The censuses are recorded in intervals of a few years and the information of individuals in each household is quite stable from a point in time to the next one. This redundancy is used to assist the transcriber, so the redundant information is transferred from the census already transcribed to the next one. Household records are aligned from one year to the next one using the knowledge of the ordering by street address. Given an already transcribed census, a query by string word spotting is applied. Thus, names from the census in time t are used as queries in the corresponding home record in time t+1. Since the search is constrained, the obtained precision-recall values are very high, with an important reduction in the transcription time. The proposed system has been tested in a real citizen-science experience where non expert users transcribe the census data of their home town. Joan Mas Romeu, Alicia Fornés, Josep Lladós 0001 |
DAS | 3 |
| 2016 | Election Tally Sheets Processing SystemabstractIn paper based elections, manual tallies at polling station level produce myriads of documents. These documents share a common form-like structure and a reduced vocabulary worldwide. On the other hand, each tally sheet is filled by a different writer and on different countries, different scripts are used. We present a complete document analysis system for electoral tally sheet processing combining state of the art techniques with a new handwriting recognition subprocess based on unsupervised feature discovery with Variational Autoencoders and sequence classification with BLSTM neural networks. The whole system is designed to be script independent and allows a fast and reliable results consolidation process with reduced operational cost. Juan Ignacio Toledo, Alicia Fornés, Jordi Cucurull-Juan, Josep Lladós 0001 |
DAS | 4 |
| 2015 | A semi-automatic groundtruthing tool for mobile-captured document segmentationabstractThis paper presents a novel way to generate ground-truth data for the evaluation of mobile document capture systems, focusing on the first stage of the image processing pipeline involved: document object detection and segmentation in low-quality preview frames. We introduce and describe a simple, robust and fast technique based on color markers which enables a semi-automated annotation of page corners. We also detail a technique for marker removal. Methods and tools presented in the paper were successfully used to annotate, in few hours, 24889 frames in 150 video files for the smartDOC competition at ICDAR 2015. Joseph Chazalon, Marçal Rusiñol, Jean-Marc Ogier, Josep Lladós 0001 |
ICDAR | 4 |
| 2015 | Hidden Markov model topology optimization for handwriting recognitionabstractIn this paper we present a method to optimize the topology of linear left-to-right hidden Markov models. These models are very popular for sequential signals modeling on tasks such as handwriting recognition. Many topology definition methods select the number of states for a character model based on character length. This can be a drawback when characters are shorter than the minimum allowed by the model, since they can not be properly trained nor recognized. The proposed method optimizes the number of states per model by automatically including convenient skip-state transitions and therefore it avoids the aforementioned problem. We discuss and compare our method with other character length-based methods such the Fixed, Bakis and Quantile methods. Our proposal performs well on off-line handwriting recognition task. Núria Cirera, Alicia Fornés, Josep Lladós 0001 |
ICDAR | 3 |
| 2015 | Novel line verification for multiple instance focused retrieval in document collectionsabstractSpatial verification is typically employed to check the spatial consistency among matched local features and to remove outliers. However, when looking for multiple instances of the query within a target image, RANSAC algorithms which are widely applied in many one-to-one matching applications might fail due to the large proportion of “outliers” - correct matches corresponding to other instances. On the other hand, geometrical verification methods are more robust to outliers but usually suffer from high computational costs. In this paper, we introduce a novel two-step line verification method which is more flexible than existing methods and leads to lower computational complexity especially when multiple instances of a query are sought. We study this approach within an information extraction scenario, where the objective is to locate document structures indicative of certain type of information (e.g. different records on invoices). Hongxing Gao, Marçal Rusiñol, Dimosthenis Karatzas, Josep Lladós 0001, Rajiv Jain, David S. Doermann |
ICDAR | 4 |
| 2015 | Attributed Graph Grammar for floor plan analysisabstractIn this paper, we propose the use of an Attributed Graph Grammar as unique framework to model and recognize the structure of floor plans. This grammar represents a building as a hierarchical composition of structurally and semantically related elements, where common representations are learned stochastically from annotated data. Given an input image, the parsing consists on constructing that graph representation that better agrees with the probabilistic model defined by the grammar. The proposed method provides several advantages with respect to the traditional floor plan analysis techniques. It uses an unsupervised statistical approach for detecting walls that adapts to different graphical notations and relaxes strong structural assumptions such are straightness and orthogonality. Moreover, the independence between the knowledge model and the parsing implementation allows the method to learn automatically different building configurations and thus, to cope the existing variability. These advantages are clearly demonstrated by comparing it with the most recent floor plan interpretation techniques on 4 datasets of real floor plans with different notations. Lluís-Pere de las Heras, Oriol Ramos Terrades, Josep Lladós 0001 |
ICDAR | 3 |
| 2015 | Use case visual Bag-of-Words techniques for camera based identity document classificationabstractNowadays, automatic identity document recognition, including passport and driving license recognition, is at the core of many applications within the administrative and service sectors, such as police, hospitality, car renting, etc. In former years, the document information was manually extracted whereas today this data is recognized automatically from images obtained by flat-bed scanners. Yet, since these scanners tend to be expensive and voluminous, companies in the sector have recently turned their attention to cheaper, small and yet computationally powerful scanners: the mobile devices. The document identity recognition from mobile images enclose several new difficulties w.r.t traditional scanned images, such as the loss of a controlled background, perspective, blurring, etc. In this paper we present a real application for identity document classification of images taken from mobile devices. This classification process is of extreme importance since a prior knowledge of the document type and origin strongly facilitates the subsequent information extraction. The proposed method is based on a traditional Bagof-Words in which we have taken into consideration several key aspects to enhance recognition rate. The method performance has been studied on three datasets containing more than 2000 images from 129 different document classes. Lluís-Pere de las Heras, Oriol Ramos Terrades, Josep Lladós 0001, David Fernández Mota, Cristina Cañero Morales |
ICDAR | 3 |
| 2015 | Handwritten word spotting by inexact matching of grapheme graphsabstractThis paper presents a graph-based word spotting for handwritten documents. Contrary to most word spotting techniques, which use statistical representations, we propose a structural representation suitable to be robust to the inherent deformations of handwriting. Attributed graphs are constructed using a part-based approach. Graphemes extracted from shape convexities are used as stable units of handwriting, and are associated to graph nodes. Then, spatial relations between them determine graph edges. Spotting is defined in terms of an error-tolerant graph matching using bipartite-graph matching algorithm. To make the method usable in large datasets, a graph indexing approach that makes use of binary embeddings is used as preprocessing. Historical documents are used as experimental framework. The approach is comparable to statistical ones in terms of time and memory requirements, especially when dealing with large document collections. Pau Riba, Josep Lladós 0001, Alicia Fornés |
ICDAR | 2 |
| 2015 | Towards query-by-speech handwritten keyword spottingabstractIn this paper, we present a new querying paradigm for handwritten keyword spotting. We propose to represent handwritten word images both by visual and audio representations, enabling a query-by-speech keyword spotting system. The two representations are merged together and projected to a common sub-space in the training phase. This transform allows to, given a spoken query, retrieve word instances that were only represented by the visual modality. In addition, the same method can be used backwards at no additional cost to produce a handwritten text-to-speech system. We present our first results on this new querying mechanism using synthetic voices over the George Washington dataset. Marçal Rusiñol, David Aldavert, Ricardo Toledo, Josep Lladós 0001 |
ICDAR | 4 |
| 2015 | A comparative study of local detectors and descriptors for mobile document classificationabstractIn this paper we conduct a comparative study of local key-point detectors and local descriptors for the specific task of mobile document classification. A classification architecture based on direct matching of local descriptors is used as baseline for the comparative study. A set of four different key-point detectors and four different local descriptors are tested in all the possible combinations. The experiments are conducted in a database consisting of 30 model documents acquired on 6 different backgrounds, totaling more than 36.000 test images. Marçal Rusiñol, Joseph Chazalon, Jean-Marc Ogier, Josep Lladós 0001 |
ICDAR | 4 |
| 2014 | Sequential Word Spotting in Historical Handwritten DocumentsabstractIn this work we present a handwritten word spotting approach that takes advantage of the a priori known order of appearance of the query words. Given an ordered sequence of query word instances, the proposed approach performs a sequence alignment with the words in the target collection. Although the alignment is quite sparse, i.e. the number of words in the database is higher than the query set, the improvement in the overall performance is sensitively higher than isolated word spotting. As application dataset, we use a collection of handwritten marriage licenses taking advantage of the ordered index pages of family names. David Fernández Mota, R. Manmatha, Alicia Fornés, Josep Lladós 0001 |
Document Analysis Systems | 4 |
| 2014 | A Novel Learning-Free Word Spotting Approach Based on Graph RepresentationabstractEffective information retrieval on handwritten document images has always been a challenging task. In this paper, we propose a novel handwritten word spotting approach based on graph representation. The presented model comprises both topological and morphological signatures of handwriting. Skeleton-based graphs with the Shape Context labelled vertexes are established for connected components. Each word image is represented as a sequence of graphs. In order to be robust to the handwriting variations, an exhaustive merging process based on DTW alignment result is introduced in the similarity measure between word images. With respect to the computation complexity, an approximate graph edit distance approach using bipartite matching is employed for graph matching. The experiments on the George Washington dataset and the marriage records from the Barcelona Cathedral dataset demonstrate that the proposed approach outperforms the state-of-the-art structural methods. Peng Wang 0006, Véronique Eglin, Christophe Garcia, Christine Largeron, Josep Lladós 0001, Alicia Fornés |
Document Analysis Systems | 5 |
| 2013 | Near Convex Region Adjacency Graph and Approximate Neighborhood String Matching for Symbol Spotting in Graphical DocumentsabstractThis paper deals with a sub graph matching problem in Region Adjacency Graph (RAG) applied to symbol spotting in graphical documents. RAG is a very important, efficient and natural way of representing graphical information with a graph but this is limited to cases where the information is well defined with perfectly delineated regions. What if the information we are interested in is not confined within well defined regions? This paper addresses this particular problem and solves it by defining near convex grouping of oriented line segments which results in near convex regions. Pure convexity imposes hard constraints and can not handle all the cases efficiently. Hence to solve this problem we have defined a new type of convexity of regions, which allows convex regions to have concavity to some extend. We call this kind of regions Near Convex Regions (NCRs). These NCRs are then used to create the Near Convex Region Adjacency Graph (NCRAG) and with this representation we have formulated the problem of symbol spotting in graphical documents as a sub graph matching problem. For sub graph matching we have used the Approximate Edit Distance Algorithm (AEDA) on the neighborhood string, which starts working after finding a key node in the input or target graph and iteratively identifies similar nodes of the query graph in the neighborhood of the key node. The experiments are performed on artificial, real and distorted datasets. Anjan Dutta 0001, Josep Lladós 0001, Horst Bunke, Umapada Pal 0001 |
ICDAR | 2 |
| 2013 | Integrating Visual and Textual Cues for Query-by-String Word SpottingabstractIn this paper, we present a word spotting framework that follows the query-by-string paradigm where word images are represented both by textual and visual representations. The textual representation is formulated in terms of character n-grams while the visual one is based on the bag-of-visual-words scheme. These two representations are merged together and projected to a sub-vector space. This transform allows to, given a textual query, retrieve word instances that were only represented by the visual modality. Moreover, this statistical representation can be used together with state-of-the-art indexation structures in order to deal with large-scale scenarios. The proposed method is evaluated using a collection of historical documents outperforming state-of-the-art performances. David Aldavert, Marçal Rusiñol, Ricardo Toledo, Josep Lladós 0001 |
ICDAR | 4 |
| 2013 | Show-Through Cancellation and Image Enhancement by Multiresolution Contrast ProcessingabstractHistorical documents suffer from different types of degradation and noise such as background variation, uneven illumination or dark spots. In case of double-sided documents, another common problem is that the back side of the document usually interferes with the front side because of the transparency of the document or ink bleeding. This effect is called the show through phenomenon. Many methods are developed to solve these problems, and in the case of show-through, by scanning and matching both the front and back sides of the document. In contrast, our approach is designed to use only one side of the scanned document. We hypothesize that show-trough are low contrast components, while foreground components are high contrast ones. A Multiresolution Contrast (MC) decomposition is presented in order to estimate the contrast of features at different spatial scales. We cancel the show-through phenomenon by thresholding these low contrast components. This decomposition is also able to enhance the image removing shadowed areas by weighting spatial scales. Results show that the enhanced images improve the readability of the documents, allowing scholars both to recover unreadable words and to solve ambiguities. Alicia Fornés, Xavier Otazu, Josep Lladós 0001 |
ICDAR | 3 |
| 2013 | Key-Region Detection for Document Images - Application to Administrative Document RetrievalabstractIn this paper we argue that a key-region detector designed to take into account the special characteristics of document images can result in the detection of less and more meaningful key-regions. We propose a fast key-region detector able to capture aspects of the structural information of the document, and demonstrate its efficiency by comparing against standard detectors in an administrative document retrieval scenario. We show that using the proposed detector results to a smaller number of detected key-regions and higher performance without any drop in speed compared to standard state of the art detectors. Hongxing Gao, Marçal Rusiñol, Dimosthenis Karatzas, Josep Lladós 0001, Tomokazu Sato, Masakazu Iwamura, Koichi Kise |
ICDAR | 4 |
| 2013 | Unsupervised Wall Detector in Architectural Floor PlansabstractWall detection in floor plans is a crucial step in a complete floor plan recognition system. Walls define the main structure of buildings and convey essential information for the detection of other structural elements. Nevertheless, wall segmentation is a difficult task, mainly because of the lack of a standard graphical notation. The existing approaches are restricted to small group of similar notations or require the existence of pre-annotated corpus of input images to learn each new notation. In this paper we present an automatic wall segmentation system, with the ability to handle completely different notations without the need of any annotated dataset. It only takes advantage of the general knowledge that walls are a repetitive element, naturally distributed within the plan and commonly modeled by straight parallel lines. The method has been tested on four datasets of real floor plans with different notations, and compared with the state-of-the-art. The results show its suitability for different graphical notations, achieving higher recall rates than the rest of the methods while keeping a high average precision. Lluís-Pere de las Heras, David Fernández Mota, Ernest Valveny, Josep Lladós 0001, Gemma Sánchez |
ICDAR | 4 |
| 2011 | Interactive Trademark Image Retrieval by Fusing Semantic and Visual Content
Marçal Rusiñol, David Aldavert, Dimosthenis Karatzas, Ricardo Toledo, Josep Lladós 0001 |
ECIR | 5 |
| 2011 | Symbol Spotting in Line Drawings through Graph Paths HashingabstractIn this paper we propose a symbol spotting technique through hashing the shape descriptors of graph paths (Hamiltonian paths). Complex graphical structures in line drawings can be efficiently represented by graphs, which ease the accurate localization of the model symbol. Graph paths are the factorized substructures of graphs which enable robust recognition even in the presence of noise and distortion. In our framework, the entire database of the graphical documents is indexed in hash tables by the locality sensitive hashing (LSH) of shape descriptors of the paths. The hashing data structure aims to execute an approximate k-NN search in a sub-linear time. The spotting method is formulated by a spatial voting scheme to the list of locations of the paths that are decided during the hash table lookup process. We perform detailed experiments with various dataset of line drawings and the results demonstrate the effectiveness and efficiency of the technique. Anjan Dutta 0001, Josep Lladós 0001, Umapada Pal 0001 |
ICDAR | 2 |
| 2011 | The ICDAR 2011 Music Scores Competition: Staff Removal and Writer IdentificationabstractIn the last years, there has been a growing interest in the analysis of handwritten music scores. In this sense, our goal has been to foster the interest in the analysis of handwritten music scores by the proposal of two different competitions: Staff removal and Writer Identification. Both competitions have been tested on the CVC-MUSCIMA database: a ground-truth of handwritten music score images. This paper describes the competition details, including the dataset and ground-truth, the evaluation metrics, and a short description of the participants, their methods, and the obtained results. Alicia Fornés, Anjan Dutta 0001, Albert Gordo, Josep Lladós 0001 |
ICDAR | 4 |
| 2011 | Subgraph Spotting through Explicit Graph Embedding: An Application to Content Spotting in Graphic Document ImagesabstractWe present a method for spotting a subgraph in a graph repository. Subgraph spotting is a very interesting research problem for various application domains where the use of a relational data structure is mandatory. Our proposed method accomplishes subgraph spotting through graph embedding. We achieve automatic indexation of a graph repository during off-line learning phase, where we (i) break the graphs into 2-node sub graphs (a.k.a. cliques of order 2), which are primitive building-blocks of a graph, (ii) embed the 2-node sub graphs into feature vectors by employing our recently proposed explicit graph embedding technique, (iii) cluster the feature vectors in classes by employing a classic agglomerative clustering technique, (iv) build an index for the graph repository and (v) learn a Bayesian network classifier. The subgraph spotting is achieved during the on-line querying phase, where we (i) break the query graph into 2-node sub graphs, (ii) embed them into feature vectors, (iii) employ the Bayesian network classifier for classifying the query 2-node sub graphs and (iv) retrieve the respective graphs by looking-up in the index of the graph repository. The graphs containing all query 2-node sub graphs form the set of result graphs for the query. Finally, we employ the adjacency matrix of each result graph along with a score function, for spotting the query graph in it. The proposed subgraph spotting method is equally applicable to a wide range of domains, offering ease of query by example (QBE) and granularity of focused retrieval. Experimental results are presented for graphs generated from two repositories of electronic and architectural document images. Muhammad Muzzamil Luqman, Jean-Yves Ramel, Josep Lladós 0001, Thierry Brouard |
ICDAR | 3 |
| 2011 | Browsing Heterogeneous Document Collections by a Segmentation-Free Word Spotting MethodabstractIn this paper, we present a segmentation-free word spotting method that is able to deal with heterogeneous document image collections. We propose a patch-based framework where patches are represented by a bag-of-visual-words model powered by SIFT descriptors. A later refinement of the feature vectors is performed by applying the latent semantic indexing technique. The proposed method performs well on both handwritten and typewritten historical document images. We have also tested our method on documents written in non-Latin scripts. Marçal Rusiñol, David Aldavert, Ricardo Toledo, Josep Lladós 0001 |
ICDAR | 4 |
| 2010 | A framework for the assessment of text extraction algorithms on complex colour imagesabstractThe availability of open, ground-truthed datasets and clear performance metrics is a crucial factor in the development of an application domain. The domain of colour text image analysis (real scenes, Web and spam images, scanned colour documents) has traditionally suffered from a lack of a comprehensive performance evaluation framework. Such a framework is extremely difficult to specify, and corresponding pixel-level accurate information tedious to define. In this paper we discuss the challenges and technical issues associated with developing such a framework. Then, we describe a complete framework for the evaluation of text extraction methods at multiple levels, provide a detailed ground-truth specification and present a case study on how this framework can be used in a real-life situation. Antonio Clavelli, Dimosthenis Karatzas, Josep Lladós 0001 |
Document Analysis Systems | 3 |
| 2010 | A bag of notes approach to writer identification in old handwritten musical scoresabstractDetermining the authorship of a document, namely writer identification, can be an important source of information for document categorization. Contrary to text documents, the identification of the writer of graphical documents is still a challenge. In this paper we present a robust approach for writer identification in a particular kind of graphical documents, old music scores. This approach adapts the bag of visual terms method for coping with graphic documents. The identification is performed only using the graphical music notation. For this purpose, we generate a graphic vocabulary without recognizing any music symbols, and consequently, avoiding the difficulties in the recognition of hand-drawn symbols in old and degraded documents. The proposed method has been tested on a database of old music scores from the 17th to 19th centuries, achieving very high identification rates. Albert Gordo, Alicia Fornés, Ernest Valveny, Josep Lladós 0001 |
Document Analysis Systems | 4 |
| 2010 | Query driven word retrieval in graphical documentsabstractIn this paper, we present an approach towards the retrieval of words from graphical document images. In graphical documents, due to presence of multi-oriented characters in non-structured layout, word indexing is a challenging task. The proposed approach uses recognition results of individual components to form character pairs with the neighboring components. An indexing scheme is designed to store the spatial description of components and to access them efficiently. Given a query text word (ascii/unicode format), the character pairs present in it are searched in the document. Next the retrieved character pairs are linked sequentially to form character string. Dynamic programming is applied to find different instances of query words. A string edit distance is used here to match the query word as the objective function. Recognition of multi-scale and multi-oriented character component is done using Support Vector Machine classifier. To consider multi-oriented character strings the features used in the SVM are invariant to character orientation. Experimental results show that the method is efficient to locate a query word from multi-oriented text in graphical documents. Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001 |
Document Analysis Systems | 3 |
| 2010 | Efficient logo retrieval through hashing shape context descriptorsabstractIn this paper we present a method for organizing and indexing logo digital libraries like the ones of the patent and trademark offices. We propose an efficient queried-by-example retrieval system which is able to retrieve logos by similarity from large databases of logo images. Logos are compactly described by a variant of the shape context descriptor. These descriptors are then indexed by a locality-sensitive hashing data structure aiming to perform approximate k-NN search in high dimensional spaces in sub-linear time. The experiments demonstrate the effectiveness and efficiency of this system on realistic datasets as the Tobacco-800 logo database. Marçal Rusiñol, Josep Lladós 0001 |
Document Analysis Systems | 2 |
| 2009 | Graphological Analysis of Handwritten Text Documents for Human Resources RecruitmentabstractThe use of graphology in recruitment processes has become a popular tool in many human resources companies. This paper presents a model that links features from handwritten images to a number of personality characteristics used to measure applicant aptitudes for the job in a particular hiring scenario. In particular we propose a model of measuring active personality and leadership of the writer. Graphological features that define such a profile are measured in terms of document and script attributes like layout configuration, letter size, shape, slant and skew angle of lines, etc. After the extraction, data is classified using a neural network. An experimental framework with real samples has been constructed to illustrate the performance of the approach. Ricard Coll, Alicia Fornés, Josep Lladós 0001 |
ICDAR | 3 |
| 2009 | On the Use of Textural Features for Writer Identification in Old Handwritten Music ScoresabstractWriter identification consists in determining the writer of a piece of handwriting from a set of writers. In this paper we present a system for writer identification in old handwritten music scores which uses only music notation to determine the author. The steps of the proposed system are the following. First of all, the music sheet is preprocessed for obtaining a music score without the staff lines. Afterwards, four different methods for generating texture images from music symbols are applied. Every approach uses a different spatial variation when combining the music symbols to generate the textures. Finally, Gabor filters and Grey-scale Co-ocurrence matrices are used to obtain the features. The classification is performed using a k-NN classifier based on Euclidean distance. The proposed method has been tested on a database of old music scores from the 17th to 19th centuries, achieving encouraging identification rates. Alicia Fornés, Josep Lladós 0001, Gemma Sánchez, Horst Bunke |
ICDAR | 2 |
| 2009 | Seal Detection and Recognition: An Approach for Document IndexingabstractReliable indexing of documents having seal instances can be achieved by recognizing seal information. This paper presents a novel approach for detecting and classifying such multi-oriented seals in these documents. First, Hough Transform based methods are applied to extract the seal regions in documents. Next, isolated text characters within these regions are detected. Rotation and size invariant features and a Support Vector Machine based classifier have been used to recognize these detected text characters. Next, for each pair of character, we encode their relative spatial organization using their distance and angular position with respect to the centre of the seal, and enter this code into a hash table. Given an input seal, we recognize the individual text characters and compute the code for pair-wise character based on the relative spatial organization. The code obtained from the input seal helps to retrieve model hypothesis from the hash table. The seal model to which we get maximum hypothesis is selected for the recognition of the input seal. The methodology is tested to index seal in rotation and size invariant environment and we obtained encouraging results. Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001 |
ICDAR | 3 |
| 2009 | Multi-Oriented and Multi-Sized Touching Character Segmentation Using Dynamic ProgrammingabstractIn this paper, we present a scheme towards the segmentation of English multi-oriented touching strings into individual characters. When two or more characters touch, they generate a big cavity region at the background portion. Using Convex Hull information, we use these background information to find some initial points to segment a touching string into possible primitive segments (a primitive segment consists of a single character or a part of a character). Next these primitive segments are merged to get optimum segmentation and dynamic programming is applied using total likelihood of characters as the objective function. SVM classifier is used to find the likelihood of a character. To consider multi-oriented touching strings the features used in the SVM are invariant to character orientation. Circular ring and convex hull ring based approach has been used along with angular information of the contour pixels of the character to make the feature rotation invariant. From the experiment, we obtained encouraging results. Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001, Mathieu Delalandre |
ICDAR | 3 |
| 2009 | Logo Spotting by a Bag-of-words Approach for Document CategorizationabstractIn this paper we present a method for document categorization which processes incoming document images such as invoices or receipts. The categorization of these document images is done in terms of the presence of a certain graphical logo detected without segmentation. The graphical logos are described by a set of local features and the categorization of the documents is performed by the use of a bag-of-words model. Spatial coherence rules are added to reinforce the correct category hypothesis, aiming also to spot the logo inside the document image. Experiments which demonstrate the effectiveness of this system on a large set of real data are presented. Marçal Rusiñol, Josep Lladós 0001 |
ICDAR | 2 |
| 2008 | Performance Evaluation of Symbol Recognition and Spotting Systems: An OverviewabstractThis paper deals with the topic of performance evaluation of the symbol recognition & spotting systems. It presents an overview as a result of the work and the discussions undertaken by a working group on this subject. The paper starts by giving a general view of symbol recognition & spotting and performance evaluation. Next, the two main issues of performance evaluation are discussed: groundtruthing and performance characterization. Different problems related to both issues are addressed: groundtruthing of real documents, generation of synthetic documents, degradation models, the use of a priori knowledge, mapping of the groundtruth with the system results, and so on. Open problems arising from this overview are also discussed at the end of the paper. Mathieu Delalandre, Ernest Valveny, Josep Lladós 0001 |
Document Analysis Systems | 3 |
| 2008 | Writer Identification in Old Handwritten Music ScoresabstractThe aim of writer identification is determining the writer of a piece of handwriting from a set of writers. In this paper we present a system for writer identification in old handwritten music scores. Even though an important amount of compositions contains handwritten text in the music scores, the aim of our work is to use only music notation to determine the author. The steps of the system proposed are the following. First of all, the music sheet is preprocessed and normalized for obtaining a single binarized music line, without the staff lines. Afterwards, 100 features are extracted for every music line, which are subsequently used in a k-NN classifier that compares every feature vector with prototypes stored in a database. By applying feature selection and extraction methods on the original feature set, the performance is increased. The proposed method has been tested on a database of old music scores from the 17th to 19th centuries, achieving a recognition rate of about 95%. Alicia Fornés, Josep Lladós 0001, Gemma Sánchez, Horst Bunke |
Document Analysis Systems | 2 |
| 2008 | HistoSketch: A Semi-Automatic Annotation Tool for Archival DocumentsabstractThis article describes a sketch-based framework for semi-automatic annotation of historical document collections. It is motivated by the fact that fully automatic methods, while helpful for extracting metadata from large collections, have two main drawbacks in a real-world application: (i) they are error-prone and (ii) they only capture a subset of all the knowledge in the document base, both meaning that manual intervention is always required. Therefore, we have developed a practical framework for allowing experts to extract knowledge from document collections in a sketch-based scenario. The main possibilities of the proposed framework are: (a) browsing the collection efficiently, (b) providing gestures for metadata input, (c) supporting handwritten notes and (d) providing gestures for launching automatic extraction processes such as OCR or word spotting. Joan Mas Romeu, José A. Rodríguez 0001, Dimosthenis Karatzas, Gemma Sánchez, Josep Lladós 0001 |
Document Analysis Systems | 5 |
| 2008 | Multi-Oriented English Text Line Extraction Using Background and Foreground InformationabstractIn graphical documents (map, engineering drawing), artistic documents etc. there exist many printed materials where text lines are not parallel to each other and they are multi-oriented and curve in nature. For the OCR of such documents we need to extract individual text lines from the documents. Extraction of individual text lines from multi-oriented and/or curved text document is a difficult problem. In this paper, we propose a novel method to extract individual text lines from such document pages and the method is based on the foreground and background information of the characters of the text. To take care of background information, water reservoir concept is used here. In the proposed scheme at first, individual components are detected and grouped into 3-character clusters using their inter-component distance, size and positional information. Applying concept of graph, initial 3-character clusters are merged to have larger cluster group. Using inter-character background information, orientations of the extreme characters of a larger cluster are decided and based on these orientation, two candidate regions are formed from the cluster. Finally, with the help of these candidate regions, individual lines are extracted. From the experiment, we obtained encouraging result. Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001, Fumitaka Kimura |
Document Analysis Systems | 3 |
| 2008 | Word and Symbol Spotting Using Spatial Organization of Local DescriptorsabstractIn this paper we present a method to spot both text and graphical symbols in a collection of images of wiring diagrams. Word spotting and symbol spotting methods tend to use the most discriminative features to describe the objects to be located. This fact makes that one can not tackle with textual and symbolic information at the same time. We propose a spotting architecture able to index both words and symbols, inspired in off-the-shelf object recognition architectures. Keypoints are extracted from a document image and a local descriptor is computed at each of these points of interest. The spatial organization of these descriptors validate the hypothesis to find an object (text or symbol) in a certain location and under a certain pose. Marçal Rusiñol, Josep Lladós 0001 |
Document Analysis Systems | 2 |
| 2007 | Indexing Historical Documents by Word Shape SignaturesabstractIn this paper a word spotting approach to index archival image documents is presented. Indices are constructed from keyword images. The spotting strategy is formulated on an indexing-by-shape basis. The well known shape context descriptor is used to compute word image signatures from the skeleton points. Afterwards, codewords are extracted from thresholded shape contexts. It is a simpler and more compact representation based on bit vectors. Document images are roughly segmented into words and a lookup table is constructed. Each word subimage is taken as a bin. Keyword images are spotted into documents by a voting strategy consisting in indexing into the lookup table by codewords, and voting into the corresponding bins. The approach is illustrated by a real application scenario consisting of documents from a digital archive of the Spanish Civil War. Josep Lladós 0001, Gemma Sánchez |
ICDAR | 1 |
| 2007 | An Incremental On-line Parsing Algorithm for Recognizing Sketching DiagramsabstractThis paper presents a syntactic recognition approach for on-line drawn graphical symbols. The proposed method consists in an incremental on-line predictive parser based on symbol descriptions by an adjacency grammar. The parser analyzes input strokes as they are drawn by the user and is able to get ahead which symbols are likely to be recognized when a partial subshape is drawn in an intermediate state. In addition, the parser takes into account two issues. First, symbol strokes are drawn in any order by the user and second, since it is an on-line framework, the system requires real-time response. The method has been applied to an on-line sketching interface for architectural symbols. Joan Mas Romeu, Gemma Sánchez, Josep Lladós 0001, Bart Lamiroy |
ICDAR | 3 |
| 2007 | A Pen-Based Interface for Real-Time Document EditionabstractThis article presents a pen-based framework for manual edition of digital documents on tablet computers. In this system, the user draws certain proofreading symbols on the text parts to edit; some symbols can be accompanied by handwritten text. The input is interpreted and the corresponding editing action is executed in real time. The possibility that the input contains handwritten text is a novelty with respect to previous real-time systems that faced sketch-based edition, where usually text input is carried out via keyboard. Also, multimodal feedback mechanisms for error recovery are present. In this work we focus on the symbol recognition part. Different features are evaluated in recognition experiments using a support vector machine classifier. Experiments show that the symbol recognition is efficient enough for a real-time task and that the system can be used in real conditions with some experience. José A. Rodríguez 0001, Gemma Sánchez, Josep Lladós 0001 |
ICDAR | 3 |
| 2007 | Camera-Based Graphical Symbol DetectionabstractIn this paper we present a method to locate and recognize graphical symbols appearing in real images. A vectorial signature is defined to describe graphical symbols. It is formulated in terms of accumulated length and angular information computed from polygonal approximation of contours. The proposed method aims to locate and recognize graphical symbols in cluttered environments at the same time, without needing a segmentation step. The symbol signature is tolerant to rotation, scale, translation and to distortions such as weak perspective, blurring effect and illumination changes usually present when working with scenes acquired with low resolution cameras in open environments. Marçal Rusiñol, Josep Lladós 0001, Philippe Dosch |
ICDAR | 2 |
| 2006 | The Fuzzy-Spatial Descriptor for the Online Graphic Recognition: Overlapping Matrix Algorithm
Noorazrin Zakaria, Jean-Marc Ogier, Josep Lladós 0001 |
Document Analysis Systems | 3 |
| 2004 | A Platform to Extract Knowledge from Graphic Documents. Application to an Architectural Sketch Understanding Scenario
Gemma Sánchez, Ernest Valveny, Josep Lladós 0001, Joan Mas Romeu, Narcís Lozano |
Document Analysis Systems | 3 |
| 2001 | ICAR: Identity Card Automatic ReaderabstractThis paper describes the ICAR system, an application for automatic reading of identity cards and passports. The system acquires the image of the document by a flatbed scanner and recognizes the type of the document among a set of predefined models using color information. Textual fields are located in the image by a connected component analysis and identified in terms of their structural arrangement. A set of complementary statistical and structural OCR techniques are combined by a voting strategy to read each text image region. For unknown input documents, lines compliant with machine readable ICAO 9303 format are located and recognized. Although the system has been initially designed for Spanish documents, it allows the integration of new formats by a supervised learning procedure. The system is currently installed as a check-in application in a real environment. Josep Lladós 0001, Felipe Lumbreras, Vicente Chapaprieta, Joan Queralt |
ICDAR | 1 |
| 2001 | A Graph Grammar to Recognize Textured SymbolsabstractThis paper describes a graph grammar to modelize textured symbols in a graphics recognition framework. A textured symbol means a symbol consisting of repetitive structured patterns. We propose a method to infer a graph grammar from a structured texture detected in a document, and the subsequent parser to decide whether a symbol is accepted by the grammar. The grammar is based on a region adjacency graph representation of the vectorized document and the productions are based on the neighboring relations of the patterns forming the textured symbol. The syntactic framework is applied on an architectural plan understanding application. Gemma Sánchez, Josep Lladós 0001 |
ICDAR | 2 |
| 1999 | A Hough-based Method for Hatched Pattern Detection in Maps and DiagramsabstractA hatched area is characterized by a set of parallel straight lines placed at regular intervals. In this paper a Hough-based schema is introduced to recognize hatched areas in technical documents from attributed graph structures representing the document once it has been vectorized. Defining a Hough-based transform from a graph instead of the raster image allows to drastically reduce the processing time and second, to obtain more reliable results because straight lines have already been detected in the vectorization step. A second advantage of the proposed method is that no assumptions must be made a priori about the slope and frequency of hatching patterns, but they are computed in run time for each hatched area. Josep Lladós 0001, Enric Martí, Jaime López-Krahe |
ICDAR | 1 |