Alicia Fornés

dblp:76/4618 · also Alicia Fornés Bisquerra · DBLP profile ↗
← Back
85ranked-venue papers
11as first author
20since 2021 · last 2026
0000-0002-9692-5336ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 60 · 9 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 25 · 6 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Preserving privacy without compromising accuracy: Machine unlearning for handwritten text recognition
abstract
• Encoder-only Transformer baseline for handwriting text recognition (HTR). • Added writer-style head tracks & controls user-identifiable memorization. • Neural-activation pruning removes writer clues while preserving HTR accuracy. • Writer-ID Confusion (WIC) enforces uniform IDs, boosting efficient machine unlearning. • Membership-inference tests confirm privacy gains without costly retraining. Handwritten Text Recognition (HTR) is crucial for document digitization, but handwritten data can contain user-identifiable features, like unique writing styles, posing privacy risks. Regulations such as the “right to be forgotten” require models to remove these sensitive traces without full retraining. We introduce a practical encoder-only transformer baseline as a robust reference for future HTR research. Building on this, we propose a two-stage unlearning framework for multihead transformer HTR models. Our method combines neural pruning with machine unlearning applied to a writer classification head, ensuring sensitive information is removed while preserving the recognition head. We also present Writer-ID Confusion (WIC), a method that forces the forget set to follow a uniform distribution over writer identities, unlearning user-specific cues while maintaining text recognition performance. We compare WIC to Random Labeling, Fisher Forgetting, Amnesiac Unlearning, and DELETE within our prune-unlearn pipeline and consistently achieve better privacy and accuracy trade-offs. This is the first systematic study of machine unlearning for HTR. Using metrics such as Accuracy, Character Error Rate (CER), Word Error Rate (WER), and Membership Inference Attacks (MIA) on the IAM and CVL datasets, we demonstrate that our method achieves state-of-the-art or superior performance for effective unlearning. These experiments show that our approach effectively safeguards privacy without compromising accuracy, opening new directions for document analysis research. Our code is publicly available at https://github.com/leitro/WIC-WriterIDConfusion-MachineUnlearning .
Lei Kang 0002, Xuanshuo Fu, Lluís Gómez i Bigorda, Alicia Fornés, Ernest Valveny, Dimosthenis Karatzas
Pattern Recognit.4
2025 Recovering and Classifying Upper Limb Impairment Trajectories After Stroke
abstract
Upper limb impairment is a loss of motor function after a stroke, leading to difficulties in performing daily tasks. With a low remission rate six months after stroke, monitoring during this critical period is essential. Telemonitoring with inertial sensors has become a common approach that requires the identification and recognition of specific movements. However, data quality in clinical databases is a challenge, making data augmentation necessary. In this work, we propose a novel method that can learn a trajectory subspace from partial 3D signals in an unsupervised manner. Our method is simple yet effective, producing novel human-feasible motions compatible with the estimated trajectory subspace. To this end, three approaches are introduced from global to local models that can capture a wide variety of human motions. Our method outperforms the results in the state of the art in both control and patient subjects.
Martín Méndez-López, Alicia Fornés, Antonio Agudo
ICIP2
2025 UCL-MHTR: A unified continual learning system for multilingual handwritten text recognition
Marwa Dhiaf, Mohamed Ali Souibgui, Yousri Kessentini, Alicia Fornés, Ahmed Cheikhrouhou
Expert Syst. Appl.4
2024 GRIF-DM: Generation of Rich Impression Fonts Using Diffusion Models
abstract
Fonts are integral to creative endeavors, design processes, and artistic productions. The appropriate selection of a font can significantly enhance artwork and endow advertisements with a higher level of expressivity. Despite the availability of numerous diverse font designs online, traditional retrieval-based methods for font selection are increasingly being supplanted by generation-based approaches. These newer methods offer enhanced flexibility, catering to specific user preferences and capturing unique stylistic impressions. However, current impression font techniques based on Generative Adversarial Networks (GANs) necessitate the utilization of multiple auxiliary losses to provide guidance during generation. Furthermore, these methods commonly employ weighted summation for the fusion of impression-related keywords. This leads to generic vectors with the addition of more impression keywords, ultimately lacking in detail generation capacity. In this paper, we introduce a diffusion-based method, termed GRIF-DM, to generate fonts that vividly embody specific impressions, utilizing an input consisting of a single letter and a set of descriptive impression keywords. The core innovation of GRIF-DM lies in the development of dual cross-attention modules, which process the characteristics of the letters and impression keywords independently but synergistically, ensuring effective integration of both types of information. Our experimental results, conducted on the MyFonts dataset, affirm that this method is capable of producing realistic, vibrant, and high-fidelity fonts that are closely aligned with user specifications. This confirms the potential of our approach to revolutionize font generation by accommodating a broad spectrum of user-driven design requirements. Our code is publicly available at https://github.com/leitro/GRIF-DM.
Lei Kang 0002, Fei Yang 0004, Kai Wang 0060, Mohamed Ali Souibgui, Lluís Gómez i Bigorda, Alicia Fornés, Ernest Valveny, Dimosthenis Karatzas
ECAI6
2024 ICDAR 2024 Competition on Handwriting Recognition of Historical Ciphers
Alicia Fornés, Pau Torras, Carles Badal, Beáta Megyesi, Michelle Waldispühl, Nils Kopal, George Lasry
ICDAR (6)1
2024 A unified representation framework for the evaluation of Optical Music Recognition systems
abstract
Abstract Modern-day Optical Music Recognition (OMR) is a fairly fragmented field. Most OMR approaches use datasets that are independent and incompatible between each other, making it difficult to both combine them and compare recognition systems built upon them. In this paper we identify the need of a common music representation language and propose the Music Tree Notation format, with the idea to construct a common endpoint for OMR research that allows coordination, reuse of technology and fair evaluation of community efforts. This format represents music as a set of primitives that group together into higher-abstraction nodes, a compromise between the expression of fully graph-based and sequential notation formats. We have also developed a specific set of OMR metrics and a typeset score dataset as a proof of concept of this idea.
Pau Torras, Sanket Biswas, Alicia Fornés
Int. J. Document Anal. Recognit.3
2023 Text-DIAE: A Self-Supervised Degradation Invariant Autoencoder for Text Recognition and Document Enhancement
abstract
In this paper, we propose a Text-Degradation Invariant Auto Encoder (Text-DIAE), a self-supervised model designed to tackle two tasks, text recognition (handwritten or scene-text) and document image enhancement. We start by employing a transformer-based architecture that incorporates three pretext tasks as learning objectives to be optimized during pre-training without the usage of labelled data. Each of the pretext objectives is specifically tailored for the final downstream tasks. We conduct several ablation experiments that confirm the design choice of the selected pretext tasks. Importantly, the proposed model does not exhibit limitations of previous state-of-the-art methods based on contrastive losses, while at the same time requiring substantially fewer data samples to converge. Finally, we demonstrate that our method surpasses the state-of-the-art in existing supervised and self-supervised settings in handwritten and scene text recognition and document image enhancement. Our code and trained models will be made publicly available at https://github.com/dali92002/SSL-OCR
Mohamed Ali Souibgui, Sanket Biswas, Andrés Mafla, Ali Furkan Biten, Alicia Fornés, Yousri Kessentini, Josep Lladós 0001, Lluís Gómez i Bigorda, Dimosthenis Karatzas
AAAI5
2022 Musigraph: Optical Music Recognition Through Object Detection and Graph Neural Network
Arnau Baró, Pau Riba, Alicia Fornés
ICFHR3
2022 A Few Shot Multi-representation Approach for N-Gram Spotting in Historical Manuscripts
Giuseppe De Gregorio, Sanket Biswas, Mohamed Ali Souibgui, Asma Bensalah, Josep Lladós 0001, Alicia Fornés, Angelo Marcelli
ICFHR6
2022 DocEnTr: An End-to-End Document Image Enhancement Transformer
abstract
Document images can be affected by many degradation scenarios, which cause recognition and processing difficulties. In this age of digitization, it is important to denoise them for proper usage. To address this challenge, we present a new encoder-decoder architecture based on vision transformers to enhance both machine-printed and handwritten document images, in an end-to-end fashion. The encoder operates directly on the pixel patches with their positional information without the use of any convolutional layers, while the decoder reconstructs a clean image from the encoded patches. Conducted experiments show a superiority of the proposed model compared to the state-of-the-art methods on several DIBCO benchmarks. Code and models will be publicly available at: https://github.com/dali92002/DocEnTR.
Mohamed Ali Souibgui, Sanket Biswas, Sana Khamekhem Jemni, Yousri Kessentini, Alicia Fornés, Josep Lladós 0001, Umapada Pal 0001
ICPR5
2022 One-shot Compositional Data Generation for Low Resource Handwritten Text Recognition
abstract
Low resource Handwritten Text Recognition (HTR) is a hard problem due to the scarce annotated data and the very limited linguistic information (dictionaries and language models). For example, in the case of historical ciphered manuscripts, which are usually written with invented alphabets to hide the message contents. Thus, in this paper we address this problem through a data generation technique based on Bayesian Program Learning (BPL). Contrary to traditional generation approaches, which require a huge amount of annotated images, our method is able to generate human-like handwriting using only one sample of each symbol in the alphabet. After generating symbols, we create synthetic lines to train state-of-the-art HTR architectures in a segmentation free fashion. Quantitative and qualitative analyses were carried out and confirm the effectiveness of the proposed method.
Mohamed Ali Souibgui, Ali Furkan Biten, Sounak Dey, Alicia Fornés, Yousri Kessentini, Lluís Gómez i Bigorda, Dimosthenis Karatzas, Josep Lladós 0001
WACV4
2022 Advances in handwriting recognition
Utkarsh Porwal, Alicia Fornés, Faisal Shafait
Int. J. Document Anal. Recognit.2
2022 Content and Style Aware Generation of Text-Line Images for Handwriting Recognition
abstract
Handwritten Text Recognition has achieved an impressive performance in public benchmarks. However, due to the high inter- and intra-class variability between handwriting styles, such recognizers need to be trained using huge volumes of manually labeled training data. To alleviate this labor-consuming problem, synthetic data produced with TrueType fonts has been often used in the training loop to gain volume and augment the handwriting style variability. However, there is a significant style bias between synthetic and real data which hinders the improvement of recognition performance. To deal with such limitations, we propose a generative method for handwritten text-line images, which is conditioned on both visual appearance and textual content. Our method is able to produce long text-line samples with diverse handwriting styles. Once properly trained, our method can also be adapted to new target data by only accessing unlabeled text-line images to mimic handwritten styles and produce images with any textual content. Extensive experiments have been done on making use of the generated samples to boost Handwritten Text Recognition performance. Both qualitative and quantitative results demonstrate that the proposed approach outperforms the current state of the art.
Lei Kang 0002, Pau Riba, Marçal Rusiñol, Alicia Fornés, Mauricio Villegas
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Enhance to read better: A Multi-Task Adversarial Network for Handwritten Document Image Enhancement
Sana Khamekhem Jemni, Mohamed Ali Souibgui, Yousri Kessentini, Alicia Fornés
Pattern Recognit.4
2022 Pay attention to what you read: Non-recurrent handwritten text-Line recognition
Lei Kang 0002, Pau Riba, Marçal Rusiñol, Alicia Fornés, Mauricio Villegas
Pattern Recognit.4
2022 Table detection in business document images by message passing networks
Pau Riba, Lutz Goldmann, Oriol Ramos Terrades, Diede Rusticus, Alicia Fornés, Josep Lladós 0001
Pattern Recognit.5
2022 Few shots are all you need: A progressive learning approach for low resource handwritten text recognition
abstract
Handwritten text recognition in low resource scenarios, such as manuscripts with rare alphabets, is a challenging problem. In this paper, we propose a few-shot learning-based handwriting recognition approach that significantly reduces the human annotation process, by requiring only a few images of each alphabet symbols. The method consists of detecting all the symbols of a given alphabet in a textline image and decoding the obtained similarity scores to the final sequence of transcribed symbols. Our model is first pretrained on synthetic line images generated from an alphabet, which could differ from the alphabet of the target domain. A second training step is then applied to reduce the gap between the source and the target data. Since this retraining would require annotation of thousands of handwritten symbols together with their bounding boxes, we propose to avoid such human effort through an unsupervised progressive learning approach that automatically assigns pseudo-labels to the unlabeled data. The evaluation on different datasets shows that our model can lead to competitive results with a significant reduction in human effort. The code will be publicly available in the following repository: https://github.com/dali92002/HTRbyMatching
Mohamed Ali Souibgui, Alicia Fornés, Yousri Kessentini, Beáta Megyesi
Pattern Recognit. Lett.2
2021 Deep learning for graphics recognition: document understanding and beyond
Jean-Christophe Burie, Alicia Fornés, KC Santosh, Muhammad Muzzamil Luqman
Int. J. Document Anal. Recognit.2
2021 Candidate fusion: Integrating language modelling into a sequence-to-sequence handwritten word recognition architecture
Lei Kang 0002, Pau Riba, Mauricio Villegas, Alicia Fornés, Marçal Rusiñol
Pattern Recognit.4
2021 Learning graph edit distance by graph neural networks
Pau Riba, Andreas Fischer 0002, Josep Lladós 0001, Alicia Fornés
Pattern Recognit.4
2020 GANwriting: Content-Conditioned Generation of Styled Handwritten Word Images
Lei Kang 0002, Pau Riba, Yaxing Wang, Marçal Rusiñol, Alicia Fornés, Mauricio Villegas
ECCV (23)5
2020 Handwritten Historical Music Recognition by Sequence-to-Sequence with Attention Mechanism
abstract
Despite decades of research in Optical Music Recognition (OMR), the recognition of old handwritten music scores remains a challenge because of the variabilities in the handwriting styles, paper degradation, lack of standard notation, etc. Therefore, the research in OMR systems adapted to the particularities of old manuscripts is crucial to accelerate the conversion of music scores existing in archives into digital libraries, fostering the dissemination and preservation of our music heritage. In this paper we explore the adaptation of sequence-to-sequence models with attention mechanism (used in translation and handwritten text recognition) and the generation of specific synthetic data for recognizing old music scores. The experimental validation demonstrates that our approach is promising, especially when compared with long short-term memory neural networks.
Arnau Baró, Carles Badal, Alicia Fornés
ICFHR3
2020 Distilling Content from Style for Handwritten Word Recognition
abstract
Despite the latest transcription accuracies reached using deep neural network architectures, handwritten text recognition still remains a challenging problem, mainly because of the large inter-writer style variability. Both augmenting the training set with artificial samples using synthetic fonts, and writer adaptation techniques have been proposed to yield more generic approaches aimed at dodging style unevenness. In this work, we take a step closer to learn style independent features from handwritten word images. We propose a novel method that is able to disentangle the content and style aspects of input images by jointly optimizing a generative process and a handwritten word recognizer. The generator is aimed at transferring writing style features from one sample to another in an image-to-image translation approach, thus leading to a learned content-centric features that shall be independent to writing style attributes. Our proposed recognition model is able then to leverage such writer-agnostic features to reach better recognition performances. We advance over prior training strategies and demonstrate with qualitative and quantitative evaluations the performance of both the generative process and the recognition efficiency in the IAM dataset.
Lei Kang 0002, Pau Riba, Marçal Rusiñol, Alicia Fornés, Mauricio Villegas
ICFHR4
2020 Named Entity Recognition and Relation Extraction with Graph Neural Networks in Semi Structured Documents
abstract
The use of administrative documents to communicate and leave record of business information requires of methods able to automatically extract and understand the content from such documents in a robust and efficient way. In addition, the semi-structured nature of these reports is specially suited for the use of graph-based representations which are flexible enough to adapt to the deformations from the different document templates. Moreover, Graph Neural Networks provide the proper methodology to learn relations among the data elements in these documents. In this work we study the use of Graph Neural Network architectures to tackle the problem of entity recognition and relation extraction in semi-structured documents. Our approach achieves state of the art results in the three tasks involved in the process. Additionally, the experimentation with two datasets of different nature demonstrates the good generalization ability of our approach.
Manuel Carbonell, Pau Riba, Mauricio Villegas, Alicia Fornés, Josep Lladós 0001
ICPR4
2020 A Few-shot Learning Approach for Historical Ciphered Manuscript Recognition
abstract
Encoded (or ciphered) manuscripts are a special type of historical documents that contain encrypted text. The automatic recognition of this kind of documents is challenging because: 1) the cipher alphabet changes from one document to another, 2) there is a lack of annotated corpus for training and 3) touching symbols make the symbol segmentation difficult and complex. To overcome these difficulties, we propose a novel method for handwritten ciphers recognition based on few-shot object detection. Our method first detects all symbols of a given alphabet in a line image, and then a decoding step maps the symbol similarity scores to the final sequence of transcribed symbols. By training on synthetic data, we show that the proposed architecture is able to recognize handwritten ciphers with unseen alphabets. In addition, if few labeled pages with the same alphabet are used for fine tuning, our method surpasses existing unsupervised and supervised HTR methods for ciphers recognition.
Mohamed Ali Souibgui, Alicia Fornés, Yousri Kessentini, Crina Tudor
ICPR2
2020 Unsupervised Adaptation for Synthetic-to-Real Handwritten Word Recognition
abstract
Handwritten Text Recognition (HTR) is still a challenging problem because it must deal with two important difficulties: the variability among writing styles, and the scarcity of labelled data. To alleviate such problems, synthetic data generation and data augmentation are typically used to train HTR systems. However, training with such data produces encouraging but still inaccurate transcriptions in real words. In this paper, we propose an unsupervised writer adaptation approach that is able to automatically adjust a generic handwritten word recognizer, fully trained with synthetic fonts, towards a new incoming writer. We have experimentally validated our proposal using five different datasets, covering several challenges (i) the document source: modern and historic samples, which may involve paper degradation problems; (ii) different handwriting styles: single and multiple writer collections; and (iii) language, which involves different character combinations. Across these challenging collections, we show that our system is able to maintain its performance, thus, it provides a practical and generic approach to deal with new document collections without requiring any expensive and tedious manual annotation step.
Lei Kang 0002, Marçal Rusiñol, Alicia Fornés, Pau Riba, Mauricio Villegas
WACV3
2020 Hierarchical stochastic graphlet embedding for graph-based pattern recognition
abstract
Abstract Despite being very successful within the pattern recognition and machine learning community, graph-based methods are often unusable because of the lack of mathematical operations defined in graph domain. Graph embedding, which maps graphs to a vectorial space, has been proposed as a way to tackle these difficulties enabling the use of standard machine learning techniques. However, it is well known that graph embedding functions usually suffer from the loss of structural information. In this paper, we consider the hierarchical structure of a graph as a way to mitigate this loss of information. The hierarchical structure is constructed by topologically clustering the graph nodes and considering each cluster as a node in the upper hierarchical level. Once this hierarchical structure is constructed, we consider several configurations to define the mapping into a vector space given a classical graph embedding, in particular, we propose to make use of the stochastic graphlet embedding (SGE). Broadly speaking, SGE produces a distribution of uniformly sampled low-to-high-order graphlets as a way to embed graphs into the vector space. In what follows, the coarse-to-fine structure of a graph hierarchy and the statistics fetched by the SGE complements each other and includes important structural information with varied contexts. Altogether, these two techniques substantially cope with the usual information loss involved in graph embedding techniques, obtaining a more robust graph representation. This fact has been corroborated through a detailed experimental evaluation on various benchmark graph datasets, where we outperform the state-of-the-art methods.
Anjan Dutta 0001, Pau Riba, Josep Lladós 0001, Alicia Fornés
Neural Comput. Appl.4
2020 A neural model for text localization, transcription and named entity recognition in full pages
Manuel Carbonell, Alicia Fornés, Mauricio Villegas, Josep Lladós 0001
Pattern Recognit. Lett.2
2020 Hierarchical graphs for coarse-to-fine error tolerant matching
Pau Riba, Josep Lladós 0001, Alicia Fornés
Pattern Recognit. Lett.3
2019 Table Detection in Invoice Documents by Graph Neural Networks
abstract
Tabular structures in documents offer a complementary dimension to the raw textual data, representing logical or quantitative relationships among pieces of information. In digital mail room applications, where a large amount of administrative documents must be processed with reasonable accuracy, the detection and interpretation of tables is crucial. Table recognition has gained interest in document image analysis, in particular in unconstrained formats (absence of rule lines, unknown information of rows and columns). In this work, we propose a graph-based approach for detecting tables in document images. Instead of using the raw content (recognized text), we make use of the location, context and content type, thus it is purely a structure perception approach, not dependent on the language and the quality of the text reading. Our framework makes use of Graph Neural Networks (GNNs) in order to describe the local repetitive structural information of tables in invoice documents. Our proposed model has been experimentally validated in two invoice datasets and achieved encouraging results. Additionally, due to the scarcity of benchmark datasets for this task, we have contributed to the community a novel dataset derived from the RVL-CDIP invoice data. It will be publicly released to facilitate future research.
Pau Riba, Anjan Dutta 0001, Lutz Goldmann, Alicia Fornés, Oriol Ramos Terrades, Josep Lladós 0001
ICDAR4
2019 Training-Free and Segmentation-Free Word Spotting using Feature Matching and Query Expansion
abstract
Historical handwritten text recognition is an interesting yet challenging problem. In recent times, deep learning based methods have achieved significant performance in handwritten text recognition. However, handwriting recognition using deep learning needs training data, and often, text must be previously segmented into lines (or even words). These limitations constrain the application of HTR techniques in document collections, because training data or segmented words are not always available. Therefore, this paper proposes a training-free and segmentation-free word spotting approach that can be applied in unconstrained scenarios. The proposed word spotting framework is based on document query word expansion and relaxed feature matching algorithm, which can easily be parallelised. Since handwritten words posses distinct shape and characteristics, this work uses a combination of different keypoint detectors and Fourier-based descriptors to obtain a sufficient degree of relaxed matching. The effectiveness of the proposed method is empirically evaluated on well-known benchmark datasets using standard evaluation measures. The use of informative features along with query expansion significantly contributed in efficient performance of the proposed method.
Ekta Vats, Anders Hast, Alicia Fornés
ICDAR3
2019 Information extraction from historical handwritten document images with a context-aware neural model
Juan Ignacio Toledo, Manuel Carbonell, Alicia Fornés, Josep Lladós 0001
Pattern Recognit.3
2019 From Optical Music Recognition to Handwritten Music Recognition: A baseline
Arnau Baró, Pau Riba, Jorge Calvo-Zaragoza, Alicia Fornés
Pattern Recognit. Lett.4
2018 Joint Recognition of Handwritten Text and Named Entities with a Neural End-to-End Model
abstract
When extracting information from handwritten documents, text transcription and named entity recognition are usually faced as separate subsequent tasks. This has the disadvantage that errors in the first module affect heavily the performance of the second module. In this work we propose to do both tasks jointly, using a single neural network with a common architecture used for plain text recognition. Experimentally, the work has been tested on a collection of historical marriage records. Results of experiments are presented to show the effect on the performance for different configurations: different ways of encoding the information, doing or not transfer learning and processing at text line or multi-line region level. The results are comparable to state of the art reported in the ICDAR 2017 Information Extraction competition, even though the proposed technique does not use any dictionaries, language modeling or post processing.
Manuel Carbonell, Mauricio Villegas, Alicia Fornés, Josep Lladós 0001
DAS3
2018 Word-Hunter: A Gamesourcing Experience to Validate the Transcription of Historical Manuscripts
abstract
Nowadays, there are still many handwritten historical documents in archives waiting to be transcribed and indexed. Since manual transcription is tedious and time consuming, the automatic transcription seems the path to follow. However, the performance of current handwriting recognition techniques is not perfect, so a manual validation is mandatory. Crowdsourcing is a good strategy for manual validation, however it is a tedious task. In this paper we analyze experiences based in gamification in order to propose and design a gamesourcing framework that increases the interest of users. Then, we describe and analyze our experience when validating the automatic transcription using the gamesourcing application. Moreover, thanks to the combination of clustering and handwriting recognition techniques, we can speed up the validation while maintaining the performance.
Pau Riba, Alicia Fornés, Joan Mas Romeu, Josep Lladós 0001, Joana Maria Pujadas-Mora
ICFHR3
2018 Learning Graph Distances with Message Passing Neural Networks
abstract
Graph representations have been widely used in pattern recognition thanks to their powerful representation formalism and rich theoretical background. A number of error-tolerant graph matching algorithms such as graph edit distance have been proposed for computing a distance between two labelled graphs. However, they typically suffer from a high computational complexity, which makes it difficult to apply these matching algorithms in a real scenario. In this paper, we propose an efficient graph distance based on the emerging field of geometric deep learning. Our method employs a message passing neural network to capture the graph structure and learns a metric with a siamese network approach. The performance of the proposed graph distance is validated in two application cases, graph classification and graph retrieval of handwritten words, and shows a promising performance when compared with (approximate) graph edit distance benchmarks.
Pau Riba, Andreas Fischer 0002, Josep Lladós 0001, Alicia Fornés
ICPR4
2017 Pyramidal Stochastic Graphlet Embedding for Document Pattern Classification
abstract
Document pattern classification methods using graphs have received a lot of attention because of its robust representation paradigm and rich theoretical background. However, the way of preserving and the process for delineating documents with graphs introduce noise in the rendition of underlying data, which creates instability in the graph representation. To deal with such unreliability in representation, in this paper, we propose Pyramidal Stochastic Graphlet Embedding (PSGE). Given a graph representing a document pattern, our method first computes a graph pyramid by successively reducing the base graph. Once the graph pyramid is computed, we apply Stochastic Graphlet Embedding (SGE) for each level of the pyramid and combine their embedded representation to obtain a global delineation of the original graph. The consideration of pyramid of graphs rather than just a base graph extends the representational power of the graph embedding, which reduces the instability caused due to noise and distortion. When plugged with support vector machine, our proposed PSGE has outperformed the state-of-the-art results in recognition of handwritten words as well as graphical symbols.
Anjan Dutta 0001, Pau Riba, Josep Lladós 0001, Alicia Fornés
ICDAR4
2017 ICDAR2017 Competition on Information Extraction in Historical Handwritten Records
abstract
The extraction of relevant information from historical handwritten document collections is one of the key steps in order to make these manuscripts available for access and searches. In this competition, the goal is to detect the named entities and assign each of them a semantic category, and therefore, to simulate the filling in of a knowledge database. This paper describes the dataset, the tasks, the evaluation metrics, the participants methods and the results.
Alicia Fornés, Verónica Romero 0001, Arnau Baró, Juan Ignacio Toledo, Joan-Andreu Sánchez, Enrique Vidal 0001, Josep Lladós 0001
ICDAR1
2017 Improving Information Retrieval in Multiwriter Scenario by Exploiting the Similarity Graph of Document Terms
abstract
Information Retrieval (IR) is the activity of obtaining information resources relevant to a questioned information. It usually retrieves a set of objects ranked according to the relevancy to the needed fact. In document analysis, information retrieval receives a lot of attention in terms of symbol and word spotting. However, through decades the community mostly focused either on printed or on single writer scenario, where the state-of-the-art results have achieved reasonable performance on the available datasets. Nevertheless, the existing algorithms do not perform accordingly on multiwriter scenario. A graph representing relations between a set of objects is a structure where each node delineates an individual element and the similarity between them is represented as a weight on the connecting edge. In this paper, we explore different analytics of graphs constructed from words or graphical symbols, such as diffusion, shortest path, etc. to improve the performance of information retrieval methods in multiwriter scenario.
Pau Riba, Anjan Dutta 0001, Sounak Dey, Josep Lladós 0001, Alicia Fornés
ICDAR5
2017 Handwriting Recognition by Attribute Embedding and Recurrent Neural Networks
abstract
Handwriting recognition consists in obtaining the transcription of a text image. Recent word spotting methods based on attribute embedding have shown good performance when recognizing words. However, they are holistic methods in the sense that they recognize the word as a whole (i.e. they find the closest word in the lexicon to the word image). Consequently, these kinds of approaches are not able to deal with out of vocabulary words, which are common in historical manuscripts. Also, they cannot be extended to recognize text lines. In order to address these issues, in this paper we propose a handwriting recognition method that adapts the attribute embedding to sequence learning. Concretely, the method learns the attribute embedding of patches of word images with a convolutional neural network. Then, these embeddings are presented as a sequence to a recurrent neural network that produces the transcription. We obtain promising results even without the use of any kind of dictionary or language model.
Juan Ignacio Toledo, Sounak Dey, Alicia Fornés, Josep Lladós 0001
ICDAR3
2017 Large-scale graph indexing using binary embeddings of node contexts for information spotting in document image databases
Pau Riba, Josep Lladós 0001, Alicia Fornés, Anjan Dutta 0001
Pattern Recognit. Lett.3
2016 A Segmentation-Free Handwritten Word Spotting Approach by Relaxed Feature Matching
abstract
The automatic recognition of historical handwritten documents is still considered a challenging task. For this reason, word spotting emerges as a good alternative for making the information contained in these documents available to the user. Word spotting is defined as the task of retrieving all instances of the query word in a document collection, becoming a useful tool for information retrieval. In this paper we propose a segmentation-free word spotting approach able to deal with large document collections. Our method is inspired on feature matching algorithms that have been applied to image matching and retrieval. Since handwritten words have different shape, there is no exact transformation to be obtained. However, the sufficient degree of relaxation is achieved by using a Fourier based descriptor and an alternative approach to RANSAC called PUMA. The proposed approach is evaluated on historical marriage records, achieving promising results.
Anders Hast, Alicia Fornés
DAS2
2016 An Interactive Transcription System of Census Records Using Word-Spotting Based Information Transfer
abstract
This paper presents a system to assist in the transcription of historical handwritten census records in a crowdsourcing platform. Census records have a tabular structured layout. They consist in a sequence of rows with information of homes ordered by street address. For each household snippet in the page, the list of family members is reported. The censuses are recorded in intervals of a few years and the information of individuals in each household is quite stable from a point in time to the next one. This redundancy is used to assist the transcriber, so the redundant information is transferred from the census already transcribed to the next one. Household records are aligned from one year to the next one using the knowledge of the ordering by street address. Given an already transcribed census, a query by string word spotting is applied. Thus, names from the census in time t are used as queries in the corresponding home record in time t+1. Since the search is constrained, the obtained precision-recall values are very high, with an important reduction in the transcription time. The proposed system has been tested in a real citizen-science experience where non expert users transcribe the census data of their home town.
Joan Mas Romeu, Alicia Fornés, Josep Lladós 0001
DAS2
2016 Election Tally Sheets Processing System
abstract
In paper based elections, manual tallies at polling station level produce myriads of documents. These documents share a common form-like structure and a reduced vocabulary worldwide. On the other hand, each tally sheet is filled by a different writer and on different countries, different scripts are used. We present a complete document analysis system for electoral tally sheet processing combining state of the art techniques with a new handwriting recognition subprocess based on unsupervised feature discovery with Variational Autoencoders and sequence classification with BLSTM neural networks. The whole system is designed to be script independent and allows a fast and reliable results consolidation process with reduced operational cost.
Juan Ignacio Toledo, Alicia Fornés, Jordi Cucurull-Juan, Josep Lladós 0001
DAS2
2016 Towards the Recognition of Compound Music Notes in Handwritten Music Scores
abstract
The recognition of handwritten music scores still remains an open problem. The existing approaches can only deal with very simple handwritten scores mainly because of the variability in the handwriting style and the variability in the composition of groups of music notes (i.e. compound music notes). In this work we focus on this second problem and propose a method based on perceptual grouping for the recognition of compound music notes. Our method has been tested using several handwritten music scores of the CVC-MUSCIMA database and compared with a commercial Optical Music Recognition (OMR) software. Given that our method is learning-free, the obtained results are promising.
Arnau Baró, Pau Riba, Alicia Fornés
ICFHR3
2016 Using the MGGI Methodology for Category-Based Language Modeling in Handwritten Marriage Licenses Books
abstract
Handwritten marriage licenses books have been used for centuries by ecclesiastical and secular institutions to register marriages. The information contained in these historical documents is useful for demography studies and genealogical research, among others. Despite the generally simple structure of the text in these documents, automatic transcription and semantic information extraction is difficult due to the distinct and evolutionary vocabulary, which is composed mainly of proper names that change along the time. In previous works we studied the use of category-based language models to both improve the automatic transcription accuracy and make easier the extraction of semantic information. Here we analyze the main causes of the semantic errors observed in previous results and apply a Grammatical Inference technique known as MGGI to improve the semantic accuracy of the language model obtained. Using this language model, full handwritten text recognition experiments have been carried out, with results supporting the interest of the proposed approach.
Verónica Romero 0001, Alicia Fornés, Enrique Vidal 0001, Joan-Andreu Sánchez
ICFHR2
2015 Hidden Markov model topology optimization for handwriting recognition
abstract
In this paper we present a method to optimize the topology of linear left-to-right hidden Markov models. These models are very popular for sequential signals modeling on tasks such as handwriting recognition. Many topology definition methods select the number of states for a character model based on character length. This can be a drawback when characters are shorter than the minimum allowed by the model, since they can not be properly trained nor recognized. The proposed method optimizes the number of states per model by automatically including convenient skip-state transitions and therefore it avoids the aforementioned problem. We discuss and compare our method with other character length-based methods such the Fixed, Bakis and Quantile methods. Our proposal performs well on off-line handwriting recognition task.
Núria Cirera, Alicia Fornés, Josep Lladós 0001
ICDAR2
2015 Handwritten word spotting by inexact matching of grapheme graphs
abstract
This paper presents a graph-based word spotting for handwritten documents. Contrary to most word spotting techniques, which use statistical representations, we propose a structural representation suitable to be robust to the inherent deformations of handwriting. Attributed graphs are constructed using a part-based approach. Graphemes extracted from shape convexities are used as stable units of handwriting, and are associated to graph nodes. Then, spatial relations between them determine graph edges. Spotting is defined in terms of an error-tolerant graph matching using bipartite-graph matching algorithm. To make the method usable in large datasets, a graph indexing approach that makes use of binary embeddings is used as preprocessing. Historical documents are used as experimental framework. The approach is comparable to statistical ones in terms of time and memory requirements, especially when dealing with large document collections.
Pau Riba, Josep Lladós 0001, Alicia Fornés
ICDAR3
2014 Sequential Word Spotting in Historical Handwritten Documents
abstract
In this work we present a handwritten word spotting approach that takes advantage of the a priori known order of appearance of the query words. Given an ordered sequence of query word instances, the proposed approach performs a sequence alignment with the words in the target collection. Although the alignment is quite sparse, i.e. the number of words in the database is higher than the query set, the improvement in the overall performance is sensitively higher than isolated word spotting. As application dataset, we use a collection of handwritten marriage licenses taking advantage of the ordered index pages of family names.
David Fernández Mota, R. Manmatha, Alicia Fornés, Josep Lladós 0001
Document Analysis Systems3
2014 A Novel Learning-Free Word Spotting Approach Based on Graph Representation
abstract
Effective information retrieval on handwritten document images has always been a challenging task. In this paper, we propose a novel handwritten word spotting approach based on graph representation. The presented model comprises both topological and morphological signatures of handwriting. Skeleton-based graphs with the Shape Context labelled vertexes are established for connected components. Each word image is represented as a sequence of graphs. In order to be robust to the handwriting variations, an exhaustive merging process based on DTW alignment result is introduced in the similarity measure between word images. With respect to the computation complexity, an approximate graph edit distance approach using bipartite matching is employed for graph matching. The experiments on the George Washington dataset and the marriage records from the Barcelona Cathedral dataset demonstrate that the proposed approach outperforms the state-of-the-art structural methods.
Peng Wang 0006, Véronique Eglin, Christophe Garcia, Christine Largeron, Josep Lladós 0001, Alicia Fornés
Document Analysis Systems6
2014 On the Influence of Key Point Encoding for Handwritten Word Spotting
abstract
In this paper we evaluate the influence of the selection of key points and the associated features in the performance of word spotting processes. In general, features can be extracted from a number of characteristic points like corners, contours, skeletons, maxima, minima, crossings, etc. A number of descriptors exist in the literature using different interest point detectors. But the intrinsic variability of handwriting vary strongly on the performance if the interest points are not stable enough. In this paper, we analyze the performance of different descriptors for local interest points. As benchmarking dataset we have used the Barcelona Marriage Database that contains handwritten records of marriages over five centuries.
David Fernández Mota, Pau Riba, Alicia Fornés, Josep Lladós 0001
ICFHR3
2014 e-Crowds: A Mobile Platform for Browsing and Searching in Historical Demography-Related Manuscripts
abstract
This paper presents a prototype system running on portable devices for browsing and word searching through historical handwritten document collections. The platform adapts the paradigm of eBook reading, where the narrative is not necessarily sequential, but centered on the user actions. The novelty is to replace digitally born books by digitized historical manuscripts of marriage licenses, so document analysis tasks are required in the browser. With an active reading paradigm, the user can cast queries of people names, so he/she can implicitly follow genealogical links. In addition, the system allows combined searches: the user can refine a search by adding more words to search. As a second contribution, the retrieval functionality involves as a core technology a word spotting module with an unified approach, which allows combined query searches, and also two input modalities: query-by-example, and query-by-string.
Pau Riba, Jon Almazán, Alicia Fornés, David Fernández Mota, Ernest Valveny, Josep Lladós 0001
ICFHR3
2014 BH2M: The Barcelona Historical, Handwritten Marriages Database
abstract
This paper presents an image database of historical handwritten marriages records stored in the archives of Barcelona cathedral, and the corresponding meta-data addressed to evaluate the performance of document analysis algorithms. The contribution of this paper is twofold. First, it presents a complete ground truth which covers the whole pipeline of handwriting recognition research, from layout analysis to recognition and understanding. Second, it is the first dataset in the emerging area of genealogical document analysis, where documents are manuscripts pseudo-structured with specific lexicons and the interest is beyond pure transcriptions but context dependent.
David Fernández Mota, Jon Almazán, Núria Cirera, Alicia Fornés, Josep Lladós 0001
ICPR4
2014 A Coarse-to-Fine Word Spotting Approach for Historical Handwritten Documents Based on Graph Embedding and Graph Edit Distance
abstract
Effective information retrieval on handwritten document images has always been a challenging task, especially historical ones. In the paper, we propose a coarse-to-fine handwritten word spotting approach based on graph representation. The presented model comprises both the topological and morphological signatures of the handwriting. Skeleton-based graphs with the Shape Context labelled vertexes are established for connected components. Each word image is represented as a sequence of graphs. Aiming at developing a practical and efficient word spotting approach for large-scale historical handwritten documents, a fast and coarse comparison is first applied to prune the regions that are not similar to the query based on the graph embedding methodology. Afterwards, the query and regions of interest are compared by graph edit distance based on the Dynamic Time Warping alignment. The proposed approach is evaluated on a public dataset containing 50 pages of historical marriage license records. The results show that the proposed approach achieves a compromise between efficiency and accuracy.
Peng Wang 0006, Véronique Eglin, Christophe Garcia, Christine Largeron, Josep Lladós 0001, Alicia Fornés
ICPR6
2014 A graph-based approach for segmenting touching lines in historical handwritten documents
David Fernández Mota, Josep Lladós 0001, Alicia Fornés
Int. J. Document Anal. Recognit.3
2014 Word Spotting and Recognition with Embedded Attributes
abstract
This paper addresses the problems of word spotting and word recognition on images. In word spotting, the goal is to find all instances of a query word in a dataset of images. In recognition, the goal is to recognize the content of the word image, usually aided by a dictionary or lexicon. We describe an approach in which both word images and text strings are embedded in a common vectorial subspace. This is achieved by a combination of label embedding and attributes learning, and a common subspace regression. In this subspace, images and strings that represent the same word are close together, allowing one to cast recognition and retrieval tasks as a nearest neighbor problem. Contrary to most other existing methods, our representation has a fixed length, is low dimensional, and is very fast to compute and, especially, to compare. We test our approach on four public datasets of both handwritten documents and natural images showing results comparable or better than the state-of-the-art on spotting and recognition tasks.
Jon Almazán, Albert Gordo, Alicia Fornés, Ernest Valveny
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Segmentation-free word spotting with exemplar SVMs
Jon Almazán, Albert Gordo, Alicia Fornés, Ernest Valveny
Pattern Recognit.3
2013 Handwritten Word Spotting with Corrected Attributes
abstract
We propose an approach to multi-writer word spotting, where the goal is to find a query word in a dataset comprised of document images. We propose an attributes-based approach that leads to a low-dimensional, fixed-length representation of the word images that is fast to compute and, especially, fast to compare. This approach naturally leads to an unified representation of word images and strings, which seamlessly allows one to indistinctly perform query-by-example, where the query is an image, and query-by-string, where the query is a string. We also propose a calibration scheme to correct the attributes scores based on Canonical Correlation Analysis that greatly improves the results on a challenging dataset. We test our approach on two public datasets showing state-of-the-art results.
Jon Almazán, Albert Gordo, Alicia Fornés, Ernest Valveny
ICCV3
2013 Deformable HOG-Based Shape Descriptor
abstract
In this paper we deal with the problem of recognizing handwritten shapes. We present a new deformable feature extraction method that adapts to the shape to be described, dealing in this way with the variability introduced in the handwriting domain. It consists in a selection of the regions that best define the shape to be described, followed by the computation of histograms of oriented gradients-based features over these points. Our results significantly outperform other descriptors in the literature for the task of hand-drawn shape recognition and handwritten word retrieval.
Jon Almazán, Alicia Fornés, Ernest Valveny
ICDAR2
2013 Show-Through Cancellation and Image Enhancement by Multiresolution Contrast Processing
abstract
Historical documents suffer from different types of degradation and noise such as background variation, uneven illumination or dark spots. In case of double-sided documents, another common problem is that the back side of the document usually interferes with the front side because of the transparency of the document or ink bleeding. This effect is called the show through phenomenon. Many methods are developed to solve these problems, and in the case of show-through, by scanning and matching both the front and back sides of the document. In contrast, our approach is designed to use only one side of the scanned document. We hypothesize that show-trough are low contrast components, while foreground components are high contrast ones. A Multiresolution Contrast (MC) decomposition is presented in order to estimate the contrast of features at different spatial scales. We cancel the show-through phenomenon by thresholding these low contrast components. This decomposition is also able to enhance the image removing shadowed areas by weighting spatial scales. Results show that the enhanced images improve the readability of the documents, allowing scholars both to recover unreadable words and to solve ambiguities.
Alicia Fornés, Xavier Otazu, Josep Lladós 0001
ICDAR1
2013 ICDAR 2013 Music Scores Competition: Staff Removal
abstract
The first competition on music scores that was organized at ICDAR in 2011 awoke the interest of researchers, who participated both at staff removal and writer identification tasks. In this second edition, we focus on the staff removal task and simulate a real case scenario: old music scores. For this purpose, we have generated a new set of images using two kinds of degradations: local noise and 3D distortions. This paper describes the dataset, distortion methods, evaluation metrics, the participant's methods and the obtained results.
Muriel Visani, Van Cuong Kieu, Alicia Fornés, Nicholas Journet
ICDAR3
2013 Writer identification in handwritten musical scores with bags of notes
Albert Gordo, Alicia Fornés, Ernest Valveny
Pattern Recognit.2
2013 The ESPOSALLES database: An ancient marriage license corpus for off-line handwriting recognition
Verónica Romero 0001, Alicia Fornés, Joan-Andreu Sánchez, Alejandro H. Toselli, Volkmar Frinken, Enrique Vidal 0001, Josep Lladós 0001
Pattern Recognit.2
2012 Efficient Exemplar Word Spotting
abstract
In this paper we propose an unsupervised segmentation-free method for word spotting in document images. Documents are represented with a grid of HOG descriptors, and a sliding window approach is used to locate the document regions that are most similar to the query. We use the Exemplar SVM framework to produce a better representation of the query in an unsupervised way. Finally, the document descriptors are precomputed and compressed with Product Quantization. This offers two advantages: first, a large number of documents can be kept in RAM memory at the same time. Second, the sliding window becomes significantly faster since distances between quantized HOG descriptors can be precomputed. Our results significantly outperform other segmentation-free methods in the literature, both in accuracy and in speed and memory usage.
Jon Almazán, Albert Gordo, Alicia Fornés, Ernest Valveny
BMVC3
2012 A Coarse-to-Fine Approach for Handwritten Word Spotting in Large Scale Historical Documents Collection
abstract
In this paper we propose an approach for word spotting in handwritten document images. We state the problem from a focused retrieval perspective, i.e. locating instances of a query word in a large scale dataset of digitized manuscripts. We combine two approaches, namely one based on word segmentation and another one segmentation-free. The first approach uses a hashing strategy to coarsely prune word images that are unlikely to be instances of the query word. This process is fast but has a low precision due to the errors introduced in the segmentation step. The regions containing candidate words are sent to the second process based on a state of the art technique from the visual object detection field. This discriminative model represents the appearance of the query word and computes a similarity score. In this way we propose a coarse-to-fine approach achieving a compromise between efficiency and accuracy. The validation of the model is shown using a collection of old handwritten manuscripts. We appreciate a substantial improvement in terms of precision regarding the previous proposed method with a low computational cost increase.
Jon Almazán, David Fernández Mota, Alicia Fornés, Josep Lladós 0001, Ernest Valveny
ICFHR3
2012 On Influence of Line Segmentation in Efficient Word Segmentation in Old Manuscripts
abstract
The objective of this work is to show the importance of a good line segmentation to obtain better results in the segmentation of words of historical documents. We have used the approach developed by Manmatha and Rothfeder to segment words in old handwritten documents. In their work the lines of the documents are extracted using projections. In this work, we have developed an approach to segment lines more efficiently. The new line segmentation algorithm tackles with skewed, touching and noisy lines, so it is significantly improves word segmentation. Experiments using Spanish documents from the Marriages Database of the Barcelona Cathedral show that this approach reduces the error rate by more than 20%.
Josep Lladós 0001, Alicia Fornés, R. Manmatha
ICFHR3
2012 CVC-MUSCIMA: a ground truth of handwritten music score images for writer identification and staff removal
Alicia Fornés, Anjan Dutta 0001, Albert Gordo, Josep Lladós 0001
Int. J. Document Anal. Recognit.1
2012 On the Influence of Word Representations for Handwritten Word Spotting in Historical Documents
abstract
Word spotting is the process of retrieving all instances of a queried keyword from a digital library of document images. In this paper we evaluate the performance of different word descriptors to assess the advantages and disadvantages of statistical and structural models in a framework of query-by-example word spotting in historical documents. We compare four word representation models, namely sequence alignment using DTW as a baseline reference, a bag of visual words approach as statistical model, a pseudo-structural model based on a Loci features representation, and a structural approach where words are represented by graphs. The four approaches have been tested with two collections of historical data: the George Washington database and the marriage records from the Barcelona Cathedral. We experimentally demonstrate that statistical representations generally give a better performance, however it cannot be neglected that large descriptors are difficult to be implemented in a retrieval scenario where word spotting requires the indexation of data with million word images.
Josep Lladós 0001, Marçal Rusiñol, Alicia Fornés, David Fernández Mota, Anjan Dutta 0001
Int. J. Pattern Recognit. Artif. Intell.3
2012 A non-rigid appearance model for shape description and recognition
Jon Almazán, Alicia Fornés, Ernest Valveny
Pattern Recognit.2
2011 A Non-rigid Feature Extraction Method for Shape Recognition
abstract
This paper presents a methodology for shape recognition that focuses on dealing with the difficult problem of large deformations. The proposed methodology consists in a novel feature extraction technique, which uses a non-rigid representation adaptable to the shape. This technique employs a deformable grid based on the computation of geometrical centroids that follows a region partitioning algorithm. Then, a feature vector is extracted by computing pixel density measures around these geometrical centroids. The result is a shape descriptor that adapts its representation to the given shape and encodes the pixel density distribution. The validity of the method when dealing with large deformations has been experimentally shown over datasets composed of handwritten shapes. It has been applied to signature verification and shape recognition tasks demonstrating high accuracy and low computational cost.
Jon Almazán, Alicia Fornés, Ernest Valveny
ICDAR2
2011 The ICDAR 2011 Music Scores Competition: Staff Removal and Writer Identification
abstract
In the last years, there has been a growing interest in the analysis of handwritten music scores. In this sense, our goal has been to foster the interest in the analysis of handwritten music scores by the proposal of two different competitions: Staff removal and Writer Identification. Both competitions have been tested on the CVC-MUSCIMA database: a ground-truth of handwritten music score images. This paper describes the competition details, including the dataset and ground-truth, the evaluation metrics, and a short description of the participants, their methods, and the obtained results.
Alicia Fornés, Anjan Dutta 0001, Albert Gordo, Josep Lladós 0001
ICDAR1
2011 Co-training for Handwritten Word Recognition
abstract
To cope with the tremendous variations of writing styles encountered between different individuals, unconstrained automatic handwriting recognition systems need to be trained on large sets of labeled data. Traditionally, the training data has to be labeled manually, which is a laborious and costly process. Semi-supervised learning techniques offer methods to utilize unlabeled data, which can be obtained cheaply in large amounts in order, to reduce the need for labeled data. In this paper, we propose the use of Co-Training for improving the recognition accuracy of two weakly trained handwriting recognition systems. The first one is based on Recurrent Neural Networks while the second one is based on Hidden Markov Models. On the IAM off-line handwriting database we demonstrate a significant increase of the recognition accuracy can be achieved with Co-Training for single word recognition.
Volkmar Frinken, Andreas Fischer 0002, Horst Bunke, Alicia Fornés
ICDAR4
2011 Circular Blurred Shape Model for Multiclass Symbol Recognition
abstract
In this paper, we propose a circular blurred shape model descriptor to deal with the problem of symbol detection and classification as a particular case of object recognition. The feature extraction is performed by capturing the spatial arrangement of significant object characteristics in a correlogram structure. The shape information from objects is shared among correlogram regions, where a prior blurring degree defines the level of distortion allowed in the symbol, making the descriptor tolerant to irregular deformations. Moreover, the descriptor is rotation invariant by definition. We validate the effectiveness of the proposed descriptor in both the multiclass symbol recognition and symbol detection domains. In order to perform the symbol detection, the descriptors are learned using a cascade of classifiers. In the case of multiclass categorization, the new feature space is learned using a set of binary classifiers which are embedded in an error-correcting output code design. The results over four symbol data sets show the significant improvements of the proposed descriptor compared to the state-of-the-art descriptors. In particular, the results are even more significant in those cases where the symbols suffer from elastic deformations.
Sergio Escalera, Alicia Fornés, Oriol Pujol, Josep Lladós 0001, Petia Radeva
IEEE Trans. Syst. Man Cybern. Part B2
2010 A bag of notes approach to writer identification in old handwritten musical scores
abstract
Determining the authorship of a document, namely writer identification, can be an important source of information for document categorization. Contrary to text documents, the identification of the writer of graphical documents is still a challenge. In this paper we present a robust approach for writer identification in a particular kind of graphical documents, old music scores. This approach adapts the bag of visual terms method for coping with graphic documents. The identification is performed only using the graphical music notation. For this purpose, we generate a graphic vocabulary without recognizing any music symbols, and consequently, avoiding the difficulties in the recognition of hand-drawn symbols in old and degraded documents. The proposed method has been tested on a database of old music scores from the 17th to 19th centuries, achieving very high identification rates.
Albert Gordo, Alicia Fornés, Ernest Valveny, Josep Lladós 0001
Document Analysis Systems2
2010 A Symbol-Dependent Writer Identification Approach in Old Handwritten Music Scores
abstract
Writer identification consists in determining the writer of a piece of handwriting from a set of writers. In this paper we introduce a symbol-dependent approach for identifying the writer of old music scores, which is based on two symbol recognition methods. The main idea is to use the Blurred Shape Model descriptor and a DTW-based method for detecting, recognizing and describing the music clefs and notes. The proposed approach has been evaluated in a database of old music scores, achieving very high writer identification rates.
Alicia Fornés, Josep Lladós 0001
ICFHR1
2010 An Efficient Staff Removal Approach from Printed Musical Documents
abstract
Staff removal is an important preprocessing step of the Optical Music Recognition (OMR). The process aims to remove the stafflines from a musical document and retain only the musical symbols, later these symbols are used effectively to identify the music information. This paper proposes a simple but robust method to remove stafflines from printed musical scores. In the proposed methodology we have considered a staffline segment as a horizontal linkage of vertical black runs with uniform height. We have used the neighbouring properties of a staffline segment to validate it as a true segment. We have considered the dataset along with the deformations described in for evaluation purpose. From experimentation we have got encouraging results.
Anjan Dutta 0001, Umapada Pal 0001, Alicia Fornés, Josep Lladós 0001
ICPR3
2010 Symbol Classification Using Dynamic Aligned Shape Descriptor
abstract
Shape representation is a difficult task because of several symbol distortions, such as occlusions, elastic deformations, gaps or noise. In this paper, we propose a new descriptor and distance computation for coping with the problem of symbol recognition in the domain of Graphical Document Image Analysis. The proposed D-Shape descriptor encodes the arrangement information of object parts in a circular structure, allowing different levels of distortion. The classification is performed using a cyclic Dynamic Time Warping based method, allowing distortions and rotation. The methodology has been validated on different data sets, showing very high recognition rates.
Alicia Fornés, Sergio Escalera, Josep Lladós 0001, Ernest Valveny
ICPR1
2010 Rotation invariant hand-drawn symbol recognition based on a dynamic time warping model
Alicia Fornés, Josep Lladós 0001, Gemma Sánchez, Dimosthenis Karatzas
Int. J. Document Anal. Recognit.1
2010 A combination of features for symbol-independent writer identification in old music scores
Alicia Fornés, Josep Lladós 0001, Gemma Sánchez, Xavier Otazu, Horst Bunke
Int. J. Document Anal. Recognit.1
2009 Graphological Analysis of Handwritten Text Documents for Human Resources Recruitment
abstract
The use of graphology in recruitment processes has become a popular tool in many human resources companies. This paper presents a model that links features from handwritten images to a number of personality characteristics used to measure applicant aptitudes for the job in a particular hiring scenario. In particular we propose a model of measuring active personality and leadership of the writer. Graphological features that define such a profile are measured in terms of document and script attributes like layout configuration, letter size, shape, slant and skew angle of lines, etc. After the extraction, data is classified using a neural network. An experimental framework with real samples has been constructed to illustrate the performance of the approach.
Ricard Coll, Alicia Fornés, Josep Lladós 0001
ICDAR2
2009 On the Use of Textural Features for Writer Identification in Old Handwritten Music Scores
abstract
Writer identification consists in determining the writer of a piece of handwriting from a set of writers. In this paper we present a system for writer identification in old handwritten music scores which uses only music notation to determine the author. The steps of the proposed system are the following. First of all, the music sheet is preprocessed for obtaining a music score without the staff lines. Afterwards, four different methods for generating texture images from music symbols are applied. Every approach uses a different spatial variation when combining the music symbols to generate the textures. Finally, Gabor filters and Grey-scale Co-ocurrence matrices are used to obtain the features. The classification is performed using a k-NN classifier based on Euclidean distance. The proposed method has been tested on a database of old music scores from the 17th to 19th centuries, achieving encouraging identification rates.
Alicia Fornés, Josep Lladós 0001, Gemma Sánchez, Horst Bunke
ICDAR1
2009 Circular Blurred Shape Model for symbol spotting in documents
abstract
Symbol spotting problem requires feature extraction strategies able to generalize from training samples and to localize the target object while discarding most part of the image. In the case of document analysis, symbol spotting techniques have to deal with a high variability of symbols' appearance. In this paper, we propose the Circular Blurred Shape Model descriptor. Feature extraction is performed capturing the spatial arrangement of significant object characteristics in a correlogram structure. Shape information from objects is shared among correlogram regions, being tolerant to the irregular deformations. Descriptors are learnt using a cascade of classifiers and Abadoost as the base classifier. Finally, symbol spotting is performed by means of a windowing strategy using the learnt cascade over plan and old musical score documents. Spotting and multi-class categorization results show better performance comparing with the state-of-the-art descriptors.
Sergio Escalera, Alicia Fornés, Oriol Pujol, Alberto Escudero, Petia Radeva
ICIP2
2009 Blurred Shape Model for binary and grey-level symbol recognition
Sergio Escalera, Alicia Fornés, Oriol Pujol, Petia Radeva, Gemma Sánchez, Josep Lladós 0001
Pattern Recognit. Lett.2
2008 Writer Identification in Old Handwritten Music Scores
abstract
The aim of writer identification is determining the writer of a piece of handwriting from a set of writers. In this paper we present a system for writer identification in old handwritten music scores. Even though an important amount of compositions contains handwritten text in the music scores, the aim of our work is to use only music notation to determine the author. The steps of the system proposed are the following. First of all, the music sheet is preprocessed and normalized for obtaining a single binarized music line, without the staff lines. Afterwards, 100 features are extracted for every music line, which are subsequently used in a k-NN classifier that compares every feature vector with prototypes stored in a database. By applying feature selection and extraction methods on the original feature set, the performance is increased. The proposed method has been tested on a database of old music scores from the 17th to 19th centuries, achieving a recognition rate of about 95%.
Alicia Fornés, Josep Lladós 0001, Gemma Sánchez, Horst Bunke
Document Analysis Systems1
2007 Multi-class Binary Object Categorization Using Blurred Shape Models
Sergio Escalera, Alicia Fornés, Oriol Pujol, Josep Lladós 0001, Petia Radeva
CIARP2