VLDB 2026 Research / reviewers in the wild / expert
Alejandro H. Toselli
dblp:49/942 · also Alejandro Héctor Toselli, Alejandro Héctor Toselli Rossi
· DBLP profile ↗
76ranked-venue papers
13as first author
17since 2021 · last 2026
0000-0001-6955-9249ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 10 first-author · 16 since 2021Databases, data management, data science and information retrieval · 30 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Simple handwritten text recognition techniques for highly accurate writer identificationabstractAbstract The theater of the Spanish Early Modern period encompasses thousands of textual works and hundreds of playwrights and is one of the greatest examples of Spanish literature. These works were usually copied and altered, changing sometimes the meaning of the original works written by the authors. Therefore, identifying autograph testimonies written directly by the playwrights themselves are of particular importance. In this paper, we address an approach for writer identification based on Deep Convolutional-Recurrent Neural Networks and n -gram language models. The identification task is posed as a classification problem, introducing a probabilistic framework that goes beyond the plain transcription of handwritten text. Experiments are conducted to validate our proposal to distinguish between Lope de Vega ’s manuscripts and other non-Lope hands. The good results achieved will ultimately allow to provide modern researchers with a useful tool for cultural heritage recovery. Alejandro H. Toselli, Álvaro Cuéllar, Sònia Boadas, Enrique Vidal 0001, Joan-Andreu Sánchez |
Pattern Anal. Appl. | 1 |
| 2025 | The PARES Database: Information Extraction over Historical Parish RecordsabstractAbstract Historical census records convey information that is key to perform genealogical research and demographic studies. Given the large number of documents of this type that exist, it is crucial to research methods that allow the automatic extraction of information from this type of document. In this work, we present a new corpus of this kind, comprising 535 historical census tables from French archives. Alongside this dataset, we have assessed three different baseline methods for information extraction. The first two methods employ a traditional sequential approach, where table rows are detected before extracting information. The third baseline uses an end-to-end model that directly extracts information from the table images without prior row detection. Our results demonstrate the effectiveness of all three baselines in tackling the information extraction task. José Andrés, Casey Wall, Solène Tarride, Mickaël Coustaty, Alejandro H. Toselli, Enrique Vidal 0001 |
Int. J. Document Anal. Recognit. | 5 |
| 2024 | Mining and Analyzing Statistical Information from Untranscribed Form Images
José Andrés, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR (5) | 2 |
| 2024 | BRESSAY: A Brazilian Portuguese Dataset for Offline Handwritten Text Recognition
Arthur Flor de Sousa Neto, Byron L. D. Bezerra, Sávio S. Araújo, Wiliane M. A. S. Souza, Kléberson F. Alves, Macileide F. Oliveira, Samara V. S. Lins, Hugo J. F. Hazin, Pedro H. V. Rocha, Alejandro H. Toselli |
ICDAR (2) | 10 |
| 2024 | ICDAR 2024 Competition on Handwritten Text Recognition in Brazilian Essays - BRESSAY
Arthur Flor de Sousa Neto, Byron L. D. Bezerra, Sávio S. Araújo, Wiliane M. A. S. Souza, Kléberson F. Alves, Macileide F. Oliveira, Samara V. S. Lins, Hugo J. F. Hazin, Pedro H. V. Rocha, Alejandro H. Toselli |
ICDAR (6) | 10 |
| 2024 | Zipf Curves and Basic Text Analytics from Untranscribed Manuscript Images
Enrique Vidal 0001, Alejandro H. Toselli |
ICDAR (3) | 2 |
| 2024 | What distinguishes conspiracy from critical narratives? A computational analysis of oppositional discourseabstractAbstract The current prevalence of conspiracy theories on the internet is a significant issue, tackled by many computational approaches. However, these approaches fail to recognize the relevance of distinguishing between texts which contain a conspiracy theory and texts which are simply critical and oppose mainstream narratives. Furthermore, little attention is usually paid to the role of inter‐group conflict in oppositional narratives. We contribute by proposing a novel topic‐agnostic annotation scheme that differentiates between conspiracies and critical texts, and that defines span‐level categories of inter‐group conflict. We also contribute with the multilingual XAI‐DisInfodemics corpus (English and Spanish), which contains a high‐quality annotation of Telegram messages related to COVID‐19 (5000 messages per language). We also demonstrate the feasibility of an NLP‐based automatization by performing a range of experiments that yield strong baseline solutions. Finally, we perform an analysis which demonstrates that the promotion of intergroup conflict and the presence of violence and anger are key aspects to distinguish between the two types of oppositional narratives, that is, conspiracy versus critical. Damir Korencic, Berta Chulvi, Xavier Bonet Casals, Alejandro H. Toselli, Mariona Taulé, Paolo Rosso |
Expert Syst. J. Knowl. Eng. | 4 |
| 2024 | Segmenting large historical notarial manuscripts into multi-page deedsabstractAbstract Archives around the world hold vast digitized series of historical manuscript books or “bundles” containing, among others, notarial records also known as “deeds” or “acts”. One of the first steps to provide metadata which describe the contents of those bundles is to segment them into their individual deeds. Even if deeds are often page-aligned, as in the bundles considered in the present work, this is a time-consuming task, often prohibitive given the huge scale of the manuscript series involved. Unlike traditional Layout Analysis methods for page-level segmentation, our approach goes beyond the realm of a single-page image, providing consistent deed detection results on full bundles. This is achieved in two tightly integrated steps: first, we estimate the class-posterior at the page level for the “initial”, “middle”, and “final” classes; then we “decode” these posteriors applying a series of sequentiality consistency constraints to obtain a consistent book segmentation. Experiments are presented for four large historical manuscripts, varying the number of “deeds” used for training. Two metrics are introduced to assess the quality of book segmentation, one of them taking into account the loss of information entailed by segmentation errors. The problem formalization, the metrics and the empirical work significantly extend our previous works on this topic. José Ramón Prieto, David Camilo Becerra Romero, Alejandro H. Toselli, Carlos Alonso, Enrique Vidal 0001 |
Pattern Anal. Appl. | 3 |
| 2023 | Search for Hyphenated Words in Probabilistic Indices: A Machine Learning Approach
José Andrés, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR (1) | 2 |
| 2023 | Lexicon-based probabilistic indexing of handwritten text imagesabstractAbstract Keyword Spotting (KWS) is here considered as a basic technology for Probabilistic Indexing (PrIx) of large collections of handwritten text images to allow fast textual access to the contents of these collections. Under this perspective, a probabilistic framework for lexicon-based KWS in text images is presented. The presentation aims at providing formal insights which help understanding classical statements of KWS (from which PrIx borrows fundamental concepts), as well as the relative challenges entailed by these statements. The development of the proposed framework makes it clear that word recognition or classification implicitly or explicitly underlies any formulation of KWS. Moreover, it suggests that the same statistical models and training methods successfully used for handwriting text recognition can advantageously be used also for PrIx, even though PrIx does not generally require or rely on any kind of previously produced image transcripts. Experiments carried out using these approaches support the consistency and the general interest of the proposed framework. Results on three datasets traditionally used for KWS benchmarking are significantly better than those previously published for these datasets. In addition, good results are also reported on two new, larger handwritten text image datasets (B entham and P lantas ), showing the great potential of the methods proposed in this paper for indexing and textual search in large collections of untranscribed handwritten documents. Specifically, we achieved the following Average Precision values: IAMDB: 0.89, G eorge W ashington : 0.91, P arzival : 0.95, B entham : 0.91 and P lantas : 0.92. Enrique Vidal 0001, Alejandro H. Toselli, Joan Puigcerver |
Neural Comput. Appl. | 2 |
| 2023 | End-to-End page-Level assessment of handwritten text recognitionabstractThe evaluation of Handwritten Text Recognition (HTR) systems has traditionally used metrics based on the edit distance between HTR and ground truth (GT) transcripts, at both the character and word levels. This is very adequate when the experimental protocol assumes that both GT and HTR text lines are the same, which allows edit distances to be independently computed to each given line. Driven by recent advances in pattern recognition, HTR systems increasingly face the end-to-end page-level transcription of a document, where the precision of locating the different text lines and their corresponding reading order (RO) play a key role. In such a case, the standard metrics do not take into account the inconsistencies that might appear. In this paper, the problem of evaluating HTR systems at the page level is introduced in detail. We analyse the convenience of using a two-fold evaluation, where the transcription accuracy and the RO goodness are considered separately. Different alternatives are proposed, analysed and empirically compared both through partially simulated and through real, full end-to-end experiments. Results support the validity of the proposed two-fold evaluation approach. An important conclusion is that such an evaluation can be adequately achieved by just two simple and well-known metrics: the Word Error Rate (WER), that takes transcription sequentiality into account, and the here re-formulated Bag of Words Word Error Rate (bWER), that ignores order. While the latter directly and very accurately assess intrinsic word recognition errors, the difference between both metrics (ΔWER) gracefully correlates with the Normalised Spearman’s Foot Rule Distance (NSFD), a metric which explicitly measures RO errors associated with layout analysis flaws. To arrive to these conclusions, we have introduced another metric called Hungarian Word Word Rate (hWER), based on a here proposed regularised version of the Hungarian Algorithm. This metric is shown to be always almost identical to bWER and both bWER and hWER are also almost identical to WER whenever HTR transcripts and GT references are guarantee to be in the same RO. Enrique Vidal 0001, Alejandro H. Toselli, Antonio Ríos-Vila, Jorge Calvo-Zaragoza |
Pattern Recognit. | 2 |
| 2023 | Open set classification of untranscribed handwritten text image documentsabstractContent-based classification of manuscripts is an important task that is generally carried out by expert archivists. Nevertheless, many historical manuscript collections are so vast that in most cases this task is hardly feasible, even for large, well staffed archives. Nowadays, manuscripts are generally preserved in the form of sets of digital images. Therefore, the technical problem we are interested in is automatic classification of “‘image documents”, each consisting of a set of untranscribed handwritten text images, by the textual contents of the images. The traditional Pattern Recognition classification paradigm does provide the basic tools to deal with this problem. However, in practice, the set of relevant classes of a large documental series is seldom known in advance. Therefore, a classifier trained with a predefined set of classes will systematically fail when new image documents arrive which do not belong to any of the classes assumed in training. Here we adopt the “Open Set Classification” framework to extend and consolidate our previous work on image document classification in order to adequately handle new documents from unknown classes. The proposed approaches are based on a relatively novel technology for text image representation known as “probabilistic indexing”, which proves very effective to characterise the intrinsic word-level uncertainty exhibited by historical handwritten text images. We assess the performance of this approach on a moderately sized but representative dataset extracted from a huge series of complex notarial manuscripts from the Spanish Archivo Histórico Provincial de Cádiz, with good results. José Ramón Prieto, Juan José Flores Arellano, Enrique Vidal 0001, Alejandro H. Toselli |
Pattern Recognit. Lett. | 4 |
| 2022 | Approximate Search for Keywords in Handwritten Text Images
José Andrés, Alejandro H. Toselli, Enrique Vidal 0001 |
DAS | 2 |
| 2022 | A robust handwritten recognition system for learning on different data restriction scenarios
Arthur Flor de Sousa Neto, Byron L. D. Bezerra, Alejandro H. Toselli, Estanislau Lima |
Pattern Recognit. Lett. | 3 |
| 2021 | Probabilistic Indexing and Search for Hyphenated Words
Enrique Vidal 0001, Alejandro H. Toselli |
ICDAR (2) | 2 |
| 2021 | ICDAR 2021 Competition on Components Segmentation Task of Document Photos
Celso A. M. Lopes Junior, Ricardo Batista das Neves Junior, Byron L. D. Bezerra, Alejandro H. Toselli, Donato Impedovo |
ICDAR (4) | 4 |
| 2021 | Digital Editions as Distant Supervision for Layout Analysis of Printed Books
Alejandro H. Toselli, David A. Smith |
ICDAR (2) | 1 |
| 2020 | HTR-Flor++: A Handwritten Text Recognition System Based on a Pipeline of Optical and Language ModelsabstractOffline Handwritten Text Recognition (HTR) is a task that offers a challenge in computer vision, where images are the only source of information. In fact, several approaches to optical models have been developed, such as through of Hidden Markov Model (HMM) or recurrent Bidirectional/Multidimensional layers. The current state-of-the-art consists of combined deep learning techniques, the Convolutional Recurrent Neural Networks (CRNN), in which recurrent layers still suffer from vanishing gradient problem when processing very long texts. In a way, high-performance models generally have millions of trainable parameters and a high computational cost. However, recently a new optical model architecture, Gated-CNN, demonstrated improvements to complement CRNN modeling. Thus, in this work, we present a new small architecture for HTR (based on Gated-CNN) integrated with two steps of language model at the character and word levels, respectively. Therefore, we used 9 state-of-the-art approaches and validated the results using the IAM public dataset. Finally, the proposed model surpasses the results obtained by different approaches in the literature, reaching recognition rates of CER 2.7% and WER 5.6%, which means an improvement of 13% over the best results on IAM dataset. Arthur Flor de Sousa Neto, Byron L. D. Bezerra, Alejandro H. Toselli, Estanislau Lima |
DocEng | 3 |
| 2020 | The Carabela Project and Manuscript Collection: Large-Scale Probabilistic Indexing and Content-based ClassificationabstractThe main aim of the Carabela project was to develop and apply techniques that allow textual searching on massive Spanish collections of 15th-19th century manuscripts. The project focused on a relatively small subset of 125 000 images of collections of interest to underwater archaeology. For this type of manuscripts, state-of-the-art automatic transcription techniques, generally fail to achieve usable transcription accuracy. Therefore, rather than insisting in actual transcription, methodologies for probabilistic indexing of handwritten text images have been adopted. This has allowed us to effectively cope with the intrinsically high degree of uncertainty of the text contained in most historical manuscripts, leading to highly effective systems for textual search and retrieval. Carabela has gone one step further by developing new techniques to classify probabilistically indexed, but otherwise untranscribed, text images according to their textual content. These techniques have been successfully used to automatically classify Carabela bundels (each containing hundreds or thousands of pages) according to their “level of risk” of public exposure, in order to control their access and avoid as much as possible the plundering of Spanish underwater heritage. Enrique Vidal 0001, Verónica Romero 0001, Alejandro H. Toselli, Joan-Andreu Sánchez, Vicente Bosch, Lorenzo Quirós, José-Miguel Benedí, José Ramón Prieto, Moisés Pastor, Francisco Casacuberta, Carlos Alonso, Carmen García, Lourdes Márquez, Carmen Orcero |
ICFHR | 3 |
| 2020 | HU-PageScan: a fully convolutional neural network for document page cropabstractNovember The offer of online, automated, and impersonal services demand users to upload scanned copies of their documents to the organisations. As a consequence of this decentralisation, the documents present more challenges to the already complex process of image processing and information extraction. To address this problem, the authors presented an optimised fully convolutional neural network model for document segmentation that works on mobile devices to detect the region of the document in the captured image. They performed experiments in three representative datasets comparing the proposed method with the Geodesic object Proposals, U‐net, Mask R‐CNN, and OctHU‐PageScan algorithms. They also compared the proposed model with all competitors of the ICDAR2015 Competition on smartphone document capture. Furthermore, they performed a qualitative and comparative analysis with the CamScanner software, a popular app for Android and iOS smartphones used for more than 100 million users in over 200 countries. The proposed approach achieved a significant performance compared with the current state‐of‐the‐art methods, providing a powerful approach for document segmentation in photos and scanned images. Ricardo Batista das Neves Junior, Estanislau Lima, Byron L. D. Bezerra, Cleber Zanchettin, Alejandro H. Toselli |
IET Image Process. | 5 |
| 2019 | Music Symbol Sequence Indexing in Medieval Plainchant ManuscriptsabstractHuge amounts of musical manuscripts are preserved in cathedrals, abbeys, and archives. However, without reliable transcripts, their contents are inaccessible. Manual transcription is unaffordable for large collections, and current automatic technologies-such as Optical Music Recognition or Handwritten Music Recognition-do not provide sufficient accuracy for a fully-automatic scenario. In many cases, perfect transcripts are not really needed, given that content-based search with some degree of reliability would already be extremely useful. Spotting just single music symbols is rather useless (most of the symbols generally appear in all pages); instead, helpful search targets are melodic patterns, which typically correspond to music symbol sequences. We explore approaches for accurate retrieval of melodic patterns, represented by music symbol sequences, from collections of Medieval plainchant manuscripts. Our statistical framework, based on the use of convolutional recurrent neural networks and probabilistic indices, is shown to be useful for retrieving music patterns which appear frequently in this untranscribed images, yielding an Average Precision of 86 %. Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001, Joan-Andreu Sánchez |
ICDAR | 2 |
| 2019 | Making Two Vast Historical Manuscript Collections Searchable and Extracting Meaningful Textual Features Through Large-Scale Probabilistic IndexingabstractTextual access to large collections of digitized images remains unfeasible because usually they lack transcripts. Transcribing such collections is in turn typically unattainable in terms of costs. However, the use of probabilistic indices can facilitate textual accessing with only moderate demands of resources. Besides allowing effortless information retrieval, it will be shown that probabilistic indices can also be used to estimate textual features of the indexed but otherwise untranscribed collections, such as running words and Zipf's curves. Complete probabilistic indices have been recently produced for two iconic large collections: "Bentham" (90K images) and "Spanish Golden Age Theater" (40K images). To show the repercussion of making these collections searchable, we provide accessing statistics gathered through their corresponding search interfaces. To the best of our knowledge this is the first publication of large collections of untranscribed manuscripts which are now publicly accessible for effective and efficient textual access. Alejandro H. Toselli, Verónica Romero 0001, Joan-Andreu Sánchez, Enrique Vidal 0001 |
ICDAR | 1 |
| 2019 | Hybrid hidden Markov models and artificial neural networks for handwritten music recognition in mensural notation
Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001 |
Pattern Anal. Appl. | 2 |
| 2019 | Probabilistic multi-word spotting in handwritten text images
Alejandro H. Toselli, Enrique Vidal 0001, Joan Puigcerver, Ernesto Noya-García |
Pattern Anal. Appl. | 1 |
| 2019 | A set of benchmarks for Handwritten Text Recognition on historical documents
Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Mauricio Villegas, Enrique Vidal 0001 |
Pattern Recognit. | 3 |
| 2019 | Handwritten Music Recognition for Mensural notation with convolutional recurrent neural networks
Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 2 |
| 2018 | Automatic Alignment of Handwritten Images and Transcripts for Training Handwritten Text Recognition SystemsabstractState-of-the-art Handwritten Text Recognition techniques are based on statistical models such as hidden Markov models or recurrent neural networks for optical modeling of characters and N-grams for language modeling. These models are trained using well known, learning techniques: Expectation-Maximization, backpropagation, etc. Therefore, training data is needed to build these models. In the case of the optical models the training data consist of text line images with their corresponding transcripts. When the transcript of a handwritten document is available, putting in correspondence automatically the physical lines in the images with the lines of the transcripts is not an easy task. We present a method for automatically aligning handwritten text images and their respective transcripts. The approach automatically segments the images into lines and then recognizes them. An alignment confidence is obtained using the Levenshtein distance between the recognition results and the transcripts. The most confident lines are then used for training. Experiments carried out using a historical document present encouraging results. Verónica Romero 0001, Alejandro H. Toselli, Vicente Bosch, Joan-Andreu Sánchez, Enrique Vidal 0001 |
DAS | 2 |
| 2018 | Probabilistic Music-Symbol Spotting in Handwritten ScoresabstractContent-based search on musical manuscripts is usually performed assuming that there are accurate transcripts of the sources in a symbolic, structured format. Given that current systems for Handwritten Music Recognition are far from offering guarantees about their accuracy, this traditional approach does not represent a scalable scenario. In this work we propose a probabilistic framework for Music-Symbol Spotting (MSS), that allows for content-based music search directly over the images of the manuscripts. By means of statistical recognition systems, a probabilistic index is built upon which the search can be carried out efficiently. Our experiments over a dataset of an Early handwritten music manuscript in Mensural notation demonstrates that this MSS framework can be presented as a promising alternative to the traditional approach for content-based music search. Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 2 |
| 2018 | Text Line Extraction Based on Distance Map Features and Dynamic ProgrammingabstractText Line Segmentation is a basic document layout task that consists in detecting and extracting the text lines present in a document page image. Although considered a basic task, generally, it is a necessary step for Handwritten Text Recognition (HTR) higher level tasks. Most state of the art automatic text recognition, text-to-line image alignment and key word spotting systems require it due to their need for isolated text line images as input. Traditionally most Text Line Segmentation approaches cover both detection and extraction sub steps. However, the community has recently shifted its focus to tackle independently the baseline detection in document images. This shift generates the need for extraction methods that use these detected baselines as input. In this paper, a binarization free dynamic programming approach that generates an equidistant text line extraction polygon is presented. The approach performs this calculation, based on the information provided by priorly detected text baselines and automatically generated foreground pixels distance maps. We evaluate our approach both in a synthetic competition corpus and in a challenging real handwritten text recognition task corpus. We evaluate it not only at the graphical error level but also the impact it produces on an HTR task trained with the line images it yields. We compare our solution with other solutions ranging from the actual human reviewed ground-truth polygons to simpler automatic generated rectangle areas. Vicente Bosch, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 3 |
| 2018 | Probabilistic Indexing and Search for Information Extraction on Handwritten German Parish RecordsabstractWe endeavor to perform very large scale indexing of an ancient German collection of manuscript parish records. To this end we will compute "probabilistic indexes" (PIs), which are known to allow for very accurate and efficient implementation of (single-)keyword spotting. PIs may become prohibitively large for vast manuscript collections. Therefore we analyze simple index pruning methods to achieve adequate tradeoffs between memory requirements and search performance. We also study how to adequately deal with the large variety of non-ASCII symbols and handwritten word spelling variations (accents, umlauts, etc.) which appear in this kind of historical collections. Finally, and most importantly, since most of the images of the collection we aim to index are handwritten tables, we explore the use of PIs to support structured queries for information extraction from untranscribed handwritten images containing tabular data. Empirical results on a small, but complex and representative dataset extracted from the collection considered confirm the viability and adequateness of the chosen approaches. Eva Lang, Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 3 |
| 2018 | From HMMs to RNNs: Computer-Assisted Transcription of a Handwritten Notarial Records CollectionabstractWe present the process which is being followed for the transcription of a large XVIII Century Manuscript collection with the help of Handwritten Text Recognition (HTR) Technology. The documents are being processed in batches of 50 pages each. For each batch we perform two semi-supervised processes: one in order to analyze the layout and detect the text lines and another to provide the full transcripts of the text. As per users request, both diplomatic and modernized transcripts, as well as semantically tagged versions are being produced. Layout analysis supervision is performed by means of a conventional layout editing tool. On the other hand, transcripts, including automatic modernization and tagging, are being produced by means of a web based computer-assisted interactive-predictive tool (CATTI). We present results of the performance of this process through 12 image batches processed so far. These results show the impact caused by an optical modelling technological transition: from classical HMM-based methods to new technology based on recurrent neural networks. Lorenzo Quirós, Vicente Bosch, Lluis Serrano, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 4 |
| 2018 | Active Learning in Handwritten Text Recognition using the Derivational EntropyabstractHandwritten Text Recognition systems are based on statistical models such as recurrent neural networks or hidden Markov models for optical modeling of characters. These models need large corpora for training, consisting in text line images with their corresponding transcripts. The manual annotation of this training data is expensive because it is carried out by experts in paleography, who are specialized in reading ancient scripts. An alternative to reduce the annotation human effort is to use Active Learning techniques to selecting the most informative samples to be used for training. In this paper we study an Active Learning technique to selecting the most informative samples in an HTR scenario. The expert paleographer transcribes only the most informative samples in each stage. The technique followed here is based in the derivational entropy computed from word-graphs obtained from the recognition process. Verónica Romero 0001, Joan-Andreu Sánchez, Alejandro H. Toselli |
ICFHR | 3 |
| 2017 | Preparatory KWS Experiments for Large-Scale Indexing of a Vast Medieval Manuscript Collection in the HIMANIS ProjectabstractMaking large-scale collections of digitized historical documents searchable is being earnestly demanded by many archives and libraries. Probabilistically indexing the text images of these collections by means of keyword spotting techniques is currently seen as perhaps the only feasible approach to meet this demand. A vast medieval manuscript collection, written in both Latin and French, called "Chancery", is currently being considered for indexing at large. In addition to its bilingual nature, one of the major difficulties of this collection is the very high rate of abbreviated words which, on the other hand, are completely expanded in the ground truth transcripts available. In preparation to undertake full indexing of Chancery, experiments have been carried out on a relatively small but fully representative subset of this collection. To this end, a keyword spotting approach has been adopted which computes word relevance probabilities using character lattices produced by a recurrent neural network and a N-gram character language model. Results confirm the viability of the chosen approach for the large-scale indexing aimed at and show the ability of the proposed modeling and training approaches to properly deal with the abbreviation difficulties mentioned. Théodore Bluche, Sébastien Hamel, Christopher Kermorvant, Joan Puigcerver, Dominique Stutzmann, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 6 |
| 2017 | Handwritten Music Recognition for Mensural Notation: Formulation, Data and Baseline ResultsabstractMusic is a key element for cultural transmission, and so large collections of music manuscripts have been preserved over the centuries. In order to develop computational tools for analysis, indexing and retrieval from these sources, it is necessary to transcribe the content to some machine-readable format. In this paper we discuss the Handwritten Music Recognition problem, which refers to the development of automatic transcription systems for musical manuscripts. We focus on mensural notation, one of the most widespread varieties of Western classical music. For that, we present a labeled corpus containing 576 staves, along with a baseline recognition system based on a combination of hidden Markov models and N-gram language models. The baseline error obtained at symbol level is about 40 % which, given the difficulty of the task, can be considered a good starting point for future developments. Our aim is that these data and preliminary results help to promote this research field, serving as a reference in future developments. Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 2 |
| 2017 | ICDAR2017 Competition on Handwritten Text Recognition on the READ DatasetabstractThis paper describes the fourth edition of the Handwritten Text Recognition (HTR) competition that was prepared this time in the context of the International Conference on Document Analysis and Recognition (ICDAR) 2017. Previous editions of this competition were conducted, first, with datasets from the tranScriptorium project in ICFHR 2014, and ICDAR 2015, and then, with datasets from the "Recognition and Enrichment of Archival Documents (READ)" European project in ICFHR 2016. This competition aims to bring together researchers working on off-line HTR and provides them a suitable benchmark to compare their techniques on the task of transcribing typical and difficult historical handwritten documents. The competition proposed for ICDAR 2017 aims at introducing a usual scenario for some collections in which there exist transcripts at page level for many pages useful for training, but these transcripts are not aligned with line images. Two tracks with different conditions on the use of training data were proposed. Most of the data comes from the Alfred Escher Letter Collection. But handwritten images were drawn from other German collections written by several hands. Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Mauricio Villegas, Enrique Vidal 0001 |
ICDAR | 3 |
| 2017 | Querying out-of-vocabulary words in lexicon-based keyword spotting
Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
Neural Comput. Appl. | 2 |
| 2017 | Word graphs size impact on the performance of handwriting document applications
Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001 |
Neural Comput. Appl. | 1 |
| 2016 | Handwriting Transcription and Keyword Spotting in Historical Daily Records DocumentsabstractHistorical records of daily activities provide an intriguing look into the historic life. These documents have interesting information, useful for demography studies and genealogical research. However, automatic processing of historical documents, has mostly been focused on single works of literature and less on daily records, which tend to have a distinct layout, structure, and vocabulary. This paper presents a study about the capability of state-of-the-art handwritten text recognition and key word spotting systems, when applied to this kind of documents. A relatively small set of handwritten birth records registered in Wien in the 16th century is used in the experiments. A word accuracy of about 70% and an AP of 0.74 are achieved for plain image transcription and key word spotting respectively. Taking into account the many difficulties exhibited by these handwritten documents, these preliminary results are quite encouraging. Verónica Romero 0001, Alejandro H. Toselli, Joan-Andreu Sánchez, Enrique Vidal 0001 |
DAS | 2 |
| 2016 | Early Handwritten Music Recognition with Hidden Markov ModelsabstractThis work presents a statistical method to tackle the Handwritten Music Recognition task for Early notation, which comprises more than 200 different symbols. Unlike previous approaches to deal with music notation, our strategy is to perform a holistic recognition without any previous segmentation or staff removal process. The input consists of a page of a music book, which is processed to extract and normalize the staves contained. Then, a feature extraction process is applied to define such sections as a sequence of numerical vectors. The recognition is based on the use of Hidden Markov Models for the optical processing and smoothed N-grams as language model. Experimentation results over a historical archive of Hispanic music reported an error around 40 %, which confirms our proposal as a good starting point taking into account the difficulty of the task. Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 2 |
| 2016 | Sheet Music Statistical Layout AnalysisabstractIn order to provide access to the contents of ancient music scores to researchers, the transcripts of both the lyrics and the musical notation is required. Before attempting any type of automatic or semi-automatic transcription of sheet music, an adequate layout analysis (LA) is needed. This LA must provide not only the locations of the different image regions, but also adequate region labels to distinguish between different region types such as staff, lyric, etc. To this end, we adapt a stochastic framework for LA based on Hidden Markov Models that we had previously introduced for detection and classification of text lines in typical handwritten text images. The proposed approach takes a scanned music score image as input and, after basic preprocessing, simultaneously performs region detection and region classification in an integrated way. To assess this statistical LA approach several experiments were carried out on a representative sample of a historical music archive, under different difficulty settings. The results show that our approach is able to tackle these structured documents providing good results not only for region detection but also for classification of the different regions. Vicente Bosch, Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 3 |
| 2016 | ICFHR2016 Handwritten Keyword Spotting Competition (H-KWS 2016)abstractThe H-KWS 2016, organized in the context of the ICFHR 2016 conference aims at setting up an evaluation framework for benchmarking handwritten keyword spotting (KWS) examining both the Query by Example (QbE) and the Query by String (QbS) approaches. Both KWS approaches were hosted into two different tracks, which in turn were split into two distinct challenges, namely, a segmentation-based and a segmentation-free to accommodate different perspectives adopted by researchers in the KWS field. In addition, the competition aims to evaluate the submitted training-based methods under different amounts of training data. Four participants submitted at least one solution to one of the challenges, according to the capabilities and/or restrictions of their systems. The data used in the competition consisted of historical German and English documents with their own characteristics and complexities. This paper presents the details of the competition, including the data, evaluation metrics and results of the best run of each participating methods. Ioannis Pratikakis, Konstantinos Zagoris, Basilios Gatos, Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 5 |
| 2016 | ICFHR2016 Competition on Handwritten Text Recognition on the READ DatasetabstractThis paper describes the Handwritten Text Recognition (HTR) competition on the READ dataset that has been held in the context of the International Conference on Frontiers in Handwriting Recognition 2016. This competition aims to bring together researchers working on off-line HTR and provide them a suitable benchmark to compare their techniques on the task of transcribing typical historical handwritten documents. Two tracks with different conditions on the use of training data were proposed. Ten research groups registered in the competition but finally five submitted results. The handwritten images for this competition were drawn from the German document Ratsprotokolle collection composed of minutes of the council meetings held from 1470 to 1805, used in the READ project. The selected dataset is written by several hands and entails significant variabilities and difficulties. The five participants achieved good results with transcriptions word error rates ranging from 21% to 47%. Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 3 |
| 2016 | Two Methods to Improve Confidence Scores for Lexicon-Free Word Spotting in Handwritten TextabstractTwo methods are presented to improve word confidence scores for Line-Level Query-by-String Lexicon-Free Keyword Spotting (KWS) in handwritten text images. The first one approaches true relevance probabilities by means of computations directly carried out on character lattices obtained from the lines images considered. The second method uses the same character lattices, but it obtains relevance scores by first computing frame-level character sequence scores which resemble the word posteriorgrams used in previous approaches for lexicon-based KWS. The first method results from a formal probabilistic derivation, which allow us to better understand and further develop the underlying ideas. The second one is less formal but, according with experiments presented in the paper, it obtains almost identical results with much lower computational cost. Moreover, in contrast with the first method, the second one allows to directly obtain accurate bounding boxes for the spotted words. Alejandro H. Toselli, Joan Puigcerver, Enrique Vidal 0001 |
ICFHR | 1 |
| 2016 | Exploiting Existing Modern Transcripts for Historical Handwritten Text RecognitionabstractExisting transcripts for historic manuscripts are a very valuable resource for training models useful for automatic recognition, aided transcription, and/or indexing of the remaining untranscribed parts of these collections. However, these existing transcripts generally exhibit two main problems which hinder their convenience: a) text of the transcripts is seldom aligned with manuscript lines, and b) text often deviate very significantly from what can be seen in the manuscript, either because writing style has been modernized or abbreviations have been expanded, or both. This work presents an analysis of these problems and discusses possible solutions for minimizing human effort needed to adapt existing transcripts in order to render them usable. Empirical results presented show the huge performance gain that can be obtained by adequately adapting the transcripts, thus motivating future development of the proposed solutions. Mauricio Villegas, Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001 |
ICFHR | 2 |
| 2016 | HMM word graph based keyword spotting in handwritten document images
Alejandro H. Toselli, Enrique Vidal 0001, Verónica Romero 0001, Volkmar Frinken |
Inf. Sci. | 1 |
| 2015 | Probabilistic interpretation and improvements to the HMM-filler for handwritten keyword spottingabstractTraditionally, the HMM-Filler approach has been widely used in the fields of speech recognition and handwritten text recognition to tackle lexicon-free, query-by-string keyword spotting (KWS). It computes a score to determine whether a given keyword is written in a certain image region. It is conjectured, that this score is related to the confidence of the system, respect to the previous question. However, it is still not clear what this relationship is. In this paper, the HMM-Filler score is derived from a probabilistic formulation of KWS, which gives a better understanding of its behavior and limits. Additionally, the same probabilistic framework is used to present a new algorithm to compute the KWS scores, which results in better average precision (AP), for a keyword spotting task in the widely used IAM database. We show that the new algorithm can improve the HMM-filler results up to 10.4% relative (5.3% absolute) points in AP, in the considered task. Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 2 |
| 2015 | ICDAR2015 Competition on Keyword Spotting for Handwritten DocumentsabstractThe principal goal of the Competition on Keyword Spotting for Handwritten Documents was to promote different approaches used in the field of Keyword Spotting and to fairly compare them using uniform data and metrics. To accommodate different perspectives adopted by researches in this field, the competition was divided into two distinct tracks, namely, a training-free and a training-based track, and each track entailed two optional assignments. Six participants submitted solutions to one or both assignments, depending on the capabilities and/or restrictions of their systems. The data used in the competition consisted of historical documents in English with different levels of complexity. This paper presents the details of the competition, including the data, evaluation metrics and results of the best participant methods. Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 2 |
| 2015 | ICDAR 2015 competition HTRtS: Handwritten Text Recognition on the tranScriptorium datasetabstractThis paper describes the second edition of the Handwritten Text Recognition (HTR) contest on the tranScriptorium datasets that has been held in the context of the International Conference on Document Analysis and Recognition 2015. Two tracks with different conditions on the use of training data were proposed. Nine research groups registered in the contest but finally three research submitted results. The handwritten images for this contest were drawn from the English “Bentham collection” dataset used in the tranScriptorium project. A small subset of this collection has been chosen for the present HTR competition. The selected subset has been written by several hands and entails significant variabilities and difficulties regarding the quality of text images, writing styles and crossed-out text. This contest is clearly more difficult than the the first edition both for training and for testing. A portion of the training dataset and the full test dataset were provided in the form of carefully segmented line images, along with the corresponding transcripts. Another portion of the training dataset was provided as raw images and their corresponding transcripts at region level. The three participants achieved good results, with transcription word error rates ranging from 31% down to 44%. Joan-Andreu Sánchez, Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001 |
ICDAR | 2 |
| 2015 | Context-aware lattice based filler approach for key word spotting in handwritten documentsabstractThe so-called filler or garbage Hidden Markov Models (HMM-Filler) are among the most widely used models for lexicon-free, query by string key word spotting (KWS) in the fields of speech recognition and (lately) handwritten text recognition. However, it has important drawbacks. First, the keyword-specific HMM Viterbi decoding process needed to obtain the confidence scores of each spotted word involves a large computational cost. Second, in its traditional conception, the model does not take into account any context information - and more recent works where simple character bi-gram context is used show that not only the computational cost becomes even larger, but also the required keyword-specific language model becomes quite intricate to build. In a previous work we introduced KWS methods based on character lattices which proved very much simpler and faster than the traditional HMM-Filler, while providing practically identical results. Here we extend our previous work by using context-aware character lattices obtained by means of Viterbi decoding with high-order character N-gram models. Experimental results show that, as compared with a direct 2-gram HMM-filler implementation, the proposed approach requires between one and two orders of magnitude less query computing time. Moreover, for the first time in the field of handwritten text KWS, Filler-based results for N-grams up to N = 6 are reported, clearly showing a great impact of context on precision-recall performance. Alejandro H. Toselli, Joan Puigcerver, Enrique Vidal 0001 |
ICDAR | 1 |
| 2015 | High performance Query-by-Example keyword spotting using Query-by-String techniquesabstractKeyword Spotting (KWS) has been traditionally considered under two distinct frameworks: Query-by-Example (QbE) and Query-by-String (QbS). In both cases the user of the system wished to find occurrences of a particular keyword in a collection of document images. The difference is that, in QbE, the keyword is given as an exemplar image while, in QbS the keyword is given as a text string. In several works, the QbS scenario has been approached using QbE techniques; but the converse has not been studied in depth yet, despite of the fact that QbS systems typically achieve higher accuracy. In the present work, we present a very effective probabilistic approach to QbE KWS, based on highly accurate QbS KWS techniques which rely on models which need to be trained from annotated data. To assess the effectiveness of this approach, we tackle the segmentation-free QbE task of the ICFHR-2014 Competition on Handwritten KWS. Our approach achieves a mean average precision (mAP) as high as 0.715, which improves by more than 70% the best mAP achieved in this competition (0.419 under the same experimental conditions). Enrique Vidal 0001, Alejandro H. Toselli, Joan Puigcerver |
ICDAR | 2 |
| 2015 | Context-Aware Gestures for Mixed-Initiative Text Editing UIsabstractThis work is focused on enhancing highly interactive text-editing applications with gestures. Concretely, we study Computer Assisted Transcription of Text Images (CATTI), a handwriting transcription system that follows a corrective feedback paradigm, where both the user and the system collaborate efficiently to produce a high-quality text transcription. CATTI-like applications demand fast and accurate gesture recognition, for which we observed that current gesture recognizers are not adequate enough. In response to this need we developed MinGestures, a parametric context-aware gesture recognizer. Our contributions include a number of stroke features for disambiguating copy-mark gestures from handwritten text, plus the integration of these gestures in a CATTI application. It becomes finally possible to create highly interactive stroke-based text-editing interfaces, without worrying to verify the user intent on-screen. We performed a formal evaluation with 22 e-pen users and 32 mouse users using a gesture vocabulary of 10 symbols. MinGestures achieved an outstanding accuracy (<1% error rate) with very high performance (<1 ms of recognition time). We then integrated MinGestures in a CATTI prototype and tested the performance of the interactive handwriting system when it is driven by gestures. Our results show that using gestures in interactive handwriting applications is both advantageous and convenient when gestures are simple but context-aware. Taken together, this work suggests that text-editing interfaces not only can be easily augmented with simple gestures, but also may substantially improve user productivity. Luis A. Leiva, Vicente Alabau, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
Interact. Comput. | 4 |
| 2014 | Ground-Truth Production in the Transcriptorium ProjectabstractTran Scriptorium is a 3-years project that aims to develop innovative, cost-effective solutions for the indexing, search and full transcription of historical handwritten document images, using Handwritten Text Recognition (HTR) technology. The production of ground-truth (GT) of a dataset of handwritten document images is among the first tasks. We address novel approaches for the faster production of this GT based on crowd-sourcing and on prior-knowledge methods. We also address here a novel low-cost semi-supervised procedure for obtaining pairs of correct line-level aligned detected/extracted text line images and text line transcripts, specially suitable for training models of the HTR technology employed in Tran Scriptorium. Basilios Gatos, Georgios Louloudis, Tim Causer, Kris Grint, Verónica Romero 0001, Joan-Andreu Sánchez, Alejandro H. Toselli, Enrique Vidal 0001 |
Document Analysis Systems | 7 |
| 2014 | Word-Graph Based Handwriting Key-Word Spotting: Impact of Word-Graph Size on PerformanceabstractKey-Word Spotting (KWS) in handwritten documents is approached here by means of Word Graphs (WG) obtained using segmentation-free handwritten text recognition technology based on N-gram Language Models and Hidden Markov Models. Linguistic context significantly boost KWS performance with respect to methods which ignore word contexts and/or rely on image-matching with pre-segmented isolated words. On the other hand, WG-based KWS can be significantly faster than other KWS approaches which directly work on the original images where, in general, computational demands are exceedingly high. A large WG contains most of the relevant information of the original text (line) image needed for KWS but, if it is too large, the computational advantages over traditional, image matching-based KWS become diminished. Conversely, if it is too small, relevant information may be lost, leading to degraded KWS precision/recall performance. We study the trade off between WG size and KWS information retrieval performance. Results show that small, computationally cheap WGs can be used without loosing the excellent KWS performance achieved with huge WGs. Alejandro H. Toselli, Enrique Vidal 0001 |
Document Analysis Systems | 1 |
| 2014 | Semiautomatic Text Baseline Detection in Large Historical Handwritten DocumentsabstractA semiautomatic iterative process for the detection of text baselines in historical handwritten document images is presented. It relies on the use of Hidden Markov Models (HMM) to provide initial text baselines hypotheses, followed by user review in order to produce ground-truth quality results. Using the set of revised baselines as ground truth, the HMM's are re-trained before processing the next batch of pages. This process has been evaluated in the context of a real transcription task which, as a by-product, has produced line-detection ground truth. We show that the usage of a formal, HMM-based line-detection approach which requires training data, not only yields good detection results but is also of practical use in large handwritten image collections. Through experiments with real users we show that the proposed approach has interesting features, namely, accuracy, scalability and ease of use, as well as low overall human effort requirements. Vicente Bosch, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 2 |
| 2014 | Word-Graph and Character-Lattice Combination for KWS in Handwritten DocumentsabstractWe present a handwritten text Keyword Spotting (KWS) approach based on the combination of KWS methods using word-graphs (WGs) and character-lattices (CLs). It aims to solve the problem that WG-based models present for out of vocabulary (OOV) keywords: since there is no available information about them in the lexicon or the language model, null scores are assigned. OOV keywords may have a significant impact on the global performance of KWS systems, as we show. By using a CL approach, which does not suffer from the previous problem, to estimate the OOV scores, we take advantage of both models, using the speed and accuracy that WGs provide for in-vocabulary keywords and the flexibility of the CL approach. This combination improves significantly both average precision and mean average precision over the two methods. Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 2 |
| 2014 | ICFHR2014 Competition on Handwritten Text Recognition on Transcriptorium Datasets (HTRtS)abstractA contest on Handwritten Text Recognition organised in the context of the ICFHR 2014 conference is described. Two tracks with increased freedom on the use of training data were proposed and three research groups participated in these two tracks. The handwritten images for this contest were drawn from an English data set which is currently being considered in the Tran scriptorium project. The goal of this project is to develop innovative, efficient and cost-effective solutions for the transcription of historical handwritten document images, focusing on four languages: English, Spanish, German and Dutch. For the English language, the so-called "Bentham collection" is being considered in Tran scriptorium. It encompasses a large set of manuscripts written by the renowned English philosopher and reformer Jeremy Bentham (1748-1832). A small subset of this collection has been chosen for the present HTR competition. The selected subset has been written by several hands (Bentham himself and his secretaries) and entails significant variabilities and difficulties regarding the quality of text images and writing styles. Training and test data were provided in the form of carefully segmented line images, along with the corresponding transcripts. The three participants achieved very good results, with transcription word error rates ranging from 15.0% down to 8.6%. Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 3 |
| 2014 | Bleed-Through Removal by Learning a Discriminative Color ChannelabstractThis paper proposes a novel bleed-through removal technique based on learning a color channel that is optimized so that the foreground text is enhanced while at the same time the variability of the background (including the bleed-through) is diminished. The technique is intended to be part of an interactive transcription system in which the objective is obtaining high quality transcriptions with the least human effort. Thus, instead of training the bleed-through removal to work in general for any document, the technique requires a user to label regions both as foreground text and as bleed-through, with the aim that the method is adapted to the characteristics of each document. The proposal is assessed using the handwritten recognition performance on a real 17th century manuscript. Mauricio Villegas, Alejandro H. Toselli |
ICFHR | 2 |
| 2014 | Word-Graph-Based Handwriting Keyword Spotting of Out-of-Vocabulary QueriesabstractThanks to the use of lexical and syntactic information, Word Graphs (WG) have shown to provide a competitive Precision-Recall performance, along with fast lookup times, in comparison to other techniques used for Key-Word Spotting (KWS) in handwritten text images. However, a problem of WG approaches is that they assign a null score to any keyword that was not part of the training data, i.e. Out-of-Vocabulary (OOV) keywords, whereas other techniques are able to estimate a reasonable score even for these kind of keywords. We present a smoothing technique which estimates the score of an OOV keyword based on the scores of similar keywords. This makes the WG-based KWS as flexible as other techniques with the benefit of having much faster lookup times. Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
ICPR | 2 |
| 2013 | Fast HMM-Filler Approach for Key Word Spotting in Handwritten DocumentsabstractThe so-called filler or garbage Hidden Markov Models (HMM) are among the most widely used models for lexicon-free, query by string key word spotting in the fields of speech recognition and (lately) handwritten text recognition. An important drawback of this approach is the large computational cost of the keyword-specific HMM Viterbi decoding process needed to obtain the confidence scores of each word to be spotted. This paper presents a novel way to compute such confidence scores, directly from character lattices produced during a single Viterbi decoding process using only the "filler" model (i.e. no explicit keyword-specific decoding is needed). Experiments show that, as compared with the classical HMM-filler approach, the proposed method obtains essentially the same spotting results, while requiring between one and two orders of magnitude less query computing time. Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 1 |
| 2013 | The ESPOSALLES database: An ancient marriage license corpus for off-line handwriting recognition
Verónica Romero 0001, Alicia Fornés, Joan-Andreu Sánchez, Alejandro H. Toselli, Volkmar Frinken, Enrique Vidal 0001, Josep Lladós 0001 |
Pattern Recognit. | 5 |
| 2012 | Statistical Text Line Analysis in Handwritten DocumentsabstractIn this paper we present an approach for text line analysis and detection in handwritten documents based on Hidden Markov Models, a technique widely used in other handwritten and speech recognition tasks. It is shown that text line analysis and detection can be solved using a more formal methodology in contraposition to most of the proposed heuristic approaches found in the literature. Our approach not only provides the best position coordinates for each of the vertical page regions but also labels them, in this manner surpassing the traditional heuristic methods. In our experiments we demonstrate the performance of the approach (both in line analysis and detection) and study the impact of increasingly constrained "vertical layout language models" on text line detection accuracy. Through this experimentation we also show the improvement in quality of the baselines yielded by our approach in comparison with a state-of-the-art heuristic method based on vertical projection profiles. Vicente Bosch, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 2 |
| 2012 | Multimodal Computer-Assisted transcription of Text Images at Character-Level InteractionabstractCurrently, automatic handwriting recognition systems are ineffectual in unconstrained handwriting documents. Therefore, to obtain perfect transcriptions, heavy human intervention is required to validate and correct the results of such systems. Given that this post-editing process is inefficient and uncomfortable, a multimodal interactive approach has been proposed in previous works, which aims at obtaining correct transcriptions with the minimum human effort. In this approach, the user interacts with the system by means of an e-pen and/or more traditional methods such as keyboard or mouse. This user's feedback allows to improve system accuracy and multimodality increases system ergonomics and user acceptability. Until now, multimodal interaction has been considered only at whole-word level. In this work, multimodal interaction at character-level is studied, that may lead to more effective interactivity, since it is faster and easier to write only one character rather than a whole word. Here we study this kind of fine-grained multimodal interaction and present developments that allow taking advantage of interaction-derived context to significantly improve feedback decoding accuracy. Empirical tests on three cursive handwritten tasks suggest that, despite losing the deterministic accuracy of traditional peripherals, this approach can save significant amounts of user effort with respect to fully manual transcription as well as to noninteractive post-editing correction. Daniel Martín-Albo, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2011 | Evaluating an Interactive-Predictive Paradigm on Handwriting Transcription: A Case Study and Lessons LearnedabstractTranscribing handwritten text is a laborious task which currently is carried out manually. As the accuracy of automatic handwritten text recognizers improves, post-editing the output of these recognizers could be foreseen as a possible alternative. Alas, the state-of-the-art technology is not suitable to perform this kind of work, since current approaches are not accurate enough and the process is usually both inefficient and uncomfortable for the user. As alternative, an interactive-predictive paradigm has gained recently an increasing popularity, mainly due to promising empirical results that estimate considerable reductions of user effort. In order to assess whether these empirical results can lead indeed to actual benefits, we developed a working prototype and conducted a field study remotely. Thirteen regular computer users tested two different transcription engines through the above-mentioned prototype. We observed that the interactive-predictive version allowed to transcribe better (less errors and fewer iterations to achieve a high-quality output) in comparison to the manual engine. Additionally, participants ranked higher such an interactive-predictive system in a usability questionnaire. We describe the evaluation methodology and discuss our preliminary results. While acknowledging the known limitations of our experimentation, we conclude that the interactive-predictive paradigm is an efficient approach for transcribing handwritten text. Luis A. Leiva, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
COMPSAC | 3 |
| 2011 | Study of different interactive editing operations in an assisted transcription systemabstractTo date, automatic handwriting recognition systems are far from being perfect. Therefore, once the full recognition process of a handwritten text image has finished, heavy human intervention is required in order to correct the results of such systems. As an alternative, an interactive system has been proposed in previous works. This alternative follows an Interactive Predictive paradigm and the results show that significant amounts of human effort can be saved. So far only word substitutions and pointer actions have been considered in this interactive system. In this work, we study different interactive editing operations that can allow for more effective, ergonomic and friendly interfaces. Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICMI | 2 |
| 2010 | Interactive layout analysis and transcription systems for historic handwritten documentsabstractThe amount of digitized legacy documents has been rising dramatically over the last years due mainly to the increasing number of on-line digital libraries publishing this kind of documents, waiting to be classified and finally transcribed into a textual electronic format (such as ASCII or PDF). Nevertheless, most of the available fully-automatic applications addressing this task are far from being perfect and heavy and inefficient human intervention is often required to check and correct the results of such systems. In contrast, multimodal interactive-predictive approaches may allow the users to participate in the process helping the system to improve the overall performance. With this in mind, two sets of recent advances are introduced in this work: a novel interactive method for text block detection and two multimodal interactive handwritten text transcription systems which use active learning and interactive-predictive technologies in the recognition process. Oriol Ramos Terrades, Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001, Alfons Juan-Císcar |
ACM Symposium on Document Engineering | 2 |
| 2010 | Handwritten Word Verification by SVM-Based Hypotheses Re-scoring and Multiple Thresholds RejectionabstractIn the field of isolated handwritten word recognition, the development of verification systems that optimize the trade-off between performance and reliability is still an active research topic. To minimize the recognition errors, usually, a verification system is used to accept or reject the hypotheses output by an existing recognition system. In this paper, a novel verification architecture is presented. In essence, the recognition hypotheses, re-scored by a set of the support vector machines, are validated by a verification mechanism based on multiple rejection thresholds. In order to tune these (class-dependent) rejection thresholds, an algorithm based on dynamic programming is proposed which focus on maximizing the recognition rate for a given prefixed error rate. Preliminary reported results of experiments carried out on RIMES database show that this approach performs equal or superior to other state-of-the-art rejection methods. Laurent Guichard, Alejandro H. Toselli, Bertrand Coüasnon |
ICFHR | 2 |
| 2010 | Character-Level Interaction in Computer-Assisted Transcription of Text ImagesabstractTo date, automatic handwriting recognition systems are far from being perfect and heavy human intervention is often required to check and correct the results of such systems. As an alternative, an interactive framework that integrates the human knowledge into the transcription process has been presented in previous works. This new approach follows an Interactive Predictive paradigm and our results show that significant amounts of human effort can be saved. Until now only whole-word interactions with this system have been considered. In this work, character-level keystroke interactions, that can allow for a more ergonomic and friendly interfaces, are proposed. Empirical results show that this allows for further improvements in user productivity. Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 2 |
| 2010 | A Novel Verification System for Handwritten Words RecognitionabstractIn the field of isolated handwritten word recognition, the development of highly effective verification systems to reject words presenting ambiguities is still an active research topic. In this paper, a novel verification system based on support vector machine scoring and multiple reject class-dependent thresholds is presented. In essence, a set of support vector machines appended to a standard HMM-based recognition system provides class-dependent confidence measures employed by the verification mechanism to accept or reject the recognized hypotheses. Experimental results on RIMES database show that this approach outperforms other state-of-the-art approaches. Laurent Guichard, Alejandro H. Toselli, Bertrand Coüasnon |
ICPR | 2 |
| 2010 | A Bi-modal Handwritten Text Corpus: Baseline ResultsabstractHandwritten text is generally captured through two main modalities: off-line and on-line. Smart approaches to handwritten text recognition (HTR) may take advantage of both modalities if they are available. This is for instance the case in computer-assisted transcription of text images, where on-line text can be used to interactively correct errors made by a main off-line HTR system. We present here baseline results on the biMod-IAM-PRHLT corpus, which was recently compiled for experimentation with techniques aimed at solving the proposed multi-modal HTR problem, and is being used in one of the official ICPR-2010 contests. Moisés Pastor, Alejandro H. Toselli, Francisco Casacuberta, Enrique Vidal 0001 |
ICPR | 2 |
| 2010 | Computer Assisted Transcription of Text Images: Results on the GERMANA Corpus and Analysis of Improvements Needed for Practical UseabstractWe present a study of the application of Computer Assisted Transcription of Text Images (CATTI) to a task which is much closer to real applications than other tasks previously studied. The new task consists in the transcription of a new publicly available historic handwritten document, called GERMANA. A detailed analysis of the main factors influencing the system performance are exposed and some strategies to circumvent them are proposed. Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICPR | 2 |
| 2010 | Multimodal interactive transcription of text images
Alejandro H. Toselli, Verónica Romero 0001, Moisés Pastor, Enrique Vidal 0001 |
Pattern Recognit. | 1 |
| 2009 | Using Mouse Feedback in Computer Assisted Transcription of Handwritten Text ImagesabstractTo date, automatic handwriting recognition systems are far from being perfect and heavy human intervention is often required to check and correct the results of such systems. In order to achieve correct transcriptions, human knowledge can be integrated into the transcription process, following an Interactive Predictive paradigm. We have recently proposed Mouse Actions as a significant feedback information source for the underlying interactive system to improve the productivity of the human transcriptor. In this paper we review this way to interact with the system and report comparative results using the publicly available IAMDB dataset. Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 2 |
| 2009 | Interactive multimodal transcription of text images using a web-based demo systemabstractThis document introduces a web based demo of an interactive framework for transcription of handwritten text, where the user feedback is provided by means of pen strokes on a touchscreen. Here, the automatic handwriting text recognition system and the user both cooperate to generate the final transcription. Verónica Romero 0001, Luis A. Leiva, Alejandro H. Toselli, Enrique Vidal 0001 |
IUI | 3 |
| 2007 | Computer Assisted Transcription of Handwritten Text ImagesabstractTo date, automatic handwriting recognition systems are far from being perfect and often they need a post editing where a human intervention is required to check and correct the results of such systems. We propose to have a new interactive, on-line framework which, rather than full automation, aims at assisting the human in the proper recognition- transcription process; that is, facilitate and speed up their transcription task of handwritten texts. This framework combines the efficiency of automatic handwriting recognition systems with the accuracy of the human transcriptor. The best result is a cost-effective perfect transcription of the handwriting text images. Alejandro H. Toselli, Verónica Romero 0001, Luis Rodríguez, Enrique Vidal 0001 |
ICDAR | 1 |
| 2005 | Writing Speed Normalization for On-Line Handwritten Text RecognitionabstractPen-based interfaces aim at improving the man-machine interaction of many portable systems. While statistical models can be used to learn pen position sequences, they suffer from the huge variability exhibited by the speed of writing. To improve performance, invariance to the writing speed is needed. Trace segmentation is a technique that can be used to normalize the writing speed. This method is controlled by a parameter called resampling distance. A study of the resampling distance is presented here, along with another approximation to the writing speed normalization called "derivatives normalization". The improvement using trace segmentation was 193% relative to the baseline, whilst the improvement using derivatives normalization was 47.3% relative. Moisés Pastor, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 2 |
| 2004 | Integrated Handwriting Recognition And Interpretation Using Finite-State ModelsabstractThe interpretation of handwritten sentences is carried out using a holistic approach in which both text image recognition and the interpretation itself are tightly integrated. Conventional approaches follow a serial, first-recognition then-interpretation scheme which cannot adequately use semantic–pragmatic knowledge to recover from recognition errors. Stochastic finite-sate transducers are shown to be suitable models for this integration, permitting a full exploitation of the final interpretation constraints. Continuous-density hidden Markov models are embedded in the edges of the transducer to account for lexical and morphological constraints. Robustness with respect to stroke vertical variability is achieved by integrating tangent vectors into the emission densities of these models. Experimental results are reported on a syntax-constrained interpretation task which show the effectiveness of the proposed approaches. These results are also shown to be comparatively better than those achieved with other conventional, N-gram-based techniques which do not take advantage of full integration. Alejandro H. Toselli, Alfons Juan-Císcar, Ismael Salvador, Enrique Vidal 0001, Francisco Casacuberta, Daniel Keysers, Hermann Ney |
Int. J. Pattern Recognit. Artif. Intell. | 1 |