VLDB 2026 Research / reviewers in the wild / expert
Joan-Andreu Sánchez
dblp:15/2848 · also Joan Andreu Sánchez Peiró
· DBLP profile ↗
65ranked-venue papers
15as first author
16since 2021 · last 2026
0000-0003-0423-2020ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 9 first-author · 12 since 2021Databases, data management, data science and information retrieval · 27 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Full-page recognition and alignment of historical musical documentsabstractAbstract Optical Music Recognition aims to transcribe musical manuscript images into digital formats by using automatic methods for enhanced accessibility and preservation. This task is challenging for handwritten historical musical pieces from the Late Middle Ages, Early Renaissance, and previous time periods. This music has the interesting characteristic that both musical and lyrical elements are present with an implicit time alignment between them. This paper introduces techniques for simultaneously transcribing the musical and lyrical elements. We research how to automatically obtain the time alignment for an accurate musicological interpretation. Convolutional and Recurrent Neural Networks and Transformer models are explored for holistically transcribing and aligning historical pieces. This paper explores different techniques to improve the training of the models in limited data scenarios. Experiments are conducted on two different datasets from the same time period. Our findings highlight the potential of Transformer models in overcoming the alignment challenge, providing the best alignment capabilities without compromising the quality of transcriptions and offering a promising direction for future research in the automatic recognition of historical musical documents. Manuel Villarreal, Joan-Andreu Sánchez, Daniel Parres |
Int. J. Document Anal. Recognit. | 2 |
| 2026 | Simple handwritten text recognition techniques for highly accurate writer identificationabstractAbstract The theater of the Spanish Early Modern period encompasses thousands of textual works and hundreds of playwrights and is one of the greatest examples of Spanish literature. These works were usually copied and altered, changing sometimes the meaning of the original works written by the authors. Therefore, identifying autograph testimonies written directly by the playwrights themselves are of particular importance. In this paper, we address an approach for writer identification based on Deep Convolutional-Recurrent Neural Networks and n -gram language models. The identification task is posed as a classification problem, introducing a probabilistic framework that goes beyond the plain transcription of handwritten text. Experiments are conducted to validate our proposal to distinguish between Lope de Vega ’s manuscripts and other non-Lope hands. The good results achieved will ultimately allow to provide modern researchers with a useful tool for cultural heritage recovery. Alejandro H. Toselli, Álvaro Cuéllar, Sònia Boadas, Enrique Vidal 0001, Joan-Andreu Sánchez |
Pattern Anal. Appl. | 5 |
| 2025 | PARDES: Automatic Generation of Descriptive Terms for Logical Units in Historical Handwritten Collections
Josepa Raventós-Pajares, Joan-Andreu Sánchez, Enrique Vidal 0001 |
IEEE Big Data | 2 |
| 2024 | Speed-Up Pre-trained Vision Encoder-Decoder Transformers by Leveraging Lightweight Mixer Layers for Text Recognition
Daniel Parres, Dan Anitei, Roberto Paredes, Joan-Andreu Sánchez, José-Miguel Benedí |
DAS | 4 |
| 2024 | Improving Efficiency and Performance Through CTC-Based Transformers for Mathematical Expression Recognition
Dan Anitei, Daniel Parres, Joan-Andreu Sánchez, José-Miguel Benedí |
ICDAR (5) | 3 |
| 2024 | Enhancing Recognition of Historical Musical Pieces with Synthetic and Composed Images
Manuel Villarreal, Joan-Andreu Sánchez |
ICDAR (3) | 2 |
| 2024 | Ground-truth generation through crowdsourcing with probabilistic indexesabstractAbstract Automatic transcription of large series of historical handwritten documents generally aims at allowing to search for textual information in these documents. However, automatic transcripts often lack the level of accuracy needed for reliable text indexing and search purposes. Probabilistic Indexing (PrIx) offers a unique alternative to raw transcripts. Since it needs training data to achieve good search performance, PrIx-based crowdsourcing techniques are introduced in this paper to gather the required data. In the proposed approach, PrIx confidence measures are used to drive a correction process in which users can amend errors and possibly add missing text. In a further step, corrected data are used to retrain the PrIx models. Results on five large series are reported which show consistent improvements after retraining. However, it can be argued whether the overall costs of the crowdsourcing operation pay off for the improvements, or perhaps it would have been more cost-effective to just start with a larger and cleaner amount of professionally produced training transcripts. Joan-Andreu Sánchez, Enrique Vidal 0001, Vicente Bosch, Lorenzo Quirós |
Neural Comput. Appl. | 1 |
| 2023 | Synchronous Recognition of Music Images Using Coupled N-Gram ModelsabstractHandwritten music recognition researches the use of technologies to automatically transcribe handwritten music pieces that are only found in image format, and make them available to the general public. Many historical music pieces are composed by a music part and a lyrics part. Handwritten music recognition has focused mainly on transcribing the music elements in historical images, but there exist many pieces where both music and lyrics are present and of relevance. The recognition of both music and lyrics is generally carried out as separate tasks. Both parts are synchronized in many historical documents at line level and loosely at word level. These two elements are strongly related having each one affecting the other. Discovering this relation may be very relevant to improve recognition results in both parts and to further steps like music analysis, composition analysis, etc. This paper introduces a preliminary system that transcribes synchronously and simultaneously both the music and lyrics elements of handwritten historical music images. The results obtained over a historical manuscript dataset show that this system obtains an improvement of up to 15.4% at symbol rate on stave recognition and up to an approximately average 7.6% improvement when both the music and lyrics part are jointly considered. Manuel Villarreal, Joan-Andreu Sánchez |
DocEng | 2 |
| 2023 | Discriminative estimation of probabilistic context-free grammars for mathematical expression recognition and retrievalabstractAbstract We present a discriminative learning algorithm for the probabilistic estimation of two-dimensional probabilistic context-free grammars (2D-PCFG) for mathematical expressions recognition and retrieval. This algorithm is based on a generalization of the H-criterion as the objective function and the growth transformations as the optimization method. For the development of the discriminative estimation algorithm, the N-best interpretations provided by the 2D-PCFG have been considered. Experimental results are reported on two available datasets: Im2Latex and IBEM. The first experiment compares the proposed discriminative estimation method with the classic Viterbi-based estimation method. The second one studies the performance of the estimated models depending on the length of the mathematical expressions and the number of admissible errors in the metric used. Ernesto Noya, José-Miguel Benedí, Joan-Andreu Sánchez, Dan Anitei |
Pattern Anal. Appl. | 3 |
| 2023 | The IBEM dataset: A large printed scientific image dataset for indexing and searching mathematical expressionsabstractSearching for information in printed scientific documents is a challenging problem that has recently received special attention from the Pattern Recognition research community. Mathematical expressions are complex elements that appear in scientific documents, and developing techniques for locating and recognizing them requires the preparation of datasets that can be used as benchmarks. Most current techniques for dealing with mathematical expressions are based on Machine Learning techniques which require a large amount of annotated data. These datasets must be prepared with ground-truth information for automatic training and testing. However, preparing large datasets with ground-truth is a very expensive and time-consuming task. This paper introduces the IBEM dataset, consisting of scientific documents that have been prepared for mathematical expression recognition and searching. This dataset consists of 600 documents, more than 8200 page images with more than 160000 mathematical expressions. It has been automatically generated from the version of the documents and can be enlarged easily. The ground-truth includes the position at the page level and the transcript for mathematical expressions both embedded in the text and displayed. This paper also reports a baseline classification experiment with mathematical symbols and a baseline experiment of Mathematical Expression Recognition performed on the IBEM dataset. These experiments aim to provide some benchmarks for comparison purposes so that future users of the IBEM dataset can have a baseline framework. Dan Anitei, Joan-Andreu Sánchez, José-Miguel Benedí, Ernesto Noya |
Pattern Recognit. Lett. | 2 |
| 2023 | Processing a large collection of historical tabular imagesabstractProcessing automatically historical document images to allow the search of textual information requires the preparation of ground-truth data for training and evaluation. This process is an expensive and arduous task, especially when the historical document images contain specialized vocabulary and/or tabular information. In the latter case, relevant decisions have to be taken to annotate the tabular parts. This paper presents a complex collection of historical document images and the resulting database, which is called HisClima. In this database, half of the images are in tabular format and half as running text. Both types of images contain pre-printed and handwritten text. The textual information is plenty of abbreviations and specific vocabulary related to weather conditions and old ships. This database can be used to research technologies related to historical document image processing and analysis, both for tabular and running text recognition. Baseline results are presented for Document Layout Analysis, Text Recognition, and Probabilistic Indexing. Although these results are good, there is still room for improvement and some indications are provided in this direction. Emilio Granell, Verónica Romero 0001, José Ramón Prieto, José Andrés, Lorenzo Quirós, Joan-Andreu Sánchez, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 6 |
| 2023 | Information extraction in handwritten historical logbooksabstractDocument Image Understanding is a demanding Pattern Recognition problem that requires complex recognition models. This problem is even more difficult for document images with complicated layouts like tables, where the reading order is often intrinsically ambiguous, and consequently, the context is generally ambiguous as well. In this paper, we compare two machine learning approaches for extracting information in pre-printed historical tables with handwritten information. We analyze the performance of each approach at each step of the extraction process over different corpora, up to a realistic scenario where documents with different table layouts written by different hands are used. The results are good in general and show that a model based on Multilayer Perceptrons yields better results on more homogeneous documents, while another model based on Graph Neural Networks generalizes better on heterogeneous corpora. José Ramón Prieto, José Andrés, Emilio Granell, Joan-Andreu Sánchez, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 4 |
| 2022 | Information Extraction from Handwritten Tables in Historical Documents
José Andrés, José Ramón Prieto, Emilio Granell, Verónica Romero 0001, Joan-Andreu Sánchez, Enrique Vidal 0001 |
DAS | 5 |
| 2022 | Effective Crowdsourcing in the EDT Project with Probabilistic Indexes
Joan-Andreu Sánchez, Enrique Vidal 0001, Vicente Bosch |
DAS | 1 |
| 2021 | ICDAR 2021 Competition on Mathematical Formula Detection
Dan Anitei, Joan-Andreu Sánchez, José Manuel Fuentes, Roberto Paredes, José-Miguel Benedí |
ICDAR (4) | 2 |
| 2021 | Reducing the Human Effort in Text Line Segmentation for Historical Documents
Emilio Granell, Lorenzo Quirós, Verónica Romero 0001, Joan-Andreu Sánchez |
ICDAR (3) | 4 |
| 2020 | Two Semi-Supervised Training Approaches for Automated Text RecognitionabstractAutomated text recognition is a fundamental problem in Document Image Analysis. Optical models are used for modeling characters while language models are used for composing sentences. Since the scripts and linguistic context differ widely, it is mandatory to specialize the models by training on task-dependent ground-truth. However, to create a sufficient amount of ground-truth, at least for historical handwritten scripts, well-qualified persons have to mark and transcribe text lines, which is very time-consuming. On the other hand, in many cases unassigned transcripts are already available on page-level from another process chain, or at least transcripts from similar linguistic context are available. In this work we present two approaches that make use of such transcripts: whereas the first one creates training data by automatically assigning page-dependent transcripts to text lines, the second one uses a task-specific language model to generate highly confident training data. Both approaches are successfully applied on a very challenging historical handwritten collection. Gundram Leifert, Roger Labahn, Joan-Andreu Sánchez |
ICFHR | 3 |
| 2020 | The Carabela Project and Manuscript Collection: Large-Scale Probabilistic Indexing and Content-based ClassificationabstractThe main aim of the Carabela project was to develop and apply techniques that allow textual searching on massive Spanish collections of 15th-19th century manuscripts. The project focused on a relatively small subset of 125 000 images of collections of interest to underwater archaeology. For this type of manuscripts, state-of-the-art automatic transcription techniques, generally fail to achieve usable transcription accuracy. Therefore, rather than insisting in actual transcription, methodologies for probabilistic indexing of handwritten text images have been adopted. This has allowed us to effectively cope with the intrinsically high degree of uncertainty of the text contained in most historical manuscripts, leading to highly effective systems for textual search and retrieval. Carabela has gone one step further by developing new techniques to classify probabilistically indexed, but otherwise untranscribed, text images according to their textual content. These techniques have been successfully used to automatically classify Carabela bundels (each containing hundreds or thousands of pages) according to their “level of risk” of public exposure, in order to control their access and avoid as much as possible the plundering of Spanish underwater heritage. Enrique Vidal 0001, Verónica Romero 0001, Alejandro H. Toselli, Joan-Andreu Sánchez, Vicente Bosch, Lorenzo Quirós, José-Miguel Benedí, José Ramón Prieto, Moisés Pastor, Francisco Casacuberta, Carlos Alonso, Carmen García, Lourdes Márquez, Carmen Orcero |
ICFHR | 4 |
| 2020 | Handwritten Music Recognition Improvement through Language Model Re-interpretation for Mensural NotationabstractHandwritten Music Recognition studies techniques for computers to transcribe handwritten musical notation that is registered in document images into electronic format, and to make this music available to the public. This task has been of great interest lately, as the technologies improve and can get better and better results on this problem. Recent machine intelligent approaches based on Deep and Recurrent Neural Networks have already shown how they work significantly better in the problem than traditional HMM-based approaches, especially when we are talking about Mensural Notation. These Neural Network-based researches have investigated the task of recognizing Mensural Notation as another written text recognition task, but have not explored the characteristics of musical elements in depth. Other papers have tried to dig deeper into analyzing musical elements and the extraction of their characteristics from segmented symbols, without reflecting this in holistic way. In this paper, we will try to make a complete recognition system directly from the scores, using techniques that enhance information obtained from symbols. We explore other language model interpretations and test our proposal on a publicly available dataset. In our experiments, we have made a 31% relative improvement in regards to error at the symbol level. With this, we have gone from a 3.91% absolute error rate, using Neural Network-based technology, to a 2.70% absolute error rate, by using language model re-interpretations. Manuel Villarreal, Joan-Andreu Sánchez |
ICFHR | 2 |
| 2020 | The HisClima database: historical weather logs for automatic transcription and information extractionabstractKnowing the weather and atmospheric conditions from the past can help weather researchers to generate models like the ones used to predict how weather conditions are likely to change as global temperatures continue to rise. Many historical weather records are available from the past registered on a systemic basis. Historical weather logs were registered in ships, when they were on the high seas, recording daily weather conditions such as: wind speed, temperature, coordinates, etc. These historical documents represent an important source of knowledge with valuable information to extract climatic information of several centuries ago. This paper presents a database for researching about the capability of state-of-the-art handwritten text recognition systems to extract relevant information from this source of knowledge. This database is composed mainly by handwritten tables that contain mostly numerical information. The problems on these documents are explained and baseline experiments are introduced on them. Verónica Romero 0001, Joan-Andreu Sánchez |
ICPR | 2 |
| 2020 | Generation of Hypergraphs from the N-Best Parsing of 2D-Probabilistic Context-Free Grammars for Mathematical Expression RecognitionabstractWe consider hypergraphs as a tool obtained with bidimensional Probabilistic Context-Free Grammars to compactly represent the result of the n-best parse trees for an input image that represents a mathematical expression. More specifically, in this paper we propose: i) an algorithm to compute the N-best parse trees from a 2D-PCFGs, ii) an algorithm to represent the n-best parse trees using a compact representation in the form of hypergraphs, and iii) a formal framework for the development of inference algorithms (inside and outside) and normalization strategies of hypergraphs. Ernesto Noya, Joan-Andreu Sánchez, José-Miguel Benedí |
ICPR | 2 |
| 2020 | Computation of moments for probabilistic finite-state automata
Joan-Andreu Sánchez, Verónica Romero 0001 |
Inf. Sci. | 1 |
| 2019 | Music Symbol Sequence Indexing in Medieval Plainchant ManuscriptsabstractHuge amounts of musical manuscripts are preserved in cathedrals, abbeys, and archives. However, without reliable transcripts, their contents are inaccessible. Manual transcription is unaffordable for large collections, and current automatic technologies-such as Optical Music Recognition or Handwritten Music Recognition-do not provide sufficient accuracy for a fully-automatic scenario. In many cases, perfect transcripts are not really needed, given that content-based search with some degree of reliability would already be extremely useful. Spotting just single music symbols is rather useless (most of the symbols generally appear in all pages); instead, helpful search targets are melodic patterns, which typically correspond to music symbol sequences. We explore approaches for accurate retrieval of melodic patterns, represented by music symbol sequences, from collections of Medieval plainchant manuscripts. Our statistical framework, based on the use of convolutional recurrent neural networks and probabilistic indices, is shown to be useful for retrieving music patterns which appear frequently in this untranscribed images, yielding an Average Precision of 86 %. Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001, Joan-Andreu Sánchez |
ICDAR | 4 |
| 2019 | Making Two Vast Historical Manuscript Collections Searchable and Extracting Meaningful Textual Features Through Large-Scale Probabilistic IndexingabstractTextual access to large collections of digitized images remains unfeasible because usually they lack transcripts. Transcribing such collections is in turn typically unattainable in terms of costs. However, the use of probabilistic indices can facilitate textual accessing with only moderate demands of resources. Besides allowing effortless information retrieval, it will be shown that probabilistic indices can also be used to estimate textual features of the indexed but otherwise untranscribed collections, such as running words and Zipf's curves. Complete probabilistic indices have been recently produced for two iconic large collections: "Bentham" (90K images) and "Spanish Golden Age Theater" (40K images). To show the repercussion of making these collections searchable, we provide accessing statistics gathered through their corresponding search interfaces. To the best of our knowledge this is the first publication of large collections of untranscribed manuscripts which are now publicly accessible for effective and efficient textual access. Alejandro H. Toselli, Verónica Romero 0001, Joan-Andreu Sánchez, Enrique Vidal 0001 |
ICDAR | 3 |
| 2019 | A set of benchmarks for Handwritten Text Recognition on historical documents
Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Mauricio Villegas, Enrique Vidal 0001 |
Pattern Recognit. | 1 |
| 2018 | Automatic Alignment of Handwritten Images and Transcripts for Training Handwritten Text Recognition SystemsabstractState-of-the-art Handwritten Text Recognition techniques are based on statistical models such as hidden Markov models or recurrent neural networks for optical modeling of characters and N-grams for language modeling. These models are trained using well known, learning techniques: Expectation-Maximization, backpropagation, etc. Therefore, training data is needed to build these models. In the case of the optical models the training data consist of text line images with their corresponding transcripts. When the transcript of a handwritten document is available, putting in correspondence automatically the physical lines in the images with the lines of the transcripts is not an easy task. We present a method for automatically aligning handwritten text images and their respective transcripts. The approach automatically segments the images into lines and then recognizes them. An alignment confidence is obtained using the Levenshtein distance between the recognition results and the transcripts. The most confident lines are then used for training. Experiments carried out using a historical document present encouraging results. Verónica Romero 0001, Alejandro H. Toselli, Vicente Bosch, Joan-Andreu Sánchez, Enrique Vidal 0001 |
DAS | 4 |
| 2018 | Active Learning in Handwritten Text Recognition using the Derivational EntropyabstractHandwritten Text Recognition systems are based on statistical models such as recurrent neural networks or hidden Markov models for optical modeling of characters. These models need large corpora for training, consisting in text line images with their corresponding transcripts. The manual annotation of this training data is expensive because it is carried out by experts in paleography, who are specialized in reading ancient scripts. An alternative to reduce the annotation human effort is to use Active Learning techniques to selecting the most informative samples to be used for training. In this paper we study an Active Learning technique to selecting the most informative samples in an HTR scenario. The expert paleographer transcribes only the most informative samples in each stage. The technique followed here is based in the derivational entropy computed from word-graphs obtained from the recognition process. Verónica Romero 0001, Joan-Andreu Sánchez, Alejandro H. Toselli |
ICFHR | 2 |
| 2018 | On the Derivational Entropy of Left-to-Right Probabilistic Finite-State Automata and Hidden Markov ModelsabstractProbabilistic finite-state automata are a formalism that is widely used in many problems of automatic speech recognition and natural language processing. Probabilistic finite-state automata are closely related to other finite-state models as weighted finite-state automata, word lattices, and hidden Markov models. Therefore, they share many similar properties and problems. Entropy measures of finite-state models have been investigated in the past in order to study the information capacity of these models. The derivational entropy quantifies the uncertainty that the model has about the probability distribution it represents. The derivational entropy in a finite-state automaton is computed from the probability that is accumulated in all of its individual state sequences. The computation of the entropy from a weighted finite-state automaton requires a normalized model. This article studies an efficient computation of the derivational entropy of left-to-right probabilistic finite-state automata, and it introduces an efficient algorithm for normalizing weighted finite-state automata. The efficient computation of the derivational entropy is also extended to continuous hidden Markov models. Joan-Andreu Sánchez, Martha-Alicia Rocha, Verónica Romero 0001, Mauricio Villegas |
Comput. Linguistics | 1 |
| 2017 | ICDAR2017 Competition on Information Extraction in Historical Handwritten RecordsabstractThe extraction of relevant information from historical handwritten document collections is one of the key steps in order to make these manuscripts available for access and searches. In this competition, the goal is to detect the named entities and assign each of them a semantic category, and therefore, to simulate the filling in of a knowledge database. This paper describes the dataset, the tasks, the evaluation metrics, the participants methods and the results. Alicia Fornés, Verónica Romero 0001, Arnau Baró, Juan Ignacio Toledo, Joan-Andreu Sánchez, Enrique Vidal 0001, Josep Lladós 0001 |
ICDAR | 5 |
| 2017 | ICDAR2017 Competition on Handwritten Text Recognition on the READ DatasetabstractThis paper describes the fourth edition of the Handwritten Text Recognition (HTR) competition that was prepared this time in the context of the International Conference on Document Analysis and Recognition (ICDAR) 2017. Previous editions of this competition were conducted, first, with datasets from the tranScriptorium project in ICFHR 2014, and ICDAR 2015, and then, with datasets from the "Recognition and Enrichment of Archival Documents (READ)" European project in ICFHR 2016. This competition aims to bring together researchers working on off-line HTR and provides them a suitable benchmark to compare their techniques on the task of transcribing typical and difficult historical handwritten documents. The competition proposed for ICDAR 2017 aims at introducing a usual scenario for some collections in which there exist transcripts at page level for many pages useful for training, but these transcripts are not aligned with line images. Two tracks with different conditions on the use of training data were proposed. Most of the data comes from the Alfred Escher Letter Collection. But handwritten images were drawn from other German collections written by several hands. Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Mauricio Villegas, Enrique Vidal 0001 |
ICDAR | 1 |
| 2016 | Handwriting Transcription and Keyword Spotting in Historical Daily Records DocumentsabstractHistorical records of daily activities provide an intriguing look into the historic life. These documents have interesting information, useful for demography studies and genealogical research. However, automatic processing of historical documents, has mostly been focused on single works of literature and less on daily records, which tend to have a distinct layout, structure, and vocabulary. This paper presents a study about the capability of state-of-the-art handwritten text recognition and key word spotting systems, when applied to this kind of documents. A relatively small set of handwritten birth records registered in Wien in the 16th century is used in the experiments. A word accuracy of about 70% and an AP of 0.74 are achieved for plain image transcription and key word spotting respectively. Taking into account the many difficulties exhibited by these handwritten documents, these preliminary results are quite encouraging. Verónica Romero 0001, Alejandro H. Toselli, Joan-Andreu Sánchez, Enrique Vidal 0001 |
DAS | 3 |
| 2016 | Using the MGGI Methodology for Category-Based Language Modeling in Handwritten Marriage Licenses BooksabstractHandwritten marriage licenses books have been used for centuries by ecclesiastical and secular institutions to register marriages. The information contained in these historical documents is useful for demography studies and genealogical research, among others. Despite the generally simple structure of the text in these documents, automatic transcription and semantic information extraction is difficult due to the distinct and evolutionary vocabulary, which is composed mainly of proper names that change along the time. In previous works we studied the use of category-based language models to both improve the automatic transcription accuracy and make easier the extraction of semantic information. Here we analyze the main causes of the semantic errors observed in previous results and apply a Grammatical Inference technique known as MGGI to improve the semantic accuracy of the language model obtained. Using this language model, full handwritten text recognition experiments have been carried out, with results supporting the interest of the proposed approach. Verónica Romero 0001, Alicia Fornés, Enrique Vidal 0001, Joan-Andreu Sánchez |
ICFHR | 4 |
| 2016 | Handwritten Text Recognition for BengaliabstractHandwritten text recognition of Bengali is a difficult task because of complex character shapes due to the presence of modified/compound characters as well as zone-wise writing styles of different individuals. Most of the research published so far on Bengali handwriting recognition deals with either isolated character recognition or isolated word recognition, and just a few papers have researched on recognition of continuous handwritten Bengali. In this paper we present a research on continuous handwritten Bengali. We follow a classical line-based recognition approach with a system based on hidden Markov models and n-gram language models. These models are trained with automatic methods from annotated data. We research both on the maximum likelihood approach and the minimum error phone approach for training the optical models. We also research on the use of word-based language models and character-based language models. This last approach allow us to deal with the out-of-vocabulary word problem in the test when the training set is of limited size. From the experiments we obtained encouraging results. Joan-Andreu Sánchez, Umapada Pal 0001 |
ICFHR | 1 |
| 2016 | ICFHR2016 Competition on Handwritten Text Recognition on the READ DatasetabstractThis paper describes the Handwritten Text Recognition (HTR) competition on the READ dataset that has been held in the context of the International Conference on Frontiers in Handwriting Recognition 2016. This competition aims to bring together researchers working on off-line HTR and provide them a suitable benchmark to compare their techniques on the task of transcribing typical historical handwritten documents. Two tracks with different conditions on the use of training data were proposed. Ten research groups registered in the competition but finally five submitted results. The handwritten images for this competition were drawn from the German document Ratsprotokolle collection composed of minutes of the council meetings held from 1470 to 1805, used in the READ project. The selected dataset is written by several hands and entails significant variabilities and difficulties. The five participants achieved good results with transcriptions word error rates ranging from 21% to 47%. Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 1 |
| 2016 | An integrated grammar-based approach for mathematical expression recognition
Francisco Alvaro, Joan-Andreu Sánchez, José-Miguel Benedí |
Pattern Recognit. | 2 |
| 2015 | Influence of text line segmentation in Handwritten Text RecognitionabstractText line segmentation is the process by which text lines in a document image are localized and extracted. It is an important step in off-line Handwritten Text Recognition (HTR) given that the input of these systems is the line image of the text to be transcribed. A myriad of solutions to the text line segmentation problem have been proposed in the literature. Although these solutions may differ greatly on what is actually applied to perform the segmentation, they can be classified by the level of precision and detail in the final extracted lines. In this paper we study the influence and real needs of different levels of precision and detail in the segmentation solutions in a real HTR task. We test three technics of text line segmentation whose output range from a simple rectangle for each line to a perfect fitted polygon surrounding the detected lines. Experiments have been carried out with a historical collection and results show that good HTR accuracy can be obtained with simple extraction algorithms. Verónica Romero 0001, Joan-Andreu Sánchez, Vicente Bosch, Katrien Depuydt, Jesse de Does |
ICDAR | 2 |
| 2015 | ICDAR 2015 competition HTRtS: Handwritten Text Recognition on the tranScriptorium datasetabstractThis paper describes the second edition of the Handwritten Text Recognition (HTR) contest on the tranScriptorium datasets that has been held in the context of the International Conference on Document Analysis and Recognition 2015. Two tracks with different conditions on the use of training data were proposed. Nine research groups registered in the contest but finally three research submitted results. The handwritten images for this contest were drawn from the English “Bentham collection” dataset used in the tranScriptorium project. A small subset of this collection has been chosen for the present HTR competition. The selected subset has been written by several hands and entails significant variabilities and difficulties regarding the quality of text images, writing styles and crossed-out text. This contest is clearly more difficult than the the first edition both for training and for testing. A portion of the training dataset and the full test dataset were provided in the form of carefully segmented line images, along with the corresponding transcripts. Another portion of the training dataset was provided as raw images and their corresponding transcripts at region level. The three participants achieved good results, with transcription word error rates ranging from 31% down to 44%. Joan-Andreu Sánchez, Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001 |
ICDAR | 1 |
| 2015 | Crossing the lines: making optimal use of context in line-based Handwritten Text RecognitionabstractHand-written text recognition (HTR) is often carried out line-by-line: the decoding of text lines is carried out independently. This approach is known to deteriorate recognition accuracy of words and characters close to the line boundaries. The present study investigates this issue from the point of view of the language modeling component of the HTR system. Obviously, lack of linguistic context may be one of the reasons for loss of accuracy, but it certainly is not the only factor in play. We seek to clarify to which extent the problem can be influenced by the language modeling component of the system. We first discuss how to develop adapted language models which significantly improve HTR performance in general. We then focus on the deployment of methods to improve accuracy at line boundaries. The final result is an efficient approach which significantly improves HTR accuracy without changing the basic HTR system setup. Jafar Tanha, Jesse de Does, Katrien Depuydt, Joan-Andreu Sánchez |
ICDAR | 4 |
| 2015 | Optical modelling and language modelling trade-off for Handwritten Text RecognitionabstractTraining the models needed for Automatic Handwritten Text Recognition of historical documents generally requires a significant amount of human effort. This is mainly due to the great differences that often exist between collections and to the lack of linguistic resources from the period when the documents were written, which results in a need of manual data labelling effort. This paper presents a study on the reuse of models trained with data from a different collection, focusing on the contribution that the language model and the optical models have on the performance. An empirical evaluation is performed using data from Jeremy Bentham manuscripts with the aim of recognising a manuscript about a very different topic written by Jane Austen. Mauricio Villegas, Joan-Andreu Sánchez, Enrique Vidal 0001 |
ICDAR | 2 |
| 2015 | Structure detection and segmentation of documents using 2D stochastic context-free grammars
Francisco Alvaro, Francisco Cruz 0003, Joan-Andreu Sánchez, Oriol Ramos Terrades, José-Miguel Benedí |
Neurocomputing | 3 |
| 2014 | Ground-Truth Production in the Transcriptorium ProjectabstractTran Scriptorium is a 3-years project that aims to develop innovative, cost-effective solutions for the indexing, search and full transcription of historical handwritten document images, using Handwritten Text Recognition (HTR) technology. The production of ground-truth (GT) of a dataset of handwritten document images is among the first tasks. We address novel approaches for the faster production of this GT based on crowd-sourcing and on prior-knowledge methods. We also address here a novel low-cost semi-supervised procedure for obtaining pairs of correct line-level aligned detected/extracted text line images and text line transcripts, specially suitable for training models of the HTR technology employed in Tran Scriptorium. Basilios Gatos, Georgios Louloudis, Tim Causer, Kris Grint, Verónica Romero 0001, Joan-Andreu Sánchez, Alejandro H. Toselli, Enrique Vidal 0001 |
Document Analysis Systems | 6 |
| 2014 | ICFHR2014 Competition on Handwritten Text Recognition on Transcriptorium Datasets (HTRtS)abstractA contest on Handwritten Text Recognition organised in the context of the ICFHR 2014 conference is described. Two tracks with increased freedom on the use of training data were proposed and three research groups participated in these two tracks. The handwritten images for this contest were drawn from an English data set which is currently being considered in the Tran scriptorium project. The goal of this project is to develop innovative, efficient and cost-effective solutions for the transcription of historical handwritten document images, focusing on four languages: English, Spanish, German and Dutch. For the English language, the so-called "Bentham collection" is being considered in Tran scriptorium. It encompasses a large set of manuscripts written by the renowned English philosopher and reformer Jeremy Bentham (1748-1832). A small subset of this collection has been chosen for the present HTR competition. The selected subset has been written by several hands (Bentham himself and his secretaries) and entails significant variabilities and difficulties regarding the quality of text images and writing styles. Training and test data were provided in the form of carefully segmented line images, along with the corresponding transcripts. The three participants achieved very good results, with transcription word error rates ranging from 15.0% down to 8.6%. Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 1 |
| 2014 | Offline Features for Classifying Handwritten Math Symbols with Recurrent Neural NetworksabstractIn mathematical expression recognition, symbol classification is a crucial step. Numerous approaches for recognizing handwritten math symbols have been published, but most of them are either an online approach or a hybrid approach. There is an absence of a study focused on offline features for handwritten math symbol recognition. Furthermore, many papers provide results difficult to compare. In this paper we assess the performance of several well-known offline features for this task. We also test a novel set of features based on polar histograms and the vertical repositioning method for feature extraction. Finally, we report and analyze the results of several experiments using recurrent neural networks on a large public database of online handwritten math expressions. The combination of online and offline features significantly improved the recognition rate. Francisco Alvaro, Joan-Andreu Sánchez, José-Miguel Benedí |
ICPR | 2 |
| 2014 | Recognition of on-line handwritten mathematical expressions using 2D stochastic context-free grammars and hidden Markov models
Francisco Alvaro, Joan-Andreu Sánchez, José-Miguel Benedí |
Pattern Recognit. Lett. | 2 |
| 2013 | tranScriptorium: a european project on handwritten text recognitionabstractThe tranScriptorium project aims to develop innovative, efficient and cost-effective solutions for annotating handwritten historical documents using modern, holistic Handwritten Text Recognition (HTR) technology. Three actions are planned in tranScriptorium: i) improve basic image preprocessing and holistic HTR techniques; ii) develop novel indexing and keyword searching approaches; and iii) capitalize on new, user-friendly interactive-predictive HTR approaches for computer-assisted operation. Joan-Andreu Sánchez, Günter Mühlberger, Basilios Gatos, Philip Schofield, Katrien Depuydt, Richard M. Davis, Enrique Vidal 0001, Jesse de Does |
ACM Symposium on Document Engineering | 1 |
| 2013 | Classification of On-Line Mathematical Symbols with Hybrid Features and Recurrent Neural NetworksabstractRecognition of on-line handwritten mathematical symbols has been tackled using different methods, but the recognition rates achieved until now still leave room for improvement. Many of the published approaches are based on hidden Markov models, and some of them use off-line information extracted from the on-line data. In this paper, we present a set of hybrid features that combine both on-line and off-line information. Lately, recurrent neural networks have demonstrated to obtain good results and they have outperformed hidden Markov models in several sequence learning tasks, including handwritten text recognition. Hence, we also studied a state-of-the-art recurrent neural network classifier and we compared its performance with a classifier based on hidden Markov models. Experiments using a large public database showed that both the new proposed features and recurrent neural network classifier improved significantly the classification results. Francisco Alvaro, Joan-Andreu Sánchez, José-Miguel Benedí |
ICDAR | 2 |
| 2013 | Category-Based Language Models for Handwriting Recognition of Marriage License BooksabstractHandwritten marriage licenses books have been used for centuries by ecclesiastical institutions to register marriages. These documents have interesting information, useful for demography studies, organized in a list of individual marriage license records, such as an accounting book. The information in these books is usually collected by expert demographers that devote a lot of time to transcribe them. Despite the structure of the text, the automatic transcription and semantic information extraction of these documents is quite difficult due to the distinct and evolutionary vocabulary, which is composed mainly of proper names that change along the time. In this paper, we have defined some categories taking into account the semantic information included in the licenses. Then a category-based language model has been generated and integrated into the handwritten text recognition system. We study how the use of these categories can benefit not only the handwriting recognition step, but also the posterior semantic information extraction and knowledge discovery. Verónica Romero 0001, Joan-Andreu Sánchez |
ICDAR | 2 |
| 2013 | Human Evaluation of the Transcription Process of a Marriage License BookabstractHandwriting Text Recognition (HTR) of historical documents is a very important research field of Document Image Analysis. Currently, the most well-accepted technology for off-line HTR is based on holistic, segmentation-free techniques that do not need any kind of character or word segmentation. This HTR technology is based in stochastic models that are trained with annotated data. The performance of this technology is still far from being perfect and therefore the user intervention is necessary to obtain perfect transcripts. The user intervention can be carried out in a post-editing process, in which the user corrects the errors produced by an automatic HTR system. Interactive techniques have been proposed in the past few years to obtain the correct transcript as an alternative to post-editing the transcripts. In these interactive approaches, the user and the system work interactively in tight mutual collaboration to obtain the perfect transcript of the data. In this interactive scenario, the feedback provided by the user is used to improve interactively the system output. In the post-editing scenario and in the interactive scenario, the transcribed material can be used for retraining the models as the data is processed. In this research we carried out a study with a real transcriber about how the performance of an HTR system improved with respect to the amount of training data, and how the human efficiency improved during the transcription process in both transcription scenarios. Verónica Romero 0001, Joan-Andreu Sánchez |
ICDAR | 2 |
| 2013 | The ESPOSALLES database: An ancient marriage license corpus for off-line handwriting recognition
Verónica Romero 0001, Alicia Fornés, Joan-Andreu Sánchez, Alejandro H. Toselli, Volkmar Frinken, Enrique Vidal 0001, Josep Lladós 0001 |
Pattern Recognit. | 4 |
| 2012 | Unbiased Evaluation of Handwritten Mathematical Expression RecognitionabstractSeveral approaches have been proposed to tackle the problem of mathematical expression recognition, and automatic methods for performance evaluation are required. Mathematical expressions are usually encoded as a LaTeX string or a tree (MathML) for evaluation purpose, but these formats do not enforce uniqueness. Consequently, given that there can be several representations syntactically different but semantically equivalent, the automatic performance evaluation of mathematical expressions can be biased. Given a mathematical expression recognition tree and its ground-truth tree, the error is usually computed by comparing them. In this paper we propose to obtain a new tree, equivalent to the ground-truth tree, according to the model representation criteria. Then, we can compute an error by comparing the recognized tree with the obtained by using the model, both with the same bias. Several experiments were carried out in order to evaluate this approach and results showed that representation criteria had a significative effect in the evaluation results. Francisco Alvaro, Joan-Andreu Sánchez, José-Miguel Benedí |
ICFHR | 2 |
| 2011 | Recognition of Printed Mathematical Expressions Using Two-Dimensional Stochastic Context-Free GrammarsabstractIn this work, a system for recognition of printed mathematical expressions has been developed. Hence, a statistical framework based on two-dimensional stochastic context-free grammars has been defined. This formal framework allows to jointly tackle the segmentation, symbol recognition and structural analysis of a mathematical expression by computing its most probable parsing. In order to test this approach a reproducible and comparable experiment has been carried out over a large publicly available (InftyCDB-1) database. Results are reported using a well-defined global dissimilitude measure. Experimental results show that this technique is able to properly recognize mathematical expressions, and that the structural information improves the symbol recognition step. Francisco Alvaro, Joan-Andreu Sánchez, José-Miguel Benedí |
ICDAR | 2 |
| 2011 | Handwritten Text Recognition for Marriage Register BooksabstractMarriage register books are documents that were used for centuries by ecclesiastical institutions to register marriages. Most of these books were handwritten. These documents have interesting information, useful for demography studies. The information in these books is usually collected by expert demographers that devote a lot of time to transcribe them. The automatic transcription of these documents by using Handwritten Text Recognition techniques is difficult since the vocabulary is large, given that it is composed mainly of proper names. In this work, interactive Handwritten Text Recognition techniques were studied for the assisted transcription of these documents. Verónica Romero 0001, Joan-Andreu Sánchez, Enrique Vidal 0001 |
ICDAR | 2 |
| 2010 | Syntax Augmented Inversion Transduction Grammars for Machine Translation
Guillem Gascó i Mora, Joan-Andreu Sánchez |
CICLing | 2 |
| 2010 | Comparing Several Techniques for Offline Recognition of Printed Mathematical SymbolsabstractAutomatic recognition of printed mathematical symbols is a fundamental problem for recognition of mathematical expressions. Several classification techniques has been previously used, but there are very few works that compare different classification techniques on the same database and with the same experimental conditions. In this work we have tested classical and novelty classification techniques for mathematical symbol recognition on two databases. Francisco Alvaro, Joan-Andreu Sánchez |
ICPR | 2 |
| 2010 | Enlarged Search Space for SITG Parsing
Guillem Gascó i Mora, Joan-Andreu Sánchez, José-Miguel Benedí |
HLT-NAACL | 2 |
| 2008 | Using Parsed Corpora for Estimating Stochastic Inversion Transduction Grammars
Germán Sanchis-Trilles, Joan-Andreu Sánchez |
LREC | 2 |
| 2006 | Obtaining Word Phrases with Stochastic Inversion Translation Grammars for Phrase-based Statistical Machine Translation
Joan-Andreu Sánchez, José-Miguel Benedí |
EAMT | 1 |
| 2005 | Estimation of stochastic context-free grammars and their use as language models
José-Miguel Benedí, Joan-Andreu Sánchez |
Comput. Speech Lang. | 2 |
| 2004 | A hybrid language model based on a combination of N-grams and stochastic context-free grammarsabstractIn this paper, a hybrid language model is defined as a combination of a word-based n-gram, which is used to capture the local relations between words, and a category-based stochastic context-free grammar (SCFG) with a word distribution into categories, which is defined to represent the long-term relations between these categories. The problem of unsupervised learning of a SCFG in General Format and in Chomsky Normal Form by means of estimation algorithms is studied. Moreover, a bracketed version of the classical estimation algorithm based on the Earley algorithm is proposed. This paper also explores the use of SCFGs obtained from a treebank corpus as initial models for the estimation algorithms. Experiments on the UPenn Treebank corpus are reported. These experiments have been carried out in terms of the test set perplexity and the word error rate in a speech recognition experiment. Diego Linares, José-Miguel Benedí, Joan-Andreu Sánchez |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2000 | Combination Of N-Grams And Stochastic Context-Free Grammars For Language Modeling
José-Miguel Benedí, Joan-Andreu Sánchez |
COLING | 2 |
| 1999 | Acoustic and syntactical modeling in the ATROS systemabstractCurrent speech technology allows us to build efficient speech recognition systems. However, model learning of knowledge sources in a speech recognition system is not a closed problem. In addition, lower demand of computational requirements are crucial to building real-time systems. ATROS is an automatic speech recognition system whose acoustic, lexical, and syntactical models can be learnt automatically from training data by using similar techniques. In this paper, an improved version of ATROS which can deal with large smoothed language models and with large vocabularies is presented. This version supports acoustic and syntactical models trained with advanced grammatical inference techniques. It also incorporates new data structures and improved search algorithms to reduce the computational requirements for decoding. The system has been tested on a Spanish task of queries to a geographical database (with a vocabulary of 1,208 words). David Llorens, Francisco Casacuberta, Encarna Segarra, Joan-Andreu Sánchez, Pablo Aibar, María José Castro Bleda |
ICASSP | 4 |
| 1999 | A fast version of the atros system
María José Castro Bleda, David Llorens, Joan-Andreu Sánchez, Francisco Casacuberta, Pablo Aibar, Encarna Segarra |
EUROSPEECH | 3 |
| 1999 | Learning of stochastic context-free grammars by means of estimation algorithms
Joan-Andreu Sánchez, José-Miguel Benedí |
EUROSPEECH | 1 |
| 1998 | Estimation of the probability distributions of stochastic context-free grammars from the k-best derivationsabstractThe use of the Inside-Outside (IO) algorithm for the estimation of the probability distributions of Stochastic Context-Free Grammars (SCFGs) in Natural-Language processing is restricted due to the time complexity per iteration and the large number of iterations that it needs to converge. Alternatively, an algorithm based on the Viterbi score (VS) is used. This VS algorithm converges more rapidly, but obtains less competitive models. We describe here a new algorithm that only considers the k-best derivations in the estimation process. The experimental results show that this algorithm achieves faster convergence than the IO and better models than the VS algorithm. Joan-Andreu Sánchez, José-Miguel Benedí |
ICSLP | 1 |
| 1997 | Consistency of Stochastic Context-Free Grammars From Probabilistic Estimation Based on Growth TransformationsabstractAn important problem related to the probabilistic estimation of stochastic context-free grammars (SCFGs) is guaranteeing the consistency of the estimated model. This problem was considered by Booth-Thompson (1973) and Wetherell (1980) and studied by Maryanski (1974) and Chaudhuri et al. (1983) for unambiguous SCFGs only, when the probability distributions were estimated by the relative frequencies in a training sample. In this work, we extend this result by proving that the property of consistency is guaranteed for all SCFGs without restrictions, when the probability distributions are learned from the classical inside-outside and Viterbi algorithms, both of which are based on growth transformations. Other important probabilistic properties which are related to these results are also proven. Joan-Andreu Sánchez, José-Miguel Benedí |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |