VLDB 2026 Research / reviewers in the wild / expert
Verónica Romero 0001
dblp:85/1037-1 · also Verónica Romero-Gomez
· DBLP profile ↗
55ranked-venue papers
15as first author
9since 2021 · last 2025
0000-0002-1721-5732ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 8 first-author · 5 since 2021Databases, data management, data science and information retrieval · 26 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Lightweight Named Entity Recognition in Handwritten Documents by Predicting Pyramidal Histograms of CharactersabstractNamed Entity Recogniton (NER) consists of tagging parts of an unstructured text containing particular semantic information. When applied to handwritten documents, it is possible to do it as a two-step approach in which Handwritten Text Recognition (HTR) is performed prior to tagging the automatic transcription. However, it is also possible to do both tasks simultaneously by using an HTR model that learns to output the transcription and the tagging symbols. In this paper, we focus on improving the one-step approach by introducing the auxiliary task of predicting Pyramidal Histograms of Characters (PHOC) in a Convolutional Recurrent Neural Network (CRNN) model. Moreover, given the recent rise of models that digest large amounts of data, we also study the usage of synthetic data to pretrain the proposed architecture. Our experiments show that by pretraining the PHOC-based architecture on synthetic data substantial improvements can be made in both transcription and tagging quality without compromising the computational cost of the decoding step. The resulting model matches the NER performance of the state-of-the-art while keeping its lightweight nature. David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor |
DocEng | 3 |
| 2024 | Reading Order Independent Metrics for Information Extraction in Handwritten Documents
David Villanova-Aparisi, Solène Tarride, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Christopher Kermorvant, Moisés Pastor |
ICDAR (2) | 4 |
| 2024 | Towards a general user model to develop intelligent user interfacesabstractAbstract The way end-users interact with a system plays a crucial role in the high acceptance of software. Related to this, the concept of Intelligent User Interfaces has emerged as a solution to learn from user interactions with the system and adapt interfaces to the user’s characteristics and preferences. However, existing approaches to designing intelligent user interfaces are limited by their user models, which are not capable of representing each and every user characteristic valid for any context. This work aims to address this limitation by presenting a user model that can abstractly represent a wide set of user characteristics in any context of interaction. The model is based on a synthesis of previous works that have proposed specific user models. After the analysis of these works, a more sophisticated user model has been defined, including some required characteristics not existing in previous works. This model has been validated with 62 real end-users who have expressed the users’ characteristics that they consider as relevant to adapt the interaction. The results show that most of these characteristics can be represented by the proposed user model. This user model is the first step towards creating intelligent user interfaces that can adapt interactions to users with similar characteristics and preferences in similar contexts. Alberto Gaspar, Miriam Gil, José Ignacio Panach, Verónica Romero 0001 |
Multim. Tools Appl. | 4 |
| 2023 | Consistent Nested Named Entity Recognition in Handwritten Documents via Lattice Rescoring
David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor |
ICDAR (1) | 3 |
| 2023 | Evaluation of Different Tagging Schemes for Named Entity Recognition in Handwritten Documents
David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor |
ICDAR (3) | 3 |
| 2023 | Processing a large collection of historical tabular imagesabstractProcessing automatically historical document images to allow the search of textual information requires the preparation of ground-truth data for training and evaluation. This process is an expensive and arduous task, especially when the historical document images contain specialized vocabulary and/or tabular information. In the latter case, relevant decisions have to be taken to annotate the tabular parts. This paper presents a complex collection of historical document images and the resulting database, which is called HisClima. In this database, half of the images are in tabular format and half as running text. Both types of images contain pre-printed and handwritten text. The textual information is plenty of abbreviations and specific vocabulary related to weather conditions and old ships. This database can be used to research technologies related to historical document image processing and analysis, both for tabular and running text recognition. Baseline results are presented for Document Layout Analysis, Text Recognition, and Probabilistic Indexing. Although these results are good, there is still room for improvement and some indications are provided in this direction. Emilio Granell, Verónica Romero 0001, José Ramón Prieto, José Andrés, Lorenzo Quirós, Joan-Andreu Sánchez, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | Information Extraction from Handwritten Tables in Historical Documents
José Andrés, José Ramón Prieto, Emilio Granell, Verónica Romero 0001, Joan-Andreu Sánchez, Enrique Vidal 0001 |
DAS | 4 |
| 2022 | Evaluation of Named Entity Recognition in Handwritten Documents
David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor |
DAS | 3 |
| 2021 | Reducing the Human Effort in Text Line Segmentation for Historical Documents
Emilio Granell, Lorenzo Quirós, Verónica Romero 0001, Joan-Andreu Sánchez |
ICDAR (3) | 3 |
| 2020 | A comparison of sequential and combined approaches for named entity recognition in a corpus of handwritten medieval chartersabstractThis paper introduces a new corpus of multilingual medieval handwritten charter images, annotated with full transcription and named entities. The corpus is used to compare two approaches for named entity recognition in historical document images in several languages: on the one hand, a sequential approach, more commonly used, that sequentially applies handwritten text recognition (HTR) and named entity recognition (NER), on the other hand, a combined approach that simultaneously transcribes the image text line and extracts the entities. Experiments conducted on the charter corpus in Latin, early new high German and old Czech for name, date and location recognition demonstrate a superior performance of the combined approach. Emanuela Boros, Verónica Romero 0001, Martin Maarand, Katerina Zenklová, Jitka Krecková, Enrique Vidal 0001, Dominique Stutzmann, Christopher Kermorvant |
ICFHR | 2 |
| 2020 | The Carabela Project and Manuscript Collection: Large-Scale Probabilistic Indexing and Content-based ClassificationabstractThe main aim of the Carabela project was to develop and apply techniques that allow textual searching on massive Spanish collections of 15th-19th century manuscripts. The project focused on a relatively small subset of 125 000 images of collections of interest to underwater archaeology. For this type of manuscripts, state-of-the-art automatic transcription techniques, generally fail to achieve usable transcription accuracy. Therefore, rather than insisting in actual transcription, methodologies for probabilistic indexing of handwritten text images have been adopted. This has allowed us to effectively cope with the intrinsically high degree of uncertainty of the text contained in most historical manuscripts, leading to highly effective systems for textual search and retrieval. Carabela has gone one step further by developing new techniques to classify probabilistically indexed, but otherwise untranscribed, text images according to their textual content. These techniques have been successfully used to automatically classify Carabela bundels (each containing hundreds or thousands of pages) according to their “level of risk” of public exposure, in order to control their access and avoid as much as possible the plundering of Spanish underwater heritage. Enrique Vidal 0001, Verónica Romero 0001, Alejandro H. Toselli, Joan-Andreu Sánchez, Vicente Bosch, Lorenzo Quirós, José-Miguel Benedí, José Ramón Prieto, Moisés Pastor, Francisco Casacuberta, Carlos Alonso, Carmen García, Lourdes Márquez, Carmen Orcero |
ICFHR | 2 |
| 2020 | The HisClima database: historical weather logs for automatic transcription and information extractionabstractKnowing the weather and atmospheric conditions from the past can help weather researchers to generate models like the ones used to predict how weather conditions are likely to change as global temperatures continue to rise. Many historical weather records are available from the past registered on a systemic basis. Historical weather logs were registered in ships, when they were on the high seas, recording daily weather conditions such as: wind speed, temperature, coordinates, etc. These historical documents represent an important source of knowledge with valuable information to extract climatic information of several centuries ago. This paper presents a database for researching about the capability of state-of-the-art handwritten text recognition systems to extract relevant information from this source of knowledge. This database is composed mainly by handwritten tables that contain mostly numerical information. The problems on these documents are explained and baseline experiments are introduced on them. Verónica Romero 0001, Joan-Andreu Sánchez |
ICPR | 1 |
| 2020 | Study of the influence of lexicon and language restrictions on computer assisted transcription of historical manuscripts
Emilio Granell, Verónica Romero 0001, Carlos D. Martínez-Hinarejos |
Neurocomputing | 2 |
| 2020 | Computation of moments for probabilistic finite-state automata
Joan-Andreu Sánchez, Verónica Romero 0001 |
Inf. Sci. | 2 |
| 2019 | Making Two Vast Historical Manuscript Collections Searchable and Extracting Meaningful Textual Features Through Large-Scale Probabilistic IndexingabstractTextual access to large collections of digitized images remains unfeasible because usually they lack transcripts. Transcribing such collections is in turn typically unattainable in terms of costs. However, the use of probabilistic indices can facilitate textual accessing with only moderate demands of resources. Besides allowing effortless information retrieval, it will be shown that probabilistic indices can also be used to estimate textual features of the indexed but otherwise untranscribed collections, such as running words and Zipf's curves. Complete probabilistic indices have been recently produced for two iconic large collections: "Bentham" (90K images) and "Spanish Golden Age Theater" (40K images). To show the repercussion of making these collections searchable, we provide accessing statistics gathered through their corresponding search interfaces. To the best of our knowledge this is the first publication of large collections of untranscribed manuscripts which are now publicly accessible for effective and efficient textual access. Alejandro H. Toselli, Verónica Romero 0001, Joan-Andreu Sánchez, Enrique Vidal 0001 |
ICDAR | 2 |
| 2019 | Image-speech combination for interactive computer assisted transcription of handwritten documents
Emilio Granell, Verónica Romero 0001, Carlos D. Martínez-Hinarejos |
Comput. Vis. Image Underst. | 2 |
| 2019 | A set of benchmarks for Handwritten Text Recognition on historical documents
Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Mauricio Villegas, Enrique Vidal 0001 |
Pattern Recognit. | 2 |
| 2018 | Comparing Different Feedback Modalities in Assisted Transcription of ManuscriptsabstractTranscription of handwritten text can be speed-up by using off-line Handwritten Text Recognition techniques, that allow the obtention of an initial draft transcription of an image with handwritten text. However, this draft transcription usually contains errors that must be amended by the transcriber by providing a feedback signal. The usual approach is post-edition, where each error is corrected without modifying the rest of the current transcription. A more sophisticated approach can employ the current modification to provide a new whole transcription, hopefully with less errors. Apart from that, feedback can be provided in different modalities: keyboard input, on-line handwritten text, or speech. Each of these modalities presents different features with respect to ambiguity, derived errors, and final transcription time. In this work we study how the different modalities behave in the assisted transcription of a historical handwritten text document in Spanish and we evaluate their transcription productivity. Carlos D. Martínez-Hinarejos, Emilio Granell, Verónica Romero 0001 |
DAS | 3 |
| 2018 | Automatic Alignment of Handwritten Images and Transcripts for Training Handwritten Text Recognition SystemsabstractState-of-the-art Handwritten Text Recognition techniques are based on statistical models such as hidden Markov models or recurrent neural networks for optical modeling of characters and N-grams for language modeling. These models are trained using well known, learning techniques: Expectation-Maximization, backpropagation, etc. Therefore, training data is needed to build these models. In the case of the optical models the training data consist of text line images with their corresponding transcripts. When the transcript of a handwritten document is available, putting in correspondence automatically the physical lines in the images with the lines of the transcripts is not an easy task. We present a method for automatically aligning handwritten text images and their respective transcripts. The approach automatically segments the images into lines and then recognizes them. An alignment confidence is obtained using the Levenshtein distance between the recognition results and the transcripts. The most confident lines are then used for training. Experiments carried out using a historical document present encouraging results. Verónica Romero 0001, Alejandro H. Toselli, Vicente Bosch, Joan-Andreu Sánchez, Enrique Vidal 0001 |
DAS | 1 |
| 2018 | Text Line Extraction Based on Distance Map Features and Dynamic ProgrammingabstractText Line Segmentation is a basic document layout task that consists in detecting and extracting the text lines present in a document page image. Although considered a basic task, generally, it is a necessary step for Handwritten Text Recognition (HTR) higher level tasks. Most state of the art automatic text recognition, text-to-line image alignment and key word spotting systems require it due to their need for isolated text line images as input. Traditionally most Text Line Segmentation approaches cover both detection and extraction sub steps. However, the community has recently shifted its focus to tackle independently the baseline detection in document images. This shift generates the need for extraction methods that use these detected baselines as input. In this paper, a binarization free dynamic programming approach that generates an equidistant text line extraction polygon is presented. The approach performs this calculation, based on the information provided by priorly detected text baselines and automatically generated foreground pixels distance maps. We evaluate our approach both in a synthetic competition corpus and in a challenging real handwritten text recognition task corpus. We evaluate it not only at the graphical error level but also the impact it produces on an HTR task trained with the line images it yields. We compare our solution with other solutions ranging from the actual human reviewed ground-truth polygons to simpler automatic generated rectangle areas. Vicente Bosch, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 2 |
| 2018 | Active Learning in Handwritten Text Recognition using the Derivational EntropyabstractHandwritten Text Recognition systems are based on statistical models such as recurrent neural networks or hidden Markov models for optical modeling of characters. These models need large corpora for training, consisting in text line images with their corresponding transcripts. The manual annotation of this training data is expensive because it is carried out by experts in paleography, who are specialized in reading ancient scripts. An alternative to reduce the annotation human effort is to use Active Learning techniques to selecting the most informative samples to be used for training. In this paper we study an Active Learning technique to selecting the most informative samples in an HTR scenario. The expert paleographer transcribes only the most informative samples in each stage. The technique followed here is based in the derivational entropy computed from word-graphs obtained from the recognition process. Verónica Romero 0001, Joan-Andreu Sánchez, Alejandro H. Toselli |
ICFHR | 1 |
| 2018 | Multimodality, interactivity, and crowdsourcing for document transcriptionabstractAbstract Knowledge mining from documents usually use document engineering techniques that allow the user to access the information contained in documents of interest. In this framework, transcription may provide efficient access to the contents of handwritten documents. Manual transcription is a time‐consuming task that can be sped up by using different mechanisms. A first possibility is employing state‐of‐the‐art handwritten text recognition systems to obtain an initial draft transcription that can be manually amended. A second option is employing crowdsourcing to obtain a massive but not error‐free draft transcription. In this case, when collaborators employ mobile devices, speech dictation can be used as a transcription source, and speech and handwritten text recognition can be fused to provide a better draft transcription, which can be amended with even less effort. A final option is using interactive assistive frameworks, where the automatic system that provides the draft transcription and the transcriber cooperate to generate the final transcription. The novel contributions presented in this work include the study of the data fusion on a multimodal crowdsourcing framework and its integration with an interactive system. The use of the proposed solutions reduces the required transcription effort and optimizes the overall performance and usability, allowing for a better transcription process. Emilio Granell, Verónica Romero 0001, Carlos D. Martínez-Hinarejos |
Comput. Intell. | 2 |
| 2018 | On the Derivational Entropy of Left-to-Right Probabilistic Finite-State Automata and Hidden Markov ModelsabstractProbabilistic finite-state automata are a formalism that is widely used in many problems of automatic speech recognition and natural language processing. Probabilistic finite-state automata are closely related to other finite-state models as weighted finite-state automata, word lattices, and hidden Markov models. Therefore, they share many similar properties and problems. Entropy measures of finite-state models have been investigated in the past in order to study the information capacity of these models. The derivational entropy quantifies the uncertainty that the model has about the probability distribution it represents. The derivational entropy in a finite-state automaton is computed from the probability that is accumulated in all of its individual state sequences. The computation of the entropy from a weighted finite-state automaton requires a normalized model. This article studies an efficient computation of the derivational entropy of left-to-right probabilistic finite-state automata, and it introduces an efficient algorithm for normalizing weighted finite-state automata. The efficient computation of the derivational entropy is also extended to continuous hidden Markov models. Joan-Andreu Sánchez, Martha-Alicia Rocha, Verónica Romero 0001, Mauricio Villegas |
Comput. Linguistics | 3 |
| 2017 | ICDAR2017 Competition on Information Extraction in Historical Handwritten RecordsabstractThe extraction of relevant information from historical handwritten document collections is one of the key steps in order to make these manuscripts available for access and searches. In this competition, the goal is to detect the named entities and assign each of them a semantic category, and therefore, to simulate the filling in of a knowledge database. This paper describes the dataset, the tasks, the evaluation metrics, the participants methods and the results. Alicia Fornés, Verónica Romero 0001, Arnau Baró, Juan Ignacio Toledo, Joan-Andreu Sánchez, Enrique Vidal 0001, Josep Lladós 0001 |
ICDAR | 2 |
| 2017 | ICDAR2017 Competition on Handwritten Text Recognition on the READ DatasetabstractThis paper describes the fourth edition of the Handwritten Text Recognition (HTR) competition that was prepared this time in the context of the International Conference on Document Analysis and Recognition (ICDAR) 2017. Previous editions of this competition were conducted, first, with datasets from the tranScriptorium project in ICFHR 2014, and ICDAR 2015, and then, with datasets from the "Recognition and Enrichment of Archival Documents (READ)" European project in ICFHR 2016. This competition aims to bring together researchers working on off-line HTR and provides them a suitable benchmark to compare their techniques on the task of transcribing typical and difficult historical handwritten documents. The competition proposed for ICDAR 2017 aims at introducing a usual scenario for some collections in which there exist transcripts at page level for many pages useful for training, but these transcripts are not aligned with line images. Two tracks with different conditions on the use of training data were proposed. Most of the data comes from the Alfred Escher Letter Collection. But handwritten images were drawn from other German collections written by several hands. Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Mauricio Villegas, Enrique Vidal 0001 |
ICDAR | 2 |
| 2017 | Word graphs size impact on the performance of handwriting document applications
Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001 |
Neural Comput. Appl. | 2 |
| 2016 | An Interactive Approach with Off-Line and On-Line Handwritten Text Recognition Combination for Transcribing Historical DocumentsabstractAutomatic transcription of historical documents is becoming an important research topic, specially because of the increasing number of digitised historical documents that libraries and archives are publishing. However, state-of-the-art handwritten text recognition systems are far from being perfect. Therefore, to have perfect transcriptions, human expert revision is required to really produce a transcription of standard quality. In this context, an interactive assistive scenario, where the automatic system and the human transcriber cooperate to generate the perfect transcription, would allow for a more effective approach. In this paper we present a multimodal interactive transcription system where user feedback is provided by means of touchscreen pen strokes, traditional keyboard and mouse operations. The combination of both the main and the feedback data stream is based on the use of Confusion Networks derived from the output of the on-line and off-line handwritten text recognition systems. The use of the proposed combination help to optimise overall performance and usability. Emilio Granell, Verónica Romero 0001, Carlos D. Martínez-Hinarejos |
DAS | 2 |
| 2016 | Handwriting Transcription and Keyword Spotting in Historical Daily Records DocumentsabstractHistorical records of daily activities provide an intriguing look into the historic life. These documents have interesting information, useful for demography studies and genealogical research. However, automatic processing of historical documents, has mostly been focused on single works of literature and less on daily records, which tend to have a distinct layout, structure, and vocabulary. This paper presents a study about the capability of state-of-the-art handwritten text recognition and key word spotting systems, when applied to this kind of documents. A relatively small set of handwritten birth records registered in Wien in the 16th century is used in the experiments. A word accuracy of about 70% and an AP of 0.74 are achieved for plain image transcription and key word spotting respectively. Taking into account the many difficulties exhibited by these handwritten documents, these preliminary results are quite encouraging. Verónica Romero 0001, Alejandro H. Toselli, Joan-Andreu Sánchez, Enrique Vidal 0001 |
DAS | 1 |
| 2016 | Using the MGGI Methodology for Category-Based Language Modeling in Handwritten Marriage Licenses BooksabstractHandwritten marriage licenses books have been used for centuries by ecclesiastical and secular institutions to register marriages. The information contained in these historical documents is useful for demography studies and genealogical research, among others. Despite the generally simple structure of the text in these documents, automatic transcription and semantic information extraction is difficult due to the distinct and evolutionary vocabulary, which is composed mainly of proper names that change along the time. In previous works we studied the use of category-based language models to both improve the automatic transcription accuracy and make easier the extraction of semantic information. Here we analyze the main causes of the semantic errors observed in previous results and apply a Grammatical Inference technique known as MGGI to improve the semantic accuracy of the language model obtained. Using this language model, full handwritten text recognition experiments have been carried out, with results supporting the interest of the proposed approach. Verónica Romero 0001, Alicia Fornés, Enrique Vidal 0001, Joan-Andreu Sánchez |
ICFHR | 1 |
| 2016 | ICFHR2016 Competition on Handwritten Text Recognition on the READ DatasetabstractThis paper describes the Handwritten Text Recognition (HTR) competition on the READ dataset that has been held in the context of the International Conference on Frontiers in Handwriting Recognition 2016. This competition aims to bring together researchers working on off-line HTR and provide them a suitable benchmark to compare their techniques on the task of transcribing typical historical handwritten documents. Two tracks with different conditions on the use of training data were proposed. Ten research groups registered in the competition but finally five submitted results. The handwritten images for this competition were drawn from the German document Ratsprotokolle collection composed of minutes of the council meetings held from 1470 to 1805, used in the READ project. The selected dataset is written by several hands and entails significant variabilities and difficulties. The five participants achieved good results with transcriptions word error rates ranging from 21% to 47%. Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 2 |
| 2016 | Exploiting Existing Modern Transcripts for Historical Handwritten Text RecognitionabstractExisting transcripts for historic manuscripts are a very valuable resource for training models useful for automatic recognition, aided transcription, and/or indexing of the remaining untranscribed parts of these collections. However, these existing transcripts generally exhibit two main problems which hinder their convenience: a) text of the transcripts is seldom aligned with manuscript lines, and b) text often deviate very significantly from what can be seen in the manuscript, either because writing style has been modernized or abbreviations have been expanded, or both. This work presents an analysis of these problems and discusses possible solutions for minimizing human effort needed to adapt existing transcripts in order to render them usable. Empirical results presented show the huge performance gain that can be obtained by adequately adapting the transcripts, thus motivating future development of the proposed solutions. Mauricio Villegas, Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001 |
ICFHR | 3 |
| 2016 | HMM word graph based keyword spotting in handwritten document images
Alejandro H. Toselli, Enrique Vidal 0001, Verónica Romero 0001, Volkmar Frinken |
Inf. Sci. | 3 |
| 2015 | Influence of text line segmentation in Handwritten Text RecognitionabstractText line segmentation is the process by which text lines in a document image are localized and extracted. It is an important step in off-line Handwritten Text Recognition (HTR) given that the input of these systems is the line image of the text to be transcribed. A myriad of solutions to the text line segmentation problem have been proposed in the literature. Although these solutions may differ greatly on what is actually applied to perform the segmentation, they can be classified by the level of precision and detail in the final extracted lines. In this paper we study the influence and real needs of different levels of precision and detail in the segmentation solutions in a real HTR task. We test three technics of text line segmentation whose output range from a simple rectangle for each line to a perfect fitted polygon surrounding the detected lines. Experiments have been carried out with a historical collection and results show that good HTR accuracy can be obtained with simple extraction algorithms. Verónica Romero 0001, Joan-Andreu Sánchez, Vicente Bosch, Katrien Depuydt, Jesse de Does |
ICDAR | 1 |
| 2015 | ICDAR 2015 competition HTRtS: Handwritten Text Recognition on the tranScriptorium datasetabstractThis paper describes the second edition of the Handwritten Text Recognition (HTR) contest on the tranScriptorium datasets that has been held in the context of the International Conference on Document Analysis and Recognition 2015. Two tracks with different conditions on the use of training data were proposed. Nine research groups registered in the contest but finally three research submitted results. The handwritten images for this contest were drawn from the English “Bentham collection” dataset used in the tranScriptorium project. A small subset of this collection has been chosen for the present HTR competition. The selected subset has been written by several hands and entails significant variabilities and difficulties regarding the quality of text images, writing styles and crossed-out text. This contest is clearly more difficult than the the first edition both for training and for testing. A portion of the training dataset and the full test dataset were provided in the form of carefully segmented line images, along with the corresponding transcripts. Another portion of the training dataset was provided as raw images and their corresponding transcripts at region level. The three participants achieved good results, with transcription word error rates ranging from 31% down to 44%. Joan-Andreu Sánchez, Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001 |
ICDAR | 3 |
| 2015 | Context-Aware Gestures for Mixed-Initiative Text Editing UIsabstractThis work is focused on enhancing highly interactive text-editing applications with gestures. Concretely, we study Computer Assisted Transcription of Text Images (CATTI), a handwriting transcription system that follows a corrective feedback paradigm, where both the user and the system collaborate efficiently to produce a high-quality text transcription. CATTI-like applications demand fast and accurate gesture recognition, for which we observed that current gesture recognizers are not adequate enough. In response to this need we developed MinGestures, a parametric context-aware gesture recognizer. Our contributions include a number of stroke features for disambiguating copy-mark gestures from handwritten text, plus the integration of these gestures in a CATTI application. It becomes finally possible to create highly interactive stroke-based text-editing interfaces, without worrying to verify the user intent on-screen. We performed a formal evaluation with 22 e-pen users and 32 mouse users using a gesture vocabulary of 10 symbols. MinGestures achieved an outstanding accuracy (<1% error rate) with very high performance (<1 ms of recognition time). We then integrated MinGestures in a CATTI prototype and tested the performance of the interactive handwriting system when it is driven by gestures. Our results show that using gestures in interactive handwriting applications is both advantageous and convenient when gestures are simple but context-aware. Taken together, this work suggests that text-editing interfaces not only can be easily augmented with simple gestures, but also may substantially improve user productivity. Luis A. Leiva, Vicente Alabau, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
Interact. Comput. | 3 |
| 2014 | Ground-Truth Production in the Transcriptorium ProjectabstractTran Scriptorium is a 3-years project that aims to develop innovative, cost-effective solutions for the indexing, search and full transcription of historical handwritten document images, using Handwritten Text Recognition (HTR) technology. The production of ground-truth (GT) of a dataset of handwritten document images is among the first tasks. We address novel approaches for the faster production of this GT based on crowd-sourcing and on prior-knowledge methods. We also address here a novel low-cost semi-supervised procedure for obtaining pairs of correct line-level aligned detected/extracted text line images and text line transcripts, specially suitable for training models of the HTR technology employed in Tran Scriptorium. Basilios Gatos, Georgios Louloudis, Tim Causer, Kris Grint, Verónica Romero 0001, Joan-Andreu Sánchez, Alejandro H. Toselli, Enrique Vidal 0001 |
Document Analysis Systems | 5 |
| 2014 | ICFHR2014 Competition on Handwritten Text Recognition on Transcriptorium Datasets (HTRtS)abstractA contest on Handwritten Text Recognition organised in the context of the ICFHR 2014 conference is described. Two tracks with increased freedom on the use of training data were proposed and three research groups participated in these two tracks. The handwritten images for this contest were drawn from an English data set which is currently being considered in the Tran scriptorium project. The goal of this project is to develop innovative, efficient and cost-effective solutions for the transcription of historical handwritten document images, focusing on four languages: English, Spanish, German and Dutch. For the English language, the so-called "Bentham collection" is being considered in Tran scriptorium. It encompasses a large set of manuscripts written by the renowned English philosopher and reformer Jeremy Bentham (1748-1832). A small subset of this collection has been chosen for the present HTR competition. The selected subset has been written by several hands (Bentham himself and his secretaries) and entails significant variabilities and difficulties regarding the quality of text images and writing styles. Training and test data were provided in the form of carefully segmented line images, along with the corresponding transcripts. The three participants achieved very good results, with transcription word error rates ranging from 15.0% down to 8.6%. Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 2 |
| 2014 | An iterative multimodal framework for the transcription of handwritten historical documents
Vicente Alabau, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Antonio L. Lagarda |
Pattern Recognit. Lett. | 3 |
| 2013 | Interactive Off-Line Handwritten Text Transcription Using On-Line Handwritten Text as FeedbackabstractHandwritten Text Recognition is a problem that has gained attention in the last years mainly due to the interest in the transcription of historical documents. However, the automatic transcription is ineffectual in unconstrained handwritten documents. Thus, human intervention is typically needed to correct the results. Given that a post-editing approach is inefficient and uncomfortable, multimodal interactive approaches have begun to emerge in the last years. In this scheme, the user interacts with the system by means of an e-pen. This multimodal feedback, on the one hand, allows to improve the accuracy of the system and, on the other hand, increases user acceptability. In this work, we present a new approach on interaction based on character sequences. Here we present developments that allow taking advantage of interaction-derived context to significantly improve feedback decoding accuracy. Empirical tests suggest that, despite the loss of the deterministic accuracy of traditional peripherals, this approach can save significant amounts of user effort with respect to non-interactive post-editing correction. Daniel Martín-Albo, Verónica Romero 0001, Enrique Vidal 0001 |
ICDAR | 2 |
| 2013 | Category-Based Language Models for Handwriting Recognition of Marriage License BooksabstractHandwritten marriage licenses books have been used for centuries by ecclesiastical institutions to register marriages. These documents have interesting information, useful for demography studies, organized in a list of individual marriage license records, such as an accounting book. The information in these books is usually collected by expert demographers that devote a lot of time to transcribe them. Despite the structure of the text, the automatic transcription and semantic information extraction of these documents is quite difficult due to the distinct and evolutionary vocabulary, which is composed mainly of proper names that change along the time. In this paper, we have defined some categories taking into account the semantic information included in the licenses. Then a category-based language model has been generated and integrated into the handwritten text recognition system. We study how the use of these categories can benefit not only the handwriting recognition step, but also the posterior semantic information extraction and knowledge discovery. Verónica Romero 0001, Joan-Andreu Sánchez |
ICDAR | 1 |
| 2013 | Human Evaluation of the Transcription Process of a Marriage License BookabstractHandwriting Text Recognition (HTR) of historical documents is a very important research field of Document Image Analysis. Currently, the most well-accepted technology for off-line HTR is based on holistic, segmentation-free techniques that do not need any kind of character or word segmentation. This HTR technology is based in stochastic models that are trained with annotated data. The performance of this technology is still far from being perfect and therefore the user intervention is necessary to obtain perfect transcripts. The user intervention can be carried out in a post-editing process, in which the user corrects the errors produced by an automatic HTR system. Interactive techniques have been proposed in the past few years to obtain the correct transcript as an alternative to post-editing the transcripts. In these interactive approaches, the user and the system work interactively in tight mutual collaboration to obtain the perfect transcript of the data. In this interactive scenario, the feedback provided by the user is used to improve interactively the system output. In the post-editing scenario and in the interactive scenario, the transcribed material can be used for retraining the models as the data is processed. In this research we carried out a study with a real transcriber about how the performance of an HTR system improved with respect to the amount of training data, and how the human efficiency improved during the transcription process in both transcription scenarios. Verónica Romero 0001, Joan-Andreu Sánchez |
ICDAR | 1 |
| 2013 | The ESPOSALLES database: An ancient marriage license corpus for off-line handwriting recognition
Verónica Romero 0001, Alicia Fornés, Joan-Andreu Sánchez, Alejandro H. Toselli, Volkmar Frinken, Enrique Vidal 0001, Josep Lladós 0001 |
Pattern Recognit. | 1 |
| 2012 | Multimodal Computer-Assisted transcription of Text Images at Character-Level InteractionabstractCurrently, automatic handwriting recognition systems are ineffectual in unconstrained handwriting documents. Therefore, to obtain perfect transcriptions, heavy human intervention is required to validate and correct the results of such systems. Given that this post-editing process is inefficient and uncomfortable, a multimodal interactive approach has been proposed in previous works, which aims at obtaining correct transcriptions with the minimum human effort. In this approach, the user interacts with the system by means of an e-pen and/or more traditional methods such as keyboard or mouse. This user's feedback allows to improve system accuracy and multimodality increases system ergonomics and user acceptability. Until now, multimodal interaction has been considered only at whole-word level. In this work, multimodal interaction at character-level is studied, that may lead to more effective interactivity, since it is faster and easier to write only one character rather than a whole word. Here we study this kind of fine-grained multimodal interaction and present developments that allow taking advantage of interaction-derived context to significantly improve feedback decoding accuracy. Empirical tests on three cursive handwritten tasks suggest that, despite losing the deterministic accuracy of traditional peripherals, this approach can save significant amounts of user effort with respect to fully manual transcription as well as to noninteractive post-editing correction. Daniel Martín-Albo, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2011 | Evaluating an Interactive-Predictive Paradigm on Handwriting Transcription: A Case Study and Lessons LearnedabstractTranscribing handwritten text is a laborious task which currently is carried out manually. As the accuracy of automatic handwritten text recognizers improves, post-editing the output of these recognizers could be foreseen as a possible alternative. Alas, the state-of-the-art technology is not suitable to perform this kind of work, since current approaches are not accurate enough and the process is usually both inefficient and uncomfortable for the user. As alternative, an interactive-predictive paradigm has gained recently an increasing popularity, mainly due to promising empirical results that estimate considerable reductions of user effort. In order to assess whether these empirical results can lead indeed to actual benefits, we developed a working prototype and conducted a field study remotely. Thirteen regular computer users tested two different transcription engines through the above-mentioned prototype. We observed that the interactive-predictive version allowed to transcribe better (less errors and fewer iterations to achieve a high-quality output) in comparison to the manual engine. Additionally, participants ranked higher such an interactive-predictive system in a usability questionnaire. We describe the evaluation methodology and discuss our preliminary results. While acknowledging the known limitations of our experimentation, we conclude that the interactive-predictive paradigm is an efficient approach for transcribing handwritten text. Luis A. Leiva, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
COMPSAC | 2 |
| 2011 | Handwritten Text Recognition for Marriage Register BooksabstractMarriage register books are documents that were used for centuries by ecclesiastical institutions to register marriages. Most of these books were handwritten. These documents have interesting information, useful for demography studies. The information in these books is usually collected by expert demographers that devote a lot of time to transcribe them. The automatic transcription of these documents by using Handwritten Text Recognition techniques is difficult since the vocabulary is large, given that it is composed mainly of proper names. In this work, interactive Handwritten Text Recognition techniques were studied for the assisted transcription of these documents. Verónica Romero 0001, Joan-Andreu Sánchez, Enrique Vidal 0001 |
ICDAR | 1 |
| 2011 | Study of different interactive editing operations in an assisted transcription systemabstractTo date, automatic handwriting recognition systems are far from being perfect. Therefore, once the full recognition process of a handwritten text image has finished, heavy human intervention is required in order to correct the results of such systems. As an alternative, an interactive system has been proposed in previous works. This alternative follows an Interactive Predictive paradigm and the results show that significant amounts of human effort can be saved. So far only word substitutions and pointer actions have been considered in this interactive system. In this work, we study different interactive editing operations that can allow for more effective, ergonomic and friendly interfaces. Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICMI | 1 |
| 2011 | A Multimodal Approach to Dictation of Handwritten Historical Documents
Vicente Alabau, Verónica Romero 0001, Antonio L. Lagarda, Carlos D. Martínez-Hinarejos |
INTERSPEECH | 2 |
| 2010 | Interactive layout analysis and transcription systems for historic handwritten documentsabstractThe amount of digitized legacy documents has been rising dramatically over the last years due mainly to the increasing number of on-line digital libraries publishing this kind of documents, waiting to be classified and finally transcribed into a textual electronic format (such as ASCII or PDF). Nevertheless, most of the available fully-automatic applications addressing this task are far from being perfect and heavy and inefficient human intervention is often required to check and correct the results of such systems. In contrast, multimodal interactive-predictive approaches may allow the users to participate in the process helping the system to improve the overall performance. With this in mind, two sets of recent advances are introduced in this work: a novel interactive method for text block detection and two multimodal interactive handwritten text transcription systems which use active learning and interactive-predictive technologies in the recognition process. Oriol Ramos Terrades, Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001, Alfons Juan-Císcar |
ACM Symposium on Document Engineering | 4 |
| 2010 | Character-Level Interaction in Computer-Assisted Transcription of Text ImagesabstractTo date, automatic handwriting recognition systems are far from being perfect and heavy human intervention is often required to check and correct the results of such systems. As an alternative, an interactive framework that integrates the human knowledge into the transcription process has been presented in previous works. This new approach follows an Interactive Predictive paradigm and our results show that significant amounts of human effort can be saved. Until now only whole-word interactions with this system have been considered. In this work, character-level keystroke interactions, that can allow for a more ergonomic and friendly interfaces, are proposed. Empirical results show that this allows for further improvements in user productivity. Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 1 |
| 2010 | Computer Assisted Transcription of Text Images: Results on the GERMANA Corpus and Analysis of Improvements Needed for Practical UseabstractWe present a study of the application of Computer Assisted Transcription of Text Images (CATTI) to a task which is much closer to real applications than other tasks previously studied. The new task consists in the transcription of a new publicly available historic handwritten document, called GERMANA. A detailed analysis of the main factors influencing the system performance are exposed and some strategies to circumvent them are proposed. Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICPR | 1 |
| 2010 | Multimodal interactive transcription of text images
Alejandro H. Toselli, Verónica Romero 0001, Moisés Pastor, Enrique Vidal 0001 |
Pattern Recognit. | 2 |
| 2009 | Using Mouse Feedback in Computer Assisted Transcription of Handwritten Text ImagesabstractTo date, automatic handwriting recognition systems are far from being perfect and heavy human intervention is often required to check and correct the results of such systems. In order to achieve correct transcriptions, human knowledge can be integrated into the transcription process, following an Interactive Predictive paradigm. We have recently proposed Mouse Actions as a significant feedback information source for the underlying interactive system to improve the productivity of the human transcriptor. In this paper we review this way to interact with the system and report comparative results using the publicly available IAMDB dataset. Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 1 |
| 2009 | A multimodal predictive-interactive application for computer assisted transcription and translationabstractTraditionally, Natural Language Processing (NLP) technologies have mainly focused on full automation. However, full automation often proves unnatural in many applications, where technology is expected to assist rather than replace the human agents. Vicente Alabau, Daniel Ortiz-Martínez, Verónica Romero 0001, Jorge Ocampo |
ICMI | 3 |
| 2009 | Interactive multimodal transcription of text images using a web-based demo systemabstractThis document introduces a web based demo of an interactive framework for transcription of handwritten text, where the user feedback is provided by means of pen strokes on a touchscreen. Here, the automatic handwriting text recognition system and the user both cooperate to generate the final transcription. Verónica Romero 0001, Luis A. Leiva, Alejandro H. Toselli, Enrique Vidal 0001 |
IUI | 1 |
| 2007 | Computer Assisted Transcription of Handwritten Text ImagesabstractTo date, automatic handwriting recognition systems are far from being perfect and often they need a post editing where a human intervention is required to check and correct the results of such systems. We propose to have a new interactive, on-line framework which, rather than full automation, aims at assisting the human in the proper recognition- transcription process; that is, facilitate and speed up their transcription task of handwritten texts. This framework combines the efficiency of automatic handwriting recognition systems with the accuracy of the human transcriptor. The best result is a cost-effective perfect transcription of the handwriting text images. Alejandro H. Toselli, Verónica Romero 0001, Luis Rodríguez, Enrique Vidal 0001 |
ICDAR | 2 |