VLDB 2026 Research / reviewers in the wild / expert
Enrique Vidal 0001
dblp:39/3758 · also Enrique Vidal-Ruiz
· DBLP profile ↗
176ranked-venue papers
19as first author
20since 2021 · last 2026
0000-0003-4579-5196ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 123 · 14 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 70 · 7 first-authorDatabases, data management, data science and information retrieval · 36 · 3 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 7Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Simple handwritten text recognition techniques for highly accurate writer identificationabstractAbstract The theater of the Spanish Early Modern period encompasses thousands of textual works and hundreds of playwrights and is one of the greatest examples of Spanish literature. These works were usually copied and altered, changing sometimes the meaning of the original works written by the authors. Therefore, identifying autograph testimonies written directly by the playwrights themselves are of particular importance. In this paper, we address an approach for writer identification based on Deep Convolutional-Recurrent Neural Networks and n -gram language models. The identification task is posed as a classification problem, introducing a probabilistic framework that goes beyond the plain transcription of handwritten text. Experiments are conducted to validate our proposal to distinguish between Lope de Vega ’s manuscripts and other non-Lope hands. The good results achieved will ultimately allow to provide modern researchers with a useful tool for cultural heritage recovery. Alejandro H. Toselli, Álvaro Cuéllar, Sònia Boadas, Enrique Vidal 0001, Joan-Andreu Sánchez |
Pattern Anal. Appl. | 4 |
| 2025 | PARDES: Automatic Generation of Descriptive Terms for Logical Units in Historical Handwritten Collections
Josepa Raventós-Pajares, Joan-Andreu Sánchez, Enrique Vidal 0001 |
IEEE Big Data | 3 |
| 2025 | The PARES Database: Information Extraction over Historical Parish RecordsabstractAbstract Historical census records convey information that is key to perform genealogical research and demographic studies. Given the large number of documents of this type that exist, it is crucial to research methods that allow the automatic extraction of information from this type of document. In this work, we present a new corpus of this kind, comprising 535 historical census tables from French archives. Alongside this dataset, we have assessed three different baseline methods for information extraction. The first two methods employ a traditional sequential approach, where table rows are detected before extracting information. The third baseline uses an end-to-end model that directly extracts information from the table images without prior row detection. Our results demonstrate the effectiveness of all three baselines in tackling the information extraction task. José Andrés, Casey Wall, Solène Tarride, Mickaël Coustaty, Alejandro H. Toselli, Enrique Vidal 0001 |
Int. J. Document Anal. Recognit. | 6 |
| 2024 | Mining and Analyzing Statistical Information from Untranscribed Form Images
José Andrés, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR (5) | 3 |
| 2024 | Zipf Curves and Basic Text Analytics from Untranscribed Manuscript Images
Enrique Vidal 0001, Alejandro H. Toselli |
ICDAR (3) | 1 |
| 2024 | Ground-truth generation through crowdsourcing with probabilistic indexesabstractAbstract Automatic transcription of large series of historical handwritten documents generally aims at allowing to search for textual information in these documents. However, automatic transcripts often lack the level of accuracy needed for reliable text indexing and search purposes. Probabilistic Indexing (PrIx) offers a unique alternative to raw transcripts. Since it needs training data to achieve good search performance, PrIx-based crowdsourcing techniques are introduced in this paper to gather the required data. In the proposed approach, PrIx confidence measures are used to drive a correction process in which users can amend errors and possibly add missing text. In a further step, corrected data are used to retrain the PrIx models. Results on five large series are reported which show consistent improvements after retraining. However, it can be argued whether the overall costs of the crowdsourcing operation pay off for the improvements, or perhaps it would have been more cost-effective to just start with a larger and cleaner amount of professionally produced training transcripts. Joan-Andreu Sánchez, Enrique Vidal 0001, Vicente Bosch, Lorenzo Quirós |
Neural Comput. Appl. | 2 |
| 2024 | Segmenting large historical notarial manuscripts into multi-page deedsabstractAbstract Archives around the world hold vast digitized series of historical manuscript books or “bundles” containing, among others, notarial records also known as “deeds” or “acts”. One of the first steps to provide metadata which describe the contents of those bundles is to segment them into their individual deeds. Even if deeds are often page-aligned, as in the bundles considered in the present work, this is a time-consuming task, often prohibitive given the huge scale of the manuscript series involved. Unlike traditional Layout Analysis methods for page-level segmentation, our approach goes beyond the realm of a single-page image, providing consistent deed detection results on full bundles. This is achieved in two tightly integrated steps: first, we estimate the class-posterior at the page level for the “initial”, “middle”, and “final” classes; then we “decode” these posteriors applying a series of sequentiality consistency constraints to obtain a consistent book segmentation. Experiments are presented for four large historical manuscripts, varying the number of “deeds” used for training. Two metrics are introduced to assess the quality of book segmentation, one of them taking into account the loss of information entailed by segmentation errors. The problem formalization, the metrics and the empirical work significantly extend our previous works on this topic. José Ramón Prieto, David Camilo Becerra Romero, Alejandro H. Toselli, Carlos Alonso, Enrique Vidal 0001 |
Pattern Anal. Appl. | 5 |
| 2023 | Search for Hyphenated Words in Probabilistic Indices: A Machine Learning Approach
José Andrés, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR (1) | 3 |
| 2023 | Lexicon-based probabilistic indexing of handwritten text imagesabstractAbstract Keyword Spotting (KWS) is here considered as a basic technology for Probabilistic Indexing (PrIx) of large collections of handwritten text images to allow fast textual access to the contents of these collections. Under this perspective, a probabilistic framework for lexicon-based KWS in text images is presented. The presentation aims at providing formal insights which help understanding classical statements of KWS (from which PrIx borrows fundamental concepts), as well as the relative challenges entailed by these statements. The development of the proposed framework makes it clear that word recognition or classification implicitly or explicitly underlies any formulation of KWS. Moreover, it suggests that the same statistical models and training methods successfully used for handwriting text recognition can advantageously be used also for PrIx, even though PrIx does not generally require or rely on any kind of previously produced image transcripts. Experiments carried out using these approaches support the consistency and the general interest of the proposed framework. Results on three datasets traditionally used for KWS benchmarking are significantly better than those previously published for these datasets. In addition, good results are also reported on two new, larger handwritten text image datasets (B entham and P lantas ), showing the great potential of the methods proposed in this paper for indexing and textual search in large collections of untranscribed handwritten documents. Specifically, we achieved the following Average Precision values: IAMDB: 0.89, G eorge W ashington : 0.91, P arzival : 0.95, B entham : 0.91 and P lantas : 0.92. Enrique Vidal 0001, Alejandro H. Toselli, Joan Puigcerver |
Neural Comput. Appl. | 1 |
| 2023 | A proxy learning curve for the Bayes classifierabstractIn this paper, a theoretical learning curve is derived for the multi-class Bayes classifier. This curve fits general multivariate parametric models of the class-conditional probability density. The derivation uses a proxy approach based on analyzing the convergence of a statistic which is proportional to the posterior probability of the true class. By doing so, the curve depends only on the training set size and on the dimension of the feature vector; it does not depend on the model parameters. Essentially, the learning curve provides an estimate of the reduction in the excess of the probability of error that can be obtained by increasing the training set size. This makes it attractive in order to deal with the practical problems of defining appropriate training set sizes. Addisson Salazar, Luis Vergara, Enrique Vidal 0001 |
Pattern Recognit. | 3 |
| 2023 | End-to-End page-Level assessment of handwritten text recognitionabstractThe evaluation of Handwritten Text Recognition (HTR) systems has traditionally used metrics based on the edit distance between HTR and ground truth (GT) transcripts, at both the character and word levels. This is very adequate when the experimental protocol assumes that both GT and HTR text lines are the same, which allows edit distances to be independently computed to each given line. Driven by recent advances in pattern recognition, HTR systems increasingly face the end-to-end page-level transcription of a document, where the precision of locating the different text lines and their corresponding reading order (RO) play a key role. In such a case, the standard metrics do not take into account the inconsistencies that might appear. In this paper, the problem of evaluating HTR systems at the page level is introduced in detail. We analyse the convenience of using a two-fold evaluation, where the transcription accuracy and the RO goodness are considered separately. Different alternatives are proposed, analysed and empirically compared both through partially simulated and through real, full end-to-end experiments. Results support the validity of the proposed two-fold evaluation approach. An important conclusion is that such an evaluation can be adequately achieved by just two simple and well-known metrics: the Word Error Rate (WER), that takes transcription sequentiality into account, and the here re-formulated Bag of Words Word Error Rate (bWER), that ignores order. While the latter directly and very accurately assess intrinsic word recognition errors, the difference between both metrics (ΔWER) gracefully correlates with the Normalised Spearman’s Foot Rule Distance (NSFD), a metric which explicitly measures RO errors associated with layout analysis flaws. To arrive to these conclusions, we have introduced another metric called Hungarian Word Word Rate (hWER), based on a here proposed regularised version of the Hungarian Algorithm. This metric is shown to be always almost identical to bWER and both bWER and hWER are also almost identical to WER whenever HTR transcripts and GT references are guarantee to be in the same RO. Enrique Vidal 0001, Alejandro H. Toselli, Antonio Ríos-Vila, Jorge Calvo-Zaragoza |
Pattern Recognit. | 1 |
| 2023 | Processing a large collection of historical tabular imagesabstractProcessing automatically historical document images to allow the search of textual information requires the preparation of ground-truth data for training and evaluation. This process is an expensive and arduous task, especially when the historical document images contain specialized vocabulary and/or tabular information. In the latter case, relevant decisions have to be taken to annotate the tabular parts. This paper presents a complex collection of historical document images and the resulting database, which is called HisClima. In this database, half of the images are in tabular format and half as running text. Both types of images contain pre-printed and handwritten text. The textual information is plenty of abbreviations and specific vocabulary related to weather conditions and old ships. This database can be used to research technologies related to historical document image processing and analysis, both for tabular and running text recognition. Baseline results are presented for Document Layout Analysis, Text Recognition, and Probabilistic Indexing. Although these results are good, there is still room for improvement and some indications are provided in this direction. Emilio Granell, Verónica Romero 0001, José Ramón Prieto, José Andrés, Lorenzo Quirós, Joan-Andreu Sánchez, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 7 |
| 2023 | Information extraction in handwritten historical logbooksabstractDocument Image Understanding is a demanding Pattern Recognition problem that requires complex recognition models. This problem is even more difficult for document images with complicated layouts like tables, where the reading order is often intrinsically ambiguous, and consequently, the context is generally ambiguous as well. In this paper, we compare two machine learning approaches for extracting information in pre-printed historical tables with handwritten information. We analyze the performance of each approach at each step of the extraction process over different corpora, up to a realistic scenario where documents with different table layouts written by different hands are used. The results are good in general and show that a model based on Multilayer Perceptrons yields better results on more homogeneous documents, while another model based on Graph Neural Networks generalizes better on heterogeneous corpora. José Ramón Prieto, José Andrés, Emilio Granell, Joan-Andreu Sánchez, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 5 |
| 2023 | Open set classification of untranscribed handwritten text image documentsabstractContent-based classification of manuscripts is an important task that is generally carried out by expert archivists. Nevertheless, many historical manuscript collections are so vast that in most cases this task is hardly feasible, even for large, well staffed archives. Nowadays, manuscripts are generally preserved in the form of sets of digital images. Therefore, the technical problem we are interested in is automatic classification of “‘image documents”, each consisting of a set of untranscribed handwritten text images, by the textual contents of the images. The traditional Pattern Recognition classification paradigm does provide the basic tools to deal with this problem. However, in practice, the set of relevant classes of a large documental series is seldom known in advance. Therefore, a classifier trained with a predefined set of classes will systematically fail when new image documents arrive which do not belong to any of the classes assumed in training. Here we adopt the “Open Set Classification” framework to extend and consolidate our previous work on image document classification in order to adequately handle new documents from unknown classes. The proposed approaches are based on a relatively novel technology for text image representation known as “probabilistic indexing”, which proves very effective to characterise the intrinsic word-level uncertainty exhibited by historical handwritten text images. We assess the performance of this approach on a moderately sized but representative dataset extracted from a huge series of complex notarial manuscripts from the Spanish Archivo Histórico Provincial de Cádiz, with good results. José Ramón Prieto, Juan José Flores Arellano, Enrique Vidal 0001, Alejandro H. Toselli |
Pattern Recognit. Lett. | 3 |
| 2022 | Information Extraction from Handwritten Tables in Historical Documents
José Andrés, José Ramón Prieto, Emilio Granell, Verónica Romero 0001, Joan-Andreu Sánchez, Enrique Vidal 0001 |
DAS | 6 |
| 2022 | Approximate Search for Keywords in Handwritten Text Images
José Andrés, Alejandro H. Toselli, Enrique Vidal 0001 |
DAS | 3 |
| 2022 | Effective Crowdsourcing in the EDT Project with Probabilistic Indexes
Joan-Andreu Sánchez, Enrique Vidal 0001, Vicente Bosch |
DAS | 2 |
| 2022 | Reading order detection on handwritten documentsabstractAbstract Recent advances in Handwritten Text Recognition and Document Layout Analysis have made it possible to convert digital images of manuscripts into electronic text. However, providing this text with the correct structure and context is still an open problem that needs to be solved to actually enable extracting the relevant information conveyed by the text. The most important structure needed for a set of text elements is their reading order. Most of the studies on the reading order problem are rule-based approaches and focus on printed documents. Much less attention has been paid so far to handwritten text documents, where the problem becomes particularly important—and challenging. In this work, we propose a new approach to automatically determine the reading order of text regions and lines in handwritten text documents. The task is approached as a sorting problem where the order-relation operator is automatically learned from examples. We experimentally demonstrate the effectiveness of our method on three different datasets at different hierarchical levels. Lorenzo Quirós, Enrique Vidal 0001 |
Neural Comput. Appl. | 2 |
| 2021 | Probabilistic Indexing and Search for Hyphenated Words
Enrique Vidal 0001, Alejandro H. Toselli |
ICDAR (2) | 1 |
| 2021 | Improved Graph Methods for Table Layout Understanding
José Ramón Prieto, Enrique Vidal 0001 |
ICDAR (2) | 2 |
| 2020 | A comparison of sequential and combined approaches for named entity recognition in a corpus of handwritten medieval chartersabstractThis paper introduces a new corpus of multilingual medieval handwritten charter images, annotated with full transcription and named entities. The corpus is used to compare two approaches for named entity recognition in historical document images in several languages: on the one hand, a sequential approach, more commonly used, that sequentially applies handwritten text recognition (HTR) and named entity recognition (NER), on the other hand, a combined approach that simultaneously transcribes the image text line and extracts the entities. Experiments conducted on the charter corpus in Latin, early new high German and old Czech for name, date and location recognition demonstrate a superior performance of the combined approach. Emanuela Boros, Verónica Romero 0001, Martin Maarand, Katerina Zenklová, Jitka Krecková, Enrique Vidal 0001, Dominique Stutzmann, Christopher Kermorvant |
ICFHR | 6 |
| 2020 | Text Content Based Layout AnalysisabstractState-of-the-art Document Layout Analysis methods rely on graphical appearance features in order to detect and classify the different layout regions present in a scanned text image. In many cases, however, performing this task using only graphical information is problematic or impossible. Only by actually reading some text in the boundaries of the problematic regions it becomes possible to reliably detect and separate these regions. In these situations, textual, content-based features would be required, but since transcription is usually performed after layout analysis, a vicious circle arises. In this work, we circumvent this deadlock by making use of the recently introduced concept of Probabilistic Index Map. We use the word relevance probabilities provided by this map to calculate relevant text content based features at the pixel level. We assess the impact of these new features on a historical document complex paragraph classification task. The experiments are performed using both a classical Hidden Markov Model approach and Deep Neural Networks. The obtained results are encouraging and showcase the positive impact text content based features will have on the Document Layout Analysis research field. José Ramón Prieto, Vicente Bosch, Enrique Vidal 0001, Dominique Stutzmann, Sébastien Hamel |
ICFHR | 3 |
| 2020 | The Carabela Project and Manuscript Collection: Large-Scale Probabilistic Indexing and Content-based ClassificationabstractThe main aim of the Carabela project was to develop and apply techniques that allow textual searching on massive Spanish collections of 15th-19th century manuscripts. The project focused on a relatively small subset of 125 000 images of collections of interest to underwater archaeology. For this type of manuscripts, state-of-the-art automatic transcription techniques, generally fail to achieve usable transcription accuracy. Therefore, rather than insisting in actual transcription, methodologies for probabilistic indexing of handwritten text images have been adopted. This has allowed us to effectively cope with the intrinsically high degree of uncertainty of the text contained in most historical manuscripts, leading to highly effective systems for textual search and retrieval. Carabela has gone one step further by developing new techniques to classify probabilistically indexed, but otherwise untranscribed, text images according to their textual content. These techniques have been successfully used to automatically classify Carabela bundels (each containing hundreds or thousands of pages) according to their “level of risk” of public exposure, in order to control their access and avoid as much as possible the plundering of Spanish underwater heritage. Enrique Vidal 0001, Verónica Romero 0001, Alejandro H. Toselli, Joan-Andreu Sánchez, Vicente Bosch, Lorenzo Quirós, José-Miguel Benedí, José Ramón Prieto, Moisés Pastor, Francisco Casacuberta, Carlos Alonso, Carmen García, Lourdes Márquez, Carmen Orcero |
ICFHR | 1 |
| 2020 | Textual-Content-Based Classification of Bundles of Untranscribed Manuscript ImagesabstractContent-based classification of manuscripts is an important task that is generally performed in archives and libraries by experts with a wealth of knowledge on the manuscript's contents. Unfortunately, many manuscript collections are so vast that it is not feasible to rely solely on experts to perform this task. Current approaches for textual-content-based manuscript classification generally require the handwritten images to be first transcribed into text - but achieving sufficiently accurate transcripts are generally unfeasible for large sets of historical manuscripts. We propose a new approach to perform automatically this classification task which does not rely on any explicit image transcripts. It is based on “probabilistic indexing”, a relatively novel technology which allows to effectively represent the intrinsic word-level uncertainty generally exhibited by handwritten text images. We assess the performance of this approach on a large collection of complex manuscripts from the Spanish Archivo General de Indias, with promising results. To the best of our knowledge, this is the first published work proposing, developing and assessing a successful approach for content-based classification of untranscribed manuscript images. José Ramón Prieto, Vicente Bosch, Enrique Vidal 0001, Carlos Alonso, M. Carmen Orcero, Lourdes Márquez |
ICPR | 3 |
| 2020 | Writer Identification Using Deep Neural Networks: Impact of Patch Size and Number of PatchesabstractTraditional approaches for the recognition or identification of the writer of a handwritten text image used to relay on heuristic knowledge about the shape and other features of the strokes of previously segmented characters. However, recent works have done significantly advances on the state of the art thanks to the use of various types of deep neural networks. In most of all of these works, text images are decomposed into patches, which are processed by the networks without any previous character or word segmentation. In this paper, we study how the way images are decomposed into patches impact recognition accuracy, using three publicly available datasets. The study also includes a simpler architecture where no patches are used at all - a single deep neural network inputs a whole text image and directly provides a writer recognition hypothesis. Results show that bigger patches generally lead to improved accuracy, achieving in one of the datasets a significant improvement over the best results reported so far. Akshay Punjabi, José Ramón Prieto, Enrique Vidal 0001 |
ICPR | 3 |
| 2020 | Learning to Sort Handwritten Text Lines in Reading Order through Estimated Binary Order RelationsabstractRecent advances in Handwritten Text Recognition and Document Layout Analysis make it possible to extract information from digitized documents and make them accessible beyond the archive shelves. But the reading order of the elements in those documents still is an open problem that has to be solved in order to provide that information with the correct structure. Most of the studies on the reading order task are rule-base approaches that focus on printed documents, while less attention has been paid to handwritten text documents. In this work we propose a new approach to automatically determine the reading order of text lines in handwritten text documents. The task is approached as a sorting problem where the order-relation operator is learned directly from examples. We demonstrate the effectiveness of our method on three different datasets. Lorenzo Quirós, Enrique Vidal 0001 |
ICPR | 2 |
| 2020 | Pattern recognition techniques for provenance classification of archaeological ceramics using ultrasounds
Addisson Salazar, Gonzalo Safont, Luis Vergara, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 4 |
| 2019 | Music Symbol Sequence Indexing in Medieval Plainchant ManuscriptsabstractHuge amounts of musical manuscripts are preserved in cathedrals, abbeys, and archives. However, without reliable transcripts, their contents are inaccessible. Manual transcription is unaffordable for large collections, and current automatic technologies-such as Optical Music Recognition or Handwritten Music Recognition-do not provide sufficient accuracy for a fully-automatic scenario. In many cases, perfect transcripts are not really needed, given that content-based search with some degree of reliability would already be extremely useful. Spotting just single music symbols is rather useless (most of the symbols generally appear in all pages); instead, helpful search targets are melodic patterns, which typically correspond to music symbol sequences. We explore approaches for accurate retrieval of melodic patterns, represented by music symbol sequences, from collections of Medieval plainchant manuscripts. Our statistical framework, based on the use of convolutional recurrent neural networks and probabilistic indices, is shown to be useful for retrieving music patterns which appear frequently in this untranscribed images, yielding an Average Precision of 86 %. Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001, Joan-Andreu Sánchez |
ICDAR | 3 |
| 2019 | Making Two Vast Historical Manuscript Collections Searchable and Extracting Meaningful Textual Features Through Large-Scale Probabilistic IndexingabstractTextual access to large collections of digitized images remains unfeasible because usually they lack transcripts. Transcribing such collections is in turn typically unattainable in terms of costs. However, the use of probabilistic indices can facilitate textual accessing with only moderate demands of resources. Besides allowing effortless information retrieval, it will be shown that probabilistic indices can also be used to estimate textual features of the indexed but otherwise untranscribed collections, such as running words and Zipf's curves. Complete probabilistic indices have been recently produced for two iconic large collections: "Bentham" (90K images) and "Spanish Golden Age Theater" (40K images). To show the repercussion of making these collections searchable, we provide accessing statistics gathered through their corresponding search interfaces. To the best of our knowledge this is the first publication of large collections of untranscribed manuscripts which are now publicly accessible for effective and efficient textual access. Alejandro H. Toselli, Verónica Romero 0001, Joan-Andreu Sánchez, Enrique Vidal 0001 |
ICDAR | 4 |
| 2019 | Hybrid hidden Markov models and artificial neural networks for handwritten music recognition in mensural notation
Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001 |
Pattern Anal. Appl. | 3 |
| 2019 | Probabilistic multi-word spotting in handwritten text images
Alejandro H. Toselli, Enrique Vidal 0001, Joan Puigcerver, Ernesto Noya-García |
Pattern Anal. Appl. | 2 |
| 2019 | A set of benchmarks for Handwritten Text Recognition on historical documents
Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Mauricio Villegas, Enrique Vidal 0001 |
Pattern Recognit. | 5 |
| 2019 | Handwritten Music Recognition for Mensural notation with convolutional recurrent neural networks
Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 3 |
| 2018 | Automatic Alignment of Handwritten Images and Transcripts for Training Handwritten Text Recognition SystemsabstractState-of-the-art Handwritten Text Recognition techniques are based on statistical models such as hidden Markov models or recurrent neural networks for optical modeling of characters and N-grams for language modeling. These models are trained using well known, learning techniques: Expectation-Maximization, backpropagation, etc. Therefore, training data is needed to build these models. In the case of the optical models the training data consist of text line images with their corresponding transcripts. When the transcript of a handwritten document is available, putting in correspondence automatically the physical lines in the images with the lines of the transcripts is not an easy task. We present a method for automatically aligning handwritten text images and their respective transcripts. The approach automatically segments the images into lines and then recognizes them. An alignment confidence is obtained using the Levenshtein distance between the recognition results and the transcripts. The most confident lines are then used for training. Experiments carried out using a historical document present encouraging results. Verónica Romero 0001, Alejandro H. Toselli, Vicente Bosch, Joan-Andreu Sánchez, Enrique Vidal 0001 |
DAS | 5 |
| 2018 | Probabilistic Music-Symbol Spotting in Handwritten ScoresabstractContent-based search on musical manuscripts is usually performed assuming that there are accurate transcripts of the sources in a symbolic, structured format. Given that current systems for Handwritten Music Recognition are far from offering guarantees about their accuracy, this traditional approach does not represent a scalable scenario. In this work we propose a probabilistic framework for Music-Symbol Spotting (MSS), that allows for content-based music search directly over the images of the manuscripts. By means of statistical recognition systems, a probabilistic index is built upon which the search can be carried out efficiently. Our experiments over a dataset of an Early handwritten music manuscript in Mensural notation demonstrates that this MSS framework can be presented as a promising alternative to the traditional approach for content-based music search. Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 3 |
| 2018 | Text Line Extraction Based on Distance Map Features and Dynamic ProgrammingabstractText Line Segmentation is a basic document layout task that consists in detecting and extracting the text lines present in a document page image. Although considered a basic task, generally, it is a necessary step for Handwritten Text Recognition (HTR) higher level tasks. Most state of the art automatic text recognition, text-to-line image alignment and key word spotting systems require it due to their need for isolated text line images as input. Traditionally most Text Line Segmentation approaches cover both detection and extraction sub steps. However, the community has recently shifted its focus to tackle independently the baseline detection in document images. This shift generates the need for extraction methods that use these detected baselines as input. In this paper, a binarization free dynamic programming approach that generates an equidistant text line extraction polygon is presented. The approach performs this calculation, based on the information provided by priorly detected text baselines and automatically generated foreground pixels distance maps. We evaluate our approach both in a synthetic competition corpus and in a challenging real handwritten text recognition task corpus. We evaluate it not only at the graphical error level but also the impact it produces on an HTR task trained with the line images it yields. We compare our solution with other solutions ranging from the actual human reviewed ground-truth polygons to simpler automatic generated rectangle areas. Vicente Bosch, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 4 |
| 2018 | Probabilistic Indexing and Search for Information Extraction on Handwritten German Parish RecordsabstractWe endeavor to perform very large scale indexing of an ancient German collection of manuscript parish records. To this end we will compute "probabilistic indexes" (PIs), which are known to allow for very accurate and efficient implementation of (single-)keyword spotting. PIs may become prohibitively large for vast manuscript collections. Therefore we analyze simple index pruning methods to achieve adequate tradeoffs between memory requirements and search performance. We also study how to adequately deal with the large variety of non-ASCII symbols and handwritten word spelling variations (accents, umlauts, etc.) which appear in this kind of historical collections. Finally, and most importantly, since most of the images of the collection we aim to index are handwritten tables, we explore the use of PIs to support structured queries for information extraction from untranscribed handwritten images containing tabular data. Empirical results on a small, but complex and representative dataset extracted from the collection considered confirm the viability and adequateness of the chosen approaches. Eva Lang, Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 4 |
| 2018 | From HMMs to RNNs: Computer-Assisted Transcription of a Handwritten Notarial Records CollectionabstractWe present the process which is being followed for the transcription of a large XVIII Century Manuscript collection with the help of Handwritten Text Recognition (HTR) Technology. The documents are being processed in batches of 50 pages each. For each batch we perform two semi-supervised processes: one in order to analyze the layout and detect the text lines and another to provide the full transcripts of the text. As per users request, both diplomatic and modernized transcripts, as well as semantically tagged versions are being produced. Layout analysis supervision is performed by means of a conventional layout editing tool. On the other hand, transcripts, including automatic modernization and tagging, are being produced by means of a web based computer-assisted interactive-predictive tool (CATTI). We present results of the performance of this process through 12 image batches processed so far. These results show the impact caused by an optical modelling technological transition: from classical HMM-based methods to new technology based on recurrent neural networks. Lorenzo Quirós, Vicente Bosch, Lluis Serrano, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 5 |
| 2017 | Preparatory KWS Experiments for Large-Scale Indexing of a Vast Medieval Manuscript Collection in the HIMANIS ProjectabstractMaking large-scale collections of digitized historical documents searchable is being earnestly demanded by many archives and libraries. Probabilistically indexing the text images of these collections by means of keyword spotting techniques is currently seen as perhaps the only feasible approach to meet this demand. A vast medieval manuscript collection, written in both Latin and French, called "Chancery", is currently being considered for indexing at large. In addition to its bilingual nature, one of the major difficulties of this collection is the very high rate of abbreviated words which, on the other hand, are completely expanded in the ground truth transcripts available. In preparation to undertake full indexing of Chancery, experiments have been carried out on a relatively small but fully representative subset of this collection. To this end, a keyword spotting approach has been adopted which computes word relevance probabilities using character lattices produced by a recurrent neural network and a N-gram character language model. Results confirm the viability of the chosen approach for the large-scale indexing aimed at and show the ability of the proposed modeling and training approaches to properly deal with the abbreviation difficulties mentioned. Théodore Bluche, Sébastien Hamel, Christopher Kermorvant, Joan Puigcerver, Dominique Stutzmann, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 7 |
| 2017 | Handwritten Music Recognition for Mensural Notation: Formulation, Data and Baseline ResultsabstractMusic is a key element for cultural transmission, and so large collections of music manuscripts have been preserved over the centuries. In order to develop computational tools for analysis, indexing and retrieval from these sources, it is necessary to transcribe the content to some machine-readable format. In this paper we discuss the Handwritten Music Recognition problem, which refers to the development of automatic transcription systems for musical manuscripts. We focus on mensural notation, one of the most widespread varieties of Western classical music. For that, we present a labeled corpus containing 576 staves, along with a baseline recognition system based on a combination of hidden Markov models and N-gram language models. The baseline error obtained at symbol level is about 40 % which, given the difficulty of the task, can be considered a good starting point for future developments. Our aim is that these data and preliminary results help to promote this research field, serving as a reference in future developments. Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 3 |
| 2017 | ICDAR2017 Competition on Information Extraction in Historical Handwritten RecordsabstractThe extraction of relevant information from historical handwritten document collections is one of the key steps in order to make these manuscripts available for access and searches. In this competition, the goal is to detect the named entities and assign each of them a semantic category, and therefore, to simulate the filling in of a knowledge database. This paper describes the dataset, the tasks, the evaluation metrics, the participants methods and the results. Alicia Fornés, Verónica Romero 0001, Arnau Baró, Juan Ignacio Toledo, Joan-Andreu Sánchez, Enrique Vidal 0001, Josep Lladós 0001 |
ICDAR | 6 |
| 2017 | ICDAR2017 Competition on Handwritten Text Recognition on the READ DatasetabstractThis paper describes the fourth edition of the Handwritten Text Recognition (HTR) competition that was prepared this time in the context of the International Conference on Document Analysis and Recognition (ICDAR) 2017. Previous editions of this competition were conducted, first, with datasets from the tranScriptorium project in ICFHR 2014, and ICDAR 2015, and then, with datasets from the "Recognition and Enrichment of Archival Documents (READ)" European project in ICFHR 2016. This competition aims to bring together researchers working on off-line HTR and provides them a suitable benchmark to compare their techniques on the task of transcribing typical and difficult historical handwritten documents. The competition proposed for ICDAR 2017 aims at introducing a usual scenario for some collections in which there exist transcripts at page level for many pages useful for training, but these transcripts are not aligned with line images. Two tracks with different conditions on the use of training data were proposed. Most of the data comes from the Alfred Escher Letter Collection. But handwritten images were drawn from other German collections written by several hands. Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Mauricio Villegas, Enrique Vidal 0001 |
ICDAR | 5 |
| 2017 | Querying out-of-vocabulary words in lexicon-based keyword spotting
Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
Neural Comput. Appl. | 3 |
| 2017 | Word graphs size impact on the performance of handwriting document applications
Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001 |
Neural Comput. Appl. | 3 |
| 2016 | Handwriting Transcription and Keyword Spotting in Historical Daily Records DocumentsabstractHistorical records of daily activities provide an intriguing look into the historic life. These documents have interesting information, useful for demography studies and genealogical research. However, automatic processing of historical documents, has mostly been focused on single works of literature and less on daily records, which tend to have a distinct layout, structure, and vocabulary. This paper presents a study about the capability of state-of-the-art handwritten text recognition and key word spotting systems, when applied to this kind of documents. A relatively small set of handwritten birth records registered in Wien in the 16th century is used in the experiments. A word accuracy of about 70% and an AP of 0.74 are achieved for plain image transcription and key word spotting respectively. Taking into account the many difficulties exhibited by these handwritten documents, these preliminary results are quite encouraging. Verónica Romero 0001, Alejandro H. Toselli, Joan-Andreu Sánchez, Enrique Vidal 0001 |
DAS | 4 |
| 2016 | Early Handwritten Music Recognition with Hidden Markov ModelsabstractThis work presents a statistical method to tackle the Handwritten Music Recognition task for Early notation, which comprises more than 200 different symbols. Unlike previous approaches to deal with music notation, our strategy is to perform a holistic recognition without any previous segmentation or staff removal process. The input consists of a page of a music book, which is processed to extract and normalize the staves contained. Then, a feature extraction process is applied to define such sections as a sequence of numerical vectors. The recognition is based on the use of Hidden Markov Models for the optical processing and smoothed N-grams as language model. Experimentation results over a historical archive of Hispanic music reported an error around 40 %, which confirms our proposal as a good starting point taking into account the difficulty of the task. Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 3 |
| 2016 | Sheet Music Statistical Layout AnalysisabstractIn order to provide access to the contents of ancient music scores to researchers, the transcripts of both the lyrics and the musical notation is required. Before attempting any type of automatic or semi-automatic transcription of sheet music, an adequate layout analysis (LA) is needed. This LA must provide not only the locations of the different image regions, but also adequate region labels to distinguish between different region types such as staff, lyric, etc. To this end, we adapt a stochastic framework for LA based on Hidden Markov Models that we had previously introduced for detection and classification of text lines in typical handwritten text images. The proposed approach takes a scanned music score image as input and, after basic preprocessing, simultaneously performs region detection and region classification in an integrated way. To assess this statistical LA approach several experiments were carried out on a representative sample of a historical music archive, under different difficulty settings. The results show that our approach is able to tackle these structured documents providing good results not only for region detection but also for classification of the different regions. Vicente Bosch, Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 4 |
| 2016 | ICFHR2016 Handwritten Keyword Spotting Competition (H-KWS 2016)abstractThe H-KWS 2016, organized in the context of the ICFHR 2016 conference aims at setting up an evaluation framework for benchmarking handwritten keyword spotting (KWS) examining both the Query by Example (QbE) and the Query by String (QbS) approaches. Both KWS approaches were hosted into two different tracks, which in turn were split into two distinct challenges, namely, a segmentation-based and a segmentation-free to accommodate different perspectives adopted by researchers in the KWS field. In addition, the competition aims to evaluate the submitted training-based methods under different amounts of training data. Four participants submitted at least one solution to one of the challenges, according to the capabilities and/or restrictions of their systems. The data used in the competition consisted of historical German and English documents with their own characteristics and complexities. This paper presents the details of the competition, including the data, evaluation metrics and results of the best run of each participating methods. Ioannis Pratikakis, Konstantinos Zagoris, Basilios Gatos, Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 6 |
| 2016 | Using the MGGI Methodology for Category-Based Language Modeling in Handwritten Marriage Licenses BooksabstractHandwritten marriage licenses books have been used for centuries by ecclesiastical and secular institutions to register marriages. The information contained in these historical documents is useful for demography studies and genealogical research, among others. Despite the generally simple structure of the text in these documents, automatic transcription and semantic information extraction is difficult due to the distinct and evolutionary vocabulary, which is composed mainly of proper names that change along the time. In previous works we studied the use of category-based language models to both improve the automatic transcription accuracy and make easier the extraction of semantic information. Here we analyze the main causes of the semantic errors observed in previous results and apply a Grammatical Inference technique known as MGGI to improve the semantic accuracy of the language model obtained. Using this language model, full handwritten text recognition experiments have been carried out, with results supporting the interest of the proposed approach. Verónica Romero 0001, Alicia Fornés, Enrique Vidal 0001, Joan-Andreu Sánchez |
ICFHR | 3 |
| 2016 | ICFHR2016 Competition on Handwritten Text Recognition on the READ DatasetabstractThis paper describes the Handwritten Text Recognition (HTR) competition on the READ dataset that has been held in the context of the International Conference on Frontiers in Handwriting Recognition 2016. This competition aims to bring together researchers working on off-line HTR and provide them a suitable benchmark to compare their techniques on the task of transcribing typical historical handwritten documents. Two tracks with different conditions on the use of training data were proposed. Ten research groups registered in the competition but finally five submitted results. The handwritten images for this competition were drawn from the German document Ratsprotokolle collection composed of minutes of the council meetings held from 1470 to 1805, used in the READ project. The selected dataset is written by several hands and entails significant variabilities and difficulties. The five participants achieved good results with transcriptions word error rates ranging from 21% to 47%. Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 4 |
| 2016 | Two Methods to Improve Confidence Scores for Lexicon-Free Word Spotting in Handwritten TextabstractTwo methods are presented to improve word confidence scores for Line-Level Query-by-String Lexicon-Free Keyword Spotting (KWS) in handwritten text images. The first one approaches true relevance probabilities by means of computations directly carried out on character lattices obtained from the lines images considered. The second method uses the same character lattices, but it obtains relevance scores by first computing frame-level character sequence scores which resemble the word posteriorgrams used in previous approaches for lexicon-based KWS. The first method results from a formal probabilistic derivation, which allow us to better understand and further develop the underlying ideas. The second one is less formal but, according with experiments presented in the paper, it obtains almost identical results with much lower computational cost. Moreover, in contrast with the first method, the second one allows to directly obtain accurate bounding boxes for the spotted words. Alejandro H. Toselli, Joan Puigcerver, Enrique Vidal 0001 |
ICFHR | 3 |
| 2016 | Exploiting Existing Modern Transcripts for Historical Handwritten Text RecognitionabstractExisting transcripts for historic manuscripts are a very valuable resource for training models useful for automatic recognition, aided transcription, and/or indexing of the remaining untranscribed parts of these collections. However, these existing transcripts generally exhibit two main problems which hinder their convenience: a) text of the transcripts is seldom aligned with manuscript lines, and b) text often deviate very significantly from what can be seen in the manuscript, either because writing style has been modernized or abbreviations have been expanded, or both. This work presents an analysis of these problems and discusses possible solutions for minimizing human effort needed to adapt existing transcripts in order to render them usable. Empirical results presented show the huge performance gain that can be obtained by adequately adapting the transcripts, thus motivating future development of the proposed solutions. Mauricio Villegas, Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001 |
ICFHR | 4 |
| 2016 | HMM word graph based keyword spotting in handwritten document images
Alejandro H. Toselli, Enrique Vidal 0001, Verónica Romero 0001, Volkmar Frinken |
Inf. Sci. | 2 |
| 2015 | Improving sigma-lognormal parameter extractionabstractA fully automatic framework based on the kinematic theory of rapid human movements was recently introduced for analyzing and modeling complex human movements patterns such as those involved in handwriting. In this paper, we present a new approach to better extract and estimate the lognormal primitives and parameters. Through a comprehensive evaluation using 32,000 words from a public database, we show that our approach greatly improves the state-of-the-art extractor. Daniel Martín-Albo, Réjean Plamondon, Enrique Vidal 0001 |
ICDAR | 3 |
| 2015 | Probabilistic interpretation and improvements to the HMM-filler for handwritten keyword spottingabstractTraditionally, the HMM-Filler approach has been widely used in the fields of speech recognition and handwritten text recognition to tackle lexicon-free, query-by-string keyword spotting (KWS). It computes a score to determine whether a given keyword is written in a certain image region. It is conjectured, that this score is related to the confidence of the system, respect to the previous question. However, it is still not clear what this relationship is. In this paper, the HMM-Filler score is derived from a probabilistic formulation of KWS, which gives a better understanding of its behavior and limits. Additionally, the same probabilistic framework is used to present a new algorithm to compute the KWS scores, which results in better average precision (AP), for a keyword spotting task in the widely used IAM database. We show that the new algorithm can improve the HMM-filler results up to 10.4% relative (5.3% absolute) points in AP, in the considered task. Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 3 |
| 2015 | ICDAR2015 Competition on Keyword Spotting for Handwritten DocumentsabstractThe principal goal of the Competition on Keyword Spotting for Handwritten Documents was to promote different approaches used in the field of Keyword Spotting and to fairly compare them using uniform data and metrics. To accommodate different perspectives adopted by researches in this field, the competition was divided into two distinct tracks, namely, a training-free and a training-based track, and each track entailed two optional assignments. Six participants submitted solutions to one or both assignments, depending on the capabilities and/or restrictions of their systems. The data used in the competition consisted of historical documents in English with different levels of complexity. This paper presents the details of the competition, including the data, evaluation metrics and results of the best participant methods. Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 3 |
| 2015 | ICDAR 2015 competition HTRtS: Handwritten Text Recognition on the tranScriptorium datasetabstractThis paper describes the second edition of the Handwritten Text Recognition (HTR) contest on the tranScriptorium datasets that has been held in the context of the International Conference on Document Analysis and Recognition 2015. Two tracks with different conditions on the use of training data were proposed. Nine research groups registered in the contest but finally three research submitted results. The handwritten images for this contest were drawn from the English “Bentham collection” dataset used in the tranScriptorium project. A small subset of this collection has been chosen for the present HTR competition. The selected subset has been written by several hands and entails significant variabilities and difficulties regarding the quality of text images, writing styles and crossed-out text. This contest is clearly more difficult than the the first edition both for training and for testing. A portion of the training dataset and the full test dataset were provided in the form of carefully segmented line images, along with the corresponding transcripts. Another portion of the training dataset was provided as raw images and their corresponding transcripts at region level. The three participants achieved good results, with transcription word error rates ranging from 31% down to 44%. Joan-Andreu Sánchez, Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001 |
ICDAR | 4 |
| 2015 | Context-aware lattice based filler approach for key word spotting in handwritten documentsabstractThe so-called filler or garbage Hidden Markov Models (HMM-Filler) are among the most widely used models for lexicon-free, query by string key word spotting (KWS) in the fields of speech recognition and (lately) handwritten text recognition. However, it has important drawbacks. First, the keyword-specific HMM Viterbi decoding process needed to obtain the confidence scores of each spotted word involves a large computational cost. Second, in its traditional conception, the model does not take into account any context information - and more recent works where simple character bi-gram context is used show that not only the computational cost becomes even larger, but also the required keyword-specific language model becomes quite intricate to build. In a previous work we introduced KWS methods based on character lattices which proved very much simpler and faster than the traditional HMM-Filler, while providing practically identical results. Here we extend our previous work by using context-aware character lattices obtained by means of Viterbi decoding with high-order character N-gram models. Experimental results show that, as compared with a direct 2-gram HMM-filler implementation, the proposed approach requires between one and two orders of magnitude less query computing time. Moreover, for the first time in the field of handwritten text KWS, Filler-based results for N-grams up to N = 6 are reported, clearly showing a great impact of context on precision-recall performance. Alejandro H. Toselli, Joan Puigcerver, Enrique Vidal 0001 |
ICDAR | 3 |
| 2015 | High performance Query-by-Example keyword spotting using Query-by-String techniquesabstractKeyword Spotting (KWS) has been traditionally considered under two distinct frameworks: Query-by-Example (QbE) and Query-by-String (QbS). In both cases the user of the system wished to find occurrences of a particular keyword in a collection of document images. The difference is that, in QbE, the keyword is given as an exemplar image while, in QbS the keyword is given as a text string. In several works, the QbS scenario has been approached using QbE techniques; but the converse has not been studied in depth yet, despite of the fact that QbS systems typically achieve higher accuracy. In the present work, we present a very effective probabilistic approach to QbE KWS, based on highly accurate QbS KWS techniques which rely on models which need to be trained from annotated data. To assess the effectiveness of this approach, we tackle the segmentation-free QbE task of the ICFHR-2014 Competition on Handwritten KWS. Our approach achieves a mean average precision (mAP) as high as 0.715, which improves by more than 70% the best mAP achieved in this competition (0.419 under the same experimental conditions). Enrique Vidal 0001, Alejandro H. Toselli, Joan Puigcerver |
ICDAR | 1 |
| 2015 | Optical modelling and language modelling trade-off for Handwritten Text RecognitionabstractTraining the models needed for Automatic Handwritten Text Recognition of historical documents generally requires a significant amount of human effort. This is mainly due to the great differences that often exist between collections and to the lack of linguistic resources from the period when the documents were written, which results in a need of manual data labelling effort. This paper presents a study on the reuse of models trained with data from a different collection, focusing on the contribution that the language model and the optical models have on the performance. An empirical evaluation is performed using data from Jeremy Bentham manuscripts with the aim of recognising a manuscript about a very different topic written by Jane Austen. Mauricio Villegas, Joan-Andreu Sánchez, Enrique Vidal 0001 |
ICDAR | 3 |
| 2015 | Context-Aware Gestures for Mixed-Initiative Text Editing UIsabstractThis work is focused on enhancing highly interactive text-editing applications with gestures. Concretely, we study Computer Assisted Transcription of Text Images (CATTI), a handwriting transcription system that follows a corrective feedback paradigm, where both the user and the system collaborate efficiently to produce a high-quality text transcription. CATTI-like applications demand fast and accurate gesture recognition, for which we observed that current gesture recognizers are not adequate enough. In response to this need we developed MinGestures, a parametric context-aware gesture recognizer. Our contributions include a number of stroke features for disambiguating copy-mark gestures from handwritten text, plus the integration of these gestures in a CATTI application. It becomes finally possible to create highly interactive stroke-based text-editing interfaces, without worrying to verify the user intent on-screen. We performed a formal evaluation with 22 e-pen users and 32 mouse users using a gesture vocabulary of 10 symbols. MinGestures achieved an outstanding accuracy (<1% error rate) with very high performance (<1 ms of recognition time). We then integrated MinGestures in a CATTI prototype and tested the performance of the interactive handwriting system when it is driven by gestures. Our results show that using gestures in interactive handwriting applications is both advantageous and convenient when gestures are simple but context-aware. Taken together, this work suggests that text-editing interfaces not only can be easily augmented with simple gestures, but also may substantially improve user productivity. Luis A. Leiva, Vicente Alabau, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
Interact. Comput. | 5 |
| 2014 | Ground-Truth Production in the Transcriptorium ProjectabstractTran Scriptorium is a 3-years project that aims to develop innovative, cost-effective solutions for the indexing, search and full transcription of historical handwritten document images, using Handwritten Text Recognition (HTR) technology. The production of ground-truth (GT) of a dataset of handwritten document images is among the first tasks. We address novel approaches for the faster production of this GT based on crowd-sourcing and on prior-knowledge methods. We also address here a novel low-cost semi-supervised procedure for obtaining pairs of correct line-level aligned detected/extracted text line images and text line transcripts, specially suitable for training models of the HTR technology employed in Tran Scriptorium. Basilios Gatos, Georgios Louloudis, Tim Causer, Kris Grint, Verónica Romero 0001, Joan-Andreu Sánchez, Alejandro H. Toselli, Enrique Vidal 0001 |
Document Analysis Systems | 8 |
| 2014 | Word-Graph Based Handwriting Key-Word Spotting: Impact of Word-Graph Size on PerformanceabstractKey-Word Spotting (KWS) in handwritten documents is approached here by means of Word Graphs (WG) obtained using segmentation-free handwritten text recognition technology based on N-gram Language Models and Hidden Markov Models. Linguistic context significantly boost KWS performance with respect to methods which ignore word contexts and/or rely on image-matching with pre-segmented isolated words. On the other hand, WG-based KWS can be significantly faster than other KWS approaches which directly work on the original images where, in general, computational demands are exceedingly high. A large WG contains most of the relevant information of the original text (line) image needed for KWS but, if it is too large, the computational advantages over traditional, image matching-based KWS become diminished. Conversely, if it is too small, relevant information may be lost, leading to degraded KWS precision/recall performance. We study the trade off between WG size and KWS information retrieval performance. Results show that small, computationally cheap WGs can be used without loosing the excellent KWS performance achieved with huge WGs. Alejandro H. Toselli, Enrique Vidal 0001 |
Document Analysis Systems | 2 |
| 2014 | Semiautomatic Text Baseline Detection in Large Historical Handwritten DocumentsabstractA semiautomatic iterative process for the detection of text baselines in historical handwritten document images is presented. It relies on the use of Hidden Markov Models (HMM) to provide initial text baselines hypotheses, followed by user review in order to produce ground-truth quality results. Using the set of revised baselines as ground truth, the HMM's are re-trained before processing the next batch of pages. This process has been evaluated in the context of a real transcription task which, as a by-product, has produced line-detection ground truth. We show that the usage of a formal, HMM-based line-detection approach which requires training data, not only yields good detection results but is also of practical use in large handwritten image collections. Through experiments with real users we show that the proposed approach has interesting features, namely, accuracy, scalability and ease of use, as well as low overall human effort requirements. Vicente Bosch, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 3 |
| 2014 | Training of On-Line Handwriting Text Recognizers with Synthetic Text Generated Using the Kinematic Theory of Rapid Human MovementsabstractA method for automatic generation of synthetic handwritten words is presented which is based in the Kinematic Theory and its Sigma-lognormal model. To generate a new synthetic sample, first a real word is modelled using the Sigma-lognormal model. Then the Sigma-lognormal parameters are randomly perturbed within a range, introducing human-like variations in the sample. Finally, the velocity function is recalculated taking into account the new parameters. The synthetic words are then used as training data for a Hidden Markov Model based on-line handwritten recognizer. The experimental results confirm the great potential of the kinematic theory of rapid human movements applied to writer adaptation. Daniel Martín-Albo, Réjean Plamondon, Enrique Vidal 0001 |
ICFHR | 3 |
| 2014 | Word-Graph and Character-Lattice Combination for KWS in Handwritten DocumentsabstractWe present a handwritten text Keyword Spotting (KWS) approach based on the combination of KWS methods using word-graphs (WGs) and character-lattices (CLs). It aims to solve the problem that WG-based models present for out of vocabulary (OOV) keywords: since there is no available information about them in the lexicon or the language model, null scores are assigned. OOV keywords may have a significant impact on the global performance of KWS systems, as we show. By using a CL approach, which does not suffer from the previous problem, to estimate the OOV scores, we take advantage of both models, using the speed and accuracy that WGs provide for in-vocabulary keywords and the flexibility of the CL approach. This combination improves significantly both average precision and mean average precision over the two methods. Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 3 |
| 2014 | ICFHR2014 Competition on Handwritten Text Recognition on Transcriptorium Datasets (HTRtS)abstractA contest on Handwritten Text Recognition organised in the context of the ICFHR 2014 conference is described. Two tracks with increased freedom on the use of training data were proposed and three research groups participated in these two tracks. The handwritten images for this contest were drawn from an English data set which is currently being considered in the Tran scriptorium project. The goal of this project is to develop innovative, efficient and cost-effective solutions for the transcription of historical handwritten document images, focusing on four languages: English, Spanish, German and Dutch. For the English language, the so-called "Bentham collection" is being considered in Tran scriptorium. It encompasses a large set of manuscripts written by the renowned English philosopher and reformer Jeremy Bentham (1748-1832). A small subset of this collection has been chosen for the present HTR competition. The selected subset has been written by several hands (Bentham himself and his secretaries) and entails significant variabilities and difficulties regarding the quality of text images and writing styles. Training and test data were provided in the form of carefully segmented line images, along with the corresponding transcripts. The three participants achieved very good results, with transcription word error rates ranging from 15.0% down to 8.6%. Joan-Andreu Sánchez, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 4 |
| 2014 | Word-Graph-Based Handwriting Keyword Spotting of Out-of-Vocabulary QueriesabstractThanks to the use of lexical and syntactic information, Word Graphs (WG) have shown to provide a competitive Precision-Recall performance, along with fast lookup times, in comparison to other techniques used for Key-Word Spotting (KWS) in handwritten text images. However, a problem of WG approaches is that they assign a null score to any keyword that was not part of the training data, i.e. Out-of-Vocabulary (OOV) keywords, whereas other techniques are able to estimate a reasonable score even for these kind of keywords. We present a smoothing technique which estimates the score of an OOV keyword based on the scores of similar keywords. This makes the WG-based KWS as flexible as other techniques with the benefit of having much faster lookup times. Joan Puigcerver, Alejandro H. Toselli, Enrique Vidal 0001 |
ICPR | 3 |
| 2014 | Interactive translation prediction versus conventional post-editing in practice: a study with the CasMaCat workbench
Germán Sanchis-Trilles, Vicente Alabau, Christian Buck, Michael Carl, Francisco Casacuberta, Mercedes García-Martínez, Ulrich Germann, Jesús González-Rubio, Robin L. Hill, Philipp Koehn, Luis A. Leiva, Bartolomé Mesa-Lao, Daniel Ortiz-Martínez, Herve Saint-Amand, Chara Tsoukala, Enrique Vidal 0001 |
Mach. Transl. | 16 |
| 2013 | tranScriptorium: a european project on handwritten text recognitionabstractThe tranScriptorium project aims to develop innovative, efficient and cost-effective solutions for annotating handwritten historical documents using modern, holistic Handwritten Text Recognition (HTR) technology. Three actions are planned in tranScriptorium: i) improve basic image preprocessing and holistic HTR techniques; ii) develop novel indexing and keyword searching approaches; and iii) capitalize on new, user-friendly interactive-predictive HTR approaches for computer-assisted operation. Joan-Andreu Sánchez, Günter Mühlberger, Basilios Gatos, Philip Schofield, Katrien Depuydt, Richard M. Davis, Enrique Vidal 0001, Jesse de Does |
ACM Symposium on Document Engineering | 7 |
| 2013 | Interactive Off-Line Handwritten Text Transcription Using On-Line Handwritten Text as FeedbackabstractHandwritten Text Recognition is a problem that has gained attention in the last years mainly due to the interest in the transcription of historical documents. However, the automatic transcription is ineffectual in unconstrained handwritten documents. Thus, human intervention is typically needed to correct the results. Given that a post-editing approach is inefficient and uncomfortable, multimodal interactive approaches have begun to emerge in the last years. In this scheme, the user interacts with the system by means of an e-pen. This multimodal feedback, on the one hand, allows to improve the accuracy of the system and, on the other hand, increases user acceptability. In this work, we present a new approach on interaction based on character sequences. Here we present developments that allow taking advantage of interaction-derived context to significantly improve feedback decoding accuracy. Empirical tests suggest that, despite the loss of the deterministic accuracy of traditional peripherals, this approach can save significant amounts of user effort with respect to non-interactive post-editing correction. Daniel Martín-Albo, Verónica Romero 0001, Enrique Vidal 0001 |
ICDAR | 3 |
| 2013 | Fast HMM-Filler Approach for Key Word Spotting in Handwritten DocumentsabstractThe so-called filler or garbage Hidden Markov Models (HMM) are among the most widely used models for lexicon-free, query by string key word spotting in the fields of speech recognition and (lately) handwritten text recognition. An important drawback of this approach is the large computational cost of the keyword-specific HMM Viterbi decoding process needed to obtain the confidence scores of each word to be spotted. This paper presents a novel way to compute such confidence scores, directly from character lattices produced during a single Viterbi decoding process using only the "filler" model (i.e. no explicit keyword-specific decoding is needed). Experiments show that, as compared with the classical HMM-filler approach, the proposed method obtains essentially the same spotting results, while requiring between one and two orders of magnitude less query computing time. Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 2 |
| 2013 | Warped K-Means: An algorithm to cluster sequentially-distributed data
Luis A. Leiva, Enrique Vidal 0001 |
Inf. Sci. | 2 |
| 2013 | The ESPOSALLES database: An ancient marriage license corpus for off-line handwriting recognition
Verónica Romero 0001, Alicia Fornés, Joan-Andreu Sánchez, Alejandro H. Toselli, Volkmar Frinken, Enrique Vidal 0001, Josep Lladós 0001 |
Pattern Recognit. | 7 |
| 2012 | Statistical Text Line Analysis in Handwritten DocumentsabstractIn this paper we present an approach for text line analysis and detection in handwritten documents based on Hidden Markov Models, a technique widely used in other handwritten and speech recognition tasks. It is shown that text line analysis and detection can be solved using a more formal methodology in contraposition to most of the proposed heuristic approaches found in the literature. Our approach not only provides the best position coordinates for each of the vertical page regions but also labels them, in this manner surpassing the traditional heuristic methods. In our experiments we demonstrate the performance of the approach (both in line analysis and detection) and study the impact of increasingly constrained "vertical layout language models" on text line detection accuracy. Through this experimentation we also show the improvement in quality of the baselines yielded by our approach in comparison with a state-of-the-art heuristic method based on vertical projection profiles. Vicente Bosch, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 3 |
| 2012 | Simple, fast, and accurate clustering of data sequencesabstractMany devices generate large amounts of data that follow some sort of sequentiality, e.g., motion sensors, e-pens, or eye trackers, and therefore these data often need to be compressed for classification, storage, and/or retrieval purposes. This paper introduces a simple, accurate, and extremely fast technique inspired by the well-known K-means algorithm to properly cluster sequential data. We illustrate the feasibility of our algorithm on a web-based prototype that works with trajectories derived from mouse and touch input. As can be observed, our proposal outperforms the classical K-means algorithm in terms of accuracy (better, well-formed segmentations) and performance (less computation time). Luis A. Leiva, Enrique Vidal 0001 |
IUI | 2 |
| 2012 | Multimodal Computer-Assisted transcription of Text Images at Character-Level InteractionabstractCurrently, automatic handwriting recognition systems are ineffectual in unconstrained handwriting documents. Therefore, to obtain perfect transcriptions, heavy human intervention is required to validate and correct the results of such systems. Given that this post-editing process is inefficient and uncomfortable, a multimodal interactive approach has been proposed in previous works, which aims at obtaining correct transcriptions with the minimum human effort. In this approach, the user interacts with the system by means of an e-pen and/or more traditional methods such as keyboard or mouse. This user's feedback allows to improve system accuracy and multimodality increases system ergonomics and user acceptability. Until now, multimodal interaction has been considered only at whole-word level. In this work, multimodal interaction at character-level is studied, that may lead to more effective interactivity, since it is faster and easier to write only one character rather than a whole word. Here we study this kind of fine-grained multimodal interaction and present developments that allow taking advantage of interaction-derived context to significantly improve feedback decoding accuracy. Empirical tests on three cursive handwritten tasks suggest that, despite losing the deterministic accuracy of traditional peripherals, this approach can save significant amounts of user effort with respect to fully manual transcription as well as to noninteractive post-editing correction. Daniel Martín-Albo, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2012 | A Word-Based Naïve Bayes Classifier for Confidence Estimation in Speech RecognitionabstractConfidence estimation has been largely used in speech recognition to detect words in the recognized sentence that have been likely misrecognized. Confidence estimation can be seen as a conventional pattern classification problem in which a set of features is obtained for each hypothesized word in order to classify it as either correct or incorrect. We propose a smoothed naïve Bayes classification model to profitably combine these features. The model itself is a combination of word-dependent (specific) and word-independent (generalized) naïve Bayes models. As in statistical language modeling, the purpose of the generalized model is to smooth the (class posterior) estimates given by the specific models. Our classification model is empirically compared with confidence estimation based on posterior probabilities computed on word graphs. Empirical results clearly show that the good performance of word graph-based posterior probabilities can be improved by using the naïve Bayes combination of features. Alberto Sanchís, Alfons Juan-Císcar, Enrique Vidal 0001 |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Evaluating an Interactive-Predictive Paradigm on Handwriting Transcription: A Case Study and Lessons LearnedabstractTranscribing handwritten text is a laborious task which currently is carried out manually. As the accuracy of automatic handwritten text recognizers improves, post-editing the output of these recognizers could be foreseen as a possible alternative. Alas, the state-of-the-art technology is not suitable to perform this kind of work, since current approaches are not accurate enough and the process is usually both inefficient and uncomfortable for the user. As alternative, an interactive-predictive paradigm has gained recently an increasing popularity, mainly due to promising empirical results that estimate considerable reductions of user effort. In order to assess whether these empirical results can lead indeed to actual benefits, we developed a working prototype and conducted a field study remotely. Thirteen regular computer users tested two different transcription engines through the above-mentioned prototype. We observed that the interactive-predictive version allowed to transcribe better (less errors and fewer iterations to achieve a high-quality output) in comparison to the manual engine. Additionally, participants ranked higher such an interactive-predictive system in a usability questionnaire. We describe the evaluation methodology and discuss our preliminary results. While acknowledging the known limitations of our experimentation, we conclude that the interactive-predictive paradigm is an efficient approach for transcribing handwritten text. Luis A. Leiva, Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
COMPSAC | 4 |
| 2011 | Handwritten Text Recognition for Marriage Register BooksabstractMarriage register books are documents that were used for centuries by ecclesiastical institutions to register marriages. Most of these books were handwritten. These documents have interesting information, useful for demography studies. The information in these books is usually collected by expert demographers that devote a lot of time to transcribe them. The automatic transcription of these documents by using Handwritten Text Recognition techniques is difficult since the vocabulary is large, given that it is composed mainly of proper names. In this work, interactive Handwritten Text Recognition techniques were studied for the assisted transcription of these documents. Verónica Romero 0001, Joan-Andreu Sánchez, Enrique Vidal 0001 |
ICDAR | 4 |
| 2011 | Study of different interactive editing operations in an assisted transcription systemabstractTo date, automatic handwriting recognition systems are far from being perfect. Therefore, once the full recognition process of a handwritten text image has finished, heavy human intervention is required in order to correct the results of such systems. As an alternative, an interactive system has been proposed in previous works. This alternative follows an Interactive Predictive paradigm and the results show that significant amounts of human effort can be saved. So far only word substitutions and pointer actions have been considered in this interactive system. In this work, we study different interactive editing operations that can allow for more effective, ergonomic and friendly interfaces. Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICMI | 3 |
| 2010 | The APP Oracle - An Interactive Student Competition on Pattern Recognition
Alfons Juan-Císcar, Jesús Andrés-Ferrer, Adrià Giménez, Jorge Civera, Roberto Paredes, Enrique Vidal 0001 |
CSEDU (2) | 6 |
| 2010 | Interactive layout analysis and transcription systems for historic handwritten documentsabstractThe amount of digitized legacy documents has been rising dramatically over the last years due mainly to the increasing number of on-line digital libraries publishing this kind of documents, waiting to be classified and finally transcribed into a textual electronic format (such as ASCII or PDF). Nevertheless, most of the available fully-automatic applications addressing this task are far from being perfect and heavy and inefficient human intervention is often required to check and correct the results of such systems. In contrast, multimodal interactive-predictive approaches may allow the users to participate in the process helping the system to improve the overall performance. With this in mind, two sets of recent advances are introduced in this work: a novel interactive method for text block detection and two multimodal interactive handwritten text transcription systems which use active learning and interactive-predictive technologies in the recognition process. Oriol Ramos Terrades, Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001, Alfons Juan-Císcar |
ACM Symposium on Document Engineering | 5 |
| 2010 | Character-Level Interaction in Computer-Assisted Transcription of Text ImagesabstractTo date, automatic handwriting recognition systems are far from being perfect and heavy human intervention is often required to check and correct the results of such systems. As an alternative, an interactive framework that integrates the human knowledge into the transcription process has been presented in previous works. This new approach follows an Interactive Predictive paradigm and our results show that significant amounts of human effort can be saved. Until now only whole-word interactions with this system have been considered. In this work, character-level keystroke interactions, that can allow for a more ergonomic and friendly interfaces, are proposed. Empirical results show that this allows for further improvements in user productivity. Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICFHR | 3 |
| 2010 | A Bi-modal Handwritten Text Corpus: Baseline ResultsabstractHandwritten text is generally captured through two main modalities: off-line and on-line. Smart approaches to handwritten text recognition (HTR) may take advantage of both modalities if they are available. This is for instance the case in computer-assisted transcription of text images, where on-line text can be used to interactively correct errors made by a main off-line HTR system. We present here baseline results on the biMod-IAM-PRHLT corpus, which was recently compiled for experimentation with techniques aimed at solving the proposed multi-modal HTR problem, and is being used in one of the official ICPR-2010 contests. Moisés Pastor, Alejandro H. Toselli, Francisco Casacuberta, Enrique Vidal 0001 |
ICPR | 4 |
| 2010 | Computer Assisted Transcription of Text Images: Results on the GERMANA Corpus and Analysis of Improvements Needed for Practical UseabstractWe present a study of the application of Computer Assisted Transcription of Text Images (CATTI) to a task which is much closer to real applications than other tasks previously studied. The new task consists in the transcription of a new publicly available historic handwritten document, called GERMANA. A detailed analysis of the main factors influencing the system performance are exposed and some strategies to circumvent them are proposed. Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICPR | 3 |
| 2010 | Multimodal interactive transcription of text images
Alejandro H. Toselli, Verónica Romero 0001, Moisés Pastor, Enrique Vidal 0001 |
Pattern Recognit. | 4 |
| 2009 | Using Mouse Feedback in Computer Assisted Transcription of Handwritten Text ImagesabstractTo date, automatic handwriting recognition systems are far from being perfect and heavy human intervention is often required to check and correct the results of such systems. In order to achieve correct transcriptions, human knowledge can be integrated into the transcription process, following an Interactive Predictive paradigm. We have recently proposed Mouse Actions as a significant feedback information source for the underlying interactive system to improve the productivity of the human transcriptor. In this paper we review this way to interact with the system and report comparative results using the publicly available IAMDB dataset. Verónica Romero 0001, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 3 |
| 2009 | Interactive multimodal transcription of text images using a web-based demo systemabstractThis document introduces a web based demo of an interactive framework for transcription of handwritten text, where the user feedback is provided by means of pen strokes on a touchscreen. Here, the automatic handwriting text recognition system and the user both cooperate to generate the final transcription. Verónica Romero 0001, Luis A. Leiva, Alejandro H. Toselli, Enrique Vidal 0001 |
IUI | 4 |
| 2009 | Statistical Approaches to Computer-Assisted TranslationabstractCurrent machine translation (MT) systems are still not perfect. In practice, the output from these systems needs to be edited to correct errors. A way of increasing the productivity of the whole translation process (MT plus human work) is to incorporate the human correction activities within the translation process itself, thereby shifting the MT paradigm to that of computer-assisted translation. This model entails an iterative process in which the human translator activity is included in the loop: In each iteration, a prefix of the translation is validated (accepted or amended) by the human and the system computes its best (or n-best) translation suffix hypothesis to complete this prefix. A successful framework for MT is the so-called statistical (or pattern recognition) framework. Interestingly, within this framework, the adaptation of MT systems to the interactive scenario affects mainly the search process, allowing a great reuse of successful techniques and models. In this article, alignment templates, phrase-based models, and stochastic finite-state transducers are used to develop computer-assisted translation systems. These systems were assessed in a European project (TransType2) in two real tasks: The translation of printer manuals; manuals and the translation of the Bulletin of the European Union. In each task, the following three pairs of languages were involved (in both translation directions): English-Spanish, English-German, and English-French. Sergio Barrachina 0001, Oliver Bender, Francisco Casacuberta, Jorge Civera, Elsa Cubel, Shahram Khadivi, Antonio L. Lagarda, Hermann Ney, Jesús Tomás, Enrique Vidal 0001, Juan Miguel Vilar |
Comput. Linguistics | 10 |
| 2008 | Improving Interactive Machine Translation via Mouse Actions
Germán Sanchis-Trilles, Daniel Ortiz-Martínez, Jorge Civera, Francisco Casacuberta, Enrique Vidal 0001, Hieu Hoang |
EMNLP | 5 |
| 2008 | Maximum entropy models for speech confidence estimationabstractIn this work we implement a confidence estimation system based on a Naive Bayes classifier, by using the maximum entropy paradigm. The model takes information from various sources including a set of scores which have proved to be useful in confidence estimation tasks. Two different approaches are modeled. First a basic model which takes advantages of smoothing techniques used in a previous work, and second an optimized model, which is designed to hold a set of very few but essential characteristics of the model, without decrease in the performance. A considerably reduction in the number of parameters is obtained compared to the basic model. Both models are evaluated with two different corpora and compared to a model previously developed. Claudio Estienne, Alberto Sanchís, Alfons Juan-Císcar, Enrique Vidal 0001 |
ICASSP | 4 |
| 2008 | Learning weighted distances for relevance feedback in image retrievalabstractWe present a new method for relevance feedback in image retrieval and a scheme to learn weighted distances which can be used in combination with different relevance feedback methods. User feedback is a crucial step in image retrieval to maximise retrieval performance as was shown in recent image retrieval evaluations. Machine learning is expected to be able to learn how to rank images according to users needs. Most image retrieval systems incorporate user feedback using rather heuristic means and only few groups have formally investigated how to maximise the benefit from it using machine learning techniques. We incorporate our distance-learning method into our new relevance feedback scheme and into two different approaches from the literature. The methods are compared on two publicly available databases, one which is purely content-based and one which uses additional textual information. It is shown that the new relevance feedback scheme outperforms the other methods and that all methods benefit from weighted distance learning. Thomas Deselaers, Roberto Paredes, Enrique Vidal 0001, Hermann Ney |
ICPR | 3 |
| 2007 | Computer Assisted Transcription of Handwritten Text ImagesabstractTo date, automatic handwriting recognition systems are far from being perfect and often they need a post editing where a human intervention is required to check and correct the results of such systems. We propose to have a new interactive, on-line framework which, rather than full automation, aims at assisting the human in the proper recognition- transcription process; that is, facilitate and speed up their transcription task of handwritten texts. This framework combines the efficiency of automatic handwriting recognition systems with the accuracy of the human transcriptor. The best result is a cost-effective perfect transcription of the handwriting text images. Alejandro H. Toselli, Verónica Romero 0001, Luis Rodríguez, Enrique Vidal 0001 |
ICDAR | 4 |
| 2007 | Estimation of confidence measures for machine translation
Alberto Sanchís, Alfons Juan-Císcar, Enrique Vidal 0001 |
MTSummit | 3 |
| 2007 | Learning finite-state models for machine translationabstractIn formal language theory, finite-state transducers are well-know models for simple “input-output” mappings between two languages. Even if more powerful, recursive models can be used to account for more complex mappings, it has been argued that the input-output relations underlying most usual natural language pairs can essentially be modeled by finite-state devices. Moreover, the relative simplicity of these mappings has recently led to the development of techniques for learning finite-state transducers from a training set of input-output sentence pairs of the languages considered. In the last years, these techniques have lead to the development of a number of machine translation systems. Under the statistical statement of machine translation, we overview here how modeling, learning and search problems can be solved by using stochastic finite-state transducers. We also review the results achieved by the systems we have developed under this paradigm. As a main conclusion of this review we argue that, as task complexity and training data scarcity increase, those systems which rely more on statistical techniques tend produce the best results. Francisco Casacuberta, Enrique Vidal 0001 |
Mach. Learn. | 2 |
| 2006 | A Computer-Assisted Translation Tool based on Finite-State Technology
Jorge Civera, Antonio L. Lagarda, Elsa Cubel, Francisco Casacuberta, Enrique Vidal 0001, Juan Miguel Vilar, Sergio Barrachina 0001 |
EAMT | 5 |
| 2006 | Learning Weighted Metrics to Minimize Nearest-Neighbor Classification ErrorabstractIn order to optimize the accuracy of the Nearest-Neighbor classification rule, a weighted distance is proposed, along with algorithms to automatically learn the corresponding weights. These weights may be specific for each class and feature, for each individual prototype, or for both. The learning algorithms are derived by (approximately) minimizing the Leaving-One-Out classification error of the given training set. The proposed approach is assessed through a series of experiments with UCI/STATLOG corpora, as well as with a more specific task of text classification which entails very sparse data representation and huge dimensionality. In all these experiments, the proposed approach shows a uniformly good behavior, with results comparable to or better than state-of-the-art results published with the same data so far. Roberto Paredes, Enrique Vidal 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Learning prototypes and distances: A prototype reduction technique based on nearest neighbor error minimization
Roberto Paredes, Enrique Vidal 0001 |
Pattern Recognit. | 2 |
| 2006 | Computer-assisted translation using speech recognitionabstractCurrent machine translation systems are far from being perfect. However, such systems can be used in computer-assisted translation to increase the productivity of the (human) translation process. The idea is to use a text-to-text translation system to produce portions of target language text that can be accepted or amended by a human translator using text or speech. These user-validated portions are then used by the text-to-text translation system to produce further, hopefully improved suggestions. There are different alternatives of using speech in a computer-assisted translation system: From pure dictated translation to simple determination of acceptable partial translations by reading parts of the suggestions made by the system. In all the cases, information from the text to be translated can be used to constrain the speech decoding search space. While pure dictation seems to be among the most attractive settings, unfortunately perfect speech decoding does not seem possible with the current speech processing technology and human error-correcting would still be required. Therefore, approaches that allow for higher speech recognition accuracy by using increasingly constrained models in the speech recognition process are explored here. All these approaches are presented under the statistical framework. Empirical results support the potential usefulness of using speech within the computer-assisted translation paradigm. Enrique Vidal 0001, Francisco Casacuberta, Luis Rodríguez, Jorge Civera, Carlos D. Martínez-Hinarejos |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Writing Speed Normalization for On-Line Handwritten Text RecognitionabstractPen-based interfaces aim at improving the man-machine interaction of many portable systems. While statistical models can be used to learn pen position sequences, they suffer from the huge variability exhibited by the speed of writing. To improve performance, invariance to the writing speed is needed. Trace segmentation is a technique that can be used to normalize the writing speed. This method is controlled by a parameter called resampling distance. A study of the resampling distance is presented here, along with another approximation to the writing speed normalization called "derivatives normalization". The improvement using trace segmentation was 193% relative to the baseline, whilst the improvement using derivatives normalization was 47.3% relative. Moisés Pastor, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 3 |
| 2005 | On the use of speech recognition in computer assisted translation
Luis Rodríguez, Jorge Civera, Enrique Vidal 0001, Francisco Casacuberta, César Ernesto Martínez |
INTERSPEECH | 3 |
| 2005 | Probabilistic Finite-State Machines-Part IabstractProbabilistic finite-state machines are used today in a variety of areas in pattern recognition, or in fields to which pattern recognition is linked: computational linguistics, machine learning, time series analysis, circuit testing, computational biology, speech recognition, and machine translation are some of them. In Part I of this paper, we survey these generative objects and study their definitions and properties. In Part II, we will study the relation of probabilistic finite-state automata with other well-known devices that generate strings as hidden Markov models and n-grams and provide theorems, algorithms, and properties that represent a current state of the art of these objects. Enrique Vidal 0001, Franck Thollard, Colin de la Higuera, Francisco Casacuberta, Rafael C. Carrasco |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Probabilistic Finite-State Machines-Part IIabstractProbabilistic finite-state machines are used today in a variety of areas in pattern recognition or in fields to which pattern recognition is linked. In part I of this paper, we surveyed these objects and studied their properties. In this part, we study the relations between probabilistic finite-state automata and other well-known devices that generate strings like hidden Markov models and n-grams and provide theorems, algorithms, and properties that represent a current state of the art of these objects. Enrique Vidal 0001, Franck Thollard, Colin de la Higuera, Francisco Casacuberta, Rafael C. Carrasco |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Inference of finite-state transducers from regular languages
Francisco Casacuberta, Enrique Vidal 0001, David Picó |
Pattern Recognit. | 2 |
| 2004 | Finite-State Models for Computer Assisted Translation
Elsa Cubel, Jorge Civera, Juan Miguel Vilar, Antonio L. Lagarda, Francisco Casacuberta, Enrique Vidal 0001, David Picó, Luis Rodríguez |
ECAI | 6 |
| 2004 | From Machine Translation to Computer Assisted Translation using Finite-State Models
Jorge Civera, Elsa Cubel, Antonio L. Lagarda, David Picó, Enrique Vidal 0001, Francisco Casacuberta, Juan Miguel Vilar, Sergio Barrachina 0001 |
EMNLP | 6 |
| 2004 | New features based on multiple word graphs for utterance verificationabstractThe goal of Utterance Verification is to estimate a confidence measure which helps detecting words in the hypothesized sentence that are likely to have been missrecognized. Word graphs have been extensively employed for directly estimating the confidence measure and for extracting important predictor features. In all the cases, a single word graph which is obtained through the recognition process. In this paper we propose the use of multiple word graphs to compute new features. The experimental study shows that these proposed features outperform those computed on a single word graph and other well-known predictor features. Moreover, the combination of the proposed features along with other kind of features provides improvements in the verification accuracy. Alberto Sanchís, Alfons Juan-Císcar, Enrique Vidal 0001 |
INTERSPEECH | 3 |
| 2004 | Pattern Recognition Approaches for Speech-To-Speech TranslationabstractWe propose a statistical approach to speech-to-speech translation that uses finite-state models in all levels. Acoustic hidden Markov models (HMMs) model the pronunciation of the input-language phonemes and words, while the input–output word mapping, along with the syntax of the output language, are jointly modeled by means a large stochastic finite-state transducer. This allows for a complete integration of all the models so that the translation process can be performed by searching for an optimal path of states through the integrated network. As in speech recognition, HMMs can be trained from an input-language speech corpus, and the translation model is learned automatically from a parallel (text) training corpus. This approach has been assessed in the framework of the EuTrans project, funded by the European Union. Extensive experiments have been carried out with speech-input translations from Spanish to English and from Italian to English in applications involving the interaction (by telephone) of a customer with the front desk of a hotel. A summary of the most relevant results is presented. Francisco Casacuberta, Enrique Vidal 0001, Alberto Sanchís, Juan Miguel Vilar |
Cybern. Syst. | 2 |
| 2004 | Machine Translation with Inferred Stochastic Finite-State TransducersabstractFinite-state transducers are models that are being used in different areas of pattern recognition and computational linguistics. One of these areas is machine translation, in which the approaches that are based on building models automatically from training examples are becoming more and more attractive. Finite-state transducers are very adequate for use in constrained tasks in which training samples of pairs of sentences are available. A technique for inferring finite-state transducers is proposed in this article. This technique is based on formal relations between finite-state transducers and rational grammars. Given a training corpus of source-target pairs of sentences, the proposed approach uses statistical alignment methods to produce a set of conventional strings from which a stochastic rational grammar (e.g., an n-gram) is inferred. This grammar is finally converted into a finite-state transducer. The proposed methods are assessed through a series of machine translation experiments within the framework of the E u Trans project. Francisco Casacuberta, Enrique Vidal 0001 |
Comput. Linguistics | 2 |
| 2004 | Some approaches to statistical and finite-state speech-to-speech translation
Francisco Casacuberta, Hermann Ney, Franz Josef Och, Enrique Vidal 0001, Juan Miguel Vilar, Sergio Barrachina 0001, Ismael García-Varea, David Llorens, Carlos D. Martínez-Hinarejos, Sirko Molau |
Comput. Speech Lang. | 4 |
| 2004 | Integrated Handwriting Recognition And Interpretation Using Finite-State ModelsabstractThe interpretation of handwritten sentences is carried out using a holistic approach in which both text image recognition and the interpretation itself are tightly integrated. Conventional approaches follow a serial, first-recognition then-interpretation scheme which cannot adequately use semantic–pragmatic knowledge to recover from recognition errors. Stochastic finite-sate transducers are shown to be suitable models for this integration, permitting a full exploitation of the final interpretation constraints. Continuous-density hidden Markov models are embedded in the edges of the transducer to account for lexical and morphological constraints. Robustness with respect to stroke vertical variability is achieved by integrating tangent vectors into the emission densities of these models. Experimental results are reported on a syntax-constrained interpretation task which show the effectiveness of the proposed approaches. These results are also shown to be comparatively better than those achieved with other conventional, N-gram-based techniques which do not take advantage of full integration. Alejandro H. Toselli, Alfons Juan-Císcar, Ismael Salvador, Enrique Vidal 0001, Francisco Casacuberta, Daniel Keysers, Hermann Ney |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2003 | Musical Style Identification Using Grammatical Inference: The Encoding Problem
Pedro P. Cruz-Alcázar, Enrique Vidal 0001, Juan-Carlos Perez-Cortes |
CIARP | 2 |
| 2003 | Improving utterance verification using a smoothed naive Bayes modelabstractUtterance verification can be seen as a conventional pattern classification problem in which a feature vector is obtained for each hypothesized word in order to classify it as either correct or incorrect. It is unclear, however, which predictor (pattern) features and classification model should be used. Regarding the features, we have proposed a new feature, called word trellis stability (WTS), that can be profitably used in conjunction with more or less standard features such as acoustic stability. This is confirmed in this paper, where a smoothed naive Bayes classification model is proposed to adequately combine predictor features. On a series of experiments with this classification model and several features, we have found that the results provided by each feature alone are outperformed by certain combinations. In particular, the combination of the two above-mentioned features has been consistently found to give the most accurate result in two verification tasks. Alberto Sanchís, Alfons Juan-Císcar, Enrique Vidal 0001 |
ICASSP (1) | 3 |
| 2003 | Utterance verification using an optimized k-nearest neighbour classifier
Roberto Paredes, Alberto Sanchís, Enrique Vidal 0001, Alfons Juan-Císcar |
INTERSPEECH | 3 |
| 2002 | Cyclic Sequence Alignments: Approximate Versus Optimal TechniquesabstractThe problem of cyclic sequence alignment is considered. Most existing optimal methods for comparing cyclic sequences are very time consuming. For applications where these alignments are intensively used, optimal methods are seldom a feasible choice. The alternative to an exact and costly solution is to use a close-to-optimal but cheaper approach. In previous works, we have presented three suboptimal techniques inspired on the quadratic-time suboptimal algorithm proposed by Bunke and Bühler. Do these approximate approaches come sufficiently close to the optimal solution, with a considerable reduction in computing time? Is it thus worthwhile investigating these approximate methods? This paper shows that approximate techniques are good alternatives to optimal methods. Ramón Alberto Mollineda, Enrique Vidal 0001, Francisco Casacuberta |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2002 | On the use of Bernoulli mixture models for text classification
Alfons Juan-Císcar, Enrique Vidal 0001 |
Pattern Recognit. | 2 |
| 2002 | An efficient prototype merging strategy for the condensed 1-NN rule through class-conditional hierarchical clustering
Ramón Alberto Mollineda, Francesc J. Ferri, Enrique Vidal 0001 |
Pattern Recognit. | 3 |
| 2002 | A merge-based condensing strategy for multiple prototype classifiersabstractA class-conditional hierarchical clustering framework has been used to generalize and improve previously proposed condensing schemes to obtain multiple prototype classifiers. The proposed method conveniently uses geometric properties and clusters to efficiently obtain reduced sets of prototypes that accurately represent the data while significantly keeping its discriminating power. The benefits of the proposed approach are empirically assessed with regard to other previously proposed algorithms which are similar in their foundations. Other well-known multiple prototype classifiers have also been taken into account in the comparison. Ramón Alberto Mollineda, Francesc J. Ferri, Enrique Vidal 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2001 | Speech-to-speech translation based on finite-state transducersabstractNowadays, the most successful speech recognition systems are based on stochastic finite-state networks (hidden Markov models and n-grams). Speech translation can be accomplished in a similar way as speech recognition. Stochastic finite-state transducers, which are specific stochastic finite-state networks, have proved very adequate for translation modeling. In this work a speech-to-speech translation system, the EuTRANS system, is presented. The acoustic, language and translation models are finite-state networks that are automatically learnt from training samples. This system was assessed in a series of translation experiments from Spanish to English and from Italian to English in an application involving the interaction (by telephone) of a customer with a receptionist at the front-desk of a hotel. Francisco Casacuberta, David Llorens, Carlos D. Martínez-Hinarejos, Sirko Molau, Francisco Nevado, Hermann Ney, Moisés Pastor, David Picó, Alberto Sanchís, Enrique Vidal 0001, Juan Miguel Vilar |
ICASSP | 10 |
| 2001 | Eutrans: a speech-to-speech translator prototype
Moisés Pastor, Alberto Sanchís, Francisco Casacuberta, Enrique Vidal 0001 |
INTERSPEECH | 4 |
| 2001 | A morphological analyser for machine translation based on finite-state transducersabstractA finite-state, rule-based morphological analyser is presented here, within the framework of machine translation system TAVAL. This morphological analyser introduces specific features which are particularly useful for translation, such as the detection and morphological tagging of word groups that act as a single lexical unit for translation purposes. The case where words in one such group are not strictly contiguous is also covered. A brief description of the Spanish-to-Catalan and Catalan-to-Spanish translation system TAVAL is given in the paper. Alberto Sanchís, David Picó, Joan M. de Val, Ferran Fabregat, Jesús Tomás, Moisés Pastor, Francisco Casacuberta, Enrique Vidal 0001 |
MTSummit | 8 |
| 2001 | Language Simplification through Error-Correcting and Grammatical Inference Techniques
Juan-Carlos Amengual, Alberto Sanchís, Enrique Vidal 0001, José-Miguel Benedí |
Mach. Learn. | 3 |
| 2000 | On the Estimation of Error-Correcting ParametersabstractError-correcting (EC) techniques allow for coping with divergences in pattern strings with regard to their "standard" form as represented by the language /spl Lscr/ accepted by a regular or context-free grammar. There are two main types of EC parsers: minimum-distance and stochastic. The latter apply the maximum likelihood rule: classification into the classes of the strings in /spl Lscr/ that have the greatest probability given the strings representing unknown patterns. Stochastic models are important in pattern recognition if good estimations for their parameters are provided. The problem of parameter estimation has been well studied for stochastic grammars, but this is not the case of EC parameters. This work is aimed at providing solutions to adequately solve it. Juan-Carlos Amengual, Enrique Vidal 0001 |
ICPR | 2 |
| 2000 | On the Use of Normalized Edit Distances and an Efficient k-NN Search Technique (k-AESA) for Fast and Accurate String ClassificationabstractClassification based on nearest neighbours (NN) is a uniformly good approach to many pattern recognition tasks. However, two important aspects need to be taken into account to actually achieve good performance in practice: 1) the metric or dissimilarity measure adopted to compare the considered patterns; and 2) the computational cost incurred by the NN searching operation. As it is shown in this paper, by using adequate techniques to cope with these two issues, the NN-based classification leads to better results than those obtained by other approaches that have been applied to a task of human banded chromosomes classification. Alfons Juan-Císcar, Enrique Vidal 0001 |
ICPR | 2 |
| 2000 | A Cluster-Based Merging Strategy for Nearest Prototype ClassifiersabstractA generalized prototype-based learning scheme founded on hierarchical clustering is proposed. The basic idea is to obtain a condensed nearest neighbor classification rule by replacing a group of prototypes by a representative while approximately keeping their original classification power. The algorithm improves and generalizes previous works by explicitly introducing the concept of cluster and cluster consistency. The proposed scheme also permits a very efficient implementation based on geometric cluster properties. Empirical results demonstrate the merits of the proposed algorithm taking into account the size of the condensed sets of prototypes, the accuracy of the corresponding condensed 1-NN classification rule and the computation time. Ramón Alberto Mollineda, Francesc J. Ferri, Enrique Vidal 0001 |
ICPR | 3 |
| 2000 | Weighting Prototypes. A New Editing Approach
Roberto Paredes, Enrique Vidal 0001 |
ICPR | 2 |
| 2000 | Efficient Use of the Grammar Scale Factor to Classify Incorrect Words in Speech Recognition VerificationabstractThe goal of verification in speech recognition systems is to detect words in the hypothesized sentence that are likely to have been misrecognized. This decision can be based on the persistence of the different words in the output of the speech recognizer when some recognition parameter is varied. To this end, a parameter that proves particularly adequate is the so called grammar scale factor (which balances acoustic and language model scores). The main disadvantage of this method is that it needs to repeat the recognition process many times. In the paper, after formulating it as a statistical pattern classification problem, we show how to speed-up this method, so that less than two average repetitions of the recognition process are enough to achieve essentially the same verification performance as with the many more repetitions needed by the original proposal. Alberto Sanchís, Enrique Vidal 0001, Víctor M. Jiménez |
ICPR | 2 |
| 2000 | The EuTrans Spoken Language Translation System
Juan-Carlos Amengual, M. Asunción Castaño, Antonio Castellanos, Víctor M. Jiménez, David Llorens, Andrés Marzal, Federico Prat, Juan Miguel Vilar, José-Miguel Benedí, Francisco Casacuberta, Moisés Pastor, Enrique Vidal 0001 |
Mach. Transl. | 12 |
| 2000 | A class-dependent weighted dissimilarity measure for nearest neighbor classification problems
Roberto Paredes, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 2 |
| 1999 | Considerations about sample-size sensitivity of a family of edited nearest-neighbor rulesabstractThe edited nearest neighbor classification rules constitute a valid alternative to k-NN rules and other nonparametric classifiers. Experimental results with synthetic and real data from various domains and from different researchers and practitioners suggest that some editing algorithms (especially, the optimal ones) are very sensitive to the total number of prototypes considered. This paper investigates the possibility of modifying optimal editing to cope with a broader range of practical situations. Most previously introduced editing algorithms are presented in a unified form and their different properties (acid not just their asymptotic behavior) are intuitively analyzed. The results show the relative limits in the applicability of different editing algorithms. Francesc J. Ferri, Jesús V. Albert, Enrique Vidal 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 1998 | Fast k-nearest-neighbours searching through extended versions of the approximating and eliminating search algorithm (AESA)abstractThe approximating and eliminating search algorithm (AESA) is probably the technique requiring the fewest distance computations for nearest-neighbour searching in general metric spaces. In this paper we propose direct and refined extensions to the AESA for finding k-nearest-neighbours. Results of a number of experiments involving synthetic data are reported, showing that both extensions, and especially the last one, lead to computational savings similar to that of the original (1-NN) AESA. Alfons Juan-Císcar, Enrique Vidal 0001, Pablo Aibar |
ICPR | 2 |
| 1998 | The extended general spacefilling curves heuristicabstractAn extended general spacefilling curves heuristic (EGSH) is introduced as an extension to the general spacefilling curves heuristic (GSH) proposed by Bartholdi and Platzman (1988). These are generic methods directly applicable to many problems in which data is represented in a multidimensional real vector space. A mapping is established between a region of the multidimensional space and an interval of the real line, and then the problem is solved in one dimension. This becomes quite useful if the problem has an easier, faster or more reliable solution in the real line. The proposed extension allows accurate solutions to many problems not reliably solvable by the original heuristic. A successful application to function approximation is presented. Juan-Carlos Perez-Cortes, Enrique Vidal 0001 |
ICPR | 2 |
| 1998 | Language understanding and subsequential transducer learning
Antonio Castellanos, Enrique Vidal 0001, Miguel Angel Varó, José Oncina |
Comput. Speech Lang. | 2 |
| 1998 | Efficient Error-Correcting Viterbi ParsingabstractThe problem of error-correcting parsing (ECP) using an insertion-deletion-substitution error model and a finite state machine is examined. The Viterbi algorithm can be straightforwardly extended to perform ECP, though the resulting computational complexity can become prohibitive for many applications. We propose three approaches in order to achieve an efficient implementation of Viterbi-like ECP which are compatible with beam search acceleration techniques. Language processing and shape recognition experiments which assess the performance of the proposed algorithms are presented. Juan-Carlos Amengual, Enrique Vidal 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | Finite-state speech-to-speech translationabstractA fully integrated approach to speech input language translation in limited domain applications is presented. The mapping from the input to the output language is modeled in terms of a finite state translation model which is learned from examples of input output sentences of the task considered. This model is tightly integrated with standard acoustic phonetic models of the input language and the resulting global model directly supplies, through Viterbi search, an optimal output language sentence for each input language utterance. Several extensions to this framework, recently developed to cope with the increasing difficulty of translation tasks, are reviewed. Finally, results for a task in the framework of hotel front desk communication, with a vocabulary of about 700 words, are reported. Enrique Vidal 0001 |
ICASSP | 1 |
| 1997 | Speech translation based on automatically trainable finite-state modelsabstractThis paper extends previous work exploring the use of Subsequential Transducers to perform speech-input translation in limited-domain tasks. This is done following an integrated approach in which a Subsequential Transducer replaces the input-language model of a conventional speech recognition system, and is used both as language and translation model. This way, the search for the recognised sentence also produces the corresponding translation. A corpus-based approach is adopted in order to build the required models from training data. Experimental results are presented for the translation task considered in the EUTRANS project: one in the hotel domain with more than 500 words per language and language perplexities near to 10. Juan-Carlos Amengual, José-Miguel Benedí, Klaus Beulen, Francisco Casacuberta, M. Asunción Castaño, Antonio Castellanos, Víctor M. Jiménez, David Llorens, Andrés Marzal, Hermann Ney, Federico Prat, Enrique Vidal 0001, Juan Miguel Vilar |
EUROSPEECH | 12 |
| 1996 | Simplifying language through error-correcting decodingabstractIn many speech processing tasks, most of the sentences generally convey rather simple meanings.In these tasks, the "wordrecognition" problem is much more difficult than the underlying "speech understanding" problem would be.Accordingly we try to develop an adequate framework to focus on a properly defined "understanding" of the sentences rather than "recognizing" the (possibly) superfluous words.This can be seen as closely related with Spontaneous Language Understanding and Disfluence Modeling.In our approach, these problems are placed under the framework of Error-Correcting Decoding (ECD).A complex task is modeled in terms of a basic stochastic grammar, G, and an Error Model, E (taking insertions, substitutions and deletions into account).G should account for the basic (syntactic) structures underlying this task which would convey the semantics.E should account for general vocabulary variations, speech disfluencies, word disappearance, superfluous words, and so on.Each "complex" user sentence, x, will thus be considered as a corrupted version (according to E) of some "simple" sentence y of LG.Recognition can then be seen as an ECD process: given x, find a sentence y of LG with maximum posterior probability.We introduce fast ECD techniques and adequate procedures for simultaneously training G and E and apply these ideas to a simple task with results showing the potential of the proposed approach. Juan-Carlos Amengual, Enrique Vidal 0001, José-Miguel Benedí |
ICSLP | 2 |
| 1996 | Text and speech translation by means of subsequential transducersabstractThe full paper explores the possibility of using Subsequential Transducers (SST), a finite state model, in limited domain translation tasks, both for text and speech input. A distinctive advantage of SSTs is that they can be efficiently learned from sets of input-output examples by means of OSTIA, the Onward Subsequential Transducer Inference Algorithm (Oncina et al. 1993). In this work a technique is proposed to increase the performance of OSTIA by reducing the asynchrony between the input and output sentences, the use of error correcting parsing to increase the robustness of the models is explored, and an integrated architecture for speech input translation by means of SSTs is described. Juan Miguel Vilar, Víctor M. Jiménez, Juan-Carlos Amengual, Antonio Castellanos, David Llorens, Enrique Vidal 0001 |
Nat. Lang. Eng. | 6 |
| 1995 | QWI: a method for improved smoothing in language modellingabstractN-grams have been extensively and successfully used for language modelling in continuous speech recognition tasks. On the other hand, it has been shown that k-testable stochastic languages (k-TS) are strictly equivalent to N-grams. A major problem to be solved when using a language model is the estimation of the probabilities of events not represented in the training corpus, i.e. unseen events. The aim of this work is to improve other well established smoothing procedures by interpolating models with different levels of complexity (quality weighted interpolation-QWI). The effect of QWI was experimentally evaluated over a set of back-off smoothed k-TS language models. These experiments were carried out over several corpora using the test-set perplexity as an evaluation criterion. In all the cases the introduction of QWI resulted in a reduction of the test-set perplexity. Germán Bordel, M. Inés Torres, Enrique Vidal 0001 |
ICASSP | 3 |
| 1995 | Some results with a trainable speech translation and understanding systemabstractThe problems of limited-domain spoken language translation and understanding are considered. A standard continuous speech recognizer is extended for using automatically learnt finite-state transducers as translation models. Understanding is considered as a particular case of translation where the target language is a formal language. From the different approaches compared, the best results are obtained with a fully integrated approach, in which the input language acoustic and lexical models, and (N-gram) language models of input and output languages, are embedded into the learnt transducers. Optimal search through this global network obtains the best translation for a given input acoustic signal. Víctor M. Jiménez, Antonio Castellanos, Enrique Vidal 0001 |
ICASSP | 3 |
| 1995 | Preliminary experiments for automatic speech understanding through simple recurrent networks
M. Asunción Castaño, Enrique Vidal 0001, Francisco Casacuberta |
EUROSPEECH | 2 |
| 1995 | Learning language translation in limited domains using finite-state models: some extensions and improvements
Juan Miguel Vilar, Andrés Marzal, Enrique Vidal 0001 |
EUROSPEECH | 3 |
| 1995 | Fast Computation of Normalized Edit DistancesabstractThe normalized edit distance (NED) between two strings X and Y is defined as the minimum quotient between the sum of weights of the edit operations required to transform X into Y and the length of the editing path corresponding to these operations. An algorithm for computing the NED was introduced by Marzal and Vidal (1993) that exhibits 0(mn/sup 2/) computing complexity, where m and n are the lengths of X and Y. We propose here an algorithm that is observed to require in practice the same 0(mn) computing resources as the conventional unnormalized edit distance algorithm does. The performance of this algorithm is illustrated through computational experiments with synthetic data, as well as with real data consisting of OCR chain-coded strings.> Enrique Vidal 0001, Andrés Marzal, Pablo Aibar |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1994 | Back-off smoothing in a syntactic approach to language modelling
Germán Bordel, M. Inés Torres, Enrique Vidal 0001 |
ICSLP | 3 |
| 1994 | Fast K-means-like clustering in metric spaces
Alfons Juan-Císcar, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 2 |
| 1994 | A new version of the nearest-neighbour approximating and eliminating search algorithm (AESA) with linear preprocessing time and memory requirements
Luisa Micó, José Oncina, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 3 |
| 1994 | Optimum polygonal approximation of digitized curves
Juan-Carlos Perez-Cortes, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 2 |
| 1994 | New formulation and improvements of the nearest-neighbour approximating and eliminating search algorithm (AESA)
Enrique Vidal 0001 |
Pattern Recognit. Lett. | 1 |
| 1993 | Learning direct acoustic-to-semantic mappings through simple recurrent networks
M. Asunción Castaño, Enrique Vidal 0001, Francisco Casacuberta |
EUROSPEECH | 2 |
| 1993 | Efficient enumeration of sentence hypotheses in connected word recognition
Víctor M. Jiménez, Andrés Marzal, Enrique Vidal 0001 |
EUROSPEECH | 3 |
| 1993 | Learning how to understand languageabstractIn this paper we discuss learning paradigms for the problem of understanding spoken language. The basic idea consists in redefining the language understanding problem in terms of translation between a natural language and a formal language that represents the meaning of sentences. Within this framework, with the assumption that input and output sentences can be put into sequential correspondence, understanding can be seen as a problem of sequential transduction. In this case several techniques exist for learning the corresponding transducers, some of which can be properly stated in terms of Hidden Markov modeling (conceptual HMMs). If the sequential assumption does not hold, there are new algorithms that also seem able to solve the learning problem. This view of a language understanding system opens new perspectives in the field of automatic learning of language models. Roberto Pieraccini, Esther Levin, Enrique Vidal 0001 |
EUROSPEECH | 3 |
| 1993 | Learning associations between grammars: a new approach to natural language understanding
Enrique Vidal 0001, Roberto Pieraccini, Esther Levin |
EUROSPEECH | 1 |
| 1993 | Computation of Normalized Edit Distance and ApplicationsabstractGiven two strings X and Y over a finite alphabet, the normalized edit distance between X and Y, d(X,Y) is defined as the minimum of W(P)/L(P), where P is an editing path between X and Y, W(P) is the sum of the weights of the elementary edit operations of P, and L(P) is the number of these operations (length of P). It is shown that in general, d(X,Y) cannot be computed by first obtaining the conventional (unnormalized) edit distance between X and Y and then normalizing this value by the length of the corresponding editing path. In order to compute normalized edit distances, an algorithm that can be implemented to work in O(m*n/sup 2/) time and O(n/sup 2/) memory space is proposed, where m and n are the lengths of the strings under consideration, and m>or=n. Experiments in hand-written digit recognition are presented, revealing that the normalized edit distance consistently provides better results than both unnormalized or post-normalized classical edit distances.> Andrés Marzal, Enrique Vidal 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1993 | Learning Subsequential Transducers for Pattern Recognition Interpretation TasksabstractA formalization of the transducer learning problem and an effective and efficient method for the inductive learning of an important class of transducers, the class of subsequential transducers, are presented. The capabilities of subsequential transductions are illustrated through a series of experiments that also show the high effectiveness of the proposed learning method in obtaining very accurate and compact transducers for the corresponding tasks.> José Oncina, Pedro García 0001, Enrique Vidal 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1992 | Small sample size effects in the use of editing techniquesabstractEditing is recognized as a useful, convenient and even mandatory preprocessing to be carried out prior to using nearest-neighbor (NN) classification techniques. The performance of editing has only been shown for large sets of data but, in practice, one can seldom afford such large sets; either because of the cost of data collection and/or because of computational costs of the adopted editing technique. The authors present evidence showing that editing, and multiedit in particular, dramatically degrade the results as the size of the data set becomes smaller. In conclusion, a modification of the multiedit which can behave better under the small sample assumption is proposed.> Francesc J. Ferri, Enrique Vidal 0001 |
ICPR (2) | 2 |
| 1992 | A N-best sentence hypotheses enumeration algorithm with duration constraints based on the two level algorithmabstractA new N-best sentence hypotheses enumeration algorithm for continuous speech recognition that correctly incorporates duration constraints and which is very efficient in practice is presented. Based on the most general two-level procedure, this algorithm provides greater flexibility in acoustic modeling than previous approaches and is very well suited for large vocabulary applications.> Andrés Marzal, Enrique Vidal 0001 |
ICPR (3) | 2 |
| 1992 | An algorithm for finding nearest neighbours in constant average time with a linear space complexityabstractGiven a set of n points or 'prototypes' and another point or 'test sample'. The authors present an algorithm that finds a prototype that is a nearest neighbour of the test sample, by computing only a constant number of distances on the average. This is achieved through a preprocessing procedure that computes only a number of distances and uses an amount of memory that grows lineally with n. The algorithm is an improvement of the previously introduced AESA algorithm and, as such, does not assume the data to be structured into a vector space, making only use of the metric properties of the given distance.> Luisa Micó, José Oncina, Enrique Vidal 0001 |
ICPR (2) | 3 |
| 1992 | Transducer learning in pattern recognitionabstract'Interpretation' is a general and interesting pattern recognition framework in which a system is considered to input object representations, and output the corresponding interpretations in terms of 'semantic messages' specifying the actions to be carried out as system's responses. From the syntactic pattern recognition viewpoint, interpretation reduces to formal transduction. The authors propose an efficient and effective algorithm to automatically infer a finite state transducer from a training set of input-output examples of the interpretation problem considered. The proposed algorithm has been shown to identify an important class of transductions known as 'subsequential transductions.' Experimental results are presented showing the performance and capabilities of the proposed method.> José Oncina, Pedro García 0001, Enrique Vidal 0001 |
ICPR (2) | 3 |
| 1992 | An algorithm for the optimum piecewise linear approximation of digitized curvesabstractA dynamic programming algorithm for piecewise linear approximation of a digitized curve is presented. This approximation yields a global minimization of the total error between the approximating straight-line segments and the original points. Many error measures can be used with this algorithm although a lower computational cost can be achieved if error measures that can be computed incrementally are used. Two incremental formulations of classical error measures, along with the results of a sample experiment, are also presented.> Juan-Carlos Perez-Cortes, Enrique Vidal 0001 |
ICPR (3) | 2 |
| 1992 | Font-independent mixed-size digit recognition through error-correcting grammatical inference (ECGI)abstractThe application of the structural learning technique known as error correcting grammatical inference to planar shape recognition is discussed and illustrated with a non-trivial printed digit recognition task. Experimental results are presented and compared with those of other more conventional (non-structural) techniques, showing the new technique to provide significantly improved performance.> Enrique Vidal 0001, Hector Rulot Segovia, José Miguel Valiente González, Gabriela Andreu |
ICPR (2) | 1 |
| 1992 | Problems and algorithms in optimal linguistic decoding: a unified formulation
Pablo Aibar, Andrés Marzal, Enrique Vidal 0001, Francisco Casacuberta |
ICSLP | 3 |
| 1992 | Colour image segmentation and labeling through multiedit-condensing
Francesc J. Ferri, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 2 |
| 1992 | Learning language models through the ECGI method
Natividad Prieto, Enrique Vidal 0001 |
Speech Commun. | 2 |
| 1991 | Automatic learning of structural language modelsabstractA novel approach to adaptive language acquisition is proposed. This approach is based on the pattern recognition framework of interpretation, and models the acoustic, lexical, syntactic, and semantic constraints of a given continuous speech task through the concept of sequential finite-state transduction. In order to automatically learn the required finite-state models from training data, a grammatical inference procedure is applied which directly uses a previously introduced error-correcting grammatical inference algorithm. Experiments with relatively simple but nontrivial continuous speech understanding tasks are presented, with results showing both the viability and appropriateness of the proposed approach.> Natividad Prieto, Enrique Vidal 0001 |
ICASSP | 2 |
| 1991 | Learning language models through the ECGI method
Natividad Prieto, Enrique Vidal 0001 |
EUROSPEECH | 2 |
| 1990 | Learning the structure of HMM's through grammatical inference techniquesabstractA technique is described in which all the components of a hidden Markov model are learnt from training speech data. The structure or topology of the model (i.e. the number of states and the actual transitions) is obtained by means of an error-correcting grammatical inference algorithm (ECGI). This structure is then reduced by using an appropriate state pruning criterion. The statistical parameters that are associated with the obtained topology are estimated from the same training data by means of the standard Baum-Welch algorithm. Experimental results showing the applicability of this technique to speech recognition are presented.> Francisco Casacuberta, Enrique Vidal 0001, B. Mas, Hector Rulot Segovia |
ICASSP | 2 |
| 1990 | On the Use of the Morphic Generator Grammatical Inference (MGGI) Methodology in Automatic speech RecognitionabstractRecently, a new methodology, referred to as “Morphic Generator Grammatical Inference” (MGGI), has been introduced as a step towards a general methodology for the inference of regular languages. In this paper we consider the application of this methodology to a real problem of automatic speech recognition, thus allowing (and also requiring) the proposed problem to be properly formulated within the canonical framework of syntactic pattern recognition. The results show both the viability and appropriateness of the application of MGGI to the problem considered. Pedro García 0001, Encarna Segarra, Enrique Vidal 0001, Isabel Galiano |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 1990 | Inference of k-Testable Languages in the Strict Sense and Application to Syntactic Pattern RecognitionabstractThe inductive inference of the class of k-testable languages in the strict sense (k-TSSL) is considered. A k-TSSL is essentially defined by a finite set of substrings of length k that are permitted to appear in the strings of the language. Given a positive sample R of strings of an unknown language, a deterministic finite-state automation that recognizes the smallest k-TSSL containing R is obtained. The inferred automation is shown to have a number of transitions bounded by O(m) where m is the number of substrings defining this k-TSSL, and the inference algorithm works in O(kn log m) where n is the sum of the lengths of all the strings in R. The proposed methods are illustrated through syntactic pattern recognition experiments in which a number of strings generated by ten given (source) non-k-TSSL grammars are used to infer ten k-TSSL stochastic automata, which are further used to classify new strings generated by the same source grammars. The results of these experiments are consistent with the theory and show the ability of (stochastic) k-TSSLs to approach other classes of regular languages.> Pedro García 0001, Enrique Vidal 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1989 | Learning accurate finite-state structural models of words through the ECGI algorithmabstractThe ECGI (error correcting grammatical inference) algorithm is a recently introduced automatic learning method. The authors present improvements to the ECGI and the corresponding experiments to test these improvements. The results show the ability of the ECGI to obtain structural models of (isolated) words which enable a very high speaker-independent recognition rate to be achieved.> Hector Rulot Segovia, Natividad Prieto, Enrique Vidal 0001 |
ICASSP | 3 |
| 1989 | A Hybrid Framework Combining Structural and Decision-Theoretic Pattern Recognition and ApplicationsabstractA new framework is introduced which allows the formulation of difficult structural classification tasks in terms of decision-theoretic-based pattern recognition. It is based on extending the classical formulation of generalized linear discriminant functions so as to permit each given object to have a different vector representation in each class. The proposed extension properly accounts for the corresponding extension of the classical learning techniques of linear discriminant functions in a way such that the convergence of the extended techniques can still be proved. The proposed framework can be considered as a hybrid methodology in which both structural and decision-theoretic pattern recognition are integrated. Furthermore, it can be considered as a means to achieve convenient tradeoffs between the inductive and deductive ways of knowledge acquisition, which can result in rendering tractable the possibly hard original inductive learning problem associated with the given task. The proposed framework and methods are illustrated through their use in two difficult structural classification tasks, showing both the appropriateness and the capability of these methods to obtain useful results. Enrique Vidal 0001, Francisco Casacuberta |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1988 | On the verification of triangle inequality by dynamic time-warping dissimilarity measures
Enrique Vidal 0001, Francisco Casacuberta, José-Miguel Benedí, Maria-José Lloret, Hector Rulot Segovia |
Speech Commun. | 1 |
| 1988 | Fast speaker-independent DTW recognition of isolated words using a metric-space search algorithm (AESA)
Enrique Vidal 0001, Maria-José Lloret |
Speech Commun. | 1 |
| 1987 | Local Languages, the Succesor Method, and a Step Towards a General Methodology for the Inference of Regular GrammarsabstractA methodology is proposed for the inference of regular grammars from positive samples of their languages. It is mainly based on the generative mechanism associated with local languages, which allows us to obtain arbitrary regular languages by applying morphic operators to local languages. The actual inference procedure of this methodology consists of obtaining a local language associated with the given positive sample. This procedure, which is very simple, is always the same, regardless of the problem considered, while the task-dependent features that are desired for the inferred languages, are specified through the definition of certain task-appropriate symbol renaming functions (morphisms). Pedro García 0001, Enrique Vidal 0001, Francisco Casacuberta |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1986 | An algorithm for finding nearest neighbours in (approximately) constant average time
Enrique Vidal 0001 |
Pattern Recognit. Lett. | 1 |
| 1985 | Is the DTW "distance" really a metric? An algorithm reducing the number of DTW comparisons in isolated word recognition
Enrique Vidal 0001, Francisco Casacuberta, Hector Rulot Segovia |
Speech Commun. | 1 |