VLDB 2026 Research / reviewers in the wild / expert
José Andrés
dblp:300/1111
· DBLP profile ↗
7ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-5541-8231ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The PARES Database: Information Extraction over Historical Parish RecordsabstractAbstract Historical census records convey information that is key to perform genealogical research and demographic studies. Given the large number of documents of this type that exist, it is crucial to research methods that allow the automatic extraction of information from this type of document. In this work, we present a new corpus of this kind, comprising 535 historical census tables from French archives. Alongside this dataset, we have assessed three different baseline methods for information extraction. The first two methods employ a traditional sequential approach, where table rows are detected before extracting information. The third baseline uses an end-to-end model that directly extracts information from the table images without prior row detection. Our results demonstrate the effectiveness of all three baselines in tackling the information extraction task. José Andrés, Casey Wall, Solène Tarride, Mickaël Coustaty, Alejandro H. Toselli, Enrique Vidal 0001 |
Int. J. Document Anal. Recognit. | 1 |
| 2024 | Mining and Analyzing Statistical Information from Untranscribed Form Images
José Andrés, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR (5) | 1 |
| 2023 | Search for Hyphenated Words in Probabilistic Indices: A Machine Learning Approach
José Andrés, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR (1) | 1 |
| 2023 | Processing a large collection of historical tabular imagesabstractProcessing automatically historical document images to allow the search of textual information requires the preparation of ground-truth data for training and evaluation. This process is an expensive and arduous task, especially when the historical document images contain specialized vocabulary and/or tabular information. In the latter case, relevant decisions have to be taken to annotate the tabular parts. This paper presents a complex collection of historical document images and the resulting database, which is called HisClima. In this database, half of the images are in tabular format and half as running text. Both types of images contain pre-printed and handwritten text. The textual information is plenty of abbreviations and specific vocabulary related to weather conditions and old ships. This database can be used to research technologies related to historical document image processing and analysis, both for tabular and running text recognition. Baseline results are presented for Document Layout Analysis, Text Recognition, and Probabilistic Indexing. Although these results are good, there is still room for improvement and some indications are provided in this direction. Emilio Granell, Verónica Romero 0001, José Ramón Prieto, José Andrés, Lorenzo Quirós, Joan-Andreu Sánchez, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 4 |
| 2023 | Information extraction in handwritten historical logbooksabstractDocument Image Understanding is a demanding Pattern Recognition problem that requires complex recognition models. This problem is even more difficult for document images with complicated layouts like tables, where the reading order is often intrinsically ambiguous, and consequently, the context is generally ambiguous as well. In this paper, we compare two machine learning approaches for extracting information in pre-printed historical tables with handwritten information. We analyze the performance of each approach at each step of the extraction process over different corpora, up to a realistic scenario where documents with different table layouts written by different hands are used. The results are good in general and show that a model based on Multilayer Perceptrons yields better results on more homogeneous documents, while another model based on Graph Neural Networks generalizes better on heterogeneous corpora. José Ramón Prieto, José Andrés, Emilio Granell, Joan-Andreu Sánchez, Enrique Vidal 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | Information Extraction from Handwritten Tables in Historical Documents
José Andrés, José Ramón Prieto, Emilio Granell, Verónica Romero 0001, Joan-Andreu Sánchez, Enrique Vidal 0001 |
DAS | 1 |
| 2022 | Approximate Search for Keywords in Handwritten Text Images
José Andrés, Alejandro H. Toselli, Enrique Vidal 0001 |
DAS | 1 |