EDBT 2026 Demo / reviewers in the wild / expert
Moisés Pastor
dblp:27/573 · also Moisés Pastor-i-Gadea
· DBLP profile ↗
7ranked-venue papers in the field
1as first author
5since 2021 · last 2025
0000-0002-1833-7440ORCID · reported
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 5 (1 first)Information Retrieval & Web Search · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Lightweight Named Entity Recognition in Handwritten Documents by Predicting Pyramidal Histograms of CharactersabstractNamed Entity Recogniton (NER) consists of tagging parts of an unstructured text containing particular semantic information. When applied to handwritten documents, it is possible to do it as a two-step approach in which Handwritten Text Recognition (HTR) is performed prior to tagging the automatic transcription. However, it is also possible to do both tasks simultaneously by using an HTR model that learns to output the transcription and the tagging symbols. In this paper, we focus on improving the one-step approach by introducing the auxiliary task of predicting Pyramidal Histograms of Characters (PHOC) in a Convolutional Recurrent Neural Network (CRNN) model. Moreover, given the recent rise of models that digest large amounts of data, we also study the usage of synthetic data to pretrain the proposed architecture. Our experiments show that by pretraining the PHOC-based architecture on synthetic data substantial improvements can be made in both transcription and tagging quality without compromising the computational cost of the decoding step. The resulting model matches the NER performance of the state-of-the-art while keeping its lightweight nature. David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor |
DocEng | 4 |
| 2024 | Reading Order Independent Metrics for Information Extraction in Handwritten Documents
David Villanova-Aparisi, Solène Tarride, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Christopher Kermorvant, Moisés Pastor |
ICDAR (2) | 6 |
| 2023 | Consistent Nested Named Entity Recognition in Handwritten Documents via Lattice Rescoring
David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor |
ICDAR (1) | 4 |
| 2023 | Evaluation of Different Tagging Schemes for Named Entity Recognition in Handwritten Documents
David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor |
ICDAR (3) | 4 |
| 2022 | Evaluation of Named Entity Recognition in Handwritten Documents
David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor |
DAS | 4 |
| 2017 | Baseline Detection on Arabic Handwritten DocumentsabstractDocument processing comprises different steps depending on the nature of the documents. For text documents, specially for handwritten documents, transcription of their contents is one of the main tasks. Handwritten Text Recognition (HTR) is the process of automatically obtaining the transcription of the content of a handwritten text document. In document processing, the basic unit for the acquisition process is the page image, whilst line image is the basic form for the HTR process. This is a bottle-neck which is holding back the massive industrial document processing. Baseline detection can be used not only to segment page images into line images but also for many other document processing steps. Baseline detection problem can be formulated as a clustering problem over a set of interest points. In this work, we study the use of an automatic baseline detection technique, based on interest point clustering, in Arabic handwritten documents. The experiments reveal that this technique provides promising results for this task. Ahmed Fawzi, Moisés Pastor, Carlos D. Martínez-Hinarejos |
DocEng | 2 |
| 2005 | Writing Speed Normalization for On-Line Handwritten Text RecognitionabstractPen-based interfaces aim at improving the man-machine interaction of many portable systems. While statistical models can be used to learn pen position sequences, they suffer from the huge variability exhibited by the speed of writing. To improve performance, invariance to the writing speed is needed. Trace segmentation is a technique that can be used to normalize the writing speed. This method is controlled by a parameter called resampling distance. A study of the resampling distance is presented here, along with another approximation to the writing speed normalization called "derivatives normalization". The improvement using trace segmentation was 193% relative to the baseline, whilst the improvement using derivatives normalization was 47.3% relative. Moisés Pastor, Alejandro H. Toselli, Enrique Vidal 0001 |
ICDAR | 1 |