Carlos D. Martínez-Hinarejos

dblp:04/1927 · also Carlos David Martínez, Carlos David Martínez-Hinarejos · DBLP profile ↗
← Back
11ranked-venue papers in the field
1as first author
6since 2021 · last 2025
0000-0002-6139-2891ORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 7 (1 first)Information Retrieval & Web Search · 4
YearPublicationVenuePosition
2025 Improving Lightweight Named Entity Recognition in Handwritten Documents by Predicting Pyramidal Histograms of Characters
abstract
Named Entity Recogniton (NER) consists of tagging parts of an unstructured text containing particular semantic information. When applied to handwritten documents, it is possible to do it as a two-step approach in which Handwritten Text Recognition (HTR) is performed prior to tagging the automatic transcription. However, it is also possible to do both tasks simultaneously by using an HTR model that learns to output the transcription and the tagging symbols. In this paper, we focus on improving the one-step approach by introducing the auxiliary task of predicting Pyramidal Histograms of Characters (PHOC) in a Convolutional Recurrent Neural Network (CRNN) model. Moreover, given the recent rise of models that digest large amounts of data, we also study the usage of synthetic data to pretrain the proposed architecture. Our experiments show that by pretraining the PHOC-based architecture on synthetic data substantial improvements can be made in both transcription and tagging quality without compromising the computational cost of the decoding step. The resulting model matches the NER performance of the state-of-the-art while keeping its lightweight nature.
David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor
DocEng2
2024 Reading Between the Frames: Multi-modal Depression Detection in Videos from Non-verbal Cues
David Gimeno-Gómez, Ana-Maria Bucur, Adrian Cosma, Carlos D. Martínez-Hinarejos, Paolo Rosso
ECIR (1)4
2024 Reading Order Independent Metrics for Information Extraction in Handwritten Documents
David Villanova-Aparisi, Solène Tarride, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Christopher Kermorvant, Moisés Pastor
ICDAR (2)3
2023 Consistent Nested Named Entity Recognition in Handwritten Documents via Lattice Rescoring
David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor
ICDAR (1)2
2023 Evaluation of Different Tagging Schemes for Named Entity Recognition in Handwritten Documents
David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor
ICDAR (3)2
2022 Evaluation of Named Entity Recognition in Handwritten Documents
David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor
DAS2
2018 Comparing Different Feedback Modalities in Assisted Transcription of Manuscripts
abstract
Transcription of handwritten text can be speed-up by using off-line Handwritten Text Recognition techniques, that allow the obtention of an initial draft transcription of an image with handwritten text. However, this draft transcription usually contains errors that must be amended by the transcriber by providing a feedback signal. The usual approach is post-edition, where each error is corrected without modifying the rest of the current transcription. A more sophisticated approach can employ the current modification to provide a new whole transcription, hopefully with less errors. Apart from that, feedback can be provided in different modalities: keyboard input, on-line handwritten text, or speech. Each of these modalities presents different features with respect to ambiguity, derived errors, and final transcription time. In this work we study how the different modalities behave in the assisted transcription of a historical handwritten text document in Spanish and we evaluate their transcription productivity.
Carlos D. Martínez-Hinarejos, Emilio Granell, Verónica Romero 0001
DAS1
2017 Baseline Detection on Arabic Handwritten Documents
abstract
Document processing comprises different steps depending on the nature of the documents. For text documents, specially for handwritten documents, transcription of their contents is one of the main tasks. Handwritten Text Recognition (HTR) is the process of automatically obtaining the transcription of the content of a handwritten text document. In document processing, the basic unit for the acquisition process is the page image, whilst line image is the basic form for the HTR process. This is a bottle-neck which is holding back the massive industrial document processing. Baseline detection can be used not only to segment page images into line images but also for many other document processing steps. Baseline detection problem can be formulated as a clustering problem over a set of interest points. In this work, we study the use of an automatic baseline detection technique, based on interest point clustering, in Arabic handwritten documents. The experiments reveal that this technique provides promising results for this task.
Ahmed Fawzi, Moisés Pastor, Carlos D. Martínez-Hinarejos
DocEng3
2016 An Interactive Approach with Off-Line and On-Line Handwritten Text Recognition Combination for Transcribing Historical Documents
abstract
Automatic transcription of historical documents is becoming an important research topic, specially because of the increasing number of digitised historical documents that libraries and archives are publishing. However, state-of-the-art handwritten text recognition systems are far from being perfect. Therefore, to have perfect transcriptions, human expert revision is required to really produce a transcription of standard quality. In this context, an interactive assistive scenario, where the automatic system and the human transcriber cooperate to generate the perfect transcription, would allow for a more effective approach. In this paper we present a multimodal interactive transcription system where user feedback is provided by means of touchscreen pen strokes, traditional keyboard and mouse operations. The combination of both the main and the feedback data stream is based on the use of Confusion Networks derived from the output of the on-line and off-line handwritten text recognition systems. The use of the proposed combination help to optimise overall performance and usability.
Emilio Granell, Verónica Romero 0001, Carlos D. Martínez-Hinarejos
DAS3
2016 A Multimodal Crowdsourcing Framework for Transcribing Historical Handwritten Documents
abstract
Transcription of handwritten historical documents is one of the main topics in document analysis systems, due to cultural reasons. State-of-the-art handwritten text recognition systems allow to speed up the transcription task. Currently, this automatic transcription is far from perfect, and human expert revision is required in order to obtain the actual transcription. In this context, crowdsourcing emerged as a powerful tool for massive transcription at a relatively low cost, since the supervision effort of professional transcribers may be dramatically reduced. However, current transcription crowdsourcing platforms are mainly limited to the use of non-mobile devices, since the use of keyboards in mobile devices is not friendly enough for most users. This work presents the alternative of using speech dictation of handwritten text lines as transcription source in a crowdsourcing platform. The experiments explore how an initial handwritten text recognition hypothesis can be improved by using the contribution of speech recognition from several speakers, providing as a final result a better hypothesis to be amended by a professional transcriber with less effort.
Emilio Granell, Carlos D. Martínez-Hinarejos
DocEng2
2015 Combining handwriting and speech recognition for transcribing historical handwritten documents
abstract
Transcription of historical documents is an interesting task for libraries in order to make available their funds. In the lasts years, the use of Handwritten Text Recognition allowed paleographs to speed up the manual transcription process, since they are able to correct on a draft transcription. Another alternative is obtaining the draft transcription by dictating the contents to an Automatic Speech Recognition system. When both sources (image and speech) are available, a multimodal combination is possible, and an iterative process can be used in order to refine the final hypothesis. In this work, a multimodal combination based on confusion networks is presented. Results on two different sets of data, with different difficulty level, show that the proposed technique provides similar or better draft transcriptions than a previously proposed approach, allowing for a faster transcription process.
Emilio Granell, Carlos D. Martínez-Hinarejos
ICDAR2