Jorge Calvo-Zaragoza

dblp:136/2163 · DBLP profile ↗
← Back
73ranked-venue papers
20as first author
40since 2021 · last 2026
0000-0003-3183-2232ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 62 · 16 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 11 · 3 first-author · 8 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music
abstract
Abstract Optical Music Recognition (OMR) has made significant progress since its inception, with various approaches now capable of accurately transcribing music scores into digital formats. Despite these advancements, most so-called end-to-end OMR approaches still rely on multi-stage processing pipelines for transcribing full-page score images, which entails challenges such as the need for dedicated layout analysis and specific annotated data, thereby limiting the general applicability of such methods. In this paper, we present the first truly end-to-end approach for page-level OMR in complex layouts. Our system, which combines convolutional layers with autoregressive Transformers, processes an entire music score page and outputs a complete transcription in a music encoding format. This is made possible by both the architecture and the training procedure, which utilizes curriculum learning through incremental synthetic data generation. We evaluate the proposed system using pianoform corpora, which is one of the most complex sources in the OMR literature. This evaluation is conducted first in a controlled scenario with synthetic data, and subsequently against two real-world corpora of varying conditions. Our approach is compared with leading commercial OMR software. The results demonstrate that our system not only successfully transcribes full-page music scores but also outperforms the commercial tool in both zero-shot settings and after fine-tuning with the target domain, representing a significant contribution to the field of OMR.
Antonio Ríos-Vila, Jorge Calvo-Zaragoza, David Rizo, Thierry Paquet
Int. J. Comput. Vis.2
2026 Exploring federated learning in optical music recognition
abstract
Optical Music Recognition (OMR) technology plays a crucial role in the preservation of cultural heritage by automating the digitization of music documents, enabling their storage in symbolic formats and their subsequent analysis through digital tools. Progress in this field is, however, constrained by limited data availability: access to or distribution of music collections is frequently restricted due to legal or ownership barriers. This work investigates a potential mitigation of this issue through Federated Learning (FL) strategies, which enable decentralized training and eliminate the need to release restricted collections. Our methodology assumes a setting in which client nodes operate on small, heterogeneous corpora, while evaluation is conducted on established OMR benchmark collections. We examine widely used FL aggregation techniques, such as FedAvg, FedProx, and SCAFFOLD; as well as modern methodologies, such as FedKT. In addition, we introduce two modules specifically designed for OMR: FedClassPrior , which integrates class-prior information to improve symbol balance, and FedNGram , which supports decentralized language modeling to exploit notational regularities during decoding. Experimental results show consistent accuracy improvements when using FL compared to purely local training. Furthermore, the proposed modules substantially narrow the performance gap between standard FL and the ideal scenario in which centralized training is feasible.
Eric Ayllon, Beatriz Serrano Sánchez, Jorge Calvo-Zaragoza
Neurocomputing3
2026 Handwritten Text Recognition: A Survey
abstract
Handwritten Text Recognition (HTR) has become an essential field within pattern recognition and machine learning, with applications spanning historical document preservation to modern data entry and accessibility solutions. The complexity of HTR lies in the high variability of handwriting, which makes it challenging to develop robust recognition systems. This survey examines the evolution of HTR models, tracing their progression from early heuristic-based approaches to contemporary state-of-the-art neural models, which leverage deep learning techniques. The scope of the field has also expanded, with models initially capable of recognizing only word-level content progressing to recent end-to-end document-level approaches. Our paper categorizes existing work into two primary levels of recognition: (1) up to line-level, encompassing word and line recognition, and (2) beyond line-level, addressing paragraph- and document-level challenges. We provide a unified framework that examines research methodologies, recent advances in benchmarking, key datasets in the field, and a discussion of the results reported in the literature. Finally, we identify pressing research challenges and outline promising future directions, aiming to equip researchers and practitioners with a roadmap for advancing the field.
Carlos Garrido-Munoz, Antonio Ríos-Vila, Jorge Calvo-Zaragoza
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Aligned music notation and lyrics transcription
abstract
The digitization of vocal music scores presents unique challenges that go beyond traditional Optical Music Recognition and Optical Character Recognition, as it necessitates preserving the critical alignment between music notation and lyrics. This alignment is essential for proper interpretation and processing in practical applications. This paper introduces and formalizes, for the first time, the Aligned Music Notation and Lyrics Transcription (AMNLT) challenge, which addresses the complete transcription of vocal scores by jointly considering music symbols, lyrics, and their synchronization. We analyze different approaches to address this challenge, ranging from traditional divide-and-conquer methods that handle music and lyrics separately, to novel end-to-end solutions including direct transcription, unfolding mechanisms, and language modeling. To evaluate these methods, we introduce four datasets of Gregorian chants, comprising both real and synthetic sources, along with custom metrics specifically designed to assess both transcription and alignment accuracy. Our experimental results demonstrate that end-to-end approaches generally outperform heuristic methods in the alignment challenge, with language models showing particular promise in scenarios where sufficient training data is available. This work establishes the first comprehensive framework for AMNLT, providing both theoretical foundations and practical solutions for preserving and digitizing vocal music heritage.
Eliseo Fuentes-Martínez, Antonio Ríos-Vila, Juan C. Martinez-Sevilla, David Rizo, Jorge Calvo-Zaragoza
Pattern Recognit.5
2026 TriScore: Aligning audio, symbolic scores, and sheet music images in a shared embedding space
abstract
Multimodal representation learning has attracted increasing attention in Music Information Retrieval (MIR), yet score-based multimodality is still constrained by the lack of datasets that jointly provide audio, notation-level symbolic scores, and sheet music images with reliable alignment and balanced coverage. To address this gap, we introduce TriScore , a compact tri-modal collection of excerpt-aligned triplets annotated with composer , instrument , and piece title . We further propose a scalable fusion approach that keeps modality-specific pretrained encoders frozen and learns a shared shallow projection that maps heterogeneous embeddings into a common latent space. Extensive experiments study the impact of normalization, class-imbalance mitigation, projection dimensionality, and k NN neighborhood size. The learned space achieves strong excerpt-level classification performance across modalities, reaching 90.38 F1 for composer and 95.88 F1 for instrument on average, and attains 88.29/92.43 Top-5/Top-10 accuracy for piece retrieval . In-depth analyses show that the projection yields consistent global semantic alignment while preserving local modality-specific submanifolds: neighborhoods are dominated by same-modality samples in the full search space, yet cross-modal exclusion results remain far above distinct-modality pretrained baselines, indicating non-trivial cross-modal consistency. Overall, TriScore and the proposed projection framework establish a controlled benchmark and strong baselines for tri-modal, score-based MIR.
Antonio Hidalgo-Centeno, Eliseo Fuentes-Martínez, Jorge Calvo-Zaragoza, Antonio Javier Gallego 0001
Pattern Recognit.3
2026 Insights into imbalance-aware Multilabel Prototype Generation mechanisms for k-Nearest Neighbor classification in noisy scenarios
abstract
Prototype Generation (PG) techniques enhance the efficiency of the k -Nearest Neighbor ( k NN) classifier by condensing datasets through the use of specific rules. More precisely, these strategies work on the premise of merging the elements in the reference data collection to generate an alternative and more compact data assortment that substitutes the former one without remarkably affecting the recognition performance. Nevertheless, despite being widely studied in multiclass scenarios, PG is still underexplored in multilabel contexts, leading to limitations, notably in the handling of label imbalance and noise. In this regard, this work introduces a reduction framework that allows for multilabel PG methods to handle these challenges of label imbalance and noise. The proposed mechanisms comprise a selection strategy that exclusively preserves noise-free samples in the process, a mechanism to avoid severely imbalanced samples from being inadequately processed, and two new merging policies for the PG methods to generate novel samples. These enhancements are considered along with three established multilabel PG methods: Multilabel Reduction through Homogeneous Clustering, Multilabel Chen, and Multilabel Reduction through Space Partitioning. Evaluations are conducted using three k NN-based multilabel classifiers and 12 diverse datasets with different levels of label imbalance. We additionally study the performance with varying values of k under different label-noise scenarios. The results are assessed through statistical tests and indicate that our proposals outperform the original methods that disregard label imbalance, even in the presence of noise, thus validating these approaches and fostering further research in the field.
Jose J. Valero-Mas, Carlos Peñarrubia, Francisco J. Castellanos 0001, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
Pattern Recognit.5
2025 On the Generalization of Handwritten Text Recognition Models
abstract
Recent advances in Handwritten Text Recognition (HTR) have led to significant reductions in transcription errors on standard benchmarks under the i.i.d. assumption, thus focusing on minimizing in-distribution (ID) errors. However, this assumption does not hold in real-world applications, which has motivated HTR research to explore Transfer Learning and Domain Adaptation techniques. In this work, we investigate the unaddressed limitations of HTR models in generalizing to out-of-distribution (OOD) data. We adopt the challenging setting of Domain Generalization, where models are expected to generalize to OOD data without any prior access. To this end, we analyze 336 OOD cases from eight state-of-the-art HTR models across seven widely used datasets, spanning five languages. Additionally, we study how HTR models leverage synthetic data to generalize. We reveal that the most significant factor for generalization lies in the textual divergence between domains, followed by visual divergence. We demonstrate that the error of HTR models in OOD scenarios can be reliably estimated, with discrepancies falling below 10 points in 70% of cases. We identify the underlying limitations of HTR models, laying the foundation for future research to address this challenge. Code is available at github.com/carlos10garrido/HTR-OOD.
Carlos Garrido-Munoz, Jorge Calvo-Zaragoza
CVPR2
2025 TextSAM-LoRA: Efficient Fine-Tuning of Segment Anything Model for Text Detection with Low-Rank Adaptation
Carlos de la Fuente, Adrián Sánchez-Hernández, Jorge Calvo-Zaragoza
ICDAR (5)3
2025 On the use of synthetic data for body detection in maritime search and rescue operations
abstract
Time is a critical factor in maritime Search And Rescue (SAR) missions, during which promptly locating survivors is paramount. Unmanned Aerial Vehicles (UAVs) are a useful tool with which to increase the success rate by rapidly identifying targets. While this task can be performed using other means, such as helicopters, the cost-effectiveness of UAVs makes them an effective choice. Moreover, these vehicles allow the easy integration of automatic systems that can be used to assist in the search process. Despite the impact of artificial intelligence on autonomous technology, there are still two major drawbacks to overcome: the need for sufficient training data to cover the wide variability of scenes that a UAV may encounter and the strong dependence of the generated models on the specific characteristics of the training samples. In this work, we address these challenges by proposing a novel approach that leverages computer-generated synthetic data alongside novel modifications to the You Only Look Once (YOLO) architecture that enhance its robustness, adaptability to new environments, and accuracy in detecting small targets. Our method introduces a new patch-sample extraction technique and task-specific data augmentation, ensuring robust performance across diverse weather conditions. The results demonstrate our proposal’s superiority, showing an average 28% relative improvement in mean Average Precision (mAP) over the best-performing state-of-the-art baseline under training conditions with sufficient real data, and a remarkable 218% improvement when real data is limited. The proposal also presents a favorable balance between efficiency, effectiveness, and resource requirements. • Small target detection architecture for rapid and precise maritime SAR missions. • Fusion of real and synthetic data simulating real imagery, boosting model robustness. • Transforming data to simulate weather conditions like rain, fog, and sunsets. • In-depth hyperparameter analysis, effect of data scarcity, and data combination. • Proven efficiency in variable weather scenarios, comparison with state of the art.
Juan Pedro Martinez-Esteso, Francisco J. Castellanos 0001, Adrian Rosello, Jorge Calvo-Zaragoza, Antonio Javier Gallego 0001
Eng. Appl. Artif. Intell.4
2025 Self-Supervised Learning for Text Recognition: A Critical Survey
abstract
Abstract Text Recognition (TR) refers to the research area that focuses on retrieving textual information from images, a topic that has seen significant advancements in the last decade due to the use of Deep Neural Networks (DNN). However, these solutions often necessitate vast amounts of manually labeled or synthetic data. Addressing this challenge, Self-Supervised Learning (SSL) has gained attention by utilizing large datasets of unlabeled data to train DNN, thereby generating meaningful and robust representations. Although SSL was initially overlooked in TR because of its unique characteristics, recent years have witnessed a surge in the development of SSL methods specifically for this field. This rapid development, however, has led to many methods being explored independently, without taking previous efforts in methodology or comparison into account, thereby hindering progress in the field of research. This paper, therefore, seeks to consolidate the use of SSL in the field of TR, offering a critical and comprehensive overview of the current state of the art. We will review and analyze the existing methods, compare their results, and highlight inconsistencies in the current literature. This thorough analysis aims to provide general insights into the field, propose standardizations, identify new research directions, and foster its proper development.
Carlos Peñarrubia, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
Int. J. Comput. Vis.3
2025 Spatial context-based Self-Supervised Learning for Handwritten Text Recognition
abstract
Handwritten Text Recognition (HTR) is a relevant problem in computer vision, and implies unique challenges owing to its inherent variability and the rich contextualization required for its interpretation. Despite the success of Self-Supervised Learning (SSL) in computer vision, its application to HTR has been rather scattered, leaving key SSL methodologies unexplored. This work specifically focuses on Spatial Context-based SSL. We investigate how this family of approaches can be adapted and optimized for HTR and propose new workflows that leverage the unique features of handwritten text. Our experiments demonstrate that the methods considered lead to advancements in the state-of-the-art of SSL for HTR in a number of benchmark cases. • We investigate the performance of existing spatial context-based SSL methods for HTR. • We propose new methods within this family of SSL. • We advance the state of the art in SSL for HTR. • We highlight that HTR contains rich spatial information due to font and stroke style.
Carlos Peñarrubia, Carlos Garrido-Munoz, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
Pattern Recognit. Lett.4
2024 Contrastive Self-Supervised Learning for Optical Music Recognition
Carlos Peñarrubia, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
DAS3
2024 SDFR: Synthetic Data for Face Recognition Competition
abstract
Large-scale face recognition datasets are collected by crawling the Internet and without individuals' consent, raising legal, ethical, and privacy concerns. With the recent advances in generative models, recently several works proposed generating synthetic face recognition datasets to mitigate concerns in web-crawled face recognition datasets. This paper presents the summary of the Synthetic Data for Face Recognition (SDFR) Competition held in conjunction with the 18th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2024) and established to investigate the use of synthetic data for training face recognition models. The SDFR competition was split into two tasks, allowing participants to train face recognition systems using new synthetic datasets and/or existing ones. In the first task, the face recognition backbone was fixed and the dataset size was limited, while the second task provided almost complete freedom on the model backbone, the dataset, and the training pipeline. The submitted models were trained on existing and also new synthetic datasets and used clever methods to improve training with synthetic data. The submissions were evaluated and ranked on a diverse set of seven benchmarking datasets. The paper gives an overview of the submitted face recognition models and reports achieved performance compared to baseline models trained on real and synthetic datasets. Furthermore, the evaluation of submissions is extended to bias assessment across different demography groups. Lastly, an outlook on the current state of the research in training face recognition models using synthetic data is presented, and existing problems as well as potential future directions are also discussed.
Hatef Otroshi-Shahreza, Christophe Ecabert, Anjith George, Alexander Unnervik, Sébastien Marcel, Nicolò Di Domenico, Guido Borghi, Davide Maltoni, Fadi Boutros, Julia Vogel, Naser Damer, Ángela Sánchez-Pérez, Enrique Mas-Candela, Jorge Calvo-Zaragoza, Bernardo Biesseck, Pedro Vidal 0001, Roger Granada, David Menotti, Ivan DeAndres-Tame, Simone Maurizio La Cava, Sara Concas, Pietro Melzi, Ruben Tolosana, Rubén Vera-Rodríguez, Gianpaolo Perelli, Giulia Orrù, Gian Luca Marcialis, Julian Fierrez
FG14
2024 A Transformer Approach for Polyphonic Audio-to-Score Transcription
abstract
End-to-end Audio-to-Score (A2S) transcription aims to derive a score that represents the music content of an audio recording in a single step. While current state-of-the-art methods, which rely on Convolutional Recurrent Neural Networks trained with the Connectionist Temporal Classification loss function, have shown promising results under constrained circumstances, these approaches still exhibit fundamental limitations, especially when dealing with complex sequence modeling tasks, such as polyphonic music. To address these conditions, this work introduces an alternative learning scheme based on a Transformer decoder, specifically tailored for A2S by incorporating a two-dimensional positional encoding to preserve frequency-time relationships when processing the audio signal. The results obtained over three datasets of polyphonic string music confirm the adequacy of the method, which improves the transcription rate by an average of 44% compared to previous approaches.
María Alfaro-Contreras, Antonio Ríos-Vila, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
ICASSP4
2024 Evaluating Face Recognition Performance on Synthetic Data: A Comprehensive Analysis of Methodologies and Benchmarks
abstract
Synthetic data has become increasingly important for biometrics as an alternative to real data, thereby preventing issues associated with data protection regulations. In the realm of face recognition, synthetic generation methods are emerging to provide useful databases for training. However, a standardized protocol to compare the performance of these methods is lacking, along with an assessment framework for evaluating the performance of images generated when training face recognition models. In this paper, we report the outcomes of our efforts to assess the performance of face recognition when trained with synthetic data. We utilized 5 face recognition models and 7 synthetic datasets (plus 1 real dataset as a reference). Furthermore, for each training, four strategies for model selection—an issue typically neglected—were considered, to study their influence on final performance. Each case was evaluated under known face recognition benchmarks, with different conditions. Our results provide valuable conclusions regarding the influence of each part of the workflow, empirically confirming both known takeaways and unveiling underexplored aspects, notably the importance of model selection.
Ángela Sánchez-Pérez, Enrique Mas-Candela, Jorge Calvo-Zaragoza
IJCB3
2024 Analysis of the Calibration of Handwriting Text Recognition Models
Eric Ayllon, Francisco J. Castellanos 0001, Jorge Calvo-Zaragoza
ICDAR (2)3
2024 Sheet Music Transformer: End-To-End Optical Music Recognition Beyond Monophonic Transcription
Antonio Ríos-Vila, Jorge Calvo-Zaragoza, Thierry Paquet
ICDAR (6)2
2024 Source-Free Domain Adaptation for Optical Music Recognition
Adrian Rosello, Eliseo Fuentes-Martínez, María Alfaro-Contreras, David Rizo, Jorge Calvo-Zaragoza
ICDAR (6)5
2024 MUSCAT: A Multimodal mUSic Collection for Automatic Transcription of Real Recordings and Image Scores
Alejandro Galán-Cuenca, Jose J. Valero-Mas, Juan C. Martinez-Sevilla, Antonio Hidalgo-Centeno, Antonio Pertusa, Jorge Calvo-Zaragoza
ACM Multimedia6
2024 Exploring recursive neural networks for compact handwritten text recognition models
abstract
Abstract This paper addresses the challenge of deploying recognition models in specific scenarios in which memory size is relevant, such as in low-cost devices or browser-based applications. We specifically focus on developing memory-efficient approaches for Handwritten Text Recognition (HTR) by leveraging recursive networks. These networks reuse learned weights across successive layers, thus enabling the maintenance of depth, a critical factor associated with model accuracy, without an increase in memory footprint. We apply neural recursion techniques to models typically used in HTR that contain convolutional and recurrent layers. We additionally study the impact of kernel scaling, which allows the activations of these recursive layers to be modified for greater expressiveness with little cost to memory. Our experiments on various HTR benchmarks demonstrate that recursive networks are, indeed, a good alternative. It is noteworthy that these recursive networks not only preserve but in some instances also enhance accuracy, making them a promising solution for memory-efficient HTR applications. This research establishes the utility of recursive networks in addressing memory constraints in HTR models. Their ability to sustain or improve accuracy while being memory-efficient positions them as a promising solution for practical deployment, especially in contexts where memory size is a critical consideration, such as low-cost devices and browser-based applications.
Enrique Mas-Candela, Jorge Calvo-Zaragoza
Int. J. Document Anal. Recognit.2
2023 A Holistic Approach for Aligned Music and Lyrics Transcription
Juan C. Martinez-Sevilla, Antonio Ríos-Vila, Francisco J. Castellanos 0001, Jorge Calvo-Zaragoza
ICDAR (1)4
2023 Insights into end-to-end audio-to-score transcription with real recordings: A case study with saxophone works
abstract
Neural end-to-end Audio-to-Score (A2S) transcription aims to retrieve a score that encodes the music content of an audio recording in a single step. Due to the recentness of this formulation, the existing works have exclusively addressed controlled scenarios with synthetic data that fail to provide conclusions applicable to real-world cases. In response to this gap in the literature, this work introduces a novel assortment of real saxophone recordings---together with their digital scores---and poses several experimental scenarios involving real and synthetic data. The obtained results confirm the adequacy of this A2S framework to deal with real data as well as proving the relevance of leveraging synthetic interpretations to improve the recognition rate in scenarios with real-data scarcity.
Juan C. Martinez-Sevilla, María Alfaro-Contreras, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
INTERSPEECH4
2023 Late multimodal fusion for image and audio music transcription
abstract
Music transcription, which deals with the conversion of music sources into a structured digital format, is a key problem for Music Information Retrieval (MIR). When addressing this challenge in computational terms, the MIR community follows two lines of research: music documents, which is the case of Optical Music Recognition (OMR), or audio recordings, which is the case of Automatic Music Transcription (AMT). The different nature of the aforementioned input data has conditioned these fields to develop modality-specific frameworks. However, their recent definition in terms of sequence labeling tasks leads to a common output representation, which enables research on a combined paradigm. In this respect, multimodal image and audio music transcription comprises the challenge of effectively combining the information conveyed by image and audio modalities. In this work, we explore this question at a late-fusion level: we study four combination approaches in order to merge, for the first time, the hypotheses regarding end-to-end OMR and AMT systems in a lattice-based search space. The results obtained for a series of performance scenarios–in which the corresponding single-modality models yield different error rates–showed interesting benefits of these approaches. In addition, two of the four strategies considered significantly improve the corresponding unimodal standard recognition frameworks.
María Alfaro-Contreras, Jose J. Valero-Mas, José Manuel Iñesta Quereda, Jorge Calvo-Zaragoza
Expert Syst. Appl.4
2023 End-to-end optical music recognition for pianoform sheet music
abstract
Abstract End-to-end solutions have brought about significant advances in the field of Optical Music Recognition. These approaches directly provide the symbolic representation of a given image of a musical score. Despite this, several documents, such as pianoform musical scores, cannot yet benefit from these solutions since their structural complexity does not allow their effective transcription. This paper presents a neural method whose objective is to transcribe these musical scores in an end-to-end fashion. We also introduce the GrandStaff dataset, which contains 53,882 single-system piano scores in common western modern notation. The sources are encoded in both a standard digital music representation and its adaptation for current transcription technologies. The method proposed in this paper is trained and evaluated using this dataset. The results show that the approach presented is, for the first time, able to effectively transcribe pianoform notation in an end-to-end manner.
Antonio Ríos-Vila, David Rizo, José Manuel Iñesta Quereda, Jorge Calvo-Zaragoza
Int. J. Document Anal. Recognit.4
2023 Multimodal recognition of frustration during game-play with deep neural networks
abstract
Abstract Frustration, which is one aspect of the field of emotional recognition, is of particular interest to the video game industry as it provides information concerning each individual player’s level of engagement. The use of non-invasive strategies to estimate this emotion is, therefore, a relevant line of research with a direct application to real-world scenarios. While several proposals regarding the performance of non-invasive frustration recognition can be found in literature, they usually rely on hand-crafted features and rarely exploit the potential inherent to the combination of different sources of information. This work, therefore, presents a new approach that automatically extracts meaningful descriptors from individual audio and video sources of information using Deep Neural Networks (DNN) in order to then combine them, with the objective of detecting frustration in Game-Play scenarios. More precisely, two fusion modalities, namelydecision-levelandfeature-level, are presented and compared with state-of-the-art methods, along with different DNN architectures optimized for each type of data. Experiments performed with a real-world audiovisual benchmarking corpus revealed that the multimodal proposals introduced herein are more suitable than those of a unimodal nature, and that their performance also surpasses that of other state-of-the–art approaches, with error rate improvements of between 40%and 90%.
Carlos de la Fuente, Francisco J. Castellanos 0001, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
Multim. Tools Appl.4
2023 Kurcuma: a kitchen utensil recognition collection for unsupervised domain adaptation
abstract
Abstract The use of deep learning makes it possible to achieve extraordinary results in all kinds of tasks related to computer vision. However, this performance is strongly related to the availability of training data and its relationship with the distribution in the eventual application scenario. This question is of vital importance in areas such as robotics, where the targeted environment data are barely available in advance. In this context, domain adaptation (DA) techniques are especially important to building models that deal with new data for which the corresponding label is not available. To promote further research in DA techniques applied to robotics, this work presents Kurcuma (Kitchen Utensil Recognition Collection for Unsupervised doMain Adaptation), an assortment of seven datasets for the classification of kitchen utensils—a task of relevance in home-assistance robotics and a suitable showcase for DA. Along with the data, we provide a broad description of the main characteristics of the dataset, as well as a baseline using the well-known domain-adversarial training of neural networks approach. The results show the challenge posed by DA on these types of tasks, pointing to the need for new approaches in future work.
Adrian Rosello, Jose J. Valero-Mas, Antonio Javier Gallego 0001, Javier Sáez-Pérez, Jorge Calvo-Zaragoza
Pattern Anal. Appl.5
2023 End-to-End page-Level assessment of handwritten text recognition
abstract
The evaluation of Handwritten Text Recognition (HTR) systems has traditionally used metrics based on the edit distance between HTR and ground truth (GT) transcripts, at both the character and word levels. This is very adequate when the experimental protocol assumes that both GT and HTR text lines are the same, which allows edit distances to be independently computed to each given line. Driven by recent advances in pattern recognition, HTR systems increasingly face the end-to-end page-level transcription of a document, where the precision of locating the different text lines and their corresponding reading order (RO) play a key role. In such a case, the standard metrics do not take into account the inconsistencies that might appear. In this paper, the problem of evaluating HTR systems at the page level is introduced in detail. We analyse the convenience of using a two-fold evaluation, where the transcription accuracy and the RO goodness are considered separately. Different alternatives are proposed, analysed and empirically compared both through partially simulated and through real, full end-to-end experiments. Results support the validity of the proposed two-fold evaluation approach. An important conclusion is that such an evaluation can be adequately achieved by just two simple and well-known metrics: the Word Error Rate (WER), that takes transcription sequentiality into account, and the here re-formulated Bag of Words Word Error Rate (bWER), that ignores order. While the latter directly and very accurately assess intrinsic word recognition errors, the difference between both metrics (ΔWER) gracefully correlates with the Normalised Spearman’s Foot Rule Distance (NSFD), a metric which explicitly measures RO errors associated with layout analysis flaws. To arrive to these conclusions, we have introduced another metric called Hungarian Word Word Rate (hWER), based on a here proposed regularised version of the Hungarian Algorithm. This metric is shown to be always almost identical to bWER and both bWER and hWER are also almost identical to WER whenever HTR transcripts and GT references are guarantee to be in the same RO.
Enrique Vidal 0001, Alejandro H. Toselli, Antonio Ríos-Vila, Jorge Calvo-Zaragoza
Pattern Recognit.4
2023 Few-shot symbol classification via self-supervised learning and nearest neighbor
abstract
The recognition of symbols within document images is one of the most relevant steps involved in the Document Analysis field. While current state-of-the-art methods based on Deep Learning are capable of adequately performing this task, they generally require a vast amount of data that has to be manually labeled. In this paper, we propose a self-supervised learning-based method that addresses this task by training a neural-based feature extractor with a set of unlabeled documents and performs the recognition task considering just a few reference samples. Experiments on different corpora comprising music, text, and symbol documents report that the proposal is capable of adequately tackling the task with high accuracy rates of up to 95% in few-shot settings. Moreover, results show that the presented strategy outperforms the base supervised learning approaches trained with the same amount of data that, in some cases, even fail to converge. This approach, hence, stands as a lightweight alternative to deal with symbol classification with few annotated data.
María Alfaro-Contreras, Antonio Ríos-Vila, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
Pattern Recognit. Lett.4
2023 An experimental study on marine debris location and recognition using object detection
abstract
The large amount of debris in our oceans is a global problem that dramatically impacts marine fauna and flora. While a large number of human-based campaigns have been proposed to tackle this issue, these efforts have been deemed insufficient due to the insurmountable amount of existing litter. In response to that, there exists a high interest in the use of autonomous underwater vehicles (AUV) that may locate, identify, and collect this garbage automatically. To perform such a task, AUVs consider state-of-the-art object detection techniques based on deep neural networks due to their reported high performance. Nevertheless, these techniques generally require large amounts of data with fine-grained annotations. In this work, we explore the capabilities of the reference object detector Mask Region-based Convolutional Neural Networks for automatic marine debris location and classification in the context of limited data availability. Considering the recent CleanSea corpus, we pose several scenarios regarding the amount of available train data and study the possibility of mitigating the adverse effects of data scarcity with synthetic marine scenes. Our results achieve a new state of the art in the task, establishing a new reference for future research. In addition, it is shown that the task still has room for improvement and that the lack of data can be somehow alleviated, yet to a limited extent.
Alejandro Sánchez-Ferrer, Jose J. Valero-Mas, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
Pattern Recognit. Lett.4
2022 Neural Audio-To-Score Music Transcription For Unconstrained Polyphony Using Compact Output Representations
abstract
Neural Audio-to-Score (A2S) Music Transcription systems have shown promising results with pieces containing a fixed number of voices. However, they still exhibit fundamental limitations that constrain their applicability in wider scenarios. This work aims at tackling two of them: we introduce a novel output representation which addresses shortcomings related to the sequence-based A2S recognition framework and we report a first approximation to dealing with unconstrained polyphony. This is validated on a Convolutional Recurrent Neural Network (CRNN) with Connectionist Temporal Classification (CTC) A2S scheme using synthetic audio from string quartets and piano sonatas with intricate polyphonic mixtures. Our results, which improve fixed-polyphony state-of-the-art rates, may be considered a reference for future A2S works dealing with an unconstrained number of voices.
Víctor Arroyo, Jose J. Valero-Mas, Jorge Calvo-Zaragoza, Antonio Pertusa
ICASSP3
2022 Continual Learning for Document Image Binarization
abstract
In the field of Document Image Analysis (DIA), it is common to find great heterogeneity in terms of the possible graphic domains. In this sense, it is interesting to build neural models that can be sequentially adapted to new domains without losing the knowledge from the domains already learned. This learning paradigm is known as Continual (or Lifelong) Learning (CL). Although the adaptation comes along with a training set of the new domain, neural networks suffer what is known as "catastrophic forgetting". Therefore, assuming the constraint of not keeping data from the domains already addressed, this paradigm represents a challenge yet to be solved. This work presents an approach for CL in document image binarization, one of the most considered tasks within the DIA field. Our results report that it is indeed feasible to address CL in this field, given that the approach is successfully implemented and outperforms the baseline by a wide margin in most of the analyzed scenarios.
Carlos Garrido-Munoz, Adrián Sánchez-Hernández, Francisco J. Castellanos 0001, Jorge Calvo-Zaragoza
ICPR4
2022 Region-based layout analysis of music score images
abstract
The Layout Analysis (LA) stage is of vital importance to the correct performance of an Optical Music Recognition (OMR) system. It identifies the regions of interest, such as staves or lyrics, which must then be processed in order to transcribe their content. Despite the existence of modern approaches based on deep learning, an exhaustive study of LA in OMR has not yet been carried out with regard to the performance of different models, their generalization to different domains or, more importantly, their impact on subsequent stages of the pipeline. This work focuses on filling this gap in the literature by means of an experimental study of different neural architectures, music document types, and evaluation scenarios. The need for training data has also led to a proposal for a new semi-synthetic data-generation technique that enables the efficient applicability of LA approaches in real scenarios. Our results show that: (i) the choice of the model and its performance are crucial for the entire transcription process; (ii) the metrics commonly used to evaluate the LA stage do not always correlate with the final performance of the OMR system, and (iii) the proposed data-generation technique enables state-of-the-art results to be achieved with a limited set of labeled data.
Francisco J. Castellanos 0001, Carlos Garrido-Munoz, Antonio Ríos-Vila, Jorge Calvo-Zaragoza
Expert Syst. Appl.4
2022 Domain adaptation for staff-region retrieval of music score images
abstract
Abstract Optical music recognition (OMR) is the field that studies how to automatically read music notation from score images. One of the relevant steps within the OMR workflow is the staff-region retrieval. This process is a key step because any undetected staff will not be processed by the subsequent steps. This task has previously been addressed as a supervised learning problem in the literature; however, ground-truth data are not always available, so each new manuscript requires a preliminary manual annotation. This situation is one of the main bottlenecks in OMR, because of the countless number of existing manuscripts , and the associated manual labeling cost. With the aim of mitigating this issue, we propose the application of a domain adaptation technique, the so-called Domain-Adversarial Neural Network (DANN), based on a combination of a gradient reversal layer and a domain classifier in the inference neural architecture. The results from our experiments support the benefits of our proposed solution, obtaining improvements of approximately 29% in the F-score.
Francisco J. Castellanos 0001, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza, Ichiro Fujinaga
Int. J. Document Anal. Recognit.3
2022 A holistic approach for image-to-graph: application to optical music recognition
abstract
Abstract A number of applications would benefit from neural approaches that are capable of generating graphs from images in an end-to-end fashion. One of these fields is optical music recognition (OMR), which focuses on the computational reading of music notation from document images. Given that music notation can be expressed as a graph, the aforementioned approach represents a promising solution for OMR. In this work, we propose a new neural architecture that retrieves a certain representation of a graph—identified by a specific order of its vertices—in an end-to-end manner. This architecture works by means of a double output: It sequentially predicts the possible categories of the vertices, along with the edges between each of their pairs. The experiments carried out prove the effectiveness of our proposal as regards retrieving graph structures from excerpts of handwritten musical notation. Our results also show that certain design decisions, such as the choice of graph representations, play a fundamental role in the performance of this approach.
Carlos Garrido-Munoz, Antonio Ríos-Vila, Jorge Calvo-Zaragoza
Int. J. Document Anal. Recognit.3
2022 Decoupling music notation to improve end-to-end Optical Music Recognition
abstract
Inspired by the Text Recognition field, end-to-end schemes based on Convolutional Recurrent Neural Networks (CRNN) trained with the Connectionist Temporal Classification (CTC) loss function are considered one of the current state-of-the-art techniques for staff-level Optical Music Recognition (OMR). Unlike text symbols, music-notation elements may be defined as a combination of (i) a shape primitive located in (ii) a certain position in a staff. However, this double nature is generally neglected in the learning process, as each combination is treated as a single token. In this work, we study whether exploiting such particularity of music notation actually benefits the recognition performance and, if so, which approach is the most appropriate. For that, we thoroughly review existing specific approaches that explore this premise and propose different combinations of them. Furthermore, considering the limitations observed in such approaches, a novel decoding strategy specifically designed for OMR is proposed. The results obtained with four different corpora of historical manuscripts show the relevance of leveraging this double nature of music notation since it outperforms the standard approaches where it is ignored. In addition, the proposed decoding leads to significant reductions in the error rates with respect to the other cases.
María Alfaro-Contreras, Antonio Ríos-Vila, Jose J. Valero-Mas, José Manuel Iñesta Quereda, Jorge Calvo-Zaragoza
Pattern Recognit. Lett.5
2021 Sequential Next-Symbol Prediction for Optical Music Recognition
Enrique Mas-Candela, María Alfaro-Contreras, Jorge Calvo-Zaragoza
ICDAR (3)3
2021 Complete Optical Music Recognition via Agnostic Transcription and Machine Translation
Antonio Ríos-Vila, David Rizo, Jorge Calvo-Zaragoza
ICDAR (3)3
2021 Unsupervised neural domain adaptation for document image binarization
Francisco J. Castellanos 0001, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
Pattern Recognit.3
2021 Prototype generation in the string space via approximate median for data reduction in nearest neighbor classification
abstract
Abstract The k-nearest neighbor (kNN) rule is one of the best-known distance-based classifiers, and is usually associated with high performance and versatility as it requires only the definition of a dissimilarity measure. Nevertheless, kNN is also coupled with low-efficiency levels since, for each new query, the algorithm must carry out an exhaustive search of the training data, and this drawback is much more relevant when considering complex structural representations, such as graphs, trees or strings, owing to the cost of the dissimilarity metrics. This issue has generally been tackled through the use of data reduction (DR) techniques, which reduce the size of the reference set, but the complexity of structural data has historically limited their application in the aforementioned scenarios. A DR algorithm denominated as reduction through homogeneous clusters (RHC) has recently been adapted to string representations but as obtaining the exact median value of a set of string data is known to be computationally difficult, its authors resorted to computing the set-median value. Under the premise that a more exact median value may be beneficial in this context, we, therefore, present a new adaptation of the RHC algorithm for string data, in which an approximate median computation is carried out. The results obtained show significant improvements when compared to those of the set-median version of the algorithm, in terms of both classification performance and reduction rates.
Francisco J. Castellanos 0001, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
Soft Comput.3
2021 Incremental Unsupervised Domain-Adversarial Training of Neural Networks
abstract
In the context of supervised statistical learning, it is typically assumed that the training set comes from the same distribution that draws the test samples. When this is not the case, the behavior of the learned model is unpredictable and becomes dependent upon the degree of similarity between the distribution of the training set and the distribution of the test set. One of the research topics that investigates this scenario is referred to as domain adaptation (DA). Deep neural networks brought dramatic advances in pattern recognition and that is why there have been many attempts to provide good DA algorithms for these models. Herein we take a different avenue and approach the problem from an incremental point of view, where the model is adapted to the new domain iteratively. We make use of an existing unsupervised domain-adaptation algorithm to identify the target samples on which there is greater confidence about their true label. The output of the model is analyzed in different ways to determine the candidate samples. The selected samples are then added to the source training set by self-labeling, and the process is repeated until all target samples are labeled. This approach implements a form of adversarial training in which, by moving the self-labeled samples from the target to the source set, the DA algorithm is forced to look for new features after each iteration. Our results report a clear improvement with respect to the non-incremental case in several data sets, also outperforming other state-of-the-art DA algorithms.
Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza, Robert B. Fisher
IEEE Trans. Neural Networks Learn. Syst.2
2020 Exploring the two-dimensional nature of music notation for score recognition with end-to-end approaches
abstract
Optical Music Recognition workflows perform several steps to retrieve the content in music score images, being symbol recognition one of the key stages. State-of-the-art approaches for this stage currently address the coding of the output symbols as if they were plain text characters. However, music symbols have a two-dimensional nature that is ignored in these approaches. In this paper, we explore alternative output representations to perform music symbol recognition with state-of-the-art end-to-end neural technologies. We propose and describe new output representations which take into account the mentioned two-dimensional nature. We seek answers to the question of whether it is possible to obtain better recognition results in both printed and handwritten music scores. In this analysis, we compare the results given using three output encodings and two neural approaches. We found that one of the proposed encodings outperforms the results obtained by the standard one. This permits us to conclude that it is interesting to keep researching on this topic to improve end-to-end music score recognition.
Antonio Ríos-Vila, Jorge Calvo-Zaragoza, José Manuel Iñesta Quereda
ICFHR2
2020 Automatic scale estimation for music score images
abstract
Optical Music Recognition (OMR) is the research field focused on the automatic reading of music from scanned images. Its main goal is to encode the content into a digital and structured format with the advantages that this entails. This discipline is traditionally aligned to a workflow whose first step is the document analysis. This step is responsible of recognizing and detecting different sources of information—e.g. music notes, staff lines and text—to extract them and then processing automatically the content in the following steps of the workflow. One of the most difficult challenges it faces is to provide a generic solution to analyze documents with diverse resolutions. The endless number of existing music sources does not meet a standard that normalizes the data collections, giving complete freedom for a wide variety of image sizes and scales, thereby making this operation unsustainable. In the literature, this question is commonly overlooked and a uniform scale is assumed. In this paper, a machine learning-based approach to estimate the scale of music documents with respect to a reference scale is presented. Our goal is to propose a robust and generalizable method to adapt the input image to the requirements of an OMR system. For this, two goal-directed case studies are included to evaluate the proposed approach over common task within the OMR workflow, comparing the behavior with other state-of-the-art methods. Results suggest that it is necessary to perform this additional step in the first stage of the workflow to correct the scale of the input images. In addition, it is empirically demonstrated that our specialized approach is more promising than image augmentation strategies for the multi-scale challenge.
Francisco J. Castellanos 0001, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
Expert Syst. Appl.3
2020 Data representations for audio-to-score monophonic music transcription
abstract
This work presents an end-to-end method based on deep neural networks for audio-to-score music transcription of monophonic excerpts. Unlike existing music transcription methods, which normally perform pitch estimation, the proposed approach is formulated as an end-to-end task that outputs a notation-level music score. Using an audio file as input, modeled as a sequence of frames, a deep neural network is trained to provide a sequence of music symbols encoding a score, including key and time signatures, barlines, notes (with their pitch spelling and duration) and rests. Our framework is based on a Convolutional Recurrent Neural Network (CRNN) with Connectionist Temporal Classification (CTC) loss function trained in an end-to-end fashion, without requiring to align the input frames with the output symbols. A total of 246,870 incipits from the Répertoire International des Sources Musicales online catalog were synthesized using different timbres and tempos to build the training data. Alternative input representations (raw audio, Short-Time Fourier Transform (STFT), log-spaced STFT and Constant-Q transform) were evaluated for this task, as well as different output representations (Plaine & Easie Code, Kern, and a purpose-designed output). Results show that it is feasible to directly infer score representations from audio files and most errors come from music notation ambiguities and metering (time signatures and barlines).
Miguel A. Román, Antonio Pertusa, Jorge Calvo-Zaragoza
Expert Syst. Appl.3
2020 Ensemble classification from deep predictions with test data augmentation
Jorge Calvo-Zaragoza, Juan Ramón Rico-Juan, Antonio Javier Gallego 0001
Soft Comput.1
2019 Music Symbol Sequence Indexing in Medieval Plainchant Manuscripts
abstract
Huge amounts of musical manuscripts are preserved in cathedrals, abbeys, and archives. However, without reliable transcripts, their contents are inaccessible. Manual transcription is unaffordable for large collections, and current automatic technologies-such as Optical Music Recognition or Handwritten Music Recognition-do not provide sufficient accuracy for a fully-automatic scenario. In many cases, perfect transcripts are not really needed, given that content-based search with some degree of reliability would already be extremely useful. Spotting just single music symbols is rather useless (most of the symbols generally appear in all pages); instead, helpful search targets are melodic patterns, which typically correspond to music symbol sequences. We explore approaches for accurate retrieval of melodic patterns, represented by music symbol sequences, from collections of Medieval plainchant manuscripts. Our statistical framework, based on the use of convolutional recurrent neural networks and probabilistic indices, is shown to be useful for retrieving music patterns which appear frequently in this untranscribed images, yielding an Average Precision of 86 %.
Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001, Joan-Andreu Sánchez
ICDAR1
2019 Hybrid hidden Markov models and artificial neural networks for handwritten music recognition in mensural notation
Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001
Pattern Anal. Appl.1
2019 A selectional auto-encoder approach for document image binarization
Jorge Calvo-Zaragoza, Antonio Javier Gallego 0001
Pattern Recognit.1
2019 From Optical Music Recognition to Handwritten Music Recognition: A baseline
Arnau Baró, Pau Riba, Jorge Calvo-Zaragoza, Alicia Fornés
Pattern Recognit. Lett.3
2019 Handwritten Music Recognition for Mensural notation with convolutional recurrent neural networks
Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001
Pattern Recognit. Lett.1
2018 Data Augmentation via Variational Auto-Encoders
Unai Garay-Maestre, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
CIARP3
2018 Probabilistic Music-Symbol Spotting in Handwritten Scores
abstract
Content-based search on musical manuscripts is usually performed assuming that there are accurate transcripts of the sources in a symbolic, structured format. Given that current systems for Handwritten Music Recognition are far from offering guarantees about their accuracy, this traditional approach does not represent a scalable scenario. In this work we propose a probabilistic framework for Music-Symbol Spotting (MSS), that allows for content-based music search directly over the images of the manuscripts. By means of statistical recognition systems, a probabilistic index is built upon which the search can be carried out efficiently. Our experiments over a dataset of an Early handwritten music manuscript in Mensural notation demonstrates that this MSS framework can be presented as a promising alternative to the traditional approach for content-based music search.
Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001
ICFHR1
2018 Clustering-based k-nearest neighbor classification for large-scale data with neural codes representation
Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza, Jose J. Valero-Mas, Juan Ramón Rico-Juan
Pattern Recognit.2
2018 Oversampling imbalanced data in the string space
Francisco J. Castellanos 0001, Jose J. Valero-Mas, Jorge Calvo-Zaragoza, Juan Ramón Rico-Juan
Pattern Recognit. Lett.3
2017 Recognition of Handwritten Music Symbols with Convolutional Neural Codes
abstract
There are large collections of music manuscripts preserved over the centuries. In order to analyze these documents it is necessary to transcribe them into a machine-readable format. This process can be done automatically using Optical Music Recognition (OMR) systems, which typically consider segmentation plus classification workflows. This work is focused on the latter stage, presenting a comprehensive study for classification of handwritten musical symbols using Convolutional Neural Networks (CNN). The power of these models lies in their ability to transform the input into a meaningful representation for the task at hand, and that is why we study the use of these models to extract features (Neural Codes) for other classifiers. For the evaluation we consider four datasets containing different configurations and notation styles, along with a number of network models, different image preprocessing techniques and several supervised learning classifiers. Our results show that a remarkable accuracy can be achieved using the proposed framework, which significantly outperforms the state of the art in all datasets considered.
Jorge Calvo-Zaragoza, Antonio Javier Gallego 0001, Antonio Pertusa
ICDAR1
2017 Handwritten Music Recognition for Mensural Notation: Formulation, Data and Baseline Results
abstract
Music is a key element for cultural transmission, and so large collections of music manuscripts have been preserved over the centuries. In order to develop computational tools for analysis, indexing and retrieval from these sources, it is necessary to transcribe the content to some machine-readable format. In this paper we discuss the Handwritten Music Recognition problem, which refers to the development of automatic transcription systems for musical manuscripts. We focus on mensural notation, one of the most widespread varieties of Western classical music. For that, we present a labeled corpus containing 576 staves, along with a baseline recognition system based on a combination of hidden Markov models and N-gram language models. The baseline error obtained at symbol level is about 40 % which, given the difficulty of the task, can be considered a good starting point for future developments. Our aim is that these data and preliminary results help to promote this research field, serving as a reference in future developments.
Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001
ICDAR1
2017 Recognition of Handwritten Music Symbols using Meta-features Obtained from Weak Classifiers based on Nearest Neighbor
Jorge Calvo-Zaragoza, Jose J. Valero-Mas, Juan Ramón Rico-Juan
ICPRAM1
2017 Recognition of pen-based music notation with finite-state machines
Jorge Calvo-Zaragoza, José Oncina
Expert Syst. Appl.1
2017 Staff-line removal with selectional auto-encoders
Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
Expert Syst. Appl.2
2017 Staff-line detection and removal using a convolutional neural network
Jorge Calvo-Zaragoza, Antonio Pertusa, José Oncina
Mach. Vis. Appl.1
2017 Prototype generation on structural data using dissimilarity space representation
Jorge Calvo-Zaragoza, Jose J. Valero-Mas, Juan Ramón Rico-Juan
Neural Comput. Appl.1
2017 An efficient approach for Interactive Sequential Pattern Recognition
Jorge Calvo-Zaragoza, José Oncina
Pattern Recognit.1
2017 Selecting promising classes from generated data for an efficient multi-class nearest neighbor classification
Jorge Calvo-Zaragoza, Jose J. Valero-Mas, Juan Ramón Rico-Juan
Soft Comput.1
2017 An experimental study on rank methods for prototype selection
Jose J. Valero-Mas, Jorge Calvo-Zaragoza, Juan Ramón Rico-Juan, José Manuel Iñesta Quereda
Soft Comput.2
2016 Early Handwritten Music Recognition with Hidden Markov Models
abstract
This work presents a statistical method to tackle the Handwritten Music Recognition task for Early notation, which comprises more than 200 different symbols. Unlike previous approaches to deal with music notation, our strategy is to perform a holistic recognition without any previous segmentation or staff removal process. The input consists of a page of a music book, which is processed to extract and normalize the staves contained. Then, a feature extraction process is applied to define such sections as a sequence of numerical vectors. The recognition is based on the use of Hidden Markov Models for the optical processing and smoothed N-grams as language model. Experimentation results over a historical archive of Hispanic music reported an error around 40 %, which confirms our proposal as a good starting point taking into account the difficulty of the task.
Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001
ICFHR1
2016 Sheet Music Statistical Layout Analysis
abstract
In order to provide access to the contents of ancient music scores to researchers, the transcripts of both the lyrics and the musical notation is required. Before attempting any type of automatic or semi-automatic transcription of sheet music, an adequate layout analysis (LA) is needed. This LA must provide not only the locations of the different image regions, but also adequate region labels to distinguish between different region types such as staff, lyric, etc. To this end, we adapt a stochastic framework for LA based on Hidden Markov Models that we had previously introduced for detection and classification of text lines in typical handwritten text images. The proposed approach takes a scanned music score image as input and, after basic preprocessing, simultaneously performs region detection and region classification in an integrated way. To assess this statistical LA approach several experiments were carried out on a representative sample of a historical music archive, under different difficulty settings. The results show that our approach is able to tackle these structured documents providing good results not only for region detection but also for classification of the different regions.
Vicente Bosch, Jorge Calvo-Zaragoza, Alejandro H. Toselli, Enrique Vidal 0001
ICFHR2
2016 Computing the Expected Edit Distance from a String to a PFA
Jorge Calvo-Zaragoza, Colin de la Higuera, José Oncina
CIAA1
2016 Music staff removal with supervised pixel classification
Jorge Calvo-Zaragoza, Luisa Micó, José Oncina
Int. J. Document Anal. Recognit.1
2016 On the suitability of Prototype Selection methods for kNN classification with distributed data
Jose J. Valero-Mas, Jorge Calvo-Zaragoza, Juan Ramón Rico-Juan
Neurocomputing2
2015 Improving classification using a Confidence Matrix based on weak classifiers applied to OCR
Juan Ramón Rico-Juan, Jorge Calvo-Zaragoza
Neurocomputing2
2015 Avoiding staff removal stage in optical music recognition: application to scores written in white mensural notation
Jorge Calvo-Zaragoza, Isabel Barbancho, Lorenzo J. Tardón, Ana M. Barbancho
Pattern Anal. Appl.1
2015 Improving kNN multi-label classification in Prototype Selection scenarios using class proposals
Jorge Calvo-Zaragoza, Jose J. Valero-Mas, Juan Ramón Rico-Juan
Pattern Recognit.1
2014 Recognition of Pen-Based Music Notation: The HOMUS Dataset
abstract
A profitable way of digitizing a new musical composition is by using a pen-based (online) system, in which the score is created with the sole effort of the composition itself. However, the development of such systems is still largely unexplored. Some studies have been carried out but the use of particular little datasets has led to avoid objective comparisons between different approaches. To solve this situation, this work presents the Handwritten Online Musical Symbols (HOMUS) dataset, which consists of 15200 samples of 32 types of musical symbols from 100 different musicians. Several alternatives of recognition for the two modalities -online, using the strokes drawn by the pen, and offline, using the image generated after drawing the symbol- are also presented. Some experiments are included aimed to draw main conclusions about the recognition of these data. It is expected that this work can establish a binding point in the field of recognition of online handwritten music notation and serve as a baseline for future developments.
Jorge Calvo-Zaragoza, José Oncina
ICPR1
2014 Multi-objective adaptive evolutionary strategy for tuning compilations
Antonio Martínez-Álvarez, Jorge Calvo-Zaragoza, Sergio Cuenca-Asensi, Andrés Ortiz 0001, Antonio Jimeno-Morenilla
Neurocomputing2