Jose J. Valero-Mas

dblp:159/7607 · DBLP profile ↗
← Back
27ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0001-8667-4070ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploring the impact of label-level noise on multi-label k-Nearest Neighbor classification
abstract
Abstract Multi-label classification methods based on the k -Nearest Neighbor ( k NN) rule are widely used due to their simplicity and competitive performance, but their behavior under label-level noise remains insufficiently understood, especially when combined with data reduction techniques. This paper presents a comprehensive empirical study of the impact of label-level noise on multi-label k NN classification and on Multi-label Prototype Generation (MPG) methods. We formalize six label-level noise induction policies—Additive, Subtractive, Additive-Subtractive, Distribution-Aware Additive-Subtractive, Partial Uniform Multi-label, and Swap—parameterized by both the proportion of affected instances and a severity parameter. Their effect is analyzed on three representative k NN-based multi-label classifiers (BR k NN, LP k NN, and ML k NN) and five MPG strategies (MRHC, MChen, MRSP1-3) across eight benchmark datasets with varying label cardinality and imbalance, comprising an extensive experimental grid of 816,480 configurations. The results reveal that Additive and Partial Uniform noise are the most detrimental, whereas Subtractive and cardinality-preserving policies are comparatively less harmful. Moderate neighborhood sizes (around $$k=7$$ ) provide a good trade-off between robustness and accuracy, while ML k NN is consistently the most resilient classifier under severe noise. Among MPG methods, MRSP3 emerges as the most robust reduction strategy, whereas aggressive reductions, particularly with MRHC, can amplify the negative effects of noise. The code and complete experimental results are publicly released to support reproducibility and further research.
Antonio Requena, Alejandro Galán-Cuenca, Antonio Javier Gallego 0001, Jose J. Valero-Mas
Pattern Anal. Appl.4
2026 Insights into imbalance-aware Multilabel Prototype Generation mechanisms for k-Nearest Neighbor classification in noisy scenarios
abstract
Prototype Generation (PG) techniques enhance the efficiency of the k -Nearest Neighbor ( k NN) classifier by condensing datasets through the use of specific rules. More precisely, these strategies work on the premise of merging the elements in the reference data collection to generate an alternative and more compact data assortment that substitutes the former one without remarkably affecting the recognition performance. Nevertheless, despite being widely studied in multiclass scenarios, PG is still underexplored in multilabel contexts, leading to limitations, notably in the handling of label imbalance and noise. In this regard, this work introduces a reduction framework that allows for multilabel PG methods to handle these challenges of label imbalance and noise. The proposed mechanisms comprise a selection strategy that exclusively preserves noise-free samples in the process, a mechanism to avoid severely imbalanced samples from being inadequately processed, and two new merging policies for the PG methods to generate novel samples. These enhancements are considered along with three established multilabel PG methods: Multilabel Reduction through Homogeneous Clustering, Multilabel Chen, and Multilabel Reduction through Space Partitioning. Evaluations are conducted using three k NN-based multilabel classifiers and 12 diverse datasets with different levels of label imbalance. We additionally study the performance with varying values of k under different label-noise scenarios. The results are assessed through statistical tests and indicate that our proposals outperform the original methods that disregard label imbalance, even in the presence of noise, thus validating these approaches and fostering further research in the field.
Jose J. Valero-Mas, Carlos Peñarrubia, Francisco J. Castellanos 0001, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
Pattern Recognit.1
2025 Self-Supervised Learning for Text Recognition: A Critical Survey
abstract
Abstract Text Recognition (TR) refers to the research area that focuses on retrieving textual information from images, a topic that has seen significant advancements in the last decade due to the use of Deep Neural Networks (DNN). However, these solutions often necessitate vast amounts of manually labeled or synthetic data. Addressing this challenge, Self-Supervised Learning (SSL) has gained attention by utilizing large datasets of unlabeled data to train DNN, thereby generating meaningful and robust representations. Although SSL was initially overlooked in TR because of its unique characteristics, recent years have witnessed a surge in the development of SSL methods specifically for this field. This rapid development, however, has led to many methods being explored independently, without taking previous efforts in methodology or comparison into account, thereby hindering progress in the field of research. This paper, therefore, seeks to consolidate the use of SSL in the field of TR, offering a critical and comprehensive overview of the current state of the art. We will review and analyze the existing methods, compare their results, and highlight inconsistencies in the current literature. This thorough analysis aims to provide general insights into the field, propose standardizations, identify new research directions, and foster its proper development.
Carlos Peñarrubia, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
Int. J. Comput. Vis.2
2025 Spatial context-based Self-Supervised Learning for Handwritten Text Recognition
abstract
Handwritten Text Recognition (HTR) is a relevant problem in computer vision, and implies unique challenges owing to its inherent variability and the rich contextualization required for its interpretation. Despite the success of Self-Supervised Learning (SSL) in computer vision, its application to HTR has been rather scattered, leaving key SSL methodologies unexplored. This work specifically focuses on Spatial Context-based SSL. We investigate how this family of approaches can be adapted and optimized for HTR and propose new workflows that leverage the unique features of handwritten text. Our experiments demonstrate that the methods considered lead to advancements in the state-of-the-art of SSL for HTR in a number of benchmark cases. • We investigate the performance of existing spatial context-based SSL methods for HTR. • We propose new methods within this family of SSL. • We advance the state of the art in SSL for HTR. • We highlight that HTR contains rich spatial information due to font and stroke style.
Carlos Peñarrubia, Carlos Garrido-Munoz, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
Pattern Recognit. Lett.3
2024 Contrastive Self-Supervised Learning for Optical Music Recognition
Carlos Peñarrubia, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
DAS2
2024 A Transformer Approach for Polyphonic Audio-to-Score Transcription
abstract
End-to-end Audio-to-Score (A2S) transcription aims to derive a score that represents the music content of an audio recording in a single step. While current state-of-the-art methods, which rely on Convolutional Recurrent Neural Networks trained with the Connectionist Temporal Classification loss function, have shown promising results under constrained circumstances, these approaches still exhibit fundamental limitations, especially when dealing with complex sequence modeling tasks, such as polyphonic music. To address these conditions, this work introduces an alternative learning scheme based on a Transformer decoder, specifically tailored for A2S by incorporating a two-dimensional positional encoding to preserve frequency-time relationships when processing the audio signal. The results obtained over three datasets of polyphonic string music confirm the adequacy of the method, which improves the transcription rate by an average of 44% compared to previous approaches.
María Alfaro-Contreras, Antonio Ríos-Vila, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
ICASSP3
2024 MUSCAT: A Multimodal mUSic Collection for Automatic Transcription of Real Recordings and Image Scores
Alejandro Galán-Cuenca, Jose J. Valero-Mas, Juan C. Martinez-Sevilla, Antonio Hidalgo-Centeno, Antonio Pertusa, Jorge Calvo-Zaragoza
ACM Multimedia2
2024 An overview of ensemble and feature learning in few-shot image classification using siamese networks
abstract
Abstract Siamese Neural Networks (SNNs) constitute one of the most representative approaches for addressing Few-Shot Image Classification. These schemes comprise a set of Convolutional Neural Network (CNN) models whose weights are shared across the network, which results in fewer parameters to train and less tendency to overfit. This fact eventually leads to better convergence capabilities than standard neural models when considering scarce amounts of data. Based on a contrastive principle, the SNN scheme jointly trains these inner CNN models to map the input image data to an embedded representation that may be later exploited for the recognition process. However, in spite of their extensive use in the related literature, the representation capabilities of SNN schemes have neither been thoroughly assessed nor combined with other strategies for boosting their classification performance. Within this context, this work experimentally studies the capabilities of SNN architectures for obtaining a suitable embedded representation in scenarios with a severe data scarcity, assesses the use of train data augmentation for improving the feature learning process, introduces the use of transfer learning techniques for further exploiting the embedded representations obtained by the model, and uses test data augmentation for boosting the performance capabilities of the SNN scheme by mimicking an ensemble learning process. The results obtained with different image corpora report that the combination of the commented techniques achieves classification rates ranging from 69% to 78% with just 5 to 20 prototypes per class whereas the CNN baseline considered is unable to converge. Furthermore, upon the convergence of the baseline model with the sufficient amount of data, still the adequate use of the studied techniques improves the accuracy in figures from 4% to 9%.
Jose J. Valero-Mas, Antonio Javier Gallego 0001, Juan Ramón Rico-Juan
Multim. Tools Appl.1
2023 Insights into end-to-end audio-to-score transcription with real recordings: A case study with saxophone works
abstract
Neural end-to-end Audio-to-Score (A2S) transcription aims to retrieve a score that encodes the music content of an audio recording in a single step. Due to the recentness of this formulation, the existing works have exclusively addressed controlled scenarios with synthetic data that fail to provide conclusions applicable to real-world cases. In response to this gap in the literature, this work introduces a novel assortment of real saxophone recordings---together with their digital scores---and poses several experimental scenarios involving real and synthetic data. The obtained results confirm the adequacy of this A2S framework to deal with real data as well as proving the relevance of leveraging synthetic interpretations to improve the recognition rate in scenarios with real-data scarcity.
Juan C. Martinez-Sevilla, María Alfaro-Contreras, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
INTERSPEECH3
2023 Late multimodal fusion for image and audio music transcription
abstract
Music transcription, which deals with the conversion of music sources into a structured digital format, is a key problem for Music Information Retrieval (MIR). When addressing this challenge in computational terms, the MIR community follows two lines of research: music documents, which is the case of Optical Music Recognition (OMR), or audio recordings, which is the case of Automatic Music Transcription (AMT). The different nature of the aforementioned input data has conditioned these fields to develop modality-specific frameworks. However, their recent definition in terms of sequence labeling tasks leads to a common output representation, which enables research on a combined paradigm. In this respect, multimodal image and audio music transcription comprises the challenge of effectively combining the information conveyed by image and audio modalities. In this work, we explore this question at a late-fusion level: we study four combination approaches in order to merge, for the first time, the hypotheses regarding end-to-end OMR and AMT systems in a lattice-based search space. The results obtained for a series of performance scenarios–in which the corresponding single-modality models yield different error rates–showed interesting benefits of these approaches. In addition, two of the four strategies considered significantly improve the corresponding unimodal standard recognition frameworks.
María Alfaro-Contreras, Jose J. Valero-Mas, José Manuel Iñesta Quereda, Jorge Calvo-Zaragoza
Expert Syst. Appl.2
2023 Multimodal recognition of frustration during game-play with deep neural networks
abstract
Abstract Frustration, which is one aspect of the field of emotional recognition, is of particular interest to the video game industry as it provides information concerning each individual player’s level of engagement. The use of non-invasive strategies to estimate this emotion is, therefore, a relevant line of research with a direct application to real-world scenarios. While several proposals regarding the performance of non-invasive frustration recognition can be found in literature, they usually rely on hand-crafted features and rarely exploit the potential inherent to the combination of different sources of information. This work, therefore, presents a new approach that automatically extracts meaningful descriptors from individual audio and video sources of information using Deep Neural Networks (DNN) in order to then combine them, with the objective of detecting frustration in Game-Play scenarios. More precisely, two fusion modalities, namelydecision-levelandfeature-level, are presented and compared with state-of-the-art methods, along with different DNN architectures optimized for each type of data. Experiments performed with a real-world audiovisual benchmarking corpus revealed that the multimodal proposals introduced herein are more suitable than those of a unimodal nature, and that their performance also surpasses that of other state-of-the–art approaches, with error rate improvements of between 40%and 90%.
Carlos de la Fuente, Francisco J. Castellanos 0001, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
Multim. Tools Appl.3
2023 Kurcuma: a kitchen utensil recognition collection for unsupervised domain adaptation
abstract
Abstract The use of deep learning makes it possible to achieve extraordinary results in all kinds of tasks related to computer vision. However, this performance is strongly related to the availability of training data and its relationship with the distribution in the eventual application scenario. This question is of vital importance in areas such as robotics, where the targeted environment data are barely available in advance. In this context, domain adaptation (DA) techniques are especially important to building models that deal with new data for which the corresponding label is not available. To promote further research in DA techniques applied to robotics, this work presents Kurcuma (Kitchen Utensil Recognition Collection for Unsupervised doMain Adaptation), an assortment of seven datasets for the classification of kitchen utensils—a task of relevance in home-assistance robotics and a suitable showcase for DA. Along with the data, we provide a broad description of the main characteristics of the dataset, as well as a baseline using the well-known domain-adversarial training of neural networks approach. The results show the challenge posed by DA on these types of tasks, pointing to the need for new approaches in future work.
Adrian Rosello, Jose J. Valero-Mas, Antonio Javier Gallego 0001, Javier Sáez-Pérez, Jorge Calvo-Zaragoza
Pattern Anal. Appl.2
2023 Multilabel Prototype Generation for data reduction in K-Nearest Neighbour classification
abstract
Prototype Generation (PG) methods are typically considered for improving the efficiency of the k-Nearest Neighbour (kNN) classifier when tackling high-size corpora. Such approaches aim at generating a reduced version of the corpus without decreasing the classification performance when compared to the initial set. Despite their large application in multiclass scenarios, very few works have addressed the proposal of PG methods for the multilabel space. In this regard, this work presents the novel adaptation of four multiclass PG strategies to the multilabel case. These proposals are evaluated with three multilabel kNN-based classifiers, 12 corpora comprising a varied range of domains and corpus sizes, and different noise scenarios artificially induced in the data. The results obtained show that the proposed adaptations are capable of significantly improving—both in terms of efficiency and classification performance—the only reference multilabel PG work in the literature as well as the case in which no PG method is applied, also presenting statistically superior robustness in noisy scenarios. Moreover, these novel PG strategies allow prioritising either the efficiency or efficacy criteria through its configuration depending on the target scenario, hence covering a wide area in the solution space not previously filled by other works.
Jose J. Valero-Mas, Antonio Javier Gallego 0001, Pablo Alonso-Jiménez, Xavier Serra
Pattern Recognit.1
2023 Few-shot symbol classification via self-supervised learning and nearest neighbor
abstract
The recognition of symbols within document images is one of the most relevant steps involved in the Document Analysis field. While current state-of-the-art methods based on Deep Learning are capable of adequately performing this task, they generally require a vast amount of data that has to be manually labeled. In this paper, we propose a self-supervised learning-based method that addresses this task by training a neural-based feature extractor with a set of unlabeled documents and performs the recognition task considering just a few reference samples. Experiments on different corpora comprising music, text, and symbol documents report that the proposal is capable of adequately tackling the task with high accuracy rates of up to 95% in few-shot settings. Moreover, results show that the presented strategy outperforms the base supervised learning approaches trained with the same amount of data that, in some cases, even fail to converge. This approach, hence, stands as a lightweight alternative to deal with symbol classification with few annotated data.
María Alfaro-Contreras, Antonio Ríos-Vila, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
Pattern Recognit. Lett.3
2023 An experimental study on marine debris location and recognition using object detection
abstract
The large amount of debris in our oceans is a global problem that dramatically impacts marine fauna and flora. While a large number of human-based campaigns have been proposed to tackle this issue, these efforts have been deemed insufficient due to the insurmountable amount of existing litter. In response to that, there exists a high interest in the use of autonomous underwater vehicles (AUV) that may locate, identify, and collect this garbage automatically. To perform such a task, AUVs consider state-of-the-art object detection techniques based on deep neural networks due to their reported high performance. Nevertheless, these techniques generally require large amounts of data with fine-grained annotations. In this work, we explore the capabilities of the reference object detector Mask Region-based Convolutional Neural Networks for automatic marine debris location and classification in the context of limited data availability. Considering the recent CleanSea corpus, we pose several scenarios regarding the amount of available train data and study the possibility of mitigating the adverse effects of data scarcity with synthetic marine scenes. Our results achieve a new state of the art in the task, establishing a new reference for future research. In addition, it is shown that the task still has room for improvement and that the lack of data can be somehow alleviated, yet to a limited extent.
Alejandro Sánchez-Ferrer, Jose J. Valero-Mas, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
Pattern Recognit. Lett.2
2022 Neural Audio-To-Score Music Transcription For Unconstrained Polyphony Using Compact Output Representations
abstract
Neural Audio-to-Score (A2S) Music Transcription systems have shown promising results with pieces containing a fixed number of voices. However, they still exhibit fundamental limitations that constrain their applicability in wider scenarios. This work aims at tackling two of them: we introduce a novel output representation which addresses shortcomings related to the sequence-based A2S recognition framework and we report a first approximation to dealing with unconstrained polyphony. This is validated on a Convolutional Recurrent Neural Network (CRNN) with Connectionist Temporal Classification (CTC) A2S scheme using synthetic audio from string quartets and piano sonatas with intricate polyphonic mixtures. Our results, which improve fixed-polyphony state-of-the-art rates, may be considered a reference for future A2S works dealing with an unconstrained number of voices.
Víctor Arroyo, Jose J. Valero-Mas, Jorge Calvo-Zaragoza, Antonio Pertusa
ICASSP2
2022 Efficient k-nearest neighbor search based on clustering and adaptive k values
abstract
The k -Nearest Neighbor ( k NN) algorithm is widely used in the supervised learning field and, particularly, in search and classification tasks , owing to its simplicity, competitive performance, and good statistical properties. However, its inherent inefficiency prevents its use in most modern applications due to the vast amount of data that the current technological evolution generates, being thus the optimization of k NN-based search strategies of particular interest. This paper introduces the caKD+ algorithm, which tackles this limitation by combining the use of feature learning techniques, clustering methods , adaptive search parameters per cluster, and the use of pre-calculated K-Dimensional Tree structures, and results in a highly efficient search method. This proposal has been evaluated using 10 datasets and the results show that caKD+ significantly outperforms 16 state-of-the-art efficient search methods while still depicting such an accurate performance as the one by the exhaustive k NN search.
Antonio Javier Gallego 0001, Juan Ramón Rico-Juan, Jose J. Valero-Mas
Pattern Recognit.3
2022 Decoupling music notation to improve end-to-end Optical Music Recognition
abstract
Inspired by the Text Recognition field, end-to-end schemes based on Convolutional Recurrent Neural Networks (CRNN) trained with the Connectionist Temporal Classification (CTC) loss function are considered one of the current state-of-the-art techniques for staff-level Optical Music Recognition (OMR). Unlike text symbols, music-notation elements may be defined as a combination of (i) a shape primitive located in (ii) a certain position in a staff. However, this double nature is generally neglected in the learning process, as each combination is treated as a single token. In this work, we study whether exploiting such particularity of music notation actually benefits the recognition performance and, if so, which approach is the most appropriate. For that, we thoroughly review existing specific approaches that explore this premise and propose different combinations of them. Furthermore, considering the limitations observed in such approaches, a novel decoding strategy specifically designed for OMR is proposed. The results obtained with four different corpora of historical manuscripts show the relevance of leveraging this double nature of music notation since it outperforms the standard approaches where it is ignored. In addition, the proposed decoding leads to significant reductions in the error rates with respect to the other cases.
María Alfaro-Contreras, Antonio Ríos-Vila, Jose J. Valero-Mas, José Manuel Iñesta Quereda, Jorge Calvo-Zaragoza
Pattern Recognit. Lett.3
2021 Prototype generation in the string space via approximate median for data reduction in nearest neighbor classification
abstract
Abstract The k-nearest neighbor (kNN) rule is one of the best-known distance-based classifiers, and is usually associated with high performance and versatility as it requires only the definition of a dissimilarity measure. Nevertheless, kNN is also coupled with low-efficiency levels since, for each new query, the algorithm must carry out an exhaustive search of the training data, and this drawback is much more relevant when considering complex structural representations, such as graphs, trees or strings, owing to the cost of the dissimilarity metrics. This issue has generally been tackled through the use of data reduction (DR) techniques, which reduce the size of the reference set, but the complexity of structural data has historically limited their application in the aforementioned scenarios. A DR algorithm denominated as reduction through homogeneous clusters (RHC) has recently been adapted to string representations but as obtaining the exact median value of a set of string data is known to be computationally difficult, its authors resorted to computing the set-median value. Under the premise that a more exact median value may be beneficial in this context, we, therefore, present a new adaptation of the RHC algorithm for string data, in which an approximate median computation is carried out. The results obtained show significant improvements when compared to those of the set-median version of the algorithm, in terms of both classification performance and reduction rates.
Francisco J. Castellanos 0001, Jose J. Valero-Mas, Jorge Calvo-Zaragoza
Soft Comput.2
2018 Clustering-based k-nearest neighbor classification for large-scale data with neural codes representation
Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza, Jose J. Valero-Mas, Juan Ramón Rico-Juan
Pattern Recognit.3
2018 Oversampling imbalanced data in the string space
Francisco J. Castellanos 0001, Jose J. Valero-Mas, Jorge Calvo-Zaragoza, Juan Ramón Rico-Juan
Pattern Recognit. Lett.2
2017 Recognition of Handwritten Music Symbols using Meta-features Obtained from Weak Classifiers based on Nearest Neighbor
Jorge Calvo-Zaragoza, Jose J. Valero-Mas, Juan Ramón Rico-Juan
ICPRAM2
2017 Prototype generation on structural data using dissimilarity space representation
Jorge Calvo-Zaragoza, Jose J. Valero-Mas, Juan Ramón Rico-Juan
Neural Comput. Appl.2
2017 Selecting promising classes from generated data for an efficient multi-class nearest neighbor classification
Jorge Calvo-Zaragoza, Jose J. Valero-Mas, Juan Ramón Rico-Juan
Soft Comput.2
2017 An experimental study on rank methods for prototype selection
Jose J. Valero-Mas, Jorge Calvo-Zaragoza, Juan Ramón Rico-Juan, José Manuel Iñesta Quereda
Soft Comput.1
2016 On the suitability of Prototype Selection methods for kNN classification with distributed data
Jose J. Valero-Mas, Jorge Calvo-Zaragoza, Juan Ramón Rico-Juan
Neurocomputing1
2015 Improving kNN multi-label classification in Prototype Selection scenarios using class proposals
Jorge Calvo-Zaragoza, Jose J. Valero-Mas, Juan Ramón Rico-Juan
Pattern Recognit.2