Antonio Javier Gallego 0001

dblp:54/1887 · also Antonio Javier Gallego Sánchez · DBLP profile ↗
← Back
39ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0003-3148-6886ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 6 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 1Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Enhancing maritime search and rescue: Incremental unsupervised domain adaptation with synthetic data and pseudo-labeling
abstract
Maritime search and rescue operations are critical for saving lives in emergencies, where time is a decisive factor since delays can drastically reduce the chances of survival for those in distress. These missions are particularly challenging due to the inherent complexity of the maritime environment, marked by changing weather, dynamic sea states, and limited visibility. Developing reliable machine learning systems for this task typically requires large amounts of labeled data that capture all possible operating conditions. However, collecting and annotating such data is costly and often unfeasible in real-world maritime scenarios. To address this limitation, we propose a domain adaptation strategy for a segmentation-based detection model that estimates a probability map indicating the presence of human bodies at sea. The method enables unsupervised learning to adapt from a labeled synthetic domain to a real, unlabeled domain by employing a Domain-Adversarial Neural Network that aligns feature representations across domains, and an iterative pseudo-labeling process that selects high-confidence predictions on the target data to progressively refine the model. By leveraging synthetic data—automatically generated and labeled—our approach adapts effectively to real-world conditions without requiring manual annotation. Experimental results show that our method outperforms several state-of-the-art detectors while maintaining a lightweight architecture. Moreover, it generalizes well under diverse and adverse environmental conditions, including fog, rain, and low-light scenes, demonstrating its robustness and suitability for real-world deployment in critical rescue operations.
Juan Pedro Martinez-Esteso, Francisco J. Castellanos 0001, Antonio Javier Gallego 0001
Expert Syst. Appl.3
2026 Exploring the impact of label-level noise on multi-label k-Nearest Neighbor classification
abstract
Abstract Multi-label classification methods based on the k -Nearest Neighbor ( k NN) rule are widely used due to their simplicity and competitive performance, but their behavior under label-level noise remains insufficiently understood, especially when combined with data reduction techniques. This paper presents a comprehensive empirical study of the impact of label-level noise on multi-label k NN classification and on Multi-label Prototype Generation (MPG) methods. We formalize six label-level noise induction policies—Additive, Subtractive, Additive-Subtractive, Distribution-Aware Additive-Subtractive, Partial Uniform Multi-label, and Swap—parameterized by both the proportion of affected instances and a severity parameter. Their effect is analyzed on three representative k NN-based multi-label classifiers (BR k NN, LP k NN, and ML k NN) and five MPG strategies (MRHC, MChen, MRSP1-3) across eight benchmark datasets with varying label cardinality and imbalance, comprising an extensive experimental grid of 816,480 configurations. The results reveal that Additive and Partial Uniform noise are the most detrimental, whereas Subtractive and cardinality-preserving policies are comparatively less harmful. Moderate neighborhood sizes (around $$k=7$$ ) provide a good trade-off between robustness and accuracy, while ML k NN is consistently the most resilient classifier under severe noise. Among MPG methods, MRSP3 emerges as the most robust reduction strategy, whereas aggressive reductions, particularly with MRHC, can amplify the negative effects of noise. The code and complete experimental results are publicly released to support reproducibility and further research.
Antonio Requena, Alejandro Galán-Cuenca, Antonio Javier Gallego 0001, Jose J. Valero-Mas
Pattern Anal. Appl.3
2026 TriScore: Aligning audio, symbolic scores, and sheet music images in a shared embedding space
abstract
Multimodal representation learning has attracted increasing attention in Music Information Retrieval (MIR), yet score-based multimodality is still constrained by the lack of datasets that jointly provide audio, notation-level symbolic scores, and sheet music images with reliable alignment and balanced coverage. To address this gap, we introduce TriScore , a compact tri-modal collection of excerpt-aligned triplets annotated with composer , instrument , and piece title . We further propose a scalable fusion approach that keeps modality-specific pretrained encoders frozen and learns a shared shallow projection that maps heterogeneous embeddings into a common latent space. Extensive experiments study the impact of normalization, class-imbalance mitigation, projection dimensionality, and k NN neighborhood size. The learned space achieves strong excerpt-level classification performance across modalities, reaching 90.38 F1 for composer and 95.88 F1 for instrument on average, and attains 88.29/92.43 Top-5/Top-10 accuracy for piece retrieval . In-depth analyses show that the projection yields consistent global semantic alignment while preserving local modality-specific submanifolds: neighborhoods are dominated by same-modality samples in the full search space, yet cross-modal exclusion results remain far above distinct-modality pretrained baselines, indicating non-trivial cross-modal consistency. Overall, TriScore and the proposed projection framework establish a controlled benchmark and strong baselines for tri-modal, score-based MIR.
Antonio Hidalgo-Centeno, Eliseo Fuentes-Martínez, Jorge Calvo-Zaragoza, Antonio Javier Gallego 0001
Pattern Recognit.4
2026 Insights into imbalance-aware Multilabel Prototype Generation mechanisms for k-Nearest Neighbor classification in noisy scenarios
abstract
Prototype Generation (PG) techniques enhance the efficiency of the k -Nearest Neighbor ( k NN) classifier by condensing datasets through the use of specific rules. More precisely, these strategies work on the premise of merging the elements in the reference data collection to generate an alternative and more compact data assortment that substitutes the former one without remarkably affecting the recognition performance. Nevertheless, despite being widely studied in multiclass scenarios, PG is still underexplored in multilabel contexts, leading to limitations, notably in the handling of label imbalance and noise. In this regard, this work introduces a reduction framework that allows for multilabel PG methods to handle these challenges of label imbalance and noise. The proposed mechanisms comprise a selection strategy that exclusively preserves noise-free samples in the process, a mechanism to avoid severely imbalanced samples from being inadequately processed, and two new merging policies for the PG methods to generate novel samples. These enhancements are considered along with three established multilabel PG methods: Multilabel Reduction through Homogeneous Clustering, Multilabel Chen, and Multilabel Reduction through Space Partitioning. Evaluations are conducted using three k NN-based multilabel classifiers and 12 diverse datasets with different levels of label imbalance. We additionally study the performance with varying values of k under different label-noise scenarios. The results are assessed through statistical tests and indicate that our proposals outperform the original methods that disregard label imbalance, even in the presence of noise, thus validating these approaches and fostering further research in the field.
Jose J. Valero-Mas, Carlos Peñarrubia, Francisco J. Castellanos 0001, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
Pattern Recognit.4
2025 On the use of synthetic data for body detection in maritime search and rescue operations
abstract
Time is a critical factor in maritime Search And Rescue (SAR) missions, during which promptly locating survivors is paramount. Unmanned Aerial Vehicles (UAVs) are a useful tool with which to increase the success rate by rapidly identifying targets. While this task can be performed using other means, such as helicopters, the cost-effectiveness of UAVs makes them an effective choice. Moreover, these vehicles allow the easy integration of automatic systems that can be used to assist in the search process. Despite the impact of artificial intelligence on autonomous technology, there are still two major drawbacks to overcome: the need for sufficient training data to cover the wide variability of scenes that a UAV may encounter and the strong dependence of the generated models on the specific characteristics of the training samples. In this work, we address these challenges by proposing a novel approach that leverages computer-generated synthetic data alongside novel modifications to the You Only Look Once (YOLO) architecture that enhance its robustness, adaptability to new environments, and accuracy in detecting small targets. Our method introduces a new patch-sample extraction technique and task-specific data augmentation, ensuring robust performance across diverse weather conditions. The results demonstrate our proposal’s superiority, showing an average 28% relative improvement in mean Average Precision (mAP) over the best-performing state-of-the-art baseline under training conditions with sufficient real data, and a remarkable 218% improvement when real data is limited. The proposal also presents a favorable balance between efficiency, effectiveness, and resource requirements. • Small target detection architecture for rapid and precise maritime SAR missions. • Fusion of real and synthetic data simulating real imagery, boosting model robustness. • Transforming data to simulate weather conditions like rain, fog, and sunsets. • In-depth hyperparameter analysis, effect of data scarcity, and data combination. • Proven efficiency in variable weather scenarios, comparison with state of the art.
Juan Pedro Martinez-Esteso, Francisco J. Castellanos 0001, Adrian Rosello, Jorge Calvo-Zaragoza, Antonio Javier Gallego 0001
Eng. Appl. Artif. Intell.5
2025 Enhancing music score analysis with Monte Carlo dropout: a probabilistic approach to staff-region detection
abstract
Abstract Layout Analysis (LA) is a critical process for detecting and isolating different components within a scanned document, allowing for more straightforward and precise processing of each part independently. In Optical Music Recognition (OMR), LA is essential for identifying and extracting music staves, which enables effective music notation recognition and processing. While the literature includes several studies exploring methods for staff retrieval, there remains room for improvement in terms of robustness and accuracy. In this work, we introduce a methodology that integrates Monte Carlo Dropout (MCD) into a neural network model in order to improve reliability in staff retrieval from scanned sheet music. Our approach leverages multiple non-deterministic predictions using standard dropout layers during inference and aggregates them through pixel-level combination policies. We extend the MCD technique, originally designed for classification and regression tasks using averaged predictions, to the LA task and introduce new combination strategies: maximum and voting criteria. Experiments on three diverse music score corpora, including printed and handwritten documents, demonstrated the effectiveness of our approach. The averaging and voting (with 25% and 50% of votes) criteria reduced the relative error by 63.6% compared to the baseline and achieved a 32.1% improvement over state-of-the-art methods. Our methodology notably enhanced detection accuracy without requiring modifications to the neural architecture, especially at the edges of staves, where conventional models tend to show higher error rates.
Samuel B. Oliva-Bulpitt, Juan Pedro Martinez-Esteso, Alejandro Galán-Cuenca, Francisco J. Castellanos 0001, Antonio Javier Gallego 0001
Int. J. Document Anal. Recognit.5
2024 A Region-Based Approach for Layout Analysis of Music Score Images in Scarce Data Scenarios
Francisco J. Castellanos 0001, Juan Pedro Martinez-Esteso, Alejandro Galán-Cuenca, Antonio Javier Gallego 0001
ICDAR (4)4
2024 Multi-label logo recognition and retrieval based on weighted fusion of neural features
abstract
Abstract Classifying logo images is a challenging task as they contain elements such as text or shapes that can represent anything from known objects to abstract shapes. While the current state of the art for logo classification addresses the problem as a multi‐class task focusing on a single characteristic, logos can have several simultaneous labels, such as different colours. This work proposes a method that allows visually similar logos to be classified and searched from a set of data according to their shape, colour, commercial sector, semantics, general characteristics, or a combination of features selected by the user. Unlike previous approaches, the proposal employs a series of multi‐label deep neural networks specialized in specific attributes and combines the obtained features to perform the similarity search. To delve into the classification system, different existing logo topologies are compared and some of their problems are analysed, such as the incomplete labelling that trademark registration databases usually contain. The proposal is evaluated considering 76,000 logos (seven times more than previous approaches) from the European Union Trademarks dataset, which is organized hierarchically using the Vienna ontology. Overall, experimentation attains reliable quantitative and qualitative results, reducing the normalized average rank error of the state‐of‐the‐art from 0.040 to 0.018 for the Trademark Image Retrieval task. Finally, given that the semantics of logos can often be subjective, graphic design students and professionals were surveyed. Results show that the proposed methodology provides better labelling than a human expert operator, improving the label ranking average precision from 0.53 to 0.68.
Marisa Bernabeu, Antonio Javier Gallego 0001, Antonio Pertusa
Expert Syst. J. Knowl. Eng.2
2024 An overview of ensemble and feature learning in few-shot image classification using siamese networks
abstract
Abstract Siamese Neural Networks (SNNs) constitute one of the most representative approaches for addressing Few-Shot Image Classification. These schemes comprise a set of Convolutional Neural Network (CNN) models whose weights are shared across the network, which results in fewer parameters to train and less tendency to overfit. This fact eventually leads to better convergence capabilities than standard neural models when considering scarce amounts of data. Based on a contrastive principle, the SNN scheme jointly trains these inner CNN models to map the input image data to an embedded representation that may be later exploited for the recognition process. However, in spite of their extensive use in the related literature, the representation capabilities of SNN schemes have neither been thoroughly assessed nor combined with other strategies for boosting their classification performance. Within this context, this work experimentally studies the capabilities of SNN architectures for obtaining a suitable embedded representation in scenarios with a severe data scarcity, assesses the use of train data augmentation for improving the feature learning process, introduces the use of transfer learning techniques for further exploiting the embedded representations obtained by the model, and uses test data augmentation for boosting the performance capabilities of the SNN scheme by mimicking an ensemble learning process. The results obtained with different image corpora report that the combination of the commented techniques achieves classification rates ranging from 69% to 78% with just 5 to 20 prototypes per class whereas the CNN baseline considered is unable to converge. Furthermore, upon the convergence of the baseline model with the sufficient amount of data, still the adequate use of the studied techniques improves the accuracy in figures from 4% to 9%.
Jose J. Valero-Mas, Antonio Javier Gallego 0001, Juan Ramón Rico-Juan
Multim. Tools Appl.2
2024 Global point cloud registration network for large transformations
abstract
Abstract Three-dimensional registration is an established yet challenging problem that is key in many different applications, such as mapping the environment for autonomous vehicles, or modeling people for avatar creation, among others. Registration refers to the process of mapping multiple data into the same coordinate system by means of matching correspondences and transformation estimation. Novel proposals exploit the benefits of deep learning architectures for this purpose, as they learn the best features for the data, providing better matches and hence results. However, the state of the art is usually focused on cases of relatively small transformations, although in certain applications and in a real and practical environment, large transformations are very common. In this paper, we present ReLaTo (Registration for Large Transformations), an architecture that addresses the cases where large transformations happen while maintaining good performance for local transformations. This proposal uses a novel Softmax pooling layer to find correspondences in a bilateral consensus manner between two point sets, sampling the most confident matches. These matches estimate a coarse and global registration using weighted Singular Value Decomposition (SVD). A target-guided denoising step is applied to both the obtained matches and latent features to estimate the final fine registration considering the local geometry. All these steps are carried out following an end-to-end approach, which has been shown to perform better than 7 state-of-the-art registration methods in two datasets commonly used for this task (ModelNet40 and the Karlsruhe Institute of Technology and Toyota Technological Institute dataset, KITTI), especially in the case of large transformations. Graphic abstract
Hanz Cuevas-Velasquez, Alejandro Galán-Cuenca, Antonio Javier Gallego 0001, Marcelo Saval-Calvo, Robert B. Fisher
Pattern Anal. Appl.3
2024 Few-shot learning for COVID-19 chest X-ray classification with imbalanced data: an inter vs. intra domain study
abstract
Abstract Medical image datasets are essential for training models used in computer-aided diagnosis, treatment planning, and medical research. However, some challenges are associated with these datasets, including variability in data distribution, data scarcity, and transfer learning issues when using models pre-trained from generic images. This work studies the effect of these challenges at the intra- and inter-domain level in few-shot learning scenarios with severe data imbalance. For this, we propose a methodology based on Siamese neural networks in which a series of techniques are integrated to mitigate the effects of data scarcity and distribution imbalance. Specifically, different initialization and data augmentation methods are analyzed, and four adaptations to Siamese networks of solutions to deal with imbalanced data are introduced, including data balancing and weighted loss, both separately and combined, and with a different balance of pairing ratios. Moreover, we also assess the inference process considering four classifiers, namely Histogram, k NN, SVM, and Random Forest. Evaluation is performed on three chest X-ray datasets with annotated cases of both positive and negative COVID-19 diagnoses. The accuracy of each technique proposed for the Siamese architecture is analyzed separately. The results are compared to those obtained using equivalent methods on a state-of-the-art CNN, achieving an average F1 improvement of up to 3.6%, and up to 5.6% of F1 for intra-domain cases. We conclude that the introduced techniques offer promising improvements over the baseline in almost all cases and that the technique selection may vary depending on the amount of data available and the level of imbalance.
Alejandro Galán-Cuenca, Antonio Javier Gallego 0001, Marcelo Saval-Calvo, Antonio Pertusa
Pattern Anal. Appl.2
2024 Guest Editorial: Special issue on IbPRIA 2023
Antonio Javier Gallego 0001, Manuel J. Marín-Jiménez, Raquel Justo, Hélder Oliveira, Antonio Pertusa
Pattern Anal. Appl.1
2024 Efficient multi-task progressive learning for semantic segmentation and disparity estimation
abstract
Scene understanding is an important area in robotics and autonomous driving. To accomplish these tasks, the 3D structures in the scene have to be inferred to know what the objects and their locations are. To this end, semantic segmentation and disparity estimation networks are typically used, but running them individually is inefficient since they require high-performance resources. A possible solution is to learn both tasks together using a multi-task approach. Some current methods address this problem by learning semantic segmentation and monocular depth together. However, monocular depth estimation from single images is an ill-posed problem. A better solution is to estimate the disparity between two stereo images and take advantage of this additional information to improve the segmentation. This work proposes an efficient multi-task method that jointly learns disparity and semantic segmentation. Employing a Siamese backbone architecture for multi-scale feature extraction, the method integrates specialized branches for disparity estimation and coarse and refined segmentations, leveraging progressive task-specific feature sharing and attention mechanisms to enhance accuracy for solving both tasks concurrently. The proposal achieves state-of-the-art results for joint segmentation and disparity estimation on three distinct datasets: Cityscapes, TrimBot2020 Garden, and S-ROSeS, using only 1/3 of the parameters of previous approaches.
Hanz Cuevas-Velasquez, Alejandro Galán-Cuenca, Robert B. Fisher, Antonio Javier Gallego 0001
Pattern Recognit.4
2024 Editorial: Special session on IbPRIA 2023
Antonio Javier Gallego 0001, Manuel J. Marín-Jiménez, Raquel Justo, Hélder Oliveira, Antonio Pertusa
Pattern Recognit. Lett.1
2023 Kurcuma: a kitchen utensil recognition collection for unsupervised domain adaptation
abstract
Abstract The use of deep learning makes it possible to achieve extraordinary results in all kinds of tasks related to computer vision. However, this performance is strongly related to the availability of training data and its relationship with the distribution in the eventual application scenario. This question is of vital importance in areas such as robotics, where the targeted environment data are barely available in advance. In this context, domain adaptation (DA) techniques are especially important to building models that deal with new data for which the corresponding label is not available. To promote further research in DA techniques applied to robotics, this work presents Kurcuma (Kitchen Utensil Recognition Collection for Unsupervised doMain Adaptation), an assortment of seven datasets for the classification of kitchen utensils—a task of relevance in home-assistance robotics and a suitable showcase for DA. Along with the data, we provide a broad description of the main characteristics of the dataset, as well as a baseline using the well-known domain-adversarial training of neural networks approach. The results show the challenge posed by DA on these types of tasks, pointing to the need for new approaches in future work.
Adrian Rosello, Jose J. Valero-Mas, Antonio Javier Gallego 0001, Javier Sáez-Pérez, Jorge Calvo-Zaragoza
Pattern Anal. Appl.3
2023 Multilabel Prototype Generation for data reduction in K-Nearest Neighbour classification
abstract
Prototype Generation (PG) methods are typically considered for improving the efficiency of the k-Nearest Neighbour (kNN) classifier when tackling high-size corpora. Such approaches aim at generating a reduced version of the corpus without decreasing the classification performance when compared to the initial set. Despite their large application in multiclass scenarios, very few works have addressed the proposal of PG methods for the multilabel space. In this regard, this work presents the novel adaptation of four multiclass PG strategies to the multilabel case. These proposals are evaluated with three multilabel kNN-based classifiers, 12 corpora comprising a varied range of domains and corpus sizes, and different noise scenarios artificially induced in the data. The results obtained show that the proposed adaptations are capable of significantly improving—both in terms of efficiency and classification performance—the only reference multilabel PG work in the literature as well as the case in which no PG method is applied, also presenting statistically superior robustness in noisy scenarios. Moreover, these novel PG strategies allow prioritising either the efficiency or efficacy criteria through its configuration depending on the target scenario, hence covering a wide area in the solution space not previously filled by other works.
Jose J. Valero-Mas, Antonio Javier Gallego 0001, Pablo Alonso-Jiménez, Xavier Serra
Pattern Recognit.2
2023 An experimental study on marine debris location and recognition using object detection
abstract
The large amount of debris in our oceans is a global problem that dramatically impacts marine fauna and flora. While a large number of human-based campaigns have been proposed to tackle this issue, these efforts have been deemed insufficient due to the insurmountable amount of existing litter. In response to that, there exists a high interest in the use of autonomous underwater vehicles (AUV) that may locate, identify, and collect this garbage automatically. To perform such a task, AUVs consider state-of-the-art object detection techniques based on deep neural networks due to their reported high performance. Nevertheless, these techniques generally require large amounts of data with fine-grained annotations. In this work, we explore the capabilities of the reference object detector Mask Region-based Convolutional Neural Networks for automatic marine debris location and classification in the context of limited data availability. Considering the recent CleanSea corpus, we pose several scenarios regarding the amount of available train data and study the possibility of mitigating the adverse effects of data scarcity with synthetic marine scenes. Our results achieve a new state of the art in the task, establishing a new reference for future research. In addition, it is shown that the task still has room for improvement and that the lack of data can be somehow alleviated, yet to a limited extent.
Alejandro Sánchez-Ferrer, Jose J. Valero-Mas, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
Pattern Recognit. Lett.3
2022 Efficient gesture recognition for the assistance of visually impaired people using multi-head neural networks
abstract
Existing research for the assistance of visually impaired people mainly focus on solving a single task (such as reading a text or detecting an obstacle), hence forcing the user to switch applications to perform other actions. This paper proposes an interactive system for mobile devices controlled by hand gestures that allow the user to control the device and use several assistance tools by making simple static and dynamic hand gestures (e.g., pointing a finger at an object will show a description of it). The system is based on a multi-head neural network, which initially detects and classifies the gestures, and subsequently, depending on the gesture detected, performs a second stage that carries out the corresponding action. This architecture optimizes the resources required to perform different tasks, it takes advantage of the information obtained from an initial backbone to perform different processes in a second stage. To train and evaluate the system, a dataset with about 40k images was manually compiled and labeled including different types of hand gestures, backgrounds (indoors and outdoors), lighting conditions, etc. This dataset contains synthetic gestures (whose objective is to pre-train the system to improve the results) and real images captured using different mobile phones. The comparison made with nearly 50 state-of-the-art methods shows competitive results as regards the different actions performed by the system, such as the accuracy of classification and localization of gestures, or the generation of descriptions for objects and scenes.
Samer Alashhab, Antonio Javier Gallego 0001, Miguel Angel Lozano
Eng. Appl. Artif. Intell.2
2022 Domain adaptation for staff-region retrieval of music score images
abstract
Abstract Optical music recognition (OMR) is the field that studies how to automatically read music notation from score images. One of the relevant steps within the OMR workflow is the staff-region retrieval. This process is a key step because any undetected staff will not be processed by the subsequent steps. This task has previously been addressed as a supervised learning problem in the literature; however, ground-truth data are not always available, so each new manuscript requires a preliminary manual annotation. This situation is one of the main bottlenecks in OMR, because of the countless number of existing manuscripts , and the associated manual labeling cost. With the aim of mitigating this issue, we propose the application of a domain adaptation technique, the so-called Domain-Adversarial Neural Network (DANN), based on a combination of a gradient reversal layer and a domain classifier in the inference neural architecture. The results from our experiments support the benefits of our proposed solution, obtaining improvements of approximately 29% in the F-score.
Francisco J. Castellanos 0001, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza, Ichiro Fujinaga
Int. J. Document Anal. Recognit.2
2022 Efficient k-nearest neighbor search based on clustering and adaptive k values
abstract
The k -Nearest Neighbor ( k NN) algorithm is widely used in the supervised learning field and, particularly, in search and classification tasks , owing to its simplicity, competitive performance, and good statistical properties. However, its inherent inefficiency prevents its use in most modern applications due to the vast amount of data that the current technological evolution generates, being thus the optimization of k NN-based search strategies of particular interest. This paper introduces the caKD+ algorithm, which tackles this limitation by combining the use of feature learning techniques, clustering methods , adaptive search parameters per cluster, and the use of pre-calculated K-Dimensional Tree structures, and results in a highly efficient search method. This proposal has been evaluated using 10 datasets and the results show that caKD+ significantly outperforms 16 state-of-the-art efficient search methods while still depicting such an accurate performance as the one by the exhaustive k NN search.
Antonio Javier Gallego 0001, Juan Ramón Rico-Juan, Jose J. Valero-Mas
Pattern Recognit.1
2021 Two Heads are Better than One: Geometric-Latent Attention for Point Cloud Classification and Segmentation
Hanz Cuevas-Velasquez, Antonio Javier Gallego 0001, Robert B. Fisher
BMVC2
2021 Unsupervised neural domain adaptation for document image binarization
Francisco J. Castellanos 0001, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
Pattern Recognit.2
2021 Incremental Unsupervised Domain-Adversarial Training of Neural Networks
abstract
In the context of supervised statistical learning, it is typically assumed that the training set comes from the same distribution that draws the test samples. When this is not the case, the behavior of the learned model is unpredictable and becomes dependent upon the degree of similarity between the distribution of the training set and the distribution of the test set. One of the research topics that investigates this scenario is referred to as domain adaptation (DA). Deep neural networks brought dramatic advances in pattern recognition and that is why there have been many attempts to provide good DA algorithms for these models. Herein we take a different avenue and approach the problem from an incremental point of view, where the model is adapted to the new domain iteratively. We make use of an existing unsupervised domain-adaptation algorithm to identify the target samples on which there is greater confidence about their true label. The output of the model is analyzed in different ways to determine the candidate samples. The selected samples are then added to the source training set by self-labeling, and the process is repeated until all target samples are labeled. This approach implements a form of adversarial training in which, by moving the self-labeled samples from the target to the source set, the DA algorithm is forced to look for new features after each iteration. Our results report a clear improvement with respect to the non-incremental case in several data sets, also outperforming other state-of-the-art DA algorithms.
Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza, Robert B. Fisher
IEEE Trans. Neural Networks Learn. Syst.1
2020 Real-time Stereo Visual Servoing for Rose Pruning with Robotic Arm
abstract
The paper presents a working pipeline which integrates hardware and software in an automated robotic rose cutter. To the best of our knowledge, this is the first robot able to prune rose bushes in a natural environment. Unlike similar approaches like tree stem cutting, the proposed method does not require to scan the full plant, have multiple cameras around the bush, or assume that a stem does not move. It relies on a single stereo camera mounted on the end-effector of the robot and real-time visual servoing to navigate to the desired cutting location on the stem. The evaluation of the whole pipeline shows a good performance in a garden with unconstrained conditions, where finding and approaching a specific location on a stem is challenging due to occlusions caused by other stems and dynamic changes caused by the wind.
Hanz Cuevas-Velasquez, Antonio Javier Gallego 0001, Radim Tylecek, Jochen Hemming, B. A. J. van Tuijl, Angelo Mencarelli, Robert B. Fisher
ICRA2
2020 Automatic scale estimation for music score images
abstract
Optical Music Recognition (OMR) is the research field focused on the automatic reading of music from scanned images. Its main goal is to encode the content into a digital and structured format with the advantages that this entails. This discipline is traditionally aligned to a workflow whose first step is the document analysis. This step is responsible of recognizing and detecting different sources of information—e.g. music notes, staff lines and text—to extract them and then processing automatically the content in the following steps of the workflow. One of the most difficult challenges it faces is to provide a generic solution to analyze documents with diverse resolutions. The endless number of existing music sources does not meet a standard that normalizes the data collections, giving complete freedom for a wide variety of image sizes and scales, thereby making this operation unsustainable. In the literature, this question is commonly overlooked and a uniform scale is assumed. In this paper, a machine learning-based approach to estimate the scale of music documents with respect to a reference scale is presented. Our goal is to propose a robust and generalizable method to adapt the input image to the requirements of an OMR system. For this, two goal-directed case studies are included to evaluate the proposed approach over common task within the OMR workflow, comparing the behavior with other state-of-the-art methods. Results suggest that it is necessary to perform this additional step in the first stage of the workflow to correct the scale of the input images. In addition, it is empirically demonstrated that our specialized approach is more promising than image augmentation strategies for the multi-scale challenge.
Francisco J. Castellanos 0001, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
Expert Syst. Appl.2
2020 Ensemble classification from deep predictions with test data augmentation
Jorge Calvo-Zaragoza, Juan Ramón Rico-Juan, Antonio Javier Gallego 0001
Soft Comput.3
2019 A selectional auto-encoder approach for document image binarization
Jorge Calvo-Zaragoza, Antonio Javier Gallego 0001
Pattern Recognit.2
2018 Multimodal Object Recognition Using Deep Learning Representations Extracted from Images and Smartphone Sensors
Javier Ortega Bastida, Antonio Javier Gallego 0001, Antonio Pertusa
CIARP2
2018 Data Augmentation via Variational Auto-Encoders
Unai Garay-Maestre, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
CIARP2
2018 MirBot: A collaborative object recognition system for smartphones using convolutional neural networks
Antonio Pertusa, Antonio Javier Gallego 0001, Marisa Bernabeu
Neurocomputing2
2018 Grammatical inference of directed acyclic graph languages with polynomial time complexity
Antonio Javier Gallego 0001, Damián López, Jorge Calera-Rubio
J. Comput. Syst. Sci.1
2018 Clustering-based k-nearest neighbor classification for large-scale data with neural codes representation
Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza, Jose J. Valero-Mas, Juan Ramón Rico-Juan
Pattern Recognit.1
2018 Two-Stage Convolutional Neural Network for Ship and Spill Detection Using SLAR Images
abstract
This paper presents a system for the detection of ships and oil spills using side-looking airborne radar (SLAR) images. The proposed method employs a two-stage architecture composed of three pairs of convolutional neural networks (CNNs). Each pair of networks is trained to recognize a single class (ship, oil spill, and coast) by following two steps: a first network performs a coarse detection, and then, a second specialized CNN obtains the precise localization of the pixels belonging to each class. After classification, a postprocessing stage is performed by applying a morphological opening filter in order to eliminate small look-alikes, and removing those oil spills and ships that are surrounded by a minimum amount of coast. Data augmentation is performed to increase the number of samples, owing to the difficulty involved in obtaining a sufficient number of correctly labeled SLAR images. The proposed method is evaluated and compared to a single multiclass CNN architecture and to previous state-of-the-art methods using accuracy, precision, recall, F-measure, and intersection over union. The results show that the proposed method is efficient and competitive, and outperforms the approaches previously used for this task.
Mario Nieto-Hidalgo, Antonio Javier Gallego 0001, Pablo Gil, Antonio Pertusa
IEEE Trans. Geosci. Remote. Sens.2
2017 Recognition of Handwritten Music Symbols with Convolutional Neural Codes
abstract
There are large collections of music manuscripts preserved over the centuries. In order to analyze these documents it is necessary to transcribe them into a machine-readable format. This process can be done automatically using Optical Music Recognition (OMR) systems, which typically consider segmentation plus classification workflows. This work is focused on the latter stage, presenting a comprehensive study for classification of handwritten musical symbols using Convolutional Neural Networks (CNN). The power of these models lies in their ability to transform the input into a meaningful representation for the task at hand, and that is why we study the use of these models to extract features (Neural Codes) for other classifiers. For the evaluation we consider four datasets containing different configurations and notation styles, along with a number of network models, different image preprocessing techniques and several supervised learning classifiers. Our results show that a remarkable accuracy can be achieved using the proposed framework, which significantly outperforms the state of the art in all datasets considered.
Jorge Calvo-Zaragoza, Antonio Javier Gallego 0001, Antonio Pertusa
ICDAR2
2017 Staff-line removal with selectional auto-encoders
Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza
Expert Syst. Appl.1
2014 Construction of Intelligent Virtual Worlds Using a Grammatical Framework
abstract
The potential of integrating multiagent systems and virtual environments has not been exploited to its whole extent. This paper proposes a model based on grammars, called Minerva, to construct complex virtual environments that integrate the features of agents. A virtual world is described as a set of dynamic and static elements. The static part is represented by a sequence of primitives and transformations and the dynamic elements by a series of agents. Agent activation and communication is achieved using events, created by the so-called event generators. The grammar defines a descriptive language with a simple syntax and a semantics, defined by functions. The semantics functions allow the scene to be displayed in a graphics device, and the description of the activities of the agents, including artificial intelligence algorithms and reactions to physical phenomena. To illustrate the use of Minerva, a practical example is presented: a simple robot simulator that considers the basic features of a typical robot. The result is a functional simple simulator. Minerva is a reusable, integral, and generic system, which can be easily scaled, adapted, and improved. The description of the virtual scene is independent from its representation and the elements that it interacts with.
Gabriel López-García, Antonio Javier Gallego 0001, Rafael Molina-Carmona, Patricia Compañ-Rosique
Int. J. Intell. Syst.2
2010 Formal Model to Integrate Multi-Agent Systems and Interactive Graphic Systems
Gabriel López-García, Rafael Molina-Carmona, Antonio Javier Gallego 0001
ICAART (2)3
2009 Improving edge detection in highly noised sheet-metal images
abstract
This article proposes a new method for robust and accurate detection of the orientation and the location of an object on low-contrast surfaces in an industrial context. To be more efficient and effective, our method employs only artificial vision. Therefore, productivity is increased since it avoids the use of additional mechanical devices to ensure the accuracy of the system. The technical core is the treatment of straight line contours that occur in close neighbourhood to each other and with similar orientations. It is a particular problem in stacks of objects but can also occur in other applications. New techniques are introduced to ensure the robustness of the system and to tackle the problem of noise, such as an auto-threshold segmentation process, a new type of histogram and a robust regression method used to compute the result with a higher precision.
Antonio Javier Gallego 0001, Jorge Calera-Rubio
WACV1
2007 Rectified Reconstruction from Stereo Pairs and Robot Mapping
Antonio Javier Gallego 0001, Rafael Molina 0001, Patricia Compañ-Rosique, Carlos Villagrá
CAIP1