VLDB 2026 Research / reviewers in the wild / expert
Francisco J. Castellanos 0001
dblp:216/1535
· DBLP profile ↗
15ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0001-9949-5522ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 7 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing maritime search and rescue: Incremental unsupervised domain adaptation with synthetic data and pseudo-labelingabstractMaritime search and rescue operations are critical for saving lives in emergencies, where time is a decisive factor since delays can drastically reduce the chances of survival for those in distress. These missions are particularly challenging due to the inherent complexity of the maritime environment, marked by changing weather, dynamic sea states, and limited visibility. Developing reliable machine learning systems for this task typically requires large amounts of labeled data that capture all possible operating conditions. However, collecting and annotating such data is costly and often unfeasible in real-world maritime scenarios. To address this limitation, we propose a domain adaptation strategy for a segmentation-based detection model that estimates a probability map indicating the presence of human bodies at sea. The method enables unsupervised learning to adapt from a labeled synthetic domain to a real, unlabeled domain by employing a Domain-Adversarial Neural Network that aligns feature representations across domains, and an iterative pseudo-labeling process that selects high-confidence predictions on the target data to progressively refine the model. By leveraging synthetic data—automatically generated and labeled—our approach adapts effectively to real-world conditions without requiring manual annotation. Experimental results show that our method outperforms several state-of-the-art detectors while maintaining a lightweight architecture. Moreover, it generalizes well under diverse and adverse environmental conditions, including fog, rain, and low-light scenes, demonstrating its robustness and suitability for real-world deployment in critical rescue operations. Juan Pedro Martinez-Esteso, Francisco J. Castellanos 0001, Antonio Javier Gallego 0001 |
Expert Syst. Appl. | 2 |
| 2026 | Insights into imbalance-aware Multilabel Prototype Generation mechanisms for k-Nearest Neighbor classification in noisy scenariosabstractPrototype Generation (PG) techniques enhance the efficiency of the k -Nearest Neighbor ( k NN) classifier by condensing datasets through the use of specific rules. More precisely, these strategies work on the premise of merging the elements in the reference data collection to generate an alternative and more compact data assortment that substitutes the former one without remarkably affecting the recognition performance. Nevertheless, despite being widely studied in multiclass scenarios, PG is still underexplored in multilabel contexts, leading to limitations, notably in the handling of label imbalance and noise. In this regard, this work introduces a reduction framework that allows for multilabel PG methods to handle these challenges of label imbalance and noise. The proposed mechanisms comprise a selection strategy that exclusively preserves noise-free samples in the process, a mechanism to avoid severely imbalanced samples from being inadequately processed, and two new merging policies for the PG methods to generate novel samples. These enhancements are considered along with three established multilabel PG methods: Multilabel Reduction through Homogeneous Clustering, Multilabel Chen, and Multilabel Reduction through Space Partitioning. Evaluations are conducted using three k NN-based multilabel classifiers and 12 diverse datasets with different levels of label imbalance. We additionally study the performance with varying values of k under different label-noise scenarios. The results are assessed through statistical tests and indicate that our proposals outperform the original methods that disregard label imbalance, even in the presence of noise, thus validating these approaches and fostering further research in the field. Jose J. Valero-Mas, Carlos Peñarrubia, Francisco J. Castellanos 0001, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza |
Pattern Recognit. | 3 |
| 2025 | On the use of synthetic data for body detection in maritime search and rescue operationsabstractTime is a critical factor in maritime Search And Rescue (SAR) missions, during which promptly locating survivors is paramount. Unmanned Aerial Vehicles (UAVs) are a useful tool with which to increase the success rate by rapidly identifying targets. While this task can be performed using other means, such as helicopters, the cost-effectiveness of UAVs makes them an effective choice. Moreover, these vehicles allow the easy integration of automatic systems that can be used to assist in the search process. Despite the impact of artificial intelligence on autonomous technology, there are still two major drawbacks to overcome: the need for sufficient training data to cover the wide variability of scenes that a UAV may encounter and the strong dependence of the generated models on the specific characteristics of the training samples. In this work, we address these challenges by proposing a novel approach that leverages computer-generated synthetic data alongside novel modifications to the You Only Look Once (YOLO) architecture that enhance its robustness, adaptability to new environments, and accuracy in detecting small targets. Our method introduces a new patch-sample extraction technique and task-specific data augmentation, ensuring robust performance across diverse weather conditions. The results demonstrate our proposal’s superiority, showing an average 28% relative improvement in mean Average Precision (mAP) over the best-performing state-of-the-art baseline under training conditions with sufficient real data, and a remarkable 218% improvement when real data is limited. The proposal also presents a favorable balance between efficiency, effectiveness, and resource requirements. • Small target detection architecture for rapid and precise maritime SAR missions. • Fusion of real and synthetic data simulating real imagery, boosting model robustness. • Transforming data to simulate weather conditions like rain, fog, and sunsets. • In-depth hyperparameter analysis, effect of data scarcity, and data combination. • Proven efficiency in variable weather scenarios, comparison with state of the art. Juan Pedro Martinez-Esteso, Francisco J. Castellanos 0001, Adrian Rosello, Jorge Calvo-Zaragoza, Antonio Javier Gallego 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Enhancing music score analysis with Monte Carlo dropout: a probabilistic approach to staff-region detectionabstractAbstract Layout Analysis (LA) is a critical process for detecting and isolating different components within a scanned document, allowing for more straightforward and precise processing of each part independently. In Optical Music Recognition (OMR), LA is essential for identifying and extracting music staves, which enables effective music notation recognition and processing. While the literature includes several studies exploring methods for staff retrieval, there remains room for improvement in terms of robustness and accuracy. In this work, we introduce a methodology that integrates Monte Carlo Dropout (MCD) into a neural network model in order to improve reliability in staff retrieval from scanned sheet music. Our approach leverages multiple non-deterministic predictions using standard dropout layers during inference and aggregates them through pixel-level combination policies. We extend the MCD technique, originally designed for classification and regression tasks using averaged predictions, to the LA task and introduce new combination strategies: maximum and voting criteria. Experiments on three diverse music score corpora, including printed and handwritten documents, demonstrated the effectiveness of our approach. The averaging and voting (with 25% and 50% of votes) criteria reduced the relative error by 63.6% compared to the baseline and achieved a 32.1% improvement over state-of-the-art methods. Our methodology notably enhanced detection accuracy without requiring modifications to the neural architecture, especially at the edges of staves, where conventional models tend to show higher error rates. Samuel B. Oliva-Bulpitt, Juan Pedro Martinez-Esteso, Alejandro Galán-Cuenca, Francisco J. Castellanos 0001, Antonio Javier Gallego 0001 |
Int. J. Document Anal. Recognit. | 4 |
| 2024 | Analysis of the Calibration of Handwriting Text Recognition Models
Eric Ayllon, Francisco J. Castellanos 0001, Jorge Calvo-Zaragoza |
ICDAR (2) | 2 |
| 2024 | A Region-Based Approach for Layout Analysis of Music Score Images in Scarce Data Scenarios
Francisco J. Castellanos 0001, Juan Pedro Martinez-Esteso, Alejandro Galán-Cuenca, Antonio Javier Gallego 0001 |
ICDAR (4) | 1 |
| 2023 | A Holistic Approach for Aligned Music and Lyrics Transcription
Juan C. Martinez-Sevilla, Antonio Ríos-Vila, Francisco J. Castellanos 0001, Jorge Calvo-Zaragoza |
ICDAR (1) | 3 |
| 2023 | Multimodal recognition of frustration during game-play with deep neural networksabstractAbstract Frustration, which is one aspect of the field of emotional recognition, is of particular interest to the video game industry as it provides information concerning each individual player’s level of engagement. The use of non-invasive strategies to estimate this emotion is, therefore, a relevant line of research with a direct application to real-world scenarios. While several proposals regarding the performance of non-invasive frustration recognition can be found in literature, they usually rely on hand-crafted features and rarely exploit the potential inherent to the combination of different sources of information. This work, therefore, presents a new approach that automatically extracts meaningful descriptors from individual audio and video sources of information using Deep Neural Networks (DNN) in order to then combine them, with the objective of detecting frustration in Game-Play scenarios. More precisely, two fusion modalities, namelydecision-levelandfeature-level, are presented and compared with state-of-the-art methods, along with different DNN architectures optimized for each type of data. Experiments performed with a real-world audiovisual benchmarking corpus revealed that the multimodal proposals introduced herein are more suitable than those of a unimodal nature, and that their performance also surpasses that of other state-of-the–art approaches, with error rate improvements of between 40%and 90%. Carlos de la Fuente, Francisco J. Castellanos 0001, Jose J. Valero-Mas, Jorge Calvo-Zaragoza |
Multim. Tools Appl. | 2 |
| 2022 | Continual Learning for Document Image BinarizationabstractIn the field of Document Image Analysis (DIA), it is common to find great heterogeneity in terms of the possible graphic domains. In this sense, it is interesting to build neural models that can be sequentially adapted to new domains without losing the knowledge from the domains already learned. This learning paradigm is known as Continual (or Lifelong) Learning (CL). Although the adaptation comes along with a training set of the new domain, neural networks suffer what is known as "catastrophic forgetting". Therefore, assuming the constraint of not keeping data from the domains already addressed, this paradigm represents a challenge yet to be solved. This work presents an approach for CL in document image binarization, one of the most considered tasks within the DIA field. Our results report that it is indeed feasible to address CL in this field, given that the approach is successfully implemented and outperforms the baseline by a wide margin in most of the analyzed scenarios. Carlos Garrido-Munoz, Adrián Sánchez-Hernández, Francisco J. Castellanos 0001, Jorge Calvo-Zaragoza |
ICPR | 3 |
| 2022 | Region-based layout analysis of music score imagesabstractThe Layout Analysis (LA) stage is of vital importance to the correct performance of an Optical Music Recognition (OMR) system. It identifies the regions of interest, such as staves or lyrics, which must then be processed in order to transcribe their content. Despite the existence of modern approaches based on deep learning, an exhaustive study of LA in OMR has not yet been carried out with regard to the performance of different models, their generalization to different domains or, more importantly, their impact on subsequent stages of the pipeline. This work focuses on filling this gap in the literature by means of an experimental study of different neural architectures, music document types, and evaluation scenarios. The need for training data has also led to a proposal for a new semi-synthetic data-generation technique that enables the efficient applicability of LA approaches in real scenarios. Our results show that: (i) the choice of the model and its performance are crucial for the entire transcription process; (ii) the metrics commonly used to evaluate the LA stage do not always correlate with the final performance of the OMR system, and (iii) the proposed data-generation technique enables state-of-the-art results to be achieved with a limited set of labeled data. Francisco J. Castellanos 0001, Carlos Garrido-Munoz, Antonio Ríos-Vila, Jorge Calvo-Zaragoza |
Expert Syst. Appl. | 1 |
| 2022 | Domain adaptation for staff-region retrieval of music score imagesabstractAbstract Optical music recognition (OMR) is the field that studies how to automatically read music notation from score images. One of the relevant steps within the OMR workflow is the staff-region retrieval. This process is a key step because any undetected staff will not be processed by the subsequent steps. This task has previously been addressed as a supervised learning problem in the literature; however, ground-truth data are not always available, so each new manuscript requires a preliminary manual annotation. This situation is one of the main bottlenecks in OMR, because of the countless number of existing manuscripts , and the associated manual labeling cost. With the aim of mitigating this issue, we propose the application of a domain adaptation technique, the so-called Domain-Adversarial Neural Network (DANN), based on a combination of a gradient reversal layer and a domain classifier in the inference neural architecture. The results from our experiments support the benefits of our proposed solution, obtaining improvements of approximately 29% in the F-score. Francisco J. Castellanos 0001, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza, Ichiro Fujinaga |
Int. J. Document Anal. Recognit. | 1 |
| 2021 | Unsupervised neural domain adaptation for document image binarization
Francisco J. Castellanos 0001, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza |
Pattern Recognit. | 1 |
| 2021 | Prototype generation in the string space via approximate median for data reduction in nearest neighbor classificationabstractAbstract The k-nearest neighbor (kNN) rule is one of the best-known distance-based classifiers, and is usually associated with high performance and versatility as it requires only the definition of a dissimilarity measure. Nevertheless, kNN is also coupled with low-efficiency levels since, for each new query, the algorithm must carry out an exhaustive search of the training data, and this drawback is much more relevant when considering complex structural representations, such as graphs, trees or strings, owing to the cost of the dissimilarity metrics. This issue has generally been tackled through the use of data reduction (DR) techniques, which reduce the size of the reference set, but the complexity of structural data has historically limited their application in the aforementioned scenarios. A DR algorithm denominated as reduction through homogeneous clusters (RHC) has recently been adapted to string representations but as obtaining the exact median value of a set of string data is known to be computationally difficult, its authors resorted to computing the set-median value. Under the premise that a more exact median value may be beneficial in this context, we, therefore, present a new adaptation of the RHC algorithm for string data, in which an approximate median computation is carried out. The results obtained show significant improvements when compared to those of the set-median version of the algorithm, in terms of both classification performance and reduction rates. Francisco J. Castellanos 0001, Jose J. Valero-Mas, Jorge Calvo-Zaragoza |
Soft Comput. | 1 |
| 2020 | Automatic scale estimation for music score imagesabstractOptical Music Recognition (OMR) is the research field focused on the automatic reading of music from scanned images. Its main goal is to encode the content into a digital and structured format with the advantages that this entails. This discipline is traditionally aligned to a workflow whose first step is the document analysis. This step is responsible of recognizing and detecting different sources of information—e.g. music notes, staff lines and text—to extract them and then processing automatically the content in the following steps of the workflow. One of the most difficult challenges it faces is to provide a generic solution to analyze documents with diverse resolutions. The endless number of existing music sources does not meet a standard that normalizes the data collections, giving complete freedom for a wide variety of image sizes and scales, thereby making this operation unsustainable. In the literature, this question is commonly overlooked and a uniform scale is assumed. In this paper, a machine learning-based approach to estimate the scale of music documents with respect to a reference scale is presented. Our goal is to propose a robust and generalizable method to adapt the input image to the requirements of an OMR system. For this, two goal-directed case studies are included to evaluate the proposed approach over common task within the OMR workflow, comparing the behavior with other state-of-the-art methods. Results suggest that it is necessary to perform this additional step in the first stage of the workflow to correct the scale of the input images. In addition, it is empirically demonstrated that our specialized approach is more promising than image augmentation strategies for the multi-scale challenge. Francisco J. Castellanos 0001, Antonio Javier Gallego 0001, Jorge Calvo-Zaragoza |
Expert Syst. Appl. | 1 |
| 2018 | Oversampling imbalanced data in the string space
Francisco J. Castellanos 0001, Jose J. Valero-Mas, Jorge Calvo-Zaragoza, Juan Ramón Rico-Juan |
Pattern Recognit. Lett. | 1 |