Marcos Escudero-Viñolo

dblp:123/9822 · also Marcos Escudero · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-9156-3428ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Soft-Labelling for Budget-Constrained Semantic Segmentation: Bringing Coherence to label Down-Sampling
abstract
In semantic segmentation, training data down-sampling is commonly performed due to resource limitations, the need to adapt image size to the model input, or to improve data augmentation. This down-sampling typically employs different strategies for the image data and the annotated labels. Such discrepancy leads to mismatches between the down-sampled colour and ground-truth label images. Hence, the training performance significantly decreases as the down-sampling factor increases. In this paper, we bring together the down-sampling strategies for the image data and the training labels. To that aim, we propose a novel framework for label down-sampling via soft-labelling that better conserves label information after down-sampling, thereby, fully aligning soft-labels with image data to keep the distribution of the sampled pixels for down-sampling. This proposal also produces reliable annotations for under-represented semantic classes. Altogether, it allows training competitive models at lower resolutions. Experiments show that our proposal outperforms other down-sampling strategies. Moreover, state-of-the-art performance is achieved for reference benchmarks, but employing significantly fewer computational resources than foremost methods. This proposal enables competitive research for semantic segmentation under resource constraints.
Roberto Alcover-Couso, Marcos Escudero-Viñolo, Juan C. SanMiguel, José María Martínez Sanchez
Comput. Vis. Media2
2025 Gradient-based class weighting for unsupervised domain adaptation in dense prediction visual tasks
abstract
In unsupervised domain adaptation (UDA), where models are trained on source data (e.g., synthetic) and adapted to target data (e.g., real-world) without target annotations, addressing the challenge of significant class imbalance remains an open issue. Despite progress in bridging the domain gap, existing methods often experience performance degradation when confronted with highly imbalanced dense prediction visual tasks like semantic segmentation. This discrepancy becomes especially pronounced due to the lack of equivalent priors between the source and target domains, turning class imbalanced techniques used for other areas (e.g., image classification) ineffective in UDA scenarios. This paper proposes a class-imbalance mitigation strategy that incorporates class-weights into the UDA learning losses, with the novelty of estimating these weights dynamically through the gradients of the per-class losses, defining a Gradient-based class weighting (GBW) approach. The proposed GBW naturally increases the contribution of classes whose learning is hindered by highly-represented classes, and has the advantage of automatically adapting to training outcomes, avoiding explicit curricular learning patterns common in loss-weighing strategies. Extensive experimentation validates the effectiveness of GBW across architectures (Convolutional and Transformer), UDA strategies (adversarial, self-training and entropy minimization), tasks (semantic and panoptic segmentation), and datasets. Analysis shows that GBW consistently increases the recall of under-represented classes. • A novel class-imbalance algorithm that uses per-class gradients to assign effective class weights (GBW). • Class weights are computed through an optimization process that maximizes the decrease in the loss function. • The proposed method demonstrates significant and consistent performance improvements across various UDA methods. • The class weights provide a per-class complexity measure throughout training, establishing an automatic and adaptable curriculum.
Roberto Alcover-Couso, Marcos Escudero-Viñolo, Juan C. SanMiguel, Jesús Bescós
Pattern Recognit.2
2025 Per-class curriculum for Unsupervised Domain Adaptation in semantic segmentation
abstract
Abstract Accurate training of deep neural networks for semantic segmentation requires a large number of pixel-level annotations of real images, which are expensive to generate or not even available. In this context, Unsupervised Domain Adaptation (UDA) can transfer knowledge from unlimited synthetic annotations to unlabeled real images of a given domain. UDA methods are composed of an initial training stage with labeled synthetic data followed by a second stage for feature alignment between labeled synthetic and unlabeled real data. In this paper, we propose a novel approach for UDA focusing the initial training stage, which leads to increased performance after adaptation. We introduce a curriculum strategy where each semantic class is learned progressively. Thereby, better features are obtained for the second stage. This curriculum is based on: (1) a class-scoring function to determine the difficulty of each semantic class, (2) a strategy for incremental learning based on scoring and pacing functions that limits the required training time unlike standard curriculum-based training and (3) a training loss to operate at class level. We extensively evaluate our approach as the first stage of several state-of-the-art UDA methods for semantic segmentation. Our results demonstrate significant performance enhancements across all methods: improvements of up to 10% for entropy-based techniques and 8% for adversarial methods. These findings underscore the dependency of UDA on the accuracy of the initial training. The implementation is available at https://github.com/vpulab/PCCL .
Roberto Alcover-Couso, Juan C. SanMiguel, Marcos Escudero-Viñolo, Pablo Carballeira
Vis. Comput.3
2025 Layer-wise model merging for unsupervised domain adaptation in segmentation tasks
abstract
Abstract Merging parameters of multiple models has resurfaced as an effective strategy to enhance task performance and robustness, but prior work is limited by the high costs of ensemble creation and inference. In this paper, we leverage the abundance of freely accessible trained models to introduce a cost-free approach to model merging. It focuses on a layer-wise integration of merged models, aiming to maintain the distinctiveness of the task-specific final layers while unifying the initial layers, which are primarily associated with feature extraction. This approach ensures parameter consistency across all layers, essential for boosting performance. Moreover, it facilitates seamless integration of knowledge, enabling effective merging of models from different datasets and tasks. Specifically, we investigate its applicability in unsupervised domain adaptation (UDA), an unexplored area for model merging, for semantic and panoptic segmentation. Experimental results demonstrate substantial UDA improvements without additional costs for merging same-architecture models from distinct datasets ( $$\uparrow 2.6\%$$ ↑ 2.6 % mIoU) and different-architecture models with a shared backbone ( $$\uparrow 6.8\%$$ ↑ 6.8 % mIoU). Furthermore, merging semantic and panoptic segmentation models increases mPQ by 7%. These findings are validated across a wide variety of UDA strategies, architectures and datasets. The code will be publicly available upon acceptance in the LWMM repository: http://www-vpu.eps.uam.es/LWMM/ .
Roberto Alcover-Couso, Juan C. SanMiguel, Marcos Escudero-Viñolo, Jose M. Martínez
Vis. Comput.3
2024 Improved transferability of self-supervised learning models through batch normalization finetuning
Kirill Sirotkin, Marcos Escudero-Viñolo, Pablo Carballeira, Álvaro García-Martín
Appl. Intell.2
2023 On exploring weakly supervised domain adaptation strategies for semantic segmentation using synthetic data
abstract
Abstract Pixel-wise image segmentation is key for many Computer Vision applications. The training of deep neural networks for this task has expensive pixel-level annotation requirements, thus, motivating a growing interest on synthetic data to provide unlimited data and its annotations. In this paper, we focus on the generation and application of synthetic data as representative training corpuses for semantic segmentation of urban scenes. First, we propose a synthetic data generation protocol, which identifies key features affecting performance and provides datasets with variable complexity. Second, we adapt two popular weakly supervised domain adaptation approaches (combined training, fine-tuning) to employ synthetic and real data. Moreover, we analyze several backbone models, real/synthetic datasets and their proportions when combined. Third, we propose a new curriculum learning strategy to employ several synthetic and real datasets. Our major findings suggest the high performance impact of pace and order of synthetic and real data presentation, achieving state of the art results for well-known models. The results by training with the proposed dataset outperform popular alternatives, thus demonstrating the effectiveness of the proposed protocol. Our code and dataset are available at http://www-vpu.eps.uam.es/publications/WSDA_semantic/
Roberto Alcover-Couso, Juan C. SanMiguel, Marcos Escudero-Viñolo, Álvaro García-Martín
Multim. Tools Appl.3
2023 Attention-Based Knowledge Distillation in Scene Recognition: The Impact of a DCT-Driven Loss
abstract
Knowledge Distillation (KD) is a strategy for the definition of a set of transferability gangways to improve the efficiency of Convolutional Neural Networks. Feature-based Knowledge Distillation is a subfield of KD that relies on intermediate network representations, either unaltered or depth-reduced via maximum activation maps, as the source knowledge. In this paper, we propose and analyze the use of a 2D frequency transform of the activation maps before transferring them. We pose that—by using global image cues rather than pixel estimates, this strategy enhances knowledge transferability in tasks such as scene recognition, defined by strong spatial and contextual relationships between multiple and varied concepts. To validate the proposed method, an extensive evaluation of the state of the art in scene recognition is presented. Experimental results provide strong evidence that the proposed strategy enables the student network to better focus on the relevant image areas learnt by the teacher network, hence leading to better descriptive features and higher transferred performance than every other state-of-the-art alternative. We publicly release the training and evaluation framework used in this paper athttps://www-vpu.eps.uam.es/publications/DCTBasedKDForSceneRecognition.
Alejandro López-Cifuentes, Marcos Escudero-Viñolo, Jesús Bescós, Juan C. SanMiguel
IEEE Trans. Circuits Syst. Video Technol.2
2022 A study on the distribution of social biases in self-supervised learning visual models
abstract
Deep neural networks are efficient at learning the data distribution if it is sufficiently sampled. However, they can be strongly biased by non-relevant factors implicitly incorporated in the training data. These include operational biases, such as ineffective or uneven data sampling, but also ethical concerns, as the social biases are implicitly present—even inadvertently, in the training data or explicitly defined in unfair training schedules. In tasks having impact on human processes, the learning of social biases may produce discriminatory, unethical and untrustworthy consequences. It is often assumed that social biases stem from supervised learning on labelled data, and thus, Self-Supervised Learning (SSL) wrongly appears as an efficient and bias-free solution, as it does not require labelled data. However, it was recently proven that a popular SSL method also incorporates biases. In this paper, we study the biases of a varied set of SSL visual models, trained using ImageNet data, using a method and dataset designed by psychological experts to measure social biases. We show that there is a correlation between the type of the SSL model and the number of biases that it incorporates. Furthermore, the results also suggest that this number does not strictly depend on the model's accuracy and changes throughout the network. Finally, we conclude that a careful SSL model selection process can reduce the number of social biases in the deployed model, whilst keeping high performance. The code is available at https://github.com/vpulab/SB-SSL.
Kirill Sirotkin, Pablo Carballeira, Marcos Escudero-Viñolo
CVPR3
2022 CCL: Class-Wise Curriculum Learning for Class Imbalance Problems
abstract
Computer vision datasets usually present long-tailed training distributions where the classes are not represented with the same number of training samples. This so-called class imbalance problem hinders the proper learning of inference models, biasing them towards over-represented classes and decreasing their generalization. Adopted solutions to tackle the effect of class imbalance are based on weighting the training loss according to the number of class samples, leading to regimes where low-represented classes guide the learning just accounting for their cardinal number. To also incorporate class complexity in the process, we propose a novel training scheme called CCL: Class-wise Curriculum Learning. Classes are first sorted based on a difficulty criterion which not only accounts for the number of training samples but also for their training outcomes. The curriculum is then used to guide the training: easy classes are fed first and—incrementally, the more difficult ones are added. The proposed approach is validated for image classification using long-tailed datasets. Results show that when the proposed Class-wise Curriculum Learning scheme is used, trained models outperform specific state-of-the-art methods devoted to handle the class imbalance problem. The code, data and reported models described along this paper are publicly available at https://github.com/vpulab/CCL
Marcos Escudero-Viñolo, Alejandro López-Cifuentes
ICIP1
2022 Semantic-driven multi-camera pedestrian detection
abstract
Abstract In the current worldwide situation, pedestrian detection has reemerged as a pivotal tool for intelligent video-based systems aiming to solve tasks such as pedestrian tracking, social distancing monitoring or pedestrian mass counting. Pedestrian detection methods, even the top performing ones, are highly sensitive to occlusions among pedestrians, which dramatically degrades their performance in crowded scenarios. The generalization of multi-camera setups permits to better confront occlusions by combining information from different viewpoints. In this paper, we present a multi-camera approach to globally combine pedestrian detections leveraging automatically extracted scene context. Contrarily to the majority of the methods of the state-of-the-art, the proposed approach is scene-agnostic, not requiring a tailored adaptation to the target scenario–e.g., via fine-tuning. This noteworthy attribute does not require ad hoc training with labeled data, expediting the deployment of the proposed method in real-world situations. Context information, obtained via semantic segmentation, is used (1) to automatically generate a common area of interest for the scene and all the cameras, avoiding the usual need of manually defining it, and (2) to obtain detections for each camera by solving a global optimization problem that maximizes coherence of detections both in each 2D image and in the 3D scene. This process yields tightly fitted bounding boxes that circumvent occlusions or miss detections. The experimental results on five publicly available datasets show that the proposed approach outperforms state-of-the-art multi-camera pedestrian detectors, even some specifically trained on the target scenario, signifying the versatility and robustness of the proposed method without requiring ad hoc annotations nor human-guided configuration.
Alejandro López-Cifuentes, Marcos Escudero-Viñolo, Jesús Bescós, Pablo Carballeira
Knowl. Inf. Syst.2
2022 Online clustering-based multi-camera vehicle tracking in scenarios with overlapping FOVs
abstract
Abstract Multi-Target Multi-Camera (MTMC) vehicle tracking is an essential task of visual traffic monitoring, one of the main research fields of Intelligent Transportation Systems. Several offline approaches have been proposed to address this task; however, they are not compatible with real-world applications due to their high latency and post-processing requirements. This lack of suitable approaches motivates our proposal: A new low-latency online approach for MTMC tracking in scenarios with partially overlapping fields of view (FOVs), such as road intersections. Firstly, the proposed approach detects vehicles at each camera. Then, the detections are merged between cameras by applying cross-camera clustering based on appearance and location. Lastly, the clusters containing different detections of the same vehicle are temporally associated to compute the tracks on a frame-by-frame basis. The experiments show promising low-latency results while addressing real-world challenges such as the a priori unknown and time-varying number of targets and the continuous state estimation of them without performing any post-processing of the trajectories. Our code is available at http://www-vpu.eps.uam.es/publications/Online-MTMC-Tracking .
Elena Luna, Juan C. SanMiguel, José María Martínez Sanchez, Marcos Escudero-Viñolo
Multim. Tools Appl.4
2020 Semantic-aware scene recognition
Alejandro López-Cifuentes, Marcos Escudero-Viñolo, Jesús Bescós, Álvaro García-Martín
Pattern Recognit.2
2019 Accurate Segmentation and Registration of Skin Lesion Images to Evaluate Lesion Change
abstract
Skin cancer is a major health problem. There are several techniques to help diagnose skin lesions from a captured image. Computer-aided diagnosis (CAD) systems operate on single images of skin lesions, extracting lesion features to further classify them and help the specialists. Accurate feature extraction, which later on depends on precise lesion segmentation, is key for the performance of these systems. In this paper, we present a skin lesion segmentation algorithm based on a novel adaptation of superpixels techniques and achieve the best reported results for the ISIC 2017 challenge dataset. Additionally, CAD systems have paid little attention to a critical criterion in skin lesion diagnosis: the lesion's evolution. This requires operating on two or more images of the same lesion, captured at different times but with a comparable scale, orientation, and point of view; in other words, an image registration process should first be performed. We also propose in this work, an image registration approach that outperforms top image registration techniques. Combined with the proposed lesion segmentation algorithm, this allows for the accurate extraction of features to assess the evolution of the lesion. We present a case study with the lesion-size feature, paving the way for the development of automatic systems to easily evaluate skin lesion evolution.
Fulgencio Navarro Fajardo, Marcos Escudero-Viñolo, Jesús Bescós
IEEE J. Biomed. Health Informatics2
2018 Automatic Semantic Parsing of the Ground Plane in Scenarios Recorded With Multiple Moving Cameras
abstract
Nowadays, video surveillance scenarios usually rely on manually annotated focus areas to constrain automatic video analysis tasks. Although manual annotation simplifies several stages of the analysis, its use hinders the scalability of the developed solutions and might induce operational problems in scenarios recorded with multiple moving cameras (MMCs). To tackle these problems, an automatic method for the cooperative extraction of areas of interest (AoIs) is proposed. Each captured frame is segmented into regions with semantic roles using a state-of-the-art method. Semantic evidences from different junctures, cameras, and points-of-view are, then, spatio-temporally aligned on a common ground plane. Experimental results on widely used datasets recorded with multiple but static cameras suggest that this process provides broader and more accurate AoIs than those manually defined in the datasets. Moreover, the proposed method naturally determines the projection of obstacles and functional objects in the scene, paving the road towards systems focused on the automatic analysis of human behavior. To our knowledge, this is the first study dealing with this problem, as evidenced by the lack of publicly available MMC benchmarks. To also cope with this issue, we provide a new MMC dataset with associated semantic scene annotations.
Alejandro López-Cifuentes, Marcos Escudero-Viñolo, Jesús Bescós
IEEE Signal Process. Lett.2
2017 Severe-occluded 3D object identification via region-based descriptions
Marcos Escudero-Viñolo, Jesús Bescós
Signal Process. Image Commun.1
2014 A natural and synthetic corpus for benchmarking of hand gesture recognition systems
Javier Molina, José A. Pajuelo, Marcos Escudero-Viñolo, Jesús Bescós, José María Martínez Sanchez
Mach. Vis. Appl.3
2013 Real-time user independent hand gesture recognition from time-of-flight camera video using static and dynamic models
Javier Molina, Marcos Escudero-Viñolo, Alessandro Signoriello, Montse Pardàs, Christian Ferran Bennström, Jesús Bescós, Ferran Marqués, José María Martínez Sanchez
Mach. Vis. Appl.2
2010 A robust framework for region based video object segmentation
abstract
This work presents a novel approach to accurately segment video objects in complex environments, based on region- analysis. A robust-to-illumination region segmentation, a flexible and robust framework for region characterization and matching, and a multi-layer region-based background model, are key aspects of the proposed approach. Presented results show that, not requiring any post-processing process, the technique yields thigh-to-objects segmentation masks.
Marcos Escudero-Viñolo, Jesús Bescós
ICIP1
2008 MPEG video object segmentation under camera motion and multimodal backgrounds
abstract
This paper starts from a state-of-the-art efficient approach to real-time video object segmentation in the MPEG domain. It then describes several techniques to extend the algorithm's compelling behavior to a more generic set of situations. These are focused on the management of intra-coded macroblocks, the exploitation of objects motion coherence, the integrated use of color information and an approach to discriminate objects from multimodal background (e.g., water, flames), always under camera motion conditions. Results evaluation is presented over a ground truth generated with the help of chroma studio in order to reproduce almost real sequences in a controlled way.
Marcos Escudero-Viñolo, Fabricio Tiburzi, Jesús Bescós
ICIP1
2008 A ground truth for motion-based video-object segmentation
abstract
This paper describes the design procedure followed to generate a ground truth for the evaluation of motion-based algorithms for video-object segmentation. A thorough review and classification of the critical factors that affect the behavior of segmentation algorithms results in a set of video scripts which have then been filmed. Foreground objects have been recorded in a chroma studio, in order to automatically obtain pixel-level high quality segmentation masks for each generated sequence. The resulting corpus (segmentation ground-.truth plus filmed sequences mounted over different backgrounds) is available for research purposes under a license agreement.
Fabricio Tiburzi, Marcos Escudero-Viñolo, Jesús Bescós, José María Martínez Sanchez
ICIP2