VLDB 2026 Research / reviewers in the wild / expert
Roberto Alcover-Couso
dblp:341/1295
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0001-9609-4416ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Soft-Labelling for Budget-Constrained Semantic Segmentation: Bringing Coherence to label Down-SamplingabstractIn semantic segmentation, training data down-sampling is commonly performed due to resource limitations, the need to adapt image size to the model input, or to improve data augmentation. This down-sampling typically employs different strategies for the image data and the annotated labels. Such discrepancy leads to mismatches between the down-sampled colour and ground-truth label images. Hence, the training performance significantly decreases as the down-sampling factor increases. In this paper, we bring together the down-sampling strategies for the image data and the training labels. To that aim, we propose a novel framework for label down-sampling via soft-labelling that better conserves label information after down-sampling, thereby, fully aligning soft-labels with image data to keep the distribution of the sampled pixels for down-sampling. This proposal also produces reliable annotations for under-represented semantic classes. Altogether, it allows training competitive models at lower resolutions. Experiments show that our proposal outperforms other down-sampling strategies. Moreover, state-of-the-art performance is achieved for reference benchmarks, but employing significantly fewer computational resources than foremost methods. This proposal enables competitive research for semantic segmentation under resource constraints. Roberto Alcover-Couso, Marcos Escudero-Viñolo, Juan C. SanMiguel, José María Martínez Sanchez |
Comput. Vis. Media | 1 |
| 2025 | Pathology-Aware Adaptive Watermarking for Text-Driven Medical Image Synthesis
Chanyoung Kim 0001, Dayun Ju, Jinyeong Kim, Woojung Han, Roberto Alcover-Couso, Seong Jae Hwang |
MICCAI (16) | 5 |
| 2025 | Controlling semantics of diffusion-augmented data for unsupervised domain adaptationabstractAbstract Unsupervised domain adaptation (UDA) offers a compelling solution to bridge the gap between labelled synthetic data and unlabelled real‐world data for training semantic segmentation models, given the high costs associated with manual annotation. However, the visual differences between the synthetic and real images pose significant challenges to their practical applications. This work addresses these challenges through synthetic‐to‐real style transfer leveraging diffusion models. The authors’ proposal incorporates semantic controllers to guide the diffusion process and low‐rank adaptations (LoRAs) to ensure that style‐transferred images align with real‐world aesthetics while preserving semantic layout. Moreover, the authors introduce quality metrics to rank the utility of generated images, enabling the selective use of high‐quality images for training. To further enhance reliability, the authors propose a novel loss function that mitigates artefacts from the style transfer process by incorporating only pixels aligned with the original semantic labels. Experimental results demonstrate that the authors’ proposal outperforms selected state‐of‐the‐art methods for image generation and UDA training, achieving optimal performance even with a smaller set of high‐quality generated images. The authors’ code and models are available at http://www‐vpu.eps.uam.es/ControllingSem4UDA/ . Henrietta Ridley, Roberto Alcover-Couso, Juan C. SanMiguel |
IET Comput. Vis. | 2 |
| 2025 | Gradient-based class weighting for unsupervised domain adaptation in dense prediction visual tasksabstractIn unsupervised domain adaptation (UDA), where models are trained on source data (e.g., synthetic) and adapted to target data (e.g., real-world) without target annotations, addressing the challenge of significant class imbalance remains an open issue. Despite progress in bridging the domain gap, existing methods often experience performance degradation when confronted with highly imbalanced dense prediction visual tasks like semantic segmentation. This discrepancy becomes especially pronounced due to the lack of equivalent priors between the source and target domains, turning class imbalanced techniques used for other areas (e.g., image classification) ineffective in UDA scenarios. This paper proposes a class-imbalance mitigation strategy that incorporates class-weights into the UDA learning losses, with the novelty of estimating these weights dynamically through the gradients of the per-class losses, defining a Gradient-based class weighting (GBW) approach. The proposed GBW naturally increases the contribution of classes whose learning is hindered by highly-represented classes, and has the advantage of automatically adapting to training outcomes, avoiding explicit curricular learning patterns common in loss-weighing strategies. Extensive experimentation validates the effectiveness of GBW across architectures (Convolutional and Transformer), UDA strategies (adversarial, self-training and entropy minimization), tasks (semantic and panoptic segmentation), and datasets. Analysis shows that GBW consistently increases the recall of under-represented classes. • A novel class-imbalance algorithm that uses per-class gradients to assign effective class weights (GBW). • Class weights are computed through an optimization process that maximizes the decrease in the loss function. • The proposed method demonstrates significant and consistent performance improvements across various UDA methods. • The class weights provide a per-class complexity measure throughout training, establishing an automatic and adaptable curriculum. Roberto Alcover-Couso, Marcos Escudero-Viñolo, Juan C. SanMiguel, Jesús Bescós |
Pattern Recognit. | 1 |
| 2025 | Per-class curriculum for Unsupervised Domain Adaptation in semantic segmentationabstractAbstract Accurate training of deep neural networks for semantic segmentation requires a large number of pixel-level annotations of real images, which are expensive to generate or not even available. In this context, Unsupervised Domain Adaptation (UDA) can transfer knowledge from unlimited synthetic annotations to unlabeled real images of a given domain. UDA methods are composed of an initial training stage with labeled synthetic data followed by a second stage for feature alignment between labeled synthetic and unlabeled real data. In this paper, we propose a novel approach for UDA focusing the initial training stage, which leads to increased performance after adaptation. We introduce a curriculum strategy where each semantic class is learned progressively. Thereby, better features are obtained for the second stage. This curriculum is based on: (1) a class-scoring function to determine the difficulty of each semantic class, (2) a strategy for incremental learning based on scoring and pacing functions that limits the required training time unlike standard curriculum-based training and (3) a training loss to operate at class level. We extensively evaluate our approach as the first stage of several state-of-the-art UDA methods for semantic segmentation. Our results demonstrate significant performance enhancements across all methods: improvements of up to 10% for entropy-based techniques and 8% for adversarial methods. These findings underscore the dependency of UDA on the accuracy of the initial training. The implementation is available at https://github.com/vpulab/PCCL . Roberto Alcover-Couso, Juan C. SanMiguel, Marcos Escudero-Viñolo, Pablo Carballeira |
Vis. Comput. | 1 |
| 2025 | Layer-wise model merging for unsupervised domain adaptation in segmentation tasksabstractAbstract Merging parameters of multiple models has resurfaced as an effective strategy to enhance task performance and robustness, but prior work is limited by the high costs of ensemble creation and inference. In this paper, we leverage the abundance of freely accessible trained models to introduce a cost-free approach to model merging. It focuses on a layer-wise integration of merged models, aiming to maintain the distinctiveness of the task-specific final layers while unifying the initial layers, which are primarily associated with feature extraction. This approach ensures parameter consistency across all layers, essential for boosting performance. Moreover, it facilitates seamless integration of knowledge, enabling effective merging of models from different datasets and tasks. Specifically, we investigate its applicability in unsupervised domain adaptation (UDA), an unexplored area for model merging, for semantic and panoptic segmentation. Experimental results demonstrate substantial UDA improvements without additional costs for merging same-architecture models from distinct datasets ( $$\uparrow 2.6\%$$ ↑ 2.6 % mIoU) and different-architecture models with a shared backbone ( $$\uparrow 6.8\%$$ ↑ 6.8 % mIoU). Furthermore, merging semantic and panoptic segmentation models increases mPQ by 7%. These findings are validated across a wide variety of UDA strategies, architectures and datasets. The code will be publicly available upon acceptance in the LWMM repository: http://www-vpu.eps.uam.es/LWMM/ . Roberto Alcover-Couso, Juan C. SanMiguel, Marcos Escudero-Viñolo, Jose M. Martínez |
Vis. Comput. | 1 |
| 2024 | Open-Vocabulary Attention Maps with Token Optimization for Semantic Segmentation in Diffusion ModelsabstractDiffusion models represent a new paradigm in text-to-image generation. Beyond generating high-quality images from text prompts, models such as Stable Diffusion have been successfully extended to the joint generation of se-mantic segmentation pseudo-masks. However, current ex-tensions primarily rely on extracting attentions linked to prompt words used for image synthesis. This approach limits the generation of segmentation masks derived from word tokens not contained in the text prompt. In this work, we introduce Open- Vocabulary Attention Maps (OVAM)-a training-free method for text-to-image diffusion models that enables the generation of attention maps for any word. In addition, we propose a lightweight optimization process based on OVAM for finding tokens that generate accurate attention maps for an object class with a single annotation. We evaluate these tokens within existing state-of-the-art Stable Diffusion extensions. The best-performing model im-proves its mIoU from 52.1 to 86.6 for the synthetic images' pseudo-masks, demonstrating that our optimized tokens are an efficient way to improve the performance of existing methods without architectural changes or retraining. The implementation is available at github.com/vpulablovam. Pablo Marcos-Manchón, Roberto Alcover-Couso, Juan C. SanMiguel, Jose M. Martínez |
CVPR | 2 |
| 2023 | On exploring weakly supervised domain adaptation strategies for semantic segmentation using synthetic dataabstractAbstract Pixel-wise image segmentation is key for many Computer Vision applications. The training of deep neural networks for this task has expensive pixel-level annotation requirements, thus, motivating a growing interest on synthetic data to provide unlimited data and its annotations. In this paper, we focus on the generation and application of synthetic data as representative training corpuses for semantic segmentation of urban scenes. First, we propose a synthetic data generation protocol, which identifies key features affecting performance and provides datasets with variable complexity. Second, we adapt two popular weakly supervised domain adaptation approaches (combined training, fine-tuning) to employ synthetic and real data. Moreover, we analyze several backbone models, real/synthetic datasets and their proportions when combined. Third, we propose a new curriculum learning strategy to employ several synthetic and real datasets. Our major findings suggest the high performance impact of pace and order of synthetic and real data presentation, achieving state of the art results for well-known models. The results by training with the proposed dataset outperform popular alternatives, thus demonstrating the effectiveness of the proposed protocol. Our code and dataset are available at http://www-vpu.eps.uam.es/publications/WSDA_semantic/ Roberto Alcover-Couso, Juan C. SanMiguel, Marcos Escudero-Viñolo, Álvaro García-Martín |
Multim. Tools Appl. | 1 |