Julio Silva-Rodríguez

dblp:253/6209 · DBLP profile ↗
← Back
20ranked-venue papers
12as first author
18since 2021 · last 2025
0000-0002-9726-9393ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Conformal Prediction for Zero-Shot Models
abstract
Vision-Language models pre-trained at large scale have shown unprecedented adaptability and generalization to downstream tasks. Although its discriminative potential has been widely explored, its reliability and uncertainty are still overlooked. In this work, we investigate the capabilities of CLIP models under the split conformal prediction paradigm, which provides theoretical guarantees to black-box models based on a small, labeled calibration set. In contrast to the main body of literature on conformal predictors in vision classifiers, foundation models exhibit a particular characteristic: they are pre-trained on a one-time basis on an inaccessible source domain, different from the transferred task. This domain drift negatively affects the efficiency of the conformal sets and poses additional challenges. To alleviate this issue, we propose Conf-Ot, a transfer learning setting that operates transductive over the combined calibration and query sets. Solving an optimal transport problem, the proposed method bridges the domain gap between pre-training and adaptation without requiring additional data splits but still maintaining coverage guarantees. We comprehensively explore this conformal prediction strategy on a broad span of 15 datasets and three nonconformity scores. Conf-Ot provides consistent relative improvements of up to 20% on set efficiency while being ×15 faster than popular transductive approaches. We make the code available1.
Julio Silva-Rodríguez, Ismail Ben Ayed, Jose Dolz
CVPR1
2025 ViLU: Learning Vision-Language Uncertainties for Failure Prediction
abstract
Reliable Uncertainty Quantification (UQ) and failure prediction remain open challenges for Vision-Language Models (VLMs). We introduce ViLU, a new Vision-Language Uncertainty quantification framework that contextualizes uncertainty estimates by leveraging all task-relevant textual representations. ViLU constructs an uncertainty-aware multi-modal representation by integrating the visual embedding, the predicted textual embedding, and an image-conditioned textual representation via cross-attention. Unlike traditional UQ methods based on loss prediction, ViLU trains an uncertainty predictor as a binary classifier to distinguish correct from incorrect predictions using a weighted binary cross-entropy loss, making it loss-agnostic. In particular, our proposed approach is well-suited for post-hoc settings, where only vision and text embeddings are available without direct access to the model itself. Extensive experiments on diverse datasets show the significant gains of our method compared to state-of-the-art failure prediction methods. We apply our method to standard classification datasets, such as ImageNet-1k, as well as large-scale image-caption datasets like CC12M and LAION-400M. Ablation studies highlight the critical role of our architecture and training in achieving effective uncertainty quantification. Our code is publicly available and can be found here: https://github.com/ykrmm/ViLU.
Marc Lafon, Yannis Karmim, Julio Silva-Rodríguez, Paul Couairon, Clément Rambour, Raphaël Fournier-S'niehotta, Ismail Ben Ayed, Jose Dolz, Nicolas Thome
ICCV3
2025 Regularized Low-Rank Adaptation for Few-Shot Organ Segmentation
Ghassen Baklouti, Julio Silva-Rodríguez, Jose Dolz, Houda Bahig, Ismail Ben Ayed
MICCAI (7)2
2025 Trustworthy Few-Shot Transfer of Medical VLMs Through Split Conformal Prediction
Julio Silva-Rodríguez, Ismail Ben Ayed, Jose Dolz
MICCAI (7)1
2025 Few-Shot, Now for Real: Medical VLMs Adaptation Without Balanced Sets or Validation
Julio Silva-Rodríguez, Fereshteh Shakeri, Houda Bahig, Jose Dolz, Ismail Ben Ayed
MICCAI (7)1
2025 A Foundation Language-Image Model of the Retina (FLAIR): encoding expert knowledge in text supervision
Julio Silva-Rodríguez, Hadi Chakor, Riadh Kobbi, Jose Dolz, Ismail Ben Ayed
Medical Image Anal.1
2025 Towards Foundation Models and Few-Shot Parameter-Efficient Fine-Tuning for Volumetric Organ Segmentation
Julio Silva-Rodríguez, Jose Dolz, Ismail Ben Ayed
Medical Image Anal.1
2025 HRDC challenge: a public benchmark for hypertension and hypertensive retinopathy classification from fundus images
Xiangning Wang, Zhouyu Guan, An-ran Ran, Tingyao Li, Zheyuan Wang, Xinming Shu, Jinyang Xie, Shichang Liu, Guanyu Xing, Julio Silva-Rodríguez, Riadh Kobbi, Ping Li 0016, Tingli Chen, Lei Bi 0001, Jinman Kim, Weiping Jia, Huating Li, Harry Qin, Ping Zhang 0016, Ching Yu Cheng, Pheng-Ann Heng, Tien Yin Wong, Carol Y. Cheung, Nadia Magnenat-Thalmann, Bin Sheng 0001
Vis. Comput.13
2024 A Closer Look at the Few-Shot Adaptation of Large Vision-Language Models
abstract
Efficient transfer learning (ETL) is receiving increasing attention to adapt large pre-trained language-vision models on downstream tasks with a few labeled samples. While significant progress has been made, we reveal that state-of-the-art ETL approaches exhibit strong performance only in narrowly-defined experimental setups, and with a careful adjustment of hyperparameters based on a large corpus of labeled samples. In particular, we make two interesting, and surprising empirical observations. First, to out-perform a simple Linear Probing baseline, these methods require to optimize their hyper-parameters on each target task. And second, they typically underperform -sometimes dramatically- standard zero-shot predictions in the presence of distributional drifts. Motivated by the unrealistic assumptions made in the existing literature, i.e., access to a large validation set and case-specific grid-search for optimal hyperparameters, we propose a novel approach that meets the requirements of real-world scenarios. More concretely, we introduce a CLass-Adaptive linear Probe (CLAP) objective, whose balancing term is optimized via an adaptation of the general Augmented Lagrangian method tailored to this context. We comprehensively evaluate CLAP on a broad span of datasets and scenarios, demonstrating that it consistently outperforms SoTA approaches, while yet being a much more efficient alternative. Code available at https://github.com/jusiro/CLAP.
Julio Silva-Rodríguez, Sina Hajimiri, Ismail Ben Ayed, Jose Dolz
CVPR1
2024 Robust Calibration of Large Vision-Language Adapters
Balamurali Murugesan, Julio Silva-Rodríguez, Ismail Ben Ayed, Jose Dolz
ECCV (24)2
2024 Class and Region-Adaptive Constraints for Network Calibration
Balamurali Murugesan, Julio Silva-Rodríguez, Ismail Ben Ayed, Jose Dolz
MICCAI (8)2
2024 Few-Shot Adaptation of Medical Vision-Language Models
Fereshteh Shakeri, Yunshi Huang, Julio Silva-Rodríguez, Houda Bahig, An Tang, Jose Dolz, Ismail Ben Ayed
MICCAI (12)3
2023 Exploring the Transferability of a Foundation Model for Fundus Images: Application to Hypertensive Retinopathy
Julio Silva-Rodríguez, Jihed Chelbi, Waziha Kabir, Hadi Chakor, Jose Dolz, Ismail Ben Ayed, Riadh Kobbi
CGI (3)1
2022 Challenging Mitosis Detection Algorithms: Global Labels Allow Centroid Localization
Claudio Fernandez-Martín, Umay Kiraz, Julio Silva-Rodríguez, Sandra Morales, Emiel A. M. Janssen, Valery Naranjo
IDEAL3
2022 Supervised contrastive learning-guided prototypes on axle-box accelerations for railway crossing inspections
abstract
Increasing demands on railway structures have led to a need for new cost-effective maintenance strategies in recent years. Current dynamic railway track monitoring systems are usually based on the analysis of axle-box accelerations to automatically detect track singularities and defects. These methods rely on hand-crafted feature extraction and classifiers for different tasks. However, the low performance shown in previous literature makes it necessary to complement these analyses with in-situ inspections. Very recent works have proposed the use of deep learning systems that allow extracting more generalizable features from time–frequency spectrograms. However, the lack of specific public domain datasets and the finite number of track singularities in a railway structure have limited the development of deep learning based systems. In this paper, we propose a method capable of outstanding in low-data scenarios. In particular, we explore the use of supervised contrastive learning to cluster class embeddings nearly in the encoder latent space, which is used during inference for prototypical distance-based class assignment. We provide comprehensive experiments demonstrating the performance of our method in comparison to previous literature for detecting worn-out crossings.
Julio Silva-Rodríguez, Pablo Salvador, Valery Naranjo, Ricardo Insa
Expert Syst. Appl.1
2022 Constrained unsupervised anomaly segmentation
abstract
Current unsupervised anomaly localization approaches rely on generative models to learn the distribution of normal images, which is later used to identify potential anomalous regions derived from errors on the reconstructed images. However, a main limitation of nearly all prior literature is the need of employing anomalous images to set a class-specific threshold to locate the anomalies. This limits their usability in realistic scenarios, where only normal data is typically accessible. Despite this major drawback, only a handful of works have addressed this limitation, by integrating supervision on attention maps during training. In this work, we propose a novel formulation that does not require accessing images with abnormalities to define the threshold. Furthermore, and in contrast to very recent work, the proposed constraint is formulated in a more principled manner, leveraging well-known knowledge in constrained optimization. In particular, the equality constraint on the attention maps in prior work is replaced by an inequality constraint, which allows more flexibility. In addition, to address the limitations of penalty-based functions we employ an extension of the popular log-barrier methods to handle the constraint. Last, we propose an alternative regularization term that maximizes the Shannon entropy of the attention maps, reducing the amount of hyperparameters of the proposed model. Comprehensive experiments on two publicly available datasets on brain lesion segmentation demonstrate that the proposed approach substantially outperforms relevant literature, establishing new state-of-the-art results for unsupervised lesion segmentation, and without the need to access anomalous images.
Julio Silva-Rodríguez, Valery Naranjo, Jose Dolz
Medical Image Anal.1
2021 Looking at the whole picture: constrained unsupervised anomaly segmentation
Julio Silva-Rodríguez, Valery Naranjo, Jose Dolz
BMVC1
2021 Self-Learning for Weakly Supervised Gleason Grading of Local Patterns
abstract
Prostate cancer is one of the main diseases affecting men worldwide. The gold standard for diagnosis and prognosis is the Gleason grading system. In this process, pathologists manually analyze prostate histology slides under microscope, in a high time-consuming and subjective task. In the last years, computer-aided-diagnosis (CAD) systems have emerged as a promising tool that could support pathologists in the daily clinical practice. Nevertheless, these systems are usually trained using tedious and prone-to-error pixel-level annotations of Gleason grades in the tissue. To alleviate the need of manual pixel-wise labeling, just a handful of works have been presented in the literature. Furthermore, despite the promising results achieved on global scoring the location of cancerous patterns in the tissue is only qualitatively addressed. These heatmaps of tumor regions, however, are crucial to the reliability of CAD systems as they provide explainability to the system's output and give confidence to pathologists that the model is focusing on medical relevant features. Motivated by this, we propose a novel weakly-supervised deep-learning model, based on self-learning CNNs, that leverages only the global Gleason score of gigapixel whole slide images during training to accurately perform both, grading of patch-level patterns and biopsy-level scoring. To evaluate the performance of the proposed method, we perform extensive experiments on three different external datasets for the patch-level Gleason grading, and on two different test sets for global Grade Group prediction. We empirically demonstrate that our approach outperforms its supervised counterpart on patch-level Gleason grading by a large margin, as well as state-of-the-art methods on global biopsy-level scoring. Particularly, the proposed model brings an average improvement on the Cohen's quadratic kappa ( κ) score of nearly 18% compared to full-supervision for the patch-level Gleason grading task. This suggests that the absence of the annotator's bias in our approach and the capability of using large weakly labeled datasets during training leads to higher performing and more robust models. Furthermore, raw features obtained from the patch-level classifier showed to generalize better than previous approaches in the literature to the subjective global biopsy-level scoring.
Julio Silva-Rodríguez, Adrián Colomer, Jose Dolz, Valery Naranjo
IEEE J. Biomed. Health Informatics1
2020 Gleason Grading of Histology Prostate Images Through Semantic Segmentation via Residual U-Net
abstract
Worldwide, prostate cancer is one of the main cancers affecting men. The final diagnosis of prostate cancer is based on the visual detection of Gleason patterns in prostate biopsy by pathologists. Computer-aided-diagnosis systems allow to delineate and classify the cancerous patterns in the tissue via computer-vision algorithms in order to support the physicians' task. The methodological core of this work is a U-Net convolutional neural network for image segmentation modified with residual blocks able to segment cancerous tissue according to the full Gleason system. This model outperforms other well-known architectures, and reaches a pixel-level Cohen's quadratic Kappa of 0.52, at the level of previous image-level works in the literature, but providing also a detailed localisation of the patterns.
Amartya Kalapahar, Julio Silva-Rodríguez, Adrián Colomer, Fernando López-Mir, Valery Naranjo
ICIP2
2020 Prostate Gland Segmentation in Histology Images via Residual and Multi-resolution U-NET
Julio Silva-Rodríguez, Elena Payá-Bosch, Gabriel García 0001, Adrián Colomer, Valery Naranjo
IDEAL (1)1