EDBT 2026 Demo / reviewers in the wild / expert
Rodrigo Santa Cruz
dblp:183/6373
· DBLP profile ↗
14ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-5273-7296ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving the Generation of VAEs with High Dimensional Latent Spaces by the use of Hyperspherical CoordinatesabstractVariational autoencoders (VAE) encode data into lower-dimensional latent vectors before decoding those vectors back to data. Once trained, decoding a random latent vector from the prior usually does not produce meaningful data, at least when the latent space has more than a dozen dimensions. In this paper, we investigate this issue by drawing insight from high dimensional statistics: in these regimes, the latent vectors of a standard VAE are by construction distributed uniformly on a hypersphere. We propose to formulate the latent variables of a VAE using hyperspherical coordinates, which allows compressing the latent vectors towards an island on the hypersphere, thereby reducing the latent sparsity and we show that this improves the generation ability of the VAE. We propose a new parameterization of the latent space with limited computational overhead. Alejandro Ascarate, Léo Lebrat, Rodrigo Santa Cruz, Clinton Fookes, Olivier Salvado |
IJCNN | 3 |
| 2025 | Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task PlanningabstractTo enable robots to comprehend high-level human instructions and perform complex tasks, a key challenge lies in achieving comprehensive scene understanding: interpreting and interacting with the 3D environment in a meaningful way. This requires a smart map that fuses accurate geometric structure with rich, human-understandable semantics. To address this, we introduce the 3D Queryable Scene Representation (3D QSR), a novel framework built on multimedia data that unifies three complementary 3D representations: (1) 3D-consistent novel view rendering and segmentation from panoptic reconstruction, (2) precise geometry from 3D point clouds, and (3) structured, scalable organization via 3D scene graphs. Built on an object-centric design, the framework integrates with large vision-language models to enable semantic queryability by linking multimodal object embeddings, and supporting object-level retrieval of geometric, visual, and semantic information. The retrieved data are then loaded into a robotic task planner for downstream execution. Xun Li 0004, Rodrigo Santa Cruz, Mingze Xi, Hu Zhang 0005, Madhawa Perera, Ziwei Wang 0003, Ahalya Ravendran, Brandon J. Matthews, Matt Adcock, Dadong Wang, Jiajun Liu 0004 |
ACM Multimedia | 2 |
| 2025 | SALVE: A 3D Reconstruction Benchmark of Wounds from Consumer-Grade VideosabstractManaging chronic wounds is a global challenge that can be alleviated by the adoption of automatic systems for clinical wound assessment from consumer-grade videos. While 2D image analysis approaches are insufficient for handling the 3D features of wounds, existing approaches utilizing 3D reconstruction methods have not been thor-oughly evaluated. To address this gap, this paper presents a comprehensive study on 3D wound reconstruction from consumer-grade videos. Specifically, we introduce the SALVE dataset, comprising video recordings of realistic wound phantoms captured with different cameras. Using this dataset, we assess the accuracy and precision of state-of-the-art methods for 3D reconstruction, ranging from traditional photogrammetry pipelines to advanced neural rendering approaches. In our experiments, we observe that photogrammetry approaches do not provide smooth surfaces suitable for precise clinical measurements of wounds. Neural rendering approaches show promise in addressing this issue, advancing the use of this technology in wound care practices. We encourage the readers to visit the project page: https://rcmichierchia.github.io/SALVE/. Remi Chierchia, Léo Lebrat, David Ahmedt-Aristizabal, Olivier Salvado, Clinton Fookes, Rodrigo Santa Cruz |
WACV | 6 |
| 2024 | NeRF Director: Revisiting View Selection in Neural Volume RenderingabstractNeural Rendering representations have significantly contributed to the field of 3D computer vision. Given their potential, considerable efforts have been invested to improve their performance. Nonetheless, the essential question of selecting training views is yet to be thoroughly investigated. This key aspect plays a vital role in achieving high-quality results and aligns with the well-known tenet of deep learning: “garbage in, garbage out”. In this paper, we first illustrate the importance of view selection by demonstrating how a simple rotation of the test views within the most pervasive NeRF dataset can lead to consequential shifts in the performance rankings of state-of-the-art techniques. To address this challenge, we introduce a unified framework for view selection methods and devise a thorough benchmark to assess its impact. Significant improvements can be achieved without leveraging error or uncertainty estimation but focusing on uniform view coverage of the reconstructed object, resulting in a training-free approach. Using this technique, we show that high-quality renderings can be achieved faster by using fewer views. We conduct extensive experiments on both synthetic datasets and realistic data to demonstrate the effectiveness of our proposed method compared with random, conventional error-based, and uncertainty-guided view selection. Wenhui Xiao, Rodrigo Santa Cruz, David Ahmedt-Aristizabal, Olivier Salvado, Clinton Fookes, Léo Lebrat |
CVPR | 2 |
| 2023 | Bias Identification with RankPix SaliencyabstractSaliency methods are critical tools that allow the estimation of the most important features of an input image that contribute to the network’s prediction. These tools are pivotal in high-stakes applications such as medical diagnosis or autonomous driving. Additionally, these tools can help identify models’ biasedness, such as a strong prior on object placement, easily distinguishable background features, or frequent object co-occurrence. We introduce RankPix, a novel saliency method for visual bias identification in image classification tasks. RankPix is a derivative-free approach that allows the identification of a minimum subset of pixels/features at a given network layer that changes the output of a classifier. Surprisingly, this approaches provides equivalent performance to gradient-based approaches on the standard pointing game benchmark. More interestingly, RankPix outperforms traditional approaches for systematic bias identification. Salamata Konate, Léo Lebrat, Rodrigo Santa Cruz, Clinton Fookes, Andrew P. Bradley, Olivier Salvado |
ICASSP | 3 |
| 2023 | DBCE : A Saliency Method for Medical Deep Learning Through Anatomically-Consistent Free-Form DeformationsabstractDeep learning models are powerful tools for addressing challenging medical imaging problems. However, for an ever-growing range of applications, interpreting a model’s prediction remains non-trivial. Understanding decisions made by black-box algorithms is critical, and assessing their fairness and susceptibility to bias is a key step towards healthcare deployment. In this paper, we propose DBCE (Deformation Based Counterfactual Explainability). We optimise a diffeomorphic transformation that deforms a given input image to change the prediction of the model. This provides anatomically meaningful saliency maps indicating tissue atrophy and expansion, which can be easily interpreted by clinicians. In our test case, DBCE replicates the transition of a patient from healthy control (HC) to Alzheimer’s disease (AD). We benchmark DBCE against three commonly used saliency methods. We show that it provides more meaningful saliency maps when applied to one subject and disease-consistent atrophy patterns when used over a larger cohort. In addition, our method fulfils a recent sanity check and is repeatable for different model initialisations in contrast to classical sensitivity-based methods. Joshua Peters, Léo Lebrat, Rodrigo Santa Cruz, Aaron Nicolson, Gregg Belous, Salamata Konate, Parnesh Raniga, Vincent Doré, Pierrick Bourgeat, Jurgen Fripp, Clinton Fookes, Olivier Salvado |
WACV | 3 |
| 2022 | CorticalFlow++: Boosting Cortical Surface Reconstruction Accuracy, Regularity, and Interoperability
Rodrigo Santa Cruz, Léo Lebrat, Darren Fu, Pierrick Bourgeat, Jurgen Fripp, Clinton Fookes, Olivier Salvado |
MICCAI (5) | 1 |
| 2022 | Quantifiable brain atrophy synthesis for benchmarking of cortical thickness estimation methodsabstractCortical thickness (CTh) is routinely used to quantify grey matter atrophy as it is a significant biomarker in studying neurodegenerative and neurological conditions. Clinical studies commonly employ one of several available CTh estimation software tools to estimate CTh from brain MRI scans. In recent years, machine learning-based methods emerged as a faster alternative to the main-stream CTh estimation methods (e.g. FreeSurfer). Evaluation and comparison of CTh estimation methods often include various metrics and downstream tasks, but none fully covers the sensitivity to sub-voxel atrophy characteristic of neurodegeneration. In addition, current evaluation methods do not provide a framework for the intra-method region-wise evaluation of CTh estimation methods. Therefore, we propose a method for brain MRI synthesis capable of generating a range of sub-voxel atrophy levels (global and local) with quantifiable changes from the baseline scan. We further create a synthetic test set and evaluate four different CTh estimation methods: FreeSurfer (cross-sectional), FreeSurfer (longitudinal), DL+DiReCT and HerstonNet. DL+DiReCT showed superior sensitivity to sub-voxel atrophy over other methods in our testing framework. The obtained results indicate that our synthetic test set is suitable for benchmarking CTh estimation methods on both global and local scales as well as regional inter-and intra-method performance comparison. Filip Rusak, Rodrigo Santa Cruz, Léo Lebrat, Ondrej Hlinka, Jurgen Fripp, Elliot Smith 0003, Clinton Fookes, Andrew P. Bradley, Pierrick Bourgeat |
Medical Image Anal. | 2 |
| 2021 | MongeNet: Efficient Sampler for Geometric Deep LearningabstractRecent advances in geometric deep-learning introduce complex computational challenges for evaluating the distance between meshes. From a mesh model, point clouds are necessary along with a robust distance metric to assess surface quality or as part of the loss function for training models. Current methods often rely on a uniform random mesh discretization, which yields irregular sampling and noisy distance estimation. In this paper we introduce MongeNet, a fast and optimal transport based sampler that allows for an accurate discretization of a mesh with better approximation properties. We compare our method to the ubiquitous random uniform sampling and show that the approximation error is almost half with a very small computational overhead. Léo Lebrat, Rodrigo Santa Cruz, Clinton Fookes, Olivier Salvado |
CVPR | 2 |
| 2021 | CorticalFlow: A Diffeomorphic Mesh Transformer Network for Cortical Surface ReconstructionabstractIn this paper, we introduce CorticalFlow, a new geometric deep-learning model that, given a 3-dimensional image, learns to deform a reference template towards a targeted object. To conserve the template mesh’s topological properties, we train our model over a set of diffeomorphic transformations. This new implementation of a flow Ordinary Differential Equation (ODE) framework benefits from a small GPU memory footprint, allowing the generation of surfaces with several hundred thousand vertices. To reduce topological errors introduced by its discrete resolution, we derive numeric conditions which improve the manifoldness of the predicted triangle mesh. To exhibit the utility of CorticalFlow, we demonstrate its performance for the challenging task of brain cortical surface reconstruction. In contrast to the current state-of-the-art, CorticalFlow produces superior surfaces while reducing the computation time from nine and a half minutes to one second. More significantly, CorticalFlow enforces the generation of anatomically plausible surfaces; the absence of which has been a major impediment restricting the clinical relevance of such surface reconstruction methods. Léo Lebrat, Rodrigo Santa Cruz, Frédéric de Gournay, Darren Fu, Pierrick Bourgeat, Jurgen Fripp, Clinton Fookes, Olivier Salvado |
NeurIPS | 2 |
| 2021 | DeepCSR: A 3D Deep Learning Approach for Cortical Surface ReconstructionabstractThe study of neurodegenerative diseases relies on the reconstruction and analysis of the brain cortex from magnetic resonance imaging (MRI). Traditional frameworks for this task like FreeSurfer demand lengthy runtimes, while its accelerated variant FastSurfer still relies on a voxel-wise segmentation which is limited by its resolution to capture narrow continuous objects as cortical surfaces. Having these limitations in mind, we propose DeepCSR, a 3D deep learning framework for cortical surface reconstruction from MRI. Towards this end, we train a neural network model with hypercolumn features to predict implicit surface representations for points in a brain template space. After training, the cortical surface at a desired level of detail is obtained by evaluating surface representations at specific coordinates, and subsequently applying a topology correction algorithm and an isosurface extraction method. Thanks to the continuous nature of this approach and the efficacy of its hypercolumn features scheme, DeepCSR efficiently reconstructs cortical surfaces at high resolution capturing fine details in the cortical folding. Moreover, DeepCSR is as accurate, more precise, and faster than the widely used FreeSurfer toolbox and its deep learning powered variant FastSurfer on reconstructing cortical surfaces from MRI which should facilitate large-scale medical studies and new healthcare applications. Rodrigo Santa Cruz, Léo Lebrat, Pierrick Bourgeat, Clinton Fookes, Jurgen Fripp, Olivier Salvado |
WACV | 1 |
| 2019 | Visual Permutation LearningabstractWe present a principled approach to uncover the structure of visual data by solving a deep learning task coined visual permutation learning. The goal of this task is to find the permutation that recovers the structure of data from shuffled versions of it. In the case of natural images, this task boils down to recovering the original image from patches shuffled by an unknown permutation matrix. Permutation matrices are discrete, thereby posing difficulties for gradient-based optimization methods. To this end, we resort to a continuous approximation using doubly-stochastic matrices and formulate a novel bi-level optimization problem on such matrices that learns to recover the permutation. Unfortunately, such a scheme leads to expensive gradient computations. We circumvent this issue by further proposing a computationally cheap scheme for generating doubly stochastic matrices based on Sinkhorn iterations. To implement our approach we propose DeepPermNet, an end-to-end CNN model for this task. The utility of DeepPermNet is demonstrated on three challenging computer vision problems, namely, relative attributes learning, supervised learning-to-rank, and self-supervised representation learning. Our results show state-of-the-art performance on the Public Figures and OSR benchmarks for relative attributes learning, chronological and interestingness image ranking for supervised learning-to-rank, and competitive results in the classification and segmentation tasks of the PASCAL VOC dataset for self-supervised representation learning. Rodrigo Santa Cruz, Basura Fernando, Anoop Cherian, Stephen Gould |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Neural Algebra of ClassifiersabstractThe world is fundamentally compositional, so it is natural to think of visual recognition as the recognition of basic visually primitives that are composed according to well-defined rules. This strategy allows us to recognize unseen complex concepts from simple visual primitives. However, the current trend in visual recognition follows a data greedy approach where huge amounts of data are required to learn models for any desired visual concept. In this paper, we build on the compositionality principle and develop an "algebra" to compose classifiers for complex visual concepts. To this end, we learn neural network modules to perform boolean algebra operations on simple visual classifiers. Since these modules form a complete functional set, a classifier for any complex visual concept defined as a boolean expression of primitives can be obtained by recursively applying the learned modules, even if we do not have a single training sample. As our experiments show, using such a framework, we can compose classifiers for complex visual concepts outperforming standard baselines on two well-known visual recognition benchmarks. Finally, we present a qualitative analysis of our method and its properties. Rodrigo Santa Cruz, Basura Fernando, Anoop Cherian, Stephen Gould |
WACV | 1 |
| 2017 | DeepPermNet: Visual Permutation LearningabstractWe present a principled approach to uncover the structure of visual data by solving a novel deep learning task coined visual permutation learning. The goal of this task is to find the permutation that recovers the structure of data from shuffled versions of it. In the case of natural images, this task boils down to recovering the original image from patches shuffled by an unknown permutation matrix. Unfortunately, permutation matrices are discrete, thereby posing difficulties for gradient-based methods. To this end, we resort to a continuous approximation of these matrices using doubly-stochastic matrices which we generate from standard CNN predictions using Sinkhorn iterations. Unrolling these iterations in a Sinkhorn network layer, we propose DeepPermNet, an end-to-end CNN model for this task. The utility of DeepPermNet is demonstrated on two challenging computer vision problems, namely, (i) relative attributes learning and (ii) self-supervised representation learning. Our results show state-of-the-art performance on the Public Figures and OSR benchmarks for (i) and on the classification and segmentation tasks on the PASCAL VOC dataset for (ii). Rodrigo Santa Cruz, Basura Fernando, Anoop Cherian, Stephen Gould |
CVPR | 1 |