EDBT 2026 Demo / reviewers in the wild / expert
Ricardo Guerrero
dblp:76/10169 · also Ricardo Guerrero Moreno
· DBLP profile ↗
21ranked-venue papers
4as first author
10since 2021 · last 2024
0000-0001-5019-9933ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Efficient Vision-Language pre-training via domain-specific learning for human activitiesabstractCurrent Vision-Language (VL) models owe their success to large-scale pre-training on webcollected data, which in turn requires highcapacity architectures and large compute resources for training.We posit that when the downstream tasks are known in advance, which is in practice common, the pretraining process can be aligned to the downstream domain, leading to more efficient and accurate models, while shortening the pretraining step.To this end, we introduce a domain-aligned pretraining strategy that, without additional data collection, improves the accuracy on a domain of interest, herein, that of human activities, while largely preserving the generalist knowledge.At the core of our approach stands a new LLM-based method that, provided with a simple set of concept seeds, produces a concept hierarchy with high coverage of the target domain.The concept hierarchy is used to filter a large-scale webcrawled dataset and, then, enhance the resulting instances with targeted synthetic labels.We study in depth how to train such approaches and their resulting behavior.We further show generalization to video-based data by introducing a fast adaptation approach for transitioning from a static (image) model to a dynamic one (i.e. with temporal modeling).On the domain of interest, our approach significantly outperforms models trained on up to 60× more samples and between 10 -100× shorter training schedules for image retrieval, video retrieval and action recognition.Code will be released. Adrian Bulat, Yassine Ouali, Ricardo Guerrero, Brais Martínez, Georgios Tzimiropoulos |
EMNLP | 3 |
| 2023 | FS-DETR: Few-Shot DEtection TRansformer with prompting and without re-trainingabstractThis paper is on Few-Shot Object Detection (FSOD), where given a few templates (examples) depicting a novel class (not seen during training), the goal is to detect all of its occurrences within a set of images. From a practical perspective, an FSOD system must fulfil the following desiderata: (a) it must be used as is, without requiring any fine-tuning at test time, (b) it must be able to process an arbitrary number of novel objects concurrently while supporting an arbitrary number of examples from each class and (c) it must achieve accuracy comparable to a closed system. Towards satisfying (a)-(c), in this work, we make the following contributions: We introduce, for the first time, a simple, yet powerful, few-shot detection transformer (FS-DETR) based on visual prompting that can address both desiderata (a) and (b). Our system builds upon the DETR framework, extending it based on two key ideas: (1) feed the provided visual templates of the novel classes as visual prompts during test time, and (2) "stamp" these prompts with pseudo-class embeddings (akin to soft prompting), which are then predicted at the output of the decoder. Importantly, we show that our system is not only more flexible than existing methods, but also, it makes a step towards satisfying desideratum (c). Specifically, it is significantly more accurate than all methods that do not require fine-tuning and even matches and outperforms the current state-of-the-art fine-tuning based methods on the most well-established benchmarks (PASCAL VOC & MSCOCO). Adrian Bulat, Ricardo Guerrero, Brais Martínez, Georgios Tzimiropoulos |
ICCV | 2 |
| 2023 | D2F2WOD: Learning Object Proposals for Weakly-Supervised Object Detection via Progressive Domain AdaptationabstractWeakly-supervised object detection (WSOD) models attempt to leverage image-level annotations in lieu of accurate but costly-to-obtain object localization labels. This oftentimes leads to substandard object detection and lo-calization at inference time. To tackle this issue, we propose D2F2WOD, a Dual-Domain Fully-to-Weakly Supervised Object Detection framework that leverages synthetic data, annotated with precise object localization, to supplement a natural image target domain, where only image-level labels are available. In its warm-up domain adaptation stage, the model learns a fully-supervised object detector (FSOD) to improve the precision of the object proposals in the target domain, and at the same time learns target-domain-specific and detection-aware proposal features. In its main WSOD stage, a WSOD model is specifically tuned to the target domain. The feature extractor and the object proposal generator of the WSOD model are built upon the fine-tuned FSOD model. We test D2F2WOD on five dual-domain image benchmarks. The results show that our method results in consistently improved object detection and localization compared with state-of-the-art methods. Yuting Wang 0004, Ricardo Guerrero, Vladimir Pavlovic 0001 |
WACV | 2 |
| 2022 | Variational Continual Proxy-Anchor for Deep Metric LearningabstractThe recent proxy-anchor method achieved outstanding performance in deep metric learning, which can be acknowledged to its data efficient loss based on hard example mining, as well as far lower sampling complexity than pair-based approaches. In this paper we extend the proxy-anchor method by posing it within the continual learning framework, motivated from its batch-expected loss form (instead of instance-expected, typical in deep learning), which can potentially incur the catastrophic forgetting of historic batches. By regarding each batch as a task in continual learning, we adopt the Bayesian variational continual learning approach to derive a novel loss function. Interestingly the resulting loss has two key modifications to the original proxy-anchor loss: i) we inject noise to the proxies when optimizing the proxy-anchor loss, and ii) we encourage momentum update to avoid abrupt model changes. As a result, the learned model achieves higher test accuracy than proxy-anchor due to the robustness to noise in data (through model perturbation during training), and the reduced batch forgetting effect. We demonstrate the improved results on several benchmark datasets. Minyoung Kim 0001, Ricardo Guerrero, Hai Xuan Pham, Vladimir Pavlovic 0001 |
AISTATS | 2 |
| 2022 | SOS! Self-supervised Learning over Sets of Handled Objects in Egocentric Action Recognition
Victor Escorcia, Ricardo Guerrero, Xiatian Zhu, Brais Martínez |
ECCV (13) | 2 |
| 2021 | CHEF: Cross-modal Hierarchical Embeddings for Food Domain RetrievalabstractDespite the abundance of multi-modal data, such as image-text pairs, there has been little effort in understanding the individual entities and their different roles in the construction of these data instances. In this work, we endeavour to discover the entities and their corresponding importance in cooking recipes automatically as a visual-linguistic association problem. More specifically, we introduce a novel cross-modal learning framework to jointly model the latent representations of images and text in the food image-recipe association and retrieval tasks. This model allows one to discover complex functional and hierarchical relationships between images and text, and among textual parts of a recipe including title, ingredients and cooking instructions. Our experiments show that by making use of efficient tree-structured Long Short-Term Memory as the text encoder in our computational cross-modal retrieval framework, we are not only able to identify the main ingredients and cooking actions in the recipe descriptions without explicit supervision, but we can also learn more meaningful feature representations of food recipes, appropriate for challenging cross-modal retrieval and recipe adaption tasks. Hai Xuan Pham, Ricardo Guerrero, Vladimir Pavlovic 0001, Jiatong Li 0001 |
AAAI | 2 |
| 2021 | Multi-attribute Pizza Generator: Cross-domain Attribute Control with Conditional StyleGAN
Fangda Han, Guoyao Hao, Ricardo Guerrero, Vladimir Pavlovic 0001 |
BMVC | 3 |
| 2021 | Cross-modal Retrieval and Synthesis (X-MRS): Closing the Modality Gap in Shared Subspace LearningabstractComputational food analysis (CFA) naturally requires multi-modal evidence of a particular food, e.g., images, recipe text, etc. A key to making CFA possible is multi-modal shared representation learning, which aims to create a joint representation of the multiple views (text and image) of the data. In this work we propose a method for food domain cross-modal shared representation learning that preserves the vast semantic richness present in the food data. Our proposed method employs an effective transformer-based multilingual recipe encoder coupled with a traditional image embedding architecture. Here, we propose the use of imperfect multilingual translations to effectively regularize the model while at the same time adding functionality across multiple languages and alphabets. Experimental analysis on the public Recipe1M dataset shows that the representation learned via the proposed method significantly outperforms the current state-of-the-arts (SOTA) on retrieval tasks. Furthermore, the representational power of the learned representation is demonstrated through a generative food image synthesis model conditioned on recipe embeddings. Synthesized images can effectively reproduce the visual appearance of paired samples, indicating that the learned representation captures the joint semantics of both the textual recipe and its visual content, thus narrowing the modality gap. Ricardo Guerrero, Hai Xuan Pham, Vladimir Pavlovic 0001 |
ACM Multimedia | 1 |
| 2021 | AIxFood'21: 3rd Workshop on AIxFoodabstractFood and cooking analysis present exciting research and application challenges for modern AI systems, particularly in the context of multimodal data such as images or video. A meal that appears in a food image is a product of a complex progression of cooking stages, often described in the accompanying textual recipe form. In the cooking process, individual ingredients change their physical properties, become combined with other food components, all to produce a final, yet highly variable, appearance of the meal. Recognizing food items or meals on a plate from images or videos, their physical properties such as the amount, nutritional content such as the caloric value, food attributes such as the flavor, elucidating the cooking process behind it, or creating robotic assistants that help users complete that cooking process, is of essential scientific and technological value yet technically extremely challenging. The 3rd AIxFood workshop was held as a half-day workshop in conjunction with the 29th ACM International Conference on Multimedia (ACM MM 2021), in Chengdu, China and virtually. Ricardo Guerrero, Michael Spranger, Shuqiang Jiang, Chong-Wah Ngo |
ACM Multimedia | 1 |
| 2021 | Learning Disentangled Factors from Paired Data in Cross-Modal Retrieval: An Implicit Identifiable VAE ApproachabstractWe tackle the problem of learning the underlying disentangled latent factors that are shared between the paired bi-modal data in cross-modal retrieval. Typically the data in both modalities are complex, structured, and high dimensional (e.g., image and text), for which the conventional deep auto-encoding latent variable models such as the Variational Autoencoder (VAE) often suffer from difficulty of accurate decoder training or realistic synthesis. In this paper we propose a novel idea of the implicit decoder, which completely removes the ambient data decoding module from a latent variable model, via implicit encoder inversion that is achieved by Jacobian regularization of the low-dimensional embedding function. Motivated from the recent Identifiable-VAE (IVAE) model, we modify it to incorporate the query modality data as conditioning auxiliary input, which allows us to prove that the true parameters of the model can be identifiable under some regularity conditions. Tested on various datasets where the true factors are fully/partially available, our model is shown to identify the factors accurately, significantly outperforming conventional latent variable models. Minyoung Kim 0001, Ricardo Guerrero, Vladimir Pavlovic 0001 |
ACM Multimedia | 2 |
| 2020 | Picture-to-Amount (PITA): Predicting Relative Ingredient Amounts from Food ImagesabstractIncreased awareness of the impact of food consumption on health and lifestyle today has given rise to novel data-driven food analysis systems. Although these systems may recognize the ingredients, a detailed analysis of their amounts in the meal, which is paramount for estimating the correct nutrition, is usually ignored. In this paper, we study the novel and challenging problem of predicting the relative amount of each ingredient from a food image. We propose PITA, the Picture-to-Amount deep learning architecture to solve the problem. More specifically, we predict the ingredient amounts using a domain-driven Wasserstein loss from image-to-recipe cross-modal embeddings learned to align the two views of food data. Experiments on a dataset of recipes collected from the Internet show the model generates promising results and improves the baselines on this challenging task. A demo of our system and our data is available at: foodai.cs.rutgers.edu. Jiatong Li 0001, Fangda Han, Ricardo Guerrero, Vladimir Pavlovic 0001 |
ICPR | 3 |
| 2020 | CookGAN: Meal Image Synthesis from IngredientsabstractIn this work we propose a new computational framework, based on generative deep models, for synthesis of photo-realistic food meal images from textual list of its ingredients. Previous works on synthesis of images from text typically rely on pre-trained text models to extract text features, followed by generative neural networks (GAN) aimed to generate realistic images conditioned on the text features. These works mainly focus on generating spatially compact and well-defined categories of objects, such as birds or flowers, but meal images are significantly more complex, consisting of multiple ingredients whose appearance and spatial qualities are further modified by cooking methods. To generate real-like meal images from ingredients, we propose Cook Generative Adversarial Networks (CookGAN), CookGAN first builds an attention-based ingredients-image association model, which is then used to condition a generative neural network tasked with synthesizing meal images. Furthermore, a cycle-consistent constraint is added to further improve image quality and control appearance. Experiments show our model is able to generate meal images corresponding to the ingredients. Fangda Han, Ricardo Guerrero, Vladimir Pavlovic 0001 |
WACV | 2 |
| 2018 | Automatic View Planning with Multi-scale Deep Reinforcement Learning Agents
Amir Alansary, Loïc Le Folgoc, Ghislain Vaillant, Ozan Oktay, Wenjia Bai, Jonathan Passerat-Palmbach, Ricardo Guerrero, Konstantinos Kamnitsas, Benjamin Hou, Steven McDonagh 0001, Ben Glocker, Bernhard Kainz, Daniel Rueckert |
MICCAI (1) | 8 |
| 2018 | Disease prediction using graph convolutional networks: Application to Autism Spectrum Disorder and Alzheimer's disease
Sarah Parisot, Sofia Ira Ktena, Enzo Ferrante, Matthew C. H. Lee, Ricardo Guerrero, Ben Glocker, Daniel Rueckert |
Medical Image Anal. | 5 |
| 2018 | A large margin algorithm for automated segmentation of white matter hyperintensity
Chen Qin, Ricardo Guerrero, Christopher Bowles, Liang Chen 0018, David Alexander Dickie, Maria del C. Valdés Hernández, Joanna M. Wardlaw, Daniel Rueckert |
Pattern Recognit. | 2 |
| 2017 | Spectral Graph Convolutions for Population-Based Disease Prediction
Sarah Parisot, Sofia Ira Ktena, Enzo Ferrante, Matthew C. H. Lee, Ricardo Guerrero, Ben Glocker, Daniel Rueckert |
MICCAI (3) | 5 |
| 2017 | Group-constrained manifold learning: Application to AD risk assessment
Ricardo Guerrero, Christian Ledig, Alexander Schmidt-Richberg, Daniel Rueckert |
Pattern Recognit. | 1 |
| 2017 | Stratified Decision Forests for Accurate Anatomical Landmark Localization in Cardiac ImagesabstractAccurate localization of anatomical landmarks is an important step in medical imaging, as it provides useful prior information for subsequent image analysis and acquisition methods. It is particularly useful for initialization of automatic image analysis tools (e.g. segmentation and registration) and detection of scan planes for automated image acquisition. Landmark localization has been commonly performed using learning based approaches, such as classifier and/or regressor models. However, trained models may not generalize well in heterogeneous datasets when the images contain large differences due to size, pose and shape variations of organs. To learn more data-adaptive and patient specific models, we propose a novel stratification based training model, and demonstrate its use in a decision forest. The proposed approach does not require any additional training information compared to the standard model training procedure and can be easily integrated into any decision tree framework. The proposed method is evaluated on 1080 3D high-resolution and 90 multi-stack 2D cardiac cine MR images. The experiments show that the proposed method achieves state-of-the-art landmark localization accuracy and outperforms standard regression and classification based approaches. Additionally, the proposed method is used in a multi-atlas segmentation to create a fully automatic segmentation pipeline, and the results show that it achieves state-of-the-art segmentation accuracy. Ozan Oktay, Wenjia Bai, Ricardo Guerrero, Martin Rajchl, Antonio M. Simoes Monteiro de Marvao, Declan P. O'Regan, Stuart A. Cook, Mattias P. Heinrich, Ben Glocker, Daniel Rueckert |
IEEE Trans. Medical Imaging | 3 |
| 2016 | Multi-input Cardiac Image Super-Resolution Using Convolutional Neural Networksabstract3D cardiac MR imaging enables accurate analysis of cardiac morphology and physiology. However, due to the requirements for long acquisition and breath-hold, the clinical routine is still dominated by multi-slice 2D imaging, which hamper the visualization of anatomy and quantitative measurements as relatively thick slices are acquired. As a solution, we propose a novel image super-resolution (SR) approach that is based on a residual convolutional neural network (CNN) model. It reconstructs high resolution 3D volumes from 2D image stacks for more accurate image analysis. The proposed model allows the use of multiple input data acquired from different viewing planes for improved performance. Experimental results on 1233 cardiac short and long-axis MR image stacks show that the CNN model outperforms state-of-the-art SR methods in terms of image quality while being computationally efficient. Also, we show that image segmentation and motion tracking benefits more from SR-CNN when it is used as an initial upscaling method than conventional interpolation methods for the subsequent analysis. Ozan Oktay, Wenjia Bai, Matthew C. H. Lee, Ricardo Guerrero, Konstantinos Kamnitsas, Jose Caballero, Antonio M. Simoes Monteiro de Marvao, Stuart A. Cook, Declan P. O'Regan, Daniel Rueckert |
MICCAI (3) | 4 |
| 2014 | Multiple instance learning for classification of dementia in brain MRI
Tong Tong 0001, Robin Wolz, Qinquan Gao, Ricardo Guerrero, Joseph V. Hajnal, Daniel Rueckert |
Medical Image Anal. | 4 |
| 2011 | Laplacian Eigenmaps Manifold Learning for Landmark Localization in Brain MR Images
Ricardo Guerrero, Robin Wolz, Daniel Rueckert |
MICCAI (2) | 1 |