Xavier Giró-i-Nieto

dblp:12/7205 · also Xavier Giró · DBLP profile ↗
← Back
53ranked-venue papers
5as first author
15since 2021 · last 2025
0000-0002-9935-5332ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 20 · 9 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Compressive Meta-Learning
abstract
The rapid expansion in the size of new datasets has created a need for fast and efficient parameter-learning techniques.Compressive learning is a framework that enables efficient processing by using random, nonlinear features to project large-scale databases onto compact, information-preserving representations whose dimensionality is independent of the number of samples and can be easily stored, transferred, and processed.These database-level summaries are then used to decode parameters of interest from the underlying data distribution without requiring access to the original samples, offering an efficient and privacy-friendly learning framework.However, both the encoding and decoding techniques are typically randomized and data-independent, failing to exploit the underlying structure of the data.In this work, we propose a framework that meta-learns both the encoding and decoding stages of compressive learning methods by using neural networks that provide faster and more accurate systems than the current state-of-the-art approaches.To demonstrate the potential of the presented Compressive Meta-Learning framework, we explore multiple applications-including neural network-based compressive PCA, compressive ridge regression, compressive k-means, and autoencoders.
Daniel Mas Montserrat, David Bonet, Maria Perera, Xavier Giró-i-Nieto, Alexander G. Ioannidis
KDD (2)4
2024 HyperFast: Instant Classification for Tabular Data
abstract
Training deep learning models and performing hyperparameter tuning can be computationally demanding and time-consuming. Meanwhile, traditional machine learning methods like gradient-boosting algorithms remain the preferred choice for most tabular data applications, while neural network alternatives require extensive hyperparameter tuning or work only in toy datasets under limited settings. In this paper, we introduce HyperFast, a meta-trained hypernetwork designed for instant classification of tabular data in a single forward pass. HyperFast generates a task-specific neural network tailored to an unseen dataset that can be directly used for classification inference, removing the need for training a model. We report extensive experiments with OpenML and genomic data, comparing HyperFast to competing tabular data neural networks, traditional ML methods, AutoML systems, and boosting machines. HyperFast shows highly competitive results, while being significantly faster. Additionally, our approach demonstrates robust adaptability across a variety of classification tasks with little to no fine-tuning, positioning HyperFast as a strong solution for numerous applications and rapid model deployment. HyperFast introduces a promising paradigm for fast classification, with the potential to substantially decrease the computational burden of deep learning. Our code, which offers a scikit-learn-like interface, along with the trained HyperFast model, can be found at https://github.com/AI-sandbox/HyperFast.
David Bonet, Daniel Mas Montserrat, Xavier Giró-i-Nieto, Alexander G. Ioannidis
AAAI3
2023 Genomic Databases Homogenization with Machine Learning
abstract
Large-scale and increasingly diverse datasets power modern genomic studies, yet robust data integration and homogenization across varying sources remains a challenge. The multiplicity of file formats and the computational requirements imposed by large genomic datasets make it difficult to deal with multiple data sources. Furthermore, there is a lack of open-source customizable tools to merge genomic databases while providing quality control functionalities. To fill this gap, we present MergeGenome, a machine learning-based method designed to integrate DNA sequences from multiple variant call format (VCF) files while maintaining data quality. By leveraging pre-existing VCF manipulation and imputation software, MergeGenome provides a robust pipeline of comprehensive steps to standardize nomenclature, remove ambiguities, correct strand alignment, eliminate mismatches, impute missing positions, and filter and correct erroneous variants with machine learning, among other functionalities. We demonstrate MergeGenome’s ability to obtain a high-quality combined dataset by merging two databases containing dog DNA and effectively detecting and correcting imputation errors. Finally, we show that using the homogenized dataset provides a boost in phenotype prediction performance.
Míriam Barrabés, David Bonet, Víctor Novelle Moriano, Xavier Giró-i-Nieto, Daniel Mas Montserrat, Alexander G. Ioannidis
BIBM4
2023 Adversarial Learning for Feature Shift Detection and Correction
abstract
Data shift is a phenomenon present in many real-world applications, and while there are multiple methods attempting to detect shifts, the task of localizing and correcting the features originating such shifts has not been studied in depth. Feature shifts can occur in many datasets, including in multi-sensor data, where some sensors are malfunctioning, or in tabular and structured data, including biomedical, financial, and survey data, where faulty standardization and data processing pipelines can lead to erroneous features. In this work, we explore using the principles of adversarial learning, where the information from several discriminators trained to distinguish between two distributions is used to both detect the corrupted features and fix them in order to remove the distribution shift between datasets. We show that mainstream supervised classifiers, such as random forest or gradient boosting trees, combined with simple iterative heuristics, can localize and correct feature shifts, outperforming current statistical and neural network-based techniques. The code is available at https://github.com/AI-sandbox/DataFix.
Míriam Barrabés, Daniel Mas Montserrat, Margarita Geleta, Xavier Giró-i-Nieto, Alexander G. Ioannidis
NeurIPS4
2023 SIRA: Relightable Avatars from a Single Image
abstract
Recovering the geometry of a human head from a single image, while factorizing the materials and illumination, is a severely ill-posed problem that requires prior information to be solved. Methods based on 3D Morphable Models (3DMM), and their combination with differentiable renderers, have shown promising results. However, the expressiveness of 3DMMs is limited, and they typically yield over-smoothed and identity-agnostic 3D shapes limited to the face region. Highly accurate full head reconstructions have recently been obtained with neural fields that parameterize the geometry using multilayer perceptrons. The versatility of these representations has also proved effective for disentangling geometry, materials and lighting. However, these methods require several tens of input images. In this paper, we introduce SIRA, a method which, from a single image, reconstructs human head avatars with high fidelity geometry and factorized lights and surface materials. Our key ingredients are two data-driven statistical models based on neural fields that resolve the ambiguities of single-view 3D surface reconstruction and appearance factorization. Experiments show that SIRA obtains state of the art results in 3D head reconstruction while at the same time it successfully disentangles the global illumination, and the diffuse and specular albedos. Furthermore, our reconstructions are amenable to physically-based appearance editing and head model relighting.
Pol Caselles, Eduard Ramon, Jaime García 0001, Xavier Giró-i-Nieto, Francesc Moreno-Noguer, Gil Triginer
WACV4
2023 The Liver Tumor Segmentation Benchmark (LiTS)
abstract
In this work, we report the set-up and results of the Liver Tumor Segmentation Benchmark (LiTS), which was organized in conjunction with the IEEE International Symposium on Biomedical Imaging (ISBI) 2017 and the International Conferences on Medical Image Computing and Computer-Assisted Intervention (MICCAI) 2017 and 2018. The image dataset is diverse and contains primary and secondary tumors with varied sizes and appearances with various lesion-to-background levels (hyper-/hypo-dense), created in collaboration with seven hospitals and research institutions. Seventy-five submitted liver and liver tumor segmentation algorithms were trained on a set of 131 computed tomography (CT) volumes and were tested on 70 unseen test images acquired from different patients. We found that not a single algorithm performed best for both liver and liver tumors in the three events. The best liver segmentation algorithm achieved a Dice score of 0.963, whereas, for tumor segmentation, the best algorithms achieved Dices scores of 0.674 (ISBI 2017), 0.702 (MICCAI 2017), and 0.739 (MICCAI 2018). Retrospectively, we performed additional analysis on liver tumor detection and revealed that not all top-performing segmentation algorithms worked well for tumor detection. The best liver tumor detection method achieved a lesion-wise recall of 0.458 (ISBI 2017), 0.515 (MICCAI 2017), and 0.554 (MICCAI 2018), indicating the need for further research. LiTS remains an active benchmark and resource for research, e.g., contributing the liver-related segmentation tasks in http://medicaldecathlon.com/. In addition, both data and online evaluation are accessible via https://competitions.codalab.org/competitions/17094.
Patrick Bilic, Patrick Ferdinand Christ, Hongwei Li 0004, Eugene Vorontsov, Avi Ben-Cohen, Georgios Kaissis, Adi Szeskin, Colin Jacobs, Gabriel Efrain Humpire Mamani, Gabriel Chartrand, Fabian Lohöfer, Julian Walter Holch, Wieland H. Sommer, Felix Hofmann, Alexandre Hostettler, Naama Lev-Cohain, Michal Drozdzal, Michal Amitai, Refael Vivanti, Jacob Sosna, Ivan Ezhov, Anjany Sekuboyina, Fernando Navarro, Florian Kofler, Johannes C. Paetzold, Suprosanna Shit, Xiaobin Hu, Jana Lipková, Markus Rempfler, Marie Piraud, Jan Kirschke, Benedikt Wiestler, Christian Hülsemeyer, Marcel Beetz, Florian Ettlinger, Michela Antonelli, Woong Bae, Miriam Bellver, Lei Bi 0001, Hao Chen 0011, Grzegorz Chlebus, Erik Dam, Qi Dou 0001, Chi-Wing Fu, Bogdan Georgescu, Xavier Giró-i-Nieto, Felix Grün, Xu Han 0009, Pheng-Ann Heng, Jürgen Hesser, Jan Hendrik Moltz, Christian Igel, Fabian Isensee, Paul F. Jaeger, Fucang Jia, Krishna Chaitanya Kaluva, Mahendra Khened, Ildoo Kim, Jae-Hun Kim, Sungwoong Kim, Simon Kohl, Tomasz K. Konopczynski, Avinash Kori, Ganapathy Krishnamurthi, Xiaomeng Li 0001, John S. Lowengrub, Jun Ma 0016, Klaus H. Maier-Hein, Kevis-Kokitsi Maninis, Hans Meine, Dorit Merhof, Akshay Pai, Mathias Perslev, Jens Petersen, Jordi Pont-Tuset, Xiaojuan Qi 0001, Oliver Rippel, Karsten Roth, Ignacio Sarasua, Andrea Schenk, Zengming Shen, Jordi Torres, Christian Wachinger, Chunliang Wang, Leon Weninger, Daguang Xu, Xiaoping Yang 0001, Simon C. H. Yu, Yading Yuan, Miao Yue, Liping Zhang 0009, Manuel Jorge Cardoso, Spyridon Bakas, Rickmer Braren, Volker Heinemann, Christopher Joseph Pal, An Tang, Samuel Kadoury, Luc Soler, Bram van Ginneken, Hayit Greenspan, Leo Joskowicz, Bjoern Menze
Medical Image Anal.47
2023 A closer look at referring expressions for video object segmentation
abstract
Abstract The task of Language-guided Video Object Segmentation (LVOS) aims at generating binary masks for an object referred by a linguistic expression. When this expression unambiguously describes an object in the scene, it is namedreferring expression(RE). Our work argues that existing benchmarks used for LVOS are mainly composed of trivial cases, in which referents can be identified with simple phrases. Our analysis relies on a new categorization of the referring expressions in the DAVIS-2017 and Actor-Action datasets into trivial and non-trivial REs, where the non-trivial REs are further annotated with seven RE semantic categories. We leverage these data to analyze the performance of RefVOS, a novel neural network that obtains competitive results for the task of language-guided image segmentation and state of the art results for LVOS. Our study indicates that the major challenges for the task are related to understanding motion and static actions.
Miriam Bellver, Carles Ventura, Carina Silberer, Ioannis Kazakos, Jordi Torres, Xavier Giró-i-Nieto
Multim. Tools Appl.6
2022 Sign Language Video Retrieval with Free-Form Textual Queries
abstract
Systems that can efficiently search collections of sign language videos have been highlighted as a useful application of sign language technology. However, the problem of searching videos beyond individual keywords has received limited attention in the literature. To address this gap, in this work we introduce the task of sign language retrieval with free-form11The terminology “natural language query” is commonly used to describe unconstrained textual queries in spoken languages. However, since sign languages are also natural languages, we adopt for the term “free-form textual query” instead. textual queries: given a written query (e.g. a sentence) and a large collection of sign language videos, the objective is to find the signing video that best matches the written query. We propose to tackle this task by learning cross-modal embeddings on the recently introduced large-scale How2Sign dataset of American Sign Language (ASL). We identify that a key bottleneck in the performance of the system is the quality of the sign video embedding which suffers from a scarcity of labelled training data. We, therefore, propose SPOT-ALIGN, a framework for interleaving iterative rounds of sign spotting and feature alignment to expand the scope and scale of available training data. We validate the effectiveness of SPOT-ALIGN for learning a robust sign video embedding through improvements in both sign recognition and the proposed video retrieval task.
Amanda Cardoso Duarte, Samuel Albanie, Xavier Giró-i-Nieto, Gül Varol
CVPR3
2022 Pixinwav: Residual Steganography for Hiding Pixels in Audio
abstract
Steganography comprises the mechanics of hiding data in a host media that may be publicly available. While previous works focused on unimodal setups (e.g., hiding images in images, or hiding audio in audio), PixInWav targets the multimodal case of hiding images in audio. To this end, we propose a novel residual architecture operating on top of short-time discrete cosine transform (STDCT) audio spectrograms. Among our results, we find that the residual steganography setup we propose allows an encoding of the hidden image that is independent from the host audio without compromising quality. Accordingly, while previous works require both host and hidden signals to hide a signal, PixInWav can encode images offline—which can be later hidden, in a residual fashion, into any audio signal.
Margarita Geleta, Cristina Punti, Kevin McGuinness, Jordi Pons, Cristian Canton, Xavier Giró-i-Nieto
ICASSP6
2022 Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights
abstract
Learning representations of neural network weights given a model zoo is an emerg- ing and challenging area with many potential applications from model inspection, to neural architecture search or knowledge distillation. Recently, an autoencoder trained on a model zoo was able to learn a hyper-representation, which captures intrinsic and extrinsic properties of the models in the zoo. In this work, we ex- tend hyper-representations for generative use to sample new model weights. We propose layer-wise loss normalization which we demonstrate is key to generate high-performing models and several sampling methods based on the topology of hyper-representations. The models generated using our methods are diverse, per- formant and capable to outperform strong baselines as evaluated on several down- stream tasks: initialization, ensemble sampling and transfer learning. Our results indicate the potential of knowledge aggregation from model zoos to new models via hyper-representations thereby paving the avenue for novel research directions.
Konstantin Schürholt, Xavier Giró-i-Nieto, Damian Borth
NeurIPS3
2022 Model Zoos: A Dataset of Diverse Populations of Neural Network Models
abstract
In the last years, neural networks (NN) have evolved from laboratory environments to the state-of-the-art for many real-world problems. It was shown that NN models (i.e., their weights and biases) evolve on unique trajectories in weight space during training. Following, a population of such neural network models (referred to as model zoo) would form structures in weight space. We think that the geometry, curvature and smoothness of these structures contain information about the state of training and can reveal latent properties of individual models. With such model zoos, one could investigate novel approaches for (i) model analysis, (ii) discover unknown learning dynamics, (iii) learn rich representations of such populations, or (iv) exploit the model zoos for generative modelling of NN weights and biases. Unfortunately, the lack of standardized model zoos and available benchmarks significantly increases the friction for further research about populations of NNs. With this work, we publish a novel dataset of model zoos containing systematically generated and diverse populations of NN models for further research. In total the proposed model zoo dataset is based on eight image datasets, consists of 27 model zoos trained with varying hyperparameter combinations and includes 50’360 unique NN models as well as their sparsified twins, resulting in over 3’844’360 collected model states. Additionally, to the model zoo data we provide an in-depth analysis of the zoos and provide benchmarks for multiple downstream tasks. The dataset can be found at www.modelzoos.cc.
Konstantin Schürholt, Diyar Taskiran, Xavier Giró-i-Nieto, Damian Borth
NeurIPS4
2022 SALAI-Net: species-agnostic local ancestry inference network
abstract
MOTIVATION: Local ancestry inference (LAI) is the high resolution prediction of ancestry labels along a DNA sequence. LAI is important in the study of human history and migrations, and it is beginning to play a role in precision medicine applications including ancestry-adjusted genome-wide association studies (GWASs) and polygenic risk scores (PRSs). Existing LAI models do not generalize well between species, chromosomes or even ancestry groups, requiring re-training for each different setting. Furthermore, such methods can lack interpretability, which is an important element in each of these applications. RESULTS: We present SALAI-Net, a portable statistical LAI method that can be applied on any set of species and ancestries (species-agnostic), requiring only haplotype data and no other biological parameters. Inspired by identity by descent methods, SALAI-Net estimates population labels for each segment of DNA by performing a reference matching approach, which leads to an interpretable and fast technique. We benchmark our models on whole-genome data of humans and we test these models' ability to generalize to dog breeds when trained on human data. SALAI-Net outperforms previous methods in terms of balanced accuracy, while generalizing between different settings, species and datasets. Moreover, it is up to two orders of magnitude faster and uses considerably less RAM memory than competing methods. AVAILABILITY AND IMPLEMENTATION: We provide an open source implementation and links to publicly available data at github.com/AI-sandbox/SALAI-Net. Data is publicly available as follows: https://www.internationalgenome.org (1000 Genomes), https://www.simonsfoundation.org/simons-genome-diversity-project (Simons Genome Diversity Project), https://www.sanger.ac.uk/resources/downloads/human/hapmap3.html (HapMap), ftp://ngs.sanger.ac.uk/production/hgdp/hgdp_wgs.20190516 (Human Genome Diversity Project) and https://www.ncbi.nlm.nih.gov/bioproject/PRJNA448733 (Canid genomes). SUPPLEMENTARY INFORMATION: Supplementary data are available from Bioinformatics online.
Benet Oriol Sabat, Daniel Mas Montserrat, Xavier Giró-i-Nieto, Alexander G. Ioannidis
Bioinform.3
2021 How2Sign: A Large-Scale Multimodal Dataset for Continuous American Sign Language
abstract
One of the factors that have hindered progress in the areas of sign language recognition, translation, and production is the absence of large annotated datasets. Towards this end, we introduce How2Sign, a multimodal and multiview continuous American Sign Language (ASL) dataset, consisting of a parallel corpus of more than 80 hours of sign language videos and a set of corresponding modalities including speech, English transcripts, and depth. A three-hour subset was further recorded in the Panoptic studio enabling detailed 3D pose estimation. To evaluate the potential of How2Sign for real-world impact, we conduct a study with ASL signers and show that synthesized videos using our dataset can indeed be understood. The study further gives insights on challenges that computer vision should address in order to make progress in this field.
Amanda Cardoso Duarte, Shruti Palaskar, Lucas Ventura, Deepti Ghadiyaram, Kenneth DeHaan, Florian Metze, Jordi Torres, Xavier Giró-i-Nieto
CVPR8
2021 Seasonal Contrast: Unsupervised Pre-Training from Uncurated Remote Sensing Data
abstract
Remote sensing and automatic earth monitoring are key to solve global-scale challenges such as disaster prevention, land use monitoring, or tackling climate change. Although there exist vast amounts of remote sensing data, most of it remains unlabeled and thus inaccessible for supervised learning algorithms. Transfer learning approaches can reduce the data requirements of deep learning algorithms. However, most of these methods are pre-trained on ImageNet and their generalization to remote sensing imagery is not guaranteed due to the domain gap. In this work, we propose Seasonal Contrast (SeCo), an effective pipeline to leverage unlabeled data for in-domain pre-training of remote sensing representations. The SeCo pipeline is composed of two parts. First, a principled procedure to gather large-scale, unlabeled and uncurated remote sensing datasets containing images from multiple Earth locations at different timestamps. Second, a self-supervised algorithm that takes advantage of time and position invariance to learn transferable representations for remote sensing applications. We empirically show that models trained with SeCo achieve better performance than their ImageNet pre-trained counterparts and state-of-the-art self-supervised learning methods on multiple downstream tasks. The datasets and models in SeCo will be made public to facilitate transfer learning and enable rapid progress in remote sensing applications.1
Oscar Mañas, Alexandre Lacoste, Xavier Giró-i-Nieto, David Vázquez 0001, Pau Rodríguez
ICCV3
2021 H3D-Net: Few-Shot High-Fidelity 3D Head Reconstruction
abstract
Recent learning approaches that implicitly represent surface geometry using coordinate-based neural representations have shown impressive results in the problem of multi-view 3D reconstruction. The effectiveness of these techniques is, however, subject to the availability of a large number (several tens) of input views of the scene, and computationally demanding optimizations. In this paper, we tackle these limitations for the specific problem of few-shot full 3D head reconstruction, by endowing coordinate-based representations with a probabilistic shape prior that enables faster convergence and better generalization when using few input images (down to three). First, we learn a shape model of 3D heads from thousands of incomplete raw scans using implicit representations. At test time, we jointly overfit two coordinate-based neural networks to the scene, one modelling the geometry and another estimating the surface radiance, using implicit differentiable rendering. We devise a two-stage optimization strategy in which the learned prior is used to initialize and constrain the geometry during an initial optimization phase. Then, the prior is unfrozen and fine-tuned to the scene. By doing this, we achieve high-fidelity head reconstructions, including hair and shoulders, and with a high level of detail that consistently outperforms both state-of-the-art 3D Morphable Models methods in the few-shot scenario, and nonparametric methods when large sets of views are available.
Eduard Ramon, Gil Triginer, Janna Escur, Albert Pumarola, Jaime García 0001, Xavier Giró-i-Nieto, Francesc Moreno-Noguer
ICCV6
2020 Explore, Discover and Learn: Unsupervised Discovery of State-Covering Skills
abstract
Acquiring abilities in the absence of a task-oriented reward function is at the frontier of reinforcement learning research. This problem has been studied through the lens of empowerment, which draws a connection between option discovery and information theory. Information-theoretic skill discovery methods have garnered much interest from the community, but little research has been conducted in understanding their limitations. Through theoretical analysis and empirical evidence, we show that existing algorithms suffer from a common limitation – they discover options that provide a poor coverage of the state space. In light of this, we propose Explore, Discover and Learn (EDL), an alternative approach to information-theoretic skill discovery. Crucially, EDL optimizes the same information-theoretic objective derived from the empowerment literature, but addresses the optimization problem using different machinery. We perform an extensive evaluation of skill discovery methods on controlled environments and show that EDL offers significant advantages, such as overcoming the coverage problem, reducing the dependence of learned skills on the initial state, and allowing the user to define a prior over which behaviors should be learned.
Victor Campos 0001, Alexander Trott, Caiming Xiong, Richard Socher, Xavier Giró-i-Nieto, Jordi Torres
ICML5
2020 Automatic Reminiscence Therapy for Dementia
abstract
With people living longer than ever, the number of cases with dementia such as Alzheimer's disease increases steadily. It affects more than 46 million people worldwide, and it is estimated that in 2050 more than 100 million will be affected. While there are no effective treatments for these terminal diseases, therapies such as reminiscence, that stimulate memories from the past are recommended. Currently, reminiscence therapy takes place in care homes and is guided by a therapist or a carer. In this work, we present an AI-based solution to automate the reminiscence therapy. This consists of a dialogue system that uses photos of the users as input to generate questions about their life. Overall, this paper presents how reminiscence therapy can be automated by using deep learning, and deployed to smartphones and laptops, making the therapy more accessible to every person affected by dementia.
Mariona Caros, Maite Garolera, Petia Radeva, Xavier Giró-i-Nieto
ICMR4
2020 One Perceptron to Rule Them All: Language, Vision, Audio and Speech
abstract
Deep neural networks have boosted the convergence of multimedia data analytics in a unified framework shared by practitioners in natural language, vision and speech. Image captioning, lip reading or video sonorization are some of the first applications of a new and exciting field of research exploiting the generalization properties of deep neural representation. This tutorial will firstly review the basic neural architectures to encode and decode vision, text and audio, to later review the those models that have successfully translated information across modalities.
Xavier Giró-i-Nieto
ICMR1
2020 Enhancing Online Knowledge Graph Population with Semantic Knowledge
Delia Fernandez-Canellas, Joan Marco Rimmek, Joan Espadaler, Blai Garolera, Adrià Barja, Marc Codina, Marc Sastre, Xavier Giró-i-Nieto, Juan Carlos Riveiro, Elisenda Bou
ISWC (1)8
2020 Mask-guided sample selection for semi-supervised instance segmentation
Miriam Bellver, Amaia Salvador, Jordi Torres, Xavier Giró-i-Nieto
Multim. Tools Appl.4
2019 Simple vs complex temporal recurrences for video saliency prediction
Panagiotis Linardos, Eva Mohedano, Juan José Nieto 0002, Noel E. O'Connor, Xavier Giró-i-Nieto, Kevin McGuinness
BMVC5
2019 Inverse Cooking: Recipe Generation From Food Images
abstract
People enjoy food photography because they appreciate food. Behind each meal there is a story described in a complex recipe and, unfortunately, by simply looking at a food image we do not have access to its preparation process. Therefore, in this paper we introduce an inverse cooking system that recreates cooking recipes given food images. Our system predicts ingredients as sets by means of a novel architecture, modeling their dependencies without imposing any order, and then generates cooking instructions by attending to both image and its inferred ingredients simultaneously. We extensively evaluate the whole system on the large-scale Recipe1M dataset and show that (1) we improve performance w.r.t. previous baselines for ingredient prediction; (2) we are able to obtain high quality recipes by leveraging both image and ingredients; (3) our system is able to produce more compelling recipes than retrieval-based approaches according to human judgment. We make code and models publicly available.
Amaia Salvador, Michal Drozdzal, Xavier Giró-i-Nieto, Adriana Romero
CVPR3
2019 RVOS: End-To-End Recurrent Network for Video Object Segmentation
abstract
Multiple object video object segmentation is a challenging task, specially for the zero-shot case, when no object mask is given at the initial frame and the model has to find the objects to be segmented along the sequence. In our work, we propose a Recurrent network for multiple object Video Object Segmentation (RVOS) that is fully end-to-end trainable. Our model incorporates recurrence on two different domains: (i) the spatial, which allows to discover the different object instances within a frame, and (ii) the temporal, which allows to keep the coherence of the segmented objects along time. We train RVOS for zero-shot video object segmentation and are the first ones to report quantitative results for DAVIS-2017 and YouTube-VOS benchmarks. Further, we adapt RVOS for one-shot video object segmentation by using the masks obtained in previous time steps as inputs to be processed by the recurrent module. Our model reaches comparable results to state-of-the-art techniques in YouTube-VOS benchmark and outperforms all previous video object segmentation methods not using online learning in the DAVIS-2017 benchmark. Moreover, our model achieves faster inference runtimes than previous methods, reaching 44ms/frame on a P100 GPU.
Carles Ventura, Miriam Bellver, Andreu Girbau-Xalabarder, Amaia Salvador, Ferran Marqués, Xavier Giró-i-Nieto
CVPR6
2019 Wav2Pix: Speech-conditioned Face Generation Using Generative Adversarial Networks
abstract
Speech is a rich biometric signal that contains information about the identity, gender and emotional state of the speaker. In this work, we explore its potential to generate face images of a speaker by conditioning a Generative Adversarial Network (GAN) with raw speech input. We propose a deep neural network that is trained from scratch in an end-to-end fashion, generating a face directly from the raw speech waveform without any additional identity information (e.g reference image or one-hot encoding). Our model is trained in a self-supervised approach by exploiting the audio and visual signals naturally aligned in videos. With the purpose of training from video data, we present a novel dataset collected for this work, with high-quality videos of youtubers with notable expressiveness in both the speech and visual signals.
Amanda Cardoso Duarte, Francisco Roldan, Miquel Tubau, Janna Escur, Santiago Pascual, Amaia Salvador, Eva Mohedano, Kevin McGuinness, Jordi Torres, Xavier Giró-i-Nieto
ICASSP10
2019 VLX-Stories: Building an Online Event Knowledge Base with Emerging Entity Detection
Delia Fernandez-Canellas, Joan Espadaler, Blai Garolera, Gemma Canet, Aleix Colom, Joan Marco Rimmek, Xavier Giró-i-Nieto, Elisenda Bou, Juan Carlos Riveiro
ISWC (2)8
2019 Multiresolution co-clustering for uncalibrated multiview segmentation
Carles Ventura, David Varas, Verónica Vilaplana, Xavier Giró-i-Nieto, Ferran Marqués
Signal Process. Image Commun.4
2018 Saliency Weighted Convolutional Features for Instance Search
abstract
This work explores attention models to weight the contribution of local convolutional representations for the instance search task. We present a retrieval framework based on bags of local convolutional features (BLCF) that benefits from saliency weighting to build an efficient image representation. The use of human visual attention models (saliency) allows significant improvements in retrieval performance without the need to conduct region analysis or spatial verification, and without requiring any feature fine tuning. We investigate the impact of different saliency models, finding that higher performance on saliency benchmarks does not necessarily equate to improved performance when used in instance search tasks. The proposed approach outperforms the state-of-the-art on the challenging INSTRE benchmark by a large margin, and provides similar performance on the Oxford and Paris benchmarks compared to more complex methods that use off-the-shelf representations. Source code is publicly available at https:llgithub.com/imatge-upc/salbow.
Eva Mohedano, Kevin McGuinness, Xavier Giró-i-Nieto, Noel E. O'Connor
CBMI3
2018 Online Detection of Action Start in Untrimmed, Streaming Videos
Zheng Shou 0001, Junting Pan, Kazuyuki Miyazawa, Hassan Mansour, Anthony Vetro, Xavier Giró-i-Nieto, Shih-Fu Chang
ECCV (3)7
2018 Skip RNN: Learning to Skip State Updates in Recurrent Neural Networks
Victor Campos 0001, Brendan Jou, Xavier Giró-i-Nieto, Jordi Torres, Shih-Fu Chang
ICLR (Poster)3
2018 Demonstration of an Open Source Framework for Qualitative Evaluation of CBIR Systems
abstract
Evaluating image retrieval systems in a quantitative way, for example by computing measures like mean average precision, allows for objective comparisons with a ground-truth. However, in cases where ground-truth is not available, the only alternative is to collect feedback from a user. Thus, qualitative assessments become important to better understand how the system works. Visualizing the results could be, in some scenarios, the only way to evaluate the results obtained and also the only opportunity to identify that a system is failing. This necessitates developing a User Interface (UI) for a Content Based Image Retrieval (CBIR) system that allows visualization of results and improvement via capturing user relevance feedback. A well-designed UI facilitates understanding of the performance of the system, both in cases where it works well and perhaps more importantly those which highlight the need for improvement. Our open-source system implements three components to facilitate researchers to quickly develop these capabilities for their retrieval engine. We present: a web-based user interface to visualize retrieval results and collect user annotations; a server that simplifies connection with any underlying CBIR system; and a server that manages the search engine data. The software itself is described in a separate submission to the ACM MM Open Source Software Competition.
Paula Gómez Duran, Eva Mohedano, Kevin McGuinness, Xavier Giró-i-Nieto, Noel E. O'Connor
ACM Multimedia4
2018 Introduction to the special issue: Egocentric Vision and Lifelogging
Mariella Dimiccoli, Cathal Gurrin, David Crandall, Xavier Giró-i-Nieto, Petia Radeva
J. Vis. Commun. Image Represent.4
2018 Scanpath and saliency prediction on 360 degree images
Marc Assens, Xavier Giró-i-Nieto, Kevin McGuinness, Noel E. O'Connor
Signal Process. Image Commun.2
2017 Class Weighted Convolutional Features for Visual Instance Search
Albert Jimenez, Xavier Giró-i-Nieto
BMVC3
2017 Scaling a Convolutional Neural Network for classification of Adjective Noun Pairs with TensorFlow on GPU Clusters
abstract
Deep neural networks have gained popularity inrecent years, obtaining outstanding results in a wide range ofapplications such as computer vision in both academia andmultiple industry areas. The progress made in recent years cannotbe understood without taking into account the technologicaladvancements seen in key domains such as High PerformanceComputing, more specifically in the Graphic Processing Unit(GPU) domain. These kind of deep neural networks need massiveamounts of data to effectively train the millions of parametersthey contain, and this training can take up to days or weeksdepending on the computer hardware we are using. In thiswork, we present how the training of a deep neural networkcan be parallelized on a distributed GPU cluster. The effect ofdistributing the training process is addressed from two differentpoints of view. First, the scalability of the task and its performancein the distributed setting are analyzed. Second, the impact ofdistributed training methods on the training times and finalaccuracy of the models is studied. We used TensorFlow on top ofthe GPU cluster of servers with 2 K80 GPU cards, at BarcelonaSupercomputing Center (BSC). The results show an improvementfor both focused areas. On one hand, the experiments showpromising results in order to train a neural network faster. The training time is decreased from 106 hours to 16 hoursin our experiments. On the other hand we can observe howincreasing the numbers of GPUs in one node rises the throughput, images per second, in a near-linear way. Morever an additionaldistributed speedup of 10.3 is achieved with 16 nodes taking asbaseline the speedup of one node.
Victor Campos 0001, Francesc Sastre, Maurici Yagües, Jordi Torres, Xavier Giró-i-Nieto
CCGrid5
2017 LTA 2017: The Second Workshop on Lifelogging Tools and Applications
abstract
The organisation of personal data is receiving increasing research attention due to the challenges we face in gathering, enriching, searching, and visualising such data. Given the increasing ease with which personal data being gathered by individuals, the concept of a lifelog digital library of rich multimedia and sensory content for every individual is fast becoming a reality. The LTA 2017 workshop aims to bring together academics and practitioners to discuss approaches to lifelog data analytics and applications; and to debate the opportunities and challenges for researchers in this new and challenging area.
Cathal Gurrin, Xavier Giró-i-Nieto, Petia Radeva, Mariella Dimiccoli, Duc-Tien Dang-Nguyen, Hideo Joho
ACM Multimedia2
2017 From pixels to sentiment: Fine-tuning CNNs for visual sentiment prediction
Victor Campos 0001, Brendan Jou, Xavier Giró-i-Nieto
Image Vis. Comput.3
2016 Shallow and Deep Convolutional Networks for Saliency Prediction
abstract
The prediction of salient areas in images has been traditionally addressed with hand-crafted features based on neuroscience principles. This paper, however, addresses the problem with a completely data-driven approach by training a convolutional neural network (convnet). The learning process is formulated as a minimization of a loss function that measures the Euclidean distance of the predicted saliency map with the provided ground truth. The recent publication of large datasets of saliency prediction has provided enough data to train end-to-end architectures that are both fast and accurate. Two designs are proposed: a shallow convnet trained from scratch, and a another deeper solution whose first three layers are adapted from another network trained for classification. To the authors' knowledge, these are the first end-to-end CNNs trained and tested for the purpose of saliency prediction.
Junting Pan, Elisa Sayrol, Xavier Giró-i-Nieto, Kevin McGuinness, Noel E. O'Connor
CVPR3
2016 Bags of Local Convolutional Features for Scalable Instance Search
abstract
This work proposes a simple instance retrieval pipeline based on encoding the convolutional features of CNN using the bag of words aggregation scheme (BoW). Assigning each local array of activations in a convolutional layer to a visual word produces an assignment map, a compact representation that relates regions of an image with a visual word. We use the assignment map for fast spatial reranking, obtaining object localizations that are used for query expansion. We demonstrate the suitability of the BoW representation based on local CNN features for instance retrieval, achieving competitive performance on the Oxford and Paris buildings benchmarks. We show that our proposed system for CNN feature aggregation with BoW outperforms state-of-the-art techniques using sum pooling at a subset of the challenging TRECVid INS benchmark.
Eva Mohedano, Kevin McGuinness, Noel E. O'Connor, Amaia Salvador, Ferran Marqués, Xavier Giró-i-Nieto
ICMR6
2016 LTA 2016: The First Workshop on Lifelogging Tools and Applications
abstract
The organisation of personal data is receiving increasing research attention due to the challenges we face in gathering, enriching, searching, and visualising such data. Given the increasing ease with which personal data being gathered by individuals, the concept of a lifelog digital library of rich multimedia and sensory content for every individual is fast becoming a reality. The LTA~2016 workshop aims to bring together academics and practitioners to discuss approaches to lifelog data analytics and applications; and to debate the opportunities and challenges for researchers in this new and challenging area.
Cathal Gurrin, Xavier Giró-i-Nieto, Petia Radeva, Mariella Dimiccoli, Håvard D. Johansen, Hideo Joho, Vivek K. Singh 0001
ACM Multimedia2
2016 Assessment of crowdsourcing and gamification loss in user-assisted object segmentation
Axel Carlier, Amaia Salvador, Ferran Cabezas, Xavier Giró-i-Nieto, Vincent Charvillat, Oge Marques
Multim. Tools Appl.4
2015 Quality control in crowdsourced object segmentation
abstract
This paper explores processing techniques to deal with noisy data in crowdsourced object segmentation tasks. We use the data collected with Click'n'Cut, an online interactive segmentation tool, and we perform several experiments towards improving the segmentation results. First, we introduce different superpixel-based techniques to filter users' traces, and assess their impact on the segmentation result. Second, we present different criteria to detect and discard the traces from potential bad users, resulting in a remarkable increase in performance. Finally, we show a novel superpixel-based segmentation algorithm which does not require any prior filtering and is based on weighting each user's contribution according to his/her level of expertise.
Ferran Cabezas, Axel Carlier, Vincent Charvillat, Amaia Salvador, Xavier Giró-i-Nieto
ICIP5
2015 Improving spatial codification in semantic segmentation
abstract
This paper explores novel approaches for improving the spatial codification for the pooling of local descriptors to solve the semantic segmentation problem. We propose to partition the image into three regions for each object to be described: Figure, Border and Ground. This partition aims at minimizing the influence of the image context on the object description and vice versa by introducing an intermediate zone around the object contour. Furthermore, we also propose a richer visual descriptor of the object by applying a Spatial Pyramid over the Figure region. Two novel Spatial Pyramid configurations are explored: Cartesian-based and crown-based Spatial Pyramids. We test these approaches with state-of-the-art techniques and show that they improve the Figure-Ground based pooling in the Pascal VOC 2011 and 2012 semantic segmentation challenges.
Carles Ventura, Xavier Giró-i-Nieto, Verónica Vilaplana, Kevin McGuinness, Ferran Marqués, Noel E. O'Connor
ICIP2
2015 Exploring EEG for Object Detection and Retrieval
abstract
This paper explores the potential for using Brain Computer Interfaces (BCI) as a relevance feedback mechanism in content-based image retrieval. Several experiments are performed using a rapid serial visual presentation (RSVP) of images at different rates (5Hz and 10Hz) on 8 users with different degrees of familiarization with BCI and the dataset. We compare the feedback from the BCI and mouse-based interfaces in a subset of TRECVid images, finding that, when users have limited time to annotate the images, both interfaces are comparable in performance. Comparing our best users in a retrieval task, we found that EEG-based relevance feedback can outperform mouse-based feedback.
Eva Mohedano, Kevin McGuinness, Graham Healy, Noel E. O'Connor, Alan F. Smeaton, Amaia Salvador, Sergi Porta, Xavier Giró-i-Nieto
ICMR8
2015 Improving object segmentation by using EEG signals and rapid serial visual presentation
Eva Mohedano, Graham Healy, Kevin McGuinness, Xavier Giró-i-Nieto, Noel E. O'Connor, Alan F. Smeaton
Multim. Tools Appl.4
2014 Object Segmentation in Images using EEG Signals
abstract
This paper explores the potential of brain-computer interfaces in segmenting objects from images. Our approach is centered around designing an effective method for displaying the image parts to the users such that they generate measurable brain reactions. When an image region, specifically a block of pixels, is displayed we estimate the probability of the block containing the object of interest using a score based on EEG activity. After several such blocks are displayed, the resulting probability map is binarized and combined with the GrabCut algorithm to segment the image into object and background regions. This study shows that BCI and simple EEG analysis are useful in locating object boundaries in images.
Eva Mohedano, Graham Healy, Kevin McGuinness, Xavier Giró-i-Nieto, Noel E. O'Connor, Alan F. Smeaton
ACM Multimedia4
2014 From global image annotation to interactive object segmentation
Xavier Giró-i-Nieto, Manel Martos, Eva Mohedano, Jordi Pont-Tuset
Multim. Tools Appl.1
2014 Improving retrieval accuracy of Hierarchical Cellular Trees for generic metric spaces
Carles Ventura, Verónica Vilaplana, Xavier Giró-i-Nieto, Ferran Marqués
Multim. Tools Appl.3
2012 Hierarchical Navigation and Visual Search for Video Keyframe Retrieval
Carles Ventura, Manel Martos, Xavier Giró-i-Nieto, Verónica Vilaplana, Ferran Marqués
MMM3
2011 Diversity ranking for video retrieval from a broadcaster archive
abstract
Video retrieval through text queries is a very common practice in broadcaster archives. The query keywords are compared to the metadata labels that documentalists have previously associated to the video assets. This paper focuses on a ranking strategy to obtain more relevant keyframes among the top hits of the results ranked lists but, at the same time, keeping a diversity of video assets. Previous solutions based on a random walk over a visual similarity graph have been modified to increase the asset diversity by filtering the edges between keyframes depending on their asset. The random walk algorithm is applied separately for ever visual feature to avoid any normalization issue between visual similarity metrics. Finally, this work evaluates performance with two separate metrics: the relevance is measured by the Average Precision and the diversity is assessed by the Average Diversity, a new metric presented in this work.
Xavier Giró-i-Nieto, Monica Alfaro, Ferran Marqués
ICMR1
2010 GAT: a Graphical Annotation Tool for semantic regions
Xavier Giró-i-Nieto, Neus Camps, Ferran Marqués
Multim. Tools Appl.1
2009 Improving detection of acoustic events using audiovisual data and feature level fusion
abstract
The detection of the acoustic events (AEs) that are naturally \nproduced in a meeting room may help to describe the human \nand social activity that takes place in it. When applied to \nspontaneous recordings, the detection of AEs from only audio \ninformation shows a large amount of errors, which are mostly due to temporal overlapping of sounds. In this paper, a system to detect and recognize AEs using both audio and video information is presented. A feature-level fusion strategy is used, and the structure of the HMM-GMM based system considers each class separately and uses a one-against-all strategy for training. Experime ntal AED results with a new and rather spontaneous dataset are presented which show the advantage of the proposed approach.
Taras Butko, Cristian Canton, Carlos Segura, Xavier Giró-i-Nieto, Climent Nadeu, Javier Hernando, Josep R. Casas
INTERSPEECH4
2005 Detection of semantic objects using description graphs
abstract
This paper presents a technique to detect instances of classes (objects) according to their semantic definition in the form of a description graph. Classes are defined as combinations of instances of lower level semantic classes and allow the definition of a semantic tree that organizes classes in semantic levels. At the bottom level of the semantic tree, classes are defined by a perceptual model containing a list of low-level descriptors. The proposed detection algorithm follows a bottom-up/top-down approach, building semantic trees on a region-based representation of the media. The flexibility of the approach is assessed on different examples of planar objects, such as frontal faces, groups of islands, flags and traffic signs.
Xavier Giró-i-Nieto, Ferran Marqués
ICIP (1)1
2003 Wavelet Coding of Volumetric Medical Datasets
abstract
Several techniques based on the three-dimensional (3-D) discrete cosine transform (DCT) have been proposed for volumetric data coding. These techniques fail to provide lossless coding coupled with quality and resolution scalability, which is a significant drawback for medical applications. This paper gives an overview of several state-of-the-art 3-D wavelet coders that do meet these requirements and proposes new compression methods exploiting the quadtree and block-based coding concepts, layered zero-coding principles, and context-based arithmetic coding. Additionally, a new 3-D DCT-based coding scheme is designed and used for benchmarking. The proposed wavelet-based coding algorithms produce embedded data streams that can be decoded up to the lossless level and support the desired set of functionality constraints. Moreover, objective and subjective quality evaluation on various medical volumetric datasets shows that the proposed algorithms provide competitive lossy and lossless compression results when compared with the state-of-the-art.
Peter Schelkens, Adrian Munteanu 0001, Joeri Barbarien, Mihnea Galca, Xavier Giró-i-Nieto, Jan Cornelis 0001
IEEE Trans. Medical Imaging5