EDBT 2026 Demo / reviewers in the wild / expert
Gemma Roig
dblp:58/9606
· DBLP profile ↗
42ranked-venue papers
3as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Predictive Coding inspired convolutional networks can capture the neural dynamics of recurrent processing in human image recognitionabstractInspired by the robustness of human vision, various attempts have been made to incorporate brain-inspired mechanisms into artificial neural networks.A popular candidate has been predictive coding, a prominent theory in neuroscience, that theorizes that feedback connections communicate top-down predictions to earlier regions.While recurrence in models has been demonstrated to be useful when processing noisy and di cult stimuli, a direct evidence of its utility for explaining brain data under such situation was yet to be shown.Here, we investigated whether such brain-inspired mechanism actually helps to capture neural dynamics.Specifically, we measured the brain alignment between representations of a predictive version of a popular feedforward CNN often used as a computational model of the visual cortex-VGG16-and human EEG collected when viewing images that were relatively easy (Control) or di cult to classify (Challenge).We demonstrate that the recurrent dynamics significantly enhanced the model's alignment with EEG responses, underscoring the importance of recurrent connectivity in computational models of human vision, an e↵ect distinctly visible for challenging stimuli. Manshan Guo, Bhavin Choksi, Sari Sadiya, Pablo Oyarzo, Radoslaw Martin Cichy, Gemma Roig |
ESANN | 6 |
| 2026 | Weakly Supervised Shortcut Learning Mitigation Using Sparse AutoencodersabstractReliance on spurious features that coincidentally correlate with task labels (i.e., shortcut learning) remains a major barrier to the reliable deployment of machine learning models, particularly in high-stakes domains like medical diagnostics.Moreover, in such settings retraining models or collecting and labeling additional data is often impractical, limiting the applicability of many existing shortcut mitigation methods.In this paper we propose a lightweight framework that leverages sparse autoencoders to disentangle spurious from core features to mitigate shortcut learning.Our approach requires no model retraining and works even when group annotations are scarce or unavailable for certain classes.Results on standard benchmarks demonstrate that, even with as few as 50 labeled examples, reliance on spurious features can be significantly reduced. Sari Sadiya, Despina Tawadros, Phuong Quynh Le, Jörg Schlötterer, Christin Seifert, Gemma Roig |
ESANN | 7 |
| 2025 | Analysing the impact of brain-inspired predictive coding dynamics through gradient based explainability methodsabstractMultiple theories exist for the role of feedback connections in the brain and in the artificial neural networks, but remain untested using modern tools.In this work, we undertake this task by exploring the utility of explainability methods like GradCAMs [1] in investigating bio-inspired recurrent networks-provided with the predify[2] package-that perform hierarchical updates inspired by the predictive coding theory in neuroscience.We report an extensive search with different levels of feedforward and feedback information.Our preliminary results show that the dynamics are able to recover the GradCAMs on noisy images, providing promising avenues for future work aiming to understand the role of recurrence. Bhavin Choksi, Gionata Paolo Zalaffi, Giovanna Maria Dimitri, Gemma Roig |
ESANN | 4 |
| 2025 | Predictive Coding Dynamics Enhance Model-Brain SimilarityabstractPredictive coding-a popular theory in neuroscience-has garnered significant attention in the machine learning community aiming to incorporate brain-inspired components in neural networks.While various proposals have demonstrated the ability of predictive dynamics to render robustness and entail human-like perception of illusions, it remains unclear if they improve the alignment between brain and artificial representations.Here, we systematically investigate the conditions under which brain-inspired modifications in predictive processing improve alignment between model and neural representations in the brain.Our results reveal that the feedback component significantly increases similarity between model representations and those found in higher-level visual brain areas, especially when processing complex visual scenes. Manshan Guo, Michael Samjatin, Bhavin Choksi, Sari Sadiya, Radoslaw Martin Cichy, Gemma Roig |
ESANN | 6 |
| 2025 | UIDAPLE: Unsupervised Incremental Domain Adaptation through Adaptive Prompt LearningabstractContinual learning poses significant challenges for deep neural networks, notably catastrophic forgetting, particularly when faced with shifting data distributions that compromise previously acquired knowledge. This paper tackles these issues within the Unsupervised Incremental Domain Adaptation (UIDA) framework, where the initial source domain is labeled, but subsequent domains are not. Existing methods often struggle with limited cross-domain generalization and adaptation capabilities. As a remedy, we introduce UIDAPLE, a novel approach that utilizes a unified prompt across all domains, leveraging the foundation model CLIP to obviate the need for isolated domain treatments. Specifically, UIDAPLE implements supervised prompt learning in the labeled source domain and extends this learning to unlabeled domains through confidence-based adaptation. We also present an efficient parameter alignment strategy that maintains semantic coherence across domains, effectively balancing stability and plasticity to combat catastrophic forgetting. Extensive evaluations on two benchmark datasets reveal that UIDAPLE markedly surpasses other UIDA techniques in performance. Samrat Mukherjee, Tanuj Sur, Saurish Seksaria, Subhasis Chaudhuri, Gemma Roig, Biplab Banerjee |
ICASSP | 5 |
| 2025 | Efficient Unsupervised Shortcut Learning Detection and Mitigation in TransformersabstractShortcut learning, i.e., a model's reliance on undesired features not directly relevant to the task, is a major challenge that severely limits the applications of machine learning algorithms, particularly when deploying them to assist in making sensitive decisions, such as in medical diagnostics. In this work, we leverage recent advancements in machine learning to create an unsupervised framework that is capable of both detecting and mitigating shortcut learning in transformers. We validate our method on multiple datasets. Results demonstrate that our framework significantly improves both worst-group accuracy (samples misclassified due to shortcuts) and average accuracy, while minimizing human annotation effort. Moreover, we demonstrate that the detected shortcuts are meaningful and informative to human experts, and that our framework is computationally efficient, allowing it to be run on consumer hardware. Lukas Kuhn, Sari Sadiya, Jörg Schlötterer, Florian Buettner 0001, Christin Seifert, Gemma Roig |
ICCV | 6 |
| 2025 | BrainACTIV: Identifying visuo-semantic properties driving cortical selectivity using diffusion-based image manipulationabstractThe human brain efficiently represents visual inputs through specialized neural populations that selectively respond to specific categories. Advancements in generative modeling have enabled data-driven discovery of neural selectivity using brain-optimized image synthesis. However, current methods independently generate one sample at a time, without enforcing structural constraints on the generations; thus, these individual images have no explicit point of comparison, making it hard to discern which image features drive neural response selectivity. To address this issue, we introduce Brain Activation Control Through Image Variation (BrainACTIV), a method for manipulating a reference image to enhance or decrease activity in a target cortical region using pretrained diffusion models. Starting from a reference image allows for fine-grained and reliable offline identification of optimal visuo-semantic properties, as well as producing controlled stimuli for novel neuroimaging studies. We show that our manipulations effectively modulate predicted fMRI responses and agree with hypothesized preferred categories in established regions of interest, while remaining structurally close to the reference image. Moreover, we demonstrate how our method accentuates differences between brain regions that are selective to the same category, and how it could be used to explore neural representation of brain regions with unknown selectivities. Hence, BrainACTIV holds the potential to formulate robust hypotheses about brain representation and to facilitate the production of naturalistic stimuli for neuroscientific experiments. Diego Garcia Cerdas, Christina Sartzetaki, Magnus Petersen, Gemma Roig, Pascal Mettes, Iris I. A. Groen |
ICLR | 4 |
| 2025 | One Hundred Neural Networks and Brains Watching Videos: Lessons from AlignmentabstractWhat can we learn from comparing video models to human brains, arguably the most efficient and effective video processing systems in existence? Our work takes a step towards answering this question by performing the first large-scale benchmarking of deep video models on representational alignment to the human brain, using publicly available models and a recently released video brain imaging (fMRI) dataset. We disentangle four factors of variation in the models (temporal modeling, classification task, architecture, and training dataset) that affect alignment to the brain, which we measure by conducting Representational Similarity Analysis across multiple brain regions and model layers. We show that temporal modeling is key for alignment to brain regions involved in early visual processing, while a relevant classification task is key for alignment to higher-level regions. Moreover, we identify clear differences between the brain scoring patterns across layers of CNNs and Transformers, and reveal how training dataset biases transfer to alignment with functionally selective brain areas. Additionally, we uncover a negative correlation of computational complexity to brain alignment. Measuring a total of 99 neural networks and 10 human brains watching videos, we aim to forge a path that widens our understanding of temporal and semantic video representations in brains and machines, ideally leading towards more efficient video models and more mechanistic explanations of processing in the human brain. Christina Sartzetaki, Gemma Roig, Cees Snoek, Iris I. A. Groen |
ICLR | 2 |
| 2025 | On Explaining Knowledge Distillation: Measuring and Visualising the Knowledge Transfer ProcessabstractKnowledge distillation (KD) remains challenging due to the opaque nature of the knowledge transfer process from a Teacher to a Student, making it difficult to address certain issues related to KD. To address this, we proposed UniCAM, a novel gradient-based visual explanation method, which effectively interprets the knowledge learned during KD. Our experimental results demonstrate that with the guidance of the Teacher's knowledge, the Student model becomes more efficient, learning more relevant features while discarding those that are not relevant. We refer to the features learned with the Teacher's guidance as distilled features and the features irrelevant to the task and ignored by the Student as residual features. Distilled features focus on key aspects of the input, such as textures and parts of objects. In contrast, residual features demonstrate more diffused attention, often targeting irrelevant areas, including the backgrounds of the target objects. In addition, we proposed two novel metrics: the feature similarity score (FSS) and the rele-vance score (RS), which quantify the relevance of the dis-tilled knowledge. Experiments on the CIFAR10, ASIRRA, and Plant Disease datasets demonstrate that UniCAM and the two metrics offer valuable insights to explain the KD process. Gereziher Adhane, Mohammad Mahdi Dehshibi, Dennis Vetter, David Masip, Gemma Roig |
WACV | 5 |
| 2025 | FovEx: Human-Inspired Explanations for Vision Transformers and Convolutional Neural NetworksabstractAbstract Explainability in artificial intelligence (XAI) remains a crucial aspect for fostering trust and understanding in machine learning models. Current visual explanation techniques, such as gradient-based or class-activation-based methods, often exhibit a strong dependence on specific model architectures. Conversely, perturbation-based methods, despite being model-agnostic, are computationally expensive as they require evaluating models on a large number of forward passes. We introduce Foveation-based Explanations (FovEx), a novel XAI method inspired by human vision, which combines biologically inspired foveation-based transformations with gradient-driven overt attention to iteratively select locations of interest. These locations are selected to maximize the performance of the model to be explained with respect to the downstream task and then combined to generate an attribution map. We provide a thorough evaluation with qualitative and quantitative assessments on established benchmarks. Our method achieves state-of-the-art performance on both transformers (on 4 out of 5 metrics) and convolutional models (on 3 out of 5 metrics), demonstrating its versatility among various architectures. Furthermore, we show the alignment between the explanation map produced by FovEx and human gaze patterns (+14% in NSS compared to RISE, +203% in NSS compared to GradCAM). This comparison enhances our confidence in FovEx’s ability to close the interpretation gap between humans and machines. Mahadev Prasad Panda, Matteo Tiezzi, Martina G. Vilas, Gemma Roig, Björn M. Eskofier, Dario Zanca |
Int. J. Comput. Vis. | 4 |
| 2025 | Correction: FovEx: Human-Inspired Explanations for Vision Transformers and Convolutional Neural NetworksabstractCorrection: International Journal of Computer Vision (2025) 133:7437–7459 Mahadev Prasad Panda, Matteo Tiezzi, Martina G. Vilas, Gemma Roig, Björn M. Eskofier, Dario Zanca |
Int. J. Comput. Vis. | 4 |
| 2024 | Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive NeuroscienceabstractInner Interpretability is a promising emerging field tasked with uncovering the inner mechanisms of AI systems, though how to develop these mechanistic theories is still much debated. Moreover, recent critiques raise issues that question its usefulness to advance the broader goals of AI. However, it has been overlooked that these issues resemble those that have been grappled with in another field: Cognitive Neuroscience. Here we draw the relevant connections and highlight lessons that can be transferred productively between fields. Based on these, we propose a general conceptual framework and give concrete methodological strategies for building mechanistic explanations in AI inner interpretability research. With this conceptual framework, Inner Interpretability can fend off critiques and position itself on a productive path to explain AI systems. Martina G. Vilas, Federico Adolfi, David Poeppel, Gemma Roig |
ICML | 4 |
| 2024 | Learning Class and Domain Augmentations for Single-Source Open-Domain GeneralizationabstractSingle-source open-domain generalization (SS-ODG) addresses the challenge of labeled source domains with supervision during training and unlabeled novel target domains during testing. The target domain includes both known classes from the source domain and samples from previously unseen classes. Existing techniques for SS-ODG primarily focus on calibrating source-domain classifiers to identify open samples in the target domain. However, these methods struggle with visually fine-grained open-closed data, often misclassifying open samples as closed-set classes. Moreover, relying solely on a single source domain restricts the model’s ability to generalize. To overcome these limitations, we propose a novel framework called SODG-Net that simultaneously synthesizes novel domains and generates pseudo-open samples using a learning-based objective, in contrast to the ad-hoc mixing strategies commonly found in the literature. Our approach enhances generalization by diversifying the styles of known class samples using a novel metric criterion and generates diverse pseudo-open samples to train a unified and confident multiclass classifier capable of handling both open and closed-set data. Extensive experimental evaluations conducted on multiple benchmarks consistently demonstrate the superior performance of SODG-Net compared to the literature. Prathmesh Bele, Valay Bundele, Avigyan Bhattacharya, Ankit Jha, Gemma Roig, Biplab Banerjee |
WACV | 5 |
| 2023 | Different Algorithms (Might) Uncover Different Patterns: A Brain-Age Prediction Case StudyabstractMachine learning is a rapidly evolving field with a wide range of applications, including biological signal analysis, where novel algorithms often improve the state-of-the-art. However, robustness to algorithmic variability - measured by different algorithms, consistently uncovering similar findings - is seldom explored. In this paper we investigate whether established hypotheses in brain-age prediction from EEG research validate across algorithms. First, we surveyed literature and identified various features known to be informative for brain-age prediction. We employed diverse feature extraction techniques, processing steps, and models, and utilized the interpretative power of SHapley Additive exPlanations (SHAP) values to align our findings with the existing research in the field. Few of our models achieved state-of-the-art performance on the specific data-set we utilized. Moreover, analysis demonstrated that while most models do uncover similar patterns in the EEG signals, some variability could still be observed. Finally, a few prominent findings could only be validated using specific models. We conclude by suggesting remedies to the potential implications of this lack of robustness to model variability. Tobias Ettling, Sari Saba-Sadiya, Gemma Roig |
BIBM | 3 |
| 2023 | Analyzing Vision Transformers for Image Classification in Class Embedding SpaceabstractDespite the growing use of transformer models in computer vision, a mechanistic understanding of these networks is still needed. This work introduces a method to reverse-engineer Vision Transformers trained to solve image classification tasks. Inspired by previous research in NLP, we demonstrate how the inner representations at any level of the hierarchy can be projected onto the learned class embedding space to uncover how these networks build categorical representations for their predictions. We use our framework to show how image tokens develop class-specific representations that depend on attention mechanisms and contextual information, and give insights on how self-attention and MLP layers differentially contribute to this categorical composition. We additionally demonstrate that this method (1) can be used to determine the parts of an image that would be important for detecting the class of interest, and (2) exhibits significant advantages over traditional linear probing approaches. Taken together, our results position our proposed framework as a powerful tool for mechanistic interpretability and explainability research. Martina G. Vilas, Timothy Schaumlöffel, Gemma Roig |
NeurIPS | 3 |
| 2022 | Tafsir Dataset: A Novel Multi-Task Benchmark for Named Entity Recognition and Topic Modeling in Classical Arabic LiteratureabstractVarious historical languages, which used to be lingua franca of science and arts, deserve the attention of current NLP research. In this work, we take the first data-driven steps towards this research line for Classical Arabic (CA) by addressing named entity recognition (NER) and topic modeling (TM) on the example of CA literature. We manually annotate the encyclopedic work of Tafsir Al-Tabari with span-based NEs, sentence-based topics, and span-based subtopics, thus creating the Tafsir Dataset with over 51,000 sentences, the first large-scale multi-task benchmark for CA. Next, we analyze our newly generated dataset, which we make open-source available, with current language models (lightweight BiLSTM, transformer-based MaChAmP) along a novel script compression method, thereby achieving state-of-the-art performance for our target task CA-NER. We also show that CA-TM from the perspective of historical topic models, which are central to Arabic studies, is very challenging. With this interdisciplinary work, we lay the foundations for future research on automatic analysis of CA literature. Sajawel Ahmed, Rob van der Goot, Misbahur Rehman, Carl Kruse, Ömer Özsoy, Alexander Mehler, Gemma Roig |
COLING | 7 |
| 2022 | What do navigation agents learn about their environment?abstractToday's state of the art visual navigation agents typically consist of large deep learning models trained end to end. Such models offer little to no interpretability about the learned skills or the actions of the agent taken in response to its environment. While past works have explored interpreting deep learning models, little attention has been devoted to interpreting embodied AI systems, which often involve reasoning about the structure of the environment, target characteristics and the outcome of one's actions. In this paper, we introduce the Interpretability System for Embodied agEnts (iSEE) for Point Goal and Object Goal navigation agents. We use iSEEto probe the dynamic representations produced by these agents for the presence of information about the agent as well as the environment. We demonstrate interesting insights about navigation agents using iSEE, including the ability to encode reachable locations (to avoid obstacles), visibility of the target, progress from the initial spawn location as well as the dramatic effect on the behaviors of agents when we mask out critical individual neurons. Kshitij Dwivedi, Gemma Roig, Aniruddha Kembhavi, Roozbeh Mottaghi |
CVPR | 2 |
| 2022 | FRIDA - Generative feature replay for incremental domain adaptation
Sayan Rakshit, Anwesh Mohanty, Ruchika Chavhan, Biplab Banerjee, Gemma Roig, Subhasis Chaudhuri |
Comput. Vis. Image Underst. | 5 |
| 2021 | Unveiling functions of the visual cortex using task-specific deep neural networksabstractThe human visual cortex enables visual perception through a cascade of hierarchical computations in cortical regions with distinct functionalities. Here, we introduce an AI-driven approach to discover the functional mapping of the visual cortex. We related human brain responses to scene images measured with functional MRI (fMRI) systematically to a diverse set of deep neural networks (DNNs) optimized to perform different scene perception tasks. We found a structured mapping between DNN tasks and brain regions along the ventral and dorsal visual streams. Low-level visual tasks mapped onto early brain regions, 3-dimensional scene perception tasks mapped onto the dorsal stream, and semantic tasks mapped onto the ventral stream. This mapping was of high fidelity, with more than 60% of the explainable variance in nine key regions being explained. Together, our results provide a novel functional mapping of the human visual cortex and demonstrate the power of the computational approach. Kshitij Dwivedi, Michael F. Bonner, Radoslaw Martin Cichy, Gemma Roig |
PLoS Comput. Biol. | 4 |
| 2020 | LCD: Learned Cross-Domain Descriptors for 2D-3D MatchingabstractIn this work, we present a novel method to learn a local cross-domain descriptor for 2D image and 3D point cloud matching. Our proposed method is a dual auto-encoder neural network that maps 2D and 3D input into a shared latent space representation. We show that such local cross-domain descriptors in the shared embedding are more discriminative than those obtained from individual training in 2D and 3D domains. To facilitate the training process, we built a new dataset by collecting ≈ 1.4 millions of 2D-3D correspondences with various lighting conditions and settings from publicly available RGB-D scenes. Our descriptor is evaluated in three main experiments: 2D-3D matching, cross-domain retrieval, and sparse-to-dense depth estimation. Experimental results confirm the robustness of our approach as well as its competitive performance not only in solving cross-domain tasks but also in being able to generalize to solve sole 2D and 3D tasks. Our dataset and code are released publicly at https://hkust-vgd.github.io/lcd. Quang-Hieu Pham, Mikaela Angelina Uy, Binh-Son Hua, Duc Thanh Nguyen, Gemma Roig, Sai-Kit Yeung |
AAAI | 5 |
| 2020 | Duality Diagram Similarity: A Generic Framework for Initialization Selection in Task Transfer Learning
Kshitij Dwivedi, Radoslaw Martin Cichy, Gemma Roig |
ECCV (26) | 4 |
| 2020 | Multi-source Open-Set Deep Adversarial Domain Adaptation
Sayan Rakshit, Dipesh Tamboli, Pragati Shuddhodhan Meshram, Biplab Banerjee, Gemma Roig, Subhasis Chaudhuri |
ECCV (26) | 5 |
| 2020 | Predictive Coding Networks Meet Action RecognitionabstractAction recognition is a key problem in computer vision that labels videos with a set of predefined actions. Most of the state-of-the-art methods rely on RGB frames for extracting the semantics and pre-computed optical flow fields as a motion cue. Then, both are combined using deep neural networks. Yet, it has been argued that such models are not able to leverage the motion information extracted from the optical flow, but instead the optical flow allows for better recognition of people and objects in the video. This urges the need to explore different cues or models that can extract motion in a more informative fashion. To tackle this issue, we propose to explore the predictive coding network, so called PredNet, a recurrent neural network that propagates predictive coding errors across layers and time steps. We analyze whether PredNet can better capture motions in videos by estimating over time the representations extracted from pre-trained networks for action recognition. In this way, the model only relies on the video frames, and does not need pre-processed optical flows as input. We report the effectiveness of our proposed model on UCF101 and HMDB51 datasets. Hossein Mousavi, Gemma Roig |
ICIP | 3 |
| 2020 | AttendAffectNet: Self-Attention based Networks for Predicting Affective Responses from MoviesabstractIn this work, we propose different variants of the self-attention based network for emotion prediction from movies, which we call AttendAffectNet. We take both audio and video into account and incorporate the relation among multiple modalities by applying self-attention mechanism in a novel manner into the extracted features for emotion prediction. We compare it to the typically temporal integration of the self-attention based model, which in our case, allows to capture the relation of temporal representations of the movie while considering the sequential dependencies of emotion responses. We demonstrate the effectiveness of our proposed architectures on the extended COGNIMUSE dataset [1], [2] and the MediaEval 2016 Emotional Impact of Movies Task [3], which consist of movies with emotion annotations. Our results show that applying the self-attention mechanism on the different audio-visual features, rather than in the time domain, is more effective for emotion prediction. Our approach is also proven to outperform many state-of-the-art models for emotion prediction. The code to reproduce our results with the models' implementation is available at: https://github.com/ivyha010/AttendAffectNet. Ha Thi Phuong Thao, Balamurali B. T., Dorien Herremans, Gemma Roig |
ICPR | 4 |
| 2020 | Regression-based Music Emotion Prediction using Triplet Neural NetworksabstractIn this paper, we adapt triplet neural networks (TNNs) to a regression task, music emotion prediction. Since TNNs were initially introduced for classification, and not for regression, we propose a mechanism that allows them to provide meaningful low dimensional representations for regression tasks. We then use these new representations as the input for regression algorithms such as support vector machines and gradient boosting machines. To demonstrate the TNNs' effectiveness at creating meaningful representations, we compare them to different dimensionality reduction methods on music emotion prediction, i.e., predicting valence and arousal values from musical audio signals. Our results on the DEAM dataset show that by using TNNs we achieve 90% feature dimensionality reduction with a 9% improvement in valence prediction and 4% improvement in arousal prediction with respect to our baseline models (without TNN). Our TNN method outperforms other dimensionality reduction methods such as principal component analysis (PCA) and autoencoders (AE). This shows that, in addition to providing a compact latent space representation of audio features, the proposed approach achieves higher performance than the baseline models. Kin Wai Cheuk, Yin-Jyun Luo, Balamurali B. T., Gemma Roig, Dorien Herremans |
IJCNN | 4 |
| 2019 | Latent Space Representation for Multi-Target Speaker Detection and Identification with a Sparse Dataset Using Triplet Neural NetworksabstractWe present an approach to tackle the speaker recognition problem using Triplet Neural Networks. Currently, the i-vector representation with probabilistic linear discriminant analysis (PLDA) is the most commonly used technique to solve this problem, due to high classification accuracy with a relatively short computation time. In this paper, we explore a neural network approach, namely Triplet Neural Networks (TNNs), to built a latent space for different classifiers to solve the Multi-Target Speaker Detection and Identification Challenge Evaluation 2018 (MCE 2018) dataset. This training set contains i-vectors from 3,631 speakers, with only 3 samples for each speaker, thus making speaker recognition a challenging task. When using the train and development set for training both the TNN and baseline model (i.e., similarity evaluation directly on the i-vector representation), our proposed model outperforms the baseline by 23%. When reducing the training data to only using the train set, our method results in 309 confusions for the Multi-target speaker identification task, which is 46% better than the baseline model. These results show that the representational power of TNNs is especially evident when training on small datasets with few instances available per class. Kin Wai Cheuk, Balamurali B. T., Gemma Roig, Dorien Herremans |
ASRU | 3 |
| 2019 | Representation Similarity Analysis for Efficient Task Taxonomy & Transfer LearningabstractTransfer learning is widely used in deep neural network models when there are few labeled examples available. The common approach is to take a pre-trained network in a similar task and finetune the model parameters. This is usually done blindly without a pre-selection from a set of pre-trained models, or by finetuning a set of models trained on different tasks and selecting the best performing one by cross-validation. We address this problem by proposing an approach to assess the relationship between visual tasks and their task-specific models. Our method uses Representation Similarity Analysis (RSA), which is commonly used to find a correlation between neuronal responses from brain data and models. With RSA we obtain a similarity score among tasks by computing correlations between models trained on different tasks. Our method is efficient as it requires only pre-trained models, and a few images with no further training. We demonstrate the effectiveness and efficiency of our method to generating task taxonomy on Taskonomy dataset. We next evaluate the relationship of RSA with the transfer learning performance on Taskonomy tasks and a new task: Pascal VOC semantic segmentation. Our results reveal that models trained on tasks with higher similarity score show higher transfer learning performance. Surprisingly, the best transfer learning result for Pascal VOC semantic segmentation is not obtained from the pre-trained model on semantic segmentation, probably due to the domain differences, and our method successfully selects the high performing models. Kshitij Dwivedi, Gemma Roig |
CVPR | 2 |
| 2019 | JSIS3D: Joint Semantic-Instance Segmentation of 3D Point Clouds With Multi-Task Pointwise Networks and Multi-Value Conditional Random FieldsabstractDeep learning techniques have become the to-go models for most vision-related tasks on 2D images. However, their power has not been fully realised on several tasks in 3D space, e.g., 3D scene understanding. In this work, we jointly address the problems of semantic and instance segmentation of 3D point clouds. Specifically, we develop a multi-task pointwise network that simultaneously performs two tasks: predicting the semantic classes of 3D points and embedding the points into high-dimensional vectors so that points of the same object instance are represented by similar embeddings. We then propose a multi-value conditional random field model to incorporate the semantic and instance labels and formulate the problem of semantic and instance segmentation as jointly optimising labels in the field model. The proposed method is thoroughly evaluated and compared with existing methods on different indoor scene datasets including S3DIS and SceneNN. Experimental results showed the robustness of the proposed joint semantic-instance segmentation scheme over its single components. Our method also achieved state-of-the-art performance on semantic segmentation. Quang-Hieu Pham, Duc Thanh Nguyen, Binh-Son Hua, Gemma Roig, Sai-Kit Yeung |
CVPR | 4 |
| 2018 | DOPING: Generative Data Augmentation for Unsupervised Anomaly Detection with GANabstractRecently, the introduction of the generative adversarial network (GAN) and its variants has enabled the generation of realistic synthetic samples, which has been used for enlarging training sets. Previous work primarily focused on data augmentation for semi-supervised and supervised tasks. In this paper, we instead focus on unsupervised anomaly detection and propose a novel generative data augmentation framework optimized for this task. By using a GAN variant known as the adversarial autoencoder (AAE), we impose a distribution on the latent space of the dataset and systematically sample the latent space to generate artificial samples. To the best of our knowledge, our method is the first data augmentation technique focused on improving performance in unsupervised anomaly detection. We validate our method by demonstrating consistent improvements across several real-world datasets. Swee Kiat Lim, Yi Loo, Ngoc-Trung Tran, Ngai-Man Cheung, Gemma Roig, Yuval Elovici |
ICDM | 5 |
| 2017 | Do Deep Neural Networks Suffer from Crowding?abstractCrowding is a visual effect suffered by humans, in which an object that can be recognized in isolation can no longer be recognized when other objects, called flankers, are placed close to it. In this work, we study the effect of crowding in artificial Deep Neural Networks (DNNs) for object recognition. We analyze both deep convolutional neural networks (DCNNs) as well as an extension of DCNNs that are multi-scale and that change the receptive field size of the convolution filters with their position in the image. The latter networks, that we call eccentricity-dependent, have been proposed for modeling the feedforward path of the primate visual cortex. Our results reveal that the eccentricity-dependent model, trained on target objects in isolation, can recognize such targets in the presence of flankers, if the targets are near the center of the image, whereas DCNNs cannot. Also, for all tested networks, when trained on targets in isolation, we find that recognition accuracy of the networks decreases the closer the flankers are to the target and the more flankers there are. We find that visual similarity between the target and flankers also plays a role and that pooling in early layers of the network leads to more crowding. Additionally, we show that incorporating flankers into the images of the training set for learning the DNNs does not lead to robustness against configurations not seen at training. Anna Volokitin, Gemma Roig, Tomaso A. Poggio |
NIPS | 2 |
| 2016 | Learning to Predict Sequences of Human Visual FixationsabstractMost state-of-the-art visual attention models estimate the probability distribution of fixating the eyes in a location of the image, the so-called saliency maps. Yet, these models do not predict the temporal sequence of eye fixations, which may be valuable for better predicting the human eye fixations, as well as for understanding the role of the different cues during visual exploration. In this paper, we present a method for predicting the sequence of human eye fixations, which is learned from the recorded human eye-tracking data. We use least-squares policy iteration (LSPI) to learn a visual exploration policy that mimics the recorded eye-fixation examples. The model uses a different set of parameters for the different stages of visual exploration that capture the importance of the cues during the scanpath. In a series of experiments, we demonstrate the effectiveness of using LSPI for combining multiple cues at different stages of the scanpath. The learned parameters suggest that the low-level and high-level cues (semantics) are similarly important at the first eye fixation of the scanpath, and the contribution of high-level cues keeps increasing during the visual exploration. Results show that our approach obtains the state-of-the-art performances on two challenging data sets: 1) OSIE data set and 2) MIT data set. Ming Jiang 0019, Xavier Boix, Gemma Roig, Luc Van Gool, Qi Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Saliency Prediction with Active Semantic SegmentationabstractMing Jiang1 [email protected] Xavier Boix1,3 [email protected] Juan Xu1 [email protected] Gemma Roig2,3 [email protected] Luc Van Gool3 [email protected] Qi Zhao1 [email protected] 1 Department of Electrical and Computer Engineering National University of Singapore Singapore 2 CBMM, LCSL Massachusetts Institute of Technology Istituto Italiano di Tecnologia Cambridge, MA 3 Computer Vision Laboratory ETH Zurich Switzerland Ming Jiang 0019, Xavier Boix, Gemma Roig, Luc Van Gool, Qi Zhao 0001 |
BMVC | 4 |
| 2015 | SEEDS: Superpixels Extracted Via Energy-Driven Sampling
Michael Van den Bergh, Xavier Boix, Gemma Roig, Luc Van Gool |
Int. J. Comput. Vis. | 3 |
| 2014 | Self-Adaptable Templates for Feature Coding
Xavier Boix, Gemma Roig, Salomon Diether, Luc Van Gool |
NIPS | 2 |
| 2013 | Sparse Quantization for Patch DescriptionabstractThe representation of local image patches is crucial for the good performance and efficiency of many vision tasks. Patch descriptors have been designed to generalize towards diverse variations, depending on the application, as well as the desired compromise between accuracy and efficiency. We present a novel formulation of patch description, that serves such issues well. Sparse quantization lies at its heart. This allows for efficient encodings, leading to powerful, novel binary descriptors, yet also to the generalization of existing descriptors like SIFT or BRIEF. We demonstrate the capabilities of our formulation for both key point matching and image classification. Our binary descriptors achieve state-of-the-art results for two key point matching benchmarks, namely those by Brown and Mikolajczyk. For image classification, we propose new descriptors, that perform similar to SIFT on Caltech101 and PASCAL VOC07. Xavier Boix, Michael Gygli, Gemma Roig, Luc Van Gool |
CVPR | 3 |
| 2013 | Online Video SEEDS for Temporal Window ObjectnessabstractSuper pixel and objectness algorithms are broadly used as a pre-processing step to generate support regions and to speed-up further computations. Recently, many algorithms have been extended to video in order to exploit the temporal consistency between frames. However, most methods are computationally too expensive for real-time applications. We introduce an online, real-time video super pixel algorithm based on the recently proposed SEEDS super pixels. A new capability is incorporated which delivers multiple diverse samples (hypotheses) of super pixels in the same image or video sequence. The multiple samples are shown to provide a strong cue to efficiently measure the objectness of image windows, and we introduce the novel concept of objectness in temporal windows. Experiments show that the video super pixels achieve comparable performance to state-of-the-art offline methods while running at 30 fps on a single 2.8 GHz i7 CPU. State-of-the-art performance on objectness is also demonstrated, yet orders of magnitude faster and extended to temporal windows in video. Michael Van den Bergh, Gemma Roig, Xavier Boix, Santiago Manen, Luc Van Gool |
ICCV | 2 |
| 2013 | Active MAP Inference in CRFs for Efficient Semantic SegmentationabstractMost MAP inference algorithms for CRFs optimize an energy function knowing all the potentials. In this paper, we focus on CRFs where the computational cost of instantiating the potentials is orders of magnitude higher than MAP inference. This is often the case in semantic image segmentation, where most potentials are instantiated by slow classifiers fed with costly features. We introduce Active MAP inference 1) to on-the-fly select a subset of potentials to be instantiated in the energy function, leaving the rest of the parameters of the potentials unknown, and 2) to estimate the MAP labeling from such incomplete energy function. Results for semantic segmentation benchmarks, namely PASCAL VOC 2010 and MSRC-21, show that Active MAP inference achieves similar levels of accuracy but with major efficiency gains. Gemma Roig, Xavier Boix, Roderick de Nijs, Sebastian Ramos, Kolja Kühnlenz, Luc Van Gool |
ICCV | 1 |
| 2012 | SEEDS: Superpixels Extracted via Energy-Driven Sampling
Michael Van den Bergh, Xavier Boix, Gemma Roig, Benjamin de Capitani, Luc Van Gool |
ECCV (7) | 3 |
| 2012 | Nested Sparse Quantization for Efficient Feature Coding
Xavier Boix, Gemma Roig, Christian Leistner, Luc Van Gool |
ECCV (2) | 2 |
| 2012 | On-line semantic perception using uncertaintyabstractVisual perception capabilities are still highly unreliable in unconstrained settings, and solutions might not be accurate in all regions of an image. Awareness of the uncertainty of perception is a fundamental requirement for proper high level decision making in a robotic system. Yet, the uncertainty measure is often sacrificed to account for dependencies between object/region classifiers. This is the case of Conditional Random Fields (CRFs), the success of which stems from their ability to infer the most likely world configuration, but they do not directly allow to estimate the uncertainty of the solution. In this paper, we consider the setting of assigning semantic labels to the pixels of an image sequence. Instead of using a CRF, we employ a Perturb-and-MAP Random Field, a recently introduced probabilistic model that allows performing fast approximate sampling from its probability density function. This allows to effectively compute the uncertainty of the solution, indicating the reliability of the most likely labeling in each region of the image. We report results on the CamVid dataset, a standard benchmark for semantic labeling of urban image sequences. In our experiments, we show the benefits of exploiting the uncertainty by putting more computational effort on the regions of the image that are less reliable, and use more efficient techniques for other regions, showing little decrease of performance. Roderick de Nijs, Sebastian Ramos, Gemma Roig, Xavier Boix, Luc Van Gool, Kolja Kühnlenz |
IROS | 3 |
| 2011 | Hierarchical CRF with product label spaces for parts-based modelsabstractNon-rigid object detection is a challenging open research problem in computer vision. It is a critical part in many applications such as image search, surveillance, human-computer interaction or image auto-annotation. Most successful approaches to non-rigid object detection make use of part-based models. In particular, Conditional Random Fields (CRF) have been successfully embedded into a discriminative parts-based model framework due to its effectiveness for learning and inference (usually based on a tree structure). However, CRF-based approaches do not incorporate global constraints and only model pairwise interactions. This is especially important when modeling object classes that may have complex parts interactions (e.g. facial features or body articulations), because neglecting them yields an oversimplified model with suboptimal performance. To overcome this limitation, this paper proposes a novel hierarchical CRF (HCRF). The main contribution is to build a hierarchy of part combinations by extending the label set to a hierarchy of product label spaces. In order to keep the inference computation tractable, we propose an effective method to reduce the new label set. We test our method on two applications: facial feature detection on the Multi-PIE database and human pose estimation on the Buffy dataset. Gemma Roig, Xavier Boix, Fernando De la Torre, Joan Serrat 0002, Carles Vilella |
FG | 1 |
| 2011 | Conditional Random Fields for multi-camera object detectionabstractWe formulate a model for multi-class object detection in a multi-camera environment. From our knowledge, this is the first time that this problem is addressed taken into account different object classes simultaneously. Given several images of the scene taken from different angles, our system estimates the ground plane location of the objects from the output of several object detectors applied at each viewpoint. We cast the problem as an energy minimization modeled with a Conditional Random Field (CRF). Instead of predicting the presence of an object at each image location independently, we simultaneously predict the labeling of the entire scene. Our CRF is able to take into account occlusions between objects and contextual constraints among them. We propose an effective iterative strategy that renders tractable the underlying optimization problem, and learn the parameters of the model with the max-margin paradigm. We evaluate the performance of our model on several challenging multi-camera pedestrian detection datasets namely PETS 2009 [5] and EPFL terrace sequence [9]. We also introduce a new dataset in which multiple classes of objects appear simultaneously in the scene. It is here where we show that our method effectively handles occlusions in the multi-class case. Gemma Roig, Xavier Boix, Horesh Ben Shitrit, Pascal Fua |
ICCV | 1 |