José Oramas M.

dblp:47/9735 · also José Oramas, José Oramas Mogrovejo · DBLP profile ↗
← Back
36ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-8607-5067ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Enhancing hyperspectral image prediction with contrastive learning in low-label regimes
Salma Haidar, José Oramas M.
Appl. Intell.2
2026 A taxonomy of interpretation and explanation methods for capsule network architectures
Saja Tawalbeh, José Oramas M.
Neurocomputing2
2025 Bilinear MLPs enable weight-based mechanistic interpretability
abstract
A mechanistic understanding of how MLPs do computation in deep neural net- works remains elusive. Current interpretability work can extract features from hidden activations over an input dataset but generally cannot explain how MLP weights construct features. One challenge is that element-wise nonlinearities introduce higher-order interactions and make it difficult to trace computations through the MLP layer. In this paper, we analyze bilinear MLPs, a type of Gated Linear Unit (GLU) without any element-wise nonlinearity that neverthe- less achieves competitive performance. Bilinear MLPs can be fully expressed in terms of linear operations using a third-order tensor, allowing flexible analysis of the weights. Analyzing the spectra of bilinear MLP weights using eigendecom- position reveals interpretable low-rank structure across toy tasks, image classifi- cation, and language modeling. We use this understanding to craft adversarial examples, uncover overfitting, and identify small language model circuits directly from the weights alone. Our results demonstrate that bilinear layers serve as an interpretable drop-in replacement for current activation functions and that weight- based interpretability is viable for understanding deep-learning models.
Michael T. Pearce, Thomas Dooms, Alice Rigg, José Oramas M., Lee Sharkey
ICLR4
2025 Improving Neural Network Accuracy by Concurrently Training with a Twin Network
abstract
Recently within Spiking Neural Networks, a method called Twin Network Augmentation (TNA) has been introduced. This technique claims to improve the validation accuracy of a Spiking Neural Network simply by training two networks in conjunction and matching the logits via the Mean Squared Error loss. In this paper, we validate the viability of this method on a wide range of popular Convolutional Neural Network (CNN) benchmarks and compare this approach to existing Knowledge Distillation schemes. Next, we conduct a in-depth study of the different components that make up TNA and determine that its effectiveness is not solely situated in an increase of trainable parameters, but rather the effect of the training methodology. Finally, we analyse the representations learned by networks trained with TNA and highlight their superiority in a number of tasks, thus proving empirically the applicability of Twin Network Augmentation on CNN models.
Benjamin Vandersmissen, Lucas Deckers, José Oramas M.
ICLR3
2025 Smooth InfoMax - Towards Easier Post-Hoc Interpretability
Fabian Denoodt, Bart de Boer, José Oramas M.
ECML/PKDD (3)3
2025 SVEBI: Towards the Interpretation and Explanation of Spiking Neural Networks
Jasper De Laet, Hamed Behzadi-Khormouji, Lucas Deckers, José Oramas M.
ECML/PKDD (3)4
2025 Label-Efficient Learning for Radio Frequency Fingerprint Identification
abstract
Radio Frequency Fingerprint Identification (RF-FID) is a novel approach that aims to differentiate devices based on their unique signal transmissions, rather than their given identities. This approach has promising applications in wireless security, spectrum management, and sensing. Current RFFID research uses deep learning due to its success in various domains, which is largely attributed to the availability of massive labeled datasets for training. However, unlike other domains, labeled data in RFFID is limited, and the labeling process is expensive. In this paper, we propose a label-efficient learning approach for RFFID based on Contrastive Predictive Coding (CPC), a pre-training method that learns to predict future samples given the past without labels. Afterward, the model is fine-tuned to identify the device. We evaluate our approach on a fingerprint dataset of 20 devices. Our results show that CPC learns effective representations of RF signals and outperforms fully supervised learning in both classification performance and label efficiency, requiring up to 10 times fewer labels while maintaining competitive accuracy. Finally, we evaluate CPC's robustness against noise and observe competitive performance after fine-tuning.
Thayheng Nhem, Fabian Denoodt, Maarten Weyn, Michaël Peeters, José Oramas M., Rafael Berkvens
WCNC5
2025 Towards the characterization of representations learned via capsule-based network architectures
Saja Tawalbeh, José Oramas M.
Neurocomputing2
2025 Explaining and interpreting hyperdimensional computing classifiers on tabular data
abstract
Given the rise in the usage of artificial intelligence models and machine learning approaches in our day-to-day lives, it has become increasingly important to explain these models to increase user trust. Hyperdimensional Computing (HDC) has been introduced as a powerful, energy-efficient algorithmic framework that is intrinsically less opaque than (deep) neural networks. Nevertheless, the possibility of explaining and interpreting the HDC-based classification model has not yet been explored explicitly. Therefore, this work proposes an explanation method and an interpretation method for the HDC-based classification model working with tabular data. The proposed methods have been successfully evaluated on three tabular data sets with a diverse number of samples, features, and classes. Their faithfulness is validated with coherence checks, the deletion and insertion metrics, and a feature ablation study. The results of the proposed explanation method align well with the well-studied LIME explanations.
Laura Smets, Werner Van Leekwijck, Steven Latré, José Oramas M.
Neurocomputing4
2024 On the coherency of quantitative evaluation of visual explanations
Benjamin Vandersmissen, José Oramas M.
Comput. Vis. Image Underst.2
2024 SNIPPET: A Framework for Subjective Evaluation of Visual Explanations Applied to DeepFake Detection
abstract
Explainable Artificial Intelligence (XAI) attempts to help humans understand machine learning decisions better and has been identified as a critical component toward increasing the trustworthiness of complex black-box systems, such as deep neural networks. In this article, we propose a generic and comprehensive framework named SNIPPET and create a user interface for the subjective evaluation of visual explanations, focusing on finding human-friendly explanations. SNIPPET considers human-centered evaluation tasks and incorporates the collection of human annotations. These annotations can serve as valuable feedback to validate the qualitative results obtained from the subjective assessment tasks. Moreover, we consider different user background categories during the evaluation process to ensure diverse perspectives and comprehensive evaluation. We demonstrate SNIPPET on a DeepFake face dataset. Distinguishing real from fake faces is a non-trivial task even for humans that depends on rather subtle features, making it a challenging use case. Using SNIPPET, we evaluate four popular XAI methods which provide visual explanations: Gradient-weighted Class Activation Mapping, Layer-wise Relevance Propagation, attention rollout, and Transformer Attribution. Based on our experimental results, we observe preference variations among different user categories. We find that most people are more favorable to the explanations of rollout. Moreover, when it comes to XAI-assisted understanding, those who have no or lack relevant background knowledge often consider that visual explanations are insufficient to help them understand. We open-source our framework for continued data collection and annotation at https://github.com/XAI-SubjEvaluation/SNIPPET .
Boris Joukovsky, José Oramas M., Tinne Tuytelaars, Nikos Deligiannis
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Interpreting Convolutional Neural Networks by Explaining Their Predictions
abstract
We propose a method that exploits the feedback provided by visual explanation methods combined with pattern mining techniques to identify the relevant class-specific and class-shared internal units. In addition, we put forward a patch extraction approach to find faithfully class-specific and class-shared visual patterns. Contrary to the common practice in literature, our approach does not require pushing augmented visual patches through the model. Experiments on two CNN architectures show the effectiveness of the proposed method.
Toon Meynen, Hamed Behzadi-Khormouji, José Oramas M.
ICIP3
2023 A Protocol for Evaluating Model Interpretation Methods from Visual Explanations
abstract
With the continuous development of Convolutional Neural Networks (CNNs), there is an increasing requirement towards the understanding of the representations they internally encode. The task of studying such encoded representations is referred to as model interpretation. Efforts along this direction, despite being proved efficient, stand with two weaknesses. First, there is low semanticity on the feedback they provide which leads toward subjective visualizations. Second, there is no unified protocol for the quantitative evaluation of interpretation methods which makes the comparison between current and future methods complex.To address these issues, we propose a unified evaluation protocol for the quantitative evaluation of interpretation methods. This is achieved by enhancing existing interpretation methods to be capable of generating visual explanations and then linking these explanations with a semantic label. To achieve this, we introduce the Weighted Average Intersection-over-Union (WAIoU) metric to estimate the coverage rate between explanation heatmaps and semantic annotations. This is complemented with an analysis of several binarization techniques for heatmaps, necessary when measuring coverage. Experiments considering several interpretation methods covering different CNN architectures pre-trained on multiple datasets show the effectiveness of the proposed protocol.
Hamed BehzadiKhourmuji, José Oramas M.
WACV2
2022 AIMLAI: Advances in Interpretable Machine Learning and Artificial Intelligence
abstract
Recent technological advances rely on accurate decision support systems that can be perceived as black boxes due to their overwhelming complexity. This lack of transparency can lead to technical, ethical, legal, and trust issues. For example, if the control module of a self-driving car failed at detecting a pedestrian, it becomes crucial to know why the system erred. In some other cases, the decision system may reflect unacceptable biases that can generate distrust. The General Data Protection Regulation (GDPR), approved by the European Parliament in 2018, suggests that individuals should be able to obtain explanations of the decisions made from their data by automated processing, and to challenge those decisions. All these reasons have given rise to the domain of interpretable and explainable AI. AIMLAI aims at gathering researchers, experts and professionals, from inside and outside the domain of AI, interested in the topic of interpretable ML and interpretable AI. The workshop encourages interdisciplinary collaborations, with particular emphasis in knowledge management, Infovis, human computer interaction and psychology. It also welcomes applied research for use cases where interpretability matters. AIMLAI envisions to become a discussion venue for the advent of novel interpretable algorithms and explainability modules that mediate the communication between complex ML/AI systems and users.
Adrien Bibal, Tassadit Bouadi, Benoît Frénay, Luis Galárraga, José Oramas M.
CIKM5
2022 Object Detection To Enable Autonomous Vessels On European Inland Waterways
abstract
To enable autonomous vessels to operate on inland waterways, they need to detect, track and localize objects at close range to safely navigate. We deployed current deep learning techniques to detect and track these objects. As there are no large labeled datasets of European inland waterways, we used transfer learning to overcome the lack of data. By using preexisting similar datasets, we were able to significantly decrease the required amount of labeled data from the target distribution. Furthermore, we improved the mean Average Precision from 0.461 to 0.814 by using a limited number of labeled target data samples. We estimated the relative distance of the objects based on the generated bounding boxes. The information from the camera is then combined with LiDar data to generate a top-view map of the environment which is used as input for an object-avoidance control agent. All these methods can run in real-time on the vessel with an fps of 1.83 on a 2.7GHz vCPU.
Mattias Billast, Robin Janssens, Astrid Vanneste, Simon Vanneste, Olivier Vasseur, Ali Anwar 0002, Kevin Mets, Tom De Schepper, José Oramas M., Steven Latré, Peter Hellinckx
IECON9
2021 MinMaxCAM: Improving object coverage for CAM-based Weakly Supervised Object Localization
José Oramas M., Tinne Tuytelaars
BMVC2
2020 Multiple Exemplars-Based Hallucination for Face Super-Resolution and Editing
José Oramas M., Tinne Tuytelaars
ACCV (5)2
2020 In Defense of LSTMs for Addressing Multiple Instance Learning Problems
José Oramas M., Tinne Tuytelaars
ACCV (6)2
2020 AIMLAI'20: Third Workshop on Advances in Interpretable Machine Learning and Artificial Intelligence
abstract
The Third Workshop on "Advances in Interpretable Machine Learning and Artificial Intelligence" (AIMLAI) presents contributions in the fields of (i) interpretable ML and AI, i.e., algorithms that are natively interpretable, and (ii) interpretability modules, i.e., explanation layers on top of black-box models, also called post-hoc interpretability. AIMLAI encourages interdisciplinary collaborations with particular emphasis in knowledge management, infovis, human computer interaction and psychology. It also welcomes applied research for use cases where interpretability matters.
Adrien Bibal, Tassadit Bouadi, Benoît Frénay, Luis Galárraga, José Oramas M.
CIKM5
2020 Unpaired Image-To-Image Shape Translation Across Fashion Data
abstract
We address the problem of unpaired geometric image-to-image translation. Rather than transferring the style of an image as a whole, our goal is to translate the geometry of an object while preserving its appearance. Our model is trained without the need for paired images. It performs all steps of the shape transfer within a single model and without additional post-processing stages. Experiments on clothing-based datasets show the effectiveness of the proposed method.
Liqian Ma, José Oramas M., Luc Van Gool, Tinne Tuytelaars
ICIP3
2019 Towards Object Shape Translation Through Unsupervised Generative Deep Models
abstract
This paper focuses on the problem of unsupervised image-to-image translation. More specifically, we aim at finding a translation network such that objects and shapes that only appear in the source domain are translated to objects and shapes only appearing in the target domain, while style color features present in the source domain remain the same. To achieve this, we use a domain-specific variational autoencoder and represent each image in its latent space representation. In a second step, we learn a translation between latent spaces of different domains using generative adversarial networks. We evaluate this framework on multiple datasets and verify the effect of multiple perceptual losses. Experiments on the MNIST and SVHN datasets show the effectiveness of the proposed translation method.
Lies Bollens, Tinne Tuytelaars, José Oramas M.
ICIP3
2019 Visual Explanation by Interpretation: Improving Visual Feedback Capabilities of Deep Neural Networks
José Oramas M., Tinne Tuytelaars
ICLR (Poster)1
2018 From Pixels to Actions: Learning to Drive a Car with Deep Neural Networks
abstract
The promise of self-driving cars promotes several advantages, e.g. they have the ability to outperform human drivers while being safer. Here we take a deeper look into some aspects from algorithms aimed at making this promise a reality. More specifically, we analyze an end-to-end neural network to predict a car's steering actions on a highway based on images taken from a single car-mounted camera. We focus our analysis on several aspects which could have a significant impact on the performance of the system. These aspects are: the input data format, the temporal dependencies between consecutive inputs, and the origin of the data. We show that, for the task at hand, regression networks outperform their classifier counterparts. In addition, there seems to be a small difference between networks that use coloured images and ones that use grayscale images as input. For the second aspect, by feeding the network three concatenated images, we get a significant decrease of 30% in mean squared error. For the third aspect, by using simulation data we are able to train networks that have a performance comparable to networks trained on real-life datasets. We also qualitatively demonstrate that the standard metrics that are used to evaluate networks do not necessarily accurately reflect a system's driving behaviour. We show that a promising confusion matrix may result in poor driving behaviour while a very ill-looking confusion matrix may result in good driving behaviour.
Jonas Heylen, Seppe Iven, Bert De Brabandere, José Oramas M., Luc Van Gool, Tinne Tuytelaars
WACV4
2018 An Analysis of Human-Centered Geolocation
abstract
Online social networks contain a constantly increasing amount of images - most of them focusing on people. Due to cultural and climate factors, fashion trends and physical appearance of individuals differ from city to city. In this paper we investigate to what extent such cues can be exploited in order to infer the geographic location, i.e. the city, where a picture was taken. We conduct a user study, as well as an evaluation of automatic methods based on convolutional neural networks. Experiments on the Fashion 144k and a Pinterest-based dataset show that the automatic methods succeed at this task to a reasonable extent. As a matter of fact, our empirical results suggest that automatic methods can surpass human performance by a large margin. Further inspection of the trained models shows that human-centered characteristics, like clothing style, physical features, and accessories, are informative for the task at hand. Moreover, it reveals that also contextual features, e.g. wall type, natural environment, etc., are taken into account by the automatic methods.
Yu-Hui Huang, José Oramas M., Luc Van Gool, Tinne Tuytelaars
WACV3
2017 Context-based object viewpoint estimation: A 2D relational approach
José Oramas M., Luc De Raedt, Tinne Tuytelaars
Comput. Vis. Image Underst.1
2017 Rank Pooling for Action Recognition
abstract
We propose a function-based temporal pooling method that captures the latent structure of the video sequence data - e.g., how frame-level features evolve over time in a video. We show how the parameters of a function that has been fit to the video data can serve as a robust new video representation. As a specific example, we learn a pooling function via ranking machines. By learning to rank the frame-level features of a video in chronological order, we obtain a new representation that captures the video-wide temporal dynamics of a video, suitable for action recognition. Other than ranking functions, we explore different parametric models that could also explain the temporal changes in videos. The proposed functional pooling methods, and rank pooling in particular, is easy to interpret and implement, fast to compute and effective in recognizing a wide variety of actions. We evaluate our method on various benchmarks for generic action, fine-grained action and gesture recognition. Results show that rank pooling brings an absolute improvement of 7-10 average pooling baseline. At the same time, rank pooling is compatible with and complementary to several appearance and local motion based methods and features, such as improved trajectories and deep learning features.
Basura Fernando, Efstratios Gavves, José Oramas M., Amir Ghodrati, Tinne Tuytelaars
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Recovering hard-to-find object instances by sampling context-based object proposals
José Oramas M., Tinne Tuytelaars
Comput. Vis. Image Underst.1
2015 Modeling video evolution for action recognition
abstract
In this paper we present a method to capture video-wide temporal information for action recognition. We postulate that a function capable of ordering the frames of a video temporally (based on the appearance) captures well the evolution of the appearance within the video. We learn such ranking functions per video via a ranking machine and use the parameters of these as a new video representation. The proposed method is easy to interpret and implement, fast to compute and effective in recognizing a wide variety of actions. We perform a large number of evaluations on datasets for generic action recognition (Hollywood2 and HMDB51), fine-grained actions (MPII- cooking activities) and gestures (Chalearn). Results show that the proposed method brings an absolute improvement of 7–10%, while being compatible with and complementary to further improvements in appearance and local motion based methods.
Basura Fernando, Efstratios Gavves, José Oramas M., Amir Ghodrati, Tinne Tuytelaars
CVPR3
2015 Towards sign language recognition based on body parts relations
abstract
Over the years, hand gesture recognition has been mostly addressed considering hand trajectories in isolation. However, in most sign languages, hand gestures are defined on a particular context (body region). We propose a pipeline which models hand movements in the context of other parts of the body captured in the 3D space using the Kinect sensor. In addition, we perform sign recognition based on the different hand postures that occur during a sign. Our experiments show that considering different body parts brings improved performance when compared with methods which only consider global hand trajectories. Finally, we demonstrate that the combination of hand postures features with hand gestures features helps to improve the prediction of a given sign.
Marc Martínez-Camarena, José Oramas M., Tinne Tuytelaars
ICIP2
2014 Scene-driven Cues for Viewpoint Classification for Elongated Object Classes
José Oramas M., Tinne Tuytelaars
BMVC1
2014 Towards cautious collective inference for object verification
abstract
It is by now generally accepted that reasoning about the relationships between objects (and object hypotheses) can improve the accuracy of object detection methods. Relations between objects allow to reject inconsistent hypotheses and reduce the uncertainty of the initial hypotheses. However, most methods to date reason about object relations in a relatively crude way. In this paper we propose an alternative using cautious inference. Building on ideas from Collective Classification, we favor the most confident hypotheses as sources of contextual information and give higher relevance to the object relations observed during training. Additionally, we propose to cluster the pairwise relations into relationships. Our experiments on part of the KITTI data benchmark and the MIT StreetScenes dataset show that both steps improve the performance of relational classifiers.
José Oramas M., Luc De Raedt, Tinne Tuytelaars
WACV1
2014 There are plenty of places like home: Using relational representations in hierarchies for distance-based image understanding
Laura Antanas, Martijn van Otterlo, José Oramas M., Tinne Tuytelaars, Luc De Raedt
Neurocomputing3
2013 Allocentric Pose Estimation
abstract
The task of object pose estimation has been a challenge since the early days of computer vision. To estimate the pose (or viewpoint) of an object, people have mostly looked at object intrinsic features, such as shape or appearance. Surprisingly, informative features provided by other, external elements in the scene, have so far mostly been ignored. At the same time, contextual cues have been shown to be of great benefit for related tasks such as object detection or action recognition. In this paper, we explore how information from other objects in the scene can be exploited for pose estimation. In particular, we look at object configurations. We show that, starting from noisy object detections and pose estimates, exploiting the estimated pose and location of other objects in the scene can help to estimate the objects' poses more accurately. We explore both a camera-centered as well as an object-centered representation for relations. Experiments on the challenging KITTI dataset show that object configurations can indeed be used as a complementary cue to appearance-based pose estimation. In addition, object-centered relational representations can also assist object detection.
José Oramas M., Luc De Raedt, Tinne Tuytelaars
ICCV1
2013 Rule-based Hand Posture Recognition using Qualitative Finger Configurations Acquired with the Kinect
Lieven Billiet, José Oramas M., McElory Hoffmann, Wannes Meert, Laura Antanas
ICPRAM2
2012 A Relational Distance-based Framework for Hierarchical Image Understanding
Laura Antanas, Martijn van Otterlo, José Oramas M., Tinne Tuytelaars, Luc De Raedt
ICPRAM (2)3
2010 Not Far Away from Home: A Relational Distance-Based Approach to Understanding Images of Houses
Laura Antanas, Martijn van Otterlo, José Oramas M., Tinne Tuytelaars, Luc De Raedt
ILP3