VLDB 2026 Research / reviewers in the wild / expert
Mário Antunes 0001
dblp:130/6233 · also Mário Luis Pinto Antunes
· DBLP profile ↗
21ranked-venue papers
4as first author
14since 2021 · last 2025
0000-0002-6504-9441ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Computer networks · 4 · 1 first-author · 2 since 2021Security and privacy · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Theory of computation · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring Recommendations Attacks Through Blind Optimization
Julio Corona, Rafael Teixeira, Mário Antunes 0001, Rui L. Aguiar |
DS | 3 |
| 2025 | Impact of Resource Heterogeneity on MLOps Stages: A Computational Efficiency Study
Julio Corona, Mário Antunes 0001, Rui L. Aguiar |
ICSOFT | 3 |
| 2025 | Modeling and Predicting Machine Learning PerformanceabstractModern Machine Learning (ML) models pose challenges such as long development cycles, high costs, and difficult model selection. Understanding how dataset properties influence performance in a model-agnostic and interpretable way remains limited. This work analyzes the relationship between dataset characteristics and model behavior across diverse algorithms, introducing symbolic regression models that transparently estimate accuracy from dataset features and simple model identifiers. Results show that features such as noisiness, redundancy, and class distribution consistently affect performance. A symbolic model using only dataset features achieved modest accuracy ($R^{2}=0.361$), while including model identification improved predictions ($R^{2}=0.534$). These findings provide interpretable insights into the data-model relationship, supporting preprocessing decisions (e.g., feature selection, class balancing) and informing meta-learning strategies for algorithm recommendation. Julio Corona, Rafael Teixeira, Mário Antunes 0001, Rui L. Aguiar |
ICTAI | 3 |
| 2025 | Understanding what Federated Learning Models Learn: a Comparative Study with Traditional ModelsabstractFederated Learning (FL) offers a robust framework for training Machine Learning (ML) models across distributed devices while preserving data privacy. However, concerns remain about whether FL-trained models learn in the same way as their centralized-data counterparts. This paper examines the impact of FL on model behavior, utilizing Explainable AI (XAI) techniques and correlation metrics to compare feature importance. By applying two XAI metrics across four public datasets and evaluating multiple FL strategies, we assess how closely federated models align with traditionally trained ones. Although models often achieve similar classification performance (Matthews's Correlation Coefficient (MCC) > 0.9), we observe significant differences in attribution patterns, especially in async approaches, with correlation scores below 0.7 Spearman's Rank Correlation Coefficient (SRCC) on several occasions and below 0.4 SRCC for the worst scenario. These findings suggest that identical performance does not always imply equivalent learning. Furthermore, since these findings were obtained under Independent and Identically Distributed (IID) settings, it is expected that under non-IID settings, the results might be even worse, underscoring the need for explainability-driven validation tools to ensure the reliability, fairness, and trustworthiness of FL models in practice. Rafael Teixeira, Leonardo Almeida, Julio Corona, Mário Antunes 0001, Rui L. Aguiar |
WiMob | 5 |
| 2025 | Beyond performance comparing the costs of applying Deep and Shallow LearningabstractThe rapid growth of mobile network traffic and the emergence of complex applications, such as self-driving cars and augmented reality, demand ultra-low latency, high throughput, and massive device connectivity, which traditional network design approaches struggle to meet. These issues were initially addressed in 5th Generation (5G) and Beyond-5G (B5G) networks, where Artificial Intelligence (AI), particularly Deep Learning (DL), is proposed to optimize the network and to meet these demanding requirements. However, the resource constraints and time limitations inherent in telecommunication networks raise questions about the practicality of deploying large Deep Neural Networks (DNNs) in these contexts. This paper analyzes the costs of implementing DNNs by comparing them with shallow ML models across multiple datasets and evaluating factors such as execution time and model interpretability. Our findings demonstrate that shallow ML models offer comparable performance to DNNs, with significantly reduced training and inference times, achieving up to 90% acceleration. Moreover, shallow models are more interpretable, as explainability metrics struggle to agree on feature importance values even for high-performing DNNs. Rafael Teixeira, Leonardo Almeida, Mário Antunes 0001, Diogo Gomes 0001, Rui L. Aguiar |
Comput. Commun. | 4 |
| 2025 | Leveraging decentralized communication for privacy-preserving federated learning in 6G NetworksabstractArtificial intelligence (AI) is a fundamental pillar in developing next-generation networks. Federated learning (FL) emerges as a promising solution to address data privacy concerns during AI model training within the network. However, training AI models on user equipment raises challenges regarding battery consumption, unreliable connections, and communication overhead. This paper proposes Zenoh, a data-centric communication middleware, as an alternative to the traditional Message Passing Interface (MPI) for FL applications. Zenoh’s decentralized nature and low communication overhead make it suitable for resource-constrained devices and unreliable network connections. The paper compares Zenoh and MPI in a realistic FL scenario, demonstrating Zenoh’s potential to outperform MPI in terms of flexibility, communication efficiency, and system complexity. Rafael Teixeira, Gabriele Baldoni, Mário Antunes 0001, Diogo Gomes 0001, Rui L. Aguiar |
Comput. Commun. | 3 |
| 2024 | Rethinking Security: The Resilience of Shallow ML Models (Extended Abstract)abstractThe growth of Machine Learning (ML) has led to the commercialization of applications like data analytics, autonomous systems, and security diagnostics. These models are becoming widespread across various domains. However, security and privacy issues accompany this growth. Although actively researched, there's fragmentation in analyzing and defining ML models' resilience. This work examines the resilience of shallow ML models against typical data poisoning attacks. Our study assessed their strengths in adversarial scenarios using the MNIST dataset in a CAPTCHA context. Results show notable resilience, with accuracy and generalization maintained despite malicious inputs, offering insights to strengthen future ML systems. Understanding the mechanisms enabling this resilience can aid in fortifying the security of future ML systems. Rafael Teixeira, Mário Antunes 0001, João Paulo Barraca, Diogo Gomes 0001, Rui L. Aguiar |
DSAA | 2 |
| 2024 | Optimising Data Processing in Industrial Settings: A Comparative Evaluation of Dimensionality Reduction Approaches
José Cação, Mário Antunes 0001, Miguel Monteiro |
IoTBDS | 2 |
| 2024 | Accelerating multi-tier storage cache simulations using knee detection
Tyler Estro, Mário Antunes 0001, Pranav Bhandari, Anshul Gandhi, Geoffrey H. Kuenning, Carl A. Waldspurger, Avani Wildani, Erez Zadok |
Perform. Evaluation | 2 |
| 2023 | Exploring the Intricacies of Neural Network Optimization
Rafael Teixeira, Mário Antunes 0001, Rúben Sobral, Diogo Gomes 0001, Rui L. Aguiar |
DS | 2 |
| 2023 | Towards Improved Indoor Location with Unmodified RFID Systems
Ricardo Alexandre, Mário Antunes 0001, João Paulo Barraca |
ICPRAM | 4 |
| 2023 | Guiding Simulations of Multi-Tier Storage Caches Using Knee DetectionabstractSimulating storage cache hierarchies enables efficient exploration of their configuration space, including diverse topologies, parameters and policies, and devices with varied performance characteristics, while avoiding expensive physical experiments. Miss Ratio Curves (MRCs) efficiently characterize the performance of a cache over a range of cache sizes. These useful tools reveal “key points” for cache simulation, such as knees in the curve that immediately follow sharp cliffs. Unfortunately, there are no automated techniques for efficiently finding key points in MRCs, and the cross-application of existing knee-detection algorithms yields inaccurate results. We present a multi-stage framework that identifies key points in any MRC, for both stack-based (e.g., LRU) and more sophis-ticated eviction algorithms (e.g., ARC). Our approach quickly locates candidates using efficient hash-based sampling, curve simplification, knee detection, and novel post-processing filters. We introduce Z-Method, a new multi-knee detection algorithm that employs statistical outlier detection to choose promising points robustly and efficiently. We evaluate our framework against seven other knee-detection algorithms, using both ARC and LRU MRCs from 106 diverse real-world workloads, and apply it to identify key points in multi-tier MRCs. Compared to naive approaches, our framework reduces the total number of points needed to accurately identify the best two-tier cache hierarchies by an average factor of approximately$5.5\times$for ARC and$7.7\times$for LRU. Tyler Estro, Mário Antunes 0001, Pranav Bhandari, Anshul Gandhi, Geoffrey H. Kuenning, Carl A. Waldspurger, Avani Wildani, Erez Zadok |
MASCOTS | 2 |
| 2021 | Misalignment problem in matrix decomposition with missing valuesabstractData collection within a real-world environment may be compromised by several factors such as data-logger malfunctions and communication errors, during which no data is collected. As a consequence, appropriate tools are required to handle the missing values when analysing and processing such data. This problem is often tackled via matrix decomposition. While it has been successfully applied in a wide range of applications, in this work we report an issue that has been neglected in literature and “degenerates” the quality of the imputations obtained by matrix decomposition in multivariate time-series (with smooth evolution). Briefly, the problem consists of the misalignment of the matrix decomposition result: the missing values imputations fall within an incorrect range of values and the transitions between observed and imputed values are not smooth. We address this problem by proposing a postprocessing alignment strategy. According to our experiments, the post-processing adjustment substantially improves the accuracy of the imputations (when the misalignment occurs). Moreover, the results also suggest that the misalignment occurs mostly when dealing with a small number of time-series due to lack of generalisation ability. Sofia Fernandes 0001, Mário Antunes 0001, Diogo Gomes 0001, Rui L. Aguiar |
DSAA | 2 |
| 2021 | Misalignment problem in matrix decomposition with missing values
Sofia Fernandes 0001, Mário Antunes 0001, Diogo Gomes 0001, Rui L. Aguiar |
Mach. Learn. | 2 |
| 2019 | Stream Generation: Markov Chains vs GANs
Ricardo Jesus, Mário Antunes 0001, Petia Georgieva, Diogo Gomes 0001, Rui L. Aguiar |
IoTBDS | 2 |
| 2018 | The Impact of Clustering for Learning Semantic Categories
Mário Antunes 0001, Diogo Gomes 0001, Rui L. Aguiar |
IoTBDS | 1 |
| 2018 | Towards IoT data classification through semantic features
Mário Antunes 0001, Diogo Gomes 0001, Rui L. Aguiar |
Future Gener. Comput. Syst. | 1 |
| 2017 | Extracting Knowledge from Stream Behavioural PatternsabstractThe increasing number of small, cheap devices full of sensing capabilities lead to an untapped source of information that can be explored to improve and optimize several systems. Yet, as this number grows it becomes increasingly difficult to manage and organize all this new information. The lack of a standard context representation scheme is one of the main difficulties in this research area (Antunes et al., 2016b). With this in mind we propose a stream characterization model which aims to provide the foundations of a new stream similarity metric. Complementing previous work on context organization, we aim to provide an automatic organizational model without enforcing specific representations. Ricardo Jesus, Mário Antunes 0001, Diogo Gomes 0001, Rui L. Aguiar |
IoTBDS | 2 |
| 2016 | On the application of contextual IoT service discovery in Information Centric Networks
José Quevedo, Mário Antunes 0001, Daniel Corujo, Diogo Gomes 0001, Rui L. Aguiar |
Comput. Commun. | 2 |
| 2016 | Scalable semantic aware context storage
Mário Antunes 0001, Diogo Gomes 0001, Rui L. Aguiar |
Future Gener. Comput. Syst. | 1 |
| 2014 | Context storage for M2M scenariosabstractAs the number of environmental sensors grows, it becomes increasingly difficult to manage, store and process all these sources of information. Several context representation schemes try to standardize this information, however none of them have been widely adopted. Instead of proposing yet another context representation scheme, we discuss efficient ways to deal with this diversity of representation schemes. We defined the basic requirements for flexible context storage systems, proposed an implementation and compared our implementation against two other approaches. Our solution provides more value than the remaining solutions without suffering a significant decrease in performance. Mário Antunes 0001, Diogo Gomes 0001, Rui L. Aguiar |
ICC | 1 |