Gérôme Bovet

dblp:132/0962 · DBLP profile ↗
← Back
53ranked-venue papers
2as first author
49since 2021 · last 2026
0000-0002-4534-3483ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 21 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Security and privacy · 7 · 7 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Asynchronous Cache-based Aggregation with Fairness and Filtering for Decentralized Federated Learning
abstract
Decentralized Federated Learning (DFL) offers a scalable paradigm for collaborative intelligence at the edge, yet its practical efficacy is severely constrained by system heterogeneity. Traditional synchronous protocols enforce rigid, lockstep aggregation barriers, where the training velocity of the entire collective is strictly dictated by the slowest straggler node, inevitably leading to significant idle time and resource underutilization. While asynchronous strategies mitigate latency, they often introduce complex pathologies, such as unbounded staleness and systemic unfairness, because high-performance nodes disproportionately bias the global model toward local data distributions, thereby marginalizing slower contributors. To rigorously reconcile these conflicting trade-offs, this work presents CAFF , a novel asynchronous communication framework for DFL that decouples local optimization from global synchronization via a topology-aware, event-driven protocol. By implementing a topology-aware cache with a strict per-neighbor replacement policy, the mechanism limits per-peer dominance by enforcing a one-slot-per-neighbor cache and exclusive replacement, preventing any peer from contributing multiple updates within a single aggregation event. Furthermore, a configurable staleness filter and a dynamic aggregation threshold ensure robust convergence stability across diverse federation topologies. Extensive empirical evaluations using MNIST, FashionMNIST, CIFAR-10, and SVHN, conducted on a high-fidelity, virtualized testbed across fully connected, star, and ring topologies, demonstrate that CAFF significantly outperforms synchronous baselines. Specifically, in dense network configurations, the framework reduces wall-clock training time by up to 39% and network traffic by up to 75%, while maintaining competitive predictive fidelity with controlled accuracy degradation. These results position CAFF as a robust and scalable efficiency-oriented solution for heterogeneous peer-to-peer learning environments.
Enrique Tomás Martínez Beltrán, Eduard Gash, Gérôme Bovet, Alberto Huertas Celdrán, Burkhard Stiller
Comput. Networks3
2026 FedEnD: Communication-efficient Federated Learning for non-IID data via decentralized ensemble distillation
abstract
Federated Learning (FL) offers a paradigm for collaborative AI that mitigates raw data exposure, yet the statistical heterogeneity of client data severely constrains its practical application. This non-independent and identically distributed (non-IID) setting induces client drift, leading to unstable optimization and degraded generalization, particularly for under-represented classes. Existing solutions present a difficult trade-off: iterative, regularization-based methods suffer from high communication overhead and a centralized bottleneck, while knowledge-distillation-based approaches rely on impractical artifacts, such as shared public datasets. This work introduces FedEnD , a novel framework that addresses the previous challenge through an efficient, fully decentralized architecture. FedEnD employs a two-stage protocol that decouples local specialist training from a collaborative fusion stage. Following a communication-free training phase, clients perform a one-shot peer-to-peer broadcast that shares (optionally privatized) specialist parameters and lightweight class-count statistics. Each client then uses these statistics to construct a class-distribution-weighted teacher ensemble from the received specialists. Crucially, this ensemble’s knowledge is distilled into a robust global model on each client, using only their local data as unlabeled inputs, obviating the need for a central server or auxiliary data. Extensive experiments on MNIST, FashionMNIST, SVHN, and CIFAR-10 demonstrate that FedEnD outperforms baselines, surpassing robust methods such as SCAFFOLD by +5.7% on complex datasets in pathologically skewed settings. This superior accuracy is achieved while reducing communication bandwidth by 68.6% compared to standard iterative averaging, and by up to 84% compared to gradient-correction methods like SCAFFOLD, highlighting a favorable trade-off between accuracy and communication bandwidth in decentralized learning under severe non-IID partitions.
Enrique Tomás Martínez Beltrán, Philip Giryes, Gérôme Bovet, Burkhard Stiller, Gregorio Martínez Pérez, Alberto Huertas Celdrán
Future Gener. Comput. Syst.3
2026 GreenDFL: A framework for assessing the sustainability of Decentralized Federated Learning systems
abstract
Context: Decentralized Federated Learning (DFL) is an emerging paradigm that enables collaborative model training without centralized data and model aggregation, enhancing privacy and resilience. However, its sustainability remains underexplored, as energy consumption and carbon emissions vary across different system configurations. Understanding the environmental impact of DFL is crucial for optimizing its design and deployment. Objective: This work aims to develop a comprehensive and operational framework for assessing the sustainability of DFL systems. To address it, this work provides a systematic method for quantifying energy consumption and carbon emissions, offering insights into improving the sustainability of DFL. Methods: This work proposes GreenDFL , a fully implementable framework that has been integrated into a real-world DFL platform. GreenDFL systematically analyzes the impact of various factors, including hardware accelerators, model architecture, communication medium, data distribution, network topology, and federation size, on the sustainability of DFL systems. Besides, a sustainability-aware aggregation algorithm ( GreenDFL-SA ) and a node selection algorithm ( GreenDFL-SN ) are developed to optimize energy efficiency and reduce carbon emissions in DFL training. Results: Empirical experiments are conducted on multiple datasets, measuring energy consumption and carbon emissions at different phases of the DFL lifecycle. Results indicate that local training dominates energy consumption and carbon emissions, while communication has a relatively minor impact. Optimizing model complexity, using GPUs instead of CPUs, and strategically selecting participating nodes significantly improve sustainability. Additionally, using wired communication, particularly optical fiber, effectively reduces energy consumption during the communication phase, while integrating early stopping mechanisms further minimizes overall emissions. Conclusion: The proposed GreenDFL provides a comprehensive and practical approach for assessing the sustainability of DFL systems. Furthermore, it offers best practices for improving environmental efficiency in DFL, making sustainability considerations more actionable in real-world deployments.
Chao Feng 0001, Alberto Huertas Celdrán, Gérôme Bovet, Burkhard Stiller
Inf. Softw. Technol.4
2025 ColNet: Collaborative Optimization in Decentralized Federated Multi-Task Learning Systems
Chao Feng 0001, Nicolas Fazli Kohler, Weijie Niu, Alberto Huertas Celdrán, Gérôme Bovet, Burkhard Stiller
IEEE Big Data6
2025 Assessing the Sustainability and Trustworthiness of Federated Learning Models
abstract
Artificial intelligence (AI) increasingly influences critical decision-making across sectors. Federated Learning (FL), as a privacy-preserving collaborative AI paradigm, not only enhances data protection but also holds significant promise for intelligent network management, including distributed monitoring, adaptive control, and edge intelligence. Although the trustworthiness of FL systems has received growing attention, the sustainability dimension remains insufficiently explored, despite its importance for scalable real-world deployment. To address this gap, this work introduces sustainability as a distinct pillar within a comprehensive trustworthy FL taxonomy, consistent with AIHLEG guidelines. This pillar includes three key aspects: hardware efficiency, federation complexity, and the carbon intensity of energy sources. Experiments using the FederatedScope framework under diverse scenarios, including varying participants, system complexity, hardware, and energy configurations, validate the practicality of the approach. Results show that incorporating sustainability into FL evaluation supports environmentally responsible deployment, enabling more efficient, adaptive, and trustworthy network services and management AI models.
Chao Feng 0001, Alberto Huertas Celdrán, Pedro Miguel Sánchez Sánchez, Lynn Zumtaugwald, Gérôme Bovet, Burkhard Stiller
CNSM5
2025 From Models to Network Topologies: A Topology Inference Attack in Decentralized Federated Learning
abstract
Federated Learning (FL) is widely recognized as a privacy-preserving Machine Learning paradigm due to its model-sharing mechanism that avoids direct data exchange. Nevertheless, model training leaves exploitable traces that can be used to infer sensitive information. In Decentralized FL (DFL), the topology, defining how participants are connected, plays a crucial role in shaping the model’s privacy, robustness, and convergence. However, the topology introduces an unexplored vulnerability: attackers can exploit it to infer participant relationships and launch targeted attacks. This work uncovers the hidden risks of DFL topologies by proposing a novel Topology Inference Attack that infers the topology solely from model behavior. A taxonomy of topology inference attacks is introduced, categorizing them by the attacker’s capabilities and knowledge. Practical attack strategies are designed for various scenarios, and experiments are conducted to identify key factors influencing attack success. The results demonstrate that analyzing only the model of each node can accurately infer the DFL topology, highlighting a critical privacy risk in DFL systems. These findings offer insights for improving privacy preservation in DFL environments.
Chao Feng 0001, Yuanzhe Gao, Alberto Huertas Celdrán, Gérôme Bovet, Burkhard Stiller
ECAI4
2025 ProFe: Communication-Efficient Decentralized Federated Learning via Distillation and Prototypes
abstract
Decentralized Federated Learning (DFL) trains models in a collaborative and privacy-preserving manner while removing model centralization risks and improving communication bottlenecks. However, DFL faces challenges in efficient communication management and model aggregation within decentralized environments, especially with heterogeneous data distributions. Thus, this paper introduces ProFe, a novel communication optimization algorithm for DFL that combines knowledge distillation, prototype learning, and quantization techniques. ProFe utilizes knowledge from large local models to train smaller ones for aggregation, incorporates prototypes to better learn unseen classes, and applies quantization to reduce data transmitted during communication rounds. The performance of ProFe has been validated and compared to the literature by using benchmark datasets like MNIST, CIFAR10, and CIFAR100. Results showed that the proposed algorithm reduces communication costs by up to$\approx 40-50 \%$while maintaining or improving model performance. In addition, it adds$\approx 20 \%$training time due to increased complexity, generating a trade-off.
Pedro Miguel Sánchez Sánchez, Enrique Tomás Martínez Beltrán, Miguel Fernández Llamas, Gérôme Bovet, Gregorio Martínez Pérez, Alberto Huertas Celdrán
ICC4
2025 S-VOTE: Similarity-based Voting for Client Selection in Decentralized Federated Learning
abstract
Decentralized Federated Learning (DFL) enables collaborative, privacy-preserving model training without relying on a central server. This decentralized approach reduces bottlenecks and eliminates single points of failure, enhancing scalability and resilience. However, DFL also introduces challenges, such as suboptimal models with non-IID data distributions, increased communication overhead, and resource usage. Thus, this work proposes S-VOTE, a voting-based client selection mechanism that optimizes resource usage and enhances model performance in federations with non-IID data conditions. S-VOTE considers an adaptive strategy for spontaneous local training that addresses participation imbalance, allowing underutilized clients to contribute without significantly increasing resource costs. Extensive experiments on benchmark datasets demonstrate the S-VOTE effectiveness. More in detail, it achieves lower communication costs by up to 21%, 4-6% faster convergence, and improves local performance by 9-17% compared to baseline methods in some configurations, all while achieving a 14-24% energy consumption reduction. These results highlight the potential of S-VOTE to address DFL challenges in heterogeneous environments.
Pedro Miguel Sánchez Sánchez, Enrique Tomás Martínez Beltrán, Chao Feng 0001, Gérôme Bovet, Gregorio Martínez Pérez, Alberto Huertas Celdrán
IJCNN4
2025 Demo: A Practical Testbed for Decentralized Federated Learning on Physical Edge Devices
abstract
Federated Learning (FL) enables collaborative model training without sharing raw data, preserving participant privacy. Decentralized FL (DFL) eliminates reliance on a central server, mitigating the single point of failure inherent in the traditional FL paradigm, while introducing deployment challenges on resource-constrained devices. To evaluate real-world applicability, this work designs and deploys a physical testbed using edge devices such as Raspberry Pi and Jetson Nano. The testbed is built upon a DFL training platform, NEBULA, and extends it with a power monitoring module to measure energy consumption during training. Experiments across multiple datasets show that model performance is influenced by the communication topology, with denser topologies leading to better outcomes in DFL settings.
Chao Feng 0001, Nicolas Huber, Alberto Huertas Celdrán, Gérôme Bovet, Burkhard Stiller
LCN4
2025 De-VertiFL: A Solution for Decentralized Vertical Federated Learning
abstract
Federated Learning (FL), introduced in 2016, was designed to enhance data privacy in collaborative model training environments. Among the FL paradigm, horizontal FL, where clients share the same set of features but different data samples, has been extensively studied in both centralized and decentralized settings. In contrast, Vertical Federated Learning (VFL), which is crucial in real-world decentralized scenarios where clients possess different, yet sensitive, data about the same entity, remains underexplored. Thus, this work introduces De-VertiFL, a novel solution for training models in a decentralized VFL setting. De-VertiFL contributes by introducing a new network architecture distribution, an innovative knowledge exchange scheme, and a distributed federated training process. Specifically, De-VertiFL enables the sharing of hidden layer outputs among federation clients, allowing participants to benefit from intermediate computations, thereby improving learning efficiency. De-VertiFL has been evaluated using a variety of well-known datasets, including both image and tabular data, across binary and multiclass classification tasks. The results demonstrate that De-VertiFL generally surpasses state-of-the-art methods in F1-score performance, while maintaining a decentralized and privacy-preserving framework.
Alberto Huertas Celdrán, Chao Feng 0001, Sabyasachi Banik, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
NOMS4
2025 GuardFS: A file system for integrated detection and mitigation of Linux-based ransomware
abstract
Although ransomware has received broad attention in media and research, this evolving threat vector still poses a systematic threat. Related literature has explored their detection using various approaches leveraging Machine and Deep Learning. While these approaches are effective in detecting malware, they do not answer how to use this intelligence to protect against threats, raising concerns about their applicability in a hostile environment. Solutions that focus on mitigation rarely explore how to prevent and not just alert or halt its execution, especially when considering Linux-based samples. This paper presents GuardFS , a file system-based approach to investigate the integration of detection and mitigation of ransomware. Using a bespoke overlay file system, data is extracted before files are accessed. Models trained on this data are used by three novel defense configurations that obfuscate, delay, or track access to the file system. The experiments on GuardFS test the configurations in a reactive setting. The results demonstrate that although data loss cannot be completely prevented, it can be significantly reduced. Usability and performance analysis demonstrate that the defense effectiveness of the configurations relates to their impact on resource consumption and usability.
Jan von der Assen, Chao Feng 0001, Alberto Huertas Celdrán, Róbert Oles, Gérôme Bovet, Burkhard Stiller
J. Inf. Secur. Appl.5
2025 CyberForce: A Federated Reinforcement Learning Framework for Malware Mitigation
abstract
Recent research has shown that the integration of Reinforcement Learning (RL) with Moving Target Defense (MTD) can enhance cybersecurity in Internet-of-Things (IoT) devices. Nevertheless, the practicality of existing work is hindered by data privacy concerns associated with centralized data processing in RL, and the unsatisfactory time needed to learn right MTD techniques that are effective against a rising number of heterogeneous zero-day attacks. Thus, this work presents CyberForce, a framework that combines Federated and Reinforcement Learning (FRL) to collaboratively and privately learn suitable MTD techniques for mitigating zero-day attacks. CyberForce integrates device fingerprinting and anomaly detection to reward or penalize MTD mechanisms chosen by an FRL-based agent. The framework has been deployed and evaluated in a scenario consisting of ten physical devices of a real IoT platform affected by heterogeneous malware samples. A pool of experiments has demonstrated that CyberForce learns the MTD technique mitigating each attack faster than existing RL-based centralized approaches. In addition, when various devices are exposed to different attacks, CyberForce benefits from knowledge transfer, leading to enhanced performance and reduced learning time in comparison to recent works. Finally, different aggregation algorithms used during the agent learning process provide CyberForce with notable robustness to malicious attacks.
Chao Feng 0001, Alberto Huertas Celdrán, Pedro Miguel Sánchez Sánchez, Jan Kreischer, Jan von der Assen, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
IEEE Trans. Dependable Secur. Comput.6
2024 FEVER: Intelligent Behavioral Fingerprinting for Anomaly Detection in P4-Based Programmable Networks
Matheus Saueressig, Muriel Figueredo Franco, Eder J. Scheid, Alberto Huertas Celdrán, Gérôme Bovet, Burkhard Stiller, Lisandro Z. Granville
AINA (3)5
2024 Leveraging MTD to Mitigate Poisoning Attacks in Decentralized FL with Non-IID Data
abstract
Decentralized Federated Learning (DFL), a paradigm for managing big data in a privacy-preserving and distributed manner, is vulnerable to poisoning attacks where malicious clients tamper with data or models. Current defense methods often assume Independently and Identically Distributed (IID) data across participants, which is unrealistic in real-world applications. In more realistic non-IID contexts, existing defensive strategies face challenges when distinguishing between models that have been compromised and those that have been trained on heterogeneous data distributions (non-IID), leading to diminished efficacy. In response, this paper proposes a framework that employs the Moving Target Defense (MTD) approach to bolster the robustness of DFL models. By continuously modifying the attack surface of the DFL system, the framework aims to mitigate poisoning attacks effectively. The proposed solution includes both proactive and reactive modes, utilizing a reputation system that combines metrics of model similarity and loss, alongside various defensive techniques. Comprehensive experimental evaluations indicate that the MTD-based mechanism significantly mitigates a range of poisoning attack types across multiple datasets with different federation topologies.
Chao Feng 0001, Alberto Huertas Celdrán, Zien Zeng, Jan von der Assen, Gérôme Bovet, Burkhard Stiller
IEEE Big Data6
2024 ThreatFinderAI: Automated Threat Modeling Applied to LLM System Integration
abstract
Artificial Intelligence (AI) is a rapidly integrated technology, significantly contributing to advancements like 6G. However, its swift adoption raises considerable security concerns. Large Language Models (LLMs) pose risks such as spear phishing, code injections, and remote code execution. Conventional threat modeling, used in secure software development, faces challenges when applied to AI systems, as existing methodologies are designed for traditional software. Furthermore, AI-specific threat modeling research is sparse and lacks approaches providing practical support or automation. Thus, this demo paper presents ThreatFinderAI, an asset-centric threat modeling and risk assessment framework. ThreatFinderAI fulfills seven steps aligned with AI system design and transforms AI threat and control knowledge bases into a queryable knowledge graph for automated asset identification and threat elicitation. It also proposes business impact analysis and expert estimates for AI threat impact quantification. In the demonstration, ThreatFinderAI is illustrated by securing a customer care application relying on LLMs. Through this, it is demonstrated how the proposed framework can be used to identify relevant threats and practical countermeasures and communicate strategic risk.
Jan von der Assen, Alberto Huertas Celdrán, Jamo Sharif, Chao Feng 0001, Gérôme Bovet, Burkhard Stiller
CNSM5
2024 Sentinel: An Aggregation Function to Secure Decentralized Federated Learning
abstract
Decentralized Federated Learning (DFL) emerges as an innovative paradigm to train collaborative models, addressing the single point of failure limitation. However, the security and trustworthiness of FL and DFL are compromised by poisoning attacks, negatively impacting its performance. Existing defense mechanisms have been designed for centralized FL and they do not adequately exploit the particularities of DFL. Thus, this work introduces Sentinel, a defense strategy to counteract poisoning attacks in DFL. Sentinel leverages the accessibility of local data and defines a three-step aggregation protocol consisting of similarity filtering, bootstrap validation, and normalization to safeguard against malicious model updates. Sentinel has been evaluated with diverse datasets and data distributions. Besides, various poisoning attack types and threat levels have been verified. The results improve the state-of-the-art performance against both untargeted and targeted poisoning attacks when data follows an IID (Independent and Identically Distributed) configuration. Besides, under non-IID configuration, it is analyzed how performance degrades both for Sentinel and other state-of-the-art robust aggregation methods.
Chao Feng 0001, Alberto Huertas Celdrán, Janosch Baltensperger, Enrique Tomás Martínez Beltrán, Pedro Miguel Sánchez Sánchez, Gérôme Bovet, Burkhard Stiller
ECAI6
2024 Next Generation of AI-based Ransomware
abstract
In the era of the Internet of Things and Artificial Intelligence (AI), cybersecurity has become a critical challenge. Existing AI-based detection systems have shown merit in detecting heterogeneous malware, but their effectiveness against the next generation of AI-powered ransomware encrypting data intelligently is under discussion. In this context, this work introduces a novel AI-powered ransomware that combines Reinforcement Learning and fingerprints based on system calls to evade AI-based cybersecurity detection systems. More in detail, an RL-based agent utilizes Deep Q-Learning as a learning algorithm, system calls to model device states, and the result of unsupervised learning as input for the reward function. The efficacy of this approach has been evaluated on a Raspberry Pi acting as a real spectrum sensor that hosts a behavioral AI-based anomaly detector. Experiments have compared the AI-based ransomware evasion performance with the literature and analyzed how various benign and realistic behaviors affect evasion performance. Results showed that precise and adaptive evasion capabilities can be learned in a few minutes, motivating the need for better cybersecurity systems.
Alberto Huertas Celdrán, Jan von der Assen, Chao Feng 0001, Sandro Padovan, Gérôme Bovet, Burkhard Stiller
GLOBECOM5
2024 ORAN-Sense: Localizing Non-cooperative Transmitters with Spectrum Sensing and 5G O-RAN
abstract
Crowdsensing networks for the sole purpose of performing spectrum measurements have resulted in prior initiatives that have failed primarily due to their costs for maintenance. In this paper, we take a different view and propose ORAN-Sense, a novel architecture of Internet of Things (IoT) spectrum crowd-sensing devices integrated into the Next Generation of cellular networks. We use this framework to extend the capabilities of 5G networks and localize a transmitter that does not collaborate in the process of positioning. While 5G signals can not be applied to this scenario as the transmitter does not participate in the localization process through dedicated pilot symbols and data, we show how to use Time Difference of Arrival-based positioning using low-cost spectrum sensors, minimizing hardware impairments of low-cost spectrum receivers, introducing methods to address errors caused by over-the-air signal propagation, and proposing a low-cost synchronization technique. We have deployed our localization network in two major cities in Europe. Our experimental results indicate that signal localization of non-collaborative transmitters is feasible even using low-cost radio receivers with median accuracies of tens of meters with just a few sensors spanning cities, which makes it suitable for its integration in the Next Generation of cellular networks.
Yago Lizarribar 0001, Roberto Calvo-Palomino, Alessio Scalingi, Giuseppe Santaromita, Gérôme Bovet, Domenico Giustiniano
INFOCOM5
2024 MTFS: a Moving Target Defense-Enabled File System for Malware Mitigation
abstract
Ransomware has remained one of the most notorious threats in the cybersecurity field, for which Moving Target Defense (MTD) has been proposed as a novel defense paradigm. Although various approaches leverage MTD, few of them rely on the operating system and, specifically, the file system, thereby making them dependent on other computing devices, rendering defense against certain threats unrealistic. File-based approaches are less studied here while showing limitations in resource usage and defense effectiveness. Furthermore, existing ransomware defenses merely restore data or detect attacks without preventing them. Thus, this paper introduces the MTFS file system and the design and implementation of three novel MTD techniques – one delaying attackers, one trapping recursive directory traversal, and another one hiding file types. The effectiveness of the techniques is shown in three experiments. First, it is demonstrated that the techniques can delay and mitigate ransomware on real IoT devices. Secondly, in a broader scope, the solution was confronted with 13 ransomware samples, highlighting that it can save 97% of the files. Regarding overhead, the defense system consumes only a small amount of resources, highlighting the feasibility of proactive defense.
Jan von der Assen, Alberto Huertas Celdrán, Rinor Sefa, Burkhard Stiller, Gérôme Bovet
LCN5
2024 Voyager: MTD-Based Aggregation Protocol for Mitigating Poisoning Attacks on DFL
abstract
The growing concern over malicious attacks targeting the robustness of both Centralized and Decentralized Federated Learning (FL) necessitates novel defensive strategies. In contrast to the centralized approach, Decentralized FL (DFL) has the advantage of utilizing network topology and local dataset information, enabling the exploration of Moving Target Defense (MTD) based approaches.This work presents a theoretical analysis of the influence of network topology on the robustness of DFL models. Drawing inspiration from these findings, a three-stage MTD-based aggregation protocol, called Voyager, is proposed to improve the robustness of DFL models against poisoning attacks by manipulating network topology connectivity. Voyager has three main components: an anomaly detector, a network topology explorer, and a connection deployer. When an abnormal model is detected in the network, the topology explorer responds strategically by forming connections with more trustworthy participants to secure the model. Experimental evaluations show that Voyager effectively mitigates various poisoning attacks without imposing significant resource and computational burdens on participants. These findings highlight the proposed reactive MTD as a potent defense mechanism in the context of DFL.
Chao Feng 0001, Alberto Huertas Celdrán, Michael Vuong, Gérôme Bovet, Burkhard Stiller
NOMS4
2024 Analyzing the robustness of decentralized horizontal and vertical federated learning architectures in a non-IID scenario
abstract
Abstract Federated learning (FL) enables participants to collaboratively train machine and deep learning models while safeguarding data privacy. However, the FL paradigm still has drawbacks that affect its trustworthiness, as malicious participants could launch adversarial attacks against the training process. Previous research has examined the robustness of horizontal FL scenarios under various attacks. However, there is a lack of research evaluating the robustness of decentralized vertical FL and comparing it with horizontal FL architectures affected by adversarial attacks. Therefore, this study proposes three decentralized FL architectures: HoriChain, VertiChain, and VertiComb. These architectures feature different neural networks and training protocols suitable for horizontal and vertical scenarios. Subsequently, a decentralized, privacy-preserving, and federated use case with non-IID data to classify handwritten digits is deployed to assess the performance of the three architectures. Finally, a series of experiments computes and compares the robustness of the proposed architectures when they are affected by different data poisoning methods, including image watermarks and gradient poisoning adversarial attacks. The experiments demonstrate that while specific configurations of both attacks can undermine the classification performance of the architectures, HoriChain is the most robust one.
Pedro Miguel Sánchez Sánchez, Alberto Huertas Celdrán, Enrique Tomás Martínez Pérez, Daniel Demeter, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
Appl. Intell.5
2024 Single-board device individual authentication based on hardware performance and autoencoder transformer models
abstract
The proliferation of the Internet of Things (IoT) has led to the emergence of crowdsensing applications, where a multitude of interconnected devices collaboratively collect and analyze data. Ensuring the authenticity and integrity of the data collected by these devices is crucial for reliable decision-making and maintaining trust in the system. Traditional authentication methods are often vulnerable to attacks or can be easily duplicated, posing challenges to securing crowdsensing applications. Besides, current solutions leveraging device behavior are mostly focused on device identification, which is a simpler task than authentication. To address these issues, an individual IoT device authentication framework based on hardware behavior fingerprinting and Transformer autoencoders is proposed in this work. To support the design, a threat model details the security problems faced when performing hardware-based authentication in IoT. This solution leverages the inherent imperfections and variations in IoT device hardware to differentiate between devices with identical specifications. By monitoring and analyzing the behavior of key hardware components, such as the CPU, GPU, RAM, and Storage on devices, unique fingerprints for each device are created. The performance samples are considered as time series data and used to train outlier detection transformer models, one per device and aiming to model its normal data distribution. Then, the framework is validated within a spectrum crowdsensing system leveraging Raspberry Pi devices. After a pool of experiments, the model from each device is able to individually authenticate it between the 45 devices employed for validation. An average True Positive Rate (TPR) of 0.74±0.13 and an average maximum False Positive Rate (FPR) of 0.06±0.09 demonstrate the effectiveness of this approach in enhancing authentication, security, and trust in crowdsensing applications.
Pedro Miguel Sánchez Sánchez, Alberto Huertas Celdrán, Gérôme Bovet, Gregorio Martínez Pérez
Comput. Secur.3
2024 Corrigendum to "Fedstellar: A platform for decentralized federated learning" [Expert Syst. Appl. 242 (2024) 122861]
Enrique Tomás Martínez Beltrán, Ángel Luis Perales Gómez, Chao Feng 0001, Pedro Miguel Sánchez Sánchez, Pedro Guijas Bravo, Sergio López Bernal, Gérôme Bovet, Manuel Gil Pérez, Gregorio Martínez Pérez, Alberto Huertas Celdrán
Expert Syst. Appl.7
2024 Fedstellar: A Platform for Decentralized Federated Learning
abstract
In 2016, Google proposed Federated Learning (FL) as a novel paradigm to train Machine Learning (ML) models across the participants of a federation while preserving data privacy. Since its birth, Centralized FL (CFL) has been the most used approach, where a central entity aggregates participants’ models to create a global one. However, CFL presents limitations such as communication bottlenecks, single point of failure, and reliance on a central server. Decentralized Federated Learning (DFL) addresses these issues by enabling decentralized model aggregation and minimizing dependency on a central entity. Despite these advances, current platforms training DFL models struggle with key issues such as managing heterogeneous federation network topologies, adapting the FL process to virtualized or physical deployments, and using a limited number of metrics to evaluate different federation scenarios for efficient implementation. To overcome these challenges, this paper presents Fedstellar, a novel platform designed to train FL models in a decentralized, semi-decentralized, and centralized fashion across diverse federations of physical or virtualized devices. Fedstellar allows users to create federations by customizing parameters like the number and type of devices training FL models, the network topology connecting them, the machine and deep learning algorithms, or the datasets of each participant, among others. Additionally, it offers real-time monitoring of model and network performance. The Fedstellar implementation encompasses a web application with an interactive graphical interface, a controller for deploying federations of nodes using physical or virtual devices, and a core deployed on each device, which provides the logic needed to train, aggregate, and communicate in the network. The effectiveness of the platform has been demonstrated in two scenarios: a physical deployment involving single-board devices such as Raspberry Pis for detecting cyberattacks and a virtualized deployment comparing various FL approaches in a controlled environment using MNIST and CIFAR-10 datasets. In both scenarios, Fedstellar demonstrated consistent performance and adaptability, achieving F1scores of 91%, 98%, and 91.2% using DFL for detecting cyberattacks and classifying MNIST and CIFAR-10, respectively, reducing training time by 32% compared to centralized approaches.
Enrique Tomás Martínez Beltrán, Ángel Luis Perales Gómez, Chao Feng 0001, Pedro Miguel Sánchez Sánchez, Sergio López Bernal, Gérôme Bovet, Manuel Gil Pérez, Gregorio Martínez Pérez, Alberto Huertas Celdrán
Expert Syst. Appl.6
2024 A Big Data architecture for early identification and categorization of dark web sites
abstract
The dark web has become notorious for its association with illicit activities and there is a growing need for systems to automate the monitoring of this space. This paper proposes an end-to-end scalable architecture for the early identification of new Tor sites and the daily analysis of their content. The solution is built using an Open Source Big Data stack for data serving with Kubernetes, Kafka, Kubeflow, and MinIO, continuously discovering onion addresses in different sources (threat intelligence, code repositories, web-Tor gateways, and Tor repositories), downloading the HTML from Tor and deduplicating the content using MinHash LSH, and categorizing with the BERTopic modeling (SBERT embedding, UMAP dimensionality reduction, HDBSCAN document clustering and c-TF-IDF topic keywords). In 93 days, the system identified 80,049 onion services and characterized 90% of them, addressing the challenge of Tor volatility. A disproportionate amount of repeated content is found, with only 6.1% unique sites. From the HTML files of the dark sites, 31 different low-topics are extracted, manually labeled, and grouped into 11 high-level topics. The five most popular included sexual and violent content, repositories, search engines, carding, cryptocurrencies, and marketplaces. During the experiments, we identified 14 sites with 13,946 clones that shared a suspiciously similar mirroring rate per day, suggesting an extensive common phishing network. Among the related works, this study is the most representative characterization of onion services based on topics to date.
Javier Pastor-Galindo, Hông-Ân Sandlin, Félix Gómez Mármol, Gérôme Bovet, Gregorio Martínez Pérez
Future Gener. Comput. Syst.4
2024 Adversarial attacks and defenses on ML- and hardware-based IoT device fingerprinting and identification
abstract
In the last years, the number of IoT devices deployed has suffered an undoubted explosion, reaching the scale of billions. However, some new cybersecurity issues have appeared together with this development. Some of these issues are the deployment of unauthorized devices, malicious code modification, malware deployment, or vulnerability exploitation. This fact has motivated the requirement for new device identification mechanisms based on behavior monitoring. Besides, these solutions have recently leveraged Machine and Deep Learning (ML/DL) techniques due to the advances in this field and the increase in processing capabilities. In contrast, attackers do not stay stalled and have developed adversarial attacks focused on context modification and ML/DL evaluation evasion applied to IoT device identification solutions. However, literature has not yet analyzed in detail the impact of these attacks on individual identification solutions and their countermeasures. This work explores the performance of hardware behavior-based individual device identification, how it is affected by possible context- and ML/DL-focused attacks, and how its resilience can be improved using defense techniques. In this sense, it proposes an LSTM-CNN architecture based on hardware performance behavior for individual device identification. Then, the most usual ML/DL classification techniques have been compared with the proposed architecture using a hardware performance dataset collected from 45 Raspberry Pi devices running identical software. The LSTM-CNN improves previous solutions achieving a +0.96 average F1-Score and 0.8 minimum TPR for all devices. Afterward, context- and ML/DL-focused adversarial attacks were applied against the previous model to test its robustness. A temperature-based context attack was not able to disrupt the identification, but some ML/DL state-of-the-art evasion attacks were successful. Finally, adversarial training and model distillation defense techniques are selected to improve the model resilience to evasion attacks, improving its robustness from up to 0.88 attack success ratio to 0.17 in the worst attack case, without degrading its performance in an impactful manner.
Pedro Miguel Sánchez Sánchez, Alberto Huertas Celdrán, Gérôme Bovet, Gregorio Martínez Pérez
Future Gener. Comput. Syst.3
2024 FederatedTrust: A solution for trustworthy federated learning
abstract
The rapid expansion of the Internet of Things (IoT) and Edge Computing has presented challenges for centralized Machine and Deep Learning (ML/DL) methods due to the presence of distributed data silos that hold sensitive information. To address concerns regarding data privacy, collaborative and privacy-preserving ML/DL techniques like Federated Learning (FL) have emerged. FL ensures data privacy by design, as the local data of participants remains undisclosed during the creation of a global and collaborative model. However, data privacy and performance are insufficient since a growing need demands trust in model predictions. Existing literature has proposed various approaches dealing with trustworthy ML/DL (excluding data privacy), identifying robustness, fairness, explainability, and accountability as important pillars. Nevertheless, further research is required to identify trustworthiness pillars and evaluation metrics specifically relevant to FL models, as well as to develop solutions that can compute the trustworthiness level of FL models. This work examines the existing requirements for evaluating trustworthiness in FL and introduces a comprehensive taxonomy consisting of six pillars (privacy, robustness, fairness, explainability, accountability, and federation), along with over 30 metrics for computing the trustworthiness of FL models. Subsequently, an algorithm named FederatedTrust is designed based on the pillars and metrics identified in the taxonomy to compute the trustworthiness score of FL models. A prototype of FederatedTrust is implemented and integrated into the learning process of FederatedScope, a well-established FL framework. Finally, five experiments are conducted using different configurations of FederatedScope (with different participants, selection rates, training rounds, and differential privacy) to demonstrate the utility of FederatedTrust in computing the trustworthiness of FL models. Three experiments employ the FEMNIST dataset, and two utilize the N-BaIoT dataset, considering a real-world IoT security use case.
Pedro Miguel Sánchez Sánchez, Alberto Huertas Celdrán, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
Future Gener. Comput. Syst.4
2024 SkyPos: Real-World Evaluation of Self-Positioning With Aircraft Signals for IoT Devices
abstract
Positioning based on aircraft signals has been proposed as an alternative to satellite-based positioning systems (e.g. GPS). However, so far, no deployment of this technique exists, and the real-world performance remains unclear. This paper contributes at better understanding the performance tradeoffs under realistic conditions. We implement SkyPos, a localization system for GPS-denied areas or location integrity that opportunistically uses the large availability of aircraft signals to self-localize receivers. We analyze SkyPos with data collected from hundreds of sensors and thousands of aircraft around Europe. Our results show that we can achieve median accuracy down to 10m in seconds, enabling almost real-time positioning or location verification using aircraft signals at scale.
Yago Lizarribar 0001, Domenico Giustiniano, Gérôme Bovet, Vincent Lenders
IEEE J. Sel. Areas Commun.3
2024 CyberSpec: Behavioral Fingerprinting for Intelligent Attacks Detection on Crowdsensing Spectrum Sensors
abstract
Integrated sensing and communication is a novel paradigm using crowdsensing spectrum sensors to help with the management of spectrum scarcity. However, well-known vulnerabilities of resource-constrained spectrum sensors and the possibility of being manipulated by users with physical access complicate their protection against spectrum sensing data falsification (SSDF) attacks. Most recent literature suggests using behavioral fingerprinting and Machine/Deep Learning (ML/DL) for improving similar cybersecurity issues. Nevertheless, the applicability of these techniques in resource-constrained devices, the impact of attacks affecting spectrum data integrity, and the performance and scalability of models suitable for heterogeneous sensors types are still open challenges. To improve limitations, this work presents seven SSDF attacks affecting spectrum sensors and introduces CyberSpec, an ML/DL-oriented framework using device behavioral fingerprinting to detect anomalies produced by SSDF attacks. CyberSpec has been implemented and validated in ElectroSense, a real crowdsensing RF monitoring platform where several configurations of the proposed SSDF attacks have been executed in different sensors. A pool of experiments with different unsupervised ML/DL-based models has demonstrated the suitability of CyberSpec detecting the previous attacks within an acceptable timeframe.
Alberto Huertas Celdrán, Pedro Miguel Sánchez Sánchez, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
IEEE Trans. Dependable Secur. Comput.3
2024 Studying the Robustness of Anti-Adversarial Federated Learning Models Detecting Cyberattacks in IoT Spectrum Sensors
abstract
Device fingerprinting combined with Machine and Deep Learning (ML/DL) report promising performance when detecting spectrum sensing data falsification (SSDF) attacks. However, the amount of data needed to train models and the scenario privacy concerns limit the applicability of centralized ML/DL. Federated learning (FL) addresses these drawbacks but is vulnerable to adversarial participants and attacks. The literature has proposed countermeasures, but more effort is required to evaluate the performance of FL detecting SSDF attacks and their robustness against adversaries. Thus, the first contribution of this work is to create an FL-oriented dataset modeling the behavior of resource-constrained spectrum sensors affected by SSDF attacks. The second contribution is a pool of experiments analyzing the robustness of FL models according to i) three families of sensors, ii) eight SSDF attacks, iii) four FL scenarios dealing with anomaly detection and binary classification, iv) up to 33% of participants implementing data and model poisoning attacks, and v) four aggregation functions acting as anti-adversarial mechanisms. In conclusion, FL achieves promising performance when detecting SSDF attacks. Without anti-adversarial mechanisms, FL models are particularly vulnerable with$>$16% of adversaries. Coordinate-wise-median is the best mitigation for anomaly detection, but binary classifiers are still affected with$>$33% of adversaries.
Pedro Miguel Sánchez Sánchez, Alberto Huertas Celdrán, Timo Schenk, Adrian Lars Benjamin Iten, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
IEEE Trans. Dependable Secur. Comput.5
2024 RL and Fingerprinting to Select Moving Target Defense Mechanisms for Zero-Day Attacks in IoT
abstract
Moving Target Defense (MTD) is a promising approach to mitigate attacks by dynamically altering target attack surfaces. Still, selecting suitable MTD techniques for zero-day attacks is an open challenge. Reinforcement Learning (RL) could be an effective approach to optimize the MTD selection through trial and error, but the literature fails when i) evaluating the performance of RL and MTD solutions in real-world scenarios, ii) studying whether behavioral fingerprinting is suitable for RL, and iii) calculating the consumption of resources in single-board computers (SBC). Thus, the work at hand proposes an online RL-based framework that learns correct MTD mechanisms mitigating heterogeneous zero-day attacks in SBC. The framework considers behavioral fingerprinting to represent SBCs’ states and RL to learn MTD techniques that mitigate each malicious state. It has been deployed on a real IoT crowdsensing scenario with a Raspberry Pi acting as a spectrum sensor. The Raspberry Pi has been infected with different samples of command and control malware, rootkits, and ransomware to later select between four existing MTD techniques. A set of experiments demonstrated the suitability of the framework to learn proper MTD techniques mitigating all attacks (except a harmfulness rootkit) while consuming$\approx 10$% of RAM, and negligible CPU.
Alberto Huertas Celdrán, Pedro Miguel Sánchez Sánchez, Jan von der Assen, Timo Schenk, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
IEEE Trans. Inf. Forensics Secur.5
2024 Mitigating communications threats in decentralized federated learning through moving target defense
abstract
Abstract The rise of Decentralized Federated Learning (DFL) has enabled the training of machine learning models across federated participants, fostering decentralized model aggregation and reducing dependence on a server. However, this approach introduces unique communication security challenges that have yet to be thoroughly addressed in the literature. These challenges primarily originate from the decentralized nature of the aggregation process, the varied roles and responsibilities of the participants, and the absence of a central authority to oversee and mitigate threats. Addressing these challenges, this paper first delineates a comprehensive threat model focused on DFL communications. In response to these identified risks, this work introduces a security module to counter communication-based attacks for DFL platforms. The module combines security techniques such as symmetric and asymmetric encryption with Moving Target Defense (MTD) techniques, including random neighbor selection and IP/port switching. The security module is implemented in a DFL platform, Fedstellar, allowing the deployment and monitoring of the federation. A DFL scenario with physical and virtual deployments have been executed, encompassing three security configurations: (i) a baseline without security, (ii) an encrypted configuration, and (iii) a configuration integrating both encryption and MTD techniques. The effectiveness of the security module is validated through experiments with the MNIST dataset and eclipse attacks.The results showed an average F1 score of 95%, with the most secure configuration resulting in CPU usage peaking at 68% (± 9%) in virtual deployments and network traffic reaching 480.8 MB (± 18 MB), effectively mitigating risks associated with eavesdropping or eclipse attacks.
Enrique Tomás Martínez Beltrán, Pedro Miguel Sánchez Sánchez, Sergio López Bernal, Gérôme Bovet, Manuel Gil Pérez, Gregorio Martínez Pérez, Alberto Huertas Celdrán
Wirel. Networks4
2023 RansomAI: AI-Powered Ransomware for Stealthy Encryption
abstract
Cybersecurity solutions have shown promising performance when detecting ransomware samples that use fixed algorithms and encryption rates. However, due to the current explosion of Artificial Intelligence (AI), sooner than later, ransomware, and malware in general, will incorporate AI techniques to intelligently and dynamically adapt its behavior to be undetected. It might result in ineffective and obsolete cybersecurity solutions, but the literature lacks AI-powered ransomware samples to verify it. Thus, this work proposes RansomAI, a Reinforcement Learning-based framework that can be integrated into existing ransomware samples to adapt their encryption behavior and stay stealthy while encrypting files. RansomAI presents an agent that learns the best encryption algorithm, rate, and duration that minimizes its detection (using a reward mechanism and a fingerprinting intelligent detection system) while maximizing its damage. The proposed framework was validated with Ransomware-PoC, a ransomware that infected a Raspberry Pi 4 acting as a crowdsensor. A pool of experiments with Deep Q-Learning and Isolation Forest (deployed on the agent and detection system, respectively) has demonstrated that RansomAI evades the detection of Ransomware-PoC affecting the Raspberry Pi 4 in a few minutes with >90% accuracy.
Jan von der Assen, Alberto Huertas Celdrán, Janik Luechinger, Pedro Miguel Sánchez Sánchez, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
GLOBECOM5
2023 Maximum Likelihood Distillation for Robust Modulation Classification
abstract
Deep Neural Networks are being extensively used in communication systems and Automatic Modulation Classification (AMC) in particular. However, they are very susceptible to small adversarial perturbations that are carefully crafted to change the network decision. In this work, we build on knowledge distillation ideas and adversarial training in order to build more robust AMC systems. We first outline the importance of the quality of the training data in terms of accuracy and robustness of the model. We then propose to use the Maximum Likelihood function, which could solve the AMC problem in offline settings, to generate better training labels. Those labels teach the model to be uncertain in challenging conditions, which permits to increase the accuracy, as well as the robustness of the model when combined with adversarial training. Interestingly, we observe that this increase in performance transfers to online settings, where the Maximum Likelihood function cannot be used in practice. Overall, this work highlights the potential of learning to be uncertain in difficult scenarios, compared to directly removing label noise.
Javier Maroto, Gérôme Bovet, Pascal Frossard
ICASSP2
2023 A Lightweight Moving Target Defense Framework for Multi-purpose Malware Affecting IoT Devices
abstract
Malware affecting Internet of Things (IoT) devices is rapidly growing due to the relevance of this paradigm in real-world scenarios. Specialized literature has also detected a trend towards multi-purpose malware able to execute different malicious actions such as remote control, data leakage, encryption, or code hiding, among others. Protecting IoT devices against this kind of malware is challenging due to their well-known vulnerabilities and limitation in terms of CPU, memory, and storage. To improve it, the moving target defense (MTD) paradigm was proposed a decade ago and has shown promising results, but there is a lack of IoT MTD solutions dealing with multi-purpose malware. Thus, this work proposes four MTD mechanisms changing IoT devices' network, data, and runtime environment to mitigate multi-purpose malware. Furthermore, it presents a lightweight and IoT-oriented MTD framework to decide what, when, and how the MTD mechanisms are deployed. Finally, the efficiency and effectiveness of the framework and MTD mechanisms are evaluated in a real-world scenario with one IoT spectrum sensor affected by multi-purpose malware.
Jan von der Assen, Alberto Huertas Celdrán, Pedro Miguel Sánchez Sánchez, Jordan Cedeño, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
ICC5
2023 Fedstellar: A Platform for Training Models in a Privacy-preserving and Decentralized Fashion
abstract
This paper presents Fedstellar, a platform for training decentralized Federated Learning (FL) models in heterogeneous topologies in terms of the number of federation participants and their connections. Fedstellar allows users to build custom topologies, enabling them to control the aggregation of model parameters in a decentralized manner. The platform offers a Web application for creating, managing, and connecting nodes to ensure data privacy and provides tools to measure, monitor, and analyze the performance of the nodes. The paper describes the functionalities of Fedstellar and its potential applications. To demonstrate the applicability of the platform, different use cases are presented in which decentralized, semi-decentralized, and centralized architectures are compared in terms of model performance, convergence time, and network overhead when collaboratively classifying hand-written digits using the MNIST dataset.
Enrique Tomás Martínez Beltrán, Pedro Miguel Sánchez Sánchez, Sergio López Bernal, Gérôme Bovet, Manuel Gil Pérez, Gregorio Martínez Pérez, Alberto Huertas Celdrán
IJCAI4
2023 Contrastive learning with self-reconstruction for channel-resilient modulation classification
abstract
Despite the substantial success of deep learning for Automatic Modulation Classification (AMC), models trained on a specific transmitter configuration and channel model often fail to generalize well to other scenarios with different transmitter configurations, wireless fading channels, or receiver impairments such as clock offset. This paper proposes Contrastive Learning with Self-Reconstruction called CLSR-AMC to learn good representations of signals resilient to channel changes. While contrastive loss focuses on the differences between individual modulations, the reconstruction loss captures representative features of the signal. Additionally, we develop three data augmentation operators to emulate the impact of channel and hardware impairments without exhaustive modeling of different channel profiles. We perform extensive experimentation with commonly used realistic datasets. We show that CLSR-AMC outperforms its counterpart based on contrastive learning for the same amount of labeled data by significant average accuracy gains of 24.29%, 17.01%, and 15.97% in the Additive White Gaussian Noise (AWGN), Rayleigh, and Rician channels, respectively.
Erma Perenda, Sreeraj Rajendran, Gérôme Bovet, Mariya Zheleva, Sofie Pollin
INFOCOM3
2023 A Framework for Wireless Technology Classification using Crowdsensing Platforms
abstract
Spectrum crowdsensing systems do not provide labeled data near real-time yet. We propose a framework that addresses this challenge and relies solely on Power Spectrum Density (PSD) data collected by low-cost receivers. A major hurdle is to design a system that is computationally efficient for near real-time operation, yet using only the limited 2 MHz bandwidth of low-cost spectrum sensors. First, we present a method for unsupervised transmission detection that works with PSD data already collected by the backend of the crowdsensing platform, and that provides stable detection of transmission boundaries. Second, we introduce a data-driven deep learning solution to classify the wireless technology used by the transmitter, using transmission features in a compressed space extracted from single PSD measurements over at most 2 MHz band. We build an experimental platform, and evaluate our framework with real-world data collected from 47 different sensors deployed across Europe. We show that our framework yields an average classification accuracy close to 94.25% over the testing dataset, with a maximum latency of 3.4 seconds when integrated in the backend of a major crowdsensing network. Code and data have been released for reproducibility and further studies.
Alessio Scalingi, Domenico Giustiniano, Roberto Calvo-Palomino, Nikolaos Apostolakis, Gérôme Bovet
INFOCOM5
2023 SecBox: A Lightweight Container-based Sandbox for Dynamic Malware Analysis
abstract
Cybersecurity solutions based on machine learning (ML) and behavioral fingerprinting have demonstrated their suitability when detecting heterogeneous malware. However, most solutions are black boxes missing explainable and visual capabilities needed to analyze relevant metrics and malicious behaviors to be collected. In this demonstration, SecBox, a dynamic malware analysis platform with integrated data collection and visualization for malware execution, is presented. To provide a lightweight sandboxing approach, the architecture relies on Linux containers for isolation. The sandboxing and data analysis components of the SecBox architecture are deployed in a test bed to show the analysis of two malware families. In the presented scenario, the Monti ransomware and CoinMiner, a Monero-based cryptojacker are analyzed after obtaining them from a public database.
Jan von der Assen, Alberto Huertas Celdrán, Adrian Zermin, Raffael Mogicato, Gérôme Bovet, Burkhard Stiller
NOMS5
2023 Early Detection of Cryptojacker Malicious Behaviors on IoT Crowdsensing Devices
abstract
Traditionally, IoT crowdsensing devices have been outside the cryptomining domain due to their limitations in terms of computational power. In 2014, Monero (XNR) changed this situation forever. Monero is an open-source digital payment token that can be mined in resource-constrained devices like IoT and single-board computers. Despite the Monero advantages, it opened the door for cryptojackers illicitly mining cryptocurrencies by exploiting well-known vulnerabilities of IoT devices. Existing detection solutions provide good performance while detecting the mining phase of cryptojackers, but early detection is desired to avoid malware spreading and resource misuse. Thus, this work proposes a framework that combines device behavioral fingerprinting and machine learning to detect and classify preparatory phases of cryptojackers. The framework has been deployed in a crowdsensing IoT spectrum sensor, Raspberry Pi, infected by a recent cryptojacker called Linux.MulDrop.14. Promising detection results demonstrate the framework’s suitability while detecting early phases of cryptojackers.
Alberto Huertas Celdrán, Jan von der Assen, Konstantin Moser, Pedro Miguel Sánchez Sánchez, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
NOMS5
2023 Behavioral fingerprinting to detect ransomware in resource-constrained devices
abstract
The Internet of Things (IoT), a network of interconnected devices, has grown and gained traction over the last few years. This paradigm can impact our lives while also providing significant economic benefits. However, although resource-constrained IoT devices offer numerous advantages, they are also vulnerable to cyberattacks. As a result, ransomware severely threatens IoT devices managing sensitive and relevant information. Solutions based on Machine and Deep Learning (ML/DL) that consider behavioral data have been identified as promising. However, most detection solutions have been developed for Windows-based systems, which generally have more resources than IoT devices. As a result, these solutions are not suitable for resource-constrained components. In addition, no solution compares the pros and cons of different behavioral dimensions of resource-constrained devices. Thus, this work presents a framework that combines three different behavioral sources with supervised and unsupervised ML/DL algorithms to detect and classify heterogeneous ransomware impacting resource-constrained spectrum sensors. A pool of experiments has demonstrated the suitability of the proposed solution and compared its performance with a rule-based system. In conclusion, the usage of resources combined with local outlier factor and decision tree are the most promising combinations to detect anomalies and classify ransomware while consuming CPU, RAM, and time of devices in a reduced manner.
Alberto Huertas Celdrán, Pedro Miguel Sánchez Sánchez, Jan von der Assen, Dennis Shushack, Ángel Luis Perales Gómez, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
Comput. Secur.6
2023 Privacy-Preserving and Syscall-Based Intrusion Detection System for IoT Spectrum Sensors Affected by Data Falsification Attacks
abstract
Crowdsensing platforms collect, process, transmit, and analyze spectrum data worldwide to optimize radio frequency spectrum usage. However, Internet of Things (IoT) spectrum sensors, performing some of the previous tasks, are exposed to software manipulation aiming to execute spectrum sensing data falsification (SSDF) attacks to compromise data integrity and spectrum optimization. Novel intrusion detection systems (IDSs) combining device fingerprinting with machine and deep learning (ML/DL) improve the limitation of traditional solutions and remove the necessity of redundant sensors and reputation mechanisms. However, they fail when detecting SSDF attacks accurately while protecting sensors privacy. This work proposes a novel host-based and federated learning-oriented IDS for IoT spectrum sensors that consider unsupervised ML/DL and fingerprints based on system calls. The framework detection performance and consumption of resources are analyzed in local and federated scenarios with six spectrum sensors deployed on Raspberry Pis. The obtained results significantly improve related work when detecting SSDF attacks while protecting sensors privacy, and consuming CPU, memory, and storage of sensors in a reduced manner.
Alberto Huertas Celdrán, Pedro Miguel Sánchez Sánchez, Chao Feng 0001, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
IEEE Internet Things J.4
2023 A methodology to identify identical single-board computers based on hardware behavior fingerprinting
abstract
The connectivity and resource-constrained nature of single-board devices open the door to cybersecurity concerns affecting Internet of Things (IoT) scenarios. One of the most important issues is the presence of unauthorized IoT devices that want to impersonate legitimate ones by using identical hardware and software specifications. This situation can provoke sensitive information leakages, data poisoning, or privilege escalation in IoT scenarios. Combining behavioral fingerprinting and Machine/Deep Learning (ML/DL) techniques is a promising approach to identify these malicious spoofing devices by detecting minor performance differences generated by imperfections in manufacturing. However, existing solutions are not suitable for single-board devices since they do not consider their hardware and software limitations, underestimate critical aspects such as fingerprint stability or context changes, and do not explore the potential of ML/DL techniques. To improve it, this work first identifies the essential properties for single-board device identification: uniqueness, stability, diversity, scalability, efficiency, robustness, and security. Then, a novel methodology relies on behavioral fingerprinting to identify identical single-board devices and meet the previous properties. The methodology leverages the different built-in components of the system and ML/DL techniques, comparing the device internal behavior with each other to detect variations that occurred in manufacturing processes. The methodology validation has been performed in a real environment composed of 15 identical Raspberry Pi 4 Model B and 10 Raspberry Pi 3 Model B+ devices, obtaining a 91.9% average TPR with an XGBoost model and achieving the identification for all devices by setting a 50% threshold in the evaluation process. Finally, a discussion compares the proposed solution with related work, highlighting the fingerprint properties not met, and provides important lessons learned and limitations.
Pedro Miguel Sánchez Sánchez, José María Jorquera Valero, Alberto Huertas Celdrán, Gérôme Bovet, Manuel Gil Pérez, Gregorio Martínez Pérez
J. Netw. Comput. Appl.4
2023 Adaptive Uplink Data Compression in Spectrum Crowdsensing Systems
abstract
Understanding spectrum activity is challenging when attempted at scale. The wireless community has recently risen to this challenge in designing spectrum monitoring systems that utilize many low-cost spectrum sensors to gather large volumes of sampled data across space, time, and frequencies. These crowdsensing systems are limited by the uplink bandwidth available to backhaul the raw in-phase and quadrature (IQ) samples and power spectrum density (PSD) data needed to run various applications. This paper presents FlexSpec, a framework based on the Walsh-Hadamard transform to compress spectrum data collected from distributed and low-cost sensors for real-time applications. This transformation allows sensors to significantly save uplink bandwidth thanks to its inherent properties both when it is applied to IQ and PSD data. Additionally, by leveraging a feedback loop between the sensor and the edge device it connects to, FlexSpec carefully adapts the compression ratio over time to changes in the spectrum and different applications, jointly considering data size, application performance, and spectrum variations. We experimentally evaluate FlexSpec in several applications. Our results show that FlexSpec is particularly suitable for IoT transmissions and signals close to the noise floor. Compared with prior work, FlexSpec provides up to$7\times $more reduction of uplink data size for signal detection based on PSD data, and reduces up to$6\times $to$8\times $the number of undecodable messages for IQ sample decoding.
Yijing Zeng, Roberto Calvo-Palomino, Domenico Giustiniano, Gérôme Bovet, Suman Banerjee 0001
IEEE/ACM Trans. Netw.4
2022 RITUAL: a Platform Quantifying the Trustworthiness of Supervised Machine Learning
abstract
This demo presents RITUAL, a platform composed of a novel algorithm and a Web application quantifying the trustworthiness level of supervised Machine and Deep Learning (ML/DL) models according to their fairness, explainability, robustness, and accountability. The algorithm is deployed on a Web application to allow users to quantify and compare the trustworthiness of their ML/DL models. Finally, a scenario with ML/DL models classifying network cyberattacks demonstrates the platform applicability.
Alberto Huertas Celdrán, Melike Demirci, Joel Leupp, Muriel Figueredo Franco, Pedro Miguel Sánchez Sánchez, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
CNSM7
2022 Intelligent Fingerprinting to Detect Data Leakage Attacks on Spectrum Sensors
abstract
Data confidentiality protection is a must for IoT and crowdsensing platforms, and a challenge due to the constrained nature of their sensors. Currently, the combination of device fingerprinting and anomaly detection systems based on Machine and Deep Learning (ML/DL) techniques is one of the most promising approaches to detect zero-day cyberattacks. However, most of existing work is not suitable for resource-constrained devices or does not deal with cyberattacks affecting data confidentiality of spectrum sensors. Thus, this paper proposes a framework that monitors network interface events of sensors, uses unsupervised learning to create fingerprints, and detects anomalies produced by such cyberattacks. The framework validation has been performed in the crowdsensing platform ElectroSense, where a sensor has been infected by a backdoor leaking different sensitive data during an experiment. A set of unsupervised learning algorithms has been evaluated, being Autoencoder the one showing the best balance when detecting normal behavior and data leakages of different sizes and at frequencies, while providing a reduced detection time and sensor resources consumption.
Alberto Huertas Celdrán, Pedro Miguel Sánchez Sánchez, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
ICC3
2022 Policy-based and Behavioral Framework to Detect Ransomware Affecting Resource-constrained Sensors
abstract
Traditionally, data centers have been the preferred target for ransomware attacks. However, the increasing number of IoT (Internet-of-Things) devices managing valuable data is attracting the attention of cybercriminals and ransomware towards resource-constrained devices. So far, literature has demonstrated the suitability of monitoring the behavior of devices to detect some malware infections. However, most of these existing solutions have been designed and validated in Windows-based systems without computational restrictions.Thus, this work presents a lightweight policy-based framework that uses behavioral fingerprinting to detect anomalies and classify ransomware affecting resource-constrained and Linux-based sensors. The framework detection capabilities have been validated in a resource-constrained spectrum sensor belonging to ElectroSense, a real crowdsensing platform. In particular, three policies, created as a proof-of-concept, resulted in promising findings in terms of detection performance and time, when identifying anomalies by classifying two recent ransomware samples affecting a Raspberry Pi acting as sensor.
Alberto Huertas Celdrán, Pedro Miguel Sánchez Sánchez, Eder J. Scheid, Timucin Besken, Gérôme Bovet, Gregorio Martínez Pérez, Burkhard Stiller
NOMS5
2022 Federated learning for malware detection in IoT devices
abstract
Billions of IoT devices lacking proper security mechanisms have been manufactured and deployed for the last years, and more will come with the development of Beyond 5G technologies. Their vulnerability to malware has motivated the need for efficient techniques to detect infected IoT devices inside networks. With data privacy and integrity becoming a major concern in recent years, increasing with the arrival of 5G and Beyond networks, new technologies such as federated learning and blockchain emerged. They allow training machine learning models with decentralized data while preserving its privacy by design. This work investigates the possibilities enabled by federated learning concerning IoT malware detection and studies security issues inherent to this new learning paradigm. In this context, a framework that uses federated learning to detect malware affecting IoT devices is presented. N-BaIoT, a dataset modeling network traffic of several real IoT devices while affected by malware, has been used to evaluate the proposed framework. Both supervised and unsupervised federated models (multi-layer perceptron and autoencoder) able to detect malware affecting seen and unseen IoT devices of N-BaIoT have been trained and evaluated. Furthermore, their performance has been compared to two traditional approaches. The first one lets each participant locally train a model using only its own data, while the second consists of making the participants share their data with a central entity in charge of training a global model. This comparison has shown that the use of more diverse and large data, as done in the federated and centralized methods, has a considerable positive impact on the model performance. Besides, the federated models, while preserving the participant’s privacy, show similar results as the centralized ones. As an additional contribution and to measure the robustness of the federated approach, an adversarial setup with several malicious participants poisoning the federated model has been considered. The baseline model aggregation averaging step used in most federated learning algorithms appears highly vulnerable to different attacks, even with a single adversary. The performance of other model aggregation functions acting as countermeasures is thus evaluated under the same attack scenarios. These functions provide a significant improvement against malicious participants, but more efforts are still needed to make federated approaches robust.
Valerian Rey, Pedro Miguel Sánchez Sánchez, Alberto Huertas Celdrán, Gérôme Bovet
Comput. Networks4
2021 Learning the unknown: Improving modulation classification performance in unseen scenarios
abstract
Automatic Modulation Classification (AMC) is significant for the practical support of a plethora of emerging spectrum applications, such as Dynamic Spectrum Access (DSA) in 5G and beyond, resource allocation, jammer identification, intruder detection, and in general, automated interference analysis. Although a well-known problem, most of the existing AMC work has been done under the assumption that the classifier has prior knowledge about the signal and channel parameters. This paper shows that unknown signal and channel parameters significantly degrade the performance of two of the most popular research streams in modulation classification: expert feature-based and data-driven. By understanding why and where those methods fail, in such unknown scenarios, we propose two possible directions to make AMC more robust to signal shape transformations introduced by unknown signal and channel parameters. We show that Spatial Transformer Networks (STN) and Transfer Learning (TL) embedded into a light ResNeXt-based classifier can improve average classification accuracy up to 10-30% for specific unseen scenarios with only 5% labeled data for a large dataset of 20 complex higher-order modulations.
Erma Perenda, Sreeraj Rajendran, Gérôme Bovet, Sofie Pollin, Mariya Zheleva
INFOCOM3
2020 SkySense: terrestrial and aerial spectrum use analysed using lightweight sensing technology with weather balloons
abstract
Given the availability of lightweight radio and processing technology, it becomes feasible to imagine spectrum sensing systems using weather balloons. Such balloons navigate the airspace up to 40 km, and can provide a bird's eye and clear view of terrestrial, as well as aerial spectrum use. In this paper, we present SkySense, which is an extension of the Electrosense sensing framework with mobile GPS-located sensors and local data logging. In addition, we present 6 different sensing campaigns, targeting multiple terrestrial or aerial technologies such as ADS-B, AIS or LTE. For instance, for ADS-B, we can clearly conclude that the number of airplanes that are detected is the same for each balloon altitude, but the message reception rate decreases strongly with altitude because of collisions. For each sensing campaign, the dataset is described, and some example spectrum analysis results are presented. In addition, we analyse and quantify important trends visible when sensing from the sky, such as temperature and hardware variations, increased ambient interference levels, as well as hardware limitations of the lightweight system. A key challenge is the automatic gain control and dynamic range of the system, as a radio navigating over 30km, sees a very wide range of possible signal levels. All data is publicly available through the Electrosense framework, to encourage the spectrum sensing community to further analyse the data or motivate further measurement campaigns using weather balloons.
Brecht Reynders, Franco Minucci, Erma Perenda, Hazem Sallouha, Roberto Calvo-Palomino, Yago Lizarribar 0001, Markus Fuchs, Matthias Schäfer 0002, Markus Engel, Bertold Van den Bergh, Sofie Pollin, Domenico Giustiniano, Gérôme Bovet, Vincent Lenders
MobiSys13
2020 Short: LSTM-based GNSS Spoofing Detection Using Low-cost Spectrum Sensors
abstract
GNSS/GPS is a positioning system widely used nowadays in our lives for real-time localization in Earth. This technology is highly vulnerable to spoofing/jamming attacks caused by malicious intruders. In the recent years, commodity and low-cost radio-frequency hardware have been used to interfere with the legitimate GPS signal. Existing spoofing detection solutions use costly receivers and computationally expensive algorithms which limit the large-scale deployment. In this work we propose a GNSS spoofing detection system that can run on spectrum sensors with Software-Defined Radio (SDR) capabilities and cost in the order of 20 euros. Our approach exploits the predictability of the Doppler characteristics of the received GPS signals to determine the presence of anomalies or malicious attackers. We propose an artificial recurrent neural network (RNN) based on Long short-term memory (LSTM) for anomaly detection. We use data received by low-cost SDR receivers that are processed locally by low-cost embedded machines such as Nvidia Jetson Nano to provide inference capabilities. We show that our solution predicts very accurately the Doppler shift of GNSS signals and can determine the presence of a spoofing transmitter.
Roberto Calvo-Palomino, Arani Bhattacharya, Gérôme Bovet, Domenico Giustiniano
WoWMoM3
2014 A distributed web-based naming system for smart buildings
abstract
Nowadays, pervasive application scenarios relying on sensor networks are gaining momentum. The field of smart buildings is a promising playground where the use of sensors allows a reduction of the overall energy consumption. Most of current applications are using the classical DNS which is not suited for the Internet-of-Things because of requiring humans to get it working. From another perspective, Web technologies are pushing in sensor networks following the Web-of-Things paradigm advocating to use RESTful APIs for manipulating resources representing device capabilities. Being aware of these two observations, we propose to build on top of Web technologies leading to a novel naming system that is entirely autonomous. In this work, we describe the architecture supporting what can be called an autonomous Web-oriented naming system. As proof of concept, we simulate a rather large building and compare the behaviour of our approach to the legacy DNS and Multicast DNS (mDNS).
Gérôme Bovet, Jean Hennebert
WoWMoM1
2013 Introducing the Web-of-Things in Building Automation - A Gateway for KNX Installations
abstract
Due to increasing energy costs and the importance of the comfort, smart buildings tend to democratize both in new and renovated constructions, based on management systems relying on dedicated networks. Network heterogeneity leads to complex building management systems having to implement all the protocols of the building networks, resulting in low system integration and heavy maintenance efforts. Those building networks offer no common standardized application layer to build applications. To remedy this, we propose in this paper to leverage on the Web-of-Things (WoT) framework, using well-known technologies like HTTP and RESTful APIs. We outline the implementation of a gateway using the principles of the WoT to expose capabilities of the KNX building network asWeb services, allowing a fast integration in management systems.
Gérôme Bovet, Jean Hennebert
ICINCO (1)1