Han Wang 0031

dblp:67/1771-31 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
5since 2021 · last 2024
0000-0002-2772-4661ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2024 BMI: Bounded Mutual Information for Efficient Privacy-Preserving Feature Selection
David Eklund, Alfonso Iacovazzi, Han Wang 0031, Apostolos Pyrgelis, Shahid Raza
ESORICS (2)3
2024 CLEVER: Crafting Intelligent MISP for Cyber Threat Intelligence
abstract
Cyber Threat Intelligence (CTI) is crucial for modern cybersecurity because it provides the knowledge and insights needed to defend against a wide range of cyber threats. However, there are issues associated with incomplete and inconsistent CTI data that can lead to inaccurate threat assessments, increasing the risk of both false alarms and undetected threats. This paper introduces CLEVER, an extended version of the Malware Information Sharing Platform (MISP) platform that includes machine learning (ML) models to support the management and processing of CTI data. The models are designed to address specific challenges such as (i) prioritizing and ranking Indicators of Compromise (IoCs) based on severity and potential impact, (ii) classifying IoCs by attack type or threat, and (iii) aggregating similar IoCs into clusters. The effectiveness of the ML models employed in CLEVER has been thoroughly tested on three public CTI datasets, and the results provide encouraging outcomes in enhancing CTI management and analysis.
Han Wang 0031, Alfonso Iacovazzi, Seonghyun Kim, Shahid Raza
LCN1
2023 SparSFA: Towards robust and communication-efficient peer-to-peer federated learning
abstract
Federated Learning (FL) has emerged as a powerful paradigm to train collaborative machine learning (ML) models, preserving the privacy of the participants’ datasets. However, standard FL approaches present some limitations that can hinder their applicability in some applications. Thus, the need of a server or aggregator to orchestrate the learning process may not be possible in scenarios with limited connectivity, as in some IoT applications, and offer less flexibility to personalize the ML models for the different participants. To sidestep these limitations, peer-to-peer FL (P2PFL) provides more flexibility, allowing participants to train their own models in collaboration with their neighbors. However, given the huge number of parameters of typical Deep Neural Network architectures, the communication burden can also be very high. On the other side, it has been shown that standard aggregation schemes for FL are very brittle against data and model poisoning attacks. In this paper, we propose SparSFA, an algorithm for P2PFL capable of reducing the communication costs. We show that our method outperforms competing sparsification methods in P2P scenarios, speeding the convergence and enhancing the stability during training. SparSFA also includes a mechanism to mitigate poisoning attacks for each participant in any random network topology. Our empirical evaluation on real datasets for intrusion detection in IoT, considering both balanced and imbalanced-dataset scenarios, shows that SparSFA is robust to different indiscriminate poisoning attacks launched by one or multiple adversaries, outperforming other robust aggregation methods whilst reducing the communication costs through sparsification.
Han Wang 0031, Luis Muñoz-González, Muhammad Zaid Hameed, David Eklund, Shahid Raza
Comput. Secur.1
2023 FL4IoT: IoT Device Fingerprinting and Identification Using Federated Learning
abstract
Unidentified devices in a network can result in devastating consequences. It is, therefore, necessary to fingerprint and identify IoT devices connected to private or critical networks. With the proliferation of massive but heterogeneous IoT devices, it is getting challenging to detect vulnerable devices connected to networks. Current machine learning-based techniques for fingerprinting and identifying devices necessitate a significant amount of data gathered from IoT networks that must be transmitted to a central cloud. Nevertheless, private IoT data cannot be shared with the central cloud in numerous sensitive scenarios. Federated learning (FL) has been regarded as a promising paradigm for decentralized learning and has been applied in many different use cases. It enables machine learning models to be trained in a privacy-preserving way. In this article, we propose a privacy-preserved IoT device fingerprinting and identification mechanisms using FL; we call it FL4IoT. FL4IoT is a two-phased system combining unsupervised-learning-based device fingerprinting and supervised-learning-based device identification. FL4IoT shows its practicality in different performance metrics in a federated and centralized setup. For instance, in the best cases, empirical results show that FL4IoT achieves ∼99% accuracy and F1-Score in identifying IoT devices using a federated setup without exposing any private data to a centralized cloud entity. In addition, FL4IoT can detect spoofed devices with over 99% accuracy .
Han Wang 0031, David Eklund, Alina Oprea, Shahid Raza
ACM Trans. Internet Things1
2021 Non-IID data re-balancing at IoT edge with peer-to-peer federated learning for anomaly detection
abstract
The increase of the computational power in edge devices has enabled the penetration of distributed machine learning technologies such as federated learning, which allows to build collaborative models performing the training locally in the edge devices, improving the efficiency and the privacy for training of machine learning models, as the data remains in the edge devices. However, in some IoT networks the connectivity between devices and system components can be limited, which prevents the use of federated learning, as it requires a central node to orchestrate the training of the model. To sidestep this, peer-to-peer learning appears as a promising solution, as it does not require such an orchestrator. On the other side, the security challenges in IoT deployments have fostered the use of machine learning for attack and anomaly detection. In these problems, under supervised learning approaches, the training datasets are typically imbalanced, i.e. the number of anomalies is very small compared to the number of benign data points, which requires the use of re-balancing techniques to improve the algorithms' performance. In this paper, we propose a novel peer-to-peer algorithm,P2PK-SMOTE, to train supervised anomaly detection machine learning models in non-IID scenarios, including mechanisms to locally re-balance the training datasets via synthetic generation of data points from the minority class. To improve the performance in non-IID scenarios, we also include a mechanism for sharing a small fraction of synthetic data from the minority class across devices, aiming to reduce the risk of data de-identification. Our experimental evaluation in real datasets for IoT anomaly detection across a different set of scenarios validates the benefits of our proposed approach.
Han Wang 0031, Luis Muñoz-González, David Eklund, Shahid Raza
WISEC1
2017 Timeline Summarization for Event-Related Discussions on a Chinese Social Media Platform
Han Wang 0031, Jia-Ling Koh
IEA/AIE (1)1