Gbadebo Ayoade

dblp:169/8110 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
3since 2021 · last 2024
0000-0002-7567-876XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2024 Heterogeneous Domain Adaptation for Multistream Classification on Cyber Threat Data
abstract
Under a newly introduced setting of multistream classification, two data streams are involved, which are referred to as source and target streams. The source stream continuously generates data instances from a certain domain with labels, while the target stream does the same task without labels from another domain. Existing approaches assume that domains for both data streams are identical, which is not quite true, since data streams from different sources may contain distinct features. Indeed, they may even have different numbers of features. Furthermore, obtaining labels for every instance in a data stream is often expensive and time-consuming. Therefore, it has become an important topic to explore if classes of labeled instances from other related streams are helpful to predict the classes of unlabeled instances in a different stream. Note that domains of source and target streams may have distinct feature spaces and data distributions. Our objective is to predict class labels of data instances in the target stream by using the classifiers trained by the source stream. We propose a framework of multistream classification by using projected data from a common latent feature space, which is embedded from both source and target domains. This framework is also crucial for enterprise system defenders to detect cross-platform attacks, such as Advanced Persistent Threats (APTs). Empirical valuation and analysis on both real-world and synthetic datasets are performed to validate the effectiveness of our proposed algorithm, comparing to state-of-the-art techniques. Experimental results show that our approach significantly outperforms other existing approaches.
Yifan Li 0003, Yang Gao 0027, Gbadebo Ayoade, Latifur Khan, Anoop Singhal, Bhavani Thuraisingham
IEEE Trans. Dependable Secur. Comput.3
2023 Advanced Persistent Threat Detection Using Data Provenance and Metric Learning
abstract
Advanced persistent threats (APT) have increased in recent times as a result of the rise in interest by nation-states and sophisticated corporations to obtain high-profile information. Typically, APT attacks are more challenging to detect since they leverage zero-day attacks and common benign tools. Furthermore, these attack campaigns are often prolonged to evade detection. We leverage an approach that uses a provenance graph to obtain execution traces of host nodes in order to detect anomalous behavior. By using the provenance graph, we extract features that are then used to train an online adaptive metric learning. Online metric learning is a deep learning method that learns a function to minimize the separation between similar classes and maximizes the separation between dis- similar instances. We compare our approach with baseline models and we show our method outperforms the baseline models by increasing detection accuracy on average by 11.3% and increases True positive rate (TPR) on average by 18.3%. We also show that our method outperforms several state-of-the-art models performances in comprehensive attack datasets in both binary and multi-class settings.
Khandakar Ashrafi Akbar, Yigong Wang, Gbadebo Ayoade, Yang Gao 0027, Anoop Singhal, Latifur Khan, Bhavani Thuraisingham, Kangkook Jee
IEEE Trans. Dependable Secur. Comput.3
2021 Crook-sourced intrusion detection as a service
Frederico Araujo, Gbadebo Ayoade, Khaled Al-Naami, Yang Gao 0027, Kevin W. Hamlen, Latifur Khan
J. Inf. Secur. Appl.2
2019 Improving intrusion detectors by crook-sourcing
abstract
Conventional cyber defenses typically respond to detected attacks by rejecting them as quickly and decisively as possible; but aborted attacks are missed learning opportunities for intrusion detection. A method of reimagining cyber attacks as free sources of live training data for machine learning-based intrusion detection systems (IDSes) is proposed and evaluated. Rather than aborting attacks against legitimate services, adversarial interactions are selectively prolonged to maximize the defender's harvest of useful threat intelligence. Enhancing web services with deceptive attack-responses in this way is shown to be a powerful and practical strategy for improved detection, addressing several perennial challenges for machine learning-based IDS in the literature, including scarcity of training data, the high labeling burden for (semi-)supervised learning, encryption opacity, and concept differences between honeypot attacks and those against genuine services. By reconceptualizing software security patches as feature extraction engines, the approach conscripts attackers as free penetration testers, and coordinates multiple levels of the software stack to achieve fast, automatic, and accurate labeling of live web data streams.
Frederico Araujo, Gbadebo Ayoade, Khaled Al-Naami, Yang Gao 0027, Kevin W. Hamlen, Latifur Khan
ACSAC2
2019 Multistream Classification for Cyber Threat Data with Heterogeneous Feature Space
abstract
Under a newly introduced setting of multistream classification, two data streams are involved, which are referred to as source and target streams. The source stream continuously generates data instances from a certain domain with labels, while the target stream does the same task without labels from another domain. Existing approaches assume that domains for both data streams are identical, which is not quite true in real world scenario, since data streams from different sources may contain distinct features. Furthermore, obtaining labels for every instance in a data stream is often expensive and time-consuming. Therefore, it has become an important topic to explore whether labeled instances from other related streams can be helpful to predict those unlabeled instances in a given stream. Note that domains of source and target streams may have distinct features spaces and data distributions. Our objective is to predict class labels of data instances in the target stream by using the classifiers trained by the source stream.
Yifan Li 0003, Yang Gao 0027, Gbadebo Ayoade, Hemeng Tao, Latifur Khan, Bhavani Thuraisingham
WWW3
2019 Secure data processing for IoT middleware systems
Gbadebo Ayoade, Amir El-Ghamry, Vishal Karande, Latifur Khan, Mohammed F. Alrahmawy, Magdi Zakria Rashad
J. Supercomput.1
2017 Unsupervised deep embedding for novel class detection over data stream
abstract
Data streams are continuous flows of data points. Novel class detection is an important part of data stream mining. A novel class is a newly emerged class that has not previously been modeled by the classifier over the input stream. This paper proposes deep embedding for novel class detection - a novel approach that combines feature learning using denoising autoencoding with novel class detection. A denoising autoencoder is a neural network with hidden layers aiming to reconstruct the input vector from a corrupted version. A nonparametric multidimensional change point detection approach is also proposed, to detect concept-drift (the change of data feature values over time). Experiments on several real datasets show that the approach significantly improves the performance of novel class detection.
Ahmad Mustafa 0001, Gbadebo Ayoade, Khaled Al-Naami, Latifur Khan, Kevin W. Hamlen, Bhavani Thuraisingham, Frederico Araujo
IEEE BigData2