Nur Zincir-Heywood

dblp:303/3845 · also Ayse Nur Zincir-Heywood · DBLP profile ↗
← Back
148ranked-venue papers
8as first author
32since 2021 · last 2025
0000-0003-2796-7265ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 1 since 2021Security and privacy · 26 · 4 since 2021Computer networks · 22 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 3 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 4
YearPublicationVenuePosition
2025 IoT Botnet Detection with Drift-Aligned Learning and DNS-Based C2 Identification
abstract
Botnet detection in Internet of Things (IoT) environments remain challenging due to evolving attack behaviors and limited generalizability of traditional detection methods. This paper introduces a lightweight, two-stage detection framework that combines a Random Forest classifier with a Botnet Validation Layer. The proposed system leverages a Drift-Aligned Learning Curve (DALC) to adapt to data drifts by incorporating Domain Name System (DNS) based analysis for botnet validation particularly in Command & Control (C2) activities. We utilize a feature set comprising of 20 flow level and 4 DNS level features. Evaluated on diverse datasets, the proposed system achieves 99.7% F1-Score during testing while demonstrating strong generalization.
Jeffrey Adjei, Nur Zincir-Heywood, Malcolm I. Heywood, Biswajit Nandy, Nabil Seddigh
CNSM2
2025 Unsupervised Anomaly Detection for Wi-Fi Networks using RFFI
Samer Lahoud, Nur Zincir-Heywood
CNSM3
2025 SQL-GENIE: SQL Protection using GENerative Modeling for Anomaly Detection against Injection and Evolved Adversarial Attacks
abstract
In an age where data drives innovation and online interactions are integral to daily life, ensuring the security of web applications and databases has never been more critical. The growing surge and sophistication of large-scale SQL injection (SQLi) attacks highlight the urgent need for advanced detection mechanisms to protect sensitive information, especially in cloud-based environments. This paper presents SQL-GENIE, a novel approach that leverages generative modeling to strengthen modern application security, improve anomaly detection, and address emerging challenges in data protection. SQL-GENIE leverages two feature embedding techniques across two different datasets and contrasts their performance against Generative Adversarial Networks (GAN)— under various contamination rates to analyze and detect SQLi attacks, including typical and sophisticated adversarial forms. Our proposed GAN model performs the best with FastText when applied to our benchmark dataset of typical SQLI, achieving F1-score of 92.7% on attack data with a 10% contamination rate. Additionally, it demonstrates an F1-score of 98.6% on the adversarial dataset, highlighting its robustness against evolved SQLi threats.
Marwa Elsayed, Nur Zincir-Heywood
COMPSAC3
2025 Lightweight Early-Warning Bot Detection on X (Twitter): Temporal Patterns and Entropy Insights
abstract
In the evolving landscape of social media, distinguishing between bots (automated accounts) and real human users remains a significant challenge. This research aims to address this challenge by focusing on the behavioral patterns of users on X (formerly Twitter). We propose an early-warning system based on representing user tweeting behavior as binary sequences. Thus, the approach does not require content analysis, enabling a lightweight yet effective solution. We evaluated the proposed approach on a wide spectrum of datasets, with promising results. This method offers new insights that could help to improve bot-detection tools on social media platforms such as X.
Sanaz Adel Alipour, Jeannette Janssen, Rita Orji, Nur Zincir-Heywood
COMPSAC4
2025 Network Identity Management: Application, Action and Device Aware Monitoring
abstract
Instant Message Applications (IMAs) are often used in sensitive environments, where they pose a significant risk of accidental policy breaches or security lapses. To address these risks, this paper introduces Network Identity Management (NIM), a set of techniques specifically designed for such high-risk applications. NIM provides a novel approach that enables network-layer, identity-aware access control for encrypted applications. To support NIM, we introduce techniques to establish the essential visibility layer. These techniques use Machine Learning (ML) models to analyze encrypted traffic metadata. The ML models are trained on data produced by our scalable, cloud-native Android traffic generation framework. Our ML framework accurately identifies: (1) the application in use, (2) the specific user action being performed, and (3) the originating device—using only encrypted traffic metadata from multi-user environments. For its primary task, our method, using emulated traffic from eight IMAs, yielded a 98.6% F1-score for application identification with Gradient Boosting. Additionally, we demonstrated near-perfect device identification based exclusively on encrypted traffic metadata as well as supplementary evaluations providing initial validation for user action classification (group vs. one-on-one messaging) with promising performance.
Cenab Batu Bora, Julia Silva Weber, Nur Zincir-Heywood
COMPSAC3
2025 Can Flow Metadata Based Signatures Generalize for Identifying Attacks on IoT Devices?
abstract
In this research, we investigate the impact of four prevalent types of attacks, namely Portscan, Slowloris, Synflood, and Vulnerability Scan, on nine distinct Internet of Things (IoT) devices. These attacks are very common on the IoT eco-systems because they often serve as precursors to more sophisticated attack vectors. By analyzing attack vector traffic characteristics and IoT device responses, we aim to shed light on IoT eco-system vulnerabilities. To achieve this, we utilize and evaluate two feature sets extracted from the network traffic metadata using a flow analyzer, avoiding deep packet inspection. The goal of this research is to evaluate the impact of traffic flow metadata for identifying attacks on IoT devices. We further analyze the two flow feature sets in terms of generalizability of machine learning based attack detection from one IoT network to another. Results show that while generalizability is possible, it also depends on several factors including the characteristics of Iot traffic.
Jeffrey Adjei, Nur Zincir-Heywood, Malcolm I. Heywood, Biswajit Nandy, Nabil Seddigh
NOMS2
2025 Identifying Synchronous and Asynchronous Communications in IMA Traffic
abstract
Instant messaging applications (IMAs) rely on message encryption to preserve privacy and security, which makes network traffic inspection difficult. For better network monitoring and analysis, we explore network traffic signatures in IMAs. We develop a framework to automatically generate and capture traffic from seven popular IMAs. We discover patterns from text messaging behavior, such as synchronous and asynchronous as well as group and private communications. We analyze the resulting end-to-end encrypted traffic using a machine learning-based approach to traffic metadata without using deep packet inspection. The evaluations show that it is possible to identify between groups with different numbers of users for asynchronous and synchronous communication models. This, in turn, can help better plan and manage network operations for better quality of service.
Srivathsan Thirumurugan, Julia Silva Weber, Riyad Alshammari, Nur Zincir-Heywood, Biswajit Nandy, Nabil Seddigh
NOMS4
2025 Encrypted Network Traffic Analysis (ENTA) Platform: IMA VoIP Traffic Identification
abstract
This paper introduces the Encrypted Network Traffic Analysis (ENTA) platform, a scalable AI-driven system de-signed for traffic analysis with support for identifying encrypted VoIP traffic generated by instant messaging applications (IMAs). User behaviors, such as exchanging audio messages through different IMAs, are emulated, and the resulting network traffic is captured for analysis. The ENTA platform's capabilities are demonstrated in feature extraction, data pre-processing, AI model training and testing, and generating key performance indi-cators (KPIs) to assist network operations teams. This demonstration showcases the ENTA system's end-to-end functionality with a particular emphasis on accurately classifying and identifying IMA VoIP traffic.
Julia Silva Weber, Srivathsan Thirumurugan, Riyad Alshammari, Nur Zincir-Heywood, Manjinder Nir, Don Bennett, Delfin Y. Montuno, Biswajit Nandy, Nabil Seddigh
NOMS4
2024 IoT Device and State Identification based on Usage Patterns
abstract
In this paper, we explore usage patterns for the identification of IoT devices and their corresponding states. Machine Learning (ML) methods are trained on IoT device traffic patterns to recognize the state that the device is in. Three device states are the focus of this study - Power-up, Idle and Active. Devices are visible and open to cyber attacks from the moment they are powered on. Previous studies have focused primarily on identifying IoT devices which are in the active state. This study advances the research domain by exploring all three states of an IoT device. Eight different ML algorithms are evaluated using three different feature sets extracted from device network traffic, using flow analysis tools - Tranalyzer2, NFStream and Zeek. They are rigorously assessed to accurately identify diverse IoT devices under normal operational conditions over the aforementioned three states. .
Jeffrey Adjei, Nur Zincir-Heywood, Biswajit Nandy, Nabil Seddigh
CNSM2
2024 Evaluating the Robustness of ADVENT on the VeReMi-Extension Dataset
abstract
In this paper, we extend and evaluate the effectiveness of ADVENT (Attack/Anomaly Detection in VANETs), a machine learning-based system designed for early attack detection and malicious node identification in Vehicular Ad Hoc Networks (VANETs). ADVENT combines machine learning with federated learning to detect the onset of attacks while preserving user privacy. The system detects and reports malicious nodes to neighboring vehicles, allowing proactive defense against attacks. We focus on its robustness against various Distributed Denial-of-Service (DDoS) attacks. Using the Vehicular Reference Misbehavior Extension (VeReMi-Extension) dataset, we assess ADVENT across five distinct types of (D)DoS attacks, each representing diverse attack characteristics. Based on our findings, we enhance ADVENT by refining its malicious node detection step through a time slicing mechanism, improving both False Positive Rate (FPR) and F1-score metrics. Our evaluation shows that ADVENT consistently excels in detecting attack onsets and identifying attackers, even under different attack types. The results emphasize its adaptability and effectiveness in strengthening VANET security.
Hamideh Baharlouei, Adetokunbo Makanju, Nur Zincir-Heywood
CNSM3
2024 Improving Real-Time Anomaly Detection using Multiple Instances of Micro-Cluster Detection
abstract
Analysis of incoming packets in deployed systems is one of the main methods used for detection of anomalous behaviour. Techniques utilizing supervised learning subject to the need of retraining if the observed behaviour in the system changes over time. Unsupervised techniques mitigate this problem but are not always capable of real-time analysis. Real-time unsupervised techniques bring to the table both the adaptability to dynamic behaviour as well as the ability to detect and alert about anomalies in real-time. A recent state-of-the-art technique, MIDAS, shows real-time capabilities while being unsupervised, but recent works have showed that it still had some shortcomings regarding its performance over more specific datasets. An alternative method has been proposed, namely MIMC, that builds on the foundation set by MIDAS. In this paper it is shown that, for the datasets of interest, there is always a way to setup MIMC that yields a higher performance than MIDAS. Furthermore, a method for determining parameters for the technique is also presented, and it is shown that it improves the yielded performance even further in a majority of cases.
Rafael Copstein, Nur Zincir-Heywood, Malcolm I. Heywood
CNSM2
2024 Identifying IoT Devices: A Machine Learning Analysis Using Traffic Flow Metadata
abstract
Deployment of IoT (Internet of Things) devices in homes, corporate networks and industrial settings continues to rise. The weak security in many such devices leaves them susceptible to targeted attacks. As a result, there is a strong requirement to easily discover and identify such devices, even when they utilize encrypted communications. In this paper, we explore a machine learning (ML) based approach for identification of IoT devices. We contribute to this area of active research by developing a testbed consisting of nine IoT devices and create a labeled dataset which is publicly available. We further investigate eight different ML algorithms with features derived from two different open source network traffic flow analysis tools - Tranalyzer2 and NFStream. Three different flow metadata based feature sets were evaluated to determine their efficacy in identifying the different IoT devices during regular operation. The results show that the Random Forest ML model achieves the highest performance using the Tranalyzer2 feature set with a 99.68% (approximately 100%) F1-Score for identifying IoT devices.
Jeffrey Adjei, Nur Zincir-Heywood, Biswajit Nandy, Nabil Seddigh
NOMS2
2024 Guest Editorial: Special section on Networks, Systems, and Services Operations and Management Through Intelligence
abstract
Machine Learning (ML) and Artificial Intelligence (AI) can harness the immense amount of operational data from clouds to services, to social and communication networks. In the era of data science and connected devices of all varieties, Intelligence have found ways to improve operations and management of next generation networks, systems, and services. Further research is therefore needed to understand and improve the potential and suitability of ML/AI in the context of network, system, and service operations and management. This will provide deeper understanding and better decision making based on largely collected and available operational and management data. It will also present opportunities for improving ML/AI algorithms on aspects such as reliability, dependability, and scalability, as well as demonstrate the benefits of these methods in control and management systems. Moreover, there is an opportunity to define novel platforms that can harness the vast operational data and advance ML/AI algorithms to drive management decisions in open and highly programmable networks, clouds, and data centers.
Nur Zincir-Heywood, Robert Birke, Elias Bou-Harb, Takeru Inoue, Neeraj Kumar 0001, Hanan Lutfiyya, Deepak Puthal, Abdallah Shami, Natalia Stakhanova
IEEE Trans. Netw. Serv. Manag.1
2023 MIMC: Anomaly Detection in Network Data via Multiple Instances of Micro-Cluster Detection
abstract
This paper proposes and explores new attribute correlations and combined effort of multiple instances of microcluster-based anomaly detection on port scans, distributed denial of service and botnet attacks. To this end, the proposed system for micro-clustering based anomaly detection is compared against the state-of-the-art technique on three different network datasets, namely CTU-IoT, CTU-13 and UNSW-NB15. Evaluations not only show the effectiveness and high performance of the proposed system on all three datasets but also demonstrate the generalizability of the newly proposed attribute correlations and combination strategies.
Rafael Copstein, Brad Niblett, Andrew Johnston, Jeff Schwartzentruber, Malcolm I. Heywood, Nur Zincir-Heywood
CNSM6
2023 Preliminary Results on Exploring Data Exhaust of Consumer Internet of Things Devices
abstract
In this paper, we apply a machine learning classifier to the publicly available consumer Internet of Things (IoT) traffic traces to explore the nature and extent of any potential data exhaust. To this end, we propose two feature sets and compare them against the baseline flow feature set and the results from the previous works. Evaluations show the improvement in performance obtained using the proposed feature sets and the variety of information that can be extracted from the captured IoT traffic regardless of encryption.
Alexander Loginov, Jeffrey Adjei, Nur Zincir-Heywood, Srinivas Sampalli, Kevin de Snayer, Terri Dougall
CNSM3
2023 A Boosting Approach to Constructing an Ensemble Stack
Zhilei Zhou, Ziyu Qiu, Brad Niblett, Andrew Johnston, Jeff Schwartzentruber, Nur Zincir-Heywood, Malcolm I. Heywood
EuroGP6
2023 Depicting Instant Messaging Encrypted Traffic Characteristics through an Empirical Study
abstract
Instant Messaging Applications (IMAs), such as Discord and WhatsApp, have become one of the main communication tools for mobile device users. Network traffic analysis is a method of monitoring network activity to identify operational and security issues. There is limited research on network traffic analysis of IMAs on mobile devices due to the challenges of end-to-end encryption, user privacy, and dynamic port usage. In this paper, we design, develop and evaluate a framework to generate end-to-end IMA traffic on mobile devices, employ feature selection and conduct traffic analysis that can cope with encrypted traffic while identifying different IMAs. Results show a performance evaluation workbench as well as highlight the key characterictis of six popular IMAs.
Zolboo Erdenebaatar, Riyad Alshammari, Biswajit Nandy, Nabil Seddigh, Marwa Elsayed, Nur Zincir-Heywood
ICCCN6
2023 Analyzing Traffic Characteristics of Instant Messaging Applications on Android Smartphones
abstract
Instant Messaging Applications (IMAs), such as WhatsApp and Messenger, have become one of the main communication tools for smartphone users. However, there is limited research analyzing the nature of encrypted network traffic produced by IMAs. In this paper, we employ a data driven approach using machine learning classification models to analyze and identify encrypted traffic from six different IMAs. Our results show that it is possible to distinguish the behaviour of different IMAs with high F1 scores.
Zolboo Erdenebaatar, Riyad Alshammari, Nur Zincir-Heywood, Marwa Elsayed, Biswajit Nandy, Nabil Seddigh
NOMS3
2023 Instant Messaging Application Encrypted Traffic Generation System
abstract
Instant Messaging Applications (IMAs) have become the leading communication tool for smartphone users. While it is insightful for network operators and security researchers to monitor and analyze the network traffic of their organization, there is a lack of research on IMA encrypted traffic analysis. In a companion work [1], we introduced a flow-based encrypted IMA traffic analysis using a data driven approach. Given the lack of publicly available data in this area, a new encrypted IMA traffic generation system is designed and implemented to automatically generate and label encrypted IMA traffic including Discord, Facebook Messenger, Signal, Microsoft Teams, Telegram, and WhatsApp. The new system utilizes a combination of open-source tools to emulate user behavior, to capture, filter and label the resulting traffic directly on an Android device. This demonstration shows the functionality of the proposed system via data generation, capture, and analysis of the six IMAs.
Zolboo Erdenebaatar, Biswajit Nandy, Nabil Seddigh, Riyad Alshammari, Marwa Elsayed, Nur Zincir-Heywood
NOMS6
2023 On the Fence: Anomaly Detection in IoT Networks
abstract
The Internet of Things (IoT) is increasingly impacting every aspect of life, with deployment in various societal applications. This paper explores anomaly detection via novelty and outlier detection approaches for IoT networks. To this end, three unsupervised learning algorithms, namely Isolation Forest (IF), Local Outlier Factor (LOF), and One-Class Support Vector Machine (OSVM), are evaluated on three publicly available IoT datasets. The results demonstrate that when the proposed solution leverages LOF to embrace the novelty approach by considering only pure benign data for training, it achieves high performance with Fl-scores within the range of 84% to 94%.
Patrick Russell, Marwa Elsayed, Biswajit Nandy, Nabil Seddigh, Nur Zincir-Heywood
NOMS5
2023 Exploring Anomaly Detection Techniques for Enhancing VANET Availability
abstract
In VANETs the quicker an anomaly can be detected and properly classified, the faster the issue it causes can be dealt with. In this paper, we propose deploying Vehicular Edge Computing (VEC) to detect anomalies characterized by the absence of message exchange between vehicles. The detection of anomalies spread across an urban area can possibly benefit from the processing and storage capacity of the VAC. VANETs are dynamic networks, whose vehicle density varies considerably over time. VANETs components do not usually store much information, making it difficult to efficiently detect the loss of messages from multiple vehicles across the neighborhoods of a city. The article aims to investigate whether the use of VEC and conventional anomaly detection techniques benefits loss of messages detection by increasing its fault coverage capacity. To measure fault coverage, fault injection experiments were conducted using simulation.
Julia Silva Weber, Tiago Ferreto, Nur Zincir-Heywood
VTC2023-Spring3
2023 Guest Editorial: Special Section on Machine Learning and Artificial Intelligence for Managing Networks, Systems, and Services - Part II
abstract
Machine learning and artificial intelligence can harness the immense stream of operational data from clouds, to services, to social and communication networks. In the era of big data and connected devices of all varieties, machine learning and artificial intelligence have found ways to improve operations and management of information technology and communications.
Nur Zincir-Heywood, Robert Birke, Elias Bou-Harb, Giuliano Casale, Khalil El-Khatib, Takeru Inoue, Neeraj Kumar 0001, Hanan Lutfiyya, Deepak Puthal, Abdallah Shami, Natalia Stakhanova, Farhana Zulkernine
IEEE Trans. Netw. Serv. Manag.1
2022 Security of Social Networks: Lessons Learned on Twitter Bot Analysis in the Literature
abstract
Twitter is one of the popular social network platforms used by both humans and bots to share information, and to distribute misinformation or disinformation. The goal of this research is to explore state-of-the-art Twitter bot detection systems: Botometer, and Tweetbotornot. In doing so, we aim to understand their characteristics, similarities, and differences as well as to identify additional methods to improve their performances. Evaluations performed on 15 datasets show that the proposed methods used as add-ons to Botometer were able to improve the detection of bot/human accounts on Twitter using the simple characteristics of Twitter accounts.
Sanaz Adel Alipour, Rita Orji, Nur Zincir-Heywood
ARES3
2022 BoostGuard: Interpretable Misbehavior Detection in Vehicular Communication Networks
abstract
Wireless Communication and Artificial Intelligence are at the heart of driving the evolution in the transportation industry. Cooperative Intelligent Transportation Systems adopt vehicle-to-vehicle (V2V) technology to allow vehicles to exchange real-time information about speed, heading, and location wirelessly with their surrounding vehicles. Such technology has remarkable benefits for improving vehicles’ safety and awareness, albeit imposing many security risks. Despite the evolving efforts to employ authentication mechanisms, there is no guarantee that the exchanged data is trustworthy. Security breaches causing falsified data can aggressively lead to severe safety damages within vehicular networks. This paper proposes, BoostGuard, a novel interpretable framework for detecting falsified data exchanged as part of five different types of position forging attacks against vehicular networks. BoostGuard mainly adopts data science principles and leverages advanced machine learning techniques (i.e., boosting decision tree ensemble) to boost its generalization capabilities for precisely detecting and classifying attack types. Extensive experiments are conducted over an open-source dataset, reflecting dynamic real-world vehicular environments. The evaluation results demonstrate that our solution outperforms existing solutions with high detection effectiveness and computational time efficiency.
Marwa Elsayed, Nur Zincir-Heywood
NOMS2
2022 Exploring Realistic VANET Simulations for Anomaly Detection of DDoS Attacks
abstract
Simulation is widely accepted in Vehicular Ad hoc Network (VANET) research due to the cost, safety and security issues associated with real world implementations and experimentation. However, several important factors must be considered if we expect the simulation results to be realistic, comparable and extendable to the real world, especially when it comes to security issues. These factors can largely be classed under three broad categories i.e. the Grid Pattern, the Communication Settings and the Mobility Pattern. Building on prior work, in this paper, we extend the simulation results of a VANET-based DDoS attack and an anomaly detection mechanism designed to detect the attack. We show that taken these factors into consideration leads to different results, affirming the need for considering these factors in simulations. We also discuss future research directions that result directly from our observations.
Hamideh Baharlouei, Adetokunbo Makanju, Nur Zincir-Heywood
VTC Spring3
2022 Guest Editorial: Special Issue on Machine Learning and Artificial Intelligence for Managing Networks, Systems, and Services - Part I
abstract
Machine learning and artificial intelligence can harness the immense stream of operational data from clouds, to services, to social and communication networks. In the era of big data and connected devices of all varieties, machine learning and artificial intelligence have found ways to improve operations and management of information technology and communications.
Nur Zincir-Heywood, Robert Birke, Elias Bou-Harb, Giuliano Casale, Khalil El-Khatib, Takeru Inoue, Neeraj Kumar 0001, Hanan Lutfiyya, Deepak Puthal, Abdallah Shami, Natalia Stakhanova, Farhana Zulkernine
IEEE Trans. Netw. Serv. Manag.1
2021 Log Abstraction for Information Security: Heuristics and Reproducibility
abstract
The collection of log messages regarding the operation of deployed services and application is an integral component to the forensic analysis for the identification and understanding of security incidents. Approaches for parsing and abstraction of such logs, despite widespread use and study, do not directly account for the individualities of the domain of information security. This, in return, limits their applicability on the field. In this work, we analyze the state-of-the-art log parsing and abstraction algorithms from the perspective of information security. First, we reproduce/replicate previous analysis of such algorithms from the literature. Then, we evaluate their ability for parsing and abstraction of log files for forensic analysis purposes. Our study demonstrates that while the state-of-the-art techniques are accurate in log parsing, improvements are necessary in terms of achieving a holistic view to aid in forensic analysis for the identification and understanding of security incidents.
Rafael Copstein, Jeff Schwartzentruber, Nur Zincir-Heywood, Malcolm I. Heywood
ARES3
2021 Network Flow Entropy for Identifying Malicious Behaviours in DNS Tunnels
abstract
In this paper, we propose the concept of ”entropy of a flow” to augment flow statistical features for identifying malicious behaviours in DNS tunnels, specifically DNS over HTTPS traffic. In order to achieve this, we explore the use of three flow exporters, namely Argus, DoHlyzer and Tranalyzer2 to extract flow statistical features. We then augment these features using different ways of calculating the entropy of a flow. To this end, we investigate three entropy calculation approaches: Entropy over all packets of a flow, Entropy over the first 96 bytes of a flow, and Entropy over the first n-packets of a flow. We evaluate five machine learning classifiers, namely Decision Tree, Random Forest, Logistic Regression, Support Vector Machine and Naive Bayes using these features in order to identify malicious behaviours in different publicly available datasets. The evaluations show that the Decision Tree classifier achieves an F-measure of 99.7% when flow statistical features are augmented with entropy of a flow calculated over the first 4 packets.
Yulduz Khodjaeva, Nur Zincir-Heywood
ARES2
2021 Modelling and visualising SSH brute force attack behaviours through a hybrid learning framework
Xiao Luo 0002, Chengchao Yao, Nur Zincir-Heywood
Int. J. Inf. Comput. Secur.3
2021 Anomaly Detection for Insider Threats Using Unsupervised Ensembles
abstract
Insider threat represents a major cybersecurity challenge to companies, organizations, and government agencies. Insider threat detection involves many challenges, including unbalanced data, limited ground truth, and possible user behavior changes. This research presents an unsupervised learning based anomaly detection approach for insider threat detection. We employ four unsupervised learning methods with different working principles, and explore various representations of data with temporal information. Furthermore, different computational intelligence schemes are explored to combine these models to create anomaly detection ensembles for improving the detection performance. Evaluation results show that the approach allows learning from unlabelled data under challenging conditions for insider threat detection. Insider threats are detected with high detection and low false positive rates. For example, 60% of malicious insiders are detected under 0.1% investigation budget, and all malicious insiders are detected at less than 5% investigation budget. Furthermore, we explore the ability of the proposed approach to generalize for detecting new anomalous behaviors in different datasets, i.e., robustness. Finally, results demonstrate that a voting-based ensemble of anomaly detection can be used to improve detection performance as well as the robustness. Comparisons with the state-of-the-art confirm the effectiveness of the proposed approach.
Duc C. Le, Nur Zincir-Heywood
IEEE Trans. Netw. Serv. Manag.2
2021 Guest Editorial: Special Section on Embracing Artificial Intelligence for Network and Service Management
abstract
Artificial Intelligence (AI) has the potential to leverage the immense amount of operational data of clouds, services, and social and communication networks. As a concrete example, AI techniques have been adopted by telcom operators to develop virtual assistants based on advances in natural language processing (NLP) for interaction with customers and machine learning (ML) to enhance the customer experience by improving customer flow. Machine learning has also been applied to finding fraud patterns which enables operators to focus on dealing with the activity as opposed to the previous focus on detecting fraud.
Hanan Lutfiyya, Robert Birke, Giuliano Casale, Amogh Dhamdhere, Jinho Hwang, Takeru Inoue, Neeraj Kumar 0001, Deepak Puthal, Nur Zincir-Heywood
IEEE Trans. Netw. Serv. Manag.9
2021 Guest Editorial: Special Issue on Data Analytics and Machine Learning for Network and Service Management - Part II
abstract
Network and Service analytics can harness the immense stream of operational data from clouds, to services, to social and communication networks. In the era of big data and connected devices of all varieties, analytics and machine learning have found ways to improve reliability, configuration, performance, fault and security management. In particular, we see a growing trend towards using machine learning, artificial intelligence and data analytics to improve operations and management of information technology services, systems and networks.
Nur Zincir-Heywood, Giuliano Casale, David Carrera 0001, Lydia Y. Chen, Amogh Dhamdhere, Takeru Inoue, Hanan Lutfiyya, Taghrid Samak
IEEE Trans. Netw. Serv. Manag.1
2020 Exploring data leakage in encrypted payload using supervised machine learning
abstract
Data security includes but not limited to, data encryption and key management practices that protect data across all applications and platforms. In this paper, we aim to explore whether any data leakage takes place in data encryption when encrypted data is analyzed using supervised machine learning techniques. To this end, we analyze four encryption algorithms with different key sizes using five supervised learning techniques on two different datasets. The results show that as the encryption algorithms get stronger, the data leakage decreases, even though the data leakage is never zero percent.
Amir Khaleghi Moghaddam, Nur Zincir-Heywood
ARES2
2020 Temporal Representations for Detecting BGP Blackjack Attacks
abstract
Even though BGP blackholes are used to mitigate denial of service attacks, they also represent a major cybersecurity challenge to organizations. These challenges include abuse of route selection algorithms, lack of host verification, and maliciously triggering a blackhole, i.e. BGP blackjack. This research presents a supervised machine learning based approach for blackjack detection. We employ Naive Bayes and Decision Tree classifiers with three different temporal representations: (i) packets with/without timestamps; (ii) buffer of packets with/without timestamps; and (iii) overlapping / non-overlapping buffer of packets with/without timestamps. Our goal is to understand the effect of temporal data and context in the detection of blackjack attacks. Furthermore, we explore the most suitable attributes and solution complexity. Evaluations show that using overlapping buffer data with times-tamps achieves the highest accuracy/recall using five of the seven BGP attributes. We also observe that high performance is not correlated with complex solutions.
Rafael Copstein, Nur Zincir-Heywood
CNSM2
2020 COUGAR: clustering of unknown malware using genetic algorithm routines
abstract
Through malware, cyber criminals can leverage our computing resources to disrupt our work, steal our information, and even hold it hostage. Security professionals seek to classify these malicious software so as to prevent their distribution and execution, but the sheer volume of malware complicates these efforts. In response, machine learning algorithms are actively employed to alleviate the workload. One such approach is evolutionary computation, where solutions are bred, rather than built. In this paper, we design, develop and evaluate a system, COUGAR, to reduce high-dimensional malware behavioural data, and optimize clustering behaviour using a multi-objective genetic algorithm. Evaluations demonstrate that each of our chosen clustering algorithms can successfully highlight groups of malware. We also present an example real-world scenario, based on the testing data, to demonstrate practical applications.
Zachary Wilkins, Nur Zincir-Heywood
GECCO2
2020 Analyzing Data Granularity Levels for Insider Threat Detection Using Machine Learning
abstract
Malicious insider attacks represent one of the most damaging threats to networked systems of companies and government agencies. There is a unique set of challenges that come with insider threat detection in terms of hugely unbalanced data, limited ground truth, as well as behaviour drifts and shifts. This work proposes and evaluates a machine learning based system for user-centered insider threat detection. Using machine learning, analysis of data is performed on multiple levels of granularity under realistic conditions for identifying not only malicious behaviours, but also malicious insiders. Detailed analysis of popular insider threat scenarios with different performance measures are presented to facilitate the realistic estimation of system performance. Evaluation results show that the machine learning based detection system can learn from limited ground truth and detect new malicious insiders in unseen data with a high accuracy. Specifically, up to 85% of malicious insiders are detected at only 0.78% false positive rate. The system is also able to quickly detect the malicious behaviours, as low as 14 minutes after the first malicious action. Comprehensive result reporting allows the system to provide valuable insights to analysts in investigating insider threat cases.
Duc C. Le, Nur Zincir-Heywood, Malcolm I. Heywood
IEEE Trans. Netw. Serv. Manag.2
2020 Guest Editorial: Special Section on Data Analytics and Machine Learning for Network and Service Management-Part I
Nur Zincir-Heywood, Giuliano Casale, David Carrera 0001, Lydia Y. Chen, Amogh Dhamdhere, Takeru Inoue, Hanan Lutfiyya, Taghrid Samak
IEEE Trans. Netw. Serv. Manag.1
2019 Are There Bots even in FIFA World Cup 2018 Tweets?
abstract
Social media is an important communication medium in these days. Twitter is famous as microblogging service. It has been reported that social bots have been used in Twitter widely. In this research, we aim to understand whether bots are selective in the topics they participate or not. To this end, we explore tweets on FIFA World Cup. Our analysis indicate that there are bot activities even in tweets related to soccer (football) events but not just political topics.
Moath Bagarish, Riyad Alshammari, Nur Zincir-Heywood
CNSM3
2019 Exploring Feature Normalization and Temporal Information for Machine Learning Based Insider Threat Detection
abstract
Insider threat is one of the most damaging cyber security attacks to companies and organizations. In this paper, we explore different techniques to leverage spatial and temporal characteristics of user behaviours for insider threat detection. In particular, feature normalization (scaling) techniques and a scheme for representing explicit temporal information are explored to improve the performance of the machine learning based insider threat detection. The results show that these data characteristics have different effects on different classifiers, where Standard Scaler with Random Forest classifier produces the best performance.
Duc C. Le, Nur Zincir-Heywood
CNSM3
2019 Compromised Tweet Detection Using Siamese Networks and fastText Representations
abstract
The aim of this work is to detect compromised users of tweets based on their writing styles. In this paper, we use Siamese Networks to learn a representation of user tweets that allows us to classify them based on a limited amount of ground truth data. We propose the employment of this classification model to identify compromised user accounts of tweets.
Mihir Joshi, Parmeet Singh, Nur Zincir-Heywood
CNSM3
2019 Exploring NAT Detection and Host Identification Using Machine Learning
abstract
The usage of Network Address Translation (NAT) devices is common among end users, organizations, and Internet Service Providers. NAT provides anonymity for users within an organization by replacing their internal IP addresses with a single external wide area network address. While such anonymity provides an added measure of security for legitimate users, it can also be taken advantage of by malicious users hiding behind NAT devices. Thus, identifying NAT devices and hosts behind them is essential to detect malicious behaviors in traffic and application usage. In this paper, we propose a machine learning based solution to detect hosts behind NAT devices by using flow level statistics (excluding IP addresses, port numbers, and application layer information) from passive traffic measurements. We capture a large dataset and perform an extensive evaluation of our proposed approach with four existing approaches from the literature. Our results show that the proposed approach could identify NAT behaviors and hosts not only with higher accuracy but also demonstrates the impact of parameter sensitivity of the proposed approach.
Ali Safari Khatouni, Khurram Aziz, Ibrahim Zincir, Nur Zincir-Heywood
CNSM5
2019 Learning From Evolving Network Data for Dependable Botnet Detection
abstract
This work presents an emerging problem in real-world applications of machine learning (ML) in cybersecurity, particularly in botnet detection, where the dynamics and the evolution in the deployment environments may render the ML solutions inadequate. We propose an approach to tackle this challenge using Genetic Programming (GP) - an evolutionary computation based approach. Preliminary results show that GP is able to evolve pre-trained classifiers to work under evolved (expanded) feature space conditions. This indicates the potential use of such an approach for botnet detection under non-stationary environments, where much less data and training time are required to obtain a reliable classifier as new network conditions arise.
Duc C. Le, Nur Zincir-Heywood
CNSM2
2019 Network Analytics for Streaming Traffic Analysis
Sara Khanchi, Nur Zincir-Heywood, Malcolm I. Heywood
IM2
2019 Machine learning based Insider Threat Modelling and Detection
Duc C. Le, Nur Zincir-Heywood
IM2
2019 Integrating Machine Learning with Off-the-Shelf Traffic Flow Features for HTTP/HTTPS Traffic Classification
abstract
Accurate traffic classification is a key requirement for different network and security monitoring/planning tools. The evolution of Internet protocols and applications has caused traditional traffic classification approaches to be ineffective in certain cases. Key causes of the inaccuracy include: (i) the increase in the encrypted traffic; (ii) the rise in the usage of dynamic port numbers for different applications; and (iii) multiple applications running over HTTP/HTTPS protocols. Traditional solutions for traffic analysis, classification, and measurement fall short in providing visibility in users' activities - a key requirement for network and security monitoring tools. In this paper, we evaluate an automatic classifier for encrypted Social media, Video and Audio traffic without relying on particular application layer header fields that can be easily modified. We leverage machine learning algorithms together with the features provided by the well-known off-the-shelf traffic flow exporters. We evaluate the performance of such a system also for generalization (robustness) purposes on different networks. Experimental results show promising performances in terms of generating robust traffic classification on large traffic data when the trained model is moved to different networks.
Ali Safari Khatouni, Nur Zincir-Heywood
ISCC2
2019 Guest Editorial: Special Issue on Novel Techniques in Big Data Analytics for Management
abstract
Cloud and network analytics can harness the immense stream of operational data from clouds and networks, and can perform analytics processing to improve reliability, configuration, performance, fault and security management. In particular, we see a growing trend towards using statistical analysis, Artificial Intelligence (AI) and machine learning to improve operations and management of IT systems and networks.
David Carrera 0001, Giuliano Casale, Takeru Inoue, Hanan Lutfiyya, Nur Zincir-Heywood
IEEE Trans. Netw. Serv. Manag.6
2018 Benchmarking evolutionary computation approaches to insider threat detection
abstract
Insider threat detection represents a challenging problem to companies and organizations where malicious actions are performed by authorized users. This is a highly skewed data problem, where the huge class imbalance makes the adaptation of learning algorithms to the real world context very difficult. In this work, applications of genetic programming (GP) and stream active learning are evaluated for insider threat detection. Linear GP with lexicase/multi-objective selection is employed to address the problem under a stationary data assumption. Moreover, streaming GP is employed to address the problem under a non-stationary data assumption. Experiments conducted on a publicly available corporate data set show the capability of the approaches in dealing with extreme class imbalance, stream learning and adaptation to the real world context.
Duc C. Le, Sara Khanchi, Nur Zincir-Heywood, Malcolm I. Heywood
GECCO3
2018 Streaming Botnet traffic analysis using bio-inspired active learning
abstract
Non-stationary network traffic, together with stealth occurrences of malicious behaviors, make analyzing network traffic challenging. In this research, a machine learning framework is used to incrementally learn the network behavior and adapt to the changes in the traffic. This framework works under two main constraints: 1) label budget, 2) class imbalance; which makes it suitable for real-world network scenarios. Evaluations are performed on a public dataset with multiple Botnet scenarios under 0.5% and 5% label budgets; only around 2.2% of traffic is Botnet. Our results demonstrate the significance of the proposed Stream Genetic Programming solution and a general robustness to factors such as long latencies between instances of the same Botnet.
Sara Khanchi, Nur Zincir-Heywood, Malcolm I. Heywood
NOMS2
2018 How far can we push flow analysis to identify encrypted anonymity network traffic?
abstract
Anonymity networks provide privacy to the users by relaying their data to multiple destinations in order to reach the final destination anonymously. Multilayer of encryption is used to protect the users' privacy from attacks or even from the operators of the stations. In this research, we showed how flow analysis could be used to identify encrypted anonymity network traffic under four scenarios: (i) Identifying anonymity networks compared to normal background traffic; (ii) Identifying the type of applications used on the anonymity networks; (iii) Identifying traffic flow behaviors of the anonymity network users; and (iv) Identifying / profiling the users on an anonymity network based on the traffic flow behavior. In order to study these, we employ a machine learning based flow analysis approach and explore how far we can push such an approach.
Khalid Shahbar, Nur Zincir-Heywood
NOMS2
2018 A language model for compromised user analysis
abstract
Identifying compromised accounts on online social networks that are used for phishing attacks or sending spam messages is still one of the most challenging problems of cyber security. In this paper, we explore a language model that is based on artificial neural networks to differentiate the writing styles of different users on short text messages. In doing so, our aim is to be able to identify compromised user accounts. Our results indicate that we can learn the language model on one dataset and can generalize it to different datasets with approximately 85% accuracy without any modifications to the language model.
Nur Zincir-Heywood, Tien D. Phan
NOMS1
2018 Guest Editorial: Special Section on Advances in Big Data Analytics for Management
abstract
Cloud and network analytics can harness the immense stream of operational data from clouds and networks, and can perform analytics processing to improve reliability, automated configuration, performance, and optimized network management in general. In this area, we have witnessed a growing trend towards using statistical analysis and machine learning techniques to improve operations and management of IT systems and networks.
Giuliano Casale, Yixin Diao, Marco Mellia, Rajiv Ranjan 0001, Nur Zincir-Heywood
IEEE Trans. Netw. Serv. Manag.5
2017 Properties of a GP active learning framework for streaming data with class imbalance
abstract
Active learning algorithms attempt to interactively develop a subset of data from which fitness evaluation is performed. Moreover, the distribution of labeled content within the data subset may adapt over time as genetic programming (GP) individuals improve. The basic goal is therefore to identify the most meaningful subset of data to improve the current model. Under a streaming data context additional challenges exist relative to the non-streaming scenario: non-stationary processes, partial observability anytime operation. This means that it is not possible to guarantee that the content of the data subset even provides exemplars for each class that could appear in the stream (i.e., different classes appear/disappear at different parts of the stream). With this in mind, an investigation is performed into the impact of adopting different policies for controlling the development of data subset content. To do so, a generic framework is defined in terms of sampling and archiving policies. The resulting evaluation under several large multi-class datasets with class imbalance indicates that adopting random sampling with a biased archiving policy is sufficient for evolving GP classifiers that match or better the current state-of-the-art, particularly when detecting minor classes.
Sara Khanchi, Malcolm I. Heywood, Nur Zincir-Heywood
GECCO3
2017 Exploring a service-based normal behaviour profiling system for botnet detection
abstract
Effective detection of botnet traffic becomes difficult as the attackers use encrypted payload and dynamically changing port numbers (protocols) to bypass signature based detection and deep packet inspection. In this paper, we build a normal profiling-based botnet detection system using three unsupervised learning algorithms on service-based flow-based data, including self-organizing map, local outlier, and k-NN outlier factors. Evaluations on publicly available botnet data sets show that the proposed system could reach up to 91% detection rate with a false alarm rate of 5%.
Weikeng Chen, Xiao Luo 0002, Nur Zincir-Heywood
IM3
2016 On the Impact of Class Imbalance in GP Streaming Classification with Label Budgets
Sara Khanchi, Malcolm I. Heywood, Nur Zincir-Heywood
EuroGP3
2016 Autonomous system based flow marking scheme for IP-Traceback
abstract
Tracing IP packets to their sources, known as IP-Traceback, is a critical task in defending against IP spoofing and DoS attacks. There are several solutions to traceback to the origin of the attack. However, all these solutions require either all routers or ISPs to support the same IP-Traceback mechanism. To address this limitation, we propose an IP-Traceback approach at the level of autonomous systems, called Autonomous System-based Flow Marking, ASFM, to identify some key locations in the path where attacker packets are being forwarded. ASFM employs the BGP update message community attribute that enables information to be passed across ASs even if they are not necessarily involved in the IP-Traceback scheme. We also propose an authentication method, so a downstream AS can examine the correctness of the marking provided by the upstream ASs, thus eliminating the fake marking embedded by subverted routers. Finally, we evaluate and analyze the performance of our proposal, using real life datasets.
Vahid Aghaei Foroushani, Nur Zincir-Heywood
NOMS2
2016 How to choose from different botnet detection systems?
abstract
Given that botnets represent one of the most aggressive threats against cybersecurity, various detection approaches have been studied. However, whichever approach is used, the evolving nature of botnets and the required pre-defined botnet detection rule sets employed may affect the performance of detection systems. In this work, we explore the effectiveness two rule based systems and two machine learning (ML) based techniques with different feature extraction methods (packet payload based and traffic flow based). The performance of these detection systems range from 0% to 100% on thirteen public botnet data sets (i.e. CTU-13). We further analyze the performances of these systems in order to understand which type of a detection system is more effective for which type of an application.
Fariba Haddadi, Duong-Tien Phan, Nur Zincir-Heywood
NOMS3
2015 Highlights on analyzing one-way traffic using different tools
abstract
In this paper, we present our analysis using four different systems on two different one-way network traffic data sets. Specifically, we have explored the usage of two network traffic analyzers, namely Corsaro and Cisco ASA 5515-X, and two machine learning based systems, namely the C4.5 Decision Tree classifier and the AdaBoost.M1 classifier. We have employed these four systems on two publicly available one-way network data sets provided by CAIDA from 2008 and 2012. Our analysis on these systems are based on the detection rate, false alarm rate, computational cost and ease of use of these systems. To the best of our knowledge, this work is the first one performing such an analysis and evaluating machine learning based systems against well known commercial as well as open source ones on one-way network traffic data sets.
Eray Balkanli, Nur Zincir-Heywood
CISDA2
2015 Deterministic flow marking for IPv6 traceback (DFM6)
abstract
Although some security threats were taken into consideration in the IPv6 design, DDoS attacks still exist in the IPv6 networks. The main difficulty to counter the DDoS attacks is to trace the source of such attacks, as the attackers often use spoofed source IP addresses to hide their identity. This makes the IP traceback schemes very relevant to the security of the IPv6 networks. Given that most of the current IP traceback approaches are based on the IPv4, they are not suitable to be applied directly on the IPv6 networks. In this research, a modified version of the Deterministic Flow Marking (DFM) approach for the IPv6 networks, called DFM6, is presented. DFM6 embeds a fingerprint in only one packet of each flow to identify the origin of the IPv6 traffic traversing through the network. DFM6 requires only a small amount of marked packets to complete the process of traceback with high traceback rate and no false positives.
Vahid Aghaei Foroushani, Nur Zincir-Heywood
CNSM2
2015 Traffic flow analysis of tor pluggable transports
abstract
Tor provides the users the ability to use the Internet anonymously. On the Tor network, the users connect to three relays run by volunteers. The addresses of these relays are publicly available. Some organizations prevent access to Tor by blocking the addresses of these relays. To mitigate this, Tor has introduced the concept of bridges and pluggable transports. Bridges are relays that do not have publicly available addresses so that they can evade the blocking. Pluggable transports are used to obfuscate the connection to these bridges. In this paper, we investigate the robustness of these pluggable transports in evading the flow based traffic analysis and blocking systems.
Khalid Shahbar, Nur Zincir-Heywood
CNSM2
2015 Predictive Analysis on Tracking Emails for Targeted Marketing
Xiao Luo 0002, Revanth Nadanasabapathy, Nur Zincir-Heywood, Keith Gallant, Janith Peduruge
Discovery Science3
2015 Benchmarking Stream Clustering for Churn Detection in Dynamic Networks
Serdar Baran Tatar, Andrew R. McIntyre, Nur Zincir-Heywood, Malcolm I. Heywood
Discovery Science3
2015 Tapped Delay Lines for GP Streaming Data Classification with Label Budgets
Ali Vahdat, Jillian Morgan, Andrew R. McIntyre, Malcolm I. Heywood, Nur Zincir-Heywood
EuroGP5
2015 Investigating unique flow marking for tracing back DDoS attacks
abstract
In this paper, we outline the recent efforts of our research in defense against Distributed Denial of Service (DDoS) attacks. In particular, we present a novel approach to IP traceback, namely Unique Flow Marking (UFM), and we evaluate UFM against other marking schemes. Our results show that the UFM can reduce the number of marked packets compared to the other marking schemes, while achieving a better performance in terms of its ability to trace back the attack.
Vahid Aghaei Foroushani, Nur Zincir-Heywood
IM2
2015 On the Effectiveness of Different Botnet Detection Approaches
Fariba Haddadi, Duc C. Le, Laura Porter, Nur Zincir-Heywood
ISPEC4
2015 EMITS: An Experience Management System for IT Management Support
abstract
This research focuses on the identification of relevant experience required for solving IT (Information Technology) problems in small- to medium-sized enterprises. To achieve this, we integrated information retrieval techniques with clustering and optimization techniques to design and develop a custom-built Experience Management system for IT management support. We have built and evaluated our system on three different publicly available data sets: Princeton, Parallels and GoDaddy. Results support that it is possible to provide the right mix of automation and manual activity for IT experience management while achieving a high accuracy.
Can Bozdogan, Nur Zincir-Heywood, Ibrahim Zincir
Int. J. Softw. Eng. Knowl. Eng.2
2014 TDFA: Traceback-Based Defense against DDoS Flooding Attacks
abstract
Distributed Denial of Service (DDoS) attacks are one of the challenging network security problems to address. The existing defense mechanisms against DDoS attacks usually filter the attack traffic at the victim side. The problem is exacerbated when there are spoofed IP addresses in the attack packets. In this case, even if the attacking traffic can be filtered by the victim, the attacker may reach the goal of blocking the access to the victim by consuming the computing resources or by consuming a big portion of the bandwidth to the victim. This paper proposes a Trace back-based Defense against DDoS Flooding Attacks (TDFA) approach to counter this problem. TDFA consists of three main components: Detection, Trace back, and Traffic Control. In this approach, the goal is to place the packet filtering as close to the attack source as possible. In doing so, the traffic control component at the victim side aims to set up a limit on the packet forwarding rate to the victim. This mechanism effectively reduces the rate of forwarding the attack packets and therefore improves the throughput of the legitimate traffic. Our results based on real world data sets show that TDFA is effective to reduce the attack traffic and to defend the quality of service for the legitimate traffic.
Vahid Aghaei Foroushani, Nur Zincir-Heywood
AINA2
2014 Supervised learning to detect DDoS attacks
abstract
In this research, we explore the performances of two supervised learning techniques and two open-source network intrusion detection systems (NIDS) on backscatter darknet traffic. We employ Bro and Corsaro open-source systems as well as the CART Decision Tree and Naive Bayes machine learning classifiers. While designing our machine learning classifiers, we used different sizes of training/test sets and different feature sets to understand the importance of data pre-processing. Our results show that a machine learning base approach can achieve very high performance on such backscatter darknet traffic without using IP addresses and port numbers.
Eray Balkanli, Jander Alves, Nur Zincir-Heywood
CICS3
2014 Benchmarking two techniques for Tor classification: Flow level and circuit level classification
abstract
Recently, many Internet users, who seek anonymity, use Tor, which is one of the most popular anonymity software solutions. Tor provides this anonymity by hiding the identity of the user from the destination that the user aims to reach. It also hides the user activities into encrypted cells. In this work, we investigate up to what level we can define what the user in Tor is doing. To this end, we extended on the previous work to classify the user activities using information extracted from Tor circuits and cells. Moreover, we developed a classification system to identify user activities based on traffic flow features. Our results show that flow based classification can reach up to the accuracy of the cell level classification as well as being more flexible.
Khalid Shahbar, Nur Zincir-Heywood
CICS2
2014 A case study for a secure and robust geo-fencing and access control framework
abstract
The growing prevalence of Smartphones and Tablets has introduced new challenges and a heterogeneous hardware and software ecosystem. One of the interesting aspects of this phenomenon is how to control the user owned devices' access to organization/retail resources. This research focuses on enhancing a proposed indoor geo-fencing and access control framework aiming for retail environments as a case study. We investigate various techniques to improve robustness and security of such a system. The focus of these improvements is to build a system that is able to operate properly in noisy, heterogeneous and less controlled environments where the presence of attackers is a high probability. As a result statistical measures have been introduced that improve the system's robustness and positioning accuracy along with mechanisms that effectively detect and prevent domain specific attacks.
Hossein Rahimi, Tuerxun Maimaiti, Nur Zincir-Heywood
NOMS3
2013 Deterministic and Authenticated Flow Marking for IP Traceback
abstract
In this paper, we present a novel approach to IP trace back - Deterministic Flow Marking (DFM) - which allows the victim to trace back the origin of incorrect or spoofed source addresses up to the attacker node, even if the attack has been originated from a network behind a NAT or a proxy server. DFM is scalable and simple to implement, it is capable of tracing thousands of simultaneous distributed attacks in near real time. Moreover, it has a small footprint, resulting in low processing and memory overhead at the victim machines and edge routers. Additionally, DFM provides an optional authentication, so that a compromised router cannot forge markings of other uncompromised routers. Our results show that DFM can reach to ~99% trace back rate with no false positives.
Vahid Aghaei Foroushani, Nur Zincir-Heywood
AINA2
2013 Analyzing string format-based classifiers for botnet detection: GP and SVM
abstract
The domain name system (DNS) is an essential component of Internet. As it is expected to be used by all legitimate users and applications, generally there are less inspections, restrictions and filters on it. Botnets rely on this open component to accomplish their malicious operation. Therefore, to defeat the single point of failure and evade static blacklists and firewalls, they employ DNS-based methods to frequently generate new automatic domain names. Stateful-SBB, which is a form of genetic programming (GP), was previously designed and developed by the authors to detect these automatically generated domain names based on minimum a priori information which was shown efficient. In this paper, we compare Stateful-SBB against the String Subsequence Kernel (SSK) and SSK with Lambda Pruning (SSK-LP), which are based on support vector machines (SVM) and also use string format inputs. Analyzing the domain names that each of the classifiers chooses as a part of their solutions in the classification process, we notice that 50% to 63% of the Stateful-SBBs' frequently selected points on the Pareto-front are also used by SSK and SSK-LP, respectively. By analyzing these common domain names, we identify some of the characteristics of the botnet domain names. Moreover, we introduce a pruned version of the Stateful-SBB that resulted in reducing the solution complexity by 83% with the same high accuracy.
Fariba Haddadi, Nur Zincir-Heywood
IEEE Congress on Evolutionary Computation2
2013 How far an evolutionary approach can go for protocol state analysis and discovery
abstract
Securing todays computer networks requires numerous technologies to constantly be developed, refined and challenged. One area of research aiding in this process is that of protocol analysis, the study of the methods with which networks communicate. Our specific area of interest, the interaction with different protocol implementations, is a crucial component of this domain. Our work aims to identify and highlight a protocols states and state transitions, while minimizing the required a priori knowledge known about the protocol and its different versions (implementations). To this end, our approach uses a Genetic Programming (GP) based technique in order to analyze a client or a server of a given protocol via interacting with it with minimum a priori information. We evaluate our system against another well-known system from the literature on two different protocols, namely Dynamic Host Configuration Protocol (DHCP) and File Transfer Protocol (FTP). We measure the performances of these two systems in terms of the similarities and differences seen in the state diagrams produced for the protocols under testing. Results show that, by using our approach, it is possible to identify the different versions of a given protocol.
Patrick LaRoche, Aimee Burrows, Nur Zincir-Heywood
IEEE Congress on Evolutionary Computation3
2013 Indoor geo-fencing and access control for wireless networks
abstract
Having an idea of a user's location when he/she is using network services has been an area of interest ever since wireless networks became very popular. As the costs of wireless technologies decrease more and more, we observe the rise of an extremely diverse market of wireless capable devices. However, the field of indoor positioning is still wide open. In this field, most of the existing technologies are dependent on additional hardware and/or infrastructure, which increases the requirements for users. In this research, we investigate the ways of coupling indoor geo-fencing with access control including authentication and registration. To achieve this, we apply a classification based geo-fencing approach using received signal strength indicator. Consequently, we are mainly focusing on associating accurate geo-fencing with secure communication and computing. Experimental results show that we have achieved considerable positioning accuracy while providing a secure way of communication. Favouring diversity, our implementation does not mandate users to undergo any system software modification or adding new hardware components.
Hossein Rahimi, Nur Zincir-Heywood, Bharat Gadher
CICS2
2013 Investigating application behavior in network traffic traces
abstract
Identifying encrypted application traffic is an important issue for many network tasks including quality of service, firewall enforcement and security. This paper presents a machine learning based approach to identify high level application behavior in a given traffic trace using a holistic approach without looking into the content or without checking a static attribute. We demonstrate the effectiveness of our approach as a forensic analysis tool on five encrypted applications namely SSH, Skype, Gtalk, SSL (No Web) and HTTPS (Web Browsing), using traces captured from different networks. Results indicate that it is possible to identify high level application behavior such as unencrypted versus encrypted as well as identifying services running in encrypted tunnels.
Vahid Aghaei Foroushani, Nur Zincir-Heywood
CISDA2
2013 Beyond term clusters: assigning Wikipedia concepts to scientific documents
abstract
We propose a model for assigning Wikipedia Concepts as scientific category labels to scientific documents where their terms are first grouped together using the well-known topic modelling method, Latent Dirichlet Allocation (LDA) and then assigned to Wikipedia Concepts by wikification. We wikify the terms of the topic model of a document to extract related concepts from Wikipedia. We experiment on two different datasets: the abstracts of the documents from the ACM Digital Library and the full papers of the UvT Collection. The ACM dataset includes Computer Science publications whereas UvT includes scientific publications from a range of topics. Domain specific taxonomies are used for evaluation. Results show that our approach is able to assign Wikipedia Concepts to the scientific publications in an automated manner, removing any need for human supervision.
Ozge Yeloglu, Evangelos E. Milios, Nur Zincir-Heywood
ACM Symposium on Document Engineering3
2013 Malicious Automatically Generated Domain Name Detection Using Stateful-SBB
Fariba Haddadi, Hilmi Günes Kayacik, Nur Zincir-Heywood, Malcolm I. Heywood
EvoApplications3
2013 Automatic optimization for a clustering based approach to support IT management
Can Bozdogan, Nur Zincir-Heywood, Yasemin Gokcen
IM2
2013 Investigating event log analysis with minimum apriori information
Adetokunbo Makanju, Nur Zincir-Heywood, Evangelos E. Milios
IM2
2013 IP traceback through (authenticated) deterministic flow marking: an empirical evaluation
abstract
In this paper, we present a novel approach to IP traceback - deterministic flow marking (DFM). We evaluate this novel approach against two well-known IP traceback schemes. These are the probabilistic packet marking (PPM) and the deterministic packet marking (DPM) techniques. In order to do so, we analyzed these techniques in detail in terms of their performances and feasibilities on five Internet traces. These traces consist of Darpa 1999 traffic traces, CAIDA October 2012 traffic traces, MAWI December 2012 traffic traces, and Dal2010 traffic traces. We have employed 16 performance metrics to evaluate their performances. The empirical results show that the novel DFM technique can reduce the number of marked packets by 91% compared to the DPM, while achieving the same or better performance in terms of its ability to trace back the attack. Additionally, DFM provides an optional authentication so that a compromised router cannot forge markings of other uncompromised routers. Unlike PPM and DPM that trace the attack up to the ingress interface of the edge router close to the attacker, DFM allows the victim to trace the origin of incorrect or spoofed source addresses up to the attacker node, even if the attack has been originated from a network behind a network address translation (NAT) server. Our results show that DFM can reach up to approximately 99% traceback rate with no false positives.
Vahid Aghaei Foroushani, Nur Zincir-Heywood
EURASIP J. Inf. Secur.2
2012 Symbiotic evolutionary subspace clustering
abstract
New emerging high-dimensional data sets have made traditional clustering algorithms increasingly inefficient. More sophisticated approaches are required to cope with the increasing dimensionality and cardinality of such data sets. Feature selection methods are proposed as a solution to deal with this problem, however they fail for data sets where the attribute support for different clusters is not the same. For this category of data sets subspace clustering algorithms have been introduced over the past decade. We approach this problem from the perspective of Genetic Algorithms by adopting a hierarchical data structure deployed in three stages. 1) a traditional clustering algorithm is applied independently to each attribute of the data set, thus defining a grid of potential 1-d cluster centroids. 2) representing multi-dimensional cluster centroids by indexing 1-d cluster centroids. 3) converting the problem of finding the best combination of cluster centroids into that of discrete optimization and applying a multi-objective evolutionary algorithm, which uses group fitness evaluation to give a fitness to a group of clusters, as defined by process 2. Synthetic data sets with different characteristics are generated as the ground truth to evaluate the resulting algorithm for Evolutionary Subspace Clustering (ESC) as well as benchmark against alternative subspace and full-space clustering algorithms. ESC returns competitive accuracy and while typically utilizing less attributes and scaling as attribute count increases.
Ali Vahdat, Malcolm I. Heywood, Nur Zincir-Heywood
IEEE Congress on Evolutionary Computation3
2012 Network Protocol Discovery and Analysis via Live Interaction
Patrick LaRoche, Nur Zincir-Heywood, Malcolm I. Heywood
EvoApplications2
2012 GP under streaming data constraints: a case for pareto archiving?
abstract
Classification as applied to streaming data implies that only a small number of new training instances appear at each generation and are never explicitly reintroduced by the stream. Pareto competitive coevolution provides a potential framework for archiving useful training instances between generations under an archive of finite size. Such a coevolutionary framework is defined for the online evolution of classifiers under genetic programming. Benchmarking is performed under multi-class data sets with class imbalance and training partitions with between 1,000's to 100,000's of instances. The impact of enforcing different constraints for accessing the stream are investigated. The role of online adaptation is explicitly documented and tests made on the relative impact of label error on the quality of streaming classifier results.
Aaron Atwater, Malcolm I. Heywood, Nur Zincir-Heywood
GECCO3
2012 The Impact of Evasion on the Generalization of Machine Learning Algorithms to Classify VoIP Traffic
abstract
We propose a novel approach to generate well generalized signatures to classify Skype VoIP traffic using a machine learning based approach. Results show that the performance of the signatures did not degrade significantly when they were evaluated on traffic that was captured from different locations and at different times as well as employed against evasion attacks. Our results on the evasion of Skype classifier demonstrate that the performance of the signatures are very promising even if the user tries maliciously to alter the characteristics of Skype traffic to evade the classifier.
Riyad Alshammari, Nur Zincir-Heywood
ICCCN2
2012 Data mining for supporting IT management
abstract
In this paper, we focus on the identification of the experience required for solving IT problems in small to medium size enterprises. Our goal is to utilize information retrieval and data mining techniques to automatically extract information from public forums, mailing lists, and FAQs in order to automatically generate a knowledge base for dynamic system administration support. To this end, we explore two similarity-distance measures and five clustering algorithms on three different datasetsto evaluate their performances. During the evaluations, CES+ algorithm gives promising results in terms of automatically extracting the most similar past experiences (problems /solutions) to a given fault.
Can Bozdogan, Nur Zincir-Heywood
NOMS2
2012 Interactive learning of alert signatures in High Performance Cluster system logs
abstract
The ability to automatically discover error conditions with little human input is a feature lacking in most modern computer systems and networks. However, with the ever increasing size and complexity of modern systems, such a feature will become a necessity in the not too distant future. Our work proposes a hybrid framework that allows High Performance Clusters (HPC) to detect error conditions in their logs. Through the use of anomaly detection, the system is able to detect portions of the log that are likely to contain errors (anomalies). Via visualization, human administrators can inspect these anomalies and assign labels to clusters that correlate with error conditions. The system can then learn a signature from the confirmed anomalies, which it uses to detect future occurrences of the error condition. Our evaluations show the system is able to generate simple and accurate signatures using very little data.
Adetokunbo Makanju, Nur Zincir-Heywood, Evangelos E. Milios
NOMS2
2012 A Lightweight Algorithm for Message Type Extraction in System Application Logs
abstract
Message type or message cluster extraction is an important task in the analysis of system logs in computer networks. Defining these message types automatically facilitates the automatic analysis of system logs. When the message types that exist in a log file are represented explicitly, they can form the basis for carrying out other automatic application log analysis tasks. In this paper, we introduce a novel algorithm for carrying out message type extraction from event log files. IPLoM, which stands for Iterative Partitioning Log Mining, works through a 4-step process. The first three steps hierarchically partition the event log into groups of event log messages or event clusters. In its fourth and final stage, IPLoM produces a message type description or line format for each of the message clusters. IPLoM is able to find clusters in data irrespective of the frequency of its instances in the data, it scales gracefully in the case of long message type patterns and produces message type descriptions at a level of abstraction, which is preferred by a human observer. Evaluations show that IPLoM outperforms similar algorithms statistically significantly.
Adetokunbo Makanju, Nur Zincir-Heywood, Evangelos E. Milios
IEEE Trans. Knowl. Data Eng.2
2011 System State Discovery Via Information Content Clustering of System Logs
abstract
Self-awareness is an important attribute for any system to have before it is capable of self-management. A system needs to have a continuous stream of real-time data to analyze to allow it be aware of its internal state. To this end, previous approaches have utilized system performance metrics and system log data to characterize system internal state. In using system logs to characterize system internal state, the computation of strongly correlated message types is necessary. In this work, we show that strongly correlated message types can be easily discovered without much computation. Our work explores a natural behaviour of system logs where system log data partitioned using source and time information contain correlated message types. We demonstrate how the groups of partitions, which contain correlated message types, can be found by clustering the partitions based on their entropy-based information content. We evaluate our method using cluster cohesion, cluster separation and cluster conceptual purity as metrics. The results show that our proposed method not only produces well-formed clusters but also clusters that can be mapped to different alert states with a high degree of confidence.
Adetokunbo Makanju, Nur Zincir-Heywood, Evangelos E. Milios
ARES2
2011 Is machine learning losing the battle to produce transportable signatures against VoIP traffic?
abstract
Traffic classification becomes more challenging since the traditional techniques such as port numbers or deep packet inspection are ineffective against voice over IP (VoIP) applications, which uses non-standard ports and encryption. Statistical information based on network layer with the use of machine learning (ML) can achieve high classification accuracy and produce transportable signatures. However, the ability of ML to find transportable signatures depends mainly on the training data sets. In this paper, we explore the importance of sampling training data sets for the ML algorithms, specifically Genetic Programming, C5.0, Naive Bayesian and AdaBoost, to find transportable signatures. To this end, we employed two techniques for sampling network training data sets, namely random sampling and consecutive sampling. Results show that random sampling and 90-minute consecutive sampling have the best performance in terms of accuracy using C5.0 and SBB, respectively. In terms of complexity, the size of C5.0 solutions increases as the training size increases, whereas SBB finds simpler solutions.
Riyad Alshammari, Nur Zincir-Heywood
IEEE Congress on Evolutionary Computation2
2011 Genetic optimization and hierarchical clustering applied to encrypted traffic identification
abstract
An important part of network management requires the accurate identification and classification of network traffic for decisions regarding bandwidth management, quality of service, and security. This work explores the use of a Multi-Objective Genetic Algorithm (MOGA) for both, feature selection and cluster count optimization, for an unsupervised machine learning technique, K-Means, applied to encrypted traffic identification. Specifically, a hierarchical K-Means algorithm is employed, comparing its performance to the MOGA with a non-hierarchical (flat) K-Means algorithm. The latter has already been benchmarked against common unsupervised techniques found in the literature, where results have favored the proposed MOGA. The purpose of this paper is to explore the gains, if any, obtained by increasing cluster purity in the proposed model by means of a second layer of clusters. In this work, SSH is chosen as an example of an encrypted application. However, nothing prevents the proposed model to work with other types of encrypted traffic, such as SSL or Skype. Results show that with the hierarchical MOGA, significant gains are observed in terms of the classification performance of the system.
Carlos Bacquet, Nur Zincir-Heywood, Malcolm I. Heywood
CICS2
2011 Exploring the state space of an application protocol: A case study of SMTP
abstract
In this work, we explore the state space of a network application protocol by employing genetic programming techniques. To this end, we target Simple Mail Transfer Protocol (SMTP), which is a well-known and open protocol on the Internet. In order to achieve our goal, we aim to evolve the payload such that solution individuals result in an email being sent successfully through the targeted server. The proposed system implements an archive paradigm where, upon completion of the evolutionary process, a collection (archive) of solutions are presented. Specifically, they can all achieve the goal, but each does so in a unique manner. This collection allows us to examine the state space of the application protocol, giving us the ability to verify that these variations are either intended by the protocol, or should be addressed for security reasons.
Patrick LaRoche, Nur Zincir-Heywood, Malcolm I. Heywood
CICS2
2011 A Comparison of three machine learning techniques for encrypted network traffic analysis
abstract
This work evaluates three methods for encrypted traffic analysis without using the IP addresses, port number, and payload information. To this end, binary identification of SSH vs non-SSH traffic is used as a case study since the plain text initiation of the SSH protocol allows us to obtain data sets with a reliable ground truth. The methods are subject to several tests using different export options, feature sets, and training and test traffic traces for a total of 128 different configurations. Of particular interest are test cases which that use a test set from a different network than that which the model was trained on, i.e. robustness of the trained models. Results show that the multi-objective genetic algorithm (MOGA) based trained model is able to achieve the best performance among the three methods when each approach is tested on traffic traces that are captured on the same network as the training network trace. On the other hand, C4.5 achieved the best results among the three methods when tested on traffic traces which are captured on totally different networks than the training trace. Furthermore, it is shown that continuous sampling of the training data is no better than random sampling, but the training data is very important for how well the classifiers will perform on traffic traces captured from different networks. Moreover, the C4.5 based approach provides the fastest and the most human readable model, whereas the MOGA reduces the complexity of the k-means clustering algorithm tremendously.
Daniel J. Arndt, Nur Zincir-Heywood
CISDA2
2011 An investigation on identifying SSL traffic
abstract
The importance of knowing what type of traffic is flowing through a network is paramount to its success. Traffic engineering, quality of service, identifying critical business applications, intrusion detection systems, as well as network management activities all require the base knowledge of what traffic is flowing over a network before any further steps can be taken. With Secure Socket Layer (SSL) traffic on the rise due to applications securing or concealing their traffic via encryption, the ability to determine what applications are running within a network is getting more and more difficult. Traditional methods of traffic classification through port numbers and deep packet inspection tools have been deemed inadequate despite their continued popular usage. The purpose of this work is to investigate if a machine learning approach can be used with flow features to identify SSL traffic in a given network trace. To this end, different machine learning methods, namely AdaBoost, C4.5, RIPPER, and Naive Bayesian techniques, are investigated without the use of port numbers, Internet Protocol addresses, or payload information.
Curtis McCarthy, Nur Zincir-Heywood
CISDA2
2011 A next generation entropy based framework for alert detection in system logs
abstract
Recent research efforts have highlighted the capability of entropy based approaches in the automatic discovery of alerts in system logs. In this work, we extend this research to present the evaluations of three entropy based approaches on new datasets not utilized in previous papers. We also extend the approach with the introduction of a Cluster Membership Anomaly score. This extension of the approach is intended to reduce the false positive rates required to detect all alerts. Previous work has shown that false positive rates required for the detection of all alerts for an entropy based approach could be very high. The results show that the Cluster Membership Anomaly score has value for the reduction of false positive rates.
Adetokunbo Makanju, Nur Zincir-Heywood, Evangelos E. Milios
Integrated Network Management2
2011 Can encrypted traffic be identified without port numbers, IP addresses and payload inspection?
Riyad Alshammari, Nur Zincir-Heywood
Comput. Networks2
2011 Robust learning intrusion detection for attacks on wireless networks
abstract
We address the problem of evaluating the robustness of machine learning based detectors for deployment in real life networks. To this end, we employ Genetic Programming for evolving classifiers and Artificial Neural Networks as our machine learning paradigms under three different Denial-of-Service attacks at the Data Link layer (De-authentication, Authentication and Association attacks). We investigate their cross-platform robustness and cross-attack robustness. Cross-platform robustness is the ability to seamlessly port an Intrusion Detector trained on one network to another network with little or no change and without a drop in performance. Cross-attack robustness is the ability of a detector trained on one attack type to detect a different but similar attack on which it has not been trained. Our results show that the potential of a machine learning based detector can be significantly enhanced or limited by the representation of the training data for the learning algorithms.
Adetokunbo Makanju, Nur Zincir-Heywood, Evangelos E. Milios
Intell. Data Anal.2
2010 One Size Fits None: The Importance of Detector Parameterization
abstract
The parameterization of an administrator's intrusion detection system (IDS) is as crucial as the IDS itself. The difference between sufficient and insufficient parameterization can be the difference between a detected and undetected attack. This work focuses on identifying a methodical process for IDS parameterization. Such a process provides administrators of intrusion detection systems with the knowhow of selecting suitable parameters for the effective operation of their detector. The process stresses the importance of altering parameters for individual applications. Parameterization experiments are employed on two different open source IDSs, namely Stide and pH, and tested against three real world vulnerabilities. The results show the interesting trends that are observed during the experiments.
Natasha Bodorik, Nur Zincir-Heywood
ARES2
2010 Unveiling Skype encrypted tunnels using GP
abstract
The classification of Encrypted Traffic, namely Skype, from network traffic represents a particularly challenging problem. Solutions should ideally be both simple-therefore efficient to deploy-and accurate. Recent advances to team-based Genetic Programming provide the opportunity to decompose the original problem into a subset of classifiers with non-overlapping behaviors. Thus, in this work we have investigated the identification of Skype encrypted traffic using Symbiotic Bid-Based (SBB) paradigm of team based Genetic Programming (GP) found on flow features without using IP addresses, port numbers and payload data. Evaluation of SBB-GP against C4.5 and AdaBoost-representing current best practice-indicates that SBB-GP solutions are capable of providing simpler solutions in terms number of features used and the complexity of the solution/model without sacrificing accuracy.
Riyad Alshammari, Nur Zincir-Heywood
IEEE Congress on Evolutionary Computation2
2010 An analysis of clustering objectives for feature selection applied to encrypted traffic identification
abstract
This work explores the use of clustering objectives in a Multi-Objective Genetic Algorithm (MOGA) for both, feature selection and cluster count optimization, under the application of flow based encrypted traffic identification. We first explore whether it is possible to achieve the performance of a gold standard model (i.e., classification objectives), using a MOGA based on clustering objectives. Then, we explore the performance gain (if it exists) of applying a logarithmic transformation to the data prior to running the MOGA. Results show that MOGA trained with clustering objectives can closely reproduce the behavior of a gold standard model, not only in terms of the selected features, but also in terms of the achieved detection rate and false positives rate, above 90% and less than 1% respectively. On the other hand, no gain was observed by applying logarithmic transformation to the data.
Carlos Bacquet, Nur Zincir-Heywood, Malcolm I. Heywood
IEEE Congress on Evolutionary Computation2
2010 Bottom-up evolutionary subspace clustering
abstract
The ultimate goal of subspace clustering algorithms is to identify both the subset of attributes supporting a cluster and the location of the cluster in the subspace. In this work a generic evolutionary approach to bottom-up subspace clustering is proposed consisting of three steps. The first applies a non-evolutionary clustering algorithm attribute-wise to establish the lattice from which subspace clusters will be designed. In the second step a multi-objective Genetic Algorithm (MOGA) is used to evolve good candidate subspace clusters (CSC) through a combinatorial search w.r.t. the attribute-wise lattice from step 1. The third step then searches in the space of CSC from the population of the the first MOGA to find the best combination of subspace clusters, again under a MOGA formulation. Important properties of the approach are that a standard clustering algorithm is deployed in step one to build the initial lattice of attribute-wise clusters. This helps to decouple the computational expense of clustering using Evolutionary Computation, with the MOGA applied in steps 2 and 3 building clusters through a combinatorial search relative to the original lattice parameters. Benchmarking on data sets with tens to hundreds of attributes illustrates the feasibility of the approach.
Ali Vahdat, Malcolm I. Heywood, Nur Zincir-Heywood
IEEE Congress on Evolutionary Computation3
2010 An investigation on the identification of VoIP traffic: Case study on Gtalk and Skype
abstract
The classification of encrypted traffic on the fly from network traces represents a particularly challenging application domain. Recent advances in machine learning provide the opportunity to decompose the original problem into a subset of classifiers with non-overlapping behaviors, in effect providing further insight into the problem domain. Thus, the objective of this work is to classify VoIP encrypted traffic, where Gtalk and Skype applications are taken as good representatives. To this end, three different machine learning based approaches, namely, C4.5, AdaBoost and Genetic Programming (GP), are evaluated under data sets common and independent from the training condition. In this case, flow based features are employed without using the IP addresses, source/destination ports and payload information. Results indicate that C4.5 based machine learning approach has the best performance.
Riyad Alshammari, Nur Zincir-Heywood
CNSM2
2010 Using Code Bloat to Obfuscate Evolved Network Traffic
Patrick LaRoche, Nur Zincir-Heywood, Malcolm I. Heywood
EvoApplications (2)2
2009 Generalization of signatures for SSH encrypted traffic identification
abstract
The objective of this work is to discover generalized signatures for identifying encrypted traffic where SSH is taken as an example application. What we mean by generalized signatures is that the signatures learned by training on one network are still valid when they are applied to traffic coming from a totally different network. We identified 13 signatures and 14 flow attributes for SSH traffic classification where IP addresses, source/destination ports and payload information are not employed. The signatures are able to identify encrypted traffic with high detection rate and low false positive rate. We can achieve up to 97% DR and 0.8% FPR for identifying SSH traffic.
Riyad Alshammari, Nur Zincir-Heywood
CICS2
2009 Generating mimicry attacks using genetic programming: A benchmarking study
abstract
Mimicry attacks have been the focus of detector research where the objective of the attacker is to generate multiple attacks satisfying the same generic exploit goals for a given vulnerability. In this work, multi-objective Genetic programming is used to establish a “black-box” approach to mimicry attack generation. No knowledge is made of internal data structures of the target anomaly detector, only the anomaly rate reported by the detector. Such a “black box” methodology enables a vulnerability testing approach where both open-source and commodity anomaly detection systems can be tested. The approach successfully identifies exploits when benchmarked over four detectors and four applications.
Hilmi Günes Kayacik, Nur Zincir-Heywood, Malcolm I. Heywood, Stefan Burschka
CICS2
2009 Machine learning based encrypted traffic classification: Identifying SSH and Skype
abstract
The objective of this work is to assess the robustness of machine learning based traffic classification for classifying encrypted traffic where SSH and Skype are taken as good representatives of encrypted traffic. Here what we mean by robustness is that the classifiers are trained on data from one network but tested on data from an entirely different network. To this end, five learning algorithms — AdaBoost, Support Vector Machine, Naïe Bayesian, RIPPER and C4.5 — are evaluated using flow based features, where IP addresses, source/destination ports and payload information are not employed. Results indicate the C4.5 based approach performs much better than other algorithms on the identification of both SSH and Skype traffic on totally different networks.
Riyad Alshammari, Nur Zincir-Heywood
CISDA2
2009 Optimizing anomaly detector deployment under evolutionary black-box vulnerability testing
abstract
This work focuses on testing anomaly detectors from the perspective of a Multi-objective Evolutionary Exploit Generator (EEG). Such a framework provides users of anomaly detection systems two capabilities. Firstly, no knowledge of protected data structures need to be assumed (i.e. the detector is a black-box), where the time, knowledge and availability of tools to perform such an analysis might not be generally available. Secondly, the evolved exploits are then able to demonstrate weaknesses in the ensuing detector parameterization. Therefore, the system administrator can identify the suitable parameters for the effective operation of the detector. EEG is employed against two second generation anomaly detectors, namely pH and pH with schema mask, on four UNIX applications in order to perform a vulnerability assessment and make a comparison between the two detectors.
Hilmi Günes Kayacik, Nur Zincir-Heywood, Malcolm I. Heywood, Stefan Burschka
CISDA2
2009 Evolving TCP/IP packets: A case study of port scans
abstract
In this work, we investigate the ability of genetic programming techniques to evolve valid network packets, including all relevant header values, towards a specific goal. We see this as a first step in building a fuzzing system that can learn to adapt for vulnerability analysis. By developing a system that learns the packets that are required to be transmitted towards targets, using feedback from an external network source, we make a step towards having a system that can intelligently explore the capabilities of a given security system. In order to validate our system's capabilities we evolve a variety of port scan patterns while running the packets through an IDS, with the goal to minimizes the alarms raised during the scanning process. Results show that the system not only successfully evolves valid TCP packets, but also remains stealthy in its activity.
Patrick LaRoche, Nur Zincir-Heywood, Malcolm I. Heywood
CISDA2
2009 Clustering event logs using iterative partitioning
abstract
The importance of event logs, as a source of information in systems and network management cannot be overemphasized. With the ever increasing size and complexity of today's event logs, the task of analyzing event logs has become cumbersome to carry out manually. For this reason recent research has focused on the automatic analysis of these log files. In this paper we present IPLoM (Iterative Partitioning Log Mining), a novel algorithm for the mining of clusters from event logs. Through a 3-Step hierarchical partitioning process IPLoM partitions log data into its respective clusters. In its 4th and final stage IPLoM produces cluster descriptions or line formats for each of the clusters produced. Unlike other similar algorithms IPLoM is not based on the Apriori algorithm and it is able to find clusters in data whether or not its instances appear frequently. Evaluations show that IPLoM outperforms the other algorithms statistically significantly, and it is also able to achieve an average F-Measure performance 78% when the closest other algorithm achieves an F-Measure performance of 10%.
Adetokunbo Makanju, Nur Zincir-Heywood, Evangelos E. Milios
KDD2
2008 A Preliminary Investigation of Skype Traffic Classification Using a Minimalist Feature Set
abstract
In this work, AdaBoost and C4.5, are employed for classifying Skype direct (UDP and TCP) communications from traffic log files. Pre-processing is applied to the traffic data to express it as flows, which is later converted into a descriptive feature set. The aforementioned algorithms are then evaluated on this feature set. Results show that a 98% detection rate with6% false positive rate for UDP based Skype and a 94% detection rate with 4% false positive rate for TCP based Skype is possible to achieve.
Duffy Angevine, Nur Zincir-Heywood
ARES2
2008 Adaptabilty of a GP Based IDS on Wireless Networks
abstract
Abstract—Security and Intrusion detection in WiFi networks is currently an active area of research where WiFi specific Data Link layer attacks are an area of focus; particularly recent work has focused on producing machine learning based IDSs for these WiFi specific attacks. These proposed machine learning based IDSs come in addition to the already deployed signatures which are already in use in conventional intrusion detection systems like Snort-Wireless and Kismet. In this paper, we compare the detection capability of Snort-Wireless and a Genetic Programming (GP) based intrusion detector, based on the ability to adapt to modified attacks, ability to adapt to similar unknown attacks and infrastructure independent detection. Our results show that the GP based detection system is much more robust against modified attacks compared to Snort-Wireless. Moreover, by focusing on the method(s) used in feature preprocessing for presentation to learning algorithms, GP based IDSs can achieve infrastructure independent detection and can adapt to similar unknown attacks too. On the other hand, even though Snort-Wireless is an infrastructure independent detector, it cannot adapt to unknown attacks even if they are similar to others for which it has signatures on.
Adetokunbo Makanju, Nur Zincir-Heywood, Evangelos E. Milios
ARES2
2008 VEA-bility Security Metric: A Network Security Analysis Tool
abstract
In this work, we propose a novel quantitative security metric, VEA-bility, which measures the desirability of different network configurations. An administrator can then use the VEA-bility scores of different configurations to configure a secure network. Based on our findings, we conclude that the VEA-bility can be used to accurately estimate the comparative desirability of a specific network configuration. This information can then be used to explore alternate possible configurations and allows an administrator to select one among the given options. These tools are important to network administrators as they strive to provide secure, yet functional, network configurations.
Melanie Tupper, Nur Zincir-Heywood
ARES2
2008 Investigating Two Different Approaches for Encrypted Traffic Classification
abstract
The basic objective of this work is to compare the utility of an expert driven system and a data driven system for classifying encrypted network traffic, specifically SSH traffic from traffic log files. Pre-processing is applied to the traffic data to represent as traffic flows. Results show that the data driven system approach outperforms the expert driven system approach in terms of high detection and low false positive rates.
Riyad Alshammari, Nur Zincir-Heywood
PST2
2008 Mimicry Attacks Demystified: What Can Attackers Do to Evade Detection?
abstract
Mimicry attacks have been the focus of detector research where the objective of the attacker is to generate an attack that evades detection while achieving the attackerpsilas goals. If such an attack can be found, it implies that the target detector is vulnerable against mimicry attacks. In this work, we emphasize that there are two components of a buffer overflow attack: the preamble and the exploit. Although the attacker can modify the exploit component easily, the attacker may not be able to prevent preamble from generating anomalous behavior since during preamble stage, the attacker does not have full control. Previous work on mimicry attacks considered an attack to completely evade detection, if the exploit raises no alarms. On the other hand, in this work, we investigate the source of anomalies in both the preamble and the exploit components against two anomaly detectors that monitor four vulnerable UNIX applications. Our experiment results show that preamble can be a source of anomalies, particularly if it is lengthy and anomalous.
Hilmi Günes Kayacik, Nur Zincir-Heywood
PST2
2008 LogView: Visualizing Event Log Clusters
abstract
Event logs or log files form an essential part of any network management and administration setup. While log files are invaluable to a network administrator, the vast amount of data they sometimes contain can be overwhelming and can sometimes hinder rather than facilitate the tasks of a network administrator. For this reason several event clustering algorithms for log files have been proposed, one of which is the event clustering algorithm proposed by Risto Vaarandi, on which his simple log file clustering tool (SLCT) is based. The aim of this work is to develop a visualization tool that can be used to view log files based on the clusters produced by SLCT. The proposed visualization tool, which is called LogView, utilizes treemaps to visualize the hierarchical structure of the clusters produced by SLCT. Our results based on different application log files show that LogView can ease the summarization of vast amount of data contained in the log files. This in turn can help to speed up the analysis of event data in order to detect any security issues on a given application.
Adetokunbo Makanju, Stephen Brooks, Nur Zincir-Heywood, Evangelos E. Milios
PST3
2007 Automatically Evading IDS Using GP Authored Attacks
abstract
A mimicry attack is a type of attack where the basic steps of a minimalist 'core' attack are used to design multiple attacks achieving the same objective from the same application. Research in mimicry attacks is valuable in determining and eliminating weaknesses of detectors. In this work, we provide a genetic programming based automated process for designing all components of a mimicry attack relative to the Stide detector under a vulnerable Traceroute application. Results indicate that the automatic process is able to generate mimicry attacks that reduce the alarm rate from ~65% of the original attack, to ~2.7%, effectively making the attack indistinguishable from normal behaviors
Hilmi Günes Kayacik, Nur Zincir-Heywood, Malcolm I. Heywood
CISDA2
2007 A Comparison Between Signature and GP-Based IDSs for Link Layer Attacks on WiFi Networks
abstract
Data link layer attacks on WiFi networks are known to be one of the weakest points of WiFi networks. While these attacks are very simple in implementation, their effect on WiFi networks can be devastating. To this end, several intrusion detection systems (IDS) have been employed to detect these attacks. In this paper, we compare the ability of Snort-Wireless and a genetic programming (GP) based intrusion detector, in the detection of a particular data link layer attack, namely the deauthentication attack. We focus particularly on a scenario where the attacker stealthily injects the attack frames into the target network. Results show that the GP based detection system is much more robust against the different versions of the attack compared to Snort-Wireless and can achieve a detection rate in average 100% and a false positive rate in average 0.1%
Adetokunbo Makanju, Patrick LaRoche, Nur Zincir-Heywood
CISDA3
2007 A flow based approach for SSH traffic detection
abstract
The basic objective of this work is to assess the utility of two supervised learning algorithms AdaBoost and RIPPER for classifying SSH traffic from log files without using features such as payload, IP addresses and source/destination ports. Pre-processing is applied to the traffic data to express as traffic flows. Results of 10-fold cross validation for each learning algorithm indicate that a detection rate of 99% and a false positive rate of 0.7% can be achieved using RIPPER. Moreover, promising preliminary results were obtained when RIPPER was employed to identify which service was running over SSH. Thus, it is possible to detect SSH traffic with high accuracy without using features such as payload, IP addresses and source/destination ports, where this represents a particularly useful characteristic when requiring generic, scalable solutions.
Riyad Alshammari, Nur Zincir-Heywood
SMC2
2007 Growing recurrent self organizing map
abstract
The growing Recurrent Self-Organizing Map (GRSOM) is embedded into a standard Self-Organizing Map (SOM) hierarchy. To do so, the KDD benchmark dataset from the International Knowledge Discovery and Data Mining Tools Competition is employed. This dataset consists of 500,000 training patterns and 41 features for each pattern. Unlike most of the previous methods, only 6 of the basic features are employed. The resulting model has a capability of detection (false positive) rate of 89.6% (5.66%), where this is as good as the data-mining approaches that uses all 41 features and twice as faster than a similar hierarchical SOM architecture.
Ozge Yeloglu, Nur Zincir-Heywood, Malcolm I. Heywood
SMC2
2007 A hierarchical SOM-based intrusion detection system
Hilmi Günes Kayacik, Nur Zincir-Heywood, Malcolm I. Heywood
Eng. Appl. Artif. Intell.2
2006 Evolving Recurrent Linear-GP for Document Classification and Word Tracking
abstract
In this paper, we propose a novel document classification system where the recurrent linear Genetic Programming is employed to classify the documents that are represented in encoded word sequences. During this process, word sequences of documents are tracked, frequent patterns are detected and document is classified. We describe the word encoding model and the recurrent linear Genetic Programming based classification mechanism. The performance results on benchmark data set Reuters 21578 show that this system can analyze the temporal sequence patterns of a document and get competitive performance on classification. We expect that it can be easily applied to other application areas, where the temporal sequences are very significant.
Xiao Luo 0002, Nur Zincir-Heywood
IEEE Congress on Evolutionary Computation2
2006 802.11 De-authentication Attack Detection Using Genetic Programming
Patrick LaRoche, Nur Zincir-Heywood
EuroGP2
2006 On evolving buffer overflow attacks using genetic programming
abstract
In this work, we employed genetic programming to evolve a "white hat" attacker; that is to say, we evolve variants of an attack with the objective of providing better detectors. Assuming a generic buffer overflow exploit, we evolve variants of the generic attack, with the objective of evading detection by signature-based methods. To do so, we pay particular attention to the formulation of an appropriate fitness function and partnering instruction set. Moreover, by making use of the intron behavior inherent in the genetic programming paradigm, we are able to explicitly obfuscate the true intent of the code. All the resulting attacks defeat the widely used 'Snort' Intrusion Detection System.
Hilmi Günes Kayacik, Malcolm I. Heywood, Nur Zincir-Heywood
GECCO3
2006 Using self-organizing maps to build an attack map for forensic analysis
abstract
In this work, we focus on developing behavioral models of known attacks to help security experts to identify the similarities between attacks. Furthermore, these attack behavior models can be used to analyze zero-day attacks, which security experts have limited knowledge of. To this end, a Self Organizing Feature Map (SOM) is employed to model the relationship between known attacks and U-Matrix representation is used to create a two dimensional topological map of known attacks. The approach is evaluated on KDD'99 data set. Results show that attacks with similar behavior patterns are placed together on the map. Moreover, when new attacks are presented, SOM assigned similar labels to the attacks that are newer versions of the known attacks.
Hilmi Günes Kayacik, Nur Zincir-Heywood
PST2
2005 Evolving Successful Stack Overflow Attacks for Vulnerability Testing
abstract
The work presented in this paper is intended to test crucial system services against stack overflow vulnerabilities. The focus of the test is the user-accessible variables, that is to say, the inputs from the user as specified at the command line or in a configuration file. The tester is defined as a process for automatically generating a wide variety of user-accessible variables that result in malicious buffers (an exploit). In this work, the search for successful exploits is formulated as an optimization problem and solved using evolutionary computation. Moreover the resulting attacks are passed through the Snort misuse detection system to observe the detection (or not) of each exploit
Hilmi Günes Kayacik, Nur Zincir-Heywood, Malcolm I. Heywood
ACSAC2
2005 CasGP: building cascaded hierarchical models using niching
abstract
A cascaded model is introduced for mining large datasets using genetic programming without recourse to specialist hardware. Such an algorithm satisfies the seeming conflicting requirements of scalability and accuracy on large datasets by incrementally building GP classifiers through the use of a hierarchical dynamic subset selection algorithm. Models are built incrementally with each layer of the cascade receiving as input the original feature vector, plus the output from the previous layer(s). In order to encourage each layer to explicitly solve new aspects of the problem a combination of sum square error and niching is utilized. Thus, previous layers of the model are considered a niche, and the cost function is a shared error metric.
Peter Lichodzijewski, Malcolm I. Heywood, Nur Zincir-Heywood
Congress on Evolutionary Computation3
2005 Evolving recurrent models using linear GP
abstract
Turing complete Genetic Programming (GP) models introduce the concept of internal state, and therefore have the capacity for identifying interesting temporal properties. Surprisingly, there is little evidence of the application of such models to problems for prediction. An empirical evaluation is made of a simple recurrent linear GP model over standard prediction problems.
Xiao Luo 0002, Malcolm I. Heywood, Nur Zincir-Heywood
GECCO3
2005 Comparison of a SOM based sequence analysis system and naive Bayesian classifier for spam filtering
abstract
The problem introduced by the unsolicited bulk emails, also known as "spam" generates a need for reliable anti-spam filters. In this paper, we design and compare the performance of a newly designed SOM based sequence analysis (SBSA) system for the spam filtering task. The system is based on a SOM based sequential data representation combined with a kNN classifier designed to make use of word sequence information. We compare this system with the traditional baseline method naive Bayesian filter. Three different cost scenarios and suitable cost-sensitive measurements are employed. The results show that the SBSA system is superior to the naive Bayesian filter, particularly when the misclassification cost for non-spam message is high.
Xiao Luo 0002, Nur Zincir-Heywood
IJCNN2
2005 Training the SOFM efficiently: an example from intrusion detection
abstract
The dynamic subset selection (DSS) active learning algorithm is generalized to include the case of unsupervised learning. To do so, training set partitioning, exemplar difficulty and age, and early stopping criteria are introduced into the self organizing feature map algorithm. The resulting model is able to build a hierarchical SOFM on a large (500,000 pattern) dataset in 3 hours. In comparison, the same architecture without active learning requires 33 hours to construct. No reduction in accuracy is recorded for the DSS SOFM model.
Leigh Wetmore, Nur Zincir-Heywood, Malcolm I. Heywood
IJCNN2
2005 Analysis of Three Intrusion Detection System Benchmark Datasets Using Machine Learning Algorithms
Hilmi Günes Kayacik, Nur Zincir-Heywood
ISI2
2005 Evaluation of Two Systems on Multi-class Multi-label Document Classification
Xiao Luo 0002, Nur Zincir-Heywood
ISMIS2
2005 Selecting Features for Intrusion Detection: A Feature Relevance Analysis on KDD 99
Hilmi Günes Kayacik, Nur Zincir-Heywood, Malcolm I. Heywood
PST2
2005 Post-Supervised Template Induction for Information Extraction from Lists and Tables in Dynamic Web Sources
Zhongmin Shi, Evangelos E. Milios, Nur Zincir-Heywood
J. Intell. Inf. Syst.3
2005 Speeding up the Self-Organizing Feature Map Using Dynamic Subset Selection
Leigh Wetmore, Malcolm I. Heywood, Nur Zincir-Heywood
Neural Process. Lett.3
2005 Training genetic programming on half a million patterns: an example from anomaly detection
abstract
The hierarchical RSS-DSS algorithm is introduced for dynamically filtering large datasets based on the concepts of training pattern age and difficulty, while utilizing a data structure to facilitate the efficient use of memory hierarchies. Such a scheme provides the basis for training genetic programming (GP) on a data set of half a million patterns in 15 min. The method is generic, thus, not specific to a particular GP structure, computing platform, or application context. The method is demonstrated on the real-world KDD-99 intrusion detection data set, resulting in solutions competitive with those identified in the original KDD-99 competition, while only using a fraction of the original features. Parameters of the RSS-DSS algorithm are demonstrated to be effective over a wide range of values. An analysis of different cost functions indicates that hierarchical fitness functions provide the most effective solutions.
Dong Song, Malcolm I. Heywood, Nur Zincir-Heywood
IEEE Trans. Evol. Comput.3
2004 Cascaded GP models for data mining
abstract
The cascade architecture for incremental learning is demonstrated within the context of genetic programming. Such a scheme provides the basis for building steadily more complex models until a desired degree of accuracy is reached. The architecture is demonstrated for several data mining datasets. Efficient training on standard computing platforms is retained using the RSS-DSS algorithm for stochastically sampling datasets in proportion to exemplar 'difficulty' and 'age'. Finally, the ensuing empirical study provides the basis for recommending the utility of sum square cost functions in the datasets considered.
Peter Lichodzijewski, Malcolm I. Heywood, Nur Zincir-Heywood
IEEE Congress on Evolutionary Computation3
2004 Analyzing the Temporal Sequences for Text Categorization
Xiao Luo 0002, Nur Zincir-Heywood
KES2
2004 World Wide Web site summarization
Yongzheng Zhang 0001, Nur Zincir-Heywood, Evangelos E. Milios
Web Intell. Agent Syst.2
2003 A Linear Genetic Programming Approach to Intrusion Detection
Dong Song, Malcolm I. Heywood, Nur Zincir-Heywood
GECCO3
2003 On the capability of an SOM based intrusion detection system
abstract
An approach to network intrusion detection is investigated, based purely on a hierarchy of Self-Organizing Feature Maps. Our principle interest is to establish just how far such an approach can be taken in practice. To do so, the KDD benchmark dataset from the International Knowledge Discovery and Data Mining Tools Competition is employed. This supplies a connection-based description of a factitious computer network in which each connection is described in terms of 41 features. Unlike previous approaches, only 6 of the most basic features are employed. The resulting system is capable of detection (false positive) rates of 89% (4.6%), where this is at least as good as the alternative data-mining approaches that require all 41 features.
Hilmi Günes Kayacik, Nur Zincir-Heywood, Malcolm I. Heywood
IJCNN2
2003 A comparison of SOM based document categorization systems
abstract
This paper describes the development and evaluation of two unsupervised learning mechanisms for solving the automatic document categorization problem. Both mechanisms are based on a hierarchical structure of self-organizing feature maps. Specifically, one architecture is based on the vector space model whereas the other one is based on a code-books model. Results show that the latter architecture performs better than the first one which is based on the quality of the returned clusters.
Xiao Luo 0002, Nur Zincir-Heywood
IJCNN2
2003 A Case Study of Three Open Source Security Management Tools
Hilmi Günes Kayacik, Nur Zincir-Heywood
Integrated Network Management2
2002 The effect of routing under local information using a social insect metaphor
abstract
Although adaptive and heuristic approaches perform well under idealized conditions to the packet network routing problem, such algorithms are also dependent on global information that is not available under real-world conditions. This work benchmarks routing under local information conditions using the AntNet algorithm and makes recommendations regarding future approaches.
Suiliong Liang, Nur Zincir-Heywood, Malcolm I. Heywood
IEEE Congress on Evolutionary Computation2
2002 Intelligent Packets For Dynamic Network Routing Using Distributed Genetic Algorithm
Suihong Liang, Nur Zincir-Heywood, Malcolm I. Heywood
GECCO2
2002 Object-Orientated Design of Digital Library Platforms for Multiagent Environments
abstract
The application of an object-oriented (OO) methodology to the design of a platform for heterogeneous digital libraries with multiagent technologies is demonstrated. Emphasis is placed on maximizing the autonomy of the query processing activity. The Fusion OO paradigm is specifically employed as the basis for the design process due to the significance attributed to the development of object interfaces under static and dynamic conditions. Finally, characteristics of the proposed Domain Index Server (DIS) platform are contrasted with those of an alternative platform (the University of Michigan Digital Library, UMDL) by way of the respective Fusion descriptions. This identifies a different emphasis on the interface design between objects in the two platforms: the DIS system uses more dynamic links while the UMDL system focuses on permanent and constant links. Simulation of the two platforms provides performance data that demonstrates the higher capacity of the DIS scheme.
Nur Zincir-Heywood, Malcolm I. Heywood, Chris R. Chatwin
IEEE Trans. Knowl. Data Eng.1
2002 Dynamic page based crossover in linear genetic programming
abstract
Page-based linear genetic programming (GP) is proposed in which individuals are described in terms of a number of pages. Pages are expressed in terms of a fixed number of instructions, which is constant for all individuals in the population. Pairwise crossover results in the swapping of single pages, and thus, individuals are of a fixed number of instructions. Head-to-head comparison with Tree-structured GP and block-based linear GP indicates that the page-based approach evolves succinct solutions without penalizing generalization ability.
Malcolm I. Heywood, Nur Zincir-Heywood
IEEE Trans. Syst. Man Cybern. Part B2
2000 Register Based Genetic Programming on FPGA Computing Platforms
Malcolm I. Heywood, Nur Zincir-Heywood
EuroGP2
2000 Page-based linear genetic programming
abstract
Genetic programming arguably represents the most general form of evolutionary computation. However, such generality is not without significant computational overheads. Particularly, the cost of evaluating the fitness of individuals in any form of evolutionary computation represents the single most significant computational bottleneck. A less widely acknowledged computational overhead in GP involves the implementation of the crossover operator. To this end a page-based definition of individuals is used to restrict crossover to equal length code fragments. Moreover, by using a register-machine context, the significance of a priori internal register external output definitions is emphasized.
Malcolm I. Heywood, Nur Zincir-Heywood
SMC2
2000 Heterogeneous Digital Library Query Platform Using a Truly Distributed Multi-Agent Search
abstract
A platform for performing multi-agent searches in heterogeneous digital libraries is proposed. This differs significantly from previous approaches by completely removing the concept of a centralized search engine. Specifically, the organization of information held on domain index servers is constrained to conform to a virtual tree representation based on facets and global keyword concept schema particular to the set of information providers associated with the domain of interest (e.g. preparatory intranet). Simulation studies are used to compare this platform against a digital library platform presently in use, which employs the traditional central server scheme. Improvements in terms of query service time and robustness are demonstrated.
Nur Zincir-Heywood, Malcolm I. Heywood, Chris R. Chatwin, Emrullah Turhan Tunali
Int. J. Cooperative Inf. Syst.1
2000 Digital library query clearing using clustering and fuzzy decision-making
Malcolm I. Heywood, Nur Zincir-Heywood, Chris R. Chatwin
Inf. Process. Manag.2