Elias Bou-Harb

dblp:122/5447 · DBLP profile ↗
← Back
77ranked-venue papers
14as first author
37since 2021 · last 2026
0000-0001-8040-4635ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 31 · 6 first-author · 16 since 2021Computer networks · 29 · 6 first-author · 14 since 2021Systems, architecture and hardware · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Beyond Arbitrary Thresholds: Conformal Prediction for Trustworthy CAN Bus Intrusion Detection
Anthony Nasry Massaad, Aleksandar Avdalovic, Peyton Andras, Joseph Khoury, Elias Bou-Harb
ICC5
2025 Enhancing Network Security Management in Water Systems using FM-based Attack Attribution
abstract
Water systems are vital components of modern infrastructure, yet they are increasingly susceptible to sophisticated cyber attacks with potentially dire consequences on public health and safety. While state-of-the-art machine learning techniques effectively detect anomalies, contemporary model-agnostic attack attribution methods using LIME, SHAP, and LEMNA are deemed impractical for large-scale, interdependent water systems. This is due to the intricate interconnectivity and dynamic interactions that define these complex environments. Such methods primarily emphasize individual feature importance while falling short of addressing the crucial sensor-actuator interactions in water systems, which limits their effectiveness in identifying root cause attacks. To this end, we propose a novel model-agnostic Factorization Machines (FM)-based approach that capitalizes on water system sensor-actuator interactions to provide granular explanations and attributions for cyber attacks. For instance, an anomaly in an actuator pump activity can be attributed to a top root cause attack candidates, a list of water pressure sensors, which is derived from the underlying linear and quadratic effects captured by our approach. We validate our method using two real-world water system specific datasets, SWaT and WADI, demonstrating its superior performance over traditional attribution methods. In multi-feature cyber attack scenarios involving intricate sensor-actuator interactions, our FM-based attack attribution method effectively ranks attack root causes, achieving approximately 20% average improvement over SHAP and LEMNA. Additionally, our approach maintains strong performance in single-feature attack scenarios, demonstrating versatility across different types of cyber attacks. Notably, our approach maintains a low computational overhead equating to an O(n) time complexity, making it suitable for real-time applications in critical water system infrastructure. Our work underscores the importance of modeling feature interactions in water systems, offering a robust tool for operators to diagnose and mitigate root cause attacks more effectively.
Aleksandar Avdalovic, Joseph Khoury, Ahmad F. Taha, Elias Bou-Harb
NOMS4
2025 Internet-Wide Analysis, Characterization, and Family Attribution of IoT Malware: A Comprehensive Longitudinal Study
abstract
This study presents a large-scale empirical analysis of real-life Internet-of-Things (IoT) malware by conducting a comprehensive analysis of 160,000 malicious executables detected by specialized IoT honeypots over five years. Our findings contribute to improving the knowledge of IoT malware characteristics and inter-relationships, which in return, contribute towards strengthening cybersecurity measures for IoT threat detection/mitigation. To achieve these goals, we leverage various malware analysis techniques to extract useful information from the executable files. Our analysis demonstrate that in contrast to non-IoT malware, we were able to extract unsolicited IP addresses and command strings from the majority of the analyzed IoT malware binaries using off-the-shelf de-obfuscation techniques/tools. Additionally, by correlating the extracted information and performing consequent similarity analysis using NLP-based features, we were able to reveal closely related samples with shared implementation across the adversarial infrastructure. Thus, contributing to labeling previously unseen/unknown IoT malware samples while uncovering emerging, possibly new variants. Finally, given such findings, we discuss the applications of a real-time IoT honeypot, which enables capturing real-time commands from malware-infected IoT devices while enabling timely and effective IoT-malware detection, analysis, labeling, and mitigation.
Sadegh Torabi, Dorde Klisura, Joseph Khoury, Elias Bou-Harb, Chadi Assi, Mourad Debbabi
IEEE Trans. Dependable Secur. Comput.4
2024 Don't, Stop, Drop, Pause: Forensics of CONtainer CheckPOINTs (ConPoint)
abstract
In the rapidly evolving landscape of cloud computing, containerization technologies such as Docker and Kubernetes have become instrumental in deploying, scaling, and managing applications. However, these containers pose unique challenges for memory forensics due to their ephemeral nature. As memory forensics is a crucial aspect of incident response, our work combats these challenges by developing a deeper understanding of the containers, leading to the development of a novel, scalable tool for container memory forensics. Through experimental and computational analyses, our work investigates the forensic capabilities of container checkpoints, which capture a container’s state at a specific moment in time. We introduce ConPoint, a tool created for the collection of these checkpoints. We focused on three primary research questions: What is the most forensically sound approach for checkpointing a container’s memory and filesystem?, How long does the volatile memory evidence reside in memory?, and How long does the checkpoint process take on average to complete? Our approach successfully captured checkpoints and retrieved artifacts generated at runtime from container checkpoints. We found that digital evidence in a container’s volatile memory can persist during idle states, yet gradually diminishes over time and is entirely lost when the container shuts down. Our experiments determined the average time for checkpointing a container to be 0.537 seconds by acquiring a total of (n = 45) checkpoints from containers running different databases. The proposed work demonstrates the pragmatic feasibility of implementing checkpointing as an overarching strategy for container memory forensics and incident response.
Taha Gharaibeh, Steven Seiden 0002, Mohamed Abouelsaoud, Elias Bou-Harb, Ibrahim M. Baggili
ARES4
2024 Characterizing and Analyzing LEO Satellite Cyber Landscape: A Starlink Case Study
abstract
Ushering into the ‘New Space Era’, characterized by a reduction in launch expenses and simultaneous proliferation of commercial and governmental entities involved, the prominence of Low Earth Orbit (LEO) satellite technology in the sphere of Internet connectivity has risen to the forefront. However, due to current limitations under the overarching principle of ‘security-through-obscurity’, few to no research efforts have shed light on the intricacies of these networks. To this end, this paper harnesses a multilayer empirical approach in an effort to conduct an exploratory characterization and scrutiny of the cybersecurity landscape of Starlink, the largest LEO network. Using our built-in arsenal of data feeds, composed of large dark IP addresses, passive measurement sensors, BGP collectors, coupled with publicly available sources, we unveil on the Starlink cyberspace (i) Internet-scale exploitations, (ii) illicit scanning events originating from 8,675 unique Starlink end-users, (iii) suspicious Port 0 and IKE scans, (iv) Mirai-based infections, (v) source address spoofing, (vi) 8,714 vulnerabilities ranging between medium and critical, and (vii) interesting RTBH announcements associated with possible mitigation techniques.
Nasser Tieby, Joseph Khoury, Elias Bou-Harb
ICC3
2024 DCPsolver: Enhancing DNS Cache Poisoning Defenses with Resolver-Based SmartNICs
abstract
The Domain Name System (DNS) is a critical component of the internet infrastructure, yet it remains vulnerable to DNS Cache Poisoning (DCP) attacks, which can mislead users by redirecting them to malicious websites. Although various countermeasures have been proposed to safeguard DNS infrastructure, they predominantly address off-path attacks, leaving DNS resolvers susceptible to on-path attacks, and often face challenges in widespread adoption due to impractical deployment requirements. To address these shortcomings, we present a novel solution leveraging Smart Network Interface Cards (SmartNICs) to detect and mitigate both on-path and off-path DCP attacks in real-time. Our approach operates exclusively within recursive DNS resolvers, independent of existing DNS protocols and infrastructure, thereby enhancing practicality and deployability. The proposed methodology was rigorously evaluated using a comprehensive, recently-released dataset of DCP attacks. Results demonstrate that the SmartNIC-based solution accurately identifies and mitigates both attack types while efficiently leveraging hardware resources. Indeed, the approach detailed herein offers a practical, effective, and efficient solution for enhancing DNS security without the need for a widespread infrastructure overhaul.
Bharath Kollanur, Kurt Friday, Elias Bou-Harb
NCA3
2024 Jbeil: Temporal Graph-Based Inductive Learning to Infer Lateral Movement in Evolving Enterprise Networks
abstract
Lateral Movement (LM) is one of the core stages of advanced persistent threats which continues to compromise the security posture of enterprise networks at large. Recent research work have employed Graph Neural Network (GNN) techniques to detect LM in intricate networks. Such approaches employ transductive graph learning, where fixed graphs with full nodes' visibility are employed in the training phase, along with ingesting benign data. These two assumptions in real-world setups (i) do not take into consideration the evolving nature of enterprise networks where dynamic features and connectivity prevail among hosts, users, virtualized environments, and applications, and (ii) hinder the effectiveness of detecting LM by solely training on normal data, especially given the evasive, stealthy, and benign-like behaviors of contemporary malicious maneuvers. Additionally, (iii) complex networks typically do not have the entire visibility of their run-time network processes, and if they do, they often fall short in dynamically tracking LM due to latency issues with passive data analysis.To this end, this paper proposes Jbeil, a data-driven framework for self-supervised deep learning on evolving networks represented as sequences of authentication timed events. The premise of the work lies in applying an encoder on a continuous-time evolving graph to produce the embedding of the visible graph nodes for each time epoch, and a decoder that leverages these embeddings to perform LM link prediction on unseen nodes. Additionally, we enclose a threat sample augmentation mechanism within Jbeil to ensure a well-informed notion on advanced LM attacks. We evaluate Jbeil using authentication timed events from the Los Alamos network which achieves an AUC score of 99.73% and a recall score of 99.25% in predicting LM paths, even when 30% of the nodes/edges are not present in the training phase. Additionally, we assess different realistic attack scenarios and demonstrate the potential of Jbeil in predicting LM paths with an AUC score of 99% in its inductive and transductive settings, out performing the state-of-the-art by a significant margin.
Joseph Khoury, Dorde Klisura, Hadi Zanddizari, Gonzalo De La Torre Parra, Peyman Najafirad, Elias Bou-Harb
SP6
2024 On DGA Detection and Classification Using P4 Programmable Switches
Ali AlSabeh, Kurt Friday, Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb
Comput. Secur.5
2024 P4BS: Leveraging Passive Measurements From P4 Switches to Dynamically Modify a Router's Buffer Size
abstract
The performance of networked applications can be dramatically impacted by the size of the buffer at the bottleneck router. Shallow buffers may increase packet losses and decrease link utilization, while deep buffers may increase the queueing delays for latency-sensitive flows. Operators nowadays configure large buffers statically without considering the characteristics of flows or dynamic traffic patterns. This paper presents P4BS, a system that dynamically modifies the buffer size of a legacy router. P4BS leverages programmable switches as passive instruments to measure various metrics that are vital when deciding on buffer size. The measured metrics include the number of long-lived flows and their round-trip times, the packet loss rates, and the queueing delays. Using these measurements, the programmable switch sequentially searches for a buffer size that minimizes the queueing delays and the packet loss rates. The system was implemented on a Tofino hardware switch and the system was tested on a wide range of network scenarios. The results show improvements in the quality of service of various applications including Web browsing, video streaming, and voice over IP.
Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb
IEEE Trans. Netw. Serv. Manag.3
2024 EV Charging Infrastructure Discovery to Contextualize Its Deployment Security
abstract
Electric Vehicle Charging Stations (EVCSs) have been shown to be susceptible to remote exploitation due to manufacturer-induced vulnerabilities, demonstrated by recent attacks on this ecosystem. What is more alarming is that compromising these high-wattage IoT systems can be leveraged to perform coordinated oscillatory load attacks against the power grid which could lead to the instability of this critical infrastructure. In this paper, we investigate a previously sidelined aspect of EVCS security. We analyze the deployment security of EVCSs and highlight operator-induced vulnerabilities rendering the ecosystem exposed to remote intrusions. We create an advanced discovery technique that leverages Web interface artifacts to dynamically discover new charging station vendors. As a result, we uncover 33,320 charging station management systems in the wild. Consequently, we study the deployment security of the charging stations and identify that 28,046 EVCSs were found to be vulnerable to eavesdropping, and around 24% of the studied EVCSs are deployed with default configurations exposing the ecosystem to a Mirai-like attack vector. Aligned with this finding, we discover that the EVCS ecosystem has been targeted by nefarious IoT malware such as Mirai and its variants. This demonstrates that further security measures should be implemented by vendors and operators to ensure the security of this vital ecosystem. Consequently, we provide a comprehensive recommendation for securing the deployment of EVCSs.
Khaled Sarieddine, Mohammad Ali Sayed, Chadi Assi, Ribal Atallah, Sadegh Torabi, Joseph Khoury, Morteza Safaei Pour, Elias Bou-Harb
IEEE Trans. Netw. Serv. Manag.8
2024 Guest Editorial: Special section on Networks, Systems, and Services Operations and Management Through Intelligence
abstract
Machine Learning (ML) and Artificial Intelligence (AI) can harness the immense amount of operational data from clouds to services, to social and communication networks. In the era of data science and connected devices of all varieties, Intelligence have found ways to improve operations and management of next generation networks, systems, and services. Further research is therefore needed to understand and improve the potential and suitability of ML/AI in the context of network, system, and service operations and management. This will provide deeper understanding and better decision making based on largely collected and available operational and management data. It will also present opportunities for improving ML/AI algorithms on aspects such as reliability, dependability, and scalability, as well as demonstrate the benefits of these methods in control and management systems. Moreover, there is an opportunity to define novel platforms that can harness the vast operational data and advance ML/AI algorithms to drive management decisions in open and highly programmable networks, clouds, and data centers.
Nur Zincir-Heywood, Robert Birke, Elias Bou-Harb, Takeru Inoue, Neeraj Kumar 0001, Hanan Lutfiyya, Deepak Puthal, Abdallah Shami, Natalia Stakhanova
IEEE Trans. Netw. Serv. Manag.3
2023 An Unbiased Transformer Source Code Learning with Semantic Vulnerability Graph
abstract
Over the years, open-source software systems have become prey to threat actors. Even highly-adopted software has been crippled by unforeseeable attacks, leaving millions of devices exposed. Even as open-source communities act quickly to patch the breach, code vulnerability screening should be an integral part of agile software development from the beginning. Unfortunately, current vulnerability screening techniques are ineffective at identifying novel vulnerabilities or providing developers with code vulnerability and classification. Furthermore, the datasets used for vulnerability learning often exhibit distribution shifts from the real-world testing distribution due to novel attack strategies deployed by adversaries and as a result, the machine learning model’s performance may be hindered or biased. To address these issues, we propose a joint interpolated multitasked unbiased vulnerability classifier comprising a transformer "RoBERTa" and graph convolution neural network (GCN). We present a training process utilizing a semantic vulnerability graph (SVG) representation from source code, created by integrating edges from a sequential flow, control flow, and data flow, as well as a novel flow dubbed Poacher Flow (PF). Poacher flow edges reduce the gap between dynamic and static program analysis and handle complex long-range dependencies. Moreover, our approach reduces biases of classifiers regarding unbalanced datasets by integrating Focal Loss objective function along with SVG. Remarkably, experimental results show that our classifier outperforms state-of-the-art results on vulnerability detection with fewer false negatives and false positives. After testing our model across multiple datasets, it shows an improvement of at least 2.41% and 18.75% in the best-case scenario. Evaluations using N-day program samples demonstrate that our proposed approach achieves a 93% accuracy and was able to detect 4, zero-day vulnerabilities from popular GitHub repositories. Our code and data are available at https://github.com/pial08/SemVulDet
Nafis Tanveer Islam, Gonzalo De La Torre Parra, Dylan Manuel, Elias Bou-Harb, Peyman Najafirad
EuroS&P4
2023 Effective DGA Family Classification Using a Hybrid Shallow and Deep Packet Inspection Technique on P4 Programmable Switches
abstract
Domain Generation Algorithms (DGAs) are one of the most effective strategies for malware to obtain a connection with the adversary's Command and Control (C2) server. Moreover, the growing number of DGA families makes it increasingly challenging for defense strategies to promptly identify the DGA family behind a given compromise. State-of-the-art high-dimensional DGA detection models perform poorly in such multiclass classification scenarios because their domain name-based features fail to distinguish between DGA families. To this extent, this paper proposes a novel framework that harnesses the flexibility, per-packet granularity, and Terabits per second (Tbps) processing capabilities of P4 Programmable Data Plane (PDP) switches to swiftly and accurately classify DGA families. In particular, the P4 PDP switch is leveraged to extract a combination of unique network heuristics and domain name features through shallow and Deep Packet Inspection (DPI) with minimal throughput reduction. Such collected features cannot be tracked on commodity hardware without significantly degrading the throughput in high-speed networks, nor on traditional layer 2/3 switches due to their limited and fixed functionalities. We crawled hundreds of Gigabytes (GBs) of malware samples from different sources to obtain instances of 50 DGA families and show that the proposed approach can promptly classify each family with high accuracy. Such a reliable multiclass classification enables the immediate halting of malicious communications while allowing network operators to initiate appropriate mitigation, incident management, and provisioning strategies.
Ali AlSabeh, Kurt Friday, Jorge Crichigno, Elias Bou-Harb
ICC4
2023 P4CCI: P4-Based Online TCP Congestion Control Algorithm Identification for Traffic Separation
abstract
Congestion Control Algorithms (CCAs) regulate the sending rates of hosts to avoid congestion in the network. Studies have shown that when flows belonging to different CCAs coexist on the same link, their shares on that link are significantly different. If the CCAs of active flows can be determined on live traffic, then flows belonging to the same CCA can be allocated into a dedicated queue. Unfortunately, identifying the CCA at line rate is not straightforward since the CCA is not advertised in the header fields of a packet. Moreover, with Gigabits per second (Gbps) traffic crossing a network, analyzing each packet to infer the CCA is not possible, especially with general-purpose CPUs. This paper proposes P4CCI, a system that detects the CCA of a flow at line rate by leveraging Programmable Data Planes (PDP). The PDP computes and extracts the flow's bytes-in-flight and sends them to a Deep Learning model for classification. Once classified, the flows are allocated into dedicated queues based on their CCA type. The system was implemented and tested on real hardware that uses Intel's Tofino ASIC. The experiments were executed on traffic provided by CAIDA. Results show that P4CCI can detect the CCAs with high accuracy. Furthermore, the performance of the network is greatly improved when the flows are separated by their CCAs.
Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb
ICC3
2023 RPM: Ransomware Prevention and Mitigation Using Operating Systems' Sensing Tactics
abstract
Ransomware, an extortion type of malware, continues to create havoc targeting critical infrastructure and organizations at large, causing an estimated $20 Billion in direct and collateral damages in 2022. While significant efforts from both academia and industry are being pledged to address this debilitating and disrupting phenomena, the ransomware pandemic continues to expand rapidly in frequency, spread and stealthiness. To this end, in this work, we propose RPM, a Ransomware Prevention and Mitigation scheme. RPM is rooted in the proactive analysis of operating systems' API artifacts through the exploitation of a neat observation related to ransomware behavior, namely, activities generated prior to the actual execution of the malicious payloads. RPM employs OS-centric process hooking tactics to develop an offensive approach leveraging such sensing activities. To demonstrate the effectiveness of RPM, we empirically evaluated it using 100 of the most prominent ransomware samples. The results demonstrate very motivating accuracy metrics with low system footprint, asserting the rationale of the proposed scheme. We posture RPM as a strong step towards proactive mitigation, which aims at complimenting ongoing ransomware thwarting efforts.
Ricardo Misael Ayala Molina, Elias Bou-Harb, Sadegh Torabi, Chadi Assi
ICC2
2023 Unraveling Network-Based Pivoting Maneuvers: Empirical Insights and Challenges
Martin Husák, Shanchieh Jay Yang, Joseph Khoury, Dorde Klisura, Elias Bou-Harb
ICDF2C (2)5
2023 ChargePrint: A Framework for Internet-Scale Discovery and Security Analysis of EV Charging Management Systems
Tony Nasr, Sadegh Torabi, Elias Bou-Harb, Claude Fachkha, Chadi Assi
NDSS3
2023 Data-Centric Machine Learning Approach for Early Ransomware Detection and Attribution
abstract
Researchers have proposed a wide range of ransomware detection and analysis schemes. However, most of these efforts have focused on older families targeting Windows 7/8 systems. Hence there is a critical need to develop efficient solutions to tackle the latest threats, many of which may have relatively fewer samples to analyze. This paper presents a machine learning (ML) framework for early ransomware detection and attribution. The solution pursues a data-centric approach which uses a minimalist ransomware dataset and implements static analysis using portable executable (PE) files. Results for several ML classifiers confirm strong performance in terms of accuracy and zero-day threat detection.
Aldin Vehabovic, Hadi Zanddizari, Nasir Ghani, Farooq Shaikh, Elias Bou-Harb, Morteza Safaei Pour, Jorge Crichigno
NOMS5
2023 Helium-based IoT Devices: Threat Analysis and Internet-scale Exploitations
abstract
With the explosive growth of resource-constrained smart devices and the widespread deployment of Internet-of-Things (IoT) devices, there is an ever-increasing demand for low-energy and cost-effective wireless communication solutions to serve a wide variety of systems and processes. To this end, blockchain-enabled Helium devices were conceived to enable Internet services and to support third-party IoT devices. This decentralized paradigm allows individuals and entities to freely engage, monetize and deploy wireless Helium hotspots, offering Internet coverage through piggy-backing packets via their existing network and Internet infrastructure (e.g., fiber optics at home). Currently, there are close to 1M operational Helium devices deployed in 189 countries, which are owned by 425K accounts. Given this evolving paradigm, in this paper, we take a first step to explore the plausible attack vectors which could potentially impact the confidentiality, integrity, and availability of such Helium hotspots. Along this vein, we then scrutinize 2.9 TB of one-way unsolicited Internet traffic arriving at 0.5M monitored dark IP addresses to identify 869,822 darknet events pertained to 6K Helium hotspots (as infected devices and DoS victims). By further leveraging active and passive methodologies coupled with public exploitation databases, we uncover medium to critical severity vulnerabilities attributed to 62K online Helium hotspots.
Veronica Rammouz, Joseph Khoury, Dorde Klisura, Morteza Safaei Pour, Mostafa Safaei Pour, Claude Fachkha, Elias Bou-Harb
WiMob7
2023 A Comprehensive Survey of Recent Internet Measurement Techniques for Cyber Security
abstract
As the Internet has transformed into a critical infrastructure, society has become more vulnerable to its security flaws. Despite substantial efforts to address many of these vulnerabilities by industry, government, and academia, cyber security attacks continue to increase in intensity, diversity, and impact. Thus, it becomes intuitive to investigate the current cyber security threats, assess the extent to which corresponding defenses have been deployed, and evaluate the effectiveness of risk mitigation efforts. Addressing these issues in a sound manner requires large-scale empirical data to be collected and analyzed via numerous Internet measurement techniques. Although such measurements can generate comprehensive and reliable insights, doing so encompasses complex procedures involving the development of novel methodologies to ensure accuracy and completeness. Therefore, a systematic examination of recently developed Internet measurement approaches for cyber security must be conducted to enable thorough studies that employ several vantage points, correlate multiple data sources, and potentially leverage past successful techniques for more recent issues. Unfortunately, performing such an examination is challenging, as the literature is highly scattered. In large part, this is due to each research effort only focusing on a small portion of the many constituent parts of the Internet measurement domain. Moreover, to the best of our knowledge, no studies have offered an in-depth examination of this critical research domain in order to promote future advancements. To bridge these gaps, we explore all pertinent facets of utilizing Internet measurement techniques for cyber security, ranging from threats within specific application domains to threats themselves. We provide a taxonomy of cyber security-related Internet measurement studies across two dimensions. One dimension relates to the many vertical layers (and components) of the Internet ecosystem, while the other relates to internal normal functions vs. the negative impact of external parties in the Internet and physical world. A comprehensive comparison of the gathered studies is also offered in terms of measurement technique, scope, measurement size, vantage size, and the analysis approach that was leveraged. Finally, a discussion of the roadblocks to performing effective Internet measurements and possible future research directions is elaborated.
Morteza Safaei Pour, Christelle Nader, Kurt Friday, Elias Bou-Harb
Comput. Secur.4
2023 Guest Editorial: Special Section on Machine Learning and Artificial Intelligence for Managing Networks, Systems, and Services - Part II
abstract
Machine learning and artificial intelligence can harness the immense stream of operational data from clouds, to services, to social and communication networks. In the era of big data and connected devices of all varieties, machine learning and artificial intelligence have found ways to improve operations and management of information technology and communications.
Nur Zincir-Heywood, Robert Birke, Elias Bou-Harb, Giuliano Casale, Khalil El-Khatib, Takeru Inoue, Neeraj Kumar 0001, Hanan Lutfiyya, Deepak Puthal, Abdallah Shami, Natalia Stakhanova, Farhana Zulkernine
IEEE Trans. Netw. Serv. Manag.3
2022 A Near Real-Time Scheme for Collecting and Analyzing IoT Malware Artifacts at Scale
abstract
The chronic proliferation of Internet of Things (IoT) botnet malware activities coupled with an unprecedented rise in security vulnerabilities convene a new world of opportunities for perpetrators and unveil a new set of hurdles in deriving relevant IoT malware intelligence. Such shortfall within the IoT paradigm exacerbates the capabilities for largely identifying the prevailing IoT malware threats, the origin of the IoT attacks, as well as, the security deficit associated with the IoT paradigm. Previous work has vastly studied IoT malware activities in the wild but has not profiled at a large scale malicious activities to collect in near real-time central IoT artifacts much-needed to understand and eventually elevate the security posture of the IoT ecosystem.
Joseph Khoury, Morteza Safaei Pour, Elias Bou-Harb
ARES3
2022 EVOLIoT: A Self-Supervised Contrastive Learning Framework for Detecting and Characterizing Evolving IoT Malware Variants
abstract
Recent years have witnessed the emergence of new and more sophisticated malware targeting the Internet of Things. Moreover, the public release of the source code of popular malware families such as Mirai has spawned diverse variants, making it harder to disambiguate their ownership, lineage, and correct label. Such a rapidly evolving landscape makes it also harder to deploy and generalize effective learning models against retired, updated, and/or new threat campaigns. In this paper, we present EVOLIoT, a novel approach aiming at combating "concept drift" and the limitations of inter-family IoT malware classification by detecting drifting IoT malware families and understanding their diverse evolutionary trajectories. We introduce a robust and effective contrastive method that learns and compares semantically meaningful representations of IoT malware binaries and codes without the need for expensive target labels. We find that the evolution of IoT binaries can be used as an augmentation strategy to learn effective representations to contrast (dis)similar variant pairs. We discuss the impact and findings of our analysis and present several evaluation studies to highlight the tangled relationships of IoT malware, as well as the efficiency of our contrastively learned feature vectors in preserving semantics and reducing out-of-vocabulary size in cross-architecture IoT malware binaries.
Mirabelle Dib, Sadegh Torabi, Elias Bou-Harb, Nizar Bouguila, Chadi Assi
AsiaCCS3
2022 An attentive interpretable approach for identifying and quantifying malware-infected internet-scale IoT bots behind a NAT
abstract
The explosive growth of the Internet-of-Things (IoT) paradigm has brought the rise of malicious activity targeting the Internet. Indeed, the lack of basic security protocols and measures in IoT devices is allowing attackers to use exploited Internet-scale IoT devices to organize malicious botnets, and cause significant damage to the Internet through Denial of Service (DoS) attacks, illicit scraping, and cryptojacking attacks. Such IoT botnets can be Internet-facing, or can also be deployed behind Network Address Translation (NAT) gateways that provide anonymity to the exploited bots. In this paper, we aim at detecting compromised IoT bots behind NAT gateways which could possibly generate malicious activities towards the Internet by leveraging large-scale macroscopic one-way darknet data. To the best of our knowledge, we are among the first to explore the capabilities of attentive interpretable tabular transformers to capture the nature of such nodes operating on one-way network traffic. Our results, which employed 2.6GB of darknet data, show that our approach was able to classify malware-infected NATed IoT bots with an accuracy of 93%, outperforming the state-of-the-art machine learning (ML) approaches. Additionally, we were able to infer around 4 million Internet-scale Mirai-infected NATed IoT bots and 16,871 unique NATed IP addresses. Results from this work put forward interesting future work in the area of network traffic analysis of NATed IoT bots for better Internet security, while highlighting the need for addressing the notions of attention and interpretability.
Christelle Nader, Elias Bou-Harb
CF2
2022 INC: In-Network Classification of Botnet Propagation at Line Rate
Kurt Friday, Elie F. Kfoury, Elias Bou-Harb, Jorge Crichigno
ESORICS (1)3
2022 Interpretable Federated Transformer Log Learning for Cloud Threat Forensics
Gonzalo De La Torre Parra, Luis Selvera, Joseph Khoury, Hector Irizarry, Elias Bou-Harb, Peyman Najafirad
NDSS5
2022 HoneyComb: A Darknet-Centric Proactive Deception Technique For Curating IoT Malware Forensic Artifacts
abstract
Conventional IoT honeypots are known to suffer from scalability and management issues, while accumulating stringent costs. Further, their passive nature hinders the wide-scale gathering of much-needed IoT malware artifacts, impeding their measurements, analysis, and ultimately their use to infer and react to IoT maliciousness at large. To this end, in this work, we introduce HoneyComb, a proactive deception technique to curate IoT malware forensics by leveraging IoT scans captured on the darknet (i.e., Internet telescope). HoneyComb is built on the premise that we can position a large darknet network (i.e., comprising of 16.7 million IPs) as a large honeypot to interact with malware-infected IoT devices at scale. Such a large vantage point is capable of offering an incomparable hefty look into the IoT cyber security posture compared to the typical, much-restricted, currently-available IoT honeypots. In essence, the inferred IoT scans from the darknet along with the existing discrepancy in the validation algorithms of IoT malware stateless scanning modules, enable HoneyComb to initiate crafted deceiving packets (i.e., TCP SYN-ACK packets) to delude and interconnect with malware-infected IoT devices in the wild. During 48 hours of empirical measurements, the proposed scheme logged 1,432,518 interactions originating from 37,323 malware-infected IoT devices worldwide. Additionally, our findings revealed intriguing insights concerning the propagation behavior of IoT malware where 11,340 infected devices delivered the malware binaries using 1,398 unique URLs, whereas 2,114 used HexString dumping to drop their binaries, while the rest reported sensitive information (e.g., credentials) to their servers. Finally, while we observe that newly emerged IoT malware such as ZHTRAP is more capable in the takeover process due to its innovative techniques and offensive competencies, we frame HoneyComb as a complementary scheme, which would aid in addressing a number of evolving IoT-centric security endeavours, including large-scale malware attribution and C&C takedowns.
Morteza Safaei Pour, Joseph Khoury, Elias Bou-Harb
NOMS3
2022 A Learning Methodology for Line-Rate Ransomware Mitigation with P4 Switches
Kurt Friday, Elias Bou-Harb, Jorge Crichigno
NSS2
2022 A survey on security applications of P4 programmable switches and a STRIDE-based vulnerability assessment
Ali AlSabeh, Joseph Khoury, Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb
Comput. Networks5
2022 Power jacking your station: In-depth security analysis of electric vehicle charging station management systems
Tony Nasr, Sadegh Torabi, Elias Bou-Harb, Claude Fachkha, Chadi Assi
Comput. Secur.3
2022 Inferring and Investigating IoT-Generated Scanning Campaigns Targeting a Large Network Telescope
abstract
The analysis of recent large-scale cyber attacks, which leveraged insecure Internet of Things (IoT) devices to perform malicious activities on the Internet, highlighted the rise of IoT-tailored malware/botnets. These malware propagate by scanning the Internet for vulnerable, exploitable IoT devices that could be utilized for further malicious activities. In this article, we devise a multi-level methodology to investigate Internet-scale reconnaissance activities generated by infected IoT devices. We leverage theShodanIoT search engine and over 6TB of passive network traffic from a large network telescope (darknet) to infer compromised IoT devices and characterize the generated scanning campaigns. The results highlight a distinctive characteristic of IoT malware/botnets, represented by the targeted ports/services over the analysis interval. Furthermore, while these ports/services are mainly associated with well-known IoT malware/botnets (e.g.,MiraiandSatori), we uncovered newly targeted ports, which indicate emerging IoT malware/botnet. Finally, by comparing two instances of analyzed IoT-generated scanning campaigns, we highlight the persistence and evolution of IoT malware/botnets (e.g.,ADB.MinerandFbot), which exploit existing, and in some cases, possibly new vulnerabilities.
Sadegh Torabi, Elias Bou-Harb, Chadi Assi, ElMouatez Billah Karbab, Amine Boukhtouta, Mourad Debbabi
IEEE Trans. Dependable Secur. Comput.2
2022 On Ransomware Family Attribution Using Pre-Attack Paranoia Activities
abstract
Ransomware attacks are among the most disruptive cyber threats, causing significant financial losses while impacting productivity, accessibility, and reputation. Despite their end goals (encryption/locking), ransomware are often designed to evade detection by executing a series of pre-attack API calls, namely “paranoia” activities, for determining a suitable execution environment. In this work, we present a first-of-a-kind effort to utilize such paranoia activities for characterizing ransomware distinguishable behaviors. To this end, we draw-upon more than 3K samples from recent/prominent ransomware families to fingerprint their uniquely leveraged paranoia activities. Specifically, by leveraging techniques rooted in Natural Language Processing (NLP) such as Occurrence of Words (OoW), we model ransomware-generated evasion API calls while tailoring various machine and deep learning algorithms to perform ransomware classification. The thoroughly conducted evaluations demonstrate the effectiveness of the implemented approach, with the Random Forest (RF) and OoW techniques producing an optimal classification accuracy (94.92%). The insights/findings from this work not only shed light on contemporary ransomware-specific evasion methods, but also (i) indicates that such tactics could be employed effectively as features for ransomware family attribution while (ii) laying the foundation for implementing proactive and portable countermeasures for further ransomware attack detection/mitigation by solely utilizing ransomware-generated paranoia activities.
Ricardo Misael Ayala Molina, Sadegh Torabi, Khaled Sarieddine, Elias Bou-Harb, Nizar Bouguila, Chadi Assi
IEEE Trans. Netw. Serv. Manag.4
2022 Guest Editorial: Special Issue on Machine Learning and Artificial Intelligence for Managing Networks, Systems, and Services - Part I
abstract
Machine learning and artificial intelligence can harness the immense stream of operational data from clouds, to services, to social and communication networks. In the era of big data and connected devices of all varieties, machine learning and artificial intelligence have found ways to improve operations and management of information technology and communications.
Nur Zincir-Heywood, Robert Birke, Elias Bou-Harb, Giuliano Casale, Khalil El-Khatib, Takeru Inoue, Neeraj Kumar 0001, Hanan Lutfiyya, Deepak Puthal, Abdallah Shami, Natalia Stakhanova, Farhana Zulkernine
IEEE Trans. Netw. Serv. Manag.3
2021 Sanitizing the IoT Cyber Security Posture: An Operational CTI Feed Backed up by Internet Measurements
abstract
The Internet-of-Things (IoT) paradigm at large continues to be compromised, hindering the privacy, dependability, security, and safety of our nations. While the operational security communities (i.e., CERTS, SOCs, CSIRT, etc.) continue to develop capabilities for monitoring cyberspace, tools which are IoT-centric remain at its infancy. To this end, we address this gap by innovating an actionable Cyber Threat Intelligence (CTI) feed related to Internet-scale infected IoT devices. The feed analyzes, in near real-time, 3.6TB of daily streaming passive measurements ( ≈ 1M pps) by applying a custom-developed learning methodology to distinguish between compromised IoT devices and non-IoT nodes, in addition to labeling the type and vendor. The feed is augmented with third party information to provide contextual information. We report on the operation, analysis, and shortcomings of the feed executed during an initial deployment period. We make the CTI feed available for ingestion through a public, authenticated API and a front-end platform.
Morteza Safaei Pour, Dylan Watson, Elias Bou-Harb
DSN3
2021 Dynamic Router's Buffer Sizing using Passive Measurements and P4 Programmable Switches
abstract
The router's buffer size imposes significant impli-cations on the performance of the network. Network operators nowadays configure the router's buffer size manually and stati-cally. They typically configure large buffers that fill up and never go empty, increasing the Round-trip Time (RTT) of packets significantly and decreasing the application performance. Few works in the literature dynamically adjust the buffer size, but are implemented only in simulators, and therefore cannot be tested and deployed in production networks with real traffic. Previous work suggested setting the buffer size to the Bandwidth-delay Product (BDP) divided by the square root of the number of long flows. Such formula is adequate when the RTT and the number of long flows are known in advance. This paper proposes a system that leverages programmable switches as passive instruments to measure the RTT and count the number of flows traversing a legacy router. Based on the measurements, the programmable switch dynamically adjusts the buffer size of the legacy router in order to mitigate the unnecessary large queuing delays. Results show that when the buffer is adjusted dynamically, the RTT, the loss rate, and the fairness among long flows are enhanced. Additionally, the Flow Completion Time (FCT) of short flows sharing the queue is greatly improved. The system can be adopted in campus, enterprise, and service provider networks, without the need to replace legacy routers.
Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb, Gautam Srivastava 0001
GLOBECOM3
2021 A Multidimensional Network Forensics Investigation of a State-Sanctioned Internet Outage
abstract
In November 2019, the government of Iran enforced a week-long total Internet blackout that prevented the majority of Internet connectivity into and within the nation. This work elaborates upon the Iranian Internet blackout by characterizing the event through Internet-scale, near realtime network traffic measurements. Beginning with an investigation of compromised machines scanning the Internet, nearly 50 TB of network traffic data was analyzed. This work discovers 856,625 compromised IP addresses, with 17,182 attributed to the Iranian Internet space. By the second day of the Internet shut down, these numbers dropped by 18.46% and 92.81%, respectively. Empirical analysis of the Internet-of-Things (IoT) paradigm revealed that over 90% of compromised Iranian hosts were fingerprinted as IoT devices, which saw a significant drop throughout the shutdown (96.17% decrease by the blackout's second day). Further examination correlates BGP reachability metrics and related data with geolocation databases to statistically evaluate the number of reachable Iranian ASNs (dropping from approximately 1100 to under 200 reachable networks). In-depth investigation reveals the top affected ASNs, providing network forensic evidence of the longitudinal unplugging of such key networks. Lastly, the impact's interruption of the Bitcoin cryptomining market is highlighted, disclosing a massive spike in unsuccessful (i.e., pending) transactions. When combined, these network traffic measurements provide a multidimensional perspective of the Iranian Internet shutdown.
Antonio Mangino, Elias Bou-Harb
IWCMC2
2021 A Multi-Dimensional Deep Learning Framework for IoT Malware Classification and Family Attribution
abstract
The emergence of Internet of Things malware, which leverages exploited IoT devices to perform large-scale cyber attacks (e.g., Mirai botnet), is considered as a major threat to the Internet ecosystem. To mitigate such threat, there is an utmost need for effective IoT malware classification and family attribution, which provide essential steps towards initiating attack mitigation/prevention countermeasures. In this paper, motivated by the lack of sophisticated malware obfuscation in the implementation of IoT malware, we utilize features extracted from strings- and image-based representations of the executable binaries to propose a novel multi-dimensional classification approach using Deep Learning (DL) architectures. To this end, we analyze more than 70,000 recently detected IoT malware samples. Our in-depth experiments with four prominent IoT malware families highlight the significant accuracy of the approach (99.78%), which outperforms conventional single-level classifiers. Additionally, we utilize our IoT-tailored approach for labeling newly detected “unknown” malware samples, which were mainly attributed to a few predominant families. Finally, this work contributes to the security of future networks (e.g., 5G) through the implementation of effective tools/techniques for timely IoT malware classification, and attack mitigation.
Mirabelle Dib, Sadegh Torabi, Elias Bou-Harb, Chadi Assi
IEEE Trans. Netw. Serv. Manag.3
2020 Exploiting Ransomware Paranoia For Execution Prevention
abstract
Ransomware attacks cost businesses more than $75 billion/year, and it is predicted to cost $6 trillion/year by 2021. These numbers demonstrate the havoc produced by ransomware on a large number of sectors and urge security researches to tackle it. Several ransomware detection approaches have been proposed in the literature that interchange between static and dynamic analysis. Recently, ransomware attacks were shown to fingerprint the execution environment before they attack the system to counter dynamic analysis. In this paper, we exploit the behavior of contemporary ransomware to prevent its attack on real systems and thus avoid the loss of any data. We explore a set of ransomware-generated artifacts that are launched to sniff the surrounding. Furthermore, we design, develop, and evaluate an approach that monitors the behavior of a program by intercepting the called Windows APIs. Consequently, we determine in real-time if the program is trying to inspect its surrounding before the attack, and abort it immediately prior to the initiation of any malicious encryption or locking. Through empirical evaluations using real and recent ransomware samples, we study how ransomware and benign programs inspect the environment. Additionally, we demonstrate how to prevent ransomware with a low false positive rate. We make the developed approach available to the research community at large through GitHub to strongly promote cyber security defense operations and for wide-scale evaluations and enhancements.
Ali AlSabeh, Haïdar Safa, Elias Bou-Harb, Jorge Crichigno
ICC3
2020 Offloading Media Traffic to Programmable Data Plane Switches
abstract
According to estimations, approximately 80% of Internet traffic represents media traffic. Much of it is generated by end users communicating with each other (e.g., voice, video sessions). A key element that permits the communication of users that may be behind Network Address Translation (NAT) is the relay server. This paper presents a scheme for offloading media traffic from relay servers to programmable switches. The proposed scheme relies on the capability of a P4 switch with a customized parser to de-encapsulate and process packets carrying media traffic. The switch then applies multiple switch actions over the packets. As these actions are simple and collectively emulate a relay server, the scheme is capable of moving relay functionality to the data plane operating at terabits per second. Performance evaluations show that the proposed scheme not only produces optimal results regarding Quality of Service (QoS) parameters (no packet loss, minimum delay, negligible delay variation, high Mean Opinion Score) but also scales much better than current solutions. Evaluations conducted with up to 35Gbps of media traffic or its equivalent of 400,000 simultaneous G.711 media sessions (limited only by the traffic generator rather than by the switch) show an ideal operation of the switch-based solution (using$\sim \text{l}$% of the switching capacity). In contrast, a relay server with a modern CPU model used for evaluations can process up to 900 simultaneous G.711 media sessions per core.
Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb
ICC3
2020 Towards a Unified In-Network DDoS Detection and Mitigation Strategy
abstract
Distributed Denial of Service (DDoS) attacks have terrorized our networks for decades, and with attacks now reaching 1.7 Tbps, even the slightest latency in detection and subsequent remediation is enough to bring an entire network down. Though strides have been made to address such maliciousness within the context of Software Defined Networking (SDN), they have ultimately proven ineffective. Fortunately, P4 has recently emerged as a platform-agnostic language for programming the data plane and in turn allowing for customized protocols and packet processing. To this end, we propose a first-of-a-kind P4-based detection and mitigation scheme that will not only function as intended regardless of the size of the attack, but will also overcome the vulnerabilities of SDN that have characteristically been exploited by DDoS. Moreover, it successfully defends against the broad spectrum of currently relevant attacks while concurrently emphasizing the Quality of Service (QoS) of legitimate end-users and overall SDN functionality. We demonstrate the effectiveness of the proposed scheme using a software programmable P4-switch, namely, the Behavorial Model version 2 (BMv2), showing its ability to withstand a variety of DDoS attacks in real-time via three use cases that can be generalized to most contemporary attack vectors. Specifically, the results substantiate that the mechanism herein is orders of magnitude faster than traditional polling techniques (e.g., NetFlow or sFlow) while minimizing the impact on benign traffic. We concur that the approach's design particularities facilitate seamless and scalable deployments in high-speed networks requiring line-rate functionality, in addition to being generic enough to be integrated into viable network topologies.
Kurt Friday, Elie F. Kfoury, Elias Bou-Harb, Jorge Crichigno
NetSoft3
2020 An emulation-based evaluation of TCP BBRv2 Alpha for wired broadband
Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb
Comput. Commun.4
2020 On data-driven curation, learning, and analysis for inferring evolving internet-of-Things (IoT) botnets in the wild
Morteza Safaei Pour, Antonio Mangino, Kurt Friday, Matthias Rathbun, Elias Bou-Harb, Farkhund Iqbal, Sagar Samtani, Jorge Crichigno, Nasir Ghani
Comput. Secur.5
2020 A Collaborative Security Framework for Software-Defined Wireless Sensor Networks
abstract
With the advent of 5G, technologies such as Software-Defined Networks (SDNs) and Network Function Virtualization (NFV) have been developed to facilitate simple programmable control of Wireless Sensor Networks (WSNs). However, WSNs are typically deployed in potentially untrusted environments. Therefore, it is imperative to address the security challenges before they can be implemented. In this paper, we propose a software-defined security framework that combines intrusion prevention in conjunction with a collaborative anomaly detection systems. Initially, an IPS-based authentication process is designed to provide a lightweight intrusion prevention scheme in the data plane. Subsequently, a collaborative anomaly detection system is leveraged with the aim of supplying a cost-effective intrusion detection solution near the data plane. Moreover, to correlate the true positive alerts raised by the sensor nodes in the network edge, a Smart Monitoring System (SMS) is exploited in the control plane. The performance of the proposed model is evaluated under different security scenarios as well as compared with other methods, where the model's high security and reduction of false alarms are demonstrated.
Christian Miranda, Georges Kaddoum, Elias Bou-Harb, Sahil Garg, Kuljeet Kaur
IEEE Trans. Inf. Forensics Secur.3
2020 A Big Data-Enabled Consolidated Framework for Energy Efficient Software Defined Data Centers in IoT Setups
abstract
The rapidly evolving industry standards and transformative advances in the field of Internet of Things are expected to create a tsunami of Big Data shortly. This, in turn, will demand real-time data analysis and processing from cloud computing platforms. A substantial part of the computing infrastructure is supported by large-scale and geographically distributed data centers (DCs). Nevertheless, these DCs impose a substantial cost in terms of rapidly growing energy consumption, which in turn adversely affects the environment. In this context, efficient resource utilization is seen as a potential candidate to enhance energy efficiency and minimize the load on the power sector. Nevertheless, in the majority of the public clouds, the resources are idle most of the time (i.e., under-utilized) as the load of the servers is unpredictable; thereby leading to a lofty increase in the energy utilization index and wastage of resources. Thus, it is highly essential to devise a precise and efficient resource management technique. Therefore, in this article, we leverage the advantages of software defined data centers (SDDCs) to minimize energy utilization levels. Precisely, SDDC refers to the process of programmatically abstracting the logical computing, network, and storage resources; and configuring them in real-time based on workload demands. In detail, we demonstrate the possibility of 1) designing a consolidated SDDC-based model to jointly optimize the process of virtual machine (VM) deployment and network bandwidth allocation for reduced energy consumption and guaranteed quality of service (QoS), particularly for heterogeneous computing infrastructures; 2) formulating a multiobjective optimization problem to deduce the optimal allocation of resources for both critical and noncritical applications; and 3) designing an efficient scheme based on heuristics to provide suboptimal results for the formulated multiobjective optimization problem. The proposed article presents a suboptimal approach based on first fit decreasing algorithm. Further, our empirical evaluations suggest that the proposed framework leads to almost 27.9% savings in terms of energy consumptions against the existing schemes with negligible QoS violations (approximately 0.33).
Kuljeet Kaur, Sahil Garg, Georges Kaddoum, Elias Bou-Harb, Kim-Kwang Raymond Choo
IEEE Trans. Ind. Informatics4
2019 Improving Borderline Adulthood Facial Age Estimation through Ensemble Learning
abstract
Achieving high performance for facial age estimation with subjects in the borderline between adulthood and non-adulthood has always been a challenge. Several studies have used different approaches from the age of a baby to an elder adult and different datasets have been employed to measure the mean absolute error (MAE) ranging between 1.47 to 8 years. The weakness of the algorithms specifically in the borderline has been a motivation for this paper. In our approach, we have developed an ensemble technique that improves the accuracy of underage estimation in conjunction with our deep learning model (DS13K) that has been fine-tuned on the Deep Expectation (DEX) model. We have achieved an accuracy of 68% for the age group 16 to 17 years old, which is 4 times better than the DEX accuracy for such age range. We also present an evaluation of existing cloud-based and offline facial age prediction services, such as Amazon Rekognition, Microsoft Azure Cognitive Services, How-Old.net and DEX.
Felix Anda, David Lillis, Aikaterini Kanta, Brett A. Becker, Elias Bou-Harb, Nhien-An Le-Khac, Mark Scanlon
ARES5
2019 Data-driven Curation, Learning and Analysis for Inferring Evolving IoT Botnets in the Wild
abstract
The insecurity of the Internet-of-Things (IoT) paradigm continues to wreak havoc in consumer and critical infrastructure realms. Several challenges impede addressing IoT security at large, including, the lack of IoT-centric data that can be collected, analyzed and correlated, due to the highly heterogeneous nature of such devices and their widespread deployments in Internet-wide environments. To this end, this paper explores macroscopic, passive empirical data to shed light on this evolving threat phenomena. This not only aims at classifying and inferring Internet-scale compromised IoT devices by solely observing such one-way network traffic, but also endeavors to uncover, track and report on orchestrated "in the wild" IoT botnets. Initially, to prepare the effective utilization of such data, a novel probabilistic model is designed and developed to cleanse such traffic from noise samples (i.e., misconfiguration traffic). Subsequently, several shallow and deep learning models are evaluated to ultimately design and develop a multi-window convolution neural network trained on active and passive measurements to accurately identify compromised IoT devices. Consequently, to infer orchestrated and unsolicited activities that have been generated by well-coordinated IoT botnets, hierarchical agglomerative clustering is deployed by scrutinizing a set of innovative and efficient network feature sets. By analyzing 3.6 TB of recent darknet traffic, the proposed approach uncovers a momentous 440,000 compromised IoT devices and generates evidence-based artifacts related to 350 IoT botnets. While some of these detected botnets refer to previously documented campaigns such as the Hide and Seek, Hajime and Fbot, other events illustrate evolving threats such as those with cryptojacking capabilities and those that are targeting industrial control system communication and control services.
Morteza Safaei Pour, Antonio Mangino, Kurt Friday, Matthias Rathbun, Elias Bou-Harb, Farkhund Iqbal, Khaled B. Shaban, Abdelkarim Erradi
ARES5
2019 A Flow-Based Entropy Characterization of a NATed Network and Its Application on Intrusion Detection
abstract
This paper presents a flow-based entropy characterization of a small/medium-sized campus network that uses network address translation (NAT). Although most networks follow this configuration, their entropy characterization has not been previously studied. Measurements from a production network show that the entropies of flow elements (external IP address, external port, campus IP address, campus port) and tuples have particular characteristics. Findings include: i) entropies may widely vary in the course of a day. For example, in a typical weekday, the entropies of the campus and external ports may vary from below 0.2 to above 0.8 (in a normalized entropy scale 0-1). A similar observation applies to the entropy of the campus IP address; ii) building a granular entropy characterization of the individual flow elements can help detect anomalies. Data shows that certain attacks produce entropies that deviate from the expected patterns; iii) the entropy of the 3-tuple {external IP, campus IP, campus port} is high and consistent over time, resembling the entropy of a uniform distribution's variable. A deviation from this pattern is an encouraging anomaly indicator; iv) strong negative and positive correlations exist between some entropy time-series of flow elements.
Jorge Crichigno, Elie F. Kfoury, Elias Bou-Harb, Nasir Ghani, Yasmany Prieto, Christian Vega Caicedo, Jorge E. Pezoa, David Torres
ICC3
2019 Theoretic derivations of scan detection operating on darknet traffic
Morteza Safaei Pour, Elias Bou-Harb
Comput. Commun.2
2019 Big Data Sanitization and Cyber Situational Awareness: A Network Telescope Perspective
abstract
This paper addresses the problems of data sanitization and cyber situational awareness by analyzing 910 GB of real Internet-scale traffic, which has been passively collected by monitoring close to 16.5 million darknet IP addresses from a /8 and a /13 network telescopes. First, the paper offers a novel probabilistic darknet preprocessing model, which aims at sanitizing darknet data to prepare it for effective use in the task of cyber threat intelligence generation. Such model has been engineered using a distributed multithreaded approach, rendering it operational and highly effective on darknet big data. Second, the paper further contributes by presenting an innovative approach to infer large-scale orchestrated probing campaigns by leveraging darknet data, for Internet cyber situational awareness. The approach uniquely reduces the dimensionality of such big data by utilizing its artifacts, instead of processing the actual raw data. This is accomplished by extracting and analyzing probing time series using formal methods rooted in Fourier transform and Kalman filtering. Thorough empirical evaluations indeed validate the accuracy and the performance of the proposed methods and techniques. We assert that the darknet sanitization model and the probing orchestration inference approach are of significant value, given their postulated highly applicable nature to the field of Internet measurements for cyber security in the era of big data.
Elias Bou-Harb, Martin Husák, Mourad Debbabi, Chadi Assi
IEEE Trans. Big Data1
2018 Assessing Internet-wide Cyber Situational Awareness of Critical Sectors
abstract
In this short paper, we take a first step towards empirically assessing Internet-wide malicious activities generated from and targeted towards Internet-scale business sectors (i.e., financial, health, education, etc.) and critical infrastructure (i.e., utilities, manufacturing, government, etc.). Facilitated by an innovative and a collaborative large-scale effort, we have conducted discussions with numerous Internet entities to obtain rare and private information related to allocated IP blocks pertaining to the aforementioned sectors and critical infrastructure. To this end, we employ such information to attribute Internet-scale maliciousness to such sectors and realms, in an attempt to provide an in-depth analysis of the global cyber situational posture. We draw upon close to 16.8 TB of darknet data to infer probing activities (typically generated by malicious/infected hosts) and DDoS backscatter, from which we distill IP addresses of victims. By executing week-long measurements, we observed an alarming number of more than 11,000 probing machines and 300 DDoS attack victims hosted by critical sectors. We also generate rare insights related to the maliciousness of various business sectors, including financial, which typically do not report their hosted and targeted illicit activities for reputation-preservation purposes. While we treat the obtained results with strict confidence due to obvious sensitivity reasons, we postulate that such generated cyber threat intelligence could be shared with sector/critical infrastructure operators, backbone networks and Internet service providers to contribute to the overall threat remediation objective.
Martin Husák, Nataliia Neshenko, Morteza Safaei Pour, Elias Bou-Harb, Pavel Celeda
ARES4
2018 Inferring, Characterizing, and Investigating Internet-Scale Malicious IoT Device Activities: A Network Telescope Perspective
abstract
Recent attacks have highlighted the insecurity of the Internet of Things (IoT) paradigm by demonstrating the impacts of leveraging Internet-scale compromised IoT devices. In this paper, we address the lack of IoT-specific empirical data by drawing upon more than 5TB of passive measurements. We devise data-driven methodologies to infer compromised IoT devices and those targeted by denial of service attacks. We perform large-scale characterization analysis of their traffic, as well as explore a public threat repository and an in-house malware database, to underlie their malicious activities. The results expose a significant 26 thousand compromised IoT devices "in the wild," with 40% being active in critical infrastructure. More importantly, we uncover new, previously unreported malware variants that specifically target IoT devices. Our empirical results render a first attempt to highlight the large-scale insecurity of the IoT paradigm, while alarming about the rise of new generations of IoT-centric malware-orchestrated botnets.
Sadegh Torabi, Elias Bou-Harb, Chadi Assi, Mario Galluscio, Amine Boukhtouta, Mourad Debbabi
DSN2
2018 Cross-Layer Authentication Protocol Design for Ultra-Dense 5G HetNets
abstract
Creating a secure environment for communications is becoming a significantly challenging task in 5G Heterogeneous Networks (HetNets) given the stringent latency and high capacity requirements of 5G networks. This is particularly factual knowing that the infrastructure tends to be highly diversified especially with the continuous deployment of small cells. In fact, frequent handovers in these cells introduce unnecessarily recurring authentications leading to increased latency. In this paper, we propose a software-defined wireless network (SDWN)- enabled fast cross- authentication scheme which combines non- cryptographic and cryptographic algorithms to address the challenges of latency and weak security. Initially, the received radio signal strength vectors at the mobile terminal (MT) is used as a fingerprinting source to generate an unpredictable secret key. Subsequently, a cryptographic mechanism based upon the authentication and key agreement protocol by employing the generated secret key is performed in order to improve the confidentiality and integrity of the authentication handover. Further, we propose a radio trusted zone database aiming to enhance the frequent authentication of radio devices which are present in the network. In order to reduce recurring authentications, a given covered area is divided into trusted zones where each zone contains more than one small cell, thus permitting the MT to initiate a single authentication request per zone, even if it keeps roaming between different cells. Accordingly, once the RSS vectors and the encrypted mobile identification are received by the authentication slice (AS), this latter builds the authentication vector using the k- nearest neighborhood technique to estimate the kdh fingerprint distribution which is compared to the radio trusted zones database to prove the legitimacy of the MT and the network slice (NS). Cross-layer authentication protocol is consequently executed. The proposed scheme is analyzed under different attack scenarios and its complexity is compared with cryptographic and non-cryptographic approaches to demonstrate its security resilience and computational efficiency.
Christian Miranda, Georges Kaddoum, Elias Bou-Harb
ICC3
2018 Implications of Theoretic Derivations on Empirical Passive Measurements for Effective Cyber Threat Intelligence Generation
abstract
Cyber space continues to be threatened by various debilitating attacks. In this context, executing passive measurements by analyzing Internet-scale, one- way darknet traffic has proven to be an effective approach to shed the light on Internet-wide maliciousness. While typically such measurements are solely conducted from the empirical perspective on already deployed darknet IP spaces using off-the-shelf Intrusion Detection Systems (IDS), their multidimensional theoretical foundations, relations and implications continue to be obscured. In this paper, we take a first step towards comprehending the relation between attackers' behaviors, the width of the darknet vantage points, the probability of detection and the minimum detection time. We perform stochastic modeling, derivation, validation, inter-correlation and analysis of such parameters to provide numerous insightful inferences, such as the most effective IDS and the most suitable darknet IP space, given various attackers' activities in the presence of detection time/probability constraints. One of the outcomes suggests that the widely-deployed Bro IDS is ideal for inferring slow, stealthy probing activities by leveraging passive measurements. Further, the results do not recommend deploying the Snort IDS when the available darknet IP space is relatively small, which is a typical scenario when darknets are operated and employed on organizational sub-networks. We concur that the generated derivations and mathematical relations put forward a first-of-akind formal and an accurate characterization of darknet-centric notions, which possess significant implications on Internet and passive measurements. This is especially factual with the advent of evolving paradigms such as IPv6 deployments and the proliferation of highly-distributed, orchestrated, large-scale and stealthy probing botnets.
Morteza Safaei Pour, Elias Bou-Harb
ICC2
2018 On Secrecy Bounds of MIMO Wiretap Channels with ZF detectors
abstract
This paper investigates the upper and lower bounds on achievable secrecy rate of multiple-input multiple-output (MIMO) wiretap channels by employing zero-forcing (ZF) detectors at, both the legitimate and malicious receivers' sides. Particularly, both large (i.e., gamma distribution) and small scale (i.e., Rayleigh distribution) fadings are taken into considerationwith the purpose of better simulating some practical scenarios. Motivated by previous works, bounds of achievable secrecy rate for MIMO wiretap channels over uncorrelated and min semi-correlated Rayleigh fading are derived with closed-form expressions. Furthermore, an asymptotic scenario in which the number of antennas tends to infinity is evaluated. Finally, thecloseness of our analytical expressions is corroborated through simulation results.
Long Kong, Georges Kaddoum, Daniel B. da Costa 0001, Elias Bou-Harb
IWCMC4
2018 On the Collaborative Inference of DDoS: An Information-theoretic Distributed Approach
abstract
Literature contributions have shown that information theoretic techniques can effectively detect various types of Distributed Denial of Service (DDoS) attacks. However, such techniques are often centralized with a limited measurement vantage point and suffer from the issue of single point of failure. Furthermore, with the flourishing of distributed and cloudbased environments, such techniques ought to adapt to such settings for scalability and performance reasons. In this paper, we address the problem of collaborative DDoS detection using information-theoretic techniques. To this end, we propose an entropy-based detection mechanism that supports collaborative agreement to identify suitable tuning network parameters for distributed DDoS inference in real-time. Empirical evaluations with real DDoS attacks demonstrate that the proposed approach is indeed capable of cooperatively inferring DDoS attacks while achieving resiliency and scalability.
Fatima Ezzahra Ouerfelli, Khaled Barbaria, Elias Bou-Harb, Claude Fachkha, Belhassen Zouari
IWCMC3
2018 A Machine Learning Model for Classifying Unsolicited IoT Devices by Observing Network Telescopes
abstract
The Internet of Things [IoT] promises to revolutionize the way we interact with our surroundings. Smart cars, smart cities, smart homes are now being realized with the help of various embedded devices that operate with little to no human interaction. However these embedded devices bring forth a plethora of security challenges as most manufacturers still assign higher importance to the three Ps (prototyping, production and performance) than security. This inherent flaw has manifested itself in the form of various Denial of Service (DoS) attacks orchestrated with the help of unsolicited IoT devices on the Internet. We are even seeing massive throughputs without the need for amplifications affecting large scale infrastructures on the Internet. Thus, understanding the nature of these attacks and quickly identifying infected devices becomes imperative to combat this situation. In this paper we present a model to classify unsolicited IoT devices in enterprises using machine learning (ML). Namely IP header information from darknet data is collected for analysis. We then consider multiple supervised ML algorithms to classify these Layer 3 headers. We evaluate these algorithms and compare their performances in terms of accurately identifying activities of malicious IoT devices on the Internet. Our results show that Random Forest and Gradient Boosting have high recall and precision scores whereas NaiveBayes has the worst performance. We believe our model can be used by enterprises as a part of their intrusion detection system to quickly identify infected IoT devices within their own environment as well as identify scanning activities directed towards them.
Farooq Shaikh, Elias Bou-Harb, Jorge Crichigno, Nasir Ghani
IWCMC2
2018 Passive inference of attacks on CPS communication protocols
Elias Bou-Harb, Nasir Ghani, Abdelkarim Erradi, Khaled B. Shaban
J. Inf. Secur. Appl.1
2018 CSC-Detector: A System to Infer Large-Scale Probing Campaigns
abstract
This paper uniquely leverages unsolicited real darknet data to propose a novel system, CSC-Detector, that aims at identifying Cyber Scanning Campaigns. The latter define a new phenomenon of probing events that are distinguished by their orchestration (i.e., coordination) patterns. To achieve its aim, CSC-Detector adopts three engines. Its fingerprinting engine exploits a unique observation to extract probing activities from darknet traffic. The system's inference engine employs a set of behavioral analytics to generate numerous significant insights related to the machinery of the probing sources while its analysis engine exploits the previously obtained inferences to automatically infer the campaigns. CSC-Detector is empirically evaluated and validated using 240 GB of real darknet data. The outcome discloses 3 recent, previously unreported large-scale probing campaigns targeting diverse Internet services. Further, one of those inferred campaigns revealed that the sipscan campaign that was initially analyzed by CAIDA is arguably still active, yet operating in a stealthy, very low rate mode. We envision that the proposed system that is tailored towards darknet data, which is frequently, abundantly and effectively used to generate cyber threat intelligence, could be used by network security analysts, emergency response teams and/or observers of cyber events to infer large-scale orchestrated probing campaigns. This would be utilized for early cyber attack warning and notification as well as for simplified analysis and tracking of such events.
Elias Bou-Harb, Chadi Assi, Mourad Debbabi
IEEE Trans. Dependable Secur. Comput.1
2017 On the Sequential Pattern and Rule Mining in the Analysis of Cyber Security Alerts
abstract
Data mining is well-known for its ability to extract concealed and indistinct patterns in the data, which is a common task in the field of cyber security. However, data mining is not always used to its full potential among cyber security community. In this paper, we discuss usability of sequential pattern and rule mining, a subset of data mining methods, in an analysis of cyber security alerts. First, we survey the use case of data mining, namely alert correlation and attack prediction. Subsequently, we evaluate sequential pattern and rule mining methods to find the one that is both fast and provides valuable results while dealing with the peculiarities of security alerts. An experiment was performed using the dataset of real alerts from an alert sharing platform. Finally, we present lessons learned from the experiment and a comparison of the selected methods based on their performance and soundness of the results.
Martin Husák, Jaroslav Kaspar, Elias Bou-Harb, Pavel Celeda
ARES3
2017 On correlating network traffic for cyber threat intelligence: A Bloom filter approach
abstract
Internet and organizational network security is still threatened by devastating malicious activities. Given the continuous escalation of such attacks in terms of their frequency, sophistication and stealthiness, it is of paramount importance to generate effective cyber threat intelligence that aims at inferring, attributing, characterizing and mitigating such misdemeanors. Nevertheless, such imperative tasks are partially impeded by the lack of correlation approaches that can produce prompt and accurate actionable intelligence by investigating various network traffic sources. To this end, this paper proposes a simple yet effective approach to generically correlate network traffic for cyber security purposes. The approach uniquely exploits Bloom filters to infer similarities between the analyzed network traffic while eliminating false negatives and managing a very low and a measurable false positive rate. We demonstrate the effectiveness of the proposed approach by empirically evaluating it using 10 GB of real darknet data and close to 15 thousand malware traffic samples. The outcome is rendered by hundreds of inferred and attributed Internet-scale infections, which we corroborate using third-party publicly accessible threat repositories. We envision that the proposed approach could be leveraged as an effective correlation component in complex security information and event management systems to provide metrics that would aid in characterizing and comprehending various network security activities and incidents.
Adil Atifi, Elias Bou-Harb
IWCMC2
2017 Internet-scale Probing of CPS: Inference, Characterization and Orchestration Analysis
Claude Fachkha, Elias Bou-Harb, Anastasis Keliris, Nasir Memon, Mustaque Ahamad
NDSS2
2017 A first empirical look on internet-scale exploitations of IoT devices
abstract
Technological advances and innovative business models led to the modernization of the cyber-physical concept with the realization of the Internet of Things (IoT). While IoT envisions a plethora of high impact benefits in both, the consumer as well as the control automation markets, unfortunately, security concerns continue to be an afterthought. Several technical challenges impede addressing such security requirements, including, lack of empirical data related to various IoT devices in addition to the shortage of actionable attack signatures. In this paper, we present what we believe is a first attempt ever to comprehend the severity of IoT maliciousness by empirically characterizing the magnitude of Internet-scale IoT exploitations. We draw upon unique and extensive darknet (passive) data and develop an algorithm to infer unsolicited IoT devices which have been compromised and are attempting to exploit other Internet hosts. We further perform correlations by leveraging active Internet-wide scanning to identify and report on such IoT devices and their hosting environments. The generated results indicate a staggering 11 thousand exploited IoT devices that are currently in the wild. Moreover, the outcome pinpoints that IoT devices embedded deep in operational Cyber-Physical Systems (CPS) such as manufacturing plants and power utilities are the most compromised. We concur that such results highlight the wide-spread insecurities of the IoT paradigm, while the actionable generated inferences are postulated to be leveraged for prompt mitigation as well as to facilitate IoT forensic investigations using real empirical data.
Mario Galluscio, Nataliia Neshenko, Elias Bou-Harb, Yongliang Huang, Nasir Ghani, Jorge Crichigno, Georges Kaddoum
PIMRC3
2016 Passive inference of attacks on SCADA communication protocols
abstract
The security of industrial Cyber-Physical Systems (CPS) has been recently receiving significant attention from the research community. While the majority of such attention originates from the control theory domain, very few works proposed viable approaches to the problem from the practical perspective. In this work, we do not claim that we propose a particular solution to a specific problem related to CPS security, but rather present a first look into what can help shape these solutions in the future. Indeed, our vision and ultimate goal is to attempt to merge or at least diminish the gap between highly theoretical solutions and practical approaches derived from insightful empirical experimentation, for securing CPS. Towards this goal, in this work, we present what we believe is the first specimen ever of passive measurements of real attacks on CPS communication protocols. By analyzing a recent one-week dataset rendered by 20 GB of unsolicited real traffic targeting half a million routable, allocated but unused Internet Protocol (IP) addresses, we shed the light on attackers' intention and actual attacks targeting CPS. Specifically, we characterize such attacks in terms of their types, their frequency, their target protocols and possible orchestration behavior. Our results demonstrate a staggering 3 thousand scanning attempts and close to 2 thousand denial of service attacks on various CPS communication protocols. One insightful observation from our work is the fact that attackers are not interested in exploiting the Modbus protocol; in contrast to most literature works that are extensively dedicating their research efforts to devise secure models for Modbus. We hope that this paper motivates the literature to design secure and tailored CPS models that leverage tangible attacks and vulnerabilities inferred from empirical measurements, to achieve truly reliable and secure CPS.
Elias Bou-Harb
ICC1
2016 A probabilistic model to preprocess darknet data for cyber threat intelligence generation
abstract
Internet traffic destined to routable yet unallocated IP addresses is commonly referred to as telescope or darknet data. Such unsolicited traffic is frequently, abundantly and effectively exploited to generate various cyber threat intelligence related, but not limited to, scanning activities, distributed denial of service attacks and malware identification. However, such data typically contains a significant amount of misconfiguration traffic caused by network/routing or hardware/software faults. The latter not only immensely affects the purity of darknet data, which hinders the accuracy of inference algorithms that operate on such data, but also wastes valuable storage resources. This paper proposes a probabilistic model to preprocess darknet data in order to prepare it for effective use. The aim is to fingerprint darknet misconfiguration traffic and subsequently filter it out. The model is advantageous as it does not rely on arbitrary cut-off thresholds, provide separate likelihood models to distinguish between miscon-figuration and other darknet traffic, and is independent from the nature of the source of the traffic. To the best of our knowledge, the proposed model renders a first attempt ever to formally tackle the problem of preprocessing darknet traffic. Through empirical evaluations using real darknet traffic and by comparing the proposed model against the baseline and a heuristic approach, we demonstrate the accuracy and effectiveness of the model.
Elias Bou-Harb
ICC1
2016 A novel cyber security capability: Inferring Internet-scale infections by correlating malware and probing activities
Elias Bou-Harb, Mourad Debbabi, Chadi Assi
Comput. Networks1
2015 A Time Series Approach for Inferring Orchestrated Probing Campaigns by Analyzing Darknet Traffic
abstract
This paper aims at inferring probing campaigns by investigating dark net traffic. The latter probing events refer to a new phenomenon of reconnaissance activities that are distinguished by their orchestration patterns. The objective is to provide a systematic methodology to infer, in a prompt manner, whether or not the perceived probing packets belong to an orchestrated campaign. Additionally, the methodology could be easily leveraged to generate network traffic signatures to facilitate capturing incoming packets as belonging to the same inferred campaign. Indeed, this would be utilized for early cyber attack warning and notification as well as for simplified analysis and tracking of such events. To realize such goals, the proposed approach models such challenging task as a problem of interpolating and predicting time series with missing values. By initially employing trigonometric interpolation and subsequently executing state space modeling in conjunction with a time-varying window algorithm, the proposed approach is able to pinpoint orchestrated probing campaigns by only monitoring few orchestrated flows. We empirically evaluate the effectiveness of the proposed model using 330 GB of real dark net data. By comparing the outcome with a previously validated work, the results indeed demonstrate the promptness and accuracy of the proposed approach.
Elias Bou-Harb, Mourad Debbabi, Chadi Assi
ARES1
2015 Inferring distributed reflection denial of service attacks from darknet
Claude Fachkha, Elias Bou-Harb, Mourad Debbabi
Comput. Commun.2
2015 On the inference and prediction of DDoS campaigns
abstract
Abstract This work proposes a distributed denial‐of‐service (DDoS) inference and forecasting model that aims at providing insights to organizations, security operators, and emergency response teams during and after a DDoS attack. Specifically, our work strives to predict, within minutes, the attacks' features, namely intensity/rate (packets/second) and size (estimated number of used compromised machines/bots). The goal is to understand the future short‐term trend of the ongoing DDoS attack in terms of those features and thus provide the capability to recognize the current as well as future similar situations and hence appropriately respond to the threat. Further, our work aims at investigating DDoS campaigns by proposing a clustering approach to infer various victims targeted by the same campaign and predicting related features. Our analysis employs real darknet data to explore the feasibility of applying the inference and forecasting models on DDoS attacks and evaluate the accuracy of the predictions. To achieve our goal, our proposed approach leverages a number of time series and fluctuation analysis techniques, statistical methods, and forecasting approaches. The extracted inferences from various DDoS case studies exhibit a promising accuracy reaching at some points less than 1% error rate. Further, our approach could lead to a better understanding of the scale, speed, and size of DDoS attacks and generates inferences that could be adopted for immediate response and mitigation. Moreover, the accumulated insights could be used for the purpose of long‐term large‐scale DDoS analysis. Copyright © 2014 John Wiley & Sons, Ltd.
Claude Fachkha, Elias Bou-Harb, Mourad Debbabi
Wirel. Commun. Mob. Comput.2
2014 Inferring internet-scale infections by correlating malware and probing activities
abstract
This paper presents a new approach to infer malware-infected machines by solely analyzing their generated probing activities. In contrary to other adopted methods, the proposed approach does not rely on symptoms of infection to detect compromised machines. This allows the inference of malware infection at very early stages of contamination. The approach aims at detecting whether the machines are infected or not as well as pinpointing the exact malware type/family, if the machines were found to be compromised. The latter insights allow network security operators of diverse organizations, Internet service providers and backbone networks to promptly detect their clients' compromised machines in addition to effectively providing them with tailored anti-malware/patch solutions. To achieve the intended goals, the proposed approach exploits the darknet Internet space and employs statistical methods to infer large-scale probing activities. Subsequently, such activities are correlated with malware samples by leveraging fuzzy hashing and entropy based techniques. The proposed approach is empirically evaluated using 60 GB of real darknet traffic and 65 thousand real malware samples. The results concur that the rationale of exploiting probing activities for worldwide early malware infection detection is indeed very promising. Further, the results demonstrate that the extracted inferences exhibit noteworthy accuracy and can generate significant cyber security insights that could be used for effective mitigation.
Elias Bou-Harb, Claude Fachkha, Mourad Debbabi, Chadi Assi
ICC1
2014 On fingerprinting probing activities
Elias Bou-Harb, Mourad Debbabi, Chadi Assi
Comput. Secur.1
2013 A Statistical Approach for Fingerprinting Probing Activities
abstract
Probing is often the primary stage of an intrusion attempt that enables an attacker to remotely locate, target, and subsequently exploit vulnerable systems. This paper attempts to investigate whether the perceived traffic refers to probing activities and which exact scanning technique is being employed to perform the probing. Further, this work strives to examine probing traffic dimensions to infer the `machinery' of the scan, whether the probing activity is generated from a software tool or from a worm/bot net and whether the probing is random or follows a certain predefined pattern. Motivated by recent cyber attacks that were facilitated through probing, limited cyber security intelligence related to the mentioned inferences and the lack of accuracy that is provided by scanning detection systems, this paper presents a new approach to fingerprint probing activity. The approach leverages a number of statistical techniques, probabilistic distribution methods and observations in an attempt to understand and analyze probing activities. To prevent evasion, the approach formulates this matter as a change point detection problem that yielded motivating results. Evaluations performed using 55 GB of real dark net traffic shows that the extracted inferences exhibit promising accuracy and can generate significant insights that could be used for mitigation purposes.
Elias Bou-Harb, Mourad Debbabi, Chadi Assi
ARES1
2013 On detecting and clustering distributed cyber scanning
abstract
This paper proposes an approach that is composed of two techniques that respectively tackle the issues of detecting corporate cyber scanning and clustering distributed reconnaissance activity. The first employed technique is based on a non-attribution anomaly detection approach that focuses on what is being scanned rather than who is performing the scanning. The second technique adopts a statistical time series approach that is rendered by observing the correlation status of a traffic signal to perform the identification and clustering. To empirically validate both techniques, we experiment with two real network traffic datasets and implement two proof-of-concept environments. The first dataset comprises of unsolicited one-way telescope/darknet traffic while the second dataset has been captured in our lab through a customized setup. The results show, on one hand, that for a class C network with 250 active hosts and 5 monitored servers, the proposed detection technique's training period required a stabilization time of less than 1 second and a state memory of 80 bytes. Moreover, in comparison with Snort's sfPortscan technique, it was able to detect 4215 unique scans and yielded zero false negative. On the other hand, the proposed clustering technique is able to correctly identify and cluster the scanning machines with high accuracy even in the presence of legitimate traffic.
Elias Bou-Harb, Mourad Debbabi, Chadi Assi
IWCMC1
2013 Towards a Forecasting Model for Distributed Denial of Service Activities
abstract
Distributed Denial of Service (DDoS) activities continue to dominate today's attack landscape. This work proposes a DDoS forecasting model to provide significant insights to organizations, security operators and emergency response teams during and after a targeted DDoS attack. Specifically, the work strives to predict, within minutes, the attacks' impact features, namely, intensity/rate (packets/sec) and size (estimated number of used compromised machines/bots). The goal is to understand the future short term trend of the ongoing DDoS attack in terms of those features and thus provide the capability to recognize the current as well as future similar situations and hence appropriately respond to the threat. Our analysis employs real dark net data to explore the feasibility of applying the forecasting model on targeted DDoS attacks and subsequently evaluate the accuracy of the predictions. To achieve its tasks, our proposed approach leverages a number of time series fluctuation analysis and forecasting methods. The extracted inferences from various DDoS case studies exhibit promising accuracy reaching at some points less than 1% error rate. Further, our model could lead to better understanding of the scale and speed of DDoS attacks and should generate inferences that could be adopted for immediate response and hence mitigation as well as accumulated for the purpose of long term large-scale DDoS analysis.
Claude Fachkha, Elias Bou-Harb, Mourad Debbabi
NCA2
2013 A systematic approach for detecting and clustering distributed cyber scanning
Elias Bou-Harb, Mourad Debbabi, Chadi Assi
Comput. Networks1
2013 A secure, efficient, and cost-effective distributed architecture for spam mitigation on LTE 4G mobile networks
abstract
ABSTRACT The 4G of mobile networks will be a technology‐opportunistic and user‐centric system, combining the economical and technological advantages of various transmission technologies. As a part of its new architecture, LTE networks will implement an evolved packet core. Although this will provide various critical advantages, it will, on the other hand, expose telecom networks to serious IP‐based attacks. One often adopted solution to mitigate such attacks is based on a centralized security architecture. However, this approach requires large processing and memory resources to handle huge amounts of traffic, which, in turn, causes a significant over dimensioning problem in the centralized nodes. Hence, it may cause this approach to fail from achieving its security task. In this paper, we focus on a SPAM flooding attack, namely SMTP SPAM, and demonstrate, through simulations and discussion, its DoS impact on the Long Term Evolution (LTE) network and subsequent effects on the mobile network operator. Our main contribution involves proposing a distributed architecture on the LTE network that is secure and that mitigates attacks efficiently by solving the over dimensioning problem. It is also cost‐effective by utilizing ‘off‐the‐shelf’ low‐cost hardware in the distributed nodes. Through additional simulation and analysis, we demonstrate the feasibility and effectiveness of our approach. Copyright © 2012 John Wiley & Sons, Ltd.
Elias Bou-Harb, Makan Pourzandi, Mourad Debbabi, Chadi Assi
Secur. Commun. Networks1
2012 Investigating the dark cyberspace: Profiling, threat-based analysis and correlation
abstract
An effective approach to gather cyber threat intelligence is to collect and analyze traffic destined to unused Internet addresses known as darknets. In this paper, we elaborate on such capability by profiling darknet data. Such information could generate indicators of cyber threat activity as well as providing in-depth understanding of the nature of its traffic. Particularly, we analyze darknet packets distribution, its used transport, network and application layer protocols and pinpoint its resolved domain names. Furthermore, we identify its IP classes and destination ports as well as geo-locate its source countries. We further investigate darknet-triggered threats. The aim is to explore darknet embedded threats and categorize their severities. Finally, we contribute by exploring the inter-correlation of such threats, by applying association rule mining techniques, to build threat association rules. Specifically, we generate clusters of threats that co-occur targeting a specific victim. Such work proves that specific darknet threats are correlated. Moreover, it provides insights about threat patterns and allows the interpretation of threat scenarios.
Claude Fachkha, Elias Bou-Harb, Amine Boukhtouta, Son Dinh, Farkhund Iqbal, Mourad Debbabi
CRiSIS2
2012 A first look on the effects and mitigation of VoIP SPIT flooding in 4G mobile networks
abstract
The fourth generation of mobile networks is considered a technology-opportunistic and user-centric system. Part of its new architecture, 4G networks will implement an evolved packet core. Although this will provide various critical advantages, it will however expose telecom networks to serious IP-based attacks. One often adopted solution to mitigate such attacks is based on a centralized security architecture. This centralized approach nonetheless, requires large processing resources to handle large amount of traffic, which may result in a significant over dimensioning problem in the centralized nodes causing this approach to fail from achieving its security task. In this paper, we primarily contribute by presenting a first look on the DoS effects of VoIP SPIT flooding on 4G mobile networks. We further contribute by proposing a distributed architecture on the mobile network infrastructure that is secure, efficient and cost-effective.
Elias Bou-Harb, Mourad Debbabi, Chadi Assi
ICC1