Sherif Saad

dblp:40/9034 · DBLP profile ↗
← Back
30ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0002-5506-5261ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 22 · 7 first-author · 13 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Autonomous Adversary: Red-Teaming in the Age of LLM
Mohammad Saiful Islam Mamun, Mohamed Gaber, Scott Buffett, Sherif Saad
ACISP (3)4
2026 TI-NERmergerV2: Automating the integration of threat intelligence NER datasets via STIX standard
abstract
Quality-labeled data are essential for developing accurate AI models in cybersecurity, particularly for threat intelligence named entity recognition (TI-NER), which automates the extraction of threat indicators and entities from unstructured reports. While several annotated datasets exist, their isolated use hinders scalability due to inconsistent tagging schemes, label names, and non-standard entity categories. This paper introduces TI-NERmergerV2, a robust, semi-automated framework for integrating heterogeneous TI-NER datasets into a unified, high-quality corpus aligned with the structured threat information expression (STIX) standard (e.g, STIX 2.1). Building upon its predecessor, TI-NERmerger, which is limited by its reliance on strict string matching and a narrow cyber lookup space, TI-NERmergerV2 incorporates string normalization, fuzzy fallback matching, and alias expansion using the MITRE ATT&CK knowledge base to resolve lexical variation and annotation inconsistencies. We validate its effectiveness by comparing it with a manual integration of two public datasets (DNRTI and APTNER), producing a unified dataset called AAPTNER. TI-NERmergerV2 achieves over 94% alignment with the manual process, reducing months of expert effort to minutes. Evaluations using a RoBERTa-based NER model further confirm that TI-NERmergerV2 enhances annotation quality and effectively disambiguates key entity types in the resulting DNRTI-STIX2.1 and AAPTNER datasets. The framework generalizes across datasets that adopt STIX domain and observable objects, providing a scalable and reproducible foundation for cyber threat intelligence research. Both the framework and resulting datasets are publicly released to support broader efforts in standardizing and enriching TI-NER resources.
Inoussa Mouiche, Sherif Saad
Comput. Secur.2
2026 Preserving data and model privacy during inference and training
William Briguglio, Issa Traoré, Mohammad Saiful Islam Mamun, Waleed A. Yousef, Sherif Saad
Expert Syst. Appl.5
2025 Synthetic Lateral Movement Data Generation for Azure Cloud: A Hopper-Based Approach
Mohammad Saiful Islam Mamun, Hadeer Ahmed, Anas Mabrouk, Sherif Saad
CANS4
2025 Drift-RL: A Reinforcement Learning Framework for Simulating Textual Data Drift in Cybersecurity
Hadeer Ahmed, Issa Traoré, Sherif Saad, Mohammad Saiful Islam Mamun
CRiSIS3
2025 Context-Aware Entity-Relation Extraction for Threat Intelligence Knowledge Graphs
Inoussa Mouiche, Sherif Saad
CRiSIS2
2025 Adaptive Ensemble Defense: Mitigating NLP Adversarial Attacks with Data-Augmented Voting Mechanisms
Amira Abdelbaky, Sherif Saad, Mohammad Saiful Islam Mamun
ICISSP (2)2
2025 An Alternative Approach to Federated Learning for Model Security and Data Privacy
abstract
Federated learning (FL) enables machine learning on data held across multiple clients without exchanging private data. However, exchanging information for model training can compromise data privacy. Further, participants may be untrustworthy and can attempt to sabotage model performance. Also, data that is not independently and identically distributed (IID) impede the convergence of FL techniques. We present a general framework for federated learning via aggregating multivariate estimated densities (FLAMED). FLAMED aggregates density estimations of clients’ data, from which it simulates training datasets to perform centralized learning, bypassing problems arising from non-IID data and contributing to addressing privacy and security concerns. FLAMED does not require a copy of the global model to be distributed to each participant during training, meaning the aggregating server can retain sole proprietorship of the global model without the use of resource-intensive homomorphic encrypti on. We compared its performance to standard FL approaches using synthetic and real datasets and evaluated its resilience to model poisoning attacks. Our results indicate that FLAMED effectively handles non-IID data in many settings while also being more secure.
William Briguglio, Waleed A. Yousef, Issa Traoré, Mohammad Saiful Islam Mamun, Sherif Saad
ICISSP (1)5
2025 HybridMTD: Enhancing Robustness Against Adversarial Attacks with Ensemble Neural Networks and Moving Target Defense
Kimia Tahayori, Sherif Saad, Mohammad Saiful Islam Mamun, Saeed Samet
ICISSP (2)2
2025 Thoth: A Lightweight Framework for End-to-End Consumer IoT Rapid Testing
abstract
The rapid expansion of consumer IoT devices has increased the need for scalable, automated testing solutions. Manual methods are often slow, error-prone, and inadequate for capturing real-world IoT complexities. Existing frameworks typically lack comprehensiveness, quantifiable metrics, and support for cascading failure scenarios. This paper introduces Thoth, a lightweight, end-to-end IoT testing framework that addresses these limitations. Thoth enables holistic evaluation through integrated support for performance, reliability, recovery, security, and load testing. It also incorporates standardized metrics and real-time failure simulations, including cascading faults. We evaluated Thoth using eight test cases in a real-world health-monitoring setup involving a smartwatch, edge gateway, and cloud infrastructure. Key metrics—such as fault detection time, recovery speed, data loss, and energy usage—were logged and analyzed. Results show that Thoth detects faults in as little as 2.5 seconds, recovers in under 1 second, limits data loss to a few points, and maintains sub-1% energy overhead. These findings highlight its effectiveness for low-intrusion testing in resource-constrained environments. By combining scenario-driven design with reproducible, metrics-based evaluation, Thoth fills key gaps in IoT testing.
Salma Roshdy Aly, Sherif Saad, Mohammad Saiful Islam Mamun
ICSOFT2
2025 ALPACA AGAINST VICUNA: Using LLMs to Uncover Memorization of LLMs
abstract
Aly M. Kassem, Omar Mahmoud, Niloofar Mireshghallah, Hyunwoo Kim, Yulia Tsvetkov, Yejin Choi, Sherif Saad, Santu Rana. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Aly M. Kassem, Omar Mahmoud 0001, Niloofar Mireshghallah, Hyunwoo Kim 0002, Yulia Tsvetkov, Yejin Choi 0001, Sherif Saad, Santu Rana
NAACL (Long Papers)7
2025 Entity and relation extractions for threat intelligence knowledge graphs
abstract
Advanced persistent threats (APTs) represent a complex challenge in cybersecurity as they infiltrate networks stealthily to conduct espionage, steal data, and maintain a long-term presence. To combat these threats, security professionals increasingly rely on cyber knowledge graphs (CKGs), which provide scalable solutions to analyze and structure vast amounts of cyber threat intelligence (CTI) from diverse sources in real-time, enabling the automation of proactive security measures. Developing CKGs requires extracting entity and their relationships from unstructured CTI reports. However, existing approaches face significant limitations, such as difficulties with the nuances of cybersecurity language, diverse threat terminologies, and high rates of error propagation, resulting in low accuracy and poor generalizability. This paper introduces a novel Threat Intelligence Knowledge Graph (TiKG) pipeline designed to address these challenges. The TiKG framework leverages SecureBERT, a domain-specific transformer-based model optimized for cybersecurity, and integrates it with an attention-based BiLSTM to capture the context and nuances of security texts, reducing error propagation and improving extraction accuracy. Additionally, the pipeline incorporates a domain-specific ontology and inference model to ensure precise relation mapping in relation extraction. Using three large-scale TI open-source datasets (DNRTI, STUCCO, and CYNER) and a curated CTI dataset, extensive evaluations demonstrate the effectiveness of our framework, showing significant improvements over existing methods in detecting and linking cyber threats. These contributions provide a robust platform for security professionals to analyze and predict potential attacks, develop effective defenses, and enhance the strategic capabilities of cybersecurity operations.
Inoussa Mouiche, Sherif Saad
Comput. Secur.2
2025 TIJERE: A novel threat intelligence joint extraction model based on analyst expert knowledge
abstract
The extraction of entities and relationships from threat intelligence reports into structured formats, such as cybersecurity knowledge graphs, is essential for automated threat analysis, detection, and mitigation. However, existing joint extraction methods struggle with feature confusion, language ambiguity, noise propagation, and overlapping relations, resulting in low accuracy and poor model performance. This paper presents TIJERE, an innovative joint entity and relation extraction framework that formulates joint extraction as a multisequence labeling representation (MSLR) problem. Specifically, separate sequences are generated for each entity pair. Unlike prior tagging schemes, MSLR integrates expert domain features to enrich positional, contextual, and semantic representations of entities, thereby enhancing feature distinction and classification accuracy. Additionally, TIJERE reduces language ambiguity and enhances domain-specific generalization by leveraging SecureBERT+, a contextual language model fine-tuned on cybersecurity text. This improves both named entity recognition (NER) and relation extraction (RE). This paper also introduces DNRTI-JE, the first publicly available jointly labeled dataset for cybersecurity entity and RE, filling a crucial gap in cyber threat intelligence automation. Empirical evaluations on the curated DNRTI-JE dataset demonstrate that TIJERE achieves state-of-the-art performance, with F1-scores exceeding 0.93 for NER and 0.98 for RE, outperforming existing methods. Together, TIJERE and the standardized benchmarking DNRTI-JE dataset enable high-performance cybersecurity intelligence extraction, with transferable applications in healthcare, finance, and bioinformatics.
Inoussa Mouiche, Sherif Saad
Knowl. Based Syst.2
2024 Finding a Needle in the Adversarial Haystack: A Targeted Paraphrasing Approach For Uncovering Edge Cases with Minimal Distribution Distortion
abstract
are offensive/hateful in nature.
Aly M. Kassem, Sherif Saad
EACL (1)2
2024 TI-NERmerger: Semi-Automated Framework for Integrating NER Datasets in Cybersecurity
Inoussa Mouiche, Sherif Saad
SECRYPT2
2024 Effect of Text Augmentation and Adversarial Training on Fake News Detection
abstract
The action of spreading false information through fake news articles presents a significant danger to society because it has the ability to shape public opinion with inaccurate facts. This can lead to negative effects, such as reduced trust in institutions and the promotion of conflict, division, and even violence. In this article, a text augmentation technique is introduced as a means of generating new data from preexisting fake news datasets. This approach has the potential to enhance classifier performance by a range of 3%–11%. It can also be utilized to launch a successful attack on trained classifiers, with up to a 90% success rate. However, the success rate of these attacks decreased to less than 28% when the model was retrained with the generated adversarial examples. These results demonstrate the effectiveness of text augmentation as a viable method for detecting fake news and increasing classifier accuracy and performance, as well as its ability to be utilized to perform adversarial machine learning (ML) and improve the resilience of ML algorithms.
Hadeer Ahmed, Issa Traoré, Sherif Saad, Mohammad Saiful Islam Mamun
IEEE Trans. Comput. Soc. Syst.3
2023 Preserving Privacy Through Dememorization: An Unlearning Technique For Mitigating Memorization Risks In Language Models
abstract
Large Language models (LLMs) are trained on vast amounts of data, including sensitive information that poses a risk to personal privacy if exposed.LLMs have shown the ability to memorize and reproduce portions of their training data when prompted by adversaries.Prior research has focused on addressing this memorization issue and preventing verbatim replication through techniques like knowledge unlearning and data pre-processing.However, these methods have limitations regarding the number of protected samples, limited privacy types, and potentially lower-quality generative models.To tackle this challenge more effectively, we propose "DeMem," a novel unlearning approach that utilizes an efficient reinforcement learning feedback loop via proximal policy optimization.By fine-tuning the language model with a negative similarity score as a reward signal, we incentivize the LLMs to learn a paraphrasing policy to unlearn the pre-training data.Our experiments demonstrate that De-Mem surpasses strong baselines and state-ofthe-art methods in terms of its ability to generalize and strike a balance between maintaining privacy and LLM performance.
Aly M. Kassem, Omar Mahmoud 0001, Sherif Saad
EMNLP3
2023 Automated Feature Engineering for AutoML Using Genetic Algorithms
Kevin Shi, Sherif Saad
SECRYPT2
2022 iProfile: Collecting and Analyzing Keystroke Dynamics from Android Users
abstract
Keystroke dynamics is one of the most popular behavioural biometrics that are currently being used as a second factor of authentication for many web services and applications.One of the reasons that makes it really popular is that it is a resettable biometric, which meets one of the main usability requirements of authentication systems.With the recent advances in mobile technologies, developers and researchers utilized several machine learning algorithms to identify smartphone users based on their keystroke dynamics.The biggest problem that faces researchers in this area is the ability to collect datasets from smartphone users that could be used to train the machine learning algorithms and, hence, create accurate predictive model.This paper introduces iProfile, a native Android application that collects keystroke dynamics from Android smartphone users.This application opens the door for researchers to recruit participants from all over the world to contribute to the data collection of keystroke dynamics.Our iProfile application allows researchers to study the impact of several parameters, such as hardware brands, users' geolocation, native language text direction, and several other factors, on the accuracy of machine learning classifiers.It also helps maintain a standard benchmark for keystroke dynamics.Having a standard benchmark helps researchers better evaluate their work based on consistent data collection procedures and evaluation metrics.This paper explains the main building blocks of the iProfile application, the algorithms used in the implementation, the communication protocol with the database server, the structure and format of the generated dataset and the feature extraction approaches.As a proof of concept, the app was used to develop a novel feature-set that identifies Android users based on 147 features.
Haytham Elmiligi, Sherif Saad
ICISSP2
2021 Spam review detection using self-organizing maps and convolutional neural networks
Ashraf Neisari, Luis Rueda 0001, Sherif Saad
Comput. Secur.3
2020 Emerging Design Patterns for Blockchain Applications
Vijay Rajasekar, Shiv Sondhi, Sherif Saad, Shady Mohammed
ICSOFT3
2019 The Curious Case of Machine Learning in Malware Detection
abstract
In this paper, we argue that detecting malware attacks in the wild is a unique challenge for machine learning techniques. Given the current trend in malware development and the increase of unconventional malware attacks, we expect that dynamic malware analysis is the future for antimalware detection and prevention systems. A comprehensive review of machine learning for malware detection is presented. Then, we discuss how malware detection in the wild present unique challenges for the current state-of-the-art machine learning techniques. We defined three critical problems that limit the success of malware detectors powered by machine learning in the wild. Next, we discuss possible solutions to these challenges and present the requirements of next-generation malware detection. Finally, we outline potential research directions in machine learning for malware detection.
Sherif Saad, William Briguglio, Haytham Elmiligi
ICISSP1
2019 JSLess: A Tale of a Fileless Javascript Memory-Resident Malware
Sherif Saad, Farhan Mahmood, William Briguglio, Haytham Elmiligi
ISPEC1
2014 Context-aware intrusion alerts verification approach
abstract
Intrusion detection systems (IDSs) produce a massive number of intrusion alerts. A huge number of these alerts are false positives. Investigating false positive alerts is an expensive and time consuming process, and as such represents a significant problem for intrusion analysts. This shows the needs for automated approaches to eliminate false positive alerts. In this paper, we propose a novel alert verification and false positives reduction approach. The proposed approach uses context-aware and semantic similarity to filter IDS alerts and eliminate false positives. Evaluation of the approach with an IDS dataset that contains massive number of IDS alerts yields strong performance in detecting false positive alerts.
Sherif Saad, Issa Traoré, Marcelo Luiz Brocardo
IAS1
2013 Botnet detection based on traffic behavior analysis and flow intervals
Issa Traoré, Bassam Sayed, Wei Lu 0018, Sherif Saad, Ali A. Ghorbani 0001, Daniel Garant
Comput. Secur.5
2013 Semantic aware attack scenarios reconstruction
Sherif Saad, Issa Traoré
J. Inf. Secur. Appl.1
2012 Peer to Peer Botnet Detection Based on Flow Intervals
Issa Traoré, Ali A. Ghorbani 0001, Bassam Sayed, Sherif Saad, Wei Lu 0018
SEC5
2011 A semantic analysis approach to manage IDS alerts flooding
abstract
In this paper we propose a new approach to manage alerts flooding in IDSs. The proposed approach uses semantic analysis and ontology engineering techniques to combine and fuse two or more raw IDS alerts into one summarized hybrid/meta-alert. Our approach applies a new method based on measuring the semantic similarity between IDS alerts attributes to identify the alerts that are suitable for aggregation and summarization. In contrast to previous works our approach ensures that the aggregated alerts will not lose any valuable information existing in the raw alerts set. The experimental results show that our approach is effective and efficient in fusing massive number of alerts compared to previous works in the area.
Sherif Saad, Issa Traoré
IAS1
2011 Detecting P2P botnets through network behavior analysis and machine learning
abstract
Botnets have become one of the major threats on the Internet for serving as a vector for carrying attacks against organizations and committing cybercrimes. They are used to generate spam, carry out DDOS attacks and click-fraud, and steal sensitive information. In this paper, we propose a new approach for characterizing and detecting botnets using network traffic behaviors. Our approach focuses on detecting the bots before they launch their attack. We focus in this paper on detecting P2P bots, which represent the newest and most challenging types of botnets currently available. We study the ability of five different commonly used machine learning techniques to meet online botnet detection requirements, namely adaptability, novelty detection, and early detection. The results of our experimental evaluation based on existing datasets show that it is possible to detect effectively botnets during the botnet Command-and-Control (C&C) phase and before they launch their attacks using traffic behaviors only. However, none of the studied techniques can address all the above requirements at once.
Sherif Saad, Issa Traoré, Ali A. Ghorbani 0001, Bassam Sayed, Wei Lu 0018, John Felix, Payman Hakimian
PST1
2010 Method ontology for intelligent network forensics analysis
abstract
Network forensics is an after the fact process to investigate malicious activities conducted over computer networks by gathering useful intelligence. Recently, several machine learning techniques have been proposed to automate and develop intelligent network forensics systems. An intelligent network forensics system that reconstructs intrusion scenarios and makes attack attributions requires knowledge about intrusions signatures, evidences, impacts, and objectives. In addition, problem solving knowledge that describes how the system can use domain knowledge to analyze malicious activities is essential for the design of intelligent network forensics systems. In this paper we adapt recent researches in semantic-web, information architecture, and ontology engineering to design a method ontology for network forensics analysis. The proposed ontology represents both network forensics domain knowledge and problem solving knowledge. It can be used as a knowledge-base for developing sophisticated intelligent network forensics systems to support complex chain of reasoning. We use a real life network intrusion scenario to show how our ontology can be integrated and used in intelligent network forensics systems.
Sherif Saad, Issa Traoré
PST1