Saeed Salem

dblp:26/101 · DBLP profile ↗
← Back
29ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0001-6478-4674ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 2 since 2021Software engineering, systems software and programming languages · 5Security and privacy · 2 · 2 since 2021Theory of computation · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Region-Aware Compression for Healthcare Applications in O-RAN
abstract
ABSTRACT Open Radio Access Network (O‐RAN) fronthaul links face stringent bandwidth, latency, and computational constraints, which become particularly critical when transmitting high‐resolution medical images. This paper proposes an adaptive region‐aware image compression framework for healthcare imaging over O‐RAN that reduces fronthaul load while preserving diagnostically relevant information. Each image is partitioned into Region‐of‐Interest (ROI) and Non‐ROI areas and compressed using independent quantisation parameters. An optimisation model is formulated to minimise transmitted data size subject to ROI and Non‐ROI quality constraints, end‐to‐end latency bounds and computational limits at O‐RAN nodes. The framework is evaluated using two medical imaging datasets (chest X‐rays and bone fracture X‐rays), where empirical rate–distortion and quality models are derived and validated. Results demonstrate substantial fronthaul bandwidth reduction—achieving compression ratios up to 416:1—while maintaining ROI quality and diagnostic accuracy above 97%. These findings highlight the effectiveness of region‐aware optimisation for bandwidth‐efficient healthcare imaging in O‐RAN environments.
Omar Osman, Ahmed Badawy, Saeed Salem
Expert Syst. J. Knowl. Eng.3
2025 Enhancing Vulnerability Reports With Automated and Augmented Description Summarization
abstract
Public vulnerability databases, such as the National Vulnerability Database (NVD), document vulnerabilities and facilitate threat information sharing. However, they often suffer from short descriptions and outdated or insufficient information. In this paper, we introduceZad, a system designed to enrich NVD vulnerability descriptions by leveraging external resources.Zadconsists of two pipelines: one collects and filters supplementary data using two encoders to build a detailed dataset, while the other fine-tunes a pre-trained model on this dataset to generate enriched descriptions. By addressing brevity and improving content quality,Zadproduces more comprehensive and cohesive vulnerability descriptions. We evaluateZadusing standard summarization metrics and human assessments, demonstrating its effectiveness in enhancing vulnerability information.
Hattan Althebeiti, Mohammed Alkinoon, Manar Mohaisen, Saeed Salem, DaeHun Nyang, David Mohaisen
IEEE Trans. Big Data4
2025 Semantics-Preserving Node Injection Attacks Against GNN-Based ACFG Malware Classifiers
abstract
To increase security for devices connected to the internet, research has gone into using Graph Neural Networks (GNNs) to inhibit the spread of malware through detection. GNN classifiers that use Attributed Control Flow Graphs (ACFGs) have demonstrated favorable results in classifying software binaries as malicious or benign. In this work, we show that such classifiers are vulnerable to Adversarial Examples (AEs) by proposing several grey-box adversarial attacks that perform node injection and preserve the semantics of a program. We demonstrate that adversaries can take advantage of the aggregation properties of GNNs to apply effective perturbation outside of the original ACFG nodes of a software binary through node injection. We conducted experiments on our methods and compared them against two similar semantics-preserving adversarial attacks. Our results have shown that our methods of applying perturbation through node injection can result in higher evasion rates while decreasing the amount of perturbation needed to fool detectors. Namely, we deliver an evasion rate of up to 94.83% with only 2.49% of total perturbation, in comparison with a maximum evasion of 79.50% at 2.78% perturbation by a state-of-the-art approach and only 27.44% at 2.95% perturbation by the baseline attack. Our results highlight the need for creating more robust GNN malware detectors.
Dylan Zapzalka, Saeed Salem, David Mohaisen
IEEE Trans. Dependable Secur. Comput.2
2024 A Dynamic Redeployment System for Mobile Ambulances in Qatar, Empowered by Deep Reinforcement Learning
abstract
Efficiently managing ambulance deployment is crucial for the success of emergency medical services, ensuring timely responses and life-saving care delivery. The initial ambulance location problem poses a challenge, demanding optimal deployment strategies to minimize response times and maximize coverage. Traditional approaches rely on heuristics and predetermined rules, struggling to adapt to the dynamic nature of emergencies. In response, this study proposes a dynamic ambulance redeployment system to reduce ambulance response time, increasing the chances of saving lives. The system identifies available ambulances finishing patient transports and strategically redistributes them to designated spokes, enhancing readiness for future emergencies. Addressing the inherent complexity, the study introduces a deep score network, utilizing Deep Reinforcement Learning (DRL) to train the network effectively. Our proposed approach achieved approximately 75% reduction in delays in Average Response Times (AvRT), utilizing real data from Qatar in a realistic deployment scenario. The outcome is a dynamic ambulance redeployment algorithm for real-world application, supported by experimental results using real-world data.
Reem Tluli, Ahmed Badawy, Saeed Salem, Mohamed Hardan, Sailesh Chauhan, Guillaume Alinier
IWCMC3
2024 Exposing the Limitations of Machine Learning for Malware Detection Under Concept Drift
Ahmed Abusnaina, Afsah Anwar, Muhammad Saad 0001, Abdulrahman Alabduljabbar, RhongHo Jang, Saeed Salem, David Mohaisen
WISE (2)6
2024 Industry-Specific Vulnerability Assessment
Mohammed Alkinoon, Hattan Althebeiti, Ali Alkinoon, Manar Mohaisen, Saeed Salem, David Mohaisen
WISE (5)5
2024 Mining contextually meaningful subgraphs from a vertex-attributed graph
abstract
Networks have emerged as a natural data structure to represent relations among entities. Proteins interact to carry out cellular functions and protein-Protein interaction network analysis has been employed for understanding the cellular machinery. Advances in genomics technologies enabled the collection of large data that annotate proteins in interaction networks. Integrative analysis of interaction networks with gene expression and annotations enables the discovery of context-specific complexes and improves the identification of functional modules and pathways. Extracting subnetworks whose vertices are connected and have high attribute similarity have applications in diverse domains. We present an enumeration approach for mining sets of connected and cohesive subgraphs, where vertices in the subgraphs have similar attribute profile. Due to the large number of cohesive connected subgraphs and to overcome the overlap among these subgraphs, we propose an algorithm for enumerating a set of representative subgraphs, the set of all closed subgraphs. We propose pruning strategies for efficiently enumerating the search tree without missing any pattern or reporting duplicate subgraphs. On a real protein-protein interaction network with attributes representing the dysregulation profile of genes in multiple cancers, we mine closed cohesive connected subnetworks and show their biological significance. Moreover, we conduct a runtime comparison with existing algorithms to show the efficiency of our proposed algorithm.
Riyad Hakim, Saeed Salem
BMC Bioinform.2
2023 Understanding the Country-Level Security of Free Content Websites and their Hosting Infrastructure
abstract
This paper examines free content websites (FCWs) and premium content websites (PCWs) in different countries, comparing them to general websites. The focus is on the distribution of malicious websites and their correlation with the national cyber security index (NCSI), which measures a country’s cyber security maturity and its ability to deter the hosting of such malicious websites. By analyzing a dataset comprising 1,562 FCWs and PCWs, along with Alexa’s top million websites dataset sample, we discovered that a majority of the investigated websites are hosted in the United States. Interestingly, the United States has a relatively low NCSI, mainly due to a lower score in privacy policy development. Similar patterns were observed for other countries With varying NCSI criteria. Furthermore, we present the distribution of various categories of FCWs and PCWs across countries. We identify the top hosting countries for each category and provide the percentage of discovered malicious websites in those countries. Ultimately, the goal of this study is to identify regional vulnerabilities in hosting FCWs and guide policy improvements at the country level to mitigate potential cyber threats.
Mohamed Alqadhi, Ali Alkinoon, Saeed Salem, David Mohaisen
DSAA3
2022 DL-FHMC: Deep Learning-Based Fine-Grained Hierarchical Learning Approach for Robust Malware Classification
abstract
The acceptance of the Internet of Things (IoT) for both household and industrial applications is accompanied by the rapid growth of IoT malware. With the increase of their attack surface, analyzing, understanding, and detecting IoT malicious behavior are crucial. Traditionally, machine and deep learning-based approaches are used for malware detection and behavioral understanding. However, recent research has shown the susceptibility of those approaches to adversarial attacks by introducing noise to the feature space. In this work, we introduce DL-FHMC, a fine-grained hierarchical learning approach for robust IoT malware detection. DL-FHMC utilizes Control Flow Graph (CFG)-based behavioral patterns for adversarial IoT malicious software detection. In particular, we extract a comprehensive list of behavioral patterns from a large dataset of malicious IoT binaries, represented by the shared execution flows, and use them as a modality for malicious behavior detection. Leveraging machine learning and subgraph isomorphism matching algorithms, DL-FHMC provides state-of-the-art performance in detecting malware samples and adversarial examples (AEs). We first highlight the caveats of CFG-based IoT malware detection systems, showing the adversarial capabilities in generating practical functionality-preserving AEs with reduced overhead using Graph Embedding and Augmentation (GEA) techniques. We then introduce Suspicious Behavior Detector, a component that extracts comprehensive behavioral patterns from three popular IoT malicious families, Gafgyt, Mirai, and Tsunami, for AEs detection with high accuracy. The proposed detector operates as a model-independent standalone module, with no prior assumptions of the adversarial attacks nor their configurations.
Ahmed Abusnaina, Mohammed Abuhamad, Hisham Alasmary, Afsah Anwar, RhongHo Jang, Saeed Salem, DaeHun Nyang, David Mohaisen
IEEE Trans. Dependable Secur. Comput.6
2019 A linear delay algorithm for enumerating all connected induced subgraphs
abstract
BACKGROUND: Real biological and social data is increasingly being represented as graphs. Pattern-mining-based graph learning and analysis techniques report meaningful biological subnetworks that elucidate important interactions among entities. At the backbone of these algorithms is the enumeration of pattern space. RESULTS: We propose an efficient algorithm for enumerating all connected induced subgraphs of an undirected graph. Building on this enumeration approach, we propose an algorithm for mining all maximal cohesive subgraphs that integrates vertices' attributes with subgraph enumeration. To efficiently mine all maximal cohesive subgraphs, we propose two pruning techniques that remove futile search nodes in the enumeration tree. CONCLUSIONS: Experiments on synthetic and real graphs show the effectiveness of the proposed algorithm and the pruning techniques. On enumerating all connected induced subgraphs, our algorithm is several times faster than existing approaches. On dense graphs, the proposed approach is at least an order of magnitude faster than the best existing algorithm. Experiments on protein-protein interaction network with cancer gene dysregulation profile show that the reported cohesive subnetworks are biologically interesting.
Mohammed Alokshiya, Saeed Salem, Fidaa Abed
BMC Bioinform.2
2019 On the use of usage patterns from telemetry data for test case prioritization
Jeff Anderson, Maral Azizi, Saeed Salem, Hyunsook Do
Inf. Softw. Technol.3
2017 Classifying gene coexpression networks using state subnetworks
abstract
Algorithms that map graphs into feature vectors encoding the presence/absence of specific subgraphs, have shown excellent performance in various data mining tasks. Discriminative subgraphs have been successfully utilized as features for graphs classification. Most of the existing algorithms mine for discriminative subgraphs that completely appear frequently in graphs belonging to one class label and not so frequently in the other graphs. Graphs can be missing some edges due to noise in the data generation. In this paper, we propose a scoring function for discriminative subgraph and introduce a greedy algorithm for mining discriminative patterns. Experiment on large coexpression graphs show that the proposed approach has excellent classification performance.
Bassam Qormosh, Eihab El Radie, Saeed Salem
BIBM3
2017 Mining quasi frequent coexpression subnetworks
abstract
Mining multiple gene coexpressions networks allows for identifying context-specific modules, and improving biological function prediction. Frequent subnetworks represent essential biological modules. Existing algorithms for frequent subgraph mining do not scale for large networks. In this work, we propose a greedy approach for mining approximate frequent subgraphs. Experiments on two real coexpression networks demonstrate the effectiveness of the proposed algorithm. Biological enrichment analysis of the reported patterns show that the patterns are biologically relevant and enriched with known biological processes and KEGG pathways.
Eihab El Radie, Saeed Salem
BIBM2
2017 A parallel algorithm for mining maximal frequent subgraphs
abstract
Frequent graph mining has received a lot of attention from the research community because of the increasing availability of graph data in several domains, including bioinformatics, social networks, and cyber security. On large graphs such as protein-protein interaction and gene coexpression networks, frequent subgraph mining algorithms take hours to finish. In this paper, we propose a parallel algorithm for mining maximal frequent subgraphs from edge-attributed networks. Experiments on two real tissue-specific RNA-seq expression networks and synthetic data demonstrate the effectiveness of the proposed algorithm. Moreover, biological enrichment analysis of the reported patterns show that the patterns are biologically relevant and enriched with known biological processes and KEGG pathways.
Eihab El Radie, Saeed Salem
BIBM2
2016 Customized Regression Testing Using Telemetry Usage Patterns
abstract
Pervasive telemetry in modern applications is providing new possibilities in the application of regression testing techniques. Similar to how research in bioinformatics is leading to personalized medicine, tailored to individuals, usage telemetry in modern software allows for custom regression testing, tailored to the usage patterns of an installation. By customizing regression testing based on software usage, the effectiveness of regression testing techniques can be greatly improved, leading to reduced testing costs and enhanced detection of defects that are most important to that customer. In this research, we introduce the concept of fingerprinting software usage patterns through telemetry. We provide various algorithms tocompute fingerprints and conduct an empirical study that shows that fingerprints are effective in identifying distinct usage patterns. Further, we discuss how usage fingerprints can be used to improve regression test prioritization run time by over 30 percent compared to traditional prioritization techniques.
Jeff Anderson, Hyunsook Do, Saeed Salem
ICSME3
2015 Mining maximal subnetworks from interaction network with node attributes
abstract
Detecting densely connected subgraphs is of great importance in sociology, biology and computer science disciplines where systems are often represented as a large graph. Several approaches have been proposed for detecting dense connected subgraphs in large graphs. Often these large graphs have additional attribute data characterizing either the nodes or edges of a graph. Recent research has combined the problem of dense connected subgraph detection with subspace similarity over attribute data. While detecting dense and cohesive subgraphs is desirable, the density factor can prevent the existing algorithms from reporting highly cohesive subgraphs which are not particularly dense. In this paper, we introduce an algorithm for mining maximal cohesive subgraphs from node attributed graphs. Unlike other approaches for detecting dense subgraphs, this algorithm does not require any density threshold. It discovers all maximal cohesive subgraphs regardless of their density. Experiments on real world datasets show that the proposed approach is effective in mining meaningful biological subgraphs from protein-protein interaction network, where attributes are extracted from gene expression datasets. We compare the proposed approach with the baseline technique, and results show that the proposed algorithm is much faster than the baseline algorithm.
Aditya Goparaju, Bassam Qormosh, Saeed Salem
BIBM3
2015 Striving for Failure: An Industrial Case Study about Test Failure Prediction
abstract
Software regression testing is an important, yet very costly, part of most major software projects. When regression tests run, any failures that are found help catch bugs early and smooth the future development work. The act of executing large numbers of tests takes significant resources that could, otherwise, be applied elsewhere. If tests could be accurately classified as likely to pass or fail prior to the run, it could save significant time while maintaining the benefits of early bug detection. In this paper, we present a case study to build a classifier for regression tests based on industrial software, Microsoft Dynamics AX. In this study, we examine the effectiveness of this classification as well as which aspects of the software are the most important in predicting regression test failures.
Jeff Anderson, Saeed Salem, Hyunsook Do
ICSE (2)2
2015 Experience report: Mining test results for reasons other than functional correctness
abstract
Regression testing is an important part of software development projects, and it is used to ensure software quality. Traditionally, a regression test focuses primarily on functional correctness of a modified program and is examined only when it fails, meaning it found a fault that would have otherwise been undetected. For certain application domains, regression tests for non-functional quality aspects such as performance, security, and usability could be just as important. However, those regression tests are much more costly and difficult to create, and thus many applications lack adequate non-functional regression test coverage. This adds risk of regressions in these areas as changes are made over time. In this research, we propose using metrics from passing test cases to predict quality aspects of the software beyond the traditional focus of regression tests. Our industrial case study shows that metrics such as test response time from functional regression tests are good predictors of which product areas are likely to contain certain types of non-functional performance faults. Furthermore, we show that this prediction can be improved through environmental perturbation such as the use of synthetic volume datasets or data size variation.
Jeff Anderson, Hyunsook Do, Saeed Salem
ISSRE3
2014 Improving the effectiveness of test suite through mining historical data
abstract
Software regression testing is an integral part of most major software projects. As projects grow larger and the number of tests increases, performing regression testing becomes more costly. If software engineers can identify and run tests that are more likely to detect failures during regression testing, they may be able to better manage their regression testing activities. In this paper, to help identify such test cases, we developed techniques that utilizes various types of information in software repositories. To assess our techniques, we conducted an empirical study using an industrial software product, Microsoft Dynamics AX, which contains real faults. Our results show that the proposed techniques can be effective in identifying test cases that are likely to detect failures.
Jeff Anderson, Saeed Salem, Hyunsook Do
MSR2
2011 SimClus: an effective algorithm for clustering with a lower bound on similarity
Mohammad Al Hasan, Saeed Salem, Mohammed J. Zaki
Knowl. Inf. Syst.2
2010 Sequential Data Clustering
abstract
An algorithm is presented for clustering sequential data in which each unit is a collection of vectors. An example of such a type of data is speaker data in a speaker clustering problem. The algorithm first constructs affinity matrices between each pair of units, using a modified version of the Point Distribution algorithm which is initially developed for mining patterns between vector and item data. The subsequent clustering procedure is based on fitting a Gaussian mixture model on multiple random projection matrices. The final class label of each unit is determined by voting from the results of the random projection matrices.
Jianfei Wu, Loai Al Nimer, Omar Al Azzam, Charith Chitraranjan, Saeed Salem, Anne M. Denton
ICMLA5
2009 Clustering with Lower Bound on Similarity
Mohammad Al Hasan, Saeed Salem, Benjarath Pupacdi, Mohammed J. Zaki
PAKDD2
2009 FlexSnap: Flexible Non-sequential Protein Structure Alignment
Saeed Salem, Mohammed J. Zaki, Christopher Bystroff
WABI1
2009 SPARCL: an effective and efficient algorithm for mining arbitrary shape-based clusters
Vineet Chaoji, Mohammad Al Hasan, Saeed Salem, Mohammed J. Zaki
Knowl. Inf. Syst.3
2009 Robust partitional clustering by outlier and density insensitive seeding
Mohammad Al Hasan, Vineet Chaoji, Saeed Salem, Mohammed J. Zaki
Pattern Recognit. Lett.3
2008 SPARCL: Efficient and Effective Shape-Based Clustering
abstract
Clustering is one of the fundamental data mining tasks. Many different clustering paradigms have been developed over the years, which include partitional, hierarchical, mixture model based, density-based, spectral, subspace, and so on. The focus of this paper is on full-dimensional, arbitrary shaped clusters. Existing methods for this problem suffer either in terms of the memory or time complexity (quadratic or even cubic). This shortcoming has restricted these algorithms to datasets of moderate sizes. In this paper we propose SPARCL, a simple and scalable algorithm for finding clusters with arbitrary shapes and sizes, and it has linear space and time complexity. SPARCL consists of two stages - the first stage runs a carefully initialized version of the K-means algorithm to generate many small seed clusters. The second stage iteratively merges the generated clusters to obtain the final shape-based clusters. Experiments were conducted on a variety of datasets to highlight the effectiveness, efficiency, and scalability of our approach. On the large datasets SPARCL is an order of magnitude faster than the best existing approaches.
Vineet Chaoji, Mohammad Al Hasan, Saeed Salem, Mohammed J. Zaki
ICDM3
2008 An integrated, generic approach to pattern mining: data mining template library
Vineet Chaoji, Mohammad Al Hasan, Saeed Salem, Mohammed J. Zaki
Data Min. Knowl. Discov.3
2007 ORIGAMI: Mining Representative Orthogonal Graph Patterns
abstract
In this paper, we introduce the concept of alpha-orthogonal patterns to mine a representative set of graph patterns. Intuitively, two graph patterns are alpha-orthogonal if their similarity is bounded above by alpha. Each alpha-orthogonal pattern is also a representative for those patterns that are at least beta similar to it. Given user defined alpha, beta isin [0,1], the goal is to mine an alpha-orthogonal, beta-representative set that minimizes the set of unrepresented patterns. We present ORIGAMI, an effective algorithm for mining the set of representative orthogonal patterns. ORIGAMI first uses a randomized algorithm to randomly traverse the pattern space, seeking previously unexplored regions, to return a set of maximal patterns. ORIGAMI then extracts an alpha-orthogonal, beta-representative set from the mined maximal patterns. We show the effectiveness of our algorithm on a number of real and synthetic datasets. In particular, we show that our method is able to extract high quality patterns even in cases where existing enumerative graph mining methods fail to do so.
Mohammad Al Hasan, Vineet Chaoji, Saeed Salem, Jérémy Besson, Mohammed J. Zaki
ICDM3
2005 Towards Generic Pattern Mining
Mohammed J. Zaki, Nagender Parimi, Nilanjana De, Feng Gao 0016, Benjarath Pupacdi, Joe Urban, Vineet Chaoji, Mohammad Al Hasan, Saeed Salem
ICFCA9