VLDB 2026 Research / reviewers in the wild / expert
Faheem Ullah
dblp:67/9679
· DBLP profile ↗
20ranked-venue papers
12as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 5 first-author · 3 since 2021Computer networks · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Data Breaches: What Happened over the Last 20 Years?
Faheem Ullah, Uswa Fatima, Muhammad Imran Taj 0001 |
DATA | 1 |
| 2025 | Characterizing Vulnerabilities in Microservices: Source, Age and Severity
Samuel Beahan, Faheem Ullah, Lachlan Chalmers, Uswa Fatima, Mojtaba Shahin |
ICSA | 2 |
| 2025 | Design and Implementation of Fragmented Clouds for Evaluation of Distributed DatabasesabstractIn this paper, we present a Fragmented Hybrid Cloud (FHC) that provides a unified view of multiple geographically distributed private cloud datacenters. FHC leverages a fragmented usage model in which outsourcing is bi-directional across private clouds that can be hosted by static and mobile entities. The mobility aspect of private cloud nodes has important impacts on the FHC performance in terms of latency and network throughput that are reversely proportional to time-varying distances among different nodes. Mobility also results in intermittent interruption among computing nodes and network links of FHC infrastructure. To fully consider mobility and its consequences, we implemented a layered FHC that leverages Linux utilities and bash-shell programming. We also evaluated the impact of the mobility of nodes on the performance of distributed databases as a result of time-varying latency and bandwidth, downsizing and upsizing cluster nodes, and network accessibility. The findings from our extensive experiments provide deep insights into the performance of well-known big data databases, such as Cassandra, MongoDB, Redis, and MySQL, when deployed on a FHC. Yaser Mansouri, Faheem Ullah, Shagun Dhingra, Muhammad Ali Babar 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2024 | Evaluation of distributed data processing frameworks in hybrid cloudsabstractDistributed data processing frameworks (e.g., Hadoop, Spark, and Flink) are widely used to distribute data among computing nodes of a cloud. Recently, there have been increasing efforts aimed at evaluating the performance of distributed data processing frameworks hosted in private and public clouds. However, there is a paucity of research on evaluating the performance of these frameworks hosted in a hybrid cloud, which is an emerging cloud model that integrates private and public clouds to use the best of both worlds. Therefore, in this paper, we evaluate the performance of Hadoop, Spark, and Flink in a hybrid cloud in terms of execution time, resource utilization, horizontal scalability, vertical scalability, and cost. For this study, our hybrid cloud consists of OpenStack (private cloud) and MS Azure (public cloud). We use both batch and iterative workloads for the evaluation. Our results show that in a hybrid cloud (i) the execution time increases as more nodes are borrowed by the private cloud from the public cloud, (ii) Flink outperforms Spark, which in turn outperforms Hadoop in terms of execution time, (iii) Hadoop transfers the largest amount of data among the nodes during the workload execution while Spark transfers the least amount of data, (iv) all three frameworks horizontally scale better as compared to vertical scaling, and (v) Spark is found to be least expensive in terms of $ cost for data processing while Hadoop is found the most expensive. Faheem Ullah, Shagun Dhingra, Xiaoyu Xia 0001, Muhammad Ali Babar 0001 |
J. Netw. Comput. Appl. | 1 |
| 2024 | 6D object pose estimation based on dense convolutional object center voting with improved accuracy and efficiency
Faheem Ullah, Wu Wei 0001, Zhun Fan, Qiuda Yu |
Vis. Comput. | 1 |
| 2023 | Co-Tuning of Cloud Infrastructure and Distributed Data Processing PlatformsabstractDistributed Data Processing Platforms (e.g., Hadoop, Spark, and Flink) are widely used to store and process data in a cloud environment. These platforms distribute the storage and processing of data among the computing nodes of a cloud. The efficient use of these platforms requires users to (i) configure the cloud i.e., determine the number and type of computing nodes, and (ii) tune the configuration parameters (e.g., data replication factor) of the platform. However, both these tasks require in-depth knowledge of the cloud infrastructure and distributed data processing platforms. Therefore, in this paper, we first study the relationship between the configuration of the cloud and the configuration of distributed data processing platforms to determine how cloud configuration impacts platform configuration. After understanding the impacts, we propose a co-tuning approach for recommending optimal co-configuration of cloud and distributed data processing platforms. The proposed approach utilizes machine learning and optimization techniques to maximize the performance of the distributed data processing system deployed on the cloud. We evaluated our approach for Hadoop, Spark, and Flink in a cluster deployed on the OpenStack cloud. We used various benchmarking workloads in our evaluation. Our results reveal that, in comparison to default settings, our co-tuning approach reduces execution time by 17.5% and ${\$}$ cost by 14.9% solely via configuration tuning. Isuru Dharmadasa, Faheem Ullah |
IEEE Big Data | 2 |
| 2023 | An Exploratory Study of Vulnerabilities in Big Data SystemsabstractThe use of big data systems has become prevalent across sensitive domains, including health, defense and finance, among others. These big data systems are often complex and with complexity often comes vulnerabilities. Most big data systems also have Application Programming Interfaces (API) where these vulnerabilities can be present. This results in introducing security risks into the projects using these big data tools. To address this issue, this paper presents and uses a research process that can be used to assess the state of big data security in open source projects. The study examines tool versioning in various open source projects to uncover the vulnerabilities present and then assess the significant factors of these vulnerabilities. Furthermore, we provide insights into the vulnerabilities of different big data system APIs and how each of the vulnerabilities can manifest in a project using the specific tool. This study serves as a guideline for big data system developers to develop highly secure data-intensive systems. CCS CONCEPTS • Security and privacy → Database and storage security; Software and application security; Nikolas Tyllis, Faheem Ullah |
IEEE Big Data | 2 |
| 2023 | Guidance Models for Designing Big Data Cyber Security Analytics Systems
Faheem Ullah, Muhammad Ali Babar 0001 |
ECSA | 1 |
| 2023 | Defending SDN against packet injection attacks using deep learningabstractThe (logically) centralized architecture of software-defined networks makes them an easy target for packet injection attacks. In these attacks, the attacker injects malicious packets into the SDN network to affect the services and performance of the SDN controller and overflows the capacity of the SDN switches. Such attacks have been shown to ultimately stop the network functioning in real-time, leading to network breakdowns. There have been significant works on detecting and defending against similar DoS attacks in non-SDN networks, but detection and protection techniques for SDN against packet injection attacks are still in their infancy. Furthermore, many of the proposed solutions have been shown to be easily bypassed by simple modifications to the attacking packets or by altering the attacking profile. In this paper, we develop novel Graph Convolutional Neural Network models and algorithms for grouping network nodes/users into security classes by learning from network data. We start with two simple classes - nodes that engage in suspicious packet injection attacks and nodes that are not. From these classes, we then partition the network into separate segments with different security policies using distributed Ryu controllers in an SDN network. We show in experiments on an emulated SDN that our detection solution outperforms alternative approaches with above 99% detection accuracy for various types (both old and new) of injection attacks. More importantly, our mitigation solution maintains continuous functions of non-compromised nodes while isolating compromised/suspicious nodes in real-time. All code and data are publicly available for the reproducibility of our results. Anh Tuan Phu, Faheem Ullah, Tanvir Ul Huque, Ranesh Kumar Naha, Muhammad Ali Babar 0001 |
Comput. Networks | 3 |
| 2023 | Resource Utilization of Distributed Databases in Edge-Cloud EnvironmentabstractA benchmark study of modern distributed databases (DDBs) (e.g., Cassandra, MongoDB, Redis, and MySQL) is an important source of information for selecting the right technology for managing data in edge–cloud deployments. While most of the existing studies have investigated the performance and scalability of DDBs in cloud computing, there is a lack of focus on resource utilization (e.g., energy, bandwidth, and storage consumption) of workload offloading for DDBs deployed in edge–cloud environments. For this purpose, we conducted experiments on various physical and virtualized computing nodes, including variously powered servers, Raspberry Pi, and hybrid cloud (OpenStack and Azure). Our extensive experimental results reveal insights into which database under which offloading scenario is more efficient in terms of energy, bandwidth, and storage consumption. Yaser Mansouri, Victor Prokhorenko, Faheem Ullah, Muhammad Ali Babar 0001 |
IEEE Internet Things J. | 3 |
| 2022 | Scalable Containerized Pipeline for Real-time Big Data AnalyticsabstractWith the widespread usage of IoT, processing data streams in real-time have become very important. The traditional data-stream processing systems are inefficient in processing big data for detecting anomalies, classifications, clustering, and prediction in real-time using minimal resources. In this paper, we address this limitation by proposing a scalable pipeline for real-time processing of big data streams. Our proposed solution is capable of dynamically managing resources for different components of the pipeline using automatic scaling. The pipeline is containerized and deployed on a Kubernetes cluster. The proposed scalable pipeline is evaluated using a case study of anomaly detection in IoT data. The proposed solution yields a $\times 1.31$ to $\times 2.4$ increase in throughput, and $\times 32$ to $\times 80$ decreased latency compared to the commonly used static resource allocation strategy for data pipelines. Rana Aurangzaib, Waheed Iqbal, Muhammad Abdullah 0004, Faisal Bukhari, Faheem Ullah, Abdelkarim Erradi |
CloudCom | 5 |
| 2022 | Design and evaluation of adaptive system for big data cyber security analytics
Faheem Ullah, Muhammad Ali Babar 0001, Aldeida Aleti |
Expert Syst. Appl. | 1 |
| 2022 | On the scalability of Big Data Cyber Security Analytics systems
Faheem Ullah, Muhammad Ali Babar 0001 |
J. Netw. Comput. Appl. | 1 |
| 2021 | Automated Security Assessment for the Internet of ThingsabstractInternet of Things (IoT) based applications face an increasing number of potential security risks, which need to be systematically assessed and addressed. Expert-based manual assessment of IoT security is a predominant approach, which is usually inefficient. To address this problem, we propose an automated security assessment framework for IoT networks. Our framework first leverages machine learning and natural language processing to analyze vulnerability descriptions for predicting vulnerability metrics. The predicted metrics are then input into a two-layered graphical security model, which consists of an attack graph at the upper layer to present the network connectivity and an attack tree for each node in the network at the bottom layer to depict the vulnerability information. This security model automatically assesses the security of the IoT network by capturing potential attack paths. We evaluate the viability of our approach using a proof-of-concept smart building system model which contains a variety of real-world IoT devices and poten-tial vulnerabilities. Our evaluation of the proposed framework demonstrates its effectiveness in terms of automatically predicting the vulnerability metrics of new vulnerabilities with more than 90% accuracy, on average, and identifying the most vulnerable attack paths within an IoT network. The produced assessment results can serve as a guideline for cybersecurity professionals to take further actions and mitigate risks in a timely manner. Xuanyu Duan, Mengmeng Ge 0001, Triet Huynh Minh Le, Faheem Ullah, Shang Gao 0003, Xuequan Lu, Muhammad Ali Babar 0001 |
PRDC | 4 |
| 2019 | QuickAdapt: Scalable Adaptation for Big Data Cyber Security AnalyticsabstractBig Data Cyber Security Analytics (BDCA) leverages big data technologies for collecting, storing, and analyzing a large volume of security events data to detect cyber-attacks. Accuracy and response time, being the most important quality concerns for BDCA, are impacted by changes in security events data. Whilst it is promising to adapt a BDCA system's architecture to the changes in security events data for optimizing accuracy and response time, it is important to consider large search space of architectural configurations. Searching a large space of configurations for potential adaptation incurs an overwhelming adaptation time, which may cancel the benefits of adaptation. We present an adaptation approach, QuickAdapt, to enable quick adaptation of a BDCA system. QuickAdapt uses descriptive statistics (e.g., mean and variance) of security events data and fuzzy rules to (re) compose a system with a set of components to ensure optimal accuracy and response time. We have evaluated QuickAdapt for a distributed BDCA system using four datasets. Our evaluation shows that on average QuickAdapt reduces adaptation time by 105× with a competitive adaptation accuracy of 70% as compared to an existing solution. Faheem Ullah, Muhammad Ali Babar 0001 |
ICECCS | 1 |
| 2019 | An Architecture-Driven Adaptation Approach for Big Data Cyber Security AnalyticsabstractBig Data Cyber Security Analytics (BDCA) systems leverage big data technologies (e.g., Hadoop and Spark) for collecting, storing, and analyzing large volume of security event data to detect cyber-attacks. Accuracy and response time are the two most important quality concerns for BDCA systems. However, the frequent changes in the operating environment of a BDCA system (such as quality and quantity of security event data) significantly impact these qualities. In this paper, we first study the impact of such environmental changes. We then present ADABTics, an architecture-driven adaptation approach that (re)composes the system at runtime with a set of components to ensure optimal accuracy and response time. We finally evaluate our approach both in a single node and multinode settings using a Hadoop-based BDCA system and different adaptation scenarios. Our evaluation shows that on average ADABTics improves BDCA's accuracy and response time by 6.06% and 23.7%respectively. Faheem Ullah, Muhammad Ali Babar 0001 |
ICSA | 1 |
| 2019 | Quantifying the Impact of Design Strategies for Big Data Cyber Security Analytics: An Empirical InvestigationabstractBig Data Cyber Security Analytics (BDCA) systems use big data technologies (e.g., Hadoop and Spark) for collecting, storing, and analyzing a large volume of security event data to detect cyber-attacks. The state-of-the-art uses various design strategies (e.g., feature selection and alert ranking) to help BDCA systems to achieve the desired levels of accuracy and response time. However, the use of these strategies in the state-of-the-art is not consistent, which exposes a lack of consensus on "when to use (and not to use) these design strategies?" In this paper, we follow a systematic experimentation framework to quantify the impact of four design strategies on the accuracy and response time with respect to three contextual factors i.e., security data, machine learning model employed in the system, and the execution mode of the system. For the aimed quantification, we performed experiments on a Hadoop-based BDCA system using four security datasets, five machine learning models, and three execution modes. Our findings lead us to formulate a set of design guidelines that will help researchers and practitioners to decide when to use (and not to use) the design strategies. Faheem Ullah, Muhammad Ali Babar 0001 |
PDCAT | 1 |
| 2019 | Architectural Tactics for Big Data Cybersecurity Analytics Systems: A Review
Faheem Ullah, Muhammad Ali Babar 0001 |
J. Syst. Softw. | 1 |
| 2018 | Data exfiltration: A review of external attack vectors and countermeasures
Faheem Ullah, Matthew Edwards 0001, Rajiv Ramdhany, Ruzanna Chitchyan, Muhammad Ali Babar 0001, Awais Rashid |
J. Netw. Comput. Appl. | 1 |
| 2017 | Security Support in Continuous Deployment PipelineabstractContinuous Deployment (CD) has emerged as a new practice in the software industry to continuously and automatically deploy software changes into production. Continuous Deployment Pipeline (CDP) supports CD practice by transferring the changes from the repository to production. Since most of the CDP components run in an environment that has several interfaces to the Internet, these components are vulnerable to various kinds of malicious attacks. This paper reports our work aimed at designing secure CDP by utilizing security tactics. We have demonstrated the effectiveness of five security tactics in designing a secure pipeline by conducting an experiment on two CDPs - one incorporates security tactics while the other does not. Both CDPs have been analyzed qualitatively and quantitatively. We used assurance cases with goal-structured notations for qualitative analysis. For quantitative analysis, we used penetration tools. Our findings indicate that the applied tactics improve the security of the major components (i.e., repository, continuous integration server, main server) of a CDP by controlling access to the components and establishing secure connections. Faheem Ullah, Adam Johannes Raft, Mojtaba Shahin, Mansooreh Zahedi, Muhammad Ali Babar 0001 |
ENASE | 1 |