Samee Ullah Khan

dblp:k/SameeUllahKhan · also Samee Khan, Samee U. Khan · DBLP profile ↗
← Back
180ranked-venue papers
21as first author
31since 2021 · last 2026
0000-0002-5712-7796ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 92 · 12 first-author · 6 since 2021Computer networks · 24 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 20 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 3 since 2021Databases, data management, data science and information retrieval · 12 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 9Security and privacy · 7 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 BRACTIVE: A Brain Activation Approach to Human Visual Brain Learning
abstract
The human brain is a highly efficient processing unit, and understanding how it works can inspire new algorithms and architectures in machine learning. In this work, we introduce a novel framework named Brain Activation Network (BRACTIVE), a transformer-based approach to studying the human visual brain. The primary objective of BRACTIVE is to align the visual features of subjects with their corresponding brain representations using functional Magnetic Resonance Imaging (fMRI) signals. It enables us to identify the brain's Regions of Interest (ROIs) in the subjects. Unlike previous brain research methods, which can only identify ROIs for one subject at a time and are limited by the number of subjects, BRACTIVE automatically extends this identification to multiple subjects and ROIs. Our experiments demonstrate that BRACTIVE effectively identifies person-specific regions of interest, such as face and body-selective areas, aligning with neuroscience findings and indicating potential applicability to various object categories. More importantly, we found that leveraging human visual brain activity to guide deep neural networks enhances performance across various benchmarks. It encourages the potential of BRACTIVE in both neuroscience and machine intelligence studies.
Xuan-Bac Nguyen, Hojin Jang, Xin Li 0005, Samee Ullah Khan, Pawan Sinha, Khoa Luu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Multimodal feature fusion for human activity recognition using human centric temporal transformer
Samee Ullah Khan, Maryam Sultana, Sufyan Danish, Norah Saleh Alghamdi, Suchang Woo, Dong-Gyu Lee 0001, Sangtae Ahn
Eng. Appl. Artif. Intell.1
2025 Action understanding in low-light and pitch-dark conditions: A comprehensive survey
Muhammad Munsif, Samee Ullah Khan, Noman Khan, Altaf Hussain 0002, Sung Wook Baik
Eng. Appl. Artif. Intell.2
2025 Brainformer: Mimic human visual brain functions to machine vision models via fMRI
Xuan-Bac Nguyen, Xin Li 0005, Pawan Sinha, Samee Ullah Khan, Khoa Luu
Neurocomputing4
2025 Electrical Load Forecasting Over Multihop Smart Metering Networks With Federated Learning
abstract
Electric load forecasting is essential for power management and stability in smart grids. This is mainly achieved via advanced metering infrastructure, where smart meters (SMs) record household energy data. Traditional machine learning (ML) methods are often employed for load forecasting, but require data sharing, which raises data privacy concerns. Federated learning (FL) can address this issue by running distributed ML models at local SMs without data exchange. However, current FL-based approaches struggle to achieve efficient load forecasting due to imbalanced data distribution across heterogeneous SMs. This paper presents a novel personalized federated learning (PFL) method for high-quality load forecasting in metering networks. A meta-learning-based strategy is developed to address data heterogeneity at local SMs in the collaborative training of local load forecasting models. Moreover, to minimize the load forecasting delays in our PFL model, we study a new latency optimization problem based on optimal resource allocation at SMs. A theoretical convergence analysis is also conducted to provide insights into FL design for federated load forecasting. Extensive simulations from real-world datasets show that our method outperforms existing approaches regarding better load forecasting and reduced operational latency costs.
Ratun Rahman, Pablo Moriano, Samee Ullah Khan, Dinh C. Nguyen
IEEE Internet Things J.3
2025 Big Data Analysis for Industrial Activity Recognition Using Attention-Inspired Sequential Temporal Convolution Network
abstract
Deep-learning-based human activity recognition (HAR) methods have significantly transformed a wide range of domains over recent years. However, the adoption of Big Data techniques in industrial applications remains challenging due to issues such as generalized weight optimization, diverse viewpoints, and the complex spatiotemporal features of videos. To address these challenges, this work presents an industrial HAR framework consisting of two main phases. First, a squeeze bottleneck attention block (SBAB) is introduced to enhance the learning capabilities of the backbone model for contextual learning, which allows for the selection and refinement of an optimal feature vector. In the second phase, we propose an effective sequential temporal convolutional network (STCN), which is designed in parallel fashion to mitigate the issues of exploding and vanishing gradients associated with sequence learning. The high-dimensional spatiotemporal feature vectors from the STCN undergo further refinement through our proposed SBAB in a sequential manner, to optimize the features for HAR and enhance the overall performance. The efficacy of the proposed framework is validated through extensive experiments on six datasets, including data from industrial and general activities.
Altaf Hussain 0002, Tanveer Hussain 0001, Waseem Ullah, Samee Ullah Khan, Min Je Kim, Khan Muhammad 0001, Javier Del Ser, Sung Wook Baik
IEEE Trans. Big Data4
2024 High-throughput Real-time Edge Stream Processing with Topology-Aware Resource Matching
abstract
With the proliferation of Internet of Things (IoT) devices, real-time stream processing at the edge of the network has gained significant attention. However, edge stream processing systems face substantial challenges due to the heterogeneity and constraints of computational and network resources and the intricacies of multi-tenant application hosting. An optimized placement strategy for edge application topology becomes crucial to leverage the advantages offered by Edge computing and enhance the throughput and end-to-end latency of data streams. This paper presents Beaver, a resource scheduling framework designed to deploy stream processing topologies across distributed edge nodes efficiently. Its core is a novel scheduler that employs a synergistic integration of graph partitioning within application topologies and a two-sided matching technique to optimize the strategic placement of stream operators. Beaver aims to achieve optimal performance by minimizing bottlenecks in the network, memory, and CPU resources at the edge. We implemented a prototype of Beaver using Apache Storm and Kubernetes orchestration engine and evaluated its performance using an open-source real-time IoT benchmark (RIoTBench). Compared to state-of-the-art techniques, experimental evaluations demonstrate at least 1.6× improvement in the number of tuples processed within a one-second deadline under varying network delay and bandwidth scenarios.
Samee Ullah Khan, Xiaobo Zhou 0002, Palden Lama
CCGrid2
2024 Dynamic Task Offloading in Connected Vehicles: Leveraging a Graph Neural Networks Approach for Multi-hop Search
abstract
The vehicular edge computing model provides computational support to nearby vehicles requiring low-latency computation. However, existing algorithms typically only consider immediate neighbor nodes as potential vehicles for computation offloading. In this study, we address this limitation by modeling the task offloading problem using a Graph Neural Network approach. This approach incorporates dynamic properties of vehicles and edges, facilitating a more effective selection of vehicles for task offloading. By leveraging message passing and aggregate feature techniques, our proposed approach expands the search space for source vehicles, thereby distributing the workload across a wider spectrum, particularly benefiting vehicles that frequently generate computational tasks. Evaluation of our proposed system demonstrates significant improvements in the utilization factor of available vehicles, a reduction in queuing states, and an overall system efficiency of 92%.
Asad Waqar Malik, Samee Ullah Khan
IPCCC2
2024 AI-driven behavior biometrics framework for robust human activity recognition in surveillance systems
Altaf Hussain 0002, Samee Ullah Khan, Noman Khan, Mohammad Shabaz, Sung Wook Baik
Eng. Appl. Artif. Intell.2
2024 An intelligent correlation learning system for person Re-identification
Samee Ullah Khan, Noman Khan, Tanveer Hussain 0001, Sung Wook Baik
Eng. Appl. Artif. Intell.1
2024 Deep multi-scale pyramidal features network for supervised video summarization
Habib Khan, Tanveer Hussain 0001, Samee Ullah Khan, Zulfiqar Ahmad Khan 0002, Sung Wook Baik
Expert Syst. Appl.3
2024 Toward Efficient Fire Detection in IoT Environment: A Modified Attention Network and Large-Scale Data Set
abstract
Advancements in deep learning and the Internet of Things (IoT) enable early fire detection through vision-based systems, reducing ecological, social, and economic damage. These systems necessitate lightweight, cost-effective convolutional neural networks (CNNs) for real-time operation. Effective deployment on AI-assisted edge devices is crucial for optimal performance. To mitigate this problem, we present the optimized fire attention network (OFAN) for effective and efficient fire detection. In the re-engineered attention block, we swapped the convolution layers by dilated variants and integrated additional dense layers to capture global context and refine more weight optimization. We calibrate the OFAN for real-time processing using a lightweight and efficient feature extractor backbone model. Additionally, a challenging fire dataset is a critical contribution that contains extremely diverse, blazing, and non-fire, captured in lighting and foggy environments. It advances traditional fire detection samples by considering low-light and foggy conditions. A comprehensive experiment is conducted over three widely used fire detection datasets, and our proposed OFAN outperforms state-of-the-art. The proposed OFAN achieved 96.23%, 96.54% and 94.63% accuracies over BoWFire, FD and the newly proposed DiverseFire dataset, respectively. Our research sets a standard for fire detection over edge devices, offering improved accuracy and better frames per second (FPS) performance.
Naqqash Dilshad, Samee Ullah Khan, Norah Saleh Alghamdi, Tarik Taleb, Jaeseung Song
IEEE Internet Things J.2
2024 Multi-scale feature reconstruction network for industrial anomaly detection
abstract
Unsupervised anomaly detection techniques, which operate without prior knowledge of anomalies, have garnered significant attention in industrial inspection due to their adaptability and generalization. Therefore, knowledge-based computer vision techniques have been broadly applied to identify unusual image patterns. However, real-time industrial applications present challenges such as limited anomalous samples, inadequate defect knowledge, and complex background textures. These factors lead to difficulties in accurately identifying defect regions, and conventional auto-encoder networks often struggle to overcome these issues.To address these limitations, we propose a multi-scale feature reconstruction (MSFR) network specifically designed for domain shift scenarios. Our approach employs a pyramidal vision transformer network (PVTN) to reconstruct multi-scale feature maps, capturing discriminative features at various scales. Additionally, a pre-trained module extracts multi-level features at the same scale, and a dedicated feature matching module enhances accuracy by improving the alignment probability between features. The MSFR strategy surpasses conventional auto-encoders by filtering pixel-level information at multiple depths. Empirical evaluations were conducted using benchmark datasets such as MVTec AD and AeBAD-S. Furthermore, an extensive ablation study demonstrates the effectiveness and viability of the proposed MSFR approach for industrial anomaly detection tasks. The experimental results show that the proposed model significantly outperforms recent approaches, making it highly suitable for real-world industrial applications, particularly in manufacturing. • MSFR is introduced for industrial anomaly detection to handle domain shifts generalization. • A multi-scale features network is developed to efficiently detect defects. • A pyramidal network is proposed to cover the limitations of traditional autoencoders. • Comprehensive experiments are generated over AeBAD-S and MVTec IAD benchmark datasets.
Ehtesham Iqbal 0002, Samee Ullah Khan, Sajid Javed, Brain Moyo, Yahya Zweiri, Yusra Abdulrahman
Knowl. Based Syst.2
2024 Contextual visual and motion salient fusion framework for action recognition in dark environments
Muhammad Munsif, Samee Ullah Khan, Noman Khan, Altaf Hussain 0002, Min Je Kim, Sung Wook Baik
Knowl. Based Syst.2
2024 Deep-ReID: deep features and autoencoder assisted image patching strategy for person re-identification in smart cities surveillance
Samee Ullah Khan, Tanveer Hussain 0001, Amin Ullah, Sung Wook Baik
Multim. Tools Appl.1
2024 Predicting Device Anomalous Condition in a Collaborated Industrial Environment
abstract
The industrial environment augments resource-constrained devices to bring services closer to autonomous devices. However, over time, these devices get overburdened due to computational workload, which results in degraded network performance. Therefore, the devices are programmed to share resources with nearby devices. However, owing to real-time collaboration, there is the possibility that the device moves to an undefined state and starts behaving maliciously. This can impact the entire collaborative environment laid to meet the industrial product deadline. In this article, we propose an industrial simulation framework that enables the resource-sharing environment and identifies the undefined device behavior. Furthermore, our detection scheme is based on an intelligent model trained on device behavior through the machine-in-a-loop mechanism and deployed at network intersections, i.e., edge nodes. The proposed technique improves the efficiency of the collaborative network by 30%.
Unaiza Alvi, Asad Waqar Malik, Anis Ur Rahman 0001, Muazzam Ali Khan, Samee Ullah Khan
IEEE Trans. Ind. Informatics5
2023 Selecting Distinctive-Variant Training Samples Base on Intra-class Similarity
Hang Diao, Zhengchang Liu, Fan Zhang 0003, Jiaqing Huang, Feiyu Zhou, Samee Ullah Khan
ICANN (9)6
2023 Find Important Training Dataset by Observing the Training Sequence Similarity
Zhengchang Liu, Hang Diao, Fan Zhang 0003, Samee Ullah Khan
ICANN (3)4
2023 Efficient Person Reidentification for IoT-Assisted Cyber-Physical Systems
abstract
The main objective of this study is to propose a cyber–physical system (CPS)-based person reidentification (P-ReID) framework for smart surveillance. The Internet of Things (IoT)-based interconnected vision sensors in smart cities are considered essential elements of a CPS, and contribute significantly to urban security. However, the reidentification of targeted persons using emerging edge AI techniques still faces certain challenges. To improve efficiency at the edge and overcome the traditional sensing of video cameras, we employed an AI-based P-ReID framework for CPS that is functional in IoT environments. In addition, we present dual attention dilated network (DADNet), which integrates an energy-efficient convolutional neural network (CNN) with a self-attention module to substantially improve the person matching probability. Furthermore, we applied dual feature fusion to intelligently integrate discriminative and robust features using early and late fusion strategies that allow DADNet to significantly consider the foreground and marginally utilize the background information. Furthermore, we impose diversity orthogonality regularization over several CNN layers, which boosts the performance of DADNet, resulting in an appropriate usage over IoT networks. A comprehensive set of ablation studies, comparison with other state-of-the-art approaches, and a time complexity analysis confirm the strength of our DADNet for reidentification tasks in AI-enabled IoT settings that are well suited for a CPS.
Samee Ullah Khan, Ijaz Ul Haq, Noman Khan, Amin Ullah, Khan Muhammad 0001, Huiling Chen 0001, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Internet Things J.1
2023 AI-Assisted Hybrid Approach for Energy Management in IoT-Based Smart Microgrid
abstract
Power generation (PG) prediction from renewable energy sources (RESs) plays a vital role in effective energy management in smart cities. However, harnessing the potential of edge intelligence in well-controlled Internet of Things (IoT) networks poses significant challenges. To address this, we propose an IoT-based framework for intelligent and efficient PG prediction in smart microgrids. The framework begins by acquiring data from various RESs, including wind and solar. Before the training process, the data undergoes cleaning and normalization steps that use denoising and cleansing filters. For forecasting renewable energy (RE), we introduce a hybrid model that integrates a multi-head attention (MHA)-based deep autoencoder (AE) with extreme gradient boosting (XGB) algorithm. The AE’s encoder component extracts discriminative features from the cleaned data sequence, which are then learned by XGB to provide a final PG forecast. This edge computing layer facilitates information sharing through fog computing, which ensures power balancing between suppliers and consumers. Furthermore, the framework also incorporates various power consumption (PC) sectors and entities within smart cities, such as transportation and healthcare, to ensure efficient management. We evaluate the proposed hybrid model using publicly accessible benchmarks and locally gathered data sets, demonstrating state-of-the-art performance in terms of error metrics. The computational complexity of the proposed model is also suitable for resource-constrained IoT devices connected to a shared IoT-Fog setup, enabling seamless communication with smart microgrids for effective power management.
Noman Khan, Samee Ullah Khan, Fath U Min Ullah, Mi Young Lee, Sung Wook Baik
IEEE Internet Things J.2
2023 Deep Learning Assists Surveillance Experts: Toward Video Data Prioritization
abstract
Video summarization (VS) suppresses high-dimensional (HD) video data by only extracting the important information. However, prior research has not focused on the need for surveillance VS, that is used for many applications to assist video surveillance experts, including video retrieval and data storage. In addition, mainstream techniques commonly use two-dimensional (2-D) deep models for VS, ignoring event occurrences. Accordingly, we present a two-fold 3-D deep learning-assisted VS framework. First, we employ an inflated 3-D ConvNet model to extract temporal features; these features are optimized using a proposed encoder mechanism. The input video is temporally segmented using a feature comparison technique for selecting a single frame from each video segment. The segmented shots are evaluated using our novel shot segmentation evaluation scheme and are input into a saliency computation mechanism for keyframe selection in a second fold. Qualitative and quantitative analyses over VS benchmarks and surveillance videos demonstrate the superior performance of our framework, with 0.3- and 4.2-unit increases in the F1 scores for YouTube and title-based video summarization datasets, respectively. Along with accurate VS, a key contribution of our study is the novel shot segmentation criterion prior to VS, which can be used as a benchmark in future research to effectively prioritize HD visual data.
Tanveer Hussain 0001, Fath U Min Ullah, Samee Ullah Khan, Amin Ullah, Umair Haroon, Khan Muhammad 0001, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Trans. Ind. Informatics3
2022 Mining Influential Training Data by Tracing Influence on Hard Validation Samples
abstract
The ever-growing deep learning model size is constantly driven by the ever-growing dataset size. Mining the influential training data has significant payoff of either reducing the training time, model complexity as well as potentially increasing the model accuracy. In this paper, we propose a few approaches, e.g. classifying the validation dataset into easy, medium and hard levels, introducing influence value by calculating each training data on the hard validation data, to co-prune the validation dataset and the training dataset. Empirically we conclude that the portion of the hard validation data could be used to mine the most influential training data, whereby reducing the training dataset size by 50% without losing accuracy in our experiments.
Qikai Zhang, Fan Zhang 0003, Samee Ullah Khan
ICTAI3
2022 HCA Operator: A Hybrid Cloud Auto-scaling Tooling for Microservice Workloads
abstract
Elastic cloud platform, e.g. Kubernetes, enables dy-namically scale in or out computing resources in accordance with the workloads fluctuation. As the cloud evolves to hybrid, where public and private clouds co-exist as the underline substrate, autoscaling applications within a hybrid cloud is no longer straightforward. The difficulty lies in all aspects, e.g. global load balancing, hybrid-cloud monitoring and alerting, storage sharing and replication, security and privacy, etc. However, it will significantly pay off if hybrid-cloud autoscaling is supported and boundless computing resources can be utilized per request. In this paper, we design Hybrid Cloud Autoscaler Operator (HCA Operator), a customized Kubernetes Controller that leverages the Kubernetes Custom Resource to auto-scale microservice applications across hybrid clouds. HCA Operator load balances across hybrid clouds, monitors metrics, and autoscales to des-tination clusters that exist in other clouds. We discuss the implementation details and perform experiments in a hybrid cloud environment. The experimental results demonstrate that if the workload changes quickly, our Operator can properly auto-scale the microservice applications across hybrid cloud in order to meet the Service Level Agreement (SLA) requirements.
Fan Zhang 0003, Samee Ullah Khan
MSN3
2022 Learning to rank: An intelligent system for person reidentification
abstract
Person reidentification (P-Reid) is an emerging research domain in the field of information retrieval that has gained exponential growth due to its wide range of applications in pedestrian tracking and crime prevention. The primary goal of P-Reid is to recognize a person based on previous appearance in multiview surveillance videos. The mainstream approaches apply fully supervised learning techniques that have poor scalability when deployed in complex real-world scenes, due to the overfitting problem, caused by the lack of sufficient annotated data. Further, optimization of these models for unlabeled data in real-time surveillance is a challenging task. To tackle these issues, an intelligent framework (LR-Net) is proposed, consisting of three tiers including fine-tuning (FT), siamese network (SN), and fusion strategy (FS). In the first tier, a deep learning model is fine-tuned for P-Reid that can handle both labeled and unlabeled data. Next, with the assistance of transfer learning, an SN is proposed that has a strong discriminative capability in terms of similarity between a pair of images. Finally, a learning-to-rank strategy is applied to optimize the learning capability of the SN, in which a triplet network extracts spatial-temporal patterns from unlabeled samples. In addition, a bayesian fusion model (BFM) is introduced to integrate the spatiotemporal and visual features, which yields 4.4%, 9.3%, and 0.8% improvement in the matching score over Market-1501, DukeMCMT-reID, and CUHK03 data sets, respectively. The conducted experiments and ablation study on the benchmark data sets empirically validate the proposed system, which obtains a high Rank-1 score as compared with the state-of-the-art (SOTA) methods.
Samee Ullah Khan, Ijaz Ul Haq, Noman Khan, Khan Muhammad 0001, Mohammad Hijji, Sung Wook Baik
Int. J. Intell. Syst.1
2022 A modified secure hash design to circumvent collision and length extension attacks
Zeyad A. Al-Odat, Samee Ullah Khan, Eman Al-qtiemat
J. Inf. Secur. Appl.2
2022 Online Algorithms for the Interval Scheduling Problem in the Cloud: Affinity Pair Threshold Based Approaches
abstract
In the interval scheduling problem, jobs have known start and end times (referred to as job intervals) and must be assigned to processing nodes for their whole duration. Although the problem originally stems from the resource allocation demands of resident processes in operating systems, it found a renewed interest in the Cloud context, both in IaaS and SaaS, since reservations for virtual machines and services often have known activation intervals. A common objective of interval scheduling is to minimize busy time of machines which relates (among others) to minimizing the number of machines participating in the computation. As a consequence, bin packing techniques have been applied in the past. In this paper we tackle the online version of the problem, whereby future job arrivals are unknown. We propose novel algorithms that work as a pre-processing step to any bin packing scheme by offering recommendations that are enforced in all packing decisions. Job overlaps are used to characterize pairwise job affinity and subsequently provide threshold based job allocation recommendations. Thresholds are calculated using lower bound theoretical analysis upon two extreme workloads (sparse and dense). Experimental evaluation using real world workloads illustrates the merits of our approach against state-of-the-art algorithms.
Panagiotis Oikonomou, Nikos Tziritas, Thanasis Loukopoulos, Georgios Theodoropoulos 0001, Masatoshi Hanai, Samee Ullah Khan
IEEE Trans. Sustain. Comput.6
2021 SLO-aware Virtual Rebalancing for Edge Stream Processing
abstract
The Internet of Things (IoT) has enabled an abundance of geographically distributed physical devices or “things” equipped with sensors and actuators to exchange information with the Cloud. However, this paradigm remains largely under-exploited for real-time analytic applications. The benefit of realtime data acquisition at the Edge becomes fruitless as it is not readily accessible to more powerful data analytic tools in the Cloud due to wide-area network delays. In this paper, we present VRebalance, a virtual resource orchestrator that provides an end-to-end performance guarantee for concurrent stream processing workloads at the Edge. VRebalance employs Bayesian Optimization$\mathcal{BO}$to quickly identify near-optimal resource configurations. Experimental results with a real-time open-source IoT benchmark for Distributed Stream Processing Platforms (RIoTBench) and a representative stream processing engine (Apache Storm) demonstrate the superior performance, resource efficiency and adaptiveness of our$\mathcal{BO}$-based resource management system. VRebalance meets the performance SLO (service level objective) targets for stream processing workloads even in the presence of acute system dynamics. It decreases the SLO violation rate by at least 34% for static workloads and by 62.5% for dynamic workloads compared to a hill climbing method. Compared to Storm's default resource scaling mechanism, our method decreases the SLO violation rate by 83.7%.
Palden Lama, Samee Ullah Khan
IC2E3
2021 A Game-based Thermal-Aware Resource Allocation Strategy for Data Centers
abstract
Data centers (DC) host a large number of servers, computing devices and computing infrastructure, which incur significant electricity / energy. This also results in huge amount of heat produced, which if not addressed can lead to overheating of computing devices in the DC. In addition, temperature mismanagement can lead to thermal imbalance within the DC environment, which may result in the creation of hotspots. The energy consumed during the life of a hotspot is greater than the energy saved during computation. Hence, the thermal imbalance impacts on the efficiency of the cooling mechanism installed inside the DC, which can result in high energy consumption. One popular strategy to minimize energy consumption is to optimize resource allocation within the DC. However, existing scheduling strategies do not consider the ambient effect of the surrounding nodes at the time of job allocation. Moreover, thermal-aware resource scheduling as an optimization problem is a topic that is relatively understudied in the literature. Therefore, in this research, we propose a novel Game-based Thermal-Aware Resource Allocation (GTARA) strategy to reduce the thermal imbalances within the DC. Specifically, we use cooperative game theory with a Nash-bargaining solution concept to model the resource allocation as an optimization problem, where the user jobs are assigned to the computing nodes based on their thermal profiles and their potential effect on the surrounding nodes. This allows us to improve the thermal balance and avoid the hotspots. We then demonstrate the effectiveness of GTARA, TACS, TASA, and FCFS, in terms of minimizing thermal imbalance and the hotspots.
Saeed Akbar, Saif Ur Rehman Malik, Kim-Kwang Raymond Choo, Samee Ullah Khan, Naveed Ahmad 0001, Adeel Anjum
IEEE Trans. Cloud Comput.4
2021 SeSPHR: A Methodology for Secure Sharing of Personal Health Records in the Cloud
abstract
The widespread acceptance of cloud based services in the healthcare sector has resulted in cost effective and convenient exchange of Personal Health Records (PHRs) among several participating entities of the e-Health systems. Nevertheless, storing the confidential health information to cloud servers is susceptible to revelation or theft and calls for the development of methodologies that ensure the privacy of the PHRs. Therefore, we propose a methodology called SeSPHR for secure sharing of the PHRs in the cloud. The SeSPHR scheme ensures patient-centric control on the PHRs and preserves the confidentiality of the PHRs. The patients store the encrypted PHRs on the un-trusted cloud servers and selectively grant access to different types of users on different portions of the PHRs. A semi-trusted proxy called Setup and Re-encryption Server (SRS) is introduced to set up the public/private key pairs and to produce the re-encryption keys. Moreover, the methodology is secure against insider threats and also enforces a forward and backward access control. Furthermore, we formally analyze and verify the working of SeSPHR methodology through the High Level Petri Nets (HLPN). Performance evaluation regarding time consumption indicates that the SeSPHR methodology has potential to be employed for securely shar-ing the PHRs in the cloud.
Assad Abbas, Muhammad Usman Shahid Khan, Samee Ullah Khan
IEEE Trans. Cloud Comput.4
2021 A Robust Optimization Technique for Energy Cost Minimization of Cloud Data Centers
abstract
Power consumption cost constitutes significant portion of a data center's total operational cost. Variations in parameters, such as power demand, electricity price, and renewable power generation effects the data center's power consumption cost. Therefore, in this paper, a smart power management system based on a robust energy cost optimization algorithm is proposed for the data center. The designed robust optimization algorithm coordinates data center workload, battery bank, diesel generators, renewable power, and trade electricity price in real-time and day-ahead power market to reduce expected energy consumption cost. The uncertain parameters, such as data center workload and renewable power are computed using forecasting algorithms. The optimization algorithm is used to limit unbalance power purchase from real-time power market due to uncertainty of real-time electricity price. The algorithm formulation is obtained using mixed-integer linear programming. Moreover, a model to calculate prices to be used in service level agreements with data centers' clients for on-demand cloud services is designed considering operational cost of the data center. Simulations are performed on actual data center workload, weather parameters, and electricity price. The results have shown that the proposed methodology is an effective tool to minimize data center operational cost.
Muhammad Jawad 0001, Muhammad Bilal Qureshi, Muhammad Usman Shahid Khan, Sahibzada Muhammad Ali, Chaudhary Arshad Mehmood, Samee Ullah Khan
IEEE Trans. Cloud Comput.8
2021 xFogSim: A Distributed Fog Resource Management Framework for Sustainable IoT Services
abstract
Streaming large amounts of data to cloud data centers cause network congestion resulting in high network and energy consumption. The concept of fog computing is introduced to reduce workload from backbone networks and support delay-sensitive Internet of Things (IoT) applications. The concept places compute, storage, and network services closer to the source of the requests. In general fog-based simulators are used for better understanding and optimum fog resource allocation. Unfortunately, most of the simulators lack core features like network delay, latency, packet error rate, energy consumption, and distributed fog node management. In this paper, we propose a fog simulation framework termed as xFogSim to support latency-sensitive applications at the fog layer with multi-objective optimization to trade-off cost, availability, and performance among the fog federation. Moreover, during peak load, the framework provides locality-aware distributed broker node management that enables borrowing resources from nearby fog locations to meet service and energy requirements. The results show that the framework is lightweight, configurable, and scalable, capable of handling a large number of user requests using dynamic resource provisioning across the fog federation.
Asad Waqar Malik, Tariq Qayyum, Anis Ur Rahman 0001, Muazzam Ali Khan, Osman Khalid, Samee Ullah Khan
IEEE Trans. Sustain. Comput.6
2020 Graph-based Approaches for the Interval Scheduling Problem
abstract
One of the fundamental problems encountered by large-scale computing systems, such as clusters and cloud, is to schedule a set of jobs submitted by the users. Each job is characterized by resource demands, as well as start and completion time. Each job must be scheduled to execute on a machine having the required capacity between the start and completion time (referred as interval) of the job. Each machine is defined by a parallelism parameter g that indicates the maximum number of jobs that can be processed by the machine, in parallel. The above problem is referred to as the interval scheduling problem with bounded parallelism. The objective is to minimize the total busy time of all machines. Majority of the solutions proposed in the literature consider homogeneous set of jobs and machines that is a simplified assumption as in practice, heterogeneous jobs and machines are frequently encountered. In this article, we tackle the aforesaid problem with a set of heterogeneous jobs and machines. A major contribution of our work is that the problem is addressed in a novel way by combining a graph-based approach and a dynamic programming approach which is based on a variation of bin packing problem. A greedy algorithm is also proposed by employing only a graph-based approach at the aim to reduce the computational complexity. Experimental results show that the proposed algorithms can significantly reduce the cumulative busy interval over all machines compared with state-of-the-art algorithms proposed in the literature.
Panagiotis Oikonomou, Nikos Tziritas, Georgios Theodoropoulos 0001, Maria G. Koziri, Thanasis Loukopoulos, Samee Ullah Khan
ICPADS6
2020 An Ancillary Services Model for Data Centers and Power Systems
abstract
Enormous energy consumption of data centers has a major impact on power systems by significantly increasing the electrical load. Due to the increase in electrical load, power systems are facing demand and supply miss-management problems. Therefore, power systems require efficient and intelligent ancillary services to maintain robustness, reliability, and stability. Data centers can provide the computational capabilities to manage power systems; however, data centers consume a tremendous amount of energy, and energy price accounts for a significant portion of their operational cost. Power system jobs will make this situation even more critical for data centers. In our work, we seek an Ancillary Services Model (ASM) to service data centers and power systems. In ASM, we find an optimal job scheduling technique for executing power systems' jobs on data centers in terms of low power consumption, reduced makespan, and fewer preempted jobs. The power systems' jobs include Optimal Power Flow (OPF) calculation, transmission line importance index, and bus importance index. Moreover, a Service Level Agreement (SLA) between data centers and power systems is shown to provide mutual benefits.
Sahibzada Muhammad Ali, Muhammad Jawad 0001, Muhammad Usman Shahid Khan, Kashif Bilal, Jacob Glower, Scott C. Smith, Samee Ullah Khan, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Cloud Comput.7
2020 Boafft: Distributed Deduplication for Big Data Storage in the Cloud
abstract
As data progressively grows within data centers, the cloud storage systems continuously facechallenges in saving storage capacity and providing capabilities necessary to move big data within an acceptable time frame. In this paper, we present the Boafft, a cloud storage system with distributed deduplication. The Boafft achieves scalable throughput and capacity usingmultiple data servers to deduplicate data in parallel, with a minimal loss of deduplication ratio. Firstly, the Boafft uses an efficient data routing algorithm based on data similarity that reduces the network overhead by quickly identifying the storage location. Secondly, the Boafft maintains an in-memory similarity indexing in each data server that helps avoid a large number of random disk reads and writes, which in turn accelerates local data deduplication. Thirdly, the Boafft constructs hot fingerprint cache in each data server based on access frequency, so as to improve the data deduplication ratio. Our comparative analysis with EMC's stateful routing algorithm reveals that the Boafft can provide a comparatively high deduplication ratio with a low network bandwidth overhead. Moreover, the Boafft makes better usage of the storage space, with higher read/write bandwidth and good load balance.
Shengmei Luo, Guangyan Zhang, Chengwen Wu, Samee Ullah Khan, Keqin Li 0001
IEEE Trans. Cloud Comput.4
2020 Online Inter-Datacenter Service Migrations
abstract
Service migration between datacenters can reduce the network overhead within a cloud infrastructure; thereby, also improving the quality of service for the clients. Most of the algorithms in the literature assume that the client access pattern remains stable for a sufficiently long period so as to amortize such migrations. However, if such an assumption does not hold, these algorithms can take arbitrarily poor migration decisions that can substantially degrade system performance. In this paper, we approach the issue of performing service migrations for an unknown and dynamically changing client access pattern. We propose an online algorithm that minimizes the inter-datacenter network, taking into account the network load of migrating a service between two datacenters, as well as the fact that the client request pattern may change “quickly”, before such a migration is amortized. We provide a rigorous mathematical proof showing that the algorithm is 3.8-competitive for a cloud network structured as a tree of multiple datacenters. We briefly discuss how the algorithm can be modified to work on general graph networks with an O(log|V|) probabilistic approximation of the optimal algorithm. Finally, we present an experimental evaluation of the algorithm based on extensive simulations.
Nikos Tziritas, Samee Ullah Khan, Thanasis Loukopoulos, Spyros Lalis, Cheng-Zhong Xu 0001, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Cloud Comput.2
2019 Anonymous Privacy-Preserving Scheme for Big Data Over the Cloud
abstract
This paper introduces an anonymous privacy-preserving scheme for big data over the cloud. The proposed design helps to enhance the encryption/decryption time of big data by utilizing the MapReduce framework. The Hadoop distributed file system and the secure hash algorithm are employed to provide the anonymity, security and efficiency requirements for the proposed scheme. The experimental results show a significant enhancement in the computational time of data encryption and decryption.
Zeyad A. Al-Odat, Samee Ullah Khan
IEEE BigData2
2019 The Sponge Structure Modulation Application to Overcome the Security Breaches for the MD5 and SHA-1 Hash Functions
abstract
This paper presents a Sponge structure modulation of the MD5 and SHA-1 hash functions. The work employs the Keccak permutation function to build the proposed scheme. The work discusses the main two security breaches that threaten the cryptography hash standards which are collision and length extension attacks. Through analyzing several examples of collided messages of both algorithms (SHA-1 and MD5), we describe the potentials to overcome the collision and length extension attacks. Moreover, a proper replacement technique to avoid such attacks is discussed in this paper.
Zeyad A. Al-Odat, Samee Ullah Khan
COMPSAC (1)2
2019 A Framework for Dengue Surveillance and Data Collection in Pakistan
abstract
Health monitoring through smartphones applications has emerged as a popular and effective practice. Many countries including Pakistan are suffering from a viral disease called Dengue. Dengue is a mosquito-borne single positive-standard RNA virus of the family Flaviviridae. It can be identified from its symptoms, such as skin rashes, fever, headache, and nausea etc. Due to the limited connectivity and availability of information technology services in rural and far-off areas in Pakistan, timely reporting the Dengue incidents to the authorities has been a serious issue. Moreover, currently there does not exist any data about Dengue infected areas and affected patients that hinders the government authorities to timely predict the disease outbreaks. To that end, we propose a framework to collect the information of suspected patients of Dengue through smart phones. The collected information is subsequently transmitted to the doctors and authorities for further necessary measures. A key benefit of the proposed architecture is that it will result in establishing a data repository of Dengue patients that can be further used for Dengue outbreak prediction.
Komal Minhas, Munazza Tabassam, Rida Rasheed, Assad Abbas, Hasan Ali Khattak, Samee Ullah Khan
COMPSAC (2)6
2019 Online Live VM Migration Algorithms to Minimize Total Migration Time and Downtime
abstract
Virtual machine (VM) migration is a widely used technique in cloud computing systems to increase reliability. There are also many other reasons that a VM is migrated during its lifetime, such as reducing energy consumption, improving performance, maintenance, etc. During a live VM migration, the underlying VM continues being up until all or part of its data has been transmitted from source to destination. The remaining data are transmitted in an off-line manner by suspending the corresponding VM. The longer the off-line transmission time, the worse the performance of the respective VM. The above is because during the off-line data transmission, the VM service is down. Because a running VM's memory is subject to changes, already transmitted data pages may get dirtied and thus needing re-transmission. The decision of when suspending the VM is not a trivial task at all. The above is justified by the fact that when suspending the VM early we may result in transmitting off-line a significant amount of data degrading thus the VM's performance. On the other hand, a long waiting time to suspend the VM may result in re-transmitting a huge amount of dirty data, leading in that way to waste of resources. In this paper, we tackle the joint problem of minimizing both the total VM migration time (reflecting the resources spent during a migration) and the VM downtime (reflecting the performance degradation). The aforementioned objective functions are weighted according to the needs of the underlying cloud provider/user. To tackle the problem, we propose an online deterministic algorithm resulting in an strong competitive ratio, as well as a randomized online algorithm achieving significantly better results against the deterministic algorithm.
Nikos Tziritas, Thanasis Loukopoulos, Samee Ullah Khan, Cheng-Zhong Xu 0001, Albert Y. Zomaya
IPDPS3
2019 Facial Expression Recognition Based on Edge Computing
abstract
Facial action unit (AU) detection recognizes facial expressions by analyzing cues about the movement of certain atomic muscles in the local facial area. According to the detection of facial feature points, we can calculate the value of AU, and then classification algorithms are performed on these AU values for emotion detection in realtime. When this advanced system is in the actual production process, large network bandwidth is usually required to transmit a large number of frames from cameras to the backend servers. To overcome this issue, we propose to offload the computing to edge devices in which the raw image data from each camera is directly processed with our optimized and customized algorithms, and then the detected emotions are transmitted to the end-user more easily compared with less data.
Tiantian Qian, Fan Zhang 0003, Samee Ullah Khan
MSN3
2019 A V2I communication-based pipeline model for adaptive urban traffic light scheduling
Lei Nie 0004, Samee Ullah Khan, Osman Khalid, Dan Wu 0006
Frontiers Comput. Sci.3
2019 Zeus: A resource allocation algorithm for the cloud of sensors
Igor Leão dos Santos, Luci Pirmez, Flávia Coimbra Delicato, Gabriel Martins de Oliveira Costa, Claudio M. de Farias, Samee Ullah Khan, Albert Y. Zomaya
Future Gener. Comput. Syst.6
2019 Quantifying cloud elasticity with container-based autoscaling
Fan Zhang 0003, Xuxin Tang, Xiu Li 0001, Samee Ullah Khan, Zhijiang Li
Future Gener. Comput. Syst.4
2019 Privacy-preserving model and generalization correlation attacks for 1: M data with multiple sensitive attributes
Tehsin Kanwal, Sayed Ali Asjad Shaukat, Adeel Anjum, Saif Ur Rehman Malik, Kim-Kwang Raymond Choo, Abid Khan, Naveed Ahmad 0001, Mansoor Ahmad, Samee Ullah Khan
Inf. Sci.9
2019 QuantCloud: Enabling Big Data Complex Event Processing for Quantitative Finance Through a Data-Driven Execution
abstract
Quantitative Finance (QF) utilizes increasingly sophisticated mathematic models and advanced computer techniques to predict the movement of global markets, and price the derivatives and other assets. Being able to react quickly and intelligently to fast-changing markets is a decisive success factor for trading companies. To date, the rise of QF requires an integrated toolchain of enabling technologies to carry out complex event processing on the explosive growth and diversified forms of market metadata, in pursuit of a microsecond latency on an Exabyte-level dataset. Inspired by this, we present a data-driven execution paradigm that untangles the dependencies of complex processing events and integrate the paradigm with a big data infrastructure that streams time series data. This integrated platform is termed as the QuantCloud platform. Essentially, QuantCloud executes the complex event processing in a data-driven mode and manages large amounts of diversified market data in a data-parallel mode. To show its practicability and performance, we develop a prototype and benchmark by applying real-world QF research models on the New York Stock Exchange (NYSE) data. Using this prototype, we demonstrate this platform with an application to: (i) data cleaning and aggregating (including the computing of logarithmic returns from tick data and the finding the medians of grouped data) and (ii) data modeling: the autoregressive-moving average (ARMA) model. The performance results show that (a) this platform obtains a high throughput (usually in the order of millions of tick messages per second) and a sub-microsecond latency; (b) it fully executes data-dependent tasks through a data-driven execution; and (c) it implements a modular design approach for rapidly developing these data-crunching methods and QF research models. This platform resulting from an aggregated effort of the data-driven execution and big data infrastructure, offers the financial engineers with new insights and enhanced capabilities for effective and efficient incorporation of big data complex event processing technologies in their workflow.
Peng Zhang 0006, Samee Ullah Khan
IEEE Trans. Big Data3
2019 Optimal Data Center Scheduling for Quality of Service Management in Sensor-Cloud
abstract
The proposed work concentrates on the networking facets of sensor-cloud infrastructures-one of the first attempts of its kind. In a sensor-cloud, multiple sets of physical sensor nodes that are activated based on an application demand, in turn give rise to multiple distinct virtual sensors (VSs). The VSs are considered to span across multiple geographical regions; thereby, depositing the data (from each of the VS) to the closest cloud data center (DC). Quite obviously, multiple geospatial DCs get involve with an application. However, the principle of sensor-cloud is to store and conglomerate the data from various VSs, before they can be provisioned as Sensors-as-a-Service (Se-aaS). The assortation of data occurs within a single Virtual Machine (VM) (or in some cases multiple VMs) residing inside a particular DC. This work addresses the problem of scheduling a particular DC that congregates data from various VSs, and transmit the same to the end-user application. The work follows the general pairwise choice framework of the Optimal Decision Rule. The scheduling of the DC is performed under several network constraints, such as data migration cost, data delivery cost, and service delay of an application that ensures the preservation of the Quality-of-Service (QoS) and maintenance of the user satisfaction. The work quantifies the effective QoS of Se-aaS and determines an optimal decision rule for electing a particular DC. While arriving at a collective decision, the work incorporates the fallible decision making ability of a DC; thereby, excluding the loss of generality. Experimental results depict that the proposed algorithm for generating the optimal decision rule finds applicability in real-time cloud computing scenarios.
Subarna Chatterjee, Sudip Misra, Samee Ullah Khan
IEEE Trans. Cloud Comput.3
2019 Empirical Discovery of Power-Law Distribution in MapReduce Scalability
abstract
Understanding the scalability of MapReduce applications is a challenging problem. The difficulty lies in the distributed mapping of the input big data. The distribution of data and compute resources must match with fluctuating network substrates. User-defined Map and Reduce functions over application parameters further complicate the issue. Therefore, it offers great payoff to use small datasets and limited test runs to reveal the behavior of MapReduce applications over big-data. In this paper, we analyze the scaling effects of server cluster-size over varieties of Map- and Reduce-intensive applications. In our study, we discover specific conditions which lead to the power-law conformity in representative MapReduce applications. We report four major discoveries: (1) Within a range of scaling parameters, MapReduce execution time follows the power-law distribution. (2) Power-law scalability for Map-intensive applications work well even with a small cluster size. (3) Shuffle-intensive applications exhibit power-law behavior starting from larger cluster size. (4) The scaling effects may depart from power-law distribution, if the cloud resources are heavily overprovisioned than the workload demands. The above findings enable users to use bounded test runs to allocate and configure virtual and physical resources in large-scale MapReduce applications. These results can be also applied in generating business models for providing cost-effective cloud computing services.
Fan Zhang 0003, Majd F. Sakr, Kai Hwang 0001, Samee Ullah Khan
IEEE Trans. Cloud Comput.4
2019 A Double Deep Q-Learning Model for Energy-Efficient Edge Scheduling
abstract
Reducing energy consumption is a vital and challenging problem for the edge computing devices since they are always energy-limited. To tackle this problem, a deep Q-learning model with multiple DVFS (dynamic voltage and frequency scaling) algorithms was proposed for energy-efficient scheduling (DQL-EES). However, DQL-EES is highly unstable when using a single stacked auto-encoder to approximate the Q-function. Additionally, it cannot distinguish the continuous system states well since it depends on a Q-table to generate the target values for training parameters. In this paper, a double deep Q-learning model is proposed for energy-efficient edge scheduling (DDQ-EES). Specially, the proposed double deep Q-learning model includes a generated network for producing the Q-value for each DVFS algorithm and a target network for producing the target Q-values to train the parameters. Furthermore, the rectified linear units (ReLU) function is used as the activation function in the double deep Q-learning model, instead of the Sigmoid function in QDL-EES, to avoid gradient vanishing. Finally, a learning algorithm based on experience replay is developed to train the parameters of the proposed model. The proposed model is compared with DQL-EES on EdgeCloudSim in terms of energy saving and training time. Results indicate that our proposed model can save average 2%-2.4% energy and achieve a higher training efficiency than QQL-EES, proving its potential for energy-efficient edge scheduling.
Qingchen Zhang 0001, Man Lin, Laurence T. Yang, Zhikui Chen, Samee Ullah Khan, Peng Li 0027
IEEE Trans. Serv. Comput.5
2018 A Pareto-Efficient Algorithm for Data Stream Processing at Network Edges
abstract
Data stream processing has received considerable attention from both research community and industry over the last years. Since latency is a key issue in data stream processing environments, the majority of the works existing in the literature focus on minimizing the latency experienced by the users. The aforementioned minimization takes place by assigning the data stream processing components close to data sources. Server consolidation is also a key issue for drastically reducing energy consumption in computing systems. Unfortunately, energy consumption and latency are two objective functions that may be in conflict with each other. Therefore, when the target function is to minimize energy consumption, the delay experienced by users may be considerable high, and the opposite. For the above reason there is a dire need to design strategies such that by targeting the minimization of energy consumption, there is a graceful degradation in latency, as well as the opposite. To achieve the above, we propose a Pareto-efficient algorithm that tackles the problem of data processing tasks placement simultaneously in both dimensions regarding the energy consumption and latency. The proposed algorithm outputs a set of solutions that are not dominated by any solution within the set regarding energy consumption and latency. The experimental results show that the proposed approach is superior against single-solution approaches because by targeting one objective function the other one can be gracefully degraded by choosing the appropriate solution.
Thanasis Loukopoulos, Nikos Tziritas, Maria G. Koziri, Georgios I. Stamoulis, Samee Ullah Khan
CloudCom5
2018 Request Dispatching for Minimizing Service Response Time in Edge Cloud Systems
abstract
The emerging of mobile edge computing has significantly reduced the response time and Internet risk of service invocations. However, due to the distributed architecture and limited resources, balancing the load between edge servers to minimize the overall response time has become a critical objective for mobile edge computing. This problem is generally related to two aspects, request dispatching and service scheduling. To address this issue, we proposed a novel heuristic method called GASD (combined Genetic algorithm and simulated Annealing algorithm for Service request Dispatching). It tackles the problem by jointly conducting request dispatching and service scheduling. In addition, a solution combination algorithm is applied to reduce the computation complexity of the method. The experimental results show that the GASD method can achieve much lower overall response time than the compared methods. Moreover, the execution time of GASD is in a low order of magnitude and the algorithm performs excellent scalability as the experimental scale increases.
Hongyue Wu, Shuiguang Deng, Wei Li 0058, Samee Ullah Khan, Jianwei Yin, Albert Y. Zomaya
ICCCN4
2018 From Insight to Impact: Building a Sustainable Edge Computing Platform for Smart Homes
abstract
There is a growing trend for engaging edge computing to help smart homes to improve the living comfort of residents. However, the rapidly rising interest in such deployments does not normally integrate with energy management schemes, which is a central issue that smart homes have to face. In this paper, we propose a unified energy management framework for enabling a sustainable edge computing paradigm while meeting the needs of home energy management and smart home applications. This framework aims to enable the full use of renewable energy while reducing electricity bills for households. A prototype system was implemented by using low-cost and easy-to-get hardware. The experiment results demonstrated that renewable energy is fully capable of supporting the reliable running of edge computing devices and electricity bills could be cut by up to 86% when our proposed framework was employed.
Xiaomin Chang, Wei Li 0058, Chunqiu Xia, Jin Ma 0001, Samee Ullah Khan, Albert Y. Zomaya
ICPADS6
2018 Server Consolidation in Cloud Computing
abstract
Minimizing service-level agreement (SLA) violations and energy consumption through server consolidation is of paramount importance for the sustainability of cloud environments. In this paper, we propose an online method to reduce cloud SLA violations by taking into account: (a) the energy consumption of migrating virtual machines and (b) server consolidation techniques to minimize the energy consumption within the system. Rigorous mathematical competitive analysis shows that the proposed method achieves a 2.4 competitive ratio against a cognitive adversary. Our solution is superior compared to the current state of the art algorithms, such as UP-VMC, KMI, MBFD, both theoretically (competitive ratios of other alternatives are unbounded) and empirically through simulations using CloudSim. More specifically, the competitive ratios of state of the art algorithms are unbounded and our proposed methodology reduces the VM migrations and energy consumption by 85% and 45%, respectively. The aforementioned improvement comes at an expense of a small increase in terms of SLA violations. The above results are achieved without the a priori knowledge of VM utilization patterns.
Nikos Tziritas, Saad Mustafa, Maria G. Koziri, Thanasis Loukopoulos, Samee Ullah Khan, Cheng-Zhong Xu 0001, Albert Y. Zomaya
ICPADS5
2018 Potentials, trends, and prospects in edge technologies: Fog, cloudlet, mobile edge, and micro data centers
Kashif Bilal, Osman Khalid, Aiman Erbad, Samee Ullah Khan
Comput. Networks4
2018 An efficient privacy mechanism for electronic health records
Adeel Anjum, Saif Ur Rehman Malik, Kim-Kwang Raymond Choo, Abid Khan, Asma Haroon, Sangeen Khan, Samee Ullah Khan, Naveed Ahmad 0001, Basit Raza
Comput. Secur.7
2018 Very low computational complexity (VLCC) architecture for optical interconnect in data center networks
abstract
Summary Real time cloud computing applications require a low latency network. The latency of optical interconnect in Data Center Networks (DCNs) is dependent on the complexity of the routing algorithm. The routing algorithm makes decisions about the forwarding of a packet on each successive node. If a routing algorithm is more complex, it requires more hardware resources to implement, which incurs extra cost and latency to the optical interconnect. This paper analyzes the complexity of existing architectures for the first time by showing Big O notation of complexity for each architecture. Different factors affecting the computational complexity of any routing algorithm are identified. This paper proposes a new architecture named VLCC. It has very low fixed routing complexity irrespective of network size. It has no packet loss at the network layer under many to many and all to one communication pattern. The only packet loss is at the physical layer due to signal degradation caused by optical components. The level of signal degradation is analyzed in terms of received signal power, bit error rate (BER), receiver sensitivity, path loss, and blocking probability. VLCC architecture is highly suitable for real time applications that require deterministic quality of service (QoS), where network performance is not affected by the traffic pattern.
Mohsin Fayyaz, Khurram Aziz, Ghulam Mujtaba 0002, Ahmad Fayyaz, Samee Ullah Khan
Concurr. Comput. Pract. Exp.5
2018 Energy and communication aware task mapping for MPSoCs
Tahir Maqsood, Nikos Tziritas, Thanasis Loukopoulos, Sajjad Ahmad Madani, Samee Ullah Khan, Cheng-Zhong Xu 0001, Albert Y. Zomaya
J. Parallel Distributed Comput.5
2018 Thermal-Aware and DVFS-Enabled Big Data Task Scheduling for Data Centers
abstract
Big data has received considerable attentions in recent years because of massive data volumes in multifarious fields. Considering various “V” features, big data tasks are usually highly complex and computational intensive. These tasks are generally performed in parallel in data centers resulting in massive energy consumption and Green House Gases emissions. Therefore, efficient resource allocation considering the synergy of the performance and energy efficiency is one of the crucial challenges today. In this paper, we aim to achieve maximum energy efficiency by combining thermal-aware and dynamic voltage and frequency scaling (DVFS) techniques. This paper proposes: (a) a thermal-aware and power-aware hybrid energy consumption model synchronously considering the computing, cooling, and migration energy consumption; (b) a tensor-based task allocation and frequency assignment model for representing the relationship among different tasks, nodes, time slots, and frequencies; and (c) a big data Task Scheduling algorithm based on Thermal-aware and DVFS-enabled techniques (TSTD) to minimize the total energy consumption of data centers. The experimental results demonstrate that the proposed TSTD algorithm significantly outperforms the state-of-the-art energy efficient algorithms from total, computing, and cooling energy consumption perspectives, as well as cooling energy consumption proportion and total energy consumption savings.
Huazhong Liu, Baoshun Liu, Laurence T. Yang, Man Lin, Yuhui Deng 0001, Kashif Bilal, Samee Ullah Khan
IEEE Trans. Big Data7
2018 QuantCloud: Big Data Infrastructure for Quantitative Finance on the Cloud
abstract
In this paper, we present the QuantCloud infrastructure, designed for performing big data analytics in modern quantitative finance. Through analyzing market observations, quantitative finance (QF) utilizes mathematical models to search for subtle patterns and inefficiencies in financial markets to improve prospective profits. To discover profitable signals in anticipation of volatile trading patterns amid a global market, analytics are carried out on Exabyte-scale market metadata with a complex process in pursuit of a microsecond or even a nanosecond of data processing advantage. This objective motivates the development of innovative tools to address challenges for handling high volume, velocity, and variety investment instruments. Inspired by this need, we developed QuantCloud by employing large-scale SSD-backed datastore, various parallel processing algorithms, and portability in Cloud computing. QuantCloud bridges the gap between model computing techniques and financial data-driven research. The large volume of market data is structured in an SSD-backed datastore, and a daemon reacts to provide the Data-on-Demand services. Multiple client services process user requests in a parallel mode and query on-demand datasets from the datastore through Internet connections. We benchmark QuantCloud performance on a 40-core, 1TB-memory computer and a 5-TB SSD-backed datastore. We use NYSE TAQ data from the fourth quarter of 2014 as our market data. The results indicate data-access application latency as low as 3.6 nanoseconds per message, sustained throughput for parallel data processing as high as 74 million messages per second, and completion of 11 petabyte-level data analytics within 53 minutes. Our results demonstrate that the aggregated contributions of our infrastructure, parallel algorithms, and sophisticated implementations offer the algorithmic trading and financial engineering community new hope and numeric insights for their research and development.
Peng Zhang 0006, Jessica J. Yu, Samee Ullah Khan
IEEE Trans. Big Data4
2018 DROPS: Division and Replication of Data in Cloud for Optimal Performance and Security
abstract
Outsourcing data to a third-party administrative control, as is done in cloud computing, gives rise to security concerns. The data compromise may occur due to attacks by other users and nodes within the cloud. Therefore, high security measures are required to protect data within the cloud. However, the employed security strategy must also take into account the optimization of the data retrieval time. In this paper, we propose division and replication of data in the cloud for optimal performance and security (DROPS) that collectively approaches the security and performance issues. In the DROPS methodology, we divide a file into fragments, and replicate the fragmented data over the cloud nodes. Each of the nodes stores only a single fragment of a particular data file that ensures that even in case of a successful attack, no meaningful information is revealed to the attacker. Moreover, the nodes storing the fragments, are separated with certain distance by means of graph T-coloring to prohibit an attacker of guessing the locations of the fragments. Furthermore, the DROPS methodology does not rely on the traditional cryptographic techniques for the data security; thereby relieving the system of computationally expensive methodologies. We show that the probability to locate and compromise all of the nodes storing the fragments of a single file is extremely low. We also compare the performance of the DROPS methodology with 10 other schemes. The higher level of security with slight performance overhead was observed.
Kashif Bilal, Samee Ullah Khan, Bharadwaj Veeravalli, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Cloud Comput.3
2018 Segregating Spammers and Unsolicited Bloggers from Genuine Experts on Twitter
abstract
Online Social Networks (OSNs) have not only significantly reformed the social interaction pattern but have also emerged as an effective platform for recommendation of services and products. The upswing in use of the OSNs has also witnessed growth in unwanted activities on social media. On the one hand, the spammers on social media can be a high risk towards the security of legitimate users and on the other hand some of the legitimate users, such as bloggers can pollute the results of recommendation systems that work alongside the OSNs. The polluted results of recommendation systems can be precarious to the masses that track recommendations. Therefore, it is necessary to segregate such type of users from the genuine experts. We propose a framework that separates the spammers and unsolicited bloggers from the genuine experts of a specific domain. The proposed approach employs modified Hyperlink Induced Topic Search (HITS) to separate the unsolicited bloggers from the experts on Twitter on the basis of tweets. The approach considers domain specific keywords in the tweets and several tweet characteristics to identify the unsolicited bloggers. Experimental results demonstrate the effectiveness of the proposed methodology as compared to several state-of-the-art approaches and classifiers.
Muhammad Usman Shahid Khan, Assad Abbas, Samee Ullah Khan, Albert Y. Zomaya
IEEE Trans. Dependable Secur. Comput.4
2017 Cloud-of-clouds based resource provisioning strategy for continuous write applications
abstract
Nowadays, more and more online services based on cloud computing have taken the places of some traditional applications (e.g., Health Care) that continuously generate large volume of data and require data storage and analysis in time. Such applications can be categorized as “Continuous Writing Applications” (CWA) that have particular requirements on bandwidth, storage, computation, and service reliability. In the meanwhile, they are very sensitive to the cost. In this paper, we present an architecture of multiple cloud service providers (CSPs) or “Cloud-of-Clouds” to provide services to the CWA and propose a novel resource scheduling algorithm to minimize the cost of entire systems. Difference from many research efforts that focus on a single resource, we take many factors into considerations that include user's requirements of bandwidth, storage and computation, the resources of CSPs that can provide, CSPs for data backup, the configurations of Cloud-of-Clouds, system models of CSPs, and many more. We first present the system models of classic CWA applications to capture the resource requirements of users on Cloud-of-Clouds. We then present the problem formulation and our optimal strategy of user scheduling based on Minimum First Derivative Length (MFDL) of load paths among the systems. Through theoretical analysis, we prove that our proposed algorithm Optimal user Scheduling for Cloud-of-Clouds (OSCC) can achieve the optimal solution.
Zeng Zeng, Bharadwaj Veeravalli, Samee Ullah Khan, Sin G. Teo
APCC3
2017 Self-adaptation and mutual adaptation for distributed scheduling in benevolent clouds
abstract
SUMMARY Joint service involving several clouds is an emerging form of cloud computing. In hybrid clouds, the schedulers within 1 cloud must not only self‐adapt to the job arrival processes and the workload but also mutually adapt to the scheduling polices of other schedulers. However, as a combinatorial optimization problem, scheduling is challenged by the adaptation to those dynamics and uncertain behaviors of the peers. This article studies the collaboration among benevolent clouds that are cooperative in nature and willing to accept jobs from other clouds. We take advantage of machine learning and propose a distributed scheduling mechanism to learn the knowledge of job model, resource performance, and others' policies. Without explicit modeling and prediction, machine learning guides scheduling decisions based on experiences. To examine the performance of our approach, we conducted simulation using the SP2 job workload log of the San Diego Supercomputer Center under a test bed based on agent‐based systems—SWARM. The results validate that our approach has much shorter mean response time than 5 typical dynamic scheduling algorithms—opportunistic load balancing, minimum execution time, minimum completion time, switching algorithm, and k‐percent best. A better collaboration in hybrid cloud is achieved by full adaptation.
PiJun Liang, Zhao Tong 0001, Kenli Li 0001, Samee Ullah Khan, Keqin Li 0001
Concurr. Comput. Pract. Exp.5
2017 System modelling and performance evaluation of a three-tier Cloud of Things
Wei Li 0058, Igor Leão dos Santos, Flávia Coimbra Delicato, Paulo F. Pires, Luci Pirmez, Wei Wei 0006, Houbing Song, Albert Y. Zomaya, Samee Ullah Khan
Future Gener. Comput. Syst.9
2017 DaSCE: Data Security for Cloud Environment with Semi-Trusted Third Party
abstract
Off-site data storage is an application of cloud that relieves the customers from focusing on data storage system. However, outsourcing data to a third-party administrative control entails serious security concerns. Data leakage may occur due to attacks by other users and machines in the cloud. Wholesale of data by cloud service provider is yet another problem that is faced in the cloud environment. Consequently, high-level of security measures is required. In this paper, we propose data security for cloud environment with semi-trusted third party (DaSCE), a data security system that provides (a) key management (b) access control, and (c) file assured deletion. The DaSCE utilizes Shamir's (k, n) threshold scheme to manage the keys, where k out of n shares are required to generate the key. We use multiple key managers, each hosting one share of key. Multiple key managers avoid single point of failure for the cryptographic keys. We (a) implement a working prototype of DaSCE and evaluate its performance based on the time consumed during various operations, (b) formally model and analyze the working of DaSCE using high level petri nets (HLPN), and (c) verify the working of DaSCE using satisfiability modulo theories library (SMT-Lib) and Z3 solver. The results reveal that DaSCE can be effectively used for security of outsourced data by employing key management, access control, and file assured deletion.
Saif Ur Rehman Malik, Samee Ullah Khan
IEEE Trans. Cloud Comput.3
2017 Guest Editors' Introduction: Special Issue on Green and Energy-Efficient Cloud Computing Part II
abstract
The papers in this special section focus on green and energy efficient cloud computing. Cloud computing has had a huge commercial impact and has attracted the interest of the research community. Public clouds allow their customers to outsource the management of physical resources, and rent a variable amount of resources in accordance to their specific needs. Private clouds allow companies to manage on-premises resources, exploiting the capabilities offered by the cloud technologies, such as using virtualization to improve resource utilization and cloud software for resource management automation. Hybrid clouds, where private infrastructures are integrated and complemented by external resources, are becoming a common scenario as well, for example to manage load peaks. Cloud applications are hosted by data centers whose size ranges from tens to tens of thousands of servers, which raises significant challenges related to energy and cost management. It has been estimated that the Information and Communication Technology (ICT) industry alone is responsible for 2-3 percent of the global greenhouse gas emissions. Therefore, we must find innovative methods and tools to manage the energy efficiency and carbon footprint of data centers, so that they can operate and scale in a cost-effective and environmentally sustainable manner. These methods and tools are often categorized as Data Center Infrastructure Management (DCIM) to monitor, control, and optimize data centers with extensive automation. DCIM must also effectively manage the quality of service provided by the data center, since cloud customers require high reliability, availability, usability, and low response times.
Ricardo Bianchini, Samee Ullah Khan, Carlo Mastroianni
IEEE Trans. Cloud Comput.2
2017 MobiContext: A Context-Aware Cloud-Based Venue Recommendation Framework
abstract
In recent years, recommendation systems have seen significant evolution in the field of knowledge engineering. Most of the existing recommendation systems based their models on collaborative filtering approaches that make them simple to implement. However, performance of most of the existing collaborative filtering-based recommendation system suffers due to the challenges, such as: (a) cold start, (b) data sparseness, and (c) scalability. Moreover, recommendation problem is often characterized by the presence of many conflicting objectives or decision variables, such as users' preferences and venue closeness. In this paper, we proposed MobiContext, a hybrid cloud-based bi-objective recommendation framework (BORF) for mobile social networks. The MobiContext utilizes multi-objective optimization techniques to generate personalized recommendations. To address the issues pertaining to cold start and data sparseness, the BORF performs data preprocessing by using the Hub-Average (HA) inference model. Moreover, the Weighted Sum Approach (WSA) is implemented for scalar optimization and an evolutionary algorithm (NSGA-II) is applied for vector optimization to provide optimal suggestions to the users about a venue. The results of comprehensive experiments on a large-scale real dataset confirm the accuracy of the proposed recommendation framework.
Rizwana Irfan, Osman Khalid, Muhammad Usman Shahid Khan, Camelia Chira, Rajiv Ranjan 0001, Fan Zhang 0003, Samee Ullah Khan, Bharadwaj Veeravalli, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Cloud Comput.7
2017 MacroServ: A Route Recommendation Service for Large-Scale Evacuations
abstract
To respond to emergencies in a fast and an effective manner, it is of critical importance to have efficient evacuation plans that lead to minimum road congestions. Although emergency evacuation systems have been studied in the past, the existing approaches, mostly based on multi-objective optimizations, are not scalable enough when involve numerous time varying parameters, such as traffic volume, safety status, and weather conditions. In this paper, we propose a scalable emergency evacuation service, termed the MacroServ that recommends the evacuees with the most preferred routes towards safe locations during a disaster. Unlike many existing approaches that model systems with static network characteristics, our approach considers real-time road conditions to compute the maximum flow capacity of routes in the transportation network. The evacuees are directed towards those routes that are safe and have least congestion resulting in decreased evacuation time. We utilized probability distributions to model the real-life stochastic behaviors of evacuees during emergency scenarios. The results indicate that recommendation of appropriate routes during emergency scenarios play a critical role in quicker and safe evacuation of the population.
Muhammad Usman Shahid Khan, Osman Khalid, Rajiv Ranjan 0001, Fan Zhang 0003, Bharadwaj Veeravalli, Samee Ullah Khan, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Serv. Comput.8
2017 CloudNetSim++: A GUI Based Framework for Modeling and Simulation of Data Centers in OMNeT++
abstract
State-of-the-art cloud simulators in use today are limited in the number of features they provide, lack real network communication models, and do not provide extensive Graphical User Interface (GUI) to support developers and researchers to extend the behavior of the cloud environment. We propose CloudNetSim++, a comprehensive packet level simulator that enables simulation of cloud environments. CloudNetSim++ can be used to evaluate a wide spectrum of cloud components, such as processing elements, storage, networking, Service Level Agreement (SLA), scheduling algorithms, fine grained energy consumption, and VM consolidation algorithms. CloudNetSim++ offers extendibility, which means that the developers and researchers can easily incorporate own algorithms for scheduling, workload consolidation, VM migration, and SLA agreement. The simulation environment of CloudNetSim++ offers a rich GUI that provides a high level view of distributed data centers connected with various network topologies. The package also includes an energy computation module that provides a fine grained analysis of energy consumed by each component. This paper shows the flexibility and effectiveness of CloudNetSim++ through experimental results demonstrated using real-world data center workloads. Moreover, to demonstrate the correctness of CloudNetSim++, we performed formal modeling, analysis, and verification using High-level Petri Nets, Satisfiability Modulo Theories (SMT), and Z3 solver.
Asad Waqar Malik, Kashif Bilal, Saif Ur Rehman Malik, Zahid Anwar, Khurram Aziz, Dzmitry Kliazovich, Nasir Ghani, Samee Ullah Khan, Rajkumar Buyya
IEEE Trans. Serv. Comput.8
2017 Leveraging on Deep Memory Hierarchies to Minimize Energy Consumption and Data Access Latency on Single-Chip Cloud Computers
abstract
Recent advances in chip design and integration technologies have led to the development of Single-Chip Cloud computers which are a microcosm of cloud datacenters. Those computers are based on Network-on-Chip (NoC) architectures with deep memory hierarchies. Developing scheduling algorithms to reduce data access latency as well as energy consumption is a major challenge for such architectures. In this paper, we propose a set of algorithms to jointly address the problem of task scheduling and data allocation in a unified approach. Moreover, we present a feasible system model for NoC based multicores considering a three-level memory hierarchy that effectively captures the energy consumed by various elements of system including: processing cores, caches, and NoC subsystem. Simulation results show the superiority of proposed algorithms compared to two state-of-the-art algorithms found in the literature. The experimental results clearly indicate that algorithms performing data and task scheduling in a joint fashion are superior against techniques implementing task and data scheduling separately.
Tahir Maqsood, Nikos Tziritas, Thanasis Loukopoulos, Sajjad Ahmad Madani, Samee Ullah Khan, Cheng-Zhong Xu 0001
IEEE Trans. Sustain. Comput.5
2017 Data Replication and Virtual Machine Migrations to Mitigate Network Overhead in Edge Computing Systems
abstract
Several virtual machine (VM) placement algorithms have been proposed and studied in the literature with various scopes such as server consolidation or network cost minimization. In most cases, decisions on VM migrations are taken without factoring in directly the data access cost by VMs. In this paper, we investigate the use of data replication in conjunction with the VM assignment problem and target on developing algorithms that decide both on which data should be replicated where and which VM must be migrated so as to minimize the network overhead among traditional cloud and mobile cloud systems. We discuss both the un-capacitated case and the more realistic case whereby datacenters (for the traditional cloud case) and micro-datacenters (for the mobile cloud case) have limited storage and computing capacity. We propose an algorithm based on hyper-graph partitioning to solve the aforementioned problem in an optimal way regarding the unconstrained case and extend it to capture storage and computing capacity constraints. Experimental evaluation shows that the proposed algorithm yields up to 53 percent network overhead reduction when compared to state-of-the-art algorithms found in the literature.
Nikos Tziritas, Maria G. Koziri, Areti Bachtsevani, Thanasis Loukopoulos, Georgios I. Stamoulis, Samee Ullah Khan, Cheng-Zhong Xu 0001
IEEE Trans. Sustain. Comput.6
2016 Slice-based parallelization in HEVC encoding: Realizing the potential through efficient load balancing
abstract
The new video coding standard HEVC (High Efficiency Video Coding) offers the desired compression performance in the era of HDTV and UHDTV, as it achieves nearly 50% bit rate saving compared to H.264/AVC. To leverage the involved computational overhead, HEVC offers three parallelization potentials namely: wavefront parallelization, tile-based and slice-based. In this paper we study slice-based parallelization of HEVC using OpenMP on the encoding part. In particular we delve on the problem of proper slice sizing to reduce load imbalances among threads. Capitalizing on existing ideas for H.264/AVC we develop a fast dynamic approach to decide on load distribution and compare it against an alternative in the HEVC literature. Through experiments with commonly used video sequences, we highlight the merits and drawbacks of the tested heuristics. We then improve upon them for the case of Low-Delay by exploiting GOP structure. The resulting algorithm is shown to clearly outperform its counterparts achieving less than 10% load imbalance in many cases.
Maria G. Koziri, Panos Papadopoulos, Nikos Tziritas, Antonios N. Dadaliaris, Thanasis Loukopoulos, Samee Ullah Khan
MMSP6
2016 Overlay network scheduling design
Khaled B. Shaban, Mahmoud A. Khodeir, Jorge Crichigno, Samee Ullah Khan, Nasir Ghani
Comput. Commun.6
2016 Genetic algorithm in finding Pareto frontier of optimizing data transfer versus job execution in grids
abstract
Summary This work presents a genetic algorithm (GA)‐based optimization technique, called GA‐ParFnt, to find the Pareto frontier for optimizing data transfer versus job execution time in grids. As the performance of a generic GA is not suitable to find such Pareto relationship, major modifications are applied to it so that it can efficiently discover such relationship. The frontier curve representing this relationship is then matched against performance of several scheduling techniques—for both data intensive and computationally intensive applications—to measure their overall performances. Results show that few of these algorithms are far from the Pareto front despite their claims of being efficient in optimizing their targeted objectives. Results also provide invaluable insights into this formidable problem and should aid in the design of future schedulers. Copyright © 2012 John Wiley & Sons, Ltd.
Javid Taheri, Albert Y. Zomaya, Samee Ullah Khan
Concurr. Comput. Pract. Exp.3
2016 Divide-and-conquer approach for solving singular value decomposition based on MapReduce
abstract
Summary Singular value decomposition (SVD) shows strong vitality in the area of information analysis and has significant application value in most of the scientific big data fields. However, with the rapid development of Internet, the information online reveals fast growing trend. For a large‐scale matrix, applying SVD computation directly is both time consuming and memory demanding. There are many works available to speed up the computation of SVD based on the message passing interface model. However, to deal with large‐scale data processing, a MapReduce model has many advantages over a message passing interface model, such as fault tolerance, load balancing and simplicity. For a MapReduce environment, existing approaches only focus on low rank SVD approximation and tall‐and‐skinny matrix SVD computation, and there are no implementations of full rank SVD computation. In this paper, we propose a MapReduce‐based implementation for solving divide‐and‐conquer SVD algorithm. To achieve high performance, we design a two‐stage task scheduling strategy based on the mathematical characteristics of divide‐and‐conquer SVD algorithm. To further strengthen the performance, we propose a row‐index‐based divide algorithm, a pipelined task scheduling method, and revised block matrix multiplication in MapReduce framework. Experimental result shows the efficiency of our algorithm. Our implementation can accommodate full rank SVD computation of large‐scale matrix very efficiently. Copyright © 2014 John Wiley & Sons, Ltd.
Shuoyi Zhao, Ruixuan Li 0001, Weijun Xiao, Xinhua Dong, Dongjie Liao, Samee Ullah Khan, Keqin Li 0001
Concurr. Comput. Pract. Exp.7
2016 Big Data Reduction Methods: A Survey
abstract
Research on big data analytics is entering in the new phase called fast data where multiple gigabytes of data arrive in the big data systems every second. Modern big data systems collect inherently complex data streams due to the volume, velocity, value, variety, variability, and veracity in the acquired data and consequently give rise to the 6Vs of big data. The reduced and relevant data streams are perceived to be more useful than collecting raw, redundant, inconsistent, and noisy data. Another perspective for big data reduction is that the million variables big datasets cause the curse of dimensionality which requires unbounded computational resources to uncover actionable knowledge patterns. This article presents a review of methods that are used for big data reduction. It also presents a detailed taxonomic discussion of big data reduction methods including the network theory, big data compression, dimension reduction, redundancy elimination, data mining, and machine learning methods. In addition, the open research issues pertinent to the big data reduction are also highlighted.
Muhammad Habib Ur Rehman, Chee Sun Liew, Assad Abbas, Prem Prakash Jayaraman, Ying Wah Teh, Samee Ullah Khan
Data Sci. Eng.6
2016 Performance analysis of data intensive cloud systems based on data management and replication: a survey
Saif Ur Rehman Malik, Samee Ullah Khan, Sam J. Ewen, Nikos Tziritas, Joanna Kolodziej, Albert Y. Zomaya, Sajjad Ahmad Madani, Nasro Min-Allah, Lizhe Wang 0001, Cheng-Zhong Xu 0001, Qutaibah M. Malluhi, Johnatan E. Pecero, Pavan Balaji, Abhinav Vishnu, Rajiv Ranjan 0001, Sherali Zeadally, Hongxiang Li 0001
Distributed Parallel Databases2
2016 MapReduce-based fast fuzzy c-means algorithm for large-scale underwater image segmentation
Xiu Li 0001, Jingdong Song, Fan Zhang 0003, Xiaogang Ouyang, Samee Ullah Khan
Future Gener. Comput. Syst.5
2016 CA-DAG: Modeling Communication-Aware Applications for Scheduling in Cloud Computing
Dzmitry Kliazovich, Johnatan E. Pecero, Andrei Tchernykh, Pascal Bouvry, Samee Ullah Khan, Albert Y. Zomaya
J. Grid Comput.5
2016 An Energy-Efficient Task Scheduling Algorithm in DVFS-enabled Cloud Environment
Zhuo Tang, Ling Qi, Zhenzhen Cheng, Kenli Li 0001, Samee Ullah Khan, Keqin Li 0001
J. Grid Comput.5
2016 Secure and dependable software defined networks
Adnan Akhunzada, Abdullah Gani, Nor Badrul Anuar, Muhammad Khurram Khan, Amir Hayat, Samee Ullah Khan
J. Netw. Comput. Appl.7
2016 A hybrid genetic algorithm for optimization of scheduling workflow applications in heterogeneous computing systems
Saima Gulzar Ahmad, Chee Sun Liew, Ehsan Ullah Munir, Tan Fong Ang, Samee Ullah Khan
J. Parallel Distributed Comput.5
2016 Personalized healthcare cloud services for disease risk assessment and wellness management using social media
Assad Abbas, Muhammad Usman Shahid Khan, Samee Ullah Khan
Pervasive Mob. Comput.4
2016 Designing and Modeling of Covert Channels in Operating Systems
abstract
Covert channels are widely considered as a major risk of information leakage in various operating systems, such as desktop, cloud, and mobile systems. The existing works of modeling covert channels have mainly focused on using finite state machines (FSMs) and their transforms to describe the process of covert channel transmission. However, a FSM is rather an abstract model, where information about the shared resource, synchronization, and encoding/decoding cannot be presented in the model, making it difficult for researchers to realize and analyze the covert channels. In this paper, we use the high-level Petri Nets (HLPN) to model the structural and behavioral properties of covert channels. We use the HLPN to model the classic covert channel protocol. Moreover, the results from the analysis of the HLPN model are used to highlight the major shortcomings and interferences in the protocol. Furthermore, we propose two new covert channel models, namely: (a) two channel transmission protocol (TCTP) model and (b) self-adaptive protocol (SAP) model. The TCTP model circumvents the mutual inferences in encoding and synchronization operations; whereas the SAP model uses sleeping time and redundancy check to ensure correct transmission in an environment with strong noise. To demonstrate the correctness and usability of our proposed models in heterogeneous environments, we implement the TCTP and SAP in three different systems: (a) Linux, (b) Xen, and (c) Fiasco.OC. Our implementation also indicates the practicability of the models in heterogeneous, scalable and flexible environments.
Yuqi Lin, Saif Ur Rehman Malik, Kashif Bilal, Qiusong Yang, Yongji Wang 0002, Samee Ullah Khan
IEEE Trans. Computers6
2016 A Decentralized Damage Detection System for Wireless Sensor and Actuator Networks
abstract
The unprecedented capabilities of monitoring and responding to stimuli in the physical world of wireless sensor and actuator networks (WSAN) enable these networks to provide the underpinning for several Smart City applications, such as structural health monitoring (SHM). In such applications, civil structures, endowed with wireless smart devices, are able to self-monitor and autonomously respond to situations using computational intelligence. This work presents a decentralized algorithm for detecting damage in structures by using a WSAN. As key characteristics, beyond presenting a fully decentralized (in-network) and collaborative approach for detecting damage in structures, our algorithm makes use of cooperative information fusion for calculating a damage coefficient. We conducted experiments for evaluating the algorithm in terms of its accuracy and efficient use of the constrained WSAN resources. We found that our collaborative and information fusion-based approach ensures the accuracy of our algorithm and that it can answer promptly to stimuli (1.091 s), triggering actuators. Moreover, for 100 nodes or less in the WSAN, the communication overhead of our algorithm is tolerable and the WSAN running our algorithm, operating system and protocols can last as long as 468 days.
Igor Leão dos Santos, Luci Pirmez, Luiz Fernando Rust da Costa Carmo, Paulo F. Pires, Flávia Coimbra Delicato, Samee Ullah Khan, Albert Y. Zomaya
IEEE Trans. Computers6
2016 Guest Editors' Introduction: Special Issue on Green and Energy-Efficient Cloud Computing: Part I
abstract
The papers in this special section focus on green and energy efficient cloud computing. Cloud computing has had a huge commercial impact and has attracted the interest of the research community. Public clouds allow their customers to outsource the management of physical resources, and rent a variable amount of resources in accordance to their specific needs. Private clouds allow companies to manage on-premises resources, exploiting the capabilities offered by the cloud technologies, such as using virtualization to improve resource utilization and cloud software for resource management automation. Hybrid clouds, where private infrastructures are integrated and complemented by external resources, are becoming a common scenario as well, for example to manage load peaks.
Ricardo Bianchini, Samee Ullah Khan, Carlo Mastroianni
IEEE Trans. Cloud Comput.2
2016 Formal Verification of the xDAuth Protocol
abstract
Service-oriented architecture offers a flexible paradigm for information flow among collaborating organizations. As information moves out of an organization boundary, various security concerns may arise, such as confidentiality, integrity, and authenticity that needs to be addressed. Moreover, verifying the correctness of the communication protocol is also an important factor. This paper focuses on the formal verification of the xDAuth protocol, which is one of the prominent protocols for identity management in cross domain scenarios. We have modeled the information flow of xDAuth protocol using high-level Petri nets to understand the protocol information flow in a distributed environment. We analyze the rules of information flow using Z language, while Z3 SMT solver is used for the verification of the model. Our formal analysis and verification results reveal the fact that the protocol fulfills its intended purpose and provides the security for the defined protocol specific properties, e.g., secure secret key authentication, and Chinese wall security policy and secrecy specific properties, e.g., confidentiality, integrity, and authenticity.
Quratulain Alam, Saher Tabbasum, Saif Ur Rehman Malik, Masoom Alam, Tamleek Ali, Adnan Akhunzada, Samee Ullah Khan, Athanasios V. Vasilakos, Rajkumar Buyya
IEEE Trans. Inf. Forensics Secur.7
2016 On Improving Constrained Single and Group Operator Placement Using Evictions in Big Data Environments
abstract
With an ever increasing amount of data generated by scientific experiments, social networks and mobile as well as wireless sensor networks, reducing resource consumption by big data applications becomes of paramount importance. Towards this end, filtering data close to the data sources is a common strategy in order to reduce network traffic. Assuming a network of nodes, each potentially generating data and a query in the form of a single operator to be applied in these data, the basic statement of the operator placement problem is: find the best node to place the operator so that the network traffic is minimized. In this paper we study the problem of placing a set of communicating operators exhibiting a tree structure over a tree network of nodes with capacity constraints. We take advantage of our previous work on unconstrained placement in order to develop a new approach enabling both single and group operator migrations using evictions of hosted operators if free space is required. To enhance their applicability, the algorithms work in a distributed asynchronous manner, requiring only minimal knowledge at each network node. Results from simulation experiments show that the proposed algorithms reduce considerably network overhead against their counterparts.
Nikos Tziritas, Thanasis Loukopoulos, Samee Ullah Khan, Cheng-Zhong Xu 0001, Albert Y. Zomaya
IEEE Trans. Serv. Comput.3
2016 Skyline Discovery and Composition of Multi-Cloud Mashup Services
abstract
A cloud mashup is composed of multiple services with shared datasets and integrated functionalities. For example, the elastic compute cloud (EC2) provided by Amazon Web Service (AWS), the authentication and authorization services provided by Facebook, and the Map service provided by Google can all be mashed up to deliver real-time, personalized driving route recommendation service. To discover qualified services and compose them with guaranteed quality of service (QoS), we propose an integrated skyline query processing method for building up cloud mashup applications. We use a similarity test to achieve optimal localized skyline. This mashup method scales well with the growing number of cloud sites involved in the mashup applications. Faster skyline selection, reduced composition time, dataset sharing, and resources integration assure the QoS over multiple clouds. We experiment with the quality of Web service (QWS) benchmark over 10,000 Web services along six QoS dimensions. By utilizing block-elimination, data-space partitioning, and service similarity pruning, the skyline process is shortened by three times, when compared with two state-of-the-art methods.
Fan Zhang 0003, Kai Hwang 0001, Samee Ullah Khan, Qutaibah M. Malluhi
IEEE Trans. Serv. Comput.3
2015 Coordination Strategies for Agent Migrations in Wireless Sensor Networks
abstract
Agent-based middleware platforms for wireless sensor networks (WSNs) have received a lot of attention during the last years, due to their great flexibility in re-programming, monitoring, handling, and optimizing the application as well as the whole system. Even though many algorithms have been proposed for the dynamic placement of agents within the WSN, they do not take into account the coordination aspects of such migrations. This not only may result in slow convergence but also in perpetual oscillations of agent migrations, degrading application and system performance. In this paper, we propose full-coordination, semi-coordination, and non-coordination agent migration strategies, and evaluate their convergence and network overhead. We also provide proofs that convergence is guaranteed when dynamic agent placement algorithms adopt our proposed strategies. Our results show that the semi-coordination strategy is superior in terms of both network overhead and convergence rate.
Nikos Tziritas, Thanasis Loukopoulos, Spyros Lalis, Samee Ullah Khan, Cheng-Zhong Xu 0001
ICPADS4
2015 Energy efficient genetic-based schedulers in computational grids
abstract
Summary In today's highly parametrized distributed computational environments, such as green grid clusters and clouds, the growing power and cooling rates are becoming the dominant part of the users' and system managers' budgets. Computational grids, owing to their sheer sizes, still require advanced methodologies and strategies for supporting the scheduling of the users' tasks and applications to the distributed resources. The efficient resource allocation becomes even more challenging when energy utilization, beyond the conventional scheduling criteria, such as Makespan , is treated as first‐class additional scheduling objective. In this paper, we address the independent batch scheduling in computational grid as a bi‐objective global minimization problem with Makespan and energy consumption as the main criteria. We apply the dynamic voltage and frequency scaling model for the management of the cumulative power energy utilized by the grid resources. We develop three genetic algorithms as energy‐aware grid schedulers, which were empirically evaluated in three grid size scenarios in static and dynamic modes. The simulation results confirmed the effectiveness of the proposed genetic algorithm‐based schedulers in the reduction of the energy consumed by the whole system and in dynamic load balancing of the resources in grid clusters, which is sufficient to maintain the desired quality level(s). Copyright © 2012 John Wiley & Sons, Ltd.
Joanna Kolodziej, Samee Ullah Khan, Lizhe Wang 0001, Albert Y. Zomaya
Concurr. Comput. Pract. Exp.2
2015 A cloud based health insurance plan recommendation system: A user centered approach
Assad Abbas, Kashif Bilal, Samee Ullah Khan
Future Gener. Comput. Syst.4
2015 Artificial Neural Network support to monitoring of the evolutionary driven security aware scheduling in computational distributed environments
Daniel Grzonka, Joanna Kolodziej, Jie Tao 0001, Samee Ullah Khan
Future Gener. Comput. Syst.4
2015 A task-level adaptive MapReduce framework for real-time streaming data in healthcare applications
Fan Zhang 0003, Samee Ullah Khan, Keqin Li 0001, Kai Hwang 0001
Future Gener. Comput. Syst.3
2015 CloudFlow: A data-aware programming model for cloud workflow applications on modern HPC systems
Fan Zhang 0003, Qutaibah M. Malluhi, Tamer Elsayed, Samee Ullah Khan, Keqin Li 0001, Albert Y. Zomaya
Future Gener. Comput. Syst.4
2015 Image transmission using unequal error protected multi-fold turbo codes over a two-user power-line binary adder channel
abstract
Impulsive noise is one of the major challenges for reliable transmission over power lines. Interleavers provide higher protection against the impulsive noise by dispersing information across the channel and spreading the burst of errors over multiple codewords. Multi‐fold turbo (MFT) coding is a technique that improves the communication reliability using multiple interleavers. In the MFT codes, each data subsequence is equally protected. For applications in which data constitute information with various levels of importance, it is intuitive to offer the more important subsequence, a stronger protection. A modified form of the MFT codes capable of providing unequal error protection over a two‐user power‐line binary adder channel is proposed here. As a benchmark, two test images are transmitted across the channel. The trellis‐based iterative algorithm is modified for the two‐user scenario to decode the received signal. The simulation results show a gain of 1.5 dB for the modified MFT code over the conventional turbo codes for each of the transmitted images. A gain of 2 dB is also recorded for the most protected component of each image over the least protected components.
Abbas Khalid, Eraj Khan, Bamidele Adebisi, Bahram Honary, Samee Ullah Khan
IET Image Process.5
2015 Impact analysis and change propagation in service-oriented enterprises: A systematic review
Khubaib Amjad Alam, Rodina Binti Ahmad, Adnan Akhunzada, Mohd Hairul Nizam Bin Md Nasir, Samee Ullah Khan
Inf. Syst.5
2015 The rise of "big data" on cloud computing: Review and open research issues
Ibrahim Abaker Targio Hashem, Ibrar Yaqoob, Nor Badrul Anuar, Salimah Mokhtar, Abdullah Gani, Samee Ullah Khan
Inf. Syst.6
2015 Security in cloud computing: Opportunities and challenges
Samee Ullah Khan, Athanasios V. Vasilakos
Inf. Sci.2
2015 Re-Stream: Real-time and energy-efficient resource scheduling in big data stream computing environments
Dawei Sun 0001, Guangyan Zhang, Samee Ullah Khan, Keqin Li 0001
Inf. Sci.5
2015 Seamless application execution in mobile cloud computing: Motivation, taxonomy, and open challenges
Ejaz Ahmed 0003, Abdullah Gani, Muhammad Khurram Khan, Rajkumar Buyya, Samee Ullah Khan
J. Netw. Comput. Appl.5
2015 Automatic player detection and identification for sports entertainment applications
Tauseef Ali, Shahid Khattak, Laiq Hasan, Samee Ullah Khan
Pattern Anal. Appl.5
2015 Fast and Scalable Multi-Way Analysis of Massive Neural Data
abstract
Analysis of neural data with multiple modes and high density has recently become a trend with the advances in neuroscience research and practices. There exists a pressing need for an approach to accurately and uniquely capture the features without loss or destruction of the interactions amongst the modes (typically) of space, time, and frequency. Moreover, the approach must be able to quickly analyze the neural data of exponentially growing scales and sizes, in tens or even hundreds of channels, so that timely conclusions and decisions may be made. A salient approach to multi-way data analysis is the parallel factor analysis (PARAFAC) that manifests its effectiveness in the decomposition of the electroencephalography (EEG). However, the conventional PARAFAC is only suited for offline data analysis due to the high complexity, which computes to be$O(n^{2})$with the increasing data size. In this study, a large-scale PARAFAC method has been developed, which is supported by general-purpose computing on the graphics processing unit (GPGPU). Comparing to the PARAFAC running on conventional CPU-based platform, the new approach dramatically excels by${>}360$times in run-time performance, and effectively scales by${>}400$times in all dimensions. Moreover, the proposed approach forms the basis of a model for the analysis of electrocochleography (ECoG) recordings obtained from epilepsy patients, which proves to be effective in the epilepsy state detection. The time evolutions of the proposed model are well correlated with the clinical observations. Moreover, the frequency signature is stable and high in the ictal phase. Furthermore, the spatial signature explicitly identifies the propagation of neural activities among various brain regions. The model supports real-time analysis of ECoG in${>}1{,}000$channels on an inexpensive and available cyber-infrastructure.
Dan Chen 0001, Xiaoli Li 0002, Lizhe Wang 0001, Samee Ullah Khan
IEEE Trans. Computers4
2015 Survivable Cloud Network Mapping for Disaster Recovery Support
abstract
Network virtualization is a key provision for improving the scalability and reliability of cloud computing services. In recent years, various mapping schemes have been developed to reserve VN resources over substrate networks. However, many cloud providers are very concerned about improving service reliability under catastrophic disaster conditions yielding multiple system failures. To address this challenge, this work presents a novel failure region-disjoint VN mapping scheme to improve VN mapping survivability. The problem is first formulated as a mixed integer linear programming problem and then two heuristic solutions are proposed to compute a pair of failure region-disjoint VN mappings. The solution also takes into account mapping costs and load balancing concerns to help improve resource efficiencies. The schemes are then analyzed in detail for a variety of networks and their overall performances compared to some existing survivable VN mapping schemes.
Khaled B. Shaban, Nasir Ghani, Samee Ullah Khan, Mahshid Rahnamay-Naeini, Majeed M. Hayat, Chadi Assi
IEEE Trans. Computers4
2015 JEM: Just in Time/Just Enough Energy Management Methodology for Computing Systems
abstract
This paper presents the Just in Time/Just Enough Energy Management (JEM) methodology that is applicable to a broad range of computing systems. The conventional concept of a fixed voltage supply (VDD) scheme for both performance and power saving modes of computing systems is revisited and is improved with JEM. The JEM consists of an efficient DC/DC converter and a Power Management Integrated Circuit (PMIC) with a feedback to monitor the activities within a given computing system, providing a new means for dynamic voltage scaling at the system level. The JEM is tested and validated on a blade server that results in 15.11 percent power savings at the motherboard level. A significant thermal improvement of 9.0CC is measured in a 16 GB memory module of the blade server, as well. Moreover, a JEM enabled CMOS circuit depicts a remarkable reduction in the supply current. Furthermore, the JEM is compared to a conventional power supply design, with significant improvement in the processor performance and considerable power savings in the blade server.
Muhammad Jawad 0001, Sahibzada Muhammad Ali, Joel A. Jorgenson, Samee Ullah Khan
IEEE Trans. Computers4
2015 CloudGenius: A Hybrid Decision Support Method for Automating the Migration of Web Application Clusters to Public Clouds
abstract
With the increase in cloud service providers, and the increasing number of compute services offered, a migration of information systems to the cloud demands selecting the best mix of compute services and virtual machine (VM ) images from an abundance of possibilities. Therefore, a migration process for web applications has to automate evaluation and, in doing so, ensure that Quality of Service (QoS) requirements are met, while satisfying conflicting selection criteria like throughput and cost. When selecting compute services for multiple connected software components, web application engineers must consider heterogeneous sets of criteria and complex dependencies across multiple layers, which is impossible to resolve manually. The previously proposed CloudGenius framework has proven its capability to support migrations of single-component web applications. In this paper, we expand on the additional complexity of facilitating migration support for multi-component web applications. In particular, we present an evolutionary migration process for web application clusters distributed over multiple locations, and clearly identify the most important criteria relevant to the selection problem. Moreover, we present a multi-criteria-based selection algorithm based on Analytic Hierarchy Process (AHP). Because the solution space grows exponentially, we developed a Genetic Algorithm (GA)-based approach to cope with computational complexities in a growing cloud market. Furthermore, a use case example proofs CloudGenius’ applicability. To conduct experiments, we implemented CumulusGenius, a prototype of the selection algorithm and the GA deployable on hadoop clusters. Experiments with CumulusGenius give insights on time complexities and the quality of the GA.
Michael Menzel 0002, Rajiv Ranjan 0001, Lizhe Wang 0001, Samee Ullah Khan, Jinjun Chen
IEEE Trans. Computers4
2015 Adaptive Workflow Scheduling on Cloud Computing Platforms with IterativeOrdinal Optimization
abstract
The scheduling of multitask jobs on clouds is an NP-hard problem. The problem becomes even worse when complex workflows are executed on elastic clouds, such as Amazon EC2 or IBM RC2. The main difficulty lies in the large search space and high overhead of generating optimal schedules, especially for real-time applications with dynamic workloads. In this work, a new iterative ordinal optimization (IOO) method is proposed. The ordinal optimization method is applied in each iteration to achieve sub-optimal schedules. IOO aims at generating more efficient schedules from a global perspective over a long period. We prove through overhead analysis the advantages in time and space efficiency in using the IOO method. The IOO method is designed to adapt to system dynamism to yield suboptimal performance. In cloud experiments on IBM RC2 cloud, we execute 20,000 tasks in LIGO (Laser Interferometer Gravitational-wave Observatory) verification workflow on 128 virtual machines. The IOO schedule is generated in less than 1,000 seconds, while using the Monte Carlo simulation takes 27.6 hours, 100 times longer to yield an optimal schedule. The IOO-optimized schedule results in a throughput of 1,100 tasks/sec with 7 GB memory demand, compared with 60 percent decrease in throughput and 70 percent increase in memory demand in using the Monte Carlo method. Our LIGO experimental results clearly demonstrate the advantage of using the IOO-based workflow scheduling over the traditional blind-pick, ordinal optimization, or Monte Carlo methods. These numerical results are also validated by the theoretical complexity and overhead analysis provided.
Fan Zhang 0003, Kai Hwang 0001, Keqin Li 0001, Samee Ullah Khan
IEEE Trans. Cloud Comput.5
2015 Distributed Algorithms for the Operator Placement Problem
abstract
Operator placement plays a key role in reducing the aggregate network overhead within a wireless sensor network (WSN) to extend battery life and the longevity of the network. Consequently, optimal algorithms for the operator placement problem (OPP) are of paramount importance to WSN performance. Unfortunately, the OPP becomes NP-complete when capacity constraints on the WSN nodes are taken into account. There are many algorithms in the literature that tackle the OPP; however, most of them consider tree-structured query graphs without limitations regarding the operators hosted by the WSN nodes. Therefore, there is a need to propose sophisticated approaches such that the problem is solved in an effective fashion. In this paper, we propose a fully distributed approach that takes into account the WSN node capacity constraints. The proposed approach is thoroughly evaluated through simulations and the results reveal that the proposed approach is superior to several state-of-the-art algorithms, such as DRA, DBA, MCFA, dFNS, and GRAL* found in the literature.
Nikos Tziritas, Thanasis Loukopoulos, Samee Ullah Khan, Cheng-Zhong Xu 0001
IEEE Trans. Comput. Soc. Syst.3
2014 A taxonomy and survey on Green Data Center Networks
Kashif Bilal, Saif Ur Rehman Malik, Osman Khalid, Abdul Hameed, Vidura Wijayasekara, Rizwana Irfan, Sarjan Shrestha, Debjyoti Dwivedy, Muhammad Usman Shahid Khan, Assad Abbas, Nauman Jalil, Samee Ullah Khan
Future Gener. Comput. Syst.14
2014 Security, energy, and performance-aware resource allocation mechanisms for computational grids
Joanna Kolodziej, Samee Ullah Khan, Lizhe Wang 0001, Marek Kisiel-Dorohinicki, Sajjad Ahmad Madani, Ewa Niewiadomska-Szynkiewicz, Albert Y. Zomaya, Cheng-Zhong Xu 0001
Future Gener. Comput. Syst.2
2014 Multi-objective scheduling of many tasks in cloud platforms
Fan Zhang 0003, Keqin Li 0001, Samee Ullah Khan, Kai Hwang 0001
Future Gener. Comput. Syst.4
2014 Survey on Grid Resource Allocation Mechanisms
Muhammad Bilal Qureshi, Maryam Mehri Dehnavi, Nasro Min-Allah, Muhammad Shuaib Qureshi, Hameed Hussain, Ilias Rentifis, Nikos Tziritas, Thanasis Loukopoulos, Samee Ullah Khan, Cheng-Zhong Xu 0001, Albert Y. Zomaya
J. Grid Comput.9
2014 The optical character recognition of Urdu-like cursive scripts
Saeeda Naz, Khizar Hayat 0002, Muhammad Imran Razzak, Muhammad Waqas Anwar, Sajjad Ahmad Madani, Samee Ullah Khan
Pattern Recognit.6
2014 C2Detector: a covert channel detection framework in cloud computing
abstract
ABSTRACT Cloud computing is becoming increasingly popular because of the dynamic deployment of computing service. Another advantage of cloud is that data confidentiality is protected by the cloud provider with the virtualization technology. However, a covert channel can break the isolation of the virtualization platform and leak confidential information without letting it known by virtual machines. In this paper, the threat model of covert channels is analyzed. The channels are classified into three categories, and only the category that is new to cloud computing is concerned, for example, CPU load‐based, cache‐based, and shared memory‐based covert channels. The covert channel scenario is modeled into an error‐corrected four‐state automaton, and two error‐corrected algorithms are designed. A new detection framework termed C2Detector is presented. C2Detector includes a captor located in the hypervisor and a two‐phase synthesis algorithm implemented as Markov and Bayesian detectors. A prototype of C2Detector is implemented on Xen hypervisor, and its performance of detecting the covert channels is demonstrated. The experiment results show that C2Detector can detect the three types of the covert channels with an acceptable false positive rate by using a pessimistic threshold. Moreover, C2Detector is a plug‐in framework and can be easily extended. It is believed that new covert channels can be detected by C2Detector in the future. Copyright © 2013 John Wiley & Sons, Ltd.
JingZheng Wu, Liping Ding, Nasro Min-Allah, Samee Ullah Khan, Yongji Wang 0002
Secur. Commun. Networks5
2014 Single and Group Agent Migration: Algorithms, Bounds, and Optimality Issues
abstract
Recent embedded middleware platforms enable the structuring of an application as a set of collaborating agents deployed on various nodes of the underlying wireless sensor network (WSN). Of particular importance is the network cost incurred due to agent communication, which in turn depends on how the agents are placed within the WSN system. In this paper, we present two agent migration algorithms with the aim of minimizing the total network overhead. The first one takes independent single agent migration decisions, while the second one considers groups of agents for migration. Both algorithms work in a fully distributed fashion based on the knowledge available locally at each node, and can be used both for one-shot initial application deployment as well as for the continuous updating of agent placement. We also propose two methodologies to tackle the problem when WSN nodes have limited capacity. We show through theoretical analysis that one of our algorithms (called GRAL$\ast$) always results in an optimal placement, while for the rest of the algorithms, we derive approximation ratios pertaining to their performance. We evaluate the performance of our algorithms through a series of simulation experiments. Results show that group migration algorithms are superior compared to single agent migration algorithms with the performance difference reaching 34% for some settings.
Nikos Tziritas, Samee Ullah Khan, Thanasis Loukopoulos, Spyros Lalis, Cheng-Zhong Xu 0001, Petros Lampsas
IEEE Trans. Computers2
2014 CIVSched: A Communication-Aware Inter-VM Scheduling Technique for Decreased Network Latency between Co-Located VMs
abstract
Server consolidation in cloud computing environments makes it possible for multiple servers or desktops to run on a single physical server for high resource utilization, low cost, and reduced energy consumption. However, the scheduler in the virtual machine monitor (VMM), such as Xen credit scheduler, is agnostic about the communication behavior between the guest operating systems (OS). The aforementioned behavior leads to increased network communication latency in consolidated environments. In particular, the CPU resources management has a critical impact on the network latency between co-located virtual machines (VMs) when there are CPUand I/O-intensive workloads running simultaneously. This paper presents the design and implementation of a communication-aware inter-VM scheduling (CIVSched) technique that takes into account the communication behavior between inter-VMs running on the same virtualization platform. The CIVSched technique inspects the network packets transmitted between local co-resident domains to identify the target VM and process that will receive the packets. Thereafter, the target VM and process are preferentially scheduled by the VMM and the guest OS. The cooperation of these two schedulers makes the network packets to be timely received by the target application. Experimental results on the Xen virtualization platform depict that the CIVSched technique can reduce the average response time of network traffic by approximately 19 percent for the highly consolidated environment, while keeping the inherent fairness of the VMM scheduler.
Bei Guan, JingZheng Wu, Yongji Wang 0002, Samee Ullah Khan
IEEE Trans. Cloud Comput.4
2014 A Review on the State-of-the-Art Privacy-Preserving Approaches in the e-Health Clouds
abstract
Cloud computing is emerging as a new computing paradigm in the healthcare sector besides other business domains. Large numbers of health organizations have started shifting the electronic health information to the cloud environment. Introducing the cloud services in the health sector not only facilitates the exchange of electronic medical records among the hospitals and clinics, but also enables the cloud to act as a medical record storage center. Moreover, shifting to the cloud environment relieves the healthcare organizations of the tedious tasks of infrastructure management and also minimizes development and maintenance costs. Nonetheless, storing the patient health data in the third-party servers also entails serious threats to data privacy. Because of probable disclosure of medical records stored and exchanged in the cloud, the patients' privacy concerns should essentially be considered when designing the security and privacy mechanisms. Various approaches have been used to preserve the privacy of the health information in the cloud environment. This survey aims to encompass the state-of-the-art privacy-preserving approaches employed in the e-Health clouds. Moreover, the privacy-preserving approaches are classified into cryptographic and noncryptographic approaches and taxonomy of the approaches is also presented. Furthermore, the strengths and weaknesses of the presented approaches are reported and some open issues are highlighted.
Assad Abbas, Samee Ullah Khan
IEEE J. Biomed. Health Informatics2
2014 OmniSuggest: A Ubiquitous Cloud-Based Context-Aware Recommendation System for Mobile Social Networks
abstract
The evolution of mobile social networks and the availability of online check-in services, such as Foursquare and Gowalla, have initiated a new wave of research in the area of venue recommendation systems. Such systems recommend places to users closely related to their preferences. Although venue recommendation systems have been studied in recent literature, the existing approaches, mostly based on collaborative filtering, suffer from various issues, such as: 1) data sparseness, 2) cold start, and 3) scalability. Moreover, many existing schemes are limited in functionality, as the generated recommendations do not consider group of “friends” type situations. Furthermore, the traditional systems do not take into account the effect of real-time physical factors (e.g., distance from venue, traffic, and weather conditions) on recommendations. To address the aforementioned issues, this paper proposes a novel cloud-based recommendation framework OmniSuggest that utilizes: 1) Ant colony algorithms, 2) social filtering, and 3) hub and authority scores, to generate optimal venue recommendations. Unlike existing work, our approach suggests venues at a finer granularity for an individual or a “group” of friends with similar interest. Comprehensive experiments are conducted with a large-scale real dataset collected from Foursquare. The results confirm that our method offers more effective recommendations than many state of the art schemes.
Osman Khalid, Muhammad Usman Shahid Khan, Samee Ullah Khan, Albert Y. Zomaya
IEEE Trans. Serv. Comput.3
2013 CA-DAG: Communication-Aware Directed Acyclic Graphs for Modeling Cloud Computing Applications
abstract
The review of the requirements of different cloud applications identified the need to consider communication processes explicitly and equally to the computing tasks. Following this observation, we propose a new communication-aware model for cloud computing applications, called CA-DAG. This model is based on Directed Acyclic Graphs (DAGs) that in addition to computing vertices include separate vertices to represent communications. Such a representation allows making separate resource allocation decisions, assigning processors to handle computing jobs and network resources for information transmissions, such as application database requests.
Dzmitry Kliazovich, Johnatan E. Pecero, Andrei Tchernykh, Pascal Bouvry, Samee Ullah Khan, Albert Y. Zomaya
IEEE CLOUD5
2013 Genetic-Based Solutions For Independent Batch Scheduling In Data Grids
abstract
Scheduling in traditional distributed systems has been mainly studied for system performance parameters without data transmission requirements. With the emergence of Data Grids (DGs) and Data Centers, data-aware scheduling has become a major research issue. In this work we present two implementations of classical genetic-based data-aware schedulers of independent tasks submitted to the grid environment. The results of a simple. empirical analysis confirm the high effectiveness of the genetic algorithms in solving very complex data intensive combinatorial optimization problems.
Joanna Kolodziej, Magdalena Szmajduch, Samee Ullah Khan, Lizhe Wang 0001, Dan Chen 0001
ECMS3
2013 Accounting for load variation in energy-efficient data centers
abstract
The energy consumption in data centers is drastically increasing and becoming a significant portion in the data center operating expenses. Enabling a sleep mode in the idle computing servers and network hardware is the most efficient method to avoid unnecessary power consumption. However, changes in the power modes introduce considerable delays. Moreover, inability to wake up a sleeping server immediately requires an availability of a pool of idle servers able to accommodate incoming load in the short term to prevent QoS degradation. In this paper we investigate the amount of computing servers and network hardware needed to accommodate different incoming load patters in the data centers. Furthermore, we propose to build these servers on energy efficient hardware, which is costly but can scale its power consumption with the offered load levels. The evaluation results show that the proposed methodology can save up to $750 per server per year on average.
Dzmitry Kliazovich, Sisay T. Arzo, Fabrizio Granelli, Pascal Bouvry, Samee Ullah Khan
ICC5
2013 Application-Aware Workload Consolidation to Minimize Both Energy Consumption and Network Load in Cloud Environments
abstract
In this paper we tackle the problem of virtual machine (VM) placement onto physical servers to jointly optimize two objective functions. The first objective is to minimize the total energy spent within a cloud due to the servers that are commissioned to satisfy the computational demands of VMs. The second objective is to minimize the total network overhead incurred due to: (a) communicational dependencies between VMs, and (b) the VM migrations performed for the transition from an old assignment scheme to a new one. We study different methodologies for solving the aforementioned problem. The first approach is based on VM packing algorithms that optimize the above objective functions separately, reaching a single solution. The other approach is to tackle simultaneously the two optimization targets and define a set of non-dominating solutions. Performance evaluation using simulation experiments reveals interesting trade-offs between energy consumption and network load.
Nikos Tziritas, Cheng-Zhong Xu 0001, Thanasis Loukopoulos, Samee Ullah Khan, Zhibin Yu 0001
ICPP4
2013 Quantitative comparisons of the state-of-the-art data center architectures
abstract
SUMMARY Data centers are experiencing a remarkable growth in the number of interconnected servers. Being one of the foremost data center design concerns, network infrastructure plays a pivotal role in the initial capital investment and ascertaining the performance parameters for the data center. Legacy data center network (DCN) infrastructure lacks the inherent capability to meet the data centers growth trend and aggregate bandwidth demands. Deployment of even the highest‐end enterprise network equipment only delivers around 50% of the aggregate bandwidth at the edge of network. The vital challenges faced by the legacy DCN architecture trigger the need for new DCN architectures, to accommodate the growing demands of the ‘cloud computing’ paradigm. We have implemented and simulated the state of the art DCN models in this paper, namely: (a) legacy DCN architecture, (b) switch‐based, and (c) hybrid models, and compared their effectiveness by monitoring the network: (a) throughput and (b) average packet delay. The presented analysis may be perceived as a background benchmarking study for the further research on the simulation and implementation of the DCN‐customized topologies and customized addressing protocols in the large‐scale data centers. We have performed extensive simulations under various network traffic patterns to ascertain the strengths and inadequacies of the different DCN architectures. Moreover, we provide a firm foundation for further research and enhancement in DCN architectures. Copyright © 2012 John Wiley & Sons, Ltd.
Kashif Bilal, Samee Ullah Khan, Hongxiang Li 0001, Khizar Hayat 0002, Sajjad Ahmad Madani, Nasro Min-Allah, Lizhe Wang 0001, Dan Chen 0001, Majid I. Iqbal, Cheng-Zhong Xu 0001, Albert Y. Zomaya
Concurr. Comput. Pract. Exp.2
2013 Scalable optimization in grid, cloud, and intelligent network computing - foreword
abstract
Global optimization in large-scale distributed systems requires massive amounts of computations for complex objective functions. Conventional global optimization based on stochastic algorithms cannot guarantee an actual global optimum with a finite searching iteration. Therefore, scalability is a desirable feature for the optimization techniques in highly distributed dynamic environments, where the storage and computing capabilities can be spread over a wide geographical area. They must dynamically adapt to organizational relationships and real-world uncertainties. Intelligent Networks, such as grids, peer-to-peer, ad hoc networks, constellations, and clouds enable the flexible routing and charging, advanced user interactions and the aggregation and sharing of geographically distributed resources. Collectively owned and managed by distinct organizational bodies, such complex large-scale distributed systems typically encompass computational resources from different institutions, enterprises, and individuals and are governed by heterogeneous administrative policies and regulations. System management techniques must therefore be able to group, predict, and classify different sets of rules, configuration directives, and environmental conditions to impose dissimilar usage policies on various users and resources. They must effectively deal with various optimization criteria, users’ requirements, massive data processing, and, finally, uncertainties in system information that may be incomplete, imprecise, and fragmentary. Next information technology architectures, such as green cloud-to-cloud systems and green mobile clouds, provide elastic and in fact unlimited resources, including storage, as various services to cloud users with possible minimal energy utilization. However, both cloud users and cloud service providers are almost certain to be from different trust domains. Therefore, a secure user-enforced data access control mechanism must be provided before cloud users have the liberty to outsource sensitive data to the cloud for storage and further processing. With the advent of intelligent networks, where efficient interdomain operation and high scalability of the whole system are the most important features, it is arguably required to investigate novel methods and techniques to enable secure access to data and resources, flexible communication, efficient scheduling, self-adaptation, decentralization, and self-organization. This special issue herewith presents six research papers with novel concepts in the analysis, implementation, and evaluation of the next generation of intelligent scalable techniques for data-intensive processing and global optimization problems in large-scale distributed systems. The first three papers discuss novel scalable solutions of data-intensive global optimization problems in well-known large-scale network environments. The presented techniques and their implementations are based on formal mathematical and logical models with the new optimization criteria (energy conservation), semantic rules and ontology, and modern synchronization modules of parallel computational processes. Li et al. in 1 introduced a methodology for improvement of the performance of the dynamic core of Global/Regional Assimilation and Prediction System (GRAPES) – the Numerical Weather Prediction system used by Chinese Meteorology Administration. The system performance is formally modeled as a sequence of large, sparse linear systems formulated by the discretization of global 3D Helmholtz equation. The authors developed a solver that enables an effective synchronization of the numerical processes at the global units of the system. The results of simple empirical analysis show good scalability of the proposed methodology achieved by using up to 6144 active cores in GRAPES. In 2, the authors present a framework for the energy-aware system management in backbone networks. The energy optimization problem is formulated as a general mathematical programming problem with various constraints and control parameters. Dynamic voltage and frequency scaling method is implemented for minimizing the energy utilization at global and local levels of the management system along with a wide range of the resolution methodologies. All possible energy saving decisions of the system units are directly specified, together with decisions concerning traffic assignment to particular links. The results of the experiments show the best performance of the system in the case of concentration of the network traffic on a minimal subset of network components. The problem of massive processing of huge volumes of data in the Internet is discussed in 3. Dong and Hussein propose an ontology-based Web crawler and Web page classifier with an embedded semisupervised learning module. This module enables the continuous enrichment of the definitions of ontological concepts in crawling and Web page classification process. The semantic relevance of crawling topics and Web pages is specified by semantic similarity and probabilistic models. The remaining three papers address the big-data paradigm from various perspective. Bilal et al. 4 benchmark some well-known data center network architectures and categorically state their pros and cons. With this knowledge, the authors propose future advancements pertaining to the network architecture of data centers. In 5, a generic data-structure oriented programming template is discussed for supporting massive remote sensing data. The authors have built the case that the templates provide distributed abstractions for large remote sensing image data with complex data structures. The performance of their technique is improved by developing efficient parallel input/output (I/O) directly to and from the distributed data structures. Zhang et al. 6 have discussed an advanced data center architecture that harness the power of multiple data centers. The key technology that they advocate to manage such a large-scale distributed computing system is to build on both groups of distributed data centers/clusters that are equipped with data center or cluster resource manager. Additional security and access control procedures are put in place to provide a seamless interaction between various domains. In addition to a structure, the domain data centers are organized as collaborative modules, which enables processing of workflow workloads. We believe that all of the papers presented in this Special Issue ought to serve as a reference for students, researchers, and industry practitioners interested or currently working in the evolving and interdisciplinary area of scalable computing and intelligent networking. We hope that the readers will find new inspiration for their research. We are grateful to all the contributors of this issue. We thank the authors for their time and efforts in the presentation of their recent research results. We also would like to express our sincere thanks to the reviewers, who have helped us to ensure the quality of this publication. Our special thanks go to Prof Geoffrey C. Fox (Editor-in-Chief) and all of the editorial and management team of Concurrency and Computation: Practice and Experience Wiley journal for their great support throughout the entire publication process.
Joanna Kolodziej, Samee Ullah Khan, El-Ghazali Talbi
Concurr. Comput. Pract. Exp.2
2013 Hybrid modelling and simulation of huge crowd over a hierarchical Grid architecture
Dan Chen 0001, Lizhe Wang 0001, Jingying Chen 0001, Samee Ullah Khan, Joanna Kolodziej, Mingwei Tian, Fang Huang 0001, Wangyang Liu
Future Gener. Comput. Syst.5
2013 Towards secure mobile cloud computing: A survey
Abdul Nasir Khan, Miss Laiha Mat Kiah, Samee Ullah Khan, Sajjad Ahmad Madani
Future Gener. Comput. Syst.3
2013 Hopfield neural network for simultaneous job scheduling and data replication in grids
Javid Taheri, Albert Y. Zomaya, Pascal Bouvry, Samee Ullah Khan
Future Gener. Comput. Syst.4
2013 Energy-aware parallel task scheduling in a cluster
Lizhe Wang 0001, Samee Ullah Khan, Dan Chen 0001, Joanna Kolodziej, Rajiv Ranjan 0001, Cheng-Zhong Xu 0001, Albert Y. Zomaya
Future Gener. Comput. Syst.2
2013 Lowest priority first based feasibility analysis of real-time systems
Nasro Min-Allah, Samee Ullah Khan, Albert Y. Zomaya
J. Parallel Distributed Comput.2
2013 On minimizing the resource consumption of cloud applications using process migrations
Nikos Tziritas, Samee Ullah Khan, Cheng-Zhong Xu 0001, Thanasis Loukopoulos, Spyros Lalis
J. Parallel Distributed Comput.2
2013 "Security-Aware and Data Intensive Low-Cost Mobile Systems" Editorial
abstract
We are witnessing a paradigm shift in the way mobile devices are being used and operated.What was once a voice network is now predominantly a data network.As a consequence, end-users are now using mobile systems for applications that fall under the data intensive paradigm, such as Skyline queries, streaming information relays, and crowd sourced disaster management.However, this paradigm shift has opened new research directions, such as: (a) Security, as the system now has numerous distributed entry points and the behavior of a malicious entity does not really correlate with any previously known phenomenon (e.g., Internet virus attacks, DOS attacks, etc.).(b) Data interoperability that must cater to the fundamental issue that mobile devices are required to work seamlessly with Internet data, thus requiring revision of protocols, data exchange frameworks to improve data sharing among mobile devices and with the Internet.(c) Sustainable software development that entails the development of software models for mobile devices that have a longer life-cycle and require fewer updates.This also effectively translates into an economically viable mobile system.Privacy and security aspects need to be covered at all layers of mobile networks, from mobile users' devices, to privacy-respecting credentials and mobile identity management.All of the above mentioned research domains are complex on their own, which makes it a very attractive research area for academia and industry.The eventual goal is to make the mobile systems seamless integrate with Inter and Intranet devices without a measurable performance degradation.Is is arguably required to investigate novel methods and techniques to enable secure access to data, network nodes and services, flexible communication, efficient scheduling, self-adaptation, decentralization, and self-organization.This special issue herewith presents six research papers with novel concepts in the analysis, implementation, and evaluation of the next generation of intelligent scalable techniques for data intensive processing and security related problems in modern mobile environments.The first three papers span the fields of key management, power modeling and mobility modeling, but all share a relevance to security aspects.Cryptographic and key management systems in mobile networks must be computationally low-cost because of the limitations of computational and data storage capacities and battery life of most of the network nodes.Wu and Lin present non-interactive authenticated key agreement (NI-AKA) protocols based on the idea of bilinear pairingbased cryptosystem model and Elliptic Curve Encryption (ECE) scheme.The ECE allows the encryption of message multiple times with different keys that can be decrypted in
Joanna Kolodziej, Martin Gilje Jaatun, Samee Ullah Khan, Mario Köppen
Mob. Networks Appl.3
2013 Distributed Online Algorithms for the Agent Migration Problem in WSNs
Nikos Tziritas, Spyros Lalis, Samee Ullah Khan, Thanasis Loukopoulos, Cheng-Zhong Xu 0001, Petros Lampsas
Mob. Networks Appl.3
2013 A survey on resource allocation in high performance distributed computing systems
Hameed Hussain, Saif Ur Rehman Malik, Abdul Hameed, Samee Ullah Khan, Gage Bickler, Nasro Min-Allah, Muhammad Bilal Qureshi, Yongji Wang 0002, Nasir Ghani, Joanna Kolodziej, Albert Y. Zomaya, Cheng-Zhong Xu 0001, Pavan Balaji, Abhinav Vishnu, Frédéric Pinel, Johnatan E. Pecero, Dzmitry Kliazovich, Pascal Bouvry, Hongxiang Li 0001, Lizhe Wang 0001, Dan Chen 0001, Ammar Rayes
Parallel Comput.4
2013 Comparative study of trust and reputation systems for wireless sensor networks
abstract
ABSTRACT Wireless sensor networks (WSNs) are emerging as useful technology for information extraction from the surrounding environment by using numerous small‐sized sensor nodes that are mostly deployed in sensitive, unattended, and (sometimes) hostile territories. Traditional cryptographic approaches are widely used to provide security in WSN. However, because of unattended and insecure deployment, a sensor node may be physically captured by an adversary who may acquire the underlying secret keys, or a subset thereof, to access the critical data and/or other nodes present in the network. Moreover, a node may not properly operate because of insufficient resources or problems in the network link. In recent years, the basic ideas of trust and reputation have been applied to WSNs to monitor the changing behaviors of nodes in a network. Several trust and reputation monitoring (TRM) systems have been proposed, to integrate the concepts of trust in networks as an additional security measure, and various surveys are conducted on the aforementioned system. However, the existing surveys lack a comprehensive discussion on trust application specific to the WSNs. This survey attempts to provide a thorough understanding of trust and reputation as well as their applications in the context of WSNs. The survey discusses the components required to build a TRM and the trust computation phases explained with a study of various security attacks. The study investigates the recent advances in TRMs and includes a concise comparison of various TRMs. Finally, a discussion on open issues and challenges in the implementation of trust‐based systems is also presented. Copyright © 2012 John Wiley & Sons, Ltd.
Osman Khalid, Samee Ullah Khan, Sajjad Ahmad Madani, Khizar Hayat 0002, Majid Iqbal Khan, Nasro Min-Allah, Joanna Kolodziej, Lizhe Wang 0001, Sherali Zeadally, Dan Chen 0001
Secur. Commun. Networks2
2013 On the Characterization of the Structural Robustness of Data Center Networks
abstract
Data centers being an architectural and functional block of cloud computing are integral to the Information and Communication Technology (ICT) sector. Cloud computing is rigorously utilized by various domains, such as agriculture, nuclear science, smart grids, healthcare, and search engines for research, data storage, and analysis. A Data Center Network (DCN) constitutes the communicational backbone of a data center, ascertaining the performance boundaries for cloud infrastructure. The DCN needs to be robust to failures and uncertainties to deliver the required Quality of Service (QoS) level and satisfy Service Level Agreement (SLA). In this paper, we analyze robustness of the state-of-the-art DCNs. Our major contributions are: (a) we present multi-layered graph modeling of various DCNs; (b) we study the classical robustness metrics considering various failure scenarios to perform a comparative analysis; (c) we present the inadequacy of the classical network robustness metrics to appropriately evaluate the DCN robustness; and (d) we propose new procedures to quantify the DCN robustness. Currently, there is no detailed study available centering the DCN robustness. Therefore, we believe that this study will lay a firm foundation for the future DCN robustness research.
Kashif Bilal, Marc Manzano, Samee Ullah Khan, Eusebi Calle, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Cloud Comput.3
2013 Modeling and Analysis of State-of-the-art VM-based Cloud Management Platforms
abstract
Virtualization is a key aspect to achieve scalability and flexibility in a cloud. Many solutions have been proposed to monitor and deploy Virtual Machines (VM) in resource pool of cloud. However, most of the cloud management systems, such as Amazon EC2 are proprietary. In the said perspective, many open source VM-based platforms have tossed for general users to research. The existing work has mainly focused on the discussion of architecture, feature-set, and performance analysis. Other important aspects, such as formal analysis, modeling, and verification are usually ignored. In this paper, we provide formal analysis, modeling, and verification of three open source state-of-the-art VM-based cloud platforms: (a) Eucalyptus, (b) Open Nebula, and (c) Nimbus. We used High-Level Petri Nets (HLPN) to model and analyze the structural and behavioral properties of the systems. Moreover, to verify the models, we have used Satisfiability Modulo Theories Library (SMT-Lib) and Z3 Solver. We modeled about 100 VM to verify the correctness and feasibility of our models. The results reveal that the models are functioning correctly. Moreover, the increase in the number of VM does not affect the working of the models that indicates the practicability of the models in a highly scalable and flexible environment.
Saif Ur Rehman Malik, Samee Ullah Khan, Sudarshan K. Srinivasan
IEEE Trans. Cloud Comput.2
2013 Solving symbolic regression problems with uniform design-aided gene expression programming
Yunliang Chen 0002, Dan Chen 0001, Samee Ullah Khan, Jianzhong Huang 0001, Changsheng Xie 0001
J. Supercomput.3
2013 Green computing and communications
Samee Ullah Khan, Lizhe Wang 0001, Laurence T. Yang, Feng Xia 0001
J. Supercomput.1
2013 Review of performance metrics for green data centers: a taxonomy study
Lizhe Wang 0001, Samee Ullah Khan
J. Supercomput.2
2013 A Fully Distributed Scheme for Discovery of Semantic Relationships
abstract
The availability of large volumes of Semantic Web data has created the potential of discovering vast amounts of knowledge. Semantic relation discovery is a fundamental technology in analytical domains, such as business intelligence and homeland security. Because of the decentralized and distributed nature of Semantic Web development, semantic data tend to be created and stored independently in different organizations. Under such circumstances, discovering semantic relations faces numerous challenges, such as isolation, scalability, and heterogeneity. This paper proposes an effective strategy to discover semantic relationships over large-scale distributed networks based on a novel hierarchical knowledge abstraction and an efficient discovery protocol. The approach will effectively facilitate the realization of the full potential of harnessing the collective power and utilization of the knowledge scattered over the Internet.
Juan Li 0004, Wendy Hui Wang, Samee Ullah Khan, Qingrui Li, Albert Y. Zomaya
IEEE Trans. Serv. Comput.3
2013 Secure Wireless Multicast for Delay-Sensitive Data via Network Coding
abstract
Wireless multicast for delay-sensitive data is challenging because of the heterogeneity effect where each receiver may experience different packet losses. Fortunately, network coding, a new advanced routing protocol, offers significant advantages over the traditional Automatic Repeat reQuest (ARQ) protocols in that it mitigates the need for retransmission and has the potential to approach the min-cut capacity. Network-coded multicast would be, however, vulnerable to false packet injection attacks, in which the adversary injects bogus packets to prevent receivers from correctly decoding the original data. Without a right defense in place, even a single bogus packet can completely change the decoding outcome. Existing solutions either incur high computation cost or cannot withstand high packet loss. In this paper, we propose a novel scheme to defend against false packet injection attacks on network-coded multicast for delay-sensitive data. Specifically, we propose an efficient authentication mechanism based on null space properties of coded packets, aiming to enable receivers to detect any bogus packets with high probability. We further design an adaptive scheduling algorithm based on the Markov Decision Processes (MDP) to maximize the number of authenticated packets received within a given time constraint. Both analytical and simulation results have been provided to demonstrate the efficacy and efficiency of our proposed scheme.
Thuan T. Tran, Hongxiang Li 0001, Guanying Ru, Robert J. Kerczewski, Lingjia Liu 0001, Samee Ullah Khan
IEEE Trans. Wirel. Commun.6
2012 A Comparative Study Of Data Center Network Architectures
abstract
Data Centers (DCs) are experiencing a tremendous growth in the number of hosted servers. Aggregate bandwidth requirement is a major bottleneck to data center performance. New Data Center Network (DCN) architectures are proposed to handle different challenges faced by current DCN architecture. In this paper we have implemented and simulated two promising DCN architectural models, namely switch-based and hybrid models, and compared their effectiveness by monitoring the network throughputs and average packet latencies. The presented analysis may be a background for the further studies on the simulation and implementation of the DCN customized topologies, and customized addressing protocols in the large-scale data centers.
Kashif Bilal, Samee Ullah Khan, Joanna Kolodziej, Khizar Hayat 0002, Sajjad Ahmad Madani, Nasro Min-Allah, Lizhe Wang 0001, Dan Chen 0001
ECMS2
2012 A Checkpoint Based Message Forwarding Approach For Opportunistic Communication
abstract
In a Delay Tolerant Network (DTN), the nodes have intermittent connectivity and complete path(s) between the source and destination may not exist. The communication takes place opportunistically when any two nodes enter the effective range. One of the major challenges in DTNs is message forwarding when a sender must select a best neighbor that has the highest probability of forwarding the message to the actual destination. However, finding an appropriate route remains an NP-hard problem. This paper presents a concept of Checkpoint (CP) based message forwarding in DTNs. The CPs are autonomous high-end wireless devices with large buffer storage and are responsible for temporarily storing the messages to be forwarded. The CPs are deployed at various places within the city parameter that are covered by bus routes and where human meeting frequencies are higher. For the simulative analysis a synthetic human mobility model in ONE simulator is constructed for the city of Fargo, ND, USA. The model is tested over various DTN routing protocols and the results indicate that using CP overlay over the existing DTN architecture significantly decreases message delivery time as well as buffer usage.
Osman Khalid, Samee Ullah Khan, Joanna Kolodziej, Juan Li 0004, Khizar Hayat 0002, Sajjad Ahmad Madani, Lizhe Wang 0001, Dan Chen 0001
ECMS2
2012 The Median Resource Failure Checkpointing
abstract
In grid computing, the realization of an enviable fault tolerance ability is linked with the proper utilization of resources and scheduling of jobs. The literature offers two solutions to these two challenging tasks, viz. checkpointing and replication. A checkpointing strategy is being proposed that uses the median of failure intervals of the resources in deciding the checkpoint intervals for the given jobs. The strategy shows improved system throughput, job losses and job execution times while eliminating unnecessary checkpoints.
Suleman Khan 0001, Khizar Hayat 0002, Sajjad Ahmad Madani, Samee Ullah Khan, Joanna Kolodziej
ECMS4
2012 Adaptive scheduling for multicasting hard deadline constrained prioritized data via network coding
abstract
Network coding offers a promising platform for multicast transmission by approaching its min-cut capacity. However, pushing the network throughput toward this upper bound comes with a sacrifice in delivery delay due to the decoding procedure that requires performing batch of coded packets. Further, in some transmission scenarios where the receivers experience deep fading or unable to collect a full set of the transmitted data, no useful information is recovered. The effect is more severe in the networks where the transmitted information has priority structure with hard deadline constraint due to the limited delivery time and data interdependencies. In this paper, we consider single-hop wireless networks where the transmitter wishes to multicast hard deadline constrained prioritized data to many receivers over lossy channels. We first study the network performance of a variety of transmission techniques, depending on how the transmitter schedules transmission in each time slot. We then propose an adaptive encoding and scheduling technique to maximize the network throughput. To find the optimal transmission scheduling at the presence of the network dynamics, we cast the problem in the framework of Markov Decision Processes (MDP) and use backward induction method to find an optimal solution. We further propose simulation-based algorithm and greedy scheduling technique that obtain high performance with much lower time complexity. Both analytical and simulation results have been provided to corroborate the effectiveness of the proposed techniques.
Thuan T. Tran, Hongxiang Li 0001, Weiyao Lin, Lingjia Liu 0001, Samee Ullah Khan
GLOBECOM5
2012 Secure network-coded wireless multicast for delay-sensitive data
abstract
Wireless multicast for delay-sensitive data is challenging because different receivers may experience different packet losses. Network coding offers significant advantages over the traditional Automatic Repeat-reQuest (ARQ) protocols in that it mitigates the need for retransmission and has the potential to approach the min-cut capacity. Network-coded multicast would be, however, vulnerable to false packet injection attacks, in which the adversary injects bogus packets to prevent receivers from correctly decoding the original data. Without a right defense in place, even a single bogus packet can completely change the decoding outcome. Existing solutions either incur high computation cost or cannot withstand high packet loss. In this paper, we propose a novel scheme to defend against false packet injection attacks on network-coded multicast for delay-sensitive data. Specifically, we propose an efficient authentication mechanism based on null space properties of coded packets, aiming to enable receivers to detect any bogus packets with high probability. We further design an adaptive scheduling algorithm based on Markov Decision Processes (MDP) to maximize the number of authenticated packets that can be received within a given time constraint. Both analytical and simulation results have been provided to demonstrate the efficacy and efficiency of our proposed scheme.
Thuan T. Tran, Hongxiang Li 0001, Lingjia Liu 0001, Samee Ullah Khan
ICC4
2012 Diverse routing in multi-domain optical networks with correlated and probabilistic multi-failures
abstract
Multi-failure network survivability is a key concern for operators owing to the recent spate of natural and man-made disasters. Hence this paper presents a new diverse lightpath protection scheme for random correlated failure recovery in multi-domain optical networks. The solution leverages topology abstraction and hierarchical inter-domain routing and assumes a-priori link failure probability state. The path computation process also takes into account both traffic engineering and risk minimization objectives. The proposed scheme is analyzed using network simulation for a variety of multi-failure scenarios.
Nasro Min-Allah, Samee Ullah Khan, Nasir Ghani
ICC3
2012 An Optimal Fully Distributed Algorithm to Minimize the Resource Consumption of Cloud Applications
abstract
According to the pay-per-use model adopted in clouds, the more the resources consumed by an application running in a cloud computing environment, the greater the amount of money the owner of the corresponding application will be charged. Therefore, applying intelligent solutions to minimize the resource consumption is of great importance. Because centralized solutions are deemed unsuitable for large-distributed systems or large-scale applications, we propose a fully distributed algorithm (called DRA) to overcome the scalability issues. The aforementioned problem can be solved by identifying an assignment scheme between the interacting components of an application, such as processes and virtual machines, and the computing nodes of a cloud system, such that the total amount of resources consumed by the respective application is minimized. The decisions for the transition from one assignment scheme to another one are made in a dynamic way and based only on local information. It should be stressed that DRA achieves convergence and always results in the optimal solution. We also show, through an experimental evaluation, that DRA achieves up to 55% network cost reduction when compared to the most recent algorithm in the literature.
Nikos Tziritas, Samee Ullah Khan, Cheng-Zhong Xu 0001, Jue Hong
ICPADS2
2012 Parallel Processing of Massive EEG Data with MapReduce
abstract
Analysis of neural signals like electroencephalogram (EEG) is one of the key technologies in detecting and diagnosing various brain disorders. As neural signals are non-stationary and non-linear in nature, it is almost impossible to understand their true physical dynamics until the recent advent of the Ensemble Empirical Mode Decomposition (EEMD) algorithm. The neural signal processing with EEMD is highly compute-intensive due to the high complexity of the EEMD algorithm. It is also data intensive because 1) EEG signals contain massive data sets 2) EEMD has to introduce a large number of trials in processing to ensure precision. The Map Reduce programming mode is a promising parallel computing paradigm for data intensive computing. To increase the efficiency and performance of the neural signal analysis, this research develops parallel EEMD neural signal processing with Map Reduce. In this paper, we implement the parallel EEMD with Hadoop in a modern cyber infrastructure. Test results and performance evaluation show that parallel EEMD can significantly improve the performance of neural signal processing.
Lizhe Wang 0001, Dan Chen 0001, Rajiv Ranjan 0001, Samee Ullah Khan, Joanna Kolodziej, Jun Wang 0001
ICPADS4
2012 Introducing Agent Evictions to Improve Application Placement in Wireless Distributed Systems
abstract
With the development of mobile code frameworks for embedded systems, an application can be structured as a set of cooperating components (agents) that are placed on the nodes of the system in a flexible way. Reducing the network traffic caused by the application components is a crucial issue for the increase in the lifetime of wireless embedded systems, since it is widely known that the communication cost plays the most significant role in the energy consumption of embedded devices. To this end, most placement algorithms place or move an agent towards the center of gravity of the communication workload. However, if the target node does not have enough capacity, the attempt is usually aborted. In this paper, we introduce eviction-enabled algorithms that allow nodes to free capacity by forcing a locally hosted agent to move to another node, even at a loss, to accept a new and potentially more beneficial agent. To the best of our knowledge, this is the first time that agents are evicted by local hosts to enable beneficial agent migrations and eventually improve the total network cost. In this paper, we provide algorithms tackling the aforementioned problem in a fully distributed manner. We also present and discuss the results of extensive simulations, showing that eviction-enabled algorithms can outperform their counterparts by up to 300%.
Nikos Tziritas, Petros Lampsas, Spyros Lalis, Thanasis Loukopoulos, Samee Ullah Khan, Cheng-Zhong Xu 0001
ICPP5
2012 Multi-level hierarchic genetic-based scheduling of independent jobs in dynamic heterogeneous grid environment
Joanna Kolodziej, Samee Ullah Khan
Inf. Sci.2
2012 Power efficient rate monotonic scheduling for multi-core systems
Nasro Min-Allah, Hameed Hussain, Samee Ullah Khan, Albert Y. Zomaya
J. Parallel Distributed Comput.3
2012 A New Cooperative Spectrum Sensing Scheme for Cognitive Ad-Hoc Networks
Hongxiang Li 0001, Weiyao Lin, Lingjia Liu 0001, Samee Ullah Khan, Sentang Wu
Mob. Networks Appl.6
2012 A Semantics-based Approach to Large-Scale Mobile Social Networking
Juan Li 0004, Wendy Hui Wang, Samee Ullah Khan
Mob. Networks Appl.3
2012 Energy-efficient high-performance parallel and distributed computing
Samee Ullah Khan, Pascal Bouvry, Thomas Engel 0001
J. Supercomput.1
2012 A goal programming based energy efficient resource allocation in data centers
Samee Ullah Khan, Nasro Min-Allah
J. Supercomput.1
2012 Green networks
Samee Ullah Khan, Sherali Zeadally, Pascal Bouvry, Naveen K. Chilamkurti
J. Supercomput.1
2012 GreenCloud: a packet-level simulator of energy-aware cloud computing data centers
Dzmitry Kliazovich, Pascal Bouvry, Samee Ullah Khan
J. Supercomput.3
2012 Comparison and analysis of eight scheduling heuristics for the optimization of energy consumption and makespan in large-scale distributed systems
Peder Lindberg, James Leingang, Daniel Lysaker, Samee Ullah Khan, Juan Li 0004
J. Supercomput.4
2012 A comparative study of rate monotonic schedulability tests
Nasro Min-Allah, Samee Ullah Khan, Nasir Ghani, Juan Li 0004, Lizhe Wang 0001, Pascal Bouvry
J. Supercomput.2
2012 Optimal task execution times for periodic tasks using nonlinear constrained optimization
Nasro Min-Allah, Samee Ullah Khan, Yongji Wang 0002
J. Supercomput.2
2012 Thermal aware workload placement with task-temperature profiles in a data center
Lizhe Wang 0001, Samee Ullah Khan, Jai Dayal
J. Supercomput.2
2012 Energy-efficient networking: past, present, and future
Sherali Zeadally, Samee Ullah Khan, Naveen K. Chilamkurti
J. Supercomput.2
2011 A Multi-objective GRASP Algorithm for Joint Optimization of Energy Consumption and Schedule Length of Precedence-Constrained Applications
abstract
We address the problem of scheduling precedence-constrained scientific applications on a heterogeneous distributed processor system with the twin objectives of minimizing simultaneously energy consumption and schedule length. Previous research efforts on scheduling have focused on the minimization of a quality of service metric based on the completion time of applications (e.g., the schedule length). Recently, many researchers are working on the design of new scheduling algorithms that consider the minimization of energy consumption. We report a new scheduling algorithm accounting for both objectives. The new scheduling algorithm is based on a multi-start randomized adaptive search technique (GRASP framework) that adopts Dynamic Voltage Scaling technique to minimize energy consumption. This technique enables processors to operate in different voltage supply levels at the cost of sacrificing clock frequencies. This multiple voltage implies a trade-off between the quality of the schedules and energy consumption. Therefore, the new proposed approach is designed as a multi-objective algorithm that simultaneously optimize both objectives. Simulation results on a set of real-world applications emphasize the robust performance of the proposed approach.
Johnatan E. Pecero, Pascal Bouvry, Héctor J. Fraire H., Samee Ullah Khan
DASC4
2011 Energy-Aware High Performance Computing: A Taxonomy Study
abstract
To reduce the energy consumption and build a sustainable computer infrastructure now becomes a major goal of the high performance community. A number of research projects have been carried out in the field of energy-aware high performance computing. This paper is devoted to categorize energy-aware computing methods for the high-end computing infrastructures, such as servers, clusters, data centers, and Grids/Clouds. Based on a taxonomy of methods and system scales, this paper reviews the current status of energy-aware HPC research and summarizes open questions and research directions of software architecture for future energy-aware HPC studies.
Lizhe Wang 0001, Samee Ullah Khan, Jie Tao 0001
ICPADS3
2011 Mosaic-Net: a game theoretical method for selection and allocation of replicas in ad hoc networks
Samee Ullah Khan
J. Supercomput.1
2010 GreenCloud: A Packet-Level Simulator of Energy-Aware Cloud Computing Data Centers
abstract
Cloud computing data centers are becoming increasingly popular for the provisioning of computing resources. The cost and operating expenses of data centers have skyrocketed with the increase in computing capacity. Several governmental, industrial, and academic surveys indicate that the energy utilized by computing and communication units within a data center contributes to a considerable slice of the data center operational costs. In this paper, we present a simulation environment for energy-aware cloud computing data centers. Along with the workload distribution, the simulator is designed to capture details of the energy consumed by data center components (servers, switches, and links) as well as packet-level communication patterns in realistic setups. The simulation results obtained for two-tier, three- tier, and three-tier high-speed data center architectures demonstrate the effectiveness of the simulator in utilizing power management schema, such as voltage scaling, frequency scaling, and dynamic shutdown that are applied to the computing and networking components.
Dzmitry Kliazovich, Pascal Bouvry, Yury Audzevich, Samee Ullah Khan
GLOBECOM4
2010 Replicating data objects in large distributed database systems: an axiomatic game theoretic mechanism design approach
Samee Ullah Khan, Ishfaq Ahmad 0001
Distributed Parallel Databases1
2009 MobiSN: Semantics-Based Mobile Ad Hoc Social Network Framework
abstract
Mobile ad hoc social networks are self-configuring social networks that connect users using mobile devices, such as laptops, PDAs, and cellular phones. These social networks facilitate users to form virtual communities of similar interests or commonalities. This paper proposes a concrete, generalized, and novel framework to develop a fully functional mobile ad hoc social network. The proposed framework provides effective and efficient solutions to social network construction, semantics-based user profile matching, and multi-hop semantics-based routing. Moreover, the proposed framework because of its generality is applicable in applications of critical importance, such as disaster-recovery, homeland security, and personnel control. Furthermore, the proposed framework is rigorously benchmarked using an elaborate simulation setup and released as a prototype system that can be run on cellular phones.
Juan Li 0004, Samee Ullah Khan
GLOBECOM2
2009 A goal programming approach for the joint optimization of energy consumption and response time in computational grids
abstract
We study the multi-objective problem of mapping independent tasks onto a set of computational grid machines that simultaneously minimizes the energy consumption and response time (makespan) subject to the constraints of deadlines and architectural requirements. We propose an algorithm based on goal programming that effectively converges to the compromised Pareto optimal solution. Compared to other traditional multi-objective optimization techniques that require identification of the Pareto frontier, goal programming directly converges to the compromised solution. Such a property makes goal programming a very efficient multi-objective optimization technique. Moreover, simulation results show that the proposed technique achieves superior performance compared to the greedy and linear relaxation heuristics, and competitive performance relative to the optimal solution implemented in LINDO for small-scale problems.
Samee Ullah Khan
IPCCC1
2009 Robust CDN replica placement techniques
abstract
Creating replicas of frequently accessed data objects across a read-intensive content delivery network (CDN) can result in reduced user response time. Because CDNs often operate under volatile conditions, it is of the utmost importance to study replica placement techniques that can cope with uncertainties in the system parameters. We propose four CDN replica placement heuristics that guarantee a robust performance under the uncertainty of arbitrary CDN server failures. By robust performance we mean the solution quality that a heuristic guarantees given the uncertainties in system parameters. The simulation results reveal interesting characteristics of the studied heuristics. We report these characteristics with a detailed discussion on which heuristics to utilize for robust CDN data replication given a specific scenario.
Samee Ullah Khan, Anthony A. Maciejewski, Howard Jay Siegel
IPDPS1
2009 A Pure Nash Equilibrium-Based Game Theoretical Method for Data Replication across Multiple Servers
abstract
This paper proposes a non-cooperative game based technique to replicate data objects across a distributed system of multiple servers in order to reduce user perceived Web access delays. In the proposed technique computational agents represent servers and compete with each other to optimize the performance of their servers. The optimality of a non-cooperative game is typically described by Nash equilibrium, which is based on spontaneous and non-deterministic strategies. However, Nash equilibrium may or may not guarantee system-wide performance. Furthermore, there can be multiple Nash equilibria, making it difficult to decide which one is the best. In contrast, the proposed technique uses the notion of pure Nash equilibrium, which if achieved, guarantees stable optimal performance. In the proposed technique, agents use deterministic strategies that work in conjunction with their self-interested nature but ensure system-wide performance enhancement. In general, the existence of a pure Nash equilibrium is hard to achieve, but we prove the existence of such equilibrium in the proposed technique. The proposed technique is also experimentally compared against some well-known conventional replica allocation methods, such as branch and bound, greedy, and genetic algorithms.
Samee Ullah Khan, Ishfaq Ahmad 0001
IEEE Trans. Knowl. Data Eng.1
2009 A Cooperative Game Theoretical Technique for Joint Optimization of Energy Consumption and Response Time in Computational Grids
abstract
With the explosive growth in computers and the growing scarcity in electric supply, reduction of energy consumption in large-scale computing systems has become a research issue of paramount importance. In this paper, we study the problem of allocation of tasks onto a computational grid, with the aim to simultaneously minimize the energy consumption and the makespan subject to the constraints of deadlines and tasks' architectural requirements. We propose a solution from cooperative game theory based on the concept of Nash bargaining solution. In this cooperative game, machines collectively arrive at a decision that describes the task allocation that is collectively best for the system, ensuring that the allocations are both energy and makespan optimized. Through rigorous mathematical proofs we show that the proposed cooperative game in mere O(n mlog(m)) time (where n is the number of tasks and m is the number of machines in the system) produces a Nash bargaining solution that guarantees Pareto-optimally. The simulation results show that the proposed technique achieves superior performance compared to the greedy and linear relaxation (LR) heuristics, and with competitive performance relative to the optimal solution implemented in LINDO for small-scale problems.
Samee Ullah Khan, Ishfaq Ahmad 0001
IEEE Trans. Parallel Distributed Syst.1
2008 Using game theory for scheduling tasks on multi-core processors for simultaneous optimization of performance and energy
abstract
Multi-core processors are beginning to revolutionize the landscape of high-performance computing. In this paper, we address the problem of power-aware scheduling/mapping of tasks onto heterogeneous and homogeneous multi-core processor architectures. The objective of scheduling is to minimize the energy consumption as well as the makespan of computationally intensive problems. The multi-objective optimization problem is not properly handled by conventional approaches that try to maximize a single objective. Our proposed solution is based on game theory. We formulate the problem as a cooperate game. Although we can guarantee the existence of a Bargaining Point in this problem, the classical cooperative game theoretical techniques such as the Nash axiomatic technique cannot be used to identify the Bargaining Point due to low convergence rates and high complexity. Hence, we transform the problem to a max-max-min problem such that it can generate solutions with fast turnaround time.
Ishfaq Ahmad 0001, Sanjay Ranka, Samee Ullah Khan
IPDPS3
2008 A game theoretical data replication technique for mobile ad hoc networks
abstract
Adaptive replication of data items on servers of a mobile ad hoc network can alleviate access delays. The selection of data items and servers requires solving a constrained optimization problem, that is in general NP-complete. The problem is further complicated by frequent partitions of the ad hoc network. In this paper, a mathematical model for data replication in ad hoc networks is formulated. We treat the mobile servers in the ad hoc network as self-interested entities, hence they have the capability to manipulate the outcome of a resource allocation mechanism by misrepresenting their valuations. We design a game theoretic "truthful" mechanism in which replicas are allocated to mobile servers based on reported valuations. We sketch the exact properties of the truthful mechanism and derive a payment scheme that suppresses the selfish behavior of the mobile servers. The proposed technique is extensively evaluated against three ad hoc network replica allocation methods: (a) extended static access frequency, (b) extended dynamic access frequency and neighborhood, and (c) extended dynamic connectivity grouping. The experimental results reveal that the proposed approach outperforms the three techniques in solution quality and has competitive execution times.
Samee Ullah Khan, Anthony A. Maciejewski, Howard Jay Siegel, Ishfaq Ahmad 0001
IPDPS1
2008 Comparison and analysis of ten static heuristics-based Internet data replication techniques
Samee Ullah Khan, Ishfaq Ahmad 0001
J. Parallel Distributed Comput.1
2007 A cooperative game theoretical replica placement technique
abstract
Creating replicas of frequently accessed data objects across a read intensive network can lead to reduced communication cost and end-user response time. On the contrary, data replication in the presence of writes incurs extra cost due to multiple updates. The selection of data objects and servers requires solving a constraint optimization problem, which is NP-complete in general. A majority of the current state-of-the-art replica placement techniques suffer from high computational complexity issues. To circumvent such issues, we propose a cooperative game theoretical technique in which the servers (players) in the system collectively deliberate to converge at a replica schema that is beneficial to the system as a whole. In particular we make use of the Aumann-Shapley mechanism of cooperative game theory to propose an effective replica placement technique that yields good solutions when the system is very heavily loaded. Experimental comparisons are made against: (1) branch and bound, (2) greedy, (3) genetic, (4) Dutch auction, and (5) English auction. As demonstrated by the experimental results, the proposed technique maintains superior solution quality in terms of lower communication cost and reduced execution time.
Samee Ullah Khan, Ishfaq Ahmad 0001
ICPADS1
2007 A Semi-Distributed Axiomatic Game Theoretical Mechanism for Replicating Data Objects in Large Distributed Computing Systems
abstract
Replicating data objects onto servers across a system can alleviate access delays. The selection of data objects and servers requires solving a constraint optimization problem, which is NP-complete in general. A majority of conventional replica placement techniques falter on issues of scalability or solution quality. To counteract such issues, we propose a game theoretical replica placement technique, in which computational agents compete for the allocation or reallocation of replicas onto their servers in order to reduce the user perceived access delays. The technique is based upon six well-defined axioms, each guaranteeing certain basic game theoretical properties. This eccentric method of designing game theoretical techniques using axioms is unique in the literature and takes away from the designers the cumbersome mathematical details of game theory. The distinctive feature of these axioms is that when amassed together, their individual properties constrict into one system-wide performance enhancement property, which in our case is the reduction of access time. The control of the proposed technique is "semi-distributed" in nature, wherein all the heavy processing is done on the servers of the distributed system and the central body is only required to take a binary decision: (0) not to replicate or (1) to replicate. This semi-distributed approach makes the technique scalable and helps solutions to converge in a fast turn-around time without loosing much of the solution quality. Experimental comparisons are made against: 1) branch and bound, 2) greedy, 3) genetic, 4) Dutch auction, and 5) English auction. As attested by the results, the proposed technique maintains superior solution quality in terms of lower communication cost and reduced execution time.
Samee Ullah Khan, Ishfaq Ahmad 0001
IPDPS1
2007 Autonomic Power & Performance Management for Large-Scale Data Centers
abstract
With the rapid growth of servers and applications spurred by the Internet, the power consumption of servers has become critically important and must be efficiently managed. High energy consumption also translates into excessive heat dissipation which in turn, increases cooling costs and causes servers to become more prone to failure. This paper presents a theoretical and experimental framework and general methodology for hierarchical autonomic power & performance management in high performance distributed data centers. We optimize for power & performance (performance/watt) at each level of the hierarchy while maintaining scalability. We adopt mathematically-rigorous optimization approach to provide the application with the required amount of memory at runtime. This enables us to transition the unused memory capacity to a low power state. Our experimental results show a maximum performance/watt improvement of 88.48% compared to traditional techniques. We also present preliminary results of using game theory to optimize performance/watt at the cluster level of a data center. Our cooperative technique reduces the power consumption by 65% when compared to traditional techniques (min-min heuristic).
Bithika Khargharia, Salim Hariri, Ferenc Szidarovszky, Manal Houri, Hesham El-Rewini, Samee Ullah Khan, Ishfaq Ahmad 0001, Mazin S. Yousif
IPDPS6
2007 Approximate Optimal Sensor Placements in Grid Sensor Fields
abstract
This paper proposes a simple heuristic to effectively and efficiently place sensors in grid sensor fields under the constraint of complete coverage. The heuristic guarantees a solution of min(1 + alpha, 3)-optimal when an additional constraint of prioritized placement is enforced, where alpha is the maximum ratio between the weights (priorities) of the grid points. When there is no prioritized placement, a solution of 2-optimal is guaranteed. We also show that these bounds are the best possible unless P = NP. Comparisons are performed against some well known sensor placement techniques, where the proposed heuristic outperforms in solution quality and execution time.
Samee Ullah Khan
VTC Spring1
2006 Non-cooperative, semi-cooperative, and cooperative games-based grid resource allocation
abstract
In this paper we consider, compare and analyze three game theoretical grid resource allocation mechanisms. Namely, 1) the non-cooperative sealed-bid method where tasks are auctioned off to the highest bidder, 2) the semi-cooperative n-round sealed-bid method in which each site delegate its work to others if it cannot perform the work itself, and 3) the cooperative method in which all of the sites deliberate with one another to execute all the tasks as efficiently as possible. To experimentally evaluate the above mentioned techniques, we perform extensive simulation studies that effectively encapsulate the task and machine heterogeneity. The tasks are assumed to be independent and bear multiple execution time deadlines. The simulation model is built around a hierarchical grid infrastructure where machines are abstracted into larger computing centers labeled "federations", each of which are responsible for managing their own resources independently. These federations are then linked together with a primary portal to which grid tasks would be submitted. To measure the effectiveness of these game theoretical techniques, the recorded performance is evaluated against a conventional baseline method in which tasks are randomly assigned to the sites without any task execution guarantee.
Samee Ullah Khan, Ishfaq Ahmad 0001
IPDPS1