Rajiv Ranjan 0001

dblp:68/163-1 · DBLP profile ↗
← Back
207ranked-venue papers
27as first author
71since 2021 · last 2026
0000-0002-6610-1328ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 101 · 16 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 1 first-author · 12 since 2021Software engineering, systems software and programming languages · 23 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 19 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 17 · 15 since 2021Computer networks · 9 · 5 since 2021Security and privacy · 7 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Theory of computation · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards General Trace Theory
Ryszard Janicki, Maciej Koutny, Lukasz Mikulski, Rajiv Ranjan 0001
PETRI NETS4
2026 Beyond the Cloud: Extending Cloud Computing to the Extreme Edge for Autonomous, Trustworthy, and Real-Time AI
Rajiv Ranjan 0001
CLOSER1
2026 TAERM: Traffic accident emergency response management framework for detection and classification using IoT and YOLOv9
Ayman Noor, Hanan Almukhalfi, Talal H. Noor, Rajiv Ranjan 0001
Future Gener. Comput. Syst.4
2026 Pluggable AI-based real-time stragglers detection framework in Hadoop
abstract
The growing reliance on big data frameworks such as Hadoop has revolutionised data processing across various domains, enabling large-scale storage and distributed computation. Hadoop is widely employed in real-world applications such as high-performance computation tasks, e-commerce and data analysis in healthcare. However, the efficiency of Hadoop systems is often hampered by faults and anomalies, with stragglers emerging as one of the most prevalent issues. Stragglers disrupt workflows, waste resources and degrade system performance. While existing anomaly detection models employ methods like median analysis or static thresholds, they often struggle with issues such as high false positives, lack of adaptability and poor handling of complex heterogeneous environments. To address these challenges, this paper presents Plabs , a flexible stragglers detection framework for Hadoop. The framework comprises two core components: (1) a Monitoring Module providing real-time tracking of cluster resources and task progress and (2) a Pluggable AI-based straggler detection module, designed for precise straggler task identification. By leveraging advanced monitoring and AI-driven analysis, Plabs offers an automated, flexible and scalable solution for detecting stragglers at run-time in Hadoop clusters. We evaluated Plabs exhaustively with three Machine Learning (ML), two Deep Learning (DL) and two Large Language Models (LLMs) on five different applications in a real testbed environment. Our experiment evaluation shows that DL models outperform others in identifying Hadoop stragglers, achieving superior accuracy and reliability for all the applications.
Yinhao Li 0003, Rajiv Ranjan 0001, Devki Nandan Jha
High Confid. Comput.3
2026 Patch Matter: Dual Modality Patch Contrastive for Non-Stationary Radio Signals
abstract
The emergence of abundant non-stationary radio signal (NSRS) data presents significant opportunities for applications in wireless communications, radar systems, remote sensing, and healthcare. While deep learning models have shown promise in capturing sequence dependencies, deriving generic and fine-grained representations of NSRS data remains challenging due to its complex, dynamic nature and the scarcity of labeled data. The NSRS data are often frequency-sensitive and exhibit minuscule inter-class distances, posing significant challenges for precise classification. To address these issues, we propose a novelDualModalityPatchContrastive (DMPC) framework. This framework leverages a stochastic patching paradigm for diverse local pattern extraction and a time-frequency cross-view optimization for frequency-sensitive feature mining. Furthermore, an Attentive Patch Aggregation (APA) mechanism enhances fine-grained inference under few-shot conditions through patch-level feature voting. Extensive experiments demonstrate the effectiveness of our approach in addressing the unique challenges of NSRS data.
Jie Su 0001, Yuheng Ye, Zhenyu Wen, Taotao Li, Shibo He, Xiaoqin Zhang 0002, Rajiv Ranjan 0001
IEEE Trans. Mob. Comput.8
2025 Service-Oriented Evolution of Modern AI: A Position Paper
abstract
[Context]: It is well known that understanding the evolution of technologies and its cause is essential for more discoveries and innovations. In the Artificial Intelligence (AI) domain, it has also been identified that scrutinising the development context and path of AI will be able to help both academia and industry better understand the current AI limitations, reveal future AI trends, and facilitate AI/digital transformations. [Objectives]: Given the dramatic boom of modern AI, this research aims to unearth the evolution pattern along the recent three AI waves (namely predictive AI, generative AI and agentic AI), and accordingly to guide AI research and development to focus on the most promising directions. [Method]: We employed analogical reasoning as the research method and referred to the existing software architectural styles to inspire our understanding of the architectural evolution of modern AI technologies. [Results]: We see a service-oriented trend in modern AI's working mechanisms, and the offering of AI power seems to be transiting from a heavyweight and monolithic paradigm to an organisational and collaborative paradigm with more and more specific separation of concerns. Following this service-oriented evolution trend, we borrow software architecture lessons and foresee opportunities to grow the current AI wave to a further height, e.g., standardising AI agent-friendly APIs and developing serverless AI agents. [Conclusions]: What is happening in the AI domain has happened before in the software engineering domain. It is worth reusing software architecture knowledge to evolve the architecture of AI technologies.
Zheng Li 0001, Christopher McKie, Hui Wang 0001, Hamza Shakeel, Rajiv Ranjan 0001
SSE5
2025 LAGD: Local Topological-Alignment and Global Semantic-Deconstruction for Incremental 3D Semantic Segmentation
abstract
Numerous deep learning-based works focusing on 3D semantic segmentation have been proposed and have achieved impressive performance. However, due to the catastrophic forgetting, existing methods will degrade dramatically in a real-world scenario where new 3D semantic categories are arriving continually. Straightforwardly applying typical class-incremental learning methods on 3D data even aggravates forgetting due to the irregular and noisy geometric structure. Aiming to address this realistic challenge, from the perspective of capturing local topological characteristics and mitigating global semantic shift, we propose a unified framework named Local topological Alignment and Global semantic Deconstruction (LAGD) to incrementally learn semantic knowledge of novel 3D categories while maintaining performance on previously learned knowledge. Specifically, we develop a novel Interaction Topological-aware Alignment (ITA) to maintain the learned knowledge efficiently by capturing the local geometric characteristics with interacted adjacent state-specific knowledge. Besides, to mitigate the forgetting caused by the global semantic shift, we deconstruct the logits into positive and negative parts which are distilled separately, achieving an elaborate distillation process in terms of Semantic-knowledge Deconstruction Distillation (SDD). With the cooperation of ITA and SDD, LAGD achieves a sota performance, especially in the long-term incremental learning scenario. Extensive experimental results illustrate the superiority of our proposed LAGD.
Haoran Duan 0001, Rui Sun 0010, Tejal Shah, Rajiv Ranjan 0001, Bo Wei 0003
AAAI6
2025 DART: Device-Native Adaptive Real-Time Training for Lifelong Learning on IoT Boards
Shamil Al-Ameen, Bharath Sudharsan, Osamah Alzacko, Roua Al-Taie, Tejal Shah, Rajiv Ranjan 0001
IEEE Big Data6
2025 Benchmarking Confidential Computing: Application Performance Comparison of TDX v/s SEV-SNP
abstract
With the growing reliance on cloud computing for both personal and enterprise-level applications, the need for robust data security has become essential. Confidential Computing Environments (CCEs) offer a promising solution by providing hardware-based isolation to protect sensitive data and computations. It encrypts the main memory and creates an isolated execution environment for executing the workload. Among the leading CCE technologies, Intel's Trusted Domain Extensions (TDX) and AMD's Secure Encrypted Virtualisation with Secure Nested Paging (SEV-SNP) have emerged as significant solutions. However, it becomes essential to understand their performance impact across a range of applications. The challenge lies in balancing security and performance, especially for organisations deploying resource-intensive workloads in the cloud. This study is motivated by the need to offer a comprehensive, applicationlevel performance comparison of Intel TDX and AMD SEVSNP in a real-world cloud environment. By evaluating both CCEs for micro benchmarks and application microservices on the Microsoft Azure cloud environment, this paper aims to provide insights into how these technologies handle different types of workloads, thus enabling cloud users to make informed decisions based on their security and performance requirements.
Mehul Sankhe, Ronil Rodrigues, Tomasz Szydlo, Rajiv Ranjan 0001, Devki Nandan Jha
HPCC5
2025 D2R: Dual Regularization Loss with Collaborative Adversarial Generation for Model Robustness
Huizhi Liang 0001, Rajiv Ranjan 0001, Zhanxing Zhu, Václav Snásel, Varun Ojha 0001
ICANN (1)3
2025 Rehearsal-free Federated Domain-incremental Learning
abstract
We introduce a rehearsal-free federated domain incremental learning framework, RefFiL, based on a global prompt-sharing paradigm to alleviate catastrophic forgetting challenges in federated domain-incremental learning, where unseen domains are continually learned. Typical methods for mitigating forgetting, such as the use of additional datasets and the retention of private data from earlier tasks, are not viable in federated learning (FL) due to devices’ limited resources. Our method, RefFiL, addresses this by learning domain-invariant knowledge and incorporating various domain-specific prompts from the domains represented by different FL participants. A key feature of RefFiL is the generation of local finegrained prompts by our domain adaptive prompt generator, which effectively learns from local domain knowledge while maintaining distinctive boundaries on a global scale. We also introduce a domain-specific prompt contrastive learning loss that differentiates between locally generated prompts and those from other domains, enhancing RefFiL’s precision and effectiveness. Compared to existing methods, RefFiL significantly alleviates catastrophic forgetting without requiring extra memory space, making it ideal for privacy-sensitive and resource-constrained devices.
Rui Sun 0010, Haoran Duan 0001, Jiahua Dong 0001, Varun Ojha 0001, Tejal Shah, Rajiv Ranjan 0001
ICDCS6
2025 A Behavioural Fingerprinting-Based Attack Detection Framework for Smart Home Devices
abstract
Smart home systems, which integrate diverse IoT devices for automation and energy efficiency, are increasingly targeted by cyber threats such as Man-in-the-Middle (MitM) and Denial-of-Service (DoS)/Distributed DoS (DDoS) attacks. Traditional security mechanisms often lack the adaptability and responsiveness needed for real-time protection, leaving devices vulnerable to data breaches and manipulation. This paper proposes a behavioural fingerprinting-based security framework that combines threshold-based detection with machine learning (ML) classification to identify anomalies in device activity patterns, including round-trip time (RTT) and Address Resolution Protocol (ARP) behaviour. The proposed approach is evaluated in a real smart home testbed environment using an IP camera, Raspberry Pis, and an ASUS router. Experimental results demonstrate a high detection accuracy across varied attack scenarios. The results show that the proposed system can detect attacks in near real-time, while maintaining compatibility with third-party applications, making it a practical and effective solution for enhancing smart home security.
Tianpu Li, Tomasz Szydlo, Rajiv Ranjan 0001, Devki Nandan Jha
ICPADS3
2025 LINEADAPTER: Parameter-Efficient Fine-Tuning for Log Anomaly Detection and Root Cause Analysis
abstract
The growing scale and complexity of distributed systems such as Hadoop produce massive volumes of complex log data, making automated anomaly detection and root cause analysis both essential and increasingly challenging. Traditional rule-based approaches, relying on static thresholds or manual heuristics, struggle to scale due to limited adaptability and high maintenance costs. Transformer-based language models have emerged as powerful tools for modelling log sequences by capturing contextual patterns. To enable efficient adaptation with fewer parameters, techniques such as In-Context Learning (ICL) and Low-Rank Adaptation (LoRA) have been proposed. However, applying these methods to log analysis with Small Language Models (SLMs) introduces several challenges, including high memory and computational overhead (in the case of ICL), limited fine-tuning capacity (with LoRA on smaller models), and poor generalisation across heterogeneous environments. To address these limitations, we propose Lineadapter, a parameter-efficient fine-tuning framework tailored for SLMs in log anomaly detection and root cause analysis. Lineadapter extends convolutional adapter principles to sequential data by integrating lightweight 1D convolutional layers within transformer blocks. This design enables SLMs to adapt effectively to log-specific patterns with minimal computational cost, while preserving the backbone model's representational power. We evaluate LINEADAPTER on multiple real-world system log datasets, comparing its performance with ICL, LoRA, and rule-based baselines. Results show that Lineadapter achieves higher F1-scores and better precision-recall trade-offs, particularly on medium-scale models, establishing it as a scalable, robust, and practical solution for log-based anomaly detection.
Wenhao Bao, Yinhao Li 0003, Rajiv Ranjan 0001, Devki Nandan Jha
ICPADS5
2025 Data Quality Detector: Automating Data Quality Detection in Smart City Environment
abstract
Ensuring Data Quality (DQ) is crucial for the reliability of environmental monitoring systems in the smart city environment. In this work, we present an automated framework, Data Quality Detector (DQD), for detecting, classifying, and analyzing DQ issues. We leveraged the Internet of Data Fault Taxonomy (IoDFT) to categorise faults based on various metrics, including type, duration, and pitfalls, providing a structured approach to identifying anomalies. By integrating these fault characteristics with key DQ dimensions-completeness, timeliness, and consis-tency- DQD enables a more granular classification of anomalies. Instead of simply marking missing faults, DQD differentiates between general and contextual missing faults, reducing false positives and improving decision-making accuracy. The DQD framework has been evaluated on real data from Newcastle Urban Observatory (NUO) to demonstrate its ability to enhance fault detection, minimise misclassification, and distinguish normal variations from actual quality issues, ultimately improving the trustworthiness and effectiveness of Air Quality (AQ) monitoring systems.
Sultan Altarrazi, Devki Nandan Jha, Tomasz Szydlo, Rajiv Ranjan 0001
ISCC4
2025 Modular neural network for edge-based detection of early-stage IoT botnet
abstract
The Internet of Things (IoT) has led to rapid growth in smart cities. However, IoT botnet-based attacks against smart city systems are becoming more prevalent. Detection methods for IoT botnet-based attacks have been the subject of extensive research, but the identification of early-stage behaviour of the IoT botnet prior to any attack remains a largely unexplored area that could prevent any attack before it is launched. Few studies have addressed the early stages of IoT botnet detection using monolithic deep learning algorithms that could require more time for training and detection. We, however, propose an edge-based deep learning system for the detection of the early stages of IoT botnets in smart cities. The proposed system, which we call EDIT (Edge-based Detection of early-stage IoT Botnet), aims to detect abnormalities in network communication traffic caused by early-stage IoT botnets based on the modular neural network (MNN) method at multi-access edge computing (MEC) servers. MNN can improve detection accuracy and efficiency by leveraging parallel computing on MEC. According to the findings, EDIT has a lower false-negative rate compared to a monolithic approach and other studies. At the MEC server, EDIT takes as little as 16 ms for the detection of an IoT botnet.
Duaa S. Alqattan, Varun Ojha 0001, Fawzy Habib, Ayman Noor, Graham Morgan, Rajiv Ranjan 0001
High Confid. Comput.6
2025 Analysis of deep learning under adversarial attacks in hierarchical federated learning
abstract
Hierarchical Federated Learning (HFL) extends traditional Federated Learning (FL) by introducing multi-level aggregation in which model updates pass through clients, edge servers, and a global server. While this hierarchical structure enhances scalability, it also increases vulnerability to adversarial attacks — such as data poisoning and model poisoning — that disrupt learning by introducing discrepancies at the edge server level. These discrepancies propagate through aggregation, affecting model consistency and overall integrity. Existing studies on adversarial behaviour in FL primarily rely on single-metric approaches — such as cosine similarity or Euclidean distance — to assess model discrepancies and filter out anomalous updates. However, these methods fail to capture the diverse ways adversarial attacks influence model updates, particularly in highly heterogeneous data environments and hierarchical structures. Attackers can exploit the limitations of single-metric defences by crafting updates that seem benign under one metric while remaining anomalous under another. Moreover, prior studies have not systematically analysed how model discrepancies evolve over time, vary across regions, or affect clustering structures in HFL architectures. To address these limitations, we propose the Model Discrepancy Score (MDS), a multi-metric framework that integrates Dissimilarity, Distance, Uncorrelation, and Divergence to provide a comprehensive analysis of how adversarial activity affects model discrepancies. Through temporal, spatial, and clustering analyses, we examine how attacks affect model discrepancies at the edge server level in 3LHFL and 4LHFL architectures and evaluate MDS’s ability to distinguish between benign and malicious servers. Our results show that while 4LHFL effectively mitigates discrepancies in regional attack scenarios, it struggles with distributed attacks due to additional aggregation layers that obscure distinguishable discrepancy patterns over time, across regions, and within clustering structures. Factors influencing detection include data heterogeneity, attack sophistication, and hierarchical aggregation depth. These findings highlight the limitations of single-metric approaches and emphasize the need for multi-metric strategies such as MDS to enhance HFL security.
Duaa S. Alqattan, Václav Snásel, Rajiv Ranjan 0001, Varun Ojha 0001
High Confid. Comput.3
2025 Parameter Efficient Fine-Tuning for Multi-modal Generative Vision Models with Möbius-Inspired Transformation
abstract
Abstract The rapid development of multimodal generative vision models has drawn scientific curiosity. Notable advancements, such as OpenAI’s ChatGPT and Stable Diffusion, demonstrate the potential of combining multimodal data for generative content. Nonetheless, customising these models to specific domains or tasks is challenging due to computational costs and data requirements. Conventional fine-tuning methods take redundant processing resources, motivating the development of parameter-efficient fine-tuning technologies such as adapter module, low-rank factorization and orthogonal fine-tuning. These solutions selectively change a subset of model parameters, reducing learning needs while maintaining high-quality results. Orthogonal fine-tuning, regarded as a reliable technique, preserves semantic linkages in weight space but has limitations in its expressive powers. To better overcome these constraints, we provide a simple but innovative and effective transformation method inspired by Möbius geometry, which replaces conventional orthogonal transformations in parameter-efficient fine-tuning. This strategy improved fine-tuning’s adaptability and expressiveness, allowing it to capture more data patterns. Our strategy, which is supported by theoretical understanding and empirical validation, outperforms existing approaches, demonstrating competitive improvements in generation quality for key generative tasks.
Haoran Duan 0001, Bing Zhai, Tejal Shah, Jungong Han, Rajiv Ranjan 0001
Int. J. Comput. Vis.6
2025 Correction: Parameter Efficient Fine-Tuning for Multi-modal Generative Vision Models with Möbius-Inspired Transformation
Haoran Duan 0001, Bing Zhai, Tejal Shah, Jungong Han, Rajiv Ranjan 0001
Int. J. Comput. Vis.6
2025 Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields
abstract
In this work, we propose a method that leverages CLIP feature distillation, achieving efficient 3D segmentation through language guidance. Unlike previous methods that rely on multi-scale CLIP features and are limited by processing speed and storage requirements, our approach aims to streamline the workflow by directly and effectively distilling dense CLIP features, thereby achieving precise segmentation of 3D scenes using text. To achieve this, we introduce an adapter module and mitigate the noise issue in the dense CLIP feature distillation process through a self-cross-training strategy. Moreover, to enhance the accuracy of segmentation edges, this work presents a low-rank transient query attention mechanism. To ensure the consistency of segmentation for similar colors under different viewpoints, we convert the segmentation task into a classification task through label volume, which significantly improves the consistency of segmentation in color-similar areas. We also propose a simplified text augmentation strategy to alleviate the issue of ambiguity in the correspondence between CLIP features and text. Extensive experimental results show that our method surpasses current state-of-the-art technologies in both training speed and performance.
Xingyu Miao, Haoran Duan 0001, Yang Bai 0011, Tejal Shah, Jun Song 0003, Yang Long 0001, Rajiv Ranjan 0001, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 SIB: Sorted-Integers-Based Index for Compact and Fast Caching in Top-Down Logic Rule Mining Targeting KB Compression
abstract
Background Mining logic rules from structured knowledge bases is the basis of knowledge engineering. Due to the NP‐hardness of the rule mining problem, logic rules cannot be efficiently induced from knowledge bases, especially large‐scale ones. Idea In this article, we propose a compact and efficient index structure for the maintenance of the intermediate data during top‐down rule mining, such that the memory consumption can be reduced and mining efficiency can be improved. Developing Points The index is based on a mapping from constant symbols to integers and the sorting of the mapped integers. Index update has been dissembled into four basic operations. Moreover, the index itself acts as the cache during top‐down mining. Value Most contributions in existing works employ algorithmic and architectural optimizations to improve efficiency. Data‐oriented optimizations have also been explored to some extent, but the data efficiency is relatively low, and the memory consumption is thus becoming a new challenge for state‐of‐the‐art systems. We tackle this challenge in this article, and our technique has been proven more efficient than state‐of‐the‐art systems. We evaluate our method on six datasets which contain up to 160 K records and are frequently used as benchmarks in tasks related to knowledge engineering. The experimental results show that the proposed technique speeds up the rule mining procedure by on average and reduces memory consumption by up to 70%. The space overhead of the data structure is about twice that of the indexed records, which is more than 80% lower than that of the state‐of‐the‐art technique.
Ruoyu Wang 0004, Raymond K. Wong 0001, Daniel Sun 0004, Rajiv Ranjan 0001
Softw. Pract. Exp.4
2025 Nicaea: A Byzantine Fault Tolerant Consensus Under Unpredictable Message Delivery Failures for Parallel and Distributed Computing
abstract
Byzantine fault-tolerant (BFT) consensus is a critical problem in parallel and distributed computing systems, particularly with potential adversaries. Most prior work on BFT consensus assumes reliable message delivery and tolerates arbitrary failures of up to$\frac{n}{3}$nodes out of$n$total nodes. However, many systems face unpredictable message delivery failures. This paper investigates the impact of unpredictable message delivery failures on the BFT consensus problem. We propose Nicaea, a novel protocol enabling consensus among loyal nodes when the number of Byzantine nodes is below a new threshold, given by:$\frac{\left(2-\rho\right)\left(1-\rho\right)^{2n-2}-1}{\left(2-\rho\right) \left(1-\rho\right)^{2n-2}+1}n$, where$\rho$denotes the message failure rate. Theoretical proofs and experimental results validate Nicaea's Byzantine resilience. Our findings reveal a fundamental trade-off: as message delivery instability increases, a system's tolerance to Byzantine failures decreases. The well-known$\frac{n}{3}$threshold under reliable message delivery is a special case of our generalized threshold when$\rho=0$. To the best of our knowledge, this work presents the first quantitative characterization of unpredictable message delivery failures’ impact on Byzantine fault tolerance in parallel and distributed computing.
Guanlin Jing, Yifei Zou, Minghui Xu 0001, Yanqiang Zhang, Dongxiao Yu, Zhiguang Shan, Xiuzhen Cheng, Rajiv Ranjan 0001
IEEE Trans. Computers8
2025 MMCANet A Multimodal and Cross-Attention Network for Cloud Removal and Exploration of Progressive Remote Sensing Images Restoration Algorithm
abstract
In Earth observation, cloud severely affects the interpretation of optical satellites generated high-resolution images. Cloud-free optical images are vital for downstream tasks such as semantic segmentation and object detection. Thus, the elimination of clouds from optical imagery has emerged as a significant topic in remote sensing. Currently, most existing methods are proposed to leverage the texture information from auxiliary synthetic aperture radar (SAR) images to restore cloud-free images via direct channel merging. However, such a unified feature extraction approach often neglects the inherent distribution disparity between SAR and optical images—the result of differing imaging principles-potentially leading to significant feature loss. To this end, we introduce a network by jointing SAR and optical images multimodal and cross-attention network (MMCANet) to effectively extract multiscale contextual features from SAR imagery and integrate them with optical features. Specifically, instead of simple concatenation of the channels of SAR and optical images, we obtain high-dimensional features from them through independent feature extractors. The integration of these features is facilitated by a cross-attention mechanism that provides a more fine-grained amalgamation of information. Meanwhile, an atrous spatial pyramid pooling (ASPP) module is introduced into the integration of high-level features, which captures multiscale contextual information around clouded areas. In addition, we propose four advanced remote sensing image restoration algorithms that approach image restoration as a series of subtasks, gradually eliminating clouds to enhance performance. Comprehensive assessments show that MMCANet performs well on the SEN 12 MS-CR dataset with peak signal-to-noise ratio (PSNR) of 39.8871, structural similarity index (SSIM) of 0.9672, mean absolute error (MAE) of 0.0081, and spectral angle mapper (SAM) of 2.9884.
Yejian Zhou, Jiahui Suo, Yachen Wang, Jie Su 0001, Zhen Hong, Rajiv Ranjan 0001, Lizhe Wang 0001, Zhenyu Wen
IEEE Trans. Geosci. Remote. Sens.7
2025 A Semantic-Consistent Few-Shot Modulation Recognition Framework for IoT Applications
abstract
The rapid growth of the Internet of Things (IoT) has led to the widespread adoption of the IoT networks in numerous digital applications. To counter physical threats in these systems, automatic modulation classification (AMC) has emerged as an effective approach for identifying the modulation format of signals in noisy environments. However, identifying those threats can be particularly challenging due to the scarcity of labeled data, which is a common issue in various IoT applications, such as anomaly detection for unmanned aerial vehicles (UAVs) and intrusion detection in the IoT networks. Few-shot learning (FSL) offers a promising solution by enabling models to grasp the concepts of new classes using only a limited number of labeled samples. However, prevalent FSL techniques are primarily tailored for tasks in the computer vision domain and are not suitable for the wireless signal domain. Instead of designing a new FSL model, this work suggests a novel approach that enhances wireless signals to be more efficiently processed by the existing state-of-the-art (SOTA) FSL models. We present the semantic-consistent signal pretransformation (ScSP), a parameterized transformation architecture that ensures signals with identical semantics exhibit similar representations. ScSP is designed to integrate seamlessly with various SOTA FSL models for signal modulation recognition and supports commonly used deep learning backbones. Our evaluation indicates that ScSP boosts the performance of numerous SOTA FSL models, while preserving flexibility.
Jie Su 0001, Zhenyu Wen, Fangda Guo, Yiming Wu 0009, Zhen Hong, Haoran Duan 0001, Yawen Huang, Rajiv Ranjan 0001, Yefeng Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.10
2024 LEAP: Lifelong Learning Edge-Cloud Adaptive Fused Framework for Mobility Prediction
abstract
Accurate mobility prediction has become pivotal for a wide range of smart city applications including optimizing electric vehicles (EV) charging management, traffic management, infrastructure planning, etc. However, traditional mobility prediction models face significant challenges including ineffective integration of geographical information, inability to dynamically adapt to changing location popularity, and struggle with long-term dependency. Furthermore, these models are susceptible to catastrophic forgetting, losing previously learned knowledge when exposed to new data.To overcome these challenges, we propose the Lifelong Edge-cloud Adaptive Prediction (LEAP) framework, a fresh approach that integrates lifelong learning into mobility prediction. LEAP improves prediction accuracy and reliability in dynamic real-world environments by fusing a central cloud model for capturing long-term trends with multiple local edge models that process real-time data. LEAP employs Spatially Adaptive LSTM (SA-LSTM) and Temporal Adjustment LSTM (TA-LSTM) to incorporate dynamic spatio-temporal patterns, along with Global Context Operations (GC Ops) to manage long-term dependencies. To prevent catastrophic forgetting, LEAP uses Learning without Forgetting (LwF), enabling on-device continuous learning and adaptation at the edge.Extensive evaluations demonstrate that LEAP surpasses ten state-of-the-art methods, including Scikit-Learn’s LSTM, with average improvements of 2.61x in Recall-5, 2.35x in Recall-10, 2.80x in NDCG-5, and 3.06x in NDCG-10. These results highlight LEAP’s superior effectiveness, accuracy, and adaptability, proving it a worthy choice for dynamic real-world mobility prediction tasks while effectively addressing catastrophic forgetting.
Shamil Al-Ameen, Bharath Sudharsan, Roua Al-Taie, Tejal Shah, Rajiv Ranjan 0001
IEEE Big Data5
2024 Poly Instance Recurrent Neural Network for Real-time Lifelong Learning at the Low-power Edge
abstract
As machine learning moves towards edge deployment, lifelong learning becomes crucial due to evolving data distributions and new tasks. Yet, applying traditional methods to learn from vast, complex IoT data streams poses challenges. These include excessive CPU usage, RAM overflow, prolonged convergence times disrupting device operation, and difficulties in adapting to concept drift. Consequently, models trained on devices struggle to handle frequently changing data, affecting their ability to respond effectively to new inputs.To address these issues, we introduce Poly Instance Lifelong Learning (PILL), an algorithm designed for real-time on-device model training and inference at the edge under lifelong learning settings. PILL is lightweight, operating efficiently on the CPUs of low-power single-board computers (SBCs). It achieves this by partitioning input data into manageable instances, filtering out label noise, and applying early stopping for rapid predictions.PILL was evaluated on three popular low-power SBCs as well as a high-end Windows 10 machine using four datasets of different sizes and features. The results indicate that despite the superior resources of the Windows 10 machine, models trained using PILL on SBCs differ in accuracy by only ±0.05%. Additionally, PILL’s LSTM trains 2.41 - 2.85 times faster than the widely used Scikit-Learn’s LSTM. Additionally, when compared to ten state-of-the-art methods, PILL demonstrated superior performance across key metrics (Precision, Recall, and F1-Score) while minimizing computational overhead, making it an ideal choice for efficient, real-time edge deployment.
Shamil Al-Ameen, Bharath Sudharsan, Tejus Vijayakumar, Tomasz Szydlo, Tejal Shah, Rajiv Ranjan 0001
IEEE Big Data6
2024 Swarm Storm: An Automated Chaos Tool for Docker Swarm Applications
abstract
Docker Swarm facilitates the deployment of modular applications in an independent yet interconnected manner. Applications within the Swarm communicate through Docker's internal network, ensuring rapid computation and enhanced security. Despite these advantages, the complex nature of software development often leads to the occurrence of various faults, including memory leaks, application failures, and network outages. To systematically identify and address such issues, chaos engineering has emerged as a powerful approach within both development and production environments. However, conducting chaos experiments within a Docker Swarm cluster in a multi-cloud environment proves to be challenging. In this paper, we present Swarm Storm, an automated framework designed for orchestrating chaos engineering experiments within Docker Swarm-based clusters in a cloud-agnostic environment with comprehensive testing features. We validate Swarm Storm using a simple Java benchmark application, demonstrating its effectiveness in addressing faults within the system.
Travis Higgins, Devki Nandan Jha, Rajiv Ranjan 0001
HPDC3
2024 Security Assessment of Hierarchical Federated Deep Learning
Duaa S. Alqattan, Rui Sun 0010, Huizhi Liang 0001, Guiseppe Nicosia, Václav Snásel, Rajiv Ranjan 0001, Varun Ojha 0001
ICANN (6)6
2024 Verifiable Querying Framework for Multi-Blockchain Applications
abstract
Effectively and securely retrieving data from various blockchain networks remains a critical challenge. We propose a novel framework that provides enhanced capabilities for authenticated data retrieval, empowering users and regulatory bodies. Our framework provides a scalable solution for information verification across diverse blockchain systems, enabling a variety of different types of actors to be integrated. Our framework allows metadata from various systems to be combined, also providing support for secure querying.
Stanly Wilson, Kwabena Adu-Duodu, Yinhao Li 0003, Ringo W. H. Sham, Ellis Solaiman, Omer F. Rana, Rajiv Ranjan 0001
ICBC7
2024 MatchCom: Stable Matching-Based Software Services Composition in Cloud Computing Environments
Renyu Yang, Rajiv Ranjan 0001, Rami Bahsoon, Jie Xu 0007, Rajkumar Buyya
ICWE3
2024 Dual Variational Knowledge Attention for Class Incremental Vision Transformer
abstract
Class incremental learning (CIL) strives to emulate the human cognitive process of continuously learning and adapting to new tasks while retaining knowledge from past experiences. Despite significant advancements in this field, Transformer-based models have not fully leveraged the potential of attention mechanisms to balance the transferable knowledge between tokens and the associated information. This paper addresses this gap by using a dual variational knowledge attention (DVKA) mechanism within a Transformer-based encoder-decoder framework, tailored for CIL. DVKA mechanism aims to manage the information flow through the attention maps, ensuring a balanced representation of all classes, and mitigating the risk of information dilution as new classes are incrementally introduced. This method, leverage the information bottleneck and mutual information principle, selectively filters less relevant information, directing the model’s focus towards the most significant details for each class. The DVKA is designed with two distinct attentions: one focused on the feature level and the other on the token dimension. The feature-focused attention aims to purify the complex nature of various classification tasks, ensuring a comprehensive representation of both old and new tasks. The token-focused attention mechanism highlights specific tokens, facilitating local discrimination among disparate patches and fostering global coordination for a spectrum of task tokens. Our work is a major stride towards improving transformer models for class incremental learning, presenting a theoretical rationale and effective experimental results on three widely-used datasets.
Haoran Duan 0001, Rui Sun 0010, Varun Ojha 0001, Tejal Shah, Zhuoxu Huang, Zizhou Ouyang, Yawen Huang, Yang Long 0001, Rajiv Ranjan 0001
IJCNN9
2024 General data protection regulation: a study on attitude and emotional empowerment
abstract
Over the last few years, digitalisation has accelerated its pace, fuelling the creation of a massive amount of data.This has resulted in a need to introduce legal mechanisms to protect the privacy and security of data being exchanged between people and organisations.However, little is known about the individuals' perspective on such mechanisms.Given the gap in the literature, this research investigated the drivers and the implications of individuals' attitude towards GDPR compliance.To test the research model, structural equational modelling was employed using 540 responses.The result showed that perceived threat severity, self-efficacy and response efficacy determine a positive attitude towards GDPR compliance, which results in emotional empowerment.The findings contribute to the literature on legal privacy-preserving mechanisms, by providing a user's view on the coping and threat appraisal factors underpinning attitude and demonstrating the implications for driving confidence in control over personal data.The findings also contribute to the literature on protection motivation by demonstrating that attitude towards adaptive behaviour drives emotional empowerment.The study offers suggestions to policymakers on how to enhance public perception of the GDPR.The findings also provide guidelines for organisations on how to inform individuals' understanding of compliance with the legal framework.
Davit Marikyan, Savvas Papagiannidis, Omer F. Rana, Rajiv Ranjan 0001
Behav. Inf. Technol.4
2024 Wearable-based behaviour interpolation for semi-supervised human activity recognition
abstract
While traditional feature engineering for Human Activity Recognition (HAR) involves a trial-and-error process, deep learning has emerged as a preferred method for high-level representations of sensor-based human activities. However, most deep learning-based HAR requires a large amount of labelled data and extracting HAR features from unlabelled data for effective deep learning training remains challenging. We, therefore, introduce a deep semi-supervised HAR approach, MixHAR, which concurrently uses labelled and unlabelled activities. Our MixHAR employs a linear interpolation mechanism to blend labelled and unlabelled activities while addressing both inter- and intra-activity variability. A unique challenge identified is the activity-intrusion problem during mixing, for which we propose a mixing calibration mechanism to mitigate it in the feature embedding space. Additionally, we rigorously explored and evaluated the five conventional/popular deep semi-supervised technologies on HAR, acting as the benchmark of deep semi-supervised HAR. Our results demonstrate that MixHAR significantly improves performance, underscoring the potential of deep semi-supervised techniques in HAR.
Haoran Duan 0001, Varun Ojha 0001, Shizheng Wang, Yawen Huang, Yang Long 0001, Rajiv Ranjan 0001, Yefeng Zheng 0001
Inf. Sci.7
2024 BFT-DSN: A Byzantine Fault-Tolerant Decentralized Storage Network
abstract
With the rapid development of blockchain and its applications, the amount of data stored on decentralized storage networks (DSNs) has grown exponentially. DSNs bring together affordable storage resources from around the world to provide robust, decentralized storage services for tens of thousands of decentralized applications (dApps). However, existing DSNs do not offer verifiability when implementing erasure coding for redundant storage, making them vulnerable to Byzantine encoders. Additionally, there is a lack of Byzantine fault-tolerant consensus for optimal resilience in DSNs. This paper introduces BFT-DSN, a Byzantine fault-tolerant decentralized storage network designed to address these challenges. BFT-DSN combines storage-weighted BFT consensus with erasure coding and incorporates homomorphic fingerprints and weighted threshold signatures for decentralized verification. The implementation of BFT-DSN demonstrates its comparable performance in terms of storage cost and latency as well as superior performance in Byzantine resilience when compared to existing industrial decentralized storage networks.
Hechuan Guo, Minghui Xu 0001, Jiahao Zhang 0003, Chun-Chi Liu, Rajiv Ranjan 0001, Dongxiao Yu, Xiuzhen Cheng
IEEE Trans. Computers5
2024 Rapid Crowd Evacuation for Passenger Ships Using LPWAN
abstract
An emerging evacuation path planning technique that uses Low Power Wide Area Networks (LPWAN) to enable real-time danger prediction and user-oriented path planning can ensure the safe and timely navigation of evacuees in complex scenarios such as cruise ships. However, most existing LPWAN-based evacuation models assume pedestrians’ walking speed remains constant and ignore crowd congestion in corridors before exits, which is not appropriate for rocking ships. To overcome these issues, this paper proposes a congestion-relived guiding framework with dedicated path planning for emergency evacuation on passenger ships. The basic idea is to averagely minimize the total evacuation time while meeting the deadline for ship capsizing under all circumstances by selecting uncrowded paths for each passenger individually. First, we use probability distributions rather than constant numbers to represent walking time (also called delay) along passageways. A worst-case delay bound with a high level of trustworthiness is also estimated for each passageway under the boundary condition of ship capsizing. Next, we predict the congestion of corridors by modeling the spatiotemporal movement of passengers, and then distribute evacuation loads evenly among corridors to alleviate the congestion. The total expected evacuation time of all corridors is finally minimized based on the delay probability distribution and estimated congestion, and the deadline for ship evacuation under all circumstances is met with the worst-case delay bound. Simulation results show that our approach significantly reduces the total escaping time of crowd evacuation by 45% and 34% while improving the navigation success ratio by more than 20% and 80% compared with the state-of-the-art emergency evacuation systems, namely the look-up table guiding scheme and the group-based guiding evacuation scheme, respectively.
Kezhong Liu, Mozi Chen, Yinhao Li 0003, Rui Sun 0010, Rajiv Ranjan 0001
IEEE Trans. Intell. Transp. Syst.6
2024 GeoDeploy: Geo-Distributed Application Deployment Using Benchmarking
abstract
Geo-distributed web-applications (GWA) can be deployed across multiple geographically separated datacenters to reduce the latency of access for users. Finding a suitable deployment for a GWA is challenging due to the requirement to consider a number of different parameters, such as host configurations across a federated infrastructure. The ability to evaluate multiple deployment configurations enables an efficient outcome to be determined, balancing resource usage while satisfying user requirements. We proposeGeoDeploy, a framework designed for finding a deployment solution for GWA. We evaluateGeoDeployusing both a formal algorithmic model and a practical cloud-based deployment. We also compare our approach with other existing techniques.
Devki Nandan Jha, Yinhao Li 0003, Zhenyu Wen, Graham Morgan, Prem Prakash Jayaraman, Maciej Koutny, Omer F. Rana, Rajiv Ranjan 0001
IEEE Trans. Parallel Distributed Syst.8
2024 Long Live the Image: On Enabling Resilient Production Database Containers for Microservice Applications
abstract
Microservices architecture advocates decentralized data ownership for building software systems. Particularly, in the Database per Service pattern, each microservice is supposed to maintain its own database and to handle the data related to its functionality. When implementing microservices in practice, however, there seems to be a paradox: The de facto technology (i.e., containerization) for microservice implementation is claimed to be unsuitable for the microservice component (i.e., database) in production environments, mainly due to the data persistence issues (e.g., dangling volumes) and security concerns. As a result, the existing discussions generally suggest replacing database containers with cloud database services, while leaving the on-premises microservice implementation out of consideration. After identifying three statelessness-dominant application scenarios, we proposed container-native data persistence as a conditional solution to enable resilient database containers in production. In essence, this data persistence solution distinguishes stateless data access (i.e., reading) from stateful data processing (i.e., creating, updating, and deleting), and thus it aims at the development of stateless microservices for suitable applications. In addition to developing our proposal, this research is particularly focused on its validation, via prototyping the solution and evaluating its performance, and via applying this solution to two real-world microservice applications. From the industrial perspective, the validation results have proved the feasibility, usability, and efficiency of fully containerized microservices for production in applicable situations. From the academic perspective, this research has shed light on the operation-side micro-optimization of individual microservices, which fundamentally expands the scope of “software micro-optimization” and reveals new research opportunities.
Zheng Li 0001, Nicolás Saldías-Vallejos, Diego Seco Naveiras, M. Andrea Rodríguez, Rajiv Ranjan 0001
IEEE Trans. Software Eng.5
2023 Tracking Material Reuse across Construction Supply Chains
abstract
Material reuse and recycling plays a key role in reducing carbon emissions in the architecture and construction sector. A “Material Passport” (MP) is a record describing how a material is used throughout its lifetime, from genesis to termination, recording operations carried out on the material. The granularity of information recorded in a MP can vary, however ensuring that this provenance trail remains immutable is a key requirement. The benefits of using a MP, operations carried out on a MP, and recording of transactions within a distributed Blockchain (parachain) is described. A scenario is used to illustrate how the proposed approach can be used in practice.
Stanly Wilson, Kwabena Adu-Duodu, Yinhao Li 0003, Ringo W. H. Sham, Ellis Solaiman, Charith Perera, Rajiv Ranjan 0001, Omer F. Rana
e-Science8
2023 Compliance Checking of Cloud Providers: Design and Implementation
abstract
The recognition of capabilities supplied by cloud systems is presently growing. Collecting or sharing healthcare data and sensitive information especially during the Covid-19 pandemic has motivated organizations and enterprises to leverage the upsides coming from cloud-based applications. However, the privacy of electronic data in such applications remains a significant challenge for cloud vendors to adapt their solutions with existing privacy legislation standards such as general data protection regulation (GDPR). This article first proposes a formal model and verification for data usage requests of providers in a cloud composite service using a model checking tool. A cloud pharmacy scenario is presented to illustrate the connectivity of providers in the composite service and the stream of their requests for both collection and movement of patient data. A set of verifications is then undertaken over the pharmacy service in accordance with three significant GDPR obligations, namely user consent, data access, and data transfer. Following that, the article designs and implements a cloud container virtualization based on the verified formal model realizing GDPR requirements. The container makes use of some enforcement smart contracts to only proceed with the providers’ requests that are compliant with GDPR. Finally, several experiments are provided to investigate the performance of our approach in terms of time, memory, and cost.
Masoud Barati, Kwabena Adu-Duodu, Omer F. Rana, Gagangeet Singh Aujla, Rajiv Ranjan 0001
Distributed Ledger Technol. Res. Pract.5
2023 IoT-QWatch: A Novel Framework to Support the Development of Quality-Aware Autonomic IoT Applications
abstract
The unprecedented growth of Internet of Things (IoT) is leading to its increased usage in various domains, such as manufacturing, health, and smart cities. A majority of IoT applications are autonomic, i.e., they operate under minimal/no human intervention, and make decisions/actuations based on machine-to-machine communication and data analytics. A key challenge in the development of such applications is the ability to measure their quality while they are working in a diverse and heterogeneous IoT ecosystem. In this article, we propose an agent-based IoT-Quality Watch (IoT-QWatch) framework that provides the ability to measure IoT quality metrics at each stage of the autonomic IoT application life cycle running in the IoT ecosystem. We envision that IoT-QWatch will enable the development of a new generation of quality-aware autonomic IoT applications that are able to be resilient to the heterogeneous and uncertain nature of IoT ecosystems. We present architectural details and implementation of IoT-QWatch, and corresponding models used to measure IoT quality metrics at different stages. We conduct extensive experiments using a real-world IoT test bed from the domain of manufacturing to validate the efficacy of IoT-QWatch. Experimental outcomes provide promising results in realizing IoT-QWatch in real-world deployment, while the framework itself offers significant extensibility to include new models for measuring IoT quality metrics.
Kaneez Fizza, Prem Prakash Jayaraman, Abhik Banerjee, Nitin Auluck, Rajiv Ranjan 0001
IEEE Internet Things J.5
2023 Horn rule discovery with batched caching and rule identifier for proficient compressor of knowledge data
abstract
Abstract Knowledge data has been widely applied to artificial intelligence applications for interpretable and complex reasoning. Modern knowledge bases are constructed via automatic knowledge extraction from open‐accessible sources. Thus the sizes of KBs are continuously growing, heavily burdening the maintenance and application of the knowledge data. Besides the grammatical redundancies, semantically repeated information also frequently appears in knowledge bases but is still under‐explored. Existing semantic compressors fail to efficiently discover expressive patterns and thus perform unsatisfyingly on knowledge data. This article proposes SInC, a semantic inductive compressor, to efficiently induce first‐order Horn rules and semantically compress knowledge bases. SInC improves the scalability of top‐down rule mining by batching correlated records in the cache and further optimizes the pruning of duplication and specialization via an identifier structure of Horn rules. SInC was evaluated on real‐world and synthetic datasets and compared against the state‐of‐the‐art. The results show that the batched caching speed up the rule mining procedure by more than two orders while consuming fewer than three times memory space. The identifier technique speeds up the duplication and specialization pruning by orders of magnitude with less than 5‰ and 15% error rates, respectively. SInC outperforms the state‐of‐the‐art from the perspective of overall compression on both scalability and compression effect.
Ruoyu Wang 0004, Daniel Sun 0004, Raymond K. Wong 0001, Rajiv Ranjan 0001
Softw. Pract. Exp.4
2023 Cross-Channel: Scalable Off-Chain Channels Supporting Fair and Atomic Cross-Chain Operations
abstract
Cross-chain technology facilitates the interoperability among isolated blockchains on which users can freely communicate and transfer values. Existing cross-chain protocols suffer from the scalability problem when processing on-chain transactions. Off-chain channel, as a promising blockchain scaling technique, can enable micro-payment transactions without involving on-chain transaction settlement. However, existing channel schemes can only be applied to operations within a single blockchain, failing to support cross-chain services. Therefore in this paper, we propose$\mathsf {Cross}$-$\mathsf {Channel}$, the first off-chain channel to support cross-chain services. We introduce a novel hierarchical channel structure with a hierarchical interaction protocol, a new hierarchical settlement protocol, and a smart general fair exchange protocol, to ensure scalability, fairness, and atomicity of cross-chain interactions. Besides,$\mathsf {Cross}$-$\mathsf {Channel}$provides strong security and practicality by avoiding high latency in asynchronous networks.Through a 50-instance deployment of$\mathsf {Cross}$-$\mathsf {Channel}$on AliCloud, we demonstrate that$\mathsf {Cross}$-$\mathsf {Channel}$is well-suited for processing cross-chain transactions in high-frequency and large-scale, and brings a significantly enhanced throughput with a small amount of gas and delay overhead.
Minghui Xu 0001, Dongxiao Yu, Yong Yu 0002, Rajiv Ranjan 0001, Xiuzhen Cheng
IEEE Trans. Computers5
2023 OsmoticGate: Adaptive Edge-Based Real-Time Video Analytics for the Internet of Things
abstract
Edge computing has gained momentum in recent years, and can provide more immediate analysis of streaming video data. However, the edge devices often lack the computing capabilities (processing power, memory) to guarantee reasonable performance (e.g., accuracy, latency, throughput) for complex video analytics tasks. To alleviate this critical problem, the prevalent trend is to offload some video analytics tasks from the edge devices to the cloud. However, existing offloading approaches fail to consider the dynamic nature of the video analytical tasks (e.g., varying encoding format for different video content) and are unable to adapt system dynamics (e.g., varying workload between the edge and the cloud). To overcome the limitation of existing approaches, we develop an edge-cloud offloading performance model based on the concept of hierarchical queues. The resource constraints (e.g., computing capacity and network bandwidth) of each edge nodes and dynamic edge-cloud network conditions are used to parameterize the performance model. Since finding optimal solutions for the performance model is NP-hard, we develop a two-stage gradient-based algorithm and compare it with some state-of-the-art (SOTA) solutions (e.g., FastVA, DeepDecision, Hill Climbing). Experiments have shown our performance model's advantages and the stability of the proposed offloading approach given different systems (edge-cloud) and video analytics application dynamics.
Bin Qian 0002, Zhenyu Wen, Junqi Tang, Ye Yuan 0001, Albert Y. Zomaya, Rajiv Ranjan 0001
IEEE Trans. Computers6
2023 iQuery: A Trustworthy and Scalable Blockchain Analytics Platform
abstract
Blockchain, a distributed and shared ledger, provides a credible and transparent solution to increase application auditability by querying the immutable records written in the ledger. Unfortunately, existing query APIs offered by the blockchain are inflexible and unscalable. Some studies propose off-chain solutions to provide more flexible and scalable query services. However, the query service providers (SPs) may deliver fake results without executing the real computation tasks and collude to cheat users. In this article, we propose a novel intelligent blockchain analytics platform termediQuery, in which we design a game theory based smart contract to ensure the trustworthiness of the query results at a reasonable monetary cost. Furthermore, the contract introduces the second opinion game that employs a randomized SP selection approach coupled with non-ordered asynchronous querying primitive to prevent collusion. We achieve a fixed price equilibrium, destroy the economic foundation of collusion, and can incentivize all rational SPs to act diligently with proper financial rewards. In particular,iQuerycan flexibly support semantic and analytical queries for generic consortium or public blockchains, achieving query scalability to massive blockchain data. Extensive experimental evaluations show thatiQueryis significantly faster than state-of-the-art systems. Specifically, in terms of the conditional, analytical, and multi-origin query semantics,iQueryis 2 ×, 7 ×, and 1.5 × faster than advanced blockchain and blockchain databases. Meanwhile, to guarantee 100% trustworthiness, only two copies of query results need to be verified iniQuery, whileiQuery's latency is$2 \sim 134$× smaller than the state-of-the-art systems.
Lingling Lu, Zhenyu Wen, Ye Yuan 0001, Binru Dai, Changting Lin, Qinming He, Zhenguang Liu, Jianhai Chen, Rajiv Ranjan 0001
IEEE Trans. Dependable Secur. Comput.10
2023 Enhanced Bayesian Factorization With Variant Scale Partitioning for Multivariate Time Series Analysis
abstract
Multivariate time series data (Mv-TSD) portray the evolving processes of the system(s) under examination in a “multi-view” manner. Factorization methods are salient for Mv-TSD analysis with the potentials of structural feature construction correlating various data attributes. However, research challenges remain in the derivation of factors due to highly scattered data distribution of Mv-TSD and intensive interferences/outliers embedded in the source data. The proposed Enhanced Bayesian Factorization approach (Enhanced-BF) addresses the challenges in three phases: (1) variant scale partitioning applies to Mv-TSD according to degree of amplitude and obtains the blocks of variant scales; (2) hierarchical Bayesian model for tensor factorization automatically derives the factors of each block with interferences suppressed; (3) Bayesian unification model merges those block factors to construct the final structural features.Enhanced-BFhas been evaluated using a case study of brain data engineering with multivariate electroencephalogram (EEG). Experimental results indicate that the proposed method manifests robustness to the interferences and outperforms the counterparts in terms of operation efficiency and error when factorizing EEG tensor. Besides,Enhanced-BFexcels in factorization-based analysis of ongoing autism spectrum disorder (ASD) EEG: 3 times speed-up in factorization and$87.35\%$accuracy in ASD discrimination. The latent factors (“biomarkers”) can distinctly interpret the typical EEG characteristics of ASD subjects.
Yunbo Tang, Dan Chen 0001, Yiping Zuo, Xiaoqiang Lu, Rajiv Ranjan 0001, Albert Y. Zomaya, Quanming Yao, Xiaoli Li 0002
IEEE Trans. Knowl. Data Eng.5
2023 Janus: Latency-Aware Traffic Scheduling for IoT Data Streaming in Edge Environments
abstract
This article focuses on a simple, yet fundamental question of distributed edge computing: “how to handle IoT traffic with different levels of sensitivity and criticality by satisfying the application-specific latency constraints?” This question arises in the practical deployment of edge computing, where user data can arrive at a much faster rate than that they can be processed by an edge node. Addressing this question is critical for meeting the latency requirement for latency-sensitive applications, but existing approaches are inadequate to the problem. We presentJanus, a multi-level traffic scheduling system for managing multiple data streams with various degrees of latency constraints. At the edge node level,Janususes multi-level queues to manage data streams with different latency constraints. It then allocates the output bandwidth of the edge node according to the requirements of applications in different priority queues, aiming to reduce the queuing and processing delay of latency-sensitive streams while maximizing the edge-node throughput. At the network level,Janusactively redirects incoming data streams to the less-loaded ones to achieve better network-wide load balance and improve the overall throughput. Experiments show thatJanusreduces the latency to only 16.6% of a non-priority based solution and improves the throughput by 1.7x of a state-of-the-art priority-aware data stream scheduling approach.
Zhenyu Wen, Renyu Yang, Bin Qian 0002, Yubo Xuan, Lingling Lu, Zheng Wang 0001, Hao Peng 0001, Jie Xu 0007, Albert Y. Zomaya, Rajiv Ranjan 0001
IEEE Trans. Serv. Comput.10
2023 A Hybrid Accuracy- and Energy-Aware Human Activity Recognition Model in IoT Environment
abstract
Personalised health and fitness provide users with information regarding their wellbeing and an opportunity to inform healthcare services for better patient outcomes. Underpinning this industry sector is the need to establish human activity recognition (HAR) in a ubiquitous manner. For example, through the use of smartwatches and/or mobile phones gathering information such as heart rates, movement, and steps of a user. The engineering challenge is providing accurate, informative, and timely data without rapidly depleting the mobile device's battery life. This problem is compounded as a number of algorithms used to process such data require substantial, cloud-based resources, to achieve higher accuracy. Therefore, a balance is required between battery depletion, accuracy of data, and timely delivery of results through a mixture of cloud and local algorithmic execution. In this article, we proposeAE-HAR (Accuracy and Energy Aware-HAR)model that delivers engineered solutions which approach optimal combinations in the consideration of energy consumption, accuracy, and timeliness of results.AE-HARintroduces a “light-weight”machine learningon-device component identifying the probabilistic accuracy of data together with energy consumption identification requirements. A heuristic is then adopted to determine if cloud-enabled calculations are required while including possible performance costs related to the analysis of networking infrastructures. Our model is validated in a real-world environment through experimentation that demonstrates accuracy in excess of 93% and energy consumption savings in excess of 94%.
Devki Nandan Jha, Zhenghua Chen, Shudong Liu 0003, Min Wu 0008, Jiahan Zhang, Graham Morgan, Rajiv Ranjan 0001, Xiaoli Li 0001
IEEE Trans. Sustain. Comput.7
2022 SAPPARCHI: an Osmotic Platform to Execute Scalable Applications on Smart City Environments
abstract
In the Smart Cities context, a plethora of Middle-ware Platforms had been proposed to support applications execution and data processing. Despite all the progress already made, the vast majority of solutions have not met the requirements of Applications’ Runtime, Development, and Deployment when related to Scalability. Some studies point out that just 1 of 97 (1%) reported platforms reach this all this set of requirements at same time. This small number of platforms may be explained by some reasons: i) Big Data: The huge amount of processed and stored data with various data sources and data types, ii) Multi-domains: many domains involved (Economy, Traffic, Health, Security, Agronomy, etc.), iii) Multiple processing methods like Data Flow, Batch Processing, Services, and Microservices, and 4) High Distributed Degree: The use of multiple IoT and BigData tools combined with execution at various computational levels (Edge, Fog, Cloud) leads applications to present a high level of distribution. Aware of those great challenges, we propose Sapparchi, an integrated architectural model for Smart Cities applications that defines multi-processing levels (Edge, Fog, and Cloud). Also, it presents the Sapparchi middleware platform for developing, deploying, and running applications in the smart city environment with an osmotic multi-processing approach that scales applications from Cloud to Edge. Finally, an experimental evaluation exposes the main advantages of adopting Sapparchi.
Arthur Souza 0001, Nélio Cacho, Thaís Vasconcelos Batista, Rajiv Ranjan 0001
CLOUD4
2022 Age of Data Aware Internet of Things Applications
abstract
The unprecedented growth of Internet of Things (IoT) underpinned by machine to machine communication, analytics and actuation is spearheading the development of autonomic IoT applications in areas such as Smart Cities. Such autonomic IoT applications have minimal human involvement in the decision making and actuation process. A key challenge in developing such autonomic IoT applications is uncertainty in the data produced by the IoT devices with data freshness being a critical aspect. In this paper, we address this challenge by introducing Age of Data (AoD), a metric to quantify the freshness of the data produced by IoT devices. We analyse the impact of AoD on IoT applications and propose a model for computing AoD that can be used by IoT applications in the decision making process. We validate the proposed model via experimental evaluations using real-world data obtained from parking sensors. Our analysis found that in real-world scenarios, 21.4% of sensors provide data that is outdated by several hours. We show that incorporating AoD in the application logic leads to improved application decision making.
Kaneez Fizza, Prem Prakash Jayaraman, Abhik Banerjee, Dimitrios Georgakopoulos 0001, Rajiv Ranjan 0001
CCNC5
2022 TinyRL: Towards Reinforcement Learning on Tiny Embedded Devices
abstract
We observe significant interest in reinforcement learning methods for real-world sensing-control scenarios driven by the sensor data streams. However, the delay introduced to the data by the communication channels may degrade the system's performance. It is especially crucial in the internet of things (IoT), where devices with constraint resources and low throughput networks are used.
Tomasz Szydlo, Prem Prakash Jayaraman, Yinhao Li 0003, Graham Morgan, Rajiv Ranjan 0001
CIKM5
2022 TinyML-CAM: 80 FPS image recognition in 1 kB RAM
abstract
TinyML: Tiny in size, big in impact! This paper presents TinyML-CAM pipeline for real-time and memory-efficient image recognition on IoT boards. TinyML-CAM can be used by developers to implement their customized tasks in ≈ 30 minutes, with minimal code configuration. We evaluated TinyML-CAM by using it to create a RandomForestClassifier (RF) based real-time image recognition system, which ran at 80 FPS and consumed only 1 kB of RAM on ESP32.
Bharath Sudharsan, Simone Salerno, Rajiv Ranjan 0001
MobiCom3
2022 CorrDetector: A framework for structural corrosion detection from drone images using ensemble deep learning
Abdur Forkan, Yong-Bin Kang, Prem Prakash Jayaraman, Kewen Liao, Rohit Kaul, Graham Morgan, Rajiv Ranjan 0001, Samir Sinha
Expert Syst. Appl.7
2022 Adaptive density peaks clustering: Towards exploratory EEG analysis
Tengfei Gao, Dan Chen 0001, Yunbo Tang, Bo Du 0001, Rajiv Ranjan 0001, Albert Y. Zomaya, Schahram Dustdar
Knowl. Based Syst.5
2022 SInC: Semantic approach and enhancement for relational data compression
Ruoyu Wang 0004, Daniel Sun 0004, Raymond K. Wong 0001, Rajiv Ranjan 0001, Albert Y. Zomaya
Knowl. Based Syst.4
2022 IoTSim-Osmosis-RES: Towards autonomic renewable energy-aware osmotic computing
abstract
Abstract Internet of Things systems exists in various areas of our everyday life. For example, sensors installed in smart cities and homes are processed in edge and cloud computing centers providing several benefits that improve our lives. The place of data processing is related to the required system response times—processing data closer to its source results in a shorter system response time. The osmotic computing concept enables flexible deployment of data processing services and their possible movement, just like particles in the osmosis phenomenon move between regions of different densities. At the same time, the impact of complex computer architecture on the environment is increasingly being compensated by the use of renewable and low‐carbon energy sources. However, the uncertainty of supplying green energy makes the management of osmotic computing demanding, and therefore their autonomy is desirable. In the article, we present a framework enabling osmotic computing simulation based on renewable energy sources and autonomic osmotic agents, allowing the analysis of distributed management algorithms. We discuss the challenges posed to the framework and analyze various management algorithms for cooperating osmotic agents. In the evaluation we show that changing the adaptation logic of the osmotic agents, it is possible to increase the self‐consumption of renewable energy sources or increase the usage of low emission ones.
Tomasz Szydlo, Amadeusz Szabala, Nazar Kordiumov, Konrad Siuzdak, Lukasz Wolski, Khaled Alwasel, Fawzy Habeeb, Rajiv Ranjan 0001
Softw. Pract. Exp.8
2022 AutoDiagn: An Automated Real-Time Diagnosis Framework for Big Data Systems
abstract
Big data processing systems, such as Hadoop and Spark, usually work in large-scale, highly-concurrent, and multi-tenant environments that can easily cause hardware and software malfunctions or failures, thereby leading to performance degradation. Several systems and methods exist to detect big data processing systems’ performance degradation, perform root-cause analysis, and even overcome the issues causing such degradation. However, these solutions focus on specific problems such as stragglers and inefficient resource utilization. There is a lack of a generic and extensible framework to support the real-time diagnosis of big data systems. In this article, we propose, develop and validate AutoDiagn. This generic and flexible framework provides holistic monitoring of a big data system while detecting performance degradation and enabling root-cause analysis. We present an implementation and evaluation of AutoDiagn that interacts with a Hadoop cluster deployed on a public cloud and tested with real-world benchmark applications. Experimental results show that AutoDiagn can offer a high accuracy root-cause analysis framework, at the same time as offering a small resource footprint, high throughput, and low latency.
Umit Demirbaga, Zhenyu Wen, Ayman Noor, Karan Mitra, Khaled Alwasel, Saurabh Kumar Garg 0001, Albert Y. Zomaya, Rajiv Ranjan 0001
IEEE Trans. Computers8
2022 Spatial-Keyword Skyline Publish/Subscribe Query Processing Over Distributed Sliding Window Streaming Data
abstract
Current spatial-keyword publish/subscribe systems need to handle spatial-keyword skyline queries over geo-textual streams to continuously obtain good results. The skyline queries in such systems face two main problems: (1) query problems, because the powerful query capability is required for the strict limit of the response time and the large number of items concerned by the users, and (2) scalability issue, because millions of active users are maintained simultaneously with many network-connected machines. Unfortunately, the current approach is towards static data. Thus, this paper first proposes a distributed skyline query processing framework. Then, we optimize the skyline computing by introducing MF-R$^t$-tree, which is an update-efficient and space-saving indexing structure and a fast approach for processing a continuous spatial-keyword skyline query called$eager^*$. Finally, a spatial and textual signature-based communication optimization method is proposed to support scalability. The experimental results indicate that (1) MF-R$^t$-tree can significantly reduce update costs, while maintaining a low storage cost, and a query performance comparable to IL-Quadtree, (2)$eager^*$can averagely accelerate 79.72 × faster than the method based on BNL, (3) the communication optimization method significantly reduces the communication cost, and (4) the distributed framework can efficiently support large-scale skyline queries.
Ze Deng, Schahram Dustdar, Rajiv Ranjan 0001, Albert Y. Zomaya, Lizhe Wang 0001
IEEE Trans. Computers5
2022 Lime: Low-Cost and Incremental Learning for Dynamic Heterogeneous Information Networks
abstract
Understanding the interconnected relationships of large-scale information networks like social, scholar and Internet of Things networks is vital for tasks like recommendation and fraud detection. The vast majority of the real-world networks are inherently heterogeneous and dynamic, containing many different types of nodes and edges and can change drastically over time. The dynamicity and heterogeneity make it extremely challenging to reason about the network structure. Unfortunately, existing approaches are inadequate in modeling real-life dynamical networks as they either have strong assumption of a given stochastic process or fail to capture the heterogeneity of network structure, and they all require extensive computational resources. We introduceLime, a better approach for modeling dynamic and heterogeneous information networks.Limeis designed to extract high-quality network representation with significantly lower memory resources and computational time over the state-of-the-arts. Unlike prior work that uses a vector to encode each network node, we exploit the semantic relationships among network nodes to encode multiple nodes with similar semantics in shared vectors. By using many fewer node vectors, our approach significantly reduces the required memory space for encoding large-scale networks. To effectively trade information sharing for reduced memory footprint, we employ the recursive neural network (RsNN) with carefully designed optimization strategies to explore the node semantics in a novel cuboid space. We then go further by showing, for the first time, how an effective incremental learning approach can be developed – with the help of RsNN, our cuboid structure, and a set of novel optimization techniques – to allow a learning framework to quickly and efficiently adapt to a constantly evolving network. We evaluateLimeby applying it to three representative network-based tasks, node classification, node clustering and anomaly detection, performing on three large-scale datasets. We compareLimeagainst eleven prior state-of-the-art approaches for learning network representation. Our extensive experiments demonstrate thatLimenot only reduces the memory footprint by over 80 percent and the processing time over 2x when learning network representation but also delivers comparable performance for downstream processing tasks. We show that our incremental learning method can boost the learning time by up to 20x without compromising the quality of the learned network representation.
Hao Peng 0001, Renyu Yang, Zheng Wang 0001, Jianxin Li 0002, Lifang He 0001, Philip S. Yu, Albert Y. Zomaya, Rajiv Ranjan 0001
IEEE Trans. Computers8
2022 Privacy-Aware Cloud Auditing for GDPR Compliance Verification in Online Healthcare
abstract
Emerging multitenant cloud computing ecosystems allow multiple applications to share virtualized pool of computing and networking resources. As a result, such ecosystems are becoming increasingly prone to data privacy concerns (personal data leakages and unauthorized access). While cloud computing providers support robust security and privacy mechanisms (e.g., public key cryptography, firewalls, and virtual private networks, among many others), they lack mechanisms and frameworks to monitor, audit, and verify these data privacy concerns. The emergence of data protection regulations around the world, such as General Data Protection Regulation in Europe and the Data Protection Act in the U.K., further emphasizes the need to overcome these privacy limitations. In this article, a novel technique for monitoring, auditing, and verifying the operations carried out on a user’s personal data in cloud computing ecosystems is proposed. Our research methodology leverages distributed ledger technologies (e.g., blockchain and smart contracts) for developing an immutable recording technique, which transparently logs, monitors, and verifies the operations carried out on user data. Using a healthcare pharmacy scenario and extensive real-world experiments, we validate the feasibility of the proposed technique. The proposed work handles a large pool of requests ($>$13K) ensuring minimal latency ($\approx$50–60 ms) and overheads for three different service packages varied with respect to the number of actors and operations.
Masoud Barati, Gagangeet Singh Aujla, Jose Tomas Llanos, Kwabena Adu-Duodu, Omer F. Rana, Madeline Carr, Rajiv Ranjan 0001
IEEE Trans. Ind. Informatics7
2022 Dynamic Bandwidth Slicing for Time-Critical IoT Data Streams in the Edge-Cloud Continuum
abstract
Edge computing has gained momentum in recent years, as complementary to cloud computing, for supporting applications (e.g., industrial control systems) that require time-critical communication guarantees. While edge computing can provide immediate analysis of streaming data from Internet of Things devices, those devices lack computing capabilities to guarantee reasonable performance for time-critical applications. To alleviate this critical problem, the prevalent trend is to offload these data analytic tasks from the edge devices to the cloud. However, existing offloading approaches are static in nature as they are unable to adapt varying workload and network conditions. To handle these issues, we present a novel distributed and quality of services based multilevel queue traffic scheduling system that can undertake semiautomatic bandwidth slicing to process time-critical incoming traffic in the edge-cloud environments. Our developed system shows a great enhancement in latency and throughput as well as reduction in energy consumption for edge-cloud environments.
Fawzy Habeeb, Khaled Alwasel, Ayman Noor, Devki Nandan Jha, Duaa S. Alqattan, Yinhao Li 0003, Gagangeet Singh Aujla, Tomasz Szydlo, Rajiv Ranjan 0001
IEEE Trans. Ind. Informatics9
2022 Multi-scale Features Fusion for the Detection of Tiny Bleeding in Wireless Capsule Endoscopy Images
abstract
Wireless capsule endoscopy is a modern non-invasive Internet of Medical Imaging Things that has been increasingly used in gastrointestinal tract examination. With about one gigabyte image data generated for a patient in each examination, automatic lesion detection is highly desirable to improve the efficiency of the diagnosis process and mitigate human errors. Despite many approaches for lesion detection have been proposed, they mainly focus on large lesions and are not directly applicable to tiny lesions due to the limitations of feature representation. As bleeding lesions are a common symptom in most serious gastrointestinal diseases, detecting tiny bleeding lesions is extremely important for early diagnosis of those diseases, which is highly relevant to the survival, treatment, and expenses of patients. In this article, a method is proposed to extract and fuse multi-scale deep features for detecting and locating both large and tiny lesions. A feature extracting network is first used as our backbone network to extract the basic features from wireless capsule endoscopy images, and then at each layer multiple regions could be identified as potential lesions. As a result, the features maps of those potential lesions are obtained at each level and fused in a top-down manner to the fully connected layer for producing final detection results. Our proposed method has been evaluated on a clinical dataset that contains 20,000 wireless capsule endoscopy images with clinical annotation. Experimental results demonstrate that our method can achieve 98.9% prediction accuracy and 93.5% score, which has a significant performance improvement of up to 31.69% and 22.12% in terms of recall rate and score, respectively, when compared to the state-of-the-art approaches for both large and tiny bleeding lesions. Moreover, our model also has the highest AP and the best medical diagnosis performance compared to state-of-the-art multi-scale models.
Feng Lu 0003, Wei Li 0058, Chengwangli Peng, Zhiyong Wang 0001, Bin Qian 0002, Rajiv Ranjan 0001, Hai Jin 0001, Albert Y. Zomaya
ACM Trans. Internet Things7
2022 EDCSuS: Sustainable Edge Data Centers as a Service in SDN-Enabled Vehicular Environment
abstract
Cloud computing has emerged as one of the popular technologies which provide on-demand services to the end users. Such services are hosted by massive geo-distributed data centers (DCs). Nowadays, connected vehicles in a smart city can also avail cloud services through Internet using cellular technologies. But, the advent of 5G technology has posed challenges for DCs such as-low latency and higher data rate requirements. To handle these challenges, edge-DCs (EDCs) can be deployed across a smart city to provide low latency services to the connected vehicles. In lieu of this, in this paper, EDCSuS: Sustainable EDC as a service framework in software defined vehicular environment is proposed. In EDCSuS, first, a software defined controller handles the incoming requests and suggest an optimal flow path. Second, a multi-leader multi-follower Stackelberg game is presented for resource allocation. Third, to improve the resource utilization, a cooperative resource sharing scheme is designed, thereby minimizing the energy consumption of servers in the EDCs. Lastly, a caching scheme is presented to avert excessive energy consumption for retracing the lost link due to vehicular mobility. The efficacy of the proposed scheme has been evaluated using extensive simulations with respect to various parameters. The results obtained prove the effectiveness of EDCSuS.
Gagangeet Singh Aujla, Neeraj Kumar 0001, Sahil Garg, Kuljeet Kaur, Rajiv Ranjan 0001
IEEE Trans. Sustain. Comput.5
2021 Working in a Smart Home-office: Exploring the Impacts on Productivity and Wellbeing
abstract
Following the outbreak of the Coronavirus (COVID-19) pandemic, many organisations have shifted to remote working overnight. The new reality has created conditions to use smart home technologies for work purposes, for which they were not originally intended. The lack of insights into the new application of smart home technologies has led to two research objectives. First, the paper aimed to investigate the factors correlating with productivity and perceived wellbeing. Second, the study tried to explore individuals’ intentions to use smart home offices for remote work in the future. 528 responses were gathered from individuals who had smart homes and had worked from home during the pandemic. The results showed that productivity positively relates to service relevance, perceived usefulness, perceived ease of use, hedonic beliefs, control over environmental conditions, innovativeness and attitude. Task-technology fit, service relevance, attitude to smart homes, innovativeness, hedonic beliefs, perceived usefulness, perceived ease of use and control over environmental conditions correlate with perceived wellbeing. The intention to work from smart home-offices in the future is determined by perceived wellbeing. Findings contribute to the research on smart homes and remote work practices, by providing the first empirical evidence about the new applications and outcomes of smart home use in the work context.
Davit Marikyan, Savvas Papagiannidis, Rajiv Ranjan 0001, Omer F. Rana
WEBIST3
2021 A study on the evaluation of HPC microservices in containerized environment
abstract
Summary Containers are gaining popularity over virtual machines as they provide the advantages of virtualization with the performance of near bare metal. The uniformity of support provided by Docker containers across different cloud providers makes them a popular choice for developers. Evolution of microservice architecture allows complex applications to be structured into independent modular components making them easier to manage. High‐performance computing (HPC) applications are one such application to be deployed as microservices, placing significant resource requirements on the container framework. However, there is a possibility of interference between different microservices hosted within the same container (intracontainer) and different containers (intercontainer) on the same physical host. In this paper, we describe an extensive experimental investigation to determine the performance evaluation of Docker containers executing heterogeneous HPC microservices. We are particularly concerned with how intracontainer and intercontainer interference influences the performance. Moreover, we investigate the performance variations in Docker containers when control groups (cgroups) are used for resource limitation. For ease of presentation and reproducibility, we use Cloud Evaluation Experiment Methodology (CEEM) to conduct our comprehensive set of experiments. We expect that the results of evaluation can be used in understanding the behavior of HPC microservices in the interfering containerized environment.
Devki Nandan Jha, Saurabh Kumar Garg 0001, Prem Prakash Jayaraman, Rajkumar Buyya, Zheng Li 0001, Graham Morgan, Rajiv Ranjan 0001
Concurr. Comput. Pract. Exp.7
2021 IoTSim-Osmosis: A framework for modeling and simulating IoT applications over an edge-cloud continuum
Khaled Alwasel, Devki Nandan Jha, Fawzy Habeeb, Umit Demirbaga, Omer F. Rana, Thar Baker, Schahram Dustdar, Massimo Villari, Philip James 0002, Ellis Solaiman, Rajiv Ranjan 0001
J. Syst. Archit.11
2021 BigDataSDNSim: A simulator for analyzing big data applications in software-defined cloud data centers
abstract
Abstract The integration and crosscoordination of big data processing and software‐defined networking (SDN) are vital for improving the performance of big data applications. Various approaches for combining big data and SDN have been investigated by both industry and academia. However, empirical evaluations of solutions that combine big data processing and SDN are extremely costly and complicated. To address the problem of effective evaluation of solutions that combine big data processing with SDN, we present a new, self‐contained simulation tool named BigDataSDNSim that enables the modeling and simulation of the big data management system YARN, its related programming models MapReduce, and SDN‐enabled networks in a cloud computing environment. BigDataSDNSim supports cost‐effective and easy to conduct experimentation in a controllable, repeatable, and configurable manner. The article illustrates the simulation accuracy and correctness of BigDataSDNSim by comparing the behavior and results of a real environment that combines big data processing and SDN with an equivalent simulated environment. Finally, the article presents two uses cases of BigDataSDNSim, which exhibit its practicality and features, illustrate the impact of data replication mechanisms of MapReduce in Hadoop YARN, and show the superiority of SDN over traditional networks to improve the performance of MapReduce applications.
Khaled Alwasel, Rodrigo N. Calheiros, Saurabh Kumar Garg 0001, Rajkumar Buyya, Mukaddim Pathan, Dimitrios Georgakopoulos 0001, Rajiv Ranjan 0001
Softw. Pract. Exp.7
2021 BaPa: A Novel Approach of Improving Load Balance in Parallel Matrix Factorization for Recommender Systems
abstract
A simplified approach to accelerate matrix factorization of big data is to parallelize it. A commonly used method is to divide the matrix into multiple non-intersecting blocks and concurrently calculate them. This operation causes the Load balance problem, which significantly impacts parallel performance and is a big concern. A general belief is that the load balance across blocks is impossible by balancing rows and columns separately. We challenge the belief by proposing an approach of “Balanced Partitioning (BaPa)”. We demonstrate under what circumstance independently balancing rows and columns can lead to the balanced intersection of rows and columns, why, and how. We formally prove the feasibility of BaPa by observing the variance of rating numbers across blocks, and empirically validate its soundness by applying it to two standard parallel matrix factorization algorithms, DSGD and CCD++. Besides, we establish a mathematical model of “Imbalance Degree” to explain further why BaPa works well. BaPa is applied to synchronous parallel matrix factorization, but as a general load balance solution, it has significant application potential.
Ruixin Guo, Feng Zhang 0012, Lizhe Wang 0001, Wusheng Zhang, Xinya Lei, Rajiv Ranjan 0001, Albert Y. Zomaya
IEEE Trans. Computers6
2021 Detection of SLA Violation for Big Data Analytics Applications in Cloud
abstract
SLA violations do happen in real world. An SLA violation represents the failure of guaranteeing a service, which leads to unwanted consequences such as penalty payments, profit margin reduction, reputation degradation, customer churn and service interruptions. Hence, in the context of cloud-hosted big data analytics applications (BDAAs), it is paramount for providers to predict and prevent SLA violations. While machine learning-based techniques have been applied to detect SLA violations for web service or general cloud service, the study on detecting SLA violations dedicated for cloud-hosted BDAAs is still lacking. In this article, we propose four machine learning techniques and integrate 12 resampling methods to detect SLA violations for batch-based BDAAs in the cloud. We evaluate the efficiency of the proposed techniques in comparison with ideal and baseline classifiers based on a real-world trace dataset (Alibaba). Our work not only helps providers to choose the best performing prediction technique, but also provides them capabilities to uncover the hidden pattern of multiple configurations of BDAAs across layers.
Xuezhi Zeng, Saurabh Kumar Garg 0001, Mutaz Barika, Sanat Kumar Bista, Deepak Puthal, Albert Y. Zomaya, Rajiv Ranjan 0001
IEEE Trans. Computers7
2021 Running Industrial Workflow Applications in a Software-Defined Multicloud Environment Using Green Energy Aware Scheduling Algorithm
abstract
Industry 4.0 have automated the entire manufacturing sector (including technologies and processes) by adopting Internet of Things and cloud computing. To handle the workflows from Industrial Cyber-Physical systems, more and more data centers have been built across the globe to serve the growing needs of computing and storage. This has led to an enormous increase in energy usage by cloud data centers, which is not only a financial burden but also increases their carbon footprint. The private software defined wide area network (SDWAN) connects a cloud provider's data centers across the planet. This gives the opportunity to develop new scheduling strategies to manage cloud providers workload in a more energy-efficient manner. In this context, this article addresses the problem of scheduling data-driven industrial workflow applications over a set of private SDWAN connected data centers in an energy-efficient manner while managing tradeoff of a cloud provider' revenue. Our proposed algorithm aims to minimize the cloud provider's revenue and the usage of nonrenewable energy by utilizing the real-world electricity prices with the availability of green energy on different cloud data centers, where the energy consumption consists of the usage of running application over multiple data centers and transferring the data among them through SDWAN. The evaluation shows that our proposed method can increase usage of green energy for the execution of industrial workflow up to 3× times with a slight increase in the cost when compared to cost-based workflow scheduling methods.
Zhenyu Wen, Saurabh Kumar Garg 0001, Gagangeet Singh Aujla, Khaled Alwasel, Deepak Puthal, Schahram Dustdar, Albert Y. Zomaya, Rajiv Ranjan 0001
IEEE Trans. Ind. Informatics8
2021 A lightweight solution to epileptic seizure prediction based on EEG synchronization measurement
Dan Chen 0001, Rajiv Ranjan 0001, Hengjin Ke, Yunbo Tang, Albert Y. Zomaya
J. Supercomput.3
2021 Streaming Social Event Detection and Evolution Discovery in Heterogeneous Information Networks
abstract
Events are happening in real world and real time, which can be planned and organized for occasions, such as social gatherings, festival celebrations, influential meetings, or sports activities. Social media platforms generate a lot of real-time text information regarding public events with different topics. However, mining social events is challenging because events typically exhibit heterogeneous texture and metadata are often ambiguous. In this article, we first design a novel event-based meta-schema to characterize the semantic relatedness of social events and then build an event-based heterogeneous information network (HIN) integrating information from external knowledge base. Second, we propose a novel Pairwise Popularity Graph Convolutional Network, named as PP-GCN, based on weighted meta-path instance similarity and textual semantic representation as inputs, to perform fine-grained social event categorization and learn the optimal weights of meta-paths in different tasks. Third, we propose a streaming social event detection and evolution discovery framework for HINs based on meta-path similarity search, historical information about meta-paths, and heterogeneous DBSCAN clustering method. Comprehensive experiments on real-world streaming social text data are conducted to compare various social event detection and evolution discovery algorithms. Experimental results demonstrate that our proposed framework outperforms other alternative social event detection and evolution discovery techniques.
Hao Peng 0001, Jianxin Li 0002, Yangqiu Song, Renyu Yang, Rajiv Ranjan 0001, Philip S. Yu, Lifang He 0001
ACM Trans. Knowl. Discov. Data5
2021 Online Scheduling Technique To Handle Data Velocity Changes in Stream Workflows
abstract
Many IoT applications and services such as smart parking and smart traffic control contain a network of different analytical components, which are composed in the form of a workflow to make better decisions. These workflows are also known as stream workflows. The focus of existing research works is on the streaming operator graph, which differs from stream workflow application as it involves heterogeneity, multiple data sources and multiple outputs. Considering the complexity and dynamism of stream workflow, meeting real-time data analysis requirements at deployment time is not the whole story as the velocity of data changes over time. This change is the most dynamic form of stream workflow that occurs frequently during the execution of this application. In this article, we propose a new dynamic scheduling technique that manages cloud resources over time to handle data velocity changes in stream workflow while maintaining user-defined real-time data analysis requirements and minimising execution cost. The efficiency of the proposed technique is evaluated, and experimental results showed that this technique outperformed its competitors and is close to the lower bound.
Mutaz Barika, Saurabh Kumar Garg 0001, Albert Y. Zomaya, Rajiv Ranjan 0001
IEEE Trans. Parallel Distributed Syst.4
2020 TOPOSCH: Latency-Aware Scheduling Based on Critical Path Analysis on Shared YARN Clusters
abstract
Balancing resource utilization and application QoS is a long-standing research topic in cluster resource management. Big data YARN clusters need to co-schedule diverse workloads on shared resources including batch processing jobs, streaming jobs, and other long-running applications such as web services, database services, etc. Current resource managers are only responsible for resource allocation among applications/jobs but completely unaware of runtime QoS requirements of interactive and latency-sensitive applications. Prior works to maximize the QoS of monolithic applications ignore inherent dependencies and temporal-spatio performance variability of components, characteristics of distributed applications primarily driven by microservices. In this paper, we present Toposch, a new resource management system to adaptively co-locate batch tasks and microservices by harvesting runtime latency. In particular, Toposch tracks full footprints of every request across microservices over time. A latency graph is periodically generated for identifying victim microservices through an end-to-end latency critical path analysis. We then exploit per-microservice and per-node risk assessment to gauge the visible resources to the capacity scheduler in YARN. Execution of batch tasks are adaptively throttled or delayed, thereby avoiding latency increase due to node over-saturation. TOPOSCH is integrated with YARN and experiments show that the latency of DLRAs can be reduced by up to 39.8% against the default capacity scheduling in YARN.
Chunming Hu, Jianyong Zhu, Renyu Yang, Hao Peng 0001, Tianyu Wo, Shiqing Xue, Xiaoqiang Yu, Jie Xu 0007, Rajiv Ranjan 0001
CLOUD9
2020 Active Hazard Observation via Human in the Loop Social Media Analytics System
abstract
We demonstrate AHOM, a system that can Actively Observe Hazards via Monitoring Social Media Streams. AHOM proposes an active way to include the human in the loop of hazard information ac-quisition for social media. Different from state of the art, it supports bi-directional interaction between social media data processing system and social media users, which leads to the establishment of deeper and more accurate situational awareness of hazard events. We demonstrate how AHOM utilizes Twitter streams and bi-directional information exchange with social media users for enhanced hazard observation.
Zhenyu Wen, Jedsada Phengsuwan, Nipun Balan Thekkummal, Rui Sun 0010, Pooja jamathi-Chidananda, Tejal Shah, Philip James 0002, Rajiv Ranjan 0001
CIKM8
2020 New Horizons in IoT Workflows Provisioning in Edge and Cloud Datacentres for Fast Data Analytics: The Osmotic Computing Approach
Rajiv Ranjan 0001
CLOSER1
2020 New Horizons in IoT Workflows Provisioning in Edge and Cloud Datacentres for Fast Data Analytics: The Osmotic Computing Approach
Rajiv Ranjan 0001
COMPLEXIS1
2020 IoTWC: Analytic Hierarchy Process Based Internet of Things Workflow Composition System
abstract
Internet of Things (IoT) allows the creation of virtually endless connections into a global array of distributed intelligence. However, the design, development, and deployment of IoT applications are complex and complicated due to various unwarranted challenges. For instance, addressing the IoT application users' subjective and objective opinions with IoT workflow instances remains a challenge for the design of a more holistic approach. Moreover, the complexity of IoT applications increased exponentially due to the heterogeneous nature of the Edge/Cloud services, utilised with the aim of lowering latency in data transformation and increase re-usability. Hence, in this paper, we present an IoT workflow composition system (IoTWC) to allow IoT users to pipeline their workflows with proposed IoT workflow activity abstract patterns. IoTWC leverages the analytic hierarchy process (AHP) to compose the multi-level IoT workflow that satisfies the requirements of any IoT application. Moreover, the users are befitted with recommended IoT workflow configurations using an AHP based multi-level composition framework. The proposed IoTWC is validated on a user case study to evaluate the coverage of IoT workflow activity abstract patterns and a real-world scenario for smart buildings. The comprehensive analysis shows the effectiveness of IoTWC in terms of IoT workflow abstraction and composition.
Yinhao Li 0003, Devki Nandan Jha, Gagangeet Singh Aujla, Graham Morgan, Albert Y. Zomaya, Rajiv Ranjan 0001
IC2E6
2020 Priority-based Fair Scheduling in Edge Computing
abstract
Scheduling is important in Edge computing. In contrast to the Cloud, Edge resources are hardware limited and cannot support workload-driven infrastructure scaling. Hence, resource allocation and scheduling for the Edge requires a fresh perspective. Existing Edge scheduling research assumes availability of all needed resources whenever a job request is made. This paper challenges that assumption, since not all job requests from a Cloud server can be scheduled on an Edge node. Thus, guaranteeing fairness among the clients (Cloud servers offloading jobs) while accounting for priorities of the jobs becomes a critical task. This paper presents four scheduling techniques, the first is a naive first come first serve strategy and further proposes three strategies, namely a client fair, priority fair, and hybrid that accounts for the fairness of both clients and job priorities. An evaluation on a target platform under three different scenarios, namely equal, random, and Gaussian job distributions is presented. The experimental studies highlight the low overheads and the distribution of scheduled jobs on the Edge node when compared to the naive strategy. The results confirm the superior performance of the hybrid strategy and showcase the feasibility of fair schedulers for Edge computing.
Arkadiusz Madej, Nan Wang 0009, Nikolaos Athanasopoulos, Rajiv Ranjan 0001, Blesson Varghese
ICFEC4
2020 New Horizons in IoT Workflows Provisioning in Edge and Cloud Datacentres for Fast Data Analytics: The Osmotic Computing Approach
Rajiv Ranjan 0001
IoTBDS1
2020 MobDL: A Framework for Profiling Deep Learning Models: A Case Study using Mobile Digital Health Applications
abstract
Smart mobile devices coupled with the Internet of Things (IoT) and Artificial Intelligence (AI) have emerged as a key enabler of modern digital health applications. While cloud computing is now a well established paradigm for analysing IoT captured data in mobile health applications, on-board analysis of data using AI approaches such as Deep Learning (DL) is gaining significant momentum. This is driven primarily by advances in on-board resources enabling modern mobile devices to execute complex DL models, while also offering improved response time and accuracy for rapid decision-making, and enhanced user privacy. While the number of mobile digital health applications that use IoT and DL is increasing, progress is currently impeded by a lack of framework for profiling and evaluating the performance of DL models on mobile devices. To this end, we propose MobDL, a framework for profiling and evaluating DL models running on smart mobile devices. We present the architecture of this framework and devise a novel evaluation methodology for conducting quantitative comparisons of various DL models running on mobile devices. Three diverse digital health applications using heterogeneous data (e.g. image, time series) are introduced. We conduct extensive experimental evaluations using several DL models that have been developed using the data sets obtained for the three digital health applications to validate the effectiveness of the proposed MobDL framework.
Abdur Forkan, Prem Prakash Jayaraman, Rohit Kaul, Yuxin Zhang 0001, Chris McCarthy, Pari Delir Haghighi, Rajiv Ranjan 0001
MobiQuitous7
2020 Cost effective stream workflow scheduling to handle application structural changes
Mutaz Barika, Saurabh Kumar Garg 0001, Rajiv Ranjan 0001
Future Gener. Comput. Syst.3
2020 Stochastic scheduling for variation-aware virtual machine placement in a cloud computing CPS
Yunliang Chen 0002, Xiaodao Chen, Wangyang Liu, Yuchen Zhou 0003, Albert Y. Zomaya, Rajiv Ranjan 0001, Shiyan Hu 0001
Future Gener. Comput. Syst.6
2020 A note on advances in scheduling algorithms for Cyber-Physical-Social workflows
Rajiv Ranjan 0001, Lydia Y. Chen, Prem Prakash Jayaraman, Albert Y. Zomaya
Future Gener. Comput. Syst.1
2020 ESMLB: Efficient Switch Migration-Based Load Balancing for Multicontroller SDN in IoT
abstract
In software-defined networks (SDNs), the deployment of multiple controllers improves the reliability and scalability of the distributed control plane. Recently, edge computing (EC) has become a backbone to networks where computational infrastructures and services are getting closer to the end user. The unique characteristics of SDN can serve as a key enabler to lower the complexity barriers involved in EC, and provide better quality-of-services (QoS) to users. As the demand for IoT keeps growing, gradually a huge number of smart devices will be connected to EC and generate tremendous IoT traffic. Due to a huge volume of control messages, the controller may not have sufficient capacity to respond to them. To handle such a scenario and to achieve better load balancing, dynamic switch migrating is one effective approach. However, a deliberate mechanism is required to accomplish such a task on the control plane, and the migration process results in high network delay. Taking it into consideration, this article has introduced an efficient switch migration-based load balancing (ESMLB) framework, which aims to assign switches to an underutilized controller effectively. Among many alternatives for selecting a target controller, a multicriteria decision-making method, i.e., the technique for order preference by similarity to an ideal solution (TOPSIS), has been used in our framework. This framework enables flexible decision-making processes for selecting controllers having different resource attributes. The emulation results indicate the efficacy of the ESMLB.
Kshira Sagar Sahoo, Deepak Puthal, Mayank Tiwari 0003, Muhammad Usman 0015, Bibhudatta Sahoo 0001, Zhenyu Wen, B. P. S. Sahoo, Rajiv Ranjan 0001
IEEE Internet Things J.8
2020 COMITMENT: A Fog Computing Trust Management Approach
Mohammed Al-Khafajiy, Thar Baker, Muhammad Asim 0001, Zehua Guo 0001, Rajiv Ranjan 0001, Antonella Longo, Deepak Puthal, Mark Taylor 0005
J. Parallel Distributed Comput.5
2020 IoTSim-SDWAN: A simulation framework for interconnecting distributed datacenters over Software-Defined Wide Area Network (SD-WAN)
Khaled Alwasel, Devki Nandan Jha, Deepak Puthal, Mutaz Barika, Blesson Varghese, Saurabh Kumar Garg 0001, Philip James 0002, Albert Y. Zomaya, Graham Morgan, Rajiv Ranjan 0001
J. Parallel Distributed Comput.11
2020 En-ABC: An ensemble artificial bee colony based anomaly detection scheme for cloud environment
Sahil Garg, Kuljeet Kaur, Shalini Batra, Gagangeet Singh Aujla, Graham Morgan, Neeraj Kumar 0001, Albert Y. Zomaya, Rajiv Ranjan 0001
J. Parallel Distributed Comput.8
2020 A general purpose contention manager for software transactions on the GPU
Craig Sharp, Richard Davison 0001, Gary Ushaw, Rajiv Ranjan 0001, Albert Y. Zomaya, Graham Morgan
J. Parallel Distributed Comput.5
2020 IoTSim-Edge: A simulation framework for modeling the behavior of Internet of Things and edge computing environments
abstract
Summary With the proliferation of Internet of Things (IoT) and edge computing paradigms, billions of IoT devices are being networked to support data‐driven and real‐time decision making across numerous application domains, including smart homes, smart transport, and smart buildings. These ubiquitously distributed IoT devices send the raw data to their respective edge device (eg, IoT gateways) or the cloud directly. The wide spectrum of possible application use cases make the design and networking of IoT and edge computing layers a very tedious process due to the: (i) complexity and heterogeneity of end‐point networks (eg, Wi‐Fi, 4G, and Bluetooth); (ii) heterogeneity of edge and IoT hardware resources and software stack; (iv) mobility of IoT devices; and (iii) the complex interplay between the IoT and edge layers. Unlike cloud computing, where researchers and developers seeking to test capacity planning, resource selection, network configuration, computation placement, and security management strategies had access to public cloud infrastructure (eg, Amazon and Azure), establishing an IoT and edge computing testbed that offers a high degree of verisimilitude is not only complex, costly, and resource‐intensive but also time‐intensive. Moreover, testing in real IoT and edge computing environments is not feasible due to the high cost and diverse domain knowledge required in order to reason about their diversity, scalability, and usability. To support performance testing and validation of IoT and edge computing configurations and algorithms at scale, simulation frameworks should be developed. Hence, this article proposes a novel simulator IoTSim‐Edge, which captures the behavior of heterogeneous IoT and edge computing infrastructure and allows users to test their infrastructure and framework in an easy and configurable manner. IoTSim‐Edge extends the capability of CloudSim to incorporate the different features of edge and IoT devices. The effectiveness of IoTSim‐Edge is described using three test cases. Results show the varying capability of IoTSim‐Edge in terms of application composition, battery‐oriented modeling, heterogeneous protocols modeling, and mobility modeling along with the resources provisioning for IoT applications.
Devki Nandan Jha, Khaled Alwasel, Areeb Alshoshan, Xianghua Huang, Ranesh Kumar Naha, Sudheer Kumar Battula, Saurabh Kumar Garg 0001, Deepak Puthal, Philip James 0002, Albert Y. Zomaya, Schahram Dustdar, Rajiv Ranjan 0001
Softw. Pract. Exp.12
2020 Software tools and techniques for fog and edge computing
abstract
The Internet of Things (IoT) paradigm promises to make “things” such as physical objects with sensing capabilities and/or attached with tags, mobile objects such as smartphones and vehicles, consumer electronic devices, and home appliances such as fridge, television, health care devices, as part of the Internet environment. In cloud-centric IoT applications, the sensor data from these “things” is extracted, accumulated, and processed at the public/private clouds, leading to significant latencies. To satisfy the ever increasing demand for cloud computing resources from emerging applications such as IoT, academics and industry experts are now advocating for going from large-centralized cloud computing infrastructures to micro data centers located at the edge of the network. These micro data centers are often closer to a user (geographically and in access latency) compared to the centralized cloud data center. The aim of utilizing such edge resources is to off load computation that would have “traditionally” been carried out at the cloud data center to a resource that is closer to a user or edge devices. This vision also acknowledges the variation in network latency from an end-user to cloud data center. While the network around a data center is often high capacity and speed, that near the user device may have variable properties (in terms of resilience, bandwidth, latency, etc.). Referred to as “fog/edge computing,” this paradigm is expected to improve the agility of cloud service deployments in addition to bringing computing resources closer to end-users. The emergence of computing paradigms such as edge and fog computing supports the data analysis near the data sources for a wide range of applications. Edge computing is the middle layer between users and cloud data centers, and it plays an important role in the IoT use cases where applications required near real-time actions. The intermediate edge layer provides limited computing and storage resources, which consists of network gateways and micro data centers. Edge computing application orchestration is capable of big data processing and can be installed in heterogeneous hardware configurations. Due to the large-scale deployment and device heterogeneity, edge computing infrastructure designing and implementation are challenging, including model analysis, system integration, protocol designing, energy, and security modeling. In addition, since edge data centers are installed in network gateways with open network configuration, they are prone to several network threats, less trustworthy, and easy to compromise. On the one hand, the development of fog and edge clouds includes dedicated facilities, operating system, network, and middleware techniques to build and operate such micro data centers that host virtualized computing resources. On the other hand, the use of fog and edge clouds requires extension to current programming models and proposes new abstractions that will allow developers to design new applications that take benefit from such massively distributed systems. The use of this approach also opens up other challenges in security and privacy (as a user now needs to “trust” every micro data center they interact with), support for resource management for mobile users who transfer session from one micro data center to another, and support for “embedding” such micro data centers into devices (eg, cars, buildings, etc). The objective of this special issue is to disseminate original contributions and research findings concerning the challenges and changes (both evolutionary and disruptive) in edge and fog computing. It provides cutting-edge research from both academia and industry, with emphasis on current developments and future directions in security and privacy issues of emerging fog computing. The call for special issues received a number of submissions. Each paper was reviewed by at least three reviewers and went through at least two rounds of reviews. After a two-phase peer review process, we have accepted 14 high-quality papers related to the aforementioned areas of interest. The accepted papers focus on recent solutions by developing novel research ideas around edge and fog computing for several applications, such as health care, smart city, urban pollution monitoring, etc. The brief contributions of these papers are discussed in the following section. The first paper titled “Abnormal visual event detection based on multi-instance learning and autoregressive integrated moving average model in edge-based Smart City surveillance” by Xu et al proposes an abnormal event detection approach based on multi-instance learning and autoregressive integrated moving average model for video surveillance of crowded scenes in urban public places. It utilizes an unsupervised method for abnormal event detection by combining multi-instance visual feature selection and the autoregressive integrated moving average model. This approach has thoroughly experimented, and the experimental results demonstrate the efficiency of the proposed approach by achieving better abnormal event detection performance for a crowded scene of urban public places with an edge environment. The second paper titled “User allocation-aware edge cloud placement in mobile edge computing” by Guo et al studies the edge cloud placement problem, which is to place the edge clouds at the candidate locations and allocate the mobile users to the edge clouds. Further, it formulates as a multi-objective optimization problem with the objective to balance the workload between edge clouds and minimize the service communication delay of mobile users. The experiment results show the performance of the proposed approach in terms of workload balance and communication delay for validation. The third paper titled “A secure fog-based platform for SCADA-based IoT critical infrastructure” by Baker et al contributes a novel security “toolbox” to reinforce the integrity, security, and privacy of SCADA-based IoT critical infrastructure at the fog layer. The toolbox incorporates a key feature, that is, a cryptographic-based access approach to the cloud services using identity-based cryptography and signature schemes at the fog layer. This paper also presents the implementation details of a prototype for our proposed secure fog-based platform and provides performance evaluation results to demonstrate the appropriateness of the proposed platform in a real-world scenario. The results from the experiments demonstrate a superior performance of the secure fog-based platform, which is around 2.8 seconds when adding five virtual machines (VMs), 3.2 seconds when adding 10 VMs, and 112 seconds when adding 1000 VMs, compared to the multilevel user access control platform. The fourth paper titled “Developing applications in large scale, dynamic fog computing: A case study” by Giang et al presents a case study in building fog computing applications using an open-source platform distributed node-RED. It shows how applications can be decomposed and deployed to a geographically distributed infrastructure using distributed node-RED, and how existing software components can be adapted and reused to participate in fog applications. This case study is implemented in a lab-based fog infrastructure and simulated for large-scale evaluation. The fifth paper titled “An osmotic computing infrastructure for urban pollution monitoring” by Longo et al focuses on the design and development of a middleware that integrates data coming from mobile and IoT devices specifically deployed in urban contexts using the osmotic computing paradigm. Moreover, a component of the osmotic membrane has been developed in this paper for security management. The sixth paper titled “Characterizing application scheduling on edge, fog, and cloud computing resources” by Varshney and Simmhan offers a taxonomy of concepts essential for specifying and solving the problem of scheduling applications on edge, fog, and cloud computing resources. The proposed model is divided into multiple steps, initially characterized by the resource capabilities and limitations of these infrastructures and offers a taxonomy of application models, quality-of-service constraints and goals, and scheduling techniques based on a literature review, followed by tabulated key research prototypes and papers using this taxonomy. It also highlights gaps in the literature and open problems remain. The seventh paper titled “Cloud-aided online electroencephalography (EEG) classification system for brain healthcare: A case study of depression evaluation with a lightweight CNN” by Ke et al presents the design of an online EEG classification system aided by cloud centering on a lightweight convolutional neural network (CNN). The system incrementally trains the CNN on cloud and enables hot deployment of the trained classifier without the need to restart the gateway to adapt to the users' needs. The classifier maintains a high convolutional layer to gain the ability of processing high-dimensional EEG segments. Finally, the model is experimented to validate the contribution. The eighth paper titled “SEWMS: An Edge-based Smart Wearable Maintenance System in Communication Network” by Rui et al proposes a dynamic context-aware information push algorithm (DCAIP) by focusing on the current low level in information, complicated scenes, and various information on on-site maintenance. It also presents a smart wearable maintenance system (SEWMS), an edge computing-assisted IoT platform for the real-time guidance of technical experts and systems for on-site maintenance personnel, aiming to improve the efficiency and quality of on-site maintenance. The ninth paper titled “A crosswalk pedestrian recognition system by using deep learning and zebra-crossing recognition techniques” by Dow et al investigates a real-time pedestrian recognition system that ensures high accuracy by using a deep learning classifier and zebra-crossing recognition techniques. The proposed system was designed to improve pedestrian safety and reduce accidents at intersections. Environmental feature vectors were first used to detect zebra crossings and to determine crossing areas. An adaptive mapping technique was then used to map the pedestrian waiting area based on the crossing area. A dual-camera mechanism was used to maintain detection accuracy and improve system fault tolerance. Finally, the you-only-look-once model was used to recognize pedestrians at intersections. The 10th paper titled “Intelligent sentiment analysis approach using edge computing-based deep learning technique” by Sankar et al studies machine learning algorithms to extract the best features from the training review dataset. Then, the selected features are fed into the CNN and other fully connected layers for further processing. This work has also employed a pretrained sentiment analysis model over an Android application framework to classify reviews on a smartphone without the need for any cloud or server-side application programming interface. The 11th paper titled “Pipeline provenance for cloud-based big data analytics” by Wang et al proposes a solution, named LogProv toward realizing the functionalities for big data provenance, which needs to renovate data pipelines or some of big data software infrastructure to generate structured logs for pipeline events, and then stores data and logs separately in cloud space. The LogProv is implemented and deployed in Nectar Cloud, associated with Apache Pig, Hadoop ecosystem, and adopted Elasticsearch to provide query service. The 12th paper titled “Socially aware microcloud service overlay optimization in community networks” by Apolónia et al presents a model, named Select in Community Networks (SELECTinCN), which enhances the overlay creation for pub/sub systems over peer-to-peer (P2P) networks. Moreover, SELECTinCN includes social information based on cooperation within CNs by exploiting the social aspects of the community of practice. The model organizes the peers in a ring topology and provides an adaptive P2P connection establishment algorithm, where each peer identifies the number of connections needed based on the social structure and user availability. The 13th paper titled “DewSim: A trace-driven toolkit for simulating mobile device clusters in Dew computing environments” by Hirsch et al models and develops a trace-based toolkit built on modular software artifacts to speed up research in resource management techniques in Dew environments. A trace-driven methodology is adopted to assure the practical value of simulated scenarios. The toolkit comprises a device profiler application for Android to capture generic battery and central processing unit traces from real devices, a profile mixer to create user interaction baseline traces through generic ones, and an extensible engine to simulate the execution of workloads configurable via text files. The 14th paper titled “How to Place Your Apps in the Fog-State of the Art and Open Challenges” Borgi et al review the existing methodologies to solve the application placement problem in the fog, while pursuing three main objectives. First, it offers a comprehensive overview of the currently employed algorithms, on the availability of open-source prototypes, and on the size of test use cases. Second, it classifies the literature based on the application and fog infrastructure characteristics that are captured by available models, with a focus on the considered constraints and the optimized metrics. Finally, it identifies some open challenges in application placement in the fog. The 15th paper titled ‘SELFNET 5G mobile edge computing infrastructure: Design and prototyping’ by Chirivella-Perez et al. presented the design and prototype implementation of the fifth-generation (5G) mobile edge infrastructure based on a mobile edge computing paradigm. This mobile edge infrastructure SELFNET is an amlganmation of cloud computing, software-defined networking, and network function virtualization to end up with a portable 5G infrastructure testbed which enabled the realistic execution and testing. Finally, in the last paper titled ‘SDN/NFV security framework for fog-to-things computing infrastructure’ by Krishnan et al. proposed DTARS which is a System for Distributed Threat Analytics and Response for an Edge/Fog and SDN integrated architecture. In this system, the detection scheme runs at the data plane wherein a coarse-grained behavioral, anti-spoofing, flow monitoring and fine-grained traffic multi-feature entropy-based algorithms are deployed. The proposed framework has been developed for defense applications on malware testbed. We hope that the research contributions and findings in this special issue would benefit the readers in terms of enhancing their knowledge and encouraging them to work on various aspects of edge and fog computing. We express our sincere thanks to the editor-in-chief for allowing us to organize this special issue. The editorial office staffs are excellent and thanks for their support. We are also thankful to all the authors who made this special issue possible, and to the reviewers for their thoughtful contributions.
Rajiv Ranjan 0001, Massimo Villari, Haiying Shen, Omer F. Rana, Rajkumar Buyya
Softw. Pract. Exp.1
2020 A Multi-Order Distributed HOSVD with Its Incremental Computing for Big Services in Cyber-Physical-Social Systems
abstract
Big service is an extremely important application of service computing to provide predictive and needed services to humans. To operationalize big services, the heterogeneous data collected from Cyber-Physical-Social Systems (CPSS) must be processed efficiently. However, because of the rapid rise in the volume of data, faster and more efficient computational techniques are required. Therefore, in this paper, we propose a multi-order distributed high-order singular value decomposition method (MDHOSVD) with its incremental computational algorithm. To realize the MDHOSVD, a tensor blocks unfolding integration regulation is proposed. This method allows for the efficient analysis of large-scale heterogeneous data in blocks in an incremental fashion. Using simulation and experimental results from real-life, the high-efficiency of the proposed data processing and computational method, is demonstrated. Further, a case study about cyber-physical-social system data processing is illustrated. The proposed MDHOSVD method speeds up data processing, scales with data volume, improves the adaptability and extensibility over data diversity and converts low-level data into actionable knowledge.
Xiaokang Wang 0001, Laurence T. Yang, Lizhe Wang 0001, Rajiv Ranjan 0001, Xiaodao Chen, M. Jamal Deen
IEEE Trans. Big Data5
2020 Stochastic Workload Scheduling for Uncoordinated Datacenter Clouds with Multiple QoS Constraints
abstract
Cloud computing is now a well-adopted computing paradigm. With unprecedented scalability and flexibility, the computational cloud is able to carry out large scale computing tasks in parallel. The datacenter cloud is a new cloud computing model that uses multi-datacenter architectures for large scale massive data processing or computing. In datacenter cloud computing, the overall efficiency of the cloud depends largely on the workload scheduler, which allocates clients' tasks to different Cloud datacenters. Developing high performance workload scheduling techniques in Cloud computing imposes a great challenge which has been extensively studied. Most previous works aim only at minimizing the completion time of all tasks. However, timeliness is not the only concern, reliability and security are also very important. In this work, a comprehensive Quality of Service (QoS) model is proposed to measure the overall performance of datacenter clouds. An advanced Cross-Entropy based stochastic scheduling (CESS) algorithm is developed to optimize the accumulative QoS and sojourn time of all tasks. Experimental results show that our algorithm improves accumulative QoS and sojourn time by up to 56.1 and 25.4 percent respectively compared to the baseline algorithm. The runtime of our algorithm grows only linearly with the number of Cloud datacenters and tasks. Given the same arrival rate and service rate ratio, our algorithm steadily generates scheduling solutions with satisfactory QoS without sacrificing sojourn time.
Yunliang Chen 0002, Lizhe Wang 0001, Xiaodao Chen, Rajiv Ranjan 0001, Albert Y. Zomaya, Yuchen Zhou 0003, Shiyan Hu 0001
IEEE Trans. Cloud Comput.4
2020 Dynamically Partitioning Workflow over Federated Clouds for Optimising the Monetary Cost and Handling Run-Time Failures
abstract
Several real-world problems in domain of healthcare, large scale scientific simulations, and manufacturing are organised as workflow applications. Efficiently managing workflow applications on the Cloud computing data-centres is challenging due to the following problems: (i) they need to perform computation over sensitive data (e.g., Healthcare workflows) hence leading to additional security and legal risks especially considering public cloud environments and (ii) the dynamism of the cloud environment can lead to several run-time problems such as data loss and abnormal termination of workflow task due to failures of computing, storage, and network services. To tackle above challenges, this paper proposes a novel workflow management framework call Deploy on Federated Cloud Framework (DoFCF) that can dynamically partition scientific workflows across federated cloud (public/private) data-centres for minimising the financial cost, adhering to security requirements, while gracefully handling run-time failures. The framework is validated in cloud simulation tool (CloudSim) as well as in a realistic workflow-based cloud platform (e-Science Central). The results showed that our approach is practical and is successful in meeting users security requirements and reduces overall cost, and dynamically adapts to the run-time failures.
Zhenyu Wen, Rawaa Qasha, Zequn Li 0002, Rajiv Ranjan 0001, Paul Watson 0001, Alexander B. Romanovsky
IEEE Trans. Cloud Comput.4
2020 A User-centric Security Solution for Internet of Things and Edge Convergence
abstract
The Internet of Things (IoT) is becoming a backbone of sensing infrastructure to several mission-critical applications such as smart health, disaster management, and smart cities. Due to resource-constrained sensing devices, IoT infrastructures use Edge datacenters (EDCs) for real-time data processing. EDCs can be either static or mobile in nature, and this article considers both of these scenarios. Generally, EDCs communicate with IoT devices in emergency scenarios to evaluate data in real-time. Protecting data communications from malicious activity becomes a key factor, as all the communication flows through insecure channels. In such infrastructures, it is a challenging task for EDCs to ensure the trustworthiness of the data for emergency evaluations. The current communication security pattern of “communication before authentication” leaves a “black hole” for intruders to become part of communication processes without authentication. To overcome this issue and to develop security infrastructures for IoT and distributed Edge datacenters, this article proposes a user-centric security solution. The proposed security solution shifts from a network-centric approach to a user-centric security approach by authenticating users and devices before communication is established. A trusted controller is initialized to authenticate and establishes the secure channel between the devices before they start communication between themselves. The centralized controller draws a perimeter for secure communications within the boundary. Theoretical analysis and experimental evaluation of the proposed security model show that it not only secures the communication infrastructure but also improves the overall network performance.
Deepak Puthal, Laurence T. Yang, Schahram Dustdar, Zhenyu Wen, Jun Song 0003, Aad P. A. van Moorsel, Rajiv Ranjan 0001
ACM Trans. Cyber Phys. Syst.7
2020 Multiobjective Deployment of Data Analysis Operations in Heterogeneous IoT Infrastructure
abstract
The growth of Internet of Things (IoT) technology brings many new opportunities for applications in areas including smart healthcare, smart buildings, and smart agriculture. These applications must normally distribute the computations, required for extracting value from sensor data, over the IoT infrastructure platforms (e.g., sensors, phones, field-gateways, and clouds). This can be very challenging for IoT application developers due to the heterogeneity of the aforementioned platforms, potentially conflicting nonfunctional requirements (e.g., battery power, latency, and cost), and related deployment criteria, which is impossible to resolve manually. To address the above challenges, we have developed the PATH2iot framework that decomposes a complex IoT application into self-contained micro-operations. Based on the deployment criteria, PATH2iot automatically distributes the set of micro-operations across IoT infrastructure platforms, while respecting their run-time data and control flow dependencies. In our previous work, we have shown how to use the PATH2iot to optimize the battery life of a healthcare wearable. In this article, we describe a new research that significantly extends PATH2iot, which introduces a heuristic model capable of making optimal deployment decisions based on multiple conflicting nonfunctional requirements and selection criteria (user preferences). It does so by leveraging a well-known multicriteria decision-making method called the analytic hierarchical processes (AHP). The applicability of the deployment model is validated based on a real-world digital healthcare analytics use case. The results show that our model is able to find the optimal deployment solution for different user preferences.
Devki Nandan Jha, Peter Michalák, Zhenyu Wen, Rajiv Ranjan 0001, Paul Watson 0001
IEEE Trans. Ind. Informatics4
2020 GA-Par: Dependable Microservice Orchestration Framework for Geo-Distributed Clouds
abstract
Recent advances in composing Cloud applications have been driven by deployments of inter-networking heterogeneous microservices across multiple Cloud datacenters. System dependability has been of the upmost importance and criticality to both service vendors and customers. Security, a measurable attribute, is increasingly regarded as the representative example of dependability. Literally, with the increment of microservice types and dynamicity, applications are exposed to aggravated internal security threats and externally environmental uncertainties. Existing work mainly focuses on the QoS-aware composition of native VM-based Cloud application components, while ignoring uncertainties and security risks among interactive and interdependent container-based microservices. Still, orchestrating a set of microservices across datacenters under those constraints remains computationally intractable. This paper describes a new dependable microservice orchestration framework GA-Par to effectively select and deploy microservices whilst reducing the discrepancy between user security requirements and actual service provision. We adopt a hybrid (both whitebox and blackbox based) approach to measure the satisfaction of security requirement and the environmental impact of network QoS on system dependability. Due to the exponential grow of solution space, we develop a parallel Genetic Algorithm framework based on Spark to accelerate the operations for calculating the optimal or near-optimal solution. Large-scale real world datasets are utilized to validate models and orchestration approach. Experiments show that our solution outperforms the greedy-based security aware method with 42.34 percent improvement. GA-Par is roughly 4× faster than a Hadoop-based genetic algorithm solver and the effectiveness can be constantly guaranteed under different application scales.
Zhenyu Wen, Tao Lin 0004, Renyu Yang, Shouling Ji, Rajiv Ranjan 0001, Alexander B. Romanovsky, Chang-Ting Lin, Jie Xu 0007
IEEE Trans. Parallel Distributed Syst.5
2020 Holistic Technologies for Managing Internet of Things Services
abstract
The Internet of Things (IoT) is the latest Internet evolution that incorporates billions of sensors, actuators, and related software services that collectively distill high value information, perform actions that affect the physical world, and support a variety of applications controlled by different organizations and individuals. IoT's ability to observe and affect the physical world presents a unprecedented opportunity for creating IoT-based smart services and products that address grant challenges in emerging opportunities in areas such as climate change, precision agriculture, smart health, advanced manufacturing, and smart cities. This special issue identifies and addresses some of the key issues that hinder the development of IoT-based solutions. It includes articles that present the latest innovations in IoT security and privacy, IoT data quality and analysis, IoT resources and task management, as well as examples of IoT-based application services and domains.
Rajiv Ranjan 0001, Ching-Hsien Hsu, Lydia Y. Chen, Dimitrios Georgakopoulos 0001
IEEE Trans. Serv. Comput.1
2019 A Framework for Monitoring Microservice-Oriented Cloud Applications in Heterogeneous Virtualization Environments
abstract
Microservices have emerged as a new approach for developing and deploying cloud applications that require higher levels of agility, scale, and reliability. To this end, a microservice-based cloud application architecture advocates decomposition of monolithic application components into independent software components called "microservices". As the independent microservices can be developed, deployed, and updated independently of each other, it leads to complex run-time performance monitoring and management challenges. To solve this problem, we propose a generic monitoring framework, Multi-microservices Multi-virtualization Multi-cloud (M3) that monitors the performance of microservices deployed across heterogeneous virtualization platforms in a multi-cloud environment. We validated the efficacy and efficiency of M3 using a Book-Shop application executing across AWS and Azure.
Ayman Noor, Devki Nandan Jha, Karan Mitra, Prem Prakash Jayaraman, Arthur Souza 0001, Rajiv Ranjan 0001, Schahram Dustdar
CLOUD6
2019 SmartMonit: Real-Time Big Data Monitoring System
abstract
Modern big data processing systems are becoming very complex in terms of large-scale, high-concurrency and multiple talents. Thus, many failures and performance reductions only happen at run-time and are very difficult to capture. Moreover, some issues may only be triggered when some components are executed. To analyze the root cause of these types of issues, we have to capture the dependencies of each component in real-time. In this paper, we propose SmartMonit, a real-time big data monitoring system, which collects infrastructure information such as the process status of each task. At the same time, we develop a real-time stream processing framework to analyze the coordination among the tasks and the infrastructures. This coordination information is essential for troubleshooting the reasons for failures and performance reduction, especially the ones propagated from other causes.
Umit Demirbaga, Ayman Noor, Zhenyu Wen, Philip James 0002, Karan Mitra, Rajiv Ranjan 0001
SRDS6
2019 A Cost-Efficient Multi-cloud Orchestrator for Benchmarking Containerized Web-Applications
Devki Nandan Jha, Zhenyu Wen, Yinhao Li 0003, Michael Nee, Maciej Koutny, Rajiv Ranjan 0001
WISE6
2019 SmartDBO: Smart Docker Benchmarking Orchestrator for Web-application
abstract
Containerized web-applications have gained popularity recently due to the advantages provided by the containers including light-weight, packaged, fast start up and shut down and easy scalability. As there are more than 267 cloud providers, finding a flexible deployment option for containerized web-applications is very difficult as each cloud offers numerous deployment infrastructure. Benchmarking is one of the eminent options to evaluate the provisioned resources before product-level deployment. However, benchmarking the massive infrastructure resources provisioned by various cloud providers is a time consuming, tedious and costly process and is not practical to accomplish manually.
Devki Nandan Jha, Michael Nee, Zhenyu Wen, Albert Y. Zomaya, Rajiv Ranjan 0001
WWW5
2019 IoTSim-Stream: Modelling stream graph application in cloud simulation
Mutaz Barika, Saurabh Kumar Garg 0001, Andrew H. C. Chan, Rodrigo N. Calheiros, Rajiv Ranjan 0001
Future Gener. Comput. Syst.5
2019 Category Preferred Canopy-K-means based Collaborative Filtering algorithm
Jianjiang Li, Karan Mitra, Rajiv Ranjan 0001
Future Gener. Comput. Syst.7
2019 GoSharing: An intelligent incentive framework based on users' association for cooperative content sharing in mobile edge networks
Shuyun Luo, Zhenyu Wen, Xiaomei Zhang 0001, Weiqiang Xu 0001, Albert Y. Zomaya, Rajiv Ranjan 0001
Future Gener. Comput. Syst.6
2019 Implementation of a real-time network traffic monitoring service with network functions virtualization
Chao-Tung Yang, Shuo-Tsung Chen, Jung-Chun Liu, Yao-Yu Yang, Karan Mitra, Rajiv Ranjan 0001
Future Gener. Comput. Syst.6
2019 IoT-CANE: A unified knowledge management system for data-centric Internet of Things application systems
Yinhao Li 0003, Awatif Alqahtani, Ellis Solaiman, Charith Perera, Prem Prakash Jayaraman, Rajkumar Buyya, Graham Morgan, Rajiv Ranjan 0001
J. Parallel Distributed Comput.8
2019 Secure authentication and load balancing of distributed edge datacenters
abstract
Edge computing is an emerging research area to incorporate cloud computing into edge network devices. An Edge datacenter, also referred to as EDC, processes data streams and user requests in real-time and is therefore used to decrease the latency and congestion in the network. EDC is usually setup as a distributed system and is accordingly placed between the cloud datacenter and the data source . These EDCs work as an intermediate layer in the fog hierarchy between IoT and Cloud datacenter. EDC’s are aided by load balancers, responsible for distributing the workload amongst multiple EDC, in order to optimize resource utilization and response time . The load balancers make sure that the workload is equally divided amongst the available EDCs to avoid over loading of some EDCs while other remain idle as this directly impacts the user response and real-time event detection . Given the fact that EDCs are deployed in remote environments, the need for secure authentication is of major importance. In this paper we propose a novel load balancing technique that enables EDC authentication as well as identification of idle EDCs for better load balancing. The proposed load balancing technique is also compared with existing approaches and proves to be more efficient in locating EDC’s with less workload. In addition to the improved efficiency, the proposed scheme also strengthens the security of the network by incorporating destination EDC authentication.
Deepak Puthal, Rajiv Ranjan 0001, Ashish Nanda, Priyadarsi Nanda, Prem Prakash Jayaraman, Albert Y. Zomaya
J. Parallel Distributed Comput.2
2019 A note on tools and techniques for end-to-end QoS monitoring in Internet of Things
Rajiv Ranjan 0001, Ellis Solaiman, Massimo Villari, Paul Watson 0001
J. Parallel Distributed Comput.1
2019 Service level agreement specification for end-to-end IoT application ecosystems
abstract
Summary With an ever‐increasing variety and complexity of Internet of Things (IoT) applications delivered by increasing numbers of service providers, there is a growing demand for an automated mechanism that can monitor and regulate the interaction between the parties involved in IoT service provision and delivery. This mechanism needs to take the form of a contract, which, in this context, is referred to as a service level agreement (SLA). As a first step toward SLA monitoring and management, an SLA specification is essential. We believe that current SLA specification formats are unable to accommodate the unique characteristics of the IoT domain, such as its multilayered nature. Therefore, we propose a grammar for a syntactical structure of an SLA specification for IoT. The grammar is built based on a proposed conceptual model that considers the main concepts that can be used to express the requirements for hardware and software components of an IoT application on an end‐to‐end basis. We followed the goal question metric approach to evaluate the generality and expressiveness of the proposed grammar by reviewing its concepts and their predefined lists of vocabularies against two use cases with a considerable number of participants whose research interests are mainly related to IoT. The results of the analysis show that the proposed grammar achieved 91.70% of its generality goal and 93.43% of its expressiveness goal.
Awatif Alqahtani, Ellis Solaiman, Pankesh Patel, Schahram Dustdar, Rajiv Ranjan 0001
Softw. Pract. Exp.5
2019 SEEN: A Selective Encryption Method to Ensure Confidentiality for Big Sensing Data Streams
abstract
Resource constrained sensing devices are being used widely to build and deploy self-organizing wireless sensor networks for a variety of critical applications such as smart cities, smart health, precision agriculture and industrial control systems. Many such devices sense the deployed environment and generate a variety of data and send them to the server for analysis as data streams. A Data Stream Manager (DSM) at the server collects the data streams (often called big data) to perform real time analysis and decision-making for these critical applications. A malicious adversary may access or tamper with the data in transit. One of the challenging tasks in such applications is to assure the trustworthiness of the collected data so that any decisions are made on the processing of correct data. Assuring high data trustworthiness requires that the system satisfies two key security properties: confidentiality and integrity. To ensure the confidentiality of collected data, we need to prevent sensitive information from reaching the wrong people by ensuring that the right people are getting it. Sensed data are always associated with different sensitivity levels based on the sensitivity of emerging applications or the sensed data types or the sensing devices. For example, a temperature in a precision agriculture application may not be as sensitive as monitored data in smart health. Providing multilevel data confidentiality along with data integrity for big sensing data streams in the context of near real time analytics is a challenging problem. In this paper, we propose a Selective Encryption (SEEN) method to secure big sensing data streams that satisfies the desired multiple levels of confidentiality and data integrity. Our method is based on two key concepts: common shared keys that are initialized and updated by DSM without requiring retransmission, and a seamless key refreshment process without interrupting the data stream encryption/decryption. Theoretical analyses and experimental results of our SEEN method show that it can significantly improve the efficiency and buffer usage at DSM without compromising the confidentiality and integrity of the data streams.
Deepak Puthal, Xindong Wu 0001, Surya Nepal, Rajiv Ranjan 0001, Jinjun Chen
IEEE Trans. Big Data4
2019 Cross-Layer Multi-Cloud Real-Time Application QoS Monitoring and Benchmarking As-a-Service Framework
abstract
Cloud computing provides on-demand access to affordable hardware (e.g., multi-core CPUs, GPUs, disks, and networking equipment) and software (e.g., databases, application servers and data processing frameworks) platforms with features such as elasticity, pay-per-use, low upfront investment and low time to market. This has led to the proliferation of business critical applications that leverage various cloud platforms. Such applications hosted on single/multiple cloud provider platforms have diverse characteristics requiring extensive monitoring and benchmarking mechanisms to ensure run-time Quality of Service (QoS) (e.g., latency and throughput). This paper proposes, develops and validates CLAMBS-Cross-Layer Multi-Cloud Application Monitoring and Benchmarking as-a-Service for efficient QoS monitoring and benchmarking of cloud applications hosted on multi-clouds environments. The major highlight of CLAMBS is its capability of monitoring and benchmarking individual application components such as databases and web servers, distributed across cloud layers (*-aaS), spread among multiple cloud providers. We validate CLAMBS using prototype implementation and extensive experimentation and show that CLAMBS efficiently monitors and benchmarks application components on multi-cloud platforms including Amazon EC2 and Microsoft Azure.
Khalid Alhamazani, Rajiv Ranjan 0001, Prem Prakash Jayaraman, Karan Mitra, Chang Liu 0001, Fethi A. Rabhi, Dimitrios Georgakopoulos 0001, Lizhe Wang 0001
IEEE Trans. Cloud Comput.2
2019 Introduction to the Special Issue on Human-interaction-aware Data Analytics for Cyber-physical Systems
abstract
No abstract available.
Tongquan Wei, Junlong Zhou, Rajiv Ranjan 0001, Isaac Triguero, Huafeng Yu, Chun Jason Xue, Schahram Dustdar
ACM Trans. Cyber Phys. Syst.3
2019 SAFE: SDN-Assisted Framework for Edge-Cloud Interplay in Secure Healthcare Ecosystem
abstract
Improved quality of life has lead the healthcare industry to geographically expand and support real-time services. Following this trend, a surge of healthcare monitoring devices has substantially overgrown in the global market. These devices tend to generate data in humongous quantity that need real-time analysis with seamless and secure transmission to the computing nodes. The existing computing and networking infrastructures fall short to cater the services with desirable quality of service. Hence, to overcome these challenges, the proposed work presents a comprehensive platform referred as software defined network (SDN) Assisted Framework for Edge-Cloud Interplay in Secure Healthcare Ecosystem (SAFE). The objectives of SAFE include: first, an offloading scheme to support edge-cloud interplay, second, an SDN-assisted virtualized flow management scheme, and, third, a secure Lattice-based cryptosystem. Finally, the proposed scheme is validated on different performance parameters. Additionally, a security evaluation of the designed cryptosystem is also presented. The results obtained indicate the supremacy of the designed framework.
Gagangeet Singh Aujla, Rajat Chaudhary, Kuljeet Kaur, Sahil Garg, Neeraj Kumar 0001, Rajiv Ranjan 0001
IEEE Trans. Ind. Informatics6
2019 Renewable Energy-Based Multi-Indexed Job Classification and Container Management Scheme for Sustainability of Cloud Data Centers
abstract
Cloud computing has emerged as one of the most popular technologies of the modern era for providing on-demand services to the end users. Most of the computing tasks in cloud data centers are performed by geodistributed data centers which may consume a hefty amount of energy for their operations. However, the usage of renewable energy resources with appropriate server selection and consolidation can mitigate the energy related issues in cloud environment. Hence, in this paper, we propose a renewable energy-aware multi-indexed job classification and scheduling scheme using container as-a-service for data centers sustainability. In the proposed scheme, incoming workloads from different devices are transferred to the data center which has sufficient amount of renewable energy available with it. For this purpose, a renewable energy-based host selection and container consolidation scheme is also designed. The proposed scheme has been evaluated using Google workload traces. The results obtained prove 15%, 28%, and 10.55% higher energy savings in comparison to the existing schemes of its category.
Neeraj Kumar 0001, Gagangeet Singh Aujla, Sahil Garg, Kuljeet Kaur, Rajiv Ranjan 0001, Saurabh Kumar Garg 0001
IEEE Trans. Ind. Informatics5
2019 Unsupervised blocking and probabilistic parallelisation for record matching of distributed big data
Chenxiao Dou, Daniel Sun 0004, Raymond K. Wong 0001, Muhammad Atif 0003, Guoqiang Li 0001, Rajiv Ranjan 0001
J. Supercomput.7
2019 A Hybrid Deep Learning-Based Model for Anomaly Detection in Cloud Datacenter Networks
abstract
With the emergence of the Internet-of-Things (IoT) and seamless Internet connectivity, the need to process streaming data on real-time basis has become essential. However, the existing data stream management systems are not efficient in analyzing the network log big data for real-time anomaly detection. Further, the existing anomaly detection approaches are not proficient because they cannot be applied to networks, are computationally complex, and suffer from high false positives. Thus, in this paper a hybrid data processing model for network anomaly detection is proposed that leverages grey wolf optimization (GWO) and convolutional neural network (CNN). To enhance the capabilities of the proposed model, GWO and CNN learning approaches were enhanced with: 1) improved exploration, exploitation, and initial population generation abilities and 2) revamped dropout functionality, respectively. These extended variants are referred to as Improved-GWO (ImGWO) and Improved-CNN (ImCNN). The proposed model works in two phases for efficient network anomaly detection. In the first phase, ImGWO is used for feature selection in order to obtain an optimal trade-off between two objectives, i.e., reduced error rate and feature-set minimization. In the second phase, ImCNN is used for network anomaly classification. The efficacy of the proposed model is validated on benchmark (DARPA'98 and KDD'99) and synthetic datasets. The results obtained demonstrate that the proposed cloud-based anomaly detection model is superior in comparison to the other state-of-the-art models (used for network anomaly detection), in terms of accuracy, detection rate, false positive rate, and F-score. In average, the proposed model exhibits an overall improvement of 8.25%, 4.08%, and 3.62% in terms of detection rate, false positives, and accuracy, respectively; relative to standard GWO with CNN.
Sahil Garg, Kuljeet Kaur, Neeraj Kumar 0001, Georges Kaddoum, Albert Y. Zomaya, Rajiv Ranjan 0001
IEEE Trans. Netw. Serv. Manag.6
2018 Cloud computing based bushfire prediction for cyber-physical emergency applications
Saurabh Kumar Garg 0001, Jagannath Aryal, Tejal Shah, Gabor Kecskemeti, Rajiv Ranjan 0001
Future Gener. Comput. Syst.6
2018 A multi-layered performance analysis for cloud-based topic detection and tracking in Big Data applications
Meisong Wang, Prem Prakash Jayaraman, Ellis Solaiman, Lydia Y. Chen, Zheng Li 0001, Jun Song 0003, Dimitrios Georgakopoulos 0001, Rajiv Ranjan 0001
Future Gener. Comput. Syst.8
2018 Optimal Decision Making for Big Data Processing at Edge-Cloud Environment: An SDN Perspective
abstract
With the evolution of Internet and extensive usage of smart devices for computing and storage, cloud computing has become popular. It provides seamless services such as e-commerce, e-health, e-banking, etc., to the end users. These services are hosted on massive geodistributed data centers (DCs), which may be managed by different service providers. For faster response time, such a data explosion creates the need to expand DCs. So, to ease the load on DCs, some of the applications may be executed on the edge devices near to the proximity of the end users. However, such a multi-edge-cloud environment involves huge data migrations across the underlying network infrastructure, which may generate long migration delay and cost. Hence, in this paper, an efficient workload slicing scheme is proposed for handling data-intensive applications in multiedge-cloud environment using software-defined networks (SDN). To handle the inter-DC migrations efficiently, an SDN-based control scheme is presented, which provides energy-aware network traffic flow scheduling. Finally, a multileader multifollower Stackelberg game is proposed to provide cost-effective inter-DC migrations. The efficacy of the proposed scheme is evaluated on Google workload traces using various parameters. The results obtained show the effectiveness of the proposed scheme.
Gagangeet Singh Aujla, Neeraj Kumar 0001, Albert Y. Zomaya, Rajiv Ranjan 0001
IEEE Trans. Ind. Informatics4
2018 Tensor-Based Big Data Management Scheme for Dimensionality Reduction Problem in Smart Grid Systems: SDN Perspective
abstract
Smart grid (SG) is an integration of traditional power grid with advanced information and communication infrastructure for bidirectional energy flow between grid and end users. A huge amount of data is being generated by various smart devices deployed in SG systems. Such a massive data generation from various smart devices in SG systems may lead to various challenges for the networking infrastructure deployed between users and the grid. Hence, an efficient data transmission technique is required for providing desired QoS to the end users in this environment. Generally, the data generated by smart devices in SG has high dimensions in the form of multiple heterogeneous attributes, values of which are changed with time. The high dimensions of data may affect the performance of most of the designed solutions in this environment. Most of the existing schemes reported in the literature have complex operations for the data dimensionality reduction problem which may deteriorate the performance of any implemented solution for this problem. To address these challenges, in this paper, a tensor-based big data management scheme is proposed for dimensionality reduction problem of big data generated from various smart devices. In the proposed scheme, first the Frobenius norm is applied on high-order-tensors (used for data representation) to minimize the reconstruction error of the reduced tensors. Then, an empirical probability-based control algorithm is designed to estimate an optimal path to forward the reduced data using software-defined networks for minimization of the network load and effective bandwidth utilization. The proposed scheme minimizes the transmission delay incurred during the movement of the dimensionally reduced data between different nodes. The efficacy of the proposed scheme has been evaluated using extensive simulations carried out on the data traces using `R' programming and Matlab. The big data traces considered for evaluation consist of more than two million entries (2,075,259) collected at one minute sampling rate having hetrogenous features such as-voltage, energy, frequency, electric signals, etc. Moreover, a comparative study for different data traces and a real SG testbed is also presented to prove the efficacy of the proposed scheme. The results obtained depict the effectiveness of the proposed scheme with respect to the parameters such asnetwork delay, accuracy, and throughput.
Gagangeet Singh Aujla, Neeraj Kumar 0001, Albert Y. Zomaya, Charith Perera, Rajiv Ranjan 0001
IEEE Trans. Knowl. Data Eng.6
2018 Guest Editorial: Special Section on Advances in Big Data Analytics for Management
abstract
Cloud and network analytics can harness the immense stream of operational data from clouds and networks, and can perform analytics processing to improve reliability, automated configuration, performance, and optimized network management in general. In this area, we have witnessed a growing trend towards using statistical analysis and machine learning techniques to improve operations and management of IT systems and networks.
Giuliano Casale, Yixin Diao, Marco Mellia, Rajiv Ranjan 0001, Nur Zincir-Heywood
IEEE Trans. Netw. Serv. Manag.4
2018 G-ML-Octree: An Update-Efficient Index Structure for Simulating 3D Moving Objects Across GPUs
abstract
In real simulation applications, simulations often involve large volumes of three-dimensinal (3D) moving objects. With the rapid growth of the scale of simulation-problem domains, it has become a key requirement to efficiently manage massive 3D moving objects. Conventional indexing approaches for managing 3D moving objects during simulations generally sufferfrom excessive update costs. Aiming to this problem, this paper first proposes an update-efficient indexing structure by fusing a loose Octree and one update-memo structure, namely ML-Octree. ML-Octree significantly reduces the update costs of one simulation involving massive 3D moving objects. Towards providing a more efficient indexing approach, this paper has explored the feasibility of paralleling ML-Octree by employing Graphic Processing Unit (GPU). A load-balancing scheme is used to further improve the update performance of the GPU-aided ML-Octree. Finally, a distributed GPU-aided ML-Octree is proposed for large-scale simulations. The experimental results indicate that (1) ML-Octree can acquire the update-performance gain of an order of magnitude similar to that of Octree, (2) the GPU-aided ML-Octree can accelerate 5.07χ fasterthan a parallel ML-Octree with 8 CPU threads on average, (3) the load-balance scheme can improve GPU-aided ML-Octree by 2.3χ on average, and (4) the distributed GPU-aided ML-Octree can efficiently support large-scale simulations.
Ze Deng, Lizhe Wang 0001, Wei Han 0006, Rajiv Ranjan 0001, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.4
2018 Advances in Orchestrating Sustainable Smart Cities (Part 2)
abstract
This special issue asked for high quality original research papers (including smart city experience papers) that made significant contributions to the state-of-the-art in "method and techniques to build sustainable smart city solutions" research area. Rapid urbanization is a global megatrend with 66 percent of the world’s population expected to live in urban areas by 2050. The staggering exponential increase in urbanization is leading to more people migrating to major cities in the search of better opportunities and quality of life. Cities need to increase the efficiency in which they operate and use their resources sustainability in order to meet the demands imposed by rapid urbanisation. The challenge is to continue providing basic resources such as sufficient fresh water; cleaner energy; transportation alternatives to commute efficiently from one place to another; adaption to changing climatic conditions; safety and security; while also ensuring economical, social, and environment sustainability.
Rajiv Ranjan 0001, Prem Prakash Jayaraman, Massimo Villari, Dimitrios Georgakopoulos 0001
IEEE Trans. Sustain. Comput.1
2017 Towards a RISC Framework for Efficient Contextualisation in the IoT
abstract
The Internet of Things (IoT) is a new internet evolution that involves connecting billions of internet-connected devices that we refer to as IoT things. These devices can communicate directly and intelligently over the Internet, and generate a massive amount of data that needs to be consumed by a variety of IoT applications. This paper focuses on the automatic contextualisation of IoT data, which also involves distilling information and knowledge from the IoT aiming to simplify answering the following fundamental questions that often arises in IoT applications: Which data collected by IoT are relevant to myself and the IoT Things I care for? Related work around context management and contextualisation ranges from database techniques that involve query re-writing, to semantic web and rule-based context management approaches, to machine learning and data science-based solutions in mobile and ambient computing. All such existing approaches have two main aspects in common: They are highly incompatible and horribly inefficient from a scalability and performance perspective. In this paper, we discuss a new RISC Contextualisation Framework (RCF) we have developed, implemented key aspects of, and assess its scalability. RCF provides fundamental contextualisation concepts that can be mapped to all existing contextualisation approaches for IoT data (and in this sense, it provides a common denominator that unifies the contextualisation space). RCF can be easily implemented as a cloud-based service, and provides better scalability and performance that any of the existing content management and contextualisation approaches in the IoT space.
Dimitrios Georgakopoulos 0001, Ali Yavari, Prem Prakash Jayaraman, Rajiv Ranjan 0001
ICDCS4
2017 A balanced scheduler with data reuse and replication for scientific workflows in cloud computing systems
Israel Casas, Javid Taheri, Rajiv Ranjan 0001, Lizhe Wang 0001, Albert Y. Zomaya
Future Gener. Comput. Syst.3
2017 An efficient online direction-preserving compression approach for trajectory streaming data
Ze Deng, Wei Han 0006, Lizhe Wang 0001, Rajiv Ranjan 0001, Albert Y. Zomaya, Wei Jie
Future Gener. Comput. Syst.4
2017 A note on exploration of IoT generated big data using semantics
Rajiv Ranjan 0001, Dhavalkumar Thakker, Armin Haller, Rajkumar Buyya
Future Gener. Comput. Syst.1
2017 A scalable parallel algorithm for atmospheric general circulation models on a multi-core cluster
Jinrong Jiang, He Zhang 0005, Lizhe Wang 0001, Rajiv Ranjan 0001, Albert Y. Zomaya
Future Gener. Comput. Syst.6
2017 Elasticity management of Streaming Data Analytics Flows on clouds
Alireza Khoshkbarforoushha, Alireza Khosravian, Rajiv Ranjan 0001
J. Comput. Syst. Sci.3
2017 A dynamic prime number based efficient security mechanism for big sensing data streams
Deepak Puthal, Surya Nepal, Rajiv Ranjan 0001, Jinjun Chen
J. Comput. Syst. Sci.3
2017 A CPS framework based perturbation constrained buffer planning approach in VLSI design
Xiaodao Chen, Xiaohui Huang 0002, Yang Xiang 0001, Dongmei Zhang 0006, Rajiv Ranjan 0001, Chen Liao
J. Parallel Distributed Comput.5
2017 IOTSim: A simulator for analysing IoT applications
Xuezhi Zeng, Saurabh Kumar Garg 0001, Peter E. Strazdins, Prem Prakash Jayaraman, Dimitrios Georgakopoulos 0001, Rajiv Ranjan 0001
J. Syst. Archit.6
2017 Flower: A Data Analytics Flow Elasticity Manager
abstract
A data analytics flow typically operates on three layers: ingestion, analytics, and storage, each of which is provided by a data-intensive system. These systems are often available as cloud managed services, enabling the users to have pain-free deployment of data analytics flow applications such as click-stream analytics. Despite straightforward orchestration, elasticity management of the flows is challenging. This is due to: a) heterogeneity of workloads and diversity of cloud resources such as queue partitions, compute servers and NoSQL throughputs capacity, b) workload dependencies between the layers, and c) different performance behaviours and resource consumption patterns. In this demonstration, we present Flower, a holistic elasticity management system that exploits advanced optimization and control theory techniques to manage elasticity of complex data analytics flows on clouds. Flower analyzes statistics and data collected from different data-intensive systems to provide the user with a suite of rich functionalities, including: workload dependency analysis, optimal resource share analysis, dynamic resource provisioning, and cross-platform monitoring. We will showcase various features of Flower using a real-world data analytics flow. We will allow the audience to explore Flower by visually defining and configuring a data analytics flow elasticity manager and get hands-on experience with integrated data analytics flow management.
Alireza Khoshkbarforoushha, Rajiv Ranjan 0001, Qing Wang 0002, Carsten Friedrich
Proc. VLDB Endow.2
2017 Analytics-as-a-service in a multi-cloud environment through semantically-enabled hierarchical data processing
abstract
Summary A large number of cloud middleware platforms and tools are deployed to support a variety of internet‐of‐things (IoT) data analytics tasks. It is a common practice that such cloud platforms are only used by its owners to achieve their primary and predefined objectives, where raw and processed data are only consumed by them. However, allowing third parties to access processed data to achieve their own objectives significantly increases integration and cooperation and can also lead to innovative use of the data. Multi‐cloud, privacy‐aware environments facilitate such data access, allowing different parties to share processed data to reduce computation resource consumption collectively. However, there are interoperability issues in such environments that involve heterogeneous data and analytics‐as‐a‐service providers. There is a lack of both architectural blueprints that can support such diverse, multi‐cloud environments and corresponding empirical studies that show feasibility of such architectures. In this paper, we have outlined an innovative hierarchical data‐processing architecture that utilises semantics at all the levels of IoT stack in multi‐cloud environments. We demonstrate the feasibility of such architecture by building a system based on this architecture using OpenIoT as a middleware, and Google Cloud and Microsoft Azure as cloud environments. The evaluation shows that the system is scalable and has no significant limitations or overheads. Copyright © 2016 John Wiley & Sons, Ltd.
Prem Prakash Jayaraman, Charith Perera, Dimitrios Georgakopoulos 0001, Schahram Dustdar, Dhavalkumar Thakker, Rajiv Ranjan 0001
Softw. Pract. Exp.6
2017 Special issue on Big Data and Cloud of Things (CoT)
abstract
Special issue on Big Data and Cloud of Things (CoT)Cloud computing and Internet of Things (IoT) are two technologies that are already becoming part of our daily lives and are attracting significant interest from both industry and academia.The Cloud of Things (CoT) is a vision inspired from the IoT paradigm where everyday devices, namely, 'smart objects', are fully connected to the internet and are integrated with the cloud.It is expected the IoT will grow to 35 billion units by 2020, making it one of the main sources of 'Big Data' with characteristics such as volume, heterogeneity, complexity, velocity, and value.In recent years, IoT has given rise to a number of new CoT paradigms (but not limited to) including: Sensing-as-a-Service, Sensing-and Actuation-as-a-Service, Video-Surveillance-as-a-Service, Big Data Analytics-asa-Service, Data-as-a-Service, Sensor-as-a-Service, and Sensor-Event-as-a-Service. Cloud computing is a more mature technology compared to IoT.It can offer virtually unrestricted capabilities (e.g., storage and computation) to support IoT services and application that can exploit the data produced from IoT devices.The cloud essentially acts as a transparent layer between the IoT and applications providing flexibility, scalability, and hiding the complexities between the two layers (IoT and applications).However, the integration of cloud and IoT into Cloud of Things is not straightforward and imposes several challenges.These challenges include IoT device and service discovery, IoT device integration, big data management and analytics, cloud monitoring and orchestration for distributed IoT applications, mobility issues in cloud access, privacy and security, and SLA management for both cloud and IoT.Specific attention must be paid to address a range of issues from IoT data collection, storage, processing, analytics on demand to automatic provision and management of cloud resources to support the growing population of things.Hence, this special issue solicits paper related to topics including CoT architectures and models for smart provision of CoT applications, data management challenges facing CoT applications, software and tools to monitor, manage, deploy and deliver CoT applications, quality of service and related SLA management and policies for CoT applications, and security and privacy challenges facing CoT applications.The call for special issues received a number of submissions.After a two-phase peer review process, we have accepted 10 high-quality papers related to the aforementioned areas of interest.The first paper titled Using adaptive resource allocation to implement an elastic MapReduce framework by Jiaqi Zhao, Changlong Xue, Xinlin Tao, Shugong Zhang, and Jie Tao addresses the runtime resource demand challenge faced by application running on MapReduce frameworks.The proposed approach is capable of making the map reduce application, aware of overloading or under-loading situations with the resources allocated.They have extended the existing Hadoop MapReduce resource manager to implement the proposed strategy and validated the concept on an high-performance computing cluster with standard benchmark applications.Experimental results show a significant performance gain, for example, an up to 45% improvement in execution time for running multiple applications.The second paper titled A traffic hotline discovery method over cloud of things using big taxi GPS data by Xiaolong Xu, Wanchun Dou, Xuyun Zhang, Chunhua Hu, and Jinjun Chen addresses the challenge of discovering traffic hotline in CoT environments.Traffic hotlines are identified as the traffic lines with intensive traffic flows among traffic spots.They propose a hotline discovery method over CoT by establishing a hotline discovery principle.They have implemented their approach on SAP HANA cloud and tested it using big taxi global positioning system data under two application scenarios.
Rajiv Ranjan 0001, Lizhe Wang 0001, Prem Prakash Jayaraman, Karan Mitra, Dimitrios Georgakopoulos 0001
Softw. Pract. Exp.1
2017 MobiContext: A Context-Aware Cloud-Based Venue Recommendation Framework
abstract
In recent years, recommendation systems have seen significant evolution in the field of knowledge engineering. Most of the existing recommendation systems based their models on collaborative filtering approaches that make them simple to implement. However, performance of most of the existing collaborative filtering-based recommendation system suffers due to the challenges, such as: (a) cold start, (b) data sparseness, and (c) scalability. Moreover, recommendation problem is often characterized by the presence of many conflicting objectives or decision variables, such as users' preferences and venue closeness. In this paper, we proposed MobiContext, a hybrid cloud-based bi-objective recommendation framework (BORF) for mobile social networks. The MobiContext utilizes multi-objective optimization techniques to generate personalized recommendations. To address the issues pertaining to cold start and data sparseness, the BORF performs data preprocessing by using the Hub-Average (HA) inference model. Moreover, the Weighted Sum Approach (WSA) is implemented for scalar optimization and an evolutionary algorithm (NSGA-II) is applied for vector optimization to provide optimal suggestions to the users about a venue. The results of comprehensive experiments on a large-scale real dataset confirm the accuracy of the proposed recommendation framework.
Rizwana Irfan, Osman Khalid, Muhammad Usman Shahid Khan, Camelia Chira, Rajiv Ranjan 0001, Fan Zhang 0003, Samee Ullah Khan, Bharadwaj Veeravalli, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Cloud Comput.5
2017 DLSeF: A Dynamic Key-Length-Based Efficient Real-Time Security Verification Model for Big Data Stream
abstract
Applications in risk-critical domains such as emergency management and industrial control systems need near-real-time stream data processing in large-scale sensing networks. The key problem is how to ensure online end-to-end security (e.g., confidentiality, integrity, and authenticity) of data streams for such applications. We refer to this as an online security verification problem. Existing data security solutions cannot be applied in such applications as they cannot deal with data streams with high-volume and high-velocity data in real time. They introduce a significant buffering delay during security verification, resulting in a requirement for a large buffer size for the stream processing server. To address this problem, we propose a Dynamic Key-Length-Based Security Framework (DLSeF) based on a shared key derived from synchronized prime numbers; the key is dynamically updated at short intervals to thwart potential attacks to ensure end-to-end security. Theoretical analyses and experimental results of the DLSeF framework show that it can significantly improve the efficiency of processing stream data by reducing the security verification time and buffer usage without compromising security.
Deepak Puthal, Surya Nepal, Rajiv Ranjan 0001, Jinjun Chen
ACM Trans. Embed. Comput. Syst.3
2017 PSO-DS: a scheduling engine for scientific workflow managers
Israel Casas, Javid Taheri, Rajiv Ranjan 0001, Albert Y. Zomaya
J. Supercomput.3
2017 Rendezvous based routing protocol for wireless sensor networks with mobile sink
Suraj Sharma, Deepak Puthal, Sanjay Kumar Jena, Albert Y. Zomaya, Rajiv Ranjan 0001
J. Supercomput.5
2017 Erratum to: Rendezvous based routing protocol for wireless sensor networks with mobile sink
Suraj Sharma, Deepak Puthal, Sanjay Kumar Jena, Albert Y. Zomaya, Rajiv Ranjan 0001
J. Supercomput.5
2017 MacroServ: A Route Recommendation Service for Large-Scale Evacuations
abstract
To respond to emergencies in a fast and an effective manner, it is of critical importance to have efficient evacuation plans that lead to minimum road congestions. Although emergency evacuation systems have been studied in the past, the existing approaches, mostly based on multi-objective optimizations, are not scalable enough when involve numerous time varying parameters, such as traffic volume, safety status, and weather conditions. In this paper, we propose a scalable emergency evacuation service, termed the MacroServ that recommends the evacuees with the most preferred routes towards safe locations during a disaster. Unlike many existing approaches that model systems with static network characteristics, our approach considers real-time road conditions to compute the maximum flow capacity of routes in the transportation network. The evacuees are directed towards those routes that are safe and have least congestion resulting in decreased evacuation time. We utilized probability distributions to model the real-life stochastic behaviors of evacuees during emergency scenarios. The results indicate that recommendation of appropriate routes during emergency scenarios play a critical role in quicker and safe evacuation of the population.
Muhammad Usman Shahid Khan, Osman Khalid, Rajiv Ranjan 0001, Fan Zhang 0003, Bharadwaj Veeravalli, Samee Ullah Khan, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Serv. Comput.4
2017 A Survey on Modeling Energy Consumption of Cloud Applications: Deconstruction, State of the Art, and Trade-Off Debates
abstract
Given the complexity and heterogeneity in Cloud computing scenarios, the modeling approach has widely been employed to investigate and analyze the energy consumption of Cloud applications, by abstracting real-world objects and processes that are difficult to observe or understand directly. It is clear that the abstraction sacrifices, and usually does not need, the complete reflection of the reality to be modeled. Consequently, current energy consumption models vary in terms of purposes, assumptions, application characteristics and environmental conditions, with possible overlaps between different research works. Therefore, it would be necessary and valuable to reveal the state-of-the-art of the existing modeling efforts, so as to weave different models together to facilitate comprehending and further investigating application energy consumption in the Cloud domain. By systematically selecting, assessing, and synthesizing 76 relevant studies, we rationalized and organized over 30 energy consumption models with unified notations. To help investigate the existing models and facilitate future modeling work, we deconstructed the runtime execution and deployment environment of Cloud applications, and identified 18 environmental factors and 12 workload factors that would be influential on the energy consumption. In particular, there are complicated trade-offs and even debates when dealing with the combinational impacts of multiple factors.
Zheng Li 0001, Selome Kostentinos Tesfatsion, Saeed Bastani, Ahmed Ali-Eldin, Erik Elmroth, Maria Kihl, Rajiv Ranjan 0001
IEEE Trans. Sustain. Comput.7
2017 Advances in Orchestrating Sustainable Smart Cities (Part 1)
abstract
Rapid urbanization is a global megatrend with 66 percent of the world’s population expected to live in urban areas by 2050. The staggering exponential increase in urbanization is leading to more people migrating to major cities in the search of better opportunities and quality of life. Cities need to increase the efficiency in which they operate and use their resources sustainability in order to meet the demands imposed by rapid urbanization. The challenge is to continue providing basic resources such as sufficient fresh water; cleaner energy; transportation alternatives to commute efficiently from one place to another; adaption to changing climatic conditions; safety and security; while also ensuring economical, social, and environment sustainability. These challenges represent a huge opportunity for a paradigm shift that will require the need for data processing, analysis, and security close to the connected "things" i.e., towards the edge of the network in-order to support the growing smart city ecosystem. This paradigm shift will lead to an explosive growth of independent, owned and operated things and services including gateways, repeaters, smart infrastructure, and systems. Such a paradigm needs to be architected in a way that is easy to operate and dramatically simplifies the management of service offerings through scalable orchestration and proper automation. It must allow management, integration, and deployment of different tenants (such as services and things independently owned) within the smart city ecosystem in a uniform way. It should also have a suitable policy framework, letting specific stakeholders have access to data produced by other tenants, and analyze and extract values from the data. In order to address these challenges, this special issue solicits high quality original research papers (including smart city experience papers) that made significant contributions to the state-of-the-art in "method and techniques to build sustainable smart city solutions" research area. The call for papers received a number of submissions. After a two-phase peer review process, we have accepted five high-quality papers related to the aforementioned areas of interest which will be published in the October-December 2017 as Part 1. The papers in this issue are briefly summarized.
Rajiv Ranjan 0001, Prem Prakash Jayaraman, Massimo Villari, Dimitrios Georgakopoulos 0001
IEEE Trans. Sustain. Comput.1
2016 Resource and Performance Distribution Prediction for Large Scale Analytics Queries
abstract
Efficient resource consumption and performance estimation of data-intensive workloads is central to the design and development of workload management techniques. Recent work has explored the efficacy of using distribution-based estimation of workload performance as opposed to single point prediction for a number of workload management problems such as query scheduling, admission control, and the like. However, the proposed approaches lack an efficient workload performance distribution prediction in that they simply assume that the probability distribution function (pdf) of the target value is already available. This paper aims to address this problem for an inseparable portion of big data analytics workloads, Hive queries. To this end, we combine knowledge of Hive query executions with the novel usage of mixture density networks to predict the whole spectrum of resource and performance as probability density functions. We evaluate our technique using the TPC-H benchmark, showing that it not only produces accurate pdf predictions but outperforms the state of the art single point techniques in half of experiments.
Alireza Khoshkbarforoushha, Rajiv Ranjan 0001
ICPE2
2016 Performance analysis of data intensive cloud systems based on data management and replication: a survey
Saif Ur Rehman Malik, Samee Ullah Khan, Sam J. Ewen, Nikos Tziritas, Joanna Kolodziej, Albert Y. Zomaya, Sajjad Ahmad Madani, Nasro Min-Allah, Lizhe Wang 0001, Cheng-Zhong Xu 0001, Qutaibah M. Malluhi, Johnatan E. Pecero, Pavan Balaji, Abhinav Vishnu, Rajiv Ranjan 0001, Sherali Zeadally, Hongxiang Li 0001
Distributed Parallel Databases15
2016 Spot pricing in the Cloud ecosystem: A comparative investigation
Zheng Li 0001, He Zhang 0001, Liam O'Brien, Maria Kihl, Rajiv Ranjan 0001
J. Syst. Softw.7
2016 An integrated static detection and analysis framework for android
Jun Song 0003, Chunling Han, Rajiv Ranjan 0001, Lizhe Wang 0001
Pervasive Mob. Comput.5
2016 A Computing Perspective on Smart City [Guest Editorial]
abstract
The papers in this special section focus on next generation urbanization that incorporates smart city development. The development of smart cities is viewed as the key to the next generation urbanization process for improving the efficiency, reliability, and security of a traditional city. The concept of smart city includes various aspects such as environmental sustainability, social sustainability, regional competitiveness,natural resources management, cybersecurity, and quality of life improvement. With the massive deployment of networked smart devices/sensors, an unprecedentedly large amount of sensory data can be collected and processed by advanced computing paradigms, which are the enabling techniques for smart city. For example, given historical environmental, population, and economic information, salient modeling and analytics are needed to simulate the impact of potential city planning strategies, which will be critical for intelligent decision-making.
Lizhe Wang 0001, Shiyan Hu 0001, Gilles Betis, Rajiv Ranjan 0001
IEEE Trans. Computers4
2016 CEVP: Cross Entropy based Virtual Machine Placement for Energy Optimization in Clouds
Xiaodao Chen, Yunliang Chen 0002, Albert Y. Zomaya, Rajiv Ranjan 0001, Shiyan Hu 0001
J. Supercomput.4
2016 An online greedy allocation of VMs with non-increasing reservations in clouds
Yonggen Gu, Jie Tao 0001, Guoqiang Li 0001, Prem Prakash Jayaraman, Daniel Sun 0004, Rajiv Ranjan 0001, Albert Y. Zomaya, Jingti Han
J. Supercomput.7
2015 Cross-Layer SLA Management for Cloud-hosted Big Data Analytics Applications
abstract
As we come to terms with various big data challenges, one vital issue remains largely untouched. That is service level agreement (SLA) management to deliver strong Quality of Service (QoS) guarantees for big data analytics applications (BDAA) sharing the same underlying infrastructure, for example, a public cloud platform. Although SLA and QoS are not new concepts as they originated much before the cloud computing and big data era, its importance is amplified and complexity is aggravated by the emergence of time-sensitive BDAAs such as social network-based stock recommendation and environmental monitoring. These applications require strong QoS guarantees and dependability from the underlying cloud computing platform to accommodate real-time responses while handling ever-increasing complexities and uncertainties. Hence, the over-reaching goal of this PhD research is to develop novel simulation, modelling and benchmarking tools and techniques that can aid researchers and practitioners in studying the impact of uncertainties (contention, failures, anomalies, etc.) on the final SLA and QoS of a cloud-hosted BDAA.
Xuezhi Zeng, Rajiv Ranjan 0001, Peter E. Strazdins, Saurabh Kumar Garg 0001, Lizhe Wang 0001
CCGRID2
2015 A Dynamic Key Length Based Approach for Real-Time Security Verification of Big Sensing Data Stream
Deepak Puthal, Surya Nepal, Rajiv Ranjan 0001, Jinjun Chen
WISE (2)3
2015 Reporting an experience on design and implementation of e-Health systems on Azure cloud
abstract
Summary Electronic Health (e‐Health) technology has brought the world with significant transformation from traditional paper‐based medical practice to Information and Communication Technologies (ICT)‐based systems for automatic management (storage, processing, and archiving) of information. Traditionally, e‐Health systems have been designed to operate within stovepipes on dedicated networks, physical computers, and locally managed software platforms that make it susceptible to many serious limitations including: (1) lack of on‐demand scalability during critical situations, (2) high administrative overheads and costs, and (3) inefficient resource utilization and energy consumption due to lack of automation. In this paper, we present an approach to migrate the ICT systems in the e‐Health sector from traditional in‐house Client/Server (C/S) architecture to the virtualized cloud computing environment. To this end, we developed two cloud‐based e‐Health applications (Medical Practice Management System and Telemedicine Practice System) for demonstrating how cloud services can be leveraged for developing and deploying such applications. The Windows Azure cloud computing platform is selected as an example public cloud platform for our study. We conducted several performance evaluation experiments to understand the QoS tradeoffs of our applications under variable workload on Azure. Copyright © 2014 John Wiley & Sons, Ltd.
Shilin Lu, Rajiv Ranjan 0001, Peter E. Strazdins
Concurr. Comput. Pract. Exp.2
2015 A note on resource orchestration for cloud computing
abstract
A note on resource orchestration for cloud computingWelcome to the special issue of Concurrency and Computation: Practice and Experience (CCPE) journal.This special issue compiles a number of excellent technical contributions that significantly advance the state-of-the-art in the areas of orchestrating cloud resources, composing new cloud services from existing ones, increasing energy efficiency via cloud resource orchestration, and developing cloud-based image processing solutions.Over the past few years, cloud computing [1-4] has emerged as the latest and most dominant utility computing solution offering both hardware and software resources as virtualization-enabled services.Cloud computing providers such as Amazon Web Services and Microsoft Azure currently provide application owners the option of deploying their applications over a network of a virtually infinite resource pool with practically no up-front capital investment and with operating cost proportional to the actual use (i.e., implementing a pay-as-you-go model).An increasing number of cloud vendors offer information and communication technology (ICT) resources such as hardware (CPUs, GPUs, storage, and networks), software infrastructure (e.g., databases, webservers, stream-processing systems, and data-mining packages), and collaboration/communication applications (e.g., email, video on demand, and social networks) as infrastructure as a service (IAAS), platform as a service (PAAS), and software as a service (SAAS), respectively.This approach allows enterprises to easily, cost effectively, and reliably offer business services that are supported by computing and software resources that are provided and maintained by IAAS, PAAS, and SAAS providers.This makes cloud computing attractive to especially small and medium size enterprises (SMEs), as it allows them to focus more on their core business and less on ICT infrastructure.One of the fundamental issues in exploiting cloud computing in this fashion is developing better Resource Orchestration (RO) [1-5] techniques and programming frameworks.More specifically, Resource Orchestration (RO) is 'the set of operations that cloud providers (e.g., AWS) and application owners (e.g., Netflix) undertake (either manually or automatically via computer programs) for selecting, deploying, monitoring, and dynamically controlling configuration of hardware and software resources as a system of QoS assured components that can be seamlessly delivered to end-users' [1].Since RO operations span across all layers of cloud computing stack [1], an overall goal of RO is to ensure successful hosting and delivery of applications (SAAS) by managing the fulfillment of the QoS objectives of both the application owners (e.g., maximize availability, maximize throughput, minimize latency, and avoid overloading) and the Cloud resource providers (e.g., maximize utilization, maximize energy efficiency, and maximize profit).One of the main complexities in Cloud resource management is that Cloud resources are typically identified by unique functional specifications, and then evaluated via their Quality of Service (QoS) properties.However, in practice, each resource may have multiple unique functional specifications that enable serving diverse user needs.For example, a song retrieval application is a Cloud resource that can be identified and enacted by the name of a song or via its lyrics, as users do may not know or remember the names of all song they want to find.The paper titled 'A Service Evaluation Method for Cross-cloud Service Choreography' [6] addresses this challenge by proposing a multifunctional specification solution for cross-cloud service choreography.This solution is based on a mixed integer programming model for multifunctional specification/identification that decomposes global resource and application constraints into local constraints.This allows constraint evaluations to be performed for cross-cloud service choreography.Experimental verification of this approach is also provided.Provisioning cloud resources requires providers and consumers to reach an agreement on the service usage terms and conditions.Such agreements are captured as Service Level Agreements (SLAs).The paper titled 'AutoSLAM -A Policy-based Framework for Automated SLA Establishment in Cloud
Rajiv Ranjan 0001, Rajkumar Buyya, Surya Nepal, Dimitrios Georgakopoulos 0001
Concurr. Comput. Pract. Exp.1
2015 Remote sensing big data computing: Challenges and opportunities
Yan Ma 0001, Haiping Wu, Lizhe Wang 0001, Bormin Huang, Rajiv Ranjan 0001, Albert Y. Zomaya, Wei Jie
Future Gener. Comput. Syst.5
2015 A note on new trends in data-aware scheduling and resource provisioning in modern HPC systems
Jie Tao 0001, Joanna Kolodziej, Rajiv Ranjan 0001, Prem Prakash Jayaraman, Rajkumar Buyya
Future Gener. Comput. Syst.3
2015 Software Tools and Techniques for Big Data Computing in Healthcare Clouds
Lizhe Wang 0001, Rajiv Ranjan 0001, Joanna Kolodziej, Albert Y. Zomaya, Leila Alem
Future Gener. Comput. Syst.2
2015 Towards building a data-intensive index for big data computing - A case study of Remote Sensing data processing
Yan Ma 0001, Lizhe Wang 0001, Peng Liu 0024, Rajiv Ranjan 0001
Inf. Sci.4
2015 Particle Swarm Optimization based dictionary learning for remote sensing big data
Lizhe Wang 0001, Hao Geng, Peng Liu 0024, Ke Lu 0002, Joanna Kolodziej, Rajiv Ranjan 0001, Albert Y. Zomaya
Knowl. Based Syst.6
2015 MuR-DPA: Top-Down Levelled Multi-Replica Merkle Hash Tree Based Secure Public Auditing for Dynamic Big Data Storage on Cloud
abstract
Cloud computing that provides elastic computing and storage resource on demand has become increasingly important due to the emergence of “big data”. Cloud computing resources are a natural fit for processing big data streams as they allow big data application to run at a scale which is required for handling its complexities (data volume, variety and velocity). With the data no longer under users' direct control, data security in cloud computing is becoming one of the most concerns in the adoption of cloud computing resources. In order to improve data reliability and availability, storing multiple replicas along with original datasets is a common strategy for cloud service providers. Public data auditing schemes allow users to verify their outsourced data storage without having to retrieve the whole dataset. However, existing data auditing techniques suffers from efficiency and security problems. First, for dynamic datasets with multiple replicas, the communication overhead for update verifications is very large, because each update requires updating of all replicas, where verification for each update requires O(log n ) communication complexity. Second, existing schemes cannot provide public auditing and authentication of block indices at the same time. Without authentication of block indices, the server can build a valid proof based on data blocks other than the blocks client requested to verify. In order to address these problems, in this paper, we present a novel public auditing scheme named MuR-DPA. The new scheme incorporated a novel authenticated data structure (ADS) based on the Merkle hash tree (MHT), which we call MR-MHT. To support full dynamic data updates and authentication of block indices, we included rank and level values in computation of MHT nodes. In contrast to existing schemes, level values of nodes in MR-MHT are assigned in a top-down order, and all replica blocks for each data block are organized into a same replica sub-tree. Such a configuration allows efficient verification of updates for multiple replicas. Compared to existing integrity verification and public auditing schemes, theoretical analysis and experimental results show that the proposed MuR-DPA scheme can not only incur much less communication overhead for both update verification and integrity verification of cloud datasets with multiple replicas, but also provide enhanced security against dishonest cloud service providers.
Chang Liu 0001, Rajiv Ranjan 0001, Chi Yang, Xuyun Zhang, Lizhe Wang 0001, Jinjun Chen
IEEE Trans. Computers2
2015 CloudGenius: A Hybrid Decision Support Method for Automating the Migration of Web Application Clusters to Public Clouds
abstract
With the increase in cloud service providers, and the increasing number of compute services offered, a migration of information systems to the cloud demands selecting the best mix of compute services and virtual machine (VM ) images from an abundance of possibilities. Therefore, a migration process for web applications has to automate evaluation and, in doing so, ensure that Quality of Service (QoS) requirements are met, while satisfying conflicting selection criteria like throughput and cost. When selecting compute services for multiple connected software components, web application engineers must consider heterogeneous sets of criteria and complex dependencies across multiple layers, which is impossible to resolve manually. The previously proposed CloudGenius framework has proven its capability to support migrations of single-component web applications. In this paper, we expand on the additional complexity of facilitating migration support for multi-component web applications. In particular, we present an evolutionary migration process for web application clusters distributed over multiple locations, and clearly identify the most important criteria relevant to the selection problem. Moreover, we present a multi-criteria-based selection algorithm based on Analytic Hierarchy Process (AHP). Because the solution space grows exponentially, we developed a Genetic Algorithm (GA)-based approach to cope with computational complexities in a growing cloud market. Furthermore, a use case example proofs CloudGenius’ applicability. To conduct experiments, we implemented CumulusGenius, a prototype of the selection algorithm and the GA deployable on hadoop clusters. Experiments with CumulusGenius give insights on time complexities and the quality of the GA.
Michael Menzel 0002, Rajiv Ranjan 0001, Lizhe Wang 0001, Samee Ullah Khan, Jinjun Chen
IEEE Trans. Computers2
2015 Ultra-Scalable CPU-MIC Acceleration of Mesoscale Atmospheric Modeling on Tianhe-2
abstract
In this work an ultra-scalable algorithm is designed and optimized to accelerate a 3D compressible Euler atmospheric model on the CPU-MIC hybrid system of Tianhe-2. We first reformulate the mesocale model to avoid long-latency operations, and then employ carefully designed inter-node and intra-node domain decomposition algorithms to achieve balance utilization of different computing units. Proper communication-computation overlap and concurrent data transfer methods are utilized to reduce the cost of data movement at scale. A variety of optimization techniques on both the CPU side and the accelerator side are exploited to enhance the in-socket performance. The proposed hybrid algorithm successfully scales to 6,144 Tianhe-2 nodes with a nearly ideal weak scaling efficiency, and achieve over 8 percent of the peak performance in double precision. This ultra-scalable hybrid algorithm may be of interest to the community to accelerating atmospheric models on increasingly dominated heterogeneous supercomputers.
Wei Xue 0003, Chao Yang 0002, Haohuan Fu, Yangtong Xu, Junfeng Liao, Lin Gan 0001, Yutong Lu, Rajiv Ranjan 0001, Lizhe Wang 0001
IEEE Trans. Computers9
2015 Workload Prediction Using ARIMA Model and Its Impact on Cloud Applications' QoS
abstract
As companies shift from desktop applications to cloud-based software as a service (SaaS) applications deployed on public clouds, the competition for end-users by cloud providers offering similar services grows. In order to survive in such a competitive market, cloud-based companies must achieve good quality of service (QoS) for their users, or risk losing their customers to competitors. However, meeting the QoS with a cost-effective amount of resources is challenging because workloads experience variation overtime. This problem can be solved with proactive dynamic provisioning of resources, which can estimate the future need of applications in terms of resources and allocate them in advance, releasing them once they are not required. In this paper, we present the realization of a cloud workload prediction module for SaaS providers based on the autoregressive integrated moving average (ARIMA) model. We introduce the prediction based on the ARIMA model and evaluate its accuracy of future workload prediction using real traces of requests to Web servers. We also evaluate the impact of the achieved accuracy in terms of efficiency in resource utilization and QoS. Simulation results show that our model is able to achieve an average accuracy of up to 91 percent, which leads to efficiency in resource utilization with minimal impact on the QoS.
Rodrigo N. Calheiros, Enayat Masoumi, Rajiv Ranjan 0001, Rajkumar Buyya
IEEE Trans. Cloud Comput.3
2015 Recent advances in autonomic provisioning of big data applications on clouds
abstract
Cloud computing assembles large networks of virtualised ICT services such as hardware resources (such as CPU, storage, and network), software resources (such as databases, application servers, and web servers) and applications. Big Data applications have become a common phenomenon in domain of science, engineering, and commerce. Large-scale, heterogeneous, and uncertain Big Data applications are becoming increasingly common, yet current cloud resource provisioning methods do not scale well and nor do they perform well under highly unpredictable conditions (data volume, data variety, data arrival rate, etc.). Much research effort have been paid in the fundamental understanding, technologies, and concepts related to autonomic provisioning of cloud resources for Big Data applications, to make cloud-hosted Big Data applications operate more efficiently, with reduced financial and environmental costs, reduced under-utilisation of resources, and better performance at times of unpredictable workload. Targeting the aforementioned research challenges, this special issue compiles recent advances in Autonomic Provisioning of Big Data Applications on Clouds. The special issue articles are briefly summarized.
Rajiv Ranjan 0001, Lizhe Wang 0001, Albert Y. Zomaya, Dimitrios Georgakopoulos 0001, Xian-He Sun, Guojun Wang 0001
IEEE Trans. Cloud Comput.1
2015 Guest Editorial on Advances in Tools and Techniques for Enabling Cyber-Physical-Social Systems - Part I
abstract
The papers in this special section (Part I) are devoted to the topic of cyberphysical social systems (CPSS). These systems integrate computational physical elements (seamless integration of computational algorithms and physical components) capable of interacting with, reflecting and influencing each other as well as the complicated system and information exhibited by human’s social behavior. Rapid advances in mobile cloud computing, man–machine, machine–machine communications,smart phone networks will cause a new paradigm shift in CPSS design, applications, and operations by bringing improvements not only to the quality of service (QoS) but also to quality of experiment (QoE) and quality of protection (QoP) in terms of the cost efficiency, reliability, security, and energy efficiency from the human perspective of view. The integration of cyber–physical systems and social networks also provides a novel platform to addresses challenges in cyber–physical–social interactions and human-centric technologies development by using powerful tools such as social computing, social cooperation, and social sensing. Application examples of employing mobile computing in CPSS include traffic accidents detection in smart transportation, energy management in smart grid, health monitoring and evaluation.
Mianxiong Dong, Rajiv Ranjan 0001, Albert Y. Zomaya, Man Lin
IEEE Trans. Comput. Soc. Syst.2
2015 Guest Editorial on Advances in Tools and Techniques for Enabling Cyber-Physical-Social Systems - Part II
abstract
The six papers in this special section (Part II) on Cyber–Physical–Social Systems (CPSS) focus on computational social systems for emerging techniques for radio access networks, data deduplication, big data computing, smart community, cloud computing, and Internet of Things.
Mianxiong Dong, Rajiv Ranjan 0001, Albert Y. Zomaya, Man Lin
IEEE Trans. Comput. Soc. Syst.2
2015 Parallel Processing of Dynamic Continuous Queries over Streaming Data Flows
abstract
More and more real-time applications need to handle dynamic continuous queries over streaming data of high density. Conventional data and query indexing approaches generally do not apply for excessive costs in either maintenance or space. Aiming at these problems, this study first proposes a new indexing structure by fusing an adaptive cell and KDB-tree, namely CKDB-tree. A cell-tree indexing approach has been developed on the basis of the CKDB-tree that supports dynamic continuous queries. The approach significantly reduces the space costs and scales well with the increasing data size. Towards providing a scalable solution to filtering massive steaming data, this study has explored the feasibility to utilize the contemporary general-purpose computing on the graphics processing unit (GPGPU). The CKDB-tree-based approach has been extended to operate on both the CPU (host) and the GPU (device). The GPGPU-aided approach performs query indexing on the host while perform streaming data filtering on the device in a massively parallel manner. The two heterogeneous tasks execute in parallel and the latency of streaming data transfer between the host and the device is hidden. The experimental results indicate that (1) CKDB-tree can reduce the space cost comparing to the cell-based indexing structure by 60 percent on average, (2) the approach upon the CKDB-tree outperforms the traditional counterparts upon the KDB-tree by 66, 75 and 79 percent in average for uniform, skewed and hyper-skewed data in terms of update costs, and (3) the GPGPU-aided approach greatly improves the approach upon the CKDB-tree with the support of only a single Kepler GPU, and it provides real-time filtering of streaming data with 2.5M data tuples per second. The massively parallel computing technology exhibits great potentials in streaming data monitoring.
Ze Deng, Lizhe Wang 0001, Xiaodao Chen, Rajiv Ranjan 0001, Albert Y. Zomaya, Dan Chen 0001
IEEE Trans. Parallel Distributed Syst.5
2015 A Parallel File System with Application-Aware Data Layout Policies for Massive Remote Sensing Image Processing in Digital Earth
abstract
Remote sensing applications in Digital Earth are overwhelmed with vast quantities of remote sensing (RS) image data. The intolerable I/O burden introduced by the massive amounts of RS data and the irregular RS data access patterns has made the traditional cluster based parallel I/O systems no longer applicable. We propose a RS data object-based parallel file system for remote sensing applications and implement it with the OrangeFS file system. It provides application-aware data layout policies, together with RS data object based data I/O interfaces, for efficient support of various data access patterns of RS applications from the server side. With the prior knowledge of the desired RS data access patterns, HPGFS could offer relevant space-filling curves to organize the sliced 3-D data bricks and distribute them over I/O servers. In this way, data layouts consistent with expected data access patterns could be created to explore data locality and achieve performance improvement. Moreover, the multi-band RS data with complex structured geographical metadata could be accessed and managed as a single data object. Through experiments on remote sensing applications with different access patterns, we have achieved performance improvement of about 30 percent for I/O and 20 percent overall.
Lizhe Wang 0001, Yan Ma 0001, Albert Y. Zomaya, Rajiv Ranjan 0001, Dan Chen 0001
IEEE Trans. Parallel Distributed Syst.4
2014 Real-Time QoS Monitoring for Cloud-Based Big Data Analytics Applications in Mobile Environments
abstract
The service delivery model of cloud computing acts as a key enabler for big data analytics applications enhancing productivity, efficiency and reducing costs. The ever increasing flood of data generated from smart phones and sensors such as RFID readers, traffic cams etc require innovative provisioning and QoS monitoring approaches to continuously support big data analytics. To provide essential information for effective and efficient bid data analytics application QoS monitoring, in this paper we propose and develop CLAMS-Cross-Layer Multi-Cloud Application Monitoring-as-a-Service Framework. The proposed framework: (a) performs multi-cloud monitoring, and (b) addresses the issue of cross-layer monitoring of applications. We implement and demonstrate CLAMS functions on real-world multi-cloud platforms such as Amazon and Azure.
Khalid Alhamazani, Rajiv Ranjan 0001, Prem Prakash Jayaraman, Karan Mitra, Meisong Wang, Zhiqiang George Huang, Lizhe Wang 0001, Fethi A. Rabhi
MDM (1)2
2014 Running Data-Intensive Scientific Workflows in the Cloud
abstract
The scale of scientific applications becomes increasingly large not only in computation, but also in data. Many of these applications also concern inter-related tasks with data dependencies, hence, they are scientific workflows. The efficient coordination of executing/running scientific workflows is of great practical importance. The core of such coordination is scheduling and resource allocation. In this paper, we present three scheduling heuristics for running large-scale, data-intensive scientific workflows in clouds. In particular, the three heuristic algorithms are designed to leverage slot queue threshold, data locality and data prefetching, respectively. We also demonstrate how these heuristics can be collectively used to tackle different issues in running "data-intensive" workflows in clouds although each of these heuristics can be used independently. The practicality of our algorithms has been realized by actually implementing and incorporating them into our workflow execution system (DEWE). Using Montage, an astronomical image mosaic engine, as an example workflow, and Amazon EC2 as the cloud environment, we evaluate the performance of our heuristics in terms primarily of completion time (make span). We also scrutinize workflow execution showing different execution phases to identify their impact on performance. Our algorithms scale well and reduce make span by up to 27%.
Chiaki Sato, Luke M. Leslie, Young Choon Lee, Albert Y. Zomaya, Rajiv Ranjan 0001
PDCAT5
2014 Towards understanding the runtime configuration management of do-it-yourself content delivery network applications over public clouds
Zheng Li 0001, Karan Mitra, Miranda Zhang, Rajiv Ranjan 0001, Dimitrios Georgakopoulos 0001, Albert Y. Zomaya, Liam O'Brien
Future Gener. Comput. Syst.4
2014 A security framework in G-Hadoop for big data computing across distributed Cloud data centres
Jiaqi Zhao 0004, Lizhe Wang 0001, Jie Tao 0001, Jinjun Chen, Weiye Sun, Rajiv Ranjan 0001, Joanna Kolodziej, Achim Streit, Dimitrios Georgakopoulos 0001
J. Comput. Syst. Sci.6
2014 A note on software tools and techniques for monitoring and prediction of cloud services
abstract
Cloud computing is the latest computing paradigm that transparently delivers Information and Communication Technology resources as services, freeing the users of Cloud applications from dealing with low-level implementation and system administration details. Cloud provides the promise of on-demand access to affordable large-scale computing (e.g., multi-core CPUs, GPUs, and clusters of GPUs), storage (such as disks), and software (e.g., databases, application servers, and data processing frameworks) resources without substantial up-front investment. Cloud resources are hosted in large datacenters, often referred to as virtualized data farms, operated by companies such as Amazon, Apple, GoGrid, and Microsoft. While the growing ubiquity of Cloud computing is having a significant impact in many applications domains, there are still significant problems that exist with regard to efficient provisioning and delivery of applications using its Information and Communication Technology resources. These barriers are due to resource uncertainties 1 that have degradable effect on the run-time Quality of Service (e.g., access latency and number of requests being successfully served per second) of software applications deployed in the Cloud. There are many reasons for such uncertainties including (i) unpredictable application workload types (enterprise, scientific, and streaming big data analytics), (ii) fluctuations in resource capacity demands (i.e., bandwidth, memory, storage, and CPU), (iii) abrupt failures (e.g., failure of a network link), (iv) stochastic access patterns (e.g., number of end-users and their geo-location), (v) heterogeneity in device types (e.g., mobile phone, laptop, and smart TV), (v) heterogeneous resource types and their providers, and (vi) heterogeneity in data types (3D images, videos, audios, text, etc.) and network types (e.g., wired and wireless). These Cloud resource uncertainties need to be managed optimally to maintain contractual requirements defined in Service-Level Agreements (SLAs) that underlie most Cloud computing contracts. Basically, SLAs are legal documents (paper and/or electronic) that encode the nature and scope of QoS parameters (e.g., ensure availability 99.99% and ensure web application server latency to be less than 100 ms). To tackle uncertainties, recent research and industry efforts 2 have focused on developing monitoring techniques and frameworks that can assist cloud providers and application owners in (i) keeping their resources and applications operating at peak efficiency, (ii) detecting variations in resource and application performance, (iii) accounting the SLA violations of certain QoS parameters, and (iv) tracking the leave and join operations of cloud resources due to failures and other dynamic configuration changes. The rest of this editorial note is organized as follows: Section 2 gives a brief overview of the research and development work carried out for monitoring application QoS over cloud resources; Section 3 summarizes the research contributions that were accepted for this special issue; Section 3 concludes the paper with some future remarks. In last 20 years, a large body of research has focused on developing tools and techniques for monitoring the QoS status of resources and applications over distributed systems (e.g., grids, clusters, and clouds). Some QoS monitoring techniques have been investigated and implemented in computational grids, such as Network Weather Service (NWS) 3, which monitors the network and computing resource QoS and periodically forecast the QoS in a future arrival of a application workload. The current version of NWS gathers the operating system level metrics such as available CPU percentage, available non-paged memory, and TCP/IP Performance. Other monitoring tools 4, 5 that were popular in grid and cluster computing era included R-GMA, Hawkeye, Ganglia, MDS-I, and MDS-II. Aforementioned monitoring techniques and tools were designed for managing static system configuration, where numbers of hardware and software resource types were assumed to remain constant over lifecycle of an application. In other words, these tools did not consider the issue of auto scaling and de-scaling primitives supported by virtualized cloud resources. These tools were only concerned about monitoring the QoS parameters for the hardware resources (CPU, storage, and network), while being completely agnostic to application-specific QoS parameters and SLA requirements. The performance of these tools was optimized for monitoring the QoS of only one type of application (e.g., high performance computing application). On the other hand, in cloud computing datacenters, multiple application instances can be multiplexed and co-allocated on single physical resource. Clearly, the monitoring tools developed in grid and cluster computing era (while being innovative and useful) is not suitable to tackle the challenges on cloud computing environments and hosted application types. Current cloud resource and application QoS monitoring frameworks (e.g., Amazon CloudWatch 6, Azure Fabric Controller) typically monitor the entire virtual machine (VM, a software implementation of a physical CPU resource) as a black box and lacks ability to inter-operate across cloud datacenters managed by different providers (e.g., Amazon, Microsoft, GoGrid, and CA). This means that QoS of software resources (e.g., web server, and database server) contained in the application stack is not properly monitored and managed. While frameworks such as Monitis 7 and Nimsoft 8 overcome the aforementioned limitations of CloudWatch and Fabric Controller, they lack ability to monitor and enforce application-specific QoS requirements. Further, all of the aforementioned frameworks lack ability to predict and detect faults before they occur. Some of the recent research works 9 have also focused on applying large-scale data and pattern mining to the QoS monitoring history and event log data. Authors in 10 evaluated the prediction capability of Support Vector Machine, Neural Network, and Linear Regression techniques for learning the QoS behavior of cloud hosted applications. To predict the CPU usage of VMs, authors in 11 applied Markov Chain model. Authors in 12 applied prediction techniques such as Moving Average, Auto Regression, Neural Networks, Support Vector Machines, and Gene Expression Programming for predictive VM QoS monitoring and provisioning. Most of these techniques focused on monitoring and predicting QoS of VMs rather than individual application components. Further, these approaches did not reason about the interplay of QoS parameters and SLA requirements across multiple layers (software as a service, platform as a service, and infrastructure as a service) of cloud application stack. In this special issue, we present seven articles that tackle several aspects of the aforementioned resource uncertainties for monitoring QoS of applications hosted on Cloud resources. In particular, Ryckbosch and Diwan propose a Temporal Pattern Analyzer system in their paper 13 Analyzing Performance Traces Using Temporal Formulas that uses formulas in linear-temporal logic extended with variables to analyze traces to investigate long-tail performance problems at Google and reduce the manual labor involved in analyzing traces. The technique is applied on user request logs, which contain events at each stage of processing of a user request to Gmail. The authors show that the system can scale to large traces, a prerequisite considering that Gmail produces a million or more events a second. Two of the case studies presented in the paper have directly contributed to improving the performance of Google. Cao et al. also use execution trace information, in this case, CPU load traces and propose 14 a novel method for CPU load prediction for cloud environment based on a dynamic ensemble model to obtain better performances. The ensemble model proposed consists of two layers, a predictor optimization layer that can continuously incorporate new predictor instances and remove those ones with a poor performance and an ensemble layer that is responsible for producing the final prediction based on the results of multiple predictor instances. The four papers are all concerned with monitoring Cloud applications, ranging from a model and language to define design-time adaption techniques in the paper 15 by Inzinger et al. on a Generic Event-Based Monitoring and Adaptation Methodology for Heterogeneous Distributed Systems, to better visualization techniques in the monitoring process in the paper 16 A Novel Monitoring Mechanism by Event Trigger for Hadoop System Performance Analysis by Chang et al., to adapting to failed application service in a distributed environment by introducing fault avoidance service that can be called instead of the failed service by Gülcü et al. in their paper 17 Fault Masking as a Service, to a feature-based high availability mechanism that monitors data streams for a quantile feature in the paper 18 by Ding et al. on a Feature-based High Availability Mechanism for Quantile Tasks in Real-time Data Stream Processing. In particular, Inzinger et al. present 15 a novel domain-specific language termed MONINA that allows specification of system components and their monitoring and adaptation-relevant behavior for controlling Cloud systems. The authors propose a mechanism for optimal deployment of the defined control operators onto available computing resources by monitoring the cloud environment with complex-event processing queries and adapt to problems by condition action rules performed on top of a distributed knowledge base. Chang et al. propose 16 a system called Event Trigger that provides an automatic recording mechanism on the Hadoop Cloud Computing system, to check the system performance at every static time interval, and compares the variation. The performance parameters are collected during the system monitoring process and are applied onto an easy-understandable visual graph for users to adjust the hardware deployment in order to refine the Hadoop system. Gülcü et al. propose 17 an approach to prevent the occurrence of errors that result from the unavailability of partner services in the first place. They introduce a fault avoidance service to which composite services can register at will. After registration, this fault avoidance service periodically checks the partner links, detects unavailable partner services, and updates the composite service with available alternatives. Thus, in case of a partner service error, the composite service will have been updated before attempting an ill-destined request. Ding et al. focus 18 on the monitoring of data streams on the quantile tasks, a typical summary-oriented operation for aggregation, and propose a feature-based high availability mechanism to reduce related overhead and latency. With the help of a monitor module, the quantile feature is maintained incrementally through histogram synopsis over a time-based sliding window. Consequently, failed tasks can be recovered precisely with a high probability in an efficient way. Finally, the special issue is rounded off by a paper 19 on Design and Implementation of Task Scheduling Strategies for Massive Remote Sensing Data Processing Across Multiple Data Centers by Zhang et al. that proposes scheduling strategies for data processing workflows. In particular, they propose scheduling strategies in massive remote sensing data processing to reduce the total task execution time. The authors divided the data processing workflows into two categories, namely, Bag of Tasks applications that consist of a large number of independent tasks and Direct Acyclic Graph applications that contain a large number of interdependent tasks. They propose two strategies to deal with issues in either of the two categories, a Partitioning Group based on Hypergraph algorithm that partitions data into several groups to minimize the amount of sharing data transferring and an Optimized Task Tree strategy to find the key workflow path, which would be endowed with a high priority in the execution. This special issue presents, through these seven papers, several techniques that can dynamically predict and capture the relationship between an application performance targets, current hardware resource allocation, and changes in workload patterns, in order to adjust resource configuration at design-time and run-time. More work in this area is rapidly emerging, further improving the availability of massively distributed Cloud applications. This will further improve the economies of scale of Cloud applications making Cloud Computing an even more compelling paradigm in comparison to traditional in-house hosted applications. Application QoS monitoring will continue to remain an important research area for cloud-based systems. More tangible efforts are needed for developing monitoring tools and techniques that can specify, reason, and monitor QoS related to a variety of application (enterprise, scientific, and streaming big data analytics) and cloud datacenter types (private and public). Further, research should also aim to correlate events with data from many different sources (e.g., holiday schedules, job schedules, and trends from social media about application usage sentiment) in order to predict how external events can impact an application QoS. In this special issue, we have selected research papers that aim to address some of these challenges. We hope that the readers will find the articles of this special issue informative and useful.
Rajiv Ranjan 0001, Rajkumar Buyya, Philipp Leitner 0001, Armin Haller, Stefan Tai
Softw. Pract. Exp.1
2014 Authorized Public Auditing of Dynamic Big Data Storage on Cloud with Efficient Verifiable Fine-Grained Updates
abstract
Cloud computing opens a new era in IT as it can provide various elastic and scalable IT services in a pay-as-you-go fashion, where its users can reduce the huge capital investments in their own IT infrastructure. In this philosophy, users of cloud storage services no longer physically maintain direct control over their data, which makes data security one of the major concerns of using cloud. Existing research work already allows data integrity to be verified without possession of the actual data file. When the verification is done by a trusted third party, this verification process is also called data auditing, and this third party is called an auditor. However, such schemes in existence suffer from several common drawbacks. First, a necessary authorization/authentication process is missing between the auditor and cloud service provider, i.e., anyone can challenge the cloud service provider for a proof of integrity of certain file, which potentially puts the quality of the so-called ‘auditing-as-a-service’ at risk; Second, although some of the recent work based on BLS signature can already support fully dynamic data updates over fixed-size data blocks, they only support updates with fixed-sized blocks as basic unit, which we call coarse-grained updates. As a result, every small update will cause re-computation and updating of the authenticator for an entire file block, which in turn causes higher storage and communication overheads. In this paper, we provide a formal analysis for possible types of fine-grained data updates and propose a scheme that can fully support authorized auditing and fine-grained update requests. Based on our scheme, we also propose an enhancement that can dramatically reduce communication overheads for verifying small updates. Theoretical analysis and experimental results demonstrate that our scheme can offer not only enhanced security and flexibility, but also significantly lower overhead for big data applications with a large number of frequent small updates, such as applications in social media and business transactions.
Chang Liu 0001, Jinjun Chen, Laurence T. Yang, Xuyun Zhang, Chi Yang, Rajiv Ranjan 0001, Kotagiri Ramamohanarao
IEEE Trans. Parallel Distributed Syst.6
2014 Task-Tree Based Large-Scale Mosaicking for Massive Remote Sensed Imageries with Dynamic DAG Scheduling
abstract
Remote sensed imagery mosaicking at large scale has been receiving increasing attentions in regional to global research. However, when scaling to large areas, image mosaicking becomes extremely challenging for the dependency relationships among a large collection of tasks which give rise to ordering constraint, the demand of significant processing capabilities and also the difficulties inherent in organizing these enormous tasks and RS image data. We propose a task-tree based mosaicking for remote sensed imageries at large scale with dynamic DAG scheduling. It expresses large scale mosaicking as a data-driven task tree with minimal height. And also a critical path based dynamical DAG scheduling solution with status queue named CPDS-SQ is provided to offer an optimized schedule on multi-core cluster with minimal completion time. All the individual dependent tasks are run by a core parallel mosaicking program implemented with MPI to perform mosaicking on different pairs of images. Eventually, an effective but easier approach is offered to improve the large-scale processing capability by decoupling the dependence relationships among tasks from the complex parallel processing procedure. Through experiments on large-scale mosaicking, we confirmed that our approach were efficient and scalable.
Yan Ma 0001, Lizhe Wang 0001, Albert Y. Zomaya, Dan Chen 0001, Rajiv Ranjan 0001
IEEE Trans. Parallel Distributed Syst.5
2013 Early Observations on Performance of Google Compute Engine for Scientific Computing
abstract
Although Cloud computing emerged for business applications in industry, public Cloud services have been widely accepted and encouraged for scientific computing in academia. The recently available Google Compute Engine (GCE) is claimed to support high-performance and computationally intensive tasks, while little evaluation studies can be found to reveal GCE's scientific capabilities. Considering that fundamental performance benchmarking is the strategy of early-stage evaluation of new Cloud services, we followed the Cloud Evaluation Experiment Methodology (CEEM) to benchmark GCE and also compare it with Amazon EC2, to help understand the elementary capability of GCE for dealing with scientific problems. The experimental results and analyses show both potential advantages of, and possible threats to applying GCE to scientific computing. For example, compared to Amazon's EC2 service, GCE may better suit applications that require frequent disk operations, while it may not be ready yet for single VM-based parallel computing. Following the same evaluation methodology, different evaluators can replicate and/or supplement this fundamental evaluation of GCE. Based on the fundamental evaluation results, suitable GCE environments can be further established for case studies of solving real science problems.
Zheng Li 0001, Liam O'Brien, Rajiv Ranjan 0001, Miranda Zhang
CloudCom (1)3
2013 Adaptive workflow scheduling for dynamic grid and cloud computing environment
abstract
SUMMARY Effective scheduling is a key concern for the execution of performance‐driven grid applications such as workflows. In this paper, we first define the workflow scheduling problem and describe the existing heuristic‐based and metaheuristic‐based workflow scheduling strategies in grids. Then, we propose a dynamic critical‐path‐based adaptive workflow scheduling algorithm for grids, which determines efficient mapping of workflow tasks to grid resources dynamically by calculating the critical path in the workflow task graph at every step. Using simulation, we compared the performance of the proposed approach with the existing approaches, discussed in this paper for different types and sizes of workflows. The results demonstrate that the heuristic‐based scheduling techniques can adapt to the dynamic nature of resource and avoid performance degradation in dynamically changing grid environments. Finally, we outline a hybrid heuristic combining the features of the proposed adaptive scheduling technique with metaheuristics for optimizing execution cost and time as well as meeting the users requirements to efficiently manage the dynamism and heterogeneity of the hybrid cloud environment. Copyright © 2013 John Wiley & Sons, Ltd.
Mustafizur Rahman 0003, Md. Rafiul Hassan, Rajiv Ranjan 0001, Rajkumar Buyya
Concurr. Comput. Pract. Exp.3
2013 Special issue: second international workshop on workflow management in service and cloud computing (WMSC2010)
abstract
This special issue of Concurrency and Computation: Practice and Experience contains selected high-quality papers from the Second International Workshop on Workflow Management in Service and Cloud Computing (WMSC2010) that was held on 11–13 December 2010 in Hong Kong 1. The WMSC workshop series aims to provide an international forum for the presentation and discussion of research and development trends regarding workflow support in service and cloud environments. WMSC2010 attracted many international attendants, allowing deep discussion and the exchange of ideas and results related to ongoing research among attendants. Following ICWM2009 on 4 May 4 2009 in Geneva, Switzerland, WaGe2008 on 25 May 2008 in Kunming China and WaGe2007 on 17 August 2007 in Urumqi China; WMSC2010 continues to discuss workflow management in service and cloud environments from different perspectives and areas in order to tackle different potentials for further research and development. Workflow management in distributed computing environments has been under investigation for several years 2-8. In particular, the special issue titled Workflow in Grid Systems in Concurrency and Computation: Practice and Experience was a key step 7. The special issue was edited by Professor Geoffrey C. Fox and Professor Dennis Gannon from Indiana University in USA. A follow-up were the special issues in the same journal for WSGE2006 (first International Workshop on Workflow Systems in Grid Environments), WaGe2007, WaGe2008 and ICWM2009 9-12. This WMSC2010 special issue is another follow-up of those special issues in order to further boost the research and development of workflow management and applications. Many research and development efforts have been made in the field of workflow management and applications in distributed service and cloud environments such as 2-8, 13-21. More and more people from different areas are trying to facilitate the techniques from their respective areas to tackle tough issues in workflow management such as resource scheduling, security, computation reduction, service discovery and composition and data service query issues. Following the special issue of ICWM2009, this special issue continues to accommodate a range of papers from different perspectives and areas such as service computing, cloud computing, authentication/security in order to provide some different views and hints for workflow management research. This special issue contains nine papers based on those that were presented at WMSC2010. They are listed as 22-30. Research problems in these papers have been analysed systematically, and for specific approaches or models, evaluation has been performed to demonstrate their feasibility and advantages. The nine papers were selected on this basis and also peer reviewed thoroughly. They are summarised in the succeeding text. This paper 22 is a workflow scheduling in dynamic grid environment. The paper defines the workflow scheduling problem and describes the existing heuristic and meta-heuristic-based workflow scheduling strategies in Grids. Then, we propose a dynamic critical path-based adaptive workflow scheduling algorithm for Grids, which determines efficient mapping of workflow tasks to Grid resources dynamically by calculating the critical path in the workflow task graph at every step. Corresponding evaluation is conducted to demonstrate the performance. This paper 23 is about service discovery in elastic cloud computing environment. A QoS-aware service discovery method is investigated for elastic cloud computing in an unstructured P2P network. The method is deployed by two phases, that is, service registering phase and service discovery phase. More specifically, for a peer node engaged in the unstructured P2P network, it firstly registers its functional and non-functional information to its neighbours in a flooding way. The simulations are conducted to evaluate the feasibility of our method. This paper 24 is about service-oriented business ecosystem (SOBE). It presents BSNet, a model based on the service correlation networks to manage the SOBE. The model consists of a who–what–how service correlation network that captures the various relations in SOBE. Finally, a prototyping system has been developed to demonstrate the value of this model for business service management and a simulation-based case study is also provided. This paper 25 proposes the concept of graph refactoring, which transforms certain types of sequential tasks to run in parallel without changing the system's functionality. Experiments and analysis show that graph refactoring can improve the system performance scalable because of concurrent execution of previously sequential tasks. This paper 26 proposes a novel Virtual Organization (VO) creating algorithm (called Group-Choose) based on reputation system that can help initiator to minimise the operating risk and guarantee the success on Internet. In our model, a VO initiator aggregate selected partner's trust to find more appropriate new partners for VO, instead of evaluating their trust by himself only. Simulation results illustrate that VO creating with Group-Choose has more stable and higher success rates for tasks workflow executing under various kind attacks. This paper 27 proposes a service correlation context aware for composite service selection approach is proposed in this paper. Referring to the concept of ‘single-entry single-exit (SESE) region’ in compiler theory, the paper proposes the concept of ‘SESE pattern’ and uses it in composite service selection. Experimental results demonstrate that the approach can improve the quality of selected composite services effectively in the correlation context. This paper 28 proposes an approach to mine batch processing workflow models from event logs by considering the batch processing relations among activity instances in multiple workflow cases. The notion of batch processing feature and its corresponding mining algorithm are also presented for discovering the batch processing area in the model by using the input and output data information of activity instances in events. The algorithms presented in this paper can help to enhance the applicability of existing process mining approaches and broaden the process mining spectrum. This paper 29 proposes an event view model in order to better support service process collaboration. The model is composed of a set of event types and their dependency relationships. It provides a general and flexible way to define a public view of a service process model and serves as the basis for defining service process collaboration protocols. A case study is presented and some implementation issues for defining and publishing an event view are discussed. This paper 30 is about encrypted database query in service cloud environment. The paper proposes a nonlinear order-preserving scheme for indexing encrypted data, which facilitates the range queries over encrypted databases. The scheme is secure even there are a large number of duplicates in plaintexts. This scheme is suitable for long-standing databases because its use does not need any assumption on the database data such as their distribution, range and number, which may change dramatically over time.
Jinjun Chen, Rajiv Ranjan 0001
Concurr. Comput. Pract. Exp.2
2013 A scalable Helmholtz solver in GRAPES over large-scale multicore cluster
abstract
SUMMARY This paper discusses performance optimization on the dynamical core of global numerical weather prediction model in Global/Regional Assimilation and Prediction System (GRAPES). GRAPES is a new generation of numerical weather prediction system developed and currently used by Chinese Meteorology Administration. The computational performance of the dynamical core in GRAPES relies on the efficient solution of three‐dimensional Helmholtz equations, which lead to large‐scale and sparse linear systems formulated by the discretization in space and time. We choose generalized conjugate residual (GCR) algorithm to solve the corresponding linear systems and further propose algorithm optimizations for large‐scale parallelism in two aspects: (i) reduction of iteration number for solution and (ii) performance enhancement of each GCR iteration. The reduction of iteration number is achieved by advanced preconditioning techniques, combining block incomplete LU factorization‐k preconditioner over 7‐diagonals of the coefficient matrix with the restricted additive Schwarz method effectively . The improvement for GCR iteration is to reduce the global communication operations by refactoring the GCR algorithm, which decreases the communication overhead over large number of cores. Performance evaluation on the Tianhe‐1A system shows that the new preconditioning techniques reduce almost one‐third iterations for solving the linear systems, the proposed methods can obtain 25% performance improvement on average compared with the original version of Helmholtz solver in GRAPES, and the speedup with our algorithms can reach 10 using 2048 cores compared with 256 cores. Copyright © 2013 John Wiley & Sons, Ltd.
Wei Xue 0003, Rajiv Ranjan 0001, Zhiyan Jin
Concurr. Comput. Pract. Exp.3
2013 On construction of heuristic QoS bandwidth management in clouds
abstract
ABSTRACT In recent years, cloud computing has become popular and its applications widespread. Thus, there exists a common concern, that is, how to arrange and monitor various resources in the cloud computing environment. In the literature, Ganglia and Network Weather Service (NWS) were used to monitor and gather node status and network‐related data, respectively. With supports of Ganglia and NWS, one can effectively administer available resources in the cloud computing environment. In order to achieve high performance of cloud computing, comprehensive monitoring and efficient management are critical. Ganglia is often used to gather status data of resources, such as live states of hosts, CPU or memory utilizations, and surely Ganglia is also capable of monitoring network‐related information; however, instead of Ganglia, we used NWS services to gather network‐related information such as end‐to‐end transmission control protocol/Internet protocol performance data. Compared with Ganglia, NWS services offer more selections and flexibility for measurement schemes. Besides, NWS services could be deployed with nonintruding manner that makes it easier and faster in deploying services to cloud nodes. The network‐related information is acquired immediately after deployment. Although NWS services also provide measurements for CPU and memory utilizations, but less functionality is provided by them than Ganglia in these aspects. Therefore, we combine advantageous features of Ganglia and NWS to achieve the aims of effective monitoring and management of available resources in the cloud environment. Nevertheless, Ganglia and NWS services may not provide sufficient data in realistic situations due to diversified needs of users, especially application developers. For instance, users are not able to directly access utilizations or allocations of resources in the cloud environment via interfaces or channels of Ganglia or NWS. In addition, NWS services based on a domain‐based network information model could greatly decrease overheads caused by unnecessary measurements. Hence, we propose a heuristic QoS measurement approach based on the domain‐based information model. This measurement approach is capable of providing essential information to satisfy user requirements, and thus let users manage and monitor various resources in the cloud environment in a more efficient way. © 2013 Wiley Periodicals, Inc.
Chao-Tung Yang, Jung-Chun Liu, Rajiv Ranjan 0001, Wen-Chung Shih, Chih-Hao Lin
Concurr. Comput. Pract. Exp.3
2013 Model-driven provisioning of application services in hybrid computing environments
Rajiv Ranjan 0001, Rajkumar Buyya, Surya Nepal
Future Gener. Comput. Syst.1
2013 Energy-aware parallel task scheduling in a cluster
Lizhe Wang 0001, Samee Ullah Khan, Dan Chen 0001, Joanna Kolodziej, Rajiv Ranjan 0001, Cheng-Zhong Xu 0001, Albert Y. Zomaya
Future Gener. Comput. Syst.5
2013 G-Hadoop: MapReduce across distributed data centers for data-intensive computing
Lizhe Wang 0001, Jie Tao 0001, Rajiv Ranjan 0001, Holger Marten, Achim Streit, Jingying Chen 0001, Dan Chen 0001
Future Gener. Comput. Syst.3
2013 A workload-driven approach to database query processing in the cloud
Adnene Guabtni, Rajiv Ranjan 0001, Fethi A. Rabhi
J. Supercomput.2
2013 Peer-to-peer service provisioning in cloud computing environments
Rajiv Ranjan 0001, Liang Zhao 0009
J. Supercomput.1
2012 Cloud monitoring for optimizing the QoS of hosted applications
abstract
Cloud monitoring involves dynamically tracking the Quality of Service (QoS) parameters related to virtualized services (e.g., CPU, storage, network, appliances, etc.), the physical resources they share, and the applications running on them or data hosted on them. Monitoring techniques and services can help a cloud provider or application developer in regards to: (i) keeping the cloud services and hosted applications operating at peak efficiency; (ii) detecting variations in service and application performance; (iii) accounting the SLA violations of certain QoS parameters; and (iv) tracking the leave and join operations of cloud services due to failures and other dynamic configuration changes. In this paper, we describe the PhD research motivation, question, and approach and methodology related to developing novel cloud monitoring techniques and services enabling automated application QoS management under uncertainties.
Khalid Alhamazani, Rajiv Ranjan 0001, Fethi A. Rabhi, Lizhe Wang 0001, Karan Mitra
CloudCom2
2012 Investigating decision support techniques for automating Cloud service selection
abstract
The compass of Cloud infrastructure services advances steadily leaving users in the agony of choice. To be able to select the best mix of service offering from an abundance of possibilities, users must consider complex dependencies and heterogeneous sets of criteria. Therefore, we present a PhD thesis proposal on investigating an intelligent decision support system for selecting Cloud-based infrastructure services (e.g. storage, network, CPU). The outcomes of this will be decision support tools and techniques, which will automate and map users' specified application requirements to Cloud service configurations.
Miranda Zhang, Rajiv Ranjan 0001, Armin Haller, Dimitrios Georgakopoulos 0001, Peter E. Strazdins
CloudCom2
2012 An ontology-based system for Cloud infrastructure services' discovery
abstract
The Cloud infrastructure services landscape advances steadily leaving users in the agony of choice. As a result, Cloud service dentification and discovery remains a hard problem due to different service descriptions, nonstandardised naming conventions and heterogeneous types and features of Cloud
Miranda Zhang, Rajiv Ranjan 0001, Armin Haller, Dimitrios Georgakopoulos 0001, Michael Menzel 0002, Surya Nepal
CollaborateCom2
2012 Parallel Processing of Massive EEG Data with MapReduce
abstract
Analysis of neural signals like electroencephalogram (EEG) is one of the key technologies in detecting and diagnosing various brain disorders. As neural signals are non-stationary and non-linear in nature, it is almost impossible to understand their true physical dynamics until the recent advent of the Ensemble Empirical Mode Decomposition (EEMD) algorithm. The neural signal processing with EEMD is highly compute-intensive due to the high complexity of the EEMD algorithm. It is also data intensive because 1) EEG signals contain massive data sets 2) EEMD has to introduce a large number of trials in processing to ensure precision. The Map Reduce programming mode is a promising parallel computing paradigm for data intensive computing. To increase the efficiency and performance of the neural signal analysis, this research develops parallel EEMD neural signal processing with Map Reduce. In this paper, we implement the parallel EEMD with Hadoop in a modern cyber infrastructure. Test results and performance evaluation show that parallel EEMD can significantly improve the performance of neural signal processing.
Lizhe Wang 0001, Dan Chen 0001, Rajiv Ranjan 0001, Samee Ullah Khan, Joanna Kolodziej, Jun Wang 0001
ICPADS3
2012 Do-It-Yourself Content Delivery Network Orchestrator
Rajiv Ranjan 0001, Karan Mitra, Suhit Saha, Dimitrios Georgakopoulos 0001, Arkady B. Zaslavsky
WISE1
2012 CloudGenius: decision support for web server cloud migration
abstract
Cloud computing is the latest computing paradigm that delivers hardware and software resources as virtualized services in which users are free from the burden of worrying about the low-level system administration details. Migrating Web applications to Cloud services and integrating Cloud services into existing computing infrastructures is non-trivial. It leads to new challenges that often require innovation of paradigms and practices at all levels: technical, cultural, legal, regulatory, and social. The key problem in mapping Web applications to virtualized Cloud services is selecting the best and compatible mix of software images (e.g., Web server image) and infrastructure services to ensure that Quality of Service (QoS) targets of an application are achieved. The fact that, when selecting Cloud services, engineers must consider heterogeneous sets of criteria and complex dependencies between infrastructure services and software images, which are impossible to resolve manually, is a critical issue. To overcome these challenges, we present a framework (called CloudGenius) which automates the decision-making process based on a model and factors specifically for Web server migration to the Cloud. CloudGenius leverages a well known multi-criteria decision making technique, called Analytic Hierarchy Process, to automate the selection process based on a model, factors, and QoS parameters related to an application. An example application demonstrates the applicability of the theoretical CloudGenius approach. Moreover, we present an implementation of CloudGenius that has been validated through experiments.
Michael Menzel 0002, Rajiv Ranjan 0001
WWW2
2012 Special section on autonomic cloud computing: technologies, services, and applications
abstract
Special section on autonomic cloud computing: technologies,
Rajiv Ranjan 0001, Rajkumar Buyya, Manish Parashar
Concurr. Comput. Pract. Exp.1
2012 Special section: software architectures and application development environments for Cloud computing
abstract
Special section: software architectures and application development environments for Cloud computingWelcome to the special issue of Software: Practice and Experience journal on Cloud computing.This special issue compiles a number of excellent technical contributions that significantly advance the state-of-the-art of software architectures and application development environments for cloud computing.Cloud computing [1-3] is positioning itself as a promising platform for delivering infrastructureas-a-service (IaaS), platform-as-a-service (PaaS), and software-as-a-service (SaaS) as services.Clouds aim to power the next-generation data centers by architecting them as a network of virtual services (hardware, database, user-interface, application logic) so that users are able to deploy and access applications globally and on demand at competitive costs depending on users' QoS requirements.Cloud infrastructures are exposed through collections of software services at SaaS and PaaS layers designed to support creation and deployment of application services.To this end, developing scalable architectures and application development environments to build, access, manage, deploy, and maintain applications in clouds in a developer-friendly manner has become critical.Several vendors have emerged in this space including IBM, VMware, Microsoft, Manjrasoft, and Yahoo.This model of computing is quite attractive, especially for small and medium size enterprises, as it allows them to focus on consuming or offering services on top of the Cloud infrastructure.At high level, Cloud computing might not seem radically different from the existing paradigms: World Wide Web, grid computing, service computing, and cluster computing.However, key differentiators of Cloud computing are its technical characteristics such as on-demand resource pooling or rapid elasticity, self-service, almost infinite scalability, end-to-end virtualization support, and robust support of resource usage metering and billing.Additionally, nontechnical differentiators include services that are offered under pay-as-you-go-model, guaranteed SLA, faster time to deployments, lower upfront costs, little or no maintenance overhead, and environment friendliness.Public IaaS and PaaS vendors including Amazon, Microsoft, Google, and GoGrid offer different types of software programming architectures and interfaces.Next, these are implemented using different programming environments, hence should be accessed through vendor-dependent adapter interfaces.In particular current Cloud programming approaches have the following limitations: (i) requires human familiarity with different types of Cloud resources and typically rely on procedural programming in general purpose or scripting languages; (ii) interaction with Cloud resources is mainly performed through low-level APIs and command line interfaces; (iii) SaaS (application) implementation is dependent on the programming environment supported by the IaaS and Paas vendors; and (iv) lacks flexibility and efficiency of supporting generic applications that can be simultaneously deployed across multiple Cloud vendors infrastructure.Hence, it is clear that developing system architecture and application development environments that can simplify and improve the task of Cloud programming are key to harnessing the capability of clouds.In this special issue, we have featured high quality papers that deal with some of the aforementioned issues.All of the selected papers underwent a rigorous peer-review process and their contributions are briefly discussed below:Though Cloud computing infrastructure services enable the flexible creation of virtual infrastructures on demand basis, it is only a tiny step of the overall complex process required for provisioning application services.Other steps such as installation, deployment, configuration, monitoring, and management of software components are needed to fully provide services to end-users in the Cloud.To this end, in the paper titled 'Towards an Architecture for Deploying Elastic Services in the Cloud', Kirschnick et al. [4] describes a peer-to-peer architecture to automatically deploy services
Rajiv Ranjan 0001, Rajkumar Buyya, Boualem Benatallah
Softw. Pract. Exp.1
2012 Coordinated load management in Peer-to-Peer coupled federated grid systems
Rajiv Ranjan 0001, Aaron Harwood, Rajkumar Buyya
J. Supercomput.1
2011 Virtual Machine Provisioning Based on Analytical Performance and QoS in Cloud Computing Environments
abstract
Cloud computing is the latest computing paradigm that delivers IT resources as services in which users are free from the burden of worrying about the low-level implementation or system administration details. However, there are significant problems that exist with regard to efficient provisioning and delivery of applications using Cloud-based IT resources. These barriers concern various levels such as workload modeling, virtualization, performance modeling, deployment, and monitoring of applications on virtualized IT resources. If these problems can be solved, then applications can operate more efficiently, with reduced financial and environmental costs, reduced under-utilization of resources, and better performance at times of peak load. In this paper, we present a provisioning technique that automatically adapts to workload changes related to applications for facilitating the adaptive management of system and offering end-users guaranteed Quality of Services (QoS) in large, autonomous, and highly dynamic environments. We model the behavior and performance of applications and Cloud-based IT resources to adaptively serve end-user requests. To improve the efficiency of the system, we use analytical performance (queueing network system model) and workload information to supply intelligent input about system requirements to an application provisioner with limited information about the physical infrastructure. Our simulation-based experimental results using production workload models indicate that the proposed provisioning technique detects changes in workload intensity (arrival pattern, resource demands) that occur over time and allocates multiple virtualized IT resources accordingly to achieve application QoS targets.
Rodrigo N. Calheiros, Rajiv Ranjan 0001, Rajkumar Buyya
ICPP2
2011 A taxonomy and survey on autonomic management of applications in grid computing environments
abstract
Abstract In Grid computing environments, the availability, performance, and state of resources, applications, services, and data undergo continuous changes during the life cycle of an application. Uncertainty is a fact in Grid environments, which is triggered by multiple factors, including: (1) failures, (2) dynamism, (3) incomplete global knowledge, and (4) heterogeneity. Unfortunately, the existing Grid management methods, tools, and application composition techniques are inadequate to handle these resource, application and environment behaviors. The aforementioned characteristics impose serious requirements on the Grid programming and runtime systems if they wish to deliver efficient performance to scientific and commercial applications. To overcome the above challenges, the Grid programming and runtime systems must become autonomic or self‐managing in accordance with the high‐level behavior specified by system administrators. Autonomic systems are inspired by biological systems that deal with similar challenges of complexity, dynamism, heterogeneity, and uncertainty. To this end, we propose a comprehensive taxonomy that characterizes and classifies different software components and high‐level methods that are required for autonomic management of applications in Grids. We also survey several representative Grid computing systems that have been developed by various leading research groups in the academia and industry. The taxonomy not only highlights the similarities and differences of state‐of‐the‐art technologies utilized in autonomic application management from the perspective of Grid computing, but also identifies the areas that require further research initiatives. We believe that this taxonomy and its mapping to relevant systems would be highly useful for academic‐ and industry‐based researchers, who are engaged in the design of Autonomic Grid and more recently, Cloud computing systems. Copyright © 2011 John Wiley & Sons, Ltd.
Mustafizur Rahman 0003, Rajiv Ranjan 0001, Rajkumar Buyya, Boualem Benatallah
Concurr. Comput. Pract. Exp.2
2011 CloudSim: a toolkit for modeling and simulation of cloud computing environments and evaluation of resource provisioning algorithms
abstract
Abstract Cloud computing is a recent advancement wherein IT infrastructure and applications are provided as ‘services’ to end‐users under a usage‐based payment model. It can leverage virtualized services even on the fly based on requirements (workload patterns and QoS) varying with time. The application services hosted under Cloud computing model have complex provisioning, composition, configuration, and deployment requirements. Evaluating the performance of Cloud provisioning policies, application workload models, and resources performance models in a repeatable manner under varying system and user configurations and requirements is difficult to achieve. To overcome this challenge, we propose CloudSim: an extensible simulation toolkit that enables modeling and simulation of Cloud computing systems and application provisioning environments. The CloudSim toolkit supports both system and behavior modeling of Cloud system components such as data centers, virtual machines (VMs) and resource provisioning policies. It implements generic application provisioning techniques that can be extended with ease and limited effort. Currently, it supports modeling and simulation of Cloud computing environments consisting of both single and inter‐networked clouds (federation of clouds). Moreover, it exposes custom interfaces for implementing policies and provisioning techniques for allocation of VMs under inter‐networked Cloud computing scenarios. Several researchers from organizations, such as HP Labs in U.S.A., are using CloudSim in their investigation on Cloud resource provisioning and energy‐efficient management of data center resources. The usefulness of CloudSim is demonstrated by a case study involving dynamic provisioning of application services in the hybrid federated clouds environment. The result of this case study proves that the federated Cloud computing model significantly improves the application QoS requirements under fluctuating resource and service demand patterns. Copyright © 2010 John Wiley & Sons, Ltd.
Rodrigo N. Calheiros, Rajiv Ranjan 0001, Anton Beloglazov, César A. F. De Rose, Rajkumar Buyya
Softw. Pract. Exp.2
2010 InterCloud: Utility-Oriented Federation of Cloud Computing Environments for Scaling of Application Services
Rajkumar Buyya, Rajiv Ranjan 0001, Rodrigo N. Calheiros
ICA3PP (1)2
2010 A Taxonomy of Autonomic Application Management in Grids
abstract
In this paper, we propose a taxonomy that characterizes and classifies different components of autonomic application management in Grids. We also survey several representative Grid systems developed by various projects world-wide to demonstrate the comprehensiveness of the taxonomy. The taxonomy not only highlights the similarities and differences of state-of-the-art technologies utilized in autonomic application management from the perspective of Grid computing, but also identifies the areas that require further research initiatives.
Mustafizur Rahman 0003, Rajiv Ranjan 0001, Rajkumar Buyya
ICPADS2
2010 Reputation-based dependable scheduling of workflow applications in Peer-to-Peer Grids
Mustafizur Rahman 0003, Rajiv Ranjan 0001, Rajkumar Buyya
Comput. Networks2
2010 Special section: Federated resource management in grid and cloud computing systems
Rajkumar Buyya, Rajiv Ranjan 0001
Future Gener. Comput. Syst.2
2010 Cooperative and decentralized workflow scheduling in global grids
Mustafizur Rahman 0003, Rajiv Ranjan 0001, Rajkumar Buyya
Future Gener. Comput. Syst.2
2008 A Decentralized and Cooperative Workflow Scheduling Algorithm
abstract
In the current approaches to workflow scheduling, there is no cooperation between the distributed workflow brokers and as a result, the problem of conflicting schedules occur. To overcome this problem, in this paper, we propose a decentralized and cooperative workflow scheduling algorithm. The proposed approach utilizes a peer-to-peer (P2P) coordination space with respect to coordinating the application schedules among the Grid wide distributed workflow brokers. The proposed algorithm is completely decentralized in the sense that there is no central point of contact in the system, the responsibility of the key functionalities such as resource discovery and scheduling coordination are delegated to the P2P coordination space. With the implementation of our approach, not only the performance bottlenecks are likely to be eliminated but also efficient scheduling with enhanced scalability and better autonomy for the users are likely to be achieved. We prove the feasibility of our approach through an extensive trace driven simulation study.
Rajiv Ranjan 0001, Mustafizur Rahman 0003, Rajkumar Buyya
CCGRID1
2008 A case for cooperative and incentive-based federation of distributed clusters
Rajiv Ranjan 0001, Aaron Harwood, Rajkumar Buyya
Future Gener. Comput. Syst.1
2007 Decentralised Resource Discovery Service for Large Scale Federated Grids
abstract
Efficient resource discovery mechanism is one of the fundamental requirement for grid computing systems, as it aids in resource management and scheduling of applications. Resource discovery involves searching for resources that match the user's application requirements. Various kinds of solutions to grid resource discovery have been developed, including the centralised and hierarchical information server approach. However, these approaches have serious limitations in regards to scalability, fault-tolerance and network congestion. To overcome such limitations, we propose a decentralised grid resource discovery system based on a spatial publish/subscribe index. It utilises a distributed hash table (DHT) routing substrate for delegation of d-dimensional service messages. Our approach has been validated using a simulated publish/subscribe index that assigns regions of a d-dimensional resource attribute space to the grid peers in the system. We generated the resource attribute distribution using the configurations obtained from the top 500 supercomputer list. The simulation study takes into account various parameters such as resource query rate, index load distribution, number of index messages generated, overlay routing hops and system size. Our results show that grid resource query rate directly affects the performance of the decentralised resource discovery system, and that at higher rates the queries can experience considerable latencies. Further, contrary to what one can expect, system size does not have a significant impact on the performance of the system, in particular the query latency.
Rajiv Ranjan 0001, Lipo Chan, Aaron Harwood, Shanika Karunasekera, Rajkumar Buyya
eScience1
2006 SLA-Based Coordinated Superscheduling Scheme for Computational Grids
abstract
The service level agreement (SLA) based grid superscheduling approach promotes coordinated resource sharing. Superscheduling is facilitated between administratively and topologically distributed grid sites via grid schedulers such as resource brokers and workflow engines. In this work, we present a market-based SLA coordination mechanism, based on a well known contract net protocol. The key advantages of our approach are that it allows: (i) resource owners to have finer degree of control over the resource allocation which is something that is not possible with traditional mechanisms; and (ii) superschedulers to bid for SLA contracts in the contract net, with focus on completing a job within a user specified deadline. In this work, we use simulation to show the effectiveness of our proposed approach
Rajiv Ranjan 0001, Aaron Harwood, Rajkumar Buyya
CLUSTER1
2005 A Case for Cooperative and Incentive-Based Coupling of Distributed Clusters
abstract
Interest in grid computing has grown significantly over the past five years. Management of distributed cluster resources is a key issue in grid computing. Central to management of resources is the effectiveness of resource allocation, as it determines the overall utility of the system. In this paper, we propose a new grid system that consists of grid federation agents which couple together distributed cluster resources to enable a cooperative environment. The agents use a computational economy methodology, that facilitates QoS scheduling, with a cost-time scheduling heuristic based on a scalable, shared federation directory. We show by simulation, while some users that are local to popular resources can experience higher cost and/or longer delays, the overall users' QoS demands across the federation are better met. Also, the federation's average case message passing complexity is seen to be scalable, though some jobs in the system may lead to large numbers of messages before being scheduled
Rajiv Ranjan 0001, Rajkumar Buyya, Aaron Harwood
CLUSTER1
2005 A model for cooperative federation of distributed clusters
abstract
Interest in grid computing has grown significantly over the past five years. Management of distributed cluster resources is a key issue in grid computing. Central to management of resources is the effectiveness of resource allocation, as it determines the overall utility of the system. In this paper, we propose a new grid system that consists of grid federation agents which couple together distributed cluster resources to enable a cooperative environment.
Rajiv Ranjan 0001, Rajkumar Buyya, Aaron Harwood
HPDC1