Omer F. Rana

dblp:r/OmerFRana · also Omer Farooq Rana, Omer Rana 0001 · DBLP profile ↗
← Back
253ranked-venue papers
15as first author
78since 2021 · last 2026
0000-0003-3597-2646ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 116 · 13 first-author · 24 since 2021Software engineering, systems software and programming languages · 28 · 13 since 2021Computer networks · 27 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 12 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 18 · 4 since 2021Security and privacy · 17 · 7 since 2021Human-computer interaction and ubiquitous computing · 11 · 4 since 2021Theory of computation · 3Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SenseLess: Minimal Vision, Maximum Insight for Smart Homes
abstract
We present SenseLess, a hybrid anomaly detection framework for smart homes that, during the training phase, automatically labels images without manual annotation by combining sensor-guided detection, self-supervised visual clustering, and unsupervised multi-sensor delay estimation for precise alignment. During operation, the system relies primarily on non-vision sensors and activates a confidence-aware vision model only under low-confidence, thereby preserving privacy while maintaining adaptability. Evaluated in real home monitoring, SenseLess achieved an average label coverage of 97.65% with 94.9% accuracy and reduced vision usage to less than 4% of wall-clock operating time. Calibration mechanisms and minimal configuration requirements support scalability and deployment across diverse residential environments.
Norah Albazzai, Omer F. Rana, Charith Perera
PerCom2
2026 A survey of learning-based intrusion detection systems for in-vehicle networks
abstract
Connected and Autonomous Vehicles (CAVs) have advanced modern transportation by improving the efficiency, safety, and convenience of mobility through automation and connectivity, yet they remain vulnerable to cybersecurity threats, particularly through the insecure Controller Area Network (CAN) bus. Cyberattacks can have devastating consequences in connected vehicles, including the loss of control over critical systems, necessitating robust security solutions. In-vehicle Intrusion Detection Systems (IDSs) offer a promising approach by detecting malicious activities in real time. This survey provides a comprehensive review of state-of-the-art research on learning-based in-vehicle IDSs, focusing on Machine Learning (ML), Deep Learning (DL), and Federated Learning (FL) approaches. Based on the reviewed studies, we critically examine existing IDS approaches, categorising them by the types of attacks they detect—known, unknown, and combined known-unknown attacks—while identifying their limitations. We also review the evaluation metrics used in research, emphasising the need to consider multiple criteria to meet the requirements of safety-critical systems. Additionally, we analyse FL-based IDSs and highlight their limitations. By doing so, this survey helps identify effective security measures, address existing limitations, and guide future research toward more resilient and adaptive protection mechanisms, ensuring the safety and reliability of CAVs.
Muzun Althunayyan, Amir Javed, Omer F. Rana
Comput. Networks3
2026 ProfitAware: A Cost Effective Service Placement Technique for Multi-Access Edge Clouds
abstract
Cloud computing and datacentres provided by companies such as Google, Alibaba, Tencent, and Amazon Web Services still dominate datacentres industry from a business revenue perspective. However, imminent technologies and a variety of digital devices that form part of the Multi-access Edge Clouds (MECs) and Internet of Things (IoT), along with modular applications, are starting to take the central stage. In the MEC scenario, there are multiple providers such as networks, processing resources, and applications that are involved in service offerings and often have conflicting optimisation objectives. For example, the profit of the infrastructure providers would need more users, which can degrade the network performance. Therefore, to ensure service quality, all providers' goals should be kept in mind, in particular, when deciding placements. Game theory, particularly the Stackelberg framework, can be used to deal with such optimisation problems because it can model the conflicting objectives typical in MEC environments-such as communication between service providers (leaders) and users (followers); therefore, enabling efficient service placement decisions. In this paper, we model the optimisation problem as a Stackelberg game and propose a bidding strategy that ensures the objectives of all providers are met. Our evaluations and results, based on real workload traces, suggest that the proposed strategy runs services while ensuring their expected levels of performance (∼0.018%-3.27% loss), energy efficiency (∼14.43%-34.79%), reduced runtimes (or at least comparable to the no migration strategy) therefore, users' costs (∼7.88%-15.89%), and minimizing the response time (∼6.37%-8.18%). Furthermore, approximately 16.93% migrations are avoided.
Muhammad Zakarya, Lee Gillam, Omer F. Rana
IEEE Trans. Cloud Comput.3
2026 LEMON: LLM-Enabled Monitoring for Microservices Orchestration
abstract
The complexity of modern microservice architectures has surpassed the capabilities of traditional orchestration tools, which rely on static, manual workflows. This limits scalability and adaptability, especially to dynamic workloads. This paper argues for a paradigm shift towards self-managing, intent-driven microservices orchestration systems, where human operators express high-level goals in natural language. As a foundational step towards this vision of autonomous orchestration agents, we introduce LEMON: a new architecture that leverages Large Language Models (LLMs) for intelligent microservice monitoring. We also propose a comprehensive classification to structure this emerging field. Our evaluation demonstrates that fine-tuning Small Language Models (SLMs) for intent classification significantly enhances accuracy while ensuring the model's output is reliably structured for automation. Furthermore, our analysis of the trade-offs between model size, precision, and latency provides a practical guide for deploying these systems. We foresee this monitoring capability as the critical ”sense” component toward autonomous microservices orchestration loops. By diagnosing performance bottlenecks from natural language queries, LEMON enables future systems to automatically suggest and execute solutions. This work lays the groundwork for truly self-managing, intent-driven systems.
Markus Götz, Hojjat Baghban, Patrizio Dazzi, Omer F. Rana
IEEE Trans. Serv. Comput.4
2026 OntoSage: Intelligent Human-Building Smartbot for Semantic Smart Building Question Answering
abstract
Abstract Smart buildings remain heterogeneous across sensing infrastructure, metadata quality, legacy protocols, and analytics requirements, hindering reusable human–building natural language interfaces. We present OntoSage , a modular framework for ontologically grounded question answering (QA) and fulfillment of analytic intents over smart building data. The framework (i) leverages Brick Schema-based RDF model with reasoning capabilities, (ii) translates natural language (NL) questions into executable SPARQL via a fine-tuned seq2seq model (T5-Base), and (iii) orchestrates portable analytics microservices that operate on time-series sensor data referenced through ontology-linked UUIDs. A summarization component (open-weights Mistral-7B, zero-shot) converts structured SPARQL/SQL/analytic outputs into concise stakeholder-aware responses without requiring task-specific fine-tuning. We categorize QA complexity into four reasoning classes and report component-level execution metrics supporting these categories. To address portability, we formalize a lightweight adaptation workflow (ontology ingestion $$\rightarrow $$ entity enrichment for NLU $$\rightarrow $$ NL2SPARQL validity checks $$\rightarrow $$ analytics binding) designed to minimize per-building retraining. Reproducibility is enabled through public source code, synthetic and ontology-derived datasets, Docker/Compose service descriptors, and documented supporting scripts “( https://github.com/suhasdevmane/OntoBot )”. The developers’ documentation is publicly accessible “( https://ontosage-docs.github.io )”.
Suhas Devmane, Omer F. Rana, Charith Perera
World Wide Web (WWW)2
2025 Enhancing Remote Conservation in Borneo: NVIS for Wildlife Monitoring and Protection using BearWave
abstract
Biodiversity conservation in fragmented and remote ecosystems often requires labour-intensive, time-consuming fieldwork, placing staff at risk and limiting the scope of long-term monitoring. To address these challenges in a sustainable, locally adaptable manner, we present BearWave, a novel, place-based HF communication framework attuned to local infrastructure constraints and designed to support low-power sensing networks under severe radio-frequency (RF) conditions. By leveraging Near Vertical Incidence Skywave (NVIS) propagation and the FT8 digital modulation technique, BearWave achieves reliable, bidirectional data transfer—even in dense tropical rainforest conditions—while keeping costs under £200 per node and enabling extended battery-powered operation. A case study in UK woodlands, chosen as an environmental analogue to Borneo’s rainforest, demonstrated BearWave’s robustness and adaptability: despite dense vegetation and non-line-of-sight paths, the system maintained over 90% message reliability at distances of up to 25 km using only 1 W of RF power. Notably, the strongest signal propagation and best reception rates occurred during nighttime, reflecting diurnal ionospheric variations. These empirical results confirm that BearWave outperforms conventional technologies such as LoRaWAN or satellite systems in harsh, attenuating environments, offering a scalable, energy-efficient approach that lowers both ecological footprints and staff labor risks. This research advances conservation-focused communication by providing a universal, scientifically validated framework capable of supporting long-term ecological monitoring, poacher detection, and improved animal welfare. Crucially, BearWave’s low-cost, low-impact design broadens access for under-resourced organisations and community-driven conservation programs, where local knowledge and stakeholder insights help shape technology decisions on the ground. By embracing a socio-technical innovation model, BearWave exemplifies how computing can be sustainably embedded in remote ecosystems worldwide.
Mark Andrew Butterworth, Omer F. Rana, Pablo Orozco ter Wengel, Benoit Goossens, Charith Perera
COMPASS2
2025 Co-Creating Sustainable Thermal Adaptation in Place: A Card-Based Toolkit for Localized Comfort Interventions
abstract
Adaptive behavior plays a critical role in how individuals navigate discomfort, influencing their ability to respond to environmental challenges. Addressing thermal discomfort which is a pressing concern in the context of climate adaptation and sustainable living. It requires designers and developers to move beyond traditional Graphical User Interface (GUI) solutions, embracing interactive and physically engaging design approaches. This paper introduces D-FACT (Discomfort Card-Based Toolkit for Facilitating Adaptation), a participatory design tool developed to foster sustainable and inclusive adaptation strategies. Tested in four collaborative workshops with 40 participants, D-FACT enables both designers and non-designers to ideate solutions for adapting to thermal discomfort. Our findings demonstrate the toolkit’s effectiveness in nurturing creative, cooperative approaches that promote adaptive environments while advancing energy efficiency and environmental health. By integrating principles of sustainable design, the toolkit encourages diverse and inclusive participation, resulting in conceptual designs that address localized challenges in thermal adaptation. This work contributes to the discourse on computing and sustainability, offering practical methods for embedding participatory design into efforts to create resilient, equitable, and resource-efficient spaces.
Asma Irfan, Omer F. Rana, Charith Perera
COMPASS2
2025 Toward Scalable and Secure Blockchain in Internet of Things: A Preference-Driven Committee Member Auction Consensus Approach
abstract
Blockchain technology is acclaimed for eliminating the need for a central authority while ensuring stability, security, and immutability. However, its integration into Internet of Things (IoT) environments is hampered by the limited computational resources of IoT devices. Consensus algorithms, vital for blockchain safety and efficiency, often require substantial computational power and face challenges related to security, scalability, and resource demands. To address these critical issues, we propose a novel model that significantly enhances the security and performance of blockchain in IoT environments. Our model introduces three key innovations: (1) a bidirectional-linked blockchain system that strengthens security against long-range attacks by exploiting dual reference points for block validation; (2) the integration of user preferences into the Committee Member Auction (CMA) consensus algorithm, optimizing miner selection to balance resource efficiency with security; and (3) a comprehensive performance and frequency analysis that demonstrates the system’s resilience against double-spend, long-range, and eclipse attacks. The proposed model not only reduces block validation delays but also enhances overall system performance, as evidenced by simulations comparing its effectiveness with existing CMA algorithms. These advancements have the potential to significantly impact the deployment of blockchain in resource-constrained IoT environments, offering a more secure and efficient solution.
Akshaya Mathur, Masoud Barati, Gagangeet Singh Aujla, Omer F. Rana
Distributed Ledger Technol. Res. Pract.4
2025 Enhancing performance of machine learning tasks on edge-cloud infrastructures: A cross-domain Internet of Things based framework
abstract
The Internet of Things (IoT) and Edge-Cloud Computing have been trending technologies over the past few years. In this work, we introduce the Enhanced Optimized-Greedy Nominator Heuristic (EO-GNH), a framework designed to optimize machine learning (ML) and artificial intelligence (AI) application placement in edge environments, aiming to improve Quality of Service (QoS). Developed specifically for sectors such as smart agriculture, industry, and healthcare, EO-GNH integrates asynchronous MapReduce and parallel meta-heuristics to effectively manage AI applications, focusing on execution performance, resource utilization, and infrastructure resilience. The framework carefully addresses the distribution challenges of AI applications, especially Service Function Chains (SFCs), in edge-cloud infrastructures. It contains Data Flow Management, which covers aspects of data storage and data privacy, and also considers factors like regional adaptations, mobile access, and AI model refinement. EO-GNH ensures high availability for forecasting, prediction, and training AI models, operating efficiently within a geo-distributed infrastructure. The proposed strategies within EO-GNH emphasize concurrent multi-node execution, enhancing AI application placement by improving execution time, dependability, and cost-effectiveness. The efficiency of EO-GNH is demonstrated through its impact on QoS in real-time resource management across three application domains, highlighting its adaptability and potential in diverse cross-domain IoT-based environments. • EO-GNH optimizes AI application placement in edge computing environments for improved QoS performance • EO-GNH leverages Parsl for parallel meta-heuristic execution and optimized workload deployment • EO-GNH is faster than distributed NSGA-II in finding solutions to multi-objective optimization problems • EO-GNH delivers edge AI capabilities from federated learning to real-time inference • EO-GNH transforms IoT operations across healthcare, industry, and agriculture applications
Osama Almurshed, Ashish Kaushal, Souham Meshoul, Asmail Muftah, Osama Almoghamis, Ioan Petri, Nitin Auluck, Omer F. Rana
Future Gener. Comput. Syst.8
2025 ToSiM-IoT: Toward a Sustainable Optimization of Machine Learning Tasks in Internet of Things
abstract
With the rise of digital infrastructure and Internet of Things (IoT), a substantial amount of data is continuously generated that needs to be processed efficiently. While modern artificial intelligence (AI) approaches have shown good capabilities in handling large volumes of data, their excessive demands for memory and processing power result in very high utilization of resources. In this work, we propose ToSiM-IoT, an optimization framework that introduces a layer selection approach to identify an ideal mix of active, and inactive layers, using a genetic algorithm for model training. Next, we design a pruning mechanism that identifies performance-critical features using heatmap visualization, during model inference, and eliminates the remaining features. Two machine learning (ML) models: 1) InceptionV3 and 2) VGG16, have been evaluated on an agricultural weed detection scenario, using the DeepWeeds image classification dataset. Experimental results demonstrate that our framework can achieve a significant reduction in model size and training time, while maintaining high accuracy, for both models. Therefore, this approach provides the potential to be efficiently deployed on intelligent IoT systems where computational capabilities are limited.
Ashish Kaushal, Osama Almurshed, Asmail Muftah, Nitin Auluck, Omer F. Rana
IEEE Internet Things J.5
2025 Real-Time Anomaly Detection for Industrial Robotic Arms Using Edge Computing
abstract
The integration of Internet of Things (IoT) devices in industrial applications has become viable due to advancements in ubiquitous computing that enable complex machine learning (ML) tasks on resource-constrained devices. Unlike prior approaches that rely on built-in sensors, our system utilizes externally gathered Inertial Measurement Units (IMU) data for anomaly detection. In this paper, we show that simple 1D-CNN and LSTM models on an ultra-low-power device (Nicla Sense ME) optimized for edge-based industrial anomaly detection can achieve approximately 98 movement-based anomalies (e.g., collisions and joint velocity deviations) in industrial robotic arms. We analyzed an advanced manufacturing scenario where the robotic arm performs three consecutive, distinct tasks (pick-and-place, painting, and screwdriving) and demonstrated that the proposed anomaly detection system is task-independent. We implemented these models ondevice by designing a minimal model architecture and modifying source code to minimize RAM usage and Bluetooth Low Energy (BLE) overhead. Additionally, we examined the challenges of deploying ML models in resource-constrained environments by analyzing various quantization methods and the impact of hyperparameter choices on inference time, accuracy, and memory consumption. Our approach focuses on detecting anomalies directly at the data source which enables true real-time detection with a complete edge computing framework that achieves a 10Hz data frequency and a 250ms inference time when BLE is active. Furthermore, we generated a comprehensive dataset capturing quaternion and IMU data from an industrial robotic arm over 26 hours, including various anomaly scenarios, and made the source code available on GitHub for replicability.
Hakan Kayan, Ryan Heartfield, Omer F. Rana, Pete Burnap, Charith Perera
IEEE Internet Things J.3
2025 PrivacyCube: Data Physicalization for Enhancing Privacy Awareness in IoT
abstract
People are increasingly bringing Internet of Things (IoT) devices into their homes without understanding how their data is gathered, processed, and used. We describe PrivacyCube, a novel data physicalization designed to increase privacy awareness within smart home environments. PrivacyCube visualizes IoT data consumption by displaying privacy-related notices. PrivacyCube aims at assisting smart home occupants to (i) understand their data privacy better and (ii) have conversations around data management practices of IoT devices used within their homes. Using PrivacyCube, households can learn and make informed privacy decisions collectively. To evaluate PrivacyCube, we used multiple research methods throughout the different stages of design. We first conducted a focus group study in two stages with six participants to compare PrivacyCube to text and state-of-the-art privacy policies. We then deployed PrivacyCube in a 14-day-long in-home field study with eight households. Lastly, we conducted an event-based field study comparing PrivacyCube with a mobile application, engaging 26 participants with diverse demographics. Our results show that PrivacyCube helps home occupants comprehend IoT privacy better with significantly increased privacy awareness at p < .05 (p = 0.00041, t = -5.57). Participants preferred PrivacyCube over text privacy policies because it was comprehensive and easier to use. PrivacyCube, Privacy Label, and the mobile application, all received positive reviews from participants, with PrivacyCube being preferred for its interactivity and ability to encourage conversations. PrivacyCube was also considered by home occupants as a piece of home furniture , encouraging them to socialize and discuss IoT privacy implications using this device. Watch the demo ( Demo Video ) ( Source Code ).
Bayan Al Muhander, Nalin Arachchilage, Yasar Majib, Mohammed Alosaimi, Omer F. Rana, Charith Perera
ACM Trans. Internet Things5
2025 Guest Editorial:Special Section on SC22 Student Cluster Competition
abstract
Since 2015, as part of a Reproducibility Initiative, the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) has encouraged, and more recently required, all technical papers to include an Artifact Description (AD) Appendix. Each year, the conference selects one accepted paper to form the basis of a Reproducibility Challenge in the Student Cluster Competition (SCC) - a mainstay in the conference program - for the following year. In the SCC, teams of undergraduate students partner with their educational institution and vendors to build, benchmark and operate a small HPC system during the conference week. The Reproducibility Challenge component has them attempt to reproduce parts of the selected work on their own cluster. At SC22, teams worked to reproduce the SC21 paper “Productivity, Portability, Performance: Data-Centric Python” by Ziogas et al. This special section includes a selection of reports from student teams, along with an extended version of that paper, discussing the students’ results...
Omer F. Rana, Josef Spillner, Stephen Leak, Jay F. Lofstead, Rafael Tolosana-Calasanz
IEEE Trans. Parallel Distributed Syst.1
2025 BackFillMe: An Energy and Performance Efficient Virtual Machine Scheduler for IaaS Datacenters
abstract
Backfilling refers to the practice of allowing small jobs to be completed ahead of schedule as long as they do not cause the first job in the line to wait. Users are expected to offer estimates of how long jobs will take to complete in order to make these decisions possible, and these projections are often based on historical data. However, predictions are very hard and may not be accurate, particularly in cloud computing scenarios where jobs or applications run on Virtual Machines (VMs). In addition, scheduling and consolidation techniques can improve the energy efficiency and performance of applications. Consolidation involves VM migrations that can have a negative impact on workload performance and users’ costs. Backfilling can be used as an alternative technique for consolidation (short-term) and/or can be used along with consolidation (long-term). Backfilling methods are well-utilised in single computing systems, but are relatively unexplored in cloud resource allocation. A backfilling-based resource allocation and consolidation technique is proposed. Using real workloads from the Google cluster traces, we investigate the impact of backfilling on infrastructure energy efficiency and performance. For 12583 heterogeneous servers and approximately three million jobs that belong to three different applications, we observed that approximately 19% energy savings and 6% workload performance improvements are achievable using the backfilling approach. Furthermore, our evaluation suggests that using VM runtime as a criterion for the backfilling approach is approximately 3.56%–7.78% more energy and 1.91%–3.38% more performance efficient than using priority as a backfilling criterion.
Muhammad Zakarya, Lee Gillam, Mohammad Reza Chalak Qazani, Ayaz Ali Khan, Khaled Salah 0001, Omer F. Rana
IEEE Trans. Serv. Comput.6
2024 Verifiable Querying Framework for Multi-Blockchain Applications
abstract
Effectively and securely retrieving data from various blockchain networks remains a critical challenge. We propose a novel framework that provides enhanced capabilities for authenticated data retrieval, empowering users and regulatory bodies. Our framework provides a scalable solution for information verification across diverse blockchain systems, enabling a variety of different types of actors to be integrated. Our framework allows metadata from various systems to be combined, also providing support for secure querying.
Stanly Wilson, Kwabena Adu-Duodu, Yinhao Li 0003, Ringo W. H. Sham, Ellis Solaiman, Omer F. Rana, Rajiv Ranjan 0001
ICBC6
2024 Towards Enhancing Linked Data Retrieval in Conversational UIs Using Large Language Models
Omar Mussa, Omer F. Rana, Benoit Goossens, Pablo Orozco ter Wengel, Charith Perera
WISE (4)2
2024 Talking Buildings: Interactive Human-Building Smart-Bot for Smart Buildings
Devmane Suhas, Omer F. Rana, Simon Lannon, Charith Perera
WISE (1)2
2024 General data protection regulation: a study on attitude and emotional empowerment
abstract
Over the last few years, digitalisation has accelerated its pace, fuelling the creation of a massive amount of data.This has resulted in a need to introduce legal mechanisms to protect the privacy and security of data being exchanged between people and organisations.However, little is known about the individuals' perspective on such mechanisms.Given the gap in the literature, this research investigated the drivers and the implications of individuals' attitude towards GDPR compliance.To test the research model, structural equational modelling was employed using 540 responses.The result showed that perceived threat severity, self-efficacy and response efficacy determine a positive attitude towards GDPR compliance, which results in emotional empowerment.The findings contribute to the literature on legal privacy-preserving mechanisms, by providing a user's view on the coping and threat appraisal factors underpinning attitude and demonstrating the implications for driving confidence in control over personal data.The findings also contribute to the literature on protection motivation by demonstrating that attitude towards adaptive behaviour drives emotional empowerment.The study offers suggestions to policymakers on how to enhance public perception of the GDPR.The findings also provide guidelines for organisations on how to inform individuals' understanding of compliance with the legal framework.
Davit Marikyan, Savvas Papagiannidis, Omer F. Rana, Rajiv Ranjan 0001
Behav. Inf. Technol.3
2024 Learning-driven service caching in MEC networks with bursty data traffic and uncertain delays
Wenhao Ren, Zichuan Xu, Weifa Liang, Haipeng Dai 0001, Omer F. Rana, Pan Zhou 0001, Qiufen Xia, Haozhe Ren, Mingchu Li, Guowei Wu 0001
Comput. Networks5
2024 SHIELD: A Secure Heuristic Integrated Environment for Load Distribution in Rural-AI
Ashish Kaushal, Osama Almurshed, Osama Almoghamis, Areej Alabbas, Nitin Auluck, Bharadwaj Veeravalli, Omer F. Rana
Future Gener. Comput. Syst.7
2024 Serverless computing in the cloud-to-edge continuum
Carlo Puliafito, Omer F. Rana, Luiz Fernando Bittencourt, Hao Wu 0022
Future Gener. Comput. Syst.2
2024 Orchestration and Management of Adaptive IoT-Centric Distributed Applications
abstract
Current Internet of Things (IoT) devices provide a diverse range of functionalities, ranging from measurement and dissemination of sensory data observation, to computation services for real-time data stream processing. In extreme situations such as emergencies, a significant benefit of IoT devices is that they can help gain a more complete situational understanding of the environment. However, this requires the ability to utilize IoT resources while taking into account location, battery life, and other constraints of the underlying edge and IoT devices. A dynamic approach is proposed for orchestration and management of distributed workflow applications using services available in cloud data centers, deployed on servers, or IoT devices at the network edge. Our proposed approach is specifically designed for knowledge-driven business process workflows that are adaptive, interactive, evolvable and emergent. A comprehensive empirical evaluation shows that the proposed approach is effective and resilient to situational changes.
Sehrish Amjad, Ahmed Akhtar, Ayesha Afzal, Basit Shafiq, Jaideep Vaidya, Shafay Shamail, Omer F. Rana
IEEE Internet Things J.8
2024 ApMove: A Service Migration Technique for Connected and Autonomous Vehicles
abstract
Multi-access edge computing systems (MECs) bring the capabilities of cloud computing closer to the radio access network (RAN), in the context of 4G and 5G telecommunication systems, and converge with existing radio access technologies like satellite or WiFi. An MEC is a cloud server that runs at the mobile network’s edge and is installed and executed using virtual machines (VMs), containers, and/or functions. A cloudlet is similar to an MEC that consists of many servers which provide real-time, low-latency, computing services to connected users in close proximity. In connected vehicles, services may be provisioned from the cloud or edge that will be running users’ applications. As a result, when users travel across many MECs, it will be necessary to transfer their applications in a transparent manner so that performance and connectivity are not negatively affected. In this paper, we propose an effective strategy for migrating connected users’ services from one edge to another or, more likely, to a remote cloud in an MEC. A mathematical model is presented to estimate the expected times to allocate and migrate services. Our evaluations, based on real workload traces and mobility patterns, suggest that the proposed strategy “ApMove" migrates connected services while ensuring their performance ( 0.004% – 2.99% loss), reduced runtimes, therefore, users’ costs ( 4.3% – 11.63%), and minimizing the response time ( 7.45% – 9.04%). Furthermore, approximately 17.39% migrations are avoided. We also study the impacts of variations in the car’s speed and network transfer rates on service migration durations, latencies, and service execution times.
Muhammad Zakarya, Lee Gillam, Ayaz Ali Khan, Omer F. Rana, Rajkumar Buyya
IEEE Internet Things J.4
2024 Age-Aware Data Selection and Aggregator Placement for Timely Federated Continual Learning in Mobile Edge Computing
abstract
Federated continual learning (FCL) is emerging as a key technology for time-sensitive applications in highly adaptive environments including autonomous driving and industrial digital twin. Each FCL trains machine learning models using newly-generated datasets as soon as possible, to obtain a highly accurate machine learning model for new event predictions. Theage of data, defined as the time difference between the generation time of a dataset and the current time, is widely adopted as a key criterion to evaluate both timeline and quality of training. In this paper, we study the problem of age-aware FCL in a mobile edge computing (MEC) network. We not only investigate optimization techniques that optimize the data selection and aggregator placement for FCL but also implement a real system as a prototype for age-aware FCL. Specifically, we first propose an approximation algorithm with a provable approximation ratio for the age-aware data selection and aggregator placement problem for FCL with a single request. In real application scenarios, there are usually multiple FCL requests that require to train models, and delays in the MEC network are usually uncertain. We then study the problem of age-aware data selection and aggregator placement problem for FCL with uncertain delays and multiple requests, by devising an online learning algorithm with a bounded regret based on contextual bandits. We finally implement a prototype for FCL in an MEC network, with various heterogeneous user equipments (UEs) and cloudlets with different computing capabilities in the network. Experiment results show that the performance of the proposed algorithms outperform existing studies, by achieving 47% lower age of data and 12% higher model accuracy.
Zichuan Xu, Lin Wang 0093, Weifa Liang, Qiufen Xia, Wenzheng Xu, Pan Zhou 0001, Omer F. Rana
IEEE Trans. Computers7
2024 Blockchain Based Auditable Access Control for Business Processes With Event Driven Policies
abstract
The use of blockchain technology has been proposed to provide auditable access control for individual resources. Unlike the case where all resources are owned by a single organization, this work focuses on distributed applications such as business processes and distributed workflows. These applications are often composed of multiple resources/services that are subject to the security and access control policies of different organizational domains. Here, blockchains provide an attractive decentralized solution to provide auditability. However, the underlying access control policies may have event-driven constraints and can be overlapping in terms of the component conditions/rules as well as events. Existing work cannot handle event-driven constraints and does not sufficiently account for overlaps leading to significant overhead in terms of cost and computation time for evaluating authorizations over the blockchain. In this work, we propose an automata-theoretic approach for generating a cost-efficient composite access control policy. We reduce this composite policy generation problem to the standard weighted set cover problem. We show that the composite policy correctly captures all the local access control policies and reduces the policy evaluation cost over the blockchain. We have implemented the initial prototype of our approach using Ethereum as the underlying blockchain and empirically validated the effectiveness and efficiency of our approach. Ablation studies were conducted to determine the impact of changes in individual service policies on the overall cost.
Ahmed Akhtar, Masoud Barati, Basit Shafiq, Omer F. Rana, Ayesha Afzal, Jaideep Vaidya, Shafay Shamail
IEEE Trans. Dependable Secur. Comput.4
2024 Designing Privacy-Aware IoT Applications for Unregulated Domains
abstract
Internet of Things (IoT) applications (apps) are challenging to design because of the heterogeneous systems on which they are deployed. IoT devices and apps may collect and analyse sensitive personal data, which is often protected by data privacy laws, some within highly regulated domains such as healthcare. Privacy-by-design (PbD) schemes can be used by developers to consider data privacy at the design stage. However, software developers are not widely adopting these approaches due to difficulties in understanding and interpreting them. There are currently a limited number of tools available for developers to use in this context. We believe that a successful PbD tool should be able to (i) assist developers in addressing privacy requirements in less regulated domains, as well as (ii) help them learn about privacy as they use the tool. The findings of two controlled lab studies are presented, involving 42 developers. We discuss how such a PbD tool can help novice IoT developers comply with privacy laws (e.g., GDPR) and follow privacy guidelines (e.g., privacy patterns). Based on our findings, such tools can help raise awareness of data privacy requirements at design. This increases the likelihood that subsequent designs will be more aware of data privacy requirements.
Nada Alhirabi, Stephanie Beaumont, Omer F. Rana, Charith Perera
ACM Trans. Internet Things3
2024 CASPER: Context-Aware IoT Anomaly Detection System for Industrial Robotic Arms
abstract
Industrial cyber-physical systems (ICPS) are widely employed in supervising and controlling critical infrastructures, with manufacturing systems that incorporate industrial robotic arms being a prominent example. The increasing adoption of ubiquitous computing technologies in these systems has led to benefits such as real-time monitoring, reduced maintenance costs, and high interconnectivity. This adoption has also brought cybersecurity vulnerabilities exploited by adversaries disrupting manufacturing processes via manipulating actuator behaviors. Previous incidents in the industrial cyber domain prove that adversaries launch sophisticated attacks rendering network-based anomaly detection mechanisms insufficient as the “physics” involved in the process is overlooked. To address this issue, we propose an IoT-based cyber-physical anomaly detection system that can detect motion-based behavioral changes in an industrial robotic arm. We apply both statistical and state-of-the-art machine learning methods to real-time Inertial Measurement Unit data collected from an edge development board attached to an arm doing a pick-and-place operation. To generate anomalies, we modify the joint velocity of the arm. Our goal is to create an air-gapped secondary protection layer to detect “physical” anomalies without depending on the integrity of network data, thus augmenting overall anomaly detection capability. Our empirical results show that the proposed system, which utilizes 1D convolutional neural networks, can successfully detect motion-based anomalies on a real-world industrial robotic arm. The significance of our work lies in its contribution to developing a comprehensive solution for ICPS security, which goes beyond conventional network-based methods.
Hakan Kayan, Ryan Heartfield, Omer F. Rana, Pete Burnap, Charith Perera
ACM Trans. Internet Things3
2024 A scalable and flexible platform for service placement in multi-fog and multi-cloud environments
Sadoon Azizi, Pedram Farzin, Mohammad Shojafar, Omer F. Rana
J. Supercomput.4
2024 Energy or Accuracy? Near-Optimal User Selection and Aggregator Placement for Federated Learning in MEC
abstract
To unveil the hidden value in the datasets of user equipments (UEs) while preserving user privacy, federated learning (FL) is emerging as a promising technique to train a machine learning model using the datasets of UEs locally without uploading the datasets to a central location. Customers require to train machine learning models based on different datasets of UEs, through issuing FL requests that are implemented by FL services in a mobile edge computing (MEC) network. A key challenge of enabling FL in MEC networks is how to minimize the energy consumption of implementing FL requests while guaranteeing the accuracy of machine learning models, given that the availabilities of UEs usually are uncertain. In this paper, we investigate the problem of energy minimization for FL in an MEC network with uncertain availabilities of UEs. We first consider the energy minimization problem for a single FL request in an MEC network. We then propose a novel optimization framework for the problem with a single FL request, which consists of (1) an online learning algorithm with a bounded regret for the UE selection, by considering various contexts (side information) that influence energy consumption; and (2) an approximation algorithm with an approximation ratio for the aggregator placement for a single FL request. We third deal with the problem with multiple FL requests, for which we devise an online learning algorithm with a bounded regret. We finally evaluate the performance of the proposed algorithms by extensive experiments. Experimental results show that the proposed algorithms outperform their counterparts by reducing at least 13% of the total energy consumption while achieving the same accuracy.
Zichuan Xu, Dongrui Li, Weifa Liang, Wenzheng Xu, Qiufen Xia, Pan Zhou 0001, Omer F. Rana, Hao Li 0080
IEEE Trans. Mob. Comput.7
2024 Flow-Time Minimization for Timely Data Stream Processing in UAV-Aided Mobile Edge Computing
abstract
Unmanned Aerial Vehicles (UAVs) have gained increasing attention by both academic and industrial communities, due to their flexible deployment and efficient line-of-sight communication. Recently, UAVs equipped with base stations have been envisioned as a key technology to provide 5G network services for mobile users. In this article, we provide timely services on the data streams of mobile users in a UAV-aided Mobile Edge Computing (MEC) network, in which each UAV is equipped with a 5G small-cell base station for communication and data processing. Specifically, we first formulate a flow-time minimization problem by jointly caching services and offloading tasks of mobile users to the UAV-aided MEC with the aim to minimize the flow time, where the flow time of a user request is referred to the time duration from the request issuing time point to its completion point, subject to resource and energy capacity on each UAV. We then propose a spatial-temporal learning optimization framework. We also devise an online algorithm with a competitive ratio for the problem based upon the framework, by leveraging the round-robin scheduling and dual fitting techniques. Finally, we evaluate the performance of the proposed algorithms through experimental simulation. The simulation results demonstrate that the proposed algorithms outperform their comparison counterparts, by reducing the flow time no less than 19% on average.
Zichuan Xu, Haiyang Qiao, Weifa Liang, Qiufen Xia, Pan Zhou 0001, Omer F. Rana, Wenzheng Xu
ACM Trans. Sens. Networks7
2024 GeoDeploy: Geo-Distributed Application Deployment Using Benchmarking
abstract
Geo-distributed web-applications (GWA) can be deployed across multiple geographically separated datacenters to reduce the latency of access for users. Finding a suitable deployment for a GWA is challenging due to the requirement to consider a number of different parameters, such as host configurations across a federated infrastructure. The ability to evaluate multiple deployment configurations enables an efficient outcome to be determined, balancing resource usage while satisfying user requirements. We proposeGeoDeploy, a framework designed for finding a deployment solution for GWA. We evaluateGeoDeployusing both a formal algorithmic model and a practical cloud-based deployment. We also compare our approach with other existing techniques.
Devki Nandan Jha, Yinhao Li 0003, Zhenyu Wen, Graham Morgan, Prem Prakash Jayaraman, Maciej Koutny, Omer F. Rana, Rajiv Ranjan 0001
IEEE Trans. Parallel Distributed Syst.7
2024 Editorial
abstract
Carbon neutrality is a growing important objective for human activities, to prevent climate change. As for computer systems, we are urged to provide sustainability in computing to help mitigate such problems. As such, carbon neutrality shall occupy a critical role in the next-generation digital infrastructure. To this end, we have collected 15 established works towards carbon neutrality in computer systems in the Special Issue on Carbon-Neutral Computing for Next-Generation Digital Infrastructures.
Zichen Xu 0001, Shaolei Ren, Omer F. Rana
IEEE Trans. Sustain. Comput.3
2023 Performance Analysis of Apache OpenWhisk Across the Edge-Cloud Continuum
abstract
Serverless computing offers opportunities for auto-scaling, a pay-for-use cost model, quicker deployment and faster updates to support computing services. Apache OpenWhisk is one such open-source, distributed serverless platform that can be used to execute user functions in a stateless manner. We conduct a performance analysis of OpenWhisk on an edge-cloud continuum, using a function chain of video analysis applications. We consider a combination of Raspberry Pi and cloud nodes to deploy OpenWhisk, modifying a number of parameters, such as maximum memory limit and runtime, to investigate application behaviours. The five main factors considered are: cold and warm activation, memory and input size, CPU architecture, runtime packages used, and concurrent invocations. The results have been evaluated using initialization, and execution time, minimum memory requirement, inference time and accuracy.
Areej Alabbas, Ashish Kaushal, Osama Almurshed, Omer F. Rana, Nitin Auluck, Charith Perera
CLOUD4
2023 Tracking Material Reuse across Construction Supply Chains
abstract
Material reuse and recycling plays a key role in reducing carbon emissions in the architecture and construction sector. A “Material Passport” (MP) is a record describing how a material is used throughout its lifetime, from genesis to termination, recording operations carried out on the material. The granularity of information recorded in a MP can vary, however ensuring that this provenance trail remains immutable is a key requirement. The benefits of using a MP, operations carried out on a MP, and recording of transactions within a distributed Blockchain (parachain) is described. A scenario is used to illustrate how the proposed approach can be used in practice.
Stanly Wilson, Kwabena Adu-Duodu, Yinhao Li 0003, Ringo W. H. Sham, Ellis Solaiman, Charith Perera, Rajiv Ranjan 0001, Omer F. Rana
e-Science9
2023 Using the blockchain to enable transparent and auditable processing of personal data in cloud- based services: Lessons from the Privacy-Aware Cloud Ecosystems (PACE) project
abstract
The architecture of cloud-based services is typically opaque and intricate. As a result, data subjects cannot exercise adequate control over their personal data, and overwhelmed data protection authorities must spend their limited resources in costly forensic efforts to ascertain instances of non-compliance. To address these data protection challenges, a group of computer scientists and socio-legal scholars joined forces in the Privacy-Aware Cloud Ecosystems (PACE) project to design a blockchain-based privacy-enhancing technology (PET). This article presents the fruits of this collaboration, highlighting the capabilities and limits of our PET, as well as the challenges we encountered during our interdisciplinary endeavour. In particular, we explore the barriers to interdisciplinary collaboration between law and computer science that we faced, and how these two fields’ different expectations as to what technology can do for data protection law compliance had an impact on the project's development and outcome. We also explore the overstated promises of techno-regulation, and the practical and legal challenges that militate against the implementation of our PET: most industry players have no incentive to deploy it, the transaction costs of running it make it prohibitively expensive, and there are significant clashes between the blockchain's decentralised architecture and GDPR's requirements that hinder its deployability. We share the insights and lessons we learned from our efforts to overcome these challenges, hoping to inform other interdisciplinary projects that are increasingly important to shape a data ecosystem that promotes the protection of our personal data.
Jose Tomas Llanos, Madeline Carr, Omer F. Rana
Comput. Law Secur. Rev.3
2023 Detection and mitigation of field flooding attacks on oil and gas critical infrastructure communication
abstract
Industrial Cyber-Physical Systems (ICPS) are highly dependent on Supervisory Control and Data Acquisition (SCADA) for process monitoring and control. Such SCADA systems are known to communicate using various insecure protocols such as Modbus, DNP3, and Open Platform Communication (OPC) Data Access standards (providing access to real-time automation data), which are vulnerable to a range of attacks. This leads to increased cyber risks faced in critical infrastructures, especially in the Oil and Gas sector. One of the most popular and critical attacks deployed against such infrastructure is Denial of Service (DoS), as it can have severe consequences that range from financial loss to loss of life. Such attacks can disrupt the ability of an operator to control hazardous operations leading to potentially unsafe scenarios. A novel Field Flooding attack is described which takes advantage of the packet memory structure of the Modbus protocol to perform a DoS attack. This attack can cause overflowing of the memory bank allocated in the Programmable Logic Controller (PLC) for Modbus operations. The attack is deployed and evaluated on a real industrial testbed and its impact against the Mitre ATT&CK framework is assessed, in order to identify which tactics an adversary could use to compromise the system. A novel mechanism that utilises supervised machine learning to detect this attack in industrial control system networks is also described. Experimental results show that the proposed mechanism, using the XGBoost algorithm, can identify this attack with 99% accuracy.
Abubakar Sadiq Mohammed, Eirini Anthi, Omer F. Rana, Neetesh Saxena, Pete Burnap
Comput. Secur.3
2023 Predicting machine behavior from Google cluster workload traces
abstract
Summary Data centers today host a number of computational resources to support the increasing demand for computation and storage. Understanding how these physical and virtual machines transition between different states of operation (referred to as machine lifecycle) enables more efficient data center operation management. Furthermore, it helps data center operators define policies on how new computational resources can be added or existing infrastructure decommissioned. Using Google cluster trace data set version 3 collected from approximately 96 k machines, we analyze machine failure and changes in machine lifecycle over time. We observed that there is a 13% chance of another machine failure under the same network switch within 1 min of the previous machine failure. A Markov chain‐based model is proposed, that can predict machine states at any given time. Using the model and estimated probabilities, we predicted the machine state over a span of several days with a high probability. Using the predicted machine state, we reconstructed the active machines trend and compared this with the trend reported in the data set, observing an error of 1.76%.
Adnan Umer, Adnan Noor Mian, Omer F. Rana
Concurr. Comput. Pract. Exp.3
2023 Compliance Checking of Cloud Providers: Design and Implementation
abstract
The recognition of capabilities supplied by cloud systems is presently growing. Collecting or sharing healthcare data and sensitive information especially during the Covid-19 pandemic has motivated organizations and enterprises to leverage the upsides coming from cloud-based applications. However, the privacy of electronic data in such applications remains a significant challenge for cloud vendors to adapt their solutions with existing privacy legislation standards such as general data protection regulation (GDPR). This article first proposes a formal model and verification for data usage requests of providers in a cloud composite service using a model checking tool. A cloud pharmacy scenario is presented to illustrate the connectivity of providers in the composite service and the stream of their requests for both collection and movement of patient data. A set of verifications is then undertaken over the pharmacy service in accordance with three significant GDPR obligations, namely user consent, data access, and data transfer. Following that, the article designs and implements a cloud container virtualization based on the verified formal model realizing GDPR requirements. The container makes use of some enforcement smart contracts to only proceed with the providers’ requests that are compliant with GDPR. Finally, several experiments are provided to investigate the performance of our approach in terms of time, memory, and cost.
Masoud Barati, Kwabena Adu-Duodu, Omer F. Rana, Gagangeet Singh Aujla, Rajiv Ranjan 0001
Distributed Ledger Technol. Res. Pract.3
2023 Semantics-based privacy by design for Internet of Things applications
abstract
As Internet of Things (IoT) technologies become more widespread in everyday life , privacy issues are becoming more prominent. The aim of this research is to develop a personal assistant that can answer software engineers’ questions about Privacy by Design (PbD) practices during the design phase of IoT system development. Semantic web technologies are used to model the knowledge underlying PbD measurements, their intersections with privacy patterns, IoT system requirements and the privacy patterns that should be applied across IoT systems. This is achieved through the development of the PARROT ontology, developed through a set of representative IoT use cases relevant for software developers. This was supported by gathering Competency Questions (CQs) through a series of workshops, resulting in 81 curated CQs. These CQs were then recorded as SPARQL queries, and the developed ontology was evaluated using the Common Pitfalls model with the help of the Protégé HermiT Reasoner and the Ontology Pitfall Scanner (OOPS!), as well as evaluation by external experts. The ontology was assessed within a user study that identified that the PARROT ontology can answer up to 58% of privacy-related questions from software engineers.
Lamya Alkhariji, Suparna De, Omer F. Rana, Charith Perera
Future Gener. Comput. Syst.3
2023 BlockFaaS: Blockchain-enabled Serverless Computing Framework for AI-driven IoT Healthcare Applications
Muhammed Golec, Sukhpal Singh, Mustafa Golec, Minxian Xu, Soumya K. Ghosh 0001, Salil S. Kanhere, Omer F. Rana, Steve Uhlig
J. Grid Comput.7
2023 Query Interface for Smart City Internet of Things Data Marketplaces: A Case Study
abstract
Cities are increasingly becoming augmented with sensors through public, private, and academic sector initiatives. Most of the time, these sensors are deployed with a primary purpose (objective) in mind (e.g., deploy sensors to understand noise pollution) by a sensor owner (i.e., the organization that invests in sensing hardware, e.g., a city council). Over the past few years, communities undertaking smart city development projects have understood the importance of making the sensor data available to a wider community—beyond their primary usage. Different business models have been proposed to achieve this, including creating data marketplaces. The vision is to encourage new startups and small and medium-scale businesses to create novel products and services using sensor data to generate additional economic value. Currently, data are sold as pre-defined independent datasets (e.g., noise level and parking status data may be sold separately). This approach creates several challenges, such as (i) difficulties in pricing, which leads to higher prices (per dataset); (ii) higher network communication and bandwidth requirements; and (iii) information overload for data consumers (i.e., those who purchase data). We investigate the benefit of semantic representation and its reasoning capabilities toward creating a business model that offers data on demand within smart city Internet of Things data marketplaces. The objective is to help data consumers (i.e., small and medium enterprises) acquire the most relevant data they need. We demonstrate the utility of our approach by integrating it into a real-world IoT data marketplace (developed by the synchronicity-iot.eu project). We discuss design decisions and their consequences (i.e., tradeoffs) on the choice and selection of datasets. Subsequently, we present a series of data modeling principles and recommendations for implementing IoT data marketplaces.
Naeima Hamed, Andrea Gaglione, Alex Gluhak, Omer F. Rana, Charith Perera
ACM Trans. Internet Things4
2023 Interactive Privacy Management: Toward Enhancing Privacy Awareness and Control in the Internet of Things
abstract
The balance between protecting user privacy while providing cost-effective devices that are functional and usable is a key challenge in the burgeoning Internet of Things (IoT). In traditional desktop and mobile contexts, the primary user interface is a screen; however, in IoT devices, screens are rare or very small, invalidating many existing approaches to protecting user privacy. Privacy visualizations are a common approach for assisting users in understanding the privacy implications of web and mobile services. To gain a thorough understanding of IoT privacy, we examine existing web, mobile, and IoT visualization approaches. Following that, we define five major privacy factors in the IoT context: type, usage, storage, retention period, and access. We then describe notification methods used in various contexts as reported in the literature. We aim to highlight key approaches that developers and researchers can use for creating effective IoT privacy notices that improve user privacy management (awareness and control). Using a toolkit, a use case scenario, and two examples from the literature, we demonstrate how privacy visualization approaches can be supported in practice.
Bayan Al Muhander, Jason Wiese, Omer F. Rana, Charith Perera
ACM Trans. Internet Things3
2023 A Novel Image-Based Homomorphic Approach for Preserving the Privacy of Autonomous Vehicles Connected to the Cloud
abstract
Autonomous vehicles are taking a leap forward by performing operations without human intervention through continuous monitoring of their surroundings using multiple sensors. Images gathered through vehicle mounted cameras can be large, requiring specialized storage such as cloud. However, cloud data centres can be prone to security and privacy challenges. A partial image-based, homomorphic searchable encryption scheme is proposed, which uses pixel-level encryption to identify objects within encrypted images. The scheme provides Object-Trapdoor and Trapdoor-Image indistinguishability – as the trapdoors are probabilistic. The proposed scheme is deployed on a cloud data centre and tested over a real data set. The proposed scheme reduces storage overhead by approximately 20 times, and is 33 times more efficient compared to the generic Paillier homomorphic searchable encryption scheme. Security analysis demonstrates that the scheme maintains high levels of security and privacy.
Aiman Sultan, Shahzaib Tahir, Hasan Tahir, Tayyaba Anwer, Fawad Khan, Muttukrishnan Rajarajan, Omer F. Rana
IEEE Trans. Intell. Transp. Syst.7
2023 HierFedML: Aggregator Placement and UE Assignment for Hierarchical Federated Learning in Mobile Edge Computing
abstract
Federated learning (FL) is a distributed machine learning technique that enables model development on user equipments (UEs) locally, without violating their data privacy requirements. Conventional FL adopts a single parameter server to aggregate local models from UEs, and can suffer from efficiency and reliability issues – especially when multiple users issue concurrentFL requests. Hierarchical FL consisting of a master aggregator and multiple worker aggregators to collectively combine trained local models from UEs is emerging as a solution to efficient and reliable FL. The placement of worker aggregators and assignment of UEs to worker aggregators plays a vital role in minimizing the cost of implementing FL requests in a Mobile Edge Computing (MEC) network. Cost minimization associated with joint worker aggregator placement and UE assignment problem in an MEC network is investigated in this work. An optimization framework for FL and an approximation algorithm with an approximation ratio for a single FL request is proposed. Online worker aggregator placements and UE assignments for dynamic FL request admissions with uncertain neural network models, where FL requests arrive one by one without the knowledge of future arrivals, is also investigated by proposing an online learning algorithm with a bounded regret. The performance of the proposed algorithms is evaluated using both simulations and experiments in a real testbed with its hardware consisting of server edge servers and devices and software built upon an open source hierarchical FedML (HierFedML) environment. Simulation results show that the performance of the proposed algorithms outperform their benchmark counterparts, by reducing the implementation cost by at least 15% per FL request. Experimental results in the testbed demonstrate the performance gain using the proposed algorithms using real datasets for image identification and text recognition applications.
Zichuan Xu, Dapeng Zhao, Weifa Liang, Omer F. Rana, Pan Zhou 0001, Mingchu Li, Wenzheng Xu, Hao Li 0080, Qiufen Xia
IEEE Trans. Parallel Distributed Syst.4
2023 Real-Time Scheduling on Hierarchical Heterogeneous Fog Networks
abstract
Cloud computing is widely used to support offloaded data processing for various applications. However, latency constrained data processing has requirements that may not always be suitable for cloud-based processing. Fog computing brings processing closer to data generation sources, by reducing propagation and data transfer delays. It is a viable alternative for processing tasks with real-time requirements. We propose a scheduling algorithm$RTH^{2}S$(RealTimeHeterogeneousHierarchicalScheduling) for a set of real-time tasks on a heterogeneous integrated fog-cloud architecture. We consider a hierarchical model for fog nodes, with nodes at higher tiers having greater computational capacity than nodes at lower tiers, though with greater latency from data generation sources. Tasks with various profiles have been considered. For the regular profile jobs, we use least laxity first (LLF) to find the preferred fog node for scheduling. In case of “tagged” profiles, based on their tag values, the jobs are split in order to finish execution before the deadline, or the LLF heuristic is used. Using HPC2N workload traces across 3.5 years of activity, the real-time performance of$RTH^{2}S$versus comparable algorithms is demonstrated. We also consider Microsoft Azure-based costs for the proposed algorithm. Our proposed approach is validated using both simulation (to demonstrate scale up) as well as a lab-based testbed.
Amanjot Kaur, Nitin Auluck, Omer F. Rana
IEEE Trans. Serv. Comput.3
2023 CoLocateMe: Aggregation-Based, Energy, Performance and Cost Aware VM Placement and Consolidation in Heterogeneous IaaS Clouds
abstract
In many production clouds, with the notable exception of Google, aggregation-based VM placement policies are used to provision datacenter resources energy and performance efficiently. However, if VMs with similar workloads are placed onto the same machines, they might suffer from contention, particularly, if they are competing for similar resources. High levels of resource contention may degrade VMs performance, and, therefore, could potentially increase users’ costs and infrastructure's energy consumption. Furthermore, segregation-based methods result in stranded resources and, therefore, less economics. The recent industrial interest in segregating workloads opens new directions for research. In this article, we demonstrate how aggregation and segregation-based VM placement policies lead to variabilities in energy efficiency, workload performance, and users’ costs. We, then, propose various approaches to aggregation-based placement and migration. We investigate through a number of experiments, using Microsoft Azure and Google's workload traces for more than twelve thousand hosts and a million VMs, the impact of placement decisions on energy, performance, and costs. Our extensive simulations and empirical evaluation demonstrate that, for certain workloads, aggregation-based allocation and consolidation is$\sim$9.61% more energy and$\sim$20.0% more performance efficient than segregation-based policies. Moreover, various aggregation metrics, such as runtimes and workload types, offer variations in energy consumption and performance, therefore, users’ costs.
Muhammad Zakarya, Lee Gillam, Khaled Salah 0001, Omer F. Rana, Santosh Tirunagari, Rajkumar Buyya
IEEE Trans. Serv. Comput.4
2023 Edge-Enhanced QoS Aware Compression Learning for Sustainable Data Stream Analytics
abstract
Existing Cloud systems involve large volumes of data streams being sent to a centralised data centre for monitoring, storage and analytics. However, migrating all the data to the cloud is often not feasible due to cost, privacy, and performance concerns. However, Machine Learning (ML) algorithms typically require significant computational resources, hence cannot be directly deployed on resource-constrained edge devices for learning and analytics. Edge-enhanced compressive offloading becomes a sustainable solution that allows data to be compressed at the edge and offloaded to the cloud for further analysis, reducing bandwidth consumption and communication latency. The design and implementation of a learning method for discovering compression techniques that offer the best QoS for an application is described. The approach uses a novel modularisation approach that maps features to models and classifies them for a range of Quality of Service (QoS) features. An automated QoS-aware orchestrator has been designed to select the best autoencoder model in real-time for compressive offloading in edge-enhanced clouds based on changing QoS requirements. The orchestrator has been designed to have diagnostic capabilities to search appropriate parameters that give the best compression. A key novelty of this work is harnessing the capabilities of autoencoders for edge-enhanced compressive offloading based on portable encodings, latent space splitting and fine-tuning network weights. Considering how the combination of features lead to different QoS models, the system is capable of processing a large number of user requests in a given time. The proposed hyperparameter search strategy (over the neural architectural space) reduces the computational cost of search through the entire space by up to 89%. When deployed on an edge-enhanced cloud using an Azure IoT testbed, the approach saves up to 70% data transfer costs and takes 32% less time for job completion. It eliminates the additional computational cost of decompression, thereby reducing the processing cost by up to 30%.
Maryleen U. Ndubuaku, Muhammad K. Ali, Ashiq Anjum, Lu Liu 0001, Antonio Liotta, Omer F. Rana
IEEE Trans. Sustain. Comput.6
2022 Circular Economy and Construction Supply Chains
abstract
With increasing complexity of global construction supply chains and the substantial contribution such supply chains make to the environment, there is a need to understand how components of these supply chains can be reused and repurposed. Blockchains provide an important basis for recording transactions carried out in supply chains to aid potential reusability.A recent example of circularity in the built environment includes the tracking of material reuse across the building lifecycle (from construction to demolition and then reuse). The use of decentralisation and immutability in Blockchains is used to demonstrate how material passports may be supported and used as a basis to create a circular economy in construction.
Dan Incorvaja, Yasin Celik, Ioan Petri, Omer F. Rana
BDCAT4
2022 The Lannion report on Big Data and Security Monitoring Research
abstract
During the last decade, big data management has attracted increasing interest from both the industrial and academic communities. In parallel, Cyber Security has become mandatory due to various and more intensive threats. In June 2022, a group of researchers has met to reflect on their community’s impacts on current research challenges. In particular, they have considered four dimensions: (1) dedicated systems being data processing and analytic platforms or time series management systems; (2) graphs analytics and distributed computation; (3) privacy; and (4) new hardware.
Laurent d'Orazio, Jalil Boukhobza, Omer F. Rana, Juba Agoun, Le Gruenwald, Hervé Rannou, Elisa Bertino, Mohand-Said Hacid, Taofik Saïdi, Georges Bossert, Dimitri Tombroff, Makoto Onizuka
IEEE Big Data3
2022 SparkFlow: Towards High-Performance Data Analytics for Spark-based Genome Analysis
abstract
The recent advances in DNA sequencing technology triggered next-generation sequencing (NGS) research in full scale. Big Data (BD) is becoming the main driver in analyzing these large-scale bioinformatics data. However, this complicated process has become the system bottleneck, requiring an amal-gamation of scalable approaches to deliver the needed performance and hide the deployment complexity. Utilizing cutting-edge scientific workflows can robustly address these challenges. This paper presents a Spark-based alignment workflow called SparkFlow for massive NGS analysis over singularity containers. SparkFlow is highly scalable, reproducible, and capable of parallelizing computation by utilizing data-level parallelism and load balancing techniques in HPC and Cloud environments. The proposed workflow capitalizes on benchmarking two state-of-art NGS workflows, i.e., Base Recalibrator and ApplyBQSR. SparkFlow realizes the ability to accelerate large-scale cancer genomic analysis by scaling vertically (HyperThreading) and horizontally (provisions on-demand). Our result demonstrates a trade-off inevitably between the targeted applications and proces-sor architecture. SparkFlow achieves a decisive improvement in NGS computation performance, throughput, and scalability while maintaining deployment complexity. The paper's findings aim to pave the way for a wide range of revolutionary enhancements and future trends within the High-performance Data Analytics (HPDA) genome analysis realm.
Rosa Filgueira, Feras M. Awaysheh, Adam C. Carter, Darren J. White, Omer F. Rana
CCGRID5
2022 Poster: Ontology Enabled Chatbot for Applying Privacy by Design in IoT Systems
abstract
Our aim is to create a personal assistant, a chatbot, that can answer queries from software developers regarding Privacy by Design (PbD) methods and applications throughout the design phase of IoT system development. We used semantic web technologies to model the PARROT Ontology that includes knowledge underlying PbD measurements, their intersections with privacy patterns, IoT system needs, and the privacy patterns that should be applied across IoT systems. To determine the PARROT ontology's requirements, a collection of real-world IoT use cases were aided by a series of workshops to gather Competency Questions (CQs) from researchers and software engineers, resulting in 81 selected CQs. In a user study, the PARROT ontology was able to answer up to 58% of software developers' privacy-related issues. The technical report \citeorca149337 contains further analysis and results from data collecting and intermediate synthesis steps.
Lamya Alkhariji, Suparna De, Omer F. Rana, Charith Perera
CCS3
2022 FLoX: Federated Learning with FaaS at the Edge
abstract
Federated learning (FL) is a technique for distributed machine learning that enables the use of siloed and distributed data. With FL, individual machine learning models are trained separately and then only model parameters (e.g., weights in a neural network) are shared and aggregated to create a global model, allowing data to remain in its original environment. While many applications can benefit from FL, existing frameworks are incomplete, cumbersome, and environment-dependent. To address these issues, we present FLoX, an FL framework built on the funcX federated serverless computing platform. FLoX decouples FL model training/inference from infrastructure management and thus enables users to easily deploy FL models on one or more remote computers with a single line of Python code. We evaluate FLoX using three benchmark datasets deployed on ten heterogeneous and distributed compute endpoints. We show that FLoX incurs minimal overhead, especially with respect to the large communication overheads between endpoints for data transfer. We show how balancing the number of samples and epochs with respect to the capacities of participating endpoints can significantly reduce training time with minimal reduction in accuracy. Finally, we show that global models consistently outperform any single model on average by 8%.
Nikita Kotsehub, Matt Baughman, Ryan Chard, Nathaniel Hudson 0001, Panos Patros, Omer F. Rana, Ian T. Foster, Kyle Chard
e-Science6
2022 Cooperative Offloading Based on Online Auction for Mobile Edge Computing
Syed Bilal Hussain Shah, Liqaa F. Nawaf, Omer F. Rana, Jianyuan Gan
WASA (3)4
2022 Wheels on the Modbus - Attacking ModbusTCP Communications
abstract
Industrial Cyber-Physical Systems (ICPS) make significant use of Supervisory Control and Data Acquisition (SCADA) for control. Such SCADA systems are known to utilise insecure communication protocols such as Modbus, DNP3 and OPC DA. This leads to increased cyber risks faced in critical infrastructures, as these protocols allow threat actors to mount attacks like Denial of Service (DoS). We present a novel field flooding attack, compromising the structure of the ModbusTCP packet and disrupting a controller's interpretation of the commands sent to it. This can disrupt the ability of an operator to control hazardous operations leading to potentially unsafe scenarios.
Abubakar Sadiq Mohammed, Neetesh Saxena, Omer F. Rana
WISEC3
2022 Distributed hierarchical deep optimization for federated learning in mobile edge computing
Syed Bilal Hussain Shah, Ali Kashif Bashir, Raheel Nawaz, Omer F. Rana
Comput. Commun.5
2022 QoS-aware trust establishment for cloud federation
abstract
Abstract Cloud federation enables inter‐layer resource exchanges among multiple, heterogeneous cloud service providers. This article proposes a Quality of Service (QoS) aware trust model for effective resource allocation in response to the various user requests within the Clouds4Coordination (C4C) federation system. This QoS mainly comprises of nine parameters combined into three categories: (i) node profile, (ii) reliability, and (iii) competence. Numerical values for these parameters are computed every ‘t’ seconds for each cloud provider. All values measured over an interval Δt are further processed by the proposed model to evaluate the utility associated with a provider (referred to as a discipline in the presented case study). The decision about interacting with a discipline in a collaborative project is based on this utility value. The systems architecture, evaluation methodology, proposed model, and experimental evaluation on a practical test bed is outlined. The proposed QoS‐aware trust evaluation mechanism allows selection of the most useful (based on a utility value) providers. The proposed approach can be used to support federation of cloud services across a number of different application domains.
Usama Ahmed, Asma Al-Saidi, Ioan Petri, Omer F. Rana
Concurr. Comput. Pract. Exp.4
2022 Greedy Nominator Heuristic: Virtual function placement on fog resources
abstract
Abstract Fog computing is an intermediate infrastructure between edge devices (e.g., Internet of Things) and cloud systems that is used to reduce latency in real‐time applications. An application can be composed of a collection of virtual functions, between which dependency constraints can be captured in a service function chain (SFC). Virtual functions within an SFC can be executed at different geo‐distributed locations. However, virtual functions are prone to failure and often do not complete within a deadline. This results in function reallocation to other nodes within the infrastructure; causing delays, potential data loss during function migration, and increased costs. We proposed Greedy Nominator Heuristic (GNH) to address these issues. GNH is based on redundant deployment and failure tracking of virtual functions. GNH places replicas of each function at multiple locations—taking account of expected completion time, failure risk, and cost. We make use of a MapReduce‐based mechanism, where Mappers find suitable locations in parallel, and a Reducer then ranks these locations. Our results show that GNH reduces latency by up to 68%, and is more cost effective than other approaches which rely on state‐of‐the‐art optimization algorithms to allocate replicas.
Osama Almurshed, Omer F. Rana, Kyle Chard
Concurr. Comput. Pract. Exp.2
2022 Adversarial machine learning in IoT from an insider point of view
abstract
With the rapid progress and significant successes in various applications, machine learning has been considered a crucial component in the Internet of Things ecosystem. However, machine learning models have recently been vulnerable to carefully crafted perturbations, so-called adversarial attacks. A capable insider adversary can subvert the machine learning model at either the training or testing phase, causing them to behave differently. The vulnerability of machine learning to adversarial attacks becomes one of the significant risks. Therefore, there is a need to secure machine learning models enabling the safe adoption in malicious insider cases. This paper reviews and organizes the body of knowledge in adversarial attacks and defense presented in IoT literature from an insider adversary point of view. We proposed a taxonomy of adversarial methods against machine learning models that an insider can exploit. Under the taxonomy, we discuss how these methods can be applied in real-life IoT applications. Finally, we explore defensive methods against adversarial attacks. We believe this can draw a comprehensive overview of the scattered research works to raise awareness of the existing insider threats landscape and encourages others to safeguard machine learning models against insider threats in the IoT ecosystem.
Fatimah Aloraini, Amir Javed, Omer F. Rana, Pete Burnap
J. Inf. Secur. Appl.3
2022 HUNTER: AI based holistic resource management for sustainable cloud computing
Shreshth Tuli, Sukhpal Singh, Minxian Xu, Peter Garraghan, Rami Bahsoon, Schahram Dustdar, Rizos Sakellariou, Omer F. Rana, Rajkumar Buyya, Giuliano Casale, Nicholas R. Jennings
J. Syst. Softw.8
2022 Security analytics for real-time forecasting of cyberattacks
abstract
Summary Protection of networked computing infrastructures (such as Internet of Things, Industrial Control Systems, and Edge computing) is dependent on the continuous monitoring of interaction between such devices and network/Cloud‐based hosts (especially in Industry 4.0 environments). This real‐time monitoring enables an analyst to quantify evolving and emerging threats to such network infrastructures. A framework for identifying patterns in observed cyberthreats and the use of these patterns for forecasting the growth of an emerging threat to network infrastructure is proposed. This framework enables predicting the maximum threat intensity and the time period over which this maximum intensity is likely to occur. The proposed framework integrates: (a) continuous monitoring of device/network activity, (b) forecasting behavior using exponentially weighted moving averages, (c) utilizing Fibonacci retracement for estimating the potential intensity of a cyberattack, and (d) linear regression for predicting response time for high risk thresholds and a machine learning strategy to predict potential risk over a pre‐defined time window. Using this approach, we can produce time intervals between the forecast and the actual attacks using real‐world network activity data. Our results show an average lead time of around 1.75 hours, providing a window of opportunity to limit the impact of an attack and counter it.
Amir Javed, Mike Lakoju, Pete Burnap, Omer F. Rana
Softw. Pract. Exp.4
2022 RES: Real-Time Video Stream Analytics Using Edge Enhanced Clouds
abstract
With increasing availability and use of Internet of Things (IoT) devices such as sensors and video cameras, large amounts of streaming data is now being produced at high velocity. Applications which require low latency response such as video surveillance, augmented reality and autonomous vehicles demand a swift and efficient analysis of this data. Existing approaches employ cloud infrastructure to store and perform machine learning-based analytics on this data. This centralized approach has limited ability to support real-time analysis of large-scale streaming data due to network bandwidth and latency constraints between data source and cloud. We propose RealEdgeStream (RES) an edge enhanced stream analytics system for large-scale, high performance data analytics. The proposed approach investigates the problem of video stream analytics by proposing (i) filtration and (ii) identification phases. The filtration phase reduces the amount of data by filtering low-value stream objects using configurable rules. The identification phase uses deep learning inference to perform analytics on the streams of interest. The phases consist of stages which are mapped onto available in-transit and cloud resources using a placement algorithm to satisfy the Quality of Service (QoS) constraints identified by a user. We demonstrate that for a 10K element data streams, with a frame rate of 15–100 per second, the job completion in the proposed system takes 49 percent less time and saves 99 percent bandwidth compared to a centralized cloud-only based approach.
Muhammad K. Ali, Ashiq Anjum, Omer F. Rana, Ali Reza Zamani, Daniel Balouek-Thomert, Manish Parashar
IEEE Trans. Cloud Comput.3
2022 Cybersecurity Challenges in the Offshore Oil and Gas Industry: An Industrial Cyber-Physical Systems (ICPS) Perspective
abstract
There has been significant interest within the offshore oil and gas industry to utilise Industrial Internet of Things (IIoT) and Industrial Cyber-Physical Systems (ICPS) . There has also been a corresponding increase in cyberattacks targeted at oil and gas companies. Offshore oil production requires remote access to and control of large and complex hardware resources. This is achieved by integrating ICPS, Supervisory, Control and Data Acquisition (SCADA) systems, and IIoT technologies. A successful cyberattack against an oil and gas (O&G) offshore asset could have a major impact on the environment, marine ecosystem and safety of personnel. Any disruption to the world’s supply of O&G can also have an effect on oil prices and the global economy. We describe the cyberattack surface within the oil and gas industry, discussing emerging trends in the offshore sub-sector and provide a historical perspective of known cyberattacks. We also present a case study of a subsea control system architecture typically used in offshore O&G operations and highlight potential vulnerabilities affecting the components of the system. This study is the first to provide a detailed analysis of attack vectors in a subsea control system. The analysis provided can be used to understand key vulnerabilities in such systems and may be used to implement efficient mitigation methods.
Abubakar Sadiq Mohammed, Philipp Reinecke, Pete Burnap, Omer F. Rana, Eirini Anthi
ACM Trans. Cyber Phys. Syst.4
2022 Privacy-Aware Cloud Auditing for GDPR Compliance Verification in Online Healthcare
abstract
Emerging multitenant cloud computing ecosystems allow multiple applications to share virtualized pool of computing and networking resources. As a result, such ecosystems are becoming increasingly prone to data privacy concerns (personal data leakages and unauthorized access). While cloud computing providers support robust security and privacy mechanisms (e.g., public key cryptography, firewalls, and virtual private networks, among many others), they lack mechanisms and frameworks to monitor, audit, and verify these data privacy concerns. The emergence of data protection regulations around the world, such as General Data Protection Regulation in Europe and the Data Protection Act in the U.K., further emphasizes the need to overcome these privacy limitations. In this article, a novel technique for monitoring, auditing, and verifying the operations carried out on a user’s personal data in cloud computing ecosystems is proposed. Our research methodology leverages distributed ledger technologies (e.g., blockchain and smart contracts) for developing an immutable recording technique, which transparently logs, monitors, and verifies the operations carried out on user data. Using a healthcare pharmacy scenario and extensive real-world experiments, we validate the feasibility of the proposed technique. The proposed work handles a large pool of requests ($>$13K) ensuring minimal latency ($\approx$50–60 ms) and overheads for three different service packages varied with respect to the number of actors and operations.
Masoud Barati, Gagangeet Singh Aujla, Jose Tomas Llanos, Kwabena Adu-Duodu, Omer F. Rana, Madeline Carr, Rajiv Ranjan 0001
IEEE Trans. Ind. Informatics5
2022 Aggregated Capability Assessment (AgCA) For CAIQ Enabled Cross-cloud Federation
abstract
Cross-Cloud Federation (CCF) enables resource exchange among multiple, heterogeneous Cloud Service Providers (CSPs) to support the composition of services (workflow) hosted by different providers. CCF participation can either be fixed, or the types of services that can be used are limited to reduce the potential risk of service failure or secure access. Although many service selection approaches have been proposed in literature for cloud computing, their applicability to CCF i.e., cloud-to-cloud interaction, has not been adequately investigated. A key component of this cloud-to-cloud paradigm involves assessing the combined capability of contributing participants within a federation and their connectivity. A novel Aggregated Capability Assessment (AgCA) approach based on using the Consensus Assessment Initiative Questionnaire from Cloud Security Alliance is proposed for CCF. The proposed mechanism is implemented as a component of a centralized broker to enhance the quality of the selection process for participants within a federation. Our experimental results show thatAgCAis a useful tool for partner selection in a dynamic, heterogeneous and multilevel cloud federation.
Usama Ahmed, Imran Raza, Omer F. Rana, Syed Asad Hussain
IEEE Trans. Serv. Comput.3
2022 Tracking GDPR Compliance in Cloud-Based Service Delivery
abstract
The European General Data Protection Regulation (GDPR) has had a far-reaching impact on data privacy and compliance for cloud providers. GDPR influences access to, storage, processing and transmission of personal data, requiring these operations to be verified by a cloud user through explicit consent prior to execution. GDPR rules implemented for such operations can be ambiguous and often open to interpretation, making manual verification a time consuming and error prone process for cloud providers. An encoding of GDPR rules is described, with each operation carried out using these rules recorded into a Blockchain for auditing purposes. Specifically, this work shows how some GDPR rules can appear asopcodesin smart contracts to verify the operations of providers on user data in a transparent and automatic way. An abstract model is designed to demonstrate how cloud providers can access and deploy such smart contracts through a Blockchain-based virtual machine. A case study is used to demonstrate how this approach can be used in practice. The case study uses a collection of design patterns and smart contracts to verify provider operations, includingread,write,executionandtransferon user data. Validation is undertaken by deploying the smart contracts in a Blockchain test network to investigate the execution costs of GDPR compliance checking.
Masoud Barati, Omer F. Rana
IEEE Trans. Serv. Comput.2
2022 epcAware: A Game-Based, Energy, Performance and Cost-Efficient Resource Management Technique for Multi-Access Edge Computing
abstract
Internet of Things (IoT) is producing an extraordinary volume of data daily, and it is possible that the data may become useless while on its way to the cloud, due to long distances. Fog/edge computing is a new model for analysing and acting on time-sensitive data, adjacent to where it is produced. Further, cloud services provided by large companies such as Google, can also be localised to improve response time and service agility. This is accomplished through deploying small-scale datacentres in various locations, where needed in proximity of users; and connected to a centralised cloud that establish a multi-access edge computing (MEC). The MEC setup involves three parties, i.e., service providers (IaaS), application providers (SaaS), network providers (NaaS); which might have different goals, therefore, making resource management difficult. Unlike existing literature, we consider resource management with respect to all parties; and suggest game-theoretic resource management techniques to minimise infrastructure energy consumption and costs while ensuring applications’ performance. Our empirical evaluation, using Google’s workload traces, suggests that our approach could reduce up to 11.95 percent energy consumption, and$\sim$17.86% user costs with negligible loss in performance. Moreover, IaaS can reduce up to 20.27 percent energy bills and NaaS can increase their costs-savings up to 18.52 percent as compared to other methods.
Muhammad Zakarya, Lee Gillam, Hashim Ali 0001, Izaz Ur Rahman, Khaled Salah 0001, Rahim Khan, Omer F. Rana, Rajkumar Buyya
IEEE Trans. Serv. Comput.7
2021 Checking GDPR Compliance for Cloud-based Services
abstract
Accessing a cloud-hosted service may involve executing a number of sub-services which are unknown to the user. A user is only aware of the service they directly invoke, not the sub-services which may be hosted across other cloud providers (including advertising and data processing services). Each service in this chain may collect and process personal user data via read, write and transfer operations. The European General Data Protection Regulation (GDPR) enforces cloud providers to receive explicit consent from their users prior to executing any such operations. We present a Blockchain-based architecture that supports GDPR compliance verification (especially in the context of such a service chain) for enhancing the data privacy of cloud users. The architecture supports a factory of smart contracts, including user consent , GDPR compliance , container and verification , each of which is activated by an actor within a cloud environment. Figure 1 illustrates the interactions between the different components that make up our system – classified into three different phases:
Masoud Barati, Omer F. Rana
SERVICES2
2021 Scheduling Real Tim Security Aware Tasks in Fog Networks
abstract
Fog computing extends the capability of cloud services to support latency sensitive applications. Adding fog computing nodes in proximity to a data generation/ actuation source can support data analysis tasks that have stringent deadline constraints. We introduce a real time, security-aware scheduling algorithm that can execute over a fog environment [1 , 2] . The applications we consider comprise of: (i) interactive applications which are less compute intensive, but require faster response time; (ii) computationally intensive batch applications which can tolerate some delay in execution. From a security perspective, applications are divided into three categories: public, private and semi-private which must be hosted over trusted, semi-trusted and untrusted resources. We propose the architecture and implementation of a distributed orchestrator for fog computing, able to combine task requirements (both performance and security) and resource properties.
Anil Singh, Nitin Auluck, Omer F. Rana, Surya Nepal
SERVICES3
2021 Working in a Smart Home-office: Exploring the Impacts on Productivity and Wellbeing
abstract
Following the outbreak of the Coronavirus (COVID-19) pandemic, many organisations have shifted to remote working overnight. The new reality has created conditions to use smart home technologies for work purposes, for which they were not originally intended. The lack of insights into the new application of smart home technologies has led to two research objectives. First, the paper aimed to investigate the factors correlating with productivity and perceived wellbeing. Second, the study tried to explore individuals’ intentions to use smart home offices for remote work in the future. 528 responses were gathered from individuals who had smart homes and had worked from home during the pandemic. The results showed that productivity positively relates to service relevance, perceived usefulness, perceived ease of use, hedonic beliefs, control over environmental conditions, innovativeness and attitude. Task-technology fit, service relevance, attitude to smart homes, innovativeness, hedonic beliefs, perceived usefulness, perceived ease of use and control over environmental conditions correlate with perceived wellbeing. The intention to work from smart home-offices in the future is determined by perceived wellbeing. Findings contribute to the research on smart homes and remote work practices, by providing the first empirical evidence about the new applications and outcomes of smart home use in the work context.
Davit Marikyan, Savvas Papagiannidis, Rajiv Ranjan 0001, Omer F. Rana
WEBIST4
2021 Privacy-aware cloud ecosystems: Architecture and performance
abstract
Summary With an increasing number of cloud providers offering services made use of by both individual users and other providers, there is a realization that service provision now involves an “ecosystem” of providers. Some providers may be directly visible to a user, while others may be contributors to composite services and not directly known to the user—as only the provider offering the composite service is visible. Such services may include: domain specific services (eg, simulation), advertising services, or profiling/analytics services. Understanding the impact on data privacy of a user for such a composite service remains a challenge, and providing transparency (and obtaining user consent for data use) remains a key requirement of the European General Data Protection Regulation (GDPR). An architecture that makes use of blockchains and smart contracts is proposed that addresses this requirement. An implementation of the architecture is used to demonstrate how access control can be managed and audited. The scalability and cost of undertaking access control, as the number of actors (both service providers and “voters”) increases, is also described. The proposed approach can be used to support service aggregation across both private and public clouds.
Masoud Barati, Omer F. Rana
Concurr. Comput. Pract. Exp.2
2021 Quality of Service-aware matchmaking for adaptive microservice-based applications
abstract
Summary Applications that make use of Internet of Things (IoT) can capture an enormous amount of raw data from sensors and actuators, which is frequently transmitted to cloud data centers for processing and analysis. However, due to varying and unpredictable data generation rates and network latency, this can lead to a performance bottleneck for data processing. With the emergence of fog and edge computing hosted microservices, data processing could be moved towards the network edge. We propose a new method for continuous deployment and adaptation of multi‐tier applications along edge, fog, and cloud tiers by considering resource properties and non‐functional requirements (e.g., operational cost, response time and latency etc.). The proposed approach supports matchmaking of application and Cloud‐To‐Things infrastructure based on a subgraph pattern matching (P‐Match) technique. Results show that the proposed approach improves resource utilization and overall application Quality of Service. The approach can also be integrated into software engineering workbenches for the creation and deployment of cloud‐native applications, enabling partitioning of an application across the multiple infrastructure tiers outlined above.
Polona Stefanic, Petar Kochovski, Omer F. Rana, Vlado Stankovski
Concurr. Comput. Pract. Exp.3
2021 IoTSim-Osmosis: A framework for modeling and simulating IoT applications over an edge-cloud continuum
Khaled Alwasel, Devki Nandan Jha, Fawzy Habeeb, Umit Demirbaga, Omer F. Rana, Thar Baker, Schahram Dustdar, Massimo Villari, Philip James 0002, Ellis Solaiman, Rajiv Ranjan 0001
J. Syst. Archit.5
2021 An Energy and Performance Aware Consolidation Technique for Containerized Datacenters
abstract
Cloud datacenters have become a backbone for today’s business and economy, which are the fastest-growing electricity consumers, globally. Numerous studies suggest that$\sim$30% of the US datacenters are comatose and the others are grossly less-utilized, which make it possible to save energy through resource consolidation techniques. However, consolidation comprises migrations that are expensive in terms of energy consumption and performance degradation, which is mostly not accounted for in many existing models, and, possibly, it could be more energy and performance efficient not to consolidate. In this paper, we investigate how migration decisions should be taken so that the migration cost is recovered, as only when migration cost has been recovered and performance is guaranteed, will energy start to be saved. We demonstrate through several experiments, using the Google workload data for 12,583 hosts and approximately one million tasks that belong to three different kinds of workload, how different allocation policies, combined with various migration approaches, will impact on datacenter’s energy and performance efficiencies. Using several plausible assumptions for containerised datacenter set-up, we suggest, that a combination of the proposed energy-performance-aware allocation (Epc-Fu) and migration (Cper) techniques, and migrating relatively long-running containers only, offers for ideal energy and performance efficiencies.
Ayaz Ali Khan, Muhammad Zakarya, Rajkumar Buyya, Rahim Khan, Mukhtaj Khan, Omer F. Rana
IEEE Trans. Cloud Comput.6
2021 Forecasting peak energy demand for smart buildings
abstract
Abstract Predicting energy consumption in buildings plays an important part in the process of digital transformation of the built environment, and for understanding the potential for energy savings. This also contributes to reducing the impact of climate change, where buildings need to increase their adaptability and resilience while reducing energy consumption and maintain user comfort. The use of Internet of Things devices for monitoring and control of energy consumption in buildings can take into account user preferences, event monitoring and building optimization. Detecting peak energy demand from historical building data can enable users to manage their energy use more efficiently, while also enabling real-time response strategies (including control and actuation) to known or future scenarios. Several statistical, time series, and machine learning techniques are proposed in this work to predict electricity consumption for five different building types, by using peak demand forecasting to achieve energy efficiency. We have used several indigenous and exogenous variables with a view to test different energy forecasting scenarios. The suggested techniques are evaluated for creating predictive models, including linear Regression, dynamic regression, ARIMA time series, exponential smoothing time series, artificial neural network, and deep neural network. We conduct the analysis on an energy consumption dataset of five buildings from 2014 until 2019. Our results show that for a day ahead prediction, the ARIMA model outperforms the other approaches with an accuracy of 98.91% when executed over a 168 h (1 week) of uninterrupted data for five government buildings.
Mona A. Alduailij, Ioan Petri, Omer F. Rana, Mai A. Alduailij, Abdulrahman S. Aldawood
J. Supercomput.3
2021 NFV-Enabled IoT Service Provisioning in Mobile Edge Clouds
abstract
Conventional Internet of Things (IoT) applications involve data capture from various sensors in environments, and the captured data then is processed in remote clouds. However, some critical IoT applications (e.g., autonomous vehicles) require a much lower response latency and more secure guarantees than those offered by remote clouds today. Mobile edge clouds (MEC) supported by the network function virtualization (NFV) technique have been envisioned as an ideal platform for supporting such IoT applications. Specifically, MECs enable to handle IoT applications in edge networks to shorten network latency, and NFV enables agile and low-cost network functions to run in low-cost commodity servers as virtual machines (VMs). One fundamental problem for the provisioning of IoT applications in an NFV-enabled MEC is where to place virtualized network functions (VNFs) for IoT applications in the MEC, such that the operational cost of provisioning IoT applications is minimized. In this paper, we first address this fundamental problem, by considering a special case of the IoT application placement problem, where the IoT application and VNFs of each service request are consolidated into a single location (gateway or cloudlet), for which we propose an exact solution and an approximation algorithm with a provable approximation ratio. We then develop a heuristic algorithm that controls the resource violation ratios of edge clouds in the network. For the IoT application placement problem for IoT applications where their VNFs can be placed to multiple locations, we propose an efficient heuristic that jointly places the IoT application and its VNFs. We finally study the performance of the proposed algorithms by simulations and implementations in a real test-bed, Experimental results show that the performance of the proposed algorithms outperform their counterparts by at least 10 percent.
Zichuan Xu, Wanli Gong, Qiufen Xia, Weifa Liang, Omer F. Rana, Guowei Wu 0001
IEEE Trans. Mob. Comput.5
2021 Synthesising Privacy by Design Knowledge Toward Explainable Internet of Things Application Designing in Healthcare
abstract
Privacy by Design (PbD) is the most common approach followed by software developers who aim to reduce risks within their application designs, yet it remains commonplace for developers to retain little conceptual understanding of what is meant by privacy. A vision is to develop an intelligent privacy assistant to whom developers can easily ask questions to learn how to incorporate different privacy-preserving ideas into their IoT application designs. This article lays the foundations toward developing such a privacy assistant by synthesising existing PbD knowledge to elicit requirements. It is believed that such a privacy assistant should not just prescribe a list of privacy-preserving ideas that developers should incorporate into their design. Instead, it should explain how each prescribed idea helps to protect privacy in a given application design context—this approach is defined as “Explainable Privacy.” A total of 74 privacy patterns were analysed and reviewed using ten different PbD schemes to understand how each privacy pattern is built and how each helps to ensure privacy. Due to page limitations, we have presented a detailed analysis in Reference [3]. In addition, different real-world Internet of Things (IoT) use-cases, including a healthcare application, were used to demonstrate how each privacy pattern could be applied to a given application design. By doing so, several knowledge engineering requirements were identified that need to be considered when developing a privacy assistant. It was also found that, when compared to other IoT application domains, privacy patterns can significantly benefit healthcare applications. In conclusion, this article identifies the research challenges that must be addressed if one wishes to construct an intelligent privacy assistant that can truly augment software developers’ capabilities at the design phase.
Lamya Alkhariji, Nada Alhirabi, Mansour Naser Alraja, Mahmoud Barhamgi, Omer F. Rana, Charith Perera
ACM Trans. Multim. Comput. Commun. Appl.5
2021 Energy-Aware Inference Offloading for DNN-Driven Applications in Mobile Edge Clouds
abstract
With increasing focus on Artificial Intelligence (AI) applications, Deep Neural Networks (DNNs) have been successfully used in a number of application areas. As the number of layers and neurons in DNNs increases rapidly, significant computational resources are needed to execute a learned DNN model. This ever-increasing resource demand of DNNs is currently met by large-scale data centers with state-of-the-art GPUs. However, increasing availability of mobile edge computing and 5G technologies provide new possibilities for DNN-driven AI applications, especially where these application make use of data sets that are distributed in different locations. One fundamental process of a DNN-driven application in mobile edge clouds is the adoption of “inferencing” - the process of executing a pre-trained DNN based on newly generated image and video data from mobile devices. We investigate offloading DNN inference requests in a 5G-enabled mobile edge cloud (MEC), with the aim to admit as many inference requests as possible. We propose exact and approximate solutions to the problem of inference offloading in MECs. We also consider dynamic task offloading for inference requests, and devise an online algorithm that can be adapted in real time. The proposed algorithms are evaluated through large-scale simulations and using a real world test-bed implementation. The experimental results demonstrate that the empirical performance of the proposed algorithms outperform their theoretical counterparts and other similar heuristics reported in literature.
Zichuan Xu, Liqian Zhao, Weifa Liang, Omer F. Rana, Pan Zhou 0001, Qiufen Xia, Wenzheng Xu, Guowei Wu 0001
IEEE Trans. Parallel Distributed Syst.4
2021 Scheduling Real-Time Security Aware Tasks in Fog Networks
abstract
Fog computing brings the cloud closer to a user with the help of a micro data center ($mdc$), leading to lower response times for delay sensitive applications.RT-SANE(Real-TimeSecurityAware scheduling on theNetworkEdge) supports batch and interactive applications, taking account of their deadline and security constraints. RT-SANE chooses between an$mdc$(in proximity to a user) and a cloud data center ($cdc$) by taking account of network delay and security tags. Jobs submitted by a user are tagged as: private, semi-private and public, and$mdcs$and$cdcs$are classified as: trusted, semi-trusted and untrusted. RT-SANE executes private jobs on a user’s local$mdcs$or pre-trusted$cdcs$, and semi-private and public jobs on remote$mdcs$and$cdcs$. A security and performance-aware distributed orchestration architecture and protocol is made use of in RT-SANE. For evaluation, workload traces from the CERIT-SC Cloud system are used. The effect of slow executing straggler jobs on the Fog framework are also considered, involving migration of such jobs. Experiments reveal thatRT-SANEoffers a higher “success ratio” (successfully completed jobs) to comparable algorithms, including consideration of security tags.
Anil Singh, Nitin Auluck, Omer F. Rana, Andrew C. Jones, Surya Nepal
IEEE Trans. Serv. Comput.3
2020 Learning-based Online Query Evaluation for Big Data Analytics in Mobile Edge Clouds
abstract
The rise of big data brings extraordinary benefits and opportunities to businesses and governments. Enterprise users can analyze their consumers' data and infer the business value obtained, such as purchasing goods correlations, customer preferences, and hidden patterns. Meanwhile, with the emerge of big data processing frameworks, such as Hadoop and Tensor-flow, more and more mobile users are embracing big data analytics by issuing queries to analyze their data. In this paper, we investigate the problem of Quality-of-Service (QoS) aware query evaluation for big data analytics in a mobile edge cloud to maximize the system throughput while minimizing the query evaluation time of each admitted query, by exploring the materialization of intermediate query results. We consider dynamic big-data query evaluations where user queries arrive one by one without the knowledge of future arrivals, and the system needs to respond to each query by accepting or rejecting the query immediately. We propose an online algorithm for query admissions within a finite time horizon, the proposed algorithm can intelligently determine whether some immediate results during a query evaluation need to be materialized for later use of other queries, by making use of the Reinforcement Learning (RL) method with predictions. We finally investigate the performance of the proposed algorithm by simulations, and results show that the performance of the proposed algorithm is promising, by achieving a higher system throughput while reducing the average evaluation cost per query by from 20% to 52% compared to the comparison benchmarks.
Qiufen Xia, Zichuan Xu, Weifa Liang, Omer F. Rana, Guowei Wu 0001
ICC5
2020 Blockchain Based Auditable Access Control for Distributed Business Processes
abstract
The use of blockchain technology has been proposed to provide auditable access control for individual resources. However, when all resources are owned by a single organization, such expensive solutions may not be needed. In this work we focus on distributed applications such as business processes and distributed workflows. These applications are often composed of multiple resources/services that are subject to the security and access control policies of different organizational domains. Here, blockchains can provide an attractive decentralized solution to provide auditability. However, the underlying access control policies may be overlapping in terms of the component conditions/rules, and simply using existing solutions would result in repeated evaluation of user's authorization separately for each resource, leading to significant overhead in terms of cost and computation time over the blockchain. To address this challenge, we propose an approach that formulates a constraint optimization problem to generate an optimal composite access control policy. This policy is in compliance with all the local access control policies and minimizes the policy evaluation cost over the blockchain. The developed smart contract(s) can then be deployed to the blockchain, and used for access control enforcement. We also discuss how the access control enforcement can be audited using a game-theoretic approach to minimize cost. We have implemented the initial prototype of our approach using Ethereum as the underlying blockchain and experimentally validated the effectiveness and efficiency of our approach.
Ahmad Akhtar, Basit Shafiq, Jaideep Vaidya, Ayesha Afzal, Shafay Shamail, Omer F. Rana
ICDCS6
2020 Automating GDPR Compliance Verification for Cloud-hosted Services
abstract
Cloud-hosted business processes require access to customer data to complete a transaction, to improve a customer's on-line experience or provide useful product recommendations. However, privacy concerns associated with the use of this data have led to legal regulations that impose restrictions on how such data is requested or processed by an on-line service, with large penalties for violating these restrictions, e.g. the European General Data Protection Regulation (GDPR). We propose a framework for helping cloud-hosted services automate GDPR compliance checking. The framework comprises three steps: represent data flow in business processes with an appropriate abstraction (timed transition systems), formalise GDPR rules and obligations and incorporate them into the same abstraction, and implement the abstraction in a model checking tool (Uppaal) in order to automatically verify compliance of business process activities with GDPR. We demonstrate the approach using a cloud-based purchase order system.
Masoud Barati, George Theodorakopoulos 0001, Omer F. Rana
ISNCC3
2020 Enabling Multicast Slices in Edge Networks
abstract
Telecommunication networks are undergoing a disruptive transition toward distributed mobile edge networks with virtualized network functions (VNFs) [e.g., firewalls, intrusion detection systems (IDSs), and transcoders] within the proximity of users. This transition will enable network services, especially Internet-of-Things (IoT) applications, to be provisioned as network slices with sequences of VNFs, in order to guarantee the performance and security of their continuous data and control flows. In this article, we study the problems of delay-aware network slicing for multicasting traffic of IoT applications in edge networks. We first propose exact solutions by formulating the problems into integer linear programs (ILPs). We further devise an approximation algorithm with an approximation ratio for the problem of delay-aware network slicing for a single multicast slice, with the objective to minimize the implementation cost of the network slice subject to its delay requirement constraint. Given multiple multicast slicing requests, we also propose an efficient heuristic that admits as many user requests as possible, through exploring the impact of a nontrivial interplay of the total computing resource demand and delay requirements. We then investigate the problem of delay-oriented network slicing with given levels of delay guarantees, considering that different types of IoT applications have different levels of delay requirements, for which we propose an efficient heuristic based on reinforcement learning (RL). We finally evaluate the performance of the proposed algorithms through both simulations and implementations in a real testbed. The experimental results demonstrate that the proposed algorithms are promising.
Yugen Qin, Qiufen Xia, Zichuan Xu, Pan Zhou 0001, Alex Galis, Omer F. Rana, Jiankang Ren, Guowei Wu 0001
IEEE Internet Things J.6
2020 ThermoSim: Deep learning based framework for modeling and simulation of thermal-aware resource management for cloud computing environments
Sukhpal Singh, Shreshth Tuli, Adel Nadjaran Toosi, Félix Cuadrado, Peter Garraghan, Rami Bahsoon, Hanan Lutfiyya, Rizos Sakellariou, Omer F. Rana, Schahram Dustdar, Rajkumar Buyya
J. Syst. Softw.9
2020 Software tools and techniques for fog and edge computing
abstract
The Internet of Things (IoT) paradigm promises to make “things” such as physical objects with sensing capabilities and/or attached with tags, mobile objects such as smartphones and vehicles, consumer electronic devices, and home appliances such as fridge, television, health care devices, as part of the Internet environment. In cloud-centric IoT applications, the sensor data from these “things” is extracted, accumulated, and processed at the public/private clouds, leading to significant latencies. To satisfy the ever increasing demand for cloud computing resources from emerging applications such as IoT, academics and industry experts are now advocating for going from large-centralized cloud computing infrastructures to micro data centers located at the edge of the network. These micro data centers are often closer to a user (geographically and in access latency) compared to the centralized cloud data center. The aim of utilizing such edge resources is to off load computation that would have “traditionally” been carried out at the cloud data center to a resource that is closer to a user or edge devices. This vision also acknowledges the variation in network latency from an end-user to cloud data center. While the network around a data center is often high capacity and speed, that near the user device may have variable properties (in terms of resilience, bandwidth, latency, etc.). Referred to as “fog/edge computing,” this paradigm is expected to improve the agility of cloud service deployments in addition to bringing computing resources closer to end-users. The emergence of computing paradigms such as edge and fog computing supports the data analysis near the data sources for a wide range of applications. Edge computing is the middle layer between users and cloud data centers, and it plays an important role in the IoT use cases where applications required near real-time actions. The intermediate edge layer provides limited computing and storage resources, which consists of network gateways and micro data centers. Edge computing application orchestration is capable of big data processing and can be installed in heterogeneous hardware configurations. Due to the large-scale deployment and device heterogeneity, edge computing infrastructure designing and implementation are challenging, including model analysis, system integration, protocol designing, energy, and security modeling. In addition, since edge data centers are installed in network gateways with open network configuration, they are prone to several network threats, less trustworthy, and easy to compromise. On the one hand, the development of fog and edge clouds includes dedicated facilities, operating system, network, and middleware techniques to build and operate such micro data centers that host virtualized computing resources. On the other hand, the use of fog and edge clouds requires extension to current programming models and proposes new abstractions that will allow developers to design new applications that take benefit from such massively distributed systems. The use of this approach also opens up other challenges in security and privacy (as a user now needs to “trust” every micro data center they interact with), support for resource management for mobile users who transfer session from one micro data center to another, and support for “embedding” such micro data centers into devices (eg, cars, buildings, etc). The objective of this special issue is to disseminate original contributions and research findings concerning the challenges and changes (both evolutionary and disruptive) in edge and fog computing. It provides cutting-edge research from both academia and industry, with emphasis on current developments and future directions in security and privacy issues of emerging fog computing. The call for special issues received a number of submissions. Each paper was reviewed by at least three reviewers and went through at least two rounds of reviews. After a two-phase peer review process, we have accepted 14 high-quality papers related to the aforementioned areas of interest. The accepted papers focus on recent solutions by developing novel research ideas around edge and fog computing for several applications, such as health care, smart city, urban pollution monitoring, etc. The brief contributions of these papers are discussed in the following section. The first paper titled “Abnormal visual event detection based on multi-instance learning and autoregressive integrated moving average model in edge-based Smart City surveillance” by Xu et al proposes an abnormal event detection approach based on multi-instance learning and autoregressive integrated moving average model for video surveillance of crowded scenes in urban public places. It utilizes an unsupervised method for abnormal event detection by combining multi-instance visual feature selection and the autoregressive integrated moving average model. This approach has thoroughly experimented, and the experimental results demonstrate the efficiency of the proposed approach by achieving better abnormal event detection performance for a crowded scene of urban public places with an edge environment. The second paper titled “User allocation-aware edge cloud placement in mobile edge computing” by Guo et al studies the edge cloud placement problem, which is to place the edge clouds at the candidate locations and allocate the mobile users to the edge clouds. Further, it formulates as a multi-objective optimization problem with the objective to balance the workload between edge clouds and minimize the service communication delay of mobile users. The experiment results show the performance of the proposed approach in terms of workload balance and communication delay for validation. The third paper titled “A secure fog-based platform for SCADA-based IoT critical infrastructure” by Baker et al contributes a novel security “toolbox” to reinforce the integrity, security, and privacy of SCADA-based IoT critical infrastructure at the fog layer. The toolbox incorporates a key feature, that is, a cryptographic-based access approach to the cloud services using identity-based cryptography and signature schemes at the fog layer. This paper also presents the implementation details of a prototype for our proposed secure fog-based platform and provides performance evaluation results to demonstrate the appropriateness of the proposed platform in a real-world scenario. The results from the experiments demonstrate a superior performance of the secure fog-based platform, which is around 2.8 seconds when adding five virtual machines (VMs), 3.2 seconds when adding 10 VMs, and 112 seconds when adding 1000 VMs, compared to the multilevel user access control platform. The fourth paper titled “Developing applications in large scale, dynamic fog computing: A case study” by Giang et al presents a case study in building fog computing applications using an open-source platform distributed node-RED. It shows how applications can be decomposed and deployed to a geographically distributed infrastructure using distributed node-RED, and how existing software components can be adapted and reused to participate in fog applications. This case study is implemented in a lab-based fog infrastructure and simulated for large-scale evaluation. The fifth paper titled “An osmotic computing infrastructure for urban pollution monitoring” by Longo et al focuses on the design and development of a middleware that integrates data coming from mobile and IoT devices specifically deployed in urban contexts using the osmotic computing paradigm. Moreover, a component of the osmotic membrane has been developed in this paper for security management. The sixth paper titled “Characterizing application scheduling on edge, fog, and cloud computing resources” by Varshney and Simmhan offers a taxonomy of concepts essential for specifying and solving the problem of scheduling applications on edge, fog, and cloud computing resources. The proposed model is divided into multiple steps, initially characterized by the resource capabilities and limitations of these infrastructures and offers a taxonomy of application models, quality-of-service constraints and goals, and scheduling techniques based on a literature review, followed by tabulated key research prototypes and papers using this taxonomy. It also highlights gaps in the literature and open problems remain. The seventh paper titled “Cloud-aided online electroencephalography (EEG) classification system for brain healthcare: A case study of depression evaluation with a lightweight CNN” by Ke et al presents the design of an online EEG classification system aided by cloud centering on a lightweight convolutional neural network (CNN). The system incrementally trains the CNN on cloud and enables hot deployment of the trained classifier without the need to restart the gateway to adapt to the users' needs. The classifier maintains a high convolutional layer to gain the ability of processing high-dimensional EEG segments. Finally, the model is experimented to validate the contribution. The eighth paper titled “SEWMS: An Edge-based Smart Wearable Maintenance System in Communication Network” by Rui et al proposes a dynamic context-aware information push algorithm (DCAIP) by focusing on the current low level in information, complicated scenes, and various information on on-site maintenance. It also presents a smart wearable maintenance system (SEWMS), an edge computing-assisted IoT platform for the real-time guidance of technical experts and systems for on-site maintenance personnel, aiming to improve the efficiency and quality of on-site maintenance. The ninth paper titled “A crosswalk pedestrian recognition system by using deep learning and zebra-crossing recognition techniques” by Dow et al investigates a real-time pedestrian recognition system that ensures high accuracy by using a deep learning classifier and zebra-crossing recognition techniques. The proposed system was designed to improve pedestrian safety and reduce accidents at intersections. Environmental feature vectors were first used to detect zebra crossings and to determine crossing areas. An adaptive mapping technique was then used to map the pedestrian waiting area based on the crossing area. A dual-camera mechanism was used to maintain detection accuracy and improve system fault tolerance. Finally, the you-only-look-once model was used to recognize pedestrians at intersections. The 10th paper titled “Intelligent sentiment analysis approach using edge computing-based deep learning technique” by Sankar et al studies machine learning algorithms to extract the best features from the training review dataset. Then, the selected features are fed into the CNN and other fully connected layers for further processing. This work has also employed a pretrained sentiment analysis model over an Android application framework to classify reviews on a smartphone without the need for any cloud or server-side application programming interface. The 11th paper titled “Pipeline provenance for cloud-based big data analytics” by Wang et al proposes a solution, named LogProv toward realizing the functionalities for big data provenance, which needs to renovate data pipelines or some of big data software infrastructure to generate structured logs for pipeline events, and then stores data and logs separately in cloud space. The LogProv is implemented and deployed in Nectar Cloud, associated with Apache Pig, Hadoop ecosystem, and adopted Elasticsearch to provide query service. The 12th paper titled “Socially aware microcloud service overlay optimization in community networks” by Apolónia et al presents a model, named Select in Community Networks (SELECTinCN), which enhances the overlay creation for pub/sub systems over peer-to-peer (P2P) networks. Moreover, SELECTinCN includes social information based on cooperation within CNs by exploiting the social aspects of the community of practice. The model organizes the peers in a ring topology and provides an adaptive P2P connection establishment algorithm, where each peer identifies the number of connections needed based on the social structure and user availability. The 13th paper titled “DewSim: A trace-driven toolkit for simulating mobile device clusters in Dew computing environments” by Hirsch et al models and develops a trace-based toolkit built on modular software artifacts to speed up research in resource management techniques in Dew environments. A trace-driven methodology is adopted to assure the practical value of simulated scenarios. The toolkit comprises a device profiler application for Android to capture generic battery and central processing unit traces from real devices, a profile mixer to create user interaction baseline traces through generic ones, and an extensible engine to simulate the execution of workloads configurable via text files. The 14th paper titled “How to Place Your Apps in the Fog-State of the Art and Open Challenges” Borgi et al review the existing methodologies to solve the application placement problem in the fog, while pursuing three main objectives. First, it offers a comprehensive overview of the currently employed algorithms, on the availability of open-source prototypes, and on the size of test use cases. Second, it classifies the literature based on the application and fog infrastructure characteristics that are captured by available models, with a focus on the considered constraints and the optimized metrics. Finally, it identifies some open challenges in application placement in the fog. The 15th paper titled ‘SELFNET 5G mobile edge computing infrastructure: Design and prototyping’ by Chirivella-Perez et al. presented the design and prototype implementation of the fifth-generation (5G) mobile edge infrastructure based on a mobile edge computing paradigm. This mobile edge infrastructure SELFNET is an amlganmation of cloud computing, software-defined networking, and network function virtualization to end up with a portable 5G infrastructure testbed which enabled the realistic execution and testing. Finally, in the last paper titled ‘SDN/NFV security framework for fog-to-things computing infrastructure’ by Krishnan et al. proposed DTARS which is a System for Distributed Threat Analytics and Response for an Edge/Fog and SDN integrated architecture. In this system, the detection scheme runs at the data plane wherein a coarse-grained behavioral, anti-spoofing, flow monitoring and fine-grained traffic multi-feature entropy-based algorithms are deployed. The proposed framework has been developed for defense applications on malware testbed. We hope that the research contributions and findings in this special issue would benefit the readers in terms of enhancing their knowledge and encouraging them to work on various aspects of edge and fog computing. We express our sincere thanks to the editor-in-chief for allowing us to organize this special issue. The editorial office staffs are excellent and thanks for their support. We are also thankful to all the authors who made this special issue possible, and to the reviewers for their thoughtful contributions.
Rajiv Ranjan 0001, Massimo Villari, Haiying Shen, Omer F. Rana, Rajkumar Buyya
Softw. Pract. Exp.4
2020 Special Issue: Identification, Information, and Knowledge in the Internet of Things
abstract
Realizing the full potential of the Internet of Things (IoT) requires solving technical and business challenges including the identification of things, their organization, and integration. The subsequent management of large data volumes that are generated from such systems, and the effective use of knowledge-based decision systems that can make use of IoT resources remains a challenge at present. Various representation formats already exist for specifying sensors and devices that are part of the IoT ecosystem. However, many of these are either specific to use within a particular application area (e.g., environmental monitoring), or specific to a middleware platform. Overcoming device, firmware, and data format heterogeneity remains a significant challenge in real world IoT systems. Consequently, dealing with data that are generated from such systems and reasoning with these data is constrained due to these limitations. IoT-based platforms also offer a variety of different communication protocols (for both long range [at low data rates] and short range [at high data rates]), such as SigFox, LoRaWAN, NB-IoT, Wifi Direct, and so on. These protocols generally offer different decision points around energy used, distance covered, and data rates observed. Another aspect of heterogeneity in IoT systems therefore relates to dealing and switching between these protocols based on context of use and application requirements. We received 13 papers aligned with the theme of this special issue. In particular, the benefit of using deep learning to solve “traditional” problems is being recognized by a number of researchers. All papers were initially screened by the editors to ensure alignment with the theme of this special issue, and high-quality papers were sent to reviewers. In total, seven papers were selected for inclusion in the special issue. The submitted papers combine the use of novel methods and demonstrate effective use of experimental techniques. The papers included in this special issue are: A physiological data-driven model for learners' cognitive load detection using HRV-PRV feature fusion and optimized XGBoost classification
Hao Wu 0022, Rongfang Bie, Charith Pereira, Omer F. Rana
Softw. Pract. Exp.4
2020 Guest Editorial: Special Section on Advances of Utility and Cloud Computing Technologies and Services
abstract
The articles in this special section focus on advancements of utility and cloud computing technologies and services. Computing is rapidly moving towards a model where it is provided as services that are delivered in a manner similar to traditional utilities such as water, electricity, gas, and telephony. In such a model, users access services according to their requirements, without regard to where the services are hosted or how they are delivered. Several computing architectures have evolved to realize this utility computing vision, including Grid computing, Service- Oriented Architecture (SOA) and Cloud computing,which has recently shifted into the center of attention in the ICT industry. Increasing numbers of IT vendors are promising to offer applications, storage and computation hosting services with conforming Service-Level Agreements
Ching-Hsien Hsu, Manish Parashar, Omer F. Rana
IEEE Trans. Cloud Comput.3
2020 Security and Privacy Requirements for the Internet of Things: A Survey
abstract
The design and development process for internet of things (IoT) applications is more complicated than that for desktop, mobile, or web applications. First, IoT applications require both software and hardware to work together across many different types of nodes with different capabilities under different conditions. Second, IoT application development involves different types of software engineers such as desktop, web, embedded, and mobile to work together. Furthermore, non-software engineering personnel such as business analysts are also involved in the design process. In addition to the complexity of having multiple software engineering specialists cooperating to merge different hardware and software components together, the development process requires different software and hardware stacks to be integrated together (e.g., different stacks from different companies such as Microsoft Azure and IBM Bluemix). Due to the above complexities, non-functional requirements (such as security and privacy, which are highly important in the context of the IoT) tend to be ignored or treated as though they are less important in the IoT application development process. This article reviews techniques, methods, and tools to support security and privacy requirements in existing non-IoT application designs, enabling their use and integration into IoT applications. This article primarily focuses on design notations, models, and languages that facilitate capturing non-functional requirements (i.e., security and privacy). Our goal is not only to analyse, compare, and consolidate the empirical research but also to appreciate their findings and discuss their applicability for the IoT.
Nada Alhirabi, Omer F. Rana, Charith Perera
ACM Trans. Internet Things2
2020 QoS-Aware VNF Placement and Service Chaining for IoT Applications in Multi-Tier Mobile Edge Networks
abstract
Mobile edge computing and network function virtualization (NFV) paradigms enable new flexibility and possibilities of the deployment of extreme low-latency services for Internet-of-Things (IoT) applications within the proximity of their users. However, this poses great challenges to find optimal placements of virtualized network functions (VNFs) for data processing requests of IoT applications in a multi-tier cloud network, which consists of many small- or medium-scale servers, clusters, or cloudlets deployed within the proximity of IoT nodes and a few large-scale remote data centers with abundant computing and storage resources. In particular, it is challenging to jointly consider VNF instance placement and routing traffic path planning for user requests, as they are not only delay sensitive but also resource hungry. In this article, we consider admissions of NFV-enabled requests of IoT applications in a multi-tier cloud network, where users request network services by issuing service requests with service chain requirements, and the service chain enforces the data traffic of the request to pass through the VNFs in the chain one by one until it reaches its destination. To this end, we first formulate the throughput maximization problem with the aim to maximize the system throughput. We then propose an integer linear program solution if the problem size is small; otherwise, we devise an efficient heuristic that jointly takes into account VNF placements to both cloudlets and data centers and routing path finding for each request. For a special case of the problem with a set of service chains, we propose an approximation algorithm with a provable approximation ratio. Next, we also devise efficient learning-based heuristics for VNF provisioning for IoT applications by incorporating the mobility and energy conservation features of IoT devices. We finally evaluate the performance of the proposed algorithms by simulations. The simulation results show that the performance of the proposed algorithms is promising.
Zichuan Xu, Weifa Liang, Qiufen Xia, Omer F. Rana, Guowei Wu 0001
ACM Trans. Sens. Networks5
2020 Efficient Algorithms for Delay-Aware NFV-Enabled Multicasting in Mobile Edge Clouds With Resource Sharing
abstract
Stringent delay requirements of many mobile applications have led to the development of mobile edge clouds, to offer low latency network services at the network edges. Most conventional network services are implemented via hardware-based network functions, including firewalls and load balancers, to guarantee service security and performance. However, implementing hardware-based network functions usually incurs both a high capital expenditure (CAPEX) and operating expenditure (OPEX). Network Function Virtualization (NFV) exhibits a potential to reduce CAPEX and OPEX significantly, by deploying software-based network functions in virtual machines (VMs) on edge-clouds. We consider a fundamental problem of NFV-enabled multicasting in a mobile edge cloud, where each multicast request has both service function chain and end-to-end delay requirements. Specifically, each multicast request requires chaining of a sequence of network functions (referred to as a service function chain) from a source to a set of destinations within specified end-to-end delay requirements. We devise an approximation algorithm with a provable approximation ratio for a single multicast request admission if its delay requirement is negligible; otherwise, we propose an efficient heuristic. Furthermore, we also consider admissions of a given set of the delay-aware NFV-enabled multicast requests, for which we devise an efficient heuristic such that the system throughput is maximized, while the implementation cost of admitted requests is minimized. We finally evaluate the performance of the proposed algorithms in a real test-bed, and experimental results show that our algorithms outperform other similar approaches reported in literature.
Haozhe Ren, Zichuan Xu, Weifa Liang, Qiufen Xia, Pan Zhou 0001, Omer F. Rana, Alex Galis, Guowei Wu 0001
IEEE Trans. Parallel Distributed Syst.6
2020 Deadline Constrained Video Analysis via In-Transit Computational Environments
abstract
Combining edge processing (at data capture site) with analysis carried out while data is enroute from the capture site to a data center offers a variety of different processing models. Such in-transit nodes include network data centers that have generally been used to support content distribution (providing support for data multicast and caching), but have recently started to offer user-defined programmability, through Software Defined Networks (SDN) capability, e.g., OpenFlow and Network Function Visualization (NFV). We demonstrate how this multi-site computational capability can be aggregated to support video analytics, with Quality of Service and cost constraints (e.g., latency-bound analysis). The use of SDN technology enables separation of the data path from the control path, enabling in-network processing capabilities to be supported as data is migrated across the network. We propose to leverage SDN capability to gain control over the data transport service with the purpose of dynamically establishing data routes such that we can opportunistically exploit the latent computational capabilities located along the network path. Using a number of scenarios, we demonstrate the benefits and limitations of this approach for video analysis, comparing this with the baseline scenario of undertaking all such analysis at a data center located at the core of the infrastructure.
Ali Reza Zamani, Mengsong Zou, Javier Diaz Montes, Ioan Petri, Omer F. Rana, Ashiq Anjum, Manish Parashar
IEEE Trans. Serv. Comput.5
2020 Emotions Behind Drive-by Download Propagation on Twitter
abstract
Twitter has emerged as one of the most popular platforms to get updates on entertainment and current events. However, due to its 280-character restriction and automatic shortening of URLs, it is continuously targeted by cybercriminals to carry out drive-by download attacks, where a user’s system is infected by merely visiting a Web page. Popular events that attract a large number of users are used by cybercriminals to infect and propagate malware by using popular hashtags and creating misleading tweets to lure users to malicious Web pages. A drive-by download attack is carried out by obfuscating a malicious URL in an enticing tweet and used as clickbait to lure users to a malicious Web page. In this article, we answer the following two questions: Why are certain malicious tweets retweeted more than others? Do emotions reflecting in a tweet drive virality? We gathered tweets from seven different sporting events over 3 years and identified those tweets that were used to carry to out a drive-by download attack. From the malicious (N= 105, 642) and benign (N= 169, 178) data sample identified, we built models to predict information flow size and survival. We define size as the number of retweets of an original tweet, and survival as the duration of the original tweet’s presence in the study window. We selected the zero-truncated negative binomial (ZTNB) regression method for our analysis based on the distribution exhibited by our dependent size measure and the comparison of results with other predictive models. We used the Cox regression technique to model the survival of information flows as it estimates proportional hazard rates for independent measures. Our results show that both social and content factors are statistically significant for the size and survival of information flows for both malicious and benign tweets. In the benign data sample, positive emotions and positive sentiment reflected in the tweet significantly predict size and survival. In contrast, for the malicious data sample, negative emotions, especially fear, are associated with both size and survival of information flows.
Amir Javed, Pete Burnap, Matthew L. Williams, Omer F. Rana
ACM Trans. Web4
2019 Enhancing User Privacy in IoT: Integration of GDPR and Blockchain
Masoud Barati, Omer F. Rana
BlockSys2
2019 Edge-Cloud Orchestration: Strategies for Service Placement and Enactment
abstract
As devices existing at the edge of the network improve in their processing and data storage capacity, there is increasing potential to host and enact services on such devices. A workflow that was traditionally enacted on a data centre can be fragmented across both edge and data centre hosted resources. The following aspects are investigated in this work: (i) mechanisms for dividing a workflow across edge and cloud/data centre resources; (ii) service hosting environments that can be shared across edge and data centre resources; (iii) performance metrics that can influence service placement and selection. An "edge orchestrator" is a resource manager that makes such decisions on the behalf of a user application, and which may be centralised or distributed. An industry scenarios is used to illustrate decision points that influence such choices within an edge orchestrator. The overall objective considered is the completion of the workflow within some deadline constraint by the edge orchestrator.
Ioan Petri, Omer F. Rana, Ali Reza Zamani, Yacine Rezgui
IC2E2
2019 NFV-Enabled Multicasting in Mobile Edge Clouds with Resource Sharing
abstract
Driven by stringent delay requirements of mobile applications, the mobile edge cloud has emerged as a major platform to offer low latency network services from the edge of networks. Most conventional network services are implemented via hardware-based network functions, such as firewalls and load balancers, to guarantee service security and performance. However, implementing such hardware-based network functions incurs high purchase and maintenance costs. Network function virtualization (NFV) as a promising technology exhibits great potential to reduce the purchase and maintenance costs by implementing network functions as software in virtual machines (VMs). In this paper, we consider a fundamental problem of NFV-enabled multicasting in a mobile edge cloud, where each multicast request requires to process its traffic in a specified sequence of network functions (referred to as a service chain) before the traffic from a source to a set of destinations. We devise a provable approximation algorithm with an approximation ratio for the problem if requests do not have delay requirements; otherwise, we propose an efficient heuristic for it. We also evaluate the performance of the proposed algorithms against the state-of-the-art NFV-enabled multicasting algorithms, and results show that our algorithms outperform their counterparts.
Zichuan Xu, Yutong Zhang 0003, Weifa Liang, Qiufen Xia, Omer F. Rana, Alex Galis, Guowei Wu 0001, Pan Zhou 0001
ICPP5
2019 Prediction of drive-by download attacks on Twitter
abstract
The popularity of Twitter for information discovery, coupled with the automatic shortening of URLs to save space, given the 140 character limit, provides cybercriminals with an opportunity to obfuscate the URL of a malicious Web page within a tweet. Once the URL is obfuscated, the cybercriminal can lure a user to click on it with enticing text and images before carrying out a cyber attack using a malicious Web server. This is known as a drive-by download . In a drive-by download a user's computer system is infected while interacting with the malicious endpoint, often without them being made aware the attack has taken place. An attacker can gain control of the system by exploiting unpatched system vulnerabilities and this form of attack currently represents one of the most common methods employed. In this paper we build a machine learning model using machine activity data and tweet metadata to move beyond post-execution classification of such URLs as malicious, to predict a URL will be malicious with 0.99 F -measure (using 10-fold cross-validation) and 0.833 (using an unseen test set) at 1 s into the interaction with the URL. Thus, providing a basis from which to kill the connection to the server before an attack has completed and proactively blocking and preventing an attack, rather than reacting and repairing at a later date.
Amir Javed, Pete Burnap, Omer F. Rana
Inf. Process. Manag.3
2019 Getting to the root of the problem: A detailed comparison of kernel and user level data for dynamic malware analysis
abstract
Dynamic malware analysis is fast gaining popularity over static analysis since it is not easily defeated by evasion tactics such as obfuscation and polymorphism. During dynamic analysis it is common practice to capture the system calls that are made to better understand the behaviour of malware. There are several techniques to capture system calls, the most popular of which is a user-level hook. To study the effects of collecting system calls at different privilege levels and viewpoints, we collected data at a process-specific user-level using a virtualised sandbox environment and a system-wide kernel-level using a custom-built kernel driver. We then tested the performance of several state-of-the-art machine learning classifiers on the data. Random Forest was the best performing classifier with an accuracy of 95.2% for the kernel driver and 94.0% at a user-level. The combination of user and kernel level data gave the best classification results with an accuracy of 96.0% for Random Forest. This may seem intuitive but was hitherto not empirically demonstrated. Additionally, we observed that machine learning algorithms trained on data from the user-level tended to use the anti-debug/anti-vm features in malware to distinguish it from benignware. Whereas, when trained on data from our kernel driver, machine learning algorithms seemed to use the differences in the general behaviour of the system to make their prediction, which explains why they complement each other so well. Our results show that capturing data at different privilege levels will affect the classifier’s ability to detect malware, with kernel-level providing more utility than user-level for malware classification. Despite this, there exist more established user-level tools than kernel-level tools, suggesting more research effort should be directed at kernel-level. In short, this paper provides the first objective, evidence-based comparison of user and kernel level data for the purposes of malware classification.
Matthew Nunes, Pete Burnap, Omer F. Rana, Philipp Reinecke, Kaelon Lloyd
J. Inf. Secur. Appl.3
2019 Guest Editors' Introduction to the Special Issue on Fog, Edge, and Cloud Integration for Smart Environments
abstract
editorial Free Access Share on Guest Editors’ Introduction to the Special Issue on Fog, Edge, and Cloud Integration for Smart Environments Authors: Francesco Longo Università degli Studi di Messina, Italy Università degli Studi di Messina, ItalyView Profile , Antonio Puliafito Università degli Studi di Messina, Italy Università degli Studi di Messina, ItalyView Profile , Omer Rana Cardiff University, UK Cardiff University, UKView Profile Authors Info & Claims ACM Transactions on Internet TechnologyVolume 19Issue 2May 2019 Article No.: 17pp 1–4https://doi.org/10.1145/3319404Published:12 April 2019Publication History 4citation490DownloadsMetricsTotal Citations4Total Downloads490Last 12 Months56Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteView all FormatsPDF
Francesco Longo 0001, Antonio Puliafito, Omer F. Rana
ACM Trans. Internet Techn.3
2019 Fog Computing for the Internet of Things: A Survey
abstract
Research in the Internet of Things (IoT) conceives a world where everyday objects are connected to the Internet and exchange, store, process, and collect data from the surrounding environment. IoT devices are becoming essential for supporting the delivery of data to enable electronic services, but they are not sufficient in most cases to host application services directly due to their intrinsic resource constraints. Fog Computing (FC) can be a suitable paradigm to overcome these limitations, as it can coexist and cooperate with centralized Cloud systems and extends the latter toward the network edge. In this way, it is possible to distribute resources and services of computing, storage, and networking along the Cloud-to-Things continuum. As such, FC brings all the benefits of Cloud Computing (CC) closer to end (user) devices. This article presents a survey on the employment of FC to support IoT devices and services. The principles and literature characterizing FC are described, highlighting six IoT application domains that may benefit from the use of this paradigm. The extension of Cloud systems towards the network edge also creates new challenges and can have an impact on existing approaches employed in Cloud-based deployments. Research directions being adopted by the community are highlighted, with an indication of which of these are likely to have the greatest impact. An overview of existing FC software and hardware platforms for the IoT is also provided, along with the standardisation efforts in this area initiated by the OpenFog Consortium (OFC).
Carlo Puliafito, Enzo Mingozzi, Francesco Longo 0001, Antonio Puliafito, Omer F. Rana
ACM Trans. Internet Techn.5
2019 A Dynamic Data-throttling Approach to Minimize Workflow Imbalance
abstract
Scientific workflows enable scientists to undertake analysis on large datasets and perform complex scientific simulations. These workflows are often mapped onto distributed and parallel computational infrastructures to speed up their executions. Prior to its execution, a workflow structure may suffer transformations to accommodate the computing infrastructures, normally involving task clustering and partitioning. However, these transformations may cause workflow imbalance because of the difference between execution task times (runtime imbalance) or because of unconsidered data dependencies that lead to data locality issues (data imbalance). In this article, to mitigate these imbalances, we enhance the workflow lifecycle process in use by introducing a workflow imbalance phase that quantifies workflow imbalance after the transformations. Our technique is based on structural analysis of Petri nets, obtained by model transformation of a data-intensive workflow, and Linear Programming techniques. Our analysis can be used to assist workflow practitioners in finding more efficient ways of transforming and scheduling their workflows. Moreover, based on our analysis, we also propose a technique to mitigate workflow imbalance by data throttling. Our approach is based on autonomic computing principles that determine how data transmission must be throttled throughout workflow jobs. Our autonomic data-throttling approach mainly monitors the execution of the workflow and recompute data-throttling values when certain watchpoints are reached and time derivation is observed. We validate our approach by a formal proof and by simulations along with the Montage workflow. Our findings show that a dynamic data-throttling approach is feasible, does not introduce a significant overhead, and minimizes the usage of input buffers and network bandwidth.
Ricardo J. Rodríguez, Rafael Tolosana-Calasanz, Omer F. Rana
ACM Trans. Internet Techn.3
2019 Deep Learning Hyper-Parameter Optimization for Video Analytics in Clouds
abstract
A system to perform video analytics is proposed using a dynamically tuned convolutional network. Videos are fetched from cloud storage, preprocessed, and a model for supporting classification is developed on these video streams using cloud-based infrastructure. A key focus in this paper is on tuning hyper-parameters associated with the deep learning algorithm used to construct the model. We further propose an automatic video object classification pipeline to validate the system. The mathematical model used to support hyper-parameter tuning improves performance of the proposed pipeline, and outcomes of various parameters on system's performance is compared. Subsequently, the parameters that contribute toward the most optimal performance are selected for the video object classification pipeline. Our experiment-based validation reveals an accuracy and precision of 97% and 96%, respectively. The system proved to be scalable, robust, and customizable for a variety of different applications.
Muhammad Usman Yaseen, Ashiq Anjum, Omer F. Rana, Nikos Antonopoulos
IEEE Trans. Syst. Man Cybern. Syst.3
2018 Edge Enhanced Deep Learning System for Large-Scale Video Stream Analytics
abstract
Applying deep learning models to large-scale IoT data is a compute-intensive task and needs significant computational resources. Existing approaches transfer this big data from IoT devices to a central cloud where inference is performed using a machine learning model. However, the network connecting the data capture source and the cloud platform can become a bottleneck. We address this problem by distributing the deep learning pipeline across edge and cloudlet/fog resources. The basic processing stages and trained models are distributed towards the edge of the network and on in-transit and cloud resources. The proposed approach performs initial processing of the data close to the data source at edge and fog nodes, resulting in significant reduction in the data that is transferred and stored in the cloud. Results on an object recognition scenario show 71\% efficiency gain in the throughput of the system by employing a combination of edge, in-transit and cloud resources when compared to a cloud-only approach.
Muhammad K. Ali, Ashiq Anjum, Muhammad Usman Yaseen, Ali Reza Zamani, Daniel Balouek-Thomert, Omer F. Rana, Manish Parashar
ICFEC6
2018 EclipseIoT: A secure and adaptive hub for the Internet of Things
Eirini Anthi, Shazaib Ahmad, Omer F. Rana, George Theodorakopoulos 0001, Pete Burnap
Comput. Secur.3
2018 Cloud-based scalable object detection and classification in video streams
Muhammad Usman Yaseen, Ashiq Anjum, Omer F. Rana, Richard Hill
Future Gener. Comput. Syst.3
2018 A computational model to support in-network data analysis in federated ecosystems
Ali Reza Zamani, Mengsong Zou, Javier Diaz Montes, Ioan Petri, Omer F. Rana, Manish Parashar
Future Gener. Comput. Syst.5
2017 gSched: a resource aware Hadoop scheduler for heterogeneous cloud computing environments
abstract
Summary MapReduce has become a major programming model for data‐intensive applications in cloud computing environments. Hadoop, an open source implementation of MapReduce, has been adopted by an increasingly wide user community. However, Hadoop suffers from task scheduling performance degradation in heterogeneous contexts because of its homogeneous design focus. This paper presents gSched, a resource‐aware Hadoop scheduler that takes into account both the heterogeneity of computing resources and provisioning charges in task allocation in cloud computing environments. gSched is initially evaluated in an experimental Hadoop cluster and demonstrates enhanced performance compared with the default Hadoop scheduler. Further evaluations are conducted on the Amazon EC2 cloud that demonstrates the effectiveness of gSched in task allocation in heterogeneous cloud computing environments. Copyright © 2016 John Wiley & Sons, Ltd.
Godwin Caruana, Maozhen Li 0001, Man Qi, Mukhtaj Khan, Omer F. Rana
Concurr. Comput. Pract. Exp.5
2017 A social content delivery network for e-Science
abstract
Summary We are in the midst of a scientific data explosion in which the rate of data growth is rapidly increasing. While large‐scale research projects have developed sophisticated data distribution networks to share their data with researchers globally, there is no such support for the many millions of research projects generating data of interest to much smaller audiences (as exemplified by the long tail scientist). In data‐oriented research, every aspect of the research process is influenced by data access. However, sharing and accessing data efficiently as well as lowering access barriers are difficult. In the absence of dedicated large‐scale storage, many have noted that there is an enormous storage capacity available via connected peers, none more so than the storage resources of many research groups. With widespread usage of the content delivery network model for disseminating web content, we believe a similar model can be applied to distributing, sharing, and accessing long tail research data in an e‐Science context. We describe the vision and architecture of a social content delivery network – a model that leverages the social networks of researchers to automatically share and replicate data on peers' resources based upon shared interests and trust. Using this model, we describe a simulator and investigate how aspects such as user activity, geographic distribution, trust, and replica selection algorithms affect data access and storage performance. From these results, we show that socially informed replication strategies are comparable with more general strategies in terms of availability and outperform them in terms of spatial efficiency. Copyright © 2016 John Wiley & Sons, Ltd.
Kyle Chard, Simon Caton, Kai Kugler 0002, Omer F. Rana, Daniel S. Katz
Concurr. Comput. Pract. Exp.4
2017 Introducing distributed dynamic data-intensive (D3) science: Understanding applications and infrastructure
abstract
Summary A common feature across many science and engineering applications is the amount and diversity of data and computation that must be integrated to yield insights. Datasets are growing larger and becoming distributed; their location, availability, and properties are often time‐dependent. Collectively, these characteristics give rise to dynamic distributed data‐intensive applications. While “static” data applications have received significant attention, the characteristics, requirements, and software systems for the analysis of large volumes of dynamic, distributed data, and data‐intensive applications have received relatively less attention. This paper surveys several representative dynamic distributed data‐intensive application scenarios, provides a common conceptual framework to understand them, and examines the infrastructure used in support of applications.
Shantenu Jha, Daniel S. Katz, André Luckow, Neil P. Chue Hong, Omer F. Rana, Yogesh L. Simmhan
Concurr. Comput. Pract. Exp.5
2017 Computational resource management for data-driven applications with deadline constraints
abstract
Summary Recent advances in the type and variety of sensing technologies have led to an extraordinary growth in the volume of data being produced and led to a number of streaming applications that make use of this data. Sensors typically monitor environmental or physical phenomenon at predefined time intervals or triggered by user‐defined events. Understanding how such streaming content (the raw data or events) can be processed within a time threshold remains an important research challenge. We investigate how a cloud‐based computational infrastructure can autonomically respond to such streaming content, offering quality of service guarantees. In particular, we contextualize our approach using an electric vehicles (EVs) charging scenario, where such vehicles need to connect to the electrical grid to charge their batteries. There has been an emerging interest in EV aggregators (primarily intermediate brokers able to estimate aggregate charging demand for a collection of EVs) to coordinate the charging process. We consider predicting EV charging demand as a potential workload with execution time constraints. We assume that an EV aggregator manages a number of geographic areas and a pool of computational resources of a cloud computing cluster to support scheduling of EV charging. The objective is to ensure that there is enough computational capacity to satisfy the requirements for managing EV battery charging requests within specific time constraints.
Rafael Tolosana-Calasanz, Javier Diaz Montes, Omer F. Rana, Manish Parashar, Erotokritos Xydas, Charalampos E. Marmaras, Panagiotis Papadopoulos, Liana Cipcigan
Concurr. Comput. Pract. Exp.3
2017 Scalable real-time classification of data streams with concept drift
abstract
Inducing adaptive predictive models in real-time from high throughput data streams is one of the most challenging areas of Big Data Analytics. The fact that data streams may contain concept drifts (changes of the pattern encoded in the stream over time) and are unbounded, imposes unique challenges in comparison with predictive data mining from batch data. Several real-time predictive data stream algorithms exist, however, most approaches are not naturally parallel and thus limited in their scalability. This paper highlights the Micro-Cluster Nearest Neighbour (MC-NN) data stream classifier. MC-NN is based on statistical summaries of the data stream and a nearest neighbour approach, which makes MC-NN naturally parallel. In its serial version MC-NN is able to handle data streams, the data does not need to reside in memory and is processed incrementally. MC-NN is also able to adapt to concept drifts. This paper provides an empirical study on the serial algorithm’s speed, adaptivity and accuracy. Furthermore, this paper discusses the new parallel implementation of MC-NN, its parallel properties and provides an empirical scalability study.
Mark Tennant, Frederic T. Stahl, Omer F. Rana, João Bártolo Gomes
Future Gener. Comput. Syst.3
2017 Cloud Security Engineering: Theory, Practice and Future Research
abstract
The eleven papers in this special issue address security and privacy concerns associated with cloud computing. This special issue is dedicated to the identification of techniques that enable security mechanisms to be engineered and implemented in cloud services and cloud systems. A key focus is on the integration of theoretical foundations with practical deployment of security strategies that make cloud systems more secure for both end users and providers – enabling end users to increase the level of trust they have in cloud service providers – and conversely for cloud service providers to provide greater guarantees to end users about the security of their services and data.
Kim-Kwang Raymond Choo, Omer F. Rana, Muttukrishnan Rajarajan
IEEE Trans. Cloud Comput.2
2017 Can We Predict a Riot? Disruptive Event Detection Using Twitter
abstract
In recent years, there has been increased interest in real-world event detection using publicly accessible data made available through Internet technology such as Twitter, Facebook, and YouTube. In these highly interactive systems, the general public are able to post real-time reactions to “real world” events, thereby acting as social sensors of terrestrial activity. Automatically detecting and categorizing events, particularly small-scale incidents, using streamed data is a non-trivial task but would be of high value to public safety organisations such as local police, who need to respond accordingly. To address this challenge, we present an end-to-end integrated event detection framework that comprises five main components: data collection, pre-processing, classification, online clustering, and summarization. The integration between classification and clustering enables events to be detected, as well as related smaller-scale “disruptive events,” smaller incidents that threaten social safety and security or could disrupt social order. We present an evaluation of the effectiveness of detecting events using a variety of features derived from Twitter posts, namely temporal, spatial, and textual content. We evaluate our framework on a large-scale, real-world dataset from Twitter. Furthermore, we apply our event detection system to a large corpus of tweets posted during the August 2011 riots in England. We use ground-truth data based on intelligence gathered by the London Metropolitan Police Service, which provides a record of actual terrestrial events and incidents during the riots, and show that our system can perform as well as terrestrial sources, and even better in some cases.
Nasser Alsaedi, Pete Burnap, Omer F. Rana
ACM Trans. Internet Techn.3
2017 Feedback-Control & Queueing Theory-Based Resource Management for Streaming Applications
abstract
Recent advances in sensor technologies and instrumentation have led to an extraordinary growth of data sources and streaming applications. A wide variety of devices, from smart phones to dedicated sensors, have the capability of collecting and streaming large amounts of data at unprecedented rates. A number of distinct streaming data models have been proposed. Typical applications for this include smart cites & built environments for instance, where sensor-based infrastructures continue to increase in scale and variety. Understanding how such streaming content can be processed within some time threshold remains a non-trivial and important research topic. We investigate how a cloud-based computational infrastructure can autonomically respond to such streaming content, offering Quality of Service guarantees. We propose an autonomic controller (based on feedback control and queueing theory) to elastically provision virtual machines to meet performance targets associated with a particular data stream. Evaluation is carried out using a federated Cloud-based infrastructure (implemented using CometCloud)-where the allocation of new resources can be based on: (i) differences between sites, i.e., types of resources supported (e.g., GPU versus CPU only), (ii) cost of execution; (iii) failure rate and likely resilience, etc. In particular, we demonstrate how Little's Law-a widely used result in queuing theory-can be adapted to support dynamic control in the context of such resource provisioning.
Rafael Tolosana-Calasanz, Javier Diaz Montes, Omer F. Rana, Manish Parashar
IEEE Trans. Parallel Distributed Syst.3
2017 Modelling and Implementing Social Community Clouds
abstract
As the number of people who interact on social networks increases, and coupled with the greater capability made available within our computational devices, there is the potential to establish “Social Clouds”-a resource sharing infrastructure that enable people who have trust relationships to come together to share computational/ data services within a community. Social clouds can also provide the means to enhance multi-user collaboration and greatly stimulate the exchange of resources among participants. Recent research in the establishment and use of Social Clouds has raised significant interest by proposing an environment where users are able to trade resources mediated by a social networking mechanism. In such a cloud environment the incentives for sharing can represent a solution for improving resource utilisation and for making available additional capacity to friends and collaborators. In this paper we demonstrate how revenue can be earned within a social cloud community, by executing internal (intra community) and external (inter community) tasks. A number of different scenarios are first investigated through simulation, using the PeerSim simulator, in order to validate our approach. We use two key metrics: revenue and reputation, to evaluate how the system dynamics change as new tasks are added to one or more communities for execution, along with additional behaviours, such as nodes migrating from one community to another, or selectively reporting on the outcome of task execution. Subsequently, we develop a practical deployment using a federated cloud scenario using the CometCloud system-deployed over three sites: Cardiff (UK), Rutgers and Indiana. We show how approaches that have been simulated in PeerSim can be implemented in practice.
Ioan Petri, Javier Diaz Montes, Omer F. Rana, Magdalena Punceva, Ivan Rodero, Manish Parashar
IEEE Trans. Serv. Comput.3
2016 Managing QoS Constraints in a P2P-Cloud Video on Demand System
abstract
As multimedia traffic provides the dominating data flows on Internet channels nowadays, on-demand video streaming has also grown in popularity. Users watching on-demand streams are interested in receiving their video streams without interruption and at low cost. To achieve this, video providers need to build computational infrastructure that can adapt to changes in demand. This paper presents an architecture to build synergy between Peer-to-Peer and Cloud Computing systems to achieve the necessary throughput for video-on-demand providers with reduced costs. We propose heuristics that overcome lack of stability in P2P systems by the inclusion of cloud servers into the pool of video streaming servers, avoiding disruptions (stalls) even when the total number of servers are reduced. Simulation results show that the proposed heuristic strategies can help reduce costs while maintaining quality of service.
Elias De O. Granja, Luiz Fernando Bittencourt, Ioan Petri, Omer F. Rana, Cesar A. V. Melo
CLOUD4
2016 Patterns to distribute mobility simulations
abstract
Travelers mobility simulation is a powerful tool to test strategies in a virtual environment, without impacting the quality of the real traffic network. However, existing mobility multiagent and micro-simulations can only consider a sample of the real volumes of travelers, especially for big regions. With distributed simulations, it would be easier to analyze and predict the status of nowadays networks. This kind of simulations requires big computational power and methods to split the simulation between several machines. This work describes how to achieve such a distribution in a microscopic simulation context, and compare our results with a previous work on macroscopic simulation.
Matthieu Mastio, Mahdi Zargayouna, Gérard Scémama, Omer F. Rana
AICCSA4
2016 Automatic Summarization of Real World Events Using Twitter
Nasser Alsaedi, Pete Burnap, Omer F. Rana
ICWSM3
2016 Sensing Real-World Events Using Arabic Twitter Posts
Nasser Alsaedi, Pete Burnap, Omer F. Rana
ICWSM3
2016 Sensing Real-World Events Using Social Media Data and a Classification-Clustering Framework
abstract
In recent years, there has been increased interest in real-world event identification using data collected from social media, where the Web enables the general public to post real-time reactions to terrestrial events - thereby acting as social sensors of terrestrial activity. Automatically extracting and categorizing activity from streamed data is a non-trivial task. To address this task, we present a novel event detection framework which comprises five main components: data collection, pre-processing, classification, online clustering and summarization. The integration between classification and clustering allows events to be detected - including “disruptive” events - incidents that threaten social safety and security, or could disrupt the social order. We evaluate our framework on a large-scale, real-world dataset from Twitter. We also compare our results to other leading approaches using Flickr MediaEval Event Detection Benchmark.
Nasser Alsaedi, Pete Burnap, Omer F. Rana
WI3
2016 Temporal TF-IDF: A High Performance Approach for Event Summarization in Twitter
abstract
In recent years, there has been increased interest in real-world event summarization using publicly accessible data made available through social networking services such as Twitter and Facebook. People use these outlets to communicate with others, express their opinion and commentate on a wide variety of real-world events. Due to the heterogeneity, the sheer volume of text and the fact that some messages are more informative than others, automatic summarization is a very challenging task. This paper presents three techniques for summarizing microblog documents by selecting the most representative posts for real-world events (clusters). In particular, we tackle the task of multilingual summarization in Twitter. We evaluate the generated summaries by comparing them to both human produced summaries and to the summarization results of similar leading summarization systems. Our results show that our proposed Temporal TF-IDF method outperforms all the other summarization systems for both the English and non-English corpora as they lead to informative summaries.
Nasser Alsaedi, Pete Burnap, Omer F. Rana
WI3
2016 Identifying cyber risk hotspots: A framework for measuring temporal variance in computer network risk
abstract
Modern computer networks generate significant volume of behavioural system logs on a daily basis. Such networks comprise many computers with Internet connectivity, and many users who access the Web and utilise Cloud services make use of numerous devices connected to the network on an ad-hoc basis. Measuring the risk of cyber attacks and identifying the most recent modus-operandi of cyber criminals on large computer networks can be difficult due to the wide range of services and applications running within the network, the multiple vulnerabilities associated with each application, the severity associated with each vulnerability, and the ever-changing attack vector of cyber criminals. In this paper we propose a framework to represent these features, enabling real-time network enumeration and traffic analysis to be carried out, in order to produce quantified measures of risk at specific points in time. We validate the approach using data from a University network, with a data collection consisting of 462,787 instances representing threats measured over a 144 hour period. Our analysis can be generalised to a variety of other contexts.
Malik Shahzad Kaleem Awan, Pete Burnap, Omer F. Rana
Comput. Secur.3
2016 A multifaceted evaluation of the reference model of information assurance & security
Yulia Cherdantseva, Jeremy Hilton, Omer F. Rana, Wendy Ivins
Comput. Secur.3
2016 Cloud computing for mobile environments
abstract
Cloud computing provides a useful metaphor for combining capability at different scales. Such environments may therefore consist of devices ranging from handheld smart phones to supercomputers, to serve communities ranging from individuals to whole industries. Cloud computing requirements are also widely regarded as a key enabler for next generation network environments and are expected to drive the creation of both jobs and commercial products. Cloud computing is also sometimes referred synonymously as next generation data centers, with an architecture consisting of network of virtual services (hardware, database, user-interface, and application logic) so that users are able to deploy and access applications globally and on demand at competitive costs depending on user driven quality of service requirements. Developers with innovative ideas for new Internet services no longer require large capital outlays in hardware to deploy their services, or human expense to operate it 1-7. This technology is being driven by and used in a wide range of academic, research, and commercial application areas. This use is producing important new practical experience in a variety of different problem domains in each of these areas. There are also new computational methods, such as mobile agents, cellular automata, and massively parallel neural networks, which are particularly suited to concurrent execution. As an enabler, this technology is leading to a rapid growth in both scientific and information applications that will, in turn, enable additional requirements for cloud computing technologies to be identified. These will impact academia, business, and education 4-6, 8-12. Given this context, this special issue calls for high-quality research papers in the development of cloud computing technology for mobile environments. In particular, the special issue showcased the most recent achievements and developments in the realm of cloud computing. Original research articles were solicited covering theoretical studies, practical applications, new communication technology, and experimental prototypes. All submitted papers were peer reviewed and selected on the basis of both their quality and their relevance to the theme of this special issue. The published papers are expected to focus on novel approaches for advanced cloud computing of mobile environments for future computing environments and to present high-quality results for tackling problems arising from the ever-growing advanced cloud computing technologies and services for future computing environments. We have received many manuscripts. Only five manuscripts with high quality were finally selected for this special issue. Each manuscript selected was blindly reviewed by at least three reviewers consisting of guest editors and external reviewers. We present a brief overview of each manuscript in the following. Recent advances in cloud computing for mobile environments have created new topics of interest that included the following: (1) cloud economics in mobile and pervasive environments; (2) cloud computing programming and application development; (3) scalability, discovery of services, and data in cloud computing infrastructures; (4) trust and clouds for mobile infrastructure; (5)client-cloud computing challenges; (6) grid computing services and applications for cloud deployments; (7) virtualization, modeling, and metadata in grid; (8) cloud monitoring, control, and management; (9) traffic and load balancing in multi-vendor cloud environments; (10) user profile for cloud environments; (11) performance and security in pervasive and ubiquitous cloud environments; (12) fault tolerance, resilience, survivability, and robustness in cloud environments; and (13) geographical constraints for deploying clouds. And also, the special issues include manuscripts of applications of cloud computing for mobile environments as follows: (1) concurrent solutions to specific problems in academia, industry, and society; (2) concurrent algorithms and computational methods; (3) programming environments, operating systems, tools, concurrent languages, compilers, and interpreters; (4) performance prediction, analysis, models, and results; (5) applications, algorithms, and software technologies arising from the World Wide Web; (6) unification of computing and communication and unification of parallel and distributed computing; (7) social network analysis facilitated through cloud environments; (8) ad hoc and mesh networks via cloud systems; and (9) managing streaming content with cloud environments. In these several topics of cloud computing, some articles proposed the following: Park et al. 13 proposed and experimented the group classification for a dynamic environment, where information of mobile devices is changing. The existing studies, which provide arbitrary cutoff points, are not appropriate in a dynamic environment. The proposed method provides dynamic grouping by setting cutoff points and reflecting the changes of mobile devices after calculating the entropy values. We also proposed a fault tolerance technique for a dynamic environment based on our grouping algorithm. In the proposed algorithm, different fault tolerance techniques are applied to every group according to the different information of mobile devices of each group. Choi and Lim 14, in order to maximize the parallelism within a basic block, we exploit a novel register allocation approach, called as alternative register allocation. The alternative register allocation enables the target processors to improve the parallelism by eliminating the instruction dependency within a translated block. Alternative register allocation divides the register set of the target processor into two separate groups that are used to alternately handle translated blocks. This paper achieved performance improvement by using the two sets of register allocation orders. Alternative register allocation is to maximize the pipelining performance, which is a very popular technique for all modern microprocessors, eliminating the false data dependency. Thus, this paper resolves the false dependency problem and outperforms conventional approach by up to 26.3% for some cases. Sood 15 provided an efficient prediction model for adaptive resource provisioning in the cloud for dynamic resource management. The strategy used in this paper helps the cloud service provider for optimizing resource allocation as well as making the cloud cost-effective. Load-balancing algorithm is used to divide the resource requirements of application in different clouds and, at the same time, to minimize the cost to the best possible extent. The approach used in this paper is best suitable for making the decision to allocate or release a Virtual Machine (VM) in time. This paper had also evaluated the metrics described in manuscript in terms of cost and functionality by performing various experiments. These experiments validate the accuracy of the proposed method, and the results show that the prediction accuracy of the proposed technique is considerably effective. Thus, the proposed approach is highly effective in terms of both cost and functionality. Wang et al. 16 provided a dynamic power model based on division of computation and memory and a static power model based on real-time temperature perception. The evaluation of nine typical applications showed that with the method proposed in this research, about 35% of dynamic energy (within 10% error bounds) subject to performance constraint conditions can be saved, while about 21.6% of static energy could be saved with a performance loss of less than 2.8%. De and De 17 proposed a new mobile cloud architecture, where the interactions between mobile applications and mobile cloud services are decoupled during service invocation and delivery phases. The tuple space model is used in this architecture to decouple such interactions, thereby improving adaptability and robustness of mobile cloud computing. This paper also suggested an approach of modeling and formally verifying the correctness of the proposed architecture, which also validated its adaptability and robustness during interactions between mobile applications and mobile cloud services. The modeling and verification are carried out using the mobile UNITY model. Our special thanks go to Professor Geoffrey C. Fox and Professor David W. Walker who are editors-in-chief of Concurrency and Computation: Practical and Experience and all editorial staffs for their valuable supports throughout the preparation and publication of this special issue. We would like to thank all authors for their contributions to this special issue. We also extend our thanks to the external reviewers for their excellent help in reviewing the manuscripts.
Hwa-Young Jeong, Omer F. Rana, Ching-Hsien Hsu, Young-Sik Jeong
Concurr. Comput. Pract. Exp.2
2016 Analyzing Hadoop power consumption and impact on application QoS
Javier Conejero, Omer F. Rana, Pete Burnap, Jeffrey Morgan, María Blanca Caminero, Carmen Carrión 0001
Future Gener. Comput. Syst.2
2016 Automated configuration support for infrastructure migration to the cloud
Jesús García-Galán, Pablo Trinidad Martín-Arroyo, Omer F. Rana, Antonio Ruiz Cortés
Future Gener. Comput. Syst.3
2016 Cloud Market Maker: An automated dynamic pricing marketplace for cloud users
Barkha Javed, Peter Bloodsworth, Raihan Ur Rasool, Kamran Munir, Omer F. Rana
Future Gener. Comput. Syst.5
2016 Resource management for bursty streams on multi-tenancy cloud environments
Rafael Tolosana-Calasanz, José A. Bañares, CongDuc Pham, Omer F. Rana
Future Gener. Comput. Syst.4
2016 Digital Wildfires: Propagation, Verification, Regulation, and Responsible Innovation
abstract
Social media platforms provide an increasingly popular means for individuals to share content online. Whilst this produces undoubted societal benefits, the ability for content to be spontaneously posted and reposted creates an ideal environment for rumour and false/malicious information to spread rapidly. When this occurs it can cause significant harm and can be characterised as a “digital wildfire.” In this article, we demonstrate that the propagation and regulation of digital wildfires form important topics for research and conduct an overview of existing work in this area. We outline the relevance of a range of work from the computational and social sciences, including a series of insights into the propagation of rumour and false/malicious information. We argue that significant research gaps remain—for instance, there is an absence of systematic studies on the effects of digital wildfires and there is a need to combine empirical research with a consideration of how the responsible governance of social media can be determined. We propose an agenda for research that establishes a methodology to explore in full the propagation and regulation of unverified content on social media. This agenda promotes high-quality interdisciplinary research that will also inform policy debates.
Helena Webb, Pete Burnap, Rob Procter, Omer F. Rana, Bernd C. Stahl, Matthew L. Williams, William Housley, Adam Edwards, Marina Jirotka
ACM Trans. Inf. Syst.4
2015 Identifying Disruptive Events from Social Media to Enhance Situational Awareness
abstract
Decision makers use information from a range of terrestrial and online sources to help underpin the processes through which they develop policies and react to events as they unfold. One such source of online information is social media. Twitter, as a form of social media, is a popular micro-blogging Web application serving hundreds of millions of users. User-generated content can be exploited as a rich source of information for identifying 'real-world' disruptive events. In this paper, we present an in-depth comparison of three types of features that could be useful for identifying disruptive events: temporal, spatial and textual. We make several interesting observations: first, disruptive events are identifiable regardless of the "influence of the user" discussing them, and over a variety of topics. Second, temporal features are the best event identifiers and hence should not be disregarded or ignored. Third, a combination of optimum textual features with temporal and spatial features achieves best performance in the event detection task. We believe that these findings provide new insights for gathering information around real-world events as well as a useful resource for improving situational awareness and decision support.
Nasser Alsaedi, Pete Burnap, Omer F. Rana
ASONAM3
2015 Real-time Classification of Malicious URLs on Twitter using Machine Activity Data
abstract
Massive online social networks with hundreds of millions of active users are increasingly being used by Cyber criminals to spread malicious software (malware) to exploit vulnerabilities on the machines of users for personal gain. Twitter is particularly susceptible to such activity as, with its 140 character limit, it is common for people to include URLs in their tweets to link to more detailed information, evidence, news reports and so on. URLs are often shortened so the endpoint is not obvious before a person clicks the link. Cyber criminals can exploit this to propagate malicious URLs on Twitter, for which the endpoint is a malicious server that performs unwanted actions on the person's machine. This is known as a drive-by-download. In this paper we develop a machine classification system to distinguish between malicious and benign URLs within seconds of the URL being clicked (i.e. 'real-time'). We train the classifier using machine activity logs created while interacting with URLs extracted from Twitter data collected during a large global event -- the Superbowl -- and test it using data from another large sporting event -- the Cricket World Cup. The results show that machine activity logs produce precision performances of up to 0.975 on training data from the first event and 0.747 on a test data from a second event. Furthermore, we examine the properties of the learned model to explain the relationship between machine activity and malicious software behaviour, and build a learning curve for the classifier to illustrate that very small samples of training data can be used with only a small detriment to performance.
Pete Burnap, Amir Javed, Omer F. Rana, Malik Shahzad Kaleem Awan
ASONAM3
2015 Integrating Software Defined Networks within a Cloud Federation
abstract
Cloud computing has generally involved the use of specialist data centres to support computation and data storage at a central site (or a limited number of sites). The motivation for this has come from the need to provide economies of scale (and subsequent reduction in cost) for supporting large scale computation for multiple user applications over (generally) a shared, multi-tenancy infrastructure. The use of such infrastructures requires moving data to a central location (data may be pre-staged to such a location prior to processing using terrestrial delivery channels and does not always require the use of a network-based transfer), undertaking processing on the data, and subsequently enabling users to download results of analysis. We extend this model using software defined networks (SDNs), whereby capability within the network can be used to support in-transit processing while data is in movement from source to destination. Using a smart building infrastructure scenario, consisting of sensors and actuators embedded within a built environment, we describe how an SDN-based architecture can be used to support real time data processing. This significantly influences the processing times to support energy optimisation of the building and reduces costs. We describe an architecture for such a distributed, multi-layered Cloud system and discuss a prototype that has been implemented using the CometCloud system, deployed across three sites in the UK and the US. Wevalidate the prototype using data from sensors within a Sports facility and making use of EnergyPlus.
Ioan Petri, Mengsong Zou, Ali Reza Zamani, Javier Diaz Montes, Omer F. Rana, Manish Parashar
CCGRID5
2015 In-transit Analytics on Distributed Clouds - Applications and Architecture
Omer F. Rana
CLOSER1
2015 Assessing Data Breach Risk in Cloud Systems
abstract
The emerging cloud market introduces a multitude of cloud service providers, making it difficult for consumers to select providers who are likely to be a low risk from a security perspective. Recently, significant emphasis has arisen on the need to specify Service Level Agreements that address security concerns of consumers (referred to as SecSLAs) -- these are intended to clarify security support in addition to Quality of Service characteristics associated with services. It has been found that such SecSLAs are not consistent among providers, even though they offer services with similar functionality. However, measuring security service levels and the associated risk plays an important role when choosing a cloud provider. Data breaches have been identified as a high priority threat influencing the adoption of cloud computing. This paper proposes a general analysis framework which can compute risk associated with data breaches based on pre-agreed SecSLAs for different cloud providers. The framework exploits a tree based structure to identify possible attack scenarios that can lead to data breaches in the cloud and a means of assessing the use of potential mitigation strategies to reduce such breaches.
Yo Rahul, Muttukrishnan Rajarajan, Omer F. Rana, Malik Shahzad Kaleem Awan, Pete Burnap, Sajal K. Das 0001
CloudCom3
2015 Implementing Migration-Aware Virtual Machines
abstract
Virtual Machines hosted in cloud systems are susceptible to migration usually without notifying the cloud consumer. This is generally undertaken to load balance user requests across multiple data centres, often without direct awareness of the user. Migration could be to a regional site or to a data centre in another geographical area, i.e. to a country which has non-conforming laws with regards to data privacy. This concern becomes even more significant when a cloud federation is considered, where a number of different providers may need to work together. It is therefore necessary to develop a mechanism that enables a user to detect if migration of a VM has happened. More importantly, such a mechanism should be user driven and not require input from a provider. We compare various techniques to enable a VM migration to be detected, by monitoring events inside a VM that could signify whether such a migration has taken place, and subsequently notifying the owner about such an event. A review of migration detection techniques is presented followed by the proposition of a hybrid model to carry out the migration detection process.
Taimur Al Said, Omer F. Rana
CSCloud2
2015 Towards a Distributed Multiagent Travel Simulation
Matthieu Mastio, Mahdi Zargayouna, Omer F. Rana
KES-AMSTA3
2015 A classification framework for distinct cyber-attacks based on occurrence patterns
abstract
An increasingly mature, stealthy and dynamic techniques and attack vectors used by cyber criminals have made the critical network infrastructure more vulnerable to security breaches. Following 'Bring Your Own Device (BYOD)' policies and remote-work style of accessing network infra structure leaves the whole network vulnerable to new unknown malware, botnets, advanced persistent threats, coordinated attack patterns, etc., in addition to existing vulnerabilities inherent in software applications. Such an environment demands a network administrator to understand the nature and patterns of cyber-attacks targeting the network infra structure so that appropriate measures could be introduced. In this paper we propose a framework to classify cyber-attacks based on their pattern of occurrence. We validate the classification approach using real malicious traffic logs by focusing on: i) temporal behaviour of cyber-attacks; ii) correlation between cyber-attacks; and iii) targeted software applications.
Malik Shahzad Kaleem Awan, Mohammad A. Alghamdi, Sultan H. Almotiri, Pete Burnap, Omer F. Rana
SIN5
2015 Incentivising resource sharing in social clouds
abstract
Summary Social Clouds provide the capability to share resources among participants within a social network—leveraging on the trust relationships already existing between such participants. In such a system, users are able to trade resources between each other rather than make use of capability offered at a (centralized) data center. Although such an environment has significant potential for improving resource utilization and making available additional capacity that remains dormant, incentives for sharing remain an important hurdle limiting its effective. In this paper, we utilize the socioeconomic model proposed by Silvio Gesell to demonstrate how a ‘virtual currency’ can be used to incentivise sharing of resources within a ‘community’. We subsequently demonstrate, through simulations, the benefit provided to participants within such a community using a variety of economic (such as overall credits gained) and technical (number of successfully completed transactions) metrics. Further, we describe our implementation of such a Social Cloud using CometCloud. CometCloud is an autonomic computing engine for cloud and grid environments. It supports highly heterogeneous and dynamic federated cloud/Grid infrastructures, integration of public/private clouds and autonomic cloudbursts. We demonstrate the implementation of two designs on the basis of the master/worker approach: (i) one tuple space per cluster and (ii) one coordination tuple space and multiple transient spaces—one per each cluster. Finally, we discuss an extended version of our Social Cloud model where intermediary relay nodes take on more active roles as traders in a transaction. Copyright © 2013 John Wiley & Sons, Ltd.
Magdalena Punceva, Ivan Rodero, Manish Parashar, Omer F. Rana, Ioan Petri
Concurr. Comput. Pract. Exp.4
2015 CloudPick: a framework for QoS-aware and ontology-based service deployment across clouds
abstract
SUMMARY The cloud computing paradigm allows on‐demand access to computing and storage services over the Internet. Multiple providers are offering a variety of software solutions in the form of virtual appliances and computing units in the form of virtual machines with different pricing and QoS in the market. Thus, it is important to exploit the benefit of hosting virtual appliances on multiple providers to not only reduce the cost and provide better QoS but also achieve failure‐resistant deployment. This paper presents a framework called CloudPick to simplify cross‐cloud deployment and particularly focuses on QoS modeling and deployment optimization. For QoS modeling, cloud services have been automatically enriched with semantic descriptions using our translator component to increase precision and recall in discovery and benefit from descriptive QoS from multiple domains. In addition, an optimization approach for deploying networks of appliances is required to guarantee minimum cost, low latency, and high reliability. We propose and compare two different deployment optimization approaches: genetic‐based and forward‐checking‐based backtracking. They take into account QoS criteria such as reliability, data communication cost, and latency between multiple clouds to select the most appropriate combination of virtual machines and appliances. We evaluate our approach using a real case study and different request types. Experimental results suggest that both algorithms reach near‐optimal solution. Further, we investigate the effects of factors such as latency, reliability requirements, and data communication between appliances on the performance of the algorithms and placement of appliances across multiple clouds. The results show the efficiency of optimization algorithms depends on the data transfer rate between appliances. Copyright © 2014 John Wiley & Sons, Ltd.
Amir Vahid Dastjerdi, Saurabh Kumar Garg 0001, Omer F. Rana, Rajkumar Buyya
Softw. Pract. Exp.3
2015 Market Models for Federated Clouds
abstract
Multi-cloud systems have enabled resource and service providers to co-exist in a market where the relationship between clients and services depends on the nature of an application and can be subject to a variety of different Quality of Service (QoS) constraints. Deciding whether a cloud provider should host (or finds it profitable to host) a service in the long-term would be influenced by parameters such as the service price, the QoS guarantees required by customers, the deployment cost (taking into account both cost of resource provisioning and operational expenditure, e.g. energy costs) and the constraints over which these guarantees should be met. In a federated cloud system users can combine specialist capabilities offered by a limited number of providers, at particular cost bands-such as availability of specialist co-processors and software libraries. In addition, federation also enables applications to be scaled on-demand and restricts lock in to the capabilities of a particular provider. We devise a market model to support federated clouds and investigate its efficiency in two real application scenarios:(i) energy optimisation in built environments and (ii) cancer image processing both requiring significant computational resources to execute simulations. We describe and evaluate the establishment of such an application based federation and identify a cost-decision based mechanism to determine when tasks should be outsourced to external sites in the federation. The following contributions are provided: (i) understanding the criteria for accessing sites within a federated cloud dynamically, taking into account factors such as performance, cost, user perceived value, and specific application requirements; (ii) developing and deploying a cost based federated cloud framework for supporting real applications over three federated sites at Cardiff (UK), Rutgers and Indiana (USA), (iii) a performance analysis of the application scenarios to determine how task submission could be supported across these three sites, subject to particular revenue targets.
Ioan Petri, Javier Diaz Montes, Mengsong Zou, Thomas H. Beach, Omer F. Rana, Manish Parashar
IEEE Trans. Cloud Comput.5
2015 Governance Model for Cloud Computing in Building Information Management
abstract
The AEC (Architecture Engineering and Construction) sector is a highly fragmented, data intensive, project based industry, involving a number or very different professions and organisations. The industry's strong data sharing and processing requirements means that the management of building data is complex and challenging. We present a data sharing capability utilising Cloud Computing, with two key contributions: 1) a governance model for building data, based on extensive research Pand industry consultation. This governance model describes how individual data artefacts within a building information model relate to each other and how access to this data is controlled; 2) a prototype implementation of this governance model, utilising the CometCloud autonomic cloud computing engine, using the Master/Work paradigm. This prototype is able to successfully store and manage building data, provide security based on a defined policy language and demonstrate scale-out in case of increasing demand or node failure. Our prototype is evaluated both qualitatively and quantitatively. To enable this evaluation we have integrated our prototype with the 3D modelling software-Google Sketchup. We also evaluate the prototype's performance when scaling to utilise additional nodes in the Cloud and to determine its performance in case of node failures.
Thomas H. Beach, Omer F. Rana, Yacine Rezgui, Manish Parashar
IEEE Trans. Serv. Comput.2
2014 Improving Resource Matchmaking through Feedback Integration
abstract
Distributed systems in which users consume and supply different types of services and resources are becoming ever more prevalent. For the matching of consumers and providers, preference-based two-sided matching algorithms can be used to improve the efficiency of the overall match outcome. In such systems, requesting (providing) users rank others based on preferences derived from feedback. Trust can be inferred using trust networks through direct or indirect feedback based on previous actions, and influences the preference ranking that users submit to the matching algorithm. As feedback influences the preference ranking and thus the matching, in this paper we propose a novel methodology which makes use of feedback in the decision making of preference rankings, in order to avoid being matched to untrustworthy users (or provides not likely to deliver their service). We use a simulation based validation approach to determine the effects of dynamic/continuous feedback and its influence on preference ranking of providers.
Christian Haas 0003, Ioan Petri, Omer F. Rana
CCGRID3
2014 Cloud Supported Building Data Analytics
abstract
With increasing availability of instrumented infrastructures in built environments, it is necessary to understand how such data will be stored, processed and analysed in a timely manner. Many "smart cities" applications, for instance, identify how data from building sensors can be combined together to support applications such as emergency response, energy management, etc. Enabling sensor data to be transmitted to a Cloud environment for processing provides a number of benefits, such as scalability and elastic provisioning of computational resources - as the total data size may not be known apriori. In this application-based case study, we describe the integration of an in-building sensor network (both for sensing and actuation) with a distributed Cloud environment. Energy optimisation in buildings represents a class of problems that requires significant computational resources and generally is a time consuming process. We describe the use of Cloud computing for efficiently running and deploying Energy Plus simulations with sensor data in order to fulfil a number of energy related objectives for buildings. We describe and evaluate the establishment of such a sensor based application using a Comet Cloud implementation with data collection from a real building pilot. Although our focus is on a single application, the general architecture and analysis carried out can be generalised to other similar scenarios.
Ioan Petri, Omer F. Rana, Yacine Rezgui, Haijiang Li, Thomas H. Beach, Mengsong Zou, Javier Diaz Montes, Manish Parashar
CCGRID2
2014 Enforcing Quality of Service on OpenNebula-Based Shared Clouds
abstract
With an increase in the number of monitoring sensors deployed on physical infrastructures, there is a corresponding increase in data volumes that need to be processed. Data measured or collected by sensors is typically processed at destination or "in-transit" (i.e. from data capture to delivery to a user). When such data are processed in-transit over a shared distributed computing infrastructure, it is useful to provide elastic computational capability which can be adapted based on processing requirements and demand. Where Service Level Agreements (SLAs) have been pre-agreed, such available computational capacity needs to be shared in such a way that any Quality of Service related constraints in such SLAs are not violated. This is particularly challenging for time critical applications and with highly variable and unpredictable rates of data generation (e.g. in Smart Grid applications where energy usage patterns may change unpredictably). Previously, we proposed a Reference net based architectural model for supporting QoS for multiple concurrent data streams being processed (prior to delivery to a user) over a shared infrastructure. In this paper, we describe a practical realisation of this architecture using the Open Nebula Cloud platform. We consider our infrastructure to be composed of a number of nodes, each of which has multiple processing units and data buffers. We utilize the "token bucket" model for regulating, on a per stream basis, the data injection rate into each node. We subsequently demonstrate how a streaming pipeline can be supported and managed using a dynamic control strategy at each node.
Rafael Tolosana-Calasanz, José A. Bañares, Omer F. Rana, CongDuc Pham, Erotokritos Xydas, Charalampos E. Marmaras, Panagiotis Papadopoulos, Liana Cipcigan
CCGRID3
2014 Towards Real-Time Probabilistic Risk Assessment by Sensing Disruptive Events from Streamed News Feeds
abstract
Risk management has become an important concern over recent years and understanding how risk models could be developed based on the availability of real time (streaming) data has become a challenge. As the volume and velocity of event data (from news media, for instance) continues to grow, we investigate how such data can be used to inform the development of dynamic risk models. A Bayesian Belief Network based approach is adopted in this work, which is able to make use of priors derived from a variety of different news sources (based on data available in RSS feeds).
Pete Burnap, Omer F. Rana, Nargis Pauran, Phil Bowen
CISIS2
2014 Exploring Models and Mechanisms for Exchanging Resources in a Federated Cloud
abstract
One of the key benefits of Cloud systems is their ability to provide elastic, on-demand (seemingly infinite) computing capability and performance for supporting service delivery. With the resource availability in single data centres proving to be limited, the option of obtaining extra-resources from a collection of Cloud providers has appeared as an efficacious solution. The ability to utilize resources from multiple Cloud providers is also often mentioned as a means to: (i) prevent vendor lock in, (ii) to enable in house capacity to be combined with an external Cloud provider, (iii) combine specialist capability from multiple Cloud vendors (especially when one vendor does not offer such capability or where such capability may come at a higher price). Such federation of Cloud systems can therefore overcome a limit in capacity and enable providers to dynamically increase the availability of resources to serve requests. We describe and evaluate the establishment of such a federation using a CometCloud based implementation, and consider a number of federation policies with associated scenarios and determine the impact of such policies on the overall status of our system. CometCloud provides an overlay that enables multiple types of Cloud systems (both public and private) to be federated through the use of specialist gateways. We describe how two physical sites, in the UK and the US, can be federated in a seamless way using this system.
Ioan Petri, Thomas H. Beach, Mengsong Zou, Javier Diaz Montes, Omer F. Rana, Manish Parashar
IC2E5
2014 Analysing Security requirements in Cloud-based Service Level Agreements
abstract
In cloud computing, measurable services such as packet loss and memory are quantized into different levels to provide different level of services to users. Initially, there will be a service level agreement (SLA) between users and service providers (SPs) and/or SPs and infrastructure providers (IPs). However, the most crucial service required by the users and SPs in cloud computing is security and privacy. Security parameters can be used to prevent attacks and to protect data and systems. In literature, there is no comprehensive solution which quantify all the security parameters associated with the cloud computing paradigm. In this paper, for the first time, we attempt to generalize and quantify the security parameters.
Yo Rahul, Pramod S. Pawar, Pete Burnap, Muttukrishnan Rajarajan, Omer F. Rana, George Spanoudakis
SIN5
2014 Virtual Machine Introspection
abstract
Due to exposure to the Internet, virtual machines (VMs) as forms of delivering virtualized infrastructures and resources represent a first point-of-target for security attackers who want to gain access into the virtualization environment. In-VM monitoring approach can be compromised in the event of a successful VM compromise. Virtual Machine Introspection (VMI) takes a different approach of monitoring the guest VMs externally. This paper presents a review on VMI focusing on the typical usages of integrating VMI with other virtualization security techniques.
Thu Yein Win, Huaglory Tianfield, Quentin Mair, Taimur Al Said, Omer F. Rana
SIN5
2014 The management of cloud systems
Antonio Corradi, Omer F. Rana
Future Gener. Comput. Syst.2
2014 Risk assessment in service provider communities
Ioan Petri, Omer F. Rana, Gheorghe Cosmin Silaghi, Yacine Rezgui
Future Gener. Comput. Syst.2
2014 Economics of computing services
Kurt Vanmechelen, Jörn Altmann, Omer F. Rana
Future Gener. Comput. Syst.3
2014 A Social Compute Cloud: Allocating and Sharing Infrastructure Resources via Social Networks
abstract
Social network platforms have rapidly changed the way that people communicate and interact. They have enabled the establishment of, and participation in, digital communities as well as the representation, documentation and exploration of social relationships. We believe that as `apps' become more sophisticated, it will become easier for users to share their own services, resources and data via social networks. To substantiate this, we present a social compute cloud where the provisioning of cloud infrastructure occurs through “friend” relationships. In a social compute cloud, resource owners offer virtualized containers on their personal computer(s) or smart device(s) to their social network. However, as users may have complex preference structures concerning with whom they do or do not wish to share their resources, we investigate, via simulation, how resources can be effectively allocated within a social community offering resources on a best effort basis. In the assessment of social resource allocation, we consider welfare, allocation fairness, and algorithmic runtime. The key findings of this work illustrate how social networks can be leveraged in the construction of cloud computing infrastructures and how resources can be allocated in the presence of user sharing preferences.
Simon Caton, Christian Haas 0003, Kyle Chard, Kris Bubendorfer, Omer F. Rana
IEEE Trans. Serv. Comput.5
2013 Scaling Archived Social Media Data Analysis Using a Hadoop Cloud
abstract
Over recent years, there has been an emerging interest in supporting social media analysis for marketing, opinion analysis and understanding community cohesion. Social media data conforms to many of the categorisations attributed to "big-data" -- i.e. volume, velocity and variety. Generally analysis needs to be undertaken over large volumes of data in an efficient and timely manner. A variety of computational infrastructures have been reported to achieve this. We present the COSMOS platform supporting sentiment and tension analysis on Twitter data, and demonstrate how this platform can be scaled using the OpenNebula Cloud environment with Map/Reduce-based analysis using Hadoop. In particular, we describe the types of system configurations that would be most useful from a performance perspective -- i.e. how virtual machines in the infrastructure should be distributed to reduce variability in the analysis performance. We demonstrate the approach using a data set consisting of several million Twitter messages, analysed over two types of Cloud infrastructure.
Javier Conejero, Pete Burnap, Omer F. Rana, Jeffrey Morgan
IEEE CLOUD3
2013 Broker Emergence in Social Clouds
abstract
Cloud computing generally involves the use of data storage and computational resources from external providers. Although a number of commercial providers are currently on the market, it is often beneficial for a user to consider capability from a number of different ones. This would prevent vendor lock-in and more economic choice for a user. Based on this observation, work on "Social Clouds" has involved using social relationships formed between individuals and institutions to establish Peer-2-Peer resource sharing networks, enabling market forces to determine how demand for resources can be met by a number of different (often individually owned) providers. In this paper we identify how trading within such a network could be enhanced by the dynamic emergence (or identification) of brokers -- based on their social position in the network (based on connectivity metrics within a social network). We investigate how offering financial incentives to such brokers, once discovered, could help improve the number of trades that could take place with a network. A social score algorithm is described and simulated with PeerSim to validate our approach. We also compare the approach to a distributed dominating set algorithm - the closest approximation to our approach.
Ioan Petri, Magdalena Punceva, Omer F. Rana, George Theodorakopoulos 0001
IEEE CLOUD3
2013 Managing Data and Processes in Cloud-Enabled Large-Scale Sensor Networks: State-of-the-Art and Future Research Directions
abstract
In this paper we focus on state-of-the-art analysis and open research issues in the context of Cloud-enabled large-scale sensor networks, which naturally marry with the emerging Big Data paradigm. We particularly address how data and processes are represented and managed in such infrastructures, by highlighting benefits and limitations. We also provide rigorous and critical discussion on actual trends and solutions available in literature, along with future research directions in this scientific field.
Alfredo Cuzzocrea, Giancarlo Fortino, Omer F. Rana
CCGRID3
2013 A Distributed In-Transit Processing Infrastructure for Forecasting Electric Vehicle Charging Demand
abstract
With an increasing interest in Electric Vehicles (EVs), it is essential to understand how EV charging could impact demand on the Electricity Grid. Existing approaches used to achieve this make use of a centralised data collection mechanism - which often is agnostic of demand variation in a given geographical area. We present an in-transit data processing architecture that is more efficient and can aggregate a variety of different types of data. A model using Reference nets has been developed and evaluated. Our focus in this paper is primarily to introduce requirements for such an architecture.
Rafael Tolosana-Calasanz, José A. Bañares, Liana Cipcigan, Omer F. Rana, Panagiotis Papadopoulos, CongDuc Pham
CCGRID4
2013 Characterising the Power Consumption of Hadoop Clouds - A Social Media Analysis Case Study
Javier Conejero, Omer F. Rana, Pete Burnap, Jeffrey Morgan, Carmen Carrión 0001, María Blanca Caminero
CLOSER2
2013 Migrating to the Cloud - A Software Product Line based Analysis
Jesús García-Galán, Omer F. Rana, Pablo Trinidad Martín-Arroyo, Antonio Ruiz Cortés
CLOSER2
2013 Energy Efficient Cloud Computing Environment via Autonomic Meta-director Framework
Thar Baker, Yanik Ngoko, Rafael Tolosana-Calasanz, Omer F. Rana, Martin Randles
DeSE4
2013 Constructing a Social Content Delivery Network for eScience
abstract
Increases in the size of research data and the move towards citizen science, in which everyday users contribute data and analyses, have resulted in a research data deluge. Researchers must now carefully determine how to store, transfer and analyze "Big Data" in collaborative environments. This task is even more complicated when considering budget and locality constraints on data storage and access. In this paper we investigate the potential to construct a Social Content Delivery Network (S-CDN) based upon the social networks that exist between researchers. The S-CDN model builds upon the incentives of collaborative researchers within a given scientific community to address their data challenges collaboratively and in proven trusted settings. In this paper we present a prototype implementation of a S-CDN and investigate the performance of the data transfer mechanisms (using Glob us Online) and the potential cost advantages of this approach.
Kai Kugler 0002, Kyle Chard, Simon Caton, Omer F. Rana, Daniel S. Katz
e-Science4
2013 Editorial for special issue: Cloud computing and distributed data management in the AEC - Architecture, Engineering and Construction industry
Haijiang Li, Yacine Rezgui, Omer F. Rana
Adv. Eng. Informatics3
2013 Distributed computing practice for large-scale science and engineering applications
abstract
SUMMARY It is generally accepted that the ability to develop large‐scale distributed applications has lagged seriously behind other developments in cyberinfrastructure. In this paper, we provide insight into how such applications have been developed and an understanding of why developing applications for distributed infrastructure is hard. Our approach is unique in the sense that it is centered around half a dozen existing scientific applications; we posit that these scientific applications are representative of the characteristics, requirements, as well as the challenges of the bulk of current distributed applications on production cyberinfrastructure (such as the US TeraGrid). We provide a novel and comprehensive analysis of such distributed scientific applications. Specifically, we survey existing models and methods for large‐scale distributed applications and identify commonalities, recurring structures, patterns and abstractions. We find that there are many ad hoc solutions employed to develop and execute distributed applications, which result in a lack of generality and the inability of distributed applications to be extensible and independent of infrastructure details. In our analysis, we introduce the notion of application vectors: a novel way of understanding the structure of distributed applications. Important contributions of this paper include identifying patterns that are derived from a wide range of real distributed applications, as well as an integrated approach to analyzing applications, programming systems and patterns, resulting in the ability to provide a critical assessment of the current practice of developing, deploying and executing distributed applications. Gaps and omissions in the state of the art are identified, and directions for future research are outlined. Copyright © 2012 John Wiley & Sons, Ltd.
Shantenu Jha, Murray Cole, Daniel S. Katz, Manish Parashar, Omer F. Rana, Jon B. Weissman
Concurr. Comput. Pract. Exp.5
2013 Adapting scientific workflow structures using multi-objective optimization strategies
abstract
Scientific workflows have become the primary mechanism for conducting analyses on distributed computing infrastructures such as grids and clouds. In recent years, the focus of optimization within scientific workflows has primarily been on computational tasks and workflow makespan. However, as workflow-based analysis becomes ever more data intensive, data optimization is becoming a prime concern. Moreover, scientific workflows can scale along several dimensions: (i) number of computational tasks, (ii) heterogeneity of computational resources, and the (iii) size and type (static versus streamed) of data involved. Adapting workflow structure in response to these scalability challenges remains an important research objective. Understanding how a workflow graph can be restructured in an automated manner (through task merge, for instance), to address constraints of a particular execution environment is explored in this work, using a multi-objective evolutionary approach. Our approach attempts to adapt the workflow structure to achieve both compute and data optimization. The question of when to terminate the evolutionary search in order to conserve computations is tackled with a novel termination criterion. The results presented in this article demonstrate the feasibility of the termination criterion and demonstrate that significant optimization can be achieved with a multi-objective approach.
Irfan Habib, Ashiq Anjum, Richard McClatchey, Omer F. Rana
ACM Trans. Auton. Adapt. Syst.4
2012 Automating Data-Throttling Analysis for Data-Intensive Workflows
abstract
Data movement between tasks in scientific workflows has received limited attention compared to task execution. Often the staging of data between tasks is either assumed or the time delay in data transfer is considered to be negligible (compared to task execution). Where data consists of files, such file transfers are accomplished as fast as the network links allow, and once transferred, the files are buffered/stored at their destination. Where a task requires multiple files to execute (from different tasks), it must, however, remain idle until all files are available. Hence, network bandwidth and buffer/storage within a workflow are often not used effectively. We propose an automated workflow structural analysis method for Directed Acyclic Graphs (DAGs) which utilises information from previous workflow executions. The method obtains data-throttling values for the data transfer to enable network bandwidth and buffer/storage capacity to be managed more efficiently. We convert a DAG representation into a Petri net model and analyse the resulting graph using an iterative method to compute data-throttling values. Our approach is demonstrated using the Montage workflow.
Ricardo J. Rodríguez, Rafael Tolosana-Calasanz, Omer F. Rana
CCGRID3
2012 End-to-End QoS on Shared Clouds for Highly Dynamic, Large-Scale Sensing Data Streams
abstract
The increasing deployment of sensor network infrastructures has led to large volumes of data becoming available, leading to new challenges in storing, processing and transmitting such data. This is especially true when data from multiple sensors is pre-processed prior to delivery to users. Where such data is processed in-transit (i.e. from data capture to delivery to a user) over a shared distributed computing infrastructure, it is necessary to provide some Quality of Service (QoS) guarantees to each user. We propose an architecture for supporting QoS for multiple concurrent scientific workflow data streams being processed (prior to delivery to a user) over a shared infrastructure. We consider such an infrastructure to be composed of a number of nodes, each of which has multiple processing units and data buffers. We utilize the ``token bucket" model for regulating, on a per workflow stream basis, the data injection rate into such a node. We subsequently demonstrate how a streaming pipeline, with intermediate data size variation (inflation/deflation), can be supported and managed using a dynamic control strategy at each node. Such a strategy supports end-to-end QoS with variations in data size between the various nodes involved in the workflow enactment process.
Rafael Tolosana-Calasanz, José A. Bañares, CongDuc Pham, Omer F. Rana
CCGRID4
2012 Topic 1: Support Tools and Environments
Omer F. Rana, Marios D. Dikaiakos, Daniel S. Katz, Christine Morin
Euro-Par1
2012 Revenue Models for Streaming Applications over Shared Clouds
abstract
When multiple users execute their streaming applications over a shared Cloud infrastructure, the provider typically captures the Quality of Service (QoS) for each application at a Service Level Agreement (SLA). Such an SLA identifies the cost that a user must pay to achieve the required QoS, and a penalty that must be paid to the user in case the QoS cannot be met. Assuming the maximisation of the revenue as the provider's objective, then it must decide: (i) which user streams to accept for storage and analysis; (ii) how many (computational / storage) resources to allocate to each stream in order to improve overall revenue and minimise cost. In this paper, we analyse revenue models for in-transit streaming applications, executed over a shared Cloud infrastructure under the presence of faulty computational resources. We propose an architecture that features a token bucket process envelop to accept user streams; and a control loop to enable resource allocation, while minimising operational cost.
Rafael Tolosana-Calasanz, José A. Bañares, CongDuc Pham, Omer F. Rana
ISPA4
2012 Measuring the Effectiveness of Throttled Data Transfers on Data-Intensive Workflows
Ricardo J. Rodríguez, Rafael Tolosana-Calasanz, Omer F. Rana
KES-AMSTA3
2012 On-demand transmission model for remote visualization using image-based rendering
abstract
SUMMARY Interactive distributed visualization is an emerging technology with numerous applications. However, many of the present approaches to interactive distributed visualization have limited performance because they are based on the traditional polygonal processing graphics pipeline. In contrast, image‐based rendering uses multiple images of the scene instead of a three‐dimensional geometrical representation, and so has the key advantage that the final output is independent of the scene complexity and depends on the desired final image resolution. Furthermore, the discrete nature of the light field dataset maps well to a hybrid solution, which can overcome the identified drawbacks. In this paper, we propose an on‐demand solution for efficiently transmitting visualization data to remote users/clients. This is achieved through sending selected parts of the dataset based on the current client viewpoint, and is done instead of downloading a complete replica of the light field dataset to each client, or remotely sending a single rendered view back from a central server to the user each time the user updates their viewing parameters. The on‐demand approach shows stable performance as the number of clients increases because the load on the server and the network traffic are reduced. Furthermore, detailed performance studies show that the proposed scheme outperforms the current solution in terms of interactivity measured in frames per second. Copyright © 2012 John Wiley & Sons, Ltd.
Asma Al-Saidi, David W. Walker, Omer F. Rana
Concurr. Comput. Pract. Exp.3
2012 Towards autonomic management for Cloud services based upon volunteered resources
abstract
SUMMARY Many research institutions and Universities own computational capacity that is not effectively utilized, thereby providing an opportunity for such institutions to use such capacity to offer Cloud services (to both internal and external users). However, the unreliability and unpredictability of these resources mean that their use in the context of a Service Level Agreement (SLA) is high risk, leading to a reduction in reputation as well as economic penalties in case of SLA violation. We propose a methodology that addresses the issues of unreliability and unpredictability such that Cloud software services could be hosted upon volunteered resources. To enable the harnessing of these resources we rely on autonomic fault management techniques that allow such systems to independently adapt the resources they use based upon their perception of individual resource reliability. Using our approach we were able to scale out the backend infrastructure of the Cloud service elastically (min 30thinspaces per worker), opportunistically and autonomically. We address two key questions in this article: can a campus volunteer infrastructure be used in Cloud provisioning? What measures are necessary in order to ensure reliability at the resource level? Copyright © 2011 John Wiley & Sons, Ltd.
Simon Caton, Omer F. Rana
Concurr. Comput. Pract. Exp.2
2012 Preface of special issue on the economics of computing services
Jörn Altmann, Omer F. Rana, Rajkumar Buyya
Future Gener. Comput. Syst.2
2012 Network-aware summarisation for resource discovery in P2P-content networks
René Brunner, Agustín C. Caminero, Omer F. Rana, Felix Freitag, Leandro Navarro-Moldes
Future Gener. Comput. Syst.3
2012 Service level agreement as a complementary currency in peer-to-peer markets
Ioan Petri, Omer F. Rana, Gheorghe Cosmin Silaghi
Future Gener. Comput. Syst.2
2012 A GridWay-based autonomic network-aware metascheduler
Luis Tomás, Agustín C. Caminero, Omer F. Rana, Carmen Carrión 0001, María Blanca Caminero
Future Gener. Comput. Syst.3
2012 Enforcing QoS in scientific workflow systems enacted over Cloud infrastructures
Rafael Tolosana-Calasanz, José A. Bañares, CongDuc Pham, Omer F. Rana
J. Comput. Syst. Sci.4
2012 Social Cloud Computing: A Vision for Socially Motivated Resource Sharing
abstract
Online relationships in social networks are often based on real world relationships and can therefore be used to infer a level of trust between users. We propose leveraging these relationships to form a dynamic "Social Cloud,” thereby enabling users to share heterogeneous resources within the context of a social network. In addition, the inherent socially corrective mechanisms (incentives, disincentives) can be used to enable a cloud-based framework for long term sharing with lower privacy concerns and security overheads than are present in traditional cloud environments. Due to the unique nature of the Social Cloud, a social market place is proposed as a means of regulating sharing. The social market is novel, as it uses both social and economic protocols to facilitate trading. This paper defines Social Cloud computing, outlining various aspects of Social Clouds, and demonstrates the approach using a social storage cloud implementation in Facebook.
Kyle Chard, Kris Bubendorfer, Simon Caton, Omer F. Rana
IEEE Trans. Serv. Comput.4
2011 Evaluating trust in peer-to-peer service provider communities
abstract
The increasing availability of Internet services has stimulated the development of peer-to-peer markets. These electronic markets have the potential for improving the efficiency of trading by reducing search and transaction costs. They also allow buyers to choose the best possible service for every
Ioan Petri, Omer F. Rana, Yacine Rezgui, Gheorghe Cosmin Silaghi
CollaborateCom2
2011 Dynamic Workflow Adaptation over Adaptive Infrastructures
Rafael Tolosana-Calasanz, José A. Bañares, Omer F. Rana
KES-AMSTA3
2011 Summary Creation for Information Discovery in Distributed Systems
abstract
In current distributed systems, such as Grids, Clouds, or P2P systems, the amount of information to handle influences the way the system is managed. In P2P systems containing large quantities of data, or in Grid systems containing a large number of (often heterogeneous) resources, information about data or resources must be spread through the system in an efficient way in order to allow them to be found. An information discovery technique based on data summarization, via clustering, is presented. These summaries can be used to classify information to provide users with greater insight about documents or computing resources compared to raw data. Also, meta-schedulers or brokers would benefit from the proposed technique due to the fact that they would have to deal with less data from resources, thus aiding to the scalability of the system. An evaluation of the approach is subsequently provided to identify the impact of choosing particular parameters to be used as part of the summary.
Agustín C. Caminero, Eduardo Huedo, Omer F. Rana, Ignacio Martín Llorente, María Blanca Caminero, Carmen Carrión 0001
PDP3
2011 Autonomic streaming pipeline for scientific workflows
abstract
Abstract Interest in data streaming within scientific workflow has increased significantly over the recent years—mainly due to the emergence of data‐driven applications. Such applications can include data streaming from sensors and data coupling between scientific simulations. To support resource management to enact such streaming‐based workflow, autonomic computing techniques for transmission have been combined with in‐transit processing, so that data elements may be processed in advance, enroute, prior to arrival at the destination. We propose the integration of an autonomic data streaming service (ADSS) with in‐transit processing into a workflow specification. This integration may imply that the associated runtime resource allocation is dependent on environmental conditions and can change for different enactments of the same workflow. In our proposal, our workflow specifications are independent of the constraints imposed by the resource allocation. We express our solutions in terms of Reference nets. We also implement an ADSS utilizing a timed Reference net simulation for predicting future states of the ADSS. There are two advantages: the Reference net which implements the ADSS and the timed model are coincident, and second, token distribution obtained from the Petri net implementation can be utilized to better understand the demand for particular types of resources in the system. Copyright © 2011 John Wiley & Sons, Ltd.
Rafael Tolosana-Calasanz, José A. Bañares, Omer F. Rana
Concurr. Comput. Pract. Exp.3
2011 Network-aware heuristics for inter-domain meta-scheduling in Grids
Agustín C. Caminero, Omer F. Rana, María Blanca Caminero, Carmen Carrión 0001
J. Comput. Syst. Sci.2
2011 Special issue on "Theory and practice of high-performance computing, communications, and security"
Tai-Hoon Kim, Omer F. Rana, Juan Touriño, Isaac Woungang
J. Supercomput.2
2010 Social Cloud: Cloud Computing in Social Networks
abstract
With the increasingly ubiquitous nature of Social networks and Cloud computing, users are starting to explore new ways to interact with, and exploit these developing paradigms. Social networks are used to reflect real world relationships that allow users to share information and form connections between one another, essentially creating dynamic Virtual Organizations. We propose leveraging the pre-established trust formed through friend relationships within a Social network to form a dynamic “Social Cloud”, enabling friends to share resources within the context of a Social network. We believe that combining trust relationships with suitable incentive mechanisms (through financial payments or bartering) could provide much more sustainable resource sharing mechanisms. This paper outlines our vision of, and experiences with, creating a Social Storage Cloud, looking specifically at possible market mechanisms that could be used to create a dynamic Cloud infrastructure in a Social network environment.
Kyle Chard, Simon Caton, Omer F. Rana, Kris Bubendorfer
IEEE CLOUD3
2010 Distributed Systems and Algorithms
Omer F. Rana, Giandomenico Spezzano, Michael Gerndt, Daniel S. Katz
Euro-Par (1)1
2010 Clustering Client Honeypot Data to Support Malware Analysis
Yaser Alosefer, Omer F. Rana
KES (4)2
2010 Classification of Software Artifacts Based on Structural Information
Yuhanis Yusof, Omer F. Rana
KES (4)2
2010 Maximizing revenue in Grid markets using an economically enhanced resource manager
abstract
Abstract Traditional resource management has had as its main objective the optimization of throughput, based on parameters such as CPU, memory, and network bandwidth. With the appearance of Grid markets, new variables that determine economic expenditure, benefit and opportunity must be taken into account. The Self‐organizing ICT Resource Management (SORMA) project aims at allowing resource owners and consumers to exploit market mechanisms to sell and buy resources across the Grid. SORMA's motivation is to achieve efficient resource utilization by maximizing revenue for resource providers and minimizing the cost of resource consumption within a market environment. An overriding factor in Grid markets is the need to ensure that the desired quality of service levels meet the expectations of market participants. This paper explains the proposed use of an economically enhanced resource manager (EERM) for resource provisioning based on economic models. In particular, this paper describes techniques used by the EERM to support revenue maximization across multiple service level agreements and provides an application scenario to demonstrate its usefulness and effectiveness. Copyright © 2008 John Wiley & Sons, Ltd.
Mario Macías, Omer F. Rana, Garry Smith, Jordi Guitart, Jordi Torres
Concurr. Comput. Pract. Exp.2
2010 Adaptive exception handling for scientific workflows
abstract
Abstract Scientific workflow systems often operate in highly unreliable, heterogeneous and dynamic environments, and have accordingly incorporated different fault tolerance techniques. We propose an exception‐handling mechanism, based on techniques adopted in programming languages, for modifying at run‐time the structure of a workflow. In contrast to other proposals that achieve the required flexibility by means of the infrastructure, our proposal expresses the exception‐handling mechanism within the workflow language—primarily as two exception‐handling patterns that are exclusively based on the Reference Nets‐within‐Nets formalism (a specific type of Petri nets). When an exception is detected, a workflow in our approach can be re‐written (replaced), based on the particular failure condition that has been detected. This enables workflow users to have better control and understanding of the behaviour of their workflow without having to be aware of the underlying infrastructure. Copyright © 2009 John Wiley & Sons, Ltd.
Rafael Tolosana-Calasanz, José A. Bañares, Omer F. Rana, Pedro Álvarez 0001, Joaquín Ezpeleta, Andreas Hoheisel
Concurr. Comput. Pract. Exp.3
2010 Special section: Management and optimisation of P2P and Grid systems with network economics
Nick Antonopoulos, Omer F. Rana
Future Gener. Comput. Syst.2
2010 An uncoordinated asynchronous checkpointing model for hierarchical scientific workflows
Rafael Tolosana-Calasanz, José A. Bañares, Pedro Álvarez 0001, Joaquín Ezpeleta, Omer F. Rana
J. Comput. Syst. Sci.5
2009 The TRIACS analytical workflows platform for distributed clinical decision support
abstract
In this paper we discuss a flexible distributed workflow-based approach that enables researchers to study biomedical data for creating decision support pipelines. Specifically we describe the TRIACS platform, which has first been applied to supporting evidence-based decisions for optimum diabetic retinopathy screening intervals In the prioritization mechanism, pseudonymised case data is stratified for screening need by computation of outcome risk or by clustering of dasiaat-riskpsila cases with past cases of actual preventable outcomes. Workflows present a novel approach to this problem by providing an appropriate level of granularity for breaking the domain problem into a collection of reusable service-oriented components that can be applied in different ways. TRIACS is intended to make the creation of new application logic quicker and easier than bespoke development methods. Through the TRIACS workflow interface, modular code is portable and available to solve analogous domain problems including application to trial studies for mining and analysing clinical data.
Adina Riposan-Taylor, Ian J. Taylor, Omer F. Rana, David R. Owens, Edward C. Conley
CBMS3
2009 Distributed Collaborative Visualization Using Light Field Rendering
abstract
Interactive distributed visualization is an emerging technology with numerous applications. However, many of the present approaches of interactive distributed visualization are based on the traditional polygonal processing graphics pipeline. Our research is centred on investigating an alternative method using image-based rendering (IBR) which uses (multiple) images of the scene instead of a 3D geometrical representation. A key advantage to the use of IBR techniques is that the bandwidth required is independent of scene complexity and is therefore predictable given knowledge of the desired final image resolution. In this paper, we describe our IBR based interactive distributed visualization platform involving light field rendering and present results which indicate the scalability of our approach to accommodate multiple collaborative users. To our knowledge this is the first system to demonstrate deployment of interactive light field rendering to large numbers of distributed users.
Asma Al-Saidi, Nick J. Avis, Ian J. Grimstead, Omer F. Rana
CCGRID4
2009 Dynamic Service Reconfiguration and Enactment using an Open Matching Architecture
Sander van Splunter, Frances M. T. Brazier, Julian A. Padget, Omer F. Rana
ICAART4
2009 Performance evaluation of an autonomic network-aware metascheduler for Grids
abstract
Abstract Grid technologies have enabled the aggregation of geographically distributed resources in the context of a particular application. The network remains an important requirement for any Grid application, as entities involved in a Grid system (such as users, services, and data) need to communicate with each other over a network. The performance of the network must therefore be considered when carrying out tasks such as scheduling, migration or monitoring of jobs. Surprisingly, many existing quality of service efforts ignore the network and focus instead on processor workload and disk access time. Making use of the network in an efficient and fault‐tolerant manner is challenging. In a previous contribution, we proposed an autonomic network‐aware scheduling architecture that is capable of adapting its behavior to the current status of the environment. Now, we present a performance evaluation in which our proposal is compared with a conventional scheduling strategy. We present simulation results that show the benefits of our approach. Copyright © 2009 John Wiley & Sons, Ltd.
Agustín C. Caminero, Omer F. Rana, María Blanca Caminero, Carmen Carrión 0001
Concurr. Comput. Pract. Exp.2
2009 Distributed image processing over an adaptive Campus Grid
abstract
Abstract A system implemented in MATLAB is described, which may be deployed over a Campus Grid utilizing the Condor job management system. Our approach can re‐distribute jobs as node availability changes. The architecture of the system, its components, and their deployment across the Cardiff University Campus Grid (consisting of 2500 machines) are presented. Challenges in image processing applications that can be deployed over such infrastructure are presented, along with performance results that demonstrate the use of our system alongside a standard Condor deployment, demonstrating a significant increase in throughput using our approach ‡ . Copyright © 2008 John Wiley & Sons, Ltd.
Simon Caton, Omer F. Rana, Bruce G. Batchelor
Concurr. Comput. Pract. Exp.2
2009 Special Issue: Advanced Strategies in Grid Environments - Models and Techniques for Scheduling and Programming
abstract
BACKGROUNDThis special issue focuses on 'advanced scheduling strategies and Grid programming environments', and is based on selected papers from the Fifth International Workshop on Middleware for Grid Computing (MGC 2007).The authors were invited to provide extended versions of their original papers taking into account comments and suggestions raised during the peer review process and comments from the audience during the workshop.Relevant contributions have been provided by Jones, Caminero et al., Rosinha et al., Vidal et al., and Talukder et al., focusing on multisite scheduling techniques, an autonomic network-aware scheduling architecture, a work stealing programming environment, a semantic approach to improve resource scheduling, and a workflow execution planning approach using multi-objective differential evolution.Grid-enabled environments have been motivated by the work in service-oriented computing, and designed as a set of services that are provided by individuals or institutions.In this scenario, a level of automation is necessary to achieve scalability, requiring intelligent infrastructure management capabilities to be utilized alongside the existing approaches.Despite the emerging complexity and capability made available in Grid environments, it is worthy to mention that developing cross-grid applications that span different virtual organizations (VOs) has remained difficult.For instance, there is no co-scheduling of resources across more than one VO.Cloud computing concepts have recently emerged to address the limitations of Grid environments; a key focus is the use of virtualization technologies (at both platform and middleware levels) and replication through machine/service imaging.However the limitations initially found with Grid computing remain.It is useful to note that a computing Cloud can be designed on top of an existing Grid, as it also comprises heterogeneous components that have been integrated using a top-down approach.Similarly, Cloud computing systems are mostly created with the objective of providing a determined set of capabilities to a user; thus, the interface is an important part of the design.Scheduling challenges for enabling the deployment and execution of Grid-based parallel applications have been addressed by the selected papers in this special issue, with the development of models and techniques evaluated accordingly by prototyping and simulations.
Bruno Schulze, Omer F. Rana, José Neuman de Souza
Concurr. Comput. Pract. Exp.2
2008 Using Dynamic Condor-Based Services for Classifying Schizophrenia in Diffusion Tensor Images
abstract
Diffusion tensor imaging (DTI) provides insight into the white matter of the human brain, which is affected by schizophrenia. By comparing a patient group to a control group, the DTI-images are on average expected to be different for white matter regions. Principal component analysis (PCA) and linear discriminant analysis (LDA) are used to classify the groups. In this work, the number of principal components is optimised for obtaining the minimal classification error. A robust estimate of this error is computed in a cross-validation framework, using different compositions of the data into a training and a testing set Previously, sequential runs were performed in MATLAB, resulting in long execution times. In this paper we describe an experiment where this application was run on a grid with minimal modifications and user effort. We have adopted a service-based approach that autonomously launches image analysis services onto a campus-wide Condor pool comprising of volunteer resources. This allows high throughput analysis of our data in a dynamic resource pool. The challenge in adopting such an approach comes from the nature of the resources, which change randomly with time and thus require fault tolerance. Through this approach we have reduced the computation time of each dataset from 90 minutes to less than 10. A minimal classification error of 22% was obtained, using 15 principal components.
Simon Caton, Matthan W. A. Caan, Sílvia Delgado Olabarriaga, Omer F. Rana, Bruce G. Batchelor
CCGRID4
2008 Optimizing Decentralized Grid Markets through Group Selection
abstract
Automatic coordination mechanisms for the Grid are required due to the increasing complexity exhibited in large scale distributed systems. Decentralized economic models are being considered as scalable coordination mechanisms for the management of service allocations to clients. However, decentralization incorporates further dynamicity and unpredictability into the system. Introducing higher levels of adaptation and learning in the coordination protocols helps cope with complexity. We provide a solution based on a self-organized, emergent mechanism evolving Grid Market participants through a group selection process. Dynamic congregations organize agents into optimized market segments, maximizing utility and thereby improving system-wide performance. We provide a system model and evaluation by simulation of the group selection mechanism. We further provide a prototype showing practical feasibility of the approach.
Isaac Chao, Oscar Ardaiz-Villanueva, Ramon Sangüesa, Liviu Joita, Omer F. Rana
CISIS5
2008 Topic 1: Support Tools and Environments
Marios D. Dikaiakos, Omer F. Rana, Shmuel Ur, João Lourenço
Euro-Par2
2008 Special Issue: First International Workshop on Workflow Systems in Grid Environments (WSGE2006)
abstract
Editorial of the special issue of Concurrency and Computation: Practice and Experience, which contains selected high-quality papers from the First International Workshop on Workflow Systems in Grid Environments (WSGE2006), held in the Hunan University of Science and Technology in China on 23 October 2006.
Jinjun Chen, Omer F. Rana
Concurr. Comput. Pract. Exp.2
2008 Performance evaluation of semantic registries: OWLJessKB and instanceStore
Simone A. Ludwig, Omer F. Rana
Serv. Oriented Comput. Appl.2
2008 Grid Service Discovery with Rough Sets
abstract
The computational grid is rapidly evolving into a service-oriented computing infrastructure that facilitates resource sharing and large-scale problem solving over the Internet. Service discovery becomes an issue of vital importance in utilizing grid facilities. This paper presents ROSSE, a Rough sets-based search engine for grid service discovery. Building on the Rough sets theory, ROSSE is novel in its capability to deal with the uncertainty of properties when matching services. In this way, ROSSE can discover the services that are most relevant to a service query from a functional point of view. Since functionally matched services may have distinct nonfunctional properties related to the quality of service (QoS), ROSSE introduces a QoS model to further filter matched services with their QoS values to maximize user satisfaction in service discovery. ROSSE is evaluated from the aspects of accuracy and efficiency in discovery of computing services.
Maozhen Li 0001, Bin Yu 0005, Omer F. Rana, Zidong Wang 0001
IEEE Trans. Knowl. Data Eng.3
2007 Dynamic Condor-based Services for Distributed Image Analysis
abstract
Interactive image processing is an important requirement in many industrial applications, such as the inspection of industrial parts within a manufacturing environment, or the processing of images from surveillance cameras. Being able to achieve this quickly and accurately is often essential for the success of such industrial applications. A service-based approach that autonomously launches Image Analysis Services (accessible through a Central Service Manager) onto spare network resources through a Condor system is presented. This allows high throughput analysis of these images in a dynamic resource pool. The Central Service Manager reacts to new tasks submitted to the Image Analysis Services and is able to add new service instances to manage these tasks dynamically. Each service instance here corresponds to a computational resource that is able to execute image processing algorithms. New service instances may be requested by the Central Service Manager from the Condor system, based on the number of tasks that need to be processed. This enables entire image repositories to be acted upon interactively and in parallel, as opposed to the analysis of single images individually. The approach is demonstrated through a campus-wide test bed utilising a Condor system with 90 machines.
Simon Caton, Omer F. Rana, Bruce G. Batchelor
CCGRID2
2007 Matchmaking Support for Dynamic Workflow Composition
abstract
Service description and discovery offer complementary challenges, but in both cases, the problem is finding the right trade-off between accuracy and generality that will result in a positive service identification. Discovery systems have historically tended to focus on domain-specific techniques using single sources of knowledge to help classify queries against services, making both maintenance and extension difficult. The primary contribution of this paper is the presentation of a generic brokerage framework based on the use of plug-in components, that are themselves Web services. The framework has been developed in the context of the KNOOGLE project, where the focus has been on demonstrating support for (i) the discovery of grid services for the GridSAM job submission system and (ii) integration with the Taverna workflow enactment system. However, the broker itself is domain independent and it is the multiple user-specified matchmaker plug-ins that act as sources of domain-specific knowledge. The broker collects the results of the matchmakers' comparison of the query and service and then applies a user-specified selection policy to determine the final choice of service. Thus a range of comprehensive packaging of brokerage functionality becomes possible through the use of supplied and user-defined matchers and supplied or user-defined selection policies.
Neil Chapman, Simone A. Ludwig, William Naylor, Julian A. Padget, Omer F. Rana
eScience5
2007 Automatic Assertion of Actor State in Service Oriented Architectures
abstract
Documentation of provenance, the process which was taken to create a particular data item, is critical within service oriented architectures where loose coupling between services and clients (actors) is employed. Such a process may involve interaction between multiple services, with each service being managed by a particular actor. The recording of actor state may provide critical contextual information regarding the state of a particular actor at a point during a client-service interaction, but is however typically left undocumented. We discuss the issues that are encountered with the automatic documentation of actor state and the types of resources which may be used to provide it. An architecture which allows monitoring tools to be related to user needs is presented and evaluated against a number of scenarios in a Web Services environment.
Ian Wootten, Shrija Rajbhandari, Omer F. Rana
ICWS3
2007 Integration of Descriptors for Software Component Retrieval
Yuhanis Yusof, Omer F. Rana
KSEM2
2007 A workflow portal supporting multi-language interoperation and optimization
abstract
Abstract In this paper we present a workflow portal for Grid applications, which supports different workflow languages and workflow optimization. We present an XSLT converter that converts from one workflow language to another and enables the interoperation between different workflow languages. We discuss strategies for choosing the optimal service from several semantically equivalent Web services in a Grid application. The dynamic selection of Web services involves discovering a set of semantically equivalent services by filtering the available services based on metadata, and selecting an optimal service based on real‐time data and/or historical data recorded during prior executions. Finally, we describe the framework and implementation of the workflow portal which aggregates different components of the project using Java portlets. Copyright © 2007 John Wiley & Sons, Ltd.
Lican Huang, Asif Akram, Rob Allan, David W. Walker, Omer F. Rana, Coral Walker
Concurr. Comput. Pract. Exp.5
2007 A catallactic market for data mining services
Liviu Joita, Omer F. Rana, Felix Freitag, Isaac Chao, Pablo Chacin, Leandro Navarro-Moldes, Oscar Ardaiz-Villanueva
Future Gener. Comput. Syst.2
2006 Deriving Ratings Through Social Network Structures
abstract
A review of existing approaches to recommendation in e-commerce systems is provided. A recommendation system is primarily used to identify services which may be of interest to a user based on a similarity in purchasing (or browsing) patterns with another user, or to filter services that have been returned as a result of a search. Existing systems primarily make use of collaborative filtering approaches or a semantic-annotation approach which tries to find similarity by matching on the definition of a service. However, such systems suffer from "sparseness" of ratings - as it is difficult to find enough ratings to help make a recommendation for a user. We therefore propose the use of a social network as the basis for defining how ratings can be aggregated, based on the structure of the network. We also suggest the use of product categories as the basis for aggregating ratings - and define this as a "context" in which a particular service is used. A model for a recommendation system that combines context-based rating with the structure of a social network has been suggested, along with an architecture for a system that implements the model.
Hameeda Alshabib, Omer F. Rana, Ali Shaikh Ali
ARES2
2006 Dynamic Workflow Management Using Performance Data
abstract
An approach to dynamic workflow management and optimisation using performance data is presented. We discuss strategies for choosing an optimal service (based on user specified criteria) from several semantically equivalent Web Services. Such an approach involves finding "similar" services, by first pruning the set of discovered services based on service metadata, and subsequently selecting an optimal service based on data recorded during prior executions of a service and/or current machine loads. We describe the current implementation of the system, and demonstrate this by a BLAST (used in BioInformatics for Protein-alignment) example by using the Ganglia monitoring tool to get performance data.
Lican Huang, David W. Walker, Omer F. Rana, Coral Walker
CCGRID3
2006 Evaluating Provenance-based Trust for Scientific Workflows
abstract
Provenance is the documentation concerning the origin of a result generated by a process, and provides explanations about who, how, what resources were used in a process, and the processing steps that occurred to produce the result. Such provenance information is important to improve a scientist’s ability to judge and place certain amount of trust on the generated data. We illustrate how provenance information associated with a workflow can be used to evaluate trust. This work is based on several use cases from a Bio-Diversity application. We also propose a simple architecture to illustrate our trust framework.
Shrija Rajbhandari, Ian Wootten, Ali Shaikh Ali, Omer F. Rana
CCGRID4
2006 Actor Provenance Capture With Ganglia
abstract
Provenance is generally defined as the documentation of a process that leads to some result, and has long been recognised as being fundamental to the development of problem solving mechanisms within Grid environments. The knowledge of how a particular result has been derived is just as important as the result itself within many e-Science experiments. We present the concept of "actor" provenance, which provides detailed information concerning the state of an actor at a particular time. We also demonstrate how actor provenance differs from "interaction" provenance between actors. We describe how actor provenance may be represented and recorded using monitoring tools such as ganglia. This is explained using a number of use cases in a Bio-Diversity application.
Ian Wootten, Shrija Rajbhandari, Omer F. Rana, Jaspreet Singh Pahwa
CCGRID3
2006 Choosing a Load Balancing Scheme for Agent-Based Digital Libraries
Christos Georgousopoulos, Omer F. Rana
ISPA2
2006 Navigating Provenance Information for Distributed Healthcare Management
abstract
Provenance information provides a useful basis to verify whether a particular application behavior has been adhered to. This is particularly useful to evaluate the basis for a particular outcome, as a result of a process, and to verify if the process involved in making the decision conforms to some pre-defined set of rules. This is significant in a healthcare scenario, where it is necessary to demonstrate that patient data has been processed in a particular way. Understanding how provenance information may be recorded, stored, and subsequently analyzed by a decision maker is therefore significant in a service oriented architecture, which involves the use of third party services over which the decision maker does not have control. The aggregation of data from multiple sources of patient information plays an important part in subsequent treatments that are proposed for a patient. A tool to navigate through and analyze such provenance information is proposed, based on the use of a portal framework that allows different views on provenance information to co-exist. The portal enables users to add custom portlets enabling application specific views that would facilitate particular decision making
Vikas Deora, Arnaud Contes, Omer F. Rana, Shrija Rajbhandari, Ian Wootten, Tamás Kifor, László Z. Varga
Web Intelligence3
2006 Distributed Storage of High-Volume Environmental Simulation Data: Mantle Modelling
abstract
A feasibility study of a peer-to-peer distributed storage system for the archiving of large datasets produced by the Earth mantle modelling code TERRA is presented. The manner in which the nature of such data affects the indexing, duplication and performance requirements of such a system is analysed and a data-oriented overlay network to improve efficiency is proposed. Duplication methods are analysed and test bed performance measured
Martin Wolstencroft, Omer F. Rana, J. Huw Davies
Web Intelligence2
2006 Reputation-based semantic service discovery
abstract
Abstract An important component of Semantic Grid services is the support for dynamic service discovery. Dynamic service discovery requires the provision of rich and flexible metadata that is not supported by current registry services such as UDDI. We present a framework to facilitate reputation‐based service selection in Semantic Grids. Our framework has two key features that distinguish it from other work in this area. First, we propose a dynamic, adaptive, and highly fault‐tolerant reputation‐aware service discovery algorithm. Second, we present a service‐oriented distributed reputation assessment algorithm. In this paper, we describe the main components of our framework and report on our experience of developing the prototype. Copyright © 2005 John Wiley & Sons, Ltd.
Ali Shaikh Ali, Shalil Majithia, Omer F. Rana, David W. Walker
Concurr. Comput. Pract. Exp.3
2006 Matchmaking Framework for Mathematical Web Services
Simone A. Ludwig, Omer F. Rana, Julian A. Padget, William Naylor
J. Grid Comput.2
2006 An abstraction model for a Grid execution framework
Kaizar Amin, Gregor von Laszewski, Mihael Hategan, Rashid J. Al-Ali, Omer F. Rana, David W. Walker
J. Syst. Archit.5
2005 Message from the programme committee chair
abstract
Presents the welcome message from the conference proceedings.
Omer F. Rana
CCGRID1
2005 Agent Based Computational Grids: Research Issues and Challenges
Omer F. Rana
Euro-Par1
2005 Future trends in distributed applications and problem-solving environments
José C. Cunha, Omer F. Rana, Pedro D. Medeiros
Future Gener. Comput. Syst.2
2004 QoS support for high-performance scientific Grid applications
abstract
The Grid approach provides the ability to access and use distributed resources as part of virtual organizations. The emerging Grid infrastructure gives rise to a class of scientific applications and services to support collaborative and distributed resource-sharing requirements as part of teleimmersion, visualization, and simulation services. Because such applications operate in a collaborative mode, data must be stored and delivered in timely manner to meet deadlines. Hence, this class of applications has stringent real-time constraints and quality-of-service (QoS) requirements. A QoS management approach is required to orchestrate and guarantee the interaction between such applications and services. In this paper we discuss the design and prototype implementation of a QoS system and show how we enable Grid applications to become QoS compliant. We validate this approach through a case study of nanomaterials. Our approach enhances the current Open Grid Services Architecture. We demonstrate the usefulness of the approach on a nanomaterials application.
Rashid J. Al-Ali, Gregor von Laszewski, Kaizar Amin, Mihael Hategan, Omer F. Rana, David W. Walker, Nestor J. Zaluzec
CCGRID5
2004 Pattern/Operator Based Problem Solving Environments
Maria Cecilia Gomes, Omer F. Rana, José C. Cunha
Euro-Par2
2004 A Double Auction Economic Model for Grid Services
Liviu Joita, Omer F. Rana, W. Alex Gray, John C. Miles
Euro-Par2
2004 An approach for quality of service adaptation in service-oriented Grids
abstract
Abstract Some applications utilizing Grid computing infrastructure require the simultaneous allocation of resources, such as compute servers, networks, memory, disk storage and other specialized resources. Collaborative working and visualization is one example of such applications. In this context, quality of service (QoS) is related to Grid services, and not just to the network connecting these services. With the emerging interest in service‐oriented Grids, resources may be advertised and traded as services based on a service level agreement (SLA). Such a SLA must include both general and technical specifications, including pricing policy and properties of the resources required to execute the service, to ensure QoS requirements are satisfied. An approach for QoS adaptation is presented to enable the dynamic adjustment of behavior of an application based on changes in the pre‐defined SLA. The approach is particularly useful if workload or network traffic changes in unpredictable ways during an active session. Copyright © 2004 John Wiley & Sons, Ltd.
Rashid J. Al-Ali, Abdelhakim Hafid, Omer F. Rana, David W. Walker
Concurr. Pract. Exp.3
2004 SGrid: a service-oriented model for the Semantic Grid
Maozhen Li 0001, P. van Santen, David W. Walker, Omer F. Rana, Mark A. Baker
Future Gener. Comput. Syst.4
2004 Analysis and Provision of QoS for Distributed Grid Applications
Rashid J. Al-Ali, Kaizar Amin, Gregor von Laszewski, Omer F. Rana, David W. Walker, Mihael Hategan, Nestor J. Zaluzec
J. Grid Comput.4
2004 Migrating legacy codes to distributed computing environments: a CORBA approach
Maozhen Li 0001, David W. Walker, Omer F. Rana, Coral Walker
Inf. Softw. Technol.3
2004 Agent-based computer vision
Paul L. Rosin, Omer F. Rana
Pattern Recognit.2
2003 PortalLab: A Web Services Toolkit for Building Semantic Grid Portals
abstract
Grid is computer-based infrastructure that provides dependable, consistent, pervasive access to distributed resources. Built on top of a Grid, a Semantic Grid is a service-oriented infrastructure that provides a range of computation, information and knowledge services. A purpose of a Grid portal is to provide easy and seamless access to Grid heterogeneous resources and services through a Web-based user interface. This paper presents PortalLab, a Web Services oriented toolkit for designing, integrating and building Semantic Grid portals. Portals built from PortalLab are composed from a collection of reusable Web Services oriented portlets that are themselves semantic Grid services. Each portlet has a WSDL interface and a semantic registry defined in a domain ontology repository. The use of software agents assists end users in formulating domain problems, searching possible solutions (solvers) and submitting user tasks to the Grid. Multiple agents work in a peer-to-peer environment to allow users to access federated Grid services across different domains to improve fault tolerance and quality of service in user job submission and execution on the Grid. Since portlets are context independent, a PortalLab portal provides the ability to interoperate with different Grid systems at a portal level.
Maozhen Li 0001, P. van Santen, David W. Walker, Omer F. Rana, Mark A. Baker
CCGRID4
2003 Topic Introduction
Luc Bougé, Franck Cappello, Omer F. Rana, Bernard Traversat
Euro-Par3
2003 Supporting Peer-2-Peer Interactions in the Consumer Grid
abstract
A "Consumer Grid" provides the individual-based counterpart to the organisation-based computational grid. We describe a peer-to-peer system for utilising computational resources on the Grid - extending existing work undertaken in systems such as Entropia and SETI@home. The potential of such a distributed computing resource has been in some ways demonstrated recently by the SETI@home project, having used over 650,000 years of CPU time at the time of writing. A user develops applications within such an environment using a visual workflow system called "Triana" - which automatically generates suitable code for distribution, and can support the user in making placement decisions for their modules. Triana will also be deployed as the workflow enactment engine along with the grid application toolkit (GAT) within the European GridLab project.
Ian J. Taylor, Omer F. Rana, Roger Philp, Ian Wang, Matthew S. Shields
HIPS2
2003 Organizing Service-Oriented Peer Collaborations
Asif Akram, Omer F. Rana
ICSOC2
2003 Structuring Peer-2-Peer Communities
abstract
Locating suitable resources within a peer-2-peer (P2P) system is a computationally intensive process, with no guarantee of quality and suitability of the discovered resources. An alternative approach is to categorise peers based on the services they provide, leading to the interaction between peers with common goals to form societies/communities. Organization of peers in different communities is suggested to be useful for efficient resource discovery. We analyse the types of communities that may be useful, and how they may be structured.
Asif Akram, Omer F. Rana
Peer-to-Peer Computing2
2003 Object-oriented distributed computing based on remote class reference
abstract
Abstract Java RMI, Jini and CORBA provide effective mechanisms for implementing a distributed computing system. Recently many numeral libraries have been developed that take advantage of Java as an object‐oriented and portable language. The widely‐used client‐server method limits the extent to which the benefits of the object‐oriented approach can be exploited because of the difficulties arising when a remote object is the argument or return value of a remote or local method. In this paper this problem is solved by introducing a data object that stores the data structure of the remote object and related access methods. By using this data object, the client can easily instantiate a remote object, and use it as the argument or return value of either a local or remote method. Copyright © 2003 John Wiley & Sons, Ltd.
Coral Walker, David W. Walker, Omer F. Rana
Concurr. Comput. Pract. Exp.3
2003 Triana Applications within Grid Computing and Peer to Peer Environments
Ian J. Taylor, Matthew S. Shields, Ian Wang, Omer F. Rana
J. Grid Comput.4
2003 Engineering high-performance legacy codes as CORBA components for problem-solving environments
Maozhen Li 0001, David W. Walker, Omer F. Rana, Coral Walker, P. T. Williams, R. C. Ward
J. Parallel Distributed Comput.3
2002 Coordinating Learning Agents via Utility Assignment
Steven J. Lynden, Omer F. Rana
IDEAL2
2002 Coordinated Learning to Support Resource Management in Computational Grids
abstract
Managing resources in large scale distributed systems is an important concern for both peer-2-peer and computational grid systems, and is a complex and time sensitive process. Although existing peer-2-peer systems are divided into those that support computation (CPU) sharing or data sharing, users in a computational grid generally need to share both. Identifying which resources to select is important to guarantee reasonable execution time and cost to a given user or group of users. We provide first a description of a framework to support learning in the context a community of interacting peers, and show how this can be used to support resource sharing in computational grids.
Steven J. Lynden, Omer F. Rana
Peer-to-Peer Computing2
2002 Agent based data management in digital libraries
Omer F. Rana, David W. Walker, Christos Georgousopoulos, Giovanni Aloisio, Roy Williams
Parallel Comput.2
2001 Topic 17: Metacomputing and Grid Computing
Alexander Reinefeld, Omer F. Rana, Jarek Nabrzyski, David W. Walker
Euro-Par2
2001 Resource Discovery for Dynamic Clusters in Computational Grids
abstract
We describe a de-centralised approach to resource management and discovery, based on a community of interacting software agents. Each agent either represents a user application, a resource, or a MatchMaking service. The proposed approach can support dynamic registration of resources and user tasks, facilitating the establishment of dynamic clusters. Resource capability and task requirements are described using an object based data model, enabling new types of devices or new features in existing devices to be identified. A comparison with the Discovery and LookUp services in Jini and TSpaces is also provided.
Omer F. Rana, Daniel Bunford-Jones, Kenneth A. Hawick, David W. Walker, Matthew Addis, Mike Surridge
IPDPS1
2001 Special Issue: High Performance Agent Systems
abstract
High Performance
Omer F. Rana, David Kotz
Concurr. Comput. Pract. Exp.1
2001 Wrapping MPI-based legacy codes as Java/CORBA components
Maozhen Li 0001, Omer F. Rana, David W. Walker
Future Gener. Comput. Syst.2
2000 Emergent coordination for distributed information management
abstract
With an increase in the complexity of information systems, decentralised management techniques are becoming important. Such information systems will generally involve various computational and data nodes, managed by various organisations with differing policies and needs. Furthermore, in the absence of a centralised manager, co-ordination and interaction policies between such nodes become crucial. Aspects of co-ordination can range from data formats describing shared information to convergence time to a desired goal. Each entity within such a networked environment operates with incomplete knowledge about the state of others, and makes use of an asynchronous mode of operation. An overview of techniques for supporting coordination in such an environment is first provided, followed by a proposed technique that makes use of Q-learning and data sharing between a collection of nodes, each of which is modelled as a neural network. We propose a decentralised technique for managing a cluster of nodes, where emergence is used to converge towards a 'useful' result.
Steven J. Lynden, Omer F. Rana, Steve Margetts, Antonia J. Jones
CEC2
2000 Tutorial F: Implementing Agent Applications in Java: Using Mobile and Intelligent Agents
Omer F. Rana
CLUSTER1
2000 Agent based Resource Discovery for Dynamic Clusters
Omer F. Rana, Daniel Bunford-Jones, David W. Walker, Matthew Addis, Mike Surridge, Kenneth A. Hawick
CLUSTER1
2000 Automating Performance Analysis from UML Design Patterns (Research Note)
Omer F. Rana, Dave Jennings
Euro-Par1
2000 Implementing Problem Solving Environments for Computational Science (Research Note)
Omer F. Rana, Maozhen Li 0001, Matthew S. Shields, David W. Walker, David Golby
Euro-Par1
2000 PaDDMAS: Parallel and Distributed Data Mining Application Suite
abstract
Discovering complex associations, anomalies and patterns in distributed data sets is gaining popularity in a range of scientific, medical and business applications. Various algorithms are employed to perform data analysis within a domain, and range from statistical to machine learning and AI based techniques. Several issues need to be addressed however to scale such approaches to large data sets, particularly when these are applied to data distributed at various sites. As new analysis techniques are identified, the core tool set must enable easy integration of such analytical components. Similarly, results from an analysis engines must be sharable, to enable storage, visualisation or further analysis of results. We describe the architecture of PaDDMAS, a component based system for developing distributed data mining applications. PaDDMAS provides a tool set for combining pre-developed or custom components using a dataflow approach, with components performing analysis, data extraction or data management and translation. Each component is wrapped as a Java/CORBA object, and has an interface defined in XML. Components can be serial or parallel objects, and may be binary or contain a more complex internal structure. We demonstrate a prototype using a neural network analysis algorithm.
Omer F. Rana, David W. Walker, Maozhen Li 0001, Steven J. Lynden, Mike Ward
IPDPS1
2000 A Wrapper Generator for Wrapping High Performance Legacy Codes as Java/CORBA Components
abstract
This paper describes a Wrapper Generator for wrapping high performance legacy codes as Java/CORBA components for use in a distributed component-based problem- solving environment. Using the Wrapper Generator we ave automatically wrapped an MPI-based legacycode as a single CORBA object, and implemented a problem- solving environment for molecular dynamic simulations. Performance comparisons between runs of the CORBA object and the original legacy code on a cluster of workstations and on a parallel computer are also presented.
Maozhen Li 0001, Omer F. Rana, Matthew S. Shields, David W. Walker
SC2
2000 A Java/CORBA-based visual program composition environment for PSEs
abstract
A problem solving environment (PSE) is a complete, integrated computing environment for composing, compiling and running applications in a specific problem area or domain. A visual programming composition environment (VPCE) is described, which serves as a user interface for a PSE, and uses Java and CORBA to provide a framework of tools to enable the construction of scientific applications from components. The VPCE consists of a component repository, from which the user can select off-the-shelf or in-house components, a graphical composition area on which components can be combined, various tools that facilitate the configuration of components, the integration of legacy codes into components and the design and building of new components. The VPCE produces output using dataflow techniques in the form of a task graph, annotated with a performance model plus constraints for each component, expressed in XML. In addition, the VPCE supports a domain specific expert system based on JESS (Ernest Friedman-Hill, JESS: The Java Expert System Shell. See web site at: http://herzberg.ca.sandia.gov/jess/, 1999) to guide the user in component selection and to perform integrity checking. Copyright © 2000 John Wiley & Sons, Ltd.
Matthew S. Shields, Omer F. Rana, David W. Walker, Maozhen Li 0001, David Golby
Concurr. Pract. Exp.2
2000 The software architecture of a distributed problem-solving environment
abstract
This paper describes the functionality and software architecture of a generic problem-solving environment (PSE) for collaborative computational science and engineering. A PSE is designed to provide transparent access to heterogeneous distributed computing resources, and is intended to enhance research productivity by making it easier to construct, run, and analyze the results of computer simulations. Although implementation details are not discussed in depth, the role of software technologies such as CORBA, Java, and XML is outlined. An XML-based component model is presented. The main features of a Visual Component Composition Environment for software development, and an Intelligent Resource Management System for scheduling components, are described. Some prototype implementations of PSE applications are also presented. Copyright © 2000 John Wiley & Sons, Ltd.
David W. Walker, Maozhen Li 0001, Omer F. Rana, Matthew S. Shields, Coral Walker
Concurr. Pract. Exp.3
2000 Automating Parallel Implementation of Neural Learning Algorithms
abstract
Neural learning algorithms generally involve a number of identical processing units, which are fully or partially connected, and involve an update function, such as a ramp, a sigmoid or a Gaussian function for instance. Some variations also exist, where units can be heterogeneous, or where an alternative update technique is employed, such as a pulse stream generator. Associated with connections are numerical values that must be adjusted using a learning rule, and and dictated by parameters that are learning rule specific, such as momentum, a learning rate, a temperature, amongst others. Usually, neural learning algorithms involve local updates, and a global interaction between units is often discouraged, except in instances where units are fully connected, or involve synchronous updates. In all of these instances, concurrency within a neural algorithm cannot be fully exploited without a suitable implementation strategy. A design scheme is described for translating a neural learning algorithm from inception to implementation on a parallel machine using PVM or MPI libraries, or onto programmable logic such as FPGAs. A designer must first describe the algorithm using a specialised Neural Language, from which a Petri net (PN) model is constructed automatically for verification, and building a performance model. The PN model can be used to study issues such as synchronisation points, resource sharing and concurrency within a learning rule. Specialised constructs are provided to enable a designer to express various aspects of a learning rule, such as the number and connectivity of neural nodes, the interconnection strategies, and information flows required by the learning algorithm. A scheduling and mapping strategy is then used to translate this PN model onto a multiprocessor template. We demonstrate our technique using a Kohonen and backpropagation learning rules, implemented on a loosely coupled workstation cluster, and a dedicated parallel machine, with PVM libraries.
Omer F. Rana
Int. J. Neural Syst.1
1999 A Design and Management Framework for Mobile Agent Systems
abstract
A framework for developing 'mobile agent' systems is described based on Petri net models of design patterns in the Aglets workbench. The 'Meeting' and 'Interaction' design patterns are investigated. The models can be automatically configured from Java source programs, and can account for stochastic parameters such as port contention at hosts, agent registration, background workloads and scalability issues, such as bounds on total number of agents that may be supported at a host. The models are verified by comparing the performance of an application in Aglets with Petri net simulations on the GreatSPN simulator.
Omer F. Rana
MASCOTS1