EDBT 2026 Demo / reviewers in the wild / expert
Blesson Varghese
dblp:32/7786
· DBLP profile ↗
55ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0001-8392-832XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 4Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mosaic: Composite projection pruning for resource-efficient LLMsabstractExtensive compute and memory requirements limit the deployment of large language models (LLMs) on any hardware. Compression methods, such as pruning, can reduce model size, which in turn reduces resource requirements. State-of-the-art pruning is based on coarse-grained methods. They are time-consuming and inherently remove critical model parameters, adversely impacting the quality of the pruned model. This paper introduces projection pruning, a novel fine-grained method for pruning LLMs. In addition, LLM projection pruning is enhanced by a new approach we refer to as composite projection pruning — the synergistic combination of unstructured pruning that retains accuracy and structured pruning that reduces model size. We develop Mosaic , a novel system to create and deploy pruned LLMs using composite projection pruning. Mosaic is evaluated using a range of performance and quality metrics on multiple hardware platforms, LLMs, and datasets. Mosaic is 7.19 × faster in producing models than existing approaches. Mosaic models achieve up to 84.2% lower perplexity and 31.4% higher accuracy than models obtained from coarse-grained pruning. Up to 67% faster inference and 68% lower GPU memory use is noted for Mosaic models. • LLMs are difficult to optimize without hardware and software accelerators. • Pruning makes LLMs smaller and faster for Edge devices. • Existing pruning methods rely on accelerators and reduce LLM quality. • Mosaic uses composite pruning to make LLMs resource-efficient without accelerators. • Mosaic has better accuracy and perplexity than existing pruning methods. Bailey J. Eccles, Leon Wong, Blesson Varghese |
Future Gener. Comput. Syst. | 3 |
| 2026 | FedFreeze: A dual-phase layer freezing framework for federated learningabstractRunning Federated Learning (FL) on resource-constrained devices is challenging due to the resources required for training. Layer freezing has been proposed to reduce the computational costs and thus accelerate training. However, we identify that existing layer freezing approaches either learn quickly or learn effectively, but do not balance them. Specifically, aggressive early-stage layer freezing (e.g., AutoFreeze) accelerates training but achieves a lower final accuracy. On the other hand, accuracy-guaranteed layer freezing (e.g., ALF) obtains higher final accuracy but with marginal training time improvement. This article proposes FedFreeze – a dual-phase layer freezing federated learning framework that, for the first time, combines early-stage and accuracy-guaranteed layer freezing into a unified mechanism. FedFreeze designs a novel regularization-based layer freezing strategy on the device to apply early-stage layer freezing even during the initial stages for improving training speedup. In addition, FedFreeze develops a convergence-based layer freezing strategy to achieve a high final accuracy. Experimental results show that the proposed FedFreeze framework achieves up to 1.3 × training speedup while limiting the accuracy drop to no more than 1.67% compared to vanilla FL. In contrast to state-of-the-art early-stage and accuracy-guaranteed layer freezing methods, FedFreeze consistently strikes a better balance between efficiency and accuracy across a wide range of settings, including different hardware platforms (Raspberry Pi and Jetson Nano clusters), datasets (FMNIST, CIFAR-10, CIFAR-100), model architectures (AlexNet, VGG11, ResNet12), and initialization strategies. The results demonstrate that FedFreeze outperforms state-of-the-art layer freezing techniques in accelerating FL training on resource-constrained devices, without incurring significant accuracy loss. Di Wu 0065, Leon Wong, Blesson Varghese |
Future Gener. Comput. Syst. | 3 |
| 2026 | FedOptima: Optimizing resource utilization in federated learning
Leon Wong, Blesson Varghese |
Future Gener. Comput. Syst. | 3 |
| 2025 | Special Issue on Intelligent Architectures and Platforms for Private Edge Cloud Systems
Sayed Chhattan Shah, Taehong Kim, Blesson Varghese |
Future Gener. Comput. Syst. | 3 |
| 2024 | Rapid Deployment of DNNs for Edge Computing via Structured Pruning at InitializationabstractEdge machine learning (ML) enables localized processing of data on devices and is underpinned by deep neural networks (DNNs). However, DNNs cannot be easily run on devices due to their substantial computing, memory and energy requirements for delivering performance that is comparable to cloud-based ML. Therefore, model compression techniques, such as pruning, have been considered. Existing pruning methods are problematic for edge ML since they: (1) Create compressed models that have limited runtime performance benefits (using unstructured pruning) or compromise the final model accuracy (using structured pruning), and (2) Require substantial compute resources and time for identifying a suitable compressed DNN model (using neural architecture search). In this paper, we explore a new avenue, referred to as Pruning-at-Initialization (PaI), using structured pruning to mitigate the above problems. We develop Reconvene, a system for rapidly generating pruned models suited for edge deployments using structured PaI. Reconvene systematically identifies and prunes DNN convolution layers that are least sensitive to structured pruning. Reconvene rapidly creates pruned DNNs within seconds that are up to 16.21× smaller and 2× faster while maintaining the same accuracy as an unstructured PaI counterpart. Bailey J. Eccles, Leon Wong, Blesson Varghese |
CCGrid | 3 |
| 2024 | NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local LearningabstractEfficient on-device Convolutional Neural Network (CNN) training in resource-constrained mobile and edge environments is an open challenge. Backpropagation is the standard approach adopted, but it is GPU memory intensive due to its strong inter-layer dependencies that demand intermediate activations across the entire CNN model to be retained in GPU memory. This necessitates smaller batch sizes to make training possible within the available GPU memory budget, but in turn, results in substantially high and impractical training time. We introduce NeuroFlux, a novel CNN training system tailored for memory-constrained scenarios. We develop two novel opportunities: firstly, adaptive auxiliary networks that employ a variable number of filters to reduce GPU memory usage, and secondly, block-specific adaptive batch sizes, which not only cater to the GPU memory constraints but also accelerate the training process. NeuroFlux segments a CNN into blocks based on GPU memory usage and further attaches an auxiliary network to each layer in these blocks. This disrupts the typical layer dependencies under a new training paradigm - 'adaptive local learning'. Moreover, NeuroFlux adeptly caches intermediate activations, eliminating redundant forward passes over previously trained blocks, further accelerating the training process. The results are twofold when compared to Backpropagation: on various hardware platforms, NeuroFlux demonstrates training speed-ups of 2.3× to 6.1× under stringent GPU memory budgets, and NeuroFlux generates streamlined models that have 10.9× to 29.4× fewer parameters. Dhananjay Saikumar, Blesson Varghese |
EuroSys | 2 |
| 2024 | Accelerator virtualizationabstractWelcome to this special issue on accelerator virtualization in Concurrency and Computation: Practice and Experience. Virtualization is a key technique developed for sharing the underlying physical resources of a computer, such as the processor or memory and the network. This enables effective use of resources by improving their utilization and reduces costs. A well-known example of virtualization are virtual machines that have become prominent with the advent of cloud technologies. Virtual machines are an abstraction of the underlying physical computer that can be made available to different users. The underpinning technology ensures data security by isolating the environment in which each user works. Creating and executing virtual machines requires both software and hardware support. It is thought that one of the earliest forms of virtualization was time-multiplexing a single processor for different applications. Although this significantly varies from the current notion of virtualization, processor multiplexing laid the groundwork for modern operating systems to facilitate the concurrent sharing of an expensive hardware resource among several applications as if they each used the resource exclusively. A network file system is another example of virtualization in which the file system is physically available to several nodes of a computer cluster while it is simultaneously accessed by different client nodes. In this case, storage is the common resource shared via multiplexing mechanisms. Recently, virtualization has been adopted for hardware accelerators, specialized hardware such as graphics processing units (GPUs), field programmable gate arrays (FPGAs), and tensor processing units (TPUs). Accelerators reduce the execution time of certain applications by allowing programmers to offload compute intensive components of their applications to accelerators. They are also known to improve energy efficiency. We hope that the contents of this special issue are useful to you and that you enjoy them as much as we did. Carlos Reaño, Federico Silla, Blesson Varghese |
Concurr. Comput. Pract. Exp. | 3 |
| 2024 | Enabling privacy-aware interoperable and quality IoT data sharing with contextabstractSharing Internet of Things (IoT) data across different sectors, such as in smart cities, becomes complex due to heterogeneity. This poses challenges related to a lack of interoperability, data quality issues and lack of context information, and a lack of data veracity (or accuracy). In addition, there are privacy concerns as IoT data may contain personally identifiable information. To address the above challenges, this paper presents a novel semantic technology-based framework that enables data sharing in a GDPR-compliant manner while ensuring that the data shared is interoperable, contains required context information, is of acceptable quality, and is accurate and trustworthy. The proposed framework also accounts for the edge/fog, an upcoming computing paradigm for the IoT to support real-time decisions. We evaluate the performance of the proposed framework with two different edge and fog-edge scenarios using resource-constrained IoT devices, such as the Raspberry Pi. In addition, we also evaluate shared data quality, interoperability and veracity. Our key finding is that the proposed framework can be employed on IoT devices with limited resources due to its low CPU and memory utilization for analytics operations and data transformation and migration operations. The low overhead of the framework supports real-time decision making. In addition, the 100% accuracy of our evaluation of the data quality and veracity based on 180 different observations demonstrates that the proposed framework can guarantee both data quality and veracity. Tek Raj Chhetri, Chinmaya Kumar Dehury, Blesson Varghese, Anna Fensel, Satish Narayana Srirama, Rance J. DeLong |
Future Gener. Comput. Syst. | 3 |
| 2024 | DNNShifter: An efficient DNN pruning system for edge computingabstractDeep neural networks (DNNs) underpin many machine learning applications. Production quality DNN models achieve high inference accuracy by training millions of DNN parameters which has a significant resource footprint. This presents a challenge for resources operating at the extreme edge of the network, such as mobile and embedded devices that have limited computational and memory resources. To address this, models are pruned to create lightweight, more suitable variants for these devices. Existing pruning methods are unable to provide similar quality models compared to their unpruned counterparts without significant time costs and overheads or are limited to offline use cases. Our work rapidly derives suitable model variants while maintaining the accuracy of the original model. The model variants can be swapped quickly when system and network conditions change to match workload demand. This paper presents DNNShifter , an end-to-end DNN training, spatial pruning, and model switching system that addresses the challenges mentioned above. At the heart of DNNShifter is a novel methodology that prunes sparse models using structured pruning - combining the accuracy-preserving benefits of unstructured pruning with runtime performance improvements of structured pruning. The pruned model variants generated by DNNShifter are smaller in size and thus faster than dense and sparse model predecessors, making them suitable for inference at the edge while retaining near similar accuracy as of the original dense model. DNNShifter generates a portfolio of model variants that can be swiftly interchanged depending on operational conditions. DNNShifter produces pruned model variants up to 93x faster than conventional training methods. Compared to sparse models, the pruned model variants are up to 5.14x smaller and have a 1.67x inference latency speedup, with no compromise to sparse model accuracy. In addition, DNNShifter has up to 11.9x lower overhead for switching models and up to 3.8x lower memory utilisation than existing approaches. DNNShifter is available for public use from https://github.com/blessonvar/DNNShifter. Bailey J. Eccles, Philip Rodgers, Peter Kilpatrick, Ivor T. A. Spence, Blesson Varghese |
Future Gener. Comput. Syst. | 5 |
| 2024 | PiPar: Pipeline parallelism for collaborative machine learningabstractCollaborative machine learning (CML) techniques, such as federated learning, have been proposed to train deep learning models across multiple mobile devices and a server. CML techniques are privacy-preserving as a local model that is trained on each device instead of the raw data from the device is shared with the server. However, CML training is inefficient due to low resource utilization. We identify idling resources on the server and devices due to sequential computation and communication as the principal cause of low resource utilization. A novel framework PiPar that leverages pipeline parallelism for CML techniques is developed to substantially improve resource utilization. A new training pipeline is designed to parallelize the computations on different hardware resources and communication on different bandwidth resources, thereby accelerating the training process in CML. A low overhead automated parameter selection method is proposed to optimize the pipeline, maximizing the utilization of available resources. The experimental results confirm the validity of the underlying approach of PiPar and highlight that when compared to federated learning: (i) the idle time of the server can be reduced by up to 64.1×, and (ii) the overall training time can be accelerated by up to 34.6× under varying network conditions for a collection of six small and large popular deep neural networks and four datasets without sacrificing accuracy. It is also experimentally demonstrated that PiPar achieves performance benefits when incorporating differential privacy methods and operating in environments with heterogeneous devices and changing bandwidths. Philip Rodgers, Peter Kilpatrick, Ivor T. A. Spence, Blesson Varghese |
J. Parallel Distributed Comput. | 5 |
| 2024 | ScissionLite: Accelerating Distributed Deep Learning With Lightweight Data Compression for IIoTabstractIndustrial Internet of Things (IIoT) applications can greatly benefit from leveraging edge computing. For instance, applications relying on deep neural network (DNN) models can be sliced and distributed across IIoT devices and the network edge to reduce inference latency. However, low network performance between IIoT devices and the edge often becomes a bottleneck. In this study, we propose ScissionLite, a holistic framework designed to accelerate distributed DNN inference using lightweight data compression. Our compression method features a novel lightweight down/upsampling network tailored for performance-limited IIoT devices, which is inserted at the slicing point of a DNN model to reduce outbound network traffic without causing a significant drop in accuracy. In addition, we have developed a benchmarking tool to accurately identify the optimal slicing point of the DNN for the best inference latency. ScissionLite improves inference latency by up to 15.7× with minimal accuracy degradation. Hyunho Ahn, Munkyu Lee, Sihoon Seong, Gap-Joo Na, In-Geol Chun, Blesson Varghese, Cheol-Ho Hong |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | EcoFed: Efficient Communication for DNN Partitioning-Based Federated LearningabstractEfficiently running federated learning (FL) on resource-constrained devices is challenging since they are required to train computationally intensive deep neural networks (DNN) independently. DNN partitioning-based FL (DPFL) has been proposed as one mechanism to accelerate training where the layers of a DNN (or computation) are offloaded from the device to the server. However, this creates significant communication overheads since the intermediate activation and gradient need to be transferred between the device and the server during training. While current research reduces the communication introduced by DNN partitioning using local loss-based methods, we demonstrate that these methods are ineffective in improving the overall efficiency (communication overhead and training speed) of a DPFL system. This is because they suffer from accuracy degradation and ignore the communication costs incurred when transferring the activation from the device to the server. This article proposesEcoFed– a communication efficient framework for DPFL systems.EcoFedeliminates the transmission of the gradient by developing pre-trained initialization of the DNN model on the device for the first time. This reduces the accuracy degradation seen in local loss-based methods. In addition,EcoFedproposes a novel replay buffer mechanism and implements a quantization-based compression technique to reduce the transmission of the activation. It is experimentally demonstrated thatEcoFedcan reduce the communication cost by up to 133× and accelerate training by up to 21× when compared to classic FL. Compared to vanilla DPFL,EcoFedachieves a 16× communication reduction and 2.86× training time speed-up. Di Wu 0065, Rehmat Ullah 0001, Philip Rodgers, Peter Kilpatrick, Ivor T. A. Spence, Blesson Varghese |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2023 | ROMA: Run-Time Object Detection To Maximize Real-Time AccuracyabstractThis paper analyzes the effects of dynamically varying video contents and detection latency on the real-time detection accuracy of a detector and proposes a new run-time accuracy variation model, ROMA, based on the findings from the analysis. ROMA is designed to select an optimal detector out of a set of detectors in real time without label information to maximize real-time object detection accuracy. ROMA utilizing four YOLOv4 detectors on an NVIDIA Jetson Nano shows real-time accuracy improvements by 4 to 37% for a scenario of dynamically varying video con-tents and detection latency consisting of MOT17Det and MOT20Det datasets, compared to individual YOLOv4 detectors and two state-of-the-art runtime techniques. Blesson Varghese, Hans Vandierendonck |
WACV | 2 |
| 2023 | Multi-Tier GPU Virtualization for Deep Learning in Cloud-Edge SystemsabstractAccelerator virtualization offers several advantages in the context of cloud-edge computing. Relatively weak user devices can enhance performance when running workloads by accessing virtualized accelerators available on other resources in the cloud-edge continuum. However, cloud-edge systems are heterogeneous, often leading to compatibility issues arising from various hardware and software stacks present in the system. One mechanism to alleviate this issue is using containers for deploying workloads. Containers isolate applications and their dependencies and store them as images that can run on any device. In addition, user devices may move during the course of application execution, and thus mechanisms such as container migration are required to move running workloads from one resource to another in the network. Furthermore, an optimal destination will need to be determined when migrating between virtual accelerators. Scheduling and placement strategies are incorporated to choose the best possible location depending on the workload requirements. This paper presentsAVEC, a framework for accelerator virtualization in cloud-edge computing. The AVEC framework enables the offloading of deep learning workloads for inference from weak user devices to computationally more powerful devices in a cloud-edge network. AVEC incorporates a mechanism that efficiently manages and schedules the virtualization of accelerators. It also supports migration between accelerators to enable stateless container migration. The experimental analysis highlights that AVEC can achieve up to 7x speedup by offloading applications to remote resources. Furthermore, AVEC features a low migration downtime that is less than 5 seconds. Jason Kennedy, Vishal Sharma 0001, Blesson Varghese, Carlos Reaño |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | FedAdapt: Adaptive Offloading for IoT Devices in Federated LearningabstractApplying federated learning (FL) on Internet of Things (IoT) devices is necessitated by the large volumes of data they produce and growing concerns of data privacy. However, there are three challenges that need to be addressed to make FL efficient: 1) execution on devices with limited computational capabilities; 2) accounting for stragglers due to computational heterogeneity of devices; and 3) adaptation to the changing network bandwidths. This article presentsFedAdapt, an adaptive offloading FL framework to mitigate the aforementioned challenges.FedAdaptaccelerates local training in computationally constrained devices by leveraging layer offloading of deep neural networks (DNNs) to servers. Furthermore,FedAdaptadopts reinforcement learning (RL)-based optimization and clustering to adaptively identify which layers of the DNN should be offloaded for each individual device on to a server to tackle the challenges of computational heterogeneity and changing network bandwidth. The experimental studies are carried out on a lab-based testbed and it is demonstrated that by offloading a DNN from the device to the serverFedAdaptreduces the training time of a typical IoT device by over half compared to classic FL. The training time of extreme stragglers and the overall training time can be reduced by up to 57%. Furthermore, with changing network bandwidth,FedAdaptis demonstrated to reduce the training time by up to 40% when compared to classic FL, without sacrificing accuracy. Di Wu 0065, Rehmat Ullah 0001, Paul Harvey 0002, Peter Kilpatrick, Ivor T. A. Spence, Blesson Varghese |
IEEE Internet Things J. | 6 |
| 2022 | Context-aware distribution of fog applications using deep reinforcement learningabstractFog computing is an emerging paradigm that aims to meet the increasing computation demands arising from the billions of devices connected to the Internet. Offloading services of an application from the Cloud to the edge of the network can improve the overall latency of the application since it can process data closer to user devices. Diverse Fog nodes ranging from Wi-Fi routers to mini-clouds with varying resource capabilities makes it challenging to determine which services of an application need to be offloaded. In this paper, a context-aware mechanism for distributing applications across the Cloud and the Fog is proposed. The mechanism dynamically generates (re)deployment plans for the application to maximise the performance efficiency of the application by taking operational conditions, such as hardware utilisation and network state, and running costs into account. The mechanism relies on deep Q-networks to generate a distribution plan without prior knowledge of the available resources on the Fog node, the network condition, and the application. The feasibility of the proposed context-aware distribution mechanism is demonstrated on two use-cases, namely a face detection application and a location-based mobile game. The benefits are increased utility of dynamic distribution by 50% and 20% for the two use-cases respectively when compared to a static distribution approach used in existing research. Nan Wang 0009, Blesson Varghese |
J. Netw. Comput. Appl. | 2 |
| 2021 | The Case for Adaptive Deep Neural Networks in Edge ComputingabstractDeep Neural Networks (DNNs) are an application class that benefit from being distributed across the edge and cloud. A DNN is partitioned such that specific layers of the DNN are deployed onto the edge and the cloud to meet performance and privacy objectives. However, there is limited understanding of: whether and how evolving operational conditions (increased CPU and memory utilization at the edge or reduced data transfer rates between the edge and cloud) affect the performance of already deployed DNNs, and whether a new partition configuration is required to maximize performance. A DNN that adapts to changing operational conditions is referred to as an ‘adaptive DNN’. This paper investigates whether there is a case for adaptive DNNs by considering four questions: (i) Are DNNs sensitive to operational conditions? (ii) How sensitive are DNNs to operational conditions? (iii) Do individual or a combination of operational conditions equally affect DNNs? (iv) Is DNN partitioning sensitive to hardware architectures? The exploration is carried out in the context of 8 pre-trained DNN models and the results presented are from analyzing nearly 8 million data points. The results highlight that network conditions affect DNN performance more than CPU or memory related operational conditions. Repartitioning is noted to provide a performance gain in a number of cases, but a specific trend is not noted in relation to the underlying hardware architecture. Nonetheless, the need for adaptive DNNs is confirmed. Francis McNamee, Schahram Dustdar, Peter Kilpatrick, Weisong Shi, Ivor T. A. Spence, Blesson Varghese |
CLOUD | 6 |
| 2021 | NEUKONFIG: Reducing Edge Service Downtime When Repartitioning DNNsabstractDeep Neural Networks (DNNs) may be partitioned across the edge and the cloud to improve the performance efficiency of inference. DNN partitions are determined based on operational conditions such as network speed. When operational conditions change DNNs will need to be repartitioned to maintain the overall performance. However, repartitioning using existing approaches, such as Pause and Resume, will incur a service downtime on the edge. This paper presents the NEUKONFIG framework that identifies the service downtime incurred when repartitioning DNNs and proposes approaches for reducing edge service downtime. The proposed approaches are based on ‘Dynamic Switching’ in which, when the network speed changes and given an existing edge-cloud pipeline, a new edge-cloud pipeline is initialised with new DNN partitions. Incoming inference requests are switched to the new pipeline for processing data. Experimental studies are carried out on a lab-based testbed to demonstrate that Dynamic Switching reduces the downtime by at least an order of magnitude when compared to a baseline using Pause and Resume that has a downtime of 6 seconds. A trade-off in the edge service downtime and memory required is noted. The Dynamic Switching approach that requires the same amount of memory as the baseline reduces the edge service downtime to 0.6 seconds and to less than 1 millisecond in the best case when twice the amount of memory as the baseline is available. Ayesha Abdul Majeed, Peter Kilpatrick, Ivor T. A. Spence, Blesson Varghese |
IC2E | 4 |
| 2021 | AVEC: Accelerator Virtualization in Cloud-Edge Computing for Deep Learning LibrariesabstractEdge computing offers the distinct advantage of harnessing compute capabilities on resources located at the edge of the network to run workloads of relatively weak user devices. This is achieved by offloading computationally intensive workloads, such as deep learning from user devices to the edge. Using the edge reduces the overall communication latency of applications as workloads can be processed closer to where data is generated on user devices rather than sending them to geographically distant clouds. Specialised hardware accelerators, such as Graphics Processing Units (GPUs) available in the cloud-edge network can enhance the performance of computationally intensive workloads that are offloaded from devices on to the edge. The underlying approach required to facilitate this is virtualization of GPUs. This paper therefore sets out to investigate the potential of GPU accelerator virtualization to improve the performance of deep learning workloads in a cloud-edge environment. The AVEC accelerator virtualization framework is proposed that incurs minimum overheads and requires no source-code modification of the workload. AVEC intercepts local calls to a GPU on a device and forwards them to an edge resource seamlessly. The feasibility of AVEC is demonstrated on a real-world application, namely OpenPose using the Caffe deep learning library. It is observed that on a lab-based experimental test-bed AVEC delivers up to 7.48x speedup despite communication overheads incurred due to data transfers. Jason Kennedy, Blesson Varghese, Carlos Reaño |
ICFEC | 2 |
| 2021 | TOD: Transprecise Object Detection to Maximise Real-Time Accuracy on the EdgeabstractReal-time video analytics on the edge is challenging as the computationally constrained resources typically cannot analyse video streams at full fidelity and frame rate, which results in loss of accuracy. This paper proposes a Transprecise Object Detector (TOD) which maximises the real-time object detection accuracy on an edge device by selecting an appropriate Deep Neural Network (DNN) on the fly with negligible computational overhead. TOD makes two key contributions over the state of the art: (1) TOD leverages characteristics of the video stream such as object size and speed of movement to identify networks with high prediction accuracy for the current frames; (2) it selects the best-performing network based on projected accuracy and computational demand using an effective and low-overhead decision mechanism. Experimental evaluation on a Jetson Nano demonstrates that TOD improves the average object detection precision by 34.7 % over the YOLOv4-tiny-288 model on average over the MOT17Det dataset. In the MOT17-05 test dataset, TOD utilises only 45.1 % of GPU resource and 62.7 % of the GPU board power without losing accuracy, compared to YOLOv4-416 model. We expect that TOD will maximise the application of edge devices to real-time object detection, since TOD maximises real-time object detection accuracy given edge devices according to dynamic input features without increasing inference latency in practice. Blesson Varghese, Roger F. Woods, Hans Vandierendonck |
ICFEC | 2 |
| 2021 | Editorial
Carlos Reaño, Federico Silla, Blesson Varghese |
J. Parallel Distributed Comput. | 3 |
| 2020 | Cross Architectural Power ModellingabstractExisting power modelling research focuses on the model rather than the process for developing models. An automated power modelling process that can be deployed on different processors for developing power models with high accuracy is developed. For this, (i) an automated hardware performance counter selection method that selects counters best correlated to power on both ARM and Intel processors, (ii) a noise filter based on clustering that can reduce the mean error in power models, and (iii) a two stage power model that surmounts challenges in using existing power models across multiple architectures are proposed and developed. The key results are: (i) the automated hardware performance counter selection method achieves comparable selection to the manual method reported in the literature, (ii) the noise filter reduces the mean error in power models by up to 55%, and (iii) the two stage power model can predict dynamic power with less than 8% error on both ARM and Intel processors, which is an improvement over classic models. Peter Kilpatrick, Dimitrios S. Nikolopoulos, Blesson Varghese |
CCGRID | 4 |
| 2020 | Priority-based Fair Scheduling in Edge ComputingabstractScheduling is important in Edge computing. In contrast to the Cloud, Edge resources are hardware limited and cannot support workload-driven infrastructure scaling. Hence, resource allocation and scheduling for the Edge requires a fresh perspective. Existing Edge scheduling research assumes availability of all needed resources whenever a job request is made. This paper challenges that assumption, since not all job requests from a Cloud server can be scheduled on an Edge node. Thus, guaranteeing fairness among the clients (Cloud servers offloading jobs) while accounting for priorities of the jobs becomes a critical task. This paper presents four scheduling techniques, the first is a naive first come first serve strategy and further proposes three strategies, namely a client fair, priority fair, and hybrid that accounts for the fairness of both clients and job priorities. An evaluation on a target platform under three different scenarios, namely equal, random, and Gaussian job distributions is presented. The experimental studies highlight the low overheads and the distribution of scheduled jobs on the Edge node when compared to the naive strategy. The results confirm the superior performance of the hybrid strategy and showcase the feasibility of fair schedulers for Edge computing. Arkadiusz Madej, Nan Wang 0009, Nikolaos Athanasopoulos, Rajiv Ranjan 0001, Blesson Varghese |
ICFEC | 5 |
| 2020 | Modelling Fog Offloading PerformanceabstractFog computing has emerged as a computing paradigm aimed at addressing the issues of latency, bandwidth and privacy when mobile devices are communicating with remote cloud services. The concept is to offload compute services closer to the data. However many challenges exist in the realisation of this approach. During offloading, (part of) the application underpinned by the services may be unavailable, which the user will experience as down time. This paper describes work aimed at building models to allow prediction of such down time based on metrics (operational data) of the underlying and surrounding infrastructure. Such prediction would be invaluable in the context of automated Fog offloading and adaptive decision making in Fog orchestration. Models that cater for four container-based stateless and stateful offload techniques, namely Save and Load, Export and Import, Push and Pull and Live Migration, are built using four (linear and non-linear) regression techniques. Experimental results comprising over 42 million data points from multiple lab-based Fog infrastructure are presented. The results highlight that reasonably accurate predictions (measured by the coefficient of determination for regression models, mean absolute percentage error, and mean absolute error) may be obtained when considering 25 metrics relevant to the infrastructure. Ayesha Abdul Majeed, Peter Kilpatrick, Ivor T. A. Spence, Blesson Varghese |
ICFEC | 4 |
| 2020 | DYVERSE: DYnamic VERtical Scaling in multi-tenant Edge environments
Nan Wang 0009, Michail Matthaiou, Dimitrios S. Nikolopoulos, Blesson Varghese |
Future Gener. Comput. Syst. | 4 |
| 2020 | IoTSim-SDWAN: A simulation framework for interconnecting distributed datacenters over Software-Defined Wide Area Network (SD-WAN)
Khaled Alwasel, Devki Nandan Jha, Deepak Puthal, Mutaz Barika, Blesson Varghese, Saurabh Kumar Garg 0001, Philip James 0002, Albert Y. Zomaya, Graham Morgan, Rajiv Ranjan 0001 |
J. Parallel Distributed Comput. | 6 |
| 2020 | New generation cloud computingabstractWe are pleased to present a special issue that focuses on the trends in the next generation of cloud computing. An obvious question that may arise in the mind of the reader is why another issue on clouds when there is plenty of discourse on the topic. Undoubtedly, the cloud computing landscape is rapidly changing to meet the challenges of emerging paradigms, such as Fog/Edge computing and the Internet‐of‐Things (IoT). These paradigms rely on services offered by the cloud but are demanding—need to scale for billions of heterogeneous devices and sensors while operating efficiently in real‐time. Blesson Varghese, Marco Aurélio Stelmar Netto, Ignacio Martín Llorente, Rajkumar Buyya |
Softw. Pract. Exp. | 1 |
| 2020 | ENORM: A Framework For Edge NOde Resource ManagementabstractCurrent computing techniques using the cloud as a centralised server will become untenable as billions of devices get connected to the Internet. This raises the need for fog computing, which leverages computing at the edge of the network on nodes, such as routers, base stations and switches, along with the cloud. However, to realise fog computing the challenge of managing edge nodes will need to be addressed. This paper is motivated to address the resource management challenge. We develop the first framework to manage edge nodes, namely the Edge NOde Resource Management (ENORM) framework. Mechanisms for provisioning and auto-scaling edge node resources are proposed. The feasibility of the framework is demonstrated on a PokéMon Go-like online game use-case. The benefits of using ENORM are observed by reduced application latency between 20-80 percent and reduced data transfer and communication frequency between the edge node and the cloud by up to 95 percent. These results highlight the potential of fog computing for improving the quality of service and experience. Nan Wang 0009, Blesson Varghese, Michail Matthaiou, Dimitrios S. Nikolopoulos |
IEEE Trans. Serv. Comput. | 2 |
| 2019 | Cloud Benchmarking for Maximising Performance of Scientific ApplicationsabstractHow can applications be deployed on the cloud to achieve maximum performance? This question is challenging to address with the availability of a wide variety of cloud Virtual Machines (VMs) with different performance capabilities. The research reported in this paper addresses the above question by proposing a six step benchmarking methodology in which a user provides a set of weights that indicate how important memory, local communication, computation and storage related operations are to an application. The user can either provide a set of four abstract weights or eight fine grain weights based on the knowledge of the application. The weights along with benchmarking data collected from the cloud are used to generate a set of two rankings-one based only on the performance of the VMs and the other takes both performance and costs into account. The rankings are validated on three case study applications using two validation techniques. The case studies on a set of experimental VMs highlight that maximum performance can be achieved by the three top ranked VMs and maximum performance in a cost-effective manner is achieved by at least one of the top three ranked VMs produced by the methodology. Blesson Varghese, Özgür Akgün, Ian Miguel, Long Thai, Adam Barker |
IEEE Trans. Cloud Comput. | 1 |
| 2018 | A survey and taxonomy of resource optimisation for executing bag-of-task applications on public clouds
Long Thai, Blesson Varghese, Adam Barker |
Future Gener. Comput. Syst. | 2 |
| 2018 | Next generation cloud computing: New trends and research directions
Blesson Varghese, Rajkumar Buyya |
Future Gener. Comput. Syst. | 1 |
| 2018 | Intra-Node Memory Safe GPU Co-SchedulingabstractGPUs in High-Performance Computing systems remain under-utilised due to the unavailability of schedulers that can safely schedule multiple applications to share the same GPU. The research reported in this paper is motivated to improve the utilisation of GPUs by proposing a framework, we refer to as schedGPU, to facilitate intra-node GPU co-scheduling such that a GPU can be safely shared among multiple applications by taking memory constraints into account. Two approaches, namely a client-server and a shared memory approach are explored. However, the shared memory approach is more suitable due to lower overheads when compared to the former approach. Four policies are proposed in schedGPU to handle applications that are waiting to access the GPU, two of which account for priorities. The feasibility of schedGPU is validated on three real-world applications. The key observation is that a performance gain is achieved. For single applications, a gain of over 10 times, as measured by GPU utilisation and GPU memory utilisation, is obtained. For workloads comprising multiple applications, a speed-up of up to 5x in the total execution time is noted. Moreover, the average GPU utilisation and average GPU memory utilisation is increased by 5 and 12 times, respectively. Carlos Reaño, Federico Silla, Dimitrios S. Nikolopoulos, Blesson Varghese |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | Plug and play bench: Simplifying big data benchmarking using containersabstractThe recent boom of big data, coupled with the challenges of its processing and storage gave rise to the development of distributed data processing and storage paradigms like MapReduce, Spark, and NoSQL databases. With the advent of cloud computing, processing and storing such massive datasets on clusters of machines is now feasible with ease. However, there are limited tools and approaches, which users can rely on to gauge and comprehend the performance of their big data applications deployed locally on clusters, or in the cloud. Researchers have started exploring this area by providing benchmarking suites suitable for big data applications. However, many of these tools are fragmented, complex to deploy and manage, and do not provide transparency with respect to the monetary cost of benchmarking an application. In this paper, we present Plug And Play Bench (PAPB1): an infrastructure aware abstraction built to integrate and simplify the deployment of big data benchmarking tools on clusters of machines. PAPB automates the tedious process of installing, configuring and executing common big data benchmark workloads by containerising the tools and settings based on the underlying cluster deployment framework. Our proof of concept implementation utilises HiBench as the benchmark suite, HDP as the cluster deployment framework and Azure as the cloud platform. The paper further illustrates the inclusion of cost metrics based on the underlying Microsoft Azure cloud platform. Sheriffo Ceesay, Adam Barker, Blesson Varghese |
IEEE BigData | 3 |
| 2017 | Multi-tenant virtual GPUs for optimising performance of a financial risk application
Javier Prades, Blesson Varghese, Carlos Reaño, Federico Silla |
J. Parallel Distributed Comput. | 2 |
| 2016 | DocLite: A Docker-Based Lightweight Cloud Benchmarking ToolabstractExisting benchmarking methods are time consuming processes as they typically benchmark the entire Virtual Machine (VM) in order to generate accurate performance data, making them less suitable for real-time analytics. The research in this paper is aimed to surmount the above challenge by presenting DocLite - Docker Container-based Lightweight benchmarking tool. DocLite explores lightweight cloud benchmarking methods for rapidly executing benchmarks in near real-time. DocLite is built on the Docker container technology, which allows a user-defined memory size and number of CPU cores of the VM to be benchmarked. The tool incorporates two benchmarking methods - the first referred to as the native method employs containers to benchmark a small portion of the VM and generate performance ranks, and the second uses historic benchmark data along with the native method as a hybrid to generate VM ranks. The proposed methods are evaluated on three use-cases and are observed to be up to 91 times faster than benchmarking the entire VM. In both methods, small containers provide the same quality of rankings as a large container. The native method generates ranks with over 90% and 86% accuracy for sequential and parallel execution of an application compared against benchmarking the whole VM. The hybrid method did not improve the quality of the rankings significantly. Blesson Varghese, Lawan Thamsuhang Subba, Long Thai, Adam Barker |
CCGrid | 1 |
| 2016 | Algorithms for Optimising Heterogeneous Cloud Virtual Machine ClustersabstractIt is challenging to execute an application in a heterogeneous cloud cluster, which consists of multiple types of virtual machines with different performance capabilities and prices. This paper aims to mitigate this challenge by proposing a scheduling mechanism to optimise the execution of Bag-of-Task jobs on a heterogeneous cloud cluster. The proposed scheduler considers two approaches to select suitable cloud resources for executing a user application while satisfying pre-defined Service Level Objectives (SLOs) both in terms of execution deadline and minimising monetary cost. Additionally, a mechanism for dynamic re-assignment of jobs during execution is presented to resolve potential violation of SLOs. Experimental studies are performed both in simulation and on a public cloud using real-world applications. The results highlight that our scheduling approaches result in cost saving of up to 31% in comparison to naive approaches that only employ a single type of virtual machine in a homogeneous cluster. Dynamic reassignment completely prevents deadline violation in the best-case and reduces deadline violations by 95% in the worst-case scenario. Long Thai, Blesson Varghese, Adam Barker |
CloudCom | 2 |
| 2016 | A machine learning analysis of Twitter sentiment to the Sandy Hook shootingsabstractGun related violence is a complex issue and accounts for a large proportion of violent incidents. In the research reported in this paper, we set out to investigate the pro-gun and anti-gun sentiments expressed on a social media platform, namely Twitter, in response to the 2012 Sandy Hook Elementary School shooting in Connecticut, USA. Machine learning techniques are applied to classify a data corpus of over 700,000 tweets. The sentiments are captured using a public sentiment score that considers the volume of tweets as well as population. A web-based interactive tool is developed to visualise the sentiments and is available at http://www.gunsontwitter.com. The key findings from this research are: (i) There are elevated rates of both pro-gun and anti-gun sentiments on the day of the shooting. Surprisingly, the pro-gun sentiment remains high for a number of days following the event but the anti-gun sentiment quickly falls to pre-event levels. (ii) There is a different public response from each state, with the highest pro-gun sentiment not coming from those with highest gun ownership levels but rather from California, Texas and New York. Nan Wang 0009, Blesson Varghese, Peter D. Donnelly |
eScience | 2 |
| 2016 | Container-Based Cloud Virtual Machine BenchmarkingabstractWith the availability of a wide range of cloud Virtual Machines (VMs) it is difficult to determine which VMs can maximise the performance of an application. Benchmarking is commonly used to this end for capturing the performance of VMs. Most cloud benchmarking techniques are typically heavyweight - time consuming processes which have to benchmark the entire VM in order to obtain accurate benchmark data. Such benchmarks cannot be used in real-time on the cloud and incur extra costs even before an application is deployed. In this paper, we present lightweight cloud benchmarking techniques that execute quickly and can be used in near real-time on the cloud. The exploration of lightweight benchmarking techniques are facilitated by the development of DocLite - Docker Container-based Lightweight Benchmarking. DocLite is built on the Docker container technology which allows a user-defined portion (such as memory size and the number of CPU cores) of the VM to be benchmarked. DocLite operates in two modes, in the first mode, containers are used to benchmark a small portion of the VM to generate performance ranks. In the second mode, historic benchmark data is used along with the first mode as a hybrid to generate VM ranks. The generated ranks are evaluated against three scientific high-performance computing applications. The proposed techniques are up to 91 times faster than a heavyweight technique which benchmarks the entire VM. It is observed that the first mode can generate ranks with over 90% and 86% accuracy for sequential and parallel execution of an application. The hybrid mode improves the correlation slightly but the first mode is sufficient for benchmarking cloud VMs. Blesson Varghese, Lawan Thamsuhang Subba, Long Thai, Adam Barker |
IC2E | 1 |
| 2016 | Computing probable maximum loss in catastrophe reinsurance portfolios on multi-core and many-core architecturesabstractSummary In the reinsurance market, the risks natural catastrophes pose to portfolios of properties must be quantified, so that they can be priced, and insurance offered. The analysis of such risks at a portfolio level requires a simulation of up to 800 000 trials with an average of 1000 catastrophic events per trial. This is sufficient to capture risk for a global multi‐peril reinsurance portfolio covering a range of perils including earthquake, hurricane, tornado, hail, severe thunderstorm, wind storm, storm surge and riverine flooding, and wildfire. Such simulations are both computation and data intensive, making the application of high‐performance computing techniques desirable. In this paper, we explore the design and implementation of portfolio risk analysis on both multi‐core and many‐core computing platforms. Given a portfolio of property catastrophe insurance treaties, key risk measures, such as probable maximum loss, are computed by taking both primary and secondary uncertainties into account. Primary uncertainty is associated with whether or not an event occurs in a simulated year, while secondary uncertainty captures the uncertainty in the level of loss due to the use of simplified physical models and limitations in the available data. A combination of fast lookup structures, multi‐threading and careful hand tuning of numerical operations is required to achieve good performance. Experimental results are reported for multi‐core processors and systems using NVIDIA graphics processing unit and Intel Phi many‐core accelerators. Copyright © 2015 John Wiley & Sons, Ltd. Neil Burke, Andrew Rau-Chaplin, Blesson Varghese |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Accelerating R-based analytics on the cloudabstractSummary This paper addresses how the benefits of cloud‐based infrastructure can be harnessed for analytical workloads. Often, the software handling analytical workloads is not developed by a professional programmer but on an ad hoc basis by analysts in high‐level programming environments such as R or MATLAB. The goal of this research is to allow Analysts to take an analytical job that executes on their personal workstations and with minimum effort execute it on cloud infrastructure and manage both the resources and the data required by the job. If this can be facilitated gracefully, then the Analyst benefits from on‐demand resources, low maintenance cost and scalability of computing resources, all of which are offered by the cloud. In this paper, a Platform for Parallel R‐based Analytics on the Cloud (P2RAC) that is placed between an Analyst and a cloud infrastructure is proposed and implemented. P2RAC offers a set of command‐line tools for managing the resources, such as instances and clusters, the data and the execution of the software on the Amazon Elastic Computing Cloud infrastructure. Experimental studies are pursued using two parallel problems and the results obtained confirm the feasibility of employing P2RAC for solving large‐scale analytical problems on the cloud.Copyright © 2013 John Wiley & Sons, Ltd. Ishan Patel, Andrew Rau-Chaplin, Blesson Varghese |
Concurr. Comput. Pract. Exp. | 3 |
| 2015 | Cloud Services Brokerage: A Survey and Research RoadmapabstractA Cloud Services Brokerage (CSB) acts as an intermediary between cloud service providers (e.g., Amazon and Google) and cloud service end users, providing a number of value adding services. CSBs as a research topic are in there infancy. The goal of this paper is to provide a concise survey of existing CSB technologies in a variety of areas and highlight a roadmap, which details five future opportunities for research. Adam Barker, Blesson Varghese, Long Thai |
CLOUD | 2 |
| 2015 | Budget Constrained Execution of Multiple Bag-of-Tasks Applications on the CloudabstractOptimising the execution of Bag-of-Tasks (BoT) applications on the cloud is a hard problem due to the trade-offs between performance and monetary cost. The problem can be further complicated when multiple BoT applications need to be executed. In this paper, we propose and implement a heuristic algorithm that schedules tasks of multiple applications onto different cloud virtual machines in order to maximise performance while satisfying a given budget constraint. Current approaches are limited in task scheduling since they place a limit on the number of cloud resources that can be employed by the applications. However, in the proposed algorithm there are no such limits, and in comparison with other approaches, the algorithm on average achieves an improved performance of 10%. The experimental results also highlight that the algorithm yields consistent performance even with low budget constraints which cannot be achieved by competing approaches. Long Thai, Blesson Varghese, Adam Barker |
CLOUD | 2 |
| 2015 | Executing Bag of Distributed Tasks on Virtually Unlimited Cloud ResourcesabstractThis research is supported by the EPSRC grant ‘Working Together: Constraint Programming and Cloud Computing’ (EP/K015745/1), a Royal Society Industry Fellowship, an Impact Acceleration Account Grant (IAA) and an Amazon Web Services (AWS) Education Research Grant. Long Thai, Blesson Varghese, Adam Barker |
CLOSER | 2 |
| 2015 | Acceleration-as-a-Service: Exploiting Virtualised GPUs for a Financial ApplicationabstractHow can GPU acceleration be obtained as a service in a cluster? This question has become increasingly significant due to the inefficiency of installing GPUs on all nodes of a cluster. The research reported in this paper is motivated to address the above question by employing rCUDA (remote CUDA), a framework that facilitates Acceleration-as-a-Service (AaaS), such that the nodes of a cluster can request the acceleration of a set of remote GPUs on demand. The rCUDA framework exploits virtualisation and ensures that multiple nodes can share the same GPU. In this paper we test the feasibility of the rCUDA framework on a real-world application employed in the financial risk industry that can benefit from AaaS in the production setting. The results confirm the feasibility of rCUDA and highlight that rCUDA achieves similar performance compared to CUDA, provides consistent results, and more importantly, allows for a single application to benefit from all the GPUs available in the cluster without loosing efficiency. Blesson Varghese, Javier Prades, Carlos Reaño, Federico Silla |
e-Science | 1 |
| 2015 | Cloud-based E-Infrastructure for Scheduling Astronomical ObservationsabstractGravitational microlensing exploits a transient phenomenon where an observed star is brightened due to deflection of its light by the gravity of an intervening foreground star. It is conjectured that this technique can be used to measure the abundance of planets throughout the Milky Way. In order to undertake efficient gravitational microlensing an observation schedule must be constructed such that various targets are observed while undergoing a microlensing event. In this paper, we propose a cloud-based e-Infrastructure that currently supports four methods to compute candidate schedules via the application of local search and probabilistic meta-heuristics. We then validate the feasibility of the e-Infrastructure by evaluating the methods on historic data. The experiments demonstrate that the use of on-demand cloud resources for the e-Infrastructure can allow better schedules to be found more rapidly. James Wetter, Özgür Akgün, Adam Barker, Martin Dominik, Ian Miguel, Blesson Varghese |
e-Science | 6 |
| 2015 | Task Scheduling on the Cloud with Hard ConstraintsabstractScheduling Bag-of-Tasks (BoT) applications on the cloud can be more challenging than grid and cluster environments. This is because a user may have a budgetary constraint or a deadline for executing the BoT application in order to keep the overall execution costs low. The research in this paper is motivated to investigate task scheduling on the cloud, given two hard constraints based on a user-defined budget and a deadline. A heuristic algorithm is proposed and implemented to satisfy the hard constraints for executing the BoT application in a cost effective manner. The proposed algorithm is evaluated using four scenarios that are based on the trade-off between performance and the cost of using different cloud resource types. The experimental evaluation confirms the feasibility of the algorithm in satisfying the constraints. The key observation is that multiple resource types can be a better alternative to using a single type of resource. Long Thai, Blesson Varghese, Adam Barker |
SERVICES | 2 |
| 2015 | RBioCloud: A Light-Weight Framework for Bioconductor and R-based Jobs on the CloudabstractLarge-scale ad hoc analytics of genomic data is popular using the R-programming language supported by over 700 software packages provided by Bioconductor. More recently, analytical jobs are benefitting from on-demand computing and storage, their scalability and their low maintenance cost, all of which are offered by the cloud. While biologists and bioinformaticists can take an analytical job and execute it on their personal workstations, it remains challenging to seamlessly execute the job on the cloud infrastructure without extensive knowledge of the cloud dashboard. How analytical jobs can not only with minimum effort be executed on the cloud, but also how both the resources and data required by the job can be managed is explored in this paper. An open-source light-weight framework for executing R-scripts using Bioconductor packages, referred to as `RBioCloud', is designed and developed. RBioCloud offers a set of simple command-line tools for managing the cloud resources, the data and the execution of the job. Three biological test cases validate the feasibility of RBioCloud. The framework is available from http://www.rbiocloud.com. Blesson Varghese, Ishan Patel, Adam Barker |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2014 | BigExcel: A web-based framework for exploring big data in social sciencesabstractThis paper argues that there are three fundamental challenges that need to be overcome in order to foster the adoption of big data technologies in non-computer science related disciplines: addressing issues of accessibility of such technologies for non-computer scientists, supporting the ad hoc exploration of large data sets with minimal effort and the availability of lightweight web-based frameworks for quick and easy analytics. In this paper, we address the above three challenges through the development of `BigExcel', a three tier web-based framework for exploring big data to facilitate the management of user interactions with large data sets, the construction of queries to explore the data set and the management of the infrastructure. The feasibility of BigExcel is demonstrated through two Yahoo Sandbox datasets. The first dataset is the Yahoo Buzz Score data set we use for quantitatively predicting trending technologies and the second is the Yahoo n-gram corpus we use for qualitatively inferring the coverage of important events. A demonstration of the BigExcel framework and source code is available at http://bigdata.cs.st-andrews.ac. uk/projects/bigexcel-exploring-big-data-for-social-sciences/. Muhammed Asif Saleem, Blesson Varghese, Adam Barker |
IEEE BigData | 2 |
| 2014 | Optimal Deployment of Geographically Distributed Workflow Engines on the CloudabstractWhen orchestrating Web service workflows, the geographical placement of the orchestration engine (s) can greatly affect workflow performance. Data may have to be transferred across long geographical distances, which in turn increases execution time and degrades the overall performance of a workflow. In this paper, we present a framework that, given a DAG-based workflow specification, computes the optimal Amazon EC2 cloud regions to deploy the orchestration engines and execute a workflow. The framework incorporates a constraint model that solves the workflow deployment problem, which is generated using an automated constraint modelling system. The feasibility of the framework is evaluated by executing different sample workflows representative of scientific workloads. The experimental results indicate that the framework reduces the workflow execution time and provides a speed up of 1.3x-2.5x over centralised approaches. Long Thai, Adam Barker, Blesson Varghese, Özgür Akgün, Ian Miguel |
CloudCom | 3 |
| 2014 | Executing Bag of Distributed Tasks on the Cloud: Investigating the Trade-Offs between Performance and CostabstractBag of Distributed Tasks (BoDT) can benefit from decentralised execution on the Cloud. However, there is a trade-off between the performance that can be achieved by employing a large number of Cloud VMs for the tasks and the monetary constraints that are often placed by a user. The research reported in this paper is motivated towards investigating this trade-off so that an optimal plan for deploying BoDT applications on the cloud can be generated. A heuristic algorithm, which considers the user's preference of performance and cost is proposed and implemented. The feasibility of the algorithm is demonstrated by generating execution plans for a sample application. The key result is that the algorithm generates optimal execution plans for the application over 91% of the time. Long Thai, Blesson Varghese, Adam Barker |
CloudCom | 2 |
| 2014 | Cloud Benchmarking for PerformanceabstractHow can applications be deployed on the cloud to achieve maximum performance? This question has become significant and challenging with the availability of a wide variety of Virtual Machines (VMs) with different performance capabilities in the cloud. The above question is addressed by proposing a six step benchmarking methodology in which a user provides a set of four weights that indicate how important each of the following groups: memory, processor, computation and storage are to the application that needs to be executed on the cloud. The weights along with cloud benchmarking data are used to generate a ranking of VMs that can maximise performance of the application. The rankings are validated through an empirical analysis using two case study applications, the first is a financial risk application and the second is a molecular dynamics simulation, which are both representative of workloads that can benefit from execution on the cloud. Both case studies validate the feasibility of the methodology and highlight that maximum performance can be achieved on the cloud by selecting the top ranked VMs produced by the methodology. Blesson Varghese, Özgür Akgün, Ian Miguel, Long Thai, Adam Barker |
CloudCom | 1 |
| 2013 | The royal birth of 2013: Analysing and visualising public sentiment in the UK using TwitterabstractAnalysis of information retrieved from microblog-ging services such as Twitter can provide valuable insight into public sentiment in a geographic region. This insight can be enriched by visualising information in its geographic context. Two underlying approaches for sentiment analysis are dictionary-based and machine learning. The former is popular for public sentiment analysis, and the latter has found limited use for aggregating public sentiment from Twitter data. The research presented in this paper aims to extend the machine learning approach for aggregating public sentiment. To this end, a framework for analysing and visualising public sentiment from a Twitter corpus is developed. A dictionary-based approach and a machine learning approach are implemented within the framework and compared using one UK case study, namely the royal birth of 2013. The case study validates the feasibility of the framework for analysis and rapid visualisation. One observation is that there is good correlation between the results produced by the popular dictionary-based approach and the machine learning approach when large volumes of tweets are analysed. However, for rapid analysis to be possible faster methods need to be developed using big data techniques and parallel methods. Vu Dung Nguyen, Blesson Varghese, Adam Barker |
IEEE BigData | 2 |
| 2013 | QuPARA: Query-driven large-scale portfolio aggregate risk analysis on MapReduceabstractModern insurance and reinsurance companies use stochastic simulation techniques for portfolio risk analysis. Their risk portfolios may consist of thousands of reinsurance contracts covering millions of individually insured locations. To quantify risk and to help ensure capital adequacy, each portfolio must be evaluated in up to a million simulation trials, each capturing a different possible sequence of catastrophic events (e.g., earthquakes, hurricanes, etc.) over the course of a contractual year. We present a flexible framework for portfolio risk analysis that can answer a rich variety of catastrophic risk queries. Rather than aggregating simulation data in order to produce a small set of high-level risk metrics efficiently (as done in production risk management systems), our focus is on queries on unaggregated or partially aggregated data. The goal is to allow analysts to obtain answers to a wide variety of unanticipated but natural ad hoc queries, which can help actuaries or underwriters to better understand the multiple dimensions (e.g., spatial correlation, seasonality, peril features, construction features, financial terms, etc.) that can impact portfolio risk and thus company solvency. We implemented a prototype system, called QuPARA, using Apache's Hadoop implementation of the MapReduce paradigm. This allows the user to utilize large parallel compute servers in order to answer ad hoc queries efficiently even on very large data sets typically encountered in practice. We describe the design and implementation of QuPARA and present experimental results that demonstrate its feasibility. Andrew Rau-Chaplin, Blesson Varghese, Duane Wilson, Zhimin Yao, Norbert Zeh |
IEEE BigData | 2 |
| 2013 | Achieving Speedup in Aggregate Risk Analysis Using Multiple GPUsabstractStochastic simulation techniques employed for the analysis of portfolios of insurance/reinsurance risk, often referred to as `Aggregate Risk Analysis', can benefit from exploiting state-of-the-art high-performance computing platforms. In this paper, parallel methods to speed-up aggregate risk analysis for supporting real-time pricing are explored. An algorithm for analysing aggregate risk is proposed and implemented for multi-core CPUs and for many-core GPUs. Experimental studies indicate that GPUs offer a feasible alternative solution over traditional high-performance computing systems. A simulation of 1,000,000 trials with 1,000 catastrophic events per trial on a typical exposure set and contract structure is performed in less than 5 seconds on a multiple GPU platform. The key result is that the multiple GPU implementation can be used in real-time pricing scenarios as it is approximately 77x times faster than the sequential counterpart implemented on a CPU. Aman K. Bahl, Oliver Baltzer, Andrew Rau-Chaplin, Blesson Varghese, Aaron Whiteway |
ICPP | 4 |
| 2009 | Swarm pattern transformation methodologiesabstractThe work reported in this paper is motivated by the need for developing swarm pattern transformation methodologies. Two methods, namely a macroscopic method and a mathematical method are investigated for pattern transformation. The first method is based on macroscopic parameters while the second method is based on both microscopic and macroscopic parameters. A formal definition to pattern transformation considering four special cases of transformation is presented. Simulations on a physics simulation engine are used to confirm the feasibility of the proposed transformation methods. A brief comparison between the two methods is also presented. Blesson Varghese, Gerard T. McKee |
SIS | 1 |