VLDB 2026 Research / reviewers in the wild / expert
Ivona Brandic
dblp:71/2136
· DBLP profile ↗
74ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0001-7424-0208ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 2 first-author · 10 since 2021Software engineering, systems software and programming languages · 14 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 2 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Characterizing LLM Inference Energy-Performance Tradeoffs Across Workloads and GPU ScalingabstractLLM inference exhibits substantial variability across queries and execution phases, yet inference configurations are often applied uniformly. We present a measurement-driven characterization of workload heterogeneity and energy-performance behavior of LLM inference under GPU dynamic voltage and frequency scaling (DVFS). We evaluate five decoder-only LLMs (1B-32B parameters) across four NLP benchmarks using a controlled offline setup. We show that lightweight semantic features predict inference difficulty better than input length, with 44.5% of queries achieving comparable quality across model sizes. At the hardware level, the decode phase dominates inference time (77-91%) and is largely insensitive to GPU frequency. Consequently, reducing GPU frequency from 2842 MHz to 180 MHz achieves an average of 42% energy savings with only a 1-6% latency increase. We further provide a use case with an upper-bound analysis of the potential benefits of combining workload-aware model selection with phase-aware DVFS, motivating future energy-efficient LLM inference systems. Paul Joe Maliakel, Shashikant Ilager, Ivona Brandic |
CCGrid | 3 |
| 2026 | A Decentralized and Self-Adaptive Approach for Monitoring Volatile Edge EnvironmentsabstractEdge computing provides resources for IoT workloads at the network edge. Monitoring systems are vital for efficiently managing resources and application workloads by collecting, storing, and providing relevant information about the state of the resources. However, traditional monitoring systems have a centralized architecture for both data plane and control plane, which increases latency, creates a failure bottleneck, and faces challenges in providing quick and trustworthy data in volatile edge environments, especially where infrastructures are often built upon failure-prone, unsophisticated computing and network resources. Thus, we propose DEMon, a decentralized, self-adaptive monitoring system for edge. DEMon leverages the stochastic gossip communication protocol at its core. It develops efficient protocols for information dissemination, communication, and retrieval, avoiding a single point of failure and ensuring fast and trustworthy data access. Its decentralized control enables self-adaptive management of monitoring parameters, addressing the tradeoffs between the quality of service of monitoring and resource consumption. We implement the proposed system as a lightweight and portable container-based system and evaluate it through experiments. We also present a use case demonstrating its feasibility. The results show that DEMon efficiently disseminates and retrieves the monitoring information, addressing the challenges of edge monitoring. Shashikant Ilager, Jakob Fahringer, Alessandro Tundo, Ivona Brandic |
ACM Trans. Auton. Adapt. Syst. | 4 |
| 2025 | Green-Code: Learning to Optimize Energy Efficiency in Llm-Based Code GenerationabstractLarge Language Models (LLMs) are becoming integral to daily life, showcasing their vast potential across various Natural Language Processing (NLP) tasks. Beyond NLP, LLMs are increasingly used in software development tasks, such as code completion, modification, bug fixing, and code translation. Software engineers widely use tools like GitHub Copilot and Amazon Q, streamlining workflows and automating tasks with high accuracy. While the resource and energy intensity of LLM training is often highlighted, inference can be even more resourceintensive over time, as it's a continuous process with a high number of invocations. Therefore, developing resource-efficient alternatives for LLM inference is crucial for sustainability. This work proposes GREEN-CODE, a framework for energy-aware code generation in LLMs. GREEN-CODE performs dynamic early exit during LLM inference. We train a Reinforcement Learning (RL) agent that learns to balance the trade-offs between accuracy, latency, and energy consumption. Our approach is evaluated on two open-source LLMs, Llama 3.2 3B and OPT 2.7 B, using the JavaCorpus and PY150 datasets. Results show that our method reduces the energy consumption between 2350 % on average for code generation tasks without significantly affecting accuracy. Shashikant Ilager, Lukas Florian Briem, Ivona Brandic |
CCGrid | 3 |
| 2025 | On-Device Federated Learning for Remote Alpine Livestock Monitoring
Sabtain Ahmad, Thomas Schneidergruber, Ivona Brandic, Johannes Scholz |
Euro-Par (2) | 3 |
| 2025 | Exploring channel distinguishability in local neighborhoods of the model space in quantum neural networksabstractWith the increasing interest in Quantum Machine Learning, Quantum Neural Networks (QNNs) have emerged and gained significant attention. These models have, however, been shown to be notoriously difficult to train, which we hypothesize is partially due to the architectures, called ansatzes, that are hardly studied at this point. Therefore, in this paper, we take a step back and analyze ansatzes. We initially consider their expressivity, i.e., the space of operations they are able to express, and show that the closeness to being a 2-design, the primarily used measure, fails at capturing this property. Hence, we look for alternative ways to characterize ansatzes, unrelated to expressivity, by considering the local neighborhood of the model space, in particular, analyzing model distinguishability upon small perturbation of parameters. We derive an upper bound on their distinguishability, showcasing that QNNs using the Hardware Efficient Ansatz with few parameters are hardly discriminable upon update. Our numerical experiments support our bounds and further indicate that there is a significant degree of variability, which stresses the need for warm-starting or clever initialization. Altogether, our work provides an ansatz-centric perspective on training dynamics and difficulties in QNNs, ultimately suggesting that iterative training of small quantum models may not be effective, which contrasts their initial motivation. Sabrina Herbst, Sandeep Suresh Cranganore, Vincenzo De Maio, Ivona Brandic |
ICLR | 4 |
| 2025 | FRESCO: Fast and Reliable Edge Offloading With Reputation-Based Hybrid Smart ContractsabstractMobile devices offload latency-sensitive application tasks to edge servers to satisfy applications' Quality of Service (QoS) deadlines. Consequently, ensuring reliable offloading without QoS violations is challenging in distributed and unreliable edge environments with diverse resource and reliability levels. We propose FRESCO, a fast and reliable edge offloading framework that utilizes a blockchain-based reputation system, which enhances the reliability of offloading in the distributed edge. The distributed reputation system tracks the historical performance of edge servers, while blockchain through a consensus mechanism ensures that sensitive reputation information is secured against tampering. However, blockchain consensus typically has high latency, and therefore we employ a Hybrid Smart Contract (HSC) as areputation state managerthat automatically computes and stores reputation securely on-chain (i.e., on the blockchain) while allowing fast offloading decisions off-chain (i.e., outside of blockchain). Theoffloading decision engineuses a reputation score from HSC to derive fast offloading decisions, which are based on Satisfiability Modulo Theory (SMT). The SMT can formally guarantee a feasible solution that is valuable for latency-sensitive applications that require high reliability. With a combination of an on-chain HSC reputation state manager and an off-chain SMT decision engine, FRESCO offloads tasks to reliable servers without being hindered by blockchain consensus. In our experiment, FRESCO reduces response time by up to 7.86 times and saves energy by up to 5.4% compared to all baselines while minimizing QoS violations to 0.4% and achieving an average decision time of just 5.05 milliseconds. Josip Zilic, Vincenzo De Maio, Shashikant Ilager, Ivona Brandic |
IEEE Trans. Serv. Comput. | 4 |
| 2024 | Training Computer Scientists for the Challenges of Hybrid Quantum-Classical ComputingabstractAs we enter the post-Moore era, we experience the rise of various non-von-Neumann-architectures to address the increasing computational demand for modern applications, with quantum computing being among the most prominent and promising technologies. However, this development creates a gap in current computer science curricula since most quantum computing lectures are strongly physics-oriented and have little intersection with the remaining curriculum of computer science. This fact makes designing an appealing course very difficult, in particular for non-physicists. Furthermore, in the academic community, there is consensus that quantum computers are going to be used only for specific computational tasks (e.g., in computational science), where hybrid systems - combined classical and quantum computers - facilitate the execution of an application on both quantum and classical computing resources. A hybrid system thus executes only certain suitable parts of an application on the quantum machine, while other parts are executed on the classical components of the system. To fully exploit the capabilities of hybrid systems and to meet future requirements in this emerging field, we need to prepare a new generation of computer scientists with skills in both distributed computing and quantum computing. To bridge this existing gap in standard computer science curricula, we designed a new lecture and exercise series on Hybrid Quantum-Classical Systems, where students learn how to decompose applications and implement computational tasks on a hybrid quantum-classical computational continuum. While learning the inherent concepts underlying quantum systems, students are obligated to apply techniques and methods they are already familiar with, making the entrance to the field of quantum computing comprehensive yet appealing and accessible to students of computer science. Vincenzo De Maio, Meerzhan Kanatbekova, Felix Zilk, Nicolai Friis, Tobias Guggemos, Ivona Brandic |
CCGrid | 6 |
| 2024 | Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job AnalysisabstractHPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc. Xiaoyu Chu, Daniel Hofstätter, Shashikant Ilager, Sacheendra Talluri, Duncan Kampert, Damian Podareanu, Dmitry Duplyakin, Ivona Brandic, Alexandru Iosup |
ICPADS | 8 |
| 2024 | ABBA-VSM: Time Series Classification Using Symbolic Representation on the Edge
Meerzhan Kanatbekova, Shashikant Ilager, Ivona Brandic |
ICSOC (1) | 3 |
| 2024 | Paving the way to hybrid quantum-classical scientific workflows
Sandeep Suresh Cranganore, Vincenzo De Maio, Ivona Brandic, Ewa Deelman |
Future Gener. Comput. Syst. | 3 |
| 2023 | SymED: Adaptive and Online Symbolic Representation of Data on the Edge
Daniel Hofstätter, Shashikant Ilager, Ivan Lujic, Ivona Brandic |
Euro-Par | 4 |
| 2023 | An Energy-Aware Approach to Design Self-Adaptive AI-based Applications on the EdgeabstractThe advent of edge devices dedicated to machine learning tasks enabled the execution of AI-based applications that efficiently process and classify the data acquired by the resource-constrained devices populating the Internet of Things. The proliferation of such applications (e.g., critical monitoring in smart cities) demands new strategies to make these systems also sustainable from an energetic point of view. In this paper, we present an energy-aware approach for the design and deployment of self-adaptive AI-based applications that can balance application objectives (e.g., accuracy in object detection and frames processing rate) with energy consumption. We address the problem of determining the set of configurations that can be used to self-adapt the system with a meta-heuristic search procedure that only needs a small number of empirical samples. The final set of configurations are selected using weighted gray relational analysis, and mapped to the operation modes of the self-adaptive application. We validate our approach on an AI-based application for pedestrian detection. Results show that our self-adaptive application can outperform non-adaptive baseline configurations by saving up to 81% of energy while loosing only between 2% and 6 % in accuracy. Alessandro Tundo, Marco Mobilio, Shashikant Ilager, Ivona Brandic, Ezio Bartocci, Leonardo Mariani |
ASE | 4 |
| 2023 | Sustainable Environmental Monitoring via Energy and Information Efficient Multinode PlacementabstractThe Internet of Things is gaining traction for sensing and monitoring outdoor environments such as water bodies, forests, or agricultural lands. Sustainable deployment of sensors for environmental sampling is a challenging task because of the spatial and temporal variation of the environmental attributes to be monitored, the lack of the infrastructure to power the sensors for uninterrupted monitoring, and the large continuous target environment despite the sparse and limited sampling locations. In this paper, we present an environment monitoring framework that deploys a network of sensors and gateways connected through low-power, long-range networking to perform reliable data collection. The three objectives correspond to the optimization of information quality, communication capacity, and sustainability. Therefore, the proposed environment monitoring framework consists of three main components: (i) to maximize the information collected, we propose an optimal sensor placement method based on QR decomposition that deploys sensors at information-and communication-critical locations; (ii) to facilitate the transfer of big streaming data and alleviate the network bottleneck caused by low bandwidth, we develop a gateway configuration method with the aim to reduce the deployment and communication costs; and (iii) to allow sustainable environmental monitoring, an energy-aware optimization component is introduced. We validate our method by presenting a case study for monitoring the water quality of the Ergene River in Turkey. Detailed experiments subject to real-world data show that the proposed method is both accurate and efficient in monitoring a large environment and catching up with dynamic changes. Sabtain Ahmad, Halit Uyanik, Tolga Ovatman, Mehmet Tahir Sandikkaya, Vincenzo De Maio, Ivona Brandic, Atakan Aral |
IEEE Internet Things J. | 6 |
| 2022 | Molecular Dynamics Workflow Decomposition for Hybrid Classic/Quantum SystemsabstractSince we are entering the Post-Moore Law era and consequently the limit of Von Neumann's architecture, the scientific community is looking for alternatives to satisfy the growing computing power demands of scientific applications. Quantum computing promises to achieve a computational advantage over the classic Von Neumann architecture. However, the limited capabilities of current noisy intermediate-scale quantum (NISQ) devices require quantum computers to interoperate with classic systems, forming the so-called hybrid quantum systems. Research on hybrid quantum systems led to the design of Variational Quantum Algorithms, currently the most promising way to move towards quantum advantage. However, execution time and accuracy of variational quantum algorithms are affected by different hyperparameters, including selected cost functions and parametrized quantum circuits. Consequently, providing developers with methods to select the right set of parameters is of paramount importance. In this work, we provide a formal method for the selection of hyperparameters in variational quantum algorithms, which will support quantum algorithms developers in the design of quantum applications, and evaluate it on a real-world scientific application, showing a reduction of error up to 31%. Sandeep Suresh Cranganore, Vincenzo De Maio, Ivona Brandic, Tu Mai Anh Do, Ewa Deelman |
e-Science | 3 |
| 2022 | Edge Workload Trace Gathering and Analysis for BenchmarkingabstractThe emerging field of edge computing is suffering from a lack of representative data to evaluate rapidly introduced new algorithms or techniques. That is a critical issue as this complex paradigm has numerous different use cases which translate into a highly diverse set of workload types.In this work, within the context of the edge computing activity of SPEC RG Cloud, we continue working towards an edge benchmark by defining high-level workload classes as well as collecting and analyzing traces for three real-world edge applications, which, according to the existing literature, are the representatives of those classes. Moreover, we propose a practical and generic methodology for workload definition and gathering. The traces and gathering tool are provided open-source.In the analysis of the collected workloads, we detect discrepancies between the literature and the traces obtained, thus highlighting the need for a continuing effort into gathering and providing data from real applications, which can be done using the proposed trace gathering methodology. Additionally, we discuss various insights and future directions that rise to the surface through our analysis. Klervie Toczé, Norbert Schmitt, Ulf Kargén, Atakan Aral, Ivona Brandic |
ICFEC | 5 |
| 2022 | Data Science Driven Methods for Sustainable and Failure Tolerant Edge SystemsabstractNowadays we experience a paradigm shift in our society, where every item around us is becoming a computer facilitating life-changing applications like self-driving cars, tele-medicine, precision agriculture or virtual reality. On one hand, for the execution of such resource demanding applications we need powerful IT facilities. On the other hand, the requirements often include latencies below 100 ms or even below 10 ms -- what is called ''tactile internet''. To facilitate low latency computation has to be placed in the vicinity of the end users by utilizing the concept of Edge Computing. In this talk we explain the challenges of Edge systems in combination with tactile internet. We discuss the recent problems of geographically distributed machine learning applications and novel approaches to balance competing priorities like the energy efficiency and the staleness of the machine learning models. Available failure resilience mechanisms designed for Cloud computing or generic distributed systems cannot be applied to Edge systems due to timeliness, hyper heterogeneity and resource scarcity. Therefore, we discuss a novel machine learning based mechanism that evaluates the failure resilience of a service deployed redundantly on the edge infrastructure. Our approach learns the spatiotemporal dependencies between edge server failures and combines them with the topological information to incorporate link failures by utilizing the concept of the Dynamic Bayesian Networks (DBNs). Eventually, we infer the probability that a certain set of servers fails or disconnects concurrently during service runtime. Ivona Brandic |
ICPE | 1 |
| 2022 | Multiagent Bayesian Deep Reinforcement Learning for Microgrid Energy Management Under Communication FailuresabstractMicrogrids (MGs) are important players for the future transactive energy systems where a number of intelligent Internet of Things (IoT) devices interact for energy management in the smart grid. Although there have been many works on MG energy management, most studies assume a perfect communication environment, where communication failures are not considered. In this article, we consider the MG as a multiagent environment with IoT devices in which AI agents exchange information with their peers for collaboration. However, the collaboration information may be lost due to communication failures or packet loss. Such events may affect the operation of the whole MG. To this end, we propose a multiagent Bayesian deep reinforcement learning (BA-DRL) method for MG energy management under communication failures. We first define a multiagent partially observable Markov decision process (MA-POMDP) to describe agents under communication failures, in which each agent can update its beliefs on the actions of its peers. Then, we apply a double deep$Q$-learning (DDQN) architecture for$Q$-value estimation in BA-DRL, and propose a belief-based correlated equilibrium for the joint-action selection of multiagent BA-DRL. Finally, the simulation results show that BA-DRL is robust to both power supply uncertainty and communication failure uncertainty. BA-DRL has 4.1% and 10.3% higher reward than Nash deep$Q$-learning (Nash-DQN) and alternating direction method of multipliers (ADMM), respectively, under 1% communication failure probability. Hao Zhou 0013, Atakan Aral, Ivona Brandic, Melike Erol-Kantarci |
IEEE Internet Things J. | 3 |
| 2022 | SEA-LEAP: Self-Adaptive and Locality-Aware Edge Analytics PlacementabstractNear real-time edge analytics requires dealing with the rapidly growing amount of data, limited resources, and high failure probabilities of edge nodes. Therefore, data replication is of vital importance to meet SLOs such as service availability and failure resilience. Consequently, specific input datasets, requested by on-demand analytics (e.g., object detection), can be present at different locations over time. This can prevent exploitation of data locality and timely decision-making processes. State-of-the-art solutions for on-demand edge analytics placement either fail in providing low-latency access to user-requested input data or do not consider data locality. We propose SEA-LEAP (Self-adaptive and Locality-aware Edge Analytics Placement), a framework including a new mechanism for tracking data movements, on top of which we devise a generic control mechanism. SEA-LEAP enables on-the-fly placement of on-demand analytics considering the most appropriate dataset location that minimizes overall analytics requests execution time. We conduct experiments using real-world (i) object detection application, (ii) image datasets as input, (iii) self-designed benchmarks, and (iv) heterogeneous edge infrastructure using Kubernetes. Experimental results show the ability to efficiently deploy on-demand analytics and reduce total latency by 65.85 percent on average by performing adaptive data movements, indicating a promising solution for edge multi-cluster and hybrid environments. Ivan Lujic, Vincenzo De Maio, Srikumar Venugopal, Ivona Brandic |
IEEE Trans. Serv. Comput. | 4 |
| 2022 | ARES: Reliable and Sustainable Edge Provisioning for Wireless Sensor NetworksabstractWireless sensor networks have wide applications in monitoring applications. However, sensors’ energy and processing power constraints, as well as the limited network bandwidth, constitute significant obstacles to near-real-time requirements of modern IoT applications. Offloading sensor data on an edge computing infrastructure instead of in-cloud or in-network processing is a promising solution to these issues. Nevertheless, due to geographical dispersion, ad-hoc deployment, and rudimentary support systems compared to cloud data centers, reliability is a critical issue. This forces edge service providers to deploy a huge amount of edge nodes over an urban area, with catastrophic effects on environmental sustainability. In this work, we propose ARES, a two-stage optimization algorithm for sustainable and reliable deployment of edge nodes in an urban area. Initially, ARES applies multi-objective optimization to identify a set of Pareto-optimal solutions for transmission time and energy; then it augments these candidates in the second stage to identify a solution that guarantees the desired level of reliability using a dynamic Bayesian network based reliability model. ARES is evaluated through simulations using data from the urban area of Vienna. Results demonstrate that it can achieve a better trade-off between transmission time, energy-efficiency, and reliability than the state-of-the-art solutions. Atakan Aral, Vincenzo De Maio, Ivona Brandic |
IEEE Trans. Sustain. Comput. | 3 |
| 2021 | On the Future of Cloud EngineeringabstractEver since the commercial offerings of the Cloud started appearing in 2006, the landscape of cloud computing has been undergoing remarkable changes with the emergence of many different types of service offerings, developer productivity enhancement tools, and new application classes as well as the manifestation of cloud functionality closer to the user at the edge. The notion of utility computing, however, has remained constant throughout its evolution, which means that cloud users always seek to save costs of leasing cloud resources while maximizing their use. On the other hand, cloud providers try to maximize their profits while assuring service-level objectives of the cloud-hosted applications and keeping operational costs low. All these outcomes require systematic and sound cloud engineering principles. The aim of this paper is to highlight the importance of cloud engineering, survey the landscape of best practices in cloud engineering and its evolution, discuss many of the existing cloud engineering advances, and identify both the inherent technical challenges and research opportunities for the future of cloud computing in general and cloud engineering in particular. David Bermbach, Abhishek Chandra, Chandra Krintz, Aniruddha S. Gokhale, Aleksander Slominski, Lauritz Thamsen, Everton Cavalcante, Tian Guo 0001, Ivona Brandic, Richard Wolski |
IC2E | 9 |
| 2021 | Learning Spatiotemporal Failure Dependencies for Resilient Edge Computing ServicesabstractEdge computing services are exposed to infrastructural failures due to geographical dispersion, ad hoc deployment, and rudimentary support systems. Two unique characteristics of the edge computing paradigm necessitate a novel failure resilience approach. First, edge servers, contrary to cloud counterparts with reliable data center networks, are typically connected via ad hoc networks. Thus, link failures need more attention to ensure truly resilient services. Second, network delay is a critical factor for the deployment of edge computing services. This restricts replication decisions to geographical proximity and necessitates joint consideration of delay and resilience. In this article, we propose a novel machine learning based mechanism that evaluates the failure resilience of a service deployed redundantly on the edge infrastructure. Our approach learns the spatiotemporal dependencies between edge server failures and combines them with the topological information to incorporate link failures. Ultimately, we infer the probability that a certain set of servers fails or disconnects concurrently during service runtime. Furthermore, we introduce Dependency- and Topology-aware Failure Resilience (DTFR), a two-stage scheduler that minimizes either failure probability or redundancy cost, while maintaining low network delay. Extensive evaluation with various real-world failure traces and workload configurations demonstrate superior performance in terms of availability, number of failures, network delay, and cost with respect to the state-of-the-art schedulers. Atakan Aral, Ivona Brandic |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | Reliability Management for Blockchain-Based Decentralized Multi-CloudabstractBlockchain-based decentralized multi-cloud has the potential to reduce cloud infrastructure costs and to enable geographically distributed providers of any size to monetize their computational resources. In this context, guarantees that the computational results are delivered within the promised time and budget must be provided despite the limited information available about the location and ownership of resources. Providers might claim to execute the services to get compensated for the computation even though returning incomplete or incorrect results. In this paper, we define a model to predict provider reliability, that is, the probability of failure-free execution of computational tasks and correctness of the computed outputs, by extracting the potential dependencies between providers from historical log traces. This model can then be utilized in the definition of provider reputation or the scheduling of new services. Indeed, we propose a probabilistic scheduler that chooses the providers that meet the reliability constraints among others. Finally, we validate the proposed solutions with real traces from a decentralized cloud provider and hint at the benefits of predicting reliability in this context. Atakan Aral, Rafael Brundo Uriarte, Anthony Simonet, Ivona Brandic |
CCGRID | 4 |
| 2020 | Performance-Based Pricing in Multi-Core Geo-Distributed Cloud ComputingabstractNew pricing policies are emerging where cloud providers charge resource provisioning based on the allocated CPU frequencies. As a result, resources are offered to users as combinations of different performance levels and prices which can be configured at runtime. With such new pricing schemes and the increasing energy costs in data centres, balancing energy savings with performance and revenue losses is a challenging problem for cloud providers. CPU frequency scaling can be used to reduce power dissipation, but also impacts virtual machine (VM) performance and therefore revenue. In this paper, we first propose a non-linear power model that estimates power dissipation of a multi-core CPU physical machine (PM) and second a pricing model that adjusts the pricing based on the VM's CPU-boundedness characteristics. Finally, we present a cloud controller that uses these models to allocate VM and scale CPU frequencies of the physical machine (PM) to achieve energy cost savings that exceed service revenue losses. We evaluate the proposed approach using simulations with realistic VM workloads, electricity price and temperature traces and estimate energy savings of up to 14.57 percent. Drazen Lucanin, Ilia Pietri, Simon Holmbacka, Ivona Brandic, Johan Lilius, Rizos Sakellariou |
IEEE Trans. Cloud Comput. | 4 |
| 2020 | Resilient Edge Data Management FrameworkabstractTransferring and processing huge amounts of data in the cloud can violate the low latency requirements of modern IoT applications, considering underlying network infrastructure limitations. Edge data analytics is a promising solution. However, edge resources have usually less computational capabilities than cloud nodes, resulting in a higher failure rate of IoT systems. Consequently, near-real-time decisions are often based on limited and incomplete data. State-of-the-art solutions, such as operational/workload flows, data reduction, reconstruction, focus mostly on resource and network optimization, while approaches for incomplete data recovery employ a single specific method, despite diverse data characteristics. Data quality impact on accuracy of the decision-making processes is often neglected. We propose EDMFrame, a framework featuring a generic mechanism for recovery of multiple gaps in incomplete datasets, using single-technique recovery (STR) and multiple-technique recovery (MTR) involving projection recovery maps (PRMs). We further devise an adaptive storage management mechanism for reducing data stored at the edge, keeping only the data necessary for predictive analytics. We conduct experiments using time series from smart buildings, (i) automatically recovering various multiple gaps and reducing errors up to 65.48 percent with MTR compared to STR; (ii) reducing amounts of data stored to 39.9 percent on average, keeping prediction accuracy around 98.83 percent. Ivan Lujic, Vincenzo De Maio, Ivona Brandic |
IEEE Trans. Serv. Comput. | 3 |
| 2019 | Multi-Objective Mobile Edge Provisioning in Small Cell CloudsabstractIn recent years, Mobile Cloud Computing (MCC) has been proposed as a solution to enhance the capabilities of user equipment (UE), such as smartphones, tablets and laptops. However, offloading to conventional Cloud introduces significant execution delays that are inconvenient in case of near real-time applications. Mobile Edge Computing (MEC) has been proposed as a solution to this problem. MEC brings computational and storage resources closer to the UE, enabling to offload near real-time applications from the UE while meeting strict latency requirements. However, it is very difficult for Edge providers to determine how many Edge nodes are required to provide MEC services, in order to guarantee a high QoS and to maximize their profit. In this paper, we investigate the static provisioning of Edge nodes in a area representing a cellular network in order to guarantee the required QoS to the user without affecting providers' profits. First, we design a model for MEC offloading considering user satisfaction and provider's costs. Then, we design a simulation framework based on this model. Finally, we design a multi-objective algorithm to identify a deployment solution that is a trade-off between user satisfaction and provider profit. Results show that our algorithm can guarantee a user satisfaction above 80%, with a profit for the provider of up 4 times their cost. Vincenzo De Maio, Ivona Brandic |
ICPE | 2 |
| 2019 | Addressing Application Latency Requirements through Edge SchedulingabstractAbstract Latency-sensitive and data-intensive applications, such as IoT or mobile services, are leveraged by Edge computing, which extends the cloud ecosystem with distributed computational resources in proximity to data providers and consumers. This brings significant benefits in terms of lower latency and higher bandwidth. However, by definition, edge computing has limited resources with respect to cloud counterparts; thus, there exists a trade-off between proximity to users and resource utilization. Moreover, service availability is a significant concern at the edge of the network, where extensive support systems as in cloud data centers are not usually present. To overcome these limitations, we propose a score-based edge service scheduling algorithm that evaluates network, compute, and reliability capabilities of edge nodes. The algorithm outputs the maximum scoring mapping between resources and services with regard to four critical aspects of service quality. Our simulation-based experiments on live video streaming services demonstrate significant improvements in both network delay and service time. Moreover, we compare edge computing with cloud computing and content delivery networks within the context of latency-sensitive and data-intensive applications. The results suggest that our edge-based scheduling algorithm is a viable solution for high service quality and responsiveness in deploying such applications. Atakan Aral, Ivona Brandic, Rafael Brundo Uriarte, Rocco De Nicola, Vincenzo Scoca |
J. Grid Comput. | 2 |
| 2018 | First Hop Mobile Offloading of DAG ComputationsabstractIn recent years, Mobile Cloud Computing (MCC) has been proposed to increase battery lifetime of mobile devices. However, offloading on Cloud infrastructures may be infeasible for latency critical applications, due to the geographical distribution of Cloud data centers that increases offloading time. In this paper, we investigate the use of Mobile Edge Cloud Offloading (MECO), namely offloading to a heterogeneous computing infrastructure featuring both Cloud and Edge nodes, where Edge nodes are geographically closer to the mobile device. We evaluate improvements of MECO in comparison with MCC for objectives such as applications' runtime, mobile device battery lifetime and cost for the user. Afterwards, we propose the Edge Cloud Heuristic Offloading (ECHO) approach to find a trade-off solution between the aforementioned objectives, according to user's preferences. We evaluate our approach by simulating offloading of Directed Acyclic Graphs (DAGs) representing mobile applications through the use of Monte-Carlo simulations. The results show that (1) MECO can reduce application runtime by up to 70.7% and cost by up to 70.6% in comparison to MCC and (2) ECHO allows user to select a trade-off solution with at most 18% MAPE for runtime, 16% for cost and 0.5% for battery lifetime, according to user's preferences. Vincenzo De Maio, Ivona Brandic |
CCGrid | 2 |
| 2018 | Scheduling Latency-Sensitive Applications in Edge Computing
Vincenzo Scoca, Atakan Aral, Ivona Brandic, Rocco De Nicola, Rafael Brundo Uriarte |
CLOSER | 3 |
| 2018 | Adaptive Recovery of Incomplete Datasets for Edge AnalyticsabstractThe Internet of Things (IoT) has attracted significant attention from both academia and industry, thanks to applications such as smart cities, smart buildings and intelligent traffic management. These systems rely on data, collected from IoT devices, that are sent to the cloud for analytics. Data are either used for near real-time decisions or stored for long-term analysis. However, in highly distributed IoT systems, missing or invalid data may appear because of different reasons including sensor failures, monitoring system failures and network failures. Analyzing incomplete datasets can lead to inaccurate results and imprecise decisions, with negative effects on the target systems. Also, due to the increasing size of such systems and the consequently increasing amount of data generated from sensors, recovery of incomplete datasets for analytics on the cloud is often infeasible, due to the limited bandwidth available and the strict latency constraints of IoT applications. We propose a novel semi-automatic recursive mechanism for recovery of incomplete datasets on the edge that is closer to the source of data. This mechanism enables efficient recovery of incomplete datasets employing different forecasting techniques for multiple gaps, based on user specifications. We evaluate our approach on datasets coming from the context of smart buildings and smart homes. The experimental results show that our approach is able to identify multiple gaps, then recover incomplete datasets, decreasing forecasting error by up to 82.68%, and reducing running time by up to 52.38%. Ivan Lujic, Vincenzo De Maio, Ivona Brandic |
ICFEC | 3 |
| 2017 | Efficient Edge Storage Management Based on Near Real-Time ForecastsabstractNowadays, data analytics is utilized on edge based systems to perform near real-time decisions in proximity of the user. When performing near real-time decisions on the Edge, we need historical data to perform accurate data analytics. Since storage capacities on the Edge are limited, we are faced with a challenge to balance the quantity of data stored with the quality of near real-time decisions. In this paper, we present a three-layer architecture model for data storage management on the Edge including an adaptive algorithm that dynamically finds a trade-off between providing high forecast accuracy necessary for efficient real-time decisions, and minimizing the amount of data stored in the space-limited storage. We focus on time series data, typical in the context of sensor-based monitoring in IoT environments. By using the proposed approach it is possible to reduce the amount of stored data by an average 80.27% without affecting specified threshold for prediction accuracy. Ivan Lujic, Vincenzo De Maio, Ivona Brandic |
ICFEC | 3 |
| 2016 | A hybrid cloud controller for vertical memory elasticity: A control-theoretic approach
Soodeh Farokhi, Pooyan Jamshidi, Ewnetu Bayuh Lakew, Ivona Brandic, Erik Elmroth |
Future Gener. Comput. Syst. | 4 |
| 2016 | Towards uniform management of multi-layered cloud services by applying model-driven development
Toni Mastelic, Andrés García-García, Ivona Brandic |
J. Syst. Softw. | 3 |
| 2016 | Pervasive Cloud Controller for Geotemporal InputsabstractThe rapid cloud computing growth has turned data center energy consumption into a global problem. At the same time, modern cloud providers operate multiple geographically-distributed data centers. Distributed data center infrastructure changes the rules of cloud control, as energy costs depend on current regional electricity prices and temperatures. Furthermore, to account for emerging technologies surrounding the cloud ecosystem, a maintainable control solution needs to be forward-compatible. Existing cloud controllers are focused on VM consolidation methods suitable only for a single data center or consider migration just in case of workload peaks, not accounting for all the aspects of geographically distributed data centers. In this paper, we propose a pervasive cloud controller for dynamic resource reallocation adapting to volatile time- and location-dependent factors, while considering the QoS impact of too frequent migrations and the data quality limits of time series forecasting methods. The controller is designed with extensible decision support components. We evaluate it in a simulation using historical traces of electricity prices and temperatures. By optimising for these additional factors, we estimate 28.6 percent energy cost savings compared to baseline dynamic VM consolidation. We provide a range of guidelines for cloud providers, showing the environment conditions necessary to achieve significant cost savings and we validate the controllers extensibility. Drazen Lucanin, Ivona Brandic |
IEEE Trans. Cloud Comput. | 2 |
| 2015 | A Cloud Controller for Performance-Based PricingabstractNew dynamic cloud pricing options are emerging with cloud providers offering resources as a wide range of CPU frequencies and matching prices that can be switched at runtime. On the other hand, cloud providers are facing the problem of growing operational energy costs. This raises a trade-off problem between energy savings and revenue loss when performing actions such as CPU frequency scaling. Although existing cloud controllers for managing cloud resources deploy frequency scaling, they only consider fixed virtual machine (VM) pricing. In this paper we propose a performance-based pricing model adapted for VMs with different CPU-bounded ness properties. We present a cloud controller that scales CPU frequencies to achieve energy cost savings that exceed service revenue losses. We evaluate the approach in a simulation based on real VM workload, electricity price and temperature traces, estimating energy cost savings up to 32% in certain scenarios. Drazen Lucanin, Ilia Pietri, Ivona Brandic, Rizos Sakellariou |
CLOUD | 3 |
| 2015 | Data Velocity Scaling via Dynamic Monitoring Frequency on Ultrascale InfrastructuresabstractMonitoring ultrascale systems such as Clouds requires collecting enormous amount of data by periodically reading metric values from a system. Current approaches tend to select a static frequency for sampling monitoring data. On one hand, over-sampling the data by collecting it at high frequencies results in data redundancy during steady runs of the system. On the other hand, under-sampling with low monitoring frequencies results in information loss during volatile behaviour of the system as data is significantly diluted. Therefore, choosing an optimal monitoring frequency represents a challenging research issue. In this paper, we propose a dynamic monitoring frequency algorithm for collecting monitoring data from ultrascale systems such as Clouds. The algorithm deterministically reduces data velocity by self-adapting the monitoring frequency to the volatility of data being collected. Consequently, it collects less data due to fewer readings, while keeping the same data value as the equivalent static monitoring frequency. The proposed approach is evaluated using Google traces where it is able to reduce the velocity of monitoring data by up to 85% without diluting information quality. Toni Mastelic, Ivona Brandic |
CloudCom | 2 |
| 2015 | Predicting Resource Allocation and Costs for Business Processes in the CloudabstractBy moving business processes into the cloud, business partners can benefit from lower costs, more flexibility and greater scalability in terms of resources offered by the cloud providers. In order to execute a process or a part of it, a business process owner selects and leases feasible resources while considering different constraints, e.g., Optimizing resource requirements and minimizing their costs. In this context, utilizing information about the process models or the dependencies between tasks can help the owner to better manage leased resources. In this paper, we propose a novel resource allocation technique based on the execution path of the process, used to assist the business process owner in efficiently leasing computing resources. The technique comprises three phases, namely process execution prediction, resource allocation and cost estimation. The first exploits the business process model metrics and attributes in order to predict the process execution and the requires resources, while the second utilizes this prediction for efficient allocation of the cloud resources. The final phase estimates and optimizes costs of leased resources by combining different pricing models offered by the provider. Toni Mastelic, Walid Fdhila, Ivona Brandic, Stefanie Rinderle-Ma |
SERVICES | 3 |
| 2015 | SLA enactment for large-scale healthcare workflows on multi-Cloud
Foued Jrad, Jie Tao 0001, Ivona Brandic, Achim Streit |
Future Gener. Comput. Syst. | 3 |
| 2014 | HS4MC - Hierarchical SLA-based Service Selection for Multi-Cloud EnvironmentsabstractCloud computing popularity is growing rapidly and consequently the number of companies offering their services in the form of Software-as-a-Service (SaaS) or Infrastructure-as-a-Service (IaaS) is increasing. The diversity and usage benefits of IaaS offers are encouraging SaaS providers to lease resources from the Cloud instead of operating their own data centers. However, the question remains for them how to, on the one hand, exploit Cloud benefits to gain less maintenance overheads and on the other hand, maximize the satisfactions of customers with a wide range of requirements. The complexity of addressing these issues prevent many SaaS providers to benefit from the Cloud infrastructures. In this paper, we propose HS4MC approach for automatic service selection by considering SLA claims of SaaS providers. The novelty of our approach lies in the utilization of prospect theory for the service ranking that represents a natural choice for scoring of comparable services due to the users preferences. The HS4MC approach first constructs a set of SLAs based on the given accumulated SaaS provider requirements. Then, it selects a set of services that best fulfills the SLAs. We evaluate our approach in a simulated environment by comparing it with a state-of-the-art utility- based algorithm. The evaluation results show that our approach selects services that more effectively satisfy the SLAs. Soodeh Farokhi, Foued Jrad, Ivona Brandic, Achim Streit |
CLOSER | 3 |
| 2014 | Multi-dimensional Resource Allocation for Data-intensive Large-scale Cloud ApplicationsabstractLarge scale applications are emerged as one of the important applications in distributed computing. Today, the economic and technical benefits offered by the Cloud computing technology encouraged many users to migrate their applications to Cloud. On the other hand, the variety of the existing Clouds requires them to make decisions about which providers to choose in order to achieve the expected performance and service quality while keeping the payment low. In this paper, we present a multi-dimensional resource allocation scheme to automate the deployment of data-intensive large scale applications in Mutli-Cloud environments. The scheme applies a two level approach in which the target Clouds are matched with respect to the Service Level Agreement (SLA) requirements and user payment at first and then the application workloads are distributed to the selected Clouds using a data locality driven scheduling policy. Using an implemented Multi-Cloud simulation environment, we evaluated our approach with a real data-intensive workflow application in different scenarios. The experimental results demonstrate the effectiveness of the implemented matching and scheduling policies in improving the workflow execution performance and reducing the amount and costs of Intercloud data transfers. Foued Jrad, Jie Tao 0001, Ivona Brandic, Achim Streit |
CLOSER | 3 |
| 2014 | CPU Performance Coefficient (CPU-PC): A Novel Performance Metric Based on Real-Time CPU Resource Provisioning in Time-Shared Cloud EnvironmentsabstractThe Cloud represents an emerging paradigm that provides on-demand computing resources, such as CPU. The resources are customized in quantity through various virtual machine (VM) flavours, which are deployed on top of time-shared infrastructure, where a single server can host several VMs. However, their Quality of Service (QoS) is limited and boils down to the VM availability, which does not provide any performance guarantees for the shared underlying resources. Consequently, the providers usually over-provision their resources trying to increase utilization, while the customers can suffer from poor performance due to increased concurrency. In this paper, we introduce CPU Performance Coefficient (CPU-PC), a novel performance metric used for measuring the real-time quality of CPU provisioning in virtualized environments. The metric isolates an impact of the provisioned CPU on the performance of the customer's application, hence allowing the provider to measure the quality of provisioned resources and manage them accordingly. Additionally, we provide a measurement of the proposed metric for the customer as well, thus enabling the latter to monitor the quality of rented resources. As evaluation, we utilize three real world applications used in existing Cloud services, and correlate the CPU-PC metric with the response time of the applications. An R-squared correlation of over 0.9557 indicates the applicability of our approach in the real world. Toni Mastelic, Ivona Brandic, Jasmina Jaarevic |
CloudCom | 2 |
| 2014 | Towards Uniform Management of Cloud Services by Applying Model-Driven DevelopmentabstractPopularity of Cloud Computing produced the birth of Everything-as-a-Service (XaaS) concept, where each service can comprise large variety of software and hardware elements. Although having the same concept, each of these services represent complex system that have to be deployed and managed by a provider using individual tools for almost every element. This usually leads to a combination of different deployment tools that are unable to interact with each other in order to provide an unified and automatic service deployment procedure. Therefore, the tools are usually used manually or specifically integrated for a single cloud service, which on the other hand requires changing the entire deployment procedure in case the service gets modified. In this paper we utilize Model-driven development (MDD) approach for building and managing arbitrary cloud services. We define a metamodel of a cloud service called CoPS, which describes a cloud service as a composition of software and hardware elements by using three sequential models, namely Component, Product and Service. We also present an architecture of a Cloud Management System (CMS) that is able to manage such services by automatically transforming the service models from the abstract representation to the actual deployment. Finally, we validate our approach by realizing three real world use cases using a prototype implementation of the proposed CMS architecture. Toni Mastelic, Ivona Brandic, Andrés García-García |
COMPSAC | 2 |
| 2014 | Democratization in Science and Technology through Cloud Computing
Ivona Brandic |
SECRYPT | 1 |
| 2014 | An interoperable and self-adaptive approach for SLA-based service virtualization in heterogeneous Cloud environments
Attila Kertész, Gabor Kecskemeti, Ivona Brandic |
Future Gener. Comput. Syst. | 3 |
| 2013 | TimeCap: Methodology for Comparing IT Infrastructures Based on Time and Capacity MetricsabstractScientific community is one of the major driving forces for developing and utilizing IT technologies such as Supercomputers and Grid. Although, the main race has always been for bigger and faster infrastructures, an easier access to such infrastructures in recent years created a demand for more customizable and scalable environments. However, introducing new technologies and paradigms such as Cloud computing requires a comprehensive analysis of its benefits before the actual implementation. In this paper we introduce the TimeCap, a methodology for comparing IT infrastructures based on time requirements and resource capacity wastage. We go beyond comparing just the execution time by introducing the Time Complexity as part of TimeCap, a methodology used for comparing arbitrary time related tasks required for completing a procedure, i.e., obtaining the scientific results. Moreover, a resource capacity wastage is compared using the Discrete Capacity, a second methodology as part of TimeCap used for analyzing resource assignment and utilization. We evaluate our methodology by comparing a traditional physical infrastructure and a Cloud infrastructure using our local IT resources. We use a real world scientific application for calculating plasma instabilities for analyzing time and capacity required for computation. Toni Mastelic, Ivona Brandic |
IEEE CLOUD | 2 |
| 2013 | Reducing Energy Consumption by using Clouds
Ivona Brandic |
CLOSER | 1 |
| 2013 | Cloud resource provisioning and SLA enforcement via LoM2HiS frameworkabstractSUMMARY Cloud computing represents a novel on‐demand computing technology where resources are provisioned in compliance to a set of predefined non‐functional properties specified and negotiated by means of service level agreements (SLAs). Currently, cloud providers strive to achieve efficient SLA enforcement strategies to avoid costly SLA violations during application provisioning and to timely react to failures and environmental changes. These strategies include advanced application deployment mechanisms and appropriate resource monitoring concepts. In terms of cloud resource monitoring, providers tend to adopt existing monitoring tools, such as those from grid environments. However, those tools are usually restricted to locality and homogeneity of monitored objects, are not scalable, and do not support mapping of low‐level resource metrics (e.g., system uptime and downtime) to high‐level application‐specific SLA parameters (e.g., system availability). In this paper, we present a novel low‐level metrics to high‐level SLA (LoM2HiS) framework for managing the monitoring of low‐level resource metrics and mapping them to high‐level SLAs and an application deployment mechanism for scheduling and provisioning applications in clouds. The LoM2HiS framework provides the application deployment mechanism with monitored information and SLA violation prevention techniques, thereby ensuring the performance of the applications and thus increasing the revenue of the cloud provider by avoiding SLA violation penalty cost. This framework is the building block of the Foundations of Self‐governing ICT Infrastructures project, which intends to facilitate autonomic SLA management and enforcement. Thus, the LoM2HiS framework detects future SLA violation threats and can notify the knowledge component to act so as to avert the threats. We discuss in detail the conceptual design of the LoM2HiS framework and the application deployment mechanism including their implementations. Finally, we present our evaluation results based on a use‐case scenario demonstrating the usage of the LoM2HiS framework in a real cloud environment. Copyright © 2012 John Wiley & Sons, Ltd. Vincent C. Emeakaroha, Ivona Brandic, Michael Maurer, Schahram Dustdar |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Creating standardized products for electronic markets
Ivan Breskovic, Jörn Altmann, Ivona Brandic |
Future Gener. Comput. Syst. | 3 |
| 2013 | Adaptive resource configuration for Cloud infrastructure managementabstractTo guarantee the vision of Cloud Computing QoS goals between the Cloud provider and the customer have to be dynamically met. This so-called Service Level Agreement (SLA) enactment should involve little human-based interaction in order to guarantee the scalability and efficient resource utilization of the system. To achieve this we start from Autonomic Computing, examine the autonomic control loop and adapt it to govern Cloud Computing infrastructures. We first hierarchically structure all possible adaptation actions into so-called escalation levels. We then focus on one of these levels by analyzing monitored data from virtual machines and making decisions on their resource configuration with the help of knowledge management (KM). The monitored data stems both from synthetically generated workload categorized in different workload volatility classes and from a real-world scenario: scientific workflow applications in bioinformatics. As KM techniques, we investigate two methods, Case-Based Reasoning and a rule-based approach. We design and implement both of them and evaluate them with the help of a simulation engine. Simulation reveals the feasibility of the CBR approach and major improvements by the rule-based approach considering SLA violations, resource utilization, the number of necessary reconfigurations and time performance for both, synthetically generated and real-world data. Michael Maurer, Ivona Brandic, Rizos Sakellariou |
Future Gener. Comput. Syst. | 2 |
| 2013 | Managing and Optimizing Bioinformatics Workflows for Data Analysis in Clouds
Vincent C. Emeakaroha, Michael Maurer, Patrick Stern, Pawel P. Labaj, Ivona Brandic, David P. Kreil |
J. Grid Comput. | 5 |
| 2012 | Self-Adaptive and Resource-Efficient SLA Enactment for Cloud Computing InfrastructuresabstractCloud providers aim at guaranteeing Service Level Agreements (SLAs) in a resource-efficient way. This, amongst others, means that resources of virtual (VMs) and physical machines (PMs) have to be autonomically allocated responding to external influences as workload or environmental changes. Thereby, workload volatility (WV) is one of the crucial factors that influence the quality of suggested allocations. In this paper we devise a novel approach for self-adaptive and resource-efficient decision-making considering the three conflicting goals of minimizing the number of SLA violations, maximizing resource utilization, and minimizing the number of necessary time- and energy-consuming reconfiguration actions. We propose self-adaptive rule-based knowledge management for autonomic VM reconfiguration considering the rapidness of changes in the workload, i.e., WV. We introduce a novel WV categorization and present cost and volatility based methods for self-tuning. We evaluate these methods by a large variety of synthetically generated workloads, and by real-world measurements gathered from an image rendering application and a scientific workflow for RNA sequencing. Evaluation shows that in most cases the self-adaptive approach outperforms the static approach. Michael Maurer, Ivona Brandic, Rizos Sakellariou |
IEEE CLOUD | 2 |
| 2012 | M4Cloud - Generic Application Level Monitoring for Resource-shared Cloud Environments
Toni Mastelic, Vincent C. Emeakaroha, Michael Maurer, Ivona Brandic |
CLOSER | 4 |
| 2012 | Methodology for trade-off analysis when moving scientific applications to cloudabstractScientific applications have always been one of the major driving forces for the development and efficient utilization of large scale distributed systems - computational Grids represent one of the prominent examples. While these infrastructures, such as Grids or Clusters, are widely used for running most of the scientific applications, they still use bare physical machines with fixed configurations and very little customizability. Today, Clouds represent another step forward in advanced utilization of distributed computing. They provide a fully customizable and self-managing infrastructure with scalable on-demand resources. However, true benefits and trade-offs of running scientific applications on a cloud infrastructure are still obscure, due to the lack of decision making support, which would provide a systematic approach for comparing these infrastructures. In this paper we introduce a comprehensive methodology for comparing the costs of using both infrastructures based on resource and energy usage, as well as their performance. We also introduce a novel approach for comparing the complexity of setting up and administrating such an infrastructure. Toni Mastelic, Drazen Lucanin, Andreas Ipp, Ivona Brandic |
CloudCom | 4 |
| 2012 | CASViD: Application Level Monitoring for SLA Violation Detection in CloudsabstractCloud resources and services are offered based on Service Level Agreements (SLAs) that state usage terms and penalties in case of violations. Although, there is a large body of work in the area of SLA provisioning and monitoring at infrastructure and platform layers, SLAs are usually assumed to be guaranteed at the application layer. However, application monitoring is a challenging task due to monitored metrics of the platform or infrastructure layer that cannot be easily mapped to the required metrics at the application layer. Sophisticated SLA monitoring among those layers to avoid costly SLA penalties and maximize the provider profit is still an open research challenge. This paper proposes an application monitoring architecture named CASViD, which stands for Cloud Application SLA Violation Detection architecture. CASViD architecture monitors and detects SLA violations at the application layer, and includes tools for resource allocation, scheduling, and deployment. Different from most of the existing monitoring architectures, CASViD focuses on application level monitoring, which is relevant when multiple customers share the same resources in a Cloud environment. We evaluate our architecture in a real Cloud testbed using applications that exhibit heterogeneous behaviors in order to investigate the effective measurement intervals for efficient monitoring of different application types. The achieved results show that our architecture, with low intrusion level, is able to monitor, detect SLA violations, and suggest effective measurement intervals for various workloads. Vincent C. Emeakaroha, Tiago Ferreto, Marco Aurélio Stelmar Netto, Ivona Brandic, César A. F. De Rose |
COMPSAC | 4 |
| 2012 | Topic 6: Grid, Cluster and Cloud Computing
Erik Elmroth, Paraskevi Fragopoulou, Artur Andrzejak 0001, Ivona Brandic, Karim Djemame, Paolo Romano 0002 |
Euro-Par | 4 |
| 2012 | Facilitating Self-Adaptable Inter-cloud ManagementabstractCloud Computing infrastructures have been developed as individual islands, and mostly proprietary solutions so far. However, as more and more infrastructure providers apply the technology, users face the inevitable question of using multiple infrastructures in parallel. Federated cloud management systems offer a simplified use of these infrastructures by hiding their proprietary solutions. As the infrastructure becomes more complex underneath these systems, the situations (like system failures, handling of load peaks and slopes) that users cannot easily handle, occur more and more frequently. Therefore, federations need to manage these situations autonomously without user interactions. This paper introduces a methodology to autonomously operate cloud federations by controlling their behavior with the help of knowledge management systems. Such systems do not only suggest reactive actions to comply with established Service Level Agreements (SLA) between provider and consumer, but they also find a balance between the fulfillment of established SLAs and resource consumption. The paper adopts rule-based techniques as its knowledge management solution and provides an extensible rule set for federated clouds built on top of multiple infrastructures. Gabor Kecskemeti, Michael Maurer, Ivona Brandic, Attila Kertész, Zsolt Németh, Schahram Dustdar |
PDP | 3 |
| 2012 | Special section: Recent advances in utility and cloud computing
Ivona Brandic, Rajkumar Buyya |
Future Gener. Comput. Syst. | 1 |
| 2012 | Towards autonomic detection of SLA violations in Cloud infrastructures
Vincent C. Emeakaroha, Marco Aurélio Stelmar Netto, Rodrigo N. Calheiros, Ivona Brandic, Rajkumar Buyya, César A. F. De Rose |
Future Gener. Comput. Syst. | 4 |
| 2012 | Cost-benefit analysis of an SLA mapping approach for defining standardized Cloud computing goods
Michael Maurer, Vincent C. Emeakaroha, Ivona Brandic, Jörn Altmann |
Future Gener. Comput. Syst. | 3 |
| 2011 | Towards Autonomic Market Management in Cloud Computing Infrastructures
Ivan Breskovic, Michael Maurer, Vincent C. Emeakaroha, Ivona Brandic, Jörn Altmann |
CLOSER | 4 |
| 2011 | Towards Self-Awareness in Cloud Markets: A Monitoring MethodologyabstractCurrently, the Cloud landscape is a fragmented, static and shapeless market that hinders the paradigm's ability to fulfil its promise of ubiquitous computing on tap and as a commodity. In this paper, we present our vision of an autonomic self-aware Cloud market platform, and argue that autonomic market platforms for Clouds can step up to the challenge of today's status quo. As our first steps towards achieving this vision, we present a market monitoring methodology, which includes a series of realistic market goals, sets of extractable metrics from a market platform and how to map (i.e. combine and transform) metrics to access goal performance such that autonomic adaption of the market could be undertaken. We have extended a known market simulator for distributed infrastructures (GridSim) with relevant sensors. To demonstrate the usefulness of our approach, we simulate a sudden cease in demand for goods in our market platform. Ivan Breskovic, Christian Haas 0003, Simon Caton, Ivona Brandic |
DASC | 4 |
| 2011 | Enacting SLAs in Clouds Using Rules
Michael Maurer, Ivona Brandic, Rizos Sakellariou |
Euro-Par (1) | 2 |
| 2011 | Revealing the MAPE loop for the autonomic management of Cloud infrastructuresabstractCloud computing is the result of the convergence of several concepts, ranging from virtualization, distributed application design, Grid computing, and enterprise IT management. Efficient management of Cloud computing infrastructures faces with the contradicting goals like unlimited scalability, provision of Service Level Agreements (SLAs), extensive use of virtualization, energy efficiency and minimization of the administration overhead by humans. Thus, autonomic computing seems to be one of the promising paradigms for the implementation of the management infrastructures for Clouds. However, currently available autonomic systems do not consider the characteristics of Clouds, e.g., virtualization layer, and thus are not easily applicable to Cloud infrastructures. In this paper we discuss first steps towards revealing the current MAPE (Monitoring, Analysis, Planning, Execution) loops for the application to Cloud infrastructures. We present novel techniques for the adequate monitoring of Clouds, discuss the approach for the knowledge management and present our solutions for facilitating SLA generation and management. Michael Maurer, Ivan Breskovic, Vincent C. Emeakaroha, Ivona Brandic |
ISCC | 4 |
| 2011 | Autonomic SLA-Aware Service Virtualization for Distributed SystemsabstractCloud Computing builds on the latest achievements of diverse research areas, such as Grid Computing, Service-oriented computing, business processes and virtualization. Managing such heterogeneous environments requires sophisticated interoperation of adaptive coordinating components. In this paper we introduce an SLA-aware Service Virtualization architecture that provides non-functional guarantees in the form of Service Level Agreements and consists of a three-layered infrastructure including agreement negotiation, service brokering and on demand deployment. In order to avoid costly SLA violations, flexible and adaptive SLA attainment strategies are used with a failure propagation approach. We demonstrate the advantages of our proposed solution with a biochemical case study in a Cloud simulation environment. Attila Kertész, Gabor Kecskemeti, Ivona Brandic |
PDP | 3 |
| 2010 | Compliant Cloud Computing (C3): Architecture and Language Support for User-Driven Compliance Management in CloudsabstractCloud computing represents a promising computing paradigm, where computational power is provided similar to utilities like water, electricity or gas. While most of the Cloud providers can guarantee some measurable non-functional performance metrics e.g., service availability or throughput, there is lack of adequate mechanisms for guaranteeing certifiable and auditable security, trust, and privacy of the applications and the data they process. This lack represents an obstacle for moving most business relevant applications into the Cloud. In this paper we devise a novel approach for compliance management in Clouds, which we termed Compliant Cloud Computing (C3). On one hand, we propose novel languages for specifying compliance requirements concerning security, privacy, and trust by leveraging domain specific languages and compliance level agreements. On the other hand, we propose the C3 middleware responsible for the deployment of certifiable and auditable applications, for provider selection in compliance with the user requirements, and for enactment and enforcement of compliance level agreements. We underpin our approach with a use case discussing various techniques necessary for achieving security, privacy, and trust in Clouds as for example data fragmentation among different protection domains or among different geographical regions. Ivona Brandic, Schahram Dustdar, Tobias Anstett, David Schumm, Frank Leymann, Ralf Konrad |
IEEE CLOUD | 1 |
| 2010 | Towards Knowledge Management in Self-Adaptable CloudsabstractCloud computing represents a promising computing paradigm where resources have to be dynamically allocated to software that needs to be executed. Self-manageable Cloud infrastructures are required to achieve that level of flexibility on the one hand, and to comply to users' requirements specified by means of Service Level Agreements (SLAs) on the other. Such infrastructures should automatically respond to changing component, workload, and environmental conditions minimizing user interactions with the system and preventing violations of SLAs. However, identification of system states where reactive actions are necessary for the prevention of SLA violations is far from trivial. In this paper we investigate how current knowledge management systems can be used for the prevention of SLA violations in Clouds. First, we define a typical SLA use case and formulate the expected behavior of the knowledge management system in order to prevent possible SLA violations. Second, we investigate different methods for knowledge management, e.g., situation calculus and case based reasoning (CBR). We discuss how these methods match the expected behavior for SLA violation prevention. In particular we examine the CBR method and devise several approaches for knowledge management in Clouds based on CBR. Finally, we evaluate our approach based on the presented use case. Michael Maurer, Ivona Brandic, Vincent C. Emeakaroha, Schahram Dustdar |
SERVICES | 2 |
| 2009 | Towards Self-Manageable Cloud ServicesabstractCloud computing represents a promising computing paradigm, where computational power is provided as a utility. An important characteristic of Cloud computing, other than in similar paradigms like Grid or HPC computing, is the provision of non-functional guarantees to users. Thereby, applications can be executed considering predefined execution time, price, security or privacy standards, which are guaranteed in real time in form of Service Level Agreements (SLAs). However, due to changing components, workload, external conditions, hardware, and software failures, established SLAs may be violated. Thus, frequent user interactions with the system, which are usually necessary in case of failures, might turn out to be an obstacle for the success of Cloud computing. In this paper we discuss self-manageable Cloud services. In case of failures, environmental changes, and similar, services manage themselves automatically following the principles of autonomic computing. Based on the life cycle of a self-manageable Cloud service we derive a resource submission taxonomy. Furthermore, we present an architecture for the implementation of self-manageable Cloud services. Finally, we discuss the application of autonomic computing to Cloud services based on service mediation and negotiation bootstrapping case study. Ivona Brandic |
COMPSAC (2) | 1 |
| 2009 | Monitoring and Analyzing Influential Factors of Business Process PerformanceabstractBusiness activity monitoring enables continuous observation of key performance indicators (KPIs). However, if things go wrong, a deeper analysis of process performance becomes necessary. Business analysts want to learn about the factors that influence the performance of business processes and most often contribute to the violation of KPI target values, and how they relate to each other. We provide a framework for performance monitoring and analysis of WS-BPEL processes, which consolidates process events and Quality of Service measurements. The framework uses machine learning techniques in order to construct tree structures, which represent the dependencies of a KPI on process and QoS metrics. These dependency trees allow business analysts to analyze how the process KPIs depend on lower-level process metrics and QoS characterisitics of the IT infrastructure. Deeper knowledge about the structure of dependencies can be gained by drill-down analysis of single factors of influence. Branimir Wetzstein, Philipp Leitner 0001, Florian Rosenberg, Ivona Brandic, Schahram Dustdar, Frank Leymann |
EDOC | 4 |
| 2009 | Cloud computing and emerging IT platforms: Vision, hype, and reality for delivering computing as the 5th utility
Rajkumar Buyya, Chee Shin Yeo, Srikumar Venugopal, James Broberg, Ivona Brandic |
Future Gener. Comput. Syst. | 5 |
| 2008 | Specification, planning, and execution of QoS-aware Grid workflows within the Amadeus environmentabstractAbstract Commonly, at a high level of abstraction Grid applications are specified based on the workflow paradigm. However, majority of Grid workflow systems either do not support Quality of Service (QoS), or provide only partial QoS support for certain phases of the workflow lifecycle. In this paper we present Amadeus, which is a holistic service‐oriented environment for QoS‐aware Grid workflows. Amadeus considers user requirements, in terms of QoS constraints, during workflow specification, planning, and execution. Within the Amadeus environment workflows and the associated QoS constraints are specified at a high level using an intuitive graphical notation. A distinguishing feature of our system is the support of a comprehensive set of QoS requirements, which considers in addition to performance and economical aspects also legal and security aspects. A set of QoS‐aware service‐oriented components is provided for workflow planning to support automatic constraint‐based service negotiation and workflow optimization. For improving the efficiency of workflow planning we introduce a QoS‐aware workflow reduction technique. Furthermore, we present our static and dynamic planning strategies for workflow execution in accordance with user‐specified requirements. For each phase of the workflow lifecycle we experimentally evaluate the corresponding Amadeus components. Copyright © 2007 John Wiley & Sons, Ltd. Ivona Brandic, Sabri Pllana, Siegfried Benkner |
Concurr. Comput. Pract. Exp. | 1 |
| 2007 | Performance Modeling and Prediction of Parallel and Distributed Computing Systems: A Survey of the State of the ArtabstractPerformance is one of the key features of parallel and distributed computing systems. Therefore, in the past a significant research effort was invested in the development of approaches for performance modeling and prediction of parallel and distributed computing systems. In this paper we identify the trends, contributions, and drawbacks of the state of the art approaches. We describe a wide range of the performance modeling approaches that spans from the high-level mathematical modeling to the detailed instruction-level simulation. For each approach we describe how the program and machine are modeled and estimate the model development and evaluation effort, the efficiency, and the accuracy. Furthermore, we present an overall evaluation of the presented approaches Sabri Pllana, Ivona Brandic, Siegfried Benkner |
CISIS | 2 |
| 2007 | Nonadiabatic Ab Initio Surface-Hopping Dynamics Calculation in a Grid Environment - First Experiences
Matthias Ruckenbauer, Ivona Brandic, Siegfried Benkner, Wilfried N. Gansterer, Osvaldo Gervasi, Mario Barbatti, Hans Lischka |
ICCSA (1) | 2 |
| 2007 | Component-oriented application construction for a Web service-based GridabstractAbstract We present the architecture and prototype implementation of a component‐oriented programming environment for a Web service based computational Grid. As middleware, we utilize the Vienna Grid Environment (VGE), a framework that enables the provision of compute‐intensive parallel applications as configurable, QoS‐aware Grid services. Our component model follows the Common Component Architecture (CCA) and models application Web services as distributed components. We describe a component framework that integrates VGE services with a component model allowing to express and dynamically manage application and performance meta‐data as well as dependencies on the infrastructure or other components. Furthermore, we show how the client programming interface is used to compose Grid applications from abstract application components that are mapped against available Grid services by the component framework at runtime. Copyright © 2006 John Wiley & Sons, Ltd. Rainer Schmidt 0003, Siegfried Benkner, Ivona Brandic, Gerhard Engelbrecht |
Concurr. Comput. Pract. Exp. | 3 |
| 2005 | QoS Support for Time-Critical Grid Workflow ApplicationsabstractTime critical grid applications as for example simulations for medical surgery or disaster recovery have special quality of service requirements. The Vienna Grid Environment, developed and evaluated in the context of the EU Project GEMSS, facilitates the provision of HPC applications as QoS-aware grid services by providing support for dynamic negotiation of various QoS guarantees like required execution time and price. In this paper, we extend the QoS mechanisms offered by the Vienna Grid Environment to workflow applications. We describe QoS extensions of the business process execution language and present a first prototype of a corresponding QoS-aware workflow engine which implements different strategies in order to bind the tasks of a workflow to adequate grid services subject to user-specified QoS constraints. We present different grid workflow planning approaches as well as first experimental results Ivona Brandic, Siegfried Benkner, Gerhard Engelbrecht, Rainer Schmidt 0003 |
e-Science | 1 |
| 2003 | Performance of Java Web Services Implementations
Siegfried Benkner, Ivona Brandic, Aleksandar Dimitrov, Gerhard Engelbrecht, Rainer Schmidt 0003, Nikolay Terziev |
ICWS | 2 |