VLDB 2026 Research / reviewers in the wild / expert
Erik Elmroth
dblp:36/2690
· DBLP profile ↗
112ranked-venue papers
11as first author
36since 2021 · last 2026
0000-0002-2633-6798ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 49 · 8 first-author · 11 since 2021Computer networks · 13 · 7 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Security and privacy · 5 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LogLearners: Identifying Compromised AI Functions in Serverless Systems
Adil Bin Bhutto, Erik Elmroth, Monowar Bhuyan |
CCGrid | 2 |
| 2026 | Reinforced model selection for resource efficient anomaly detection in edge cloudsabstract• Adaptive Deep Q-Networks (DQN) Integration: Introduced an innovative model selection strategy that adapts DQN for efficient and effective anomaly detection in edge cloud environments, optimizing resource usage and detection accuracy. • Dynamic Resource Management: Demonstrated how the adapted DQN approach significantly reduces computational resource usage by up to 45%, ensuring efficient operation within resource-constrained edge clouds. • Enhanced Real-Time Detection: Achieved up to 85% reduction in detection time, enabling swift anomaly detection through the adapted DQN strategy, while maintaining acceptable accuracy levels. • Robust Experimental Validation: Implemented a realistic testbed setup and tested the proposed approach, providing a comprehensive evaluation of its feasibility and performance across multiple edge cloud scenarios. • Comprehensive Ablation Analysis: Conducted detailed ablation studies on the reward function and state representation, revealing the critical role of resource-awareness in achieving balanced detection performance. Web application services and networks encounter a broad range of security and performance anomalies, necessitating sophisticated detection strategies. However, performing anomaly detection in edge cloud environments, often constrained by limited resources, presents significant computational challenges and demands minimized detection time for real-time response. In this paper, we propose a model selection approach for resource efficient anomaly detection in edge clouds by leveraging an adapted Deep Q-Network (DQN) reinforcement learning technique. The primary objective is to minimize the computational resources required for accurate anomaly detection while achieving low latency and high detection accuracy. Through extensive experimental evaluation in our testbed setup over different representative scenarios, we demonstrate that our adapted DQN approach can reduce resource usage by up to 45% and detection time by up to 85% while incurring less than an 8% drop in F1 score. These results highlight the potential of the adapted DQN model selection strategy to enable efficient, low-latency anomaly detection in resource-constrained edge cloud environments. Javad Forough, Monowar Bhuyan, Erik Elmroth |
Future Gener. Comput. Syst. | 3 |
| 2026 | Efficient retraining of machine learning algorithms in cloud management systemsabstractCloud management systems performing capacity autoscaling, application orchestration, server consolidation, and service differentiation increasingly rely on machine learning (ML) models for predictive decision-making. However, changes in user behavior, software updates, and hardware upgrades cause monitoring data to deviate from the training distribution, leading to model performance degradation. This phenomenon—known as concept drift poses a significant challenge to maintaining prediction accuracy in dynamic cloud environments. In this work, we propose a hybrid concept drift detection approach that combines statistical data monitoring with model performance metrics to inform efficient model retraining. The proposed method minimizes false positives, reduces adaptation delays, and avoids unnecessary retraining during stable periods. We conduct extensive experiments on both synthetic and real-world cloud workload datasets collected from multiple data centers, evaluating the hybrid approach against five established drift detection algorithms. To ensure statistical rigor, all experiments are repeated ten times, and the results are reported with 95% confidence intervals and significance tests. The results show that our proposed approach and implemented methods improve drift detection efficiency by eliminating unnecessary retraining and lead to more than 60% improvements over the baseline prediction accuracy. We Edited the abstract and removed the inconsistencies Lidia Kidane, Paul Townend, Thijs Metsch, Erik Elmroth |
Future Gener. Comput. Syst. | 4 |
| 2025 | ESTHER: Application-First Hardware-Level QoS-Enforcement for Cloud Native EnvironmentsabstractRecent advances in multi-core chip technology have enabled the dynamic tuning of shared memory resources, such as last-level cache and memory bus bandwidth. However, despite proven performance benefits, the complexity of effectively utilizing these hardware-level QoS enforcement features has limited their adoption in real-world cloud computing environments. In this paper, we introduce ESTHER, a novel approach to autonomously fine-tune QoS enforcement features in cloud environments using extremum seeking control, focusing on applications needs and operator ease-of-use. We demonstrate that ESTHER effectively maintains latency-critical workload SLOs and rapidly resolves any infringements by prioritizing shared memory resources. Such fast node-level resolution of SLO violations ensures that costly cluster-level scaling events may be avoided. Furthermore, ESTHER improves best-effort job throughput without impacting latency-critical workloads, achieving performance gains without utilizing workload profiling or prior knowledge of system dynamics. Oliver Larsson, Thijs Metsch, Cristian Klein, Erik Elmroth |
CLOUD | 4 |
| 2025 | Balancing Compression and Prediction: A Hybrid Autoencoder-LSTM Framework for Cloud WorkloadsabstractAccurate future workload prediction is an essential step for proactive resource allocation and efficient provisioning in cloud computing environments. Deep learning strategies have proven successful for this task, but they face challenges due to the high dimensionality of monitoring data, extensive preprocessing requirements, and computational overhead. In this paper, we propose a hybrid framework that integrates autoencoders for workload compression with Long Short-Term Memory (LSTM) networks for time-series forecasting. Unlike prior studies, our approach systematically analyzes the trade-off between compression ratio and predictive accuracy, demonstrating how dimensionality reduction can improve both scalability and robustness. Thereby reducing the computational burden associated with processing massive-scale monitoring data. Experiments conducted on both synthetic and real-world datasets demonstrate that the proposed method achieves up to 60% data compression with minimal reconstruction loss, while also improving prediction accuracy compared to baseline LSTM models. We evaluate the overall performance of the framework using various metrics, including data reduction ratio, prediction accuracy, and the effects of different compression stages on predictive performance. Additionally, we quantify the computational savings in terms of CPU usage, memory footprint, and training/inference times, confirming the framework’s feasibility for real-world deployment. These results underscore the potential of integrating compression and prediction to achieve scalable, accurate, and resource-efficient management of cloud workloads. Lidia Kidane, Paul Townend, Thijs Metsch, Erik Elmroth |
BDCAT | 4 |
| 2025 | FaLSE: A Failure and Latency-Aware Scheduling for Mission-Critical Applications at the EdgeabstractMission-critical applications, such as real-time emergency response, healthcare, and transport systems, depend heavily on the low latency and reliability provided by Mobile Edge Computing (MEC). The failure of such applications can lead to high latency and severe consequences, including loss of life, financial catastrophe, or operational disruption. However, the dependability of edge clusters is often overlooked, particularly in terms of fault awareness and recovery strategies, which are crucial to these applications. In this work, we focus on loosely coupled IoT applications and propose a Failure and Latency-aware Scheduling approach for Edge (FaLSE) that balances the trade-off between the availability of edge clusters and the latency of containerized mission-critical applications. We used a decentralized network coordinate system to estimate latency between IoT devices/users and nodes. To validate the proposed approach, we compare it with the standard Kubernetes scheduler, which is currently among the most widely used workload orchestration platforms. The results indicate that FaLSE reduced the failure request rate by 87.9% while maintaining a 71.97% lower 95th percentile latency for mission-critical applications and a 10.63% lower latency for normal applications compared to the standard Kubernetes scheduler. Nayereh Rasouli, Cristian Klein, Erik Elmroth |
CloudCom | 3 |
| 2025 | Energy Efficient and QoS-Aware Model Selection for DNN Inference in Edge IntelligenceabstractEdge intelligence is about enabling deep learning applications to run on edge platforms, often under strict Quality of Service (QoS) constraints (e.g., deadlines and accuracy). The heterogeneity and limited computational and energy capacities of edge servers necessitate further study on energy-efficient Deep Neural Network (DNN) inference. While availability of DNN model variants enables adaptive selection without compromising accuracy, it increases the complexity of the solution space. Also, existing research on model selection for DNN inference lacks efficient estimation of energy consumption. This paper proposes a polynomial-time joint strategy for QoS-aware model instance provisioning and selection, based on many-to-many stable matching. Our novel formulation uses the number of floating-point operations (FLOPs) of each model along with hardware-level characteristics of edge servers to minimize total energy usage while maximizing successful inference completions. The proposed strategy is evaluated under different preference functions. Experimental results, compared to optimal and evolutionary algorithms, demonstrate the runtime efficiency of our strategy. Furthermore, extensive evaluations against baselines highlights its superior performance and the importance of jointly considering both system- and application-level parameters in the solution. Hajar Siar, Erik Elmroth |
CloudCom | 2 |
| 2025 | Taming Cold Starts: Proactive Serverless Scheduling with Model Predictive ControlabstractServerless computing has transformed cloud application deployment by introducing a fine-grained, event-driven execution model that abstracts away infrastructure management. Its on-demand nature makes it especially appealing for latency-sensitive and bursty workloads. However, the cold start problem, i.e., where the platform incurs significant delay when provisioning new containers, remains the Achilles’ heel of such platforms.This paper presents a predictive serverless scheduling framework based on Model Predictive Control to proactively mitigate cold starts, thereby improving end-to-end response time. By forecasting future invocations, the controller jointly optimizes container prewarming and request dispatching, improving latency while minimizing resource overhead.We implement our approach on Apache OpenWhisk, deployed on a Kubernetes-based testbed. Experimental results using real-world function traces and synthetic workloads demonstrate that our method significantly outperforms state-of-the-art baselines, achieving up to $85 \%$ lower tail latency and a $34 \%$ reduction in resource usage. Chanh Nguyen 0001, Monowar Bhuyan, Erik Elmroth |
MASCOTS | 3 |
| 2025 | tinyKube: A Middleware for Dynamic Resource Management in Cloud-Edge Platforms for Large-Scale Cloud RoboticsabstractWith the rise of ubiquitous networking and distributed computing, integrating robots with cloud-edge infrastructures offers significant potential. However, challenges remain in resource allocation and scheduling across distributed environments to meet robotics applications' performance demands. This paper introduces tinyKube, a middleware tailored for dynamic resource management across the cloud-edge platform for large-scale cloud robotics deployments. Leveraging Kubernetes for orchestration and Prometheus for monitoring, tinyKube enables unified monitoring, task dispatching, and resource provisioning across cloud-edge infrastructures. We evaluate tinyKube using a robotic gripper application on the CloudGripper testbed in a real-world cloud-edge setup. Results demonstrate its ability to automate task dispatching and resource allocation, dynamically adapting to QoS requirements and workload variations. By simplifying resource management, tinyKube accelerates the development, testing, and deployment of large-scale cloud robotics applications, facilitating more efficient real-world implementation. Chanh Nguyen 0001, Eunil Seo, Oliver Larsson, Florian T. Pokorny, Erik Elmroth |
NOMS | 6 |
| 2025 | MTF-Grasp: A Multi-tier Federated Learning Approach for Robotic GraspingabstractFederated Learning (FL) is a promising machine learning paradigm that enables participating devices to train privacy-preserved and collaborative models. FL has proven its benefits for robotic manipulation tasks. However, grasping tasks lack exploration in such settings where robots train a global model without moving data and ensuring data privacy. The main challenge is that each robot learns from data that is nonindependent and identically distributed (non-IID) and of low quantity. This exhibits performance degradation, particularly in robotic grasping. Thus, in this work, we propose MTF-Grasp, a multi-tier FL approach for robotic grasping, acknowledging the unique challenges posed by the non-IID data distribution across robots, including quantitative skewness. MTF-Grasp harnesses data quality and quantity across robots to select a set of "top-level" robots with better data distribution and higher sample count. It then utilizes top-level robots to train initial seed models and distribute them to the remaining "low-level" robots, reducing the risk of model performance degradation in low-level robots. Our approach outperforms the conventional FL setup by up to 8% on the quantity-skewed Cornell and Jacquard grasping datasets. Obaidullah Zaland, Erik Elmroth, Monowar Bhuyan |
SMC | 2 |
| 2025 | Pioneering Eco-Efficiency in Cloud Computing: The Carbon-Conscious Federated Reinforcement Learning (CCFRL) ApproachabstractIn response to the growing emphasis on sustainability in federated learning (FL), this research introduces a dynamic, dual-objective optimization framework called carbon-conscious federated reinforcement learning (CCFRL). By leveraging reinforcement learning (RL), CCFRL continuously adapts client allocation and resource usage in real time, optimizing both carbon efficiency and model performance. Unlike static or greedy methods that prioritize short-term carbon constraints, existing approaches often suffer from either degrading model performance by excluding high-quality, energy-intensive clients or failing to adequately balance carbon emissions with long-term efficiency. CCFRL addresses these limitations by taking a more sustainable method, balancing immediate resource needs with long-term sustainability, and ensuring that energy consumption and carbon emissions are minimized without compromising model quality, even with nonindependent and identically distributed (non-IID) and large-scale datasets. We overcome the shortcomings of existing methods by integrating advanced state representations, adaptive exploration and exploitation transitions, and stagnating detection using t-tests to better manage real-world data heterogeneity and complex, nonlinear datasets. Extensive experiments demonstrate that CCFRL significantly reduces both energy consumption and carbon emissions while maintaining or enhancing performance. With up to a 61.78% improvement in energy conservation and a 64.23% reduction in carbon emissions, CCFRL proves the viability of aligning resource management with sustainability goals, paving the way for a more environmentally responsible future in cloud computing. Eunil Seo, Erik Elmroth |
IEEE Internet Things J. | 2 |
| 2024 | State-Aware Application Placement in Mobile Edge CloudsabstractPlacing applications within Mobile Edge Clouds (MEC) poses challenges due to dynamic user mobility. Maintaining optimal Quality of Service may require frequent application migration in response to changing user locations, potentially leading to bandwidth wastage. This paper addresses application placement challenges in MEC environments by developing a comprehensive model covering workloads, applications, and MEC infrastructures. Following this, various costs associated with application operation, including resource utilization, migration overhead, and potential service quality degradation, are systematically formulated. An online application placement algorithm, App EDC Match, inspired by the Gale-Shapley matching algorithm, is introduced to optimize application placement considering these cost factors. Through experiments that employ real mobility traces to simulate workload dynamics, the results demonstrate that the proposed algorithm efficiently determines near-optimal application placements within Edge Data Centers. It achieves total operating costs within a narrow margin of 8% higher than the approximate global optimum attained by the offline precognition algorithm, which assumes access to future user locations. Additionally, the proposed placement algorithm effectively mitigates resource scarcity in MEC. Chanh Nguyen 0001, Cristian Klein, Erik Elmroth |
CLOSER | 3 |
| 2024 | Enhancing Machine Learning Performance in Dynamic Cloud Environments with Auto-Adaptive ModelsabstractAutonomous resource management is essential for large-scale cloud data centers, where Machine Learning (ML) enables intelligent decision-making. However, shifts in data patterns within operational streams pose significant challenges to sustaining model accuracy and system efficiency. This paper proposes an auto-adaptive ML approach to miti-gate the impact of data drift in cloud systems. A knowledge base of distinct time-series batches and corresponding ML models is constructed and clustered using the Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) algorithm. When model performance degrades, the system uses Dynamic Time Warping (DTW) to retrieve matching hyper-parameters from the knowledge base and apply them to the deployed model, optimizing inference accuracy on new data streams. Experiments with two real-world cloud data traces - representing both stable and highly fluctuating environments - demonstrate that the proposed approach maintains high model accuracy (over 89%) while minimizing retraining costs. Specifically, for the Wikipedia trace with frequent data drift, retraining overhead is reduced by 22.9% compared to drift detection-based retraining and by 97% compared to incremental retraining. In stable environments, like the Google cluster trace, retraining costs decrease by 96.3% and 88.9%, respectively. Chanh Nguyen 0001, Monowar Bhuyan, Erik Elmroth |
CloudCom | 3 |
| 2024 | FloRa: Flow Table Low-Rate Overflow Reconnaissance and Detection in SDNabstractSDN has evolved to revolutionize next-generation networks, offering programmability for on-the-fly service provisioning, primarily supported by the OpenFlow (OF) protocol. The limited storage capacity of Ternary Content Addressable Memory (TCAM) for storing flow tables in OF switches introduces vulnerabilities, notably the Low-Rate Flow Table Overflow (LOFT) attacks. LOFT exploits the flow table’s storage capacity by occupying a substantial amount of space with malicious flow, leading to a gradual degradation in the flow-forwarding performance of OF switches. To mitigate this threat, we propose FloRa, a machine learning-based solution designed for monitoring and detecting LOFT attacks in SDN. FloRa continuously examines and determines the status of the flow table by closely examining the features of the flow table entries. When suspicious activity is identified, FloRa promptly activates the machine-learning based detection module. The module monitors flow properties, identifies malicious flows, and blacklists them, facilitating their eviction from the flow table. Incorporating novel features such as Packet Arrival Frequency, Content Relevance Score, and Possible Spoofed IP along with Cat Boost employed as the attack detection method. The proposed method reduces CPU overhead, memory overhead, and classification latency significantly and achieves a detection accuracy of 99.49% which is more than the state-of-the-art methods to the best of our knowledge. This approach not only protects the integrity of the flow tables but also guarantees the uninterrupted flow of legitimate traffic. Experimental results indicate the effectiveness of FloRa in LOFT attack detection, ensuring uninterrupted data forwarding and continuous availability of flow table resources in SDN. Ankur Mudgal, Abhishek Verma 0003, Munesh Singh, Kshira Sagar Sahoo, Erik Elmroth, Monowar Bhuyan |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2024 | Robust Procedural Learning for Anomaly Detection and Observability in 5G RANabstractMost existing large distributed systems have poor observability and cannot use the full potential of machine learning-based behavior analysis. The system logs, which contain the primary source of information, are unstructured and lack the context needed to track procedures and learn the system’s behavior. This work presents a new trace guideline that enables a component-and procedure-based split of the system logs for the future 5G Radio Access Network (RAN). As the system can be broken into smaller pieces, models can more accurately learn the system’s behavior and use the context to improve anomaly detection and observability. The evaluation result is astonishing; where previously state-of-the-art methods struggle to learn the behavior, a fast, dictionary-based algorithm can detect all anomalies and keep false positives close to zero. Troubleshooters can also more quickly identify anomalies and gain useful insights into the component interaction in RAN. Tobias Sundqvist, Monowar Bhuyan, Erik Elmroth |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2023 | HydraGen: A Microservice Benchmark GeneratorabstractMicroservice-based architectures have become ubiq-uitous in large-scale software systems. Experimental cloud re-searchers constantly propose enhanced resource management mechanisms for such systems. These mechanisms need to be eval-uated using both realistic and flexible microservice benchmarks to study in which ways diverse application characteristics can affect their performance and scalability. However, current mi-croservice benchmarks have limitations including static compu-tational complexity, limited architectural scale, and fixed topology (i.e., number of tiers, fan-in, and fan-out characteristics). We therefore propose HydraGen, a tool that enables re-searchers to systematically generate benchmarks with different computational complexities and topologies, to tackle experimental evaluation of performance at scale for web-serving applications, with a focus on inter-service communication. To illustrate the potential of our open-source tool, we demonstrate how it can reproduce an existing microservice benchmark with preserved architectural properties. We also demonstrate how HydraGen can enrich the evaluation of cloud management systems based on a case study related to traffic engineering. Mohammad Reza Saleh Sedghpour, Aleksandra Obeso Duque, Xuejun Cai, Björn Skubic, Erik Elmroth, Cristian Klein, Johan Tordsson |
CLOUD | 5 |
| 2023 | Bottleneck identification and failure prevention with procedural learning in 5G RANabstractTo meet the low latency requirements of 5G Radio Access Networks (RAN), it is essential to learn where performance bottlenecks occur. As parts are distributed and virtualized, it becomes troublesome to identify where unwanted delays occur. Today, vendors spend huge manual effort analyzing key performance indicators (KPIs) and system logs to detect these bottlenecks. The 5G architecture allows a flexible scaling of microservices to handle the variation in traffic. But knowing how, when, and where to scale is difficult without a detailed latency analysis. In this article, we propose a novel method that combines procedural learning with latency analysis of system log events. The method, which we call LogGenie, learns the latency pattern of the system at different load scenarios and automatically identifies the parts with the most significant increase in latency. Our evaluation in an advanced 5G testbed shows that LogGenie can provide a more detailed analysis than previous research has achieved and help troubleshooters locate bottlenecks faster. Finally, through experiments, we show how a latency prediction model can dynamically fine-tune the behavior where bottlenecks occur. This lowers resource utilization, makes the architecture more flexible, and allows the system to fulfill its latency requirements. Tobias Sundqvist, Monowar Bhuyan, Erik Elmroth |
CCGrid | 3 |
| 2023 | RAVAS: Interference-Aware Model Selection and Resource Allocation for Live Edge Video AnalyticsabstractNumerous edge applications that rely on video analytics demand precise, low-latency processing of multiple video streams from cameras. When these cameras are mobile, such as when mounted on a car or a robot, the processing load on the shared edge GPU can vary considerably. Provisioning the edge with GPUs for the worst-case load can be expensive and, for many applications, not feasible. Ali Rahmanian, Ahmed Ali-Eldin, Selome Kostentinos Tesfatsion, Björn Skubic, Harald Gustafsson, Prashant J. Shenoy, Erik Elmroth |
SEC | 7 |
| 2023 | Detecting DDoS Attacks on the Network Edge: An Information-Theoretic Correlation AnalysisabstractNowadays, edge computing has become part of the Internet of Things (IoT) that plays a vital role in developing smart applications. As the usage of IoT devices significantly increases, at the same time, network edge infrastructure faces several security challenges. Distributed Denial-of-Service (DDoS) attack is one of the most severe threats to edge-cloud services. Therefore, designing a robust mitigating system is unavoidable for the network edge, and it must be able to recognize emerging attacks. This work proposes an anomaly-based DDoS detection approach that combines information-theoretic metrics and multivariate correlation analysis. The information-theoretic metric captures the randomness and complex nature of traffic behaviour. Similarly, multivariate correlation analysis identifies the relationship among traffic features. Combining information metrics and correlation analysis, we generate normal and attack traffic profiles for the training base to estimate density. The generated profiles build on the metrics including Triangle Area Mapping (TAM) with correlation analysis, Renyi’s divergence, covariance, mean, and standard deviation, which enhances the detection performance of the proposed approach. The effectiveness of the proposed approach is evaluated using testbed and benchmark datasets. The results show that the proposed approach achieves 0.17% and 2.32%, and 0.50% higher accuracy compared to the baseline approaches on the testbed, UNSW and CIC-DDoS datasets, respectively. Ryosuke Araki, Kshira Sagar Sahoo, Yuzo Taenaka, Youki Kadobayashi, Erik Elmroth, Monowar Bhuyan |
TrustCom | 5 |
| 2023 | Unified Identification of Anomalies on the Edge: A Hybrid Sequential PGM ApproachabstractEdge cloud resources, just as many other computing resources, are prone to both performance and security anomalies due to their decentralized nature and real-time requirements for processing of data. Their behaviour initially observed as anomalous may, however, in many cases be rather generic and hard to detect. To be able to address such anomalies, it is instrumental to determine whether the anomaly is a "Security" threat or only a "Performance" concern. Therefore, in this paper, we develop an anomaly detection model capable of distinguishing between security and performance anomalies. The model is based on sequential modeling and Probabilistic Graphical Model (PGM), which leverage historical information and dependencies between previous predictions to classify future anomalies accurately. The evaluation of our proposed model shows its superior performance on our testbed and benchmark datasets. Accordingly, the model achieves an average 5%, and 3% higher F1 score compared to state-of-the-art methods in binary and multi-label anomaly detection cases, respectively. Moreover, our testing time analysis demonstrates the ability of the proposed model in early detection of such anomalies on the edge cloud. Javad Forough, Monowar Bhuyan, Erik Elmroth |
TrustCom | 3 |
| 2023 | An ICN-Based Data Marketplace Model Based on a Game Theoretic Approach Using Quality-Data Discovery and Profit OptimizationabstractIn the age of data and machine learning, massive amounts of data produced throughout our society can be rapidly delivered to various applications through a broad spectrum of cloud services. However, the spectrum of applications has vastly different data quality requirements and Willingness-To-Pay(WTP), creating a general and complex problem matching consumer quality requirements and budgets with providers’ data quality and price. This paper proposes the Information-Centric Networking(ICN)-based data marketplace to foster quality-data trading service to address the challenge above. We embed a WTP mechanism into an ICN-based data broker service running on cloud computing; therefore, a data consumer can request its desired data with a data name and quality requirement. By specifying nominal WTPs, data consumers can acquire data of the desired quality at the range of maximum nominal WTP. At the same time, a data broker can offer data of a suitable quality based on the profit-optimized price and the proposed service quality using ground-truth accuracy trained by data. We demonstrate that the data broker’s profit can be almost doubled by using the optimal data size and budget determined by considering the one-leader-multiple-followers Stackelberg game. These results show that a value-added data brokering service can profitably facilitate data trading. Eunil Seo, Hyoungshick Kim, Bhaskar Krishnamachari, Erik Elmroth |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | Semi-Supervised Range-Based Anomaly Detection for Cloud SystemsabstractThe inherent characteristics of cloud systems often lead to anomalies, which pose challenges for high availability, reliability, and high performance. Detecting anomalies in cloud key performance indicators (KPI) is a critical step towards building a secure and trustworthy system with early mitigation features. This work is motivated by (i) the efficacy of recent reconstruction-based anomaly detection (AD), (ii) the misrepresentation of the accuracy of time series anomaly detection because point-basedPrecisionandRecallare used to evaluate the efficacy for range-based anomalies, and (iii) detects performance and security anomalies when distributions shift and overlaps. In this paper, we propose a novel semi-supervised dynamic density-based detection rule that uses the reconstruction error vectors in order to detect anomalies. We use long short-term memory networks based on encoder-decoder (LSTM-ED) architecture to reconstruct the normal KPI time series. We experiment with both testbed and a diverse set of real-world datasets. The experimental results show that the dynamic density approach exhibits better performance compared to other detection rules using both standard and range-based evaluation metrics. We also compare the performance of our approach with state-of-the-art methods, outperforms in detecting both performance and security anomalies. Pratyush Kr. Deka, Yash Verma, Adil Bin Bhutto, Erik Elmroth, Monowar Bhuyan |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2022 | A Qualitative Evaluation of Service Mesh-based Traffic Management for Mobile Edge CloudabstractService mesh is getting widely adopted as the cloud-native mechanism for traffic management in microservice-based applications, in particular for generic IT workloads hosted in more centralized cloud environments. Performance-demanding applications continue to drive the decentralization of modern application execution environments, as in the case of mobile edge cloud. This paper presents a systematic and qualitative analysis of state-of-the-art service mesh to evaluate how suitable its design is for addressing the traffic management needs of performance-demanding application workloads hosted in a mobile edge cloud environment. With this analysis, we argue that today's dependability-centric service mesh design fails at addressing the needs of the different types of emerging mobile edge cloud workloads and motivate further research in the directions of performance-efficient architectures, stronger QoS guarantees and higher complexity abstractions of cloud-native traffic manage-ment frameworks. Aleksandra Obeso Duque, Cristian Klein, Jinhua Feng, Xuejun Cai, Björn Skubic, Erik Elmroth |
CCGRID | 6 |
| 2022 | Unsupervised root-cause identification of software bugs in 5G RANabstractDevelopers of complex system like 5G Radio Access Networks (RAN) need algorithms that can automatically locate the root causes of software bugs. Existing methods mainly use supervised learning to track down root causes and only a few of these provide enough information to identify the function in which a software bug occurs. Supervised learning methods work well when scenarios can be repeated, and the normal behavior is somewhat similar. In RAN, where thousands of different configurations are used, software is updated frequently, and each node has its own traffic intensity, using unsupervised learning that does not require any pre-training can be more suitable. The few existing methods that use unsupervised learning to locate the root cause of software bugs can only detect delays or software hangs, and are not able to identify the many types of bugs that occur in RAN. We propose a multi-step method that uses unsupervised learning to analyze kernel and user space traces in system logs. The methods can guide developers by suggesting top-k candidate functions that are likely to contain a software bug. Our methods, MultiSpace and CallGraph were evaluated using an advanced 5G testbed in which many different software bugs that are common in RAN, were injected. The results shows that MultiSpace and CallGraph, can detect a wider range of software bugs than previous methods and only adds an average CPU load of 1.3% on the testbed. An important aspect is also that our methods scale well with large amount of data produced by real time systems, like RAN, and can analyze the data much faster. Tobias Sundqvist, Monowar Bhuyan, Erik Elmroth |
CCNC | 3 |
| 2022 | 15 Years of Cloud Control
Erik Elmroth |
CLOSER | 1 |
| 2022 | DELA: A Deep Ensemble Learning Approach for Cross-layer VSI-DDoS Detection on the EdgeabstractWeb application services and networks become a major target of low-rate Distributed Denial of Service (DDoS) attacks such as Very Short Intermittent DDoS (VSI-DDoS). These threats exploit the TCP congestion control mechanism to cause transient resource outage and impute delays for legitimate users’ requests, while they bypass the secure systems. Besides that, cross-layer VSI-DDoS attacks, where the performed attacks are towards the different layers of the edge cloud infrastructures, are able to cause violation of customers’ Service-Level Agreements (SLAs) with less visible behavioral patterns. In this work, we propose a novel Deep Ensemble Learning Approach named DELA for detection of cross-layer VSI-DDoS on the edge cloud. This approach is developed based on Long Short-Term Memory (LSTM), ensemble learning, and a new voting mechanism based on Feed-Forward Neural Network (FFNN). In addition, it employs a novel training and detection algorithm to combat such attacks in web services and networks. The model shows improved results due to the utilization of historical information in decision- making and also the usage of neural network as aggregator instead of a static threshold-based aggregation. Moreover, we propose a novel overlapped data chunking algorithm that is able to ameliorate the detection performance. Furthermore, the evaluation of DELA shows its superior performance over our testbed and benchmark datasets. Accordingly, DELA achieves on average 4.88% higher F 1 score compared to state-of-the-art methods. Javad Forough, Monowar Bhuyan, Erik Elmroth |
ICDCS | 3 |
| 2022 | MicroSplit: Efficient Splitting of Microservices on Edge CloudsabstractEdge cloud systems reduce the latency between users and applications by offloading computations to a set of small-scale computing resources deployed at the edge of the network. However, since edge resources are constrained, they can become saturated and bottlenecked due to increased load, resulting in an exponential increase in response times or failures. In this paper, we argue that an application can be split between the edge and the cloud, allowing for better performance compared to full migration to the cloud, releasing precious resources at the edge. We model an application's internal call-Graph as a Directed-Acyclic-Graph. We use this model to develop MicroSplit, a tool for efficient splitting of microservices between constrained edge resources and large-scale distant backend clouds. MicroSplit analyzes the dependencies between the microservices of an application, and using the Louvain method for community detection-a popular algorithm from Network Science-decides how to split the microservices between the constrained edge and distant data centers. We test MicroSplit with four microservice based applications in various realistic cloud-edge settings. Our results show that Microsplit migrates up to 60 % of the microservices of an application with a slight increase in the mean-response time compared to running on the edge, and a latency reduction of up to 800 % compared to migrating the entire application to the cloud. Compared to other methods from the State-of-the-Art, MicroSplit reduces the total number of services on the edge by up to five times, with minimal reduction in response times. Ali Rahmanian, Ahmed Ali-Eldin, Björn Skubic, Erik Elmroth |
SEC | 4 |
| 2022 | Special issue on co-design of data and computation management in Fog Computing
Monica Vitali, Pierluigi Plebani, David Bermbach, Erik Elmroth |
Future Gener. Comput. Syst. | 4 |
| 2022 | Resource-Efficient Federated Learning With Non-IID Data: An Auction Theoretic ApproachabstractFederated learning (FL) has gained significant importance for intelligent applications, following data produced on a massive scale by numerous distributed IoT devices. From an FL perspective, the key aspect is that this data is not identically and independently distributed (IID) across different data sources and locations. This distribution-skewness leads to significant quality degradation. Moreover, an intrinsic consequence of using such non-IID data in decentralized learning is increasing costs that would be mitigated if using IID data. As a remedy, we propose a resource-efficient method for training an FL-based application with non-IID data, effectively minimizing cost through an auction approach and mitigating quality degradation through data sharing. In an experimental evaluation, we investigate the FL performance using real-world non-IID data and use the resulting ground-truth outputs to develop functions for estimating the utility of non-IID data, computational resource costs, and data generation costs. These functions are used to optimize the costs of model training, ensuring resource efficiency. It is further demonstrated that using shared-IID data significantly increases the resource efficiency of FL with local non-IID data. This holds true even when the shared IID data size is less than 1% of the size of the local non-IID data. Moreover, this work demonstrates that the profitability of the stakeholders can be maximized using the proposed auction procedure. The integration of the auction procedure and a resource-efficient training strategy allows FL service providers to create practical trading strategies by minimizing the FL clients’ resources and payments in a machine learning marketplace. Eunil Seo, Dusit Niyato, Erik Elmroth |
IEEE Internet Things J. | 3 |
| 2021 | Detection of VSI-DDoS Attacks on the Edge: A Sequential Modeling ApproachabstractThe advent of crucial areas such as smart healthcare and autonomous transportation, bring in new requirements on the computing infrastructure, including higher demand for real-time processing capability with minimized latency and maximized availability. The traditional cloud infrastructure has several deficiencies when meeting such requirements due to its centralization. Edge clouds seems to be the solution for the aforementioned requirements, in which the resources are much closer to the edge devices and provides local computing power and high Quality of Service (QoS). However, there are still security issues that endanger the functionality of edge clouds. One of the recent types of such issues is Very Short Intermittent Distributed Denial of Service (VSI-DDoS) which is a new category of low-rate DDoS attacks that targets both small and large-scale web services. This attack generates very short bursts of HTTP request intermittently towards target services to encounter unexpected degradation of QoS at edge clouds. In this paper, we formulate the problem with a sequence modeling approach to address short intermittent intervals of DDoS attacks during the rendering of services on edge clouds using Long Short-Term Memory (LSTM) with local attention. The proposed approach ameliorates the detection performance by learning from the most important discernible patterns of the sequence data rather than considering complete historical information and hence achieves a more sophisticated model approximation. Experimental results confirm the feasibility of the proposed approach for VSI-DDoS detection on edge clouds and it achieves 2% more accuracy when compared with baseline methods. Javad Forough, Monowar Bhuyan, Erik Elmroth |
ARES | 3 |
| 2021 | Auction-based Federated Learning using Software-defined Networking for resource efficiencyabstractThe training of global models using federated learning (FL) strategies is complicated by variations in local model quality arising from variation in data distribution across individual clients. A wide range of training strategies could be created by varying the size and distribution of the training data and the number of training iterations to be performed. All these variables affect both model quality and resource consumption. To facilitate the selection of good training strategies, we propose an auction-based FL method that can identify a training strategy that is optimal in terms of resource management efficiency subject to a given model quality requirement. An auction method is used to dynamically select resource-efficient FL clients and local models to minimize resource usage. This is enabled by using Software-defined Networking (SDN) to support the dynamic management of FL clients. We show that resource-optimal FL strategies can be implemented in the cloud/edge services market; dynamic quality-based model selection can reduce resource costs by up to 17% from the FL server's perspective. Moreover, the client utility function presented herein helps FL clients adopt practical trading strategies to cooperate efficiently with FL servers. Eunil Seo, Dusit Niyato, Erik Elmroth |
CNSM | 3 |
| 2021 | Model-based Stream Processing Auto-scaling in Geo-Distributed EnvironmentsabstractData stream processing is an attractive paradigm for analyzing IoT data at the edge of the Internet before transmitting processed results to a cloud. However, the relative scarcity of fog computing resources combined with the workloads’ non-stationary properties make it impossible to allocate a static set of resources for each application. We propose Gesscale, a resource auto-scaler which guarantees that a stream processing application maintains a sufficient Maximum Sustainable Throughput to process its incoming data with no undue delay, while not using more resources than strictly necessary. Gesscale derives its decisions about when to rescale and which geo-distributed resource(s) to add or remove on a performance model that gives precise predictions about the future maximum sustainable throughput after reconfiguration. We show that this auto-scaler uses 17% less resources, generates 52% fewer reconfigurations, and processes more input data than baseline auto-scalers based on threshold triggers or a simpler performance model. HamidReza Arkian, Guillaume Pierre, Johan Tordsson, Erik Elmroth |
ICCCN | 4 |
| 2021 | mck8s: An orchestration platform for geo-distributed multi-cluster environmentsabstractFollowing the adoption of cloud computing, the proliferation of cloud data centers in multiple regions, and the emergence of computing paradigms such as fog computing, there is a need for integrated and efficient management of geo-distributed clusters. Geo-distributed deployments suffer from resource fragmentation, as the resources in certain locations are over-allocated while others are under-utilized. Orchestration platforms such as Kubernetes and Kubernetes Federation offer the conceptual models and building blocks that can be used to build integrated solutions that address the resource fragmentation challenge. In this work, we propose mck8s – an orchestration platform for multi-cluster applications on multiple geo-distributed Kubernetes clusters. It offers controllers that automatically place, scale, and burst multi-cluster applications across multiple geo-distributed Kubernetes clusters. mck8s allocates the requested resources to all incoming applications while making efficient use of resources. We designed mck8s to be easy to use by development and operation teams by adopting Kubernetes’ design principles and manifest files. We evaluated mck8s in a geo-distributed experimental testbed in Grid’5000. Our results show that mck8s balances the resource allocation across multiple clusters and reduces the fraction of pending pods to 6% as opposed to 65% in the case of Kubernetes Federation for the same workload. Mulugeta Ayalew Tamiru, Guillaume Pierre, Johan Tordsson, Erik Elmroth |
ICCCN | 4 |
| 2021 | Fed-FiS: a Novel Information-Theoretic Federated Feature Selection for Learning Stability
Sourasekhar Banerjee, Erik Elmroth, Monowar Bhuyan |
ICONIP (5) | 2 |
| 2021 | A 1D-CNN Based Deep Learning for Detecting VSI-DDoS Attacks in IoT Applications
Enkhtur Tsogbaatar, Monowar Bhuyan, Doudou Fall, Yuzo Taenaka, Gonchigsumlaa Khishigjargal, Erik Elmroth, Youki Kadobayashi |
IEA/AIE (1) | 6 |
| 2021 | Adaptive and Application-agnostic Caching in Service Meshes for Resilient Cloud ApplicationsabstractService meshes factor out code dealing with inter-micro-service communication. The overall resilience of a cloud application is improved if constituent micro-services return stale data, instead of no data at all. This paper proposes and implements application agnostic caching for micro services. While caching is widely employed for serving web service traffic, its usage in inter-micro-service communication is lacking. Micro-services responses are highly dynamic, which requires carefully choosing adaptive time-to-life caching algorithms. Our approach is application agnostic, is cloud native, and supports gRPC. We evaluate our approach and implementation using the micro-service benchmark by Google Cloud called Hipster Shop. Our approach results in caching of about 80% of requests. Results show the feasibility and efficiency of our approach, which encourages implementing caching in service meshes. Additionally, we make the code, experiments, and data publicly available. Lars Larsson 0001, William Tärneberg, Cristian Klein, Maria Kihl, Erik Elmroth |
NetSoft | 5 |
| 2020 | An Experimental Evaluation of the Kubernetes Cluster Autoscaler in the CloudabstractInternational audience Mulugeta Ayalew Tamiru, Johan Tordsson, Erik Elmroth, Guillaume Pierre |
CloudCom | 3 |
| 2020 | Elasticity Control for Latency-Intolerant Mobile Edge ApplicationsabstractElasticity is a fundamental property required for Mobile Edge Clouds (MECs) to become mature computing platforms hosting software applications. However, MECs must cope with several challenges that do not arise in the context of conventional cloud platforms. These include the potentially highly distributed geographical deployment, heterogeneity, and limited resource capacity of Edge Data Centers (EDCs), and end-user mobility. In this paper, we present an elasticity controller to help MECs overcome these challenges by automatic proactive resource scaling. The controller utilizes information on the physical locations of EDCs and the correlation of workload changes in physically neighboring EDCs to predict request arrival rates at EDCs. These predictions are used as inputs for a queueing theory-driven performance model that estimates the number of resources that should be provisioned to EDCs in order to meet predefined Service Level Objectives (SLOs) while maximizing resource utilization. The controller also incorporates a group-level load balancer that is responsible for redirecting requests among EDCs during runtime so as to minimize the request rejection rate. We evaluate our approach by performing simulations with an emulated MEC deployed over a metropolitan area and a simulated application workload using a real-world user mobility trace. The results show that our proposed pro-active controller exhibits better scaling behavior than a state-of-the-art re-active controller and increases the efficiency of resource provisioning, thereby helping MECs to sustain resource utilization and rejection rates that satisfy predefined SLOs while maintaining system stability. Chanh Nguyen 0001, Cristian Klein, Erik Elmroth |
SEC | 3 |
| 2020 | Voilà: Tail-Latency-Aware Fog Application Replicas AutoscalerabstractLatency-sensitive fog computing applications may use replication both to scale their capacity and to place application instances as close as possible to their end users. In such geo-distributed environments, a good replica placement should maintain the tail network latency between end-user devices and their closest replica within acceptable bounds while avoiding overloaded replicas. When facing non-stationary workloads it is essential to dynamically adjust the number and locations of a fog application's replicas. We propose Voilà, a tail-Iatency-aware auto-scaler integrated in the Kubernetes orchestration system. Voila maintains a fine-grained view of the volumes of traffic generated from different user locations, and uses simple yet highly-effective procedures to maintain suitable application resources in terms of size and location. Ali J. Fahs, Guillaume Pierre, Erik Elmroth |
MASCOTS | 3 |
| 2020 | Instability in Geo-Distributed Kubernetes Federation: Causes and MitigationabstractAs resources in geo-distributed environments are typically located in remote sites characterized by high latency and intermittent network connectivity, delays and transient network failures are common between the management layer and the remote resources. In this paper, we show that delays and transient network failures coupled with static configuration, including the default configuration parameter values, can lead to instability of application deployments in Kubernetes Federation, making applications unavailable for long periods of time. Leveraging on the benefits of configuration tuning, we propose a feedback controller to dynamically adjust the concerned configuration parameter to improve the stability of application deployments without slowing down the detection of hard failures. We show the effectiveness of our approach in a geo-distributed setup across five sites of Grid'5000, bringing system stability from 83-92% with no controller to 99.5-100% using the controller. Mulugeta Ayalew Tamiru, Guillaume Pierre, Johan Tordsson, Erik Elmroth |
MASCOTS | 4 |
| 2020 | MicroRCA: Root Cause Localization of Performance Issues in MicroservicesabstractSoftware architecture is undergoing a transition from monolithic architectures to microservices to achieve resilience, agility and scalability in software development. However, with microservices it is difficult to diagnose performance issues due to technology heterogeneity, large number of microservices, and frequent updates to both software features and infrastructure. This paper presents MicroRCA, a system to locate root causes of performance issues in microservices. MicroRCA infers root causes in real time by correlating application performance symptoms with corresponding system resource utilization, with-out any application instrumentation. The root cause localization is based on an attributed graph that model anomaly propagation across services and machines. Our experimental evaluation where common anomalies are injected to a microservice benchmark running in a Kubernetes cluster shows that MicroRCA locates root causes well, with 89% precision and 97% mean average precision, outperforming several state-of-the-art methods. Johan Tordsson, Erik Elmroth, Odej Kao |
NOMS | 3 |
| 2020 | Modeling and Simulation of QoS-Aware Power Budgeting in Cloud Data CentersabstractPower budgeting is a commonly employed solution to reduce the negative consequences of high power consumption of large scale data centers. While various power budgeting techniques and algorithms have been proposed at different levels of data center infrastructures to optimize the power allocation to servers and hosted applications, testing them has been challenging with no available simulation platform that enables such testing for different scenarios and configurations. To facilitate evaluation and comparison of such techniques and algorithms, we introduce a simulation model for Quality-of-Service aware power budgeting and its implementation in CloudSim. We validate the proposed simulation model against a deployment on a real testbed, showcase simulator capabilities, and evaluate its scalability. Jakub Krzywda, Vinícius Meyer, Miguel G. Xavier, Ahmed Ali-Eldin, Per-Olov Östberg, César A. F. De Rose, Erik Elmroth |
PDP | 7 |
| 2020 | Impact of etcd deployment on Kubernetes, Istio, and application performanceabstractSummary This experience article describes lessons learned as we conducted experiments in a Kubernetes‐based environment, the most notable of which was that the performance of both the Kubernetes control plane and the deployed application depends strongly and in unexpected ways on the performance of the etcd database. The article contains (a) detailed descriptions of how networking with and without Istio works in Kubernetes, based on the Flannel Container Networking Interface (CNI) provider in VXLAN mode with IP Virtual Server (IPVS)‐backed Kubernetes Services, (b) a comprehensive discussion about how to conduct load and performance testing using a closed‐loop workload generator, and (c) an open source experiment framework useful for executing experiments in a shared cloud environment and exploring the resulting data. It also shows that statistical analysis may reveal the data resulting from such experiments to be misleading even when careful preparations are made, and that nondeterministic behavior stemming from etcd can affect both the platform as a whole and the deployed application. Finally, it is demonstrated that using high‐performance backing storage for etcd can reduce the occurrence of such nondeterministic behaviors by a statistically significant (P < .05) margin. The implication of this experience article is that systems researchers studying the performance of applications deployed on Kubernetes cannot simply consider their specific application to be under test. Instead, the particularities of the underlying Kubernetes and cloud platform must be taken into account, in particular because their performance can impact that of etcd. Lars Larsson 0001, William Tärneberg, Cristian Klein, Erik Elmroth, Maria Kihl |
Softw. Pract. Exp. | 4 |
| 2019 | Multivariate LSTM-Based Location-Aware Workload Prediction for Edge Data CentersabstractMobile Edge Clouds (MECs) is a promising computing platform to overcome challenges for the success of bandwidth-hungry, latency-critical applications by distributing computing and storage capacity in the edge of the network as Edge Data Centers (EDCs) within the close vicinity of end-users. Due to the heterogeneous distributed resource capacity in EDCs, the application deployment flexibility coupled with the user mobility, MECs bring significant challenges to control resource allocation and provisioning. In order to develop a self-managed system for MECs which efficiently decides how much and when to activate scaling, where to place and migrate services, it is crucial to predict its workload characteristics, including variations over time and locality. To this end, we present a novel location-aware workload predictor for EDCs. Our approach leverages the correlation among workloads of EDCs in a close physical distance and applies multivariate Long Short-Term Memory network to achieve on-line workload predictions for each EDC. The experiments with two real mobility traces show that our proposed approach can achieve better prediction accuracy than a state-of-the art location-unaware method (up to 44%) and a location-aware method (up to 17%). Further, through an intensive performance measurement using various input shaking methods, we substantiate that the proposed approach achieves a reliable and consistent performance. Chanh Nguyen 0001, Cristian Klein, Erik Elmroth |
CCGRID | 3 |
| 2019 | Power Shepherd: Application Performance Aware Power ShiftingabstractConstantly growing power consumption of data centers is a major concern from environmental and economical reasons. Current approaches to reduce negative consequences of high power consumption focus on limiting the peak power consumption. During high workload periods, power consumption of highly utilized servers is throttled to stay within the power budget. However, the peak power reduction affects performance of hosted applications and thus leads to Quality of Service violations. In this paper, we introduce Power Shepherd, a hierarchical system for application performance aware power shifting. Power Shepherd reduces the data center operational costs by redistributing the available power among applications hosted in the cluster. This is achieved by, assigning server power budgets by the cluster controller, enforcing these power budgets using Running Average Power Limit (RAPL), and prioritizing applications within each server by adjusting the CPU scheduling configuration. We implement a prototype of the proposed solution and evaluate it in a real testbed equipped with power meters and using representative cloud applications. Our experiments show that Power Shepherd has potential to manage a cluster consisting of thousands of servers and limit the increase of operational costs by a significant amount when the cluster power budget is limited and the system is overutilized. Finally, we identify some outstanding challenges regarding model sensitivity and the fact that this approach in its current from is not beneficial to be used in all situations, e.g., when the system is underutilized. Jakub Krzywda, Ahmed Ali-Eldin, Eddie Wadbro, Per-Olov Östberg, Erik Elmroth |
CloudCom | 5 |
| 2019 | Information-Theoretic Ensemble Learning for DDoS Detection with Adaptive BoostingabstractDDoS (Distributed Denial of Service) attacks pose a serious threat to the Internet as they use large numbers of zombie hosts to forward massive numbers of packets to the target host. Here, we present an adaptive boosting-based ensemble learning model for detecting low-and high-rate DDoS attacks by combining information divergence measures. Our model is trained against a baseline model that does not use labeled traffic data and draws on multiple baseline models developed in parallel to improve its accuracy. Incoming traffic is sampled time-periodically to characterize the normal behavior of input traffic. The model's performance is evaluated using the UmU testbed, MIT legitimate, and CAIDA DDoS datasets. We demonstrate that our model offers superior accuracy to established alternatives, reducing the incidence of false alarms and achieving an F1-score that is around 3% better than those of current state-of-the-art DDoS detection models. Monowar Bhuyan, Maode Ma, Youki Kadobayashi, Erik Elmroth |
ICTAI | 4 |
| 2019 | Adversarial Impact on Anomaly Detection in Cloud DatacentersabstractCloud datacenters are engineered to meet the requirements of generalised and specialised workloads including mission-critical applications that not only generate tremendous amounts of data traces but also opens opportunities for attackers. The increasing volume and rapid changing behaviour of metric streams (e.g., CPU, network, latency, memory) in the cloud datacenters create difficulties to ensure high availability, security, and performance to cloud service providers. Several anomaly detection techniques have been developed to combat system anomalies in cloud datacenters. By injecting a fraction of well-crafted malicious samples in cloud datacenter traces, attackers can subvert the learning process and results in unacceptable false alarms. These security issues cause threats to all categories of anomaly detection. Hence, it is crucial to assess these techniques against adversaries to improve scalability and robustness. We propose a linear regression-based optimisation framework with the ability to poison data in the training phase and demonstrate its effectiveness on cloud datacenter traces. Finally, we investigate the worst-case analysis of poisoning attacks on robust statistics-based anomaly detection techniques to quantify and assess the detection accuracy. We validate this framework using benchmark resource traces obtained from Yahoo's service cluster as well as traces collected from an experimental testbed with realistic service composition. Pratyush Kr. Deka, Monowar Bhuyan, Youki Kadobayashi, Erik Elmroth |
PRDC | 4 |
| 2019 | Graph-based Interactive Data Federation System for Heterogeneous Data Retrieval and AnalyticsabstractGiven the increasing number of heterogeneous data stored in relational databases, file systems or cloud environment, it needs to be easily accessed and semantically connected for further data analytic. The potential of data federation is largely untapped, this paper presents an interactive data federation system (https://vimeo.com/319473546) by applying large-scale techniques including heterogeneous data federation, natural language processing, association rules and semantic web to perform data retrieval and analytics on social network data. The system first creates a Virtual Database (VDB) to virtually integrate data from multiple data sources. Next, a RDF generator is built to unify data, together with SPARQL queries, to support semantic data search over the processed text data by natural language processing (NLP). Association rule analysis is used to discover the patterns and recognize the most important co-occurrences of variables from multiple data sources. The system demonstrates how it facilitates interactive data analytic towards different application scenarios (e.g., sentiment analysis, privacy-concern analysis, community detection). Xuan-Son Vu, Addi Ait-Mlouk, Erik Elmroth, Lili Jiang 0002 |
WWW | 3 |
| 2018 | Multi-scale Low-Rate DDoS Attack Detection Using the Generalized Total Variation MetricabstractWe propose a mechanism to detect multi-scale low-rate DDoS attacks which uses a generalized total variation metric. The proposed metric is highly sensitive towards detecting different variations in the network traffic and evoke more distance between legitimate and attack traffic as compared to the other detection mechanisms. Most low-rate attackers invade the security system by scale-in-and-out of periodic packet burst towards the bottleneck router which severely degrades the Quality of Service (QoS) of TCP applications. Our proposed mechanism can effectively identify attack traffic of this natures, despite its similarity to legitimate traffic, based on the spacing value of our metric. We evaluated our mechanism using datasets from CAIDA DDoS, MIT Lincoln Lab, and real-time testbed traffic. Our results demonstrate that our mechanism exhibits good accuracy and scalability in the detection of multi-scale low-rate DDoS attacks. Monowar Bhuyan, Erik Elmroth |
ICMLA | 2 |
| 2018 | Utility-based Allocation of Industrial IoT Applications in Mobile Edge CloudsabstractMobile Edge Clouds (MECs) create new opportunities and challenges in terms of scheduling and running applications that have a wide range of latency requirements, such as intelligent transportation systems, process automation, and smart grids. We propose a two-tier scheduler for allocating runtime resources to Industrial Internet of Things (IIoT) applications in MECs. The scheduler at the higher level runs periodically - monitors system state and the performance of applications - and decides whether to admit new applications and migrate existing applications. In contrast, the lower-level scheduler decides which application will get the runtime resource next. We use performance based metrics that tells the extent to which the runtimes are meeting the Service Level Objectives (SLOs) of the hosted applications. The Application Happiness metric is based on a single application's performance and SLOs. The Runtime Happiness metric is based on the Application Happiness of the applications the runtime is hosting. These metrics may be used for decision-making by the scheduler, rather than runtime utilization, for example. We evaluate four scheduling policies for the high-level scheduler and five for the low-level scheduler. The objective for the schedulers is to minimize cost while meeting the SLO of each application. The policies are evaluated with respect to the number of runtimes, the impact on the performance of applications and utilization of the runtimes. The results of our evaluation show that the high-level policy based on Runtime Happiness combined with the low-level policy based on Application Happiness outperforms other policies for the schedulers, including the bin packing and random strategies. In particular, our combined policy requires up to 30% fewer runtimes than the simple bin packing strategy and increases the runtime utilization up to 40% for the Edge Data Center (DC) in the scenarios we evaluated. Amardeep Mehta, Ewnetu Bayuh Lakew, Johan Tordsson, Erik Elmroth |
IPCCC | 4 |
| 2018 | Power-performance tradeoffs in data center servers: DVFS, CPU pinning, horizontal, and vertical scaling
Jakub Krzywda, Ahmed Ali-Eldin, Trevor E. Carlson, Per-Olov Östberg, Erik Elmroth |
Future Gener. Comput. Syst. | 5 |
| 2018 | Towards understanding HPC users and systems: A NERSC case study
Gonzalo Pedro Rodrigo Álvarez, Per-Olov Östberg, Erik Elmroth, Katie Antypas, Richard A. Gerber, Lavanya Ramakrishnan |
J. Parallel Distributed Comput. | 3 |
| 2018 | Adaptive Anomaly Detection in Performance Metric StreamsabstractContinuous detection of performance anomalies such as service degradations has become critical in cloud and Internet services due to impact on quality of service and end-user experience. However, the volume and fast changing behavior of metric streams have rendered it a challenging task. Many diagnosis frameworks often rely on thresholding with stationarity or normality assumption, or on complex models requiring extensive offline training. Such techniques are known to be prone to spurious false-alarms in online settings as metric streams undergo rapid contextual changes from known baselines. Hence, we propose two unsupervised incremental techniques following a two-step strategy. First, we estimate an underlying temporal property of the stream via adaptive learning and, then we apply statistically robust control charts to recognize deviations. We evaluated our techniques by replaying over 40 time-series streams from the Yahoo! Webscope S5 datasets as well as four other traces of real Web service QoS and ISP traffic measurements. Our methods achieve high detection accuracy and few false-alarms, and better performance in general compared to an open-source package for time-series anomaly detection. Olumuyiwa Ibidunmoye, Ali-Reza Rezaie, Erik Elmroth |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2017 | KPI-agnostic Control for Fine-Grained Vertical ElasticityabstractApplications hosted in the cloud have become indispensable in several contexts, with their performance often being key to business operation and their running costs needing to be minimized. To minimize running costs, most modern virtualization technologies such as Linux Containers, Xen, and KVM offer powerful resource control primitives for individual provisioning - that enable adding or removing of fraction of cores and/or megabytes of memory for as short as few seconds. Despite the technology being ready, there is a lack of proper techniques for fine-grained resource allocation, because there is an inherent challenge in determining the correct composition of resources an application needs, with varying workload, to ensure deterministic performance. This paper presents a control-based approach for the management of multiple resources, accounting for the resource consumption, together with the application performance, enabling fine-grained vertical elasticity. The control strategy ensures that the application meets the target performance indicators, consuming as less resources as possible. We carried out an extensive set of experiments using different applications - interactive with response-time requirements, as well as noninteractive with throughput desires - by varying the workload mixes of each application over time. The results demonstrate that our solution precisely provides guaranteed performance while at the same time avoiding both resource over-and underprovisioning. Ewnetu Bayuh Lakew, Alessandro Vittorio Papadopoulos, Martina Maggio, Cristian Klein, Erik Elmroth |
CCGrid | 5 |
| 2017 | Incentivizing self-capping to increase cloud utilizationabstractCloud Infrastructure as a Service (IaaS) providers continually seek higher resource utilization to better amortize capital costs. Higher utilization not only can enable higher profit for IaaS providers but also provides a mechanism to raise energy efficiency; therefore creating greener cloud services. Unfortunately, achieving high utilization is difficult mainly due to infrastructure providers needing to maintain spare capacity to service demand fluctuations. Mohammad Shahrad, Cristian Klein, Liang Zheng 0002, Mung Chiang, Erik Elmroth, David Wentzlaff |
SoCC | 5 |
| 2017 | Enabling Workflow-Aware Scheduling on HPC SystemsabstractScientific workflows are increasingly common in the workloads of current High Performance Computing (HPC) systems. However, HPC schedulers do not incorporate workflow-specific mechanisms beyond the capacity to declare dependencies between their jobs. Thus, workflows are run as sets of batch jobs with dependencies, which induces long intermediate wait times and, consequently, long workflow turnaround times. Alternatively, to reduce their turnaround time, workflows may be submitted as single pilot jobs that are allocated their maximum required resources for their entire runtime. Pilot jobs achieve shorter turnaround times but reduce the HPC system's utilization because resources may idle during the workflow's execution. We present a workflow-aware scheduling (WoAS) system that enables existing scheduling algorithms to exploit fine-grained information on a workflow's resource requirements and structure without modification. The current implementation of WoAS is integrated into Slurm, a widely used HPC batch scheduler. We evaluate the system using a simulator using real and synthetic workflows and a synthetic baseline workload that captures job patterns observed over three years of workload data from Edison, a large supercomputer hosted at the National Energy Research Scientific Computing Center. Our results show that WoAS reduces workflow turnaround times and improves system utilization without significantly slowing down conventional jobs. Gonzalo Pedro Rodrigo Álvarez, Erik Elmroth, Per-Olov Östberg, Lavanya Ramakrishnan |
HPDC | 2 |
| 2017 | ACTiCLOUD: Enabling the Next Generation of Cloud ApplicationsabstractDespite their proliferation as a dominant computing paradigm, cloud computing systems lack effective mechanisms to manage their vast amounts of resources efficiently. Resources are stranded and fragmented, ultimately limiting cloud systems' applicability to large classes of critical applications that pose non-moderate resource demands. Eliminating current technological barriers of actual fluidity and scalability of cloud resources is essential to strengthen cloud computing's role as a critical cornerstone for the digital economy. ACTiCLOUD proposes a novel cloud architecture that breaks the existing scale-up and share-nothing barriers and enables the holistic management of physical resources both at the local cloud site and at distributed levels. Specifically, it makes advancements in the cloud resource management stacks by extending state-of-the-art hypervisor technology beyond the physical server boundary and localized cloud management system to provide a holistic resource management within a rack, within a site, and across distributed cloud sites. On top of this, ACTiCLOUD will adapt and optimize system libraries and runtimes (e.g., JVM) as well as ACTiCLOUD-native applications, which are extremely demanding, and critical classes of applications that currently face severe difficulties in matching their resource requirements to state-of-the-art cloud offerings. Georgios I. Goumas, Konstantinos Nikas, Ewnetu Bayuh Lakew, Christos Kotselidis, Andrew Attwood, Erik Elmroth, Michail Flouris, Nikos Foutris, John Goodacre, Davide Grohmann, Vasileios Karakostas, Panagiotis Koutsourakis, Martin L. Kersten, Mikel Luján, Einar Rustad, John Thomson, Luis Tomás, Atle Vesterkjaer, Jim Webber, Ying Zhang 0027, Nectarios Koziris |
ICDCS | 6 |
| 2017 | Calvin Constrained - A Framework for IoT Applications in Heterogeneous EnvironmentsabstractCalvin is an IoT framework for application development, deployment and execution in heterogeneous environments, that includes clouds, edge resources, and embedded or constrained resources. Inside Calvin, all the distributed resources are viewed as one environment by the application. The framework provides multi-tenancy and simplifies development of IoT applications, which are represented using a dataflow of application components (named actors) and their communication. The idea behind Calvin poses similarity with the serverless architecture and can be seen as Actor as a Service instead of Function as a Service. This makes Calvin very powerful as it does not only scale actors quickly but also provides an easy actor migration capability. In this work, we propose Calvin Constrained, an extension to the Calvin framework to cover resource-constrained devices. Due to limited memory and processing power of embedded devices, the constrained side of the framework can only support a limited subset of the Calvin features. The current implementation of Calvin Constrained supports actors implemented in C as well as Python, where the support for Python actors is enabled by using MicroPython as a statically allocated library, by this we enable the automatic management of state variables and enhance code re-usability. As would be expected, Python-coded actors demand more resources over C-coded ones. We show that the extra resources needed are manageable on current off-the-shelve micro-controller-equipped devices when using the Calvin framework. Amardeep Mehta, Rami Baddour, Fredrik Svensson, Harald Gustafsson, Erik Elmroth |
ICDCS | 5 |
| 2017 | ScSF: A Scheduling Simulation Framework
Gonzalo Pedro Rodrigo Álvarez, Erik Elmroth, Per-Olov Östberg, Lavanya Ramakrishnan |
JSSPP | 2 |
| 2017 | Personality-based Knowledge Extraction for Privacy-preserving Data AnalysisabstractIn this paper, we present a differential privacy preserving approach, which extracts personality-based knowledge to serve privacy guarantee data analysis on personal sensitive data. Based on the approach, we further implement an end-to-end privacy guarantee system, KaPPA, to provide researchers iterative data analysis on sensitive data. The key challenge for differential privacy is determining a reasonable amount of privacy budget to balance privacy preserving and data utility. Most of the previous work applies unified privacy budget to all individual data, which leads to insufficient privacy protection for some individuals while over-protecting others. In KaPPA, the proposed personality-based privacy preserving approach automatically calculates privacy budget for each individual. Our experimental evaluations show a significant trade-off of sufficient privacy protection and data utility. Xuan-Son Vu, Lili Jiang 0002, Anders Brändström, Erik Elmroth |
K-CAP | 4 |
| 2017 | Dynamic application placement in the Mobile Cloud Network
William Tärneberg, Amardeep Mehta, Eddie Wadbro, Johan Tordsson, Johan Eker, Maria Kihl, Erik Elmroth |
Future Gener. Comput. Syst. | 7 |
| 2017 | A Survey on Modeling Energy Consumption of Cloud Applications: Deconstruction, State of the Art, and Trade-Off DebatesabstractGiven the complexity and heterogeneity in Cloud computing scenarios, the modeling approach has widely been employed to investigate and analyze the energy consumption of Cloud applications, by abstracting real-world objects and processes that are difficult to observe or understand directly. It is clear that the abstraction sacrifices, and usually does not need, the complete reflection of the reality to be modeled. Consequently, current energy consumption models vary in terms of purposes, assumptions, application characteristics and environmental conditions, with possible overlaps between different research works. Therefore, it would be necessary and valuable to reveal the state-of-the-art of the existing modeling efforts, so as to weave different models together to facilitate comprehending and further investigating application energy consumption in the Cloud domain. By systematically selecting, assessing, and synthesizing 76 relevant studies, we rationalized and organized over 30 energy consumption models with unified notations. To help investigate the existing models and facilitate future modeling work, we deconstructed the runtime execution and deployment environment of Cloud applications, and identified 18 environmental factors and 12 workload factors that would be influential on the energy consumption. In particular, there are complicated trade-offs and even debates when dealing with the combinational impacts of multiple factors. Zheng Li 0001, Selome Kostentinos Tesfatsion, Saeed Bastani, Ahmed Ali-Eldin, Erik Elmroth, Maria Kihl, Rajiv Ranjan 0001 |
IEEE Trans. Sustain. Comput. | 5 |
| 2016 | Towards Understanding Job Heterogeneity in HPC: A NERSC Case StudyabstractThe high performance computing (HPC) scheduling landscape is changing. Increasingly, there are large scientific computations that include high-throughput, data-intensive, and stream-processing compute models. These jobs increase the workload heterogeneity, which presents challenges for classical tightly coupled MPI job oriented HPC schedulers. Thus, it is important to define new analyses methods to understand the heterogeneity of the workload, and its possible effect on the performance of current systems. In this paper, we present a methodology to assess the job heterogeneity in workloads and scheduling queues. We apply the method on the workloads of three current National Energy Research Scientific Computing Center (NERSC) systems in 2014. Finally, we present the results of such analysis, with an observation that heterogeneity might reduce predictability in the jobs' wait time. Gonzalo Pedro Rodrigo Álvarez, Per-Olov Östberg, Erik Elmroth, Katie Antypas, Richard A. Gerber, Lavanya Ramakrishnan |
CCGrid | 3 |
| 2016 | DieHard: Reliable Scheduling to Survive Correlated Failures in Cloud Data CentersabstractIn large scale data centers, a single fault can lead to correlated failures of several physical machines and the tasks running on them, simultaneously. Such correlated failures can severely damage the reliability of a service or a job. This paper models the impact of stochastic and correlated failures on job reliability in a data center. We focus on correlated failures caused by power outages or failures of network components, on jobs running multiple replicas of identical tasks. We present a statistical reliability model and an approximation technique for computing a job's reliability in the presence of correlated failures. In addition, we address the problem of scheduling a job with reliability constraints. We formulate the scheduling problem as an optimization problem, with the aim being to achieve the desired reliability with the minimum number of extra tasks. We present a scheduling algorithm that approximates the minimum number of required tasks and a placement to achieve a desired job reliability. We study the efficiency of our algorithm using an analytical approach and by simulating a cluster with different failure sources and reliabilities. The results show that the algorithm can effectively approximate the minimum number of extra tasks required to achieve the job's reliability. Mina Sedaghat, Eddie Wadbro, John Wilkes, Sara de Luna, Oleg Seleznjev, Erik Elmroth |
CCGrid | 6 |
| 2016 | Service Level and Performance Aware Dynamic Resource Allocation in Overbooked Data CentersabstractMany cloud computing providers use overbooking to increase their low utilization ratios. This however increases the risk of performance degradation due to interference among co-located VMs. To address this problem we present a service level and performance aware controller that: (1) provides performance isolation for high QoS VMs, and (2) reduces the VM interference between low QoS VMs by dynamically mapping virtual cores to physical cores, thus limiting the amount of resources that each VM can access depending on their performance. Our evaluation based on real cloud applications and both stress, synthetic and realistic workloads demonstrates that a more efficient use of the resources is achieved, dynamically allocating the available capacity to the applications that need it more, which in turn lead to a more stable and predictable performance over time. Luis Tomás, Ewnetu Bayuh Lakew, Erik Elmroth |
CCGrid | 3 |
| 2016 | Real-time detection of performance anomalies for cloud servicesabstractService performance degradation and downtimes are a common on the Internet today. Many on-line services (e.g. Amazon.com, Spotify, and Netflix, etc.) report huge loss in revenue and traffic per episode. This is perhaps due to the correlation between performance and end-users's satisfaction. Olumuyiwa Ibidunmoye, Thijs Metsch, Erik Elmroth |
IWQoS | 3 |
| 2016 | A hybrid cloud controller for vertical memory elasticity: A control-theoretic approach
Soodeh Farokhi, Pooyan Jamshidi, Ewnetu Bayuh Lakew, Ivona Brandic, Erik Elmroth |
Future Gener. Comput. Syst. | 5 |
| 2016 | Decentralized cloud datacenter reconsolidation through emergent and topology-aware behavior
Mina Sedaghat, Francisco Hernández-Rodriguez, Erik Elmroth |
Future Gener. Comput. Syst. | 3 |
| 2016 | Modeling and Placement of Cloud Services with Internal StructureabstractVirtual machine placement is the process of mapping virtual machines to available physical hosts within a data center or on a remote data center in a cloud federation. Normally, service owners cannot influence the placement of service components beyond choosing data center provider and deployment zone at that provider. For some services, however, this lack of influence is a hindrance to cloud adoption. For example, services that require specific geographical deployment (due e.g. to legislation), or require redundancy by avoiding co-location placement of critical components. We present an approach for service owners to influence placement of their service components by explicitly specifying service structure, component relationships, and placement constraints between components. We show how the structure and constraints can be expressed and subsequently formulated as constraints that can be used in placement of virtual machines in the cloud. We use an integer linear programming scheduling approach to illustrate the approach, show the corresponding mathematical formulation of the model, and evaluate it using a large set of simulated input. Our experimental evaluation confirms the feasibility of the model and shows how varying amounts of placement constraints and data center background load affects the possibility for a solver to find a solution satisfying all constraints within a certain time-frame. Our experiments indicate that the number of constraints affects the ability of finding a solution to a higher degree than background load, and that for a high number of hosts with low capacity, component affinity is the dominating factor affecting the possibility to find a solution. Daniel Espling, Lars Larsson 0001, Wubin Li, Johan Tordsson, Erik Elmroth |
IEEE Trans. Cloud Comput. | 5 |
| 2015 | Performance-Based Service Differentiation in CloudsabstractDue to fierce competition, cloud providers need to run their data-centers efficiently. One of the issues is to increase data-center utilization while maintaining applications' performance targets. Achieving high data-center utilization while meeting applications' performance is difficult, as data-center overload may lead to poor performance of hosted services. Service differentiation has been proposed to control which services get degraded. However, current approaches are capacity-based, which are oblivious to the observed performance of each service and cannot divide the available capacity among hosted services so as to minimize overall performance degradation. In this paper we propose performance-based service differentiation. In case enough capacity is available, each service is automatically allocated the right amount of capacity that meets its target performance, expressed either as response time or throughput. In case of overload, we propose two service differentiation schemes that dynamically decide which services to degrade and to what extent. We carried out an extensive set of experiments using different services -- interactive as well as non-interactive -- by varying the workload mixes of each service over time. The results demonstrate that our solution precisely provides guaranteed performance or service differentiation depending on available capacity. Ewnetu Bayuh Lakew, Cristian Klein, Francisco Hernández-Rodriguez, Erik Elmroth |
CCGRID | 4 |
| 2015 | Telco Clouds - Modelling and SimulationabstractIn this paper, we propose a telco cloud meta-model that can be used to simulate different infrastructure configurations and explore their consequences for system performance and costs. To achieve this, we analyse current telecommunication and data centre infrastructure paradigms, describe the architecture of the telco cloud, and detail the benefits of merging both infrastructures in a unified system. Next, we detail the dynamics of the telco cloud and identify the components that are the most relevant from the perspective of modelling performance and cost. As a number of well established simulation technologies exist for most of the telco cloud components, we survey existing models in an attempt to construct a suitable composite meta-model. Finally, we present a showcase scenario to demonstrate the scope of our telco cloud simulator. Jakub Krzywda, William Tärneberg, Per-Olov Östberg, Maria Kihl, Erik Elmroth |
CLOSER | 5 |
| 2015 | Continuous Datacenter ConsolidationabstractEfficient mapping of Virtual Machines~(VMs) onto physical servers is a key problem for cloud infrastructure providers as hardware utilization directly impacts profit. Today, this mapping is commonly only performed when new VMs are created, but as VM workloads fluctuate and server availability varies, any initial mapping is bound to become suboptimal over time. We introduce a set of heuristic methods for continuous optimization of the VM-to-server mapping based on combinations of fundamental management actions, namely suspending and resuming physical machines, migrating VMs, and suspending and resuming VMs. By using these methods, cloud infrastructure providers can continuously optimize their server resources regardless of the predictability of the workload. To verify that our approach is applicable in real-world scenarios, we build a proof-of-concept datacenter management system that implements the proposed algorithms. The feasibility of our approach is evaluated through a combination of simulations and real experiments where our system provisions a workload of benchmark applications. Our results indicate that the proposed algorithms are feasible, that the combined management approach achieves the best results, and that the VM suspend and resume mechanism has the largest impact on provider profit. Petter Svärd, Wubin Li, Eddie Wadbro, Johan Tordsson, Erik Elmroth |
CloudCom | 5 |
| 2015 | HPC System Lifetime Story: Workload Characterization and Evolutionary Analyses on NERSC SystemsabstractHigh performance computing centers have traditionally served monolithic MPI applications. However, in recent years, many of the large scientific computations have included high throughput and data-intensive jobs. HPC systems have mostly used batch queue schedulers to schedule these workloads on appropriate resources. There is a need to understand future scheduling scenarios that can support the diverse scientific workloads in HPC centers. In this paper, we analyze the workloads on two systems (Hopper, Carver) at the National Energy Research Scientific Computing (NERSC) Center. Specifically, we present a trend analysis towards understanding the evolution of the workload over the lifetime of the two systems. Gonzalo Pedro Rodrigo Álvarez, Per-Olov Östberg, Erik Elmroth, Katie Antypas, Richard A. Gerber, Lavanya Ramakrishnan |
HPDC | 3 |
| 2015 | Online Spike Detection in Cloud WorkloadsabstractWe investigate methods for detection of rapid workload increases (load spikes) for cloud workloads. Such rapid and unexpected workload spikes are a main cause for poor performance or even crashing applications as the allocated cloud resources become insufficient. To detect the spikes early is fundamental to perform corrective management actions, like allocating additional resources, before the spikes become large enough to cause problems. For this, we propose a number of methods for early spike detection, based on established techniques from adaptive signal processing. A comparative evaluation shows, for example, to what extent the different methods manage to detect the spikes, how early the detection is made, and how frequently they falsely report spikes. Amardeep Mehta, Jonas Durango, Johan Tordsson, Erik Elmroth |
IC2E | 4 |
| 2015 | Analysis and characterization of a video-on-demand service workloadabstractVideo-on-Demand (VoD) and video sharing services account for a large percentage of the total downstream Internet traffic. In order to provide a better understanding of the load on these services, we analyze and model a workload trace from a VoD service provided by a major Swedish TV broadcaster. The trace contains over half a million requests generated by more than 20000 unique users. Among other things, we study the request arrival rate, the inter-arrival time, the spikes in the workload, the video popularity distribution, the streaming bit-rate distribution and the video duration distribution. Our results show that the user and the session arrival rates for the TV4 workload does not follow a Poisson process. The arrival rate distribution is modeled using a lognormal distribution while the inter-arrival time distribution is modeled using a stretched exponential distribution. We observe the "impatient user" behavior where users abandon streaming sessions after minutes or even seconds of starting them. Both very popular videos and non-popular videos are particularly affected by impatient users. We investigate if this behavior is an invariant for VoD workloads. Ahmed Ali-Eldin, Maria Kihl, Johan Tordsson, Erik Elmroth |
MMSys | 4 |
| 2014 | The CACTOS Vision of Context-Aware Cloud Topology Optimization and SimulationabstractRecent advances in hardware development coupled with the rapid adoption and broad applicability of cloud computing have introduced widespread heterogeneity in data centers, significantly complicating the management of cloud applications and data center resources. This paper presents the CACTOS approach to cloud infrastructure automation and optimization, which addresses heterogeneity through a combination of in-depth analysis of application behavior with insights from commercial cloud providers. The aim of the approach is threefold: to model applications and data center resources, to simulate applications and resources for planning and operation, and to optimize application deployment and resource use in an autonomic manner. The approach is based on case studies from the areas of business analytics, enterprise applications, and scientific computing. Per-Olov Östberg, Henning Groenda, Stefan Wesner, James Byrne, Dimitrios S. Nikolopoulos, Craig Sheridan, Jakub Krzywda, Ahmed Ali-Eldin, Johan Tordsson, Erik Elmroth, Christian Stier, Klaus Krogmann, Jörg Domaschka, Christopher B. Hauser, Peter J. Byrne, Sergej Svorobej, Barry McCollum, Zafeirios C. Papazachos, Darren Whigham, Stephan Ruth, Dragana Paurevic |
CloudCom | 10 |
| 2014 | Divide the Task, Multiply the Outcome: Cooperative VM ConsolidationabstractEfficient resource utilization is one of the main concerns of cloud providers, as it has a direct impact on energy costs and thus their revenue. Virtual machine (VM) consolidation is one the common techniques, used by infrastructure providers to efficiently utilize their resources. However, when it comes to large-scale infrastructures, consolidation decisions become computationally complex, since VMs are multi-dimensional entities with changing demand and unknown lifetime, and users often overestimate their actual demand. These uncertainties urges the system to take consolidation decisions continuously in a real time manner. In this work, we investigate a decentralized approach for VM consolidation using Peer to Peer (P2P) principles. We investigate the opportunities offered by P2P systems, as scalable and robust management structures, to address VM consolidation concerns. We present a P2P consolidation protocol, considering the dimensionality of resources and dynamicity of the environment. The protocol benefits from concurrency and decentralization of control and it uses a dimension aware decision function for efficient consolidation. We evaluate the protocol through simulation of 100,000 physical machines and 200,000 VM requests. Results demonstrate the potentials and advantages of using a P2P structure to make resource management decisions in large scale data centers. They show that the P2P approach is feasible and scalable and produces resource utilization of 75% when the consolidation aim is 90%. Mina Sedaghat, Francisco Hernández-Rodriguez, Erik Elmroth, Sarunas Girdzijauskas |
CloudCom | 3 |
| 2014 | How will Your Workload Look Like in 6 Years? Analyzing Wikimedia's WorkloadabstractAccurate understanding of workloads is key to efficient cloud resource management as well as to the design of large-scale applications. We analyze and model the workload of Wikipedia, one of the world's largest web sites. With descriptive statistics, time-series analysis, and polynomial splines, we study the trend and seasonality of the workload, its evolution over the years, and also investigate patterns in page popularity. Our results indicate that the workload is highly predictable with a strong seasonality. Our short term prediction algorithm is able to predict the workload with a Mean Absolute Percentage Error of around 2%. Ahmed Ali-Eldin, Ali Rezaie, Amardeep Mehta, Stanislav Razroev, Sara de Luna, Oleg Seleznjev, Johan Tordsson, Erik Elmroth |
IC2E | 8 |
| 2014 | Priority Operators for Fairshare Scheduling
Gonzalo Pedro Rodrigo Álvarez, Per-Olov Östberg, Erik Elmroth |
JSSPP | 3 |
| 2014 | A Tree-Based Protocol for Enforcing Quotas in CloudsabstractServices are increasingly being hosted on cloud nodes to enhance their performance and increase their availability. The virtually unlimited availability of cloud resources enables service owners to consume resources without quantitative restrictions, paying only for what they use. To avoid cost overruns, resource consumption must be controlled and capped when necessary. We present a distributed tree-based protocol for managing quotas in clouds that minimizes communication overheads and reduces the time required to determine whether a quota has been exhausted. Experimental evaluation shows that our protocol reduces communication costs by 42% relative to a distributed baseline solution and is up to 15 times faster. Ewnetu Bayuh Lakew, Lei Xu 0004, Francisco Hernández-Rodriguez, Erik Elmroth, Claus Pahl |
SERVICES | 4 |
| 2014 | Improving Cloud Service Resilience Using Brownout-Aware Load-BalancingabstractWe focus on improving resilience of cloud services (e.g., e-commerce website), when correlated or cascading failures lead to computing capacity shortage. We study how to extend the classical cloud service architecture composed of a load-balancer and replicas with a recently proposed self-adaptive paradigm called brownout. Such services are able to reduce their capacity requirements by degrading user experience (e.g., disabling recommendations). Combining resilience with the brownout paradigm is to date an open practical problem. The issue is to ensure that replica self-adaptivity would not confuse the load-balancing algorithm, overloading replicas that are already struggling with capacity shortage. For example, load-balancing strategies based on response times are not able to decide which replicas should be selected, since the response times are already controlled by the brownout paradigm. In this paper we propose two novel brownout-aware load-balancing algorithms. To test their practical applicability, we extended the popular lighttpd web server and load-balancer, thus obtaining a production-ready implementation. Experimental evaluation shows that the approach enables cloud services to remain responsive despite cascading failures. Moreover, when compared to Shortest Queue First (SQF), believed to be near-optimal in the non-adaptive case, our algorithms improve user experience by 5%, with high statistical significance, while preserving response time predictability. Cristian Klein, Alessandro Vittorio Papadopoulos, Manfred Dellkrantz, Jonas Durango, Martina Maggio, Karl-Erik Årzén, Francisco Hernández-Rodriguez, Erik Elmroth |
SRDS | 8 |
| 2013 | Decentralized Prioritization-Based Management Systems for Distributed ComputingabstractFairshare scheduling is an established technique to provide user-level differentiation in management of capacity consumption in high-performance and grid computing scheduler systems. In this paper we extend on a state-of-the-art approach to decentralized grid fairs hare and propose a generalized model for construction of decentralized prioritization-based management systems. The approach is based on (re)formulation of control problems as prioritization problems, and a proposed framework for computationally efficient decentralized priority calculation. The model is presented along with a discussion of application of decentralized management systems in distributed computing environments that outlines selected use cases and illustrates key trade-off behaviors of the proposed model. Per-Olov Östberg, Erik Elmroth |
e-Science | 2 |
| 2013 | GJMF - a composable service-oriented grid job management framework
Per-Olov Östberg, Erik Elmroth |
Future Gener. Comput. Syst. | 2 |
| 2013 | Decentralized scalable fairshare scheduling
Per-Olov Östberg, Daniel Espling, Erik Elmroth |
Future Gener. Comput. Syst. | 3 |
| 2012 | Reducing Complexity in Management of eScience ComputationsabstractIn this paper we address reduction of complexity in management of scientific computations in distributed computing environments. We explore an approach based on separation of computation design (application development) and distributed execution of computations, and investigate best practices for construction of virtual infrastructures for computational science - software systems that abstract and virtualize the processes of managing scientific computations on heterogeneous distributed resource systems. As a result we present StratUm, a toolkit for management of eScience computations. To illustrate use of the toolkit, we present it in the context of a case study where we extend the capabilities of an existing kinetic Monte Carlo software framework to utilize distributed computational resources. The case study illustrates a viable design pattern for construction of virtual infrastructures for distributed scientific computing. The resulting infrastructure is evaluated using a computational experiment from molecular systems biology. Per-Olov Östberg, Andreas Hellander, Brian Drawert, Erik Elmroth, Sverker Holmgren, Linda R. Petzold |
CCGRID | 4 |
| 2012 | Topic 6: Grid, Cluster and Cloud Computing
Erik Elmroth, Paraskevi Fragopoulou, Artur Andrzejak 0001, Ivona Brandic, Karim Djemame, Paolo Romano 0002 |
Euro-Par | 1 |
| 2012 | Management of distributed resource allocations in multi-cluster environmentsabstractWe present a fully distributed solution for managing resource allocation for services running across multiple clusters in a large-scale cloud computing environment. Our solution allows individual services running across clusters to compete dynamically for allocations based on their rate of consumption while maintaining the global cloud level allocation limits. The solution monitors resource consumption by services that are spread over a number of clusters. Global polls are triggered only when the allocated balance in a cluster decreases below a threshold and allocations are reassigned in a manner that avoids further immediate global polls. Our solution achieves scalability by minimizing global message exchanges, increases performance by distributing requests, and improves availability by avoiding a single point of failure. We perform a range of simulations to verify the accuracy of our approach, to validate our theoretical results, and to evaluate against competing approaches. Ewnetu Bayuh Lakew, Francisco Hernández-Rodriguez, Lei Xu 0004, Erik Elmroth |
IPCCC | 4 |
| 2012 | An adaptive hybrid elasticity controller for cloud infrastructuresabstractCloud elasticity is the ability of the cloud infrastructure to rapidly change the amount of resources allocated to a service in order to meet the actual varying demands on the service while enforcing SLAs. In this paper, we focus on horizontal elasticity, the ability of the infrastructure to add or remove virtual machines allocated to a service deployed in the cloud. We model a cloud service using queuing theory. Using that model we build two adaptive proactive controllers that estimate the future load on a service. We explore the different possible scenarios for deploying a proactive elasticity controller coupled with a reactive elasticity controller in the cloud. Using simulation with workload traces from the FIFA world-cup web servers, we show that a hybrid controller that incorporates a reactive controller for scale up coupled with our proactive controllers for scale down decisions reduces SLA violations by a factor of 2 to 10 compared to a regression based controller or a completely reactive controller. Ahmed Ali-Eldin, Johan Tordsson, Erik Elmroth |
NOMS | 3 |
| 2012 | OPTIMIS: A holistic approach to cloud service provisioning
Ana Juan Ferrer, Francisco Hernández-Rodriguez, Johan Tordsson, Erik Elmroth, Ahmed Ali-Eldin, Csilla Zsigri, Raül Sirvent, Jordi Guitart, Rosa M. Badia, Karim Djemame, Wolfgang Ziegler, Theodosis Dimitrakos, Srijith Krishnan Nair, George Kousiouris, Kleopatra Konstanteli, Theodora A. Varvarigou, Benoit Hudzia, Alexander Kipp, Stefan Wesner, Marcelo Corrales, Nikolaus Forgó, Tabassum Sharif, Craig Sheridan |
Future Gener. Comput. Syst. | 4 |
| 2011 | Unifying Cloud Management: Towards Overall Governance of Business Level ObjectivesabstractWe address the challenge of providing unified cloud resource management towards an overall business level objective, given the multitude of managerial tasks to be performed and the complexity of any architecture to support them. Resource level management tasks include elasticity control, virtual machine and data placement, autonomous fault management, etc, which are intrinsically difficult problems since services normally have unknown lifetime and capacity demands that varies largely over time. To unify the management of these problems, (for optimization with respect to some higher level business level objective, like optimizing revenue while breaking no more than a certain percentage of service level agreements)becomes even more challenging as the resource level managerial challenges are far from independent. After providing the general problem formulation, we review recent approaches taken by the research community, including mainly general autonomic computing technology for large-scale environments and resource level management tools equipped with some business oriented or otherwise qualitative features. We propose and illustrate a policy-driven approach where a high-level management system monitors overall system and services behavior and adjusts lower level policies (e.g., thresholds for admission control, elasticity control, server consolidation level, etc) for optimization towards the measurable business level objectives. Mina Sedaghat, Francisco Hernández-Rodriguez, Erik Elmroth |
CCGRID | 3 |
| 2011 | Increasing Flexibility and Abstracting Complexity in Service-based Grid and Cloud Software
Per-Olov Östberg, Erik Elmroth |
CLOSER | 2 |
| 2011 | A Cloud Environment for Data-intensive Storage ServicesabstractThe emergence of cloud environments has made feasible the delivery of Internet-scale services by addressing a number of challenges such as live migration, fault tolerance and quality of service. However, current approaches do not tackle key issues related to cloud storage, which are of increasing importance given the enormous amount of data being produced in today's rich digital environment (e.g. by smart phones, social networks, sensors, user generated content). In this paper we present the architecture of a scalable and flexible cloud environment addressing the challenge of providing data-intensive storage cloud services through raising the abstraction level of storage, enabling data mobility across providers, allowing computational and content-centric access to storage and deploying new data-oriented mechanisms for QoS and security guarantees. We also demonstrate the added value and effectiveness of the proposed architecture through two real-life application scenarios from the healthcare and media domains. Elliot K. Kolodner, Sivan Tal, Dimosthenis Kyriazis, Dalit Naor, Miriam Allalouf, Lucia Bonelli, Per Brand, Albert Eckert, Erik Elmroth, Spyridon V. Gogouvitis, Danny Harnik, Francisco Hernández-Rodriguez, Michael C. Jäger, Ewnetu Bayuh Lakew, José Manuel Lopez, Mirko Lorenz, Alberto Messina, Alexandra Shulman-Peleg, Roman Talyansky, Athanasios Voulodimos, Yaron Wolfsthal |
CloudCom | 9 |
| 2011 | Modeling for Dynamic Cloud Scheduling Via Migration of Virtual MachinesabstractCloud brokerage mechanisms are fundamental to reduce the complexity of using multiple cloud infrastructures to achieve optimal placement of virtual machines and avoid the potential vendor lock-in problems. However, current approaches are restricted to static scenarios, where changes in characteristics such as pricing schemes, virtual machine types, and service performance throughout the service life-cycle are ignored. In this paper, we investigate dynamic cloud scheduling use cases where these parameters are continuously changed, and propose a linear integer programming model for dynamic cloud scheduling. Our model can be applied in various scenarios through selections of corresponding objectives and constraints, and offers the flexibility to express different levels of migration overhead when restructuring an existing infrastructure. Finally, our approach is evaluated using commercial clouds parameters in selected simulations for the studied scenarios. Experimental results demonstrate that, with proper parametrizations, our approach is feasible. Wubin Li, Johan Tordsson, Erik Elmroth |
CloudCom | 3 |
| 2011 | High Performance Live Migration through Dynamic Page Transfer Reordering and CompressionabstractAlthough supported by many contemporary Virtual Machine (VM) hypervisors, live migration is impossible for certain applications. When migrating CPU and/or memory intensive VMs two problems occur, extended migration downtime that may cause service interruption or even failure, and prolonged total migration time that is harmful for the overall system performance as significant network resources must be allocated to migration. These problems become more severe for migration over slower networks, such as long distance migration between clouds. We approach this two-fold problem through a combination of techniques. A novel algorithm that dynamically adapts the transfer order of VM memory pages during live migration reduces the risk of re-transfers for frequently dirtied pages. As the amount of transferred data is thereby reduced, the total migration time is shortened. By combining this technique with a compression scheme that increases the migration bandwidth the migration downtime is also reduced. An evaluation by means of synthetic migration benchmarks shows that our combined approach reduces migration downtime by a factor 10 to 20, shortens total migration time by around 35%, as well as consumes between 26% and 39% less network bandwidth. The feasibility of our approach for real-life applications is demonstrated by migrating a streaming video server 31% faster while transferring 51% less data. Petter Svärd, Johan Tordsson, Benoit Hudzia, Erik Elmroth |
CloudCom | 4 |
| 2011 | Introduction
Ramin Yahyapour, Christian Pérez, Erik Elmroth, Ignacio Martín Llorente, Francesc Guim 0001, Karsten Oberle |
Euro-Par (1) | 3 |
| 2011 | Scheduling and monitoring of internally structured services in Cloud federationsabstractCloud infrastructure providers may form Cloud federations to cope with peaks in resource demand and to make large-scale service management simpler for service providers. To realize Cloud federations, a number of technical and managerial difficulties need to be solved. We present ongoing work addressing three related key management topics, namely, specification, scheduling, and monitoring of services. Service providers need to be able to influence how their resources are placed in Cloud federations, as federations may cross national borders or include companies in direct competition with the service provider. Based on related work in the RESERVOIR project, we propose a way to define service structure and placement restrictions using hierarchical directed acyclic graphs. We define a model for scheduling in Cloud federations that abides by the specified placement constraints and minimizes the risk of violating Service-Level Agreements. We present a heuristic that helps the model determine which virtual machines (VMs) are suitable candidates for migration. To aid the scheduler, and to provide unified data to service providers, we also propose a monitoring data distribution architecture that introduces cross-site compatibility by means of semantic metadata annotations. Lars Larsson 0001, Daniel Henriksson, Erik Elmroth |
ISCC | 3 |
| 2011 | Evaluation of delta compression techniques for efficient live migration of large virtual machinesabstractDespite the widespread support for live migration of Virtual Machines (VMs) in current hypervisors, these have significant shortcomings when it comes to migration of certain types of VMs. More specifically, with existing algorithms, there is a high risk of service interruption when migrating VMs with high workloads and/or over low-bandwidth networks. In these cases, VM memory pages are dirtied faster than they can be transferred over the network, which leads to extended migration downtime. In this contribution, we study the application of delta compression during the transfer of memory pages in order to increase migration throughput and thus reduce downtime. The delta compression live migration algorithm is implemented as a modification to the KVM hypervisor. Its performance is evaluated by migrating VMs running different type of workloads and the evaluation demonstrates a significant decrease in migration downtime in all test cases. In a benchmark scenario the downtime is reduced by a factor of 100. In another scenario a streaming video server is live migrated with no perceivable downtime to the clients while the picture is frozen for eight seconds using standard approaches. Finally, in an enterprise application scenario, the delta compression algorithm successfully live migrates a very large system that fails after migration using the standard algorithm. Finally, we discuss some general effects of delta compression on live migration and analyze when it is beneficial to use this technique. Petter Svärd, Benoit Hudzia, Johan Tordsson, Erik Elmroth |
VEE | 4 |
| 2010 | An Aspect-Oriented Approach to Consistency-Preserving Caching and Compression of Web Service Response MessagesabstractWeb Services communicate through XML-encoded messages and suffer from substantial overhead due to verbose encoding of transferred messages and extensive (de)serialization at the end-points. We demonstrate that response caching is an effective approach to reduce Internet latency and server load. Our Tantivy middleware layer reduces the volume of data transmitted without semantic interpretation of service requests or responses and thus improves the service response time. Tantivy achieves this reduction through the combined use of caching of recent responses and data compression techniques to decrease the data representation size. These benefits do not compromise the strict consistency semantics. Tantivy also decreases the overhead of message parsing via storage of application-level data objects rather than XML-representations. Furthermore, we demonstrate how the use of aspect-oriented programming techniques provides modularity and transparency in the implementation. Experimental evaluations based on the WSTest benchmark suite demonstrate that our Tantivy system gives significant performance improvements compared to non-caching techniques. Wubin Li, Johan Tordsson, Erik Elmroth |
ICWS | 3 |
| 2010 | Distributed usage logging for federated Grids
Erik Elmroth, Daniel Henriksson |
Future Gener. Comput. Syst. | 1 |
| 2010 | Three fundamental dimensions of scientific workflow interoperability: Model of computation, language, and execution environment
Erik Elmroth, Francisco Hernández-Rodriguez, Johan Tordsson |
Future Gener. Comput. Syst. | 1 |
| 2009 | RESERVOIR: Management technologies and requirements for next generation Service Oriented InfrastructuresabstractRESERVOIR project is developing an advanced system and service management approach that will serve as the infrastructure for cloud computing and communications and future Internet of services by creative coupling of service virtualization, grid computing, networking and service management techniques. This paper presents work in progress for the integration and management of such systems into a new generation of managed service infrastructure. Benny Rochwerger, Alex Galis, Eliezer Levy, Juan A. Cáceres, David Breitgand, Yaron Wolfsthal, Ignacio Martín Llorente, Mark Wusthoff, Rubén S. Montero, Erik Elmroth |
Integrated Network Management | 10 |
| 2009 | A standards-based Grid resource brokering service supporting advance reservations, coallocation, and cross-Grid interoperabilityabstractAbstract The problem of Grid‐middleware interoperability is addressed by the design and analysis of a feature‐rich, standards‐based framework for all‐to‐all cross‐middleware job submission. The architecture is designed with focus on generality and flexibility and builds on extensive use, internally and externally, of (proposed) Web and Grid services standards such as WSRF, JSDL, GLUE, and WS‐Agreement. The external use provides the foundation for easy integration into specific middlewares, which is performed by the design of a small set of plugins for each middleware. Currently, plugins are provided for integration into Globus Toolkit 4 and NorduGrid/ARC. The internal use of standard formats facilitates customization of the job submission service by replacement of custom components for performing specific well‐defined tasks. Most importantly, this enables the easy replacement of resource selection algorithms by algorithms that address the specific needs of a particular Grid environment and job submission scenario. By default, the service implements a decentralized brokering policy, striving to optimize the performance for the individual user by minimizing the response time for each job submitted. The algorithms in our implementation perform resource selection based on performance predictions, and provide support for advance reservations as well as coallocation of multiple resources for coordinated use. The performance of the system is analyzed with focus on overall service throughput (up to over 250 jobs per min) and individual job submission response time (down to under 1 s). Copyright © 2009 John Wiley & Sons, Ltd. Erik Elmroth, Johan Tordsson |
Concurr. Comput. Pract. Exp. | 1 |
| 2008 | Scalable Grid-wide capacity allocation with the SweGrid Accounting System (SGAS)abstractAbstract The SweGrid Accounting System (SGAS) allocates capacity in collaborative Grid environments by coordinating enforcement of Grid‐wide usage limits as a means to offer usage guarantees and prevent overuse. SGAS employs a credit‐based allocation model where Grid capacity is granted to projects via Grid‐wide quota allowances that can be spent across the Grid resources. The resources collectively enforce these allowances in a soft, real‐time manner. SGAS is built on service‐oriented principles with a strong focus on interoperability and Web services standards. This article covers the SGAS design and implementation, which, besides addressing inherent Grid challenges (scale, security, heterogeneity, decentralization), emphasizes generality and flexibility to produce a customizable system with lightweight integration into different middleware and scheduling system combinations. We focus the discussion around the system design, a flexible allocation model, middleware integration experiences and scalability improvements via a distributed virtual banking system, and finally, an extensive set of testbed experiments. The experiments evaluate the performance of SGAS in terms of response times, request throughput, overall system scalability, and its performance impact on the Globus Toolkit 4 job submission software. We conclude that, for all practical purposes, the quota enforcement overhead incurred by SGAS on job submissions is not a limiting factor for the job‐handling capacity of the job submission software. Copyright © 2008 John Wiley & Sons, Ltd. Peter Gardfjäll, Erik Elmroth, S. Lennart Johnsson, Olle Mulmo, Thomas Sandholm |
Concurr. Comput. Pract. Exp. | 2 |
| 2008 | Grid resource brokering algorithms enabling advance reservations and resource selection based on performance predictions
Erik Elmroth, Johan Tordsson |
Future Gener. Comput. Syst. | 1 |
| 2006 | A Service-Oriented Approach to Enforce Grid Resource AllocationsabstractWe present the SweGrid Accounting System (SGAS) — a decentralized and standards-based system for Grid resource allocation enforcement that has been developed with an emphasis on a uniform data model and easy integration into existing scheduling and workload management software. The system has been tested at the six high-performance computing centers comprising the SweGrid computational resource, and addresses the need for soft, real-time quota enforcement across the SweGrid clusters. The SGAS framework is based on state-of-the-art Web and Grid services technologies. The openness and ubiquity of Web services combined with the fine-grained resource control and cross-organizational security models of Grid services proved to be a perfect match for the SweGrid needs. Extensibility and customizability of policy implementations for the three different parties that the system serves (the user, the resource manager, and the allocation authority) are key design goals. Another goal is end-to-end security and single sign-on, to allow resources to reserve allocations and charge for resource usage on behalf of the user. We conclude this paper by illustrating the policy customization capabilities of SGAS in a simulated setting, where job streams are shaped using different modes of allocation policy enforcement. Finally, we discuss some of the early experiences from the production system. Thomas Sandholm, Peter Gardfjäll, Erik Elmroth, Olle Mulmo, S. Lennart Johnsson |
Int. J. Cooperative Inf. Syst. | 3 |
| 2005 | An advanced grid computing course for application and infrastructure developersabstractThis contribution presents our experiences from developing an advanced course in grid computing, aimed at application and infrastructure developers. The course was intended for computer science students with extensive programming experience and previous knowledge of distributed systems, parallel computing, computer networking, and security. The presentation includes brief presentations of all topics covered in the course, a list of the literature used, and descriptions of the mandatory computer assignments performed using Globus Toolkit 2 and 3. A summary of our experiences from the course and some suggestions for future directions concludes the presentation. Erik Elmroth, Peter Gardfjäll, Johan Tordsson |
CCGRID | 1 |
| 2005 | Design and Evaluation of a Decentralized System for Grid-wide Fairshare SchedulingabstractThis contribution presents a decentralized architecture for a grid-wide fairshare scheduling system and demonstrates its potential in a simulated environment. The system, which preserves local site autonomy, enforces locally and globally scoped share policies, allowing local resource capacity as well as global grid capacity to be logically divided across different groups of users. The policy model is hierarchical and subpolicy definition can be delegated so that, e.g., a VO that has been granted a resource share can partition its share across its projects, which in turn can divide their shares between project members. There is no need for a central coordinator as policies are enforced collectively by the resource schedulers. Each local scheduler adopts a grid-wide view on utilization in order to steer local resource utilization to not only maintain local resource shares but also to contribute to maintaining global shares across the entire set of grid resources. Share enforcement is addressed by an algorithm that calculates simple priority values, thus simplifying integration with local schedulers, which can remain unaware of the hierarchical share policy structure Erik Elmroth, Peter Gardfjäll |
e-Science | 1 |
| 2005 | An Interoperable, Standards-Based Grid Resource Broker and Job Submission ServiceabstractWe present the architecture and implementation of a grid resource broker and job submission service, designed to be as independent as possible of the grid middleware used on the resources. The overall architecture comprises seven general components and a few conversion and integration points where all middleware-specific issues are handled. The implementation is based on state-of-the-art grid and Web services technology as well as existing and emerging standards (WSRF, JSDL, GLUE, WS-Agreement). Features provided by the service include advance reservations and a resource selection process based on a priori estimations of the total time to delivery for the application, including a benchmark-based prediction of the execution time. The general service implementation is based on the Globus Toolkit 4. For test and evaluation, plugins and format converters are provided for use with the NorduGrid ARC middleware Erik Elmroth, Johan Tordsson |
e-Science | 1 |
| 2004 | An OGSA-based accounting system for allocation enforcement across HPC centersabstractIn this paper, we present an Open Grid Services Architecture (OGSA)-based decentralized allocation enforcement system, developed with an emphasis on a consistent data model and easy integration into existing scheduling, and workload management software at six independent high-performance computing centers forming a Grid known as SweGrid. The Swedish National Allocations Committee (SNAC) allocates resource quotas at these centers to research projects requiring substantial computer time. Our system, the SweGrid Accounting System (SGAS), addresses the need for soft real-time allocation enforcement on SweGrid for cross-domain job submission. The SGAS framework is based on state-of-the-art Web and Grid services technologies. The openness and ubiquity of Web services combined with the fine-grained resource control and cross-organizational security models of Grid services proved to be a perfect match for the SweGrid needs. Extensibility and customizability of policy implementations for the three different parties the system serves (the user, the resource manager, and the allocation authority) are key design goals. Another goal is end-to-end security and single sign-on, to allow resources-selected based on client policies-to act on behalf of the user when negotiating contracts with the bank in an environment where the six centers would continue to use their existing accounting policies and tools. We conclude this paper by showing the feasibility of SGAS, which is currently being deployed at the production sites, using simulations of reservation streams. The reservation streams are shaped using soft computing and policy-based algorithms. Thomas Sandholm, Peter Gardfjäll, Erik Elmroth, S. Lennart Johnsson, Olle Mulmo |
ICSOC | 3 |
| 2004 | Design and evaluation of a TOP100 Linux Super Cluster systemabstractAbstract The High Performance Computing Center North (HPC2N) Super Cluster is a truly self‐made high‐performance Linux cluster with 240 AMD processors in 120 dual nodes, interconnected with a high‐bandwidth, low‐latency SCI network. This contribution describes the hardware selected for the system, the work needed to build it, important software issues and an extensive performance analysis. The performance is evaluated using a number of state‐of‐the‐art benchmarks and software, including STREAM, Pallas MPI, the Atlas DGEMM, High‐Performance Linpack and NAS Parallel benchmarks. Using these benchmarks we first determine the raw memory bandwidth and network characteristics; the practical peak performance of a single CPU, a single dual‐node and the complete 240‐processor system; and investigate the parallel performance for non‐optimized dusty‐deck Fortran applications. In summary, this $500 000 system is extremely cost‐effective and shows the performance one would expect of a large‐scale supercomputing system with distributed memory architecture. According to the TOP500 list of June 2002, this cluster was the 94th fastest computer in the world. It is now fully operational and stable as the main computing facility at HPC2N. The system's utilization figures exceed 90%, i.e. all 240 processors are on average utilized over 90% of the time, 24 hours a day, seven days a week. Copyright © 2004 John Wiley & Sons, Ltd. Niklas Edmundsson, Erik Elmroth, Bo Kågström, Markus Mårtensson, Mats Nylén, Åke Sandgren, Mattias Wadenstein |
Concurr. Pract. Exp. | 2 |
| 2001 | High Performance Computations for Large Scale Simulations of Subsurface Multiphase Fluid and Heat Flow
Erik Elmroth, Chris Ding, Yu-Shu Wu |
J. Supercomput. | 1 |
| 1999 | A Parallel Implementation of the TOUGH2 Software Package for Large Scale Multiphase Fluid and Heat Flow SimulationsabstractArticle A parallel implementation of the TOUGH2 software package for large scale multiphase fluid and heat flow simulations Share on Authors: Erik Elmroth Lawrence Berkeley National Laboratory, University of California, Berkeley, CA Lawrence Berkeley National Laboratory, University of California, Berkeley, CAView Profile , Chris Ding Lawrence Berkeley National Laboratory, University of California, Berkeley, CA Lawrence Berkeley National Laboratory, University of California, Berkeley, CAView Profile , Yu-Shu Wu Lawrence Berkeley National Laboratory, University of California, Berkeley, CA Lawrence Berkeley National Laboratory, University of California, Berkeley, CAView Profile , Karsten Pruess Lawrence Berkeley National Laboratory, University of California, Berkeley, CA Lawrence Berkeley National Laboratory, University of California, Berkeley, CAView Profile Authors Info & Claims SC '99: Proceedings of the 1999 ACM/IEEE conference on SupercomputingJanuary 1999 Pages 52–eshttps://doi.org/10.1145/331532.331584Online:01 January 1999Publication History 2citation334DownloadsMetricsTotal Citations2Total Downloads334Last 12 Months1Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Erik Elmroth, Chris Ding, Yu-Shu Wu, Karsten Pruess |
SC | 1 |