Hubertus Franke

dblp:17/3453 · DBLP profile ↗
← Back
61ranked-venue papers
4as first author
22since 2021 · last 2025
0009-0005-0150-1055ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 41 · 3 first-author · 13 since 2021Software engineering, systems software and programming languages · 11 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 CPU-Limits kill Performance: Time to rethink Resource Control
Chirag C. Shetty, Sarthak Chakraborty, Hubertus Franke, Larisa Shwartz, Chandrasekhar Narayanaswami 0001, Indranil Gupta, Saurabh Jha
SoCC3
2025 Towards Continuous Integrity Attestation and Its Challenges in Practice: A Case Study of Keylime
abstract
Continuous integrity attestation is vital for cloud providers to ensure the integrity of remote systems in a continuous manner. Current solutions, such as Keylime, rely on Trusted Platform Module (TPM) and Linux’s Integrity Measurement Architecture (IMA) but often struggle to balance minimizing false alerts and maintaining effective threat detection. In this paper, we examine common causes of attestation failures in Keylime through active experiments. We find that false positives often stem from unscheduled OS updates, and we propose a dynamic policy generation scheme as a solution (validated over 66 days of experiments). Our false negative experiments reveal vulnerabilities in existing designs, including five common issues across Keylime and IMA. Exploiting these issues, attackers could evade detection in all tested scenarios. Our findings offer insights into common failures of continuous integrity attestation, and our proposed solution is being integrated into Keylime with community support, enhancing cloud security. Our code will be available at https://github.com/mruffin/Dynamic-Policy-Generator.
Margie Ruffin, Chenkai Wang 0001, Gheorghe Almási 0001, Abdulhamid Adebayo, Hubertus Franke, Gang Wang 0011
DSN5
2025 Concord: Rethinking Distributed Coherence for Software Caches in Serverless Environments
abstract
Costly accesses to global storage substantially limit the performance of serverless functions. To mitigate this overhead, data can be cached in the memory of the nodes where functions are executed. Existing caching schemes either (1) restrict a data item to be cached in a single node, causing frequent remote reads or (2) allow a data item to be cached in multiple nodes concurrently, adding substantial overhead to maintain cache coherence. Unfortunately, current approaches are suboptimal for the access patterns present in serverless workloads, which are characterized by frequent reads to small data items, strong temporal locality, and a small number of nodes that concurrently execute functions of the same application. Driven by these insights, we propose Concord, a distributed software caching system tailored to serverless environments. Concord allows multiple copies of the same data item to be cached in different nodes concurrently, allowing each cache to satisfy local reads. To maintain coherence across software caches, Concord proposes a directory-based distributed coherence protocol. The protocol is inspired by hardware cache coherence, and is enhanced to minimize coherence traffic, reduce contention points, and be robust to node failures and frequent coherence domain changes. Further, with the Concord coherence protocol, we unlock two new capabilities in serverless environments: transactional storage accesses and transparent data-aware function placement. Compared to state-of-the-art severless caching schemes, Concord running on a 16-node cluster speeds-up execution by 2.4 × and improves throughput by 1.7 ×, while using only 6.2MB of otherwise idle application memory (i.e., 4.8% of the total application memory).
Jovan Stojkovic, Chloe Alverti, Alan Andrade, Nikoleta Iliakopoulou, Hubertus Franke, Tianyin Xu, Josep Torrellas
HPCA5
2025 Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
abstract
The effectiveness of LLMs has triggered an exponential rise in their deployment, imposing substantial demands on inference clusters.Such clusters often handle numerous concurrent queries for different LLM downstream tasks.To handle multi-task settings with vast LLM parameter counts, Low-Rank Adaptation (LoRA) enables task-specific fine-tuning while sharing most of the base LLM model across tasks.Hence, it supports concurrent task serving with reduced memory requirements.However, existing designs face inefficiencies: they overlook workload heterogeneity, impose high CPU-GPU link bandwidth from frequent adapter loading, and suffer from head-of-line blocking in their schedulers.To address these challenges, we present Chameleon, a novel LLM serving system optimized for many-adapter environments.Chameleon introduces two new ideas: adapter caching and adapteraware scheduling.First, Chameleon caches popular adapters in GPU memory, minimizing adapter loading times.For caching, it uses otherwise idle GPU memory, avoiding extra memory costs.Second, Chameleon uses a non-preemptive multi-queue scheduler to efficiently account for workload heterogeneity.In this way, Chameleon simultaneously prevents head of line blocking and starvation.Under high loads, Chameleon reduces the P99 and P50 TTFT latencies by 80.7% and 48.1%, respectively, over a state-of-the-art baseline, while improving the throughput by 1.5×.
Nikoleta Iliakopoulou, Jovan Stojkovic, Chloe Alverti, Tianyin Xu, Hubertus Franke, Josep Torrellas
MICRO5
2025 cache_ext: Customizing the Page Cache with eBPF
Tal Zussman, Ioannis Zarkadas, Jeremy Carin, Andrew Cheng, Hubertus Franke, Jonas Pfefferle, Asaf Cidon
SOSP5
2025 Rex: Closing the language-verifier gap with safe and usable kernel extensions
Jinghao Jia, Ruowen Qin, Milo Craun, Egor Lukiyanov, Ayush Bansal, Minh Phan, Michael V. Le, Hubertus Franke, Hani Jamjoom, Tianyin Xu, Dan Williams 0001
USENIX ATC8
2024 UniNet: Accelerating the Container Network Data Plane in IaaS Clouds
abstract
Kubernetes ($K$8s) is a container orchestration plat-form for cloud-based IaaS environments. While it operates on either bare-metal servers or VMs, users prefer VMs for cost savings and agility reasons despite the added network overhead. This overhead, stemming from dual network tunneling at the VM and container levels, degrades performance. To address this, we present UniNet, a SmartNIC-based solution that offloads container-level network tunneling. We designed UniNet to be compatible with leading Container Network Interfaces (CNIs). This approach involves three key elements: (1) transforming VF-based NICs into a container network gateway, (2) offloading the critical path of the data plane functionalities to SmartNICs for enhanced performance and reduced latency, and (3) instituting an isolated control plane that separates VM- and container-level rule insertions, making it tenant-accessible. UniNet boosts CNI throughput by 7.08 x on average, cuts tail latency by 41.6 %, and reduces CPU usage by up to 5.6 x for the receiver and 4.02 x for the sender, respectively.
William Gropp, Hubertus Franke, Bharat Sukhwani, Sameh W. Asaad, Jinjun Xiong, Volodymyr V. Kindratenko, Deming Chen
CLOUD4
2024 Securing AI Inference in the Cloud: Is CPU-GPU Confidential Computing Ready?
abstract
Many applications have been offloaded onto cloud environments to achieve higher agility, access to more powerful computational resources, and obtain better infrastructure management. Although cloud environments provide solid security solutions, users with highly sensitive data or regulatory compliance requirements, such as HIPAA (Health Insurance Portability and Accountability Act) and GDPR (General Data Protection Regulation), still hesitate to move such application domains to the cloud. To address these concerns, cloud service providers have started to offer solutions to protect data confidentiality and integrity through trusted execution environments (TEEs). While so far these were limited to CPU TEEs only, NVIDIA's Hopper architecture has shifted the landscape by enabling confidential computing features essential to protecting confidentiality and integrity for real-world applications offloaded to GPUs, such as large language models (LLMs). However, there lacks a sufficient study on how much performance overhead confidential computing introduces in a TEE comprised of a CPU-GPU configuration. In this paper we evaluate a confidential computing environment comprised of an Intel TDX system and NVIDIA H100 GPUs through various micro benchmarks and real workloads including BERT, LLaMA, and Granite large language models and provide discussions on the overhead incurred by confidential computing when GPUs are utilized. We show that while LLMs are sensitive to the model types and batch sizes, when larger models with pipelined processing are deployed, the performance of LLM inference in CPU-GPU TEEs can be close to par with their non-confidential setups.
Apoorve Mohan, Mengmei Ye, Hubertus Franke, Mudhakar Srivatsa, Nelson Mimura Gonzalez
CLOUD3
2024 S2TAR: Shared Secure Trusted Accelerators with Reconfiguration for Machine Learning in the Cloud
abstract
The demand for hardware accelerators such as Tensor Processing Units (TPUs) and Graphics Processing Units (GPUs) is rapidly increasing due to growing Machine Learning (ML) workloads. As with any shared computing resources, there is a growing need to dynamically adjust and scale accelerator services while ensuring data privacy and confidentiality, especially in cloud environments. We propose a secure and reconfigurable TPU design with confidential computing support, achieved through a Trusted Execution Environment (TEE) framework tailored for reconfigurable TPU in a multi-tenant cloud. Our contributions include a novel TPU design based on switchbox-enabled systolic arrays to support rapid dynamic partitioning. We evaluate our TPU design with TEEs in shared environments, achieving up to 42.1 % higher performance for realistic ML inference workloads. Our remote attestation protocol extends to sub-device partitions, providing trustworthiness on a fine-grained level and decouples host and accelerator TEEs into separate attestation reports without degrading security guarantees. Our work presents a new TEE framework for secure and reconfigurable ML accelerators in a multi-tenant cloud environment.
Sandhya Koteshwara, Mengmei Ye, Hubertus Franke, Deming Chen
CLOUD4
2024 OS4C: An Open-Source SR-IOV System for SmartNIC-Based Cloud Platforms
abstract
Smart network interface cards (SmartNICs) are programmable network cards that enable the flexible offloading of network- and application-level functionality. The last several years have seen a significant rise in research related to smart, programmable NICs. Meanwhile, several open-source FPGA-based NIC and networking projects have emerged. However, these projects lack many key features necessary for strong performance in cloud settings. We identify Single Root In-put/Output Virtualization (SR-IOV) as one of the key missing features in these open-source implementations. SR-IOV enables cloud vendors to grant cloud tenants direct access to hardware resources, dramatically reducing the software overheads of device virtualization. We present OS4C, which extends the popular open-source NIC Corundum [1] with support for SR-IOV. We demonstrate that OS4C can improve virtual machine P99.9 network tail latency by up to 17x, throughput by up to 4x, and CPU effort by up to 3.9x compared to software virtualization. On top of this system, we provide a novel weighted round-robin scheduler that enables tenants and providers to control weight distributions and overhaul the Corundum simulation framework to support multi-tenant tests and performance insights.
Marissa Lanz, Bill Dai, Martin Ohmacht, Bharat Sukhwani, Hubertus Franke, Volodymyr V. Kindratenko, Deming Chen
CLOUD7
2024 OS4C: An Open-Source SR-IOV System for SmartNIC-based Cloud Platforms
abstract
OS4C is the first open-source 100 Gbps NIC research platform that supports SR-IOV. The preliminary design and corresponding throughput gains are presented.
Marissa Lanz, Bill Dai, Martin Ohmacht, Bharat Sukhwani, Hubertus Franke, Volodymyr V. Kindratenko, Deming Chen
FCCM7
2024 EcoFaaS: Rethinking the Design of Serverless Environments for Energy Efficiency
abstract
While serverless computing is increasingly popular, its energy and power consumption behavior is hardly explored. In this work, we perform a thorough characterization of the serverless environment and observe that it poses a set of challenges not effectively handled by existing energy-management schemes. Short serverless functions execute in opaque virtualized sandboxes, are idle for a large fraction of their invocation time, context switch frequently, and are co-located in a highly dynamic manner with many other functions of diverse properties. These features are a radical shift from more traditional application environments and require a new approach to manage energy and power. Driven by these insights, we design EcoFaaS, the first energy management framework for serverless environments. EcoFaaS takes a user-provided end-to-end application Service Level Objective (SLO). It then splits the SLO into per-function deadlines that minimize the total energy consumption. Based on the computed deadlines, EcoFaaS sets the optimal per-invocation core frequency using a prediction algorithm. The algorithm performs a fine-grained analysis of the execution time of each invocation, while taking into account the specific invocation inputs. To maximize efficiency, EcoFaaS splits the cores in a server into multiple Core Pools, where all the cores in a pool run at the same frequency and are controlled by a single scheduler. EcoFaaS dynamically changes the sizes and frequencies of the pools based on the current system state. We implement EcoFaaS on two open-source serverless platforms (OpenWhisk and KNative) and evaluate it using diverse serverless applications. Compared to state-of-the-art energy-management systems, EcoFaaS reduces the total energy consumption of serverless clusters by $42 \%$ while simultaneously reducing the tail latency by $34.8 \%$.
Jovan Stojkovic, Nikoleta Iliakopoulou, Tianyin Xu, Hubertus Franke, Josep Torrellas
ISCA4
2024 Power-aware Deep Learning Model Serving with μ-Serve
Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer
USENIX ATC7
2023 Free the Turtles: Removing Nested Virtualization for Performance and Confidentiality in the Cloud
abstract
With the growing popularity of containers running in cloud environments, one might believe techniques like virtual machines (VMs) and nested virtualization are no longer essential for cloud computing, but this is not true. First, they are foundational for the Infrastructure as a Service (IaaS) base that most clouds use for their Kubernetes (K8s) offerings - nodes are deployed in VMs instead of on bare-metal machines to improve flexibility and resource utilization. Second, state-of-theart technology like Kata Containers or KubeVirt deploys VMs inside K8s, to improve container isolation or to offer a cloudnative way to deploy VMs. When such techniques are coupled with VM-based K8s worker node deployments, this requires the use of nested virtualization. However, nested virtualization introduces a larger trusted computing base (TCB) and therefore raises security concerns. In addition, it introduces an important performance overhead and cannot be easily applied to confidential computing, an emerging technology to allow the execution of sensitive workloads in the cloud. In this paper, we propose the secondary-VM (secVM) framework, an alternative to nested virtualization which addresses these challenges by flattening the nested hierarchy, but maintains the resource isolation features of nested virtualization. By removing the emulation layer required for nested virtualization, the secVM framework reduces the TCB, is compatible with any confidential computing techniques, and removes the overhead associated with nested virtualization. Therefore, in reference to the Turtles Project that introduced the concept of nested virtualization in the open-source virtualization technology - Kernel-based Virtual Machine (KVM), we argue it is time to free the “turtles”.
Mengmei Ye, Angelo Ruocco, Daniele Buono, James Bottomley, Hubertus Franke
CLOUD5
2023 Remote attestation of confidential VMs using ephemeral vTPMs
abstract
Trying to address the security challenges of a cloud-centric software deployment paradigm, silicon and cloud vendors are introducing confidential computing – an umbrella term aimed at providing hardware and software mechanisms for protecting cloud workloads from the cloud provider and its software stack. Today, Intel Software Guard Extensions (SGX), AMD secure encrypted virtualization (SEV), Intel trust domain extensions (TDX), etc., provide a way to shield cloud applications from the cloud provider through encryption of the application’s memory below the hardware boundary of the CPU, hence requiring trust only in the CPU vendor. Unfortunately, existing hardware mechanisms do not automatically enable the guarantee that a protected system was not tampered with during configuration and boot time. Such a guarantee relies on a hardware root of trust, i.e., an integrity-protected location that can store measurements in a trustworthy manner, extend them, and authenticate the measurement logs to the user (remote attestation).
Vikram Narayanan, Cláudio Carvalho, Angelo Ruocco, Gheorghe Almási 0001, James Bottomley, Mengmei Ye, Tobin Feldman-Fitzthum, Daniele Buono, Hubertus Franke, Anton Burtsev
ACSAC9
2023 AccShield: a New Trusted Execution Environment with Machine-Learning Accelerators
abstract
Machine learning accelerators such as the Tensor Processing Unit (TPU) are already being deployed in the hybrid cloud, and we foresee such accelerators proliferating in the future. In such scenarios, secure access to the acceleration service and trustworthiness of the underlying accelerators become a concern. In this work, we present AccShield, a new method to extend trusted execution environments (TEEs) to cloud accelerators which takes both isolation and multi-tenancy into security consideration. We demonstrate the feasibility of accelerator TEEs by a proof of concept on an FPGA board. Experiments with our prototype implementation also provide concrete results and insights for different design choices related to link encryption, isolation using partitioning and memory encryption.
William Kozlowski, Sandhya Koteshwara, Mengmei Ye, Hubertus Franke, Deming Chen
DAC5
2023 SpecFaaS: Accelerating Serverless Applications with Speculative Function Execution
abstract
Serverless computing has emerged as a popular cloud computing paradigm. Serverless environments are convenient to users and efficient for cloud providers. However, they can induce substantial application execution overheads, especially in applications with many functions.In this paper, we propose to accelerate serverless applications with a novel approach based on software-supported speculative execution of functions. Our proposal is termed Speculative Function-as-a-Service (SpecFaaS). It is inspired by out-of-order execution in modern processors, and is grounded in a characterization analysis of FaaS applications. In SpecFaaS, functions in an application are executed early, speculatively, before their control and data dependences are resolved. Control dependences are predicted like in pipeline branch prediction, and data dependences are speculatively satisfied with memoization. With this support, the execution of downstream functions is overlapped with that of upstream functions, substantially reducing the end-to-end execution time of applications. We prototype SpecFaaS on Apache OpenWhisk, an open-source serverless computing platform. For a set of applications in a warmed-up environment, SpecFaaS attains an average speedup of 4.6×. Further, on average, the application throughput increases by 3.9× and the tail latency decreases by 58.7%.
Jovan Stojkovic, Tianyin Xu, Hubertus Franke, Josep Torrellas
HPCA3
2023 MXFaaS: Resource Sharing in Serverless Environments for Parallelism and Efficiency
abstract
Although serverless computing is a popular paradigm, current serverless environments have high overheads. Recently, it has been shown that serverless workloads frequently exhibit bursts of invocations of the same function. Such pattern is not handled well in current platforms. Supporting it efficiently can speed-up serverless execution substantially.
Jovan Stojkovic, Tianyin Xu, Hubertus Franke, Josep Torrellas
ISCA3
2023 Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task Similarity
abstract
Multi-agent reinforcement learning (MARL) has primarily focused on solving a single task in isolation, while in practice the environment is often evolving, leaving many related tasks to be solved. In this paper, we investigate the benefits of meta-learning in solving multiple MARL tasks collectively. We establish the first line of theoretical results for meta-learning in a wide range of fundamental MARL settings, including learning Nash equilibria in two-player zero-sum Markov games and Markov potential games, as well as learning coarse correlated equilibria in general-sum Markov games. Under natural notions of task similarity, we show that meta-learning achieves provable sharper convergence to various game-theoretical solution concepts than learning each task separately. As an important intermediate step, we develop multiple MARL algorithms with initialization-dependent convergence guarantees. Such algorithms integrate optimistic policy mirror descents with stage-based value updates, and their refined convergence guarantees (nearly) recover the best known results even when a good initialization is unknown. To our best knowledge, such results are also new and might be of independent interest. We further provide numerical simulations to corroborate our theoretical findings.
Weichao Mao, Haoran Qiu, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Tamer Basar
NeurIPS4
2023 AWARE: Automate Workload Autoscaling with Reinforcement Learning in Production Cloud Systems
Haoran Qiu, Weichao Mao, Chen Wang 0039, Hubertus Franke, Alaa Youssef, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer
USENIX ATC4
2022 SIMPPO: a scalable and incremental online learning framework for serverless resource management
abstract
Serverless Function-as-a-Service (FaaS) offers improved programmability for customers, yet it is not server-"less" and comes at the cost of more complex infrastructure management (e.g., resource provisioning and scheduling) for cloud providers. To maintain service-level objectives (SLOs) and improve resource utilization efficiency, recent research has been focused on applying online learning algorithms such as reinforcement learning (RL) to manage resources. Despite the initial success of applying RL, we first show in this paper that the state-of-the-art single-agent RL algorithm (S-RL) suffers up to 4.8x higher p99 function latency degradation on multi-tenant serverless FaaS platforms compared to isolated environments and is unable to converge during training. We then design and implement a scalable and incremental multi-agent RL framework based on Proximal Policy Optimization (SIMPPO). Our experiments demonstrate that in multi-tenant environments, SIMPPO enables each RL agent to efficiently converge during training and provides online function latency performance comparable to that of S-RL trained in isolation with minor degradation (<9.2%). In addition, SIMPPO reduces the p99 function latency by 4.5x compared to S-RL in multi-tenant cases.
Haoran Qiu, Weichao Mao, Archit Patke, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer
SoCC5
2022 A Mean-Field Game Approach to Cloud Resource Management with Function Approximation
abstract
Reinforcement learning (RL) has gained increasing popularity for resource management in cloud services such as serverless computing. As self-interested users compete for shared resources in a cluster, the multi-tenancy nature of serverless platforms necessitates multi-agent reinforcement learning (MARL) solutions, which often suffer from severe scalability issues. In this paper, we propose a mean-field game (MFG) approach to cloud resource management that is scalable to a large number of users and applications and incorporates function approximation to deal with the large state-action spaces in real-world serverless platforms. Specifically, we present an online natural actor-critic algorithm for learning in MFGs compatible with various forms of function approximation. We theoretically establish its finite-time convergence to the regularized Nash equilibrium under linear function approximation and softmax parameterization. We further implement our algorithm using both linear and neural-network function approximations, and evaluate our solution on an open-source serverless platform, OpenWhisk, with real-world workloads from production traces. Experimental results demonstrate that our approach is scalable to a large number of users and significantly outperforms various baselines in terms of function latency and resource utilization efficiency.
Weichao Mao, Haoran Qiu, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Tamer Basar
NeurIPS4
2020 FReaC Cache: Folded-logic Reconfigurable Computing in the Last Level Cache
abstract
The need for higher energy efficiency has resulted in the proliferation of accelerators across platforms, with custom and reconfigurable accelerators adopted in both edge devices and cloud servers. However, existing solutions fall short in providing accelerators with low-latency, high-bandwidth access to the working set and suffer from the high latency and energy cost of data transfers. Such costs can severely limit the smallest granularity of the tasks that can be accelerated and thus the applicability of the accelerators. In this work, we present FReaC Cache, a novel architecture that natively supports reconfigurable computing in the last level cache (LLC), thereby giving energy-efficient accelerators low-latency, high-bandwidth access to the working set. By leveraging the cache's existing dense memory arrays, buses, and logic folding, we construct a reconfigurable fabric in the LLC with minimal changes to the system, processor, cache, and memory architecture. FReaC Cache is a low-latency, low-cost, and low-power alternative to off-die/offchip accelerators, and a flexible, and low-cost alternative to fixed function accelerators. We demonstrate an average speedup of 3X and Perf/W improvements of 6.1X over an edge-class multi-core CPU, and add 3.5% to 15.3% area overhead per cache slice.
Ashutosh Dhar, Xiaohao Wang, Hubertus Franke, Jinjun Xiong, Jian Huang 0006, Wen-Mei W. Hwu, Nam Sung Kim, Deming Chen
MICRO3
2019 DeepStore: In-Storage Acceleration for Intelligent Queries
abstract
Recent advancements in deep learning techniques facilitate intelligent-query support in diverse applications, such as content-based image retrieval and audio texturing. Unlike conventional key-based queries, these intelligent queries lack efficient indexing and require complex compute operations for feature matching. To achieve high-performance intelligent querying against massive datasets, modern computing systems employ GPUs in-conjunction with solid-state drives (SSDs) for fast data access and parallel data processing. However, our characterization with various intelligent-query workloads developed with deep neural networks (DNNs), shows that the storage I/O bandwidth is still the major bottleneck that contributes 56%--90% of the query execution time.
Vikram S. Mailthody, Zaid Qureshi, Weixin Liang, Ziyan Feng, Simon Garcia de Gonzalo, Youjie Li, Hubertus Franke, Jinjun Xiong, Jian Huang 0006, Wen-Mei W. Hwu
MICRO7
2018 Optimization of Genomics Analysis Pipeline for Scalable Performance in a Cloud Environment
Carlos H. A. Costa, Claudia Misale, Frank Liu 0001, Marcio Silva, Hubertus Franke, Paul Crumley, Bruce D'Amora
BIBM5
2018 Capacity Optimization for Resource Pooling in Virtualized Data Centers with Composable Systems
abstract
Recent research trends exhibit a growing imbalance between the demands of tenants' software applications and the provisioning of hardware resources. Misalignment of demand and supply gradually hinders workloads from being efficiently mapped to fixed-sized server nodes in traditional data centers. The incurred resource holes not only lower infrastructure utilization but also cripple the capability of a data center for hosting large-sized workloads. This deficiency motivates the development of a new rack-wide architecture referred to as the composable system. The composable system transforms traditional server racks of static capacity into a dynamic compute platform. Specifically, this novel architecture aims to link up all compute components that are traditionally distributed on traditional server boards, such as central processing unit (CPU), random access memory (RAM), storage devices, and other application-specific processors. By doing so, a logically giant compute platform is created and this platform is more resistant against the variety of workload demands by breaking the resource boundaries among traditional server boards. In this paper, we introduce the concepts of this reconfigurable architecture and design a framework of the composable system for cloud data centers. We then develop mathematical models to describe the resource usage patterns on this platform and enumerate some types of workloads that commonly appear in data centers. From the simulations, we show that the composable system sustains nearly up to 1.6 times stronger workload intensity than that of traditional systems and it is insensitive to the distribution of workload demands. This demonstrates that this composable system is indeed an effective solution to support cloud data center services.
An-Dee Lin, Chung-Sheng Li, Wanjiun Liao, Hubertus Franke
IEEE Trans. Parallel Distributed Syst.4
2018 Learning-Based Memory Allocation Optimization for Delay-Sensitive Big Data Processing
abstract
Optimal resource provisioning is essential for scalable big data analytics. However, it has been difficult to accurately forecast the resource requirements before the actual deployment of these applications as their resource requirements are heavily application and data dependent. This paper identifies the existence of effective memory resource requirements for most of the big data analytic applications running inside JVMs in distributed Spark environments. Provisioning memory less than the effective memory requirement may result in rapid deterioration of the application execution in terms of its total execution time. A machine learning-based prediction model is proposed in this paper to forecast the effective memory requirement of an application given its service level agreement. This model captures the memory consumption behavior of big data applications and the dynamics of memory utilization in a distributed cluster environment. With an accurate prediction of the effective memory requirement, it is shown that up to 60 percent savings of the memory resource is feasible if an execution time penalty of 10 percent is acceptable. The accuracy of the model is evaluated on a physical Spark cluster with 128 cores and 1TB of total memory. The experiment results show that the proposed solution can predict the minimum required memory size for given acceptable delays with high accuracy, even if the behavior of target applications is unknown during the training of the model.
Linjiun Tsai, Hubertus Franke, Chung-Sheng Li, Wanjiun Liao
IEEE Trans. Parallel Distributed Syst.2
2017 Dynamic atomicity: optimizing swift memory management
abstract
Swift is a modern multi-paradigm programming language with an extensive developer community and open source ecosystem. Swift 3's memory management strategy is based on Automatic Reference Counting (ARC) augmented with unsafe APIs for manually-managed memory. We have seen ARC consume as much as 80% of program execution time. A significant portion of ARC's direct performance cost can be attributed to its use of atomic machine instructions to protect reference count updates from data races. Consequently, we have designed and implemented dynamic atomicity, an optimization which safely replaces atomic reference-counting operations with nonatomic ones where feasible. The optimization introduces a store barrier to detect possibly intra-thread references, compiler-generated recursive reference-tracers to find all affected objects, and a bit of state in each reference count to encode its atomicity requirements.
David M. Ungar, David Grove, Hubertus Franke
DLS3
2017 Composable architecture for rack scale big data computing
Chung-Sheng Li, Hubertus Franke, Colin Parris, Bülent Abali, Mukil Kesavan, Victor Chang 0001
Future Gener. Comput. Syst.2
2014 Pythia: Faster Big Data in Motion through Predictive Software-Defined Network Optimization at Runtime
abstract
The rise of Internet of Things sensors, social networking and mobile devices has led to an explosion of available data. Gaining insights into this data has led to the area of Big Data analytics. The MapReduce framework, as implemented in Hadoop, is one of the most popular frameworks for Big Data analysis. To handle the ever-increasing data size, Hadoop is a scalable framework that allows dedicated, seemingly unbound numbers of servers to participate in the analytics process. Response time of an analytics request is an important factor for time to value/insights. While the compute and disk I/O requirements can be scaled with the number of servers, scaling the system leads to increased network traffic. Arguably, the communication-heavy phase of MapReduce contributes significantly to the overall response time, the problem is further aggravated, if communication patterns are heavily skewed, as is not uncommon in many MapReduce workloads. In this paper we present a system that reduces the skew impact by transparently predicting data communication volume at runtime and mapping the many end-to-end flows among the various processes to the underlying network, using emerging software-defined networking technologies to avoid hotspots in the network. Dependent on the network oversubscription ratio, we demonstrate reduction in job completion time between 3% and 46% for popular MapReduce benchmarks like Sort and Nutch.
Marcelo Veiga Neves, César A. F. De Rose, Kostas Katrinis, Hubertus Franke
IPDPS4
2013 Improving virtualization in the presence of software managed translation lookaside buffers
abstract
Virtualization has become an important technology that is used across many platforms, particularly servers, to increase utilization, multi-tenancy and security. Virtualization introduces additional overhead that often relates to memory management, interrupt handling and hypervisor mode switching. Among those, memory management and translation lookaside buffer (TLB) management have been shown to have a significant impact on the performance of systems. Two principal mechanisms for TLB management exist in today's systems, namely software and hardware managed TLBs. In this paper, we analyze and quantify the overhead of a pure software virtualization that is implemented over a software managed TLB. We then describe our design of hardware extensions to support virtualization in systems with software managed TLBs to remove the most dominant overheads. These extensions were implemented in the Power embedded A2 core, which is used in the PowerEN and in the Blue Gene/Q processors. They were used to implement a KVM port. We evaluate each of these hardware extensions to determine their overall contributions to performance and efficiency. Collectively these extensions demonstrate an average improvement of 232% over a pure software implementation.
Xiaotao Chang, Hubertus Franke, Yi Ge, Kun Wang 0005, Jimi Xenidis
ISCA2
2012 Performance evaluation of interthread communicationmechanisms on multicore/multithreaded architectures
abstract
The three major solutions for increasing the nominal performance of a CPU are: multiplying the number of cores per socket, expanding the embedded cache memories and use multi-threading to reduce the impact of the deep memory hierarchy. Systems with tens or hundreds of hardware threads, all sharing a cache coherent UMA or NUMA memory space, are today the de-facto standard. While these solutions can easily provide benefits in a multi-program environment, they require recoding of applications to leverage the available parallelism. Threads must synchronize and exchange data, and the overall performance is heavily in influenced by the overhead added by these mechanisms, especially as developers try to exploit finer grain parallelism to be able to use all available resources.
Davide Pasetto, Massimiliano Meneghin, Hubertus Franke, Fabrizio Petrini, Jimi Xenidis
HPDC3
2011 Optimization of stateful hardware acceleration in hybrid architectures
abstract
In many computing domains, hardware accelerators can improve throughput and lower power consumption, instead of executing functionally equivalent software on the general-purpose micro-processors cores. While hardware accelerators often are stateless, network processing exemplifies the need for stateful hardware acceleration. The packet oriented streaming nature of current networks enables data processing as soon as packets arrive rather than when the data of the whole network flow is available. Due to the concurrence of many flows, an accelerator must maintain and switch contexts between many states of the various accelerated streams embodied in the flows, which increases overhead associated with acceleration. We propose and evaluate dynamic reordering of requests of different accelerated streams in a hybrid on-chip/memory based request queue in order to reduce the associated overhead.
Xiaotao Chang, Yike Ma, Hubertus Franke, Kun Wang 0005, Rui Hou 0001, Hao Yu 0008, Terry Nelms
DATE3
2011 Efficient data streaming with on-chip accelerators: Opportunities and challenges
abstract
The transistor density of microprocessors continues to increase as technology scales. Microprocessors designers have taken advantage of the increased transistors by integrating a significant number of cores onto a single die. However, a large number of cores are met with diminishing returns due to software and hardware scalability issues and hence designers have started integrating on-chip special-purpose logic units (i.e., accelerators) that were previously available as PCI-attached units. It is anticipated that more accelerators will be integrated on-chip due to the increasing abundance of transistors and the fact that not all logic can be powered at all times due to power budget limits. Thus, on-chip accelerator architectures deserve more attention from the research community. There is a wide spectrum of research opportunities for design and optimization of accelerators. This paper attempts to bring out some insights by studying the data access streams of on-chip accelerators that hopefully foster some future research in this area. Specifically, this paper uses a few simple case studies to show some of the common characteristics of the data streams introduced by on-chip accelerators, discusses challenges and opportunities in exploiting these characteristics to optimize the power and performance of accelerators, and then analyzes the effectiveness of some simple optimizing extensions proposed.
Rui Hou 0001, Lixin Zhang 0002, Michael C. Huang 0001, Kun Wang 0005, Hubertus Franke, Yi Ge, Xiaotao Chang
HPCA5
2010 Improving In-memory Column-Store Database Predicate Evaluation Performance on Multi-core Systems
abstract
The ability to analyze a large volume of data for the purpose of business intelligence has led to various innovations in database technology. One example is the increased interest of using column-oriented data layout to address query performance in analytical and warehousing workloads. As system architectures move towards multi-core designs, it is important to address optimizing performance for these workloads on these platforms. In this paper we present SPHINX, an architecture that utilizes multi-core systems for search-based predicate evaluation operations in analytical query workloads against in-memory column store. We discuss the natural parallelism of predicate evaluations and various bottlenecks that impact search performance. We present several performance improvement techniques and apply a scan sharing technique based on cache reuse efficiency to further improve the performance. We demonstrate the performance benefits of our scan sharing scheduler over other scheduling approaches in a workload of mixed search queries.
Hong Min, Hubertus Franke
SBAC-PAD2
2008 Stateful hardware decompression in networking environment
abstract
Compression and Decompression can significantly lower the network bandwidth requirements for common internet traffic. Driven by the demands of an enterprise network intrusion system, this paper defines and examines the requirements of popular dictionary-based decompression in the real-time network processing scenario. In particular, a "stateful" decompression is required that arises out of the packet oriented nature of current networks, where the decompression of the data of a packet depends on the decompressed contents of its preceeding packets composing the same data stream. We propose an effective hardware decompression acceleration engine, which fetches the history data into the accelerator's fast memory on-demand and hides the associated latency by exploring the parallelism of the dictionary-based decompression process. We specify and evaluate various design and implementation options of the fetch-on-demand mechanism, i.e. prefetch most frequently used history, on-accelerator history buffer management, and reuse of fetched history data. Through simulation-based performance study, we show the effectiveness of the proposed mechanism on hiding the overhead of stateful decompression. We further show the effects of the design options and the impact on the overall performance of the network service stack of an intrusion prevension system.
Hao Yu 0008, Hubertus Franke, Giora Biran, Amit Golander, Terry Nelms, Brian M. Bass
ANCS2
2007 Thermal-aware task scheduling at the system software level
abstract
Power-related issues have become important considerations in current generation microprocessor design. One of these issues is that of elevated on-chip temperatures. This has an adverse effect on cooling cost and, if not addressed suitably, on chip reliability. In this paper we investigate the general trade-offs between temporal and spatial hot spot mitigation schemes and thermal time constants, workload variations and microprocessor power distributions. By leveraging spatial and temporal heat slacks, our schemes enable lowering of on-chip unit temperatures by changing the workload in a timely manner with Operating System(OS) and existing hardware support.
Jeonghwan Choi, Chen-Yong Cher, Hubertus Franke, Hendrik F. Hamann, Alan J. Weger, Pradip Bose
ISLPED3
2004 Storage Power Management for Cluster Servers Using Remote Disk Access
Jong Hyuk Choi, Hubertus Franke
Euro-Par2
2004 Synthesizing Representative I/O Workloads for TPC-H
abstract
Synthesizing I/O requests that can accurately capture workload behavior is extremely valuable for the design, implementation and optimization of disk subsystems. This paper presents a synthetic workload generator for TPC-H, an important decision-support commercial workload, by completely characterizing the arrival and access patterns of its queries. We present a novel approach for parameterizing the behavior of inter-mingling streams of sequential requests, and exploit correlations between multiple attributes of these requests, to generate disk block-level traces that are shown to accurately mimic the behavior of a real trace in terms of response time characteristics for each TPC-H query.
Jianyong Zhang, Anand Sivasubramaniam, Hubertus Franke, Natarajan Gautam, Yanyong Zhang, Shailabh Nagar
HPCA3
2004 A Web-Based Multi-Agent System for Transportation Management to Protect Our Natural Environment
abstract
To improve the situation of wasting natural resources, the existing transportation systems have to be optimized. This means that we should not only think about new technologies for saving energy but also about better use of the existing trafficways. The most efficient way to achieve these objectives is to automate the existing means of communication and to improve transportation management. Since most of the communication channels and technologies for automation in transportation management are Web-based, we want to describe how to improve Web-based transportation management. Talking about Web-based environments means talking about the Internet and services it offers. These Web-based environments build up a good basis for an agent-based approach, because all aspects for communication and information processing are also used in agent systems. In this approach, Web servers build the agents by themselves and an agent-based interaction works with the support of Web services. Thus, we can build an agent-based structure for transportation control that is similar to the structure of the Internet. Agent-based transportation management is a possible contribution to make transportation management more effective in regard to saving energy (fuel) and protecting our environment by stopping the increase of existing trafficways.
Hubertus Franke, Wilhelm Dangelmaier
Cybern. Syst.1
2003 Decision-Support Workload Characteristics on a Clustered Database Server from the OS Perspective
abstract
A range of database services are being offered on clusters of workstations today to meet the demanding needs of applications with voluminous datasets, high computational and I/O requirements and a large number of users. The underlying database engine runs on cost-effective off-the-shelf hardware and software components that may not really be tailored/tuned for these applications. At the same time, many of these databases have legacy codes that may not be easy to modulate based on the evolving capabilities and limitations of clusters. An indepth understanding of the interaction between these database engines and the underlying operating system (OS) can identify a set of characteristics that would be extremely valuable for future research on systems support for these environments. To our knowledge, there is no prior work that has embarked on such a characterization for a clustered database server. Using IBM DB2 Universal Database (UDB) Extended Enterprise Edition (EEE) V7.2 Trial version and TPC-H like/sup 1/ decision support queries, this paper studies numerous issues by evaluating performance on an off-the-shelf Pentium/Linux cluster connected by Myrinet. These include detailed performance profiles of all kernel activities, as well as qualitative and quantitative insights on the interaction between the database engine and the operating system.
Yanyong Zhang, Jianyong Zhang, Anand Sivasubramaniam, Chun Liu 0001, Hubertus Franke
ICDCS5
2003 DRPM: Dynamic Speed Control for Power Mangagement in Server Class Disks
abstract
A large portion of the power budget in server environments goes into the I/O subsystem - the disk array in particular. Traditional approaches to disk power management involve completely stopping the disk rotation, which can take a considerable amount of time, making them less useful in cases where idle times between disk requests may not be long enough to outweigh the overheads. We present a new approach called DRPM to modulate disk speed (RPM) dynamically, and gives a practical implementation to exploit this mechanism. Extensive simulations with different workload and hardware parameters show that DRPM can provide significant energy savings without compromising much on performance. We also discuss practical issues when implementing DRPM on server disks.
Sudhanva Gurumurthi, Anand Sivasubramaniam, Mahmut T. Kandemir, Hubertus Franke
ISCA4
2003 Interplay of energy and performance for disk arrays running transaction processing workloads
abstract
The growth of business enterprises and the emergence of the Internet as a medium for data processing has led to a proliferation of applications that are server-centric. The power dissipation of such servers has a major consequence not only on the costs and environmental concerns of power generation and delivery, but also on their reliability and on the design of cooling and packaging mechanisms for these systems. This paper examines the energy and performance ramifications in the design of disk arrays which consume a major portion of the power in transaction processing environments. Using traces of TPC-C and TPC-H running on commercial servers, we conduct in-depth simulations of energy and performance behavior of disk arrays with different RAID configurations. Our results demonstrate that conventional disk power optimizations that have been previously proposed and evaluated for single disk systems' (laptops/workstations) are not very effective in server environments, even if we can design disks than have extremely fast spinup/spindown latencies and predict the idle periods accurately. On the other hand, tuning RAID parameters (RAID type, number of disks, stripe size etc.) has more impact on the power and performance behavior of these systems, sometimes having opposite effects on these two criteria.
Sudhanva Gurumurthi, Jianyong Zhang, Anand Sivasubramaniam, Mahmut T. Kandemir, Hubertus Franke, Narayanan Vijaykrishnan, Mary Jane Irwin
ISPASS5
2003 An Integrated Approach to Parallel Scheduling Using Gang-Scheduling, Backfilling, and Migration
abstract
Effective scheduling strategies to improve response times, throughput, and utilization are an important consideration in large supercomputing environments. Parallel machines in these environments have traditionally used space-sharing strategies to accommodate multiple jobs at the same time by dedicating the nodes to a single job until it completes. This approach, however, can result in low system utilization and large job wait times. This paper discusses three techniques that can be used beyond simple space-sharing to improve the performance of large parallel systems. The first technique we analyze is backfilling, the second is gang-scheduling, and the third is migration. The main contribution of this paper is an analysis of the effects of combining the above techniques. Using extensive simulations based on detailed models of realistic workloads, the benefits of combining the various techniques are shown over a spectrum of performance criteria.
Yanyong Zhang, Hubertus Franke, José E. Moreira, Anand Sivasubramaniam
IEEE Trans. Parallel Distributed Syst.2
2002 Characterizing the Scalability of Decision-Support Workloads on Clusters and SMP Systems
abstract
Using a public domain version of a commercial clustered database server and TPC-H like decision support queries, this paper studies the performance and scalability issues of a Pentium/Linux cluster and an 8-way Linux SMP. The execution profile demonstrates the dominance of the I/O subsystem in the execution, and the importance of the communication subsystem for cluster scalability. In addition to quantifying their importance, this paper provides further details on how these subsystems are exercised by the database engine. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Yanyong Zhang, Anand Sivasubramaniam, Jianyong Zhang, Shailabh Nagar, Hubertus Franke
Euro-Par5
2002 Creating and preserving locality of java applications at allocation and garbage collection times
abstract
The growing gap between processor and memory speeds is motivating the need for optimization strategies that improve data locality. A major challenge is to devise techniques suitable for pointer-intensive applications. This paper presents two techniques aimed at improving the memory behavior of pointer-intensive applications with dynamic memory allocation, such as those written in Java. First, we present an allocation time object placement technique based on the recently introduced notion of prolific (frequently instantiated) types. We attempt to co-locate, at allocation time, objects of prolific types that are connected via object references. Then, we present a novel locality based graph traversal technique. The benefits of this technique, when applied to garbage collection (GC), are twofold: (i) it improves the performance of GC due to better locality during a heap traversal and (ii) it restructures surviving objects in a way that enhances locality. On multiprocessors, this technique can further reduce overhead due to synchronization and false sharing. The experimental results, on a well-known suite of Java benchmarks (SPECjvm98 [26], SPECjbb2000 [27], and jOlden [4]), from an implementation of these techniques in the Jikes RVM [1], are very encouraging. The object co-allocation technique improves application performance by up to 21% (10% on average) in the Jikes RVM configured with a non-copying mark-and-sweep collector. The locality-based traversal technique reduces GC times by up to 20% (10% on average) and improves the performance of applications by up to 14% (6% on average) in the Jikes RVM configured with a copying semi-space collector. Both techniques combined can improve application performance by up to 22% (10% on average) in the Jikes RVM configured with a non-copying mark-and-sweep collector.
Yefim Shuf, Manish Gupta 0002, Hubertus Franke, Andrew W. Appel, Jaswinder Pal Singh
OOPSLA3
2002 Modeling and analysis of dynamic coscheduling in parallel and distributed environments
abstract
Scheduling in large-scale parallel systems has been and continues to be an important and challenging research problem. Several key factors, including the increasing use of off-the-shelf clusters of workstations to build such parallel systems, have resulted in the emergence of a new class of scheduling strategies, broadly referred to as dynamic coscheduling. Unfortunately, the size of both the design and performance spaces of these emerging scheduling strategies is quite large, due in part to the numerous dynamic interactions among the different components of the parallel computing environment as well as the wide range of applications and systems that can comprise the parallel environment. This in turn makes it difficult to fully explore the benefits and limitations of the various proposed dynamic coscheduling approaches for large-scale systems solely with the use of simulation and/or experimentation.To gain a better understanding of the fundamental properties of different dynamic coscheduling methods, we formulate a general mathematical model of this class of scheduling strategies within a unified framework that allows us to investigate a wide range of parallel environments. We derive a matrix-analytic analysis based on a stochastic decomposition and a fixed-point iteration. A large number of numerical experiments are performed in part to examine the accuracy of our approach. These numerical results are in excellent agreement with detailed simulation results. Our mathematical model and analysis is then used to explore several fundamental design and performance tradeoffs associated with the class of dynamic coscheduling policies across a broad spectrum of parallel computing environments.
Mark S. Squillante, Yanyong Zhang, Anand Sivasubramaniam, Natarajan Gautam, Hubertus Franke, José E. Moreira
SIGMETRICS5
2001 Performance of Hardware Compressed Main Memory
abstract
A new memory subsystem called Memory Expansion Technology (MXT) has been built for compressing main memory contents. MXT effectively doubles the physically available memory. This paper provides an analysis of the performance impact of memory compression using the SPEC2000 benchmarks and a database benchmark. Results show that the hardware compression of memory has a negligible performance penalty compared to a standard memory. We also show that many applications' memory contents can be compressed usually by a factor of two to one. We demonstrate this using industry benchmarks, web server benchmarks, and contents of popular web sites.
Bülent Abali, Hubertus Franke, Dan E. Poff, T. Basil Smith
HPCA2
2001 An Integrated Approach to Parallel Scheduling Using Gang-Scheduling, Backfilling, and Migration
Yanyong Zhang, Hubertus Franke, José E. Moreira, Anand Sivasubramaniam
JSSPP2
2001 Hardware Compressed Main Memory: Operating System Support and Performance Evaluation
abstract
A new memory subsystem, called Memory Xpansion Technology (MXT), has been built for compressing main memory contents. MXT effectively doubles the physically available memory transparently to the CPUs, input/output devices, device drivers, and application software. An average compression ratio of two or greater has been observed for many applications. Since compressibility of memory contents varies dynamically, the size of the memory managed by the operating system is not fixed. In this paper, we describe operating system techniques that can deal with such dynamically changing memory sizes. We also demonstrate the performance impact of memory compression using the SPEC CPU2000 and SPECweb99 benchmarks. Results show that the hardware compression of memory has a negligible performance penalty compared to a standard memory for many applications. For memory starved applications and benchmarks such as SPECweb99, memory compression improves the performance significantly. Results also show that the memory contents of many applications can be compressed, usually by a factor of two to one.
Bülent Abali, Mohammad Banikazemi, Hubertus Franke, Dan E. Poff, T. Basil Smith
IEEE Trans. Computers4
2001 Impact of Workload and System Parameters on Next Generation Cluster Scheduling Mechanisms
abstract
Scheduling of processes onto processors of a parallel machine has always been an important and challenging area of research. The issue becomes even more crucial and difficult as we gradually progress to the use of off-the-shelf workstations, operating systems, and high bandwidth networks to build cost-effective clusters for demanding applications. Clusters are gaining acceptance not just in scientific applications that need supercomputing power, but also in domains such as databases, web service, and multimedia which place diverse Quality-of-Service (QoS) demands on the underlying system. Further, these applications have diverse characteristics in terms of their computation, communication, and I/O requirements, making conventional parallel scheduling solutions, such as space sharing or gang scheduling, unattractive. At the same time, leaving it to the native operating system of each node to make decisions independently can lead to ineffective use of system resources whenever there is communication. Instead, an emerging class of dynamic coscheduling mechanisms that attempt to take remedial actions to guide the system toward coscheduled execution without requiring explicit synchronization offers a lot of promise for cluster scheduling. Using a detailed simulator, this paper evaluates the pros and cons of different dynamic coscheduling alternatives while comparing their advantages over traditional gang scheduling (and not performing any coordinated scheduling at all). The impact of dynamic job arrivals, job characteristics, and different system parameters on these alternatives is evaluated in terms of several performance criteria. In addition, heuristics to enhance one of the alternatives even further are identified, classified, and evaluated. It is shown that these heuristics can significantly outperform the other alternatives over a spectrum of workload and system parameters and is thus a much better option for clusters than conventional gang scheduling.
Yanyong Zhang, Anand Sivasubramaniam, José E. Moreira, Hubertus Franke
IEEE Trans. Parallel Distributed Syst.4
2000 The Impact of Migration on Parallel Job Scheduling for Distributed Systems
Yanyong Zhang, Hubertus Franke, José E. Moreira, Anand Sivasubramaniam
Euro-Par2
2000 A simulation-based study of scheduling mechanisms for a dynamic cluster environment
abstract
Scheduling of processes onto processors of a parallel machine has always been an important and challenging area of research. The issue becomes even more crucial and difficult as we gradually progress to the use of off-the-shelf workstations, operating systems, and high bandwidth networks to build cost-effective clusters for demanding applications. Clusters are gaining acceptance not just in scientific applications that need supercomputing power, but also in domains such as databases, web service and multimedia, which place diverse Quality-of-Service (QoS) demands on the underlying system. Further, these applications have diverse characteristics in terms of their computation, communication and I/O requirements, making conventional parallel scheduling solutions, such as space sharing or coscheduling, an unattractive option. At the same time, leaving it to the native operating system of each node to make decisions independently can lead to ineffective use of system resources whenever there is communication. Instead, an emerging class of dynamic coscheduling mechanisms, that attempt to take remedial actions to guide the system towards coscheduled execution without requiring explicit synchronization, offer a lot of promise for cluster scheduling. Using a detailed simulator, this paper evaluates the pros and cons of different dynamic coscheduling alternatives, while comparing their advantages over traditional coscheduling (and not performing any coordinated scheduling at all). The impact of dynamic job arrivals, job characteristics and different system parameters on these alternatives are evaluated in terms of several performance criteria.
Yanyong Zhang, Anand Sivasubramaniam, José E. Moreira, Hubertus Franke
ICS4
2000 Improving Parallel Job Scheduling by Combining Gang Scheduling and Backfilling Techniques
abstract
Two different approaches have been commonly used to address problems associated with space sharing scheduling strategies: (a) augmenting space sharing with backfilling, which performs out of order job scheduling; and (b) augmenting space sharing with time sharing, using a technique called coscheduling or gang scheduling. With three important experimental results-impact of priority queue order on backfilling, impact of overestimation of job execution times, and comparison of scheduling techniques-this paper presents an integrated strategy that combines backfilling with gang scheduling. Using extensive simulations based on detailed models of realistic workloads, the benefits of combining backfilling and gang scheduling are clearly demonstrated over a spectrum of performance criteria.
Yanyong Zhang, Anand Sivasubramaniam, Hubertus Franke, José E. Moreira
IPDPS3
1999 An Evaluation of Parallel Job Scheduling for ASCI Blue-Pacific
abstract
In this paper we analyze the behavior of a gang-scheduling strategy that we are developing for the ASCI Blue-Pacific machines. Using actual job logs for one of the ASCI machines we generate a statistical model of the current workload with hyper Erlang distributions. We then vary the parameters of those distributions to generate various workloads, representative of different operating points of the machine. Through simulation we obtain performance parameters for three different scheduling strategies: (i) first-come first-serve, (ii) gang-scheduling, and (iii) backfilling. Our results show that backfilling, can be very effective for the common operating points in the 60-70% utilization range. However, for higher utilization rates, time-sharing techniques such as gang-scheduling offer much better performance.
Hubertus Franke, Joefon Jann, José E. Moreira, Pratap Pattnaik, Morris A. Jette
SC1
1998 An Infrastructure for Efficient Parallel Job Execution in Terascale Computing Environments
abstract
Recent Terascale computing environments, such as those in the Department of Energy Accelerated Strategic Computing Initiative, present a new challenge to job scheduling and execution systems. The traditional way to concurrently execute multiple jobs in such large machines is through space-sharing: each job is given dedicated use of a pool of processors. Previous work in this area has demonstrated the benefits of sharing the parallel machine's resources not only spatially but also temporally. Time-sharing creates virtual processors for the execution of jobs. The scheduling is typically performed cyclically and each time-slice of the cycle can be considered an independent virtual machine. When all tasks of a parallel job are scheduled to run on the same time-slice (same virtual machine), gang-scheduling is accomplished. Research has shown that gang-scheduling can greatly improve system utilization and job response time in large parallel systems. We are developing GangLL, a research prototype system for performing gang-scheduling on the ASCI Blue-Pacific machine, an IBM RS/6000 SP to be installed at Lawrence Livermore National Laboratory. This machine consists of several hundred nodes, interconnected by a high-speed communication switch. GangLL is organized as a centralized scheduler that performs global decision-making, and a local daemon in each node that controls job execution according to those decisions. The centralized scheduler builds an Ousterhout matrix that precisely defines the temporal and spatial allocation of tasks in the system. Once the matrix is built, it is distributed to each of the local daemons using a scalable hierarchical distributions scheme. A two-phase commit is used in the distribution scheme to guarantee that all local daemons have consistent information. The local daemons enforce the schedule dedicated by the Ousterhout matrix in their corresponding nodes. This requires suspending and resuming execution of tasks and multiplexing access to the communication switch. Large supercomputing centers tend to have their own job scheduling systems, to handle site specific conditions. Therefore, we are designing GangLL so that it can interact with an external site scheduler. The goal is to let the site scheduler control spatial allocation of jobs, if so desired, and to decide when jobs run. GangLL then performs the detailed temporal allocation and controls the actual execution of jobs. The site scheduler can control the fraction of a shared processor that a job receives through an execution factor parameter. To quantify the benefits of our gang-scheduling system to job execution in a large parallel system, we simulate the system with a realistic workload. We measure performance parameters under various degrees of time-sharing, characterized by the multiprogramming level. Our results show that higher multiprogramming levels lead to higher system utilization and lower job response times. We also report some results from the initial deployment of GangLL on a small multiprocessor system.
José E. Moreira, Waiman Chan, Liana L. Fong, Hubertus Franke, Morris A. Jette
SC4
1997 Modeling of Workload in MPPs
Joefon Jann, Pratap Pattnaik, Hubertus Franke, Joseph Skovira, Joseph Riordan
JSSPP3
1996 A Gang Scheduling Design for Multiprogrammed Parallel Computing Environments
Hubertus Franke, Marios C. Papaefthymiou, Pratap Pattnaik, Larry Rudolph, Mark S. Squillante
JSSPP2
1995 MPI Programming Environment for IBM SP1/SP2
abstract
In this paper we discuss an implementation of the message passing interface standard (MPI) for the IBM Scalable Power PARALLEL 1 and 2 (SP1, SP2). Key to a reliable and efficient implementation of a message passing library on these machines is the careful design of a UNIX-Socket like layer in the user space with controlled access to the communication adapters and with adequate recovery and flow control. The performance of this implementation is at the same level as the IBM-proprietary message passing library (MPL). We also show that in the IBM SP1 and SP2 we achieve integrated tracing ability, where both system events, such as context switches and page fault etc., and MPI related activities are traced, with minimal overhead to the application program, thus presenting application programmers the trace of all the events that ultimately affect efficiency of a parallel program.
Hubertus Franke, Ching-Farn Eric Wu, Michel Riviere, Pratap Pattnaik, Marc Snir
ICDCS1
1995 Model-embedded on-line problem solving environment for chemical engineering
abstract
The building of custom monitoring, control, simulation and diagnostics applications for complex chemical plants necessitates the integration of models into the problem solving process. This paper describes a system and its practical applications that supports this activity. It is based on the Multigraph Architecture, which is a generic framework for building these model-based systems. The paper discusses the modeling paradigms used, how the applications are generated, and some practical, existing applications.
Gabor Karsai, Janos Sztipanovits, Hubertus Franke, Samir Padalkar, Frank DeCaria
ICECCS3
1994 MPI-F: An Efficient Implementation of MPI on IBM-SP1
abstract
This article introduces MPI-F an efficient implementation of MPI on the IBM-SP1 distributed memory cluster. After discussing the novel and key concepts of MPI and how they relate to an implementation, the MPI-F system architecture is outlined in detail. Although many incorrectly assume that MPI will not be efficient due to its increased functionality, MPI-F performance demonstrates efficiency as good as the best message passing library currently available on the SP1.
Hubertus Franke, Peter Hochschild, Pratap Pattnaik, Marc Snir
ICPP (3)1