VLDB 2026 Research / reviewers in the wild / expert
Cristian Klein
dblp:24/7115 · also Cristian Klein-Halmaghi
· DBLP profile ↗
24ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0003-0106-3049ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ESTHER: Application-First Hardware-Level QoS-Enforcement for Cloud Native EnvironmentsabstractRecent advances in multi-core chip technology have enabled the dynamic tuning of shared memory resources, such as last-level cache and memory bus bandwidth. However, despite proven performance benefits, the complexity of effectively utilizing these hardware-level QoS enforcement features has limited their adoption in real-world cloud computing environments. In this paper, we introduce ESTHER, a novel approach to autonomously fine-tune QoS enforcement features in cloud environments using extremum seeking control, focusing on applications needs and operator ease-of-use. We demonstrate that ESTHER effectively maintains latency-critical workload SLOs and rapidly resolves any infringements by prioritizing shared memory resources. Such fast node-level resolution of SLO violations ensures that costly cluster-level scaling events may be avoided. Furthermore, ESTHER improves best-effort job throughput without impacting latency-critical workloads, achieving performance gains without utilizing workload profiling or prior knowledge of system dynamics. Oliver Larsson, Thijs Metsch, Cristian Klein, Erik Elmroth |
CLOUD | 3 |
| 2025 | FaLSE: A Failure and Latency-Aware Scheduling for Mission-Critical Applications at the EdgeabstractMission-critical applications, such as real-time emergency response, healthcare, and transport systems, depend heavily on the low latency and reliability provided by Mobile Edge Computing (MEC). The failure of such applications can lead to high latency and severe consequences, including loss of life, financial catastrophe, or operational disruption. However, the dependability of edge clusters is often overlooked, particularly in terms of fault awareness and recovery strategies, which are crucial to these applications. In this work, we focus on loosely coupled IoT applications and propose a Failure and Latency-aware Scheduling approach for Edge (FaLSE) that balances the trade-off between the availability of edge clusters and the latency of containerized mission-critical applications. We used a decentralized network coordinate system to estimate latency between IoT devices/users and nodes. To validate the proposed approach, we compare it with the standard Kubernetes scheduler, which is currently among the most widely used workload orchestration platforms. The results indicate that FaLSE reduced the failure request rate by 87.9% while maintaining a 71.97% lower 95th percentile latency for mission-critical applications and a 10.63% lower latency for normal applications compared to the standard Kubernetes scheduler. Nayereh Rasouli, Cristian Klein, Erik Elmroth |
CloudCom | 2 |
| 2024 | State-Aware Application Placement in Mobile Edge CloudsabstractPlacing applications within Mobile Edge Clouds (MEC) poses challenges due to dynamic user mobility. Maintaining optimal Quality of Service may require frequent application migration in response to changing user locations, potentially leading to bandwidth wastage. This paper addresses application placement challenges in MEC environments by developing a comprehensive model covering workloads, applications, and MEC infrastructures. Following this, various costs associated with application operation, including resource utilization, migration overhead, and potential service quality degradation, are systematically formulated. An online application placement algorithm, App EDC Match, inspired by the Gale-Shapley matching algorithm, is introduced to optimize application placement considering these cost factors. Through experiments that employ real mobility traces to simulate workload dynamics, the results demonstrate that the proposed algorithm efficiently determines near-optimal application placements within Edge Data Centers. It achieves total operating costs within a narrow margin of 8% higher than the approximate global optimum attained by the offline precognition algorithm, which assumes access to future user locations. Additionally, the proposed placement algorithm effectively mitigates resource scarcity in MEC. Chanh Nguyen 0001, Cristian Klein, Erik Elmroth |
CLOSER | 2 |
| 2023 | HydraGen: A Microservice Benchmark GeneratorabstractMicroservice-based architectures have become ubiq-uitous in large-scale software systems. Experimental cloud re-searchers constantly propose enhanced resource management mechanisms for such systems. These mechanisms need to be eval-uated using both realistic and flexible microservice benchmarks to study in which ways diverse application characteristics can affect their performance and scalability. However, current mi-croservice benchmarks have limitations including static compu-tational complexity, limited architectural scale, and fixed topology (i.e., number of tiers, fan-in, and fan-out characteristics). We therefore propose HydraGen, a tool that enables re-searchers to systematically generate benchmarks with different computational complexities and topologies, to tackle experimental evaluation of performance at scale for web-serving applications, with a focus on inter-service communication. To illustrate the potential of our open-source tool, we demonstrate how it can reproduce an existing microservice benchmark with preserved architectural properties. We also demonstrate how HydraGen can enrich the evaluation of cloud management systems based on a case study related to traffic engineering. Mohammad Reza Saleh Sedghpour, Aleksandra Obeso Duque, Xuejun Cai, Björn Skubic, Erik Elmroth, Cristian Klein, Johan Tordsson |
CLOUD | 6 |
| 2023 | Breaking the Vicious Circle: Self-Adaptive Microservice Circuit Breaking and RetryabstractMicroservice-based architectures consist of numerous, loosely coupled services with multiple instances. Service meshes aim to simplify traffic management and prevent microservice overload through circuit breaking and request retry mechanisms. Previous studies have demonstrated that the static configuration of these mechanisms is unfit for the dynamic environment of microservices. We conduct a sensitivity analysis to understand the impact of retrying across a wide range of scenarios. Based on the findings, we propose a retry controller that can also work with dynamically configured circuit breakers. We have empirically assessed our proposed controller in various scenarios, including transient overload and noisy neighbors while enforcing adaptive circuit breaking. The results show that our proposed controller does not deviate from a well-tuned configuration while maintaining carried response time and adapting to the changes. In comparison to the default static retry configuration that is mostly used in practice, our approach improves the carried throughput up to 12x and 32x respectively in the cases of transient overload and noisy neighbors. Mohammad Reza Saleh Sedghpour, David Garlan, Bradley R. Schmerl, Cristian Klein, Johan Tordsson |
IC2E | 4 |
| 2022 | A Qualitative Evaluation of Service Mesh-based Traffic Management for Mobile Edge CloudabstractService mesh is getting widely adopted as the cloud-native mechanism for traffic management in microservice-based applications, in particular for generic IT workloads hosted in more centralized cloud environments. Performance-demanding applications continue to drive the decentralization of modern application execution environments, as in the case of mobile edge cloud. This paper presents a systematic and qualitative analysis of state-of-the-art service mesh to evaluate how suitable its design is for addressing the traffic management needs of performance-demanding application workloads hosted in a mobile edge cloud environment. With this analysis, we argue that today's dependability-centric service mesh design fails at addressing the needs of the different types of emerging mobile edge cloud workloads and motivate further research in the directions of performance-efficient architectures, stronger QoS guarantees and higher complexity abstractions of cloud-native traffic manage-ment frameworks. Aleksandra Obeso Duque, Cristian Klein, Jinhua Feng, Xuejun Cai, Björn Skubic, Erik Elmroth |
CCGRID | 2 |
| 2022 | An Empirical Study of Service Mesh Traffic Management Policies for MicroservicesabstractA microservice architecture features hundreds or even thousands of small loosely coupled services with multiple instances. Because microservice performance depends on many factors including the workload, inter-service traffic management is complex in such dynamic environments. Service meshes aim to handle this complexity and to facilitate management, observability, and communication between microservices. Service meshes provide various traffic management policies such as circuit breaking and retry mechanisms, which are claimed to protect microservices against overload and increase the robustness of communication between microservices. However, there have been no systematic studies on the effects of these mechanisms on microservice performance and robustness. Furthermore, the exact impact of various tuning parameters for circuit breaking and retries are poorly understood. This work presents a large set of experiments conducted to investigate these issues using a representative microservice benchmark in a Kubernetes testbed with the widely used Istio service mesh. Our experiments reveal effective configurations of circuit breakers and retries. The findings presented will be useful to engineers seeking to configure service meshes more systematically and also open up new areas of research for academics in the area of service meshes for (autonomic) microservice resource management. Mohammad Reza Saleh Sedghpour, Cristian Klein, Johan Tordsson |
ICPE | 2 |
| 2021 | Adaptive and Application-agnostic Caching in Service Meshes for Resilient Cloud ApplicationsabstractService meshes factor out code dealing with inter-micro-service communication. The overall resilience of a cloud application is improved if constituent micro-services return stale data, instead of no data at all. This paper proposes and implements application agnostic caching for micro services. While caching is widely employed for serving web service traffic, its usage in inter-micro-service communication is lacking. Micro-services responses are highly dynamic, which requires carefully choosing adaptive time-to-life caching algorithms. Our approach is application agnostic, is cloud native, and supports gRPC. We evaluate our approach and implementation using the micro-service benchmark by Google Cloud called Hipster Shop. Our approach results in caching of about 80% of requests. Results show the feasibility and efficiency of our approach, which encourages implementing caching in service meshes. Additionally, we make the code, experiments, and data publicly available. Lars Larsson 0001, William Tärneberg, Cristian Klein, Maria Kihl, Erik Elmroth |
NetSoft | 3 |
| 2020 | Elasticity Control for Latency-Intolerant Mobile Edge ApplicationsabstractElasticity is a fundamental property required for Mobile Edge Clouds (MECs) to become mature computing platforms hosting software applications. However, MECs must cope with several challenges that do not arise in the context of conventional cloud platforms. These include the potentially highly distributed geographical deployment, heterogeneity, and limited resource capacity of Edge Data Centers (EDCs), and end-user mobility. In this paper, we present an elasticity controller to help MECs overcome these challenges by automatic proactive resource scaling. The controller utilizes information on the physical locations of EDCs and the correlation of workload changes in physically neighboring EDCs to predict request arrival rates at EDCs. These predictions are used as inputs for a queueing theory-driven performance model that estimates the number of resources that should be provisioned to EDCs in order to meet predefined Service Level Objectives (SLOs) while maximizing resource utilization. The controller also incorporates a group-level load balancer that is responsible for redirecting requests among EDCs during runtime so as to minimize the request rejection rate. We evaluate our approach by performing simulations with an emulated MEC deployed over a metropolitan area and a simulated application workload using a real-world user mobility trace. The results show that our proposed pro-active controller exhibits better scaling behavior than a state-of-the-art re-active controller and increases the efficiency of resource provisioning, thereby helping MECs to sustain resource utilization and rejection rates that satisfy predefined SLOs while maintaining system stability. Chanh Nguyen 0001, Cristian Klein, Erik Elmroth |
SEC | 2 |
| 2020 | Security-Performance Trade-offs of Kubernetes Container RuntimesabstractThe extreme adoption rate of container technologies along with raised security concerns have resulted in the development of multiple alternative container runtimes targeting security through additional layers of indirection. In an apples-to-apples comparison, we deploy three runtimes in the same Kubernetes cluster, the security focused Kata and gVisor, as well as the default Kubernetes runtime runC. Our evaluation based on three real applications demonstrate that runC outperforms the more secure alternatives up to 5x, that gVisor deploys containers up to 2x faster than Kata, but that Kata executes container up to 1.6x faster than gVisor. Our work illustrates that alternative, more secure, runtimes can be used in a plug-and-play manner in Kubernetes, but at a significant performance penalty. Our study is useful both to practitioners - to understand the current state of the technology in order to make the right decision in the selection, operation and/or design of platforms - and to scholars to illustrate how these technologies evolved over time. William Viktorsson, Cristian Klein, Johan Tordsson |
MASCOTS | 2 |
| 2020 | Impact of etcd deployment on Kubernetes, Istio, and application performanceabstractSummary This experience article describes lessons learned as we conducted experiments in a Kubernetes‐based environment, the most notable of which was that the performance of both the Kubernetes control plane and the deployed application depends strongly and in unexpected ways on the performance of the etcd database. The article contains (a) detailed descriptions of how networking with and without Istio works in Kubernetes, based on the Flannel Container Networking Interface (CNI) provider in VXLAN mode with IP Virtual Server (IPVS)‐backed Kubernetes Services, (b) a comprehensive discussion about how to conduct load and performance testing using a closed‐loop workload generator, and (c) an open source experiment framework useful for executing experiments in a shared cloud environment and exploring the resulting data. It also shows that statistical analysis may reveal the data resulting from such experiments to be misleading even when careful preparations are made, and that nondeterministic behavior stemming from etcd can affect both the platform as a whole and the deployed application. Finally, it is demonstrated that using high‐performance backing storage for etcd can reduce the occurrence of such nondeterministic behaviors by a statistically significant (P < .05) margin. The implication of this experience article is that systems researchers studying the performance of applications deployed on Kubernetes cannot simply consider their specific application to be under test. Instead, the particularities of the underlying Kubernetes and cloud platform must be taken into account, in particular because their performance can impact that of etcd. Lars Larsson 0001, William Tärneberg, Cristian Klein, Erik Elmroth, Maria Kihl |
Softw. Pract. Exp. | 3 |
| 2019 | Multivariate LSTM-Based Location-Aware Workload Prediction for Edge Data CentersabstractMobile Edge Clouds (MECs) is a promising computing platform to overcome challenges for the success of bandwidth-hungry, latency-critical applications by distributing computing and storage capacity in the edge of the network as Edge Data Centers (EDCs) within the close vicinity of end-users. Due to the heterogeneous distributed resource capacity in EDCs, the application deployment flexibility coupled with the user mobility, MECs bring significant challenges to control resource allocation and provisioning. In order to develop a self-managed system for MECs which efficiently decides how much and when to activate scaling, where to place and migrate services, it is crucial to predict its workload characteristics, including variations over time and locality. To this end, we present a novel location-aware workload predictor for EDCs. Our approach leverages the correlation among workloads of EDCs in a close physical distance and applies multivariate Long Short-Term Memory network to achieve on-line workload predictions for each EDC. The experiments with two real mobility traces show that our proposed approach can achieve better prediction accuracy than a state-of-the art location-unaware method (up to 44%) and a location-aware method (up to 17%). Further, through an intensive performance measurement using various input shaking methods, we substantiate that the proposed approach achieves a reliable and consistent performance. Chanh Nguyen 0001, Cristian Klein, Erik Elmroth |
CCGRID | 2 |
| 2018 | Virtualization Techniques Compared: Performance, Resource, and Power Usage Overheads in CloudsabstractVirtualization solutions based on hypervisors or containers are enabling technologies for scalable, flexible, and cost-effective resource sharing. As the fundamental limitations of each technology are yet to be understood, they need to be regularly reevaluated to better understand the trade-off provided by latest technological advances. This paper presents an in-depth quantitative analysis of virtualization overheads in these two groups of systems and their gaps relative to native environments based on a diverse set of workloads that stress CPU, memory, storage, and networking resources. KVM and XEN are used to represent hypervisor-based virtualization, and LXC and Docker for container-based platforms. The systems were evaluated with respect to several cloud resource management dimensions including performance, isolation, resource usage, energy efficiency, start-up time, and density. Our study is useful both to practitioners to understand the current state of the technology in order to make the right decision in the selection, operation and/or design of platforms and to scholars to illustrate how these technologies evolved over time. Selome Kostentinos Tesfatsion, Cristian Klein, Johan Tordsson |
ICPE | 2 |
| 2017 | KPI-agnostic Control for Fine-Grained Vertical ElasticityabstractApplications hosted in the cloud have become indispensable in several contexts, with their performance often being key to business operation and their running costs needing to be minimized. To minimize running costs, most modern virtualization technologies such as Linux Containers, Xen, and KVM offer powerful resource control primitives for individual provisioning - that enable adding or removing of fraction of cores and/or megabytes of memory for as short as few seconds. Despite the technology being ready, there is a lack of proper techniques for fine-grained resource allocation, because there is an inherent challenge in determining the correct composition of resources an application needs, with varying workload, to ensure deterministic performance. This paper presents a control-based approach for the management of multiple resources, accounting for the resource consumption, together with the application performance, enabling fine-grained vertical elasticity. The control strategy ensures that the application meets the target performance indicators, consuming as less resources as possible. We carried out an extensive set of experiments using different applications - interactive with response-time requirements, as well as noninteractive with throughput desires - by varying the workload mixes of each application over time. The results demonstrate that our solution precisely provides guaranteed performance while at the same time avoiding both resource over-and underprovisioning. Ewnetu Bayuh Lakew, Alessandro Vittorio Papadopoulos, Martina Maggio, Cristian Klein, Erik Elmroth |
CCGrid | 4 |
| 2017 | Incentivizing self-capping to increase cloud utilizationabstractCloud Infrastructure as a Service (IaaS) providers continually seek higher resource utilization to better amortize capital costs. Higher utilization not only can enable higher profit for IaaS providers but also provides a mechanism to raise energy efficiency; therefore creating greener cloud services. Unfortunately, achieving high utilization is difficult mainly due to infrastructure providers needing to maintain spare capacity to service demand fluctuations. Mohammad Shahrad, Cristian Klein, Liang Zheng 0002, Mung Chiang, Erik Elmroth, David Wentzlaff |
SoCC | 2 |
| 2017 | Control Strategies for Self-Adaptive Software SystemsabstractThe pervasiveness and growing complexity of software systems are challenging software engineering to design systems that can adapt their behavior to withstand unpredictable, uncertain, and continuously changing execution environments. Control theoretical adaptation mechanisms have received growing interest from the software engineering community in the last few years for their mathematical grounding, allowing formal guarantees on the behavior of the controlled systems. However, most of these mechanisms are tailored to specific applications and can hardly be generalized into broadly applicable software design and development processes. This article discusses a reference control design process, from goal identification to the verification and validation of the controlled system. A taxonomy of the main control strategies is introduced, analyzing their applicability to software adaptation for both functional and nonfunctional goals. A brief extract on how to deal with uncertainty complements the discussion. Finally, the article highlights a set of open challenges, both for the software engineering and the control theory research communities. Antonio Filieri, Martina Maggio, Konstantinos Angelopoulos, Nicolás D'Ippolito, Ilias Gerostathopoulos, Andreas B. Hempel, Henry Hoffmann, Pooyan Jamshidi, Evangelia Kalyvianaki, Cristian Klein, Filip Krikava, Sasa Misailovic, Alessandro Vittorio Papadopoulos, Suprio Ray, Amir Molzam Sharifloo, Stepan Shevtsov, Mateusz Ujma, Thomas Vogel 0001 |
ACM Trans. Auton. Adapt. Syst. | 10 |
| 2015 | Performance-Based Service Differentiation in CloudsabstractDue to fierce competition, cloud providers need to run their data-centers efficiently. One of the issues is to increase data-center utilization while maintaining applications' performance targets. Achieving high data-center utilization while meeting applications' performance is difficult, as data-center overload may lead to poor performance of hosted services. Service differentiation has been proposed to control which services get degraded. However, current approaches are capacity-based, which are oblivious to the observed performance of each service and cannot divide the available capacity among hosted services so as to minimize overall performance degradation. In this paper we propose performance-based service differentiation. In case enough capacity is available, each service is automatically allocated the right amount of capacity that meets its target performance, expressed either as response time or throughput. In case of overload, we propose two service differentiation schemes that dynamically decide which services to degrade and to what extent. We carried out an extensive set of experiments using different services -- interactive as well as non-interactive -- by varying the workload mixes of each service over time. The results demonstrate that our solution precisely provides guaranteed performance or service differentiation depending on available capacity. Ewnetu Bayuh Lakew, Cristian Klein, Francisco Hernández-Rodriguez, Erik Elmroth |
CCGRID | 2 |
| 2015 | High performance fault-tolerance for cloudsabstractCloud computing and virtualized infrastructures are currently the baseline environments for the provision of services in different application domains. While the number of service consumers increasingly grows, service providers aim at exploiting infrastructures that enable non-disruptive service provisioning, thus minimizing or even eliminating downtime. Nonetheless, to achieve the latter current approaches are either application-specific or cost inefficient, requiring the use of dedicated hardware. In this paper we present the reference architecture of a fault-tolerance scheme, which not only enhances cloud environments with the aforementioned capabilities but also achieves high-performance as required by mission critical every day applications. To realize the proposed approach, a new paradigm for memory and I/O externalization and consolidation is introduced, while current implementation references are also provided. Dimosthenis Kyriazis, Vasileios I. Anagnostopoulos, Andrea Arcangeli, Dimitrios Kalogeras, Ronen I. Kat, Cristian Klein, Panagiotis C. Kokkinos, Yossi Kuperman, Joel Nider, Petter Svärd, Luis Tomás, Emmanouel A. Varvarigos, Theodora A. Varvarigou |
ISCC | 7 |
| 2014 | Brownout: building more robust cloud applicationsabstractSelf-adaptation is a first class concern for cloud applications, which should be able to withstand diverse runtime changes. Variations are simultaneously happening both at the cloud infrastructure level - for example hardware failures - and at the user workload level - flash crowds. However, robustly withstanding extreme variability, requires costly hardware over-provisioning. Cristian Klein, Martina Maggio, Karl-Erik Årzén, Francisco Hernández-Rodriguez |
ICSE | 1 |
| 2014 | Improving Cloud Service Resilience Using Brownout-Aware Load-BalancingabstractWe focus on improving resilience of cloud services (e.g., e-commerce website), when correlated or cascading failures lead to computing capacity shortage. We study how to extend the classical cloud service architecture composed of a load-balancer and replicas with a recently proposed self-adaptive paradigm called brownout. Such services are able to reduce their capacity requirements by degrading user experience (e.g., disabling recommendations). Combining resilience with the brownout paradigm is to date an open practical problem. The issue is to ensure that replica self-adaptivity would not confuse the load-balancing algorithm, overloading replicas that are already struggling with capacity shortage. For example, load-balancing strategies based on response times are not able to decide which replicas should be selected, since the response times are already controlled by the brownout paradigm. In this paper we propose two novel brownout-aware load-balancing algorithms. To test their practical applicability, we extended the popular lighttpd web server and load-balancer, thus obtaining a production-ready implementation. Experimental evaluation shows that the approach enables cloud services to remain responsive despite cascading failures. Moreover, when compared to Shortest Queue First (SQF), believed to be near-optimal in the non-adaptive case, our algorithms improve user experience by 5%, with high statistical significance, while preserving response time predictability. Cristian Klein, Alessandro Vittorio Papadopoulos, Manfred Dellkrantz, Jonas Durango, Martina Maggio, Karl-Erik Årzén, Francisco Hernández-Rodriguez, Erik Elmroth |
SRDS | 1 |
| 2013 | Introducing service-level awareness in the cloudabstractManaging the resources of a virtualized data-center is a key issue in cloud computing [1]. Existing research mostly assumes that applications are either allocated the required resources or fail [2--15]. Combined with the fact that most cloud applications have dynamic resource requirements [16], this imposes a fundamental limitation to cloud computing: To guarantee on-demand resource allocations, the data-center needs large spare capacity, leading to inefficient resource utilization. Cristian Klein, Martina Maggio, Karl-Erik Årzén, Francisco Hernández-Rodriguez |
SoCC | 1 |
| 2011 | An RMS for Non-predictably Evolving ApplicationsabstractNon-predictably evolving applications are applications that change their resource requirements during execution. These applications exist, for example, as a result of using adaptive numeric methods, such as adaptive mesh refinement and adaptive particle methods. Increasing interest is being shown to have such applications acquire resources on the fly. However, current HPC Resource Management Systems (RMSs) only allow a static allocation of resources, which cannot be changed after it started. Therefore, non-predictably evolving applications cannot make efficient use of HPC resources, being forced to make an allocation based on their maximum expected requirements. This paper presents CooRMv2, an RMS which supports efficient scheduling of non-predictably evolving applications. An application can make "pre-allocations" to specify its peak resource usage. The application can then dynamically allocate resources as long as the pre-allocation is not outgrown. Resources which are pre-allocated but not used, can be filled by other applications. Results show that the approach is feasible and leads to a more efficient resource usage. Cristian Klein, Christian Pérez |
CLUSTER | 1 |
| 2011 | An RMS Architecture for Efficiently Supporting Complex-Moldable ApplicationsabstractHigh-performance scientific applications are becoming increasingly complex, in particular because of the coupling of parallel codes. This results in applications having a complex structure, characterized by multiple deploy-time parameters, such as the number of processes of each code. In order to optimize the performance of these applications, the parameters have to be carefully chosen, a process which is highly resource dependent. However, the abstractions provided by current Resource Management Systems (RMS) -- either submitting rigid jobs or enumerating a list of moldable configurations -- are insufficient to efficiently select resources for such applications. This paper introduces CooRM, an RMS architecture that delegates resource selection to applications while still keeping control over the resources. The proposed architecture is evaluated using a simulator which is then validated with a proof-of-concept implementation on Grid'5000. Results show that such a system is feasible and performs well with respect to scalability and fairness. Cristian Klein, Christian Pérez |
HPCC | 1 |
| 2009 | Generating high-performance custom floating-point pipelinesabstractCustom operators, working at custom precisions, are a key ingredient to fully exploit the FPGA flexibility advantage for high-performance computing. Unfortunately, such operators are costly to design, and application designers tend to rely on less efficient off-the-shelf operators. To address this issue, an open-source architecture generator framework is introduced. Its salient features are an easy learning curve from VHDL, the ability to embed arbitrary synthesizable VHDL code, portability to mainstream FPGA targets from Xilinx and Altera, automatic management of complex pipelines with support for frequency-directed pipeline, and automatic test-bench generation. This generator is presented around the simple example of a collision detector, which it significantly improves in accuracy, DSP count, logic usage, frequency and latency with respect to an implementation using standard floating-point operators. Florent de Dinechin, Cristian Klein, Bogdan Pasca 0001 |
FPL | 2 |