Holger Fröning

dblp:56/3662 · DBLP profile ↗
← Back
46ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-9562-0680ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 33 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Characterizing Needleman-Wunsch Sequence Alignment on the Graphcore IPU: Memory Trade-Offs and Optimization Strategies
S. Kazem Shekofteh, Nils Kochendörfer, Holger Fröning
Euro-Par (1)3
2026 A Tabular Schedule Abstraction for Communication-Aware Evaluation of Pipeline-Parallel LLM Training
Daniel Barley, Jonathan Leis, Benjamin Klenk, Holger Fröning
ISPDC4
2026 Butterfly factorization for vision transformers on multi-IPU systems
abstract
Recent advances in machine learning have led to increasingly large and complex models, placing significant demands on computation and memory. Techniques such as Butterfly factorization have emerged to reduce model parameters and memory footprints while preserving accuracy. Specialized hardware accelerators, such as Graphcore’s Intelligence Processing Units (IPUs), are designed to address these challenges through massive parallelism and efficient on-chip memory utilization. In this paper, we extend our analysis of Butterfly structures for efficient utilization on single and multiple IPUs, comparing their performance with GPUs. These structures drastically reduce the number of parameters and memory footprint while preserving model accuracy. Experimental results on the Graphcore GC200 IPU chip, compared with an NVIDIA A30 GPU, demonstrate a 98.5% compression ratio, with speedups of 1.6 × and 1.3 × for Butterfly and Pixelated Butterfly structures, respectively. Extending our evaluation to Vision Transformer (ViT) models, we compare Multi-GPU and Multi-IPU systems on the M2000 machine: Multi-GPU reaches a maximum accuracy of 84.51% with a training time of 401.44 min, whereas Multi-IPU attains a higher maximum accuracy of 88.92% with a training time of 694.03 min. These results demonstrate that Butterfly factorization enables substantial compression of ViT layers (up to 97.17%) while improving model accuracy. The findings highlight the promise of IPU machines as a suitable platform for large-scale machine learning model training, especially when coupled with sparsification methods like Butterfly factorization, thanks to their efficient support for model parallelism.
S. Kazem Shekofteh, Daniel Bogacz, Christian Alles, Holger Fröning
Parallel Comput.4
2025 Variance-Aware Noisy Training: Hardening DNNs Against Unstable Analog Computations
Hendrik Borras, Bernhard Klein, Holger Fröning
ECML/PKDD (7)4
2025 GraphMatch: Subgraph Query Processing on Steroids
abstract
Recently, graphs are becoming increasingly interesting in the context of large language models and as overlays for commercial databases. Subgraph query processing is an especially challenging workload for graph analysis that is bottlenecked by slow set intersection performance on CPUs. Previous work has shown the viability of utilizing hardware acceleration for related domains like graph and relational join processing. We propose GraphMatch, a hardware-accelerated subgraph query processing system based on worst-case optimal joins (WCOJ). For efficient processing of various data and query graphs, we propose a novel set intersection algorithm, called MaxStep, that leverages hardware parallelism. GraphMatch combines MaxStep operators in a data flow architecture which efficiently solves multi-set intersections in subgraph query processing, superior to CPU-based approaches. GraphMatch achieves an average speedup of over 6.98x and 17.08x, compared to the state-of-the-art WCOJ-based systems GraphFlow and RapidMatch, respectively. On labeled graphs, GraphMatch outperforms the fastest subgraph query processing accelerator FAST by orders of magnitude.
Jonas Dann, Tobias Götz, Daniel Ritter 0001, Jana Giceva, Holger Fröning, Gustavo Alonso
Proc. ACM Manag. Data5
2024 DeepHYDRA: A Hybrid Deep Learning and DBSCAN-Based Approach to Time-Series Anomaly Detection in Dynamically-Configured Systems
abstract
Anomaly detection in distributed systems such as High-Performance Computing (HPC) clusters is vital for early fault detection, performance optimisation, security monitoring, reliability in general but also operational insights. It enables proactive measures to address issues, ensuring system reliability, resource efficiency, and protection against potential threats. Deep Neural Networks have seen successful use in detecting long-term anomalies in multidimensional data, originating for instance from industrial or medical systems, or weather prediction. A downside of such methods is that they require a static input size, or lose data through cropping, sampling, or other dimensionality reduction methods, making deployment on systems with variability on monitored data channels, such as computing clusters difficult. To address these problems, we present DeepHYDRA (Deep Hybrid DBSCAN/Reduction-Based Anomaly Detection) which combines DBSCAN and learning-based anomaly detection. DBSCAN clustering is used to find point anomalies in time-series data, mitigating the risk of missing outliers through loss of information when reducing input data to a fixed number of channels. A deep learning-based time-series anomaly detection method is then applied to the reduced data in order to identify long-term outliers. This hybrid approach reduces the chances of missing anomalies that might be made indistinguishable from normal data by the reduction process, and likewise enables the algorithm to be scalable and tolerate partial system failures while retaining its detection capabilities. Using a subset of the well-known SMD dataset family, a modified variant of the Eclipse dataset, as well as an in-house dataset with a large variability in active data channels, made publicly available with this work, we furthermore analyse computational intensity, memory footprint, and activation counts. DeepHYDRA is shown to reliably detect different types of anomalies in both large and complex datasets. At the same time, the applied reduction approach is shown to enable real-time anomaly detection of a whole computing cluster while occupying proportionally miniscule compute resources, enabling its usage on existing systems without the need for hardware changes.
Kevin Stehle, Wainer Vandelli, Felix Zahn, Giuseppe Avolio, Holger Fröning
ICS5
2024 Walking Noise: On Layer-Specific Robustness of Neural Architectures Against Noisy Computations and Associated Characteristic Learning Dynamics
Hendrik Borras, Bernhard Klein, Holger Fröning
ECML/PKDD (4)3
2024 Resource-Efficient Neural Networks for Embedded Systems
abstract
While machine learning is traditionally a resource intensive task, embedded systems, autonomous navigation, and the vision of the Internet of Things fuel the interest in resource-efficient approaches. These approaches aim for a carefully chosen trade-off between performance and resource consumption in terms of computation and energy. The development of such approaches is among the major challenges in current machine learning research and key to ensure a smooth transition of machine learning technology from a scientific environment with virtually unlimited computing resources into everyday's applications. In this article, we provide an overview of the current state of the art of machine learning techniques facilitating these real-world requirements. In particular, we focus on resource-efficient inference based on deep neural networks (DNNs), the predominant machine learning models of the past decade. We give a comprehensive overview of the vast literature that can be mainly split into three non-mutually exclusive categories: (i) quantized neural networks, (ii) network pruning, and (iii) structural efficiency. These techniques can be applied during training or as post-processing, and they are widely used to reduce the computational demands in terms of memory footprint, inference speed, and energy efficiency. We also briefly discuss different concepts of embedded hardware for DNNs and their compatibility with machine learning techniques as well as potential for energy and latency reduction. We substantiate our discussion with experiments on well-known benchmark data sets using compression techniques (quantization, pruning) for a set of resource-constrained embedded systems, such as CPUs, GPUs and FPGAs. The obtained results highlight the difficulty of finding good trade-offs between resource efficiency and prediction quality.
Wolfgang Roth, Günther Schindler, Bernhard Klein, Robert Peharz, Sebastian Tschiatschek, Holger Fröning, Franz Pernkopf, Zoubin Ghahramani
J. Mach. Learn. Res.6
2024 GraphScale: Scalable Processing on FPGAs for HBM and Large Graphs
abstract
Recent advances in graph processing on FPGAs promise to alleviate performance bottlenecks with irregular memory access patterns. Such bottlenecks challenge performance for a growing number of important application areas like machine learning and data analytics. While FPGAs denote a promising solution through flexible memory hierarchies and massive parallelism, we argue that current graph processing accelerators either use the off-chip memory bandwidth inefficiently or do not scale well across memory channels. In this work, we propose GraphScale, a scalable graph processing framework for FPGAs. GraphScale combines multi-channel memory with asynchronous graph processing (i.e., for fast convergence on results) and a compressed graph representation (i.e., for efficient usage of memory bandwidth and reduced memory footprint). GraphScale solves common graph problems like breadth-first search, PageRank, and weakly connected components through modular user-defined functions, a novel two-dimensional partitioning scheme, and a high-performance two-level crossbar design. Additionally, we extend GraphScale to scale to modern high-bandwidth memory (HBM) and reduce partitioning overhead of large graphs with binary packing.
Jonas Dann, Daniel Ritter 0001, Holger Fröning
ACM Trans. Reconfigurable Technol. Syst.3
2023 CUDAsap: Statically-Determined Execution Statistics as Alternative to Execution-Based Profiling
abstract
Today a variety of different GPU types exists, raising questions regarding high-level tasks such as provisioning and scheduling. To predict execution time on different GPU types accurately, we propose a method to obtain execution statistics based on compile-time static code analysis, in which the control flow graph for the code's basic blocks is determined. This graph is represented as an adjacency matrix and used in a system of linear equations to calculate the basic block execution frequencies. Kernel execution itself is not necessary for this analysis. We analyze the proposed method for five different benchmark suites, showing that 76 out of 79 evaluated kernels can be analyzed with an average error of 0.4 %, primarily due to different LLVM versions, with an average prediction time of 203.96 ms. Furthermore, repetitive kernels make memoization effective, and the underlying analysis is largely independent of problem size.
Yannick Emonds, Lorenz Braun, Holger Fröning
CCGrid3
2023 Characterization of data compression across CPU platforms and accelerators
abstract
Abstract The ever increasing amount of generated data makes it more and more beneficial to utilize compression to trade computations for data movement and reduced storage requirements. Lately, dedicated accelerators have been introduced to offload compression tasks from the main processor. However, research is lacking when it comes to the system costs for incorporating compression. This is especially true for the influence of the CPU platform and accelerators on the compression. This work will show that for general‐purpose lossless compression algorithms following can be recommended: (1) snappy for high throughput, but low compression ratio; (2) zstandard level 2 for moderate throughput and compression ratio; (3) xz level 5 for low throughput, but high compression ratio. And it will show that the selected platforms (ARM, IBM or Intel) have no influence on the algorithm's performance. Furthermore, it will show that the accelerator's zlib implementation achieves a comparable compression ratio as zlib level 2 on a CPU, while having up to the throughput and utilizing over 80% less CPU resources. This suggests that the overhead of offloading compression is limited but present. Overall, this work will allow system designers to identify deployment opportunities for compression while considering integration constraints.
Laura Promberger, Rainer Schwemmer, Holger Fröning
Concurr. Comput. Pract. Exp.3
2022 PipeJSON: Parsing JSON at Line Speed on FPGAs
abstract
JavaScript Object Notation (JSON) gained popularity as a data exchange and storage format. While recent advances on modern CPUs show an improved JSON parsing by using data parallelism with vector instructions, the rigid instruction set and limited pipelining of CPUs prevent parsing performance from reaching the practical limit of memory bandwidth.
Jonas Dann, Royden Wagner, Daniel Ritter 0001, Christian Färber, Holger Fröning
DaMoN5
2022 GraphScale: Scalable Bandwidth-Efficient Graph Processing on FPGAs
abstract
Recent advances in graph processing on FPGAs promise to alleviate performance bottlenecks with irregular memory access patterns. Such bottlenecks challenge performance for a growing number of important application areas like machine learning and data analytics. While FPGAs denote a promising solution through flexible memory hierarchies and massive parallelism, we argue that current graph processing accelerators either use the off-chip memory bandwidth inefficiently or do not scale well across memory channels. In this work, we propose GraphScale, a scalable graph processing framework for FPGAs. For the first time, Graph-Scale combines multi-channel memory with asynchronous graph processing (i. e., for fast convergence on results) and a com-pressed graph representation (i. e., for efficient usage of memory bandwidth and reduced memory footprint). GraphScale solves common graph problems like breadth-first search, PageRank, and weakly -connected components through modular user-defined functions, a novel two-dimensional partitioning scheme, and a high-performance two-level crossbar design.
Jonas Dann, Daniel Ritter 0001, Holger Fröning
FPL3
2022 Joint Program and Layout Transformations to Enable Convolutional Operators on Specialized Hardware Based on Constraint Programming
abstract
The success of Deep Artificial Neural Networks (DNNs) in many domains created a rich body of research concerned with hardware accelerators for compute-intensive DNN operators. However, implementing such operators efficiently with complex hardware intrinsics such as matrix multiply is a task not yet automated gracefully. Solving this task often requires joint program and data layout transformations. First solutions to this problem have been proposed, such as TVM, UNIT, or ISAMIR, which work on a loop-level representation of operators and specify data layout and possible program transformations before the embedding into the operator is performed. This top-down approach creates a tension between exploration range and search space complexity, especially when also exploring data layout transformations such as im2col, channel packing, or padding. In this work, we propose a new approach to this problem. We created a bottom-up method that allows the joint transformation of both computation and data layout based on the found embedding. By formulating the embedding as a constraint satisfaction problem over the scalar dataflow, every possible embedding solution is contained in the search space. Adding additional constraints and optimization targets to the solver generates the subset of preferable solutions. An evaluation using the VTA hardware accelerator with the Baidu DeepBench inference benchmark shows that our approach can automatically generate code competitive to reference implementations. Further, we show that dynamically determining the data layout based on intrinsic and workload is beneficial for hardware utilization and performance. In cases where the reference implementation has low hardware utilization due to its fixed deployment strategy, we achieve a geomean speedup of up to × 2.813, while individual operators can improve as much as × 170.
Dennis Rieber, Axel Acosta 0001, Holger Fröning
ACM Trans. Archit. Code Optim.3
2021 A Simple Model for Portable and Fast Prediction of Execution Time and Power Consumption of GPU Kernels
abstract
Characterizing compute kernel execution behavior on GPUs for efficient task scheduling is a non-trivial task. We address this with a simple model enabling portable and fast predictions among different GPUs using only hardware-independent features. This model is built based on random forests using 189 individual compute kernels from benchmarks such as Parboil, Rodinia, Polybench-GPU, and SHOC. Evaluation of the model performance using cross-validation yields a median Mean Average Percentage Error (MAPE) of 8.86–52.0% for time and 1.84–2.94% for power prediction across five different GPUs, while latency for a single prediction varies between 15 and 108 ms.
Lorenz Braun, Sotirios Nikas, Vincent Heuveline, Holger Fröning
ACM Trans. Archit. Code Optim.5
2020 Towards Real-Time Single-Channel Singing-Voice Separation with Pruned Multi-Scaled Densenets
abstract
Modern musical source separation systems based on deep neural networks reach unprecedented levels of separation quality. However, harnessing the power of these large-scale models in typical audio production environments, which frequently offer only limited computing resources while demanding real-time processing, remains challenging. We extend the multi-scaled DenseNet in several aspects to facilitate real-time source separation scenarios. Specifically, we reduce the computational requirements by inferring Mel-scaled masks and decrease the model size via effective use of bottleneck layers, while improving performance using a deep clustering objective. In addition, we are able to further increase the model efficiency by applying parameterized structured pruning of convolutional weights without any significant impact on the separation performance. We significantly reduce the model size and increase the computational efficiency by a factor of 1.6 and 4.3, respectively, while maintaining the separation performance.
Günther Schindler, Christian Schörkhuber, Wolfgang Roth, Franz Pernkopf, Holger Fröning
ICASSP6
2020 On Network Locality in MPI-Based HPC Applications
abstract
Data movements through interconnection networks exceed local memory accesses in terms of latency as well as energy by multiple orders of magnitude. While many optimizations make great effort to improve memory accesses, large distances in the network can easily dash these improvements, resulting in increased overall costs. Therefore, a deep understanding of network locality is key for further optimizations, such as improved mapping of ranks to physical entities.
Felix Zahn, Holger Fröning
ICPP2
2020 On Resource-Efficient Bayesian Network Classifiers and Deep Neural Networks
abstract
We present two methods to reduce the complexity of Bayesian network (BN) classifiers. First, we introduce quantization-aware training using the straight-through gradient estimator to quantize the parameters of BNs to few bits. Second, we extend a recently proposed differentiable tree-augmented naive Bayes (TAN) structure learning approach by also considering the model size. Both methods are motivated by recent developments in the deep learning community, and they provide effective means to trade off between model size and prediction accuracy, which is demonstrated in extensive experiments. Furthermore, we contrast quantized BN classifiers with quantized deep neural networks (DNNs) for small-scale scenarios which have hardly been investigated in the literature. We show Pareto optimal models with respect to model size, number of operations, and test error and find that both model classes are viable options.
Wolfgang Roth, Franz Pernkopf, Günther Schindler, Holger Fröning
ICPR4
2020 cCUDA: Effective Co-Scheduling of Concurrent Kernels on GPUs
abstract
While GPUs are meantime omnipresent for many scientific and technical computations, they still continue to evolve as processors. An important recent feature is the ability to execute multiple kernels concurrently via queue streams. However, experiments show that different parameters including the behavior of kernels, the order of kernel launches and other execution configurations, e.g., the number of concurrent thread blocks, may result in different execution time for concurrent kernel execution. Since kernels may have different resource requirements, they can be classified into different classes, which are traditionally assumed as either memory-bound or compute-bound. However, a kernel may belong to the different classes on different hardware according to the hardware resources. In this paper, the definition of kernel mix intensity is introduced. Based on this, a scheduling framework called concurrent CUDA (cCUDA) is proposed to co-schedule the concurrent kernels more efficiently. It first profiles and ranks kernels with different execution behaviors and then takes the kernel resource requirements into account to partition thread blocks of different kernels and overlap them to better utilize the GPU resources. Experimental results on real hardware demonstrate performance improvement in terms of execution time of up to 1.86x, and an average speedup of 1.28x for a wide range of kernels. cCUDA is available at https://github.com/kshekofteh/cCUDA.
S. Kazem Shekofteh, Hamid Noori, Mahmoud Naghibzadeh, Holger Fröning, Hadi Sadoghi Yazdi
IEEE Trans. Parallel Distributed Syst.4
2019 Software-Based Buffering of Associative Operations on Random Memory Addresses
abstract
An important concept for indivisible updates in parallel computing are atomic operations. For most architectures, they also provide ordering guarantees, which in practice can hurt performance. For associative and commutative updates, in this paper we present software buffering techniques that overcome the problem of ordering by combining multiple updates in a temporary buffer and by prefetching addresses before updating them. As a result, our buffering techniques reduce contention and avoid unnecessary ordering constraints, in order to increase the amount of memory parallelism. We evaluate our techniques in different scenarios, including applications like histogram and graph computations, and reason about the applicability for standard systems and multi-socket systems.
Matthias Hauck, Marcus Paradies, Holger Fröning
IPDPS3
2019 Training Discrete-Valued Neural Networks with Sign Activations Using Weight Distributions
Wolfgang Roth, Günther Schindler, Holger Fröning, Franz Pernkopf
ECML/PKDD (2)3
2019 Constructing virtual 5-dimensional tori out of lower-dimensional network cards
abstract
Summary In the Top500 and Graph500 lists of the last years, some of the most powerful systems implement a torus topology to interconnect the millions of computing nodes they include. Some of these torus networks are of five or six dimensions, which implies an additional difficulty as the node degree increases. In previous works, we proposed and evaluated the nD Twin (nDT) torus topology to virtually increase the dimensions a torus is able to implement. We showed that this new topology reduces the distances between nodes, increasing, therefore, global network performance. In this work, we present how to build a 5DT torus network using a specific commercial 6‐port network card (EXTOLL card) to interconnect those nodes. We show, using the same number of cards, that the performance of the 5DT torus network we are able to implement using our proposal is higher than the performance of the 3D torus network for the same number of compute nodes.
Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato, Holger Fröning
Concurr. Comput. Pract. Exp.6
2019 On link width scaling for energy-proportional direct interconnection networks
abstract
Summary Energy consumption is one of the most important design parameters for future large‐scale computing systems. While the end of Dennard scaling demands for increasing energy‐proportional components, interconnection networks have not received much attention regarding this topic. However, these networks are expected to contribute about 20% to the overall power consumption of these systems in the near future. Furthermore, this fraction increases if other energy‐proportional components, such as CPUs, accelerators, and memory, are not fully utilized. To avoid becoming the main contributor to power consumption and to reduce overall power consumption, it is mandatory to improve the energy‐proportionality of interconnection networks. In this work, we analyze different aspects of energy‐proportionality in interconnection networks for systems designed within current technical constraints but also for future systems that might be designed with different parameters. First, we discuss the impact of multiple design parameters and the most feasible approach for improved energy consumption, such as transition time and power state granularity. Based on this study, we introduce three different power saving policies, which try to address different requirements. While an on/off policy allows for large energy savings, it can also cause significant performance losses for adverse setups. In order to meet the demand for sustainable performance, we present two new policies that trade power saving potential for performance. For all three workload classes, we use a power‐aware network simulation to report the impact on execution time and energy consumption compared to the current situation and an idealized network. While we show that a highly regular pattern enables power saving possibilities close to the theoretical minimum, even slight deviations from such a highly iterative and temporal behavior demand for further improvements in all policies.
Felix Zahn, Steffen Lammel, Holger Fröning
Concurr. Comput. Pract. Exp.3
2019 Metric Selection for GPU Kernel Classification
abstract
Graphics Processing Units (GPUs) are vastly used for running massively parallel programs. GPU kernels exhibit different behavior at runtime and can usually be classified in a simple form as either “compute-bound” or “memory-bound.” Recent GPUs are capable of concurrently running multiple kernels, which raises the question of how to most appropriately schedule kernels to achieve higher performance. In particular, co-scheduling of compute-bound and memory-bound kernels seems promising. However, its benefits as well as drawbacks must be determined along with which kernels should be selected for a concurrent execution. Classifying kernels can be performed online by instrumentation based on performance counters. This work conducts a thorough analysis of the metrics collected from various benchmarks from Rodinia and CUDA SDK. The goal is to find the minimum number of effective metrics that enables online classification of kernels with a low overhead. This study employs a wrapper-based feature selection method based on the Fisher feature selection criterion. The results of experiments show that to classify kernels with a high accuracy, only three and five metrics are sufficient on a Kepler and a Pascal GPU, respectively. The proposed method is then utilized for a runtime scheduler. The results show an average speedup of 1.18× and 1.1× compared with a serial and a random scheduler, respectively.
S. Kazem Shekofteh, Hamid Noori, Mahmoud Naghibzadeh, Hadi Sadoghi Yazdi, Holger Fröning
ACM Trans. Archit. Code Optim.5
2018 Resource Efficient Deep Eigenvector Beamforming
abstract
We propose binary neural networks (BNN s) for acoustic beamforming. This makes the speech enhancement approach resource efficient and applicable for embedded applications. Using CHiME4 data, we use BNN s to estimate the speech presence probability mask for GEV-PAN beamformers. By doing so, we achieve audio quality and ASR scores on par to single-precision deep neural networks (DNNs), while the computational requirements and the memory footprint are significantly reduced.
Matthias Zöhrer, Lukas Pfeifenberger, Günther Schindler, Holger Fröning, Franz Pernkopf
ICASSP4
2018 Towards Efficient Forward Propagation on Resource-Constrained Systems
Günther Schindler, Matthias Zöhrer, Franz Pernkopf, Holger Fröning
ECML/PKDD (1)4
2018 Heterogeneous and unconventional cluster architectures and applications
abstract
Recent trends in cluster computing and related topics, including processor and memory design, demonstrate a continuous growing need for more processing and memory, both in terms of capacity and performance. Cluster computing has been traditionally at the forefront of such computing systems and, usually, is one of the earliest adopters of future and emerging technologies. With this special issue, we gear to gather recent related works, hoping that the reader finds these contributions helpful for a better understanding of future directions in cluster computing. Before that, we will briefly present a short rationale on our view of cluster computing. This rationale is based on two trends that are most important for such cluster architectures: first, the end of Dennard scaling has led to an era in which the growing amount of transistors, as described by Moore's law, cannot be simultaneously active because of an increasing power density. Second, applications continue to demand more processing and memory capacity. However, economy and technology laws imply that horizontal scaling is usually more cost effective than vertical scaling. Furthermore, he also described power scaling rules for CMOS silicon dies, defined by the observation that a transition to a new processing technology will decrease the feature size (ie, the size of a transistor gate) by a factor α. Based on this factor, characteristics including voltage, current, and capacity scale inversely. Formula 1 then shows that, for a given power budget, one can implement α2 more components ”a” and even increase operating frequency f by a factor of α. Thus, Dennard scaling actually gave Moore's law its teeth by enabling constant power budgets and frequency scaling. Unfortunately, since early 2015, this law is no longer applicable, mainly because voltage scaling is no longer possible because of saturated threshold voltages and because leakage power became a major contributor to the overall power consumption. As a result, a still growing amount of transistors, as described by Moore's law, now results in a growing power budget, which results intype a hard technical constraint. The usual escape path for post-Dennard performance scaling is two-fold: first, one can observe that frequency usually behaves linearly with regard to voltage, effectively making power consumption proportional to frequency cubed. For instance, reducing frequency by half would result in a relative power consumption of 1/8th, allowing to replicate one computational core 8 times while maintaining the power budget. Second, it is a common first-order approximation to assume that performance behaves proportional to frequency. To continue the previous example, such an 8-core design at half the initial frequency would now result in a relative performance improvement of a factor of 4. However, this is obviously only feasible if the application workloads exhibit enough parallelism. Given extreme examples of many-core processors with 1000s of vector units, heterogeneity is usually the chosen solution to also support the sequential parts of the workloads. As a result, we are seeing a huge interest in many-core processors, which, however, only excel in performance for massively parallel workloads. Given that not all workloads fulfill this requirement, heterogeneous architectures are being deployed. Applications continue to demand for an increasing amount of processing power and memory capacity, in particular pushed by Big Data and Machine Learning. For instance, training a recurrent deep neural network requires about 20 ExaFLOPs and still is not being trained with all data available. Similarly, in particular, deep learning required a plethora of data, leading to huge data collections for various tasks. However, the costs of resources like processors or memory do not scale linearly with capability. On the other hand, while the manufacturers are not very candid about the reasons behind, it is well known that the yield of a silicon die production is a function of the die area. While small dies are less likely to contain a manufacturing error, this probability will increase with die size. Process variation can have similar effects on operating frequency, making high-speed designs more sensitive to such variations. With the series on Heterogeneous and Unconventional Cluster Architectures and Applications, we gear to gather most recent insights and ideas from the wide area of cluster computing. This special issue of the "International Journal of Concurrency and Computation: Practice and Experience" resembles our most recent selection. In particular, this selection includes four interesting works.1-4 Two of them were contributions from the last two workshop editions (HUCAA 2015 and 2016, both collocated with the International Conference on Parallel Processing - ICPP'15 and ICPP'16). Furthermore, the two other articles are contributed by the authors based on an open call for contributions of this special issue. All contributions were peer reviewed and received in between three and five reviews. GPU accelerators have been established in the state-of-the-art clusters by offering high performance and energy efficiency. In GPU, such efficient communication among processes with their data residing in the GPU memory is of paramount importance to the application performance. This paper investigates various algorithms in conjunction with the latest GPU features to improve GPU collective operations. For clusters with multi-GPU nodes, the authors of this paper propose a hierarchical framework that allows different algorithms at each hierarchy level. By studying various combination of algorithms, the authors of this article highlight the importance of choosing the right algorithm within each level. They evaluate their framework on MPI Allreduce and show promising performance results, specifically for large message sizes which are highly in-use in deep learning and big data applications. They also show the benefit of using the Hyper-Q feature and the MPS service in jointly using different copy types to perform multiple inter-process communications. However, the authors of this paper show that efficient designs are required to further harness this potential. Accordingly, they propose Hyper-Q aware algorithms for GPU collectives. They evaluate our algorithms on MPI Allgather and MPI Allreduce operations. While their experimental results show the benefit of their algorithms, their profiling results indicate that this benefit is mainly rooted in overlapping different copy types. Virtual Screening methods (VS) simulate molecular interactions in silico to look for the best chemical compound that interacts with a given molecular target. VS are becoming increasingly popular to accelerate the drug discovery process and constitute hard optimization problems with a huge computational cost. To deal with these two challenges, the authors of this paper have created METADOCK, an application that (1) enables a wide range of metaheuristics through a parametrized schema, and (2) promotes the use of a multi-GPU environment within a heterogeneous cluster. Metaheuristics provide approximate solutions in a reasonable time frame, but given the stochastic nature of real-life procedures, the energy budget goes hand in hand with acceleration to validate the proposed solution. This paper evaluates energy trade-offs and correlations with performance for a set of metaheuristics derived from METADOCK. The authors of this paper establish a solid inference from minimal power to maximal performance in GPUs and from there to optimal energy consumption. This way, ideal heuristics can be chosen according not only to best accuracy and performance but also to energy requirements. Their study starts with a preselection of parameterized metaheuristic functions, building blocks where we will find optimal patterns from power criteria while preserving parallelism through a GPU execution. They then establish a methodology to figure out the best instances of the parameterized kernels based on energy patterns obtained, which are analyzed from different viewpoints: performance, average power, and total energy consumed. The authors of this paper also compare the best workload distributions for optimal performance and power efficiency among Pascal and Maxwell GPUs on popular Titan models. The experimental results in this paper demonstrate that the most power efficient GPU can be overloaded in order to reduce the total amount of energy required by as much as 20%, finding unique scenarios where Maxwell does it better in execution time, but with Pascal always ahead in performance per watt, reaching peaks of up to 40%. Faster, lower power, and/or less expensive computation will be a software problem forever. Hardware can only make the challenge simpler or harder, and heterogeneous approaches exacerbate it. For emerging alternative computational technologies like quantum, optical, resistive (and other forms of analog computation), and/or biological computing (among others), to be successful, they must be integrated into the existing computational infrastructure (both hardware and software) if they are to realize their full potential. The increasingly main-stream options that reconfigurable logic represents (both fine and coarse grained) will also be most useful within an infrastructure that is sympathetic to legacy memory and storage models. SAHARA is a reduction of computation into data wavefronts that, independent of the underlying technology, employs memory as the fundamental unit of computation within a simple data-flow model, essentially turning processing into a side-effect of the relevant data being made available to the logic that manipulates that data. No single aspect of SAHARA is “new”. Its foundations are more than 50 years old and started with Minsky's 1961 paper on Turing equivalence. SAHARA is an eminently useful abstraction of computation that has the potential of seamlessly integrating many disparate forms of computation behind a simple, common, architectural interface. The rise of heterogeneous systems has given place to great challenges for users, as they involve new concepts, restrictions and frameworks. Their exploitation is further complicated in the context of distributed memory systems, which require the usage of additional different programming paradigms and tools. In this paper, the authors propose a novel approach to program heterogeneous clusters that is based on high-level abstractions such as tiles and hierarchical decomposition combined with the powerful APIs that data types and embedded languages can provide in languages such as C++. Rather than building their proposal from scratch, they have implemented it as a natural integration of the existing Hierarchically Tiled Arrays (HTA) and Heterogeneous Programming Library (HPL) projects, the first one being focused on distributed computing and the second one on heterogeneous processing. The result, called Heterogeneous Hierarchically Tiled Arrays (H2TA), is very intuitive and easy to use thanks to the global view of the data and the single-threaded view of the execution that it provides at cluster level together with the transparency it provides with respect to the management of the heterogeneous devices. An evaluation comparing the proposal in this paper with previous MPI-based implementations shows its large programmability advantages and the reasonable overhead incurred. Cluster computing is currently facing a pivotal point in time as we are hitting hard constraints about the future of CMOS processors. While, currently, most energy is still spent on computations, first research results show that an increasing fraction of overall energy is spent for data movements. Given the hard constraints on power consumption, one can imagine how influential this fundamental transition will be. Still, CMOS replacements like quantum computing, neuromorphic computing and many other candidates are either still nascent or will only be helpful for certain workloads. While it seems safe to assume that this will further increase heterogeneity in the future, we will have to find out if the currently narrow workload spectrum for these architectures can be extended or if generic computing in the future will have so solely rely on CMOS processors. In this context, we hope the readers of this special issue will find the contributions interesting and inspiring for future research. We in particular acknowledge the thorough work of our review board, which did an excellent and timely work on proving opinions and helpful feedback to the submitted articles. Furthermore, we are also thankful to the authors, which submitted their research contributions to our special edition. Last, but not least, we are also indebted to the continuous support by the editor-in-chief Geoffrey Fox, who is always available for assistance and recommendations regarding the publication procedure.
Holger Fröning, Federico Silla
Concurr. Comput. Pract. Exp.1
2017 Relaxations for High-Performance Message Passing on Massively Parallel SIMT Processors
abstract
Accelerators, such as GPUs, have proven to be highly successful in reducing execution time and power consumption of compute-intensive applications. Even though they are already used pervasively, they are typically supervised by general-purpose CPUs, which results in frequent control flow switches and data transfers as CPUs are handling all communication tasks. However, we observe that accelerators are recently being augmented with peer-to-peer communication capabilities that allow for autonomous traffic sourcing and sinking. While appropriate hardware support is becoming available, it seems that the right communication semantics are yet to be identified. Maintaining the semantics of existing communication models, such as the Message Passing Interface (MPI), seems problematic as they have been designed for the CPU’s execution model, which inherently differs from such specialized processors. In this paper, we analyze the compatibility of traditional message passing with massively parallel Single Instruction Multiple Thread (SIMT) architectures, as represented by GPUs, and focus on the message matching problem. We begin with a fully MPI-compliant set of guarantees, including tag and source wildcards and message ordering. Based on an analysis of exascale proxy applications, we start relaxing these guarantees to adapt message passing to the GPU’s execution model. We present suitable algorithms for message matching on GPUs that can yield matching rates of 60M and 500M matches/s, depending on the constraints that are being relaxed. We discuss our experiments and create an understanding of the mismatch of current message passing protocols and the architecture and execution model of SIMT processors.
Benjamin Klenk, Holger Fröning, Hans Eberle, Larry Dennison
IPDPS2
2016 Heterogeneous cluster architectures and applications
abstract
Welcome to the Special Issue on Heterogeneous and Unconventional Cluster Architectures and Applications! Why such a topic for a special issue? Which is the reason for fostering recent work on cluster architectures that do not fit into the widely established categories? Are unconventional applications especially interesting so that they are worth a special issue? Do these disruptive pieces of work deserve such a focus in a journal? From the point of view of the guest editors of this special issue, the progress of technology can be seen as a two-step process. One of the steps is performing research in order to achieve a better version of the designs we are currently enjoying. In this regard, for instance, when an interconnection network is enhanced so that its bandwidth is doubled, datacenters can deliver a new level of performance to their customers. In a similar way, new generations of processors help to significantly reduce energy consumption of the servers used in such datacenter. Also, installing new versions of accelerators in such servers allows applications to experience important reductions in their execution time at the time that the overall power consumption of the facility is not impacted. All these new developments require an enormous amount of research. For example, faster interconnects require to glue together research on microelectronics, physics, communication protocols, encoding schemes, and thermal issues, to name only a few. The creation of new generations of processors and accelerators also involves large amounts of research from a variety of areas, thus making that a big team of researchers is required in order to bring all those ideas to market. Nevertheless, if those developments are observed from a very broad and long-term perspective, they could be thought to be incremental, despite of their clear significance and despite that they enable new levels of performance. Hence, which should be the non-incremental developments? Our understanding of this fundamental problem is that the non-incremental developments would be those based on ideas that could cause surprise to those researchers initially listening to them. For instance, the very first time that several computers were interconnected in order to create a cluster of computers that cooperatively work together in order to solve a problem, at that point in time an unconventional idea appeared. From a long-term perspective, all the later developments around clusters are incremental; despite those developments are the result of many hours of work and huge amounts of effort from very smart researchers that probably devoted their lives to create those improved versions of the technology behind that initial idea of putting together a few computers into a cluster so that they collaborate in order to solve a problem. One more example of these non-incremental developments could be the use of Graphics Processing Unit (GPU) or any other accelerator in general, in order to reduce the execution time of applications. In this regard, new versions of these accelerators are required so that technology makes progress. Furthermore, these new versions of these accelerators are the result of big investments in time as well as funding resources, which involve hundreds of trained researchers. However, the qualitative difference was performed when someone had the initial idea of attaching an accelerator to the system. That was the point in time that really impacted technology and made it evolve because that allowed a performance increase of several orders of magnitude for the datacenters. Hence, all these non-incremental developments are the other step of the two-step process described previously regarding the progress of technology. Actually, in this two-step process, these non-incremental ideas would be the first step, and the second step would be composed of all those developments that enhance and enrich the initial idea. In fact, the first step cannot live without the second, because both of them are equally important and both of them are required. Given that the two steps are required, and given that most of current conferences and journals include a large percentage of work devoted to the second step, in the opinion of the guest editors of this special issue, it is worth to make an effort to gather innovative work on unconventional cluster architectures and applications, which might have a big impact on future cluster architectures. This includes more or less any cluster architecture that is not based on the usual commodity components and therefore makes use of some special hardware or software components, or that is used for very special and unconventional applications. Within this fostering effort, it is particularly encouraged to work on disruptive approaches, which may show inferior performance today but can already point out their performance potential. Obviously, innovative and disruptive ideas like putting together a bunch of computers into a cluster so that they collaborate to solve a problem only happen from time to time. Actually, they happen seldom. Most innovative ideas are not so disruptive. However, this obvious fact will not discourage the editors of this special issue on their hunt. This special issue contains five pieces of work selected from eleven submissions. Four independent reviews were produced for each of the submissions. From these lines, the guest editors profoundly thank the review board members for their excellent work. We know that the review task is performed on a volunteer basis and many times taking time from other areas such as personal leisure or time with family and friends. This is why the support of our reviewers is so appreciated by us. Actually, without their time and effort, this special issue could not be completed. We are happy about receiving papers from a broad range of topics, allowing this special issue to cover aspects from hardware architecture design, from low-level software and run-times up to new system architecture concepts. The first paper shows how accelerators as traffic sources and sinks impact collective operations 1. The authors point out a set of optimizations and analyze their effectiveness. They also show that a variety of algorithms is required for best performance and that accelerators as sources and sinks significantly increase the overhead for collective operations when compared with general-purpose processors like CPUs. In summary, these improvements lead to a higher utilization of clusters composed of heterogeneous components. In the next paper, the authors propose several optimizations for schedulers of heterogeneous systems 2. First, they propose the possibility of resource range specifications to improve resource specification within batch jobs. Second, they introduce a new integer programming formulation that allows reducing the number of variables, thereby enabling faster solutions. In the third paper, the authors introduce Gaspar, an aspect-oriented framework that facilitates the development of Java applications for heterogeneous systems and clusters 3. They propose to exploit parallelism by aspect modules and a data layout that is dependent on the target platform. Their proposed aspect modules for Java have many similarities with OpenMP formulations for C/C++. The impact of different data structures and layouts for heterogeneous resources is well known, and they address this fact including support in their framework for different variants. Their evaluation shows how such an approach allows for a high-performance portability and general competitive performance. In the fourth paper, the authors show how they use a custom cluster based on reconfigurable logic to design a highly optimized solution for certain tasks, including communication-intensive molecular dynamics 4. Reconfigurable logic opens up a new degree of freedom when designing computing systems; however, this huge amount of freedom comes with significant complexity increase. The authors show how they approached the problem with a combination of modeling and abstractions, demonstrating how in the future reconfigurable logic could change the computing landscape. Last, the fifth paper describes the approach of the European Union-funded DEEP project to design a new cluster architecture for Exascale computing 5. Opposed to other systems, their concept decouples the ratio of general-purpose CPUs and specialized accelerators, giving additional freedom when upgrading such systems. Besides this, they also nicely point out the importance of accelerators directly sourcing and sinking traffic. The guest editors of this special issue hope that you enjoy the read as much as we did enjoy composing it. We believe this collection to be representative of some important problems related to heterogeneous systems and clusters today and that such a reading is highly inspiring for our own research and general interest.
Federico Silla, Holger Fröning
Concurr. Comput. Pract. Exp.2
2016 Analyzing GPU-controlled communication with dynamic parallelism in terms of performance and energy
Lena Oden, Benjamin Klenk, Holger Fröning
Parallel Comput.3
2016 Optimizing the data-collection time of a large-scale data-acquisition system through a simulation framework
Tommaso Colombo 0002, Holger Fröning, Pedro Javier García, Wainer Vandelli
J. Supercomput.2
2015 Modeling a Large Data-Acquisition Network in a Simulation Framework
abstract
The ATLAS detector at CERN records particle collision "events" delivered by the Large Hadron Collider. Its data-acquisition system identifies, selects, and stores interesting events in near real-time, with an aggregate throughput of several 10 GB/s. It is a distributed software system executed on a farm of roughly 2000 commodity worker nodes communicating via TCP/IP on an Ethernet network. Event data fragments are received from the many detector readout channels and are buffered, collected together, analyzed and either stored permanently or discarded. This system, and data-acquisition systems in general, are sensitive to the latency of the data transfer from the readout buffers to the worker nodes. Challenges affecting this transfer include the many-to-one communication pattern and the inherently bursty nature of the traffic. In this paper we introduce the main performance issues brought about by this workload, focusing in particular on the so-called TCP incast pathology. Since performing systematic studies of these issues is often impeded by operational constraints related to the mission-critical nature of these systems, we focus instead on the development of a simulation model of the ATLAS data-acquisition system, used as a case study. The simulation is based on the well-established OMNeT++ framework. Its results are compared with existing measurements of the system's behavior. The successful reproduction of the measurements by the simulations validates the modeling approach. We share some of the preliminary findings obtained from the simulation, as an example of the additional possibilities it enables, and outline the planned future investigations.
Tommaso Colombo 0002, Holger Fröning, Pedro Javier García, Wainer Vandelli
CLUSTER2
2015 Analyzing communication models for distributed thread-collaborative processors in terms of energy and time
abstract
Accelerated computing has become pervasive for increasing the computational power and energy efficiency in terms of GFLOPs/Watt. For application areas with highest demands, for instance high performance computing, data warehousing and high performance analytics, accelerators like GPUs or Intel's MICs are distributed throughout the cluster. Since current analyses and predictions show that data movement will be the main contributor to energy consumption, we are entering an era of communication-centric heterogeneous systems that are operating with hard power constraints. In this work, we analyze data movement optimizations for distributed heterogeneous systems based on CPUs and GPUs. Thread-collaborative processors like GPUs differ significantly in their execution model from generalpurpose processors like CPUs, but available communication models are still designed and optimized for CPUs. Similar to heterogeneity in processing, heterogeneity in communication can have a huge impact on energy and time. To analyze this impact, we use multiple workloads with distinct properties regarding computational intensity and communication characteristics. We show for which workloads tailored communication models are essential, not only reducing execution time but also saving energy. Exposing the impact in terms of energy and time for communication-centric heterogeneous systems is crucial for future optimizations, and this work is a first step in this direction.
Benjamin Klenk, Lena Oden, Holger Fröning
ISPASS3
2015 On the design of a new dynamic credit-based end-to-end flow control mechanism for HPC clusters
Javier Prades, Federico Silla, Holger Fröning, Mondrian Nüssle, José Duato
Parallel Comput.3
2014 Energy-Efficient Collective Reduce and Allreduce Operations on Distributed GPUs
abstract
GPUs gain high popularity in High Performance Computing, due to their massive parallelism and high performance per Watt. Despite their popularity, data transfer between multiple GPUs in a cluster remains a problem. Most communication models require the CPU to control the data flow, also intermediate staging copies to host memory are often inevitable. These two facts lead to higher CPU and memory utilization. As a result, overall performance decreases and power consumption increases. Collective operations like reduce and all reduce are very common in scientific simulations and also very sensitive to performance. Due to their massive parallelism, GPUs are very suitable for such operations, but they only excel in performance if they can process the problem in-core. Global GPU Address Spaces (GGAS) enable a direct GPU-to-GPU communication for heterogeneous clusters, which is completely in-line with the GPU's thread-collective execution model and does not require CPU assistance or staging copies in host memory. As we will see, GGAS helps to process collective operations among distributed GPUs in-core. In this paper, we introduce the implementation and optimization of collective reduce and all reduce operations using GGAS as a communication model. Compared to message passing, we get a speedup of 1.7x for small data sizes. A detailed analysis based on power measurements of CPU, host memory and GPU reveals that GGAS as communication model not only saves cycles, also the power and energy consumption is reduced dramatically. For instance, for an all reduce operation half of the energy can be saved by the reduced the power consumption in combination with the lower run time.
Lena Oden, Benjamin Klenk, Holger Fröning
CCGRID3
2013 On Achieving High Message Rates
abstract
Computer systems continue to increase in parallelism in all areas. Stagnating single thread performance as well as power constraints prevent a reversal of this trend, on the contrary, current projections show that the trend towards parallelism will accelerate. In cluster computing, scalability, and therefore the degree of parallelism, is limited by the network interconnect and more specifically by the message rate it provides. We designed an interconnection network specifically for high message rates. Among other things, it reduces the burden on the software stack by relying on communication engines that perform a large fraction of the send and receive functionality in hardware. It also supports multi-core environments very efficiently through hardware-level virtualization of the communication engines. We provide details on the overall architecture, the thin software stack, performance results for a set of MPI-based benchmarks, and an in-depth analysis of how application performance depends on the message rate. We vary the message rate by software and hardware techniques, and measure the application-level impact of different message rates. We are also using this analysis to extrapolate performance for technologies with wider data paths and higher line rates.
Holger Fröning, Mondrian Nüssle, Heiner Litz, Christian Leber, Ulrich Brüning 0001
CCGRID1
2013 GGAS: Global GPU address spaces for efficient communication in heterogeneous clusters
abstract
Modern GPUs are powerful high-core-count processors, which are no longer used solely for graphics applications, but are also employed to accelerate computationally intensive general-purpose tasks. For utmost performance, GPUs are distributed throughout the cluster to process parallel programs. In fact, many recent high-performance systems in the TOP500 list are heterogeneous architectures. Despite being highly effective processing units, GPUs on different hosts are incapable of communicating without assistance from a CPU. As a result, communication between distributed GPUs suffers from unnecessary overhead, introduced by switching control flow from GPUs to CPUs and vice versa. Most communication libraries even require intermediate copies from GPU memory to host memory. This overhead in particular penalizes small data movements and synchronization operations, reduces efficiency and limits scalability. In this work we introduce global address spaces to facilitate direct communication between distributed GPUs without CPU involvement. Avoiding context switches and unnecessary copying dramatically reduces communication overhead. We evaluate our approach using a variety of workloads including low-level latency and bandwidth benchmarks, basic synchronization primitives like barriers, and a stencil computation as an example application. We see performance benefits of up to 2× for basic benchmarks and up to 1.67× for stencil computations.
Lena Oden, Holger Fröning
CLUSTER2
2013 Oncilla: A GAS runtime for efficient resource allocation and data movement in accelerated clusters
abstract
Accelerated and in-core implementations of Big Data applications typically require large amounts of host and accelerator memory as well as efficient mechanisms for transferring data to and from accelerators in heterogeneous clusters. Scheduling for heterogeneous CPU and GPU clusters has been investigated in depth in the high-performance computing (HPC) and cloud computing arenas, but there has been less emphasis on the management of cluster resource that is required to schedule applications across multiple nodes and devices. Previous approaches to address this resource management problem have focused on either using low-performance software layers or on adapting complex data movement techniques from the HPC arena, which reduces performance and creates barriers for migrating applications to new heterogeneous cluster architectures. This work proposes a new system architecture for cluster resource allocation and data movement built around the concept of managed Global Address Spaces (GAS), or dynamically aggregated memory regions that span multiple nodes.We propose a software layer called Oncilla that uses a simple runtime and API to take advantage of non-coherent hardware support for GAS. The Oncilla runtime is evaluated using two different high-performance networks for microkernels representative of the TPC-H data warehousing benchmark, and this runtime enables a reduction in runtime of up to 81%, on average, when compared with standard disk-based data storage techniques. The use of the Oncilla API is also evaluated for a simple breadth-first search (BFS) benchmark to demonstrate how existing applications can incorporate support for managed GAS.
Jeffrey Young 0001, Se Hoon Shon, Sudhakar Yalamanchili, Alex Merritt, Karsten Schwan, Holger Fröning
CLUSTER6
2012 A New End-to-End Flow-Control Mechanism for High Performance Computing Clusters
abstract
High Performance Computing usually leverages messaging libraries such as MPI or GASNet in order to exchange data among processes in large-scale clusters. Furthermore, these libraries make use of specialized low-level networking layers in order to retrieve as much performance as possible from hardware interconnects such as Infini Band or Myrinet, for example. EXTOLL is another emerging technology targeted for high performance clusters. These specialized low-level networking layers require some kind of flow control in order to prevent buffer overflows at the received side. In this paper we present a new flow control mechanism that is able to adapt the buffering resources used by a process according to the parallel application communication pattern and the varying activity among communicating peers. The tests carried out in a 64-node 1024-core EXTOLL cluster show that our new dynamic flow-control mechanism provides extraordinarily high buffer efficiency along with very low overhead, which is reduced between 8 and 10 times.
Javier Prades, Federico Silla, José Duato, Holger Fröning, Mondrian Nüssle
CLUSTER4
2011 MEMSCALE: in-cluster-memory databases
abstract
We have developed a new memory architecture for clusters that allows automatic access from any processor to any memory module in the cluster completely by hardware. Thus, with a single assembly instruction a processor can retrieve (or update) a memory location in a remote node. The efficiency of this new paradigm makes it possible to speed-up the execution of shared-memory applications with very large memory footprints by running them across the entire cluster, thus providing them a true shared-memory environment (contrary to the emulation typically carried out by software-based distributed shared memory).
Héctor Montaner, Federico Silla, Holger Fröning, José Duato
CIKM3
2011 Highly scalable barriers for future high-performance computing clusters
abstract
Although large scale high performance computing today typically relies on message passing, shared memory can offer significant advantages, as the overhead associated with MPI is completely avoided. In this way, we have developed an FPGA-based Shared Memory Engine that allows to forward memory transactions, like loads and stores, to remote memory locations in large clusters, thus providing a single memory address space. As coherency protocols do not scale with system size we completely avoid a global coherency across the cluster. However, we maintain local coherency domains, thus keeping the cores within one node coherent. In this paper, we show the suitability of our approach by analyzing the performance of barriers, a very common synchronization primitive in parallel programs. Experiments in a real cluster prototype show that our approach allows synchronization among 1024 cores spread over 64 nodes in less than 15us, several times faster than other highly optimized barriers. We show the feasibility of this approach by executing a shared-memory implementation of FFT. Finally, note that this barrier can also be leveraged by MPI applications running on our shared memory architecture for clusters. This ensures the usefulness of this work for applications already written.
Holger Fröning, Alexander Giese, Héctor Montaner, Federico Silla, José Duato
HiPC1
2011 Unleash Your Memory-Constrained Applications: A 32-Node Non-coherent Distributed-Memory Prototype Cluster
abstract
Improvements in hardware for parallel shared-memory computing usually involve increments in the number of computing cores and in the amount of memory available for a given application. However, many shared-memory applications do not require more computing cores than available in current motherboards because their scalability is bounded to a few tens of parallel threads. Nevertheless, they may still benefit from having more memory resources. Additionally, the performance of extended systems involving more cores is typically constrained by the glueing coherency protocol, whose overhead lowers the performance of the final system. In this paper we present a 32-node prototype of a new non-coherent distributed-memory architecture for clusters, aimed to provide applications additional memory borrowed from other nodes without providing them more cores, thus avoiding the penalty of maintaining coherency among nodes of the cluster. Results from the execution of real applications in this prototype demonstrate that our proposal truly works, as well as its performance is assessed.
Héctor Montaner, Federico Silla, Holger Fröning, José Duato
HPCC3
2011 MEMSCALETM: A Scalable Environment for Databases
abstract
In this paper we propose a new memory architecture for clusters referred to as MEMSCALE. This architecture provides a distributed non-coherent shared-memory view of the memory resources present in the cluster. With this aggregation technique, a given processor can directly access any memory address located at other nodes in the cluster and, therefore, the whole memory present in the cluster can be granted to a single application. In this study we focus on in-memory databases as a memory-hungry application in order to show the possibilities of our new architecture. To prove the feasibility of our idea, a 16-node prototype cluster serves as a demonstrator. Part of the memory in each node is used to create a global memory pool of 128GB which hosts an entire database. First we show that providing more memory than usually available in a typical commodity node for a database server makes the execution of queries more than one order of magnitude faster than using regular SSD drives. After that, we go one step further and show that simultaneously accessing the database from all the nodes in the cluster converts our prototype into a powerful database server capable of beating current commercial solutions in terms of latency and throughput.
Héctor Montaner, Federico Silla, Holger Fröning, José Duato
HPCC3
2010 Getting Rid of Coherency Overhead for Memory-Hungry Applications
abstract
Current commercial solutions intended to provide additional resources to an application being executed in a cluster usually aggregate processors and memory from different nodes. In this paper we present a 16-node prototype for a shared-memory cluster architecture that follows a different approach by decoupling the amount of memory available to an application from the processing resources assigned to it. In this way, we provide a new degree of freedom so that the memory granted to a process can be expanded with the memory from other nodes in the cluster without increasing the number of processors used by the program. This feature is especially suitable for memory-hungry applications that demand large amounts of memory but present a parallelization level that prevents them from using more cores than available in a single node. The main advantage of this approach is that an application can use more memory from other nodes without involving the processors, and caches, from those nodes. As a result, using more memory no longer implies increasing the coherence protocol overhead because the number of caches involved in the coherent domain has become independent from the amount of available memory. The prototype we present in this paper leverages this idea by sharing 128GB of memory among the cluster. Real executions show the feasibility of our prototype and its scalability.
Héctor Montaner, Federico Silla, Holger Fröning, José Duato
CLUSTER3
2009 An FPGA based verification platform for HyperTransport 3.x
abstract
In this paper we present a verification platform designed for HyperTransport 3.x (HT3) applications. HyperTransport 3.x is a very low latency and high bandwidth chip-to-chip interconnect which is particularly used in AMDs novel Opteron processor series. As it is an open protocol, a broad application range exists ranging from southbridge chips over closely coupled accelerators to add in cards. Its main advantage over PCI-express is that it allows direct connection to the CPU resulting in significantly improved latency performance. To enable the development of new HyperTransport products we herein present the very first FPGA based prototyping platform for HT3.x. Such a platform is enormously valuable as new designs can be tested in real world systems before producing an costly application specific integrated circuit (ASIC). Due to the high operating frequencies of HT3.x an FPGA based solution is extremely challenging as we will describe in this paper. Our presented architecture is evaluated and implemented in the form of a printed circuit board (PCB). This add-in card represents the world's first available HyperTransport 3 device. Early adopters of HT3 benefit from the results of this work for rapid prototyping and hardware/software coverification of new HT3 designs and products.
Heiner Litz, Holger Fröning, Maximilian Thürmer, Ulrich Brüning 0001
FPL2
2008 VELO: A Novel Communication Engine for Ultra-Low Latency Message Transfers
abstract
This paper presents a novel stateless, virtualized communication engine for sub-microsecond latency. Using a field-programmable-gate-array (FPGA) based prototype we show a latency of 970 ns between two machines with our virtualized engine for low overhead (VELO). The FPGA device is directly connected to the CPUs by a hypertransport link. The described hardware architecture is optimized for small messages and avoids the overhead typically found with direct-memory access (DMA) controlled transfers. The stateless approach allows to use the hardware unit directly from many threads and processes simultaneously. It provides a secure user level communication with an extremely optimized start-up phase. Micro benchmarks results are reported both based on proprietary API and OpenMPI basis.
Heiner Litz, Holger Fröning, Mondrian Nüssle, Ulrich Brüning 0001
ICPP2