Karl Fürlinger

dblp:08/1846 · also Karl Fuerlinger · DBLP profile ↗
← Back
28ranked-venue papers
7as first author
6since 2021 · last 2026
0000-0003-0398-4087ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 7 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Cache partitioning for sparse matrix-vector multiplication on the A64FX
abstract
One of the novel features of the Fujitsu A64FX CPU is the sector cache . This feature enables hardware-supported partitioning of the L1 and L2 caches and allows the programmer control of which partition is used to place data in. This paper performs an in-depth study of applying the sector cache to sparse matrix-vector multiplication (SpMV) in the Compressed Sparse Row (CSR) format using a collection of 490 sparse matrices. A performance model based on reuse analysis is used to better understand situations in which and how the sector cache leads to improved cache reuse and to predict cache behavior. The model predicts the number of L2 cache misses within an error of 2% without cache partitioning. With sector cache enabled, depending on the configuration, the model predicts the number of L2 cache missed within 2–3% and 4–18% for sequential and parallel SpMV with 48 threads, respectively. Further experiments show the effect of various sector cache configurations on performance. A median speedup of about 1.05 × is achieved, whereas the maximum speedup is about 1.6 × .
Sergej Breiter, James D. Trotter, Karl Fürlinger
Parallel Comput.3
2025 Slotqueue: A Wait-Free Distributed Multi-producer Single-Consumer Queue with Constant Remote Operations
Do Nguyen An Huy, Thanh-Dang Diep, Karl Fürlinger, Nam Thoai
NPC (2)3
2023 From reactive to proactive load balancing for task-based parallel applications in distributed memory machines
abstract
Summary Load balancing is often a challenge in task‐parallel applications. The balancing problems are divided into static and dynamic. “Static” means that we have some prior knowledge about load information and perform balancing before execution, while “dynamic” must rely on partial information of the execution status to balance the load at runtime. Conventionally, work stealing is a practical approach used in almost all shared memory systems. In distributed memory systems, the communication overhead can make stealing tasks too late. To improve, people have proposed a reactive approach to relax communication in balancing load. The approach leaves one dedicated thread per process to monitor the queue status and offload tasks reactively from a slow to a fast process. However, reactive decisions might be mistaken in high imbalance cases. First, this article proposes a performance model to analyze reactive balancing behaviors and understand the bound leading to incorrect decisions. Second, we introduce a proactive approach to improve further balancing tasks at runtime. The approach exploits task‐based programming models with a dedicated thread as well, namely . Nevertheless, the main idea is to force not only to monitor load; it will characterize tasks and train load prediction models by online learning. “Proactive” indicates offloading tasks before each execution phase proactively with an appropriate number of tasks at once to a potential victim (denoted by an underloaded/fast process). The experimental results confirm speedup improvements from to in important use cases compared to the previous solutions. Furthermore, this approach can support co‐scheduling tasks across multiple applications.
Minh Thanh Chung, Josef Weidendorfer, Karl Fürlinger, Dieter Kranzlmüller
Concurr. Comput. Pract. Exp.3
2023 A general approach for supporting nonblocking data structures on distributed-memory systems
Thanh-Dang Diep, Phuong Hoai Ha, Karl Fürlinger
J. Parallel Distributed Comput.3
2022 A Profiling-Based Approach to Cache Partitioning of Program Data
Sergej Breiter, Josef Weidendorfer, Minh Thanh Chung, Karl Fürlinger
PDCAT4
2021 Nonblocking Data Structures for Distributed-Memory Machines: Stacks as an Example
abstract
Nonblocking data structures are an essential part of many parallel applications in that they can help to improve fault tolerance and performance. Although there are scores of nonblocking data structures, such as stacks, queues, double-ended queues (deques), lists, widely used in practice, most of them are designed to be used on shared-memory machines only, and cannot be used in a distributed-memory setting. Several recent studies focus on the development of novel tailor-made distributed nonblocking data structures and omit the potential for adapting a great wealth of existing shared-memory nonblocking ones for the distributed-memory case. Hence, we propose a general approach to bridge the gap between most existing nonblocking data structures and distributed-memory machines in this work. Several challenges, such as safe memory reclamation and solving the ABA problem, must be overcome. In this paper, we present a global memory management scheme to address these issues. The scheme takes advantage of hazard pointers which are widely used to tackle the problems in shared-memory environments. To demonstrate our general approach, we take stacks as a typical example of nonblocking data structures. This work also provides a survey of well-known nonblocking stack algorithms along with our analysis and evaluation in distributed-memory environments. Moreover, this paper depicts how to improve performance of a stack algorithm by making use of node information.
Thanh-Dang Diep, Karl Fürlinger
PDP2
2019 Engineering a Distributed Histogram Sort
abstract
Sorting is one of the most critical non-numerical algorithms and covers use cases in a wide spectrum of scientific applications. Although we can build upon excellent research over the last decades, scaling to thousands of processing units on modern many-core architectures reveals a gap between theory and practice. We adopt ideas of the well-known quickselect and sample sort algorithms to minimize data movement. Our evaluation demonstrates that we can keep up with recently proposed distribution sort algorithms in large-scale experiments, without any assumptions on the input keys. Additionally, our implementation outperforms an efficient multi-threaded merge sort on a single node. Our implementation is based on a C++ PGAS approach with an STL-like interface and can easily be integrated into many application codes. As part of the presented experiments, we further reveal challenges with multi-threaded MPI and one-sided communication.
Roger Kowalewski, Pascal Jungblut, Karl Fürlinger
CLUSTER3
2019 A time-stamping system to detect memory consistency errors in MPI one-sided applications
Thanh-Dang Diep, Kien Trung Pham, Karl Fürlinger, Nam Thoai
Parallel Comput.3
2018 Utilizing Heterogeneous Memory Hierarchies in the PGAS Model
abstract
Emerging technologies such as non-volatile or 3D-stacked memory significantly impact the design of future high performance computing systems. To keep up with the increasing core count, relying only on DRAM is inefficient due to its static power consumption. Modern HPC architectures feature a heterogeneous memory hierarchy with different capacities and capabilities. This paper addresses two challenges. First, we need programming models to abstract the complexity of the underlying heterogeneous memory hierarchy, while still giving explicit control to domain experts. We propose the concept of memory spaces to model a heterogeneous memory hierarchy and integrate it into a PGAS-like programming model. Second, we need to understand the impact of different memory capabilities on the performance of scientific applications. An experimental evaluation with a series of benchmarks, conducted on a Intel KNL platform, reveals that proper data placement on specific types of memory achieves significant speedup.
Roger Kowalewski, Tobias Fuchs, Karl Fürlinger, Tobias Guggemos
PDP3
2018 A Portable Multidimensional Coarray for C++
abstract
Fortran Coarrays are a well known data structure in High Performance Computing (HPC) applications. There have been various attempts to port the concept to other programming languages that have a wider user base outside of scientific computing. While a popular implementation of the partitioned global address space (PGAS) model is Unified Parallel C (UPC), there is currently no portable implementation of Coarrays for C++. In this paper a portable version is presented, which is closely based on the Coarray C++ implementation of the Cray Compiling Environment. In this work we focus on a common subset of all proposed features by Cray. Our implementation utilizes the distributed data structures provided by the DASH library, demonstrating their universal applicability. Finally, a performance evaluation shows that our proposed Coarray abstraction adds negligible overhead and even outperforms native Coarray Fortran.
Felix MoBbauer, Roger Kowalewski, Tobias Fuchs, Karl Fürlinger
PDP4
2018 MC-CChecker: A Clock-Based Approach to Detect Memory Consistency Errors in MPI One-Sided Applications
abstract
MPI one-sided communication decouples data movement from synchronization, which eliminates overhead from unneeded synchronization and allows for greater concurrency. On the one hand this fact is the great advantage of MPI one-sided communication, but on the other, it poses enormous challenges for programmers in preserving the reliability of programs. Memory consistency errors are notorious for degrading reliability as well as performance of MPI one-sided applications. Even an MPI expert can easily make these mistakes. The lockopts bug occurred in an RMA test case that is part of MPICH MPI implementation is an example for this situation. Hence, detecting memory consistency errors is extremely challenging. MC-Checker is the most cutting-edge debugger to address these errors effectively. MC-Checker tackles the memory consistency errors based on the happened-before relation. Taking full advantage of the relation makes DN-Analyzer of MC-Checker difficult to scale well. For that reason, MC-Checker does ignore the transitive ordering of the happened-before relation to retain scalability of DN-Analyzer. Consequently, MC-Checker is highly able to impose a potential source of false positives.
Thanh-Dang Diep, Karl Fürlinger, Nam Thoai
EuroMPI2
2017 Trends in Data Locality Abstractions for HPC Systems
abstract
The cost of data movement has always been an important concern in high performance computing (HPC) systems. It has now become the dominant factor in terms of both energy consumption and performance. Support for expression of data locality has been explored in the past, but those efforts have had only modest success in being adopted in HPC applications for various reasons. them However, with the increasing complexity of the memory hierarchy and higher parallelism in emerging HPC systems, locality management has acquired a new urgency. Developers can no longer limit themselves to low-level solutions and ignore the potential for productivity and performance portability obtained by using locality abstractions. Fortunately, the trend emerging in recent literature on the topic alleviates many of the concerns that got in the way of their adoption by application developers. Data locality abstractions are available in the forms of libraries, data structures, languages and runtime systems; a common theme is increasing productivity without sacrificing performance. This paper examines these trends and identifies commonalities that can combine various locality concepts to develop a comprehensive approach to expressing and managing data locality on future large-scale high-performance computing systems.
Didem Unat, Anshu Dubey, Torsten Hoefler, John Shalf, Mark James Abraham, Mauro Bianco, Bradford L. Chamberlain, Romain Cledat, H. Carter Edwards, Hal Finkel, Karl Fürlinger, Frank Hannig, Emmanuel Jeannot, Amir Kamil, Jeff Keasler, Paul H. J. Kelly, Vitus J. Leung, Hatem Ltaief, Naoya Maruyama, Chris J. Newburn, Miquel Pericàs
IEEE Trans. Parallel Distributed Syst.11
2016 Nasty-MPI: Debugging Synchronization Errors in MPI-3 One-Sided Applications
Roger Kowalewski, Karl Fürlinger
Euro-Par2
2015 Automatic On-Line Detection of MPI Application Structure with Event Flow Graphs
Xavier Aguilar, Karl Fürlinger, Erwin Laure
Euro-Par2
2015 DART-CUDA: A PGAS Runtime System for Multi-GPU Systems
abstract
The Partitioned Global Address Space (PGAS) approach is a promising programming model in high performance parallel computing that combines the advantages of distributed memory systems and shared memory systems. The PGAS model has been used on a variety of hardware platforms in the form of PGAS programming languages like Unified Parallel C (UPC), Chapel and Fortress. However, in spite of the increasing adoption in distributed and shared memory systems, the extension of the PGAS model to accelerator platforms is still not well supported. To exploit the immense computational power of multi-GPU systems, this work is concerned with the design and implementation of a Partitioned Global Address Space model for multi-GPU systems. Several issues related to the combination of logically separate GPU memories on multiple graphic cards are addressed. Furthermore, the execution model of modern GPU architectures is studied and a task creation mechanism with load balancing is proposed. Our work is implemented in the context of the DASH project, a C++ template library that realizes PGAS semantics through operator overloading. Experimental results suggest promising performance of the design and its implementation.
Karl Fürlinger
ISPDC2
2014 MPI Trace Compression Using Event Flow Graphs
Xavier Aguilar, Karl Fürlinger, Erwin Laure
Euro-Par2
2013 Topic 1: Support Tools and Environments - (Introduction)
Bronis R. de Supinski, Bettina Krammer, Karl Fürlinger, Jesús Labarta, Dimitrios S. Nikolopoulos
Euro-Par3
2012 A Performance Study of Virtual Machines on Multicore Architectures
abstract
Cloud computing has promoted the widespread use of virtualized machines. A question arises: How does virtualization influence the performance of running applications? The answer must be a common interest of application developers and users. This paper describes the results of our performance evaluation on a virtualized multicore machine. We tested a set of benchmark applications and detected some general features that should be considered when running applications on a virtualized multicore machine. We also studied the application execution behavior using profiling tools. We found the reason for unexpectedly poor performance of an OpenMP application in a virtualized setting and optimized the program. The optimization resulted in a significant performance gain.
Jie Tao 0001, Karl Fürlinger, Lizhe Wang 0001, Holger Marten
PDP2
2011 Parallel Aspects of OpenFOAM with Large Eddy Simulations
abstract
Open FOAM is a mainstream open-source frame-work for the simulation in several areas of CFD and engineering whose syntax is a high level representation of the mathematical notation of physical models, internal details like parallelization, tensor algebra, and mesh manipulation are hidden and automatically integrated. We used the back-facing step geometry with Large Eddy Simulations and semi-implicit methods to investigate the scalability and important MPI characteristics of Open FOAM. Moreover, this geometry provides a configuration with representative features found in current engineering and HPC problems. The algebraic multigrid solver, for example, was found to be a very powerful and fast iterative solver for the solution of PDEs compared to the BiGC during the weak and strong scaling tests. It was also determined strong relations between the sizes and shapes of sud domains with the footprint of the inter-domains communications subroutines using a graph-based partitioner. Thus, setting the bases for optimal domain decompositions. However, it was also found that the master-slave strategy introduces an unexpected bottleneck in the communication of scalar values when more than a hundred of MPI tasks were deployed. An extensive analysis revealed that this anomaly was present in few MPI tasks resulting in an severe performance reduction. Finally, we highlight the importance of tracing and profiling tools, IPM is a novel implementation that could instrument Open FOAM successfully.
Orlando Rivera, Karl Fürlinger
HPCC2
2011 Investigating the Scalability of OpenFOAM for the Solution of Transport Equations and Large Eddy Simulations
Orlando Rivera, Karl Fürlinger, Dieter Kranzlmüller
ICA3PP (2)2
2010 Effective Performance Measurement at Petascale Using IPM
abstract
As supercomputers are being built from an ever increasing number of processing elements, the effort required to achieve a substantial fraction of the system peak performance is continuously growing. Tools are needed that give developers and computing center staff holistic indicators about the resource consumption of applications and potential performance pitfalls at scale. To use the full potential of a supercomputer today, applications must incorporate multilevel parallelism (threading and message passing) and carefully orchestrate file I/O. As a consequence, performance tools must also be able to monitor these system components in an integrated way and at the full machine scales. We present IPM, a modularized monitoring approach for MPI, Open MP, file I/O, and other event sources. We describe its implementation design principles, which are targeted for efficiency and minimal application perturbation, and present an application study of using IPM at scale.
Karl Fürlinger, Nicholas J. Wright, David Skinner
ICPADS1
2010 Recording the control flow of parallel applications to determine iterative and phase-based behavior
Karl Fürlinger, Shirley Moore
Future Gener. Comput. Syst.1
2008 OpenMP-centric performance analysis of hybrid applications
abstract
Several performance analysis tools support hybrid applications. Most originated as MPI profiling or tracing tools and OpenMP capabilities were added to extend the performance analysis capabilities for the hybrid parallelization case. In this paper we describe our experience with the other path to support both programming paradigms. Our starting point is a profiling tool for OpenMP called ompP that was extended to handle MPI related data. The measured data and the method of presentation follow our focus on the OpenMP side of the performance optimization cycle. For example, the existing overhead classification scheme of ompP was extended to cover time in MPI calls as a new type of overhead.
Karl Fürlinger, Shirley Moore
CLUSTER1
2007 On Using Incremental Profiling for the Performance Analysis of Shared Memory Parallel Applications
Karl Fürlinger, Michael Gerndt, Jack J. Dongarra
Euro-Par1
2007 Specification and detection of performance problems with ASL
abstract
Abstract Performance analysis is an important step in tuning performance‐critical applications. It is a cyclic process of measuring and analyzing performance data, driven by the programmer's hypotheses on potential performance problems. Currently this process is controlled manually by the programmer. The goal of the work described in this article is to automate the performance analysis process based on a formal specification of performance properties. One result of the APART project is the APART Specification Language (ASL) for the formal specification of performance properties. Performance bottlenecks can then be identified based on the specification, since bottlenecks are viewed as performance properties with a large negative impact. We also present the overall design and an initial evaluation of the Periscope system which utilizes ASL specifications to automatically search for performance bottlenecks in a distributed manner. Copyright © 2006 John Wiley & Sons, Ltd.
Michael Gerndt, Karl Fürlinger
Concurr. Comput. Pract. Exp.2
2005 Performance Analysis of Shared-Memory Parallel Applications Using Performance Properties
Karl Fürlinger, Michael Gerndt
HPCC1
2004 Task-Queue Based Hybrid Parallelism: A Case Study
Karl Fürlinger, Olaf Schenk, Michael Hagemann
Euro-Par1
2003 Distributed Application Monitoring for Clustered SMP Architectures
Karl Fürlinger, Michael Gerndt
Euro-Par1