EDBT 2026 Demo / reviewers in the wild / expert
Michèle Weiland
dblp:79/8095
· DBLP profile ↗
9ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0003-4713-3073ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Evaluating and optimising compiler code generation for NVIDIA GraceabstractIn this paper, we explore the performance of the main optimising compiler toolchains currently available for high-performance AArch64 processors, namely the Arm Compiler for Linux (ACFL), GNU, LLVM and the NVIDIA HPC (NVHPC) compilers, on the recently released NVIDIA Grace CPU. We evaluate the performance of these compilers using the RAJA Performance Suite (RAJAPerf) to understand where each compiler does best and why. We find that compilers mostly generate well optimised code on baseline sequential runs, with the gap between the fastest and slowest being only 8% on average. However, they exhibit much larger variations on threaded parallel runs—with the gap between fastest and slowest code generated by the different compilers increasing to roughly 33%. Furthermore, we investigate in detail those kernels where LLVM performs worst relative to the remaining compilers and propose optimisations to improve code generation in those cases. We show scenarios where the default compiler behaviour produces sub-optimal code and where adjusting compiler flags, such as those explicitly controlling loop unrolling, can improve performance significantly. In cases where this is insufficient, we propose changes at the compiler level necessary to enable improved code generation and unlock further optimisations. These improvements account for speedups of over 70% in some kernels. Ricardo Jesus, Michèle Weiland |
ICPP | 2 |
| 2023 | AArch64 Atomics: Might They Be Harming Your Performance?abstractAtomic operations are indivisible operations guaranteed to execute as a whole. One of the most important and widely used atomic operations is "compare-and-swap" (CAS), which allows threads to perform concurrent read-modify-write operations on the same memory location, free of data races. On recent Arm architectures, CAS operations can be implemented either directly via CAS instructions, or via load-linked/store-conditional (LL-SC) instruction pairs. Ricardo Jesus, Michèle Weiland |
PPoPP | 2 |
| 2023 | Vectorizing and distributing number-theoretic transform to count Goldbach partitions on Arm-based supercomputersabstractSummary In this article, we explore the usage of scalable vector extension (SVE) to vectorize number‐theoretic transforms (NTTs). In particular, we show that 64‐bit modular arithmetic operations, including modular multiplication, can be efficiently implemented with SVE instructions. The vectorization of NTT loops and kernels involving 64‐bit modular operations was not possible in previous Arm‐based single instruction multiple data architectures since these architectures lacked crucial instructions to efficiently implement modular multiplication. We test and evaluate our SVE implementation on the A64FX processor in an HPE Apollo 80 system. Furthermore, we implement a distributed NTT for the computation of large‐scale exact integer convolutions. We evaluate this transform on HPE Apollo 70, Cray XC50, HPE Apollo 80, and HPE Cray EX systems, where we demonstrate good scalability to thousands of cores. Finally, we describe how these methods can be utilized to count the number of Goldbach partitions of all even numbers to large limits. We present some preliminary results concerning this problem, in particular a histogram of the number of Goldbach partitions of the even numbers up to 240. Ricardo Jesus, Tomás Oliveira e Silva, Michèle Weiland |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Performance Evaluation of Adaptive Routing on Dragonfly-based Production SystemsabstractPerformance of applications in production environments can be sensitive to network congestion. Cray Aries supports adaptively routing each network packet independently based on the load or congestion encountered as a packet traverses the network. Software can dictate different routing policies, adjusting between minimal and non-minimal bias, for each posted message. We have extensively evaluated the sensitivity of the routing bias selection on application performance as well as whole system performance in both production and controlled conditions. We show that the default routing bias used in Aries-based systems is often sub-optimal and that using a higher bias towards minimal routes will not only reduce the congestion effects on the application but also will decrease the overall congestion on the network. This routing scheme results in not only improved mean performance (by up to 12%) of most production applications but also reduced run-to-run variability. Our study prompted the two supercomputing facilities (ALCF and NERSC) to change the default routing mode on their Aries-based systems. We present the substantial improvement measured in the overall congestion management and interconnect performance in production after making this change. Sudheer Chunduri, Kevin Harms, Taylor L. Groves, Peter Mendygral, Justs Zarins, Michèle Weiland, Yasaman Ghadar |
IPDPS | 6 |
| 2021 | Usage Scenarios for Byte-Addressable Persistent Memory in High-Performance and Data Intensive Computing
Michèle Weiland, Bernhard Homölle |
J. Comput. Sci. Technol. | 1 |
| 2020 | Investigating Applications on the A64FXabstractThe A64FX processor from Fujitsu, being designed for computational simulation and machine learning applications, has the potential for unprecedented performance in HPC systems. In this paper, we evaluate the A64FX by benchmarking against a range of production HPC platforms that cover a number of processor technologies. We investigate the performance of complex scientific applications across multiple nodes, as well as single node and mini-kernel benchmarks. This paper finds that the performance of the A64FX processor across our chosen benchmarks often significantly exceeds other platforms, even without specific application optimisations for the processor instruction set or hardware. However, this is not true for all the benchmarks we have undertaken. Furthermore, the specific configuration of applications can have an impact on the runtime and performance experienced. Adrian Jackson, Michèle Weiland, Nick Brown 0002, Andrew Turner, Mark Parsons 0001 |
CLUSTER | 2 |
| 2019 | An early evaluation of Intel's optane DC persistent memory module and its impact on high-performance scientific applicationsabstractMemory and I/O performance bottlenecks in supercomputing simulations are two key challenges that must be addressed on the road to Exascale. The new byte-addressable persistent non-volatile memory technology from Intel, DCPMM, promises to be an exciting opportunity to break with the status quo, with unprecedented levels of capacity at near-DRAM speeds. Here, we explore the potential of DCPMM in the context of two high-performance scientific applications in terms of outright performance, efficiency and usability for both its Memory and App Direct modes. In Memory mode, we show equivalent performance and better efficiency for a CASTEP simulation that is limited by memory capacity on conventional DRAM-only systems without any changes to the application. For IFS, we demonstrate that a distributed object-store over NVRAM reduces the data contention created in weather forecasting data producer-consumer workflows. In addition, we also present the achievable memory bandwidth performance using STREAM. Michèle Weiland, Holger Brunst, Tiago Quintino, Nick Johnson, Olivier Iffrig, Simon D. Smart, Christian Herold, Antonino Bonanni, Adrian Jackson, Mark Parsons 0001 |
SC | 1 |
| 2019 | Leveraging MPI RMA to optimize halo-swapping communications in MONC on Cray machinesabstractSummary Remote Memory Access (RMA), also known as single‐sided communications, provides a way for reading and writing directly into the memory of other processes without having to issue explicit message passing style communication calls. Previous studies have concluded that MPI RMA can provide increased communication performance over traditional MPI Point to Point (P2P), but these are based on synthetic benchmarks rather than real‐world codes. In this work, we replace the existing non‐blocking P2P communication calls in the Met Office NERC Cloud model, a mature code for modeling the atmosphere, with MPI RMA. We describe our approach in detail and discuss the options taken for correctness and performance. Experiments are performed on ARCHER, a Cray XC30, and Cirrus, an SGI ICE machine. We demonstrate on ARCHER that, by using RMA, we can obtain between a 5% and 10% reduction in communication time at each timestep on up to 32768 cores, which over the entirety of a run (with many timesteps) results in a significant improvement in performance compared to P2P on the Cray. However, RMA is not a silver bullet, and there are challenges when integrating RMA calls into existing codes: important optimizations are necessary to achieve good performance and library support is not universally mature, as is the case on Cirrus. In this paper, we discuss, in the context of a real‐world code, the lessons learned converting P2P to RMA, explore performance and scaling challenges, and contrast alternative RMA synchronization approaches in detail. Nick Brown 0002, Michael R. Bareford, Michèle Weiland |
Concurr. Comput. Pract. Exp. | 3 |
| 2018 | In situ data analytics for highly scalable cloud modelling on Cray machinesabstractSummary MONC is a highly scalable modelling tool for the investigation of atmospheric flows, turbulence, and cloud microphysics. Typical simulations produce very large amounts of raw data, which must then be analysed for scientific investigation. For performance and scalability reasons, this analysis and subsequent writing to disk should be performed in situ on the data as it is generated; however, one does not wish to pause the computation whilst analysis is carried out. In this paper, we present the analytics approach of MONC, where cores of a node are shared between computation and data analytics. By asynchronously sending their data to an analytics core, the computational cores can run continuously without having to pause for data writing or analysis. We describe our IO server framework and analytics workflow, which is highly asynchronous, along with solutions to challenges that this approach raises and the performance implications of some common configuration choices. The result of this work is a highly scalable analytics approach, and we illustrate on up to 32 768 computational cores of a Cray XC30 that there is minimal performance impact on the runtime when enabling data analytics in MONC and also investigate the performance and suitability of our approach on the KNL. Nick Brown 0002, Michèle Weiland, Adrian Hill, Ben Shipway |
Concurr. Comput. Pract. Exp. | 2 |