EDBT 2026 Demo / reviewers in the wild / expert
Ricardo Jesus
dblp:201/3860
· DBLP profile ↗
5ranked-venue papers
5as first author
3since 2021 · last 2024
0000-0002-9651-4756ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Security and privacy · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Evaluating and optimising compiler code generation for NVIDIA GraceabstractIn this paper, we explore the performance of the main optimising compiler toolchains currently available for high-performance AArch64 processors, namely the Arm Compiler for Linux (ACFL), GNU, LLVM and the NVIDIA HPC (NVHPC) compilers, on the recently released NVIDIA Grace CPU. We evaluate the performance of these compilers using the RAJA Performance Suite (RAJAPerf) to understand where each compiler does best and why. We find that compilers mostly generate well optimised code on baseline sequential runs, with the gap between the fastest and slowest being only 8% on average. However, they exhibit much larger variations on threaded parallel runs—with the gap between fastest and slowest code generated by the different compilers increasing to roughly 33%. Furthermore, we investigate in detail those kernels where LLVM performs worst relative to the remaining compilers and propose optimisations to improve code generation in those cases. We show scenarios where the default compiler behaviour produces sub-optimal code and where adjusting compiler flags, such as those explicitly controlling loop unrolling, can improve performance significantly. In cases where this is insufficient, we propose changes at the compiler level necessary to enable improved code generation and unlock further optimisations. These improvements account for speedups of over 70% in some kernels. Ricardo Jesus, Michèle Weiland |
ICPP | 1 |
| 2023 | AArch64 Atomics: Might They Be Harming Your Performance?abstractAtomic operations are indivisible operations guaranteed to execute as a whole. One of the most important and widely used atomic operations is "compare-and-swap" (CAS), which allows threads to perform concurrent read-modify-write operations on the same memory location, free of data races. On recent Arm architectures, CAS operations can be implemented either directly via CAS instructions, or via load-linked/store-conditional (LL-SC) instruction pairs. Ricardo Jesus, Michèle Weiland |
PPoPP | 1 |
| 2023 | Vectorizing and distributing number-theoretic transform to count Goldbach partitions on Arm-based supercomputersabstractSummary In this article, we explore the usage of scalable vector extension (SVE) to vectorize number‐theoretic transforms (NTTs). In particular, we show that 64‐bit modular arithmetic operations, including modular multiplication, can be efficiently implemented with SVE instructions. The vectorization of NTT loops and kernels involving 64‐bit modular operations was not possible in previous Arm‐based single instruction multiple data architectures since these architectures lacked crucial instructions to efficiently implement modular multiplication. We test and evaluate our SVE implementation on the A64FX processor in an HPE Apollo 80 system. Furthermore, we implement a distributed NTT for the computation of large‐scale exact integer convolutions. We evaluate this transform on HPE Apollo 70, Cray XC50, HPE Apollo 80, and HPE Cray EX systems, where we demonstrate good scalability to thousands of cores. Finally, we describe how these methods can be utilized to count the number of Goldbach partitions of all even numbers to large limits. We present some preliminary results concerning this problem, in particular a histogram of the number of Goldbach partitions of the even numbers up to 240. Ricardo Jesus, Tomás Oliveira e Silva, Michèle Weiland |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | Stream Generation: Markov Chains vs GANs
Ricardo Jesus, Mário Antunes 0001, Petia Georgieva, Diogo Gomes 0001, Rui L. Aguiar |
IoTBDS | 1 |
| 2017 | Extracting Knowledge from Stream Behavioural PatternsabstractThe increasing number of small, cheap devices full of sensing capabilities lead to an untapped source of information that can be explored to improve and optimize several systems. Yet, as this number grows it becomes increasingly difficult to manage and organize all this new information. The lack of a standard context representation scheme is one of the main difficulties in this research area (Antunes et al., 2016b). With this in mind we propose a stream characterization model which aims to provide the foundations of a new stream similarity metric. Complementing previous work on context organization, we aim to provide an automatic organizational model without enforcing specific representations. Ricardo Jesus, Mário Antunes 0001, Diogo Gomes 0001, Rui L. Aguiar |
IoTBDS | 1 |