EDBT 2026 Demo / reviewers in the wild / expert
Afif Boudaoud
dblp:417/4859
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0003-8662-6353ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 33% Parallel and multicore computing · 33% High-performance computing · 33% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
autotuning |
0.9 | 1 | 2025 | PerfDojo: Automated ML Library Generation for Heterogeneous Architectures · SC 2025 |
Parallel and multicore computing
kernel optimization |
0.9 | 1 | 2025 | PerfDojo: Automated ML Library Generation for Heterogeneous Architectures · SC 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | PerfDojo: Automated ML Library Generation for Heterogeneous Architectures · SC 2025 |
High-performance computing › performance engineering
performance portability |
0.9 | 1 | 2025 | PerfDojo: Automated ML Library Generation for Heterogeneous Architectures · SC 2025 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.7large language model · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LOOPer: A Learned Automatic Code Optimizer For Polyhedral CompilersabstractWhile polyhedral compilers have shown success in implementing advanced code transformations, they still face challenges in selecting the ones that lead to the most profitable speedups. This has motivated the use of machine learning based cost models to guide the search for polyhedral optimizations. State-of-the-art polyhedral compilers have demonstrated a viable proof-of-concept of such an approach. While promising, this approach still faces significant limitations. Existing polyhedral compilers using deep learning cost models typically support only a small subset of affine transformations, limiting their ability to explore complex code transformations. Furthermore, their applicability does not scale beyond simple programs, thus excluding many program classes from their scope, such as those with non-rectangular iteration domains or multiple loop nests. These limitations significantly impact the generality of such compilers and autoschedulers, raising questions about the overall approach. In this paper, we introduce LOOPER, the first polyhedral autoscheduler that uses a deep learning based cost model and covers a large space of affine transformations and programs. LOOPER allows the optimization of an extensive set of programs while being effective at applying complex sequences of polyhedral transformations. We implement and evaluate LOOPER and show that it achieves competitive speedups over the state-of-the-art. On the PolyBench benchmarks, LOOPER achieves a geometric mean speedup of $\mathbf{1 . 8 4} \mathbf{x}$ over the Tiramisu autoscheduler and $\mathbf{1 . 4 2} \mathbf{x}$ over Pluto, two state-of-the-art polyhedral autoschedulers. Massinissa Merouani, Afif Boudaoud, Iheb Nassim Aouadj, Nassim Tchoulak, Islem Kara Bernou, Hamza Benyamina, Fatima Benbouzid-Si Tayeb, Karima Benatchba, Hugh Leather, Riyadh Baghdadi |
PACT | 2 |
| 2025 | DaCe AD: Unifying High-Performance Automatic Differentiation for Machine Learning and Scientific ComputingabstractAutomatic differentiation (AD) is a set of techniques that systematically applies the chain rule to compute the gradients of functions without requiring human intervention. Although the fundamentals of this technology were established decades ago, it is experiencing a renaissance as it plays a key role in efficiently computing gradients for backpropagation in machine learning algorithms. AD is also crucial for many applications in scientific computing domains, particularly emerging techniques that integrate machine learning models within scientific simulations and schemes. Existing AD frameworks have four main limitations: limited support of programming languages, requiring code modifications for AD compatibility, limited performance on scientific computing codes, and a naive store-all solution for forward-pass data required for gradient calculations. These limitations force domain scientists to manually compute the gradients for large problems. This work presents DaCe AD, a general, efficient automatic differentiation engine that requires no code modifications. DaCe AD uses a novel ILPbased algorithm to optimize the trade-off between storing and recomputing to achieve maximum performance within a given memory constraint. We showcase the generality of our method by applying it to NPBench, a suite of HPC benchmarks with diverse scientific computing patterns, where we outperform JAX, a Python framework with state-of-the-art general AD capabilities, by more than 92 times on average without requiring any code changes. Afif Boudaoud, Alexandru Calotoiu, Marcin Copik, Torsten Hoefler |
CLUSTER | 1 |
| 2025 | PerfDojo: Automated ML Library Generation for Heterogeneous ArchitecturesabstractThe increasing complexity of machine learning models and the proliferation of diverse hardware architectures (CPUs, GPUs, accelerators) make achieving optimal performance a significant challenge. Heterogeneity in instruction sets, specialized kernel requirements for different data types and model features (e.g., sparsity, quantization), and architecture-specific optimizations complicate performance tuning. Manual optimization is resource-intensive, while existing automatic approaches often rely on complex hardware-specific heuristics and uninterpretable intermediate representations, hindering performance portability. We introduce PerfLLM, a novel automatic optimization methodology leveraging Large Language Models (LLMs) and Reinforcement Learning (RL). Central to this is PerfDojo, an environment framing optimization as an RL game using a human-readable, mathematically-inspired code representation that guarantees semantic validity through transformations. This allows effective optimization without prior hardware knowledge, facilitating both human analysis and RL agent training. We demonstrate PerfLLM’s ability to achieve significant performance gains across diverse CPU (x86, Arm, RISC-V) and GPU architectures. Andrei Ivanov, Gioele Gottardo, Marcin Chrapek, Afif Boudaoud, Timo Schneider, Luca Benini, Torsten Hoefler |
SC | 5 |