EDBT 2026 Demo / reviewers in the wild / expert
Gioele Gottardo
dblp:408/1532
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0009-9103-1403ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 33% Parallel and multicore computing · 33% High-performance computing · 33% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
autotuning |
0.9 | 1 | 2025 | PerfDojo: Automated ML Library Generation for Heterogeneous Architectures · SC 2025 |
Parallel and multicore computing
kernel optimization |
0.9 | 1 | 2025 | PerfDojo: Automated ML Library Generation for Heterogeneous Architectures · SC 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | PerfDojo: Automated ML Library Generation for Heterogeneous Architectures · SC 2025 |
High-performance computing › performance engineering
performance portability |
0.9 | 1 | 2025 | PerfDojo: Automated ML Library Generation for Heterogeneous Architectures · SC 2025 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.7large language model · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PerfDojo: Automated ML Library Generation for Heterogeneous ArchitecturesabstractThe increasing complexity of machine learning models and the proliferation of diverse hardware architectures (CPUs, GPUs, accelerators) make achieving optimal performance a significant challenge. Heterogeneity in instruction sets, specialized kernel requirements for different data types and model features (e.g., sparsity, quantization), and architecture-specific optimizations complicate performance tuning. Manual optimization is resource-intensive, while existing automatic approaches often rely on complex hardware-specific heuristics and uninterpretable intermediate representations, hindering performance portability. We introduce PerfLLM, a novel automatic optimization methodology leveraging Large Language Models (LLMs) and Reinforcement Learning (RL). Central to this is PerfDojo, an environment framing optimization as an RL game using a human-readable, mathematically-inspired code representation that guarantees semantic validity through transformations. This allows effective optimization without prior hardware knowledge, facilitating both human analysis and RL agent training. We demonstrate PerfLLM’s ability to achieve significant performance gains across diverse CPU (x86, Arm, RISC-V) and GPU architectures. Andrei Ivanov, Gioele Gottardo, Marcin Chrapek, Afif Boudaoud, Timo Schneider, Luca Benini, Torsten Hoefler |
SC | 3 |