EDBT 2026 Demo / reviewers in the wild / expert
Guillaume Iooss
dblp:17/8138
· DBLP profile ↗
7ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-0326-1807ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Analytical Modeling of Set-Associative Caches for Optimizing Tensor OperationsabstractOptimizing for data cache memories is a difficult problem for a compiler, and an important performance bottleneck. Indeed, the behavior of a cache is complex to model statically, due to the cache policies (associativity, eviction policy), and the complex interplay with the mechanisms of a superscalar microarchitecture. Many analytical cache models exist that predict the number of cache misses at compile time, which is pertinent information for optimization. However, there is a compromise between the precision of the model, coverage of input programs, and the model’s analysis time. For example, using a fully-associative analytical cache model is one such compromise: precision is sacrificed by assuming full associativity, but the model is fast and applicable to any affine program. This article introduces a new cache model, called SARCASM (Set-Associative Rotating Cache Analytical/Simulating Model), representing a new and useful compromise, which is pertinent in the context of sampling over optimization choices (called configurations ). This was previously limited to fully-associative cache models. SARCASM is a set-associative cache model and can thus model conflict misses and achieve better precision than fully-associative cache models. It is targeted at code structures that arise with optimized implementations of tensor operations such as matrix multiplication, convolution, tensor contraction, arising in machine learning applications. These may feature several levels of tiling, but with only hyper-rectangular tile shapes. Importantly, it is fast enough to be applied at compile time. The SARCASM cache model is based on the notion of detailed footprint , a natural generalization of the notion of footprint of existing models that considers the footprint for each cache set. Once the detailed footprint is computed for each loop level, it uses a fully associative model for each cache set to obtain the number of cache misses per cache set. We show that the predictions of this model induce an ordering of configurations much closer to that of measured cache misses than a fully-associative model. We also show that it correlates much better with execution time compared to a fully-associative model. Guillaume Iooss, Christophe Guillon, Fabrice Rastello, Albert Cohen 0001, P. Sadayappan |
ACM Trans. Archit. Code Optim. | 1 |
| 2024 | Tightening I/O Lower Bounds through the Hourglass Dependency Pattern
Lionel Eyraud-Dubois, Guillaume Iooss, Julien Langou, Fabrice Rastello |
SPAA | 2 |
| 2023 | Autotuning Convolutions Is Easier Than You ThinkabstractA wide range of scientific and machine learning applications depend on highly optimized implementations of tensor computations. Exploiting the full capacity of a given processor architecture remains a challenging task, due to the complexity of the microarchitectural features that come into play when seeking near-peak performance. Among the state-of-the-art techniques for loop transformations for performance optimization, AutoScheduler [Zheng et al. 2020a ] tends to outperform other systems. It often yields higher performance as compared to vendor libraries, but takes a large number of runs to converge, while also involving a complex training environment. In this article, we define a structured configuration space that enables much faster convergence to high-performance code versions, using only random sampling of candidates. We focus on two-dimensional convolutions on CPUs. Compared to state-of-the-art libraries, our structured search space enables higher performance for typical tensor shapes encountered in convolution stages in deep learning pipelines. Compared to auto-tuning code generators like AutoScheduler, it prunes the search space while increasing the density of efficient implementations. We analyze the impact on convergence speed and performance distribution, on two Intel x86 processors and one ARM AArch64 processor. We match or outperform the performance of the state-of-the-art oneDNN library and TVM’s AutoScheduler, while reducing the autotuning effort by at least an order of magnitude. Nicolas Tollenaere, Guillaume Iooss, Stéphane Pouget, Hugo Brunie, Christophe Guillon, Albert Cohen 0001, P. Sadayappan, Fabrice Rastello |
ACM Trans. Archit. Code Optim. | 2 |
| 2022 | PALMED: Throughput Characterization for Superscalar ArchitecturesabstractIn a super-scalar architecture, the scheduler dynamically assigns micro-operations $( \mu$ OPs) to execution ports. The port mapping of an architecture describes how an instruction decomposes into $\mu$ OPs and lists for each $\mu$ OP the set of ports it can be mapped to. It is used by compilers and performance debugging tools to characterize the performance throughput of a sequence of instructions repeatedly executed as the core component of a loop.This paper introduces a dual equivalent representation: The resource mapping of an architecture is an abstract model where, to be executed, an instruction must use a set of abstract resources, themselves representing combinations of execution ports. For a given architecture, finding a port mapping is an important but difficult problem. Building a resource mapping is a more tractable problem and provides a simpler and equivalent model. This paper describes Palmed, a tool that automatically builds a resource mapping for pipelined, super-scalar, out-of-order CPU architectures. Palmed does not require hardware performance counters, and relies solely on runtime measurements.We evaluate the pertinence of our dual representation for throughput modeling by extracting a representative set of basic-blocks from the compiled binaries of the SPEC CPU 2017 benchmarks. We compared the throughput predicted by existing machine models to that produced by Palmed, and found comparable accuracy to state-of-the art tools, achieving sub-10% mean square error rate on this workload on Intel's Skylake microarchitecture. Nicolas Derumigny, Théophile Bastian, Fabian Gruber, Guillaume Iooss, Christophe Guillon, Louis-Noël Pouchet, Fabrice Rastello |
CGO | 4 |
| 2021 | IOOpt: automatic derivation of I/O complexity bounds for affine programsabstractEvaluating the complexity of an algorithm is an important step when developing applications, as it impacts both its time and energy performance. Computational complexity, which is the number of dynamic operations regardless of the execution order, is easy to characterize for affine programs. Data movement (or, I/O) complexity is more complex to evaluate as it refers, when considering all possible valid schedules, to the minimum required number of I/O between a slow (e.g. main memory) and a fast (e.g. local scratchpad) storage location. Auguste Olivry, Guillaume Iooss, Nicolas Tollenaere, Atanas Rountev, P. Sadayappan, Fabrice Rastello |
PLDI | 2 |
| 2019 | Correct-by-Construction Parallelization of Hard Real-Time Avionics Applications on Off-the-Shelf Predictable HardwareabstractWe present the first end-to-end modeling and compilation flow to parallelize hard real-time control applications while fully guaranteeing the respect of real-time requirements on off-the-shelf hardware. It scales to thousands of dataflow nodes and has been validated on two production avionics applications. Unlike classical optimizing compilation, it takes as input non-functional requirements (real time, resource limits). To enforce these requirements, the compiler follows a static resource allocation strategy, from coarse-grain tasks communicating over an interconnection network all the way to individual variables and memory accesses. It controls timing interferences resulting from mapping decisions in a precise, safe, and scalable way. Keryan Didier, Dumitru Potop-Butucaru, Guillaume Iooss, Albert Cohen 0001, Jean Souyris, Philippe Baufreton, Amaury Graillat |
ACM Trans. Archit. Code Optim. | 3 |
| 2014 | On Program Equivalence with Reductions
Guillaume Iooss, Christophe Alias, Sanjay V. Rajopadhye |
SAS | 1 |