EDBT 2026 Demo / reviewers in the wild / expert
Ettore Tiotto
dblp:22/827
· DBLP profile ↗
7ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
GPUs and heterogeneous computing · 52% Parallel and multicore computing · 42% High-performance computing · 5% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
loop transformation |
0.4 | 1 | 2019 | Memory-access-aware Safety and Profitability Analysis for Transformation of Accelerator-bound OpenMP Loops · ACM Trans. Archit. Code Optim. 2019 |
Compilers and program optimization › memory optimization
memory access optimization |
0.4 | 1 | 2019 | Memory-access-aware Safety and Profitability Analysis for Transformation of Accelerator-bound OpenMP Loops · ACM Trans. Archit. Code Optim. 2019 |
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
GPU offloading |
0.4 | 1 | 2019 | Memory-access-aware Safety and Profitability Analysis for Transformation of Accelerator-bound OpenMP Loops · ACM Trans. Archit. Code Optim. 2019 |
GPUs and heterogeneous computing › GPU memory access
memory coalescing |
0.4 | 1 | 2019 | Memory-access-aware Safety and Profitability Analysis for Transformation of Accelerator-bound OpenMP Loops · ACM Trans. Archit. Code Optim. 2019 |
Compilers and program optimization
code generation |
0.2 | 1 | 2016 | Combining Static and Dynamic Data Coalescing in Unified Parallel C · IEEE Trans. Parallel Distributed Syst. 2016 |
Parallel and multicore computing
parallel programming models |
0.2 | 1 | 2016 | Combining Static and Dynamic Data Coalescing in Unified Parallel C · IEEE Trans. Parallel Distributed Syst. 2016 |
Parallel and multicore computing › parallel programming models › distributed memory programming models
partitioned global address space |
0.2 | 1 | 2016 | Combining Static and Dynamic Data Coalescing in Unified Parallel C · IEEE Trans. Parallel Distributed Syst. 2016 |
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP |
0.1 | 1 | 2019 | Memory-access-aware Safety and Profitability Analysis for Transformation of Accelerator-bound OpenMP Loops · ACM Trans. Archit. Code Optim. 2019 |
High-performance computing
supercomputing |
0.1 | 1 | 2016 | Combining Static and Dynamic Data Coalescing in Unified Parallel C · IEEE Trans. Parallel Distributed Syst. 2016 |
Methods — techniques the papers use, named apart from their topics
static analysis · 0.8iteration point difference analysis · 0.8inspector-executor model · 0.5compiler transformation · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Experiences Building an MLIR-Based SYCL CompilerabstractSimilar to other programming models, compilers for SYCL, the open programming model for heterogeneous computing based on C++, would benefit from access to higher-level intermediate representations. The loss of high-level structure and semantics caused by premature lowering to low-level intermediate representations and the inability to reason about host and device code simultaneously present major challenges for SYCL compilers. The MLIR compiler framework, through its dialect mechanism, allows to model domain-specific, high-level intermediate representations and provides the necessary facilities to address these challenges. This work therefore describes practical experience with the design and implementation of an MLIR-based SYCL compiler. By modeling key elements of the SYCL programming model in host and device code in the MLIR dialect framework, the presented approach enables the implementation of powerful device code optimizations as well as analyses across host and device code. Compared to two LLVM-based SYCL implementations, this yields speedups of up to 4.3x on a collection of SYCL benchmark applications. Finally, this work also discusses challenges encountered in the design and implementation and how these could be addressed in the future. Ettore Tiotto, Victor Perez 0001, Whitney Tsang, Lukas Sommer, Julian Oppermann, Victor Lomüller, Mehdi Goli 0001, James Brodman |
CGO | 1 |
| 2019 | Memory-access-aware Safety and Profitability Analysis for Transformation of Accelerator-bound OpenMP LoopsabstractIteration Point Difference Analysis is a new static analysis framework that can be used to determine the memory coalescing characteristics of parallel loops that target GPU offloading and to ascertain safety and profitability of loop transformations with the goal of improving their memory access characteristics. This analysis can propagate definitions through control flow, works for non-affine expressions, and is capable of analyzing expressions that reference conditionally defined values. This analysis framework enables safe and profitable loop transformations. Experimental results demonstrate potential for dramatic performance improvements. GPU kernel execution time across the Polybench suite is improved by up to 25.5× on an Nvidia P100 with benchmark overall improvement of up to 3.2×. An opportunity detected in a SPEC ACCEL benchmark yields kernel speedup of 86.5× with a benchmark improvement of 3.3×. This work also demonstrates how architecture-aware compilers improve code portability and reduce programmer effort. Artem Chikin, Taylor Lloyd, José Nelson Amaral, Ettore Tiotto |
ACM Trans. Archit. Code Optim. | 4 |
| 2016 | Using shared-data localization to reduce the cost of inspector-execution in unified-parallel-C programs
Michail Alvanos, Ettore Tiotto, José Nelson Amaral, Montse Farreras, Xavier Martorell |
Parallel Comput. | 2 |
| 2016 | Combining Static and Dynamic Data Coalescing in Unified Parallel CabstractSignificant progress has been made in the development of programming languages and tools that are suitable for hybrid computer architectures that group several shared-memory multicores interconnected through a network. This paper addresses important limitations in the code generation for partitioned global address space (PGAS) languages. These languages allow fine-grained communication and lead to programs that perform many fine-grained accesses to data. When the data is distributed to remote computing nodes, code transformations are required to prevent performance degradation. Until now code transformations to PGAS programs have been restricted to the cases where both the physical mapping of the data or the number of processing nodes are known at compilation time. In this paper, a novel application of the inspector-executor model overcomes these limitations and allows profitable code transformations, which result in fewer and larger messages sent through the network, when neither the data mapping nor the number of processing nodes are known at compilation time. A performance evaluation reports both scaling and absolute performance numbers on up to 32,768 cores of a Power 775 supercomputer. This evaluation indicates that the compiler transformation results in speedups between 1.15x and 21x over a baseline and that these automated transformations achieve up to 63 percent the performance of the MPI versions. Michail Alvanos, Montse Farreras, Ettore Tiotto, José Nelson Amaral, Xavier Martorell |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2014 | Reducing Compiler-Inserted Instrumentation in Unified-Parallel-C Code GenerationabstractPrograms written in Partitioned Global Address Space (PGAS) languages can access any location of the entire address space via standard read/write operations. However, the compiler have to create the communication mechanisms and the runtime system to use synchronization primitives to ensure the correct execution of the programs. However, PGAS programs may have fine-grained shared accesses that lead to performance degradation. One solution is to use the inspector-executor technique to determine which accesses are indeed remote and which accesses may be coalesced in larger remote access operations. A straightforward implementation of the inspector-executor in a PGAS system may result in excessive instrumentation that hinders performance. This paper introduces a shared-data localization transformation based on linear memory descriptors (LMADs) that reduces the amount of instrumentation introduced by the compiler into programs written in the UPC language and describes a prototype implementation of the proposed transformation. A performance evaluation, using up to 2048 cores of a POWER 775 supercomputer, allows for a prediction that applications with regular accesses can achieve up to 180% of the performance of handoptimized versions while applications with irregular accesses yield performance gain from 1.12X up to 6.3X speedup. Michail Alvanos, José Nelson Amaral, Ettore Tiotto, Montse Farreras, Xavier Martorell |
SBAC-PAD | 3 |
| 2013 | Improving communication in PGAS environments: static and dynamic coalescing in UPCabstractThe goal of Partitioned Global Address Space (PGAS) languages is to improve programmer productivity in large scale parallel machines. However, PGAS programs may have many fine-grained shared accesses that lead to performance degradation. Manual code transformations or compiler optimizations are required to improve the performance of programs with fine-grained accesses. The downside of manual code transformations is the increased program complexity that hinders programmer productivity. On the other hand, most compiler optimizations of fine-grain accesses require knowledge of physical data mapping and the use of parallel loop constructs. Michail Alvanos, Montse Farreras, Ettore Tiotto, José Nelson Amaral, Xavier Martorell |
ICS | 3 |
| 2013 | Improving performance of all-to-all communication through loop scheduling in PGAS environmentsabstractNo abstract available. Michail Alvanos, Ilie Gabriel Tanase, Montse Farreras, Ettore Tiotto, José Nelson Amaral, Xavier Martorell |
ICS | 4 |