EDBT 2026 Demo / reviewers in the wild / expert
Thibaud Balem
dblp:304/1096
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Processor architecture and microarchitecture · 77% Energy-efficient computing · 23% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization › code generation
SIMD code generation |
0.6 | 1 | 2022 | Compiler-Assisted Compaction/Restoration of SIMD Instructions · IEEE Trans. Parallel Distributed Syst. 2022 |
Compilers and program optimization
vectorization |
0.6 | 1 | 2022 | Compiler-Assisted Compaction/Restoration of SIMD Instructions · IEEE Trans. Parallel Distributed Syst. 2022 |
Processor architecture and microarchitecture
SIMD |
0.6 | 1 | 2022 | Compiler-Assisted Compaction/Restoration of SIMD Instructions · IEEE Trans. Parallel Distributed Syst. 2022 |
Methods — techniques the papers use, named apart from their topics
simulation · 1.1compiler-assisted compaction/restoration · 1.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Compiler-Assisted Compaction/Restoration of SIMD InstructionsabstractVector processors (e.g., SIMD or GPUs) are ubiquitous in high performance systems. All the supercomputers in the world exploit data-level parallelism (DLP), for example by using single instructions to operate over several data elements. Improving vector processing is therefore key for exascale computing. However, despite its potential, vector code generation and execution have significant challenges. Among these challenges, control flow divergence is one of the main performance limiting factors. Most modern vector instruction sets, including SIMD, rely on predication to support divergence control. Nevertheless, the performance and energy consumption in predicated codes is usually insensitive to the number of active elements in a predicated mask. Since the trend is that vector register size increases, the energy efficiency of exascale computing systems will become sub-optimal. This article proposes a novel approach to improve execution efficiency in predicated vector codes, the Compiler-Assisted Compaction/Restoration (CACR) technique. Baseline CR delays predicated SIMD instructions with inactive elements, compacting active elements from instances of the same instruction of consecutive loop iterations. Compacted elements form an equivalentdensevector instruction. After executing the dense instructions, their results are restored to the original instructions. However, CR has a significant performance and energy penalty when it fails to find active elements, either due to lack of resources when unrolling or because of inter-loop dependencies. In CACR, the compiler analyzes the code looking for key information required to configure CR. Then, it passes this information to the processor via new instructions inserted in the code. This prevents CR from waiting for active elements on scenarios when it would fail to form dense instructions. Simulated results (gem5) show that CACR improves performance by up to 29 percent and reduces dynamic energy by up to 24.2 percent on average, for a a set of applications with predicated execution. The baseline CR only achieves 18.6 percent performance and 14 percent energy improvements for the same configuration and applications. Juan M. Cebrian, Thibaud Balem, Adrián Barredo, Marc Casas, Miquel Moretó, Alberto Ros 0001, Alexandra Jimborean |
IEEE Trans. Parallel Distributed Syst. | 2 |