EDBT 2026 Demo / reviewers in the wild / expert
Per Ekemark
dblp:175/1717
· DBLP profile ↗
4ranked-venue papers
1as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Concurrent programming · 50% Compilers and program optimization · 50% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache coherence |
0.6 | 2 | 2021 | TSOPER: Efficient Coherence-Based Strict Persistency · HPCA 2021 Automatic Detection of Large Extended Data-Race-Free Regions with Conflict Isolation · IEEE Trans. Parallel Distributed Syst. 2018 |
Memory systems › cache coherence
cache coherence protocol |
0.5 | 1 | 2021 | TSOPER: Efficient Coherence-Based Strict Persistency · HPCA 2021 |
Memory systems
non-volatile memory |
0.5 | 1 | 2021 | TSOPER: Efficient Coherence-Based Strict Persistency · HPCA 2021 |
Memory systems › non-volatile memory › persistent memory
persistency model |
0.5 | 1 | 2021 | TSOPER: Efficient Coherence-Based Strict Persistency · HPCA 2021 |
Compilers and program optimization
compiler analysis |
0.3 | 1 | 2018 | Automatic Detection of Large Extended Data-Race-Free Regions with Conflict Isolation · IEEE Trans. Parallel Distributed Syst. 2018 |
Memory systems › cache coherence
coherence protocol optimization |
0.1 | 1 | 2018 | Automatic Detection of Large Extended Data-Race-Free Regions with Conflict Isolation · IEEE Trans. Parallel Distributed Syst. 2018 |
Methods — techniques the papers use, named apart from their topics
static analysis · 0.7conflict isolation · 0.7cache coherence protocol · 0.7atomic group persist · 0.5TSO persist buffer · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | TSOPER: Efficient Coherence-Based Strict PersistencyabstractWe propose a novel approach for hardware-based strict TSO persistency, called TSOPER. We allow a TSO persistency model to freely coalesce values in the caches, by forming atomic groups of cachelines to be persisted. A group persist is initiated for an atomic group if any of its newly written values are exposed to the outside world. A key difference with prior work is that our architecture is based on the concept of a TSO persist buffer, that sits in parallel to the shared LLC, and persists atomic groups directly from private caches to NVM, bypassing the coherence serialization of the LLC. To impose dependencies among atomic groups that are persisted from the private caches to the TSO persist buffer, we introduce a sharing-list coherence protocol that naturally captures the order of coherence operations in its sharing lists, and thus can reconstruct the dependencies among different atomic groups entirely at the private cache level without involving the shared LLC. The combination of the sharing-list coherence and the TSO persist buffer allows persist operations and writes to non-volatile memory to happen in the background and trail the coherence operations. Coherence runs ahead at full speed; persistency follows belatedly. Our evaluation shows that TSOPER provides the same level of reordering as a program-driven relaxed model, hence, approximately the same level of performance, albeit without needing the programmer or compiler to be concerned about false sharing, data-race-free semantics, etc., and guaranteeing all software that can run on top of TSO, automatically persists in TSO. Per Ekemark, Yuan Yao 0009, Alberto Ros 0001, Konstantinos Sagonas, Stefanos Kaxiras |
HPCA | 1 |
| 2018 | Automatic Detection of Large Extended Data-Race-Free Regions with Conflict IsolationabstractData-race-free (DRF) parallel programming becomes a standard as newly adopted memory models of mainstream programming languages such as C++ or Java impose data-race-freedom as a requirement. We propose compiler techniques that automatically delineate extended data-race-free (xDRF) regions, namely regions of code that provide the same guarantees as the synchronization-free regions (in the context of DRF codes). xDRF regions stretch across synchronization boundaries, function calls and loop back-edges and preserve the data-race-free semantics, thus increasing the optimization opportunities exposed to the compiler and to the underlying architecture. We further enlarge xDRF regions with a conflict isolation (CI) technique, delineating what we call xDRF-CI regions while preserving the same properties as xDRF regions. Our compiler (1) precisely analyzes the threads' memory accessing behavior and data sharing in shared-memory, general-purpose parallel applications, (2) isolates data-sharing and (3) marks the limits of xDRF-CI code regions. The contribution of this work consists in a simple but effective method to alleviate the drawbacks of the compiler's conservative nature in order to be competitive with (and even surpass) an expert in delineating xDRF regions manually. We evaluate the potential of our technique by employing xDRF and xDRF-CI region classification in a state-of-the-art, dual-mode cache coherence protocol. We show that xDRF regions reduce the coherence bookkeeping and enable optimizations for performance (6.4 percent) and energy efficiency (12.2 percent) compared to a standard directory-based coherence protocol. Enhancing the xDRF analysis with the conflict isolation technique improves performance by 7.1 percent and energy efficiency by 15.9 percent. Alexandra Jimborean, Per Ekemark, Jonatan Waern, Stefanos Kaxiras, Alberto Ros 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Automatic detection of extended data-race-free regions
Alexandra Jimborean, Jonatan Waern, Per Ekemark, Stefanos Kaxiras, Alberto Ros 0001 |
CGO | 3 |
| 2016 | Multiversioned decoupled access-execute: the key to energy-efficient compilation of general-purpose programsabstractComputer architecture design faces an era of great challenges in an attempt to simultaneously improve performance and energy efficiency. Previous hardware techniques for energy management become severely limited, and thus, compilers play an essential role in matching the software to the more restricted hardware capabilities. One promising approach is software decoupled access-execute (DAE), in which the compiler transforms the code into coarse-grain phases that are well-matched to the Dynamic Voltage and Frequency Scaling (DVFS) capabilities of the hardware. While this method is proved efficient for statically analyzable codes, general-purpose applications pose significant challenges due to pointer aliasing, complex control flow and unknown runtime events. We propose a universal compile-time method to decouple general-purpose applications, using simple but efficient heuristics. Our solutions overcome the challenges of complex code and show that automatic decoupled execution significantly reduces the energy expenditure of irregular or memory-bound applications and even yields slight performance boosts. Overall, our technique achieves over 20% on average energy-delay-product (EDP) improvements (energy over 15% and performance over 5%) across 14 benchmarks from SPEC CPU 2006 and Parboil benchmark suites, with peak EDP improvements surpassing 70%. Konstantinos Koukos, Per Ekemark, Georgios Zacharopoulos 0001, Vasileios Spiliopoulos 0001, Stefanos Kaxiras, Alexandra Jimborean |
CC | 2 |