Leeor Peled

dblp:163/0025 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0003-4853-061XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Tempranillo: Non-Speculative Early Register Release
abstract
Limited by the breakdown of technology scaling, CPU architects are looking for creative solutions to deliver performance improvements while unable to traditionally scale up microarchitectural structures. One promising approach is engineering a microarchitecture that uses resources more efficiently, for example, by recycling them faster to relieve pressure on critical structures and making them look bigger than they are. The physical register file (PRF) is a key structure that faces severe area and power constraints that limit its (and the whole CPU's) scalability. Based on this observation, previous work proposed solutions to reduce the pressure on the PRF by reducing the time each register remains allocated. In this paper, we corroborate earlier findings that the early release of registers is a promising approach to reduce the occupancy of the PRF and we identify novel tight conditions for safely releasing registers. Based on this analysis, we design Tempranillo: an aggressive, non-speculative microarchitecture to release registers as early as possible without requiring additional recovery mechanisms. Tempranillo delivers up to 3.3 % and 11.8 % performance improvement over conventional release on singlethreaded and a 2 -way SMT CPUs, respectively. Additionally, Tempranillo requires modest storage overheads, translating into a performance improvement per KiB of storage of up to 2.6 % and 9.3 % for single-thread and 2-way SMT, respectively. Our evaluation shows that Tempranillo improves over both the state-of-the-art non-speculative and speculative proposals.
Carlos Escuin, Paolo Salvatore Galfano, Davide B. Bartolini, Leeor Peled, Mehdi Alipour
HPCA4
2026 Bumper: Hinting Instruction Usefulness for Robust Unified Caches
Georgios Vavouliotis, Tom Rollet, Davide B. Bartolini, Boris Grot, Leeor Peled, Lixia Yang
ISCA5
2025 To Cross, or Not to Cross Pages for Prefetching?
abstract
Despite processor vendors reporting that cache prefetchers operating with virtual addresses are permitted to cross page boundaries, academia is focused on optimizing cache prefetching for patterns within page boundaries. This work reveals that page-cross prefetching at the first-level data cache (L1D) is seldom beneficial across different execution phases and workloads while showing that state-of-the-art L1D prefetchers are not very accurate at prefetching across page boundaries. In response, we propose $M O K A$, a holistic framework for designing Page-Cross Filters, i.e., microarchitectural schemes that ensure effective and accurate prefetching across page boundaries. MOKA combines (i) hashed perceptron predictors that use prefetcher-independent program features, (ii) predictors that adapt decisions based on the system state (e.g., TLB pressure), and (iii) a scheme to dynamically optimize predictions across different execution phases and workload types. We use the MOKA framework to prototype a Page-Cross Filter, named DRIPPER, for three relevant L1D prefetchers (Berti [60], IPCP [61], BOP [57]). We show that DRIPPER accurately enables pagecross prefetching only when it is beneficial for performance. For instance, Berti [60] (state-of-the-art prefetcher) combined with DRIPPER improves single-core geomean performance over Berti that always permits page-cross prefetches and Berti that always discards page-cross prefetches by $\mathbf{1 . 7 \%}(\mathbf{1 . 2 \%})$ and $\mathbf{2 . 5 \%}(\mathbf{2 . 1 \%})$ across 218 seen (178 unseen) workloads, respectively. Across 300 8 -core mixes, the corresponding geomean speedups are $2.0 \%$ and $3.3 \%$. Finally, we show that DRIPPER provides consistent benefits when both 4KB pages and 2MB large pages are used.
Georgios Vavouliotis, Martí Torrents, Boris Grot, Kleovoulos Kalaitzidis, Leeor Peled, Marc Casas
HPCA5
2020 A Neural Network Prefetcher for Arbitrary Memory Access Patterns
abstract
Memory prefetchers are designed to identify and prefetch specific access patterns, including spatiotemporal locality (e.g., strides, streams), recurring patterns (e.g., varying strides, temporal correlation), and specific irregular patterns (e.g., pointer chasing, index dereferencing). However, existing prefetchers can only target premeditated patterns and relations they were designed to handle and are unable to capture access patterns in which they do not specialize. In this article, we propose a context-based neural network (NN) prefetcher that dynamically adapts to arbitrary memory access patterns. Leveraging recent advances in machine learning, the proposed NN prefetcher correlates program and machine contextual information with memory accesses patterns, using online-training to identify and dynamically adapt to unique access patterns exhibited by the code. By targeting semantic locality in this manner, the prefetcher can discern the useful context attributes and learn to predict previously undetected access patterns, even within noisy memory access streams. We further present an architectural implementation of our NN prefetcher, explore its power, energy, and area limitations, and propose several optimizations. We evaluate the neural network prefetcher over SPEC2006, Graph500, and several microbenchmarks and show that the prefetcher can deliver an average speedup of 21.3% for SPEC2006 (up to 2.3×) and up to 4.4× on kernels over a baseline of PC-based stride prefetcher and 30% for SPEC2006 over a baseline with no prefetching.
Leeor Peled, Uri C. Weiser, Yoav Etsion
ACM Trans. Archit. Code Optim.1
2015 Semantic locality and context-based prefetching using reinforcement learning
abstract
Most modern memory prefetchers rely on spatio-temporal locality to predict the memory addresses likely to be accessed by a program in the near future. Emerging workloads, however, make increasing use of irregular data structures, and thus exhibit a lower degree of spatial locality. This makes them less amenable to spatio-temporal prefetchers.
Leeor Peled, Shie Mannor, Uri C. Weiser, Yoav Etsion
ISCA1