Martí Torrents

dblp:145/3670 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0001-7339-8011ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Enhancing Instruction Prefetching via Cache and TLB Management
Alexandre Valentin Jamet, Georgios Vavouliotis, Martí Torrents, Dimitrios Chasapis, Marc Casas
ISCA3
2025 To Cross, or Not to Cross Pages for Prefetching?
abstract
Despite processor vendors reporting that cache prefetchers operating with virtual addresses are permitted to cross page boundaries, academia is focused on optimizing cache prefetching for patterns within page boundaries. This work reveals that page-cross prefetching at the first-level data cache (L1D) is seldom beneficial across different execution phases and workloads while showing that state-of-the-art L1D prefetchers are not very accurate at prefetching across page boundaries. In response, we propose $M O K A$, a holistic framework for designing Page-Cross Filters, i.e., microarchitectural schemes that ensure effective and accurate prefetching across page boundaries. MOKA combines (i) hashed perceptron predictors that use prefetcher-independent program features, (ii) predictors that adapt decisions based on the system state (e.g., TLB pressure), and (iii) a scheme to dynamically optimize predictions across different execution phases and workload types. We use the MOKA framework to prototype a Page-Cross Filter, named DRIPPER, for three relevant L1D prefetchers (Berti [60], IPCP [61], BOP [57]). We show that DRIPPER accurately enables pagecross prefetching only when it is beneficial for performance. For instance, Berti [60] (state-of-the-art prefetcher) combined with DRIPPER improves single-core geomean performance over Berti that always permits page-cross prefetches and Berti that always discards page-cross prefetches by $\mathbf{1 . 7 \%}(\mathbf{1 . 2 \%})$ and $\mathbf{2 . 5 \%}(\mathbf{2 . 1 \%})$ across 218 seen (178 unseen) workloads, respectively. Across 300 8 -core mixes, the corresponding geomean speedups are $2.0 \%$ and $3.3 \%$. Finally, we show that DRIPPER provides consistent benefits when both 4KB pages and 2MB large pages are used.
Georgios Vavouliotis, Martí Torrents, Boris Grot, Kleovoulos Kalaitzidis, Leeor Peled, Marc Casas
HPCA2
2025 Multi-Core Aware Evaluation of Prefetchers
abstract
Full-system simulation is key for chip design, allowing designers to evaluate pre-silicon processor performance. Given the cost of simulation, a common approach is to simulate a limited number of cores (scale-down simulation) and extrapolate results for the entire chip, assuming similar resource usage across cores. However, this can overlook critical resource competition, particularly for shared resources like the last level cache (LLC) and main memory. In this paper we quantify the error introduced by this assumption in a case study with state-of-the-art prefetchers on an industry-like 56 -core architecture. We evaluate these prefetchers running on the multi-core architecture under three interference scenarios, and show how and when scale-down simulation results can be misleading.
Martí Torrents, Paul Caheny, Stijn Eyerman, Wim Heirman
ISPASS1
2024 Exploiting Vector Code Semantics for Efficient Data Cache Prefetching
abstract
Emerging workloads from domains like high performance computing, data analytics or deep learning consume large amounts of memory bandwidth. To mitigate this problem, computing systems include large and deep memory cache hierarchies that exploit both spatial and temporal locality. In this context, hardware data cache prefetching constitutes a useful method to anticipate cache misses and boost performance. Despite their success in terms of high coverage rates, current data cache prefetchers incur a significant number of late and sometimes useless prefetches. Additionally, these state-of-the-art prefetchers are not aware of architecture trends towards larger vector units and vector-length agnostic instruction sets.
Francesc Martínez, Martí Torrents, Adrià Armejach, Marc Casas
ICS2
2019 Evaluation of a Rack-Scale Disaggregated Memory Prototype for Cloud Data Centers
abstract
Disaggregated data centers propose a modular architecture where memory and compute resources are utilized in a finer granularity, aiming for a better optimized capacity and power consumption. In contrast to a classic compute node in a data center, memory is not required to be co-located with the processor, disaggregation can enhance hardware elasticity, improve virtual machine (VM) migration, and reduce the total cost of ownership (TCO), when compared to current data center solutions.
Josue V. Quiroga, Martí Torrents, Nehir Sönmez, Dimitris Theodoropoulos 0001, Ferad Zyulkyarov, Mario Nemirovsky
RSP2
2018 dReDBox: Materializing a full-stack rack-scale system prototype of a next-generation disaggregated datacenter
abstract
Current datacenters are based on server machines, whose mainboard and hardware components form the baseline, monolithic building block that the rest of the system software, middleware and application stack are built upon. This leads to the following limitations: (a) resource proportionality of a multi-tray system is bounded by the basic building block (mainboard), (b) resource allocation to processes or virtual machines (VMs) is bounded by the available resources within the boundary of the mainboard, leading to spare resource fragmentation and inefficiencies, and (c) upgrades must be applied to each and every server even when only a specific component needs to be upgraded. The dRedBox project (Disaggregated Recursive Datacentre-in-a-Box) addresses the above limitations, and proposes the next generation, low-power, across form-factor datacenters, departing from the paradigm of the mainboard-as-a-unit and enabling the creation of function-block-as-a-unit. Hardware-level disaggregation and software-defined wiring of resources is supported by a full-fledged Type-1 hypervisor that can execute commodity virtual machines, which communicate over a low-latency and high-throughput software-defined optical network. To evaluate its novel approach, dRedBox will demonstrate application execution in the domains of network functions virtualization, infrastructure analytics, and real-time video surveillance.
Maciej Bielski, Ilias Syrigos, Kostas Katrinis, Dimitris Syrivelis, Andrea Reale, Dimitris Theodoropoulos 0001, Nikolaos Alachiotis 0001, Dionisios N. Pnevmatikatos, E. H. Pap, Georgios Zervas, Vaibhawa Mishra, Arsalan Saljoghei, Alvise Rigo, Jose Fernando Zazo, Sergio López-Buedo, Martí Torrents, Ferad Zyulkyarov, Michael Enrico, Óscar González de Dios
DATE16
2016 Facing prefetching challenges in distributed shared memories for CMPs
Martí Torrents, Raúl Martínez, Carlos Molina 0004
J. Supercomput.1