Vicent Selfa

dblp:161/8055 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
0since 2021 · last 2020
0000-0002-9732-4831ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 79% Processor architecture and microarchitecture · 18% Performance modeling and evaluation · 4%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache management
0.722020
Phase-Aware Cache Partitioning to Target Both Turnaround Time and System Performance · IEEE Trans. Parallel Distributed Syst. 2020
A Hardware Approach to Fairly Balance the Inter-Thread Interference in Shared Caches · IEEE Trans. Parallel Distributed Syst. 2017
Memory systems › cache management
cache partitioning
0.412020
Phase-Aware Cache Partitioning to Target Both Turnaround Time and System Performance · IEEE Trans. Parallel Distributed Syst. 2020
Memory systems › memory hierarchy › cache hierarchy
last-level cache
0.412020
Phase-Aware Cache Partitioning to Target Both Turnaround Time and System Performance · IEEE Trans. Parallel Distributed Syst. 2020
Processor architecture and microarchitecture
chip multiprocessor
0.422020
A Hardware Approach to Fairly Balance the Inter-Thread Interference in Shared Caches · IEEE Trans. Parallel Distributed Syst. 2017
Phase-Aware Cache Partitioning to Target Both Turnaround Time and System Performance · IEEE Trans. Parallel Distributed Syst. 2020
Memory systems › cache management › cache partitioning
shared cache partitioning
0.312017
A Hardware Approach to Fairly Balance the Inter-Thread Interference in Shared Caches · IEEE Trans. Parallel Distributed Syst. 2017
Performance modeling and evaluation › workload characterization › parallel workload analysis
multicore workload characterization
0.112017
A Hardware Approach to Fairly Balance the Inter-Thread Interference in Shared Caches · IEEE Trans. Parallel Distributed Syst. 2017

Methods — techniques the papers use, named apart from their topics

phase-aware partitioning · 0.4runtime partition adjustment · 0.3fair-progress cache partitioning · 0.3
YearPublicationVenuePosition
2020 Phase-Aware Cache Partitioning to Target Both Turnaround Time and System Performance
abstract
The Last Level Cache (LLC) plays a key role in the system performance of current multi-cores by reducing the number of long latency main memory accesses. The inter-application interference at this shared resource, however, can lead the system to undesired situations regarding performance and fairness. Recent approaches have successfully addressed fairness and turnaround time (TT) in commercial processors. Nevertheless, these approaches must face sustaining system performance, which is challenging. This work makes two main contributions. LLC behaviors regarding cache performance, data reuse and cache occupancy, that adversely impact on the final performance are identified. Second, based on these behaviors, we propose the Critical-Phase Aware Partitioning Approach (CPA), which reduces TT while sustaining (and even improving) IPC by making an effective use of the LLC space. Experimental results show that CPA outperforms CA, Dunn and KPart state-of-the-art approaches, and improves TT (over 40 percent in some workloads) over Linux default behavior while sustaining or even improving IPC by more than 3 percent in several mixes.
Lucia Pons, Julio Sahuquillo, Vicent Selfa, Salvador Petit, Julio Pons
IEEE Trans. Parallel Distributed Syst.3
2018 Improving System Turnaround Time with Intel CAT by Identifying LLC Critical Applications
Lucia Pons, Vicent Selfa, Julio Sahuquillo, Salvador Petit, Julio Pons
Euro-Par2
2018 Efficient selective multicore prefetching under limited memory bandwidth
Vicent Selfa, Julio Sahuquillo, María Engracia Gómez, Crispín Gómez Requena
J. Parallel Distributed Comput.1
2017 Application Clustering Policies to Address System Fairness with Intel's Cache Allocation Technology
abstract
Achieving system fairness is a major design concern in current multicore processors. Unfairness arises due to contention in the shared resources of the system, such as the LLC and main memory. To address this problem, many research works have proposed novel cache partitioning policies aimed at addressing system fairness without harming performance. Unfortunately, existing proposals targeting fairness require extra hardware which makes them impractical in commercial processors.Recent Intel Xeon processors feature Cache Allocation Technology (CAT), a hardware cache partitioning mechanism that can be controlled from userspace software and that allows to create partitions in the LLC and assign different groups of applications to them.In this paper we propose a family of clustering-based cache partitioning policies to address fairness in systems that feature Intel's CAT. The proposal acts at two levels: applications showing similar amount of core stalls due to LLC accesses are first grouped into clusters, after which each cluster is given a number of ways using a simple mathematical model. To the best of our knowledge, this is the first attempt to address system fairness using the cache partitioning hardware in a real product. Results show that our best performing policy reduces system unfairness by up to 80% (39% on average) for 8-application workloads and by up to 45% (25% on average) for 12-application workloads compared to a non-partitioning approach.
Vicent Selfa, Julio Sahuquillo, Lieven Eeckhout, Salvador Petit, María Engracia Gómez
PACT1
2017 A research-oriented course on Advanced Multicore Architecture: Contents and active learning methodologies
Salvador Petit, Julio Sahuquillo, María Engracia Gómez, Vicent Selfa
J. Parallel Distributed Comput.4
2017 A Hardware Approach to Fairly Balance the Inter-Thread Interference in Shared Caches
abstract
Shared caches have become the common design choice in the vast majority of modern multi-core and many-core processors, since cache sharing improves throughput for a given silicon area. Sharing the cache, however, has a downside: the requests from multiple applications compete among them for cache resources, so the execution time of each application increases over isolated execution. The degree in which the performance of each application is affected by the interference becomes unpredictable yielding the system to unfairness situations. This paper proposes Fair-Progress Cache Partitioning (FPCP), a low-overhead hardware-based cache partitioning approach that addresses system fairness. FPCP reduces the interference by allocating to each application a cache partition and adjusting the partition sizes at runtime. To adjust partitions, our approach estimates during multicore execution the time each application would have taken in isolation, which is challenging. The proposed approach has two main differences over existing approaches. First, FPCP distributes cache ways incrementally, which makes the proposal less prone to estimation errors. Second, the proposed algorithm is much less costly than the state-of-the-art ASM-Cache approach. Experimental results show that, compared to ASM-Cache, FPCP reduces unfairness by 48 percent in four-application workloads and by 28 percent in eight-application workloads, without harming the performance.
Vicent Selfa, Julio Sahuquillo, Salvador Petit, María Engracia Gómez
IEEE Trans. Parallel Distributed Syst.1
2016 Student Research Poster: A Low Complexity Cache Sharing Mechanism to Address System Fairness
abstract
Shared caches have become, de facto, the common design choice in current multi-cores, ranging from embedded devices to high-performance processors. In these systems, requests from multiple applications compete for the cache resources, degrading to different extents their progress, quantified as the performance of individual applications compared to isolated execution. The difference between the progresses of the running applications yields the system to unpredictable behavior and causes a fairness problem. This problem can be addressed by carefully partitioning cache resources among the contending applications, but to be effective, a partitioning approach needs to estimate per-application progress. This work proposes Fair-Progress Cache Partitioning (FPCP), a low-overhead cache partitioning approach which addresses fairness by distributing cache resources among applications depending on their progress. To estimate progress, we have implemented two state-of-the-art performance models, ASM and PTCA, which estimate, at runtime, the performance a given application would have if executed in isolation.
Vicent Selfa, Julio Sahuquillo, Salvador Petit, María Engracia Gómez
PACT1
2016 A Simple Activation/Deactivation Prefetching Scheme for Chip Multiprocessors
abstract
Prefetching significantly reduces the memory latencies of a wide range of applications and thus increases the system performance. However, as a speculative technique, prefetching may also noticeably increase the number of memory accesses, which in turns may negatively impact on the main memory bandwidth consumption, performance, and power. Main memory bandwidth consumption is a critical resource especially in the context of current multicore processors since memory requests from all the cores, both prefetch and demand requests, compete among them in the access to the DRAM banks. Consequently, demand requests may be delayed hurting the system performance. This work proposes the Activation/Deactivation Policies (ADP) scheme for hardware prefetchers in multicore processors. This scheme relies on activation policies that turn on the prefetcher on a given core when it is expected that prefetches will improve the performance, and turn off the prefetcher of that core when it is foreseen that performance will be scarcely improved or not improved at all. The proposed mechanism effectively reduces the memory bandwidth requirements of some cores with respect to a typical always prefetching mechanism, so making available extra bandwidth to the co-runners. Results in a four-core processor show that ADP prefetching achieves similar performance ±2.5% as always prefetching, while significantly reducing the memory bandwidth consumed by use-less prefetches. Moreover, in some applications this reduction is as much as 50%. ADP prefetching is applicable to stream-based prefetchers, global-history-buffer delta correlation prefetchers, and PC-based stride prefetchers.
Vicent Selfa, Crispín Gómez Requena, María Engracia Gómez, Julio Sahuquillo
PDP1
2015 Row Tables: Design Choices to Exploit Bank Locality in Multiprogram Workloads
abstract
Main memory is a major performance bottleneck in current chip multiprocessors. Current DRAM banks latch the last accessed row in an internal buffer, namely row buffer (RB), which allows fast subsequent accesses to that row. This throughput-oriented approach was originally designed for single-thread processors and pursues to take advantage of the spatial locality that individual applications exhibit. This paper proposes row tables, a pool of row buffers shared among threads. Depending on the needs of each thread, row buffers are dynamically allocated to threads. Two design approaches are devised differing on the table location, and referred to as BRT (Bank Row Table) and CRT (Controller Row Table), which place the table at the bank, as traditionally done in existing modules, and at the memory controller side, respectively. CRT performs better than BRT in high RB locality applications (or mixes) but performs worse in poor RB locality applications since the increase in transfer times is not later amortized. A variant of CRT referred to as CRT1/xhas been devised to reduce this performance penalty. Results for a 4-core system show that, on average, BRT and CRT1/xmechanisms save energy by 23% and 7%-16% (depending on the X value) and improve IPC by 10% and 9%-14%, respectively.
Paula Navarro, Vicent Selfa, Julio Sahuquillo, María Engracia Gómez, Crispín Gómez Requena
PDP2
2015 Methodologies and Performance Metrics to Evaluate Multiprogram Workloads
abstract
Multicore processors are dominating the microprocessor market and most research work has moved to this kind of processors. Multicore research methods are still immature and evolving from the single-threaded processor ounterparts. Three main research issues must be faced when evaluating performance and energy in multicores. First, multiple simulation methodologies are being applied to evaluate these systems, without being an agreement about which to use. Second, due to the nature of multiprogram workloads new performance metrics are required, different from those used in single-thread processors. Many metrics have been defined and distinct metrics are used across the published works. Finally, multicore processors are really complex systems which require from sophisticated and complementary (e.g. energy and performance) simulators. This paper pursues to help researchers face the three mentioned research issues. For this purpose, we compare these issues across 28 papers published in 2013 in top computer architecture conferences. Both analytical examples and experimental results are presented with the aim of providing some insights in multicore research.
Vicent Selfa, Julio Sahuquillo, Crispín Gómez Requena, María Engracia Gómez
PDP1