Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Albert Esteve

dblp:125/8617 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
0since 2021 · last 2018
0000-0001-8189-3776ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 72% Processor architecture and microarchitecture · 24% Performance modeling and evaluation · 4%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
0.932018
TokenTLB+CUP: A Token-Based Page Classification with Cooperative Usage Prediction · IEEE Trans. Parallel Distributed Syst. 2018
TLB-Based Temporality-Aware Classification in CMPs with Multilevel TLBs · IEEE Trans. Parallel Distributed Syst. 2017
Efficient TLB-Based Detection of Private Pages in Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2016
Processor architecture and microarchitecture
chip multiprocessor
0.522017
TLB-Based Temporality-Aware Classification in CMPs with Multilevel TLBs · IEEE Trans. Parallel Distributed Syst. 2017
Efficient TLB-Based Detection of Private Pages in Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2016
Memory systems › memory management › virtual memory › address translation
TLB
0.532018
TokenTLB+CUP: A Token-Based Page Classification with Cooperative Usage Prediction · IEEE Trans. Parallel Distributed Syst. 2018
TLB-Based Temporality-Aware Classification in CMPs with Multilevel TLBs · IEEE Trans. Parallel Distributed Syst. 2017
Efficient TLB-Based Detection of Private Pages in Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2016
Memory systems › cache coherence
data classification
0.312017
TLB-Based Temporality-Aware Classification in CMPs with Multilevel TLBs · IEEE Trans. Parallel Distributed Syst. 2017
Memory systems › virtual memory management
TLB hierarchy
0.312017
TLB-Based Temporality-Aware Classification in CMPs with Multilevel TLBs · IEEE Trans. Parallel Distributed Syst. 2017
Processor architecture and microarchitecture
multicore design
0.112018
TokenTLB+CUP: A Token-Based Page Classification with Cooperative Usage Prediction · IEEE Trans. Parallel Distributed Syst. 2018
Performance modeling and evaluation › parallel performance evaluation
multicore scalability
0.112018
TokenTLB+CUP: A Token-Based Page Classification with Cooperative Usage Prediction · IEEE Trans. Parallel Distributed Syst. 2018

Methods — techniques the papers use, named apart from their topics

simulation · 0.5token-based classification · 0.3cycle-accurate simulation · 0.3unicast messaging · 0.3
YearPublicationVenuePosition
2018 TokenTLB+CUP: A Token-Based Page Classification with Cooperative Usage Prediction
abstract
Discerning the private or shared condition of the data accessed by the applications is an increasingly decisive approach to achieving efficiency and scalability in multiand many-core systems. Since most memory accesses in both sequential and parallel applications are either private (accessed only by one core) or read-only (not written) data, devoting the full cost of coherence to every memory access results in sub-optimal performance and limits the scalability and efficiency of the multiprocessor. This paper introduces TokenTLB, a TLB-based page classification approach based on exchange and count of tokens. Token counting on TLBs is a natural and efficient way for classifying memory pages, and it does not require the use of complex and undesirable persistent requests or arbitration. In addition, classification is extended with Cooperative Usage Predictor (CUP), a token-based system-wide page usage predictor retrieved through TLB cooperation, in order to perform a classification unaffected by TLB size. Through cycle-accurate simulation we observed that TokenTLB spends 43.6 percent of cycles as private per page on average, and CUP further increases the time spent as private by 22.0 percent. CUP avoids 4 out of 5 TLB invalidations when compared to state-of-the-art predictors, thus proving far better prediction accuracy and making usage prediction an attractive mechanism for the first time.
Albert Esteve, Alberto Ros 0001, Antonio Robles, María Engracia Gómez
IEEE Trans. Parallel Distributed Syst.1
2017 TLB-Based Temporality-Aware Classification in CMPs with Multilevel TLBs
abstract
Recent proposals are based on classifying memory accesses into private or shared in order to process private accesses more efficiently and reduce coherence overhead. The classification mechanisms previously proposed are either not able to adapt to the dynamic sharing behavior of the applications or require frequent broadcast messages. Additionally, most of these classification approaches assume single-level translation lookaside buffers (TLBs). However, deeper and more efficient TLB hierarchies, such as the ones implemented in current commodity processors, have not been appropriately explored. This paper analyzes accurate classification mechanisms in multilevel TLB hierarchies. In particular, we propose an efficient data classification strategy for systems with distributed shared last-level TLBs. Our approach classifies data accounting for temporal private accesses and constrains TLB-related traffic by issuing unicast messages on first-level TLB misses. When our classification is employed to deactivate coherence for private data in directory-based protocols, it improves the directory efficiency and, consequently, reduces coherence traffic to merely 53.0 percent, on average. Additionally, it avoids some of the overheads of previous classification approaches for purely private TLBs, improving average execution time by nearly 9 percent for large-scale systems.
Albert Esteve, Alberto Ros 0001, María Engracia Gómez, Antonio Robles, José Duato
IEEE Trans. Parallel Distributed Syst.1
2016 TokenTLB: A Token-Based Page Classification Approach
abstract
Classifying memory accesses into private or shared data has become a fundamental approach to achieving efficiency and scalability in multi- and many-core systems. Since most memory accesses in both sequential and parallel applications are either private (accessed only by one core) or read-only (not written) data, devoting the full cost of coherence to every memory access results in sub-optimal performance and limits the scalability and efficiency of the multiprocessor.
Albert Esteve, Alberto Ros 0001, Antonio Robles, María Engracia Gómez, José Duato
ICS1
2016 Efficient TLB-Based Detection of Private Pages in Chip Multiprocessors
abstract
Most of the data referenced by sequential and parallel applications running in current chip multiprocessors are referenced by a single thread, i.e., private. Recent proposals leverage this observation to improve many aspects of chip multiprocessors, such as reducing coherence overhead or the access latency to distributed caches. The effectiveness of those proposals depends to a large extent on the amount of detected private data. However, the mechanisms proposed so far do not consider neither thread migration nor the private use of data within different application phases. As a result, a considerable amount of private data is not detected. In order to increase the detection of private data, we propose a TLB-based mechanism that is able to account for both thread migration and application phases. Simulation results show that the average number of pages detected as private significantly increases from 43 percent in previous proposals up to 79 percent in ours while keeping a reasonable TLB miss rate. Furthermore, when our proposal is used to deactivate the coherence for private data in a directory protocol, it improves execution time by 13.5 percent, on average, with respect to previous techniques.
Albert Esteve, Alberto Ros 0001, María Engracia Gómez, Antonio Robles, José Duato
IEEE Trans. Parallel Distributed Syst.1