Mahdad Davari

dblp:160/0689 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
0.422015
The Effects of Granularity and Adaptivity on Private/Shared Classification for Coherence · ACM Trans. Archit. Code Optim. 2015
Hierarchical private/shared classification: The key to simple and efficient coherence for clustered cache hierarchies · HPCA 2015
Memory systems › cache coherence
cache coherence protocol
0.212015
The Effects of Granularity and Adaptivity on Private/Shared Classification for Coherence · ACM Trans. Archit. Code Optim. 2015
Memory systems › cache coherence
data classification
0.212015
The Effects of Granularity and Adaptivity on Private/Shared Classification for Coherence · ACM Trans. Archit. Code Optim. 2015
Memory systems › cache coherence
directory-based coherence
0.212015
Hierarchical private/shared classification: The key to simple and efficient coherence for clustered cache hierarchies · HPCA 2015
Memory systems › cache coherence › invalidation
self-invalidation
0.212015
The Effects of Granularity and Adaptivity on Private/Shared Classification for Coherence · ACM Trans. Archit. Code Optim. 2015

Methods — techniques the papers use, named apart from their topics

simulation · 0.2performance simulation · 0.2coherence protocol design · 0.2
YearPublicationVenuePosition
2015 An Efficient, Self-Contained, On-chip Directory: DIR1-SISD
abstract
Directory-based cache coherence is the de-facto standard for scalable shared-memory multi/many-cores and significant effort is invested in reducing its overhead. However, directory area and complexity optimizations are often antithetical to each other. Novel directory-less coherence schemes have been introduced to remove the complexity and cost associated with directories in their entirety. However, such schemes introduce new challenges by transferring some of the directory complexity and functionality to the OS and using the page table and the TLBs to store data classification information. In this work we bridge the gap between directory-based and directory-less coherence schemes and propose a hybrid scheme called DIR1-SISD which employs self-invalidation and self-downgrade as directory policies for the shared entries. DIR1-SISD allows simultaneous optimizations in area and complexity without relying on the OS. DIR1-SISD keeps track of a single -- private -- owner, or allows multiple-readers-multiple-writers to exist simultaneously by transferring the responsibility for their coherence to the corresponding cores. A DIR1-SISD self-contained directory cache has a unique ability to minimize eviction-induced complexities by allowing directory entries to be evicted without maintaining inclusion with the cached data (thus avoiding the complexities of broadcasts) and without the need to have a backing store. Using simulation we show that a small, self-contained, DIR1-SISD cache outperforms a traditional DIR16-NB MESI protocol with a directory cache embedded in the LLC (8% in execution time and 15% in traffic) and, further, outperforms a SISD protocol that relies on the OS to provide a persistent page-based directory (4% in execution time and 20% in traffic).
Mahdad Davari, Alberto Ros 0001, Erik Hagersten, Stefanos Kaxiras
PACT1
2015 Hierarchical private/shared classification: The key to simple and efficient coherence for clustered cache hierarchies
abstract
Hierarchical clustered cache designs are becoming an appealing alternative for multicores. Grouping cores and their caches in clusters reduces network congestion by localizing traffic among several hierarchical levels, potentially enabling much higher scalability. While such architectures can be formed recursively by replicating a base design pattern, keeping the whole hierarchy coherent requires more effort and consideration. The reason is that, in hierarchical coherence, even basic operations must be recursive. As a consequence, intermediate-level caches behave both as directories and as leaf caches. This leads to an explosion of states, protocol-races, and protocol complexity. While there have been previous efforts to extend directory-based coherence to hierarchical designs their increased complexity and verification cost is a serious impediment to their adoption. We aim to address these concerns by encapsulating all hierarchical complexity in a simple function: that of determining when a data block is shared entirely within a cluster (sub-tree of the hierarchy) and is private from the outside. This allows us to eliminate complex recursive operations that span the hierarchy and instead employ simple coherence mechanisms such as self-invalidation and write-through - now restricted to operate within the cluster where a data block is shared. We examine two inclusivity options and discuss the relation of our approach to the recently proposed Hierarchical-Race-Free (HRF) memory models. Finally, comparisons to a hierarchical directory-based MOESI, VIPS-M, and TokenCMP protocols show that, despite its simplicity our approach results in competitive performance and decreased network traffic.
Alberto Ros 0001, Mahdad Davari, Stefanos Kaxiras
HPCA2
2015 The Effects of Granularity and Adaptivity on Private/Shared Classification for Coherence
abstract
Classification of data into private and shared has proven to be a catalyst for techniques to reduce coherence cost, since private data can be taken out of coherence and resources can be concentrated on providing coherence for shared data. In this article, we examine how granularity—page-level versus cache-line level—and adaptivity—going from shared to private—affect the outcome of classification and its final impact on coherence. We create a classification technique, called Generational Classification , and a coherence protocol called Generational Coherence, which treats data as private or shared based on cache-line generations. We compare two coherence protocols based on self-invalidation/self-downgrade with respect to data classification. Our findings are enlightening: (i) Some programs benefit from finer granularity, some benefit further from adaptivity, but some do not benefit from either. (ii) Reducing the amount of shared data has no perceptible impact on coherence misses caused by self-invalidation of shared data, hence no impact on performance. (iii) In contrast, classifying more data as private has implications for protocols that employ write-through as a means of self-downgrade, resulting in network traffic reduction—up to 30%—by reducing write-through traffic.
Mahdad Davari, Alberto Ros 0001, Erik Hagersten, Stefanos Kaxiras
ACM Trans. Archit. Code Optim.1