EDBT 2026 Demo / reviewers in the wild / expert
Robert Stets
dblp:94/4458
· DBLP profile ↗
6ranked-venue papers
2as first author
0since 2021 · last 2005
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Memory systems · 81% Interconnection networks and networks-on-chip · 9% Processor architecture and microarchitecture · 5% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › shared memory
distributed shared memory |
0.1 | 4 | 2005 | Shared memory computing on clusters with symmetric multiprocessors and system area networks · ACM Trans. Comput. Syst. 2005 The Effect of Network Total Order, Broadcast, and Remote-Write Capability on Network-Based Shared Memory Computing · HPCA 2000 Comparative Evaluation of Fine- and Coarse-Grain Approaches for Software Distributed Shared Memory · HPCA 1999 |
Memory systems › shared memory › distributed shared memory
software distributed shared memory |
0.1 | 4 | 2005 | Shared memory computing on clusters with symmetric multiprocessors and system area networks · ACM Trans. Comput. Syst. 2005 The Effect of Network Total Order, Broadcast, and Remote-Write Capability on Network-Based Shared Memory Computing · HPCA 2000 Comparative Evaluation of Fine- and Coarse-Grain Approaches for Software Distributed Shared Memory · HPCA 1999 |
Memory systems › cache coherence
cache coherence protocol |
0.1 | 1 | 2005 | Shared memory computing on clusters with symmetric multiprocessors and system area networks · ACM Trans. Comput. Syst. 2005 |
Memory systems
cache coherence |
0.1 | 4 | 2005 | VM-Based Shared Memory on Low-Latency, Remote-Memory-Access Networks · ISCA 1997 Shared memory computing on clusters with symmetric multiprocessors and system area networks · ACM Trans. Comput. Syst. 2005 Piranha: a scalable architecture based on single-chip multiprocessing · ISCA 2000 |
Interconnection networks and networks-on-chip
cluster interconnect |
0.0 | 2 | 2005 | The Effect of Network Total Order, Broadcast, and Remote-Write Capability on Network-Based Shared Memory Computing · HPCA 2000 Shared memory computing on clusters with symmetric multiprocessors and system area networks · ACM Trans. Comput. Syst. 2005 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2000 | Piranha: a scalable architecture based on single-chip multiprocessing · ISCA 2000 |
Memory systems › cache coherence › coherence granularity
coarse-grain coherence |
0.0 | 1 | 1999 | Comparative Evaluation of Fine- and Coarse-Grain Approaches for Software Distributed Shared Memory · HPCA 1999 |
Memory systems › cache coherence
coherence granularity |
0.0 | 1 | 1999 | Comparative Evaluation of Fine- and Coarse-Grain Approaches for Software Distributed Shared Memory · HPCA 1999 |
Memory systems › memory consistency › memory consistency model
release consistency |
0.0 | 1 | 1997 | VM-Based Shared Memory on Low-Latency, Remote-Memory-Access Networks · ISCA 1997 |
Memory systems
shared memory |
0.0 | 1 | 1997 | Cashmere-2L: Software Coherent Shared Memory on a Clustered Remote-Write Network · SOSP 1997 |
Parallel and multicore computing › parallel computing › parallel communication
user-level communication |
0.0 | 1 | 2005 | Shared memory computing on clusters with symmetric multiprocessors and system area networks · ACM Trans. Comput. Syst. 2005 |
Memory systems › memory hierarchy
cache hierarchy |
0.0 | 1 | 2000 | Piranha: a scalable architecture based on single-chip multiprocessing · ISCA 2000 |
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor |
0.0 | 1 | 1999 | Comparative Evaluation of Fine- and Coarse-Grain Approaches for Software Distributed Shared Memory · HPCA 1999 |
High-performance computing
cluster computing |
0.0 | 1 | 1997 | Cashmere-2L: Software Coherent Shared Memory on a Clustered Remote-Write Network · SOSP 1997 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.1performance measurement · 0.1protocol variant evaluation · 0.0emulation · 0.0virtual memory coherence · 0.0instrumentation · 0.0release consistency · 0.0page-size coherence blocks · 0.0directory-based coherence · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2005 | Shared memory computing on clusters with symmetric multiprocessors and system area networksabstractCashmere is a software distributed shared memory (S-DSM) system designed for clusters of server-class machines. It is distinguished from most other S-DSM projects by (1) the effective use of fast user-level messaging, as provided by modern system-area networks, and (2) a “two-level” protocol structure that exploits hardware coherence within multiprocessor nodes. Fast user-level messages change the tradeoffs in coherence protocol design; they allow Cashmere to employ a relatively simple directory-based coherence protocol. Exploiting hardware coherence within SMP nodes improves overall performance when care is taken to avoid interference with inter-node software coherence.We have implemented Cashmere on a Compaq AlphaServer/Memory Channel cluster, an architecture that provides fast user-level messages. Experiments indicate that a one-level, version of the Cashmere protocol provides performance comparable to, or slightly better than, that of TreadMarks' lazy release consistency. Comparisons to Compaq's Shasta protocol also suggest that while fast user-level messages make finer-grain software DSMs competitive, VM-based systems continue to outperform software-based access control for applications without extensive fine-grain sharing.Within the family of Cashmere protocols, we find that leveraging intranode hardware coherence provides a 37% performance advantage over a more straightforward one-level implementation. Moreover, contrary to our original expectations, noncoherent hardware support for remote memory writes, total message ordering, and broadcast, provide comparatively little in the way of additional benefits over just fast messaging for our application suite. Leonidas I. Kontothanassis, Robert Stets, Galen C. Hunt, Umit Rencuzogullari, Gautam Altekar, Sandhya Dwarkadas, Michael L. Scott |
ACM Trans. Comput. Syst. | 2 |
| 2000 | The Effect of Network Total Order, Broadcast, and Remote-Write Capability on Network-Based Shared Memory ComputingabstractEmerging system-area networks provide a variety of features that can dramatically reduce network communication overhead. In this paper, we evaluate the impact of such features on the implementation of Software Distributed Shared Memory (SDSM), and on the Cashmere system in particular. Cashmere has been implemented on the Compaq Memory Channel network, which supports low-latency messages, protected remote memory writes, in-expensive broadcast, and total ordering of network packets. Our evaluation is based on several Cashmere protocol variants, ranging from a protocol that fully leverages the Memory Channel's special features to one that uses the network only for fast messaging. We find that the special features improve performance by 18-44% for three of our applications, but less than 12% for our other seven applications. We also find that home node migration, an optimization available only in the message-based protocol, can improve performance by as much as 67%. These results suggest that for systems of modest size, low latency is much more important for SDSM performance than are remote writes, broadcast, or total ordering. At the same time, results on an emulated 32-node system indicate that broadcast based on remote writes of widely-shared data may improve performance by up to 51% for some applications. If hardware broadcast or multicast facilities can be made to scale, they can be beneficial in future system-area networks. Robert Stets, Sandhya Dwarkadas, Leonidas I. Kontothanassis, Umit Rencuzogullari, Michael L. Scott |
HPCA | 1 |
| 2000 | Piranha: a scalable architecture based on single-chip multiprocessingabstractThis paper describes the Piranha system, a research prototype being developed at Compaq that aggressively exploits chip multiprocessing by integrating eight simple Alpha processor cores along with a two-level cache hierarchy onto a single chip. Piranha also integrates further on-chip functionality to allow for scalable multiprocessor configurations to be built in a glueless and modular fashion. The use of simple processor cores combined with an industry-standard ASIC design methodology allow us to complete our prototype within a short time-frame, with a team size and investment that are an order of magnitude smaller than that of a commercial microprocessor. Our detailed simulation results show that while each Piranha processor core is substantially slower than an aggressive next-generation processor, the integration of eight cores onto a single chip allows Piranha to outperform next-generation processors by up to 2.9 times (on a per chip basis) on important workloads such as OLTP. This performance advantage can approach a factor of five by using full-custom instead of ASIC logic. In addition to exploiting chip multiprocessing, the Piranha prototype incorporates several other unique design choices including a shared second-level cache with no inclusion, a highly optimized cache coherence protocol, and a novel I/O architecture. Luiz André Barroso, Kourosh Gharachorloo, Robert McNamara, Andreas Nowatzyk, Shaz Qadeer, Barton Sano, Robert Stets, Ben Verghese |
ISCA | 8 |
| 1999 | Comparative Evaluation of Fine- and Coarse-Grain Approaches for Software Distributed Shared MemoryabstractSymmetric multiprocessors (SMPs) connected with low-latency networks provide attractive building blocks for software distributed shared memory systems. Two distinct approaches have been used: the fine-grain approach that instruments application loads and stores to support a small coherence granularity, and the coarse-grain approach based on virtual memory hardware that provides coherence at a page granularity. Fine-grain systems offer a simple migration path for applications developed on hardware multiprocessors by supporting coherence protocols similar to those implemented in hardware. On the other hand, coarse-grain systems can potentially provide higher performance through more optimized protocols and larger transfer granularities, while avoiding instrumentation overheads. Numerous studies have examined each approach individually, but major differences in experimental platforms and applications make comparison of the approaches difficult. This paper presents a detailed comparison of two mature systems, Shasta and Cashmere, representing the fine- and coarse-grain approaches, respectively. Both systems are tuned to run on the same commercially available, state-of-the-art cluster of AlphaServer SMPs connected via a Memory Channel network. As expected, our results show that Shasta provides robust performance for applications tuned for hardware multiprocessors, and can better tolerate fine-grain synchronization. In contrast, Cashmere is highly sensitive to fine-grain synchronization, but provides a performance edge for applications with coarse-grain behavior. Interestingly, we found that the performance gap between the systems can often be bridged by program modifications that address coherence and synchronization granularity. In addition, our study reveals some unexpected results related to the interaction of current compiler technology with application instrumentation, and the ability of SMP-aware protocols to avoid certain performance disadvantages of coarse-grain approaches. Sandhya Dwarkadas, Kourosh Gharachorloo, Leonidas I. Kontothanassis, Daniel J. Scales, Michael L. Scott, Robert Stets |
HPCA | 6 |
| 1997 | VM-Based Shared Memory on Low-Latency, Remote-Memory-Access NetworksabstractRecent technological advances have produced network interfaces that provide users with very low-latency access to the memory of remote machines. We examine the impact of such networks on the implementation and performance of software DSM. Specifically, we compare two DSM systems---Cashmere and TreadMarks---on a 32-processor DEC Alpha cluster connected by a Memory Channel network.Both Cashmere and TreadMarks use virtual memory to maintain coherence on pages, and both use lazy, multi-writer release consistency. The systems differ dramatically, however, in the mechanisms used to track sharing information and to collect and merge concurrent updates to a page, with the result that Cashmere communicates much more frequently, and at a much finer grain.Our principal conclusion is that low-latency networks make DSM based on fine-grain communication competitive with more coarse-grain approaches, but that further hardware improvements will be needed before such systems can provide consistently superior performance. In our experiments, Cashmere scales slightly better than TreadMarks for applications with false sharing. At the same time, it is severely constrained by limitations of the current Memory Channel hardware. In general, performance is better for TreadMarks. Leonidas I. Kontothanassis, Galen C. Hunt, Robert Stets, Nikos Hardavellas, Michal Cierniak, Srinivasan Parthasarathy 0001, Wagner Meira Jr., Sandhya Dwarkadas, Michael L. Scott |
ISCA | 3 |
| 1997 | Cashmere-2L: Software Coherent Shared Memory on a Clustered Remote-Write NetworkabstractLow-latency remote-write networks, such as DEC's Memory Channel, provide the possibility of transparent, inexpensive, huge-scale shared-memory parallel computing on clusters of shared memory multiprocessors (SMPs).The challenge is to take advantage of hardwaresharedmemoryfor sharing within an SMI: and to ensure that software overheadis incurredonly when actively sharing data across SMPs in the cluster.In this paper, we describe a 'Ywolevel" software coherent shared memory system-Cashmere-2Lthat meets this challenge.CashmereSL uses hardware to share memory within a node, while exploiting the Memory Channel's remote-write capabilities to implement "moderately lazy" release consistency with multiple concurrent writers, directories, home nodes, and page-size coherence blocks across nodes.Cashmere-2L employs a novel coherence protocol that allows a high level of asynchrony by eliminating global directory locks and the needfor TLB shootdown.Remote interrupts are minimized by exploiting the remote-write capabilities of the Memory Channel network Cashmere-2L currently runs on an &node, 32-processor DEC AlphaServersystem.Speedups rangefrom 8 to 31 on 32processors for our benchmark suite, depending on the application's characteristics.We quanhfi the importance of ourprotocol optimizations by comparing perjormance to that of several alternative protocols that do not share memory in hardware within an SMP, and require more synchronization.In comparison to a one-level protocol that does not share memory in hardware within an SMP Cashmere-2L improves performance by up to 46%. Robert Stets, Sandhya Dwarkadas, Nikos Hardavellas, Galen C. Hunt, Leonidas I. Kontothanassis, Srinivasan Parthasarathy 0001, Michael L. Scott |
SOSP | 1 |