Richard Simoni

dblp:48/2706 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
0since 2021 · last 1994
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 41% Processor architecture and microarchitecture · 22% Performance modeling and evaluation · 19%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
multiprocessor architecture
0.021994
The Stanford FLASH Multiprocessor · ISCA 1994
The Performance Impact of Flexibility in the Stanford FLASH Multiprocessor · ASPLOS 1994
Memory systems
cache coherence
0.021991
Modeling the Performance of Limited Pointers Directories for Cache Coherence · ISCA 1991
An Evaluation of Directory Schemes for Cache Coherence · ISCA 1988
Memory systems › cache coherence
directory-based coherence
0.021991
Modeling the Performance of Limited Pointers Directories for Cache Coherence · ISCA 1991
An Evaluation of Directory Schemes for Cache Coherence · ISCA 1988
Memory systems › cache coherence
cache-coherent shared memory
0.011994
The Stanford FLASH Multiprocessor · ISCA 1994
Parallel and multicore computing › parallel programming models
message passing
0.011994
The Stanford FLASH Multiprocessor · ISCA 1994
Performance modeling and evaluation
analytical modeling
0.011991
Modeling the Performance of Limited Pointers Directories for Cache Coherence · ISCA 1991
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor
0.021991
Modeling the Performance of Limited Pointers Directories for Cache Coherence · ISCA 1991
An Evaluation of Directory Schemes for Cache Coherence · ISCA 1988
Memory systems › cache coherence
cache coherence protocol
0.011994
The Performance Impact of Flexibility in the Stanford FLASH Multiprocessor · ASPLOS 1994
Interconnection networks and networks-on-chip
network interface
0.011994
The Stanford FLASH Multiprocessor · ISCA 1994
Memory systems › cache coherence › cache coherence protocol
snoopy coherence
0.011988
An Evaluation of Directory Schemes for Cache Coherence · ISCA 1988

Methods — techniques the papers use, named apart from their topics

simulation · 0.0verilog · 0.0system-level simulation · 0.0analytic modeling · 0.0trace-driven simulation · 0.0
YearPublicationVenuePosition
1994 The Performance Impact of Flexibility in the Stanford FLASH Multiprocessor
abstract
A flexible communication mechanism is a desirable feature in multiprocessors because it allows support for multiple communication protocols, expands performance monitoring capabilities, and leads to a simpler design and debug process. In the Stanford FLASH multiprocessor, flexibility is obtained by requiring all transactions in a node to pass through a programmable node controller, called MAGIC. In this paper, we evaluate the performance costs of flexibility by comparing the performance of FLASH to that of an idealized hardwired machine on representative parallel applications and a multiprogramming workload. To measure the performance of FLASH, we use a detailed simulator of the FLASH and MAGIC designs, together with the code sequences that implement the cache-coherence protocol. We find that for a range of optimized parallel applications the performance differences between the idealized machine and FLASH are small. For these programs, either the miss rates are small or the latency of the programmable protocol can be hidden behind the memory access time. For applications that incur a large number of remote misses or exhibit substantial hot-spotting, performance is poor for both machines, though the increased remote access latencies or the occupancy of MAGIC lead to lower performance for the flexible design. In most cases, however, FLASH is only 2%–12% slower than the idealized machine.
Mark A. Heinrich, Jeffrey Kuskin, David Ofelt, John Heinlein, Joel Baxter, Jaswinder Pal Singh, Richard Simoni, Kourosh Gharachorloo, David Nakahira, Mark Horowitz, Anoop Gupta, Mendel Rosenblum, John L. Hennessy
ASPLOS7
1994 The Stanford FLASH Multiprocessor
abstract
The FLASH multiprocessor efficiently integrates support for cache-coherent shared memory and high-performance message passing, while minimizing both hardware and software overhead. Each node in FLASH contains a microprocessor, a portion of the machine's global memory, a port to the interconnection network, The MAGIC chip handles all communication both within the node and among nodes, using hardwired data paths for efficient data movement and a programmable processor optimized for executing protocol operations. The use of the protocol processor makes FLASH very flexible/spl minus/it can support a variety of different communication mechanisms/spl minus/and simplifies the design and implementation. This paper presents the architecture of FLASH and MAGIC, and discusses the base cache-coherence and message-passing protocols. Latency and occupancy numbers, which are derived from our system-level simulator and our Verilog code, are given for several common protocol operations. The paper also describes our software strategy and FLASH's current status.>
Jeffrey Kuskin, David Ofelt, Mark A. Heinrich, John Heinlein, Richard Simoni, Kourosh Gharachorloo, John Chapin, David Nakahira, Joel Baxter, Mark Horowitz, Anoop Gupta, Mendel Rosenblum, John L. Hennessy
ISCA5
1991 Modeling the Performance of Limited Pointers Directories for Cache Coherence
abstract
Directory-based protocols have been proposed as an efficient means of implementing cache consistency in large-scale shamxlmemory multiprocessors.One class of these protocols utilizes a limited pointers directory, which stores the identities of a small number of caches containing a given block of data.However, the performance potential of these directories in large-scale machines has been speculative at best.In this paper we introduce an analytic model that not only explains the behavior seen in small-scale simulation studies, but also allows us to extrapolate forward to evaluate the efficiency of limited pointers directories in large-scale systems.Our model shows that miss rates inherent to invalidation-based consistency schemes me relatively high (typically 107o to 60Y0) for actively shared da~across a variety of workloads.We find that limited pointers schemes that resort to broadcasting invalidations when the pointers are exhausted perform very poorly in largescale machines, even if there are sufficient pointers most of the time.On the other han~no-broadcast strategies that limit the degree of caching to the number of pointers in an entry have only a modest impact on the cache miss rate and network traffic under a wide range of workloads, including those in which data blocks are actively accessed by a large number of processors.
Richard Simoni, Mark Horowitz
ISCA1
1988 An Evaluation of Directory Schemes for Cache Coherence
abstract
The problem of cache coherence in shared-memory multiprocessors is addressed using two basic approaches: directory schemes and snoopy cache systems. Directory schemes for cache coherence are potentially attractive in large multiprocessor systems that are beyond the scaling limits of the snoopy cache schemes. Slight modifications to directory schemes can make them competitive in performance with snoopy cache schemes for small multiprocessors. Trace-driven simulation, using data collected from several real multiprocessor applications, is used to compare the performance of standard directory schemes, modifications to these schemes, and snoopy cache protocols. In addition, the simulations show that most blocks that are written into are present in only a small number of other caches, which makes broadcast invalidates inefficient. This result suggests that a directory structure that stores with each block only a small number of pointers to caches containing the block is sufficient.>
Anant Agarwal, Richard Simoni, John L. Hennessy, Mark Horowitz
ISCA2