Basem A. Nayfeh

dblp:31/466 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
0since 2021 · last 1996
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Memory systems · 36% Processor architecture and microarchitecture · 34% Parallel and multicore computing · 10%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
chip multiprocessor
0.021996
The Case for a Single-Chip Multiprocessor · ASPLOS 1996
Exploring the Design Space for a Shared-Cache Multiprocessor · ISCA 1994
Processor architecture and microarchitecture › clustered architecture
cluster-based multiprocessor
0.021995
The Benefits of Clustering in Shared Address Space Multiprocessors: An Applications-Driven Investigation · SC 1995
Exploring the Design Space for a Shared-Cache Multiprocessor · ISCA 1994
Memory systems
cache coherence
0.021996
The Benefits of Clustering in Shared Address Space Multiprocessors: An Applications-Driven Investigation · SC 1995
Evaluation of Design Alternatives for a Multiprocessor Microprocessor · ISCA 1996
Interconnection networks and networks-on-chip › interconnection networks
bus contention
0.011996
The Impact of Shared-Cache Clustering in Small-Scale Shared-Memory Multiprocessors · HPCA 1996
Memory systems
cache design
0.011996
Evaluation of Design Alternatives for a Multiprocessor Microprocessor · ISCA 1996
Memory systems › memory hierarchy
cache hierarchy
0.011996
The Impact of Shared-Cache Clustering in Small-Scale Shared-Memory Multiprocessors · HPCA 1996
Memory systems › cache › multiprocessor cache
shared cache
0.011996
Evaluation of Design Alternatives for a Multiprocessor Microprocessor · ISCA 1996
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor
0.011996
The Impact of Shared-Cache Clustering in Small-Scale Shared-Memory Multiprocessors · HPCA 1996
Processor architecture and microarchitecture
superscalar processor
0.011996
The Case for a Single-Chip Multiprocessor · ASPLOS 1996
Interconnection networks and networks-on-chip
cluster interconnect
0.011996
The Impact of Shared-Cache Clustering in Small-Scale Shared-Memory Multiprocessors · HPCA 1996
Memory systems › cache coherence
invalidation
0.011996
Evaluation of Design Alternatives for a Multiprocessor Microprocessor · ISCA 1996
Integrated circuit design › packaging
multichip module
0.011996
The Impact of Shared-Cache Clustering in Small-Scale Shared-Memory Multiprocessors · HPCA 1996
Parallel and multicore computing › parallel architecture
on-chip parallelism
0.011996
The Case for a Single-Chip Multiprocessor · ASPLOS 1996
Performance modeling and evaluation
simulation
0.011996
Evaluation of Design Alternatives for a Multiprocessor Microprocessor · ISCA 1996
Performance modeling and evaluation › parallel performance evaluation
scientific application performance
0.011995
The Benefits of Clustering in Shared Address Space Multiprocessors: An Applications-Driven Investigation · SC 1995
Performance modeling and evaluation
workload characterization
0.011995
The Benefits of Clustering in Shared Address Space Multiprocessors: An Applications-Driven Investigation · SC 1995
Memory systems › cache management
cache partitioning
0.011994
Exploring the Design Space for a Shared-Cache Multiprocessor · ISCA 1994
Performance modeling and evaluation › design trade-off analysis
cost-performance analysis
0.011994
Exploring the Design Space for a Shared-Cache Multiprocessor · ISCA 1994

Methods — techniques the papers use, named apart from their topics

simulation · 0.0full-system simulation · 0.0application-driven performance analysis · 0.0
YearPublicationVenuePosition
1996 The Case for a Single-Chip Multiprocessor
abstract
Advances in IC processing allow for more microprocessor design options. The increasing gate density and cost of wires in advanced integrated circuit technologies require that we look for new ways to use their capabilities effectively. This paper shows that in advanced technologies it is possible to implement a single-chip multiprocessor in the same area as a wide issue superscalar processor. We find that for applications with little parallelism the performance of the two microarchitectures is comparable. For applications with large amounts of parallelism at both the fine and coarse grained levels, the multiprocessor microarchitecture outperforms the superscalar architecture by a significant margin. Single-chip multiprocessor architectures have the advantage in that they offer localized implementation of a high-clock rate processor for inherently sequential applications and low latency interprocessor communication for parallel applications.
Kunle Olukotun, Basem A. Nayfeh, Lance Hammond, Kenneth G. Wilson, Kunyung Chang
ASPLOS2
1996 The Impact of Shared-Cache Clustering in Small-Scale Shared-Memory Multiprocessors
abstract
As processor performance continues to increase, greater demands are placed on the bus and memory systems of small-scale shared-memory multiprocessors. In this paper, we investigate how to reduce these demands by organizing groups of processors into clusters which are then connected together using a shared global bus. We take advantage of the high-bandwidth, low-latency interconnections available from multichip module (MCM) technology, to build clusters with multiple high-performance processors sharing an L2 cache. The use of MCM technology allows for significantly lower shared-cache access times, and higher shared cache to processor bandwidth, than is possible using printed circuit board (PCB) designs. Our results show that for an eight processor bus-based system, bus contention can be a large portion of the overall execution time, and that clustering can eliminate much or all of it. Clustering also tends to reduce read stall times due to shared working set effects and a reduction in the effect of communication misses. The same is true for two and four processor systems, although to a lesser extent. Overall, we find that clustering can result in significant performance gains for applications which heavily utilize the memory system.
Basem A. Nayfeh, Kunle Olukotun, Jaswinder Pal Singh
HPCA1
1996 Evaluation of Design Alternatives for a Multiprocessor Microprocessor
abstract
In the future, advanced integrated circuit processing and packaging technology will allow for several design options for multiprocessor microprocessors. In this paper we consider three architectures: shared-primary cache, shared-secondary cache, and shared-memory. We evaluate these three architectures using a complete system simulation environment which models the CPU, memory hierarchy and I/O devices in sufficient detail to boot and run a commercial operating system. Within our simulation environment, we measure performance using representative hand and compiler generated parallel applications, and a multiprogramming workload. Our results show that when applications exhibit fine-grained sharing, both shared-primary and shared-secondary architectures perform similarly when the full costs of sharing the primary cache are included.
Basem A. Nayfeh, Lance Hammond, Kunle Olukotun
ISCA1
1995 The Benefits of Clustering in Shared Address Space Multiprocessors: An Applications-Driven Investigation
abstract
Clustering processors together at a level of the memory hierarchy in shared address space multiprocessors appears to be an attractive technique from several standpoints: Resources are shared, packaging technologies are exploited, and processors within a cluster can share data more effectively. We investigate the performance benefits that can be obtained by clustering on a range of important scientific and engineering applications in moderate to large scale cache coherent machines with small degrees of clustering (up to one eighth of the total number of processors in a cluster). We find that except for applications with near neighbor communication topologies this degree of clustering is not very effective in reducing the inherent communication to computation ratios. Clustering is more useful in reducing the the number of remote capacity misses in unstructured applications, and can improve performance substantially when small first-level caches are clustered in these cases. This suggests that clustering at the first level cache might be useful in highly-integrated, relatively fine-grained environments. For less integrated machines such as current distributed shared memory multiprocessors, our results suggest that clustering at the first-level caches is not very useful in improving application performance; however our results also suggest that in an machine with long interprocessor communication latencies, clustering further away from the processor can provide performance benefits.
Andrew Erlichson, Basem A. Nayfeh, Jaswinder Pal Singh, Kunle Olukotun
SC2
1994 Exploring the Design Space for a Shared-Cache Multiprocessor
abstract
In the near future, semiconductor technology will allow the integration of multiple processors on a chip or multichip-module (MCM). The authors investigate the architecture and partitioning of resources between processors and cache memory for single chip and MCM-based multiprocessors. They study the performance of a cluster-based multiprocessor architecture in which processors within a cluster are tightly coupled via a shared cluster cache for various processor-cache configurations. The results show that for parallel applications, clustering via shared caches provides an effective mechanism for increasing the total number of processors in a system, without increasing the number of invalidations. Combining these results with cost estimates for shared cluster cache implementations leads to two conclusions: 1) For a four cluster multiprocessor with single chip clusters, two processors per cluster with a smaller cache provides higher performance and better cost/performance than a single processor with a larger cache and 2) this four cluster configuration can be scaled linearly in performance by adding processors to each cluster using MCM packaging techniques.>
Basem A. Nayfeh, Kunle Olukotun
ISCA1