Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Joseph Nuzman

dblp:97/4348 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 87% Processor architecture and microarchitecture · 6% Cloud and datacenter computing · 6%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
0.212015
High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive caches · HPCA 2015
Memory systems › memory hierarchy
cache hierarchy
0.212015
High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive caches · HPCA 2015
Memory systems › memory hierarchy › cache hierarchy management
exclusive caching
0.212015
High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive caches · HPCA 2015
Memory systems › memory hierarchy › cache hierarchy
last-level cache
0.212015
High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive caches · HPCA 2015
Processor architecture and microarchitecture
chip multiprocessor
0.112015
High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive caches · HPCA 2015
Cloud and datacenter computing
server workloads
0.112015
High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive caches · HPCA 2015

Methods — techniques the papers use, named apart from their topics

simulation · 0.2
YearPublicationVenuePosition
2015 High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive caches
abstract
Increasing transistor density enables adding more on-die cache real-estate However, devoting more space to the shared last-level-cache (LLC) causes the memory latency bottleneck to move from memory access latency to shared cache access latency. As such, applications whose working set is larger than the smaller caches spend a large fraction of their execution time on shared cache access latency. To address this problem, this paper investigates increasing the size of smaller private caches in the hierarchy as opposed to increasing the shared LLC. Doing so improves average cache access latency for workloads whose working set fits into the larger private cache while retaining the benefits of a shared LLC. The consequence of increasing the size of private caches is to relax inclusion and build exclusive hierarchies. Thus, for the same total caching capacity, an exclusive cache hierarchy provides better cache access latency. We observe that server workloads benefit tremendously from an exclusive hierarchy with large private caches. This is primarily because large private caches accommodate the large code working-sets of server workloads. For a 16-core CMP, an exclusive cache hierarchy improves server workload performance by 5-12% as compared to an equal capacity inclusive cache hierarchy. The paper also presents directions for further research to maximize performance of exclusive cache hierarchies.
Aamer Jaleel, Joseph Nuzman, Adrian Moga, Simon C. Steely Jr., Joel S. Emer
HPCA2
2012 Introducing hierarchy-awareness in replacement and bypass algorithms for last-level caches
abstract
The replacement policies for the last-level caches (LLCs) are usually designed based on the access information available locally at the LLC. These policies are inherently sub-optimal due to lack of information about the activities in the inner-levels of the hierarchy. This paper introduces cache hierarchy-aware replacement (CHAR) algorithms for inclusive LLCs (or L3 caches) and applies the same algorithms to implement efficient bypass techniques for exclusive LLCs in a three-level hierarchy. In a hierarchy with an inclusive LLC, these algorithms mine the L2 cache eviction stream and decide if a block evicted from the L2 cache should be made a victim candidate in the LLC based on the access pattern of the evicted block. Ours is the first proposal that explores the possibility of using a subset of L2 cache eviction hints to improve the replacement algorithms of an inclusive LLC. The CHAR algorithm classifies the blocks residing in the L2 cache based on their reuse patterns and dynamically estimates the reuse probability of each class of blocks to generate selective replacement hints to the LLC. Compared to the static re-reference interval prediction (SRRIP) policy, our proposal offers an average reduction of 10.9% in LLC misses and an average improvement of 3.8% in instructions retired per cycle (IPC) for twelve single-threaded applications. The corresponding reduction in LLC misses for one hundred 4-way multi-programmed workloads is 6.8% leading to an average improvement of 3.9% in throughput. Finally, our proposal achieves an 11.1% reduction in LLC misses and a 4.2% reduction in parallel execution cycles for six 8-way threaded shared memory applications compared to the SRRIP policy.
Mainak Chaudhuri, Jayesh Gaur, Nithiyanandan Bashyam, Sreenivas Subramoney, Joseph Nuzman
PACT5
2003 Towards a First Vertical Prototyping of an Extremely Fine-Grained Parallel Programming Approach
Dorit Naishlos, Joseph Nuzman, Chau-Wen Tseng, Uzi Vishkin
Theory Comput. Syst.2
2001 Evaluating the XMT Parallel Programming Model
Dorit Naishlos, Joseph Nuzman, Chau-Wen Tseng, Uzi Vishkin
HIPS2
2001 Evaluating the XMT Parallel Programming Model
abstract
Explicit-multithreading (XMT) is a parallel programming model designed for exploiting on-chip parallelism. Its features include a simple thread execution model and an efficient prefix-sum instruction for synchronizing shared data accesses. By taking advantage of low-overhead parallel threads and high on-chip memory bandwidth, the XMT model tries to reduce the burden on programmers by obviating the need for explicit task assignment and thread coarsening. This paper presents features of the XMT programming model, and evaluates their utility through experiments on a prototype XMT compiler and architecture simulator. We find the lack of explicit task assignment has slight effects on performance for the XMT architecture. Despite low thread overhead, thread coarsening is still necessary to some extent, but can usually be automatically applied by the XMT compiler. The prefix-sum instruction provides more scalable synchronization than traditional locks, and the simple run-untilcompletion thread execution model (no busy-waits) does not impair performance. Finally, the combination of features in XMT can encourage simpler parallel algorithms that may be more efficient than more traditional complex approaches.
Dorit Naishlos, Joseph Nuzman, Chau-Wen Tseng, Uzi Vishkin
IPDPS2
2001 Towards a first vertical prototyping of an extremely fine-grained parallel programming approach
abstract
Explicit-multithreading (XMT) is a parallel programming approach for exploiting on-chip parallelism. XMT introduces a computational framework with 1) a simple programming style that relies on fine-grained PRAM-style algorithms; 2) hardware support for low-overhead parallel threads, scalable load balancing, and efficient synchronization. The missing link between the algorithmic-programming level and the architecture level is provided by the first prototype XMT compiler. This paper also takes this new opportunity to evaluate the overall effectiveness of the interaction between the programming model and the hardware, and enhance its performance where needed, incorporating new optimizations into the XMT compiler. We present a wide range of applications, which written in XMT obtain significant speedups relative to the best serial programs. We show that XMT is especially useful for more advanced applications with dynamic, irregular access pattern, where for regular computations we demonstrate performance gains that scale up to much higher levels than have been demonstrated before for on-chip systems.
Dorit Naishlos, Joseph Nuzman, Chau-Wen Tseng, Uzi Vishkin
SPAA2
1998 Explicit Multi-Threading (XMT) Bridging Models for Instruction Parallelism (Extended Abstract)
abstract
This paper envisions an extension to a standard instruction set which efficiently implements PRAM-style algorithms using explicit multi-threaded instruction-level parallelism (ILP); that is, Explicit Multi-Threading (XMT), a fine-grained computational paradigm covering the spectrum from algorithms through architecture to implementation is introduced; new elements are added where needed.
Uzi Vishkin, Shlomit Dascal, Efraim Berkovich, Joseph Nuzman
SPAA4