Jacob Leverich

dblp:93/999 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Memory systems · 50% Cloud and datacenter computing · 22% Energy-efficient computing · 14%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.212014
Reconciling high server utilization and sub-millisecond quality-of-service · EuroSys 2014
Cloud and datacenter computing › request scheduling
latency-sensitive scheduling
0.212014
Reconciling high server utilization and sub-millisecond quality-of-service · EuroSys 2014
Memory systems
on-chip memory
0.222008
Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008
Comparing memory systems for chip multiprocessors · ISCA 2007
Memory systems
DRAM
0.112012
Improving System Energy Efficiency with Memory Rank Subsetting · ACM Trans. Archit. Code Optim. 2012
Energy-efficient computing › power management
memory power management
0.112012
Improving System Energy Efficiency with Memory Rank Subsetting · ACM Trans. Archit. Code Optim. 2012
Memory systems › DRAM › DRAM microarchitecture
rank subsetting
0.112012
Improving System Energy Efficiency with Memory Rank Subsetting · ACM Trans. Archit. Code Optim. 2012
Processor architecture and microarchitecture
chip multiprocessor
0.132009
Comparing memory systems for chip multiprocessors · ISCA 2007
Future scaling of processor-memory interfaces · SC 2009
Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008
Energy-efficient computing
memory energy efficiency
0.112009
Future scaling of processor-memory interfaces · SC 2009
Memory systems › memory interface
processor-memory interface
0.112009
Future scaling of processor-memory interfaces · SC 2009
Memory systems
cache coherence
0.112008
Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008
Memory systems › memory access optimization
memory streaming
0.112008
Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008
Memory systems › memory management
software-managed memory
0.112008
Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008
Memory systems › on-chip memory
on-chip memory design
0.112007
Comparing memory systems for chip multiprocessors · ISCA 2007
Hardware reliability and fault tolerance › error-correcting codes for memory
chipkill correct
0.012012
Improving System Energy Efficiency with Memory Rank Subsetting · ACM Trans. Archit. Code Optim. 2012
Hardware reliability and fault tolerance
memory reliability
0.012012
Improving System Energy Efficiency with Memory Rank Subsetting · ACM Trans. Archit. Code Optim. 2012
Parallel and multicore computing › parallel programming models
stream programming
0.012007
Comparing memory systems for chip multiprocessors · ISCA 2007

Methods — techniques the papers use, named apart from their topics

simulation · 0.3process technology analysis · 0.1performance evaluation · 0.1performance comparison · 0.1
YearPublicationVenuePosition
2014 Reconciling high server utilization and sub-millisecond quality-of-service
abstract
The simplest strategy to guarantee good quality of service (QoS) for a latency-sensitive workload with sub-millisecond latency in a shared cluster environment is to never run other workloads concurrently with it on the same server. Unfortunately, this inevitably leads to low server utilization, reducing both the capability and cost effectiveness of the cluster.
Jacob Leverich, Christoforos E. Kozyrakis
EuroSys1
2012 Improving System Energy Efficiency with Memory Rank Subsetting
abstract
VLSI process technology scaling has enabled dramatic improvements in the capacity and peak bandwidth of DRAM devices. However, current standard DDR x DIMM memory interfaces are not well tailored to achieve high energy efficiency and performance in modern chip-multiprocessor-based computer systems. Their suboptimal performance and energy inefficiency can have a significant impact on system-wide efficiency since much of the system power dissipation is due to memory power. New memory interfaces, better suited for future many-core systems, are needed. In response, there are recent proposals to enhance the energy efficiency of main-memory systems by dividing a memory rank into subsets, and making a subset rather than a whole rank serve a memory request. We holistically assess the effectiveness of rank subsetting from system-wide performance, energy-efficiency, and reliability perspectives. We identify the impact of rank subsetting on memory power and processor performance analytically, compare two promising rank-subsetting proposals, Multicore DIMM and mini-rank, and verify our analysis by simulating a chip-multiprocessor system using multithreaded and consolidated workloads. We extend the design of Multicore DIMM for high-reliability systems and show that compared with conventional chipkill approaches, rank subsetting can lead to much higher system-level energy efficiency and performance at the cost of additional DRAM devices. This holistic assessment shows that rank subsetting offers compelling alternatives to existing processor-memory interfaces for future DDR systems.
Jung Ho Ahn, Norman P. Jouppi, Christoforos E. Kozyrakis, Jacob Leverich, Robert S. Schreiber
ACM Trans. Archit. Code Optim.4
2010 Evaluating impact of manageability features on device performance
abstract
Manageability is a key design constraint for IT solutions, defined as the range of operations required to maintain and administer system resources through their lifecycle phases. Emerging complex and powerful management platforms and automation software expose new tensions-between the host and management applications, between costs and performance, and between costs and complexity. In this paper, we take a systematic approach to the evaluation of manageability workloads and define metrics for evaluating manageability efficiency. We propose the Manageability Quotient (MaQ) as a holistic measure of a system's ability to deliver guarantees on both host and management application performance while minimizing cost. We evaluate a range of host and management workloads on various manageability platforms using the metrics. Our results, based on more than 4000 experiments, provides insights on the merits and demerits of different configurations.
Jacob Leverich, Vanish Talwar, Parthasarathy Ranganathan, Christoforos E. Kozyrakis
CNSM1
2009 Future scaling of processor-memory interfaces
abstract
Continuous evolution in process technology brings energy-efficiency and reliability challenges, which are harder for memory system designs since chip multiprocessors demand high bandwidth and capacity, global wires improve slowly, and more cells are susceptible to hard and soft errors. Recently, there are proposals aiming at better main-memory energy efficiency by dividing a memory rank into subsets.
Jung Ho Ahn, Norman P. Jouppi, Christoforos E. Kozyrakis, Jacob Leverich, Robert S. Schreiber
SC4
2008 Comparative evaluation of memory models for chip multiprocessors
abstract
There are two competing models for the on-chip memory in Chip Multiprocessor (CMP) systems: hardware-managed coherent caches and software-managed streaming memory . This paper performs a direct comparison of the two models under the same set of assumptions about technology, area, and computational capabilities. The goal is to quantify how and when they differ in terms of performance, energy consumption, bandwidth requirements, and latency tolerance for general-purpose CMPs. We demonstrate that for data-parallel applications on systems with up to 16 cores, the cache-based and streaming models perform and scale equally well. For certain applications with little data reuse, streaming scales better due to better bandwidth use and macroscopic software prefetching. However, the introduction of techniques such as hardware prefetching and nonallocating stores to the cache-based model eliminates the streaming advantage. Overall, our results indicate that there is not sufficient advantage in building streaming memory systems where all on-chip memory structures are explicitly managed. On the other hand, we show that streaming at the programming model level is particularly beneficial, even with the cache-based model, as it enhances locality and creates opportunities for bandwidth optimizations. Moreover, we observe that stream programming is actually easier with the cache-based model because the hardware guarantees correct, best-effort execution even when the programmer cannot fully regularize an application's code.
Jacob Leverich, Hideho Arakida, Alex Solomatnikov, Amin Firoozshahian, Mark Horowitz, Christoforos E. Kozyrakis
ACM Trans. Archit. Code Optim.1
2007 Comparing memory systems for chip multiprocessors
abstract
There are two basic models for the on-chip memory in CMP systems:hardware-managed coherent caches and software-managed streaming memory. This paper performs a direct comparison of the two modelsunder the same set of assumptions about technology, area, and computational capabilities. The goal is to quantify how and when they differ in terms of performance, energy consumption, bandwidth requirements, and latency tolerance for general-purpose CMPs. We demonstrate that for data-parallel applications, the cache-based and streaming models perform and scale equally well. For certain applications with little data reuse, streaming scales better due to better bandwidth use and macroscopic software prefetching. However, the introduction of techniques such as hardware prefetching and non-allocating stores to the cache-based model eliminates the streaming advantage. Overall, our results indicate that there is not sufficient advantage in building streaming memory systems where all on-chip memory structures are explicitly managed. On the other hand, we show that streaming at the programming model level is particularly beneficial, even with the cache-based model, as it enhances locality and creates opportunities for bandwidth optimizations. Moreover, we observe that stream programming is actually easier with the cache-based model because the hardware guarantees correct, best-effort execution even when the programmer cannot fully regularize an application's code.
Jacob Leverich, Hideho Arakida, Alex Solomatnikov, Amin Firoozshahian, Mark Horowitz, Christoforos E. Kozyrakis
ISCA1