Richard J. Eickemeyer

dblp:89/2763 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Processor architecture and microarchitecture · 31% Energy-efficient computing · 28% Hardware accelerators and domain-specific architectures · 26%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.512021
Energy Efficiency Boost in the AI-Infused POWER10 Processor · ISCA 2021
Processor architecture and microarchitecture › microprocessor design
processor core design
0.512021
Energy Efficiency Boost in the AI-Infused POWER10 Processor · ISCA 2021
Performance modeling and evaluation › simulation › architectural simulation
cycle-accurate simulation
0.112011
Abstraction and microarchitecture scaling in early-stage power modeling · HPCA 2011
Energy-efficient computing › power modeling
microarchitecture-level power modeling
0.112011
Abstraction and microarchitecture scaling in early-stage power modeling · HPCA 2011
Energy-efficient computing
power modeling
0.112011
Abstraction and microarchitecture scaling in early-stage power modeling · HPCA 2011
Performance modeling and evaluation
simulation
0.112011
Abstraction and microarchitecture scaling in early-stage power modeling · HPCA 2011
Energy-efficient computing
clock gating
0.112005
Stretching the Limits of Clock-Gating Efficiency in Server-Class Processors · HPCA 2005
Processor architecture and microarchitecture › pipelining
pipeline design
0.112005
Stretching the Limits of Clock-Gating Efficiency in Server-Class Processors · HPCA 2005
Energy-efficient computing
power management
0.112005
Stretching the Limits of Clock-Gating Efficiency in Server-Class Processors · HPCA 2005
Energy-efficient computing
leakage power
0.012005
Stretching the Limits of Clock-Gating Efficiency in Server-Class Processors · HPCA 2005
Processor architecture and microarchitecture
multithreading
0.011996
Evaluation of Multithreaded Uniprocessors for Commercial Application Environments · ISCA 1996
Processor architecture and microarchitecture
instruction-level parallelism
0.011992
Interlock collapsing ALU for increased instruction-level parallelism · MICRO 1992
Memory systems › on-chip memory
on-chip memory organization
0.011988
Performance Evaluation of On-Chip Register and Cache Organizations · ISCA 1988
Performance modeling and evaluation › workload characterization
commercial workloads
0.011996
Evaluation of Multithreaded Uniprocessors for Commercial Application Environments · ISCA 1996
Processor architecture and microarchitecture
memory latency tolerance
0.011996
Evaluation of Multithreaded Uniprocessors for Commercial Application Environments · ISCA 1996
Performance modeling and evaluation
workload characterization
0.011996
Evaluation of Multithreaded Uniprocessors for Commercial Application Environments · ISCA 1996
Processor architecture and microarchitecture › register file
register file organization
0.011987
Performance Evaluation of Multiple Register Sets · ISCA 1987
Memory systems
cache design
0.011988
Performance Evaluation of On-Chip Register and Cache Organizations · ISCA 1988
Memory systems › cache
instruction and data cache
0.011988
Performance Evaluation of On-Chip Register and Cache Organizations · ISCA 1988

Methods — techniques the papers use, named apart from their topics

pre-silicon methodology · 0.5utilization markers · 0.1abstraction · 0.1transparent pipeline clock-gating · 0.1elastic pipeline clock-gating · 0.1trace-driven simulation · 0.0area modeling · 0.0trace transformation · 0.0empirical modeling · 0.0
YearPublicationVenuePosition
2021 Energy Efficiency Boost in the AI-Infused POWER10 Processor
abstract
We present the novel micro-architectural features, supported by an innovative and novel pre-silicon methodology in the design of POWER10. The resulting projected energy efficiency boost over POWER9 is 2.6x at core level (for SPECint) and up to 3x at socket level. In addition, a new feature supporting inline AI acceleration was added to the POWER ISA and incorporated into the POWER10 processor core design. The resulting boost in SIMD/AI socket performance is projected to be up to 10x for FP32 and 21x for INT8 models of ResNet-50 and BERT-Large. In this paper, we describe the novel methodology deployed and used not only to obtain these efficiency boosts for traditional workloads, but also to infuse AI/ML/HPC capability directly into the POWER10 core.
Brian W. Thompto, Dung Q. Nguyen, José E. Moreira, Ramon Bertran Monfort, Hans M. Jacobson, Richard J. Eickemeyer, Rahul M. Rao, Michael Goulet, Marcy Byers, Christopher J. Gonzalez, Karthik Swaminathan, Nagu R. Dhanwada, Silvia M. Müller, Satish Kumar Sadasivam, Robert K. Montoye, William J. Starke, Christian G. Zoellin, Michael S. Floyd, Jeffrey Stuecheli, Nandhini Chandramoorthy, John-David Wellman, Alper Buyuktosunoglu, Matthias Pflanz, Balaram Sinharoy, Pradip Bose
ISCA6
2011 Abstraction and microarchitecture scaling in early-stage power modeling
abstract
Early-stage, microarchitecture-level power modeling methodologies have been used in industry and academic research for a decade (or more). Such methods use cycle-accurate performance simulators and deduce active power based on utilization markers. A key question faced in this context is: what key utilization metrics to monitor, and how many are needed for accuracy? Is there a systematic way to select the “best” markers? We also pose a key follow-on question: is it possible to perform accurate scaling of an abstracted model to enable exploration of new microarchitecture features? In this paper, we address these particular questions and examine the results for a range of abstraction levels. We highlight innovative insights for intelligent abstraction and microarchitecture scaling, and point out the pitfalls of abstractions that are not based on a systematic methodology or sound theory.
Hans M. Jacobson, Alper Buyuktosunoglu, Pradip Bose, Emrah Acar, Richard J. Eickemeyer
HPCA5
2005 Stretching the Limits of Clock-Gating Efficiency in Server-Class Processors
abstract
Clock-gating has been introduced as the primary means of dynamic power management in recent high-end commercial microprocessors. The temperature drop resulting from active power reduction can result in additional leakage power savings in future processors. In this paper we first examine the realistic benefits and limits of clock-gating in current generation high-performance processors (e.g. of the POWER4/spl trade/ or POWER5/spl trade/ class). We then look beyond classical clock-gating: we examine additional opportunities to avoid unnecessary clocking in real workload executions. In particular, we examine the power reduction benefits of a couple of newly invented schemes called transparent pipeline clock-gating and elastic pipeline clock-gating. Based on our experiences with current designs, we try to bound the practical limits of clock gating efficiency in future microprocessors.
Hans M. Jacobson, Pradip Bose, Alper Buyuktosunoglu, Victor V. Zyuban, Richard J. Eickemeyer, Lee Eisen, John Griswell, Doug Logan, Balaram Sinharoy, Joel M. Tendler
HPCA6
1996 Evaluation of Multithreaded Uniprocessors for Commercial Application Environments
abstract
As memory speeds grow at a considerably slower rate than processor speeds, memory accesses are starting to dominate the execution time of processors, and this will likely continue into the future. This trend will be exacerbated by growing miss rates due to commercial applications, object-oriented programming and micro-kernel based operating systems. We examine the use of coarse-grained multithreading to address this important problem in uniprocessor on-line transaction processing environments where there is a natural, coarse-grained parallelism between the tasks resulting from transactions being executed concurrently, with no application software modifications required. Our results suggest that multithreading can provide significant performance improvements for uniprocessor commercial computing environments.
Richard J. Eickemeyer, Ross E. Johnson, Steven R. Kunkel, Mark S. Squillante, Shiafun Liu
ISCA1
1992 Interlock collapsing ALU for increased instruction-level parallelism
Nadeem Malik, Richard J. Eickemeyer, Stamatis Vassiliadis
MICRO2
1988 Performance Evaluation of On-Chip Register and Cache Organizations
abstract
Several different local memory organizations applicable for single-chip processors are compared. Several cache types-instruction, data, split, unified, stack, and top-of-stack-are considered. These are compared to multiple-register-set architectures to which various caches can also be added. The performance metric of interest is effective access time, since a wide variety of register and cache organizations are used. A model for access time and a model for chip area required for each organization form the basis for comparison. Extensive simulations of several register-memory organizations are presented. Address traces from a VAX-11/780 running systems programs were used in the simulation. The data indicate that for small area (600 bytes), split or unified instruction and data caches are best. Context switch effects were measured and found to be negligible for small- and medium-size caches.>
Richard J. Eickemeyer, Janak H. Patel
ISCA1
1987 Performance Evaluation of Multiple Register Sets
abstract
In this paper a DEC VAX with multiple register sets is evaluated under many differently sized register sets. Both the number of register sets and the number of registers per set were varied. Performance, measured in terms of memory traffic, is compared to that of a standard VAX. Memory traffic is measured from many real program traces on the standard processor and from transformations of the trace for the multiple register set processors. Results are presented for each program; an empirical formula is derived which describes the average program's behavior. A decrease in memory references of approximately 16% can be expected using multiple register sets.
Richard J. Eickemeyer, Janak H. Patel
ISCA1