EDBT 2026 Demo / reviewers in the wild / expert
Thomas R. Puzak
dblp:83/2353
· DBLP profile ↗
7ranked-venue papers
1as first author
0since 2021 · last 2007
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Processor architecture and microarchitecture · 66% Energy-efficient computing · 15% Memory systems · 10% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
pipelining |
0.1 | 3 | 2007 | Pipeline spectroscopy · SIGMETRICS 2007 The Optimum Pipeline Depth for a Microprocessor · ISCA 2002 Branch History Guided Instruction Prefetching · HPCA 2001 |
Processor architecture and microarchitecture › pipelining
pipeline depth |
0.1 | 2 | 2004 | The optimum pipeline depth considering both power and performance · ACM Trans. Archit. Code Optim. 2004 Optimum Power/Performance Pipeline Depth · MICRO 2003 |
Energy-efficient computing
power-performance tradeoff |
0.1 | 2 | 2004 | The optimum pipeline depth considering both power and performance · ACM Trans. Archit. Code Optim. 2004 Optimum Power/Performance Pipeline Depth · MICRO 2003 |
Processor architecture and microarchitecture › pipelining
pipeline depth optimization |
0.1 | 2 | 2004 | The optimum pipeline depth considering both power and performance · ACM Trans. Archit. Code Optim. 2004 The Optimum Pipeline Depth for a Microprocessor · ISCA 2002 |
Processor architecture and microarchitecture › pipelining
pipeline design |
0.0 | 1 | 2004 | The optimum pipeline depth considering both power and performance · ACM Trans. Archit. Code Optim. 2004 |
Performance modeling and evaluation
workload characterization |
0.0 | 3 | 2004 | The optimum pipeline depth considering both power and performance · ACM Trans. Archit. Code Optim. 2004 Optimum Power/Performance Pipeline Depth · MICRO 2003 The Optimum Pipeline Depth for a Microprocessor · ISCA 2002 |
Memory systems
cache management |
0.0 | 1 | 2001 | Branch History Guided Instruction Prefetching · HPCA 2001 |
Memory systems › cache management › instruction cache management
instruction cache miss reduction |
0.0 | 1 | 2001 | Branch History Guided Instruction Prefetching · HPCA 2001 |
Processor architecture and microarchitecture › instruction fetch
instruction prefetching |
0.0 | 1 | 2001 | Branch History Guided Instruction Prefetching · HPCA 2001 |
Processor architecture and microarchitecture › front-end
instruction supply |
0.0 | 1 | 2001 | Branch History Guided Instruction Prefetching · HPCA 2001 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.1analytical modeling · 0.1trace-driven simulation · 0.0branch history correlation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2007 | Pipeline spectroscopyabstractNo abstract available. Thomas R. Puzak, Allan Hartstein, Vijayalakshmi Srinivasan, Philip G. Emma, Arthur Nadas |
SIGMETRICS | 1 |
| 2004 | The optimum pipeline depth considering both power and performanceabstractThe impact of pipeline length on both the power and performance of a microprocessor is explored both by theory and by simulation. A theory is presented for a range of power/performance metrics, BIPS m /W . The theory shows that the more important power is to the metric, the shorter the optimum pipeline length that results. For typical parameters neither BIPS/W nor BIPS 2 /W yield an optimum, i.e., a non-pipelined design is optimal. For BIPS 3 /W the optimum, averaged over all 55 workloads studied, occurs at a 22.5 FO4 design point, a 7 stage pipeline, but this value is highly dependent on the assumed growth in latch count with pipeline depth. As dynamic power grows, the optimal design point shifts to shorter pipelines. Clock gating pushes the optimum to deeper pipelines. Surprisingly, as leakage power grows, the optimum is also found to shift to deeper pipelines. The optimum pipeline depth varies for different classes of workloads: SPEC95 and SPEC2000 integer applications, traditional (legacy) database and on-line transaction processing applications, modern (e. g. web) applications, and floating point applications. Allan Hartstein, Thomas R. Puzak |
ACM Trans. Archit. Code Optim. | 2 |
| 2003 | Optimum Power/Performance Pipeline DepthabstractThe impact of pipeline length on both the power and performance of a microprocessor is explored both theoretically and by simulation. A theory is presented for a wide range of power/performance metrics, BIPS/sup m//W. The theory shows that the more important power is to the metric, the shorter the optimum pipeline length that results. For typical parameters neither BIPS/W nor BIPS/sup 2//W yield an optimum, i.e., a non-pipelined design is optimal. For BIPS/sup 3//W the optimum, averaged over all 55 workloads studied, occurs at a 22.5 FO4 design point, a 7 stage pipeline, but this value is highly dependent on the assumed growth in latch count with pipeline depth. As dynamic power grows, the optimal design point shifts to shorter pipelines. Clock gating pushes the optimum to deeper pipelines. Surprisingly, as leakage power grows, the optimum is also found to shift to deeper pipelines. The optimum pipeline depth varies for different classes of workloads: SPEC95 and SPEC2000 integer applications, traditional (legacy) database and on-line transaction processing applications, modern (e.g. Web) applications, and floating point applications. Allan Hartstein, Thomas R. Puzak |
MICRO | 2 |
| 2002 | The Optimum Pipeline Depth for a MicroprocessorabstractThe impact of pipeline length on the performance of a microprocessor is explored both theoretically and by simulation. An analytical theory is presented that shows two opposing architectural parameters affect the optimal pipeline length: the degree of instruction level parallelism (superscalar) decreases the optimal pipeline length, while the lack of pipeline stalls increases the optimal pipeline length. This theory is tested by analyzing the optimal pipeline length for 35 applications representing three classes of workloads. Trace tapes are collected from SPEC95 and SPEC2000 applications, traditional (legacy) database and online transaction processing (OLTP) applications, and modern applications primarily written in Java and C++. The results show that there is a clear and significant difference in the optimal pipeline length between the SPEC workloads and both the legacy and modern applications. The SPEC applications, written in C, optimize to a shorter pipeline length than the legacy applications, largely written in assembler language, with relatively little overlap in the two distributions. Additionally, the optimal pipeline length distribution for the C++ and Java workloads overlaps with the legacy applications, suggesting similar workload characteristics. These results are explored across a wide range of superscalar processors, both in-order and out-of-order. Allan Hartstein, Thomas R. Puzak |
ISCA | 2 |
| 2001 | Branch History Guided Instruction PrefetchingabstractInstruction cache misses stall the fetch stage of the processor pipeline and hence affect instruction supply to the processor. Instruction prefetching has been proposed as a mechanism to reduce instruction cache (I-cache) misses. However, a prefetch is effective only if accurate and initiated sufficiently early to cover the miss penalty. This paper presents a new hardware-based instruction prefetching mechanism, Branch History Guided Prefetching (BHGP), to improve the timeliness of instruction prefetches. BHGP correlates the execution of a branch instruction with I-cache misses and uses branch instructions to trigger prefetches of instructions that occur (N-1) branches later in the program execution, for a given N>1. Evaluations on commercial applications, windows-NT applications, and some CPU2000 applications show an average reduction of 66% in miss rate over all applications. BHGP improved the IPC bp 12 to 14% for the CPU2000 applications studied; on average 80% of the BHGP prefetches arrived in cache before their next use, even on a 4-wide issue machine with a 15 cycle L2 access penalty. Vijayalakshmi Srinivasan, Edward S. Davidson, Gary S. Tyson, Mark J. Charney, Thomas R. Puzak |
HPCA | 5 |
| 2001 | Filtering Superfluous Prefetches Using Density VectorsabstractA previous evaluation of scheduled region prefetching showed that this technique eliminates the bulk of main-memory stall time for applications with spatial locality. The downside to that aggressive prefetching scheme is that, even when it successfully improves performance, it increases enormously the amount of superfluous memory traffic generated by a program. We measure the predictability of spatial locality using density vectors, bit vectors that track the block-level access pattern within a region of memory. We evaluate a number of policies that use density vector information to filter out prefetches that are unlikely to be useful. We show that, across our benchmarks, an average of 70% of useless prefetches can be eliminated with virtually no overall performance loss from reduced coverage. Thanks to the increase in prefetch accuracy, a few benchmarks show performance improvements as high as 35% over the base region prefetching scheme. Wei-Fen Lin, Steven K. Reinhardt, Doug Burger, Thomas R. Puzak |
ICCD | 4 |
| 1992 | Contrasting instruction-fetch time and instruction-decode time branch prediction mechanisms: Achieving synergy through their cooperative operation
David R. Kaeli, Philip G. Emma, Joshua W. Knight, Thomas R. Puzak |
Microprocess. Microprogramming | 4 |