EDBT 2026 Demo / reviewers in the wild / expert
Philip G. Emma
dblp:39/1637
· DBLP profile ↗
12ranked-venue papers
4as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Processor architecture and microarchitecture · 25% Energy-efficient computing · 21% Memory systems · 17% |
Topics — the 28 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
data redundancy |
0.3 | 1 | 2018 | Duplicon Cache: Mitigating Off-Chip Memory Bank and Bank Group Conflicts Via Data Duplication · MICRO 2018 |
Memory systems
memory controller |
0.3 | 1 | 2018 | Duplicon Cache: Mitigating Off-Chip Memory Bank and Bank Group Conflicts Via Data Duplication · MICRO 2018 |
Integrated circuit design
3d integration |
0.2 | 2 | 2014 | 3D stacking of high-performance processors · HPCA 2014 Interconnects in the Third Dimension: Design Challenges for 3D ICs · DAC 2007 |
Processor architecture and microarchitecture › multi-chip architecture
3d stacking |
0.2 | 1 | 2014 | 3D stacking of high-performance processors · HPCA 2014 |
Energy-efficient computing
power density |
0.2 | 1 | 2014 | 3D stacking of high-performance processors · HPCA 2014 |
Energy-efficient computing
thermal management |
0.2 | 1 | 2014 | 3D stacking of high-performance processors · HPCA 2014 |
Processor architecture and microarchitecture › pipelining
pipeline depth optimization |
0.1 | 2 | 2004 | Integrated Analysis of Power and Performance for Pipelined Microprocessors · IEEE Trans. Computers 2004 Optimizing pipelines for power and performance · MICRO 2002 |
Processor architecture and microarchitecture › pipelining
pipeline design |
0.1 | 2 | 2004 | Integrated Analysis of Power and Performance for Pipelined Microprocessors · IEEE Trans. Computers 2004 Optimizing pipelines for power and performance · MICRO 2002 |
Energy-efficient computing
power-performance tradeoff |
0.1 | 2 | 2004 | Integrated Analysis of Power and Performance for Pipelined Microprocessors · IEEE Trans. Computers 2004 Optimizing pipelines for power and performance · MICRO 2002 |
Processor architecture and microarchitecture
pipelining |
0.1 | 2 | 2007 | Pipeline spectroscopy · SIGMETRICS 2007 Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline Performance · IEEE Trans. Computers 1987 |
Interconnection networks and networks-on-chip › die-to-die interconnect
3d interconnect |
0.1 | 1 | 2007 | Interconnects in the Third Dimension: Design Challenges for 3D ICs · DAC 2007 |
Electronic design automation
physical design |
0.1 | 1 | 2007 | Interconnects in the Third Dimension: Design Challenges for 3D ICs · DAC 2007 |
Integrated circuit design › 3d integration
through-silicon via |
0.1 | 1 | 2007 | Interconnects in the Third Dimension: Design Challenges for 3D ICs · DAC 2007 |
Memory systems
on-chip memory |
0.1 | 1 | 2006 | Industrial Perspectives: The Next Roadblocks in SOC Evolution: On-Chip Storage Capacity and Off-Chip Bandwidth · HPCA 2006 |
Processor architecture and microarchitecture › microprocessor
high-performance processors |
0.1 | 1 | 2014 | 3D stacking of high-performance processors · HPCA 2014 |
Processor architecture and microarchitecture › pipelining
pipelined processor |
0.0 | 1 | 2004 | Integrated Analysis of Power and Performance for Pipelined Microprocessors · IEEE Trans. Computers 2004 |
Energy-efficient computing
power-performance modeling |
0.0 | 1 | 2002 | Optimizing pipelines for power and performance · MICRO 2002 |
Processor architecture and microarchitecture
branch prediction |
0.0 | 2 | 1997 | Improving the Accuracy of History Based Branch Prediction · IEEE Trans. Computers 1997 Branch History Table Prediction of Moving Target Branches due to Subroutine Returns · ISCA 1991 |
Integrated circuit design
system-on-chip |
0.0 | 1 | 2006 | Industrial Perspectives: The Next Roadblocks in SOC Evolution: On-Chip Storage Capacity and Off-Chip Bandwidth · HPCA 2006 |
Integrated circuit design
technology scaling |
0.0 | 1 | 2006 | The End of Scaling? Revolutions in Technology and Microarchitecture as We Pass the 90 Nanometer Node · ISCA 2006 |
Processor architecture and microarchitecture › branch prediction
history-based branch prediction |
0.0 | 1 | 1997 | Improving the Accuracy of History Based Branch Prediction · IEEE Trans. Computers 1997 |
Performance modeling and evaluation
analytical modeling |
0.0 | 1 | 2002 | Optimizing pipelines for power and performance · MICRO 2002 |
Processor architecture and microarchitecture › branch prediction
branch target buffer |
0.0 | 1 | 1997 | Improving the Accuracy of History Based Branch Prediction · IEEE Trans. Computers 1997 |
Processor architecture and microarchitecture
data dependence |
0.0 | 1 | 1987 | Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline Performance · IEEE Trans. Computers 1987 |
Processor architecture and microarchitecture › pipelining
delayed branch |
0.0 | 1 | 1987 | Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline Performance · IEEE Trans. Computers 1987 |
Performance modeling and evaluation › tracing
execution trace |
0.0 | 1 | 1987 | Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline Performance · IEEE Trans. Computers 1987 |
Processor architecture and microarchitecture › pipelining
pipeline performance |
0.0 | 1 | 1987 | Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline Performance · IEEE Trans. Computers 1987 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 1987 | Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline Performance · IEEE Trans. Computers 1987 |
Methods — techniques the papers use, named apart from their topics
duplicon cache · 0.3data duplication · 0.3simulation · 0.1analytical modeling · 0.1sensitivity analysis · 0.0execution-driven modeling · 0.0trace-driven simulation · 0.0trace reduction · 0.0data dependency graph · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Duplicon Cache: Mitigating Off-Chip Memory Bank and Bank Group Conflicts Via Data DuplicationabstractBank and bank group conflicts are major performance bottlenecks for memory intensive workloads. Idealized experiments show removing bank and bank group conflicts collectively can improve performance by up to 37.5% and by 22.5% on average for our mix of multi-programmed memory intensive workloads. We propose the Duplicon Cache to mitigate bank and bank group conflict penalties by duplicating select lines of data to an alternative bank group, giving the memory controller the freedom to source the data from the bank group which avoids conflicts. The Duplicon Cache is entirely implemented in the memory controller and does not require changes to commodity memory. We identify and address the main challenges associated with duplication: 1) tracking duplicated data efficiently, 2) identifying which data to duplicate, and 3) replacing stale duplicated data while protecting useful ones. Our evaluations show the Duplicon Cache configured with 128MB of storage (out of 16GB of main memory) improves performance by 8.3% while reducing energy by 5.6%. Ben Lin, Michael B. Healy, Rustam Miftakhutdinov, Philip G. Emma, Yale N. Patt |
MICRO | 4 |
| 2014 | 3D stacking of high-performance processorsabstractIn most 3D work to date, people have looked at two situations: 1) a case in which power density is not a problem, and the parts of a processor and/or entire processors can be stacked atop each other, and 2) a case in which power density is limited, and storage is stacked atop processors. In this paper, we consider the case in which power density is a limitation, yet we stack processors atop processors. We also will discuss some of the physical limitations today that render many of the good ideas presented in other work impractical, and what would be required in the technology to make them feasible. In the high-performance regime, circuits are not designed to be “power efficient;” they're designed to be fast. In power-efficient design, the speed and power of a processor should be nearly proportional. In the high-performance regime, the frequency is (ever progressingly) sublinear in power. Thus, when the power density is constrained - as it is in high-performance machines, there may be opportunities to selectively exploit parallelism in workloads by running processor-on-processor systems at the same power, yet at much greater than half speed. Philip G. Emma, Alper Buyuktosunoglu, Michael B. Healy, Krishnan Kailas, Valentin Puente, Roy Yu, Allan Hartstein, Pradip Bose, Jaime H. Moreno, Eren Kursun |
HPCA | 1 |
| 2007 | Interconnects in the Third Dimension: Design Challenges for 3D ICsabstractDespite generation upon generation of scaling, computer chips have until now remained essentially 2-dimensional. Improvements in on-chip wire delay and in the maximum number of I/O per chip have not been able to keep up with transistor performance growth; it has become steadily harder to hide the discrepancy. 3D chip technologies come in a number of flavors, but are expected to enable the extension of CMOS performance. Designing in three dimensions, however, forces the industry to look at formerly-two- dimensional integration issues quite differently, and requires the re-fitting of multiple existing EDA capabilities. Kerry Bernstein, Paul S. Andry, Jerome Cann, Philip G. Emma, David Greenberg, Wilfried Haensch, Mike Ignatowski, Steven J. Koester, John Magerlein, Ruchir Puri, Albert M. Young |
DAC | 4 |
| 2007 | Pipeline spectroscopyabstractNo abstract available. Thomas R. Puzak, Allan Hartstein, Vijayalakshmi Srinivasan, Philip G. Emma, Arthur Nadas |
SIGMETRICS | 4 |
| 2006 | Industrial Perspectives: The Next Roadblocks in SOC Evolution: On-Chip Storage Capacity and Off-Chip Bandwidth
Philip G. Emma |
HPCA | 1 |
| 2006 | The End of Scaling? Revolutions in Technology and Microarchitecture as We Pass the 90 Nanometer NodeabstractThe document was not made available for publication as part of the conference proceedings. Philip G. Emma |
ISCA | 1 |
| 2004 | Integrated Analysis of Power and Performance for Pipelined MicroprocessorsabstractChoosing the pipeline depth of a microprocessor is one of the most critical design decisions that an architect must make in the concept phase of a microprocessor design. To be successful in today's cost/performance marketplace, modern CPU designs must effectively balance both performance and power dissipation. The choice of pipeline depth and target clock frequency has a critical impact on both of these metrics. We describe an optimization methodology based on both analytical models and detailed simulations for power and performance as a function of pipeline depth. Our results for a set of SPEC2000 applications show that, when both power and performance are considered for optimization, the optimal clock period is around 18 FO4. We also provide a detailed sensitivity analysis of the optimal pipeline depth against key assumptions of our energy models. Finally, we discuss the potential risks in design quality for overly aggressive or conservative choices of pipeline depth. Victor V. Zyuban, David Brooks 0001, Vijayalakshmi Srinivasan, Michael Gschwind, Pradip Bose, Philip N. Strenski, Philip G. Emma |
IEEE Trans. Computers | 7 |
| 2002 | Optimizing pipelines for power and performanceabstractDuring the concept phase and definition of next generation high-end processors, power and performance will need to be weighted appropriately to deliver competitive cost/performance. It is not enough to adopt a CPI-centric view alone in early-stage definition studies. One of the fundamental issues confronting the architect at this stage is the choice of pipeline depth and target frequency. In this paper we present an optimization methodology that starts with an analytical power-performance model to derive optimal pipeline depth for a superscalar processor. The results are validated and further refined using detailed simulation based analysis. As part of the power-modeling methodology, we have developed equations that model the variation of energy as a function of pipeline depth. Our results using a set of SPEC2000 applications show that when both power and performance are considered for optimization, the optimal clock period is around 18 FO4. We also provide a detailed sensitivity analysis of the optimal pipeline depth against key assumptions of these energy models. Vijayalakshmi Srinivasan, David Brooks 0001, Michael Gschwind, Pradip Bose, Victor V. Zyuban, Philip N. Strenski, Philip G. Emma |
MICRO | 7 |
| 1997 | Improving the Accuracy of History Based Branch PredictionabstractIn this paper, we present mechanisms that improve the accuracy and performance of history-based branch prediction. By studying the characteristics of the decision structures present in high-level languages, two mechanisms are proposed that reduce the number of wrong predictions made by a branch target buffer (BTB). Execution-driven modeling is used to evaluate the improvement in branch prediction accuracy, as well as the reduction in overall program execution. David R. Kaeli, Philip G. Emma |
IEEE Trans. Computers | 2 |
| 1992 | Contrasting instruction-fetch time and instruction-decode time branch prediction mechanisms: Achieving synergy through their cooperative operation
David R. Kaeli, Philip G. Emma, Joshua W. Knight, Thomas R. Puzak |
Microprocess. Microprogramming | 2 |
| 1991 | Branch History Table Prediction of Moving Target Branches due to Subroutine ReturnsabstractIdeally, a pipeline processor can run at a rate that is limited by its slowest stage.Branches in the instruction stream disrupt the pipeIine, and reduce processor performance to well below ideal.Since workloads contain a high percentage of taken branches, techniques are needed to reduce or eliminate thk degradation.A Branch History Table (BHT) stores past action and target for branches, and predicts that future behavior will repeat.Although past action is a good indicator of future action, the subroutine CALL/RETURN paradigm makes correct prediction of the branch target dlfflcult.We propose a new stack mechanism for reducing this type of mispredlction.Using traces of the SPEC benchmark suite running on an RS/6000, we provide an analysis of the performance enhancements possible using a BHT.We show that the proposed mechanism can reduce the number of branch wrong guesses by 18.2°/0 on average. David R. Kaeli, Philip G. Emma |
ISCA | 2 |
| 1987 | Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline PerformanceabstractThe nature by which branches and data dependencies generate delays that degrade pipeline performance is investigated in this paper. We show that for the general execution trace, few specific delays can be considered in isolation; rather, the magnitude of any specific delay may depend on the relative proximity of other delays. This phenomenon can make the task of accurately characterizing a trace tape with simple statistics intractable. We present a set of trace reductions that facilitates this task by simplifying the corresponding data-dependency graph. The reductions operate on multiple data-dependency arcs and branches in conjunction; those arcs whose performance implications are redundant with respect to the dependency graph are identified, and eliminated from the graph. We show that the reduced graph can be accurately characterized by simple statistics. We use these statistics to show that as the length of a pipeline increases, the performance degradation due to data dependencies and branches increases monotonically. However, lengthening the pipeline may correspond to decreasing the cycle time of the pipeline. These two opposing effects are used in conjunction to derive an equation for optimal pipeline length for a given trace tape. The optimal pipeline length is shown to be characterized by n = √γα where γ is the ratio of overall circuit delay to latching overhead, and a is a function of the trace statistics that accounts for the delays induced by data dependencies and branches. Philip G. Emma, Edward S. Davidson |
IEEE Trans. Computers | 1 |