Philip G. Emma

dblp:39/1637 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 4 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Processor architecture and microarchitecture · 25% Energy-efficient computing · 21% Memory systems · 17%

Topics — the 28 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
data redundancy
0.312018
Duplicon Cache: Mitigating Off-Chip Memory Bank and Bank Group Conflicts Via Data Duplication · MICRO 2018
Memory systems
memory controller
0.312018
Duplicon Cache: Mitigating Off-Chip Memory Bank and Bank Group Conflicts Via Data Duplication · MICRO 2018
Integrated circuit design
3d integration
0.222014
3D stacking of high-performance processors · HPCA 2014
Interconnects in the Third Dimension: Design Challenges for 3D ICs · DAC 2007
Processor architecture and microarchitecture › multi-chip architecture
3d stacking
0.212014
3D stacking of high-performance processors · HPCA 2014
Energy-efficient computing
power density
0.212014
3D stacking of high-performance processors · HPCA 2014
Energy-efficient computing
thermal management
0.212014
3D stacking of high-performance processors · HPCA 2014
Processor architecture and microarchitecture › pipelining
pipeline depth optimization
0.122004
Integrated Analysis of Power and Performance for Pipelined Microprocessors · IEEE Trans. Computers 2004
Optimizing pipelines for power and performance · MICRO 2002
Processor architecture and microarchitecture › pipelining
pipeline design
0.122004
Integrated Analysis of Power and Performance for Pipelined Microprocessors · IEEE Trans. Computers 2004
Optimizing pipelines for power and performance · MICRO 2002
Energy-efficient computing
power-performance tradeoff
0.122004
Integrated Analysis of Power and Performance for Pipelined Microprocessors · IEEE Trans. Computers 2004
Optimizing pipelines for power and performance · MICRO 2002
Processor architecture and microarchitecture
pipelining
0.122007
Pipeline spectroscopy · SIGMETRICS 2007
Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline Performance · IEEE Trans. Computers 1987
Interconnection networks and networks-on-chip › die-to-die interconnect
3d interconnect
0.112007
Interconnects in the Third Dimension: Design Challenges for 3D ICs · DAC 2007
Electronic design automation
physical design
0.112007
Interconnects in the Third Dimension: Design Challenges for 3D ICs · DAC 2007
Integrated circuit design › 3d integration
through-silicon via
0.112007
Interconnects in the Third Dimension: Design Challenges for 3D ICs · DAC 2007
Memory systems
on-chip memory
0.112006
Industrial Perspectives: The Next Roadblocks in SOC Evolution: On-Chip Storage Capacity and Off-Chip Bandwidth · HPCA 2006
Processor architecture and microarchitecture › microprocessor
high-performance processors
0.112014
3D stacking of high-performance processors · HPCA 2014
Processor architecture and microarchitecture › pipelining
pipelined processor
0.012004
Integrated Analysis of Power and Performance for Pipelined Microprocessors · IEEE Trans. Computers 2004
Energy-efficient computing
power-performance modeling
0.012002
Optimizing pipelines for power and performance · MICRO 2002
Processor architecture and microarchitecture
branch prediction
0.021997
Improving the Accuracy of History Based Branch Prediction · IEEE Trans. Computers 1997
Branch History Table Prediction of Moving Target Branches due to Subroutine Returns · ISCA 1991
Integrated circuit design
system-on-chip
0.012006
Industrial Perspectives: The Next Roadblocks in SOC Evolution: On-Chip Storage Capacity and Off-Chip Bandwidth · HPCA 2006
Integrated circuit design
technology scaling
0.012006
The End of Scaling? Revolutions in Technology and Microarchitecture as We Pass the 90 Nanometer Node · ISCA 2006
Processor architecture and microarchitecture › branch prediction
history-based branch prediction
0.011997
Improving the Accuracy of History Based Branch Prediction · IEEE Trans. Computers 1997
Performance modeling and evaluation
analytical modeling
0.012002
Optimizing pipelines for power and performance · MICRO 2002
Processor architecture and microarchitecture › branch prediction
branch target buffer
0.011997
Improving the Accuracy of History Based Branch Prediction · IEEE Trans. Computers 1997
Processor architecture and microarchitecture
data dependence
0.011987
Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline Performance · IEEE Trans. Computers 1987
Processor architecture and microarchitecture › pipelining
delayed branch
0.011987
Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline Performance · IEEE Trans. Computers 1987
Performance modeling and evaluation › tracing
execution trace
0.011987
Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline Performance · IEEE Trans. Computers 1987
Processor architecture and microarchitecture › pipelining
pipeline performance
0.011987
Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline Performance · IEEE Trans. Computers 1987
Performance modeling and evaluation
workload characterization
0.011987
Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline Performance · IEEE Trans. Computers 1987

Methods — techniques the papers use, named apart from their topics

duplicon cache · 0.3data duplication · 0.3simulation · 0.1analytical modeling · 0.1sensitivity analysis · 0.0execution-driven modeling · 0.0trace-driven simulation · 0.0trace reduction · 0.0data dependency graph · 0.0
YearPublicationVenuePosition
2018 Duplicon Cache: Mitigating Off-Chip Memory Bank and Bank Group Conflicts Via Data Duplication
abstract
Bank and bank group conflicts are major performance bottlenecks for memory intensive workloads. Idealized experiments show removing bank and bank group conflicts collectively can improve performance by up to 37.5% and by 22.5% on average for our mix of multi-programmed memory intensive workloads. We propose the Duplicon Cache to mitigate bank and bank group conflict penalties by duplicating select lines of data to an alternative bank group, giving the memory controller the freedom to source the data from the bank group which avoids conflicts. The Duplicon Cache is entirely implemented in the memory controller and does not require changes to commodity memory. We identify and address the main challenges associated with duplication: 1) tracking duplicated data efficiently, 2) identifying which data to duplicate, and 3) replacing stale duplicated data while protecting useful ones. Our evaluations show the Duplicon Cache configured with 128MB of storage (out of 16GB of main memory) improves performance by 8.3% while reducing energy by 5.6%.
Ben Lin, Michael B. Healy, Rustam Miftakhutdinov, Philip G. Emma, Yale N. Patt
MICRO4
2014 3D stacking of high-performance processors
abstract
In most 3D work to date, people have looked at two situations: 1) a case in which power density is not a problem, and the parts of a processor and/or entire processors can be stacked atop each other, and 2) a case in which power density is limited, and storage is stacked atop processors. In this paper, we consider the case in which power density is a limitation, yet we stack processors atop processors. We also will discuss some of the physical limitations today that render many of the good ideas presented in other work impractical, and what would be required in the technology to make them feasible. In the high-performance regime, circuits are not designed to be “power efficient;” they're designed to be fast. In power-efficient design, the speed and power of a processor should be nearly proportional. In the high-performance regime, the frequency is (ever progressingly) sublinear in power. Thus, when the power density is constrained - as it is in high-performance machines, there may be opportunities to selectively exploit parallelism in workloads by running processor-on-processor systems at the same power, yet at much greater than half speed.
Philip G. Emma, Alper Buyuktosunoglu, Michael B. Healy, Krishnan Kailas, Valentin Puente, Roy Yu, Allan Hartstein, Pradip Bose, Jaime H. Moreno, Eren Kursun
HPCA1
2007 Interconnects in the Third Dimension: Design Challenges for 3D ICs
abstract
Despite generation upon generation of scaling, computer chips have until now remained essentially 2-dimensional. Improvements in on-chip wire delay and in the maximum number of I/O per chip have not been able to keep up with transistor performance growth; it has become steadily harder to hide the discrepancy. 3D chip technologies come in a number of flavors, but are expected to enable the extension of CMOS performance. Designing in three dimensions, however, forces the industry to look at formerly-two- dimensional integration issues quite differently, and requires the re-fitting of multiple existing EDA capabilities.
Kerry Bernstein, Paul S. Andry, Jerome Cann, Philip G. Emma, David Greenberg, Wilfried Haensch, Mike Ignatowski, Steven J. Koester, John Magerlein, Ruchir Puri, Albert M. Young
DAC4
2007 Pipeline spectroscopy
abstract
No abstract available.
Thomas R. Puzak, Allan Hartstein, Vijayalakshmi Srinivasan, Philip G. Emma, Arthur Nadas
SIGMETRICS4
2006 Industrial Perspectives: The Next Roadblocks in SOC Evolution: On-Chip Storage Capacity and Off-Chip Bandwidth
Philip G. Emma
HPCA1
2006 The End of Scaling? Revolutions in Technology and Microarchitecture as We Pass the 90 Nanometer Node
abstract
The document was not made available for publication as part of the conference proceedings.
Philip G. Emma
ISCA1
2004 Integrated Analysis of Power and Performance for Pipelined Microprocessors
abstract
Choosing the pipeline depth of a microprocessor is one of the most critical design decisions that an architect must make in the concept phase of a microprocessor design. To be successful in today's cost/performance marketplace, modern CPU designs must effectively balance both performance and power dissipation. The choice of pipeline depth and target clock frequency has a critical impact on both of these metrics. We describe an optimization methodology based on both analytical models and detailed simulations for power and performance as a function of pipeline depth. Our results for a set of SPEC2000 applications show that, when both power and performance are considered for optimization, the optimal clock period is around 18 FO4. We also provide a detailed sensitivity analysis of the optimal pipeline depth against key assumptions of our energy models. Finally, we discuss the potential risks in design quality for overly aggressive or conservative choices of pipeline depth.
Victor V. Zyuban, David Brooks 0001, Vijayalakshmi Srinivasan, Michael Gschwind, Pradip Bose, Philip N. Strenski, Philip G. Emma
IEEE Trans. Computers7
2002 Optimizing pipelines for power and performance
abstract
During the concept phase and definition of next generation high-end processors, power and performance will need to be weighted appropriately to deliver competitive cost/performance. It is not enough to adopt a CPI-centric view alone in early-stage definition studies. One of the fundamental issues confronting the architect at this stage is the choice of pipeline depth and target frequency. In this paper we present an optimization methodology that starts with an analytical power-performance model to derive optimal pipeline depth for a superscalar processor. The results are validated and further refined using detailed simulation based analysis. As part of the power-modeling methodology, we have developed equations that model the variation of energy as a function of pipeline depth. Our results using a set of SPEC2000 applications show that when both power and performance are considered for optimization, the optimal clock period is around 18 FO4. We also provide a detailed sensitivity analysis of the optimal pipeline depth against key assumptions of these energy models.
Vijayalakshmi Srinivasan, David Brooks 0001, Michael Gschwind, Pradip Bose, Victor V. Zyuban, Philip N. Strenski, Philip G. Emma
MICRO7
1997 Improving the Accuracy of History Based Branch Prediction
abstract
In this paper, we present mechanisms that improve the accuracy and performance of history-based branch prediction. By studying the characteristics of the decision structures present in high-level languages, two mechanisms are proposed that reduce the number of wrong predictions made by a branch target buffer (BTB). Execution-driven modeling is used to evaluate the improvement in branch prediction accuracy, as well as the reduction in overall program execution.
David R. Kaeli, Philip G. Emma
IEEE Trans. Computers2
1992 Contrasting instruction-fetch time and instruction-decode time branch prediction mechanisms: Achieving synergy through their cooperative operation
David R. Kaeli, Philip G. Emma, Joshua W. Knight, Thomas R. Puzak
Microprocess. Microprogramming2
1991 Branch History Table Prediction of Moving Target Branches due to Subroutine Returns
abstract
Ideally, a pipeline processor can run at a rate that is limited by its slowest stage.Branches in the instruction stream disrupt the pipeIine, and reduce processor performance to well below ideal.Since workloads contain a high percentage of taken branches, techniques are needed to reduce or eliminate thk degradation.A Branch History Table (BHT) stores past action and target for branches, and predicts that future behavior will repeat.Although past action is a good indicator of future action, the subroutine CALL/RETURN paradigm makes correct prediction of the branch target dlfflcult.We propose a new stack mechanism for reducing this type of mispredlction.Using traces of the SPEC benchmark suite running on an RS/6000, we provide an analysis of the performance enhancements possible using a BHT.We show that the proposed mechanism can reduce the number of branch wrong guesses by 18.2°/0 on average.
David R. Kaeli, Philip G. Emma
ISCA2
1987 Characterization of Branch and Data Dependencies in Programs for Evaluating Pipeline Performance
abstract
The nature by which branches and data dependencies generate delays that degrade pipeline performance is investigated in this paper. We show that for the general execution trace, few specific delays can be considered in isolation; rather, the magnitude of any specific delay may depend on the relative proximity of other delays. This phenomenon can make the task of accurately characterizing a trace tape with simple statistics intractable. We present a set of trace reductions that facilitates this task by simplifying the corresponding data-dependency graph. The reductions operate on multiple data-dependency arcs and branches in conjunction; those arcs whose performance implications are redundant with respect to the dependency graph are identified, and eliminated from the graph. We show that the reduced graph can be accurately characterized by simple statistics. We use these statistics to show that as the length of a pipeline increases, the performance degradation due to data dependencies and branches increases monotonically. However, lengthening the pipeline may correspond to decreasing the cycle time of the pipeline. These two opposing effects are used in conjunction to derive an equation for optimal pipeline length for a given trace tape. The optimal pipeline length is shown to be characterized by n = √γα where γ is the ratio of overall circuit delay to latching overhead, and a is a function of the trace statistics that accounts for the delays induced by data dependencies and branches.
Philip G. Emma, Edward S. Davidson
IEEE Trans. Computers1