Peter M. Kogge

dblp:30/6540 · also Peter Michael Kogge · DBLP profile ↗
← Back
46ranked-venue papers
9as first author
2since 2021 · last 2026
0000-0002-3329-547XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 42 · 9 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
16 papers
Processor architecture and microarchitecture · 28% Performance modeling and evaluation · 24% Parallel and multicore computing · 14%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 30 heaviest of 51, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
neural operator
0.912025
Building Flexible Physics-Informed Neural Networks with Fast Fourier Transform Analysis · HPDC 2025
Machine learning › Deep learning architectures and training
physics-informed neural network
0.912025
Building Flexible Physics-Informed Neural Networks with Fast Fourier Transform Analysis · HPDC 2025
Computational science and engineering
differential equations
0.312025
Building Flexible Physics-Informed Neural Networks with Fast Fourier Transform Analysis · HPDC 2025
Processor architecture and microarchitecture › multithreading
chip multithreading
0.112011
Lightweight Chip Multi-Threading (LCMT): Maximizing Fine-Grained Parallelism On-Chip · IEEE Trans. Parallel Distributed Syst. 2011
Parallel and multicore computing › parallelization strategies
fine-grained parallelism
0.112011
Lightweight Chip Multi-Threading (LCMT): Maximizing Fine-Grained Parallelism On-Chip · IEEE Trans. Parallel Distributed Syst. 2011
Processor architecture and microarchitecture
multithreading
0.112011
Lightweight Chip Multi-Threading (LCMT): Maximizing Fine-Grained Parallelism On-Chip · IEEE Trans. Parallel Distributed Syst. 2011
Parallel and multicore computing
thread-level parallelism
0.112011
Lightweight Chip Multi-Threading (LCMT): Maximizing Fine-Grained Parallelism On-Chip · IEEE Trans. Parallel Distributed Syst. 2011
Performance modeling and evaluation
benchmarking
0.122011
On the Memory Access Patterns of Supercomputer Applications: Benchmark Selection and Its Implications · IEEE Trans. Computers 2007
Using the TOP500 to trace and project technology and architecture trends · SC 2011
Processor architecture and microarchitecture
multicore design
0.122011
Multi-core issues - Multi-Core for HPC: breakthrough or breakdown? · SC 2006
Lightweight Chip Multi-Threading (LCMT): Maximizing Fine-Grained Parallelism On-Chip · IEEE Trans. Parallel Distributed Syst. 2011
Integrated circuit design
digital circuit design
0.112008
Design of a mask-programmable memory/multiplier array using G4-FET technology · DAC 2008
Emerging computing paradigms › field-coupled nanocomputing
quantum-dot cellular automata
0.122004
Quantum-Dot Cellular Automata (QCA) circuit partitioning: problem modeling and solutions · DAC 2004
A design of and design tools for a novel quantum dot based microprocessor · DAC 2000
Performance modeling and evaluation › benchmarking
benchmark selection
0.112007
On the Memory Access Patterns of Supercomputer Applications: Benchmark Selection and Its Implications · IEEE Trans. Computers 2007
Memory systems
memory access patterns
0.112007
On the Memory Access Patterns of Supercomputer Applications: Benchmark Selection and Its Implications · IEEE Trans. Computers 2007
Performance modeling and evaluation
workload characterization
0.112007
On the Memory Access Patterns of Supercomputer Applications: Benchmark Selection and Its Implications · IEEE Trans. Computers 2007
Performance modeling and evaluation › simulation
architectural simulation
0.112006
Poster reception - The structural simulation toolkit: exploring novel architectures · SC 2006
Performance modeling and evaluation
performance prediction
0.112006
M06 - Issues for the future of supercomputing: impact of Moore's law and architecture on application performance · SC 2006
Performance modeling and evaluation
simulation
0.112006
Poster reception - The structural simulation toolkit: exploring novel architectures · SC 2006
Electronic design automation › physical design
circuit partitioning
0.012004
Quantum-Dot Cellular Automata (QCA) circuit partitioning: problem modeling and solutions · DAC 2004
Electronic design automation
physical design
0.012004
Quantum-Dot Cellular Automata (QCA) circuit partitioning: problem modeling and solutions · DAC 2004
Processor architecture and microarchitecture
pipelining
0.022001
Exploring and exploiting wire-level pipelining in emerging technologies · ISCA 2001
Maximal Rate Pipelined Solutions to Recurrance Problems · ISCA 1973
Energy-efficient computing › low-power design
power optimization
0.012001
Inherently Lower-Power High-Performance Superscalar Architectures · IEEE Trans. Computers 2001
Emerging computing paradigms › cellular automata
quantum cellular automata
0.012001
Exploring and exploiting wire-level pipelining in emerging technologies · ISCA 2001
Processor architecture and microarchitecture
superscalar processor
0.012001
Inherently Lower-Power High-Performance Superscalar Architectures · IEEE Trans. Computers 2001
Electronic design automation › physical design › interconnect optimization
wire pipelining
0.012001
Exploring and exploiting wire-level pipelining in emerging technologies · ISCA 2001
Computational geometry › curve representation
curve approximation
0.012001
Polygonal path approximation with angle constraints · SODA 2001
Computational geometry › curve representation › curve approximation
curve simplification
0.012001
Polygonal path approximation with angle constraints · SODA 2001
Graph algorithms and graph theory › graph algorithms › path problems
path approximation
0.012001
Polygonal path approximation with angle constraints · SODA 2001
Memory systems
DRAM
0.012008
Design of a mask-programmable memory/multiplier array using G4-FET technology · DAC 2008
Memory systems
memory bandwidth
0.011999
Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999
Memory systems
processing-in-memory
0.011999
Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999

Methods — techniques the papers use, named apart from their topics

fourier layer neural operator · 1.7fast fourier transform · 1.7trend analysis · 0.1synchronization · 0.1architectural support for lightweight threads · 0.1integer linear programming · 0.1heuristic algorithm · 0.1G4-FET · 0.1memory locality analysis · 0.1moore's law analysis · 0.1discrete-event simulation · 0.1simulation · 0.1dynamic programming · 0.0
YearPublicationVenuePosition
2026 Brief Announcement: Energy-Time Trajectories: A Tool to Understand Complex Parallel Efficiency
Peter M. Kogge
SPAA1
2025 Building Flexible Physics-Informed Neural Networks with Fast Fourier Transform Analysis
abstract
Physics-Informed Neural Networks (PINNs) allow incorporating differential equations to model a system's behavior, and are well able to fit easily differentiable solutions with low-frequency components. They, however, fall short when attempting to learn high-frequency regions or boundary layers. This paper addresses one approach to fixing this by the addition of a Fourier Layer Neural Operator (FNO) to improve accuracy by better fitting regions with high frequency oscillations. We use as an example an underdamped RLC circuit, which resembles that of a boundary layer. The PINN architecture is built to be generalized over other applications, making it a flexible design for further expansion, especially for multiple layers of GPUs.
Reem Shehayib, Jayden Parker Vap, Peter M. Kogge
HPDC3
2018 Optimizing for KNL Usage Modes When Data Doesn't Fit in MCDRAM
abstract
Technologies such as Multi-Channel DRAM (MCDRAM) or High Bandwidth Memory (HBM) provide significantly more bandwidth than conventional memory. This trend has raised questions about how applications should manage data transfers between levels. This paper focuses on evaluating different usage modes of the MCDRAM in Intel Knights Landing (KNL) manycore processors. We evaluate these usage modes with a sorting kernel and a sorting-based streaming benchmark. We develop a performance model for the benchmark and use experimental evidence to demonstrate the correctness of the model. The model projects near-optimal numbers of copy threads for memory bandwidth bound computations. We demonstrate on KNL up to a 1.9X speedup for sort when the problem does not fit in MCDRAM over an OpenMP GNU sort that does not use MCDRAM.
Neil Butcher, Stephen Olivier, Jonathan W. Berry, Simon D. Hammond, Peter M. Kogge
ICPP5
2014 Reading the Tea-Leaves: How Architecture Has Evolved at the High End
abstract
Summary form only given. The 2008 DARPA Exascale study was one of the first in-depth attempts to project ahead key characteristics for high-end massively parallel systems on the basis of technology trends, architectures, and computational kernels, and identified four major challenges for future systems designs. It focused on a single benchmark, Linpack, and identified two distinct classes of architectures: “heavyweight” and “lightweight.” This talk is a continuation of a series of updates to that study, and includes not only the most recent technology projections but also several new benchmarks for which significant multi-year data exists, and new classes of architectures that have emerged since then. The talk will address changes in characteristics (both before and after the seminal year of 2004 where multi-core took over), and how those characteristics are likely to project into the future. A series of vignettes on specific features will provide insight into areas where current design trends are becoming over or under-balanced. Special attention is given to both computational energy and memory.
Peter M. Kogge
IPDPS1
2013 Energy-efficient multithreading for a hierarchical heterogeneous multicore through locality-cognizant thread generation
Patrick Anthony La Fratta, Peter M. Kogge
J. Parallel Distributed Comput.2
2011 Using the TOP500 to trace and project technology and architecture trends
abstract
The TOP500 is a treasure trove of information on the leading edge of high performance computing. It was used in the 2008 DARPA Exascale technology report to isolate out the effects of architecture and technology on high performance computing, and lay the groundwork to project how current systems might mature through the coming years. Two particular classes of architectures were identified: "heavy-weight" (based on high end commodity microprocessors) and "lightweight," (primarily BlueGene variants), and projections made on performance, concurrency, memory capacity, and power. This paper updates those projections, and adds a third class of "heterogeneous" architectures (leveraging the emerging class of GPU-like chips) to the mix.
Peter M. Kogge, Timothy J. Dysart
SC1
2011 Lightweight Chip Multi-Threading (LCMT): Maximizing Fine-Grained Parallelism On-Chip
abstract
Irregular and dynamic applications, such as graph problems and agent-based simulations, often require fine-grained parallelism to achieve good performance. However, current multicore processors only provide architectural support for coarse-grained parallelism, making it necessary to use software-based multithreading environments to effectively implement fine-grained parallelism. Although these software-based environments have demonstrated superior performance over heavyweight, OS-level threads, they are still limited by the significant overhead involved in thread management and synchronization. In order to address this, we propose a Lightweight Chip Multi-Threaded (LCMT) architecture that further exploits thread-level parallelism (TLP) by incorporating direct architectural support for an “unlimited” number of dynamically created lightweight threads with very low thread management and synchronization overhead. The LCMT architecture can be implemented atop a mainstream architecture with minimum extra hardware to leverage existing legacy software environments. We compare the LCMT architecture with a Niagara-like baseline architecture. Our results show up to 1.8X better scalability, 1.91X better performance, and more importantly, 1.74X better performance per watt, using the LCMT architecture for irregular and dynamic benchmarks, when compared to the baseline architecture. The LCMT architecture delivers similar performance to the baseline architecture for regular benchmarks.
Sheng Li 0007, Shannon K. Kuntz, Jay B. Brockman, Peter M. Kogge
IEEE Trans. Parallel Distributed Syst.4
2009 Organizing wires for reliability in magnetic QCA
abstract
This article investigates, via analytic modeling, how a magnetic QCA wire should be organized to provide the highest reliability. We compare a nonredundant wire and two redundant wire organizations. For all three organizations, a fault rate per unit length is used for comparison; additionally, since extra components are necessary to implement the redundant organizations, these components are faulty as well. We show that the difference between these two fault rates is the main driver for selecting a wire organization. Lastly, we develop a guideline for selecting the most reliable wire organization during the circuit design process.
Timothy J. Dysart, Peter M. Kogge
ACM J. Emerg. Technol. Comput. Syst.2
2009 Analyzing the Inherent Reliability of Moderately Sized Magnetic and Electrostatic QCA Circuits Via Probabilistic Transfer Matrices
abstract
As computing technology delves deeper into the nanoscale regime, reliability is becoming a significant concern, and in response, Teramac-like systems will be the model for many early non-CMOS nanosystems. Engineering systems of this type requires understanding the inherent reliability of both the functional cells and the interconnect used to build the system, and which components are most critical. One particular nanodevice, quantum-dot cellular automata (QCA), offers unique challenges in understanding the reliability of its basic circuits since the device used for logic is also used for interconnect. In this paper, we analyze the reliability properties of two classes of QCA devices: molecular electrostatic-based and magnetic-domain-based. We use an analytic model, probabilistic transfer matrices (PTMs), to compute the inherent reliability of various nontrivial circuits. Additionally, linear regression is used to determine which components are most critical and estimated the reliability gains that may be achieved by improving the reliability of just a critical component. The results show the critical importance of different structures, especially interconnect, as used by the two classes of QCA.
Timothy J. Dysart, Peter M. Kogge
IEEE Trans. Very Large Scale Integr. Syst.2
2008 Design of a mask-programmable memory/multiplier array using G4-FET technology
abstract
A G4-FET is a 4 gate transistor that combines both JFET and MOS characteristics in a single device that may be fabricated in a standard silicon-on-insulator (SOI) process. In doing so, it enables the conducting channel to be controlled vertically through MOS gates, as well as horizontally, through junction gates. Further, depending upon how it is biased, a single G4-FET can serve as either a not-majority logic gate or as a charge storage-based memory cell. This unique device offers tremendous potential for innovative gate arrays, where real estate can be traded-off between logic and memory functions. In this paper, we take a first look at a mask-programmable G4-FET array that depending upon metal personalization, can function either as a DRAM array or a multiplier.
Jay B. Brockman, Sheng Li 0007, Peter M. Kogge, Amit Kashyap, Mohammad M. Mojarradi
DAC3
2008 Memory model effects on application performance for a lightweight multithreaded architecture
abstract
In this paper, we evaluate the effects of a partitioned global address space (PGAS) versus a flat, randomized distributed global address space (DGAS) in the context of a lightweight multithreaded parallel architecture. We also execute the benchmarks on the Cray MTA-2, a multithreaded architecture with a DGAS mapping. Key results demonstrate that distributing data under the PGAS mapping increases locality, effectively reducing the memory latency and the number of threads needed to achieve a given level of performance. In contrast, the DGAS mapping provides a simpler programming model by eliminating the need to distribute data and, assuming sufficient application parallelism, can achieve similar performance by leveraging large numbers of threads to hide the longer latencies.
Sheng Li 0007, Shannon K. Kuntz, Peter M. Kogge, Jay B. Brockman
IPDPS3
2007 A Heterogeneous Lightweight Multithreaded Architecture
abstract
Programs with irregular patterns of dynamic data structures and/or those with complicated control structures such as recursion are notoriously difficult to parallelize efficiently. For some highly-irregular applications, such as a SAT solver, it has been nearly impossible to obtain significant parallel speedups on conventional SMP systems over serial implementations. Lightweight multithreading, as found in the Cray MTA and the upcoming XMT (Eldorado), has been demonstrated as an effective approach to attacking these problems. In this paper, we describe a heterogeneous lightweight multithreading that extends ideas found in the Cray machines to support larger numbers of threads while reducing the cost of thread management and synchronization.
Sheng Li 0007, Amit Kashyap, Shannon K. Kuntz, Jay B. Brockman, Peter M. Kogge, Paul L. Springer, Gary Block
IPDPS5
2007 Evaluating synchronization techniques for light-weight multithreaded/multicore architectures
abstract
No abstract available.
Srinivas Sridharan 0002, Arun Rodrigues, Peter M. Kogge
SPAA3
2007 On the Memory Access Patterns of Supercomputer Applications: Benchmark Selection and Its Implications
abstract
This paper compares the system performance evaluation cooperative (SPEC) Integer and Floating-Point suites to a set of real-world applications for high-performance computing at Sandia National Laboratories. These applications focus on the high-end scientific and engineering domains; however, the techniques presented in this paper are applicable to any application domain. The applications are compared in terms of three memory properties: 1) temporal locality (or reuse over time), 2) spatial locality (or the use of data "near" data that has already been accessed), and 3) data intensiveness (or the number of unique bytes the application accesses). The results show that real-world applications exhibit significantly less spatial locality, often exhibit less temporal locality, and have much larger data sets than the SPEC benchmark suite. They further quantitatively demonstrate the memory properties of real supercomputing applications.
Richard C. Murphy, Peter M. Kogge
IEEE Trans. Computers2
2006 Fine-Grained Message Pipelining for Improved MPI Performance
abstract
By its nature, MPI leads to coarse grained communications. This is because all current MPI implementations deliver two orders of magnitude more bandwidth for large message sizes (kilobytes) than small message sizes (bytes). This translates into applications that bundle their small communications into larger communications whenever possible. In modern implementations, this sacrifice in the granularity of communication translates directly into a sacrifice in the granularity of synchronization. MPI requires that the entire message arrive before any of the data can be delivered to the application, because message completion is the only synchronization semantic the network can expose to the processor. This paper explores the implications of providing synchronization between the network and the processor at the memory word level using a mechanism such as Full/Empty Bits. This enables the application to begin computing as soon as the data for the first memory referenced has arrived without having to wait for all of the data in the message
Arun Rodrigues, Kyle B. Wheeler, Peter M. Kogge, Keith D. Underwood
CLUSTER3
2006 M06 - Issues for the future of supercomputing: impact of Moore's law and architecture on application performance
abstract
This tutorial will explain technologies driving supercomputer speed increases and supercomputers' resulting ability to solve increasingly important problems. In guided session, participants will walk through future supercomputer performance projection on a representative application.Specific topics will include:A review of scientific problems amenable to supercomputers based on the broad-based SCaLeS study. The editor of the SCaLeS report presents this session.The International Technology Roadmap for Semiconductors (ITRS) and its implications for the size, speed, power, and architecture of supercomputers.Current microprocessor-based and emerging "advanced" architectures and their implications on applications performance. There will be special emphasis on the multi-core trend and how applications can use multiple cores effectively.We will discuss the physics issues (speed and power) that will define the end of the current evolutionary trend. We will show how nanotech, reversible logic, and quantum computing may become the basis of a revolutionary change.
Erik DeBenedictis, David E. Keyes, Peter M. Kogge
SC3
2006 Poster reception - The structural simulation toolkit: exploring novel architectures
abstract
Exploring novel computer system designs requires modeling the complex interactions between processor, memory, and network. The Structural Simulation Toolkit (SST) has been developed to explore innovations in both the programming models and hardware implementation of highly concurrent systems. The Toolkit's modular design allows extensive exploration of system parameters while maximizing code reuse and provides an explicit separation of instruction interpretation from microarchitectural timing. This is built upon a high performance hybrid discrete event framework. The SST has modeled a variety of systems, from processor-in-memory to CMP and MPP. It has examined a variety of hardware and software issues in the context of HPC.This poster presents an overview of the SST. Several of its models for processors, memory systems, and networks will be detailed. Its software stack, including support for MPI and OpenMP, will also be covered. Performance results and current directions for the SST will also be shown.
Arun Rodrigues, Richard C. Murphy, Peter M. Kogge, Keith D. Underwood
SC3
2006 Multi-core issues - Multi-Core for HPC: breakthrough or breakdown?
abstract
A dramatic trend in computing is the adoption of multi-core technology by the vendors from which our current and future HPC systems are being derived. Multi-core is offered as a path to continued reliance and benefits of Moore's Law while reining in the previously unfettered growth of power consumption and design complexity. Are we saved? or is it but a fools mission, trapping us in a technical cul de sac with no long term direction and no way to reinvent an alternative future. The panel will consider the following questions:* Can multi-core span the next decade of Moore's Law progression?* Are the pins and caches a strangle hold on the future effectiveness of multi-core?* Can innovative algorithmic techniques exploit the opportunities and address the challenges of multi-core?* How will programming models and supporting system software change to accommodate the unique properties and peculiarities of multi-core structures?
Thomas L. Sterling, Peter M. Kogge, William J. Dally, Steve Scott, William Gropp, David E. Keyes, Pete Beckman
SC2
2005 The implications of working set analysis on supercomputing memory hierarchy design
abstract
Supercomputer architects strive to maximize the performance of scientific applications. Unfortunately, the large, unwieldy nature of most scientific applications has lead to the creation of artificial benchmarks, such as SPEC-FP, for architecture research. Given the impact that these benchmarks have on architecture research, this paper seeks an understanding of how they relate to real-world applications within the Department of Energy. Since the memory system has been found to be a particularly key issue for many applications, the focus of the paper is on the relationship between how the SPEC-FP benchmarks and DOE applications use the memory system. The results indicate that while the SPEC-FP suite is a well balanced suite, supercomputing applications typically demand more from the memory system and must perform more "other work" (in the form of integer computations) along with the floating point operations. The SPEC-FP suite generally demonstrates slightly more temporal locality leading to somewhat lower bandwidth demands. The most striking result is the cumulative difference between the benchmarks and the applications in terms of the requirements to sustain the floating-point operation rate: the DOE applications require significantly more data from main memory (not cache) per FLOP and dramatically more integer instructions per FLOP.
Richard C. Murphy, Arun Rodrigues, Peter M. Kogge, Keith D. Underwood
ICS3
2005 Generation of permutations for SIMD processors
abstract
Short vector (SIMD) instructions are useful in signal processing, multimedia, and scientific applications. They offer higher performance, lower energy consumption, and better resource utilization. However, compilers still do not have good support for SIMD instructions, and often the code has to be written manually in assembly language or using compiler builtin functions. Also, in some applications, higher parallelism could be achieved if compilers inserted permutation instructions that reorder the data in registers. In this paper we describe how we create SIMD instructions from regular code, and determine ordering of individual operations in the SIMD instructions to minimize the number of permutation instructions. Individual memory operations are grouped into SIMD operations based on their effective addresses. The SIMD data flow graph is then constructed by following data dependences from SIMD memory operations. Then, the orderings of operations are propagated from SIMD memory operations into the graph.We also describe our approach to compute decomposition of a given permutation into the permutation instructions of the target architecture. Experiments with our prototype compiler show that this approach scales well with the number of operations in SIMD instructions (SIMD width) and can be used to compile a number of important kernels, achieving up to 35% speedup.
Alexei Kudriavtsev, Peter M. Kogge
LCTES2
2005 Polygonal path simplification with angle constraints
Danny Ziyi Chen, Ovidiu Daescu, John Hershberger 0001, Peter M. Kogge, Ningfang Mi, Jack Snoeyink
Comput. Geom.4
2004 Quantum-Dot Cellular Automata (QCA) circuit partitioning: problem modeling and solutions
abstract
This paper presents the Quantum-Dot Cellular Automata (QCA) physical design problem, in the context of the VLSI physical design problem. The problem is divided into three subproblems: partitioning, placement, and routing of QCA circuits. This paper presents an ILP formulation and heuristic solution to the partitioning problem, and compares the two sets of results. Additionally, we compare a human-generated circuit to the ILP and Heuristic solutions. The results demonstrate that the heuristic is a practical method of reducing partitioning run time while providing a result that is close to the optimal for a given circuit.
Dominic A. Antonelli, Danny Ziyi Chen, Timothy J. Dysart, Xiaobo Sharon Hu, Andrew B. Kahng, Peter M. Kogge, Richard C. Murphy, Michael T. Niemier
DAC6
2004 Using Circuits and Systems-Level Research to Drive Nanotechnology
abstract
This paper details nano-scale devices being researched by physical scientists to build computational systems. It also reviews some existing system design work that uses the devices to be discussed. It concludes with a discussion of how the authors believe system-level research can best be used to positively affect actual device development. This work has led to a more thorough design methodology that address whether or not computationally interesting and buildable circuits are possible with the quantum-dot cellular automata (QCA), while also providing significant wins over end-of-the-roadmap CMOS.
Michael T. Niemier, Ramprasad Ravichandran, Peter M. Kogge
ICCD3
2004 Characterizing a new class of threads in scientific applications for high end supercomputers
abstract
Chip level multithreading is growing in use throughout the microprocessor world as evidenced in the Intel Pentium 4 and the upcoming innovations in the POWER architecture. These processors typically use a few coarse grain threads that can be difficult for the programmer or compiler to exploit; however, Processing in Memory (PIM) is a technology that has been explored through a long series of supercomputer projects as a facilitator for a different multithreaded execution models. In the multithreading model explored by PIMs, the threads can have radically different characteristics. Specifically, PIMs seek to exploit a large number of very fine grained threads to hide memory access latency and increase parallelism. PIM supports these small threads, or "threadlets", by providing a fast hardware synchronization mechanism, support for harware managment of creation and destruction of threads, and a "shared register" approach which extends the shared memory thread model. This paper discusses some analysis of some very large scientific codes in terms of how they might be mapped onto such a multithreading model with a focus on extremely fine grain threads.
Arun Rodrigues, Richard C. Murphy, Peter M. Kogge, Keith D. Underwood
ICS3
2004 Cache implications of aggressively pipelined high performance microprocessors
abstract
One of the major design decisions when developing a new microprocessor is determining the target pipeline depth and clock rate since both factors interact closely with one another. The optimal pipeline depth of a processor has been studied before, but the impact of the memory system on pipeline performance has received less attention. This study analyzes the affect of different level-1 cache designs across a range of pipeline depths to determine what role the memory system design plays in choosing a clock rate and pipeline depth for a microprocessor. The pipeline depths studied here range from those found in current processors to those predicted for future processors. For each pipeline depth a variety of level-1 cache sizes are simulated to explore the relationship between clock rate, pipeline depth, cache size and access latency. Results show that the larger caches afforded by shorter pipelines with slower clocks outperform longer pipelines with smaller caches and higher clock rates.
Timothy J. Dysart, Branden J. Moore, Lambert Schaelicke, Peter M. Kogge
ISPASS4
2003 Implications of a PIM Architectural Model for MPI
abstract
Memory may be the only system component that is more commoditized than a microprocessor. To simultaneously exploit this and address the impending memory wall, processing in memory (PIM) research efforts are considering ways to move processing into memory without significantly increasing the cost of the memory. As such, PIM devices may become the basis for future commodity clusters. Although these PIM devices may leverage new computational paradigms such as hardware support for multi-threading and traveling threads, they must provide support for legacy programming models if they are to supplant commodity clusters. This paper presents a prototype implementation of MPI over a traveling thread mechanism called parcels. A performance analysis indicates that the direct hardware support of a traveling thread model can lead to an efficient, lightweight MPI implementation.
Arun Rodrigues, Richard C. Murphy, Peter M. Kogge, Jay B. Brockman, Ron Brightwell, Keith D. Underwood
CLUSTER3
2003 The State of State
abstract
We are all aware of the memory wall and the deleterious effects of bandwidth and latency limitations on performance. We also watch with some degree of amazement at the relentless march of Moore's Law as ever-larger numbers of transistors are used in increasingly clever architectural and microarchitectural techniques to attempt to reduce the effects of the wall. What have not risen to the same level of consciousness, however, are the effects that all of this has on the size of program state and our ability to manipulate it, and move it, to avoid the wall. Instead, the Law of Unintended Consequences has left us with heavier and heavier state, which in turn condemns them to continued existence in the bowels of bigger and bigger microprocessor chips, and farther and farther away from the data they seek in memory. This talk is a plea to architects to reconsider what we have been doing to ourselves, and ask if there are alternatives that have been overlooked. The talk will begin by revisiting the notion of state, and plot its explosive growth over the last 30 years. The key points of expansion will be correlated with architectural and microarchitectural advances. Then we will walk through some observations gained from exploring designing in some emerging technologies, and ask the question of what alternative execution models, particularly premised on light weight states, might do to increase performance and reduce complexity.
Peter M. Kogge
HPCA1
2003 Energy-efficient issue queue design
abstract
The out-of-order issue queue (IQ), used in modern superscalar processors is a considerable source of energy dissipation. We consider design alternatives that result in significant reductions in the power dissipation of the IQ (by as much as 75%) through the use of comparators that dissipate energy mainly on a tag match, 0-B encoding of operands to imply the presence of bytes with all zeros and, bitline segmentation. Our results are validated by the execution of SPEC 95 benchmarks on a true hardware level, cycle-by-cycle simulator for a superscalar processor and SPICE measurements for actual layouts of the IQ in a 0.18-/spl mu/m CMOS process.
Dmitry V. Ponomarev, Gürhan Küçük, Oguz Ergin, Kanad Ghose, Peter M. Kogge
IEEE Trans. Very Large Scale Integr. Syst.5
2001 A Microserver View of HTMT
abstract
Hybrid technology multithreaded architecture (HTMT) is an ambitious new architecture combining cutting edge technologies to reach petaflop performance sooner than current technology trends allow. It is a massively parallel architecture with multi-threaded hardware and a multi-level memory hierarchy. Microservers provide a new perspective for viewing this memory hierarchy whereby memory is actively involved in process execution. This paper discusses the microserver memory semantics and initial HTMT execution models to analyze application at each level of the system hierarchy and to develop user-level functions for expressing this inherent concurrency and parallelism. In order to do this we studied several applications to model the control and data flow within the HTMT hierarchy and developed pseudo-code representing the user-level functions necessary to express application concurrency and parallelism.
Lilia Yerosheva, Shannon K. Kuntz, Peter M. Kogge, Jay B. Brockman
IPDPS3
2001 Exploring and exploiting wire-level pipelining in emerging technologies
abstract
Pipelining is a technique that has long since been considered fundamental by computer architects. However, the world of nanoelectronics is pushing the idea of pipelining to new and lower levels — particularly the device level. How this affects circuits and the relationship between their timing, architecture, and design will be studied in the context of an inherently self-latching nanotechnology termed Quantum Cellular Automata (QCA). Results indicate that this nanotechnology offers the potential for “free” multi-threading and “processing-in-wire”. All of this could be accomplished in a technology that could be almost three orders of magnitude denser than an equivalent design fabricated in a process at the end of the CMOS curve.
Michael T. Niemier, Peter M. Kogge
ISCA2
2001 Energy: efficient instruction dispatch buffer design for superscalar processors
abstract
The instruction dispatch buffer (DB, also known as an issue queue) used in modem superscalar processors is a considerable source of energy dissipation. We consider design alternatives that result in significant reductions in the power dissipation of the DB (by as much as 60%) through the use of: (a) fast comparators that dissipate energy mainly on a tag match, (b) zero byte encoding of operands to imply the presence of bytes with all zeros and, (c) bitline segmentation. Our results are validated by the execution of SPEC 95 benchmarks on true hardware level, cycle-by-cycle simulator for a superscalar processor and SPICE measurements for actual layouts of the DB and its variants in a 0.5 micron CMOS process.
Gürhan Küçük, Kanad Ghose, Dmitry V. Ponomarev, Peter M. Kogge
ISLPED4
2001 Polygonal path approximation with angle constraints
Danny Ziyi Chen, Ovidiu Daescu, John Hershberger 0001, Peter M. Kogge, Jack Snoeyink
SODA4
2001 Inherently Lower-Power High-Performance Superscalar Architectures
abstract
In recent years, reducing power has become an important design goal for high-performance microprocessors. This work attempts to bring the power issue to the earliest phases of microprocessor development, in particular, the stage of defining a chip microarchitecture. We investigate power-optimization techniques of superscalar microprocessors at the microarchitecture level that do not compromise performance. First, major targets for power reduction are identified within microarchitecture, where power is heavily consumed or will be heavily consumed in next-generation superscalar processors. Then, a new, energy-efficient version of a multicluster microarchitecture is developed that reduces energy the identified critical design points with minimal performance impact. A methodology is developed for energy-performance optimization at the microarchitecture level that generates, for a microarchitecture, a set of energy-efficient configurations, forming a convex hull in the power-performance space. Detailed simulation of the baseline and proposed multicluster architectures has been performed using the developed optimization methodology. A comparison of the two microarchitectures, both optimized for energy efficiency, shows that the multicluster architecture is potentially up to twice as energy efficient for wide issue processors, with an advantage that Grows with the issue width. Conversely, at the same power dissipation level, the multicluster architecture supports configurations with measurably higher performance than equivalent conventional designs.
Victor V. Zyuban, Peter M. Kogge
IEEE Trans. Computers2
2000 A design of and design tools for a novel quantum dot based microprocessor
abstract
Despite the seemingly endless upw ards spiral of modern VLSI technology, many experts are predicting a hard w all for CMOS in about a decade. Given this, researc hers con tin ue to look at alternative technologies, one of which is based on quan tumdots, called quan tumcellular automata (QCA). While the first such devices have been fabricated, little is kno wn about how to design complete systems of them. This paper summarizes one of the first such studies, namely an attempt to design a complete, albeit simple, CPU in the technology. T o design a theoretical QCA microprocessor, two things must be accomplished. First a device model of the processor must be constructed (i.e. the schematic itself). Second, methods for sim ulatingand testing QCA designs m ust be developed. This paper summarizes the beginnings of a simple QCA microprocessor (namely, its dataflow) and a QCA design and simulation tool.
Michael T. Niemier, Michael J. Kontz, Peter M. Kogge
DAC3
2000 Optimization of high-performance superscalar architectures for energy efficiency
abstract
In recent years reducing power has become a critical design goal for high-performance microprocessors. This work attempts to bring the power issue to the earliest phase of high-performance microprocessor development. We propose a methodology for power-optimization at the micro-architectural level. First, major targets for power reduction are identified within superscalar microarchitecture, then an optimization of a superscalar micro-architecture is performed that generates a set of energy-efficient configurations forming a convex hull in the power-performance space. The energy-efficient families are then compared to find configurations that dissipate the lowest power given a performance target, or, conversely, deliver the highest performance given a power budget. Application of the developed methodology to a superscalar micro-architecture shows that at the architectural level there is a potential for reducing power up to 50%, given a performance requirement, and for up to 15% performance improvement, given a power budget.
Victor V. Zyuban, Peter M. Kogge
ISLPED2
1999 Logic in Wire: Using Quantum Dots to Implement a Microprocessor
abstract
Despite the seemingly endless upwards spiral of modern VLSI technology many experts are predicting a hard wall for CMOS in about a decade. Given this, researchers continue to look at alternative technologies, one of which is based on quantum dots, called quantum cellular automata. While the first such devices have been fabricated, little is known about how to design complete systems. This paper summarizes one of the first such studies, namely an attempt to design a complete, albeit simple, CPU in the technology. The projections are striking: a projected 10 to 1 increase in circuit density when compared to a CMOS equivalent, but a design approach which is radically different from conventional "logic" design, especially in timing considerations.
Michael T. Niemier, Peter M. Kogge
Great Lakes Symposium on VLSI2
1999 Microservers: a new memory semantics for massively parallel computing
abstract
Article Microservers: a new memory semantics for massively parallel computing Share on Authors: Jay B. Brockman Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, IN Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, INView Profile , Peter M. Kogge Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, IN Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, INView Profile , Thomas L. Sterling Center for Advanced Computing Research, California Institute of Technology, Pasadena, CA Center for Advanced Computing Research, California Institute of Technology, Pasadena, CAView Profile , Vincent W. Freeh Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, IN Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, INView Profile , Shannon K. Kuntz Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, IN Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, INView Profile Authors Info & Claims ICS '99: Proceedings of the 13th international conference on SupercomputingJune 1999 Pages 454–463https://doi.org/10.1145/305138.305234Online:01 May 1999Publication History 29citation567DownloadsMetricsTotal Citations29Total Downloads567Last 12 Months9Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Jay B. Brockman, Peter M. Kogge, Thomas L. Sterling, Vincent W. Freeh, Shannon K. Kuntz
International Conference on Supercomputing2
1999 Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture
abstract
Processing-in-memory (PIM) chips that integrate processor logic into memory devices offer a new opportunity for bridging the growing gap between processor and memory speeds, especially for applications with high memory-bandwidth requirements.The Data-IntensiVe Architecture (DIVA) system combines PIM memories with one or more external host processors and a PIM-to-PIM interconnect.DIVA increases memory bandwidth through two mechanisms: (1) performing selected computation in memory, reducing the quantity of data transferred across the processor-memory interface; and (2) providing communication mechanisms called parcels for moving both data and computation throughout memory, further bypassing the processor-memory bus.DIVA uniquely supports acceleration of important irregular applications, including sparse-matrix and pointer-based computations.In this paper, we focus on several aspects of DIVA designed to effectively support such computations at very high performance levels: (1) the memory model and parcel definitions; (2) the PIM-to-PIM interconnect; and, (3) requirements for the processor-to-memory interface.We demonstrate the potential of PIMbased architectures in accelerating the performance of three irregular computations, sparse conjugate gradient, a natural-join database operation and an object-oriented database query.
Mary W. Hall, Peter M. Kogge, Jefferey G. Koller, Pedro C. Diniz, Jacqueline Chame, Jeffrey T. Draper, Jeff LaCoss, John J. Granacki, Jay B. Brockman, Apoorv Srivastava, William C. Athas, Vincent W. Freeh, Joonseok Park
SC2
1999 Accelerating object-oriented applications using method lookup caches and register windowing
Kanad Ghose, Kiran Raghavendra Desai, Peter M. Kogge
J. Syst. Archit.3
1999 Application of STD to latch-power estimation
abstract
In this paper, we use the recently developed static transition diagram technique to derive analytical formulas expressing latch power in terms of true and spurious switching activities at the data input. These formulas are verified through analog simulation and applied to a number of commonly used latch designs. The derived model will allow designers to substitute parameters of true and spurious switching activities into analytical formulas for quickly obtaining accurate latch power and then select latches which have the best power-dissipation characteristics for those parameter values.
Victor V. Zyuban, Peter M. Kogge
IEEE Trans. Very Large Scale Integr. Syst.2
1998 The energy complexity of register files
abstract
Register files (RF) represent a substantial portion of the energy budget in modern processors, and are growing rapidly with the trend towards wider instruction issue. The actual access energy costs depend greatly on the register file circuitry used. This paper compares various RF circuitry techniques for their energy ef- ficiencies, as a function of architectural parameters such as the number of registers and the number of ports. The Port Priority Selection technique was found to be the most energy efficient. The dependence of register file access energy upon technology scaling is also studied. However, as this paper shows, it appears that none of these will be enough to prevent centralized register files from becoming the dominant power component of next-generation superscalar computers, and alternative methods for inter-instruction communication need to be developed. Split register file architecture is analyzed as a possible alternative.
Victor V. Zyuban, Peter M. Kogge
ISLPED2
1994 EXECUBE - A New Architecture for Scalable MPPs
abstract
The EXECUBE chip is a new single part type building block for MPP systems that scales seamlessly from a few chips (with a few hundred mips) to thousands of chips with petaop potential. Further, the chip architecture supports directly both SIMD and MIMD modes of processing, permitting not only the best of both current parallel computing modes but also new modes not possible with more conventional designs. This paper discusses the overall architecture of the EXECUBE chip, the new computational model it represents, some comparisons against the current state of the art, how it might be used for real applications, and some extrapolations into future developments.
Peter M. Kogge
ICPP (1)1
1985 Function-based computing and parallelism: A review
Peter M. Kogge
Parallel Comput.1
1977 The Microprogramming of Pipelined Processors
Peter M. Kogge
ISCA1
1973 Maximal Rate Pipelined Solutions to Recurrance Problems
abstract
An m thorder recurrence problem is defined as the computation of X1, . . . XN, where Xi=f (ai, Xi-1, . . Xi-m) and ai is a set of parameters. On a pipelined computer, where the total stage delay in computing f is df time units, the solution output rate is one new Xi each df time unit. This paper describes a method for increasing this rate to 1 per time unit when the function f has certain simple functional properties. The total stage delay and complexity of the resulting pipelines are also described.
Peter M. Kogge
ISCA1
1973 A Parallel Algorithm for the Efficient Solution of a General Class of Recurrence Equations
abstract
An mth-order recurrence problem is defined as the computation of the series x1, x2, ..., XN, where xi= fi(xi-1, ..., xi-m) for some function fi. This paper uses a technique called recursive doubling in an algorithm for solving a large class of recurrence problems on parallel computers such as the Iliac IV. Recursive doubling involves the splitting of the computation of a function into two equally complex subfunctions whose evaluation can be performed simultaneously in two separate processors. Successive splitting of each of these subfunctions spreads the computation over more processors. This algorithm can be applied to any recurrence equation of the form xi= f(bi, g(ai, xi-1)) where f and g are functions that satisfy certain distributive and associative-like properties. Although this recurrence is first order, all linear mth-order recurrence equations can be cast into this form. Suitable applications include linear recurrence equations, polynomial evaluation, several nonlinear problems, the determination of the maximum or minimum of N numbers, and the solution of tridiagonal linear equations. The resulting algorithm computes the entire series x1, ..., xNin time proportional to [log2N] on a computer with N-fold parallelism. On a serial computer, computation time is proportional to N.
Peter M. Kogge, Harold S. Stone
IEEE Trans. Computers1