Robert K. Montoye

dblp:91/5357 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Hardware accelerators and domain-specific architectures · 38% Processor architecture and microarchitecture · 32% Energy-efficient computing · 22%
Computer networks
1 paper
Cellular and mobile networks · 100%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.512021
Energy Efficiency Boost in the AI-Infused POWER10 Processor · ISCA 2021
Processor architecture and microarchitecture › microprocessor design
processor core design
0.512021
Energy Efficiency Boost in the AI-Infused POWER10 Processor · ISCA 2021
Cellular and mobile networks › mobile networks › mobile network architecture › cellular network architecture
base station architecture
0.112012
Architectural perspectives of future wireless base stations based on the IBM PowerEN™ processor · HPCA 2012
Cellular and mobile networks
radio access networks
0.112012
Architectural perspectives of future wireless base stations based on the IBM PowerEN™ processor · HPCA 2012
Energy-efficient computing
power-performance tradeoff
0.112010
Practical Strategies for Power-Efficient Computing Technologies · Proc. IEEE 2010
Energy-efficient computing
voltage scaling
0.112010
Practical Strategies for Power-Efficient Computing Technologies · Proc. IEEE 2010
Electronic design automation
logic synthesis
0.122008
Custom is from Venus and synthesis from Mars · DAC 2008
Area-time efficient addition in charge based technology · DAC 1981
Processor architecture and microarchitecture
multithreading
0.012012
Architectural perspectives of future wireless base stations based on the IBM PowerEN™ processor · HPCA 2012
Memory systems
cache
0.012010
Practical Strategies for Power-Efficient Computing Technologies · Proc. IEEE 2010
Parallel and multicore computing
parallel algorithms
0.011982
A Practical Algorithm for the Solution of Triangular Systems on a Parallel Processing System · IEEE Trans. Computers 1982
Processor architecture and microarchitecture
arithmetic unit
0.011981
Area-time efficient addition in charge based technology · DAC 1981
Integrated circuit design
digital circuit design
0.011981
Area-time efficient addition in charge based technology · DAC 1981
Processor architecture and microarchitecture › SIMD
SIMD machine
0.011982
A Practical Algorithm for the Solution of Triangular Systems on a Parallel Processing System · IEEE Trans. Computers 1982
Electronic design automation › design optimization
area-time optimization
0.011981
Area-time efficient addition in charge based technology · DAC 1981

Methods — techniques the papers use, named apart from their topics

pre-silicon methodology · 0.5reconfigurable computing · 0.3SIMD · 0.3device-circuit-system codesign · 0.1charge-based technology · 0.0
YearPublicationVenuePosition
2021 Energy Efficiency Boost in the AI-Infused POWER10 Processor
abstract
We present the novel micro-architectural features, supported by an innovative and novel pre-silicon methodology in the design of POWER10. The resulting projected energy efficiency boost over POWER9 is 2.6x at core level (for SPECint) and up to 3x at socket level. In addition, a new feature supporting inline AI acceleration was added to the POWER ISA and incorporated into the POWER10 processor core design. The resulting boost in SIMD/AI socket performance is projected to be up to 10x for FP32 and 21x for INT8 models of ResNet-50 and BERT-Large. In this paper, we describe the novel methodology deployed and used not only to obtain these efficiency boosts for traditional workloads, but also to infuse AI/ML/HPC capability directly into the POWER10 core.
Brian W. Thompto, Dung Q. Nguyen, José E. Moreira, Ramon Bertran Monfort, Hans M. Jacobson, Richard J. Eickemeyer, Rahul M. Rao, Michael Goulet, Marcy Byers, Christopher J. Gonzalez, Karthik Swaminathan, Nagu R. Dhanwada, Silvia M. Müller, Satish Kumar Sadasivam, Robert K. Montoye, William J. Starke, Christian G. Zoellin, Michael S. Floyd, Jeffrey Stuecheli, Nandhini Chandramoorthy, John-David Wellman, Alper Buyuktosunoglu, Matthias Pflanz, Balaram Sinharoy, Pradip Bose
ISCA16
2017 Very Low Voltage (VLV) Design
abstract
This paper is a tutorial-style introduction to a special session on: Effective Voltage Scaling in the Late CMOS Era. It covers the fundamental challenges and associated solution strategies in pursuing very low voltage (VLV) designs. We discuss the performance and system reliability constraints that are key impediments to VLV. The associated trade-offs across power, performance and reliability are helpful in inferring the optimal operational voltage-frequency point. This work was performed under the auspices of an ongoing DARPA program (named PERFECT) that is focused on maximizing system-level energy efficiency.
Ramon Bertran Monfort, Pradip Bose, David Brooks 0001, Jeff Burns, Alper Buyuktosunoglu, Nandhini Chandramoorthy, Eric Cheng, Martin Cochet, Schuyler Eldridge, Daniel J. Friedman, Hans M. Jacobson, Rajiv V. Joshi, Subhasish Mitra, Robert K. Montoye, Arun Paidimarri, Pritish Parida, Kevin Skadron, Mircea R. Stan, Karthik Swaminathan, Augusto Vega, Swagath Venkataramani, Christos Vezyrtzis, Gu-Yeon Wei, John-David Wellman, Matthew M. Ziegler
ICCD14
2014 Matrix-matrix multiplication on a large register file architecture with indirection
abstract
Dense matrix-matrix multiply is an important kernel in many high performance computing applications including the emerging deep neural network based cognitive computing applications. Graphical processing units (GPU) have been very successful in handling dense matrix-matrix multiply in a variety of applications. However, recent research has shown that GPUs are very inefficient in using the available compute resources on the silicon for matrix multiply in terms of utilization of peak floating point operations per second (FLOPS). In this paper, we show that an architecture with a large register file supported by “indirection ” can utilize the floating point computing resources on the processor much more efficiently. A key feature of our proposed in-line accelerator is a bank-based very-large register file, with embedded SIMD support. This processor-in-regfile (PIR) strategy is implemented as local computation elements (LCEs) attached to each bank, overcoming the limited number of register file ports. Because each LCE is a SIMD computation element, and all of them can proceed concurrently, the PIR approach constitutes a highly-parallel super-wide-SIMD device. We show that we can achieve more than 25% better performance than the best known results for matrix multiply using GPUs. This is achieved using far lesser floating point computing units and hence lesser silicon area and power. We also show that architecture blends well with the Strassen and Winograd matrix multiply algorithms. We optimize the selective data parallelism that the LCEs enable for these algorithms and study the area-performance trade-offs.
Dheeraj Sreedhar, Jeff H. Derby, Robert K. Montoye, Charles L. Johnson
HiPC3
2013 Processor architecture for software implementation of multi-sector G-RAKE receivers for HSUPA wireless infrastructure
abstract
The high speed uplink packet access (HSUPA) wireless standard requires extremely high-performance signal processing in the baseband receiver, the most challenging being the chip rate rake receiver. In this paper we describe the architectural enhancements on the IBM's PowerEN processor, to enable it to support the computational requirements of the rake receiver in a fully programmable and scalable fashion. A key feature of these enhancements is a bank-based very-large register file, with embedded single instruction multiple data (SIMD) support. This processor-in-regfile (PIR) strategy is implemented as local computation elements (LCEs) attached to each bank. This overcomes the limitation on the number of register file ports and at the same time enables high degree of parallelism. We show that these enhancements enable the integration of multi-sector HSUPA G-RAKE receivers on a single processor.
Dheeraj Sreedhar, Jeff H. Derby, Augusto Vega, Brian Rogers, Charles L. Johnson, Robert K. Montoye
ICASSP6
2012 Architectural perspectives of future wireless base stations based on the IBM PowerEN™ processor
abstract
In wireless networks, base stations are responsible for operating on large amounts of traffic at high speed rates. With the advent of new standards, as 4G, further pressure is put in the hardware requirements to satisfy speeds of up to 1 Gbps. In this work, we study the applicability and potential benefits of the IBM PowerEN processor (a multi-core, massively multithreaded platform) in the realm of base stations for the 3G and 4G standards. The approach involves exploiting the throughput computation capabilities of the PowerEN processor, replacing the bus-attached special-function accelerators with a layer of in-line universal acceleration support, incorporated within the cores. A key feature of this in-line accelerator is a bank-based very-large register file, with embedded SIMD support. This processor-in-regfile (PIR) strategy is implemented as local computation elements (LCEs) attached to each bank, overcoming the limited number of register file ports. Because each LCE is a SIMD computation element, and all of them can proceed concurrently, the PIR approach constitutes a highly-parallel super-wide-SIMD device. To target a broad spectrum of applications for base stations, we also consider a PIR-based architecture built upon reconfigurable LCEs. In this paper, we evaluate the in-line universal accelerator and the PIR strategy focusing on two specific applications for base stations: FFT and Turbo Decoding.
Augusto Vega, Pradip Bose, Alper Buyuktosunoglu, Jeff H. Derby, Michele Franceschini, Robert K. Montoye
HPCA7
2010 Practical Strategies for Power-Efficient Computing Technologies
abstract
After decades of continuous scaling, further advancement of silicon microelectronics across the entire spectrum of computing applications is today limited by power dissipation. While the trade-off between power and performance is well-recognized, most recent studies focus on the extreme ends of this balance. By concentrating instead on an intermediate range, an ~ 8× improvement in power efficiency can be attained without system performance loss in parallelizable applications-those in which such efficiency is most critical. It is argued that power-efficient hardware is fundamentally limited by voltage scaling, which can be achieved only by blurring the boundaries between devices, circuits, and systems and cannot be realized by addressing any one area alone. By simultaneously considering all three perspectives, the major issues involved in improving power efficiency in light of performance and area constraints are identified. Solutions for the critical elements of a practical computing system are discussed, including the underlying logic device, associated cache memory, off-chip interconnect, and power delivery system. The IBM Blue Gene system is then presented as a case study to exemplify several proposed directions. Going forward, further power reduction may demand radical changes in device technologies and computer architecture; hence, a few such promising methods are briefly considered.
Leland Chang, David J. Frank, Robert K. Montoye, Steven J. Koester, Brian L. Ji, Paul Coteus, Robert H. Dennard, Wilfried Haensch
Proc. IEEE3
2008 Custom is from Venus and synthesis from Mars
abstract
Due to ever increasing cost of doing design, design productivity and more specifically, cost of design has become a major bottleneck in large scale design projects. Due to the cost crunch, automated synthesis techniques have been gaining ground on manual and cost intensive (although potentially yielding higher quality of results) custom techniques. The push towards system level performance with multiple lower performance cores as opposed to single high performance core is also making the pursuit of frequency at any cost meaningless. Synthesis vs Custom has traditionally been a lively debate in microprocessor companies but now it is becoming more widespread with the desire of fabless companies to more fully utilize the technology node in order to compete with IDMs by utilizing more custom techniques. Custom hotshots argue that we cannot afford to leave performance or power on table when you are spending 3billion $s in fabs. Industry needs to get maximum benefits possible. Synthesis fanatics will argue that designers should get over the last few ps and last few nW, as time to market, cost, and system performance are the new drivers and synthesis can not only meet the challenge but beat it many times with lower power solutions. This panel will present this internal debate in a public forum.
Ruchir Puri, William H. Joyner, Shekhar Borkar, Ty Garibay, Jonathan Lotz, Robert K. Montoye
DAC6
2005 Testing and debugging delay faults in dynamic circuits
abstract
We propose novel design for test and debug techniques to apply two patterns for delay fault test and debug in dynamic circuits. Dynamic circuits, which have traditionally been difficult to test, pose new challenges for AC tests due to the presence of a reset phase between applications of any two patterns, which impedes delay fault testing of such circuits. We present two sets of design for test and debug techniques. The first set facilitates application of two patterns to dynamic circuits in general, overcoming the issue of reset phase, and reduces the problem of test generation for dynamic circuits to test generation for pull down paths of static CMOS circuits. The second set enables application of two patterns to scan based dynamic circuits. The proposed techniques reduce the problem of delay test generation for scan based dynamic circuits to that of delay test generation for static CMOS circuits with complete accessibility to all primary inputs. The techniques have minimal area overhead and also provide significant reduction in power during scan operation
Ramyanshu Datta, Sani R. Nassif, Robert K. Montoye, Jacob A. Abraham
ITC3
2004 The four degrees of 3D
abstract
No abstract available.
Robert K. Montoye
ISPD1
2004 The four degrees of 3D
abstract
No abstract available.
Robert K. Montoye
ISPD1
1989 IBM second-generation RISC machine organization
abstract
A highly concurrent second-generation RISC (reduced-instruction-set computer) that combines a powerful RISC architecture with sophisticated hardware design techniques to achieve a short cycle time and a low cycles-per-instruction (CPI) ratio is described. Like earlier RISC processors, this design uses a register-oriented instruction set, the CPU is hardwired rather than microcoded, and it features a pipelined implementation. Unlike earlier RISC processors, however, several advanced architectural and implementation features are used, including separate instruction and data caches, zero-cycle branches, multiple-instruction dispatch, and simultaneous execution of fixed- and floating-point instructions. The CPU has a four-word data bus to main memory, a four-word instruction-fetch bus from the I-cache arrays, and a two-word data bus between the D-cache and floating-point unit. The CPU has a full 64-b floating-point engine, and thirty-two 64-b floating point registers in addition to thirty-two 32-b fixed-point registers. In a single cycle, four instructions can be executed simultaneously.>
H. B. Bakoglu, Gregory F. Grohoski, L. E. Thatcher, James A. Kahle, Charles R. Moore, David P. Tuttle, Warren E. Maule, W. R. Hardell Jr., Dwain A. Hicks, M. Nguyenphu, Robert K. Montoye, W. T. Glover, Sudhir Dhawan
ICCD11
1982 A Practical Algorithm for the Solution of Triangular Systems on a Parallel Processing System
abstract
An algorithm is presented for a more efficient and implementable solution of triangular systems on a parallel (SIMD) computer which requires 0(log (N)) fewer processing cycles than the best previous results, where N is the system size. We will also show that the data can be accessed and aligned in the same order of time using as many memory units as processors and Ω networks for data alignment. (Previous results dealing with this type of algorithm have not dealt in any detail with the problem of data access and alignment.)
Robert K. Montoye, Duncan H. Lawrie
IEEE Trans. Computers1
1981 Area-time efficient addition in charge based technology
Robert K. Montoye
DAC1